Method, system and vehicle for immersive audio reproduction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-08-11
AI Technical Summary
虽然这种方法旨在通过忠于原始混响来获得更大的保真度,但提取过程具有较高的计算要求,所述计算要求对于计算机(诸如存在于车辆中的电子控制单元ECU)来说往往要求过高且成本过高
Smart Images

Figure CN122554769A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, system, and vehicle for immersive audio reproduction. Specifically, the invention relates to a method for modifying an audio signal, a system including a plurality of speakers and a processor operable to modify the audio signal, and a vehicle including a plurality of speakers and a processor operable to modify the audio signal. Background Technology
[0002] Conventional multi-speaker systems (e.g., audio and / or high-fidelity sound systems in vehicles and other enclosed spaces) often struggle to accurately reproduce the spatial cues and reverberation characteristics present in the original recording during playback of recorded media. This deficiency results in a lack of fidelity and diminishes the listener's immersion.
[0003] Current methods for providing immersion to listeners have many drawbacks. For example, providing predefined synthetic reverberation to a medium during playback involves adding synthetically generated reverberation on top of the original audio signal (e.g., the original recording). Synthetic reverberation can be designed to mimic the characteristics of a specific physical acoustic space (such as a concert hall or jazz club). However, this approach ignores the natural reverberation already present in the original recording. Therefore, the reproduced sound may deviate from the artistic intent of the source material.
[0004] A different approach attempts to preserve the natural reverberation of the original recording by extracting the reverberation from the audio signal (e.g., using Quantum Logic Surround) and then distributing that reverberation to the system's speakers. While this method aims for greater fidelity by remaining faithful to the original reverberation, the extraction process is computationally demanding, often prohibitively expensive for computers such as the electronic control units (ECUs) found in vehicles. Therefore, using this arrangement can actually lead to a significant degradation in audio quality due to signal processing artifacts.
[0005] Therefore, the industry needs to provide high-quality reproduction of reverberation from the original recording with low computational requirements. Summary of the Invention
[0006] To achieve the above objectives, the present invention provides a method, a system, and a vehicle, as described in the appended claims.
[0007] In a preferred embodiment, a method for modifying an audio signal is provided. The method includes: receiving an audio signal; extracting reverberation characteristics from the audio signal, the reverberation characteristics including a plurality of parameters; and generating artificial reverberation, the artificial reverberation including at least one of the parameters. The method further includes: applying the artificial reverberation to the audio signal; and transmitting the modified audio signal to a plurality of speakers, the modified audio signal including the artificial reverberation applied to the audio signal.
[0008] This invention offers several advantages over conventional multi-speaker systems. By dynamically reproducing the natural reverberation characteristics of the original audio signal, the invention improves the overall fidelity of the reproduced sound. Advantageously, it enhances the fidelity and immersion of audio reproduction in multi-speaker setups (i.e., setups including two or more speaker channels). Specifically, by extracting the original reverberation characteristics (reverberation) from the originally recorded or produced audio and applying artificial reverberation based on those extracted reverberation characteristics, an artificial reverberation that closely mimics the reverberation of the original recording (e.g., mimicking the reverberation experienced in a concert hall with the original audio recording) is generated, since the reverberation is generated based on the characteristics of the original audio signal. High-quality, stable artificial reverberation is generated through artificial reverberation that is very similar to the natural reverberation characteristics extracted from the original audio signal. This is advantageous in enclosed spaces with multi-speaker setups (such as in vehicles), where audio that is very similar to the original recording and spatial cues can be reproduced, thus providing an immersive listening experience for users of the vehicle or enclosed space.
[0009] In one embodiment, the modified audio signal includes a first audio signal and a second audio signal different from the first audio signal. The method may further include: sending the first audio signal to a first channel, the first channel including one or more first speakers among the plurality of speakers; and sending the second audio signal to a second channel, the second channel including one or more second speakers among the plurality of speakers.
[0010] Advantageously, the method described above can be applied to multi-channel setups (such as stereo and other surround sound systems). By providing a first audio signal and a second audio signal, the synthesized reverberation signal can be individually assigned and processed for each channel within a multi-speaker system and for each individual speaker within each channel. This processing can involve system tuning and sound design to ensure optimal spatialization and blending with the original audio signal. By individually assigning and processing the reverberation signal for each channel and for each speaker, the phantom center and perceived front stage of the audio can be adjusted more accurately. This also provides an arrangement in which the perceived width and immersiveness of the stage size of an audio recording can be increased or decreased more accurately. For example, audio recordings for large stages (such as recordings in large concert halls) and audio recordings for small venues (such as recordings in small jazz halls), as well as any audio recordings in between, can be accurately mimicked.
[0011] In one implementation, the first audio signal includes the received audio signal, and the second audio signal includes the modified audio signal.
[0012] Advantageously, the original audio signal can be distributed through the front speakers to create a stereo foundation, while the reproduced synthetic reverberation is added to the side and / or rear speakers. This approach creates a truly immersive listening experience, allowing the listener to perceive the acoustic environment of the original recording as if they were there, while maintaining the fidelity of the original recording.
[0013] In one implementation, the method further includes sampling the audio signal at a sampling rate, wherein the sampling rate is at a predetermined frequency or a dynamically adjusted frequency. The method may also include extracting the reverberation characteristics from the audio signal at each sample of the audio signal.
[0014] By sampling the audio signal at a single sampling rate and extracting reverberation parameters at each sample of the audio signal, rather than extracting the reverberation signal continuously, the computational requirements of the computer / processor are significantly reduced. Furthermore, dynamically adjusting the sampling rate (or update frequency) of the extracted reverberation parameters allows the complexity of the extracted reverberation parameters to be adjusted according to the type of audio signal being fed into the system at a given time. For example, a church choir recording may include a large amount of reverberation, while a singer's studio recording may include a smaller amount of reverberation. Computational requirements can also be improved by dynamically adjusting the sampling rate (e.g., by reducing the number of samples of audio signals that include a lower rate of change in reverberation characteristics, and by increasing the sampling rate of audio signals that include a higher rate of change in reverberation characteristics). This makes it possible to adjust the parameter extraction frequency based on the complexity of the source audio and allows for efficient processing and optimization for different musical styles.
[0015] Therefore, the ability to dynamically adjust the sampling or update frequency (asymmetric interval) of the extracted reverberation parameters based on the complexity of the source audio allows for optimized processing efficiency, resulting in a more efficient system overall. Thus, compared to signal extraction techniques, this invention allows for the production of synthesized reverberation with superior audio quality while maintaining a high degree of fidelity to the original recording's natural reverberation. This enables the creation of unique sound design scenarios based on the received audio signal.
[0016] In the implementation scheme, the extracted reverberation parameters include reverberation time (i.e., the length of the reverberation, which can be RT60, RT30, or any other length of the reverberation length descriptor), pre-delay, attenuation shape, high-frequency and low-frequency damping, venue size, diffusion, width, modulation, spectral density, spectral envelope (frequency-dependent behavior of the reverberation), early reflection patterns in frequency and time, or any combination of the above.
[0017] Advantageously, multiple types of reverb can be considered, thereby producing a more realistic reverb reproduction from the received original audio signal.
[0018] In the implementation, the artificial reverberation further includes applying an equalizer to the audio signal, a delay to the audio signal, a gain to the audio signal, or any combination thereof.
[0019] Advantageously, a variety of different effects can be added, resulting in a more realistic reproduction of the original audio signal.
[0020] In one implementation, the artificial reverb further includes applying a predefined reverb response to the audio signal, and subsequently applying the generated artificial reverb response to the audio signal.
[0021] By first applying a predefined reverb (which may be based, for example, on music genre, style, or other metadata), a significant computational load is removed from the system. This is advantageous, for example, for vehicles that may not have high-performance computers / processors and other computing power. For instance, an audio signal identified as a first-genre (such as a live rock concert) could have a predefined artificial reverb applied to it (e.g., including standard reverb for a specific indoor or outdoor venue and providing emphasis on certain reverb parameters related to guitar, vocals, and drums). To provide a fully immersive experience, reverb characteristics can be extracted at a lower rate (e.g., at a low sampling rate, or by extracting only some parameters instead of all). Artificial reverb can then be applied on top of the predefined reverb to provide a fully immersive listening experience. Advantageously, less computation is required compared to extracting the complete reverb signal. Extracting parameters and generating reverb at smaller intervals is more efficient. Therefore, the extraction and generation process has lower computational requirements.
[0022] In one embodiment, the method further includes detecting a third audio signal corresponding to speech by one or more microphones in the vehicle. The method also includes capturing the third audio signal by the one or more microphones. The method further includes: adjusting the at least one parameter based on the third audio signal; and generating the artificial reverberation, wherein the artificial reverberation includes the adjusted at least one parameter.
[0023] Advantageously, artificial reverberation can compensate for any voice audio (e.g., one or more users speaking loudly inside the vehicle) and provide the desired artificial reverberation independent of annoying voice audio within the vehicle cabin.
[0024] In a preferred embodiment, a system is provided. The system includes a plurality of speakers and a processor. The processor is operable to: receive an audio signal; extract reverberation characteristics from the audio signal, the reverberation characteristics including a plurality of parameters; and generate artificial reverberation, the artificial reverberation including at least one of the parameters. The processor is also operable to: apply the artificial reverberation to the audio signal; and send a modified audio signal to the plurality of speakers, the modified audio signal including the artificial reverberation applied to the audio signal.
[0025] Advantageously, a system that can be placed in enclosed spaces (e.g., vehicles) is provided, offering advantages over conventional multi-speaker systems. By dynamically reproducing the natural reverberation characteristics of the original audio signal, the invention improves the overall fidelity of the reproduced sound. Advantageously, it enhances the fidelity and immersion of audio reproduction in multi-speaker setups (i.e., setups including two or more speaker channels). Specifically, by extracting the original reverberation characteristics (reverberation) from the originally recorded or produced audio and applying artificial reverberation based on those extracted reverberation characteristics, a reverberation that closely mimics the original recording (e.g., reproducing the reverberation experienced in a concert hall with the original recorded audio) is generated, since the reverberation is generated based on the characteristics of the original audio signal. High-quality, stable artificial reverberation is generated that is very similar to the natural reverberation characteristics extracted from the original audio signal. This is advantageous in enclosed spaces with multi-speaker setups (such as vehicles), where audio very similar to the original recording and accurately reproduced spatial cues can be reproduced, thus providing an immersive listening experience for users of the vehicle or enclosed space.
[0026] In the implementation, the processor is also operable to: sample the audio signal at a sampling rate, wherein the sampling rate is at a predetermined frequency or a dynamically adjusted frequency; and extract the reverberation characteristics from the audio signal at each sample of the audio signal.
[0027] By sampling the audio signal at a single sampling rate and extracting reverberation at each sample of the audio signal, rather than extracting reverberation continuously, the computational requirements of the computer / processor are significantly reduced. Furthermore, dynamically adjusting the sampling rate (or update frequency) of the extracted reverberation parameters allows the complexity of the extracted reverberation parameters to be adjusted based on the type of audio signal being fed into the system at a given time. For example, a church choir recording may include a large amount of reverberation, while a singer's studio recording may include less reverberation. Computational requirements can also be improved by dynamically adjusting the sampling rate (e.g., by reducing the number of samples of audio signals containing less reverberation and by increasing the sampling rate of audio signals containing more reverberation). This makes it possible to adjust the parameter extraction frequency based on the complexity of the source audio and allows for efficient processing and optimization for different musical styles.
[0028] Therefore, the ability to dynamically adjust the sampling or update frequency (asymmetric interval) of the extracted reverberation parameters based on the complexity of the source audio allows for optimized processing efficiency, resulting in a more efficient system overall. Thus, compared to signal extraction techniques, this invention allows for the production of synthetic reverberation with superior audio quality while maintaining a high degree of fidelity to the original recording's natural reverberation. This enables the creation of unique tuning scenarios based on the received audio signal.
[0029] In a preferred embodiment, a vehicle is provided that includes the system described above.
[0030] Advantageously, an immersive listening experience can be recreated in the vehicle. Vehicle users can enjoy music in the vehicle as if they were listening to music live (such as in a concert hall, jazz club, etc.).
[0031] In one embodiment, the plurality of speakers includes a plurality of first speakers coupled to a first channel and a plurality of second speakers coupled to a second channel, wherein the modified audio signal includes a first audio signal and a second audio signal different from the first audio signal, and the processor is also operable to send the first audio signal to the first channel and send the second audio signal to the second channel.
[0032] Advantageously, the system described above can be applied to multi-channel setups (such as stereo and other surround sound systems). By providing a first audio signal and a second audio signal, the synthesized reverberation signal can be individually assigned and processed for each channel within a multi-speaker system and for each individual speaker in each channel. This processing can involve techniques known in the fields of system tuning and sound design to ensure optimal spatialization and blending with the original audio signal. By individually assigning and processing the reverberation signal for each channel and for each speaker, the phantom center of the audio can be adjusted more accurately. This also provides an arrangement in which the perceived width of the stage size of the audio recording can be increased or decreased more accurately. For example, both audio recordings for large stages (such as recordings in large concert halls) and audio recordings for small venues (such as recordings in small jazz halls), as well as any audio recordings in between, can be accurately mimicked.
[0033] In one embodiment, the one or more first speakers are speakers facing forward relative to the user of the vehicle, and the one or more second speakers are speakers facing sideways or rearward relative to the user of the vehicle.
[0034] Advantageously, the raw audio signal can be dynamically allocated to create a more immersive listening experience for the user.
[0035] In one implementation, the first audio signal includes the received audio signal, and the second audio signal includes the modified audio signal.
[0036] Advantageously, the original audio signal can be dynamically distributed through the front speakers to create a stereo foundation, while reproduced synthetic reverberation is added to the side and / or rear speakers. This approach creates a truly immersive listening experience, allowing the listener to perceive the acoustic environment of the original recording as if they were there.
[0037] In one embodiment, the vehicle includes at least one microphone. The microphone is operable to detect and capture a third audio signal corresponding to speech. The processor is also configured to adjust the at least one parameter based on the third audio signal. The processor is further configured to generate the artificial reverberation, wherein the artificial reverberation includes the adjusted at least one parameter.
[0038] Advantageously, artificial reverberation can compensate for any voice audio (e.g., one or more users speaking loudly inside the vehicle) and provide the desired artificial reverberation independent of annoying voice audio within the vehicle cabin.
[0039] Advantageously, one or more microphones can capture speech inside the car cabin, process it using synthetic reverberation, and enhance the perception of presence inside the vehicle cabin. Attached Figure Description
[0040] The features, objects, and advantages of this disclosure will become more apparent from the following detailed description set forth in conjunction with the accompanying drawings, in which similar reference numerals refer to similar elements.
[0041] Figure 1 An example flowchart of a system and method for modifying an audio signal according to the present invention is shown, the flowchart including an input stream and an output stream for modifying the audio signal;
[0042] Figure 2 A system according to the present invention is shown, comprising a computer including a processor and memory, multiple speakers, and multiple microphones;
[0043] Figure 3 The invention includes Figure 2 The vehicle has multiple speakers and a computer;
[0044] Figure 4 A flowchart illustrating a method for modifying audio signals according to the present invention is depicted;
[0045] Figure 5 A flowchart depicts an alternative method for modifying an audio signal according to an alternative embodiment of the present invention; and
[0046] Figure 6 A flowchart depicts an alternative method for modifying an audio signal according to an alternative embodiment of the present invention. Detailed Implementation
[0047] This invention belongs to the field of audio reproduction systems, particularly those employing multiple speakers with two or more channels, designed to improve the fidelity of reproduced sound and sound field, and the perceived immersion (in other words, to enhance the perception of sound location and to mimic the experience of recording an audio signal at the point of recording). This invention is applicable to automotive audio systems and has broader implications for any multi-speaker configuration.
[0048] Figure 1 An example flowchart of a system and method for modifying an audio signal according to the present invention is shown. The flowchart includes input and output streams for modifying the audio signal. The method is executed by a computer, which includes, for example... Figure 2The processor and memory mentioned herein. The computer is coupled to a plurality of speakers 114(a), ... 114(n) (referred to herein as 114n) and may be coupled to one or more microphones 116(a), ... 116(n) (referred to herein as 116n). The method includes receiving input, such as audio signal 102 (which may be a multi-channel audio input signal), at the processor. The processor is part of the computer, as referenced below. Figure 2 As defined, the audio signal can be a live recording sensed and recorded by one or more microphones (not shown), which are connected or coupled (e.g., via wireless or wired connections) to a processor. Alternatively, the audio signal can be a media file comprising pre-recorded audio recordings from a database (e.g., an online database, or a local database, such as those stored in the reference below). Figure 2 Access is made in the local memory of the defined computer or any other computer connected or coupled to it. The media file can be any suitable audio format known in the art (e.g., MP3, AAC, WAV, or any other suitable audio format). After receiving the audio signal, the processor may perform a preprocessing step 104, which may include one or more filtering steps of the audio signal, a Fast Fourier Transform (FFT), and any other preprocessing steps required to provide frequency information about the audio signal.
[0049] After receiving the audio signal 102, the reverberation characteristics are extracted from the audio signal 102 (e.g., by means of the following text). Figure 2 The processor mentioned in the document). Extracting reverberation characteristics may include (for example, by means of processors mentioned in the document). Figure 2The processor mentioned herein applies a reverberation information retrieval (RIR) algorithm 106 to the audio signal 102. The RIR algorithm 106 combines deterministic processing techniques and a machine learning model to extract multiple parameters 108a, 108b, ... 108n (referred to herein as 108n to refer to all multiple parameters 108a, 108b, ... 108n). Deterministic processing techniques may include applying one or more of the following to the audio signal: spectral and temporal analysis, one or more audio suppression techniques, spectral editing, one or more subtraction techniques, noise gating, impulse response extraction, linear predictive coding, etc. The machine learning model may be able to clean up the extracted reverberation features, for example, by extracting reverberation features and feeding these features into one or more machine learning (and / or artificial intelligence) models to reproduce the multiple parameters 108n. The multiple parameters 108n comprehensively describe the reverberation embedded within the audio. The RIR algorithm 106 can extract reverberation based on models and processing techniques. Multiple parameters can be extracted based on IR analysis, attenuation curve calculation (e.g., by Schroeder integration), statistical processing (such as cross-correlation), short-time Fourier transform for spectrum-time analysis, peak detection for early reflections, etc. 108
[0050] The extracted reverberation characteristics include multiple parameters 108n. Parameter 108n may include reverberation time (e.g., the length of the reverberation, which may be RT60, RT30, or any other length of the reverberation length descriptor), pre-delay, attenuation shape, high-frequency and low-frequency damping, venue size, diffusion, width, modulation, spectral density, spectral envelope (the frequency-dependent behavior of the reverberation), early reflection patterns in frequency and time, or any combination of the above. Advantageously, multiple different types of reverberation can be considered, thereby producing a more realistic reverberation reproduction from the received original audio signal.
[0051] After extracting the reverberation characteristics, an artificial reverberation (e.g., a synthesized reverberation filter) is generated, which includes at least one of a plurality of parameters 108n. The artificial reverberation can be any suitable type of audio filter. Therefore, the artificial reverberation partially or completely reproduces the extracted reverberation characteristics of the original audio signal 102. The artificial reverberation can be generated by a synthesized reverberation engine / processor 110, which can be, for example,... Figure 2This refers to a portion of the computer as defined in [the specification] (i.e., utilizing a portion of the processor, memory, and input / output ports). Artificial reverberation may include only one of the parameters 108n, some of the parameters 108n (i.e., two or more parameters), or all of the parameters 108n. Therefore, reverberation filters can be generated to produce any type of desired reverberation effect. For example, this may be limited to one type of reverberation (e.g., adapting only to reverberation time and / or spectral density), or it may encompass multiple reverberation features (e.g., including spectral envelope, early reflection modes in frequency and time, RT60 or RT30, or any combination of these features).
[0052] The generated artificial mixing response is applied to audio signal 102 to produce a modified audio signal. The modified audio signal (e.g., through...) Figure 2 A processor (like the one in the example) sends signals to multiple speakers 114(a) … 114(n) (referred to herein as 114n) to reproduce the modified audio signal. Speakers 114n may be part of a single playback device (e.g., a HiFi set, a speaker array in a vehicle, or a similar device), or may be part of multiple playback devices (e.g., multiple individual HiFi sets communicating with each other). Speakers 114n may be any suitable device capable of reproducing sound. Any speaker in speakers 114n may include one or more drivers. A playback device including multiple speakers 114n may be a multi-channel playback system (e.g., a stereo system, a surround sound system, or a similar system).
[0053] This device offers several advantages over conventional multi-speaker systems. By dynamically reproducing the natural reverberation characteristics of the original audio signal, the invention improves the overall fidelity of the reproduced sound. Advantageously, it enhances the fidelity and immersion of audio reproduction in multi-speaker systems (i.e., systems comprising two or more speaker channels). Specifically, by extracting the original reverberation characteristics (reverberation) from the originally recorded or produced audio and applying artificial reverberation based on those extracted reverberation characteristics, a reverberation that closely mimics the reverberation of the original recording is generated (e.g., reproducing the reverberation experienced in a concert hall based on the characteristics of the original audio signal), since the reverberation is generated based on the characteristics of the original audio signal. High-quality, stable artificial reverberation is generated through artificial reverberation that is very similar to the natural reverberation characteristics extracted from the original audio signal. This is advantageous in enclosed spaces with multi-speaker systems (such as in vehicles), where audio that is very similar to the original recording and spatial cues can be reproduced, thus providing an immersive listening experience for users of the vehicle or enclosed space.
[0054] In implementations, the modified audio signal may include multiple different audio signals. For example, the modified audio signal may include a first audio signal and a second audio signal different from the first audio signal. As described above, multiple speakers 114n may be configured in a multi-channel arrangement (e.g., a stereo system, a surround sound system, or a similar system). The modified audio signal is then transmitted from... Figure 2 The computer sending the signal to the plurality of speakers 114n as defined herein may further include sending a first audio signal to a first channel of a multi-channel arrangement, the first channel including one or more first speakers 114a among the plurality of speakers 114n. Sending the modified audio signal to the plurality of speakers 114n may further include sending a second audio signal to a second channel of a multi-channel arrangement, the second channel including one or more second speakers 114b among the plurality of speakers 114n.
[0055] Advantageously, the modified signal, as described above, can be applied to multi-channel setups (such as stereo and other surround sound systems). By providing a first audio signal and a second audio signal, the synthesized reverberation signal can be individually assigned and processed for each channel within a multi-speaker system and for each individual speaker within each channel. This processing can involve system tuning and sound design to ensure optimal spatialization and blending with the original audio signal. By assigning and processing the reverberation signal individually for each channel and for each speaker, the phantom center of the audio can be adjusted more accurately. This also provides an arrangement in which the perceived width of the stage size of the audio recording can be increased or decreased more accurately. For example, both audio recordings for large stages (such as recordings in large concert halls) and audio recordings for small venues (such as recordings in small jazz halls), as well as any audio recordings in between, can be accurately mimicked.
[0056] The first audio signal can be the received audio signal (wherein the received audio signal does not include any spatial and / or reverberation processing), and the second audio signal can be a modified audio signal that may include virtually all artificial reverberation. Advantageously, the original audio signal can be distributed through the front speakers to create a stereo basis, while the reproduced synthetic reverberation is added to the side and / or rear speakers. This approach creates a truly immersive listening experience, allowing the listener to perceive the acoustic environment of the original recording as if they were there.
[0057] In one embodiment, the synthetic reverb engine 110 may apply at least one additional filter 112 to the audio signal before sending the modified audio signal to a plurality of speakers 114n. The additional filter 112 may include an equalizer for the audio signal, a delay of the audio signal, a gain of the audio signal, or any combination thereof. At least one additional filter 112 may be applied to a first audio signal (i.e., the received audio signal), a second audio signal (i.e., the modified signal), any other audio signal, or to any one or more audio signals within the modified audio signal. The additional filter 112 may be different for the audio signals within the modified audio signal. Advantageously, a variety of different effects and filters may be added to the audio signal to produce a more realistic reproduction of the original audio signal.
[0058] In some implementations, the system may include an audio matrix adder 113 (e.g., an audio matrix, matrix mixer, digital audio matrix, or any other suitable mixer). The audio matrix adder 113 may be coupled to a computer (e.g., as described below). Figure 2 The computer 202 described herein is connected to a plurality of speakers 114. An audio matrix adder 113 is operable to receive the input audio signal 102 and the modified audio signal before sending the input audio signal 102 or the modified audio signal to the plurality of speakers 114. The audio matrix adder 113 is operable to mix the input audio signal 102 and the modified audio signal, and subsequently distribute the mixed audio signal (which includes the input audio signal 102 and the modified audio signal) to each of the plurality of speakers. The mixed audio signal may include multiple different types of audio signals. Therefore, different audio signals may be sent to different speakers among the plurality of speakers. In an embodiment, the mixed audio signal may include a first audio signal and a second audio signal, as described above. In an embodiment, an additional filter 112 may be integrated within the audio matrix adder 113.
[0059] In one implementation, the synthetic reverb engine 110 may first apply a predefined reverb (e.g., a predefined artificial reverb) to the audio signal 102, and then apply the generated artificial reverb. The predefined reverb may take into account one or more of the parameters 108n described above. The parameters 108n of the predefined reverb may be predetermined individually, rather than based on extracted reverb characteristics. Multiple different predefined reverbs may be stored in the computer's memory (as described below). Figure 2(As mentioned in the text). Each predefined reverb can be applied to different scenarios. For example, the first predefined reverb can be used for classical music, the second for jazz, the third for studio recordings, the fourth for live concert hall recordings, and the fifth for combinations of the above, etc. Processors (such as...) Figure 2 The predefined reverb (as defined in the document) can determine which predefined reverb should be applied to the received audio signal 102 based on the metadata of the audio signal (e.g., metadata received from the network that sent the audio signal, or metadata stored on a physical copy associated with the audio signal). The metadata may be related to genre, theme, recording location, artist name, etc. Alternatively, the user can select a predefined reverb.
[0060] A predefined reverb response can be applied to an audio signal to produce a modified audio signal that includes the received audio signal and the predefined reverb. The synthetic reverb engine 110 can compare this modified audio signal with an extracted reverb characteristic of the received audio signal. The synthetic reverb engine 110 can detect one or more differences between the parameters 108n of the predefined reverb and the parameters 108n of the extracted reverb characteristic. Subsequently, the synthetic reverb engine 110 can generate artificial reverb based on the differences between the parameters 108n of the predefined reverb and the parameters 108n of the extracted reverb characteristic.
[0061] By first applying predefined reverb (which may be based, for example, on music genre, style, or other metadata), a significant computational load is removed from the system. This is advantageous, for example, for vehicles that may not have high-performance computers / processors and other computing power. For instance, an audio signal identified as a first-genre (such as a live rock concert) could have a predefined filtered reverb filter applied to it (e.g., including standard reverb for a specific indoor or outdoor venue and providing emphasis on certain reverb parameters related to guitar, vocals, and drums). To provide a fully immersive experience, reverb characteristics can be extracted at a lower rate (e.g., at a low sampling rate, or by extracting only certain parameters instead of all parameters). Artificial reverb can then be applied on top of the predefined filter to provide a fully immersive listening experience. Advantageously, the extraction and generation process has lower computational requirements.
[0062] In the implementation, the processor can sample the audio signal 102 at a sampling rate, where the sampling rate is at a predetermined frequency or at a dynamically adjusted frequency. When the load on the computer (i.e., the processor, memory, or a combination of both) is low, the frequency can be increased (i.e., more samples are taken per unit time). When the load on the computer is high, the frequency can be decreased (i.e., fewer samples are taken per unit time). The frequency can be dynamically adjusted to reflect the current load on the computer and the predicted load on the computer. For more complex reverberation in the received audio signal, the frequency can be increased. For less complex reverberation in the received audio signal, the frequency can be decreased. The RIR algorithm extracts reverberation characteristics from the audio signal at each sample of the audio signal 102. Therefore, the RIR algorithm extracts multiple parameters 108n at predetermined frequencies, which can be fixed or dynamically adjusted based on the complexity of the reverberation in the source audio. Music genres with the least variation in reverberation characteristics (such as classical or acoustic music) may require lower frequency parameter extraction compared to genres with more complex and dynamic reverberation effects.
[0063] By sampling the audio signal at a single sampling rate and extracting reverberation at each sample of the audio signal, rather than extracting reverberation continuously, the computational requirements of the computer / processor are significantly reduced. Furthermore, dynamically adjusting the sampling rate (or update frequency) of the extracted reverberation parameters allows the complexity of the extracted reverberation parameters to be adjusted based on the type of audio signal being fed into the system at a given time. For example, a church choir recording may include a large amount of reverberation, while a singer's studio recording may include less reverberation. Computational requirements can also be improved by dynamically adjusting the sampling rate (e.g., by reducing the number of samples of audio signals containing less reverberation and by increasing the sampling rate of audio signals containing more reverberation). This makes it possible to adjust the parameter extraction frequency based on the complexity of the source audio and allows for efficient processing and optimization for different musical styles.
[0064] Therefore, the ability to dynamically adjust the sampling or update frequency (asymmetric interval) of the extracted reverberation parameters based on the complexity of the source audio allows for optimized processing efficiency, resulting in a more efficient system overall. Thus, compared to signal extraction techniques, this invention allows for the production of synthetic reverberation with superior audio quality while maintaining a high degree of fidelity to the original recording's natural reverberation. This enables the creation of unique tuning scenarios based on the received audio signal.
[0065] In one implementation, one or more microphones 116n may be coupled (wirelessly or wired) to the synthetic reverberation engine 110. The one or more microphones 116n may detect a third audio signal corresponding to speech, and subsequently capture (i.e., record) the third audio signal and send it to the synthetic reverberation engine 110. The synthetic reverberation engine 110 may adjust at least one parameter 108 based on the third audio signal. For example, speech within a vehicle may adversely affect the reverberation characteristics within the vehicle cabin (e.g., by distorting the reverberation). The synthetic reverberation engine 110 may generate artificial reverberation that compensates for speech within the vehicle cabin and thus provides artificial reverberation that mimics the extracted reverberation characteristics / parameters, independent of the speech within the cabin. The synthetic reverberation engine 110 may generate artificial reverberation, wherein the artificial reverberation includes at least one adjusted parameter. Advantageously, the artificial reverberation may compensate for any speech audio (e.g., one or more users speaking loudly within the vehicle) and provide the desired artificial reverberation, independent of annoying speech audio within the vehicle cabin.
[0066] Figure 2 An exemplary system of computer 202 is shown, including processor 204, memory 206, and input / output (I / O) interface 208. Computer 202 (via wired or wireless connection) is coupled to the above-referenced... Figure 1 Multiple speakers 114n are defined. Computer 202 can be coupled (via wired or wireless connection) to the above reference. Figure 1 One or more microphones 116n are defined. Processor 204 is operable to execute a set of actions and / or instructions that can be stored in memory 206. Computer 202 may be dedicated to performing actions as described above. Figure 1 Independent units of methods and steps defined in [the document]. Besides, as mentioned above... Figure 1 In addition to the methods and steps defined herein, computer 202 may also be able to operate to perform multiple individual features. For example, computer 202 may be an electronic control unit (ECU) of the vehicle or a part of an ECU (e.g., communicating directly with the ECU via a wireless or wired connection).
[0067] Processor 204 can be applied as described above. Figure 1 The RIR algorithm 106 is defined in [the document]. The processor 204 may include [the components described above]. Figure 1 The synthesis reverb engine 110 as defined herein is available, and the steps performed by the synthesis reverb engine 110 are available. The I / O interface 208 may include one or more input ports and one or more output ports. The one or more input ports may include wired and / or wireless connections operable to receive signals as described above. Figure 1The input audio signal 102 is defined. One or more input ports are wired or wirelessly connected to a network and are operable to receive one or more instructions from the network. One or more input ports are operable to receive input signals from one or more microphones 116n. One or more output ports are wired or wirelessly connected to one or more speakers 114n and are operable to transmit or deliver audio signals (such as the input audio signal 102, the modified audio signal as defined above, or any other type of audio signal) from computer 202 to one or more speakers 114n.
[0068] In one implementation, processor 204 is operable to receive audio signal 102 and extract reverberation characteristics from audio signal 102 (e.g., by applying RIR algorithm 106), wherein the reverberation characteristics include a plurality of parameters 108n. Processor 204 is also operable to generate (e.g., by synthesizing reverberation engine 110) artificial reverberation, which includes at least one of the parameters 108n. Processor 204 is also operable to apply the artificial reverberation to audio signal 102. In other words, processor 204 is operable to generate a modified audio signal that includes the artificial reverberation applied to audio signal 102. Processor 204 is also operable to send the modified audio signal to a plurality of speakers 114n (e.g., by sending the modified audio signal to one or more output ports of I / O interface 208).
[0069] Advantageously, a system that can be placed in enclosed spaces (e.g., vehicles) is provided, offering advantages over conventional multi-speaker systems. By dynamically reproducing the natural reverberation characteristics of the original audio signal, the invention improves the overall fidelity of the reproduced sound. Advantageously, it enhances the fidelity and immersion of audio reproduction in multi-speaker setups (i.e., setups including two or more speaker channels). Specifically, by extracting the original reverberation characteristics (reverberation) from the originally recorded or produced audio and applying artificial reverberation based on those extracted reverberation characteristics, a reverberation that closely mimics the original recording (e.g., reproducing the reverberation experienced in a concert hall with the original recorded audio) is generated, since the reverberation is generated based on the characteristics of the original audio signal. High-quality, stable artificial reverberation is generated through artificial reverberation that is very similar to the natural reverberation characteristics extracted from the original audio signal. This is advantageous in enclosed spaces with multi-speaker setups (such as vehicles), where audio very similar to the original recording and accurately reproduced spatial cues can be reproduced, thus providing an immersive listening experience for users of the vehicle or enclosed space.
[0070] In the implementation scheme, processor 204 is also operable to sample audio signal 102 at a sampling rate, wherein the sampling rate is at a predetermined frequency or at a dynamically adjusted frequency, and to extract reverberation characteristics from the audio signal at each sample of the audio signal, as referenced above. Figure 1 Discussed.
[0071] By sampling the audio signal at a single sampling rate and extracting reverberation at each sample of the audio signal, rather than extracting reverberation continuously, the computational requirements of the computer / processor are significantly reduced. Furthermore, dynamically adjusting the sampling rate (or update frequency) of the extracted reverberation parameters allows the complexity of the extracted reverberation parameters to be adjusted based on the type of audio signal being fed into the system at a given time. For example, a church choir recording may include a large amount of reverberation, while a singer's studio recording may include less reverberation. Computational requirements can also be improved by dynamically adjusting the sampling rate (e.g., by reducing the number of samples of audio signals containing less reverberation and by increasing the sampling rate of audio signals containing more reverberation). This makes it possible to adjust the parameter extraction frequency based on the complexity of the source audio and allows for efficient processing and optimization for different musical styles.
[0072] Therefore, the ability to dynamically adjust the sampling or update frequency (asymmetric interval) of the extracted reverberation parameters based on the complexity of the source audio allows for optimized processing efficiency, resulting in a more efficient system overall. Thus, compared to signal extraction techniques, this invention allows for the production of synthetic reverberation with superior audio quality while maintaining a high degree of fidelity to the original recording's natural reverberation. This enables the creation of unique tuning scenarios based on the received audio signal.
[0073] Figure 3 Vehicle 302 is shown, which includes: Figure 2 The computer 202 defined in the document; multiple speakers 306a, 306b, 306c inside the vehicle compartment of vehicle 302, the multiple speakers 306a, 306b, 306c and Figure 1 and Figure 2 The multiple speakers 114n defined in the document correspond to the multiple seats 304a, 304b, 304c, and 304d for vehicle occupants. Figure 3 An exemplary arrangement of four seats 304a, 304b, 304c, and 304d in a 2x2 configuration for vehicle occupants is depicted. However, the invention is not limited to four seats or as shown in the illustration. Figure 3The arrangement shown can include a single seat for vehicle occupants, or any number of seats for multiple vehicle occupants. Therefore, a vehicle can be any land, water, or air vehicle with an enclosed space for any number of occupants, such as (but not limited to) automobiles, buses, trucks, aircraft, boats, ships, hovercraft, etc. Computer 202 in Figure 3 The computer 202 is depicted as being positioned at the center of vehicle 302. However, the invention is not limited to this arrangement, and the computer 202 can be placed anywhere within vehicle 302, enabling the computer 202 to communicate (via wired or wireless connection) with multiple speakers 306a, 306b, 306c, and enabling the computer to communicate with other speakers within the vehicle 302, such as those in the passenger compartment. Figure 1 and Figure 2 Defined in ( Figure 3 One or more microphones 116n (not shown in the image) communicate.
[0074] Figure 3 The illustration depicts two speakers 306a located at the front of the vehicle compartment (i.e., facing forward relative to one or more occupants of the vehicle 302), two speakers 306c located at the rear of the vehicle compartment (i.e., facing rearward relative to one or more occupants of the vehicle 302), two speakers 306b located on the left side of the vehicle compartment (i.e., facing left relative to one or more occupants of the vehicle 302), and two speakers 306b located on the right side of the vehicle compartment (i.e., facing right relative to one or more occupants of the vehicle 302). Figure 3 The eight-speaker configuration shown is for illustrative purposes, and the invention is not limited to this. Figure 3 The invention may include any number (i.e., one or more) of front speakers 306a, rear speakers 306b, left speakers 306b, or right speakers 306b. The invention may also include only front speakers 306a, rear speakers 306b, left speakers 306b, right speakers 306b, or any combination thereof. The invention may include speakers located in... Figure 3 Different speakers in locations not shown (e.g., below one or more vehicle occupants, above one or more vehicle occupants, diagonally opposite one or more vehicle occupants, etc.).
[0075] By providing, including, Figure 2 The computer 202 and multiple speakers 306a, 306b, 306c defined in the document can reproduce an immersive listening experience in vehicle 302. Users of vehicle 302 can enjoy music in vehicle 302 as if they were listening to music live (such as in a concert hall, jazz club, etc.).
[0076] In an implementation, the plurality of speakers 306a, 306b, 306c may include multiple channels. For example, the plurality of speakers 306a, 306b, 306c may be divided into a stereo setup, wherein a speaker on the left side of the vehicle compartment reproduces a first version of a modified audio signal (as defined above), and a speaker on the right side of the vehicle compartment reproduces a second version of the modified audio signal to reproduce the stereo setup. The plurality of speakers 306a, 306b, 306c may include a plurality of first speakers (such as left speaker 306b, one or more of the front speakers 306a on the left side of the vehicle compartment, one or more of the rear speakers 306c on the left side of the vehicle compartment, or any combination thereof), which are coupled to a first channel; and a plurality of second speakers (such as right speaker 306b, one or more of the front speakers 306a on the right side of the vehicle compartment, one or more of the rear speakers 306c on the right side of the vehicle compartment, or any combination thereof), which are coupled to a second channel. The invention is not limited to this arrangement and may include any number of channels, for example, to reproduce surround sound settings (such as 5.1, 7.1, 9.1, or similar surround sound settings). Modified audio signals (as described above) Figure 1 and Figure 2 As defined in [the document], it may include a first audio signal and a second audio signal that is different from the first audio signal, and the processor 204 of the computer 202 may be operable to send the first audio signal to the first channel and send the second audio signal to the second channel.
[0077] Advantageously, the system described above can be applied to multi-channel setups (such as stereo and other surround sound systems). By providing a first audio signal and a second audio signal, the synthesized reverberation signal can be individually assigned and processed for each channel within a multi-speaker system and for each individual speaker in each channel. This processing can involve techniques known in the fields of system tuning and sound design to ensure optimal spatialization and blending with the original audio signal. By individually assigning and processing the reverberation signal for each channel and for each speaker, the phantom center of the audio can be adjusted more accurately. This also provides an arrangement in which the perceived width of the stage size of the audio recording can be increased or decreased more accurately. For example, both audio recordings for large stages (such as recordings in large concert halls) and audio recordings for small venues (such as recordings in small jazz halls), as well as any audio recordings in between, can be accurately mimicked.
[0078] Alternatively or additionally, the plurality of speakers 306a, 306b, 306c may be divided into a plurality of forward-facing speakers relative to the vehicle occupants (e.g., one or more of a front speaker 306a, a left speaker and a right speaker 306b positioned in front of the vehicle compartment relative to the occupants of vehicle 302, or any combination thereof), said forward-facing speakers being coupled to a first channel; and a plurality of rearward-facing speakers relative to the vehicle occupants (e.g., one or more of a rear speaker 306c, a left speaker and a right speaker 306b positioned in a rearward position relative to the occupants of vehicle 302, or any combination thereof), said rearward-facing speakers being coupled to a second channel. Modified audio signals (as described above in...) Figure 1 and Figure 2 The audio signal (as defined above) may include a first audio signal and a second audio signal different from the first audio signal, and the processor 204 of the computer 202 may be operable to send the first audio signal to a first channel and the second audio signal to a second channel. Depending on the number of speaker channels, the plurality of speakers 306a, 306b, 306c may include speakers with more than two channels, and the modified audio signal may include more than two audio signals. In an embodiment, the plurality of speakers 306a, 306b, 306c may be divided into a plurality of forward-facing speakers (as defined above) coupled to a first channel; and a plurality of side-facing speakers (such as speaker 306b) relative to the vehicle occupants, the side-facing speakers coupled to a second channel. The processor 204 may be operable to send the first audio signal to a first audio channel and the second audio signal to a second channel.
[0079] Advantageously, the raw audio signal can be dynamically allocated to create a more immersive listening experience for the user.
[0080] In the implementation scheme, the first audio signal may include the received audio signal 102, and the second audio signal includes as described above. Figure 1 and Figure 2The modified audio signal is defined in [the document]. The first audio signal may include a portion of the received audio signal 102 and the modified audio signal (e.g., the modified audio signal includes the received audio signal 102 and a generated artificial reverberation applied to the received audio signal at a lower intensity). The second audio signal may include a dominant reproduction of the reverberation generated by the artificial reverberation. This may include the reproduction of the modified audio signal, which includes the generated artificial reverberation at a relatively high intensity (or full intensity) and the received audio signal at a relatively low intensity. In an embodiment, the first audio signal may be reproduced by speakers facing rearward relative to the vehicle occupants, and the second audio signal may be reproduced by speakers facing forward relative to the vehicle occupants. Advantageously, the original audio signal may be dynamically distributed via the front speakers to create a stereo basis, while the reproduced synthetic reverberation is added to the side and / or rear speakers. This approach creates a truly immersive listening experience, allowing the listener to perceive the acoustic environment of the original recording as if they were there.
[0081] Figure 4 The invention is illustrated as described above. Figure 1 , Figure 2 and Figure 3 A flowchart of the described method 400 for modifying an audio signal is provided. The method includes receiving an audio signal (such as audio signal 102 as described above) at 402. At 404, the method includes extracting reverberation characteristics from the audio signal (e.g., using RIR algorithm 106 as described above), the reverberation characteristics including multiple parameters (such as parameter 108n). The method includes generating artificial reverberation at 406 (e.g., via a synthetic reverberation engine 110, which may be part of processor 204 as described above), the artificial reverberation including at least one of the parameters. The method further includes: applying the artificial reverberation to the audio signal 102 at 408; and sending the modified audio signal, including the artificial reverberation applied to the audio signal 102, to multiple speakers at 410.
[0082] Figure 5 and Figure 6 Additional method steps for modifying an audio signal are shown, which may be related to those described above. Figure 4 The methods described and as referenced above Figure 1 , Figure 2 and Figure 3 The described arrangement is combined.
[0083] exist Figure 5The method may include sending a first audio signal at 502 to a first channel of a plurality of audio channels (as described above), the first channel including one or more first speakers of a plurality of loudspeakers (e.g., loudspeakers 114n, 306a, 306b, 306c as described above), wherein the modified audio signal includes the first audio signal and a second audio signal different from the first audio signal. The method may include sending a second audio signal at 504 to a second channel, the second channel including one or more second loudspeakers of the plurality of loudspeakers 114n, 306a, 306b, 306c. In an embodiment, the first audio signal includes the received audio signal, and the second audio signal includes the modified audio signal.
[0084] The method may further include sampling the audio signal at a sampling rate at 506, wherein the sampling rate is at a predetermined frequency or at a dynamically adjusted frequency. The method may further include extracting reverberation from the audio signal at each sample of the audio signal at 508.
[0085] Applying an artificial mixing response to an audio signal may also include: applying a predefined mixing response to the audio signal at 510; and subsequently applying the generated artificial mixing response to the audio signal at 512.
[0086] exist Figure 6 In this method, at 602, the method may detect a third audio signal corresponding to speech (e.g., by one or more microphones in the vehicle as described above). At 604, the method may capture the third audio signal (by one or more microphones). At 606, the method may adjust at least one parameter based on the third audio signal. At 608, the method may generate artificial reverberation, wherein the artificial reverberation includes the adjusted at least one parameter. Advantageously, the artificial reverberation can compensate for any speech audio (e.g., one or more users speaking loudly in the vehicle) and can provide a desired artificial reverberation independent of annoying speech audio within the vehicle cabin.
Claims
1. A method for modifying an audio signal, the method comprising: Receive audio signals; Extract reverberation characteristics from the audio signal, the reverberation characteristics including multiple parameters; Generate artificial reverberation, wherein the artificial reverberation includes at least one of the parameters; The artificial mixing response is applied to the audio signal; as well as A modified audio signal is sent to multiple speakers, the modified audio signal including the artificial reverberation applied to the audio signal.
2. The method of claim 1, wherein the modified audio signal comprises a first audio signal and a second audio signal different from the first audio signal, the method further comprising: The first audio signal is sent to a first channel, the first channel including one or more first speakers among the plurality of speakers; as well as The second audio signal is sent to a second channel, which includes one or more second speakers among the plurality of speakers.
3. The method of claim 2, wherein: The first audio signal includes the received audio signal; and The second audio signal includes the modified audio signal.
4. The method according to any one of claims 1 to 3, further comprising: The audio signal is sampled at a sampling rate, wherein the sampling rate is at a predetermined frequency or at a dynamically adjusted frequency; as well as The reverberation characteristics are extracted from the audio signal at each sample of the audio signal.
5. The method of any one of claims 1 to 4, wherein the extracted reverberation parameters include: Reverberation time; Pre-delay; Attenuation shape; High-frequency and low-frequency damping; Size of the venue; diffusion; width; modulation; Spectral density; Spectral envelope; Early reflection patterns in frequency and time; or Any combination of the above.
6. The method of any one of claims 1 to 5, wherein applying the artificial reverberation further comprises applying: The equalizer for the audio signal; The delay of the audio signal; The gain of the audio signal; or Any combination of the above.
7. The method of claim 6, wherein applying the artificial reverberation further comprises: Apply a predefined mixing response to the audio signal; as well as The generated artificial mixed response is then used for the audio signal.
8. The method according to any one of claims 1 to 7, wherein the method further comprises: A third audio signal corresponding to speech is detected by one or more microphones in the vehicle; The third audio signal is captured by the one or more microphones; Adjusting the at least one parameter based on the third audio signal; and The artificial reverberation is generated, wherein the artificial reverberation includes at least one adjusted parameter.
9. A system comprising: Multiple speakers; as well as Processor, the processor being operable to: Receive audio signals; Extract reverberation characteristics from the audio signal, the reverberation characteristics including multiple parameters; Generate artificial reverberation, wherein the artificial reverberation includes at least one of the parameters; The artificial mixing response is applied to the audio signal; as well as A modified audio signal is sent to the plurality of speakers, the modified audio signal including the artificial reverb applied to the audio signal.
10. The system of claim 9, wherein the processor is further capable of operating to: The audio signal is sampled at a sampling rate, wherein the sampling rate is at a predetermined frequency or at a dynamically adjusted frequency; and The reverberation characteristics are extracted from the audio signal at each sample of the audio signal.
11. A vehicle comprising the system as claimed in any one of claims 9 or 10.
12. The vehicle of claim 11, wherein the plurality of speakers comprises a plurality of first speakers coupled to a first channel and a plurality of second speakers coupled to a second channel, and wherein the modified audio signal comprises a first audio signal and a second audio signal different from the first audio signal, and the processor is further operable to: Send the first audio signal to the first channel; and The second audio signal is sent to the second channel.
13. The vehicle as claimed in claim 12, wherein: The one or more first speakers are speakers facing forward relative to the user of the vehicle; and The one or more second speakers are speakers facing the side or rear relative to the user of the vehicle.
14. The vehicle as claimed in claim 13, wherein: The first audio signal includes the received audio signal; and The second audio signal includes the modified audio signal.
15. The vehicle as claimed in any one of claims 11 to 14, the vehicle further comprising: At least one microphone, the at least one microphone being operable to detect and capture a third audio signal corresponding to speech, wherein the processor is further configured to: Adjusting the at least one parameter based on the third audio signal; and The artificial reverberation is generated, which includes at least one adjusted parameter and a parameter based on the third audio signal.