Generation of vibrotactile signals from audio content for playback through haptic acoustic transducers
The method generates complementary transient and steady-state haptic signals from real-time audio to enhance audio systems, addressing computational inefficiencies and providing a natural and immersive multisensory experience.
Patent Information
- Application Number
- JP2025534152
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-12-13
- Publication Date
- 2026-01-07
AI Technical Summary
Existing methods for generating haptic feedback from audio signals are computationally intensive and inefficient, particularly in systems with limited digital signal processor capabilities, leading to unnatural or unsophisticated haptic experiences.
A method for generating complementary transient and steady-state haptic signals from a real-time audio stream using transient analysis, which involves extracting transient components and generating haptic drive signals to drive haptic actuators, optimizing the haptic experience while maintaining computational efficiency.
The method provides a computationally efficient and sophisticated haptic feedback that enhances the audio experience by aligning haptic and acoustic responses, offering a more natural and immersive multisensory experience.
Smart Images

Figure 2026500494000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the generation and manipulation of drive signals for vibrotactile acoustic transducers to enhance and / or extend the use of conventional audio speaker systems as a multi-sensory experience, particularly when mounted in a seat. [Background technology]
[0002] Haptics (from the Greek word haptikos, meaning "pertaining to the sense of touch"), as described in "Haptic Technology—Comprehensive Review Study with its Applications" by T. Jaiswal, R. Yadav, and P. Kedia in the April 2018 International Journal of Advance Research in Science and Engineering, is a type of technology that uses force applications to simulate tactile sensations. This technology has been widely applied in the gaming, virtual reality (VR), and film industries, primarily enhancing the rendered perceptual experience through vibration effects rendered by acoustic transducers mounted on objects that come into direct tactile contact with the user (such as game controllers, wearable devices, or seats). These vibration effects can enhance the immersive experience of film and VR experiences, and haptic feedback can also be used to enhance the experience of musical performances.
[0003] When listening to music at high sound pressure levels (SPL), listeners can derive tactile sensations from the low frequencies of the audio, which are expressed in the form of vibrations. To enhance the listening experience, it is desirable to replicate this sensation in home and car audio solutions using seat-mounted or wearable acoustic transducers. It is generally accepted that the human audible frequency ranges from 20 Hz to 20,000 Hz. As described in "Audio-Tactile Rendering: A Review on Technology and Methods to Convey Musical Information through the Sense of Touch" by B. Remache-Vinueza et al., Sensors, September 2021, humans can detect vibrations with frequencies ranging from 0.3 Hz to 1,000 Hz.
[0004] 1, there is an overlap 13 between the audible frequency spectrum 12 and the haptic sensation range 11, meaning that an audio signal containing content in this overlapping frequency band can be rendered by a seat- or body-mounted acoustic transducer to provide haptic enhancement to the input audio. However, because different parts of the body have varying sensitivity to tactile stimuli, this approach can sound unnatural to the listener.
[0005] First, a way to achieve more natural haptic characteristics is to achieve time alignment between traditional acoustic transducers, which perform the audible portion of the signal, and haptic actuators, which render the haptic feedback. Delay lines can be introduced to correct for time mismatches that can be caused by the location of various transducers in space and their relative proximity to the listener, as well as varying latencies through different parts of the audio system. While improving the temporal alignment of haptic and acoustic responses improves the overall sense of cohesion in the system, further improvements can be achieved through signal manipulation to achieve a more convincing haptic experience.
[0006] Although humans can experience tactile sensations of vibrations up to 1,000 Hz, it is not practical for haptic actuators to render frequencies up to such high frequencies. Mechanoreception is the human ability to detect stimuli, such as changes in pressure and touch, using mechanoreceptors, a type of nerve ending. One of the primary mechanoreceptive channels in the somatosensory system (also known as somatosensation, such as touch perception), responsible for perceiving vibration, is the Pacinian channel. Pacinian corpuscles are nerve endings responsible for the skin's sensitivity to vibration. According to Birnbaum DM and Wanderley MM, A Systematic Approach to Musical Vibrotactile Feedback; Proceedings of the International Computer Music Conference; Copenhagen, Denmark. August 27–31, 2007; pp. 397–404, Pacinian corpuscles have a sensitivity range of 40 Hz–500 Hz. This means that the frequency response of haptic actuators can be band-limited to this range.
[0007] Because the human ear is most sensitive to frequencies above 250 Hz (see, e.g., H. Fletcher and W. A. Munson, "Loudness, its definition, measurement, and calculation," Journal of the Acoustical Society of America 5, 1933, pp. 82-108), it is preferable for acoustic transducers to render haptic signals below this value so as not to distract from the frequency response of other components of the audio system. Thus, a potential frequency range of interest for haptic enhancement is 40 Hz-250 Hz.
[0008] As proposed by Y. Cho et al. in "Haptic Cushion: Automatic Generation of Vibro-tactile Feedback Based on Audio Signal for Immersive Interaction with Multimedia" at the 2014 International Conference on New Actuators and published in U.S. Patent No. 11,340,704 B2, synthesis of new signals characterized by the original audio can be used to generate haptic data across a desired frequency range. However, generating new content may be seen as excessive embellishment of the original audio, which is perhaps inappropriate here.
[0009] A. Sonza et al., "A whole body vibration perception map and associated acceleration loads at the lower leg, hip, and head," Medical Engineering and Physics, 2015, shows that different parts of the human body have different perceptual sensitivities to vibration. Therefore, to accommodate different sensitivity ranges, it is preferable for a haptic rendering system to distribute and weight different levels of haptic signals to different areas of the system that are in contact with the human body.
[0010] Bone conduction is another consideration for haptic systems, as seen in "CollarBeat: Whole Body Vibrotactile Presentation via the Collarbone to Enrich Music Listening Experience" by Sakuragi R., Ikeno S., Okazaki R., Kajimoto H. et al., Proceedings of the International Conference on Artificial Reality and Telexistence and Eurographics Symposium on Virtual Environments, 2015. Because bone conduction can be a factor influencing the perception of tactile sensations, it is desirable to be able to weight signals differently for drivers near areas of the human body where the effects of bone conduction are more or less noticeable.
[0011] Decomposing a signal and processing its individual components is important for advanced control of the haptic experience. As seen in U.S. Patent No. 11,340,704 B2, artificial intelligence (AI) can be used to extract harmonic and impulsive components and process them separately. However, using AI can be disadvantageous because it requires a significant amount of processing, which can be problematic when digital signal processor (DSP) capabilities are limited. The AI must be trained based on an extensive list of content material, including various genres of audio to be rendered through the system, which can be time-consuming and costly. Furthermore, such AI training can only be done for a finite range of content material, which can lead to problems, such as omitting certain genres when rendering audio with ambiguous genres, at the expense of overall performance.
[0012] The temporal aspects of audio signals can be classified into transients and steady-state signals. According to J.O. Smith, "Introduction to Digital Filters with Audio Applications," W3K Publishing, October 2007, ISBN 978-0-9745607-1-7, a transient can be defined as a sudden, broadband phenomenon in an otherwise steady-state signal. A transient can also be classified as a phenomenon in a signal in which the broadband energy of the signal, i.e., the energy in a specific frequency band, changes rapidly. This definition is sufficient to create algorithms for tracking transients. An example of such a transient can be seen against a steady-state background in Figure 2.
[0013] In a haptic system, transients may be felt by the user as a punch sensation, while steady-state signals may be interpreted as a vibration effect. Therefore, it is preferable to separate transients from steady-state signals in the audio stream that is rendered for haptics.
[0014] For example, when this technology is used to enhance infotainment systems in the automotive industry, the on-board processing capabilities of DSPs may be limited, so it is desirable to have a lightweight solution for generating haptic signals in real time from an input audio stream. It is important that this technology be computationally inexpensive so that the process can be performed on the vehicle's on-board DSP.
[0015] Therefore, there is a need for improved, computationally efficient methods of generating haptic signals from an input audio stream while enabling more sophisticated haptic feedback that can be used to enhance and / or extend the use of audible sound from conventional audio speaker systems to provide a multisensory experience. Summary of the Invention
[0016] According to one aspect of the present invention, there is provided a method of generating one or more haptic drive signals from an input signal representing real-time audio, the method comprising: receiving an input signal; performing a transient extraction process on the input signal to determine real-time audio transient components; generating a transient haptic signal T(t) and a steady-state haptic signal S(t) from the input signal based on a transient extraction process; The transient haptic signal and the steady-state haptic signal are complementary such that S(t) + T(t) = I(t), where I(t) is the input signal on which the transient extraction process was performed. and generating one or more haptic drive signals based on one or both of the transient haptic signal and the steady-state haptic signal; Includes.
[0017] Thus, the present invention provides a method for generating complementary transient and steady-state haptic signals from a real-time audio stream using transient analysis to estimate when transients occur, which are then used to generate one or more haptic drive signals that can be used to drive haptic actuators.
[0018] There are several ways in which transient extraction analysis can be performed, but broadly speaking, the transient estimate can be expressed as the relationship between a short-term (micro) dynamic envelope over a typical time frame (0ms-100ms) and a longer-term (macro) dynamic envelope over a typical time frame (200ms-1000ms).
[0019] The received input audio stream may be processed with an algorithm to derive a transient estimate, preferably constrained to values between 0 and 1. A reference value derived from the so-called "crest factor" meets this requirement, providing a relationship between the peak value of a signal and its effective value, or root mean square (RMS). The transient estimate analyzes the relationship between the amplitude of the instantaneous peak of the real-time audio stream and the average amplitude over the previous time frame of the audio, and outputs a higher value when a transient is detected.
[0020] Thus, in a preferred embodiment, performing a transient extraction process includes deriving a real-time transient estimate C(t) from the input signal; C(t) has a value in the range 0≦C(t)≦1 and represents the transient component of real-time audio; The transient haptic signal T(t) is generated according to T(t) = C(t)I(t), and the steady-state haptic signal S(t) is generated according to S(t) = (1 - C(t))I(t).
[0021] The transient estimate can then be used to create a transient haptic signal, for example by multiplying it with the received input audio signal. The complement of the generated transient estimate signal can then be multiplied with the same band-limited audio signal to create a steady-state haptic signal. If these signals were summed, they would reproduce the generated audio signal. This allows the system to control the exaggeration or attenuation of transient or steady-state elements of the audio stream, avoiding ornamentation or coloration and maintaining transparency.
[0022] Preferably, the input audio signal is frequency band-limited prior to the transient estimation step. Such band-limiting can be performed before or after receiving the input signal. The selected band-limiting will typically be characterized by the frequency response of the drive unit of the haptic actuator, along with the perceptual sensitivity of the psychoacoustic and somatosensory systems. For example, the band-limiting can be a frequency range of 40-250 Hz for optimal tactile response.
[0023] The generated haptic drive signal can be distributed to multiple haptic actuators, characterized by factors such as transducer response, drive unit position, and other user-controlled parameters. The actuators can receive only transient signals, only steady-state signals, or a weighted sum of both signals.
[0024] These generated signals may be generated from a mono audio stream, or alternatively, the process may be applied in real time to any number of separate audio streams to create transient and steady-state haptic signals for each channel of audio.
[0025] The method may further include generating one or more acoustic drive signals in the audio frequency range based on the input signal. In this manner, both acoustic and tactile drive signals are generated that can be used to drive appropriate transducers, which provides a variety of combinations of acoustic and haptic sensory feedback to the user.
[0026] According to a second aspect of the present invention, a computer-readable medium includes computer-executable instructions that, when executed on one or more processors of an audio system, cause the system to perform the method of the first aspect. In this manner, the method of the first aspect of the present invention can be executed by one or more processors of the audio system to generate a haptic drive signal for driving a haptic actuator.
[0027] The computer readable medium of the second aspect of the present invention may update or enhance an existing digital signal processor sound source system, thus allowing an existing system to be updated.
[0028] According to a third aspect of the present invention, an audio system comprises one or more digital signal processors configured to carry out the method of the first aspect.
[0029] In some embodiments, the audio system includes a user interface for receiving user input parameters, allowing the user to control certain aspects of the haptic drive signal.
[0030] Preferably, the audio system includes one or more tactile transducers for providing haptic feedback, each of the tactile transducers being driven by one of the one or more tactile drive signals, each of which may be generated to optimally drive the respective tactile transducer.
[0031] In some embodiments, one or more of the tactile transducers are configured for use in a seat where a user sits. Tactile transducers may be placed in the backrest, under the seat, and in the leg area of the seat. Additional tactile transducers may be provided on the floor. In other embodiments, the tactile transducers may be configured for use in a wearable device worn by a user. Depending on the particular application, a particular form of actuator may be provided for optimal sensory feedback, with the drive signal appropriately adjusted.
[0032] The audio system further includes one or more acoustic transducers for providing an audible signal, each of the acoustic transducers being driven by one of the one or more acoustic drive signals. In this manner, both acoustic and tactile transducers are provided, allowing a user to experience a variety of combinations of acoustic and haptic sensory stimuli optimized for a particular location and audio type.
[0033] As will be appreciated by those skilled in the art, the present invention can be implemented in a variety of ways depending on the application. [Brief explanation of the drawings]
[0034] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings. [Figure 1] FIG. 1 shows a frequency domain representation of a portion of audio in which the overlap between the haptic and audible ranges is defined. [Figure 2] Figure 2 shows a time-domain representation of a transient occurring in an otherwise steady-state audio stream. [Figure 3] FIG. 3 illustrates an exemplary configuration for an automobile, where one acoustic transducer is mounted in the seat back. [Figure 4] FIG. 4 is a flow chart illustrating high-level transient processing of an audio signal. [Figure 5] FIG. 5 shows the time-domain analysis of a near-steady-state signal with a transient component, with (a) the original signal, (b) the real-time transient estimate, (c) the calculated transient signal, and (d) the calculated steady-state signal. [Figure 6] FIG. 6 is a schematic diagram of the process applied in the time domain analysis of FIG. [Figure 7] FIG. 7 is a flow chart showing more detailed required and optional processing steps when the transient and steady-state streams are combined before output. [Figure 8] FIG. 8 is a flow chart showing more detailed required and optional processing steps when the transient and steady-state streams are not combined before output. [Figure 9] FIG. 9 illustrates an example seat configuration with two built-in haptic transducers, one in the back and one under the seat. [Figure 10] FIG. 10 illustrates an example seat configuration with multiple small haptic transducers in the back, a large actuator under the seat, actuators in the legs, and a floor vibrator. DETAILED DESCRIPTION OF THE INVENTION
[0035] The invention can be used in several different ways depending on the audio system being used. In the following, some exemplary implementations are described with reference to the figures.
[0036] The present invention derives transient and steady-state haptic signals in real time from audio source content and distributes these signals to one or more acoustic transducers. Embodiments of the present invention can be used to enhance the audio system experience, providing a true tactile sensation consistent with the auditory perception of sound.
[0037] The goal of separating transient and steady-state haptic streams from their original sound source stems from both user experience and transducer mechanical design considerations. While Figure 3 illustrates a basic arrangement incorporating a single transducer 31 mounted on the back of seat 30, the transducer may be arbitrarily positioned and mounted depending on the use case. In this case, the single transducer must be able to effectively reproduce both steady-state and transient haptic content.
[0038] To facilitate high energy transfer at low frequencies, actuators typically drive objects with a heavy equivalent mass, which can result in a slow and weak transient response. Conversely, an actuator driving a light equivalent mass may have a sufficiently fast transient response, but may not provide enough energy at low frequencies when driven with an unprocessed input audio signal. Therefore, it is preferable to be able to weight the transient or steady-state components present in the drive signal to accommodate actuator responses that may vary from design to design. Furthermore, for user experience, it is also preferable to be able to control the balance between the steady-state (vibration) and transient (punch) components, so that the balance of the two different resulting sensations can be tailored to preference and calibrated to different audio content material.
[0039] For purposes of the present invention, audio signals containing the full range of acoustic frequencies may be processed to determine transient and steady-state components. However, it is generally preferred that the audio signal be band-limited to a frequency range of interest prior to processing. In some cases, the audio signal may be band-limited naturally. In other cases, it may be desirable to band-limit the audio signal to a desired, operable frequency range by frequency filtering.
[0040] A basic method for preparing a band-limited signal is to apply a low-pass filter with a cutoff frequency of approximately 1 kHz to the input audio signal, resulting in a band-limited signal with frequency content that matches the limits of human tactile sensory response. See "Audio-Tactile Rendering: A Review on Technology and Methods to Convey Musical Information through the Sense of Touch," by B. Remache-Vinueza et al., September 2021, Sensors.
[0041] A more considered approach to band-limited signals accounts for both the human auditory and tactile sensory responses under the influence of frequency characteristics across the entire frequency range. As previously discussed, following the discussion in "Audio-Tactile Rendering: A Review on Technology and Methods to Convey Musical Information through the Sense of Touch" by B. Remache-Vinueza et al., September 2021, Sensors, Figure 1 illustrates the overlap region 13 in the human frequency response to sound that exists between the tactile response region 11 and the auditory response region 12, whereby so-called subsonic phenomena (i.e., f < 20 Hz) can be sensed by tactile means, and frequencies in the 20 Hz-1 kHz range can be sensed by both tactile and auditory means, but the human tactile response becomes silent to frequencies above 1 kHz.
[0042] Furthermore, the primary mechanism for the tactile sensation of vibrations due to skin contact is the Pacinian channel, which has an effective response to haptic phenomena in the 40-500 Hz range, as described by Birnbaum DM and Wanderley MM, "A Systematic Approach to Musical Vibrotactile Feedback," Proceedings of the International Computer Music Conference, Copenhagen, Denmark, August 27-31, 2007, pp. 397-404. Furthermore, human hearing becomes increasingly sensitive to frequency characteristics above 250 Hz (see, e.g., H. Fletcher and W. A. Munson, "Loudness, its definition, measurement, and calculation," Journal of the Acoustical Society of America 5, 1933, pp. 82-108).
[0043] Using these considerations, a preferred implementation of the present invention involves obtaining a band-limited signal that resides in the 40 Hz-250 Hz frequency band, which allows for the excitation of an effective haptic response while reducing the likelihood that the haptic actuator output(s) will impair the quality of the auditory reception of sound in the overall system.
[0044] In some cases, for example, in multi-channel audio content, the low-frequency effects (LFE) channel may contain a frequency band suitable for driving haptic actuators and providing haptic feedback. In these cases, the LFE channel, or similar audio content containing frequencies within the human tactile perception range, may be used as the input audio haptic signal for this processing.
[0045] The resulting band-limited input signal is analyzed to calculate the aforementioned transient and steady-state haptic streams. While there are several ways such analysis can be performed, broadly speaking, the transient estimate can be expressed as a relationship between a short-term (micro) dynamic envelope, typically a time frame of 0 ms-100 ms, and a long-term (macro) dynamic envelope, typically a time frame of 200 ms-1000 ms. This process represents the core of the present invention, and a flow diagram 40 of the overall method is shown in FIG. 4. A received input audio signal 41 undergoes a transient extraction process 42, and a resulting haptic drive signal 43 is output.
[0046] To this end, the relationship between the peak value calculated from the band-limited input signal and the effective, or root-mean-square (RMS), value can be established. A useful definition for such a relationship is the crest factor, which is equivalent to the signal's peak-to-average power ratio (PAR) when expressed in decibels, as defined in TJ Rouphael, 2009, "Wireless 101: Peak to average power ratio (PAPR)" (available online at https: / / www.eetimes.com / wireless-101-peak-to-average-power-ratio-papr / , accessed July 27, 2022), and is given by:
[0047]
number
[0048] where P peak is the maximum value of the squared audio signal within a time frame of the signal, and P avg is the mean value of the squared signal calculated over the same time frame interval.
[0049] While the above provides one example of a method for obtaining a transient estimate, such values can be calculated using other methods, provided they express a relationship between peak and average signal values and can be reduced to a mapped range of 0 to 1. For example, the transient estimate C(t) can be obtained by the following relationship:
[0050]
number
[0051] In all examples of this calculation, the original signal will be an audio signal with an amplitude range of -1 to 1, representing a full-scale audio signal, so P peak (t) and P avg (t) is constrained to the range 0 to 1. Therefore, the transient estimate C(t) is also constrained to the range 0 to 1, which is particularly preferable as a scale factor.
[0052] As mentioned above, using such a transient estimate C(t), two complementary audio-haptic signals may be generated, such as: T(t) = C(t)I(t) S(t) = (1 - C(t))I(t) ∴S(t) + T(t) = I(t) where T(t) is the transient haptic signal, S(t) is the steady-state haptic signal, and I(t) is the input audio signal on which the transient extraction process is performed, allowing for optional filtering before processing. In this way, the two complementary haptic signals contain all relevant information from the input signals.
[0053] The time frame for calculating the peak and RMS values can be adjustable. A shorter dynamic envelope time frame for peak and RMS tracking makes the C(t) value more responsive to smaller changes in the signal. A longer time frame makes the resulting transient estimate smoother, thereby revealing more prominent transient events in the signal.
[0054] Figure 5 shows a time-domain progression starting with an input signal, shown in Figure 5(a), which is an apparently steady-state input signal 50 with a transient state 51 inserted into the steady-state. Applying a transient extraction process to this signal results in a transient estimate C(t) shown in Figure 5(b), where a central "sawtooth" feature 53 represents the transient state relative to a nearly constant zero background level 52. Multiplying the input signal by C(t) results in a transient haptic signal T(t) shown in Figure 5(c), and multiplying the input signal by 1 - C(t) results in a steady-state haptic signal S(t) shown in Figure 5(d).
[0055] FIG. 6 shows schematically the steps involved in this process 60, starting with an input signal 61, on which a transient tracking analysis 62 is performed, resulting in a transient estimate C(t) and its complement 1 − C(t), which are then multiplied by the input signal at 63 and 64, respectively, to generate a transient haptic signal T(t) and a steady-state haptic signal S(t).
[0056] Additional operations can be used to process the C(t) signal, specifically squaring or amplifying it to change the magnitude of its response, although it is preferable to note that the signal range should be limited to the range 0 to 1. Any functional equivalent of the above process can similarly derive corresponding transient and steady-state haptic signals.
[0057] Once the transient and steady-state haptic signals are obtained, they can be used to generate one or more haptic drive signals that are used to drive one or more haptic actuators. While the individual transient and steady-state haptic signals can be used individually as drive signals, with or without further processing, they can also be combined to generate a haptic drive signal.
[0058] For example, the signal A sent to the i-th actuator i (t) may be expressed by the following formula: Ai (t) = G ssi S(t) + G tri T(t) where G ssi is the steady-state gain (i.e., weighting) of the haptic drive signal for the i-th actuator, and G tri is the transient gain (i.e., weighting) for the i-th actuator, T(t) is the time-domain representation of the transient signal extracted from the original audio, and S(t) is the time-domain representation of the steady-state signal extracted from the original audio. In some embodiments, the steady-state and transient gain values may be fixed and optimized depending on the haptic actuator being driven, while in other embodiments, the gain values may be varied, for example, by user input.
[0059] In one implementation, the steady-state and transient gain values are:
[0060]
number
[0061] where u is an input user parameter with a value ranging from −1 to 1, b is a static scale factor that controls the strength of the user parameter, and W ssi and W tri are the static weighting values for the steady-state signal and the transient signal, respectively, for the i-th actuator.
[0062] When u = 0, the steady-state and transient-state gain values are a direct pass-through of the static weighting values. When u is positive, the transient signal level is increased and the steady-state signal level is decreased by a decibel value equal to u. This is because b represents the gain in logarithmic decibels in this implementation. In other implementations, when b represents a linear gain, the user-controlled component 10 -ub / 20 and 10 ub / 20 can be replaced by ub and 1 / ub.
[0063] Depending on the application, filtering and / or time correction may be applied to the transient and steady-state haptic signals either before or after weighting is applied, and the ordering and application of such processing may be done in the most efficient manner for a given situation.
[0064] 7 is a flow diagram illustrating step 70 of a process according to an embodiment, in which transient and steady-state haptic signals can be combined to generate one or more haptic drive signals. Key steps are indicated by solid boxes, while optional steps are indicated by dashed boxes. A received input signal 71 is optionally filtered 72 to extract haptic frequency components before performing a transient extraction analysis 73. The resulting transient and steady-state haptic signals are weighted 74t, 74s and optionally further filtered 75t, 75s before summation 76. A time correction 77 can then be applied to the combined signal before outputting one or more haptic drive signals 78.
[0065] 8 is a flow diagram illustrating step 80 of a process according to an embodiment, in which transient and steady-state haptic signals are not combined to generate one or more haptic drive signals. Major steps are again indicated by solid boxes, while optional steps are indicated by dashed boxes. A received input signal 81 is optionally filtered 82 to extract haptic frequency components before performing a transient extraction analysis 83. Appropriate weightings 84t, 84s are then applied to the resulting transient and steady-state haptic signals, after which further filtering 85t, 85s and / or time corrections 86t, 86s may optionally be performed before being summed to output one or more haptic drive signals 87.
[0066] After considering input signal analysis and touch drive signal generation, we now move on to a discussion of the configurations and characteristics of haptic actuators that can be driven in embodiments of the present invention. As previously mentioned, while FIG. 3 shows a single acoustic transducer 31 mounted on the back of a seat 30, this is only one of many actuator configurations that this technology can encompass. Any number of actuators can be placed in various locations on a seat, wearable device, or other surface, and many mounting strategies can be deployed. Furthermore, the orientation of the actuators can be configured in a variety of ways.
[0067] In an exemplary scenario, the large rear actuator 31 shown in FIG. 3 may have poor transient performance, in which case the static weighting W tri is typically a static weighting W for steady-state signals. ssi For this single actuator, both the transient and steady state signals are sent to the same drive unit. To calibrate this drive unit, W tri is set to 2 (+6 dB), and W ssi is set to 1 (+0 dB). This exemplary weighting doubles the strength of the transient components of the received signal. Furthermore, if the scale factor b is set to 6, it provides the user with a 12 dB (i.e., 2 b) range to control the signal. For example, if the user is listening to dance music, they may choose to increase the u parameter to boost transients, or if listening to orchestral music, they may choose to decrease its value.
[0068] In another example, as shown in FIG. 9, when two or more actuators are attached to the seat 90, W ssi and W triA value of σ can be used to weight the seatback actuators 91, for example, to significantly exaggerate the transient components of the signal, while the under-seat actuators 92 are weighted to exaggerate or exclusively render the steady-state components of the signal. Additionally, floor- or wearable leg-mounted actuators can be used to simulate the sensation experienced while standing in a venue where high sound pressure levels are being played. A steady-state or vibration effect can be provided at the user's feet, with the sensation of vibration transmitted through the body being similar to vibrations transmitted through the ground and transmitted through the feet from a high-volume bass driver.
[0069] In very simple configurations such as that in Figure 3, no weighting is required. The transient signal can be sent without weighting to drive the driver 31 in the seat back, which in the composite signal model is G ssi = 0 and G tri = 1. Before the content rendering stage, the transient signal may be passed through several corrective equalization filters to calibrate the actuator characteristics. Similarly, the steady-state signal can be sent without weighting to drive the under-seat actuator, which in the composite signal model is G ssi = 1 and G tri = 0. The steady-state signal may also undergo some corrective equalization before the content rendering stage.
[0070] While large mass-driven devices have been shown as potential haptic actuators for certain solutions, several small acoustic transducers may also be attached to provide tactile enhancement. Figure 10 illustrates the use of different actuators within a seat 100, depending on their location and purpose. The seat back is equipped with multiple actuators, with a medium-sized actuator 101 positioned in the occupant's spine area and supplemented by smaller actuators 102 near the periphery. Larger actuators are used under the seat 103 and in the leg area 104, with additional actuators 105 embedded in the floor to provide vibration effects at the user's feet. In this case, a narrow transient response is preferably achieved by heavily weighting the transient portion of the signal to transducers closer to the occupant, e.g., transducer 101 closest to the user's spine. Weighting the steady-state component distributes the vibration effect over a larger surface area, effectively vibrating the entire seat.
[0071] When haptic processing is implemented as part of a large, multi-speaker system, some of the drive units will render signals in the human hearing range, which can result in a time lag between the perceived auditory and haptic signals. This can be due to several factors, including the proximity of the sound source, the velocity of the actuator, and the medium of signal transmission. To correct for this, either the haptic or auditory signal can be delayed by implementing a delay line in front of either signal. Here, applying a delay for such purposes is referred to as time correction.
[0072] In summary, an audio signal containing frequencies within the human tactile perception range can be decomposed into complementary transient and steady-state components. Optionally, the initial signal can be frequency band-limited or input directly without pre-processing. The decomposed haptic signals can be weighted and distributed to multiple acoustic transducers designed to render haptic content. The individual haptic signals can be routed directly to output channels or combined before output. Optionally, time correction can be applied to facilitate integration into multimedia systems. Weighting makes the generated haptic signal system-independent, since it can be rendered by any number of actuators and accommodate their transient responses. Alternatively, the user can control the relative weighting of the transient and steady-state signals.
Claims
1. 1. A method for generating one or more haptic drive signals from an input signal representing real-time audio, the method comprising: receiving the input signal; determining a transient component of the real-time audio by performing a transient extraction process on the input signal; generating a transient haptic signal T(t) and a steady-state haptic signal S(t) from the input signal based on the transient extraction process; the transient haptic signal and the steady-state haptic signal are complementary such that S(t) + T(t) = I(t), where I(t) is the input signal on which a transient extraction process has been performed; and generating the one or more haptic drive signals based on one or both of the transient haptic signal and the steady-state haptic signal; A method comprising:
2. performing a transient extraction process includes deriving a real-time transient estimate C(t) from the input signal; C(t) has a value in the range of 0≦C(t)≦1 and represents the transient component of the real-time audio; the transient haptic signal T(t) is generated according to T(t) = C(t)I(t), and the steady-state haptic signal S(t) is generated according to S(t) = (1 - C(t))I(t); The method of claim 1.
3. C(t) is [Equation 1] is defined according to where P peak (t) ≧ P avg (t) and P peak , P avg ∈ [0:1], P peak (t) represents the maximum value of the squared audio signal within the time frame of said signal, and P avg (t) represents the mean value of the squared signal calculated over the same time frame, The method of claim 2.
4. the received input signal is band-limited, either naturally or by prior frequency filtering; The method according to any one of claims 1 to 3.
5. further comprising band-limiting the input signal prior to the transient extraction process by frequency filtering the input signal; I(t) is the input signal after the filtering; The method according to any one of claims 1 to 3.
6. the band-limited input signal is in the frequency band of 40 Hz-250 Hz; 6. The method according to claim 4 or 5.
7. the input signal is a multi-channel signal; and further comprising downmixing the input signal to a single channel signal prior to transient extraction processing. The method according to any one of claims 1 to 3.
8. the input signal includes a low-frequency effects (LFE) channel; The method according to any one of claims 1 to 3.
9. generating the one or more haptic drive signals includes applying a time correction to one or both of the transient haptic signal and the steady-state haptic signal; The method according to any one of claims 1 to 8.
10. generating the one or more haptic drive signals includes applying frequency filtering to one or both of the transient haptic signal and the steady-state haptic signal; The method according to any one of claims 1 to 9.
11. generating the one or more haptic drive signals includes combining the transient haptic signal and the steady-state haptic signal; The method according to any one of claims 1 to 10.
12. Generating the one or more haptic drive signals further includes weighting the transient haptic signal and the steady-state haptic signal before combining them, so that the i-th haptic drive signal has a weight of A i (t) = G ssi S(t) + G tri generated according to T(t), Here, the i-th haptic drive signal A i (t) to generate G ssi is the weighting applied to the steady-state haptic signal, and G tri is a weighting applied to the transient haptic signal, where i ≧ 1; The method of claim 11.
13. the weighting applied to each of the transient haptic signal and the steady-state haptic signal includes a variable weighting component determined from a user input parameter and a static weighting component. The method of claim 12.
14. The weighting G applied to the steady-state haptic signal ssi and a weighting G applied to the transient haptic signal tri teeth, [Equation 2] is calculated according to where u is the input user parameter ranging from -1 to 1, b is a fixed scale factor in decibels that controls the intensity of the user parameter, and W ssi and W tri are static weighting values for the steady-state haptic signal and the transient haptic signal, respectively, for generating the i-th haptic drive signal; The method of claim 13.
15. generating a plurality of different haptic drive signals by applying different weightings to the transient haptic signal and the steady-state haptic signal before combining them to generate respective haptic drive signals; The method according to any one of claims 11 to 14.
16. generating one or more acoustic drive signals in an audio frequency range based on the input signal; The method according to any one of claims 1 to 15.
17. A computer readable medium comprising computer executable instructions which, when executed on one or more processors of an audio system, cause the audio system to perform the method of any one of claims 1-16.
18. one or more digital signal processors configured to carry out the method of any one of claims 1 to 16, Audio system.
19. a user interface for receiving user input parameters; 20. The audio system of claim 18.
20. one or more tactile transducers for providing haptic feedback; each of the tactile transducers is driven by one of the one or more tactile drive signals; 20. An audio system according to claim 18 or 19.
21. each tactile drive signal is generated to optimally drive each of said tactile transducers; 21. The audio system of claim 20.
22. one or more of the tactile transducers are configured for use in a seat, or a wearable device, or a flooring material; 22. An audio system according to claim 20 or 21.
23. one or more acoustic transducers for providing audible feedback; each of the acoustic transducers is driven by one of the one or more acoustic drive signals; An audio system according to any one of claims 18 to 22 when dependent on claim 16.