Headrest loudspeaker and audio processing method and system therefor, and storage medium

By acquiring head position data and combining the fusion and dynamic adjustment of HRTF data, the problem of poor stereo sound hearing of the headrest speakers when the listener's head position changes is solved, achieving a personalized and immersive audio experience.

WO2025156239A1PCT designated stage Publication Date: 2025-07-31AAC ACOUSTIC TECH (SHANGHAI) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/074145
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The existing headrest speakers cannot maintain a good stereo listening experience when the listener's head position changes.

Method used

By obtaining the relative position data of the listener's head and headrest speakers, combining the head-dependent transmission function (HRTF) data of the binaural rendering algorithm and the crosstalk cancellation algorithm, the fusion process is performed and dynamic adjustment is performed to ensure the adaptability of audio processing.

Benefits of technology

Maintain a good stereo listening experience when the listener's head position changes, providing a personalized and immersive audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024074145_31072025_PF_FP_ABST
    Figure CN2024074145_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a headrest loudspeaker and an audio processing method and system therefor, and a storage medium. The method comprises: acquiring relative position data of the head of a listener and a headrest loudspeaker; on the basis of the relative position data, respectively determining first HRTF data and second HRTF data, which correspond to a binaural rendering algorithm and a crosstalk cancellation algorithm; merging the first HRTF data and the second HRTF data, so as to obtain initial merged HRTF data; performing dynamic adjustment on the initial merged HRTF data on the basis of the relative position data, so as to obtain target merged HRTF data, and using the target merged HRTF data to process initial audio, so as to obtain target audio; and playing the target audio by means of the headrest loudspeaker. In the embodiments of the present disclosure, a binaural rendering algorithm and a crosstalk cancellation algorithm are merged to process audio, and a dynamic adjustment strategy is designed, such that it is ensured that a listener can have a good auditory experience regardless of whether he / she is close to or far away from the headrest loudspeaker.
Need to check novelty before this filing date? Find Prior Art

Description

Headrest speaker, audio processing method and system, and storage medium thereof Technical Field

[0001] The embodiments of the present disclosure belong to the field of audio processing technology, and particularly relate to a headrest speaker, an audio processing method and system thereof, and a storage medium. Background Art

[0002] Headrest speakers are audio devices installed in headrests and can be used in car seats, massage chairs, and other scenarios to provide listeners with a personalized audio experience. Common playback strategies include independent playback for each headrest, synchronized playback for each headrest, zoned playback for each headrest, and interactive playback.

[0003] The above playback strategies all directly play stereo sound. The listener needs to maintain the head or ears in the appropriate position to obtain an ideal listening experience. When the listener's posture in the seat changes, such as moving the head away from or close to the speaker, it will result in a poor listening experience. Summary of the Invention

[0004] The embodiments of the present disclosure aim to solve at least one of the technical problems existing in the prior art, and provide a headrest speaker, an audio processing method and system thereof, and a storage medium.

[0005] One aspect of the present disclosure provides a headrest speaker audio processing method, comprising:

[0006] Obtaining the relative position data between the listener's head and the headrest speaker;

[0007] Determining first head-related transfer function HRTF data and second head-related transfer function HRTF data corresponding to a binaural rendering algorithm and a crosstalk cancellation algorithm respectively according to the relative position data;

[0008] fusing the first HRTF data and the second HRTF data to obtain initial fused HRTF data;

[0009] Dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data, and processing the initial audio using the target fused HRTF data to obtain target audio;

[0010] The target audio is played through the headrest speakers.

[0011] Furthermore, the crosstalk cancellation algorithm uses a delayed and inverted signal; and the fusing of the first HRTF data and the second HRTF data to obtain initial fused HRTF data includes:

[0012] The delayed and inverted signals are gain-controlled and then mixed with the binaural rendering signals to obtain the initial fused HRTF data.

[0013] Furthermore, before obtaining the initial fused HRTF data, the method further includes:

[0014] The delayed and inverted signal is subjected to frequency band division processing.

[0015] Optionally, dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data includes:

[0016] If the distance between the listener's head and the headrest speaker is less than a predetermined threshold, obtaining target fused HRTF data based on the first HRTF data;

[0017] If the distance between the listener's head and the headrest speaker is greater than the predetermined threshold, the target fused HRTF data is obtained based on the second HRTF data.

[0018] Furthermore, the dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data includes:

[0019] According to the distance between the listener's head and the headrest speaker, the gains of the first HRTF data and the second HRTF data are respectively smoothly and dynamically adjusted.

[0020] Optionally, after playing the target audio through the headrest speaker, the method further includes:

[0021] Get feedback audio data;

[0022] One or more of the first HRTF data, the second HRTF data, and the threshold are calibrated according to the feedback audio data.

[0023] Optionally, obtaining the relative position data of the listener's head and the headrest speaker includes:

[0024] Obtaining the distance between the listener's head and the headrest speaker;

[0025] The height and angle of the listener's head are obtained, and the relative angles between the listener's ears and the headrest speakers are determined based on the distance and the height and angle of the listener's head.

[0026] Furthermore, obtaining the distance between the listener's head and the headrest speaker includes:

[0027] Obtaining the distance between the listener's head and the headrest speaker through a distance sensor or a pressure sensor; and / or,

[0028] The obtaining of the height and angle of the listener's head includes:

[0029] The height and angle of the audience's head are obtained through visual sensors.

[0030] Optionally, obtaining the height and angle of the listener's head includes:

[0031] The height and angle of the listener's head are obtained based on the pre-established listener head model.

[0032] Another aspect of the present disclosure provides a headrest speaker audio processing system, comprising:

[0033] An acquisition module, used to obtain relative position data of the listener's head and the headrest speaker;

[0034] An HRTF data module is configured to determine first head-related transfer function HRTF data and second head-related transfer function HRTF data corresponding to a binaural rendering algorithm and a crosstalk cancellation algorithm, respectively, based on the relative position data;

[0035] a fusion module, configured to fuse the first HRTF data and the second HRTF data to obtain initial fused HRTF data;

[0036] an audio generation module, configured to dynamically adjust the initial fused HRTF data based on the relative position data to obtain target fused HRTF data, and process the initial audio using the target fused HRTF data to obtain target audio;

[0037] A playing module is used to play the target audio through the headrest speaker.

[0038] Yet another aspect of the present disclosure provides a headrest speaker, which adopts the headrest speaker audio processing method described above; and / or includes the headrest speaker audio processing system described above.

[0039] Yet another aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the headrest speaker audio processing method described above can be implemented.

[0040] A headrest speaker, an audio processing method and system thereof, and a storage medium in the embodiments of the present disclosure process audio by integrating a binaural rendering algorithm and a crosstalk cancellation algorithm, and designing a dynamic adjustment strategy. These methods ensure that when stereo binaural content is played on the headrest speaker, the listener can obtain a good listening experience regardless of whether they are close to or far away from the headrest. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] FIG1 is a flow chart of a headrest speaker audio processing method according to an embodiment of the present disclosure;

[0042] FIG2 is a schematic flow chart of a feedback calibration method according to another embodiment of the present disclosure;

[0043] FIG3 is a schematic structural diagram of a headrest speaker audio processing system according to another embodiment of the present disclosure. DETAILED DESCRIPTION

[0044] In the driving scenario, "headrest speakers" refer to audio devices installed in the car's headrests to provide a personalized audio experience. The specific playback strategy varies depending on the function and design of the in-car entertainment system. The following are some common playback strategies:

[0045] 1. Independent playback: Each headrest speaker can play audio independently, allowing passengers to select the music, audiobooks, or podcasts they want to listen to. This is usually achieved through the passenger's mobile device (such as a smartphone or tablet) via Bluetooth or Wi-Fi connection.

[0046] 2. Synchronous playback throughout the vehicle: All headrest speakers play the same audio at the same time, such as music or radio stations selected by the driver or a passenger.

[0047] 3. Zoned playback: The car is divided into several "audio zones", and the speakers in each zone play different audio. For example, the driver and front passenger may listen to the same audio, while the rear passengers can choose another audio.

[0048] 4. Interactive playback: This is a more complex strategy that can dynamically adjust audio playback based on passenger preferences, the ride status (such as whether the car is driving or parked), and even the passenger's mood (detected through means such as facial recognition or biometric sensors).

[0049] These playback strategies all play stereo sound directly through the headrest speakers. When the listener's head or ears are not in an ideal position, the listening experience will be poor.

[0050] Crosstalk cancellation (CTC) algorithms can be used to enhance the stereo experience, making music sound richer and more vivid. Combined with binaural technology, this can create a more immersive audio environment. To address this issue, recording audio using a binaural rendering algorithm combined with CTC technology provides passengers with the feeling that the sound source is all around them. This algorithm dynamically adjusts based on the relative position of the headrest and the listener, ensuring a consistent listening experience even when the listener's head position changes, thus achieving an optimal playback strategy for stereo binaural content played on the headrest speakers.

[0051] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.

[0052] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid blurring various aspects of the present disclosure.

[0053] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0054] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below can be referred to as the second component without departing from the teachings of the concepts of this disclosure. As used in this disclosure, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0055] Those skilled in the art will understand that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present disclosure, and therefore cannot be used to limit the scope of protection of the present disclosure.

[0056] As shown in FIG1 , an embodiment of the present disclosure provides a headrest speaker audio processing method, comprising:

[0057] Step S1: Obtain relative position data between the listener's head and the headrest speaker.

[0058] Specifically, the relative position data between the listener's head and the headrest speaker includes the distance between the listener's head and the headrest speaker and the relative angle between the listener's ears and the headrest speaker. A distance sensor, such as an infrared sensor, an ultrasonic sensor or a laser sensor, is used to detect the distance between the listener's head and the headrest speaker in the horizontal direction. Taking a typical headrest speaker with left and right speakers as an example, the distance sensor can be set between the two speakers and detect outward. Taking into account the height differences of different listeners and the possible deviation of the listener's head, the detection range of the distance sensor can diverge outward from one point, or multiple distance sensors can be set to form a detection array to form a distance sensing surface to ensure the reliable acquisition of the distance between the listener's head and the headrest speaker.

[0059] Through visual sensors, such as in-car cameras, the height of the listener's head is obtained, and then the relative height of the listener's head and the headrest speaker in the vertical direction is obtained; in view of the head posture of the listener in the car that may change at any time, the in-car camera also needs to obtain the angle of the listener's head, and then combine the distance and relative height between the listener's head and the headrest speaker to calculate the relative position of each of the listener's ears and the headrest speaker, and further accurately determine the relative angle between the listener's left and right ears and the headrest speaker.

[0060] In some embodiments, the distance between the listener's head and the headrest speaker can also be determined using a pressure sensor. For example, a pressure sensor (such as a piezoelectric sensor) can be integrated into the headrest. When the listener's head approaches or contacts the headrest, the pressure sensor detects a change in pressure. Based on the detected pressure change, it can be determined whether the listener is close to the headrest. Similarly, a pressure sensor can be installed on the seat cushion or backrest to detect changes in the listener's weight distribution. When the listener moves closer to or further away from the headrest, the pressure distribution on the seat changes. Based on these changes, the distance between the listener's head and the headrest speaker can be inferred. These two methods serve as supplementary methods for detecting the distance between the listener's head and the headrest speaker. Combined with multiple sensor technologies, they can provide more accurate and stable distance measurement.

[0061] In addition, the height and angle of the listener's head can be obtained based on a pre-established listener head model. This allows for personalized customization for individual listeners by collecting information about the distance between their ears or their relative position, even the specific relative position of their eardrums, as well as the height and angle of their head in their normal state. This model can then be used to replace the pre-established universal head model for subsequent audio processing, providing the listener with a better listening experience.

[0062] Step S2: determining first head-related transfer function (HRTF) data and second head-related transfer function (HRTF) data corresponding to a binaural rendering algorithm and a crosstalk cancellation algorithm, respectively, according to the relative position data.

[0063] Specifically, both the binaural rendering algorithm and the crosstalk cancellation algorithm use different HRTF (Head Related Transfer Functions) data based on the relative position of the listener's head and the headrest speaker.

[0064] HRTF is a function used to simulate the effect of the head on sound. It is obtained by measuring or calculating the acoustic properties between the head and the ear. HRTF measurements are typically performed using an artificial ear mold or a real human head. Some current research also uses mathematical models and computer simulations to predict the effects of the head and ear on sound. These methods calculate and predict HRTFs based on knowledge of head geometry, ear anatomy, and acoustic properties. During measurement, a multi-channel speaker system transmits different sound signals to the ear. The received sounds are then recorded and analyzed by a micro-microphone array at the ear, or performed in a simulation, to obtain information about the head's effect on the sound. The resulting data is typically presented as a set of frequency response functions that represent the variations in sound in specific directions. These frequency response functions are called HRTFs. For different directions, HRTFs can provide information about the sound source localization and spatial characteristics of the sound in that direction. Whether measured or calculated, HRTFs can be used to capture audio data based on a specific direction and simulate the effects of the head to enhance the 3D audio experience or localization of the sound source.

[0065] Binaural rendering is an algorithm that simulates how the human ear receives and processes sound to achieve three-dimensional sound effects. It uses information such as the location of the sound source and HRTF to simulate the spatial positioning and resolution of sound by calculating the difference in the audio signals received by the two ears. The HRTF data required for binaural rendering takes into account factors such as the physical properties of sound propagation, the human ear's ability to resolve sound, and the human perception of sound localization.

[0066] Crosstalk cancellation algorithms are technologies used to eliminate crosstalk during signal transmission. Crosstalk refers to the phenomenon of interference between signals during transmission, resulting in distortion or aliasing of the received signal. In audio signal transmission, such as in audio devices or speaker systems, the characteristics of audio signal propagation can cause sounds from different sources to interfere with each other, resulting in crosstalk. The goal of a crosstalk cancellation algorithm is to reduce or remove the crosstalk signal from the received signal through specific signal processing methods, thereby restoring the accuracy and clarity of the original signal. Crosstalk cancellation algorithms typically require signal analysis and modeling to estimate and extract the characteristics of the crosstalk signal. They then perform phase and power compensation with the original signal to achieve crosstalk compensation and cancellation. Common crosstalk cancellation algorithms include adaptive filters and spatial mixing matrices. In dual-speaker headrest speakers, this eliminates unwanted sound from one speaker to the opposite ear by sending a delayed and inverted signal from the other speaker.

[0067] After obtaining accurate data through various sensors in step S1, the distance and relative angle between the headrest speaker and the human ear can be calculated in real time, thereby determining the HRTF data of the binaural rendering algorithm and the crosstalk cancellation algorithm respectively.

[0068] Step S3: Fusing the first HRTF data and the second HRTF data to obtain initial fused HRTF data.

[0069] Specifically, the crosstalk cancellation algorithm uses delayed and inverted signals. According to the principle of the crosstalk cancellation algorithm, it can be used as a sub-module of the binaural rendering algorithm. The delay and inversion functions are added to the binaural rendering. The delayed and inverted signal, that is, the second HRTF data, that is, the HRTF data corresponding to the crosstalk cancellation algorithm, is gain controlled and then mixed with the binaural rendering signal, that is, the first HRTF data, that is, the HRTF data corresponding to the binaural rendering algorithm.

[0070] The time-delayed and inverted signals need to be processed in different frequency bands. Frequency division operations can use filters such as Linkwitz-Riley filters, Fast Fourier Transform (FFT), etc. The processing is divided into four different frequency bands: low frequency (<500 Hz), mid-low frequency (500 Hz < f < 1.5 kHz), mid-high frequency (1.5 kHz < f < 5 kHz), and high frequency (>5 kHz). Since sound is non-directional in the low-frequency part, the ultra-low frequency and low-frequency parts are not processed. The high-frequency part is appropriately processed according to the characteristics of high frequencies to retain more high-frequency details and maintain the sound quality after processing. The main processing of the mid-frequency part (500 Hz < f < 5 kHz) is carried out through operations such as sound gain, balance, compression, and reverberation on the audio signal. Frequency band processing can more finely adjust the audio finally played by the speaker to achieve better sound effects.

[0071] Finally, considering the distances and relative angles between the listener's two ears and the headrest speakers, a gain value is determined to amplify the HRTF data of the crosstalk cancellation algorithm and fuse and mix it with the HRTF data of binaural rendering to ensure that the crosstalk cancellation effect can be achieved. The specific fusion steps are as follows:

[0072] 1. Align the two HRTF data. First, it is necessary to ensure that the two HRTF data sets have the same sampling rate and number of data points. If the sampling rates and number of data points of the two data sets do not match, methods such as interpolation or decimation can be used for alignment;

[0073] 2. For each frequency point, average the frequency response functions of the two HRTF data sets. This can be achieved by averaging data such as the amplitude and phase of the frequency response functions of the two HRTF data sets;

[0074] 3. Normalize the averaged HRTF data to ensure that its frequency response function is within an appropriate range. Common normalization methods include setting the maximum value of the HRTF data to 1 or normalizing its total amplitude to 1.

[0075] 4. Perform mixing processing, and mix the two HRTF data sets by simple linear weighted averaging according to a ratio. The mixing ratio can be adjusted for gain according to actual needs, such as according to factors such as the position and distance of the sound source, to control the contribution degree of each HRTF data set in the final mixing.

[0076] Step S4: Dynamically adjust the initial fusion HRTF data based on the relative position data to obtain target fusion HRTF data, and use the target fusion HRTF data to process the initial audio to obtain target audio.

[0077] Specifically, after the above steps, the initial distance and relative angle between the listener's head and the headrest speaker are reflected in the initial fused HRTF data, and the initial fused HRTF data determined by the listener's typical sitting posture has been determined. In some usage scenarios, such as while driving, the relative position of the listener's head and the headrest speaker changes primarily due to distance changes, i.e., when the body leans forward or backward. The distance between the listener's head and the headrest speaker is measured in real time, and the initial fused HRTF data is dynamically adjusted based on this distance to obtain target fused HRTF data for audio processing. Specifically, an appropriate distance threshold is first set. If the distance between the listener's head and the headrest speaker is less than the predetermined threshold, the gain of the first HRTF data is increased to be greater than the gain of the second HRTF data. The target fused HRTF data is obtained based on the first HRTF data, and the binaural rendering algorithm is used as the primary algorithm. If the distance between the listener's head and the headrest speaker is greater than the predetermined threshold, the gain of the second HRTF data is increased to be greater than the gain of the first HRTF data. The target fused HRTF data is obtained based on the second HRTF data, and the crosstalk cancellation algorithm is used as the primary algorithm.

[0078] In order to avoid sudden audio changes when switching between algorithms, the gains of the first HRTF data and the second HRTF data can be smoothly and dynamically adjusted according to the distance between the listener's head and the headrest speaker. When the distance between the listener's head and the headrest speaker changes, the gains of the first HRTF data and the second HRTF data are relatively increased and decreased. Specifically, when the distance between the listener's head and the headrest speaker gradually increases, the gain of the first HRTF data gradually decreases, and the gain of the second HRTF data gradually increases; conversely, when the distance between the listener's head and the headrest speaker decreases, the gain of the first HRTF data increases, and the gain of the second HRTF data decreases. In this process, when the distance between the listener's head and the headrest speaker is equal to a predetermined threshold, the gains of the first HRTF data and the second HRTF data are equal. The above method can achieve a smooth transition between the two algorithms, so that the listener can have a more comfortable audio listening experience.

[0079] After obtaining real-time target fused HRTF data through the above process, the initial audio, i.e., the music or radio station audio selected by the user, can be processed in real time based on the target fused HRTF data, and finally the target audio generated by the above algorithm is obtained. First, the position of the headrest speaker is used as the audio source position. The relative angle between the listener's ears and the headrest speaker includes the azimuth and pitch angles. The position of the audio source is represented by the azimuth and pitch angles. The position of the audio source is then matched with the direction corresponding to the target fused HRTF to determine the final position of the sound. Finally, the initial audio data is convolved with the HRTF data to obtain the target audio data for speaker playback. The audio obtained by applying HRTF will have the characteristics of head influence, allowing the listener to feel the spatial position, direction and depth of the sound source, enhancing the immersion and realism of the audio.

[0080] Step S5: Play the target audio through the headrest speaker.

[0081] Specifically, the initial audio and the target audio obtained in the above steps are both electrical signals, and the electrical signal of the target audio is converted into a sound signal for playback through the headrest speaker. The speaker is mainly composed of an electromagnetic drive unit and a diaphragm. The electrical signal can first pass through an amplifier to amplify the low-voltage audio signal into a sufficiently large current. The amplified current then passes through the coil connected to the electromagnetic drive unit of the speaker to generate a magnetic field, which interacts with the permanent magnet in the electromagnetic drive unit. According to Ampere's law, when the coil in the electromagnet is subjected to current, it will be subjected to a reverse magnetic field force, driving the diaphragm to vibrate. The vibration of the diaphragm generates pressure waves in the air, that is, sound waves, completing the playback of the target audio.

[0082] A headrest speaker audio processing method in an embodiment of the present disclosure processes audio by integrating a binaural rendering algorithm and a crosstalk cancellation algorithm, and designs a dynamic adjustment strategy to ensure that when stereo binaural content is played on the headrest speaker, the listener can obtain a good listening experience regardless of whether the listener is close to or far away from the headrest.

[0083] Exemplarily, as shown in FIG2 , after playing the target audio through the headrest speaker, the method further includes:

[0084] Step S21: obtaining feedback audio data;

[0085] Step S22: calibrate one or more of the first HRTF data, the second HRTF data, and the threshold according to the feedback audio data.

[0086] Specifically, to ensure the effectiveness of dynamic adjustment in practical applications, feedback and calibration mechanisms can be used to optimize its performance. First, a sound recording device is used to collect the in-car audio response. For example, a tester can manually collect data using miniature microphones worn on both ears, or microphones located elsewhere in the vehicle can be used to capture feedback audio data from the sound played by the headrest speakers. This feedback audio data is then analyzed and compared with the desired target sound effect. The first and second HRTF data, or the distance threshold used for dynamic adjustment, are calibrated, and the smooth transition strategy is adjusted.

[0087] As shown in FIG3 , another embodiment of the present disclosure provides a headrest speaker audio processing system, comprising:

[0088] An acquisition module 301 is used to acquire relative position data between the listener's head and the headrest speaker;

[0089] HRTF data module 302, configured to determine first head-related transfer function HRTF data and second head-related transfer function HRTF data corresponding to a binaural rendering algorithm and a crosstalk cancellation algorithm, respectively, based on the relative position data;

[0090] A fusion module 303 is configured to fuse the first HRTF data and the second HRTF data to obtain initial fused HRTF data;

[0091] an audio generation module 304 for dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data, and processing the initial audio using the target fused HRTF data to obtain target audio;

[0092] The playing module 305 is configured to play the target audio through the headrest speaker.

[0093] The headrest speaker audio processing system of this embodiment is used to implement the headrest speaker audio processing method described in the previous embodiment. The specific implementation process has been described in the previous embodiment and will not be repeated here.

[0094] A headrest speaker audio processing system in an embodiment of the present disclosure processes audio by integrating a binaural rendering algorithm and a crosstalk cancellation algorithm, and designs a dynamic adjustment strategy to ensure that when stereo binaural content is played on the headrest speaker, the listener can obtain a good listening experience regardless of whether the listener is close to or far away from the headrest.

[0095] Yet another embodiment of the present disclosure provides a headrest speaker, which adopts the headrest speaker audio processing method described above; and / or the headrest speaker audio processing system described above.

[0096] A headrest speaker of an embodiment of the present disclosure, by having a built-in headrest speaker audio processing system of the above embodiment, or adopting a headrest speaker audio processing method of the above embodiment for audio processing, can ensure that when stereo binaural content is played on the headrest speaker, the listener can obtain a good listening experience regardless of whether the listener is close to or far away from the headrest.

[0097] Yet another embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the headrest speaker audio processing method described above can be implemented.

[0098] The computer-readable storage medium may be included in the system or electronic device of the present disclosure, or may exist independently.

[0099] Computer-readable storage media may be any tangible medium that contains or stores a program, which may be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device. More specific examples include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, optical fiber, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0100] The computer-readable storage medium may also include a data signal propagated in baseband or as part of a carrier wave, which carries the computer-readable program code. Specific examples include but are not limited to electromagnetic signals, optical signals, or any suitable combination thereof.

[0101] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present disclosure, and such modifications and improvements are also considered to be within the scope of protection of the present disclosure.

Claims

1. An audio processing method for a headrest speaker, characterized in that The method includes: Obtaining relative position data of the listener's head and the headrest speaker; According to the relative position data, respectively determining first head-related transfer function (HRTF) data and second head-related transfer function (HRTF) data corresponding to a binaural rendering algorithm and a crosstalk cancellation algorithm; Fusing the first HRTF data and the second HRTF data to obtain initial fused HRTF data; Dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data, and processing initial audio using the target fused HRTF data to obtain target audio; Playing the target audio through the headrest speaker.

2. The method according to claim 1, characterized in that, The crosstalk cancellation algorithm uses a delayed and inverted signal; the fusing the first HRTF data and the second HRTF data to obtain initial fused HRTF data includes: Performing gain control on the delayed and inverted signal and mixing it with a binaural rendering signal to obtain the initial fused HRTF data.

3. The method according to claim 2, wherein Before obtaining the initial fused HRTF data, the method further includes: Performing frequency band processing on the delayed and inverted signal.

4. The method according to any one of claims 1 to 3, characterized in that, The dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data includes: If the distance between the listener's head and the headrest speaker is less than a predetermined threshold, obtaining target fused HRTF data mainly based on the first HRTF data; If the distance between the listener's head and the headrest speaker is greater than the predetermined threshold, obtaining target fused HRTF data mainly based on the second HRTF data.

5. The method according to claim 4, wherein The dynamically adjusting the initial fused HRTF data based on the relative position data to obtain target fused HRTF data includes: Smoothly and dynamically adjusting the gains of the first HRTF data and the second HRTF data respectively according to the distance between the listener's head and the headrest speaker.

6. The method according to claim 4, characterized in that After playing the target audio through the headrest speaker, the method further includes: Obtaining feedback audio data; Calibrating one or more of the first HRTF data, the second HRTF data, and the threshold according to the feedback audio data.

7. The method according to any one of claims 1 to 3, characterized in that, The obtaining relative position data of the listener's head and the headrest speaker includes: Obtaining the distance between the listener's head and the headrest speaker; Obtaining the height and angle of the listener's head, and determining the relative angles between the listener's two ears and the headrest speaker based on the distance and the height and angle of the listener's head.

8. The method according to claim 7, wherein The obtaining the distance between the listener's head and the headrest speaker includes: Obtaining the distance between the listener's head and the headrest speaker through a distance sensor or a pressure sensor; and / or The obtaining the height and angle of the listener's head includes: Obtaining the height and angle of the listener's head through a vision sensor.

9. The method according to claim 7, wherein The obtaining the height and angle of the listener's head includes: Obtaining the height and angle of the listener's head according to a pre-established listener head model.

10. A headrest speaker audio processing system, characterized in that Includes: An obtaining module, configured to obtain relative position data of the listener's head and the headrest speaker; The HRTF data module is configured to respectively determine the first head-related transfer function (HRTF) data corresponding to the binaural rendering algorithm and the second HRTF data corresponding to the crosstalk cancellation algorithm according to the relative position data; The fusion module is configured to perform a fusion process on the first HRTF data and the second HRTF data to obtain initial fused HRTF data; The audio generation module is configured to dynamically adjust the initial fused HRTF data based on the relative position data to obtain target fused HRTF data, and process the initial audio by using the target fused HRTF data to obtain target audio; The playback module is configured to play the target audio through the headrest speakers.

11. A headrest speaker, characterized in that, The headrest speakers adopt the headrest speaker audio processing method according to any one of claims 1 to 9; and / or, include the headrest speaker audio processing system according to claim 10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it can implement the headrest speaker audio processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Headtracking for pre-rendered binaural audio

    CN109417677A

  • Near-field sound source playback method based on loudspeaker array

    CN116614761A

  • Audio processing method and device, electronic equipment, storage medium and program product

    CN117135557A

  • Zone-based sound reproduction in a vehicle

    DE102014210105A1

  • Method for audio processing

    EP4175325A1