Panoramic earphone playback method, audio equipment and computer readable storage medium
By acquiring the immersive sound reference layout and head-related transfer data of the over-ear circular open headphones, calculating the degree of difference, and determining the speaker combination mapping relationship, the problem of sound field reconstruction distortion and 3D positioning inaccuracy when playing immersive sound audio with over-ear headphones was solved, achieving high-fidelity, low-distortion 3D sound field reproduction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOLDANA TECH CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-24
AI Technical Summary
How to effectively adapt the immersive audio signal designed for spatial speaker arrays to the output of headphones, and solve the problems of sound field reconstruction distortion and 3D positioning inaccuracy when headphones play immersive audio.
By acquiring the immersive sound reference layout and head-related transfer data of the over-ear circular open headphones, the degree of difference is calculated, the speaker combination mapping relationship is determined, and the playback of immersive sound audio signals is realized.
Achieve high-fidelity, low-distortion 3D sound field reproduction on headphones, enhancing immersion and directional accuracy.
Smart Images

Figure CN121924433A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio playback technology, and in particular to a method for playing back immersive sound headphones, an audio device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of immersive audio technology, panoramic sound (such as Dolby Atmos) is widely used in film, games, virtual reality, and other scenarios, creating an immersive auditory experience through a three-dimensional sound field. Achieving high-quality panoramic sound playback on headphones is particularly important for enhancing the immersive experience in mobile scenarios.
[0003] However, due to the fundamental difference between the physical structure of headphones and the layout of traditional speakers, effectively adapting the immersive audio signals originally designed for spatial speaker arrays to the output of headphones has become a key challenge in achieving a realistic spatial audio experience. Summary of the Invention
[0004] The main objective of this application is to provide a method for reproducing immersive sound headphones, an audio device, and a computer-readable storage medium, aiming to solve the technical problem of how to effectively adapt immersive sound audio signals originally designed for spatial speaker arrays to the output of headphones.
[0005] To achieve the above objectives, this application proposes a method for reproducing panoramic sound through headphones, the method comprising: Acquire head-related transfer data for each immersive sound channel in the immersive sound reference layout, and head-related transfer data for each speaker combination in the over-ear circular open headphones; Based on the head-related transfer data corresponding to each of the panoramic sound channels and the head-related transfer data corresponding to each of the speaker combinations, the degree of difference in head-related transfer data between each of the panoramic sound channels and each of the speaker combinations is calculated. Based on the degree of difference between each of the panoramic sound channels and each of the speaker combinations, the speaker combination mapped to each of the panoramic sound channels is determined; Upon receiving a panoramic sound audio signal, the panoramic sound audio signal is played through the speaker combination mapped to each of the panoramic sound channels.
[0006] In one embodiment, the step of determining the speaker combination assigned to each of the immersive sound channels based on the degree of difference between each of the immersive sound channels and each of the speaker combinations includes: One panoramic sound channel is selected from each of the panoramic sound channels in a round-robin fashion as the target panoramic sound channel; Based on the degree of difference between the target immersive sound channel and each of the speaker combinations, the speaker combination with the smallest degree of difference between the speaker combination and the target immersive sound channel is determined as the speaker combination mapped to the target immersive sound channel. If the speaker combination mapped to the target panoramic sound channel is determined, the process returns to the step of polling and selecting a panoramic sound channel from each of the panoramic sound channels as the target panoramic sound channel, until all panoramic sound channels have been selected, thus obtaining the speaker combination mapped to each of the panoramic sound channels.
[0007] In one embodiment, the step of determining the speaker combination mapped to each of the immersive sound channels based on the degree of difference between each of the immersive sound channels and each of the speaker combinations includes: Based on each of the panoramic sound channels and each of the speaker combinations, multiple speaker allocation schemes are constructed, wherein the speaker allocation scheme refers to the scheme obtained by allocating a speaker combination to each of the panoramic sound channels from each of the speaker combinations; The total degree of difference between each speaker allocation scheme is calculated based on the degree of difference between each immersive sound channel and its corresponding allocated speaker combination in each speaker allocation scheme. The speaker combination for each of the panoramic sound channels is determined based on the speaker allocation scheme with the minimum total difference.
[0008] In one embodiment, the step of calculating the total degree of difference of each of the speaker allocation schemes based on the degree of difference between each immersive sound channel and its corresponding allocated speaker combination in each speaker allocation scheme includes: Obtain the pre-configured channel weights of each of the aforementioned panoramic sound channels; Based on the channel weights of each of the panoramic sound channels, the degree of difference between each panoramic sound channel and its corresponding allocated speaker combination is weighted and summed in each of the speaker allocation schemes to calculate the total degree of difference of each speaker allocation scheme.
[0009] In one embodiment, in the same speaker allocation scheme, the speakers in the speaker combinations allocated to different immersive sound channels are not duplicated.
[0010] In one embodiment, the speaker allocation scheme refers to a scheme in which a speaker combination is sequentially allocated to each of the panoramic sound channels according to the channel weights of each pre-configured panoramic sound channel, in descending order of channel weight.
[0011] In one embodiment, the head-related transfer data is a head-related transfer function, the panoramic sound audio signal includes channel audio signals corresponding to each of the panoramic sound channels, and the step of playing the panoramic sound audio signal through the speaker combination allocated to each of the panoramic sound channels includes: Based on the head-related transfer function corresponding to each of the panoramic sound channels and the head-related transfer function corresponding to the speaker combination mapped to each of the panoramic sound channels, the playback configuration corresponding to each of the panoramic sound channels is calculated. Based on the playback configuration of each panoramic sound channel, spatial rendering is performed on the channel audio signal corresponding to each panoramic sound channel to obtain the combined audio signal corresponding to each panoramic sound channel. Each panoramic sound channel is mapped to a speaker combination, which plays the combined audio signal corresponding to each panoramic sound channel.
[0012] In one embodiment, the playback configuration includes a playback distance, and prior to the step of playing the combined audio signal corresponding to each of the panoramic sound channels via speaker combinations mapped to each of the panoramic sound channels, the method further includes: Obtain the pre-measured binaural room impulse response corresponding to each of the aforementioned panoramic sound channels; By convolving the channel audio signals corresponding to each of the panoramic sound channels with the binaural room impulse response corresponding to each of the panoramic sound channels, the reverberation audio signals corresponding to each of the panoramic sound channels are obtained. Based on the playback distance corresponding to each of the panoramic sound channels, the dry and wet mixing ratio corresponding to each panoramic sound channel is calculated respectively. According to the dry and wet mixing ratio of each of the panoramic sound channels, the reverberation audio signal of each panoramic sound channel is added to the combined audio signal of each panoramic sound channel to obtain a new combined audio signal of each panoramic sound channel.
[0013] In addition, to achieve the above objectives, this application also proposes an audio device, the audio device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the panoramic sound headphone playback method described above.
[0014] In addition, to achieve the above objectives, this application also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the panoramic sound headphone playback method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the panoramic sound headphone playback method described above.
[0016] This application provides a method for reproducing immersive sound headphones, an audio device, and a computer-readable storage medium, relating to the field of audio playback technology. The method includes: acquiring head-related transfer data corresponding to each immersive sound channel in an immersive sound reference layout, and head-related transfer data corresponding to each speaker combination in a closed-back circular headphone; calculating the degree of difference between the head-related transfer data of each immersive sound channel and each speaker combination based on the head-related transfer data corresponding to each immersive sound channel and each speaker combination; determining the speaker combination mapped to each immersive sound channel based on the degree of difference between each immersive sound channel and each speaker combination; and playing the immersive sound audio signal through the speaker combination mapped to each immersive sound channel upon receiving the immersive sound audio signal.
[0017] This application embodiment effectively solves the key technical challenges of sound field reconstruction distortion and 3D positioning inaccuracy when adapting panoramic sound audio signals designed for multi-speaker spatial arrays to the special physical structure of over-ear circular open headphones by constructing a panoramic sound channel-speaker combination mapping mechanism based on measured head-related transfer data differences. Specifically, the method first obtains the head-related transfer data corresponding to each logical panoramic sound channel in a standard panoramic sound reference layout, as well as the head-related transfer data corresponding to several speaker combinations composed of physical speakers in the over-ear circular open headphones. Based on this, the system calculates the degree of difference in head-related transfer data between each panoramic sound channel and each speaker combination (which can be measured by frequency domain mean square error, perceptual weighted distance, or directional response similarity, etc.), thereby quantifying the auditory reproduction capability of different speaker combinations for different panoramic sound channels in the panoramic sound reference layout. Subsequently, based on these differences, a speaker combination is intelligently matched to each immersive sound channel, establishing an optimized mapping relationship from virtual speakers (i.e., immersive sound channels under the immersive sound reference layout) to physical drive units (i.e., speakers in the speaker combination). This ensures that the speaker combinations mapped to each immersive sound channel can reproduce the immersive sound audio playback effect of the immersive sound reference layout as accurately as possible. Finally, upon receiving an immersive sound audio signal containing multi-channel audio signals, the system routes the audio signals of each channel to the corresponding speaker combination for playback according to this mapping relationship, allowing the user to still perceive a three-dimensional sound field structure that matches the creator's intent on the headphones.
[0018] This application's embodiments break through the limitations of traditional binaural rendering, which relies solely on general head-related transfer functions for convolutional synthesis. Instead, it fully utilizes the hardware advantages of the physical layout of multiple speakers in a head-mounted circular open headphone. Through a channel-level matching strategy combined with a spatial rendering algorithm, it achieves high-fidelity, low-distortion spatial audio transfer from an ideal speaker sound field to a constrained headphone platform, significantly improving the immersiveness, directional accuracy, and auditory naturalness of panoramic sound playback in mobile or personal scenarios. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating the first embodiment of the panoramic sound headphone playback method of this application; Figure 2 A flowchart illustrating the second embodiment of the panoramic sound headphone playback method of this application; Figure 3 A flowchart illustrating the third embodiment of the panoramic sound headphone playback method of this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the panoramic sound headphone playback method of this application. Figure 5 This is a schematic diagram of the single-side speaker layout of a circular open-back headphone in a specific embodiment of this application; Figure 6 This is a schematic diagram of the earphone ring partition of a specific embodiment of the over-ear ring-shaped open earphone in this application; Figure 7 This is a schematic diagram illustrating the optimized mapping relationship in a specific embodiment of this application; Figure 8 This is a spatial rendering diagram of a specific embodiment of this application; Figure 9 This is a schematic diagram of the playback of the panoramic sound headphones in a specific embodiment of this application; Figure 10 This is a schematic diagram of the hardware operating environment of the electronic device involved in the panoramic sound headphone playback method in this application embodiment.
[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0024] Currently, due to the fundamental difference between the physical structure of over-ear headphones and traditional speaker layouts, the former typically outputs audio through a limited number of sound-producing units close to the ears, while the latter relies on multiple speakers distributed in three-dimensional space to construct a realistic sound wave propagation and reflection environment. When immersive audio signals designed for multi-speaker arrays (such as 5.1.4, 7.1.6, and other immersive sound reference layouts) are played directly through over-ear headphones, it easily leads to sound field reconstruction distortion and inaccurate three-dimensional positioning. Specifically, this manifests as: sound image collapse into the skull, confusion in front-back / up-down directions, decreased spatial separation, and a significant reduction in overall immersion.
[0025] Therefore, how to achieve high-fidelity reproduction of the original panoramic sound spatial intent in over-ear circular open headphones has become a technical bottleneck that urgently needs to be solved.
[0026] To address this issue, the solution provided in this application is a method for reproducing immersive sound in headphones. The method includes: acquiring head-related transfer data corresponding to each immersive sound channel in an immersive sound reference layout, and head-related transfer data corresponding to each speaker combination in a circular open-back headphone; calculating the degree of difference between the head-related transfer data of each immersive sound channel and each speaker combination based on the head-related transfer data of each immersive sound channel and each speaker combination; determining the speaker combination mapped to each immersive sound channel based on the degree of difference between each immersive sound channel and each speaker combination; and playing the immersive sound audio signal through the speaker combination mapped to each immersive sound channel upon receiving an immersive sound audio signal.
[0027] This application embodiment effectively solves the key technical challenges of sound field reconstruction distortion and 3D positioning inaccuracy when adapting panoramic sound audio signals designed for multi-speaker spatial arrays to the special physical structure of over-ear circular open headphones by constructing a panoramic sound channel-speaker combination mapping mechanism based on measured head-related transfer data differences. Specifically, the method first obtains the head-related transfer data corresponding to each logical panoramic sound channel in a standard panoramic sound reference layout, as well as the head-related transfer data corresponding to several speaker combinations composed of physical speakers in the over-ear circular open headphones. Based on this, the system calculates the degree of difference in head-related transfer data between each panoramic sound channel and each speaker combination (which can be measured by frequency domain mean square error, perceptual weighted distance, or directional response similarity, etc.), thereby quantifying the auditory reproduction capability of different speaker combinations for different panoramic sound channels in the panoramic sound reference layout. Subsequently, based on these differences, a speaker combination is intelligently matched to each immersive sound channel, establishing an optimized mapping relationship from virtual speakers (i.e., immersive sound channels under the immersive sound reference layout) to physical drive units (i.e., speakers in the speaker combination). This ensures that the speaker combinations mapped to each immersive sound channel can reproduce the immersive sound audio playback effect of the immersive sound reference layout as accurately as possible. Finally, upon receiving an immersive sound audio signal containing multi-channel audio signals, the system routes the audio signals of each channel to the corresponding speaker combination for playback according to this mapping relationship, allowing the user to still perceive a three-dimensional sound field structure that matches the creator's intent on the headphones.
[0028] This application's embodiments break through the limitations of traditional binaural rendering, which relies solely on general head-related transfer functions for convolutional synthesis. Instead, it fully utilizes the hardware advantages of the physical layout of multiple speakers in a head-mounted circular open headphone. Through a channel-level matching strategy combined with a spatial rendering algorithm, it achieves high-fidelity, low-distortion spatial audio transfer from an ideal speaker sound field to a constrained headphone platform, significantly improving the immersiveness, directional accuracy, and auditory naturalness of panoramic sound playback in mobile or personal scenarios.
[0029] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0030] This application proposes a first embodiment of a panoramic sound headphone playback method.
[0031] like Figure 1 As shown, in this embodiment, the panoramic sound headphone playback method may include steps S100~S400: Step S100: Obtain head-related transfer data corresponding to each immersive sound channel in the immersive sound reference layout, and head-related transfer data corresponding to each speaker combination in the over-ear circular open headphones. In this embodiment, the immersive sound reference layout refers to the spatial arrangement of speakers specified for a standard multi-speaker 3D audio system, which includes multiple fixed physical positions located on horizontal and vertical planes. This immersive sound reference layout is predefined by international standards organizations or audio content producers as a spatial benchmark for immersive sound audio content creation and mixing. Common immersive sound reference layouts include the immersive sound layout corresponding to 7.1.4 channels and the immersive sound layout corresponding to 5.1.4 channels specified by Dolby.
[0032] In a panoramic sound reference layout, each panoramic sound channel is a logical audio channel with a defined spatial coordinate, such as "front left," "surround right," or "center." Each panoramic sound channel corresponds to an independent audio signal stream and carries sound object or bed information in a specific direction, together forming a complete three-dimensional audio scene.
[0033] Over-ear open-back headphones refer to a type of over-ear audio device that integrates multiple physical speaker units. Instead of the traditional dual-channel structure of headphones, the speakers are arranged in a ring along the inside of the earcups (for example, 3–8 physical speakers are arranged in front, above, and behind the ear). By coordinating and driving the simulation of sound source radiation from different spatial directions, the acoustic characteristics of a three-dimensional speaker array in an approximate panoramic sound reference layout can be reconstructed within a limited wearing space.
[0034] A speaker array refers to the equivalent sound source configuration in a pair of over-ear, open-back headphones, consisting of one or more physical speaker units working together. For example, driving two speakers, "front-up" and "directly front," with a specific phase and gain combination can effectively create a virtual radiation direction pointing "left-front-up." Each speaker array represents a possible sound wave emission mode, and their number is usually greater than or equal to the number of surround sound channels to provide sufficient mapping degrees of freedom.
[0035] It should be noted that, in this embodiment, head-related transfer data (HRTD) refers to the quantitative characterization of acoustic filtering effects such as frequency response, phase delay, and amplitude attenuation caused by physiological structures such as the head, auricle, and trunk during the propagation of sound waves from a spatial sound source to the eardrum. Specifically, head-related transfer data can be head-related transfer functions (HRTF).
[0036] In this embodiment, the head-related transfer data corresponding to the panoramic sound channels can be obtained in the following ways: based on the three-dimensional coordinates (such as azimuth, elevation, and distance) of each panoramic sound channel in the standard panoramic sound reference layout, a pre-measured or simulated general HRTF database is called to extract the HRTF of the corresponding direction; or an artificial head is used to perform actual measurements on the standard speaker array under the panoramic sound reference layout to obtain the real binaural response data corresponding to each panoramic sound channel.
[0037] Correspondingly, the head-related transmission data for the speaker combination can be obtained by individually stimulating (playing a test signal) each speaker combination while the over-ear circular open-back headphones are being worn, and simultaneously recording the received signals at both ears of the wearer (or artificial head). This process can be calibrated before the over-ear circular open-back headphones leave the factory, or it can be dynamically collected through a built-in calibration program when the user first wears the over-ear circular open-back headphones, ensuring that the head-related transmission data reflects the actual hardware configuration and individual wearing differences.
[0038] Step S200: Based on the head-related transfer data corresponding to each of the panoramic sound channels and the head-related transfer data corresponding to each of the speaker combinations, calculate the degree of difference in head-related transfer data between each of the panoramic sound channels and each of the speaker combinations. It should be noted that, in this embodiment, the degree of difference in head-related transfer data between the immersive sound channel and the speaker combination actually refers to the degree of difference between the head-related transfer data corresponding to the immersive sound channel and the head-related transfer function corresponding to the speaker combination. This is used to measure "the degree to which the auditory effect perceived by the user's ears deviates from the original design intent if a speaker combination in a pair of over-ear open-back headphones is used to replace the speaker corresponding to a certain immersive sound channel in the immersive sound reference layout for audio playback." This degree of difference essentially reflects the inconsistency between the two speaker configurations at the acoustic perception level.
[0039] In practical implementations, the degree of this difference can be quantified using various mathematical or perceptual models. For example, it can be calculated using any one or a combination of the following methods: Frequency domain mean square error: Calculate the mean square error of two sets of HRTF amplitude spectra or complex responses within the key frequency band (such as 0.5 kHz-8 kHz, the sensitive area for human ear localization); Perceived weighted distance: Introducing equal loudness curves or masking thresholds to perform frequency weighting on the mean square error in the frequency domain, making the difference closer to the subjective perception of the human ear; Directional response similarity: Extract directional cue features from HRTF (such as the peak value of instantaneous cross-correlation of binaural time difference and the frequency band energy ratio of binaural sound level difference), and calculate the cosine similarity or Euclidean distance between feature vectors; This embodiment constructs a matrix of differences in head-related transmitted data by performing pairwise calculations on all panoramic sound channels and all speaker combinations. This matrix comprehensively characterizes the boundary of the headphone hardware's ability to reproduce the original panoramic sound field, providing a basis for decision-making in subsequent optimization mapping.
[0040] Step S300: Based on the degree of difference between each of the panoramic sound channels and each of the speaker combinations, determine the speaker combination mapped to each of the panoramic sound channels; It should be noted that, in this embodiment, the speaker combination mapped to the panoramic sound channel refers to the speaker combination designated to play the audio signal corresponding to that panoramic sound channel when the surround sound audio is played on the headphones. Once this mapping relationship is determined, it directly affects the accuracy and immersiveness of the final three-dimensional sound field.
[0041] After obtaining the difference matrix, this embodiment can determine the mapping relationship through various strategies. One approach is the local optimal matching strategy: for each panoramic sound channel, the speaker combination with the smallest difference is independently selected as its mapping object. This method is computationally simple, has high real-time performance, and is suitable for resource-constrained scenarios.
[0042] On the other hand, more preferably, this embodiment supports a global optimization matching strategy: First, all speaker allocation schemes that meet the constraints (such as the speaker combinations assigned to each channel not conflicting with each other, i.e., not sharing the same physical speaker) are enumerated; then, combined with the pre-set channel weights of each immersive sound channel (for example, the "center" immersive sound channel has a higher weight, and the "right surround" immersive sound channel has a lower weight), the difference in head-related transmission data between each immersive sound channel and its corresponding assigned speaker combination in each speaker allocation scheme is weighted and summed to obtain the total difference of each speaker allocation scheme; finally, the speaker allocation scheme with the smallest total difference is selected as the final mapping result. Although this strategy has a high computational complexity, it can ensure the optimal overall listening experience, especially in scenarios where key sound sources (such as vocals and main sound effects) need to be prioritized for high-fidelity reproduction.
[0043] It's worth noting that channel weights can be set based on audio content metadata, user preferences, or default industry rules, reflecting the relative importance of different surround sound channels in the overall listening experience. By introducing a weighting mechanism, the system can intelligently sacrifice the accuracy of non-critical surround sound channels with limited hardware resources in exchange for the ultimate reproduction of critical surround sound channels, achieving "perceptual optimality" rather than "mathematical average optimality."
[0044] This embodiment effectively balances computational efficiency and auditory quality by flexibly supporting local or global mapping strategies and introducing a weight-driven optimization mechanism. This allows the mapping results to approximate the original sound field intent while adapting to actual hardware limitations and content characteristics.
[0045] Step S400: Upon receiving the panoramic sound audio signal, the panoramic sound audio signal is played through the speaker combination mapped to each of the panoramic sound channels.
[0046] It should be noted that the immersive audio signal refers to a multi-channel audio stream containing multiple independent audio channels, which follows a specific immersive audio format, with each channel corresponding to a logical sound source location in the immersive audio reference layout.
[0047] Channel audio signals refer to single-channel audio data belonging to a specific panoramic sound channel in the panoramic sound audio signal, such as the WAV segment of the "left surround" panoramic sound channel.
[0048] In this embodiment, after establishing the mapping relationship, when the system receives the panoramic sound audio signal, it will automatically parse its channel structure and, according to the mapping relationship determined in step S300, route the audio signals of each channel to the corresponding speaker combination for driving output. For example, if panoramic sound channel A under the panoramic sound reference layout is mapped to a speaker combination consisting of physical speaker a and physical speaker b in a headphone ring, then the channel audio signal corresponding to panoramic sound channel A in the panoramic sound audio signal will be sent to these two physical speakers according to a preset ratio, working together to achieve a sound image that is almost equivalent to that generated by panoramic sound channel A under the panoramic sound reference layout.
[0049] This embodiment embeds the mapping relationship between the panoramic sound channel and the speaker combination into the panoramic sound audio playback process, realizing spatial audio adaptation from content decoding to physical sound generation. This allows users to experience an immersive panoramic sound effect close to that of a cinema-grade speaker array through a pair of over-ear circular open headphones, even in mobile or private scenarios.
[0050] This embodiment effectively solves the key technical challenges of sound field reconstruction distortion and 3D positioning inaccuracy when adapting panoramic sound audio signals designed for multi-speaker spatial arrays to the special physical structure of over-ear open-back headphones by constructing a panoramic sound channel-speaker combination mapping mechanism based on measured head-related transfer data differences. Specifically, the method first acquires the head-related transfer data corresponding to each logical panoramic sound channel in a standard panoramic sound reference layout, as well as the head-related transfer data corresponding to several speaker combinations composed of physical speakers in the over-ear open-back headphones. Based on this, the system calculates the degree of difference in head-related transfer data between each panoramic sound channel and each speaker combination (which can be measured by frequency domain mean square error, perceptual weighted distance, or directional response similarity, etc.), thereby quantifying the auditory reproduction capability of different speaker combinations for different panoramic sound channels in the panoramic sound reference layout. Subsequently, based on these differences, a speaker combination is intelligently matched to each immersive sound channel, establishing an optimized mapping relationship from virtual speakers (i.e., immersive sound channels under the immersive sound reference layout) to physical drive units (i.e., speakers in the speaker combination). This ensures that the speaker combinations mapped to each immersive sound channel can reproduce the immersive sound audio playback effect of the immersive sound reference layout as accurately as possible. Finally, upon receiving an immersive sound audio signal containing multi-channel audio signals, the system routes the audio signals of each channel to the corresponding speaker combination for playback according to this mapping relationship, allowing the user to still perceive a three-dimensional sound field structure that matches the creator's intent on the headphones.
[0051] This embodiment breaks through the limitations of traditional binaural rendering, which relies solely on general head-related transfer functions for convolutional synthesis. Instead, it fully utilizes the hardware advantages of the physical layout of multiple speakers in a head-mounted circular open headphone. Through a channel-level matching strategy combined with a spatial rendering algorithm, it achieves high-fidelity, low-distortion spatial audio transfer from an ideal speaker sound field to a constrained headphone platform, significantly improving the immersiveness, directional accuracy, and auditory naturalness of panoramic sound playback in mobile or personal scenarios.
[0052] Based on the first embodiment described above, this application proposes a second embodiment of a panoramic sound headphone playback method.
[0053] In the second embodiment of this application, the same or similar content as in the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0054] like Figure 2 As shown, in this embodiment, step S300 may include steps S310 to S330: Step S310: Select one panoramic sound channel from each of the panoramic sound channels in a round-robin fashion as the target panoramic sound channel. In this embodiment, polling selection refers to sequentially traversing all immersive sound channels according to a preset order (e.g., sorted by channel number, spatial location, or content importance), processing only one immersive sound channel at a time and marking it as the current target immersive sound channel to be mapped. This operation ensures that all immersive sound channels can be independently evaluated and assigned corresponding speaker combinations, avoiding omissions or duplicate processing.
[0055] It should be noted that the target immersive sound channel here does not refer to a channel with special semantics or functions, but rather to any immersive sound channel selected for mapping decisions in the current iteration. Its selection process does not depend on the mapping state of other channels, nor does it introduce global constraints (such as speaker resource mutual exclusion), thus possessing good modularity and low coupling, making it easy to deploy in embedded systems or real-time audio processing scenarios.
[0056] The technical effect of this step is that it decomposes the global problem, which may originally involve high-dimensional combinatorial optimization, into a series of atomic operations of single-channel mapping, significantly reducing the algorithm complexity while maintaining the clarity and maintainability of the implementation logic.
[0057] Step S320: Based on the degree of difference between the target panoramic sound channel and each of the speaker combinations, determine the speaker combination with the smallest degree of difference between the speaker combination and the target panoramic sound channel as the speaker combination mapped to the target panoramic sound channel. In this step, the system calls the difference matrix of head-related transfer data constructed in step S200, and extracts the difference in head-related transfer data between the current target immersive sound channel and all speaker combinations. Subsequently, by comparing these differences, the speaker combination with the smallest difference is selected as the mapping object for the target channel.
[0058] It is worth noting that this embodiment does not impose restrictions on speaker reuse, meaning that a single speaker may be mapped to multiple surround sound channels simultaneously. While this design may result in the superposition of physical speaker drive signals, it is feasible provided that the over-ear circular open-back headphones possess sufficient dynamic range and linear response. More importantly, this strategy prioritizes ensuring the auditory reproduction accuracy of each channel at the individual level, making it particularly suitable for scenarios with a sufficient number of speaker combinations or dense spatial distribution between channels.
[0059] The technical effect of this step is that it achieves optimal perceptual matching at the one-to-one channel level, ensuring that each panoramic sound channel obtains the best physical reproduction path in its independent dimension, thereby approximating the spatial structure and orientation accuracy of the original three-dimensional sound field as a whole.
[0060] Step S330: If the speaker combination mapped to the target panoramic sound channel is determined, return to the step of polling and selecting a panoramic sound channel as the target panoramic sound channel from each of the panoramic sound channels until all panoramic sound channels have been selected, and obtain the speaker combination mapped to each of the panoramic sound channels.
[0061] This step constitutes a closed-loop iterative control process: whenever a target immersive sound channel is mapped, the system automatically returns to step S310, continues to select the next unprocessed immersive sound channel, and repeats the matching logic of S320 until all channels are assigned.
[0062] This loop mechanism has a clear termination condition (i.e., all immersive sound channels have been processed), and each iteration depends only on the static difference matrix, without the need for dynamic updates or backtracking adjustments, thus possessing strong determinism and high execution efficiency.
[0063] Furthermore, since the mapping results are determined entirely by pre-calculated difference data and do not rely on user interaction or online learning, they have good reproducibility and cross-device consistency, which is beneficial for product standardization and unified user experience.
[0064] The technical effect of this step is that it transforms the complex multi-channel mapping task into a linearly scalable serialization processing flow through a structured polling-matching-iteration mechanism.
[0065] This embodiment refines step S300 and proposes a panoramic sound channel mapping method based on polling and local optimal selection. Compared to the global optimization strategy mentioned in the first embodiment, this scheme sacrifices some channel-to-channel collaborative optimization potential, but gains lower computational overhead and simpler implementation logic.
[0066] More importantly, this method is still based on the measured head-related transmission data difference measurement, ensuring that high spatial audio fidelity can be maintained even under the simplified strategy.
[0067] Based on the first embodiment described above, this application proposes a third embodiment of a panoramic sound headphone playback method.
[0068] In the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0069] like Figure 3 As shown, in this embodiment, step S300 may include steps S340-S360: Step S340: Based on each of the panoramic sound channels and each of the speaker combinations, multiple speaker allocation schemes are constructed, wherein the speaker allocation scheme refers to the scheme obtained by allocating a speaker combination to each of the panoramic sound channels from each of the speaker combinations; In this step, the system first enumerates all possible speaker allocation schemes. Each speaker allocation scheme defines a complete mapping from the set of immersive sound channels to the set of speaker combinations, that is, assigning a speaker combination as its playback medium to each immersive sound channel. If the system has N immersive sound channels and M speaker combinations, where N and M are both positive integers, then theoretically, a maximum of M combinations can be generated. N A speaker allocation scheme.
[0070] It is worth noting that this embodiment supports two different speaker allocation scheme generation strategies, corresponding to two implementation methods: Implementation Method 1 (No Duplicate Speakers): When constructing the speaker allocation scheme, it is mandatory that no two speaker combinations assigned to any two immersive sound channels contain the same physical speaker units. That is, in the same speaker allocation scheme, the speakers in the speaker combinations assigned to different immersive sound channels are not duplicated. This constraint ensures that each physical speaker drives only one audio signal corresponding to one immersive sound channel at any given time, avoiding nonlinear distortion or confusion of directional cues caused by audio signal superposition. This method is suitable for scenarios requiring high sound field separation, where the number of physical speakers in over-ear open-back headphones is sufficient, or where extreme positioning accuracy is sought.
[0071] Implementation Method Two (Allowing Speaker Redundancy): When constructing the speaker allocation scheme, physical speaker overlap between speaker combinations is not restricted, allowing multiple immersive sound channels to be mapped to speaker combinations containing the same physical speaker. In this case, the same physical speaker may simultaneously carry audio signals corresponding to multiple immersive sound channels. This method is suitable for scenarios with a small number of speakers but where it is desirable to maximize the use of existing hardware to cover all immersive sound channels.
[0072] Both methods described above can generate a set of candidate speaker allocation schemes based on the same difference matrix. The only difference lies in whether physical speaker mutual exclusion constraints are applied during the scheme selection phase. The system can automatically select the applicable mode based on device configuration files, user settings, or audio content metadata, achieving adaptive speaker allocation strategy.
[0073] Step S350: Calculate the total degree of difference between each of the above-view audio channels and their corresponding speaker combinations in each above-view audio allocation scheme. In this step, for each valid speaker allocation scheme (regardless of whether speaker duplication is allowed), the system quantifies and evaluates its overall auditory distortion level based on the difference data calculated in step S200, thereby calculating the overall distortion metric of each speaker allocation scheme, i.e., the total difference.
[0074] Specifically, the calculation method for the total degree of difference is highly configurable, supporting the following two main modes: Unweighted summation mode: Directly sum the differences in head-related transfer data between all immersive channels and their corresponding speaker combinations in the speaker allocation scheme.
[0075] Channel Weighted Summation Mode: This mode introduces pre-configured channel weights for each immersive sound channel, weighting and summing the differences in head-related transmission data between each immersive sound channel and its corresponding speaker combination. This mode prioritizes ensuring the auditory fidelity of key immersive sound channels (such as center vocals and main sound effects), achieving "perceptual optimality" rather than "mathematical average optimality," which is closer to real-world listening needs.
[0076] Furthermore, when employing an implementation that allows speaker duplication, the system can selectively introduce a multiplexing penalty term to reflect the additional auditory interference that may result from sharing physical speakers. For example, if multiple immersive sound channels are mapped to a speaker combination containing the same physical speakers, a penalty value related to the degree of sharing (such as calculated based on the number of shared speakers or signal energy overlap) can be superimposed on the total degree of difference. However, it should be emphasized that this penalty term is not mandatory. In some application scenarios (such as when speaker linearity is good and the content dynamic range is low), even if multiplexing exists, the actual auditory impact can be ignored. In such cases, the penalty term can be omitted to simplify calculations and avoid overly conservative mapping decisions.
[0077] This step, by flexibly supporting the combination of "whether to weight" and "whether to penalize", enables the evaluation of the total difference to adapt to both high-fidelity professional needs and lightweight consumer-grade implementations, fully demonstrating the scalability and practicality of the algorithm.
[0078] Step S360: Determine the speaker combination mapped to each of the panoramic sound channels based on the speaker allocation scheme with the smallest total difference.
[0079] After calculating the total difference of all candidate speaker allocation schemes, this step selects the scheme with the smallest total difference as the final mapping strategy, and assigns the corresponding speaker combination to each immersive sound channel accordingly.
[0080] It should be noted that regardless of the method used to calculate the overall degree of difference (unweighted / weighted, with penalty / without penalty), the decision-making logic in this step remains consistent: under the current constraints and evaluation criteria, select the feasible solution with the minimum overall auditory distortion.
[0081] It is worth noting that since the definition of the total degree of difference is configurable, "minimum" is always a locally optimal solution relative to the current configuration.
[0082] The technical effect of this step is that, through a unified optimization target selection mechanism, it ensures that the optimal channel-speaker mapping relationship can be output under any configuration strategy, thereby ensuring the universality of the algorithm while taking into account the diverse needs of different hardware platforms and user experience goals.
[0083] This embodiment constructs a speaker allocation scheme set that supports dual-mode constraints (speaker mutual exclusion or reusability) and combines it with a configurable total difference evaluation mechanism (supporting unweighted summation, channel-weighted summation, and optional reusability penalty) to achieve a globally optimized mapping of the immersive sound channels to the physical driver units of the over-ear circular open headphones. Compared to the local strategy of independent matching for each channel in the second embodiment, this embodiment comprehensively considers the collaborative relationship and hardware resource allocation between all channels at the system level. While ensuring high-fidelity reproduction of key sound sources, it effectively improves the spatial consistency, orientation accuracy, and immersion of the overall three-dimensional sound field. More importantly, this method is compatible with different hardware capabilities and application scenarios through parametric design. It can enable a non-repetition, weighted, and penalized mode on high-end devices to approximate the ideal sound field, and can also adopt a reusable, non-penalized, and equal-weighted mode on resource-constrained devices to maintain the basic spatial experience, thereby achieving full coverage from professional to consumer levels within a single technical framework. Therefore, this embodiment not only solves the sound field mismatch problem when multi-speaker headphones play standard immersive audio, but also significantly enhances the robustness, flexibility and user-perceived quality of spatial audio transfer, providing an efficient and scalable core playback engine for next-generation immersive personal audio devices.
[0084] Furthermore, in one feasible implementation, the speaker allocation scheme refers to a scheme obtained by allocating a speaker combination to each of the panoramic sound channels in descending order of channel weights, based on the pre-configured channel weights of each of the panoramic sound channels.
[0085] Specifically, in step S340, during the construction of the speaker allocation scheme, the system first sorts all immersive sound channels in descending order according to their preset channel weights. A higher weight indicates greater importance of the channel in the overall auditory experience (for example, the center channel typically carries vocals and has the highest weight; top or rear surround channels may have lower weights). Then, the system iterates through the sorted channel list. For the currently processed immersive sound channel, among speaker combinations whose physical speakers are not yet occupied by other allocated channels, it selects the speaker combination with the smallest difference in its head-related data transmission as its mapping object and marks this speaker combination as "occupied," prohibiting subsequent channels from using the speakers in this combination. This process continues until all immersive sound channels are allocated, thereby generating a speaker allocation scheme that satisfies the mutual exclusion constraints of physical speakers.
[0086] It should be noted that this implementation method is essentially a construction approach combining a heuristic greedy strategy with a channel priority mechanism: while it does not exhaustively explore all feasible solutions, it effectively guarantees the auditory reproduction quality of key channels by prioritizing high-weight options and selecting locally optimal solutions, thus significantly reducing computational complexity. Since each allocation selects the current optimal option from the remaining available resources, it can maximize overall auditory perception performance even with a limited number of physical speakers in a closed-back, circular headphone.
[0087] Furthermore, this approach is inherently compatible with the "no-repetition speaker" implementation mode, ensuring that each physical speaker serves only one immersive sound channel, avoiding directional cues obscuring the signal due to aliasing. Simultaneously, the channel weights it relies on can be derived from audio content metadata, industry default configurations, or user-defined settings, offering excellent flexibility and practicality.
[0088] This implementation achieves near-globally optimal mapping results by introducing channel importance ranking and exclusive resource allocation mechanisms without introducing high computational overhead from global combinatorial optimization, thus balancing algorithm efficiency, hardware constraints, and subjective listening experience.
[0089] Based on the above embodiments, this application proposes a fourth embodiment of a panoramic sound headphone playback method.
[0090] In the fourth embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0091] like Figure 4As shown, in this embodiment, the head-related transfer data is a head-related transfer function, and the panoramic sound audio signal includes the channel audio signal corresponding to each panoramic sound channel. The step S400 above, which involves playing the panoramic sound audio signal through the speaker combination allocated to each panoramic sound channel, may include steps S410 to S430: Step S410: Calculate the playback configuration corresponding to each of the panoramic sound channels based on the head-related transfer function corresponding to each of the panoramic sound channels and the head-related transfer function corresponding to the speaker combination mapped to each of the panoramic sound channels. In this embodiment, the playback configuration refers to a set of acoustic transformation parameters used to convert the original channel audio signal into a set adapted to the output of the target speaker combination. This includes not only frequency domain filtering functions for correcting direction perception, but also key parameters for restoring the spatial distance perception of the sound source, such as playback distance, gain factor, and high-frequency attenuation coefficient. Its core objective is to compensate for the auditory perception bias introduced by replacing the ideal immersive sound channel with an actual speaker combination, so that the sound image ultimately perceived by the user is as close as possible to the original intent of the content creator under the standard immersive sound reference layout.
[0092] In one example, for each immersive sound channel c i The system is known to: Its ideal spatial position in a standard immersive sound reference layout is typically expressed in three-dimensional coordinates (r). i ,θ i , i ) indicates that r i θ is the distance from the sound source to the listener's head. i It is the azimuth angle. i Angle of elevation; Its ideal HRTF under the standard reference layout is denoted as H. ref,i (f) represents the binaural auditory response that the channel should produce; The measured HRTF of the mapped speaker combination under actual headphone wearing conditions is denoted as H. dev,i (f) represents the binaural auditory response that the speaker combination can currently produce.
[0093] To ensure that the sound and image perceived by the user closely approximate the original design intent in both direction and distance dimensions, the playback configuration needs to include two parts: Direction correction filter F i (f): Used to compensate H dev,i (f) and H ref,i (f) Differences in directional cues are typically obtained through frequency domain deconvolution or optimized fitting, where F i (f)≈Href,i (f) / H dev,i (f). It is worth noting that if the loudspeaker assembly consists of multiple physical loudspeakers, then H dev,i (f) is actually the weighted sum of the HRTFs of each physical speaker in the speaker assembly, while the directional correction filter F i (f) It can be further decomposed into sub-filters distributed to each physical speaker to achieve fine spatial pointing control.
[0094] Distance rendering parameters: including orientation correction filter F i (f) The corresponding playback distance, the amplitude attenuation factor calculated based on the playback distance, and the high-frequency attenuation model (simulating air absorption effects, typically applying a roll-off proportional to the distance for frequencies above 2 kHz). Among these, the directional correction filter F... i (f) The corresponding reproduction distance refers to the F-meter distance required to compensate for the sound perceived by the wearer of the over-ear open-back headphones through a speaker array actually located at a relatively close position (e.g., 10 cm) to be auditorily equivalent to the sound emitted through an ideal immersive sound channel located at a relatively distant position (e.g., 3 meters) in the immersive sound reference layout. i (f) The equivalent distance difference corresponding to that part of the "missing sense of distance".
[0095] This playback configuration determines how the original channel audio signal is spatialized to reconstruct a virtual sound source with correct orientation and distance perception.
[0096] This embodiment improves the accuracy of spatial audio reproduction by constructing a channel-level personalized playback configuration and actively correcting the difference between the hardware acoustic characteristics and the ideal sound field, rather than simply relying on the direct playback of the original signal.
[0097] Step S420: Spatial rendering is performed on the channel audio signals corresponding to each panoramic sound channel according to the playback configuration corresponding to each panoramic sound channel, so as to obtain the combined audio signals corresponding to each panoramic sound channel. In this embodiment, the system performs multi-dimensional spatial rendering on the original channel audio signals of each panoramic sound channel according to the playback configuration corresponding to each panoramic sound channel generated in step S410, and generates a combined audio signal adapted to the combination of the speakers it maps to.
[0098] The spatial rendering described above includes directional coding based on a direction correction filter, as well as gain attenuation and high-frequency roll-off based on distance rendering parameters.
[0099] This embodiment realizes holographic spatial rendering with joint control of direction and distance, enabling the over-ear circular open headphones to not only accurately reproduce the location of the sound source, but also restore its depth position, thereby constructing a three-dimensional auditory scene with a sense of depth.
[0100] Step S430: Play the combined audio signal corresponding to each of the panoramic sound channels through the speaker combinations mapped to each of the panoramic sound channels.
[0101] After completing the spatial rendering of all channels, this step performs the final physical playback operation: the combined audio signal generated by each panoramic sound channel is routed to the speaker combination determined in step S300, and the corresponding physical speaker is driven to output synchronously.
[0102] For example, if the immersive sound channel A is mapped to a speaker combination consisting of physical speakers a and b in a headset with a circular open headphone, the combined audio signal corresponding to the immersive sound channel A will be sent to these two speakers respectively; if multiple immersive sound channels are mapped to the same speaker (in an implementation that allows repetition), real-time mixing and superposition will be performed before driving to ensure that the total output level is within a safe dynamic range.
[0103] Because each audio signal is combined with direction correction and distance rendering, the sound field perceived by the user while wearing the device not only has accurate horizontal / vertical positioning capabilities, but also presents rich front-to-back depth and near-to-far layers, truly achieving an immersive panoramic sound experience.
[0104] This embodiment completes the end-to-end high-fidelity sound field migration from the content creation space to the personal listening device, making the over-ear circular open headphones an immersive audio terminal that can simultaneously reproduce three-dimensional information of orientation, elevation angle and distance.
[0105] This embodiment constructs a complete playback configuration including a direction correction filter and distance rendering parameters, and performs multi-dimensional spatial rendering on this basis. It breaks through the limitations of traditional spatial audio, which only focuses on directionality, and achieves full-dimensional three-dimensional sound field reconstruction for the first time on a head-mounted multi-speaker platform. This method not only inherits the advantages of the previous three embodiments in intelligent channel-speaker mapping, but also introduces a refined acoustic compensation mechanism at its output end. This allows for high-fidelity reproduction of complex soundscapes under a standard immersive sound reference layout, including rich spatial cues such as airplanes flying overhead, echoes from distant valleys, and the intimacy of whispered conversations, even within a compact headphone structure. Therefore, this embodiment significantly improves the naturalness, realism, and immersive depth of immersive sound playback in mobile scenarios, providing key technical support for next-generation personal spatial audio systems.
[0106] Furthermore, in one feasible implementation, the playback configuration includes a playback distance, and prior to step S430 above, the immersive sound headphone playback method may further include steps A10~A40: Step A10: Obtain the pre-measured binaural room impulse response corresponding to each of the panoramic sound channels; In this embodiment, the Binaural Room Impulse Response (BRIR) refers to the binaural time-domain response recorded at the eardrum after a unit impulse signal is emitted from a sound source at a certain spatial location in a specific acoustic environment (such as a standard listening room, a virtual sound field, or a target playback scene). BRIR not only includes direct path information (i.e., the HRTF part) but also fully records early reflections and later reverberation components, making it a key element for achieving immersive spatial audio rendering.
[0107] In this implementation, the BRIR corresponding to each immersive sound channel can be retrieved from a pre-built BRIR database based on its spatial coordinates (azimuth, elevation, and distance) in a standard immersive sound reference layout. This database can be obtained through on-site measurements in a typical room using a human head, or generated through geometric acoustic simulation (such as ray tracing) combined with HRTF interpolation. Each BRIR corresponds one-to-one with one immersive sound channel and is used for subsequent reverberation signal synthesis.
[0108] This implementation provides each logical sound source with environmental acoustic characteristics that match its spatial location, so that the rendering results not only have a sense of direction and distance, but also a realistic sense of room immersion.
[0109] Step A20: Convolve the channel audio signals corresponding to each of the panoramic sound channels using the binaural room impulse response corresponding to each of the panoramic sound channels to obtain the reverberation audio signals corresponding to each of the panoramic sound channels. In this embodiment, the system performs temporal convolution (or efficient frequency domain multiplication) on the original channel audio signal of each immersive sound channel and its corresponding BRIR to generate a reverberant audio signal containing complete spatial reverberation characteristics. This reverberant audio signal includes direct sound, early reflections, and late reverberation, and can simulate all the auditory cues perceived by the user when the immersive sound channel plays its corresponding channel audio signal in a specific acoustic environment.
[0110] This implementation injects reverberation features into each sound source that match its spatial location and environmental semantics, significantly enhancing the spatial immersion and realism of the sound field.
[0111] Step A30: Calculate the wet / dry mixing ratio for each of the panoramic sound channels based on the playback distance corresponding to each of the panoramic sound channels. In this embodiment, the dry-wet mixing ratio is used to control the energy ratio between the direct sound (dry signal, i.e. the combined audio signal output in step S420) without reverberation processing and the reverberant sound (wet signal, i.e. the reverberant audio signal) generated by BRIR convolution in the final output signal.
[0112] Specifically, the system uses the direction correction filter F determined in step S410. i (f) The corresponding playback distance is used to calculate the wet-dry mixing ratio using a psychoacoustic model. The general rule is: the smaller the playback distance, the higher the proportion of direct sound and the lower the proportion of reverberation.
[0113] This wet-dry mixing ratio ensures that near sound sources (such as whispers) are clear and clean, while far sound sources (such as valley echoes) are spacious. This aligns with the natural human perception of the relationship between distance and reverberation, achieving an energy balance between direct sound and reverberant sound. It prevents near sound sources from being submerged by reverberation or far sound sources from sounding "dry," thereby enhancing the realism of spatial layers.
[0114] Step A40: According to the dry and wet mixing ratio of each of the panoramic sound channels, add the reverberation audio signal of each panoramic sound channel to the combined audio signal of each panoramic sound channel to obtain a new combined audio signal of each panoramic sound channel.
[0115] In this embodiment, the system weights and superimposes the dry signal (which already includes a combined audio signal with direction correction and distance attenuation) generated in step S420 and the wet signal (reverberation audio signal) generated in step A20 according to the dry-wet mixing ratio calculated in step A30, to form the final enhanced combined audio signal for playback. This makes the final immersive headphone playback effect not only accurately reproduce the direction and distance of the sound source, but also incorporate the room acoustic characteristics that match its spatial location, providing users with a more immersive and immersive listening experience. It achieves the perceptual coordination and fusion of direct sound and reverberation sound, so that the virtual sound source is both accurately located and naturally integrated into the acoustic environment, significantly improving the immersive depth and realism of immersive sound content.
[0116] This implementation further enhances the environmental realism of spatial audio by introducing BRIR-based reverberation synthesis and a playback distance-driven dry-wet mixing mechanism, building upon the fourth embodiment. This method is particularly suitable for applications such as movies, games, and VR / AR that demand a high level of sound field immersion, enabling over-ear open-back headphones not only to "locate the sound source" but also to "immerse oneself in the sound field," truly achieving a leap in auditory experience from "hearing the sound" to "being there."
[0117] To facilitate understanding of the above embodiments and implementation methods, a specific embodiment is provided below: In this specific embodiment, such as Figure 5 As shown, the multiple speakers in the over-ear circular open-back headphones are designed in a circular array. When worn, the multiple speakers are arranged in a ring around the user's ear.
[0118] The immersive audio output includes directional channels such as front, back, left, right, and height, as well as a low-frequency effects channel. The low-frequency effects channel typically contains only low-frequency components below 200 Hz and is non-directional. Therefore, this specific embodiment only describes the playback and rendering method for the non-low-frequency effects channel. The low-frequency effects can be played back simultaneously by all headphone speakers or using an additional subwoofer, without requiring special processing.
[0119] A mapping relationship is established and optimized between the immersive sound channels and the headphone speakers. Each immersive sound channel is reproduced by a speaker combination consisting of one or more speakers. The optimization strategy aims to minimize the head-related transfer function deviation and reduce the spatial localization difference of the reproduced channels.
[0120] Specifically, the immersive sound channels under the immersive sound reference layout can be divided into 5 groups: (1) left / right / center; (2) left / right surround; (3) left rear / right rear surround; (4) left front / right front height; (5) left rear / right rear height.
[0121] The earphone ring is also divided into 5 areas in a clockwise direction, with the top of the earphone ring at 0° and the front at 90°.
[0122] Preferably, the correspondence between the angle of each zone and the panoramic sound channel is as follows: Figure 6 As shown: (1) Left / Right / Center: 90°~180°; (2) Left / Right Surround: 135°~225°; (3) Left Rear / Right Rear Surround: 180°~270°; (4) Left Front / Right Front Height: 0°~90°; (5) Left Rear / Right Rear Height: 270°~360°.
[0123] Each zone may contain multiple speakers. When it is determined that multiple speakers will be used to reproduce the same immersive sound channel, signal attenuation should be set to ensure energy balance. The attenuation gain is preferably the reciprocal of the number of speakers mapped to the immersive sound channel.
[0124] Furthermore, in order to minimize the playback positioning deviation of the immersive sound channels, for each channel, it is necessary to determine the optimal number and location of speakers used in the corresponding headphone zone, that is, to optimize the mapping relationship and determine the speaker combination mapped for each immersive sound channel.
[0125] In this specific embodiment, the process of optimizing the mapping relationship is as follows: Figure 7 As shown, firstly, using a standard artificial head, such as KEMAR (Knowles Electronic Manikin for Acoustic Research), under a Dolby Atmos reference layout, the head-related transfer function from each Dolby Atmos channel to the two ears of the artificial head was measured, denoted as... Secondly, when wearing over-ear circular open-back headphones, all speaker combinations in each headphone section are statistically analyzed, denoted as M speaker combinations; for each speaker combination, the head-related transfer function to the artificial head's two ears is measured, denoted as... Finally, minimize the mean square difference between the two head-related transfer functions, i.e. This allows for the calculation of the degree of difference in head-related transmission data between each immersive sound channel and each speaker combination, thereby determining the speaker combination mapped to each immersive sound channel.
[0126] In one example, 12 identical speaker units are evenly distributed on the headphone surround at a 30-degree angular resolution, numbered 1, 2, 3...12 according to angles of 30°, 60°...360°, respectively. The two earpieces of the over-ear circular open-back headphones are symmetrically arranged. The speaker combinations mapped to each immersive sound channel under the immersive sound reference layout are as follows: the left / right channels use speakers 3 and 4 on the headphone surround, the center channel uses speakers 4 and 5 on the headphone surround, the left / right surround channels use speakers 5, 6, and 7 on the headphone surround, the left / right rear surround channels use speaker 8 on the headphone surround, the left / right front sky channels use speakers 1 and 2 on the headphone surround, and the left / right rear sky channels use speakers 9 and 10 on the headphone surround.
[0127] Optimizing the mapping relationship allows the listener's personalized auricular cues to play a greater role in spatial localization, thereby reducing confusion between front and back and resolving the lack of height perception. However, due to the inherent differences between the speaker layout of the over-ear circular open-back headphones and the immersive sound reference layout, the localization deviation between the two can only be minimized to a low level, not completely eliminated. Therefore, it is necessary to use a spatialized rendering algorithm to compensate for this.
[0128] This specific embodiment proposes a spatialized rendering algorithm that generates binaural signals for each panoramic sound channel in the panoramic sound reference layout to compensate for the positioning deviation of that panoramic sound channel.
[0129] In this specific embodiment, the spatial rendering process is as follows: Figure 8 As shown. Specifically, taking the panoramic sound channel as a unit, the three spatial rendering parameters of distance (i.e., playback distance), horizontal angle, and vertical angle are iteratively adjusted to match the sound image perception with the panoramic sound reference layout. After determining the values of the three parameters, the direct path rendering is performed: far-field attenuation and near-field correction are applied according to the distance; spatial interpolation is performed in the HRTF database of the standard artificial head according to the size of the horizontal angle and vertical angle; the convolution of the interpolated HRTF (i.e., direction correction filter) with the channel audio signal is performed using frequency domain filtering to form the output signal of the direct path (i.e., the combined audio signal determined in step S420).
[0130] Furthermore, to address the head-in-the-head effect, reverberation path rendering is also required to simulate the real-world listening environment. The reverberation path rendering employs the Ambisonics method, measuring the binaural room impulse responses in six directions (front, back, left, right, top, and bottom) in an actual room. First-order Ambisonics (FOA) encoding / decoding technology is used to convert the binaural room impulse responses into Ambisonic-to-Binaural Impulse Response (ABIR). The gain (i.e., the dry / wet mixing ratio) controlled by the direct mixing ratio is calculated based on the playback distance and applied to the channel audio signal for attenuation or amplification. The gain-processed channel audio signal is encoded into FOA format. The FOA-encoded signal is convolved with the ABIR to form the output of the reverberation path (i.e., the reverberated audio signal).
[0131] Due to room reflections, the duration of ABIR sequences can be as long as hundreds of milliseconds. To reduce computational costs, ABIR convolution can be performed using a block filtering method.
[0132] The output of the final reverberation path is superimposed on the output of the direct path to form the final binaural signal (i.e., the new combined audio signal determined in step A40).
[0133] In this specific embodiment, for different surround sound channels, the spatially rendered binaural signal output is distributed to speaker combinations in different headphone zones for playback. This zoned rendering method ensures that surround sound channels do not interfere with each other, while fully leveraging the natural positioning advantages of different zone mappings.
[0134] In this specific embodiment, the process steps for playing back the immersive sound headphones are as follows: Figure 9 As shown: Step 1: Divide the surround sound channel into five groups, and divide the headphone ring into five zones accordingly. The zone settings for the left and right ears are the same, and determine the angle range of each zone. Step 2: Optimize the mapping relationship between the immersive sound channel and the speaker combination in the headphone zone. By minimizing the head-related transfer function deviation, determine the number of speakers and speaker positions in the speaker combination required to play each immersive sound channel. Step 3: Spatially render the audio signal corresponding to each panoramic sound channel to generate multiple pairs of binaural signals, which are then distributed to the speaker combinations in the corresponding zones of the left and right ears of the headphones for playback.
[0135] It should be noted that the above embodiments / implementations are only used to assist in understanding this application and do not constitute a limitation on the panoramic sound headphone playback method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0136] In addition, this application also provides an audio device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the panoramic sound headphone playback method in the above embodiments.
[0137] The following is for reference. Figure 10 It shows a structural schematic diagram of an audio device suitable for implementing the embodiments of this application. Figure 10 The audio device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0138] like Figure 10 As shown, the audio device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the audio device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the audio device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagram shows audio equipment with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0139] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0140] The audio device provided in this application, employing the panoramic sound headphone playback method described in the above embodiments, solves the technical problem of how to effectively adapt panoramic sound audio signals originally designed for spatial speaker arrays to the output of headphones. Compared with the prior art, the beneficial effects of the audio device provided in this application are the same as those of the panoramic sound headphone playback method provided in the above embodiments, and other technical features of this audio device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0141] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above claims.
[0142] In addition, this application also provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the steps of the panoramic sound headphone playback method in the above embodiments.
[0143] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0144] The aforementioned computer-readable storage medium may be included in an audio device or may exist independently without being assembled into an audio device.
[0145] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an audio device, cause the audio device to: acquire head-related transfer data corresponding to each immersive sound channel in an immersive sound reference layout, and head-related transfer data corresponding to each speaker combination in a headset; calculate the degree of difference between the head-related transfer data of each immersive sound channel and each speaker combination based on the head-related transfer data corresponding to each immersive sound channel and the head-related transfer data corresponding to each speaker combination; determine the speaker combination mapped to each immersive sound channel based on the degree of difference between each immersive sound channel and each speaker combination; and, upon receiving an immersive sound audio signal, play the immersive sound audio signal through the speaker combination mapped to each immersive sound channel.
[0146] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0148] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0149] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for performing the steps of the above-described immersive sound headphone playback method, which solves the technical problem of how to effectively adapt immersive sound audio signals originally designed for spatial speaker arrays to the output of headphones. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the immersive sound headphone playback method provided in the above embodiments, and will not be repeated here.
[0150] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the panoramic sound headphone playback method described in the above embodiments.
[0151] The computer program product provided in this application solves the technical problem of how to effectively adapt immersive audio signals originally designed for spatial speaker arrays to the output of headphones. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the immersive audio headphone playback method provided in the above embodiments, and will not be repeated here.
[0152] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for reproducing surround sound from headphones, characterized in that, The method includes: Acquire head-related transfer data for each immersive sound channel in the immersive sound reference layout, and head-related transfer data for each speaker combination in the over-ear circular open headphones; Based on the head-related transfer data corresponding to each of the panoramic sound channels and the head-related transfer data corresponding to each of the speaker combinations, the degree of difference in head-related transfer data between each of the panoramic sound channels and each of the speaker combinations is calculated. Based on the degree of difference between each of the panoramic sound channels and each of the speaker combinations, the speaker combination mapped to each of the panoramic sound channels is determined; Upon receiving a panoramic sound audio signal, the panoramic sound audio signal is played through the speaker combination mapped to each of the panoramic sound channels.
2. The method as described in claim 1, characterized in that, The step of determining the speaker combination assigned to each of the immersive sound channels based on the degree of difference between each of the immersive sound channels and each of the speaker combinations includes: One panoramic sound channel is selected from each of the panoramic sound channels in a round-robin fashion as the target panoramic sound channel; Based on the degree of difference between the target immersive sound channel and each of the speaker combinations, the speaker combination with the smallest degree of difference between the speaker combination and the target immersive sound channel is determined as the speaker combination mapped to the target immersive sound channel. If the speaker combination mapped to the target panoramic sound channel is determined, the process returns to the step of polling and selecting a panoramic sound channel as the target panoramic sound channel from each of the panoramic sound channels until all panoramic sound channels have been selected, thus obtaining the speaker combination mapped to each of the panoramic sound channels.
3. The method as described in claim 1, characterized in that, The step of determining the speaker combination mapped to each of the panoramic sound channels based on the degree of difference between each of the panoramic sound channels and each of the speaker combinations includes: Based on each of the panoramic sound channels and each of the speaker combinations, multiple speaker allocation schemes are constructed, wherein the speaker allocation scheme refers to the scheme obtained by allocating a speaker combination to each of the panoramic sound channels from each of the speaker combinations; The total degree of difference between each speaker allocation scheme is calculated based on the degree of difference between each immersive sound channel and its corresponding allocated speaker combination in each speaker allocation scheme. The speaker combination for each of the panoramic sound channels is determined based on the speaker allocation scheme with the minimum total difference.
4. The method as described in claim 3, characterized in that, The step of calculating the total difference of each speaker allocation scheme based on the degree of difference between each immersive sound channel and its corresponding allocated speaker combination in each speaker allocation scheme includes: Obtain the pre-configured channel weights of each of the aforementioned panoramic sound channels; Based on the channel weights of each of the panoramic sound channels, the degree of difference between each panoramic sound channel and its corresponding allocated speaker combination is weighted and summed in each of the speaker allocation schemes to calculate the total degree of difference of each speaker allocation scheme.
5. The method as described in claim 3 or 4, characterized in that, In the same speaker allocation scheme, the speakers in the speaker combinations allocated to different surround sound channels are not duplicated.
6. The method as described in claim 5, characterized in that, The speaker allocation scheme refers to a scheme in which a speaker combination is allocated to each of the panoramic sound channels in descending order of channel weights, based on the pre-configured channel weights of each panoramic sound channel.
7. The method as described in claim 1, characterized in that, The head-related transfer data is a head-related transfer function, the panoramic sound audio signal includes channel audio signals corresponding to each panoramic sound channel, and the step of playing the panoramic sound audio signal through the speaker combination allocated to each panoramic sound channel includes: Based on the head-related transfer function corresponding to each of the panoramic sound channels and the head-related transfer function corresponding to the speaker combination mapped to each of the panoramic sound channels, the playback configuration corresponding to each of the panoramic sound channels is calculated. Based on the playback configuration of each panoramic sound channel, spatial rendering is performed on the channel audio signal corresponding to each panoramic sound channel to obtain the combined audio signal corresponding to each panoramic sound channel. Each panoramic sound channel is mapped to a speaker combination, which plays the combined audio signal corresponding to each panoramic sound channel.
8. The method as described in claim 7, characterized in that, The playback configuration includes a playback distance. Before the step of playing the combined audio signal corresponding to each of the panoramic sound channels through the speaker combinations mapped to each of the panoramic sound channels, the method further includes: Obtain the pre-measured binaural room impulse response corresponding to each of the aforementioned panoramic sound channels; By convolving the channel audio signals corresponding to each of the panoramic sound channels with the binaural room impulse response corresponding to each of the panoramic sound channels, the reverberation audio signals corresponding to each of the panoramic sound channels are obtained. Based on the playback distance corresponding to each of the panoramic sound channels, the dry and wet mixing ratios corresponding to each panoramic sound channel are calculated respectively. According to the dry and wet mixing ratio of each of the panoramic sound channels, the reverberation audio signal of each panoramic sound channel is added to the combined audio signal of each panoramic sound channel to obtain a new combined audio signal of each panoramic sound channel.
9. An audio device, characterized in that, The audio device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the immersive sound headphone playback method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an immersive sound headphone playback program, which, when executed by a processor, implements the steps of the immersive sound headphone playback method as described in any one of claims 1 to 8.