Binaural rendering method and device, audio equipment and computer readable storage medium

By employing a binaural rendering method that separates audio sources and configures differentiated playback, the problem of blurred vocal localization on near-ear open-back audio devices is solved, achieving clear vocals and a wide sound field, thus enhancing the spatial sense and immersive experience of audio devices.

CN121509894APending Publication Date: 2026-02-10GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511603846.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In near-ear open-back audio devices, existing technologies struggle to achieve clear voice localization while maintaining a wide sound field, resulting in problems such as blurred voice localization and diffused sound images. This is particularly problematic in applications like movie dialogues and voice broadcasts, affecting auditory clarity and subjective listening quality.

Method used

By performing source separation processing on the original audio signal, the human voice and background sound are separated into independent signals, and binaural rendering is performed using differentiated playback configurations. The human voice playback configuration is rendered with a smaller playback angle, a closer playback distance, and a lower absolute value of the elevation angle, while the background sound playback configuration is rendered with a larger playback angle, a farther playback distance, and a higher elevation angle, and finally mixed output.

Benefits of technology

It achieves clear positioning of human voices and wide diffusion of background sounds, enhancing the overall spatial sense and immersiveness of the audio, making it particularly suitable for applications centered on human voices, such as movie playback and voice calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509894A_ABST
    Figure CN121509894A_ABST
Patent Text Reader

Abstract

The invention discloses a binaural rendering method and device, audio equipment and a computer readable storage medium, and relates to the technical field of audio equipment, and the method comprises the steps: carrying out the sound source separation processing of an original audio signal received by near-ear open audio equipment, and obtaining an original human sound audio signal and an original background sound audio signal; performing binaural rendering on the original human voice audio signal based on the human voice playback configuration to obtain a target human voice audio signal, and performing binaural rendering on the background voice audio signal based on the background voice playback configuration to obtain a target background voice audio signal, the spatial perceptibility corresponding to the human voice replay configuration is higher than the spatial perceptibility corresponding to the background voice replay configuration; and carrying out sound mixing processing on the target human sound audio signal and the target background sound audio signal to obtain a target audio signal, and playing the target audio signal through near-ear open audio equipment. According to the invention, a binaural rendering effect with spatial width and clear positioning without losing voice can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a binaural rendering method, apparatus, audio device, and computer-readable storage medium. Background Technology

[0002] Stereo audio, as the most mainstream audio format today, is widely used in near-ear open-back audio devices such as headphones and smart audio glasses. Because these devices are close to the ear and do not seal the ear canal, users have higher requirements for the spatial sense and naturalness of the sound. To enhance the immersive experience, those skilled in the art often use the Head Related Transfer Function (HRTF) to perform binaural rendering of stereo audio signals to achieve a sound field expansion effect, making the sound seem to come from a wider spatial area, simulating the auditory experience of a multi-speaker system.

[0003] However, in practical applications, it has been found that although the sound field width is improved after using HRTF for binaural rendering, the vocal parts often exhibit blurred positioning and diffuse sound images, giving an unnatural feeling that the sound is "in the head" or "floating from the sides." This problem is particularly noticeable in content dominated by human voices, such as film and television dialogues and voice broadcasts, seriously affecting auditory clarity and subjective listening quality.

[0004] Therefore, how to achieve a binaural rendering effect that has both spatial width and clear vocal positioning on near-ear open-back audio devices remains a technical challenge in this field. Summary of the Invention

[0005] The main objective of this application is to provide a binaural rendering method, apparatus, audio device, and computer-readable storage medium, aiming to solve the technical problem of how to achieve a binaural rendering effect that has both spatial width and clear vocal positioning on near-ear open audio devices.

[0006] To achieve the above objectives, this application provides a binaural rendering method, which includes the following steps: The original audio signal received by the near-ear open-back audio device is processed by source separation to obtain the original human voice audio signal and the original background sound audio signal; Based on a preset human voice playback configuration, the original human voice audio signal is rendered in binaurally to obtain a target human voice audio signal. Based on a preset background sound playback configuration, the background sound audio signal is rendered in binaurally to obtain a target background sound audio signal. The spatial perception corresponding to the human voice playback configuration is higher than that corresponding to the background sound playback configuration. The target human voice audio signal and the target background sound audio signal are mixed to obtain a target audio signal, which is then played through the near-ear open-back audio device.

[0007] In one embodiment, the voice playback configuration includes a voice playback angle, the background sound playback configuration includes a background sound playback angle, and the voice playback angle is smaller than the background sound playback angle.

[0008] In one embodiment, the voice playback configuration includes a voice playback distance, the background sound playback configuration includes a background sound playback distance, and the voice playback distance is less than the background sound playback distance.

[0009] In one embodiment, the voice playback configuration includes a voice playback elevation angle, the background sound playback configuration includes a background sound playback elevation angle, and the absolute value of the voice playback elevation angle is less than the absolute value of the background sound playback elevation angle.

[0010] In one embodiment, the step of performing source separation processing on the raw audio signal received by the near-ear open-back audio device to obtain the raw human voice audio signal and the raw background sound audio signal includes: The raw audio signal received by the near-ear open-back audio device is input into a pre-trained sound source separation model to obtain the raw human voice audio signal and the raw background sound audio signal output by the sound source separation model.

[0011] In one embodiment, the steps of performing binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and performing binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, include: Get the human voice head-related transfer function corresponding to the preset human voice playback configuration, and the background sound head-related transfer function corresponding to the preset background sound playback configuration; The original human voice audio signal is binaurally rendered using the human voice head correlation transfer function to obtain the target human voice audio signal. The original background sound audio signal is rendered by binaural rendering using the background sound head correlation transfer function to obtain the target background sound audio signal.

[0012] In one embodiment, the original background sound audio signal includes a first original background sound audio signal corresponding to a first background sound source and a second original background sound audio signal corresponding to a second background sound source, wherein the first background sound source is different from the second background sound source; The background sound playback configuration includes a first background sound playback configuration corresponding to the first background sound source and a second background sound playback configuration corresponding to the second background sound source. The background sound head-related transfer function includes a first background sound head-related transfer function corresponding to the first background sound playback configuration and a second background sound head-related transfer function corresponding to the second background sound playback configuration. The target background sound audio signal includes a first target background sound audio signal corresponding to the first background sound source and a second target background sound audio signal corresponding to the second background sound source. The step of performing binaural rendering on the original background sound audio signal using the background sound head correlation transfer function to obtain the target background sound audio signal includes: The first original background sound audio signal is rendered by binaural rendering using the first background sound head correlation transfer function to obtain the first target background sound audio signal. The second original background sound audio signal is rendered by binaural rendering using the second background sound head correlation transfer function to obtain the second target background sound audio signal.

[0013] Furthermore, to achieve the above objectives, this application also provides a binaural rendering device, the binaural rendering device comprising: The separation module is used to perform source separation processing on the raw audio signal received by the near-ear open-back audio device to obtain the original human voice audio signal and the original background sound audio signal; The rendering module is used to perform binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and to perform binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration. The mixing module is used to mix the target human voice audio signal and the target background sound audio signal to obtain a target audio signal, and to play the target audio signal through the near-ear open-back audio device.

[0014] In addition, to achieve the above objectives, this application also provides an audio device, the audio device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the binaural rendering method as described above.

[0015] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of the binaural rendering method as described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the binaural rendering method described above.

[0017] This application provides a binaural rendering method, apparatus, audio device, and computer-readable storage medium, relating to the field of audio processing technology. The method includes: performing source separation processing on the original audio signal received by a near-ear open-back audio device to obtain an original human voice audio signal and an original background sound audio signal; performing binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and performing binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration; performing mixing processing on the target human voice audio signal and the target background sound audio signal to obtain a target audio signal, and playing the target audio signal through the near-ear open-back audio device.

[0018] This application embodiment effectively solves the technical challenge of balancing a wide sound field and clear vocal localization on near-ear open-back audio devices by employing a differentiated binaural rendering technology based on sound source separation. Specifically, this application embodiment first performs sound source separation processing on the original audio signal, extracting the vocal and background sound components into independent audio streams. Subsequently, a differentiated binaural rendering strategy is adopted for different sound source types: the original vocal audio signal is rendered using a vocal reproduction configuration with high spatial awareness (such as a smaller reproduction angle, a closer reproduction distance, and a lower absolute value of the elevation angle) to ensure focused vocal imaging and accurate localization, avoiding sound image blurring caused by excessive spatialization; simultaneously, the original background sound audio signal is rendered using a background sound reproduction configuration with relatively low spatial awareness, resulting in a wider and / or deeper diffusion effect, thereby creating an immersive sound field atmosphere; finally, the rendered target vocal audio signal and target background sound audio signal are mixed and output, and then played through a near-ear open-back audio device.

[0019] Through the above-mentioned technical means, the embodiments of this application achieve the auditory effect of "condensed and not scattered" human voice and "wide and not chaotic" background sound. While ensuring the clarity of speech and the stability of sound image, it significantly improves the spatial sense and immersion of the overall audio, and is especially suitable for application scenarios with human voice as the core content, such as movie playback and voice calls. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the first embodiment of the binaural rendering method of this application; Figure 2 This is a schematic diagram illustrating the training of the sound source separation model in a specific embodiment of this application; Figure 3 This is a flowchart illustrating the second embodiment of the binaural rendering method of this application; Figure 4 This is a schematic diagram of a binaural rendering scene in a specific embodiment of this application; Figure 5 This is a schematic diagram of differentiated binaural rendering in a specific embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the binaural rendering method in the embodiments of this application.

[0023] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0024] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0025] Currently, traditional binaural rendering methods mainly rely on HRTF to spatialize stereo or multi-channel audio signals, enabling the reconstruction of spatial auditory scenes in binaural headphones or near-ear open-back audio devices. This method simulates the differences in acoustic characteristics (such as time difference, intensity difference, and spectral characteristics) as sound reaches both ears from different directions, thereby achieving the perceptual restoration of sound source location and enhancing the overall sound field width and immersion.

[0026] However, in practical applications, especially when using traditional HRTF-based binaural rendering on near-ear open-back audio devices (such as open-ear clip-on headphones, bone conduction headphones, or on-ear speaker systems), a common problem is the degradation of voice localization performance. Specifically, this manifests as blurred sound image localization, diffuse sound field imaging, and insufficient center focus in the voice portion, leading to the listener's subjective auditory illusion that "the sound is floating inside the head" or "coming from unnaturally on the left and right sides"—a typical "in-head effect" or sound image instability. The root of this problem lies in the fact that traditional binaural rendering typically applies the same HRTF core or spatialization parameters to all frequency bands or the entire audio content, failing to fully consider the dominant role of the voice signal in auditory perception and its crucial role in anchoring the sound image center. Especially in application scenarios where voice is the core content, such as film and television dialogues, voice broadcasts, and video conferencing, the clarity, intelligibility, and spatial localization accuracy of the voice directly determine the quality of the user's auditory experience. When human voices are overly spatialized or treated the same as background sounds, their central sound image is easily weakened or dispersed, thereby destroying the focus and naturalness of the speech.

[0027] To achieve a binaural rendering effect on near-ear open-back audio devices that provides both spatial width and clear vocal localization, the solution of this application embodiment is a binaural rendering method, comprising: performing source separation processing on the original audio signal received by the near-ear open-back audio device to obtain an original human voice audio signal and an original background sound audio signal; performing binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and performing binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration; performing mixing processing on the target human voice audio signal and the target background sound audio signal to obtain a target audio signal, and playing the target audio signal through the near-ear open-back audio device.

[0028] This application embodiment effectively solves the technical challenge of balancing a wide sound field and clear vocal localization on near-ear open-back audio devices by employing a differentiated binaural rendering technology based on sound source separation. Specifically, this application embodiment first performs sound source separation processing on the original audio signal, extracting the vocal and background sound components into independent audio streams. Subsequently, a differentiated binaural rendering strategy is adopted for different sound source types: the original vocal audio signal is rendered using a vocal reproduction configuration with high spatial awareness (such as a smaller reproduction angle, a closer reproduction distance, and a lower absolute value of the elevation angle) to ensure focused vocal imaging and accurate localization, avoiding sound image blurring caused by excessive spatialization; simultaneously, the original background sound audio signal is rendered using a background sound reproduction configuration with relatively low spatial awareness, resulting in a wider and / or deeper diffusion effect, thereby creating an immersive sound field atmosphere; finally, the rendered target vocal audio signal and target background sound audio signal are mixed and output, and then played through a near-ear open-back audio device.

[0029] Through the above-mentioned technical means, the embodiments of this application achieve the auditory effect of "condensed and not scattered" human voice and "wide and not chaotic" background sound. While ensuring the clarity of speech and the stability of sound image, it significantly improves the spatial sense and immersion of the overall audio, and is especially suitable for application scenarios with human voice as the core content, such as movie playback and voice calls.

[0030] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0031] This application proposes a binaural rendering method according to a first embodiment.

[0032] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the binaural rendering method of this application.

[0033] In this embodiment, the binaural rendering method may include steps S100~S300: Step S100: Perform source separation processing on the original audio signal received by the near-ear open-back audio device to obtain the original human voice audio signal and the original background sound audio signal. Those skilled in the art will recognize that near-ear open-back audio devices refer to a type of audible terminal device worn on the user's ear or close to the opening of the ear canal, but not completely sealing the external auditory canal. Typical examples include open-back headphones, semi-open clip-on headphones, bone conduction headphones, and on-ear speaker systems. These devices allow ambient sound to enter the ear naturally, offering good transparency and wearing comfort, and are widely used in outdoor sports, voice interaction, and long-term wear scenarios. However, due to the lack of physical sound insulation and a sealed coupling cavity, traditional HRTF-based spatial audio rendering technology is prone to problems such as blurred sound image localization and reduced sound field focusing ability on these devices, especially when playing content dominated by human voices.

[0034] It should be noted that, in this embodiment, sound source separation processing refers to the process of decoupling and independently extracting different sound source components from a multi-source mixed audio signal in the time-frequency domain using signal processing algorithms or deep learning models. Specifically, this process identifies and separates the human voice component from the non-human voice component by analyzing the spectral characteristics, temporal envelope, spatial cues (such as phase difference and intensity difference), and semantic information (such as speech fundamental frequency and formant distribution) of the audio signal, thereby generating two independent audio streams from different sound sources: a human voice audio signal and a background sound audio signal.

[0035] It should also be noted that, in this embodiment, the original audio signal refers to the input audio data (which can be stereo audio data) received or decoded by the near-ear open-back audio device. It usually comes from film and television content, voice calls, virtual reality applications or other multimedia sources, and contains complete auditory scene information, including both human voice dialogue and background sound elements such as environmental sound effects and music.

[0036] The original human voice audio signal refers to the audio signal stream obtained after sound source separation processing, which contains only the human voice portion (such as speaking or singing). This signal retains key speech information from the original audio signal, such as pronunciation content, intonation changes, and lip-sync, and is the core object for subsequent high-precision sound image localization.

[0037] The original background audio signal refers to the collection of all sound components in the original audio signal except for human voices, obtained after source separation processing. This includes, but is not limited to, environmental noise, background music, spatial reverberation, action sound effects (such as footsteps and door opening / closing sounds), and natural sounds (such as wind and rain). This signal carries the spatial atmosphere and immersive information of the auditory scene in the original audio signal, and is used to construct a wide and layered sound field experience.

[0038] For example, in one feasible implementation, step S100 above, which performs source separation processing on the original audio signal received by the near-ear open-back audio device to obtain the original human voice audio signal and the original background sound audio signal, may include step S110: Step S110: Input the raw audio signal received by the near-ear open-ear audio device into the pre-trained sound source separation model to obtain the raw human voice audio signal and the raw background sound audio signal output by the sound source separation model.

[0039] It should be noted that sound source separation models refer to a type of audio signal processing system built on machine learning or deep learning architectures. Their function is to automatically identify and separate different sound source components from mixed stereo audio signals. This model typically uses neural networks as its core, learning time-frequency features, harmonic structures, spatial cues, and semantic patterns from large amounts of labeled audio data to establish the ability to distinguish between human and non-human voice components, thereby achieving high-precision audio demixing.

[0040] In this embodiment, the audio source separation model was trained offline using a large-scale training dataset before deployment. The training dataset contains pairs of mixed audio samples and corresponding real, clean audio labels. The mixed audio samples simulate various superpositions of human voice and background noise in real-world scenarios (such as dialogue superimposed with background music, speech superimposed with environmental noise, etc.), covering different languages, genders, speech rates, music styles, noise types, and signal-to-noise ratio conditions. The clean audio labels are the corresponding human voice audio segments and background sound audio segments, respectively. During training, the model continuously optimizes the network parameters by minimizing the temporal waveform error or time-frequency spectrogram difference between the separation result and the label as the objective function until it achieves generalization ability. The training process of this audio source separation model is as follows: Figure 2 As shown. After training, the model is embedded and integrated into near-ear open-back audio devices or paired terminal systems for real-time sound source separation tasks.

[0041] In actual operation, when the original audio signal is input into the model, it is first converted into a time-frequency representation (such as a short-time Fourier transform spectrum). Then, the model analyzes and predicts the time-frequency masks for the human voice and background sound frame by frame. These masks are then restored to the time domain signal through an inverse transform, ultimately outputting independent original human voice and background sound audio signals. This process has advantages such as fast processing speed, high separation accuracy, and strong adaptability to complex auditory scenes, and is especially suitable for dynamically changing multi-sound-source environments.

[0042] This implementation significantly improves the decoupling accuracy of human voice and background sound by introducing a deep learning-based sound source separation model. It effectively avoids the problems of "ghosting" or "crosstalk" that easily occur in high-frequency overlapping areas using traditional rule-based or simple filtering methods. Its technical effect is that it provides a high-quality, low-interference independent sound source signal foundation for subsequent differentiated binaural rendering, ensuring that the human voice playback configuration and the background sound playback configuration can function independently and accurately, thereby maximizing the optimization potential of spatial hearing.

[0043] It is easy to understand that, besides using deep learning to perform audio source separation using audio source separation models, other signal processing techniques can also be employed to achieve the same goal, such as independent component analysis, nonnegative matrix factorization, spectral subtraction, Wiener filtering, or traditional methods combining speech activity detection and time-frequency masking. Furthermore, spatially guided separation can be performed using HRTF prior knowledge, or spatial information acquired using a multi-microphone array can be used to assist the separation process. This embodiment does not impose specific limitations on these methods; the most suitable audio source separation scheme can be flexibly selected based on factors such as device computing power, latency requirements, application scenarios, and audio content types.

[0044] This embodiment achieves decoupling of human voice and background sound at the audio signal level by performing source separation processing on the original audio signal. This allows for differentiated spatial rendering strategies to be applied to different types of sound content. This preprocessing breaks through the limitation of "treating all signals uniformly" in traditional binaural rendering, laying the data foundation for achieving refined and layered sound field control. Especially for near-ear open-back audio devices, due to their severe acoustic leakage and complex coupling of sound fields inside and outside the ear canal, "in-the-head effect" or sound image drift can easily occur if key sound sources (such as human voices) are not specially processed. Therefore, this step, as a prerequisite for the entire technical solution, effectively improves the controllability and targeting of subsequent rendering stages.

[0045] Step S200: Based on the preset human voice playback configuration, the original human voice audio signal is rendered in binaurally to obtain the target human voice audio signal, and based on the preset background sound playback configuration, the background sound audio signal is rendered in binaurally to obtain the target background sound audio signal, and the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration. As those skilled in the art will know, binaural rendering refers to a technique that simulates the acoustic differences (mainly time difference, intensity difference, and frequency-related spectral changes) produced when sound travels from a specific direction to both ears. It uses digital filters (such as HRTF filter kernels) to convolve audio signals, creating an auditory perception effect from a specific location in three-dimensional space when played back through headphones. Essentially, it reconstructs the acoustic response of sound waves interacting with the human head, auricle, and torso in a free field, giving the listener an immersive spatial auditory experience.

[0046] In binaural rendering, to construct a sense of space, a virtual speaker playback system is typically pre-defined as a reference frame for rendering. This introduces the concept of "virtual speakers": virtual speakers refer to a set of spatial sound source points artificially set in the binaural rendering algorithm to simulate the layout and position of physical speakers in three-dimensional space. These virtual speakers are not real sound-producing bodies, but rather serve as geometric anchor points for the spatialization of audio signals, guiding the selection and application of HRTF (Heated Audio Frequency Rendering), thereby determining the spatial characteristics and orientation of the rendered sound. In other words, the essence of binaural rendering is to "map" the input audio signal onto a set of virtual speakers at specified locations, and then synthesize its auditory response at both ears using HRTF.

[0047] The set of parameters directly related to the spatial layout of the virtual speakers is called the playback configuration. This configuration defines the spatial geometry of the virtual speakers relative to the listener and is a key factor determining the spatial properties of the binaural rendering result. In traditional stereo or surround sound binaural rendering, the playback configuration is usually modeled based on a standard speaker system. For example, in standard stereo playback, the two virtual speakers are symmetrically arranged in front of the listener, forming an angle of ±30° with the front (0° axis), constituting a 60° playback angle, about 2-3 meters away from the listener, with an elevation angle close to 0° (ear level). This is a typical layout recommended by international standards.

[0048] Specifically, the replay configuration includes at least the following three core space parameters: Reproduction Angle: This refers to the angle between the two virtual speakers relative to the listener's front. For example, when the left and right virtual speakers are located at 30° to the left and 30° to the right of the listener, respectively, the reproduction angle is 60°. This angle directly affects the lateral width of the sound image—the larger the angle, the wider the sound field, but the less focused the sound image becomes, resulting in blurred positioning; the smaller the angle, the more concentrated the sound image, and the clearer the positioning.

[0049] Reproduction Distance: This refers to the straight-line distance from the virtual speaker to the center of the listener's head. This parameter affects the sense of depth and presence in the sound. A closer distance enhances the "externalization" and intimacy of the sound, while a greater distance tends to create a more spacious atmosphere.

[0050] Reproduction Elevation Angle: This refers to the vertical angle between the virtual speaker and the listener's ear level. This parameter controls the vertical directionality of sound, such as simulating the location of sound sources above or below the head, and is an important dimension for achieving a three-dimensional sound field.

[0051] It should be noted that in conventional binaural rendering practices in this field, especially in the processing of stereo content, technicians typically only adjust the playback angle as the primary means of spatial control, while fixing the playback distance and playback elevation angle to standard values ​​(e.g., 3m distance, 0° elevation). This approach stems from the standardized design of traditional stereo broadcasting and playback systems and is also limited by the acquisition conditions and universality considerations of the HRTF dataset. Therefore, most existing solutions, when achieving a sense of space, often only adjust the sound image width by changing the playback angle, ignoring the synergistic effect of distance and elevation angle on spatial perception, resulting in limited spatial expressiveness and difficulty in achieving refined sound field layering control.

[0052] However, this embodiment breaks through the conventional technical paradigm mentioned above, proposing a differentiated and multi-dimensionally adjustable playback configuration strategy for human voices and background sounds. Specifically: In this embodiment, the voice playback configuration refers to a set of virtual speaker parameter combinations (i.e., playback configuration) specifically designed for the original human voice audio signal, aiming to achieve a high degree of concentration and stability in voice imaging. This configuration is preferably set as follows: a smaller playback angle (e.g., within ±15°, or even close to a 0° center placement), a closer playback distance (e.g., 0.8–1.2m), and a lower absolute value of the playback elevation angle (e.g., within ±5°). Through the synergistic effect of these three factors, the virtual voice is precisely anchored in the central area directly in front of the listener, forming a "concise, externalized, and stable" sound image perception, significantly improving the intelligibility and presence of the speech.

[0053] Accordingly, the target human voice audio signal refers to the human voice audio signal with precise spatial positioning attributes generated after binaural rendering processing under human voice playback configuration. This signal has been subjected to HRTF filtering characteristics adapted to center-channel near-field human voice in both the frequency and time domains, which can reproduce a realistic and non-dispersive voice image on near-ear open-back devices.

[0054] It should also be noted that, in this embodiment, the background sound reproduction configuration refers to another set of virtual speaker parameters set for the background sound audio signal. Its purpose is to expand the width and depth of the auditory space and enhance the immersive atmosphere. This configuration typically uses a larger reproduction angle (such as ±60° to ±120°, or even extended to a surround layout), a larger reproduction distance (such as 2.5 to 5m or more), and a higher reproduction elevation angle (such as ±20° to ±45°). It further stretches the sound field boundary through multi-channel diffusion algorithms or virtual surround technology, so that the human ear can feel the environmental immersion from all directions.

[0055] Correspondingly, the target background sound audio signal refers to the background sound output signal with broad spatial diffusion characteristics generated after background sound playback configuration processing. This signal is endowed with stronger directional diversity and spatial diffusion during binaural rendering, which helps to mask the problem of local sound field collapse caused by sound leakage in near-ear open devices and enhance the realism of the overall auditory environment.

[0056] It should be noted that, in this embodiment, spatial perception refers to the listener's subjective perception of the spatial location, direction, distance, and diffusion of an audio signal corresponding to a certain sound source. Its level depends on the design tendencies of various parameters in the playback configuration. High spatial perception means clear sound image localization, strong sense of direction, and clear sense of distance; low spatial perception is characterized by a wide sound field, blurred boundaries, and lack of focus.

[0057] In this embodiment, it is explicitly stipulated that the spatial perception corresponding to the human voice playback configuration is higher than that corresponding to the background sound playback configuration. The underlying mechanism is that the human voice, as the core carrier of information transmission, must have a high degree of intelligibility and positioning consistency to avoid increased cognitive load or distraction due to excessive spatialization. The main function of the background sound is to provide contextual support and emotional enhancement, and it should aim to be "broad but not chaotic", with its spatial sharpness appropriately reduced to prevent interference with the main sound source.

[0058] The essence of this differentiated configuration strategy is to achieve a fine allocation of spatial weights for the two types of audio signals by independently adjusting the virtual speaker layout parameters (i.e., playback configuration) corresponding to human voice and background sound, thereby guiding the rational distribution of auditory attention resources—that is, making human voice "prominent and stable without being ethereal," and making background sound "receded and broad without being disruptive." This design fully aligns with the "Cocktail Party Effect" of the human auditory system, which is the ability to prioritize the central speech signal in a complex acoustic environment.

[0059] This embodiment overcomes the technical bottleneck of traditional single HRTF templates, which cannot simultaneously achieve "clear positioning" and "wide sound field," by introducing a categorized and differentiated binaural rendering mechanism. Especially in usage scenarios centered on human voice, such as film and television dialogue, video conferencing, and audiobooks, this solution can effectively suppress the "in-head effect," improve speech focus and presence, while maintaining an immersive experience of background music and environmental sound effects, achieving dual optimization of acoustic performance and user experience.

[0060] For example, in a first feasible implementation, the voice playback configuration includes a voice playback angle, the background sound playback configuration includes a background sound playback angle, and the voice playback angle is smaller than the background sound playback angle.

[0061] It should be noted that, in this embodiment, the angle between the two virtual speakers used for binaural rendering of the original human voice audio signal and the angle between them and the listener's front (0° axis) is the angle between them. This angle determines the spatial expansion of the human voice in the horizontal plane and is a key parameter affecting the focusing of the sound image.

[0062] Correspondingly, the background sound playback angle refers to the angle between the left and right virtual speakers set when performing binaural rendering of the original background sound audio signal, in order to control the spatial coverage and sound field width of the background sound.

[0063] In this embodiment, the human voice reproduction angle is specified to be smaller than the background sound reproduction angle. This means the virtual speaker layout for human voices is more concentrated, while the background sound uses a more open layout. For example, the human voice reproduction angle can be set to 30° (±15°), while the background sound reproduction angle can be set to 90° (±45°) or larger. This parameter difference results in a highly concentrated central sound image for the human voice audio signal after binaural rendering, with clear positioning and well-defined boundaries; while the background sound audio signal forms a wide and diffuse sound field, effectively expanding the lateral dimension of the auditory space.

[0064] A smaller vocal reproduction angle helps enhance the stability and externalization of the vocals, avoiding the "in-the-head effect" caused by excessive diffusion of the speech signal. This is especially suitable for scenarios where the speech clarity of near-ear open-back audio devices is easily affected by sound leakage. At the same time, a larger background sound reproduction angle can simulate the spatial characteristics of a multi-channel surround sound system, creating an immersive environment and thus achieving a natural separation of primary and secondary sound sources at the psychoacoustic level.

[0065] This implementation method achieves preliminary layered control of the spatial weights of human voice and background sound by differentiating the core spatial parameter of the playback angle. Its technical effects are twofold: firstly, it significantly improves the intelligibility and attention-grabbing ability of human voice, making it "centered, concise, and clearly distinguishable"; secondly, it maintains the broad spatial expressiveness of background sound effects, ensuring that the overall auditory experience does not feel cramped or oppressive due to voice focusing. This solution is simple in structure and low in implementation cost, making it a fundamental means of constructing a high-contrast spatial sound field.

[0066] In a second feasible implementation, the voice playback configuration includes a voice playback distance, the background sound playback configuration includes a background sound playback distance, and the voice playback distance is less than the background sound playback distance.

[0067] It should be noted that, in this embodiment, the voice reproduction distance refers to the straight-line distance from the virtual speaker to the center of the listener's head when performing binaural rendering on the original voice audio signal. This distance directly affects the depth perception and sense of intimacy of the voice.

[0068] Correspondingly, background sound reproduction distance refers to the distance between the virtual speakers when rendering the original background sound audio signal in both ears, which is used to create the spatial depth and environmental immersion characteristics of the background sound.

[0069] In this embodiment, the playback distance for human voices is specified to be smaller than that for background sounds. That is, human voices are rendered as sound sources from a closer location, while background sounds are represented as a collection of sounds from a more distant space. Specifically, the playback distance for human voices can be set within the range of 0.8 to 1.2 meters, close to the typical face-to-face conversation distance, making the speech sound more realistic and interactive. The playback distance for background sounds can be set to 3 meters or more, even reaching 5 to 10 meters, to simulate the far-field reverberation and spatial attenuation characteristics of cinema-grade or open outdoor scenes.

[0070] By designing the difference in playback distance, this implementation makes the target human voice audio signal appear as "near-field focused" in terms of hearing, enhancing the externalization and physical presence of the voice, and effectively suppressing the "internalization" listening experience common in headphone playback; at the same time, the target background sound audio signal exhibits the acoustic characteristics of "far-field diffusion", with a more uniform energy distribution and strong directional ambiguity, which is conducive to constructing a stable background sound layer without focus interference.

[0071] This implementation method, by introducing differentiated control of playback distance, overcomes the limitations of traditional binaural rendering that relies solely on horizontal angle adjustment, achieving fine layering of sound in the foreground and background depth dimensions. Its technical effects include: not only strengthening the spatial priority of human voice as the foreground subject, but also utilizing distant background sound to create realistic three-dimensional spatial depth, enhancing the overall stereoscopic and layered feel of the audio scene. Especially in applications that emphasize accurate voice reproduction, such as video conferencing and audiobooks, this solution can significantly improve subjective listening quality.

[0072] In a third feasible implementation, the voice playback configuration includes a voice playback elevation angle, the background sound playback configuration includes a background sound playback elevation angle, and the absolute value of the voice playback elevation angle is smaller than the absolute value of the background sound playback elevation angle.

[0073] It should be noted that, in this embodiment, the vocal reproduction elevation angle refers to the vertical angle between the virtual speaker and the listener's ear level when the original vocal audio signal is rendered in both ears. This parameter is used to control the spatial positioning tendency of the vocal sound in the vertical direction.

[0074] Correspondingly, the background sound playback elevation angle refers to the vertical tilt angle of the associated virtual speaker when performing binaural rendering on the original background sound audio signal, which is used to expand the spatial coverage of the sound field in the vertical dimension.

[0075] In this embodiment, the absolute value of the human voice playback elevation angle is specified to be smaller than the absolute value of the background sound playback elevation angle. That is, the virtual speaker for the human voice is basically located near the listener's ear level (close to 0°), while the virtual speakers for the background sound are distributed at higher or lower vertical positions. For example, the human voice playback elevation angle can be controlled within ±5° to ensure that the speech is always anchored in the horizontal area directly in front; while the background sound playback elevation angle can be set to ±30° to ±45° to simulate non-horizontal sound components such as ceiling reflections, ground echoes, or aerial flight effects.

[0076] By setting different playback elevation angles, this implementation method ensures that the target human voice audio signal maintains a high degree of consistency in the vertical direction, avoiding positioning confusion caused by vertical offset and further improving the stability of voice imaging; while the target background sound audio signal obtains stronger vertical spatial diffusion, which can activate the human auditory system's ability to perceive sound sources at the top and bottom, thereby breaking the spatial limitation of traditional stereo which is limited to the horizontal plane.

[0077] This implementation achieves spatial decoupling and layered rendering of sound in the vertical dimension by independently adjusting the playback elevation angle parameter. Its technical effects include: not only enhancing the positioning accuracy of human voices in three-dimensional space and preventing them from "drifting" with changes in content, but also significantly improving the immersive dimension of background sound, allowing users to experience a sense of environmental envelopment from all directions (including above and below). This is of great significance for applications such as virtual reality, panoramic audio, and film sound effects that pursue omnidirectional spatial reproduction, and is an important supplement and upgrade to horizontal planar rendering.

[0078] It should be understood that the above three implementation methods are merely examples of specific implementation paths in this embodiment and do not constitute a limitation on the scope of protection of this application. In practical applications, the above three differentiated control mechanisms can be used individually or flexibly combined according to specific application scenarios. For example, a small-angle + close-range + low-elevation human voice configuration and a large-angle + long-range + high-elevation background sound configuration can be used simultaneously to achieve multi-dimensional collaborative optimization of spatial sound field reconstruction.

[0079] Crucially, in this embodiment, regardless of the parameter combination used, a core design principle must be met: the spatial perception corresponding to the human voice playback configuration must be higher than that corresponding to the background sound playback configuration. Only in this way can we ensure that the human voice, as the core of information, always maintains its clear, stable, and focusable auditory advantage.

[0080] Preferably, in this embodiment, the spatial perception corresponding to the human voice playback configuration should be higher than that corresponding to the standard stereo playback configuration (in some standards, the standard stereo playback configuration is: playback elevation angle of 60°, playback distance of 3m, and absolute value of playback elevation angle of 0°), while the spatial perception corresponding to the background sound playback configuration should be lower than that corresponding to the standard stereo playback configuration.

[0081] In other words, the binaural rendering of human voices tends to be "more focused and more precise," while the binaural rendering of background sounds tends to be "broader and more diffuse." This inversely differentiated binaural rendering strategy is the fundamental innovation of this application—by applying spatial modulation with opposite trends to the two types of audio signals, the auditory attention hierarchy is reconstructed, ultimately achieving an ideal acoustic experience of "distinct primary and secondary elements and clear layers" on near-ear open-back devices.

[0082] Step S300: Mix the target human voice audio signal and the target background sound audio signal to obtain the target audio signal, and play the target audio signal through an open-back audio device.

[0083] As those skilled in the art will know, audio mixing refers to the technical process of linearly superimposing two or more processed audio signals according to a certain gain ratio, phase relationship, and timing alignment to generate a composite audio signal. This process must ensure that there is no obvious phase cancellation, clipping distortion, or dynamic conflict between the sub-signals, so as to guarantee that the final output sound quality is pure and layered.

[0084] It should be noted that, in this embodiment, the target audio signal refers to the final output audio stream formed by mixing the target human voice audio signal and the target background sound audio signal. This signal has integrated the human voice and background sound components after differential spatial processing, and has an ideal sound image structure and spatial balance characteristics, which can be directly sent to an open-back audio device for playback.

[0085] This embodiment re-integrates two independently spatially controlled signals through mixing processing, achieving an organic unity of "precise vocals + immersive background". Since the two signals have already completed their respective spatial modeling during the rendering stage, the mixing process does not require additional adjustment of spatial parameters; only the relative levels need to be controlled to achieve optimal auditory fusion. In addition, considering that near-ear open-back devices have certain acoustic crosstalk and environmental noise interference, the signal-to-noise ratio of the vocals can be appropriately increased during the mixing stage to further enhance speech intelligibility.

[0086] Ultimately, the target audio signal is played through an open-back audio device, providing the wearer with a superior listening experience that combines high-definition voice localization with a wide stereo sound field. Compared to traditional uniform rendering methods, this solution significantly enhances the expressiveness and practicality of spatial audio while maintaining the original transparency of the device.

[0087] In summary, this embodiment effectively solves the technical challenge of balancing a wide sound field and clear vocal localization on near-ear open-back audio devices by employing a differentiated binaural rendering technology based on sound source separation. Specifically, this embodiment first performs sound source separation processing on the original audio signal, extracting the vocal and background sound components as independent audio streams. Subsequently, differentiated binaural rendering strategies are adopted for different sound source types: the original vocal audio signal is rendered using a vocal reproduction configuration with high spatial awareness (such as a smaller reproduction angle, a closer reproduction distance, and a lower absolute value of the elevation angle) to ensure focused vocal imaging and accurate localization, avoiding sound image blurring caused by excessive spatialization; simultaneously, the original background sound audio signal is rendered using a background sound reproduction configuration with relatively low spatial awareness, resulting in a wider and / or deeper diffusion effect, thereby creating an immersive sound field atmosphere; finally, the rendered target vocal audio signal and target background sound audio signal are mixed and output, and then played through a near-ear open-back audio device.

[0088] Through the above-mentioned technical means, this embodiment achieves the auditory effect of "condensed and not scattered" human voice and "wide and not chaotic" background sound. While ensuring the clarity of speech and the stability of sound image, it significantly improves the spatial sense and immersion of the overall audio, and is especially suitable for application scenarios with human voice as the core content, such as movie playback and voice calls.

[0089] Based on the first embodiment described above, this application proposes a binaural rendering method according to a second embodiment.

[0090] In the second embodiment of this application, the same or similar content as in the above embodiments can be referred to the above description, and will not be repeated hereafter.

[0091] like Figure 3 As shown, in this embodiment, the above-mentioned step S200 performs binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain the target human voice audio signal, and performs binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain the target background sound audio signal, which may include steps S210~S220: Step S210: Obtain the human voice head-related transfer function corresponding to the preset human voice playback configuration, and the background sound head-related transfer function corresponding to the preset background sound playback configuration; As those skilled in the art will know, the Head-Related Transfer Function (HRTF) is a complex frequency response function that describes the acoustic filtering characteristics experienced by a sound wave as it propagates from a specific spatial direction to the eardrum. This function incorporates the diffraction, reflection, and resonance effects of sound waves caused by human structures such as the head, auricle, external auditory canal, and shoulder, and is a key physiological basis for determining spatial sound perception. The HRTF exhibits high directional selectivity; sound sources at different orientations (horizontal angle, elevation angle) and distances correspond to different HRTF characteristics, especially showing significant differences in spectral peaks and valleys in the high-frequency range, providing important clues for the human brain to determine the location of sound sources.

[0092] In binaural rendering, the head-related transfer function (HRTF) serves as the core processing kernel, simulating the binaural time difference, binaural intensity difference, and spectral shape changes that occur when sound emitted from a virtual speaker reaches the listener's left and right ears. By convolving the input audio signal with the corresponding HRTF, auditory stimuli consistent with real-world hearing can be reconstructed at the headphone output, thereby achieving precise localization of three-dimensional sound images.

[0093] In this embodiment, the head-related transfer function (HRTF) refers to a set of HRTF data that matches the virtual speaker spatial parameters defined in the voice playback configuration, specifically including the frequency response curves of the left and right ear channels. The selection of this function is strictly based on interpolation or retrieval according to the playback angle, playback distance, and playback elevation angle in the voice playback configuration. For example, when the voice playback configuration is set to ±10° horizontal angle, 1.0 meter distance, and 0° elevation angle, the system will extract or generate HRTF pairs corresponding to these spatial coordinates from the pre-stored HRTF database, ensuring that the rendered voice is accurately presented in the near-field center position directly in front of the listener.

[0094] Correspondingly, the background sound head-related transfer function refers to the set of HRTFs corresponding to the background sound playback configuration. Since background sound typically employs a wider spatial layout (such as wide-angle stereo or virtual surround sound), its HRTF may involve the combined use of multiple directional points, or even the use of a diffused average HRTF or a spatialized kernel with enhanced room impulse response, to create a diffuse, enveloping auditory effect. This function is also dynamically obtained based on the specific parameters of the background sound playback configuration (such as ±60° horizontal angle, 4.0 meter distance, ±30° elevation angle) to ensure the accurate reproduction of the spatial expansion characteristics of the background sound.

[0095] This embodiment achieves refined modeling of the spatial perception characteristics of different types of audio content by acquiring dedicated HRTFs adapted to the playback configurations of human voice and background sound respectively. This step provides an accurate physiological acoustic basis for subsequent differentiated binaural rendering and is a prerequisite for ensuring a reasonable distribution of the final sound image structure.

[0096] Step S220: Perform binaural rendering on the original human voice audio signal using the human voice head correlation transfer function to obtain the target human voice audio signal; In this embodiment, the binaural rendering process is completed by performing time-domain convolution or frequency-domain multiplication operations on the original human voice audio signal with the head correlation transfer functions of the left and right ears, respectively. Specifically, the system decomposes the original human voice audio signal into a time-frequency representation, multiplies it by the frequency response characteristics of the human voice HRTF, and then converts it back to the time domain to generate a two-channel output signal containing spatial information—that is, the target human voice audio signal. When played, this signal allows the listener to perceive that the speech comes from a stable point sound source at close range in front, with high focus and low diffusion, effectively avoiding the "in-head effect" or sound image drift.

[0097] Step S230: The original background sound audio signal is rendered by binaural rendering using the background sound head correlation transfer function to obtain the target background sound audio signal.

[0098] In this embodiment, the spatial properties of the original background sound audio signal are reshaped by performing a similar convolution or filtering operation on the background sound head-related transfer function. Since the background sound HRTF usually corresponds to a relatively far distance, a large angle, or a non-horizontal plane position, the rendered target background sound audio signal exhibits a broad, blurred, and deep sound field characteristic.

[0099] In addition, multi-channel diffusion algorithms or random phase perturbation techniques can be combined to further weaken its directional sharpness, making it more in line with the auditory characteristics of ambient sound that is "without a clear direction but fills the space," thus forming an effective complement to human voices rather than interference.

[0100] This embodiment achieves layered control of spatial auditory characteristics by applying dedicated HRTFs (Human Voice Rendering) to both human voice and background sound, each matched to its playback configuration. The technical effects are: it not only improves the localization accuracy and speech intelligibility of human voice, but also enhances the immersiveness and spatial inclusiveness of background sound. The synergistic effect of these two aspects significantly optimizes the overall subjective listening quality on near-ear open-back audio devices. Compared to traditional rendering methods using a uniform HRTF template, this solution is more flexible and adaptable, dynamically adjusting spatial representation strategies according to different content types, fully meeting the needs of various application scenarios such as movie playback, video calls, and virtual reality.

[0101] Furthermore, in one feasible implementation, the original background sound audio signal includes a first original background sound audio signal corresponding to a first background sound source and a second original background sound audio signal corresponding to a second background sound source, wherein the first background sound source is different from the second background sound source. The background sound playback configuration includes a first background sound playback configuration corresponding to a first background sound source and a second background sound playback configuration corresponding to a second background sound source. The background sound head-related transfer function includes a first background sound head-related transfer function corresponding to the first background sound playback configuration and a second background sound head-related transfer function corresponding to the second background sound playback configuration. The target background sound audio signal includes a first target background sound audio signal corresponding to the first background sound source and a second target background sound audio signal corresponding to the second background sound source. The above-mentioned step S230 performs binaural rendering on the original background sound audio signal using the background sound head correlation transfer function to obtain the target background sound audio signal, and may include steps S231~S232: Step S231: Perform binaural rendering on the first original background sound audio signal using the first background sound head correlation transfer function to obtain the first target background sound audio signal; Step S232: The second original background sound audio signal is binaurally rendered using the second background sound head correlation transfer function to obtain the second target background sound audio signal.

[0102] It should be noted that, in this embodiment, the first background sound source refers to a specific environmental or effect sound entity in the original background sound audio signal, such as background music, ambient reverberation, and distant traffic noise, which are sound components with continuous and diffuse characteristics. The second background sound source refers to another type of non-speech sound source in the original background sound audio signal that has clear event attributes or local spatial positioning requirements, such as footsteps, door opening and closing sounds, birdsong, and raindrop sounds, which are transient or highly directional sound effects. The two play different functional roles in the auditory scene: the former is mainly used to construct the overall acoustic atmosphere and sense of space, while the latter is used to enhance the realism of the situation and spatial narrative.

[0103] In this embodiment, the first original background sound audio signal refers to the audio stream obtained after sound source separation processing, which contains only the first background sound source component. Its spectral energy distribution is relatively uniform and its time continuity is strong. It usually exists as the "acoustic background". The second original background sound audio signal refers to the separated audio stream that contains only the second background sound source component. It has obvious start and end boundaries, large dynamic changes, and may carry local spatial clues (such as movement trajectories).

[0104] Correspondingly, the first background sound reproduction configuration refers to the combination of virtual speaker spatial parameters specifically designed for the first background sound source. Its design goal is to create a wide, deep, and unfocused sound field background. This configuration typically employs a large reproduction angle (e.g., ±90° to ±120°), a large reproduction distance (e.g., 4 to 6 m), and a high reproduction elevation angle (e.g., ±30° to ±45°). Furthermore, it uses HRTF diffusion processing or virtual surround algorithms to further weaken the sharpness of its directional perception, giving it an immersive sense of "coming from all directions."

[0105] The second background sound playback configuration refers to the combination of spatial parameters independently set for the second background sound source. Its purpose is to preserve or enhance the local spatial positioning information of this type of background sound source. For example, its virtual speaker can be placed to the side or rear (e.g., at a horizontal angle of ±60° to ±150°), at a moderate distance (2 to 3 meters), and the elevation angle can be set according to the actual scene (e.g., -10° for footsteps on the ground, +30° for birds in the air), thereby achieving spatial restoration of specific environmental events and enhancing the realism and spatial layering of the auditory narrative.

[0106] The first target background sound audio signal refers to the output signal generated after binaural rendering of the first original background sound audio signal through the first background sound head correlation transfer function. It has highly diffuse spatial characteristics and its main function is to fill the "background layer" of the auditory space, providing stable and non-interfering atmospheric support for the main sound source.

[0107] The second target background sound audio signal refers to the output signal generated after binaural rendering of the second original background sound audio signal through the second background sound head correlation transfer function. It retains a strong sense of direction and dynamic spatial changes, and is used to present the location and trajectory of specific environmental events, enriching the detail expression of the auditory scene.

[0108] In this embodiment, the system acquires a first background sound head-correlation transfer function and a second background sound head-correlation transfer function that match the first background sound playback configuration and the second background sound playback configuration, respectively, and performs independent binaural rendering processing on their respective original background sound audio signals accordingly: that is, the first original background sound audio signal is convolutionally filtered using the first background sound head-correlation transfer function to generate a first target background sound audio signal; simultaneously, the second original background sound audio signal is independently filtered using the second background sound head-correlation transfer function to generate a second target background sound audio signal. The two rendered signals will be merged into a complete target background sound audio signal in the subsequent mixing stage.

[0109] This implementation further subdivides the background sound into different types of sound sources and configures differentiated playback strategies and HRTF cores for each, achieving multi-level spatial modeling of the background sound field. Its technical effect is that it ensures both broad coverage and immersion of the overall ambient sound (achieved by the first background sound) and preserves the spatial directivity and dynamic expressiveness of key environmental sound effects (achieved by the second background sound), thereby constructing a more realistic, three-dimensional, and narrative-driven three-dimensional auditory scene. This solution is particularly suitable for applications such as film, games, and virtual reality, which have high requirements for spatial audio detail, significantly improving the information carrying capacity and artistic expression of background sound.

[0110] In addition, this layered rendering mechanism can be flexibly adjusted according to the device's computing power and latency requirements—it can be merged when resources are limited, and can be extended to more background sound source categories (such as third and fourth background sound sources) in high-performance scenarios, with good scalability and engineering adaptability.

[0111] To facilitate understanding of the binaural rendering method in the above embodiments of this application, a specific embodiment is provided: like Figure 4 As shown, in this specific embodiment, music source separation is first performed based on a deep neural network (i.e., a sound source separation model). The human voice content of the dual-path stereo audio (i.e., the original audio signal) is separated from other content to obtain a dual-path signal containing only human voice content (i.e., the original human voice audio signal) and a dual-path signal containing background content (i.e., the original background sound audio signal). Then, the dual-path signal containing only human voice content and the dual-path signal containing background content are independently and differentially rendered in stereo binaurally to achieve the effect of expanding the sound field while ensuring clear human voice localization. Finally, all left channels are superimposed to obtain the final left channel signal, and all right channels are superimposed to obtain the final right channel signal, resulting in the final dual-path signal (i.e., the target audio signal), which is then played through a near-ear open-type audio device. That is, the original audio signal received by the near-ear open-back audio device is processed by source separation to obtain the original human voice audio signal and the original background sound audio signal; the original human voice audio signal is rendered in both ears based on the preset human voice playback configuration to obtain the target human voice audio signal, and the background sound audio signal is rendered in both ears based on the preset background sound playback configuration to obtain the target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than that corresponding to the background sound playback configuration; the target human voice audio signal and the target background sound audio signal are mixed to obtain the target audio signal, and the target audio signal is played through the near-ear open-back audio device.

[0112] Music source separation refers to separating different music sources in a mixed audio stream; for example, extracting different musical components such as vocals, bass, and drums from the mixed audio. Separating music sources facilitates independent analysis and processing of the characteristics and attributes of each sound source in the input audio. In the field of music source separation, existing methods can be divided into two categories: 1. Model-based Traditional Music Source Separation Methods: Traditional model-based music source separation methods mainly include principal component analysis (PCA) and nonnegative matrix factorization (NMF). PCA can extract the main audio components from the mixed signal, but these components usually do not directly correspond to specific instruments. NMF has the advantage of strong interpretability, but its separation performance depends on understanding the characteristics of the music data and requires multi-channel audio signals as input.

[0113] 2. Deep Learning-Based Music Source Separation Methods: Common network structures include frequency-domain based music source separation networks, time-domain based music source separation networks, and network models combining the time and frequency domains. These deep neural networks, with their powerful nonlinear modeling capabilities, can automatically learn and find features and fit the relationship between input and output, demonstrating excellent music source separation performance.

[0114] like Figure 5 As shown in this specific embodiment, the stereo spatial binaural rendering algorithm logic is as follows: To ensure the stability and clarity of the human voice image, a virtual speaker with a 60° angle (i.e., a playback angle of 60°) corresponding to the standard stereo speaker playback configuration is selected. The dual-channel signal containing only human voice content is then subjected to binaural rendering, specifically using the HRTF in the ±30° direction for virtual speaker synthesis. In other words, the head-related transfer function (HRTF) corresponding to the human voice playback configuration is obtained, and the original human voice audio signal is binaurally rendered using the HRTF to obtain the target human voice audio signal.

[0115] To expand the stereo playback sound field width, virtual speakers with an included angle 2θ greater than 60° (i.e., playback angle greater than 60°) are selected. Binaural rendering is then performed on the dual-path signal containing background content; specifically, HRTFs in the ±θ directions are selected for virtual speaker synthesis. In other words, the background sound head-related transfer function corresponding to the background sound playback configuration is obtained, and the original background sound audio signal is binaurally rendered using this function to obtain the target background sound audio signal.

[0116] Considering the role of reverberation in improving the "head-in-the-head effect", this specific embodiment can add appropriate artificial reverberation to the signal path, that is, add artificial reverberation to the target audio signal.

[0117] It should be noted that the above specific embodiments are only used to assist in understanding this application and do not constitute a limitation on the binaural rendering method in this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0118] In addition, please refer to Figure 6 , Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the binaural rendering method in this application embodiment.

[0119] This application also provides an audio device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the binaural rendering method in the above embodiments.

[0120] The following is for reference. Figure 6 It shows a structural schematic diagram of an audio device suitable for implementing the embodiments of this application. Figure 6 The audio device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0121] like Figure 6 As shown, the audio device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the audio device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the audio device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagram shows audio equipment with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0122] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0123] The audio device provided in this application, employing the binaural rendering method in the above embodiments, can achieve a binaural rendering effect on near-ear open-back audio devices that possesses both spatial width and clear vocal localization. Compared with the prior art, the beneficial effects of the audio device provided in this application are the same as those of the binaural rendering method provided in the above embodiments, and other technical features of this audio device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0124] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0125] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above claims.

[0126] In addition, this application also provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the steps of the binaural rendering method in the above embodiments.

[0127] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0128] The aforementioned computer-readable storage medium may be included in an audio device or may exist independently without being assembled into an audio device.

[0129] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an audio device, cause the audio device to: perform source separation processing on the original audio signal received by the near-ear open-back audio device to obtain an original human voice audio signal and an original background sound audio signal; perform binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and perform binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration; perform mixing processing on the target human voice audio signal and the target background sound audio signal to obtain a target audio signal, and play the target audio signal through the near-ear open-back audio device.

[0130] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0132] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0133] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for performing the steps of the above-described binaural rendering method, which can achieve a binaural rendering effect with both spatial width and clear vocal localization on near-ear open-back audio devices. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the binaural rendering method provided in the above embodiments, and will not be repeated here.

[0134] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the binaural rendering method described in the above embodiments.

[0135] The computer program product provided in this application can achieve a binaural rendering effect on near-ear open-back audio devices that has both spatial width and clear vocal localization. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the binaural rendering method provided in the above embodiments, and will not be repeated here.

[0136] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A binaural rendering method, characterized in that, The binaural rendering method includes the following steps: The original audio signal received by the near-ear open-back audio device is processed by source separation to obtain the original human voice audio signal and the original background sound audio signal; Based on a preset human voice playback configuration, the original human voice audio signal is rendered in binaurally to obtain a target human voice audio signal. Based on a preset background sound playback configuration, the background sound audio signal is rendered in binaurally to obtain a target background sound audio signal. The spatial perception corresponding to the human voice playback configuration is higher than that corresponding to the background sound playback configuration. The target human voice audio signal and the target background sound audio signal are mixed to obtain a target audio signal, which is then played through the near-ear open-back audio device.

2. The binaural rendering method as described in claim 1, characterized in that, The human voice playback configuration includes a human voice playback angle, the background sound playback configuration includes a background sound playback angle, and the human voice playback angle is smaller than the background sound playback angle.

3. The binaural rendering method as described in claim 1, characterized in that, The human voice playback configuration includes a human voice playback distance, the background sound playback configuration includes a background sound playback distance, and the human voice playback distance is less than the background sound playback distance.

4. The binaural rendering method as described in claim 1, characterized in that, The human voice playback configuration includes a human voice playback elevation angle, the background sound playback configuration includes a background sound playback elevation angle, and the absolute value of the human voice playback elevation angle is less than the absolute value of the background sound playback elevation angle.

5. The binaural rendering method as described in claim 1, characterized in that, The step of performing source separation processing on the raw audio signal received by the near-ear open-back audio device to obtain the raw human voice audio signal and the raw background sound audio signal includes: The raw audio signal received by the near-ear open-back audio device is input into a pre-trained sound source separation model to obtain the raw human voice audio signal and the raw background sound audio signal output by the sound source separation model.

6. The binaural rendering method according to any one of claims 1 to 5, characterized in that, The steps of performing binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and performing binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, include: Get the human voice head-related transfer function corresponding to the preset human voice playback configuration, and the background sound head-related transfer function corresponding to the preset background sound playback configuration; The original human voice audio signal is binaurally rendered using the human voice head correlation transfer function to obtain the target human voice audio signal. The original background sound audio signal is rendered by binaural rendering using the background sound head correlation transfer function to obtain the target background sound audio signal.

7. The binaural rendering method as described in claim 6, characterized in that, The original background sound audio signal includes a first original background sound audio signal corresponding to a first background sound source and a second original background sound audio signal corresponding to a second background sound source, wherein the first background sound source is different from the second background sound source. The background sound playback configuration includes a first background sound playback configuration corresponding to the first background sound source and a second background sound playback configuration corresponding to the second background sound source. The background sound head-related transfer function includes a first background sound head-related transfer function corresponding to the first background sound playback configuration and a second background sound head-related transfer function corresponding to the second background sound playback configuration. The target background sound audio signal includes a first target background sound audio signal corresponding to the first background sound source and a second target background sound audio signal corresponding to the second background sound source. The step of performing binaural rendering on the original background sound audio signal using the background sound head correlation transfer function to obtain the target background sound audio signal includes: The first original background sound audio signal is rendered by binaural rendering using the first background sound head correlation transfer function to obtain the first target background sound audio signal. The second original background sound audio signal is rendered by binaural rendering using the second background sound head correlation transfer function to obtain the second target background sound audio signal.

8. A binaural rendering device, characterized in that, The binaural rendering device includes: The separation module is used to perform source separation processing on the raw audio signal received by the near-ear open-back audio device to obtain the original human voice audio signal and the original background sound audio signal; The rendering module is used to perform binaural rendering on the original human voice audio signal based on a preset human voice playback configuration to obtain a target human voice audio signal, and to perform binaural rendering on the background sound audio signal based on a preset background sound playback configuration to obtain a target background sound audio signal, wherein the spatial perception corresponding to the human voice playback configuration is higher than the spatial perception corresponding to the background sound playback configuration. The mixing module is used to mix the target human voice audio signal and the target background sound audio signal to obtain a target audio signal, and to play the target audio signal through the near-ear open-back audio device.

9. An audio device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the binaural rendering method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a binaural rendering program, which, when executed by a processor, implements the steps of the binaural rendering method as described in any one of claims 1 to 7.