Space rendering method and device, audio equipment and computer readable storage medium
By combining sound source separation and differentiated spatial rendering with crosstalk cancellation, the problems of blurred voice positioning and diffuse sound images in spatial rendering are solved, achieving a high-quality stereo audio experience.
Patent Information
- Application Number
- CN202511605229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies result in blurred voice localization and diffused sound images in spatial rendering, affecting the clarity and immersiveness of stereo audio, especially in applications such as film and television dialogues and voice broadcasts.
The audio signal is separated into human voice and background sound through source separation processing. Differentiated playback configurations are used for spatial rendering, and crosstalk is eliminated by combining target transfer function to ensure the stability of human voice and the diffusion of background sound.
Without increasing the number of physical speakers, it achieves synergistic optimization of spatial rendering effects and human voice focusing, improving the immersiveness and auditory experience in film and television viewing and voice interaction scenarios.
Smart Images

Figure CN121509895A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a spatial rendering method, apparatus, audio device, and computer-readable storage medium. Background Technology
[0002] Stereo audio, as the most mainstream audio format today, is widely used in audio devices such as headphones and smart audio glasses. To enhance the immersive experience, those skilled in the art often extend the sound field of stereo audio signals to achieve spatial rendering effects, making the sound seem to come from a wider spatial area, simulating the auditory experience of a multi-speaker system.
[0003] However, in practical applications, it has been found that after spatial rendering through sound field expansion, although the width of the sound image is improved, the human voice often exhibits blurred positioning and a diffuse sound image, giving the unnatural feeling that "the sound is in the head" or "drifting in from the sides." This problem is particularly noticeable in content dominated by human voices, such as film and television dialogues and voice broadcasts, seriously affecting auditory clarity and subjective listening quality.
[0004] Therefore, how to achieve spatial rendering effects while maintaining the stability and focus of human voices has become a pressing technical challenge. Summary of the Invention
[0005] The main purpose of this application is to provide a spatial rendering method, apparatus, audio device, and computer-readable storage medium, aiming to solve the technical problem of how to achieve spatial rendering effects while taking into account the stability and focus of human voices.
[0006] To achieve the above objectives, this application provides a spatial rendering method, which includes the following steps: The first audio signal received by the target audio device is subjected to source separation processing to obtain a first human voice audio signal and a first background sound audio signal; Based on a preset first playback configuration, the first human voice audio signal is spatially rendered to obtain a second human voice audio signal, and based on a preset second playback configuration, the first background sound audio signal is spatially rendered to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration. The second human voice audio signal and the second background sound audio signal are mixed to obtain the second audio signal; The target transfer function between the target audio device and the user's ears is obtained, and crosstalk cancellation is performed on the second audio signal based on the target transfer function to obtain the target audio signal.
[0007] In one embodiment, the step of obtaining the target transfer function between the target audio device and the user's ears includes: Obtain the target artificial head transfer function and the target free field transfer function between the target audio device and the user's ears; Inverting the target free field transfer function yields the target free field inverse transfer function; Multiplying the target artificial head transfer function with the target free field inverse transfer function yields the target transfer function between the target audio device and the user's ears.
[0008] In one embodiment, the step of obtaining the target artificial head transfer function and the target free field transfer function between the target audio device and the user's ears includes: When a preset artificial head is placed at the user's hearing position and the target audio device outputs an audio signal, the transfer function of the target artificial head between the target audio device and the user's ears is measured using a preset microphone placed in the ear canal of the preset artificial head; and... When the preset artificial head is removed and the target audio device outputs an audio signal, the target free field transfer function between the target audio device and the user's ears is measured by using preset microphones placed at the left and right ear positions before the preset artificial head was removed.
[0009] In one embodiment, the step of performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal includes: Invert the target transfer function to obtain the target inverse transfer function; The second audio signal is multiplied by the target inverse transfer function to obtain the target audio signal.
[0010] In one embodiment, the steps of spatially rendering the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and spatially rendering the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, include: Obtain the first head-related transfer function corresponding to the preset first replay configuration, and the second head-related transfer function corresponding to the preset second replay configuration; The first human voice audio signal is spatially rendered using the first head-related transfer function to obtain the second human voice audio signal; The second background audio signal is obtained by spatial rendering of the first background audio signal using the second head correlation transfer function.
[0011] In one embodiment, the step of performing source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal includes: The first audio signal received by the target audio device is input into a pre-trained sound source separation model to obtain the first human voice audio signal and the first background sound audio signal output by the sound source separation model.
[0012] In one embodiment, after the step of obtaining the target transfer function between the target audio device and the user's ears, the method further includes: The ambient acoustic parameters of the user's current environment are obtained, including reflected sound intensity, reverberation time, and dominant reflection direction. Generate corresponding environmental compensation coefficients based on the aforementioned environmental acoustic parameters; Multiply the environmental compensation coefficient by the target transfer function to obtain the compensated target transfer function; Based on the compensated target transfer function, the step of performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal is performed.
[0013] Furthermore, to achieve the above objectives, this application also provides a spatial rendering apparatus, the spatial rendering apparatus comprising: The separation module is used to perform source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal. The rendering module is used to perform spatial rendering of the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and to perform spatial rendering of the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration. The mixing module is used to mix the second human voice audio signal and the second background sound audio signal to obtain the second audio signal; The cancellation module is used to obtain the target transfer function between the target audio device and the user's ears, and to perform crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal.
[0014] In addition, to achieve the above objectives, this application also provides an audio device, the audio device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the spatial rendering method described above.
[0015] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of the spatial rendering method described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described spatial rendering method.
[0017] This application provides a spatial rendering method, apparatus, audio device, and computer-readable storage medium, relating to the field of audio processing technology. The method includes: performing source separation processing on a first audio signal received by a target audio device to obtain a first human voice audio signal and a first background sound audio signal; performing spatial rendering on the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and performing spatial rendering on the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration; performing mixing processing on the second human voice audio signal and the second background sound audio signal to obtain a second audio signal; obtaining a target transfer function between the target audio device and the user's two ears, and performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain a target audio signal.
[0018] This application proposes a systematic spatial rendering method by combining sound source separation with differentiated spatial rendering and introducing crosstalk cancellation processing for target audio devices, in order to solve the technical problem that human voices are prone to positioning ambiguity and sound image dispersion during spatial rendering. Specifically, this embodiment first performs sound source separation on the input first audio signal, accurately extracting the first human voice audio signal and the first background sound audio signal, thereby achieving independent control of different types of sound sources. Subsequently, different spatial perception intensity reproduction configurations are applied to the human voice and background sound respectively—the human voice is rendered using a first reproduction configuration with higher spatial perception (such as a smaller reproduction angle and a closer reproduction distance) to ensure that its sound image is focused and clearly positioned, avoiding the feeling of "the sound is in the head" or floating left and right; while the background sound is rendered using a second reproduction configuration with lower spatial perception, so that it is appropriately diffused in space to form a wide sound field surround. Afterwards, the rendered second human voice audio signal and the second background sound audio signal are mixed to generate a second audio signal, and further combined with the target transfer function between the target audio device and the user's ears, the mixed second audio signal is subjected to crosstalk cancellation processing, effectively suppressing unexpected crosstalk between the left and right channels, improving channel isolation and spatial imaging accuracy, and finally outputting the target audio signal.
[0019] Through the above-mentioned technical means, this application embodiment achieves synergistic optimization of spatial rendering effect and human voice focusing without increasing the number of physical speakers. It not only obtains a wide and natural sound field, but also ensures the clarity and intelligibility of voice content and the stability of sound image, significantly improving the user's immersion and auditory experience quality in typical scenarios such as movie watching and voice interaction. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating the first embodiment of the spatial rendering method of this application; Figure 2 This is a schematic diagram illustrating the training of the sound source separation model in a specific embodiment of this application; Figure 3 This is a flowchart illustrating the second embodiment of the spatial rendering method of this application. Figure 4 This is a flowchart illustrating the third embodiment of the spatial rendering method of this application. Figure 5 This is a schematic diagram of the spatial rendering algorithm operation in a specific embodiment of this application; Figure 6 This is a schematic diagram of a spatial rendering scene in a specific embodiment of this application; Figure 7 This is a schematic diagram of differentiated spatial rendering in a specific embodiment of this application; Figure 8 This is a schematic diagram of the spatial rendering process in a specific embodiment of this application; Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the spatial rendering method in the embodiments of this application.
[0023] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0024] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0025] Currently, traditional spatial rendering methods mainly rely on HRTF to spatialize stereo or multi-channel audio signals, enabling the reconstruction of spatial auditory scenes on audio devices. This method simulates the differences in acoustic characteristics (such as time difference, intensity difference, and spectral characteristics) as sound reaches both ears from different directions, thereby achieving the perceptual restoration of sound source location and enhancing the overall sound field width and immersion.
[0026] However, in practical applications, traditional HRTF-based spatial rendering methods generally suffer from a decline in voice localization performance. Specifically, this manifests as blurred sound image localization, diffuse sound field imaging, and insufficient center focus in the voice portion. This leads to the listener's subjective auditory illusion that the sound "floats inside the head" or "comes unnaturally from the left and right sides," a typical "in-head effect" or sound image instability. The root of this problem lies in the fact that traditional spatial rendering typically applies the same HRTF kernel or spatialization parameters to all frequency bands or full audio content, failing to fully consider the dominant role of the voice signal in auditory perception and its crucial function in anchoring the sound image center. Especially in applications where voice is the core content, such as film and television dialogues, voice broadcasting, and video conferencing, the clarity, intelligibility, and spatial localization accuracy of the voice directly determine the quality of the user's auditory experience. When the voice is over-spatialized or treated the same as background noise, its central sound image is easily weakened or dispersed, thus destroying the focus and naturalness of the speech.
[0027] Furthermore, due to the lack of physical sound insulation, the sound waves emitted by the left and right speakers inevitably cause crosstalk interference at the listener's ears during propagation—that is, the left channel signal will simultaneously reach the right ear, and the right channel signal will simultaneously reach the left ear. This unintended acoustic coupling severely damages the channel isolation of stereo imaging, distorts the originally designed spatial location information, and further exacerbates the uncertainty of voice localization.
[0028] To achieve spatial rendering effects while maintaining the stability and focus of human voices, the solution of this application embodiment is a spatial rendering method, including: performing source separation processing on a first audio signal received by a target audio device to obtain a first human voice audio signal and a first background sound audio signal; performing spatial rendering on the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and performing spatial rendering on the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration; performing mixing processing on the second human voice audio signal and the second background sound audio signal to obtain a second audio signal; obtaining the target transfer function between the target audio device and the user's two ears, and performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal.
[0029] This application proposes a systematic spatial rendering method by combining sound source separation with differentiated spatial rendering and introducing crosstalk cancellation processing for target audio devices, in order to solve the technical problem that human voices are prone to positioning ambiguity and sound image dispersion during spatial rendering. Specifically, this embodiment first performs sound source separation on the input first audio signal, accurately extracting the first human voice audio signal and the first background sound audio signal, thereby achieving independent control of different types of sound sources. Subsequently, different spatial perception intensity reproduction configurations are applied to the human voice and background sound respectively—the human voice is rendered using a first reproduction configuration with higher spatial perception (such as a smaller reproduction angle and a closer reproduction distance) to ensure that its sound image is focused and clearly positioned, avoiding the feeling of "the sound is in the head" or floating left and right; while the background sound is rendered using a second reproduction configuration with lower spatial perception, so that it is appropriately diffused in space to form a wide sound field surround. Afterwards, the rendered second human voice audio signal and the second background sound audio signal are mixed to generate a second audio signal, and further combined with the target transfer function between the target audio device and the user's ears, the mixed second audio signal is subjected to crosstalk cancellation processing, effectively suppressing unexpected crosstalk between the left and right channels, improving channel isolation and spatial imaging accuracy, and finally outputting the target audio signal.
[0030] Through the above-mentioned technical means, this application embodiment achieves synergistic optimization of spatial rendering effect and human voice focusing without increasing the number of physical speakers. It not only obtains a wide and natural sound field, but also ensures the clarity and intelligibility of voice content and the stability of sound image, significantly improving the user's immersion and auditory experience quality in typical scenarios such as movie watching and voice interaction.
[0031] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0032] This application proposes a spatial rendering method according to a first embodiment.
[0033] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the spatial rendering method of this application.
[0034] In this embodiment, the spatial rendering method may include steps S100~S400: Step S100: Perform source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal; It should be noted that the target audio device in this embodiment refers to an audio playback device where there is a certain physical distance (usually greater than 30 cm) between its speakers and the user's ears when playing audio. Such devices include, but are not limited to, smart TVs, home audio systems, built-in speakers in laptops, car audio systems, and desktop speakers. Because there is a certain distance between the left and right channel speakers of such devices and the user's head, when playing stereo content, the left channel sound will simultaneously enter the user's left and right ears, and the same applies to the right channel, resulting in significant crosstalk interference. This crosstalk interference can disrupt the spatial positioning accuracy of virtual sound images, causing sound images to collapse into the head (i.e., the "in-head effect") or exhibiting problems such as positioning blurring and drifting. Therefore, after completing spatial rendering, crosstalk cancellation processing must be performed on the unique acoustic propagation path of such devices to restore the expected spatial listening experience.
[0035] It should also be noted that, in this embodiment, the audio source separation process refers to the process of decoupling and independently extracting different audio source components from a multi-source mixed audio signal in the time-frequency domain using signal processing algorithms or deep learning models. Specifically, this process identifies and separates the human voice component from the non-human voice component by analyzing the spectral characteristics, temporal envelope, spatial cues (such as phase difference and intensity difference), and semantic information (such as speech fundamental frequency and formant distribution) of the audio signal, thereby generating two independent audio streams from different audio sources: a human voice audio signal and a background sound audio signal.
[0036] In this embodiment, the first audio signal refers to the input audio data (which can be stereo audio data) received or decoded by the target audio device. It usually comes from film and television content, voice calls, virtual reality applications or other multimedia sources, and contains complete auditory scene information, including human voice dialogue as well as environmental sound effects, music and other background sound elements.
[0037] The first human voice audio signal refers to the audio signal stream obtained after sound source separation processing, which contains only the human voice portion (such as speaking or singing). This signal retains key speech information from the first audio signal, such as pronunciation content, intonation changes, and lip-sync, and is the core object for subsequent high-precision sound image localization.
[0038] The first background sound audio signal refers to the collection of all sound components other than human voice in the first audio signal obtained after sound source separation processing, including but not limited to environmental noise, background music, spatial reverberation, action sound effects (such as footsteps, door opening and closing sounds), and natural sounds (such as wind and rain sounds). This signal carries the spatial atmosphere and immersive information of the auditory scene in the first audio signal, and is used to construct a wide and layered sound field experience.
[0039] For example, in one feasible implementation, step S100 above, which performs source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal, may include step S110: Step S110: Input the first audio signal received by the target audio device into the pre-trained sound source separation model to obtain the first human voice audio signal and the first background sound audio signal output by the sound source separation model.
[0040] It should be noted that sound source separation models refer to a type of audio signal processing system built on machine learning or deep learning architectures. Their function is to automatically identify and separate different sound source components from mixed stereo audio signals. This model typically uses neural networks as its core, learning time-frequency features, harmonic structures, spatial cues, and semantic patterns from large amounts of labeled audio data to establish the ability to distinguish between human and non-human voice components, thereby achieving high-precision audio demixing.
[0041] In this embodiment, the audio source separation model was trained offline using a large-scale training dataset before deployment. The training dataset contains pairs of mixed audio samples and corresponding real, clean audio labels. The mixed audio samples simulate various superpositions of human voice and background noise in real-world scenarios (such as dialogue superimposed with background music, speech superimposed with environmental noise, etc.), covering different languages, genders, speech rates, music styles, noise types, and signal-to-noise ratio conditions. The clean audio labels are the corresponding human voice audio segments and background sound audio segments, respectively. During training, the model continuously optimizes the network parameters by minimizing the temporal waveform error or time-frequency spectrogram difference between the separation result and the label as the objective function until it achieves generalization ability. The training process of this audio source separation model is as follows: Figure 2 As shown. After training, the model is embedded and integrated into the target audio device or its paired terminal system for real-time audio source separation tasks.
[0042] In actual operation, when the first audio signal is input into the model, it is first converted into a time-frequency representation (such as a short-time Fourier transform spectrum). Then, the model analyzes and predicts the time-frequency masks for the human voice and background sound frame by frame. These masks are then restored to the time domain signal through inverse transform, and finally output independent first human voice audio signal and first background sound audio signal. This process has advantages such as fast processing speed, high separation accuracy, and strong adaptability to complex auditory scenes, and is especially suitable for dynamically changing multi-sound-source environments.
[0043] This implementation significantly improves the decoupling accuracy of human voice and background sound by introducing a deep learning-based sound source separation model. It effectively avoids the problems of "ghosting" or "crosstalk" that easily occur in high-frequency overlapping areas using traditional rule-based or simple filtering methods. Its technical effect is that it provides a high-quality, low-interference independent sound source signal foundation for subsequent differentiated spatial rendering, ensuring that the first and second playback configurations can function independently and accurately, thereby maximizing the optimization potential of spatial listening experience.
[0044] It is easy to understand that, besides using deep learning to perform audio source separation using audio source separation models, other signal processing techniques can also be employed to achieve the same goal, such as independent component analysis, nonnegative matrix factorization, spectral subtraction, Wiener filtering, or traditional methods combining speech activity detection and time-frequency masking. Furthermore, spatially guided separation can be performed using HRTF prior knowledge, or spatial information acquired using a multi-microphone array can be used to assist the separation process. This embodiment does not impose specific limitations on these methods; the most suitable audio source separation scheme can be flexibly selected based on factors such as device computing power, latency requirements, application scenarios, and audio content types.
[0045] This embodiment achieves decoupling of human voice and background sound at the audio signal level by performing source separation processing on the first audio signal. This allows for differentiated spatial rendering strategies to be applied to different types of sound content. This preprocessing breaks through the limitation of "treating all signals uniformly" in traditional spatial rendering, laying a data foundation for achieving refined and layered sound field control. Especially for target audio devices, due to their severe acoustic leakage and complex coupling of sound fields inside and outside the ear canal, "in-the-head effect" or sound image drift can easily occur if key sound sources (such as human voices) are not specially processed. Therefore, this step, as a prerequisite for the entire technical solution, effectively improves the controllability and targeting of subsequent rendering stages.
[0046] Step S200: Spatial rendering of the first human voice audio signal is performed based on the preset first playback configuration to obtain the second human voice audio signal, and spatial rendering of the first background sound audio signal is performed based on the preset second playback configuration to obtain the second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration. As those skilled in the art will understand, spatial rendering refers to the process of converting mono or multi-channel audio signals into stereo or three-dimensional audio signals adapted to binaural hearing perception using head-related transfer functions or their equivalent acoustic models. This process simulates the spectral, phase, and intensity changes experienced by sound waves as they travel from a specific spatial location to the ear, enabling listeners, when wearing headphones or using a speaker system, to perceive the specific location, distance, and height of the sound source in virtual space, thereby constructing an auditory scene with spatial dimensions.
[0047] During spatial rendering, the system first selects or interpolates the corresponding HRTF filter based on the desired spatial location of the target sound source (usually defined by the playback configuration). Then, the input audio signal is convolved with the HRTFs corresponding to the left and right ears respectively (or achieved through efficient approximation algorithms such as frequency domain filtering or parametric HRTF modeling) to generate an output signal with binaural differences and spectral cues. These cues are used by the human auditory system to resolve the spatial properties of the sound source, including azimuth, elevation, distance, and diffusion.
[0048] The set of parameters directly related to the spatial layout of virtual speakers is called the playback configuration. This configuration defines the spatial geometry of the virtual speakers relative to the listener and is a key factor determining the spatial attributes of the rendering result. In traditional stereo or surround sound spatial rendering, the playback configuration is usually modeled based on a standard speaker system. For example, in standard stereo playback, the two virtual speakers are symmetrically arranged in front of the listener, forming an angle of ±30° with the front (0° axis), constituting a 60° playback angle, about 2-3 meters away from the listener, with an elevation angle close to 0° (ear level). This is a typical layout recommended by international standards.
[0049] Specifically, the replay configuration includes at least the following three core space parameters: Reproduction Angle: This refers to the angle between the two virtual speakers relative to the listener's front. For example, when the left and right virtual speakers are located at 30° to the left and 30° to the right of the listener, respectively, the reproduction angle is 60°. This angle directly affects the lateral width of the sound image—the larger the angle, the wider the sound field, but the less focused the sound image becomes, resulting in blurred positioning; the smaller the angle, the more concentrated the sound image, and the clearer the positioning.
[0050] Reproduction Distance: This refers to the straight-line distance from the virtual speaker to the center of the listener's head. This parameter affects the sense of depth and presence in the sound. A closer distance enhances the "externalization" and intimacy of the sound, while a greater distance tends to create a more spacious atmosphere.
[0051] Reproduction Elevation Angle: This refers to the vertical angle between the virtual speaker and the listener's ear level. This parameter controls the vertical directionality of sound, such as simulating the location of sound sources above or below the head, and is an important dimension for achieving a three-dimensional sound field.
[0052] It should be noted that in conventional spatial rendering practices in this field, especially in the processing of stereo content, technicians typically only adjust the playback angle as the primary means of spatial control, while fixing the playback distance and playback elevation angle to standard values (e.g., 3m distance, 0° elevation). This approach stems from the standardized design of traditional stereo broadcasting and playback systems and is also limited by the acquisition conditions and universality considerations of the HRTF dataset. Therefore, most existing solutions, when achieving a sense of space, often only adjust the sound image width by changing the playback angle, ignoring the synergistic effect of distance and elevation angle on spatial perception, resulting in limited spatial expressiveness and difficulty in achieving refined sound field layering control.
[0053] However, this embodiment breaks through the conventional technical paradigm mentioned above, proposing a differentiated and multi-dimensionally adjustable playback configuration strategy for human voices and background sounds. Specifically: In this embodiment, the first playback configuration refers to a set of virtual speaker parameter combinations (i.e., playback configuration) specifically designed for the first human voice audio signal, aiming to achieve a high degree of concentration and stability in human voice imaging. This configuration is preferably set as follows: a smaller playback angle (e.g., within ±15°, or even approaching a 0° center placement), a closer playback distance (e.g., 0.8–1.2 m), and a lower absolute value of the playback elevation angle (e.g., within ±5°). Through the synergistic effect of these three factors, the virtual human voice is precisely anchored in the central area directly in front of the listener, forming a "concise, externalized, and stable" sound image perception, significantly improving the intelligibility and presence of the speech.
[0054] Correspondingly, the second human voice audio signal refers to the human voice audio signal with precise spatial positioning attributes generated after spatial rendering processing under the first playback configuration. This signal has been subjected to HRTF filtering characteristics adapted to center-channel near-field human voices in both the frequency and time domains, which can reproduce a realistic and non-dispersive voice image on the target device.
[0055] It should also be noted that, in this embodiment, the second playback configuration refers to another set of virtual speaker parameters set for the background audio signal. Its purpose is to expand the width and depth of the auditory space and enhance the immersive atmosphere. This configuration typically uses a larger playback angle (such as ±60° to ±120°, or even extended to a surround layout), a larger playback distance (such as 2.5 to 5m or more), and a higher playback elevation angle (such as ±20° to ±45°). It further stretches the sound field boundary through multi-channel diffusion algorithms or virtual surround technology, so that the human ear feels a sense of environmental immersion from all directions.
[0056] Correspondingly, the second background sound audio signal refers to the background sound output signal with broad spatial diffusion characteristics generated after processing by the second playback configuration. This signal is endowed with stronger directional diversity and spatial diffusion during spatial rendering, which helps to mask the local sound field collapse problem caused by sound leakage of the target device and enhance the realism of the overall auditory environment.
[0057] It should be noted that, in this embodiment, spatial perception refers to the listener's subjective perception of the spatial location, direction, distance, and diffusion of an audio signal corresponding to a certain sound source. Its level depends on the design tendencies of various parameters in the playback configuration. High spatial perception means clear sound image localization, strong sense of direction, and clear sense of distance; low spatial perception is characterized by a wide sound field, blurred boundaries, and lack of focus.
[0058] In this embodiment, it is explicitly stipulated that the spatial perception corresponding to the first playback configuration is higher than that corresponding to the second playback configuration. The underlying mechanism is that: as the core carrier of information transmission, human voice must have high intelligibility and positioning consistency to avoid increased cognitive load or distraction due to excessive spatialization; while the main function of background sound is to provide contextual support and emotional enhancement, and should aim for "broad but not chaotic", and its spatial sharpness should be appropriately reduced to prevent interference with the main sound source.
[0059] The essence of this differentiated configuration strategy is to achieve a fine allocation of spatial weights for the two types of audio signals by independently adjusting the virtual speaker layout parameters (i.e., playback configuration) corresponding to human voice and background sound, thereby guiding the rational distribution of auditory attention resources—that is, making human voice "prominent and stable without being ethereal," and making background sound "receded and broad without being disruptive." This design fully aligns with the "Cocktail Party Effect" of the human auditory system, which is the ability to prioritize the central speech signal in a complex acoustic environment.
[0060] This embodiment overcomes the technical bottleneck of traditional single HRTF templates, which cannot simultaneously achieve "clear positioning" and "wide sound field," by introducing a categorized and differentiated spatial rendering mechanism. Especially in usage scenarios centered on human voice, such as film and television dialogue, video conferencing, and audiobooks, this solution can effectively suppress the "in-head effect," improve speech focus and presence, while maintaining an immersive experience of background music and environmental sound effects, achieving dual optimization of acoustic performance and user experience.
[0061] For example, in a first feasible implementation, the first playback configuration includes a human voice playback angle, the second playback configuration includes a background sound playback angle, and the human voice playback angle is smaller than the background sound playback angle.
[0062] It should be noted that, in this embodiment, the angle refers to the angle between the two virtual speakers (left and right) relative to the listener's front (0° axis) when spatially rendering the first human voice audio signal. This angle determines the degree of spatial expansion of the human voice in the horizontal plane and is a key parameter affecting the focusing of the sound image.
[0063] Correspondingly, the background sound playback angle refers to the angle between the two virtual speakers set when spatially rendering the first background sound audio signal, in order to control the spatial coverage and sound field width of the background sound.
[0064] In this embodiment, the human voice reproduction angle is specified to be smaller than the background sound reproduction angle. This means the virtual speaker layout for human voices is more concentrated, while the background sound uses a more open layout. For example, the human voice reproduction angle can be set to 30° (±15°), while the background sound reproduction angle can be set to 90° (±45°) or larger. This parameter difference results in a highly concentrated central sound image for the spatially rendered human voice audio signal, with clear positioning and well-defined boundaries; while the background sound audio signal forms a wide and diffuse sound field, effectively expanding the horizontal dimension of the auditory space.
[0065] A smaller vocal reproduction angle helps enhance the stability and externalization of the vocals, avoiding the "in-the-head effect" caused by excessive diffusion of the speech signal. This is especially suitable for scenarios where the speech clarity on the target audio device is easily affected by sound leakage. At the same time, a larger background sound reproduction angle can simulate the spatial characteristics of a multi-channel surround system, creating an immersive environment and thus achieving a natural separation of primary and secondary sound sources at the psychoacoustic level.
[0066] This implementation method achieves preliminary layered control of the spatial weights of human voice and background sound by differentiating the core spatial parameter of the playback angle. Its technical effects are twofold: firstly, it significantly improves the intelligibility and attention-grabbing ability of human voice, making it "centered, concise, and clearly distinguishable"; secondly, it maintains the broad spatial expressiveness of background sound effects, ensuring that the overall auditory experience does not feel cramped or oppressive due to voice focusing. This solution is simple in structure and low in implementation cost, making it a fundamental means of constructing a high-contrast spatial sound field.
[0067] In a second feasible implementation, the first playback configuration includes a human voice playback distance, the second playback configuration includes a background sound playback distance, and the human voice playback distance is less than the background sound playback distance.
[0068] It should be noted that, in this embodiment, the voice reproduction distance refers to the straight-line distance from the virtual speaker to the center of the listener's head when spatially rendering the first voice audio signal. This distance directly affects the depth perception and sense of intimacy of the voice.
[0069] Correspondingly, the background sound reproduction distance refers to the distance between the virtual speakers when spatially rendering the first background sound audio signal, which is used to create the spatial depth and environmental immersion characteristics of the background sound.
[0070] In this embodiment, the playback distance for human voices is specified to be smaller than that for background sounds. That is, human voices are rendered as sound sources from a closer location, while background sounds are represented as a collection of sounds from a more distant space. Specifically, the playback distance for human voices can be set within the range of 0.8 to 1.2 meters, close to the typical face-to-face conversation distance, making the speech sound more realistic and interactive. The playback distance for background sounds can be set to 3 meters or more, even reaching 5 to 10 meters, to simulate the far-field reverberation and spatial attenuation characteristics of cinema-grade or open outdoor scenes.
[0071] By designing the difference in playback distance, this implementation makes the second human voice audio signal appear to have a "near-field focusing" attribute in terms of hearing, which enhances the externalization and physical presence of the voice and effectively suppresses the "internalization" listening feeling common in headphone playback; at the same time, the second background sound audio signal exhibits the acoustic characteristics of "far-field diffusion", with a more uniform energy distribution and strong directional ambiguity, which is conducive to constructing a stable background sound layer without focus interference.
[0072] This implementation method, by introducing differentiated control of playback distance, overcomes the limitations of traditional spatial rendering that relies solely on horizontal angle adjustment, achieving fine layering of sound in the foreground and background depth dimensions. Its technical effects include: not only strengthening the spatial priority of human voices as the foreground subject, but also utilizing distant background sounds to create realistic three-dimensional spatial depth, enhancing the overall stereoscopic and layered feel of the audio scene. Especially in applications that emphasize accurate voice reproduction, such as video conferencing and audiobooks, this solution can significantly improve subjective listening quality.
[0073] In a third feasible implementation, the first playback configuration includes a human voice playback elevation angle, the second playback configuration includes a background sound playback elevation angle, and the absolute value of the human voice playback elevation angle is smaller than the absolute value of the background sound playback elevation angle.
[0074] It should be noted that, in this embodiment, the human voice playback elevation angle refers to the vertical angle between the virtual speaker corresponding to the first human voice audio signal and the listener's ear level when the signal is spatially rendered. This parameter is used to control the spatial positioning tendency of the human voice in the vertical direction.
[0075] Correspondingly, the background sound playback elevation angle refers to the vertical tilt angle of the virtual speaker associated with the first background sound audio signal when spatial rendering is performed, which is used to expand the spatial coverage of the sound field in the vertical dimension.
[0076] In this embodiment, the absolute value of the human voice playback elevation angle is specified to be smaller than the absolute value of the background sound playback elevation angle. That is, the virtual speaker for the human voice is basically located near the listener's ear level (close to 0°), while the virtual speakers for the background sound are distributed at higher or lower vertical positions. For example, the human voice playback elevation angle can be controlled within ±5° to ensure that the speech is always anchored in the horizontal area directly in front; while the background sound playback elevation angle can be set to ±30° to ±45° to simulate non-horizontal sound components such as ceiling reflections, ground echoes, or aerial flight effects.
[0077] By setting different playback elevation angles, this implementation method ensures that the second human voice audio signal maintains a high degree of consistency in the vertical direction, avoiding positioning confusion caused by vertical offset and further improving the stability of voice imaging; while the second background sound audio signal obtains stronger vertical spatial diffusion, which can activate the human auditory system's ability to perceive sound sources at the top and bottom, thereby breaking the spatial limitation of traditional stereo which is limited to the horizontal plane.
[0078] This implementation achieves spatial decoupling and layered rendering of sound in the vertical dimension by independently adjusting the playback elevation angle parameter. Its technical effects include: not only enhancing the positioning accuracy of human voices in three-dimensional space and preventing them from "drifting" with changes in content, but also significantly improving the immersive dimension of background sound, allowing users to experience a sense of environmental envelopment from all directions (including above and below). This is of great significance for applications such as virtual reality, panoramic audio, and film sound effects that pursue omnidirectional spatial reproduction, and is an important supplement and upgrade to horizontal planar rendering.
[0079] It should be understood that the above three implementation methods are merely examples of specific implementation paths in this embodiment and do not constitute a limitation on the scope of protection of this application. In practical applications, the above three differentiated control mechanisms can be used individually or flexibly combined according to specific application scenarios. For example, a small-angle + close-range + low-elevation human voice configuration and a large-angle + long-range + high-elevation background sound configuration can be used simultaneously to achieve multi-dimensional collaborative optimization of spatial sound field reconstruction.
[0080] Crucially, in this embodiment, regardless of the parameter combination used, a core design principle must be met: the spatial perception corresponding to the first playback configuration must be higher than that corresponding to the second playback configuration. Only in this way can we ensure that the human voice, as the core of information, always maintains its clear, stable, and focusable auditory advantage.
[0081] Preferably, in this embodiment, the spatial perception corresponding to the first playback configuration should be higher than that corresponding to the standard stereo playback configuration (in some standards, the standard stereo playback configuration is: playback elevation angle of 60°, playback distance of 3m, and absolute value of playback elevation angle of 0°), while the spatial perception corresponding to the second playback configuration should be lower than that corresponding to the standard stereo playback configuration.
[0082] In other words, the spatial rendering of human voices tends to be "more focused and more precise," while the spatial rendering of background sounds tends to be "broader and more diffuse." This reverse differentiation of spatial rendering strategy is the fundamental innovation of this application—by applying spatial modulation with opposite trends to the two types of audio signals, the auditory attention hierarchy is reconstructed, ultimately achieving an ideal acoustic experience of "distinct primary and secondary elements and clear layers" on the target device.
[0083] Step S300: Mix the second human voice audio signal and the second background sound audio signal to obtain the second audio signal; As those skilled in the art will know, audio mixing refers to the technical process of linearly superimposing two or more processed audio signals according to a certain gain ratio, phase relationship, and timing alignment to generate a composite audio signal. This process must ensure that there is no obvious phase cancellation, clipping distortion, or dynamic conflict between the sub-signals, so as to guarantee that the final output sound quality is pure and layered.
[0084] It should be noted that, in this embodiment, the second audio signal refers to the audio stream formed by mixing the second human voice audio signal and the second background sound audio signal. This signal has already incorporated the human voice and background sound components after differential spatial processing, and possesses a preliminary spatial sound field structure. However, since this signal has not yet taken into account the actual acoustic propagation characteristics of the target audio device—that is, the cross paths between the left and right speakers and the user's ears (left speaker → right ear, right speaker → left ear)—crosstalk interference will still be inevitably introduced when the signal is actually played through the target audio device. This will cause the spatial sound image, which was originally carefully constructed through differential playback configuration, to become distorted, blurred, or even collapsed, requiring further elimination of crosstalk interference.
[0085] This embodiment achieves logical separation of sound source levels and preliminary construction of the sound field structure by remixing the differentiated spatially rendered human voice with the background sound. This mixing process not only preserves the spatial attributes of both types of signals but also ensures that the auditory priority of primary and secondary sound sources conforms to human cognitive habits through a gain balancing mechanism (e.g., dynamically adjusting the human voice level according to speech intelligibility requirements). This step, as a crucial intermediate stage before crosstalk cancellation, provides a clear and well-defined input signal foundation for subsequent precise compensation based on a physical acoustic model.
[0086] Step S400: Obtain the target transfer function between the target audio device and the user's ears, and perform crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal.
[0087] As those skilled in the art will know, crosstalk cancellation is an active acoustic signal processing technique that aims to perform reverse compensation on the first audio signal at the playback end through predistortion filtering, so that after being played by the speaker of the target audio device, the user's left ear only receives the desired sound component of the left channel, and the right ear only receives the desired sound component of the right channel, thereby effectively suppressing unexpected crosstalk between the left and right ears and reconstructing accurate binaural auditory spatial cues.
[0088] It should be noted that, in this embodiment, the target transfer function refers to the set of four-channel impulse response functions that describe the acoustic propagation paths from the left and right speakers of the target audio device to the user's left and right ears, respectively.
[0089] It should also be noted that, in this embodiment, the target audio signal refers to the final output audio signal that has undergone crosstalk cancellation processing and is specifically customized and optimized for the target audio device. This signal has been subjected to inverse filtering (i.e., crosstalk cancellation filter) in the time or frequency domain. Its design goal is to ensure that when this signal is played by the target audio device, the sound field actually received by the user's ears is as close as possible to the ideal binaural signal (i.e., the left ear contains only the left virtual sound source, and the right ear contains only the right virtual sound source), thereby achieving high-fidelity spatial sound image externalization and stable positioning.
[0090] It should be noted that the transfer functions mentioned in this embodiment are all acoustic transfer functions. An acoustic transfer function is a transfer function from a sound source to the reproduction region. The transfer function refers to the ratio of the Laplace transform (or z-transform) of the response (i.e., output) of a linear system under zero initial conditions to the Laplace transform of the excitation (i.e., input).
[0091] In this embodiment, the target transfer function between the target audio device and the user's ears is the acoustic transfer function from the output sound source (i.e., speaker or loudspeaker) of the target audio device to the user's ears, which is used to reflect the changes in the output signal during the process of the output signal of the target audio device being transmitted to the user's ears.
[0092] It is understandable that, since the speakers or horns of the target audio device are not ideal sound sources, and the speakers or horns of the target audio device cannot be directly placed into the user's ear canals as audio playback devices, crosstalk problems often occur during the transmission of the target audio signal. To avoid this problem, crosstalk cancellation processing is performed on the second audio signal before the target audio signal is generated. This process cancels the crosstalk problem that occurs during the transmission of the target audio signal after playback. In other words, the target audio signal is obtained by crosstalk cancellation processing on the second audio signal. It can cancel the influence of the playback device itself and the influence of the user's head on the sound signal during the sound transmission process, while retaining the influence of the user's head on the sound transmission result.
[0093] It is understood that since the target audio signal has undergone crosstalk cancellation processing based on the crosstalk cancellation algorithm provided in this embodiment before being output through the speaker, the influence of the target audio device itself and the human head on the sound transmission result is eliminated. Therefore, when the target audio signal enters the user's ears, it can achieve the same effect as the second audio signal. If the second audio signal is the left channel, the target audio signal transmits the sound signal only to the left ear; if the second audio signal is the right channel, the target audio signal transmits the sound signal only to the right ear; if the second audio signal is stereo, the target audio signal outputs the left channel and right channel sound signals to the left ear and right ear respectively.
[0094] This embodiment cancels crosstalk in the second audio signal by utilizing the target transfer function between the target audio device and the user's ears. Essentially, it constructs an inverse system in the frequency or time domain to counteract the crosstalk path effects introduced during physical propagation. In specific implementations, minimum phase approximation, regularized inverse filtering, or robust CTC (Connectionist Temporal Classification) algorithms based on psychoacoustic masking effects can be used to eliminate crosstalk while avoiding high-frequency gain explosion or auditory distortion caused by the ill-conditioned nature of the transfer function. Especially for non-near-ear playback systems like the target audio device, crosstalk path energy is often non-negligible (sometimes even exceeding that of the direct path). Without targeted compensation, even with meticulous spatial rendering in the initial stages, the final auditory experience will still deviate significantly from expectations.
[0095] This embodiment proposes a systematic spatial rendering method by combining sound source separation with differentiated spatial rendering and introducing crosstalk cancellation processing for target audio devices, in order to solve the technical problem that human voices are prone to positioning ambiguity and sound image dispersion during spatial rendering. Specifically, this embodiment first performs source separation on the input first audio signal, accurately extracting the first human voice audio signal and the first background sound audio signal, thereby achieving independent control of different types of sound sources. Subsequently, different spatial perception intensity reproduction configurations are applied to the human voice and background sound respectively. For the human voice, a first reproduction configuration with higher spatial perception (such as a smaller reproduction angle and a closer reproduction distance) is used for spatial rendering to ensure that its sound image is focused and clearly positioned, avoiding the feeling of "sound in the head" or floating left and right. For the background sound, a second reproduction configuration with lower spatial perception is used for rendering, so that it is appropriately diffused in space to form a wide sound field sense of immersion. Afterwards, the rendered second human voice audio signal and the second background sound audio signal are mixed to generate a second audio signal. Furthermore, the target transfer function between the target audio device and the user's two ears is combined to perform crosstalk cancellation processing on the mixed second audio signal, effectively suppressing unexpected crosstalk between the left and right channels, improving channel isolation and spatial imaging accuracy, and finally outputting the target audio signal.
[0096] Through the above-mentioned technical means, this embodiment achieves synergistic optimization of spatial rendering effect and human voice focusing without increasing the number of physical speakers. It not only obtains a wide and natural sound field, but also ensures the clarity and intelligibility of voice content and the stability of sound image, significantly improving the user's immersion and auditory experience quality in typical scenarios such as movie watching and voice interaction.
[0097] Based on the above embodiments, this application proposes a spatial rendering method according to a second embodiment.
[0098] In the second embodiment of this application, the same or similar content as in the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0099] like Figure 3 As shown, in this embodiment, the above-mentioned step S200 performs spatial rendering on the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and performs spatial rendering on the background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, which may include steps S210~S220: Step S210: Obtain the first head-related transfer function corresponding to the preset first replay configuration, and the second head-related transfer function corresponding to the preset second replay configuration; As those skilled in the art will know, the Head-Related Transfer Function (HRTF) is a complex frequency response function that describes the acoustic filtering characteristics experienced by a sound wave as it propagates from a specific spatial direction to the eardrum. This function incorporates the diffraction, reflection, and resonance effects of sound waves caused by human structures such as the head, auricle, external auditory canal, and shoulder, and is a key physiological basis for determining spatial sound perception. The HRTF exhibits high directional selectivity; sound sources at different orientations (horizontal angle, elevation angle) and distances correspond to different HRTF characteristics, especially showing significant differences in spectral peaks and valleys in the high-frequency range, providing important clues for the human brain to determine the location of sound sources.
[0100] In spatial rendering, the head-related transfer function (HRTF) serves as the core processing kernel, simulating the binaural time difference, binaural intensity difference, and spectral shape changes that occur when sound emitted from a virtual speaker reaches the listener's left and right ears. By convolving the input audio signal with the corresponding HRTF, auditory stimuli consistent with real-world spatial hearing can be reconstructed at the headphone output, thereby achieving precise localization of three-dimensional sound images.
[0101] In this embodiment, the first head-related transfer function (HRTF) refers to a set of HRTF data that matches the virtual speaker spatial parameters defined by the first playback configuration, specifically including the frequency response curves of the left and right ear channels. The selection of this function is strictly based on interpolation or retrieval according to the playback angle, playback distance, and playback elevation angle in the first playback configuration. For example, when the first playback configuration is set to ±10° horizontal angle, 1.0 meter distance, and 0° elevation angle, the system will extract or generate HRTF pairs corresponding to these spatial coordinates from a pre-stored HRTF database, ensuring that the rendered voice is accurately presented in the near-center position directly in front of the listener.
[0102] Accordingly, the second-head correlation transfer function refers to the set of HRTFs corresponding to the second playback configuration. Since background noise typically employs a wider spatial layout (such as wide-angle stereo or virtual surround sound), its HRTF may involve a combination of multiple directional points, or even use a diffused average HRTF or a spatialized kernel with enhanced room impulse response, to create a diffuse, enveloping auditory effect. This function is also dynamically obtained based on the specific parameters of the second playback configuration (such as ±60° horizontal angle, 4.0 meter distance, ±30° elevation angle) to ensure the accurate reproduction of the spatial expansion characteristics of the background noise.
[0103] This embodiment achieves refined modeling of the spatial perception characteristics of different types of audio content by acquiring dedicated HRTFs adapted to the playback configurations of human voice and background sound respectively. This step provides an accurate physiological acoustic basis for subsequent differentiated spatial rendering and is a prerequisite for ensuring a reasonable distribution of the final sound image structure.
[0104] Step S220: Spatial rendering of the first human voice audio signal is performed using the first head correlation transfer function to obtain the second human voice audio signal; In this embodiment, the spatial rendering process is completed by performing time-domain convolution or frequency-domain multiplication operations on the first human voice audio signal with the first head correlation transfer functions of the left and right ears, respectively. Specifically, the system decomposes the first human voice audio signal into a time-frequency representation, multiplies it by the frequency response characteristics of the human voice HRTF, and then converts it back to the time domain to generate a two-channel output signal containing spatial information—that is, the second human voice audio signal. When played, this signal allows the listener to perceive that the speech comes from a stable point sound source at close range in front of them, with high focus and low diffusion, effectively avoiding the "in-head effect" or sound image drift.
[0105] Step S230: Spatial rendering of the first background sound audio signal is performed using the second head correlation transfer function to obtain the second background sound audio signal.
[0106] In this embodiment, the spatial properties of the first background sound audio signal are reshaped by performing a similar convolution or filtering operation on the second head-related transfer function. Since the background sound HRTF typically corresponds to a location at a relatively far distance, a large angle, or a non-horizontal plane, the rendered second background sound audio signal exhibits a broad, blurred, and deep sound field characteristic.
[0107] In addition, multi-channel diffusion algorithms or random phase perturbation techniques can be combined to further weaken its directional sharpness, making it more in line with the auditory characteristics of ambient sound that is "without a clear direction but fills the space," thus forming an effective complement to human voices rather than interference.
[0108] This embodiment achieves layered control of spatial auditory characteristics by applying dedicated HRTFs (Hypertext Transfer Functions) to both human voices and background sounds, each matched to its playback configuration. The technical effects are: it not only improves the positioning accuracy and speech intelligibility of human voices, but also enhances the immersiveness and spatial inclusiveness of background sounds. Through their synergistic effect, it significantly optimizes the overall subjective listening quality of the target audio device. Compared to traditional rendering methods using a uniform HRTF template, this solution is more flexible and adaptable, dynamically adjusting spatial representation strategies according to different content types, fully meeting the needs of various application scenarios such as movie playback, video calls, and virtual reality.
[0109] Furthermore, in one feasible implementation, the first background sound audio signal includes a third background sound audio signal corresponding to the first background sound source and a fourth background sound audio signal corresponding to the second background sound source, wherein the first background sound source is different from the second background sound source. The second playback configuration includes the third playback configuration corresponding to the first background sound source and the fourth playback configuration corresponding to the second background sound source. The second head-related transfer function includes the third head-related transfer function corresponding to the third playback configuration and the fourth head-related transfer function corresponding to the fourth playback configuration. The second background sound audio signal includes the fifth background sound audio signal corresponding to the first background sound source and the sixth background sound audio signal corresponding to the second background sound source. The above-mentioned step S230 spatially renders the first background sound audio signal using the second head correlation transfer function to obtain the second background sound audio signal, and may include steps S231~S232: Step S231: Spatial rendering of the third background sound audio signal is performed on the third background sound audio signal through the third head correlation transfer function to obtain the fifth background sound audio signal; Step S232: Spatial rendering of the fourth background sound audio signal is performed on the fourth background sound audio signal using the fourth head correlation transfer function to obtain the sixth background sound audio signal.
[0110] It should be noted that, in this embodiment, the first background sound source refers to a specific environmental or effect sound entity in the first background sound audio signal, such as background music, environmental reverberation, and distant traffic noise, which are sound components with continuous and diffuse characteristics; the second background sound source refers to another type of non-speech sound source in the first background sound audio signal that has clear event attributes or local spatial positioning requirements, such as footsteps, door opening and closing sounds, birdsong, and raindrop sounds, which are transient or highly directional sound effects. The two play different functional roles in the auditory scene: the former is mainly used to construct the overall acoustic atmosphere and sense of space, while the latter is used to enhance the realism of the situation and spatial narrative.
[0111] In this embodiment, the third background sound audio signal refers to the audio stream obtained after sound source separation processing, which contains only the first background sound source component in the first background sound audio signal. Its spectral energy distribution is relatively uniform and its time continuity is strong. It usually exists as the "acoustic background". The fourth background sound audio signal refers to the audio stream obtained after sound source separation processing, which contains only the second background sound source component in the first background sound audio signal. It has obvious start and end boundaries, large dynamic changes, and may carry local spatial clues (such as movement trajectories).
[0112] Correspondingly, the third playback configuration refers to a combination of virtual speaker spatial parameters specifically designed for the first background sound source. Its design goal is to create a wide, deep, and unfocused sound field background. This configuration typically employs a large playback angle (e.g., ±90° to ±120°), a large playback distance (e.g., 4 to 6 m), and a high playback elevation angle (e.g., ±30° to ±45°). Furthermore, it uses HRTF diffusion processing or virtual surround algorithms to further weaken the sharpness of its directional perception, giving it an immersive, enveloping feel that "comes from all directions."
[0113] The fourth playback configuration refers to a combination of spatial parameters independently set for the second background sound source. Its purpose is to preserve or enhance the local spatial positioning information of this type of background sound source. For example, its virtual speaker can be placed to the side or rear (e.g., at a horizontal angle of ±60° to ±150°), at a moderate distance (2 to 3 meters), and the elevation angle can be set according to the actual scene (e.g., -10° for footsteps on the ground, +30° for birds in the air), thereby achieving spatial restoration of specific environmental events and enhancing the realism and spatial layering of the auditory narrative.
[0114] The fifth background sound audio signal refers to the output signal generated after spatial rendering of the third background sound audio signal through the third head correlation transfer function. It has highly diffuse spatial characteristics and its main function is to fill the "background layer" of the auditory space, providing stable and non-interfering atmospheric support for the main sound source.
[0115] The sixth background sound audio signal refers to the output signal generated after spatial rendering of the fourth background sound audio signal through the fourth head correlation transfer function. It retains a strong sense of direction and dynamic spatial changes, and is used to present the location and trajectory of specific environmental events, enriching the detail expression of the auditory scene.
[0116] In this embodiment, the system acquires the third head-related transfer function and the fourth head-related transfer function that match the third and fourth playback configurations, respectively, and performs independent spatial rendering processing on the corresponding first background sound audio signals accordingly: that is, the third background sound audio signal is convolved and filtered using the third head-related transfer function to generate the fifth background sound audio signal; simultaneously, the fourth background sound audio signal is independently filtered using the fourth head-related transfer function to generate the sixth background sound audio signal. The two rendered signals will be merged into a complete second background sound audio signal in the subsequent mixing stage.
[0117] This implementation further subdivides the background sound into different types of sound sources and configures differentiated playback strategies and HRTF cores for each, achieving multi-level spatial modeling of the background sound field. Its technical effect is that it ensures both broad coverage and immersion of the overall ambient sound (achieved by the first background sound) and preserves the spatial directivity and dynamic expressiveness of key environmental sound effects (achieved by the second background sound), thereby constructing a more realistic, three-dimensional, and narrative-driven three-dimensional auditory scene. This solution is particularly suitable for applications such as film, games, and virtual reality, which have high requirements for spatial audio detail, significantly improving the information carrying capacity and artistic expression of background sound.
[0118] In addition, this layered rendering mechanism can be flexibly adjusted according to the device's computing power and latency requirements—it can be merged when resources are limited, and can be extended to more background sound source categories (such as third and fourth background sound sources) in high-performance scenarios, with good scalability and engineering adaptability.
[0119] Based on the above embodiments, this application proposes a spatial rendering method according to a third embodiment.
[0120] In the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0121] like Figure 4 As shown, in this embodiment, the step of obtaining the target transfer function between the target audio device and the user's ears in step S400 above may include steps S410 to S430: Step S410: Obtain the target artificial head transfer function and the target free field transfer function between the target audio device and the user's ears; It should be noted that the target transfer function can be understood as the influence of the user's head contour on the sound signal transmission result. In this embodiment, two different acoustic transfer functions are derived based on two different sound transmission scenarios, and then the target transfer function is calculated. Specifically, the target artificial head transfer function is the acoustic transfer function measured by a preset microphone in the ear canal of the preset artificial head when the preset artificial head is placed at the user's hearing position and the target audio device outputs an audio signal; it includes the influence of the target audio device and the preset artificial head on the sound transmission result. The free field transfer function is the acoustic transfer function measured by preset microphones placed at the left and right ear positions before the preset artificial head was removed, when the target audio device outputs a sound signal; it includes the influence of the target audio device on the sound transmission result.
[0122] In one possible implementation, the steps of obtaining the artificial head transfer function and the free field transfer function include: Step S411: When the preset artificial head is placed at the user's head listening position and the target audio device outputs an audio signal, the target artificial head transfer function between the target audio device and the user's ears is measured through the preset microphone placed in the ear canal of the preset artificial head. And in step S412, when the preset artificial head is removed and the target audio device outputs an audio signal, the target free field transfer function between the target audio device and the user's ears is measured by using preset microphones placed at the left and right ear positions before the preset artificial head was removed.
[0123] It should be noted that in this embodiment, the preset artificial head is an auxiliary device constructed to simulate the user's head for assisting in the measurement of the acoustic transfer function. It can simulate the scenario of the user receiving sound signals emitted from the test audio. The preset artificial head is equipped with left and right ears and ear canals, and microphones for receiving sound signals can be pre-placed in the ear canals.
[0124] It should also be noted that the user's head listening position refers to the position of the user's head when listening to audio using the target audio device.
[0125] As an example, the user's head position for hearing can be referenced. Figure 5 The application scenario shown is characterized by a target audio device that is typically fixed in one location. When a user listens to audio using this device, their head is usually positioned in a specific location, known as the user's head listening position. In practical applications, different target transfer functions between the target audio device and the user's ears can be measured in advance at different user head listening positions. These functions can be stored in a mapping table or through other means. Then, during crosstalk cancellation, the user's current head position relative to the target audio device can be determined using various methods to ascertain the current user head listening position. Finally, the target transfer function matching this user head listening position can be found from the pre-saved data for crosstalk cancellation.
[0126] This embodiment utilizes two microphones pre-installed in the ear canals of a pre-designed artificial head to measure the acoustic transfer function from the target audio device to both ears of the pre-designed artificial head, and records this acoustic transfer function as H1, which is the target artificial head transfer function. Next, two microphones identical to those in the ear canals of the pre-designed artificial head in step S411 are placed at the left and right ears of the pre-designed artificial head. The pre-designed artificial head is then removed, while the target audio device remains in its original position. The acoustic transfer function of the target audio device operating in a free field is measured using the two microphones unaffected by the pre-designed artificial head, and recorded as H2, which is the target free-field transfer function.
[0127] Step S420: Perform the inverse operation on the target free field transfer function to obtain the inverse transfer function of the target free field; Step S430: Multiply the target artificial head transfer function with the target free field inverse transfer function to obtain the target transfer function between the target audio device and the user's ears.
[0128] In this embodiment, the target free field transfer function H2, which includes the influence of the target audio device on the sound transmission result, is first inverted to obtain the target free field inverse transfer function, denoted as H2'. Then, the target artificial head transfer function H1, which includes the influence of the target audio device and the preset artificial head on the sound transmission result, is multiplied by H2' to obtain the target transfer function H between the target audio device and the user's ears.
[0129] It should be noted that H2' obtained after the inverse operation can eliminate the influence of the target audio device's own structure on the sound transmission result. Multiplying it by H1 can eliminate the part of H1 that affects the sound transmission result from the target audio device's own structure, while retaining the influence of the preset artificial head on the sound transmission result as the target transfer function H.
[0130] In one possible implementation, the step of performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal includes: Step S440: Perform the inverse operation on the target transfer function to obtain the target inverse transfer function; Step S450: Multiply the second audio signal with the target inverse transfer function to obtain the target audio signal.
[0131] As can be seen from the above steps, the target transfer function H represents the influence of the human head contour on the sound transmission result. It should be understood that the target inverse transfer function obtained after inverting H is equivalent to an identity matrix, which represents eliminating the influence of the human head contour on the sound transmission result. The target audio signal obtained after applying it to the second audio signal can obviously cancel the influence of the human head contour on the sound transmission result during sound signal transmission, so that the sound signal received by the user's two ears can be consistent with the second audio signal.
[0132] In one possible implementation, after obtaining the target transfer function between the target audio device and the user's ears in step S400 above, the spatial rendering method may further include step A10: Step A10: Obtain the ambient acoustic parameters of the user's current environment, including reflected sound intensity, reverberation time, and dominant reflection direction; After obtaining the target transfer function between the target audio device and the user's ears, the ambient acoustic parameters of the user's current environment are further obtained. These parameters include reflected sound intensity, reverberation time, and dominant reflection direction. Specifically, a microphone array can be placed in the user's current environment, and a set of test signals can be played through the microphone array to acquire the room impulse response (RIR) in real time. The ambient acoustic parameters are then calculated based on the acquired room impulse response.
[0133] Reflected sound intensity refers to the energy of reflected sound in the environment within a specific time window, such as the reflected sound energy within 5-50ms after the test signal is emitted; reverberation time refers to the time required for sound to decay to a certain level in the environment, such as the time required for the test signal to decay by 60dB; dominant reflection direction refers to the main propagation direction of reflected sound, such as the direction in which the reflected sound is strongest.
[0134] Step A20: Generate the corresponding environmental compensation coefficients based on the environmental acoustic parameters; After obtaining the environmental acoustic parameters, corresponding environmental compensation coefficients are generated based on these parameters. The generation of environmental compensation coefficients aims to adjust the target transfer function according to the environmental acoustic characteristics in order to compensate for the impact of the environment on audio propagation and perception.
[0135] Specifically, different correspondences between environmental acoustic parameters and environmental compensation coefficients can be pre-defined through experimental measurement or experience, and the environmental compensation coefficients corresponding to the environmental acoustic parameters can be obtained based on these correspondences. Alternatively, a prediction model can be pre-trained with environmental acoustic parameters as input and environmental compensation coefficients as output. After obtaining the environmental acoustic parameters, the environmental acoustic parameters can be input into the prediction model, and the corresponding environmental compensation coefficients can be output.
[0136] It should be noted that the transfer function can be represented as a complex matrix, and therefore, the environmental compensation coefficient can be a coefficient matrix.
[0137] Step A30: Multiply the environmental compensation coefficient by the target transfer function to obtain the compensated target transfer function; Step A40: Based on the compensated target transfer function, perform crosstalk cancellation on the second audio signal to obtain the target audio signal.
[0138] Based on the compensated target transfer function, a crosstalk cancellation process is performed on the second audio signal to obtain the target audio signal. The purpose of crosstalk cancellation is to eliminate crosstalk effects caused by mutual interference between speakers during audio signal propagation, thereby improving audio positioning accuracy and spatial perception. Performing crosstalk cancellation based on the compensated target transfer function allows for a more accurate consideration of the impact of environmental acoustic characteristics on audio propagation, ensuring that the subsequently played target audio signal meets the acoustic transmission characteristics of the current spatial environment, resulting in superior audio performance in various environments. For example, in environments with strong reflected sound, crosstalk cancellation using the compensated target transfer function can reduce the interference of reflected sound on audio positioning, enabling users to more accurately perceive the direction of audio source and enhancing the immersion and realism of the audio.
[0139] To facilitate understanding of the spatial rendering method in the above embodiments of this application, a specific embodiment is provided: like Figure 5 and 6 As shown, in this specific embodiment, music source separation is first performed based on a deep neural network (i.e., a sound source separation model). The vocal content of the dual-path stereo audio (i.e., the first audio signal) is separated from other content to obtain a dual-path signal containing only vocal content (i.e., the first vocal audio signal) and a dual-path signal containing background content (i.e., the first background audio signal). Then, the dual-path signal containing only vocal content and the dual-path signal containing background content are rendered independently and differentially in a stereo space to achieve the effect of expanding the sound field while ensuring clear vocal positioning. Next, all left channels are superimposed to obtain the final left channel signal, and all right channels are superimposed to obtain the final right channel signal, completing the mixing process. Finally, crosstalk testing is performed to obtain the target transfer function between the target audio device and the user's ears. Based on this target transfer function, crosstalk is eliminated on the mixed dual-path signal (i.e., the second audio signal) to obtain the final dual-path signal (i.e., the target audio signal), achieving the effect of simulating headphone playback. That is, the first audio signal received by the target audio device is processed by source separation to obtain a first human voice audio signal and a first background sound audio signal; the first human voice audio signal is spatially rendered based on a preset first playback configuration to obtain a second human voice audio signal, and the first background sound audio signal is spatially rendered based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration; the second human voice audio signal and the second background sound audio signal are mixed to obtain a second audio signal; the target transfer function between the target audio device and the user's two ears is obtained, and crosstalk cancellation is performed on the second audio signal based on the target transfer function to obtain the target audio signal.
[0140] Music source separation refers to separating different music sources in a mixed audio stream; for example, extracting different musical components such as vocals, bass, and drums from the mixed audio. Separating music sources facilitates independent analysis and processing of the characteristics and attributes of each sound source in the input audio. In the field of music source separation, existing methods can be divided into two categories: 1. Model-based Traditional Music Source Separation Methods: Traditional model-based music source separation methods mainly include principal component analysis (PCA) and nonnegative matrix factorization (NMF). PCA can extract the main audio components from the mixed signal, but these components usually do not directly correspond to specific instruments. NMF has the advantage of strong interpretability, but its separation performance depends on understanding the characteristics of the music data and requires multi-channel audio signals as input.
[0141] 2. Deep Learning-Based Music Source Separation Methods: Common network structures include frequency-domain based music source separation networks, time-domain based music source separation networks, and network models combining the time and frequency domains. These deep neural networks, with their powerful nonlinear modeling capabilities, can automatically learn and find features and fit the relationship between input and output, demonstrating excellent music source separation performance.
[0142] like Figure 7 As shown in this specific embodiment, the stereo spatial rendering algorithm logic is as follows: To ensure the stability and clarity of the human voice image, a virtual speaker with a 60° angle (i.e., a playback angle of 60°) corresponding to the standard stereo speaker playback configuration is selected. Spatial rendering is performed on the dual-channel signal containing only human voice content, that is, virtual speaker synthesis is performed using HRTF in the ±30° direction. In other words, the first human voice audio signal is spatially rendered based on the preset first playback configuration to obtain the second human voice audio signal.
[0143] To expand the stereo playback sound field width, a virtual speaker with an included angle 2θ greater than 60° (i.e., a playback angle greater than 60°) is selected. Spatial rendering is performed on the dual-path signal containing background content, that is, HRTF in the ±θ direction is selected for virtual speaker synthesis. In other words, the first background sound audio signal is spatially rendered based on a preset second playback configuration to obtain the second background sound audio signal.
[0144] Considering the role of reverberation in improving the "head-in-the-head effect", this specific embodiment can also add appropriate artificial reverberation to the signal path, that is, add artificial reverberation to the target audio signal.
[0145] like Figure 8 As shown in this specific embodiment, the specific processing logic for crosstalk cancellation (i.e., crosstalk cancellation) is as follows: The input signal is X, which is the first audio signal; The preset artificial head is placed on the user's head at the listening position. Using the preset microphone that is pre-placed in the ear canal of the preset artificial head, the acoustic transfer function from the target audio device to the user's ears is measured and denoted as H1, which is the target artificial head transfer function. Two additional microphones were placed at the left and right ear positions of the preset artificial head. The preset artificial head was then removed, and the acoustic transfer function of the target audio device when it was working in a free field was measured and denoted as H2, which is also the target free field transfer function. Inverting H2 yields H2', which is the inverse transfer function of the target free field. Multiplying H1 and H2' yields the target transfer function between the target audio device and the user's ears, denoted as H. This step aims to eliminate the influence of the target audio device's structure on sound propagation. Inverting H, denoted as B, yields the inverse target transfer function. After separating the audio source of the input signal X, HRTF rendering M (i.e., spatial rendering of differential playback configuration) is performed on each signal, and the rendered signals are mixed to obtain the intermediate signal, denoted as XM, which is also the second audio signal. Crosstalk cancellation is performed on XM by B, resulting in the output signal Y, which is the target audio signal.
[0146] It should be noted that the above specific embodiments are only used to assist in understanding this application and do not constitute a limitation on the spatial rendering method in this application. Any simple modifications based on this technical concept are all within the protection scope of this application.
[0147] In addition, please refer to Figure 9 , Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the spatial rendering method in this application embodiment.
[0148] This application also provides an audio device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the spatial rendering method in the above embodiments.
[0149] The following is for reference. Figure 9 It shows a structural schematic diagram of an audio device suitable for implementing the embodiments of this application. Figure 9 The audio device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0150] like Figure 9As shown, the audio device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the audio device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the audio device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagram shows audio equipment with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0151] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0152] The audio device provided in this application, employing the spatial rendering method described in the above embodiments, can achieve spatial rendering effects while maintaining the stability and focus of human voices. Compared with the prior art, the beneficial effects of the audio device provided in this application are the same as those of the spatial rendering method provided in the above embodiments, and other technical features of this audio device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0153] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0154] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the above claims.
[0155] In addition, this application also provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the steps of the spatial rendering method in the above embodiments.
[0156] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0157] The aforementioned computer-readable storage medium may be included in an audio device or may exist independently without being assembled into an audio device.
[0158] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an audio device, cause the audio device to: perform source separation processing on a first audio signal received by a target audio device to obtain a first human voice audio signal and a first background sound audio signal; perform spatial rendering on the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and perform spatial rendering on the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration; perform mixing processing on the second human voice audio signal and the second background sound audio signal to obtain a second audio signal; obtain the target transfer function between the target audio device and the user's two ears, and perform crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal.
[0159] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0161] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0162] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for performing the steps of the above-described spatial rendering method, which can achieve spatial rendering effects while also ensuring the stability and focus of human voices. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the spatial rendering method provided in the above embodiments, and will not be repeated here.
[0163] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the spatial rendering method described above.
[0164] The computer program product provided in this application can achieve spatial rendering effects while maintaining the stability and focus of human voices. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the spatial rendering method provided in the above embodiments, and will not be repeated here.
[0165] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A spatial rendering method, characterized in that, The spatial rendering method includes the following steps: The first audio signal received by the target audio device is subjected to source separation processing to obtain a first human voice audio signal and a first background sound audio signal; Based on a preset first playback configuration, the first human voice audio signal is spatially rendered to obtain a second human voice audio signal, and based on a preset second playback configuration, the first background sound audio signal is spatially rendered to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration. The second human voice audio signal and the second background sound audio signal are mixed to obtain the second audio signal; The target transfer function between the target audio device and the user's ears is obtained, and crosstalk cancellation is performed on the second audio signal based on the target transfer function to obtain the target audio signal.
2. The spatial rendering method as described in claim 1, characterized in that, The step of obtaining the target transfer function between the target audio device and the user's ears includes: Obtain the target artificial head transfer function and the target free field transfer function between the target audio device and the user's ears; Inverting the target free field transfer function yields the target free field inverse transfer function; Multiplying the target artificial head transfer function with the target free field inverse transfer function yields the target transfer function between the target audio device and the user's ears.
3. The spatial rendering method as described in claim 2, characterized in that, The steps of obtaining the target artificial head transfer function and the target free field transfer function between the target audio device and the user's ears include: When a preset artificial head is placed at the user's hearing position and the target audio device outputs an audio signal, the transfer function of the target artificial head between the target audio device and the user's ears is measured using a preset microphone placed in the ear canal of the preset artificial head; and... When the preset artificial head is removed and the target audio device outputs an audio signal, the target free field transfer function between the target audio device and the user's ears is measured by using preset microphones placed at the left and right ear positions before the preset artificial head was removed.
4. The spatial rendering method as described in claim 2, characterized in that, The step of performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal includes: Invert the target transfer function to obtain the target inverse transfer function; The second audio signal is multiplied by the target inverse transfer function to obtain the target audio signal.
5. The spatial rendering method as described in claim 1, characterized in that, The steps of spatially rendering the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and spatially rendering the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, include: Obtain the first head-related transfer function corresponding to the preset first replay configuration, and the second head-related transfer function corresponding to the preset second replay configuration; The first human voice audio signal is spatially rendered using the first head-related transfer function to obtain the second human voice audio signal; The second background audio signal is obtained by spatial rendering of the first background audio signal using the second head correlation transfer function.
6. The spatial rendering method as described in claim 1, characterized in that, The step of performing source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal includes: The first audio signal received by the target audio device is input into a pre-trained sound source separation model to obtain the first human voice audio signal and the first background sound audio signal output by the sound source separation model.
7. The spatial rendering method according to any one of claims 1 to 6, characterized in that, After the step of obtaining the target transfer function between the target audio device and the user's ears, the method further includes: The ambient acoustic parameters of the user's current environment are obtained, including reflected sound intensity, reverberation time, and dominant reflection direction. Generate corresponding environmental compensation coefficients based on the aforementioned environmental acoustic parameters; Multiply the environmental compensation coefficient by the target transfer function to obtain the compensated target transfer function; Based on the compensated target transfer function, the step of performing crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal is performed.
8. A spatial rendering device, characterized in that, The spatial rendering device includes: The separation module is used to perform source separation processing on the first audio signal received by the target audio device to obtain a first human voice audio signal and a first background sound audio signal. The rendering module is used to perform spatial rendering of the first human voice audio signal based on a preset first playback configuration to obtain a second human voice audio signal, and to perform spatial rendering of the first background sound audio signal based on a preset second playback configuration to obtain a second background sound audio signal, wherein the spatial perception corresponding to the first playback configuration is higher than the spatial perception corresponding to the second playback configuration. The mixing module is used to mix the second human voice audio signal and the second background sound audio signal to obtain the second audio signal; The cancellation module is used to obtain the target transfer function between the target audio device and the user's ears, and to perform crosstalk cancellation on the second audio signal based on the target transfer function to obtain the target audio signal.
9. An audio device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the spatial rendering method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a spatial rendering program, which, when executed by a processor, implements the steps of the spatial rendering method as described in any one of claims 1 to 7.