Signal Processing Method, Apparatus, and Electronic Device for Audio Rendering

By optimizing the response signal based on perceptual characteristics, the calculation complexity and time-consuming in audio rendering are reduced, and the problem of high computing density in the prior art is solved, and the efficient audio rendering effect is achieved in application scenarios with low hardware performance.

CN117837173BActive Publication Date: 2025-06-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280057718.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-27
Filing Date
2022-08-26
Publication Date
2025-06-13
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

The prior art has high computational complexity and time-consuming when realizing high-fidelity audio rendering, especially in application scenarios with low hardware performance, which is difficult to effectively reduce computational density.

Method used

By acquiring the response signal set, the response signal is processed based on the perceptual characteristics associated with the response signal, and the response signal is optimized to reduce the number, thereby reducing the computational complexity and improving the computational efficiency.

Benefits of technology

It realizes that in application scenarios with low hardware performance, the calculation complexity and time-consuming of audio rendering are reduced, the computing efficiency is improved, and the real spatial audio experience can be effectively provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117837173B_ABST
    Figure CN117837173B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a signal processing method, apparatus, and electronic device for audio rendering. The signal processing method for audio rendering includes: obtaining a set of response signals, the set of response signals including response signals obtained from a sound signal, where the sound signal is a signal received at a listening position; and processing the response signals in the set of response signals based on perceptual characteristics associated with the response signals to obtain response signals suitable for audio rendering, where the number of the response signals suitable for audio rendering is less than or equal to the number of the response signals in the set of response signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of audio signal processing, and particularly relates to a signal processing method, apparatus, and electronic device for audio rendering, as well as a non-transitory computer-readable storage medium. Background Art

[0002] The realism of sound in 3D spatial audio is an important consideration in spatial audio, and sound rendering or audio rendering is also crucial for high-fidelity audio effects. Sound rendering or audio rendering refers to appropriately processing the sound signal from the sound source to provide the desired listening experience for the user in the user application scenario. Sound rendering or audio rendering can often be performed by means of various appropriate acoustic models.

[0003] Currently, there are mainly two methods for modeling indoor room acoustics: One is to model through the method of wave acoustics. In wave acoustics, the wave equation is solved according to the data, the space is discretized into smaller elements and their interactions are modeled. It is computationally intensive, and the load increases rapidly with frequency. Therefore, the method of wave acoustics is more suitable for the low-frequency part. The other is to model through the method of geometric acoustics. The theory of geometric acoustics regards sound as rays and ignores the wave nature of sound, and calculates the propagation of sound through the propagation of rays. The calculation of geometric acoustics is also computationally intensive, and it is necessary to calculate a large number of rays and the energy of the rays to render the sound. However, geometric acoustics can more accurately simulate the propagation path of sound in the physical space and the attenuation of energy, and can physically simulate spatial audio to achieve the rendering effect of high-fidelity audio. Summary of the Invention

[0004] According to some embodiments of the present disclosure, there is provided a signal processing apparatus for audio rendering, which includes an acquisition module configured to acquire a set of response signals, the set of response signals including response signals obtained based on a sound signal, where the sound signal is a signal received at a listening position, and a processing module configured to process the response signals in the set of response signals based on the perceptual characteristics related to the response signals to obtain response signals suitable for audio rendering, where the number of the response signals suitable for audio rendering is less than or equal to the number of the response signals in the set of response signals.

[0005] According to some embodiments of the present disclosure, a signal processing method for audio rendering is provided, including obtaining a set of response signals, the set of response signals including response signals derived from a sound signal, where the sound signal is a signal received at a listening position, and processing the response signals in the set of response signals based on perceptual characteristics associated with the response signals to obtain response signals suitable for audio rendering, where the number of response signals suitable for audio rendering is less than or equal to the number of response signals in the set of response signals.

[0006] According to some embodiments of the present disclosure, an audio rendering apparatus is provided, including a signal processing module as described herein, configured to process response signals derived from a sound signal from a sound source to a listening position, and a rendering module, configured to perform audio rendering based on the processed response signals.

[0007] According to some embodiments of the present disclosure, an audio rendering method is provided, including processing response signals derived from a sound signal from a sound source to a listening position, and performing audio rendering based on the processed response signals.

[0008] According to still some other embodiments of the present disclosure, a chip is provided, including: at least one processor and an interface, the interface for providing computer-executable instructions for the at least one processor, and the at least one processor for executing the computer-executable instructions to implement the signal processing method for audio rendering and the audio rendering method according to any one of the embodiments described in the present disclosure.

[0009] According to still some other embodiments of the present disclosure, a computer program is provided, including: instructions that, when executed by a processor, cause the processor to execute the signal processing method for audio rendering and the audio rendering method according to any one of the embodiments described in the present disclosure.

[0010] According to still some other embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor configured to execute the signal processing method for audio rendering and the audio rendering method according to any one of the embodiments described in the present disclosure based on instructions stored in the memory device.

[0011] According to still some other embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the signal processing method for audio rendering and the audio rendering method according to any one of the embodiments described in the present disclosure.

[0012] According to still some other embodiments of the present disclosure, a computer program product is provided, including instructions that, when executed by a processor, implement the signal processing method for audio rendering and the audio rendering method according to any one of the embodiments described in the present disclosure.

[0013] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings:

[0015] Figure 1A Schematic diagrams showing some embodiments of an audio signal processing process;

[0016] Figure 1B Schematic diagrams showing a conventional audio signal rendering process;

[0017] Figure 2A Block diagrams showing a signal processing apparatus for audio rendering according to some embodiments of the present disclosure;

[0018] Figure 2B Flowcharts showing a signal processing method for audio rendering according to some embodiments of the present disclosure;

[0019] Figure 2C Block diagrams showing an audio rendering apparatus according to some embodiments of the present disclosure;

[0020] Figure 2D Flowcharts showing an audio rendering method according to some embodiments of the present disclosure;

[0021] Figure 3A Graphs showing the auditory threshold according to some embodiments of the present disclosure;

[0022] Figure 3B Schematic diagrams showing the perceptual masking effect according to some embodiments of the present disclosure;

[0023] Figure 4A Schematic diagrams showing an exemplary audio rendering process according to some embodiments of the present disclosure;

[0024] Figure 4B Flowcharts showing an exemplary processing operation according to some embodiments of the present disclosure;

[0025] Figure 5 Block diagrams showing some embodiments of an electronic device of the present disclosure;

[0026] Figure 6 Block diagrams showing other embodiments of an electronic device of the present disclosure;

[0027] Figure 7Block diagram showing some embodiments of the chip of the present disclosure. Detailed Description

[0028] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure, its application or use. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0029] Unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure. At the same time, it should be understood that, for the sake of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship. Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0030] The following refers to Figure 1A Describe some embodiments of the audio signal processing process, where Figure 1A The implementation of the various stages of an exemplary audio rendering process / system is particularly shown, exemplarily including a production stage or a manufacturing stage, and a consumption stage, and optionally further includes an intermediate processing stage, such as compression.

[0031] During the production or manufacturing stage, input audio data and audio metadata can be received and processed, particularly for authorization and metadata tagging, to obtain a production result. Exemplarily, the input for audio processing can include, but is not limited to, object-based audio signals, FOA (First-Order Ambisonics), HOA (Higher-Order Ambisonics), stereo, surround sound, etc. In some embodiments, the audio data is input to a track interface for processing, and the audio metadata is processed via general audio source data (such as ADM extensions, etc.). Optionally, normalization processing can also be performed, particularly on the results obtained from authorization and metadata tagging.

[0032] In some embodiments, during the audio content production process, the creator also needs to be able to monitor and modify the work in a timely manner. As an example, an audio rendering system can be provided to provide the monitoring function for the scene. In addition, to enable consumers to obtain the artistic intention that the creator wants to express, the rendering system provided for the creator's monitoring should be the same as the rendering system provided for consumers to ensure a consistent experience.

[0033] Optionally, according to the embodiments of the present disclosure, after the captured audio signal is produced and before being provided to the consumption stage (which may include or be referred to as the audio rendering stage), further intermediate processing can be performed on the audio signal. In some embodiments, the intermediate processing of the audio signal can include appropriate compression processing, including encoding / decoding. As an example, the produced audio content can be encoded / decoded to obtain a compression result, and then the compression result can be provided to the rendering side for rendering. The encoding / decoding in the compression can be implemented using any appropriate technology. In other embodiments, the intermediate processing of the audio signal can also include the storage and distribution of the audio signal. For example, the audio signal can be stored and distributed in appropriate formats, such as in audio storage format and audio distribution format respectively. The audio storage format and audio distribution format can be various appropriate forms in the audio processing system, which will not be described in detail here.

[0034] It should be noted that the above audio intermediate processing process, formats for storage, distribution, etc. are merely exemplary and not restrictive. The audio intermediate processing can also include any other appropriate processing and can adopt any other appropriate format, as long as the processed audio signal can be effectively transmitted to the audio rendering end for rendering.

[0035] It should be noted that the audio transmission process also includes the transmission of metadata. The metadata can be in various appropriate forms, can be applicable to all audio renderers / rendering systems, or can be respectively and correspondingly applied to each audio renderer / rendering system. Such metadata can be referred to as rendering-related metadata, and can include, for example, basic metadata and extended metadata. The basic metadata is, for example, ADM basic metadata conforming to BS.2076. The ADM metadata describing the audio format can be given in the form of XML (Extensible Markup Language). In some embodiments, the metadata can be appropriately controlled, such as hierarchical control.

[0036] Then, in the consumption stage, the audio signal from the audio production stage (and optionally, after intermediate codec processing) is processed for playback / presentation to the user. In particular, the audio signal is rendered and presented to the user with the desired effects. In particular, the audio data and metadata can be respectively recovered and rendered; then the processing results are input to the audio device after audio rendering processing. As an example, as Figure 1A shown, after receiving the audio signal from the audio production stage (and optionally, after intermediate codec processing), the track interface and general audio metadata (such as ADM extensions, etc.) can be used to respectively recover and render the data and metadata; the results after recovery and rendering are subjected to audio rendering, and the obtained results are input to the audio device for consumption by the consumer. As another example, in the case where audio signal representation compression is also performed in the intermediate stage, corresponding decompression processing can also be performed at the audio rendering end.

[0037] According to embodiments of the present disclosure, the processing in the audio rendering stage can include various appropriate types of audio rendering. In particular, for each type of audio representation, corresponding audio rendering processing can be adopted.

[0038] In some embodiments, the processing in the audio rendering stage can include scene-based audio rendering. In particular, in scene-based audio (SBA), the rendering system is independent of the capture or creation of the sound scene. The rendering of the sound scene is usually performed on the receiving device and generates real or virtual speaker signals. The vector S of the speaker array signal = [S 1 …S n T can be created in the following manner, where n represents the nth speaker.

[0039] S = D · B

[0040] where B is the vector of the SBA signal B = [B (0,0) …B (n,m) T ​​, where n and m represent the order and degree of the spherical harmonic function, and D is the rendering matrix (also called the decoding matrix) of the target speaker system.

[0041] In a more common scenario, the audio scene is presented by playing back binaural signals through headphones. The binaural signals can be obtained by convolving the virtual speaker signal S and the binaural impulse response matrix IR of the speaker positions. BIN obtained by convolution.

[0042] S BIN =(D.B)*IR BIN

[0043] In immersive applications, it is desired that the sound field rotates according to the movement of the head. Such rotation can be achieved by multiplying the SBA signal by a rotation matrix F.

[0044] B' = F.B

[0045] In other aspects, additionally or alternatively, the processing in the audio rendering stage may include channel-based audio rendering. Channel-based formats are most widely used in traditional audio production. Each channel is associated with a corresponding speaker. The positions of the speakers are standardized in, for example, ITU-R BS.2051 or MPEG CICP. In some embodiments, in the context of immersive audio, each speaker channel is regarded as a virtual sound source in a scene and rendered to the headphones; that is, the audio signal of each channel is rendered to the correct position in a virtual listening room according to the standard. The most straightforward method is to filter the audio signal of each virtual sound source with the response function measured in a reference listening room. The acoustic response function can be measured with microphones placed in the ears of a human or an artificial head. They are called binaural room impulse responses (BRIRs).

[0046] In still other aspects, additionally or alternatively, the processing in the audio rendering stage may include object-based audio rendering. In object-based audio rendering, each object sound source is presented independently along with its metadata that describes the spatial attributes of each sound source, such as position, direction, width, etc. Using these attributes, the sound sources are individually rendered in the three-dimensional audio space around the listener. The rendering can be performed for a loudspeaker array or headphones. Loudspeaker array rendering uses different types of speaker panning methods (such as VBAP, Vector Base Amplitude Panning) to present the listener with the feeling of the object sound source at a specified position by playing the sound using the loudspeaker array. There are also various different ways to render for headphones, such as directly filtering the signal of each sound source with the HRTF (Head Related Transfer Function) corresponding to the direction of the sound source. Indirect rendering methods can also be adopted, where the sound sources are rendered onto a virtual loudspeaker array and then binaural rendering is performed on each virtual loudspeaker.

[0047] It should be noted that the audio rendering processing here may include or correspond to various appropriate processes performed in the rendering stage according to embodiments of the present disclosure, including but not limited to reverberation, such as the calculation of ARIR (Acoustic Room Impulse Response), BRIR (Binaural Room Impulse Response), etc. In particular, for the realistic spatial effects of 3D spatial audio, the effect of reverberation is crucial.

[0048] Figure 1B A conventional audio rendering processing procedure involving audio spatial reverberation is shown, where first, a set of impulse responses R from a sound source is obtained, then the set of impulse responses R is time-blocked, and calculations are performed based on the time-blocked set of impulse responses R to obtain the acoustic room impulse response (ARIR).

[0049] Spatial reverberation can be achieved by various appropriate methods, such as geometric-acoustics-based spatial reverberation. In the calculation of geometric-acoustics-based spatial reverberation, mainly the ray-tracing method is used to simulate how a large number of sounds propagate in the geometric space and the environment, calculate the impulse / impulse response between the sound source and the listener through the propagation of the sound rays, then convert the ray signal into the corresponding directional spatial impulse / impulse response, and convert a large number of impulse / spatial impulse responses into binaural impulse responses, thus calculating the effect of late reverberation in the 3D space. However, to obtain a realistic sense of spatial reverberation through the ray-tracing method, a large number of spatial impulse responses need to be calculated and convolution operations need to be performed, which is very time-consuming and computationally intensive for personal computers and mobile phones. Therefore, it is very necessary to reduce the computational complexity of this method and the time consumption brought by the calculation.

[0050] In view of such problems, in some implementations, multi-process and multi-thread methods have been proposed, that is, by using high-end personal computers and mobile phones to allocate computationally intensive and complex parts to other processes or threads for calculation to reduce the computational load; and GPU and TPU calculation methods, which are similar to multi-thread methods and also allocate computationally intensive and complex parts to high-end hardware and peripherals for calculation to improve computational performance. However, as can be seen from the above, for the problem of computational intensity and complexity in the process of calculating late reverberation through the ray tracing algorithm, these optimization methods mainly use the performance of hardware to solve this problem. This hardware-dependent method cannot effectively solve the problems of computational intensity and time consumption, especially for application scenarios with low hardware performance (for example, mid- to low-end personal computers or mobile devices).

[0051] In view of this, the present disclosure proposes an improved technical solution to optimize signal processing in audio rendering, especially signal processing for reverberation processing in audio rendering. In particular, the present disclosure proposes to optimize a set of response signals obtained from a sound signal originating from a sound source to obtain optimized response signals suitable for audio rendering, especially a relatively smaller number of response signals, thereby reducing computational complexity and improving computational efficiency. In this way, for application scenarios with low hardware performance, especially for example, low-end personal computers or mobile devices, a real spatial audio experience can also be obtained.

[0052] Figure 2A The block diagram of a signal processing apparatus for audio rendering according to an embodiment of the present disclosure is shown. The signal processing apparatus 2 includes an acquisition module 21 configured to acquire a set of response signals, the set of response signals including response signals obtained from a sound signal, where the sound signal is a signal received at a listening position, and a processing module 22 configured to process the response signals in the set of response signals based on perceptual characteristics related to the response signals to obtain response signals suitable for audio rendering, where the number of response signals suitable for audio rendering is less than or equal to the number of response signals in the set of response signals. In particular, by appropriately processing the response signals, a smaller number of response signals suitable for audio rendering, especially reverberation calculation, can be obtained, thereby reducing the complexity in reverberation calculation and improving efficiency. This will be described in detail below.

[0053] According to embodiments of the present disclosure, the sound signal received at the listening position may be from a sound source. In particular, the sound signal from the sound source may include sound signals that propagate from the sound source to the listening position in various ways, such as at least one of a sound signal that propagates directly from the sound source to the listening position and a sound signal that propagates indirectly (e.g., via various reflections) from the sound source to the listening position. In some embodiments, the sound signal may be a sound signal in various suitable forms. For example, it may include a ray signal, which may be obtained by simulating the propagation of sound in a geometric space and an environment through a ray tracing method, especially a ray signal used in spatial reverberation calculations based on geometric acoustics theory.

[0054] According to embodiments of the present disclosure, the response signal may include various suitable response signals converted from the sound signal, such as impulse response, shock response, etc., especially a spatial impulse response to be utilized in reverberation calculations based on geometric acoustics theory. In particular, the response signal may indicate the response signal obtained by the sound from the sound source at the listening position. Various suitable conversion methods may be employed. In some embodiments, when the sound signal is a ray signal from the sound source to the listener, the impulse response may be a directional impulse response converted from the ray signal. The following will describe by taking the impulse response as an example, where the response signal and the impulse response may be used interchangeably, and the set of response signals will correspond to the set of impulse responses, which contains at least one impulse response or response signal. It should be noted that the embodiments of the present disclosure can also be applied to other types of response signals as long as the response signal can be converted from the sound signal and can be used for audio rendering, especially reverberation calculations.

[0055] According to some embodiments, the obtained set of impulse responses may include at least one impulse response, which may correspond to at least one sound signal arriving at the listening position from a sound source. The sound signal may include at least one of a direct signal, a reflected signal, etc. from the sound source to the listening position. For example, one impulse signal may correspond to one sound signal. On the one hand, in some embodiments, the set of impulse responses may include the impulse response derived from the direct sound signal directly propagated from the sound source to the listening position. On the other hand, in some embodiments, the set of impulse responses may further include the impulse response derived from the reflected sound signal from the sound source to the listening position. In particular, the reflected sound signal may refer to the reflected signal after the sound signal emitted from the sound source is reflected on any object or reflective position in the listening space. Thus, the set of impulse responses may include the impulse response corresponding to the sound signal from the sound source to the reflective position and then from the reflective position to the listening position. According to some embodiments, the reflected sound signal is particularly the late reflected sound signal for reverberation calculation. In particular, the late reflected sound signal may refer to the sound signal in the reflected signals that takes a longer time to reach the listening position from the sound source, for example, a sound signal exceeding a specific time length; or a sound signal that has been reflected a relatively large number of times from the sound source, for example, a sound signal exceeding a specific number of reflections.

[0056] According to embodiments of the present disclosure, the impulse response may be represented by appropriate information. In some embodiments, the impulse response may be represented by the time information of the sound signal, the sound intensity, the sound spatial orientation information, etc., where the time information may include any one of the timestamp from the sound source to the listening position, the propagation time length, etc. In some embodiments, the impulse response may be in various appropriate formats, such as a vector or vector format. Each element in the vector may correspond to the information data for representing the impulse response, for example, may include a time data element, a sound intensity element, a spatial direction element, etc. In some embodiments, the obtained set of impulse responses may be in various appropriate forms, such as a vector form, where the respective corresponding data of all impulse responses are arranged in a data string; or a matrix form, for example, the rows may correspond to the respective impulse responses, and the columns may indicate the corresponding data of each impulse response, etc.

[0057] According to embodiments of the present disclosure, the set of impulse responses may be obtained in various appropriate ways. In some embodiments, a signal processing device may obtain or receive the sound signal from the sound source to the listening position and process the sound signal, such as appropriate conversion, to obtain the set of impulse responses. In other embodiments, other appropriate devices may obtain or receive the sound signal from the sound source to the listening position to generate the set of impulse responses and provide it to the signal processing device.

[0058] According to an embodiment of the present disclosure, after obtaining the response signal set, the signal processing device will process the response signal set, especially the response signals in the response signal set, so as to obtain a response signal suitable for audio rendering. In particular, the response signal suitable for audio rendering can be obtained from the response signal set and the number is less than the number of the original response signals in the response signal set. In some embodiments, signal processing can be performed based on the perceptual characteristics related to the response signal, so as to achieve response signal reduction, reduce the number of response signals for audio rendering, and reduce the processing complexity.

[0059] According to some embodiments of the present disclosure, the perceptual characteristics related to the response signal may include characteristics related to the user's sound perception when listening to the sound corresponding to the response signal at the listening position, which may also be referred to as psychoacoustic perception characteristics, psychoacoustic characteristics, etc. The perceptual characteristics may include various appropriate information. In some embodiments, the perceptual characteristics may include the perceptual data of the user when listening to the sound at the listening position, and in particular may include information or data related to at least one of the auditory loudness of the sound signal, the mutual interference between sound signals, the proximity between sound signals, etc. The perceptual data can be calculated from the information carried by the perceptual signal, for example, calculated from the signal intensity, signal spatial orientation information, signal time information, etc. of the perceptual signal. And the perceptibility of the response signal can be judged based on the calculated perceptual data. For example, it can be judged whether the perceptual data meets the perceptual requirements, especially whether it can be effectively perceived, by comparing the perceptual data with a specific threshold, so as to determine whether the sound corresponding to the response signal can be effectively perceived.

[0060] In other embodiments, additionally or alternatively, the perceptual characteristics may include perceptual condition-related information, such as indicating the perceptual condition of the sound at the listening position, such as whether it is in a mutual influence condition (such as especially a masking condition), whether it is in a condition where the sound pressure is too low to be perceived, etc. As an example, the perceptual condition information can be indicated by corresponding bits, symbols, etc. For example, 1 bit can be used to indicate the perceptual condition information, where "1" can indicate that it can be perceived and is suitable for audio rendering, and "0" can indicate that it cannot be perceived, such as a masking condition or a condition where the sound pressure is too low to be perceived. As another example, 1 bit can be used to indicate the masking condition and 1 bit can be used to indicate the sound pressure condition respectively. It should be noted that only when both of these two bits are "1", the response signal is considered to be perceivable and suitable for audio rendering. The perceptual condition information can be obtained by comparing the corresponding perceptual data with a threshold. As an example, this especially corresponds to the following situation: the perceptual condition is determined by other devices based on the perceptual data and directly sent to the signal processing device, so that the signal processing device can more intuitively determine the perceptual condition of the signal and perform signal processing accordingly.

[0061] According to embodiments of the present disclosure, perceptual characteristics, particularly perceptual data and / or perceptual condition information, can be obtained in various suitable ways. In particular, perceptual characteristics can be obtained specifically for individual sound signals, particularly individual impulse responses. In some embodiments, they can be obtained by other suitable devices and provided to the processing module. For example, they can be obtained by a device outside the signal processing device, or a device or module outside the processing module in the signal processing device, and provided to the processing module. In other embodiments, the processing module itself can calculate the perceptual characteristics, particularly the perceptual data, of individual sound signals, particularly individual impulse responses.

[0062] In some embodiments, the above-mentioned acquisition of perceptual characteristics can be particularly performed by the perceptual characteristic acquisition module 222. The perceptual characteristic acquisition module 222 can obtain perceptual data based on the information of the acquired response signal or sound signal. For example, it can perform operations based on the information of the response signal or sound signal to obtain perceptual data. Alternatively, the perceptual characteristic acquisition module 222 can obtain perceptual data from other devices or equipment, or directly obtain perceptual condition information.

[0063] According to embodiments of the present disclosure, based on the perceptual characteristics related to the response signal, it can be determined whether the user meets the perceptual requirements when listening to the sound corresponding to the response signal at the listening position, such as whether it can be effectively perceived. Here, the perceptual requirements can correspond to the conditions or situations required for the sound corresponding to the response signal to be effectively perceived, such as non-masking situations, signal strength conditions, etc., and can be in various suitable forms. In particular, the process of determining whether the perceptual requirements are met can be particularly performed by the determination module 223. In some embodiments, the perceptual requirements can correspond to a specific perceptual condition threshold. The perceptual data of the response signals in the response signal set can be compared with the specific threshold, and based on the comparison result, it can be determined whether the perceptual requirements are met. Additionally or alternatively, in other embodiments, the perceptual requirements can correspond to indication information of an effectively perceivable situation (such as a non-masking situation, a situation where the signal sound pressure is sufficient to be perceived, etc.). It can be determined whether the information related to the perceptual situation of the response signals in the response signal set is indication information of an effectively perceivable situation. If so, it can be considered that the perceptual requirements are met; otherwise, it can be considered that the perceptual requirements cannot be met. As an example, it can be directly determined whether the information related to the perceptual situation is 1 or 0. If it is 0, the requirements cannot be met and it cannot be effectively perceived.

[0064] Thus, response signals that do not meet the perception requirements can be processed. For example, such response signals are not directly used for audio rendering, but are ignored, removed, merged, etc. As a result, compared with the set of acquired response signals, the number of response signals suitable for audio rendering can be appropriately reduced, effectively reducing the computational load and improving the computational efficiency. In particular, considering that there are multiple reflected signals, especially late reflected signals, at the listening position, such computationally intensive problems will be relatively prominent. In the embodiments according to the present disclosure, by processing the response signals (e.g., impulse responses) of the reflected signals, especially late reflected signals, at the listening position, a reduction in the impulse response of the reflected signals for audio rendering can be achieved.

[0065] Exemplary implementations of signal processing based on perceptual characteristics of embodiments of the present disclosure will be described below. Among them, exemplary implementations that apply the perceptual data included in the perceptual characteristics will be particularly described. However, it should be noted that the information related to the perceptual situation included in the perceptual characteristics can be similarly applied.

[0066] According to embodiments of the present disclosure, the perceptual characteristics related to the response signal can include various types of perceptual characteristics, especially including but not limited to relative perceptual characteristics (which can also be referred to as the first perceptual characteristics). The relative perceptual characteristics can relate to or indicate the relative perceptual situation between the response signals in the set of response signals, such as the masking situation, etc. In particular, the relative perceptual characteristics can include or indicate information related to the masking situation. In this case, correspondingly, the perception requirement is a requirement related to the corresponding perceptual characteristic, such as a requirement related to the masking situation. For example, whether the perception requirement is met can be whether the masking situation is large. And when the masking situation is large, especially when it is greater than the masking requirement corresponding to the perception requirement, it can be considered that the perception requirement is not met. Otherwise, when the masking situation is small, especially when it is less than or equal to the masking requirement corresponding to the perception requirement, it can be considered that the perception requirement is met. In this way, it is possible to determine whether there is masking of the response signal based on the relative perceptual characteristics between the response signals, and perform signal processing in the case of determining the existence of masking, such as including at least one of ignoring, removing, etc. the masked signal, or merging the signals in which the masking situation occurs. In this way, the response signals can be screened based on the masking situation. In particular, for example, sound signals with a greater mutual masking effect can be appropriately merged, so as to appropriately reduce the amount of data for audio rendering processing, in order to reduce the computational load and improve the computational efficiency.

[0067] It should be noted that the relative perceptual situation is not limited to the masking situation. It can also relate to other mutual interference and mutual influence situations of the response information. And when the mutual interference and mutual influence of the response information are large enough to cause the sound to be unable to be accurately listened to / perceived, it can be considered that the perception requirement is not met.

[0068] According to an embodiment of the present disclosure, the processing of response signals may further include comparing the relative perceptual characteristics between signals (such as in particular relative perceptual data) with a specific threshold (which may be referred to as the mutual perception threshold), and determining based on the comparison result whether the signals affect each other (in particular, for example, whether they mask each other). Thus, in the case of determining mutual masking, at least one of reduction processes such as ignoring, removing, or combining the signals may be performed.

[0069] In some embodiments of the present disclosure, masking may involve or indicate masking between adjacent signals, and may be classified into different types of masking depending on the signal adjacency type. In particular, masking may include at least one of time masking, spatial masking, frequency domain masking, etc. For example, time masking may refer to a masking situation occurring between signals that are temporally adjacent, spatial masking may refer to a masking situation occurring between signals that are spatially adjacent, and frequency domain masking may refer to a masking situation occurring between signals that are frequency adjacent.

[0070] According to an embodiment of the present disclosure, the relative perceptual characteristics between signals may involve the proximity between signals, particularly including temporal proximity relationships, spatial proximity relationships, frequency domain proximity, etc. Thus, by comparing the proximity between signals with a specific proximity threshold (which may be referred to as the first proximity threshold), and in the case of being less than the threshold, it may be considered that the signals are very close to each other so that masking may occur. For example, if the time difference between response signals is too small, for example, two response signals are very close in time, or the spatial distance between temporally adjacent response signals is too small, for example, two response signals are very close in space, then it may be considered that masking may occur between these two response signals and they will affect each other in perception. Therefore, these two signals need to be processed, for example, combined, in order to eliminate the masking and achieve signal reduction.

[0071] In other embodiments, additionally or alternatively, it may further depend on the signal strength relationship between response signals to determine whether masking may exist. For example, if the intensities between response signals within a specific time period or spatial range (such as an appropriate adjacent range) significantly affect each other, for example, the sound intensity difference between two loudness signals is very large, such as greater than a specific sound intensity threshold, then it may be determined that masking exists, and the masked signal may be removed or combined with another signal to achieve signal reduction.

[0072] Specifically, when a user listens to sound from a sound source at a listening position, the human ear's perception of the sound is affected by the masking effect. When a sound A with a relatively high sound pressure acts on the human ear, if sound B also acts on the human ear at this time, the human ear's auditory system's perception of sound B in terms of time and space will decline. For sounds below the masking threshold, the human ear can hardly perceive them, and this is when the masking effect occurs. In particular, when the signal energy of the previously presented sound A exceeds a certain threshold, it will suppress the subsequently presented low-energy signal B. The masking effect will increase as the masking sound A strengthens, and will also weaken as the masked sound B strengthens; when the energy of the subsequently presented signal B in the human ear's auditory perception is relatively large, much larger than the previously presented signal A, backward masking will also occur, as Figure 3A shown.

[0073] In particular, according to embodiments of the present disclosure, neighboring signals can be first determined, and then based on the mutual perception-related data between the neighboring signals, such as values calculated based on at least one of the spatial information, intensity information, etc. of the signals, it can be determined whether there is masking between the neighboring signals. Here, the neighboring signals can indicate signals within a specific time period or spatial range, or signals with a time difference or spatial difference between them less than a specific threshold. Here, the specific threshold can be a second proximity threshold, which is usually greater than or equal to the previous first proximity threshold, so as to more accurately determine the masking situation and perform more appropriate processing on the signals, especially merging processing.

[0074] According to some embodiments, the merging of impulse responses can be performed in various appropriate ways. In some embodiments, the merging includes performing mathematical statistics on the attribute information of two impulse responses determined to be mutually masked, such as at least one of spatial information, time information, intensity information, etc., to obtain a new impulse response. As an example, the mathematical statistics can be averaging, such as various appropriate types of average calculations, such as spatial averaging, weighted averaging, etc. For example, the merging of two impulse responses can include averaging the time information, spatial information, and intensity information of each impulse response respectively, so as to obtain an impulse response obtained by average calculation. Also for example, the mathematical statistics can be the mean of the spatial positions of the impulse responses or the weighted average of the spatial positions of the impulse responses. For example, weighted averaging can be performed based on the sound pressure level / intensity of the impulse responses.

[0075] As an example, for two impulse responses that may undergo temporal masking and / or spatial masking, the merged impulse response can be represented as follows:

[0076]

[0077] where r t,s can indicate the impulse response, Indicates the impulse response at the first time and the first spatial position, which may indicate the impulse response at the second time and the second spatial position, where when these two impulse responses are masked in time and / or in space, they can be combined to obtain a new impulse response r'. t,s . The time masking condition can be represented by t 2 -t 1 ≤τ T where τ T represents the time threshold related to time masking; the spatial masking condition can be represented by s 2 -s 1 ≤τ S where τ S represents the spatial threshold related to spatial masking. It should be noted that the merging conditions here are merely exemplary, and there can be other exemplary masking conditions, such as the signal energy difference being greater than a specific energy threshold, the signal energy ratio being less than a specific threshold, and so on.

[0078] The following will describe an exemplary implementation of the processing performed by the signal processing module according to the relative perception characteristics according to an embodiment of the present disclosure.

[0079] According to some embodiments, the signal processing module can be configured to determine, for each impulse response in the impulse response set, the proximity between the impulse response and other impulse responses in the impulse response set, including but not limited to at least one of temporal proximity, spatial proximity, and frequency-domain proximity, and process the impulse response based on the proximity. In particular, when the proximity between two impulse responses is less than a specific threshold, such as the aforementioned first proximity threshold, it can be considered that the two impulse responses are too close and masking may occur, and thus these two signals are appropriately processed, such as merged processing.

[0080] In particular, when the proximity is temporal proximity, the time difference between the impulse responses can be determined, and when the time difference is less than a specific time threshold, such as the aforementioned first proximity threshold, it can be considered that these two signals are masked. As another example, when the proximity is spatial proximity, the spatial distance between the impulse responses can be determined, and when the spatial distance is less than a specific spacing threshold, such as the aforementioned first proximity threshold, it can be considered that these two signals are masked. Here, the spatial distance between the impulse responses can include information related to the spatial interval, such as the spatial angular interval. In some embodiments, the information related to the spatial interval can be related to the spatial vector interval between the impulse responses. In some embodiments, the information related to the spatial interval is represented by the statistical characteristics of the spatial vector interval between the impulse responses, such as cosine value, sine value, etc.

[0081] According to some embodiments of the present disclosure, additionally or alternatively, based on the attribute information of the response signal, such as time information, spatial information, intensity information, etc., the mutual perception data before the response signal can be determined, and then the response signal can be processed based on the mutual perception data, such as performing the reduction processing as described above. Here, the mutual perception data mainly relates to or indicates whether a masking situation will occur between the response signals, and thus can also be referred to as masking situation-related information.

[0082] According to some embodiments, additionally or alternatively, the signal processing module can be configured to, for each impulse response in the impulse response set, determine the adjacent response set of the impulse response in the impulse response set, and for the adjacent response set, perform screening based on the masking situation-related information between the impulse responses. In particular, the adjacent response may refer to the impulse responses adjacent in the time and / or spatial dimensions. The adjacent response set of the impulse response is substantially a subset of the obtained impulse response set, which may refer to a subset of the impulse responses within a specific time range and / or spatial range including the impulse response, or a subset of the impulse responses with a time difference and / or spatial difference less than a specific threshold from the impulse response. Here, the specific range or specific threshold may correspond to, for example, the aforementioned second proximity threshold.

[0083] In some embodiments, the time adjacent response set of the impulse response is substantially a subset of the obtained impulse response set, which may refer to a subset of the impulse responses within a specific time range including the impulse response. For example, if the impulse response to be calculated is the impulse response at 2.5 seconds, its time adjacent response set may refer to the set of impulse responses within the time range between 2 seconds and 3 seconds. Or the adjacent response set may include the impulse responses with a time difference less than or equal to a specific time threshold from the impulse response, such as the aforementioned second proximity threshold, for example, it may correspond to 0.5 seconds. The time range or threshold can be set appropriately, such as by empirical setting. Preferably, the time range corresponds to the time difference between the sound signals that may mutually mask, and this time difference can be determined through experiments, empirical determination, etc. The time value here can be the time point of arrival at the listening position, or the propagation time length to the listening position, etc.

[0084] In some embodiments, for each impulse response, the obtained impulse response set can be traversed to determine whether each of the other impulse responses belongs to the time adjacent response set, such as whether it is within the time range. Or rather, for each impulse, the obtained impulse response set can be traversed to determine whether the time difference between each of the other impulse responses and the impulse is less than a specific threshold, such as the aforementioned second proximity threshold.

[0085] In particular, for facilitating the determination of the time-proximity response set of impulse responses, the impulse responses in the obtained impulse response set can also be sorted in time. In some embodiments, the processing module includes a sorting module 221 configured to sort the impulse responses in the obtained impulse response set, preferably in time order, such as sorting from the earliest to the latest arrival time at the listening position, or sorting according to the propagation time of the impulse responses from the shortest to the longest. It should be noted that other sorting methods are also possible as long as they can be appropriately sorted in time. Sorting the impulse response set can further appropriately improve the processing efficiency. As an example, for each impulse response, only the previous and the next impulse responses of the impulse response can be used as neighboring responses for judgment. As another example, only the impulse responses within a specific time range before and after the impulse response, or a specific number of impulse responses before and after the impulse response, can be used as neighboring responses for judgment. In this way, it is not necessary to traverse the entire impulse response set, thereby reducing the computational amount of the judgment process and improving the processing efficiency. It should be noted that the sorting operation can be performed by other devices / equipment, and the sorted impulse responses can be input into the signal processing device.

[0086] According to an embodiment of the present disclosure, the signal processing module is configured to determine the relative perceptual characteristics between every two impulse responses in the neighboring response set, which can be referred to as masking condition-related information. For the masking condition-related information of the impulse responses, if the masking condition indicated by the masking condition-related information between two impulse responses is large, these two impulse responses will be merged to construct a new impulse response for use in calculations in audio rendering; otherwise, the impulse responses will remain unchanged. The following gives an exemplary implementation of the calculation and application of the masking condition-related information.

[0087] As an example, depending on the implementation of the masking condition-related information, when the masking condition-related information is greater than a specific threshold, it can be considered that the masking condition indicated by the masking condition-related information is large. In this case, it can be considered that the perceptual requirements, especially the masking requirements included in the perceptual requirements, correspond to the specific threshold, and meeting the perceptual requirements can correspond to being less than or equal to the specific threshold. For example, according to the neighboring response set Calculate the intervals between the spatial vectors in the current set, such as the set of cosines of the interval angles, as the aforementioned masking condition-related information

[0088]

[0089] where and represent the vector representations of two responses in the neighboring response set Here, adding an arrow indicates the direction because each response has a direction coordinate value in space, which is equivalent to a vector. The denominator |r i|,|r j respectively indicate the magnitudes of these two responses, such as the magnitude of a vector in a specific coordinate system, which may correspond to the distance of the sound from the listener or the listening position. In this way, a cosine set between every two responses in the obtained set of neighboring responses is obtained.

[0090] Then, according to the set and the spatial cosine threshold ζ T , which can also be called a specific interval threshold, it is determined whether masking occurs. If masking occurs, a merging process is performed to generate a new set R′ t,s

[0091]

[0092] In particular, for each value in the set , it is compared with a specific threshold. And when it is greater than the threshold, that is, the angular interval / spacing between the two responses in the set is very small, which means that the two responses are too close to each other. The two responses corresponding to this value in the set are merged, such as the mean of two impulse responses. It should be noted that other merging methods can also be used. For other cases, these two impulse responses can be retained. In this way, through merging, the impulse responses included in the set of impulse responses can be reduced to obtain a new set.

[0093] Of course, the above is only exemplary, and other appropriate methods can also be used to determine the spatial interval / distance between response signals. As an example, depending on the implementation of the masking condition-related information, when the masking condition-related information is less than a specific threshold, it can be considered that the masking condition indicated by the masking condition-related information is large. For example, a sine set of spatial vectors can be determined, and when the spatial sine value is less than a specific threshold, which can also be called a specific interval threshold, this corresponds to a large masking condition, and then merging is performed. In this case, it can be considered that the perceptual requirements, especially the masking requirements included in the perceptual requirements, are corresponding to the specific interval threshold, and meeting the perceptual requirements can correspond to being greater than the specific interval threshold.

[0094] In some embodiments, the masking status related information between every two impulse responses can be calculated sequentially starting from the first impulse response in the temporally adjacent response set. In particular, the masking status related information between the first impulse response and each of the other impulse responses, and then the masking status related information between the second impulse response and each of the subsequent impulse responses, so as to obtain the masking status related information between all impulse responses in the temporally adjacent response set. Then, each piece of masking status related information is compared with a specific threshold. For two impulse responses whose masking status related information indicates a large masking status, these two impulse responses will be merged to construct a new impulse response for use in calculations in audio rendering. Otherwise, these two impulse responses can remain unchanged.

[0095] In some embodiments, in some embodiments, the masking status related information between every two impulse responses can be calculated sequentially starting from the first impulse response in the temporally adjacent response set, and a judgment process is performed along with the calculation of the masking status related information. That is, for each piece of masking status related information calculated, it is then judged whether this piece of masking status related information indicates a large masking status. If the masking status is large, a merging process is performed, and then subsequent calculations and judgment processes of the masking status related information are performed based on the impulse response obtained by the merging. This can further reduce the amount of calculation and judgment processes and improve the time processing efficiency.

[0096] It should be noted that the above calculation and judgment processes for the masking status related information in the temporally adjacent response set can equally apply to the spatially adjacent response set.

[0097] In particular, the spatially adjacent response set of an impulse response can be obtained in a manner similar to that of the temporally adjacent response set. The spatially adjacent response set of an impulse response can, for example, refer to a subset of impulse responses within a specific spatial range including this impulse response, or can be a set composed of this impulse response and impulse responses with a spatial interval less than a specific threshold from the impulse response. The spatial range or threshold can be appropriately set, for example, determined through experiments or set empirically. Preferably, the spatial range corresponds to the spatial interval between sound signals that may mutually mask, and this spatial interval can be determined through experiments, empirically determined, etc.

[0098] In some embodiments, for each impulse response, the obtained impulse response set can be traversed to judge whether each of the other impulse responses belongs to the spatially adjacent response set, for example, whether it is within this spatial range. Or rather, for each impulse, the obtained impulse response set can be traversed to judge whether the spatial interval between each of the other impulse responses and this impulse is less than a specific threshold, such as the aforementioned second proximity threshold.

[0099] In particular, for facilitating the determination of the spatially adjacent response set of impulse responses, the impulse responses in the obtained impulse response set can also be sorted spatially. In some embodiments, the sorting module 221 can also be configured to sort the impulse responses in the obtained impulse response set, preferably sorted according to the spatial interval, for example, sorted from near to far according to the spatial interval between the impulse response and the reference position in the listening environment, or based on a specific impulse response, sorted from near to far according to the spatial interval between other impulse responses and the reference impulse response, and so on. In this way, for each impulse response, the impulses adjacent to it in the sorting can be directly selected as the adjacent response set. For example, in a manner similar to sorting by time impulse, the impulse responses that are immediately adjacent to it, or a specific number of adjacent ones, or within a specific spatial range, or the spatial interval is less than a specific threshold, can be selected. This eliminates the need to traverse the entire impulse response set, thereby reducing the computational complexity of the judgment process and improving the processing efficiency.

[0100] Then, for the determined spatially adjacent response set, information related to the masking situation between the response signals in the spatially adjacent response set is determined, and the merging process is performed under the condition of judging masking, which can be performed as described above. As an example, the spatial proximity between the response signals in the spatially adjacent response set can be determined, and when the response signals are adjacent to each other, for example, less than a specific threshold, such as the aforementioned first threshold, it can be considered that masking will occur between the response signals, and then the response signals determined to be masked are processed.

[0101] In some embodiments, the above calculation and judgment process for the information related to the masking situation in the temporally adjacent response set can be extended to the entire obtained impulse response set, so that the impulse response screening can be performed for the entire obtained impulse response set.

[0102] The implementation of signal processing according to the embodiments of the present disclosure will be described below, especially the implementation for the absolute perception characteristics. According to some embodiments of the present disclosure, the absolute perception characteristics can relate to the auditory attributes of the sound related to the response signal itself, especially the perceived intensity, such as the absolute sound intensity, relative sound intensity, sound pressure, etc. In particular, the absolute perception characteristics can include information related to the intensity of the sound signal, especially the intensity-related information of the impulse response. In some embodiments, the intensity-related information is the sound pressure level of the frequency band or channel corresponding to the sound signal, especially the impulse signal. In other embodiments, the intensity-related information is the relative intensity information of the sound signal (e.g., sound pressure) relative to the reference intensity (e.g., sound pressure), especially corresponding to the auditory threshold.

[0103] As an example, whether the human ear can hear a sound depends on the frequency of the sound and whether the amplitude is higher than the absolute auditory threshold at this frequency. The absolute auditory threshold is the minimum intensity value that the human ear can feel. The auditory intensity of the human ear for sounds in different frequency bands is different. This auditory intensity, especially the auditory threshold, can correspond to the intensity at which the human ear can appropriately perceive the sound in this frequency band. The auditory threshold curve of the human ear is as shown in Figure 3A shown, and when the intensity of the sound signal is lower than the absolute hearing threshold, the human ear cannot perceive the existence of the sound. Thus, such a sound signal can be removed from the audio rendering process, which can reduce the computational load. Here, the hearing threshold can correspond to the aforementioned intensity-related information, and the absolute hearing threshold corresponds to the aforementioned intensity-related threshold.

[0104] In the embodiments according to the present disclosure, additionally or optionally, it is also possible to compare the absolute perception characteristic values of each response signal with a specific threshold (which can also be referred to as the perception threshold or the absolute perception threshold) to determine which sound signals are suitable for audio rendering. For example, sound signals above the specific threshold can be effectively perceived, while sound signals below the specific threshold may not be effectively perceived and can be filtered out, thereby further appropriately reducing the amount of data used for audio rendering processing. In particular, for the obtained set of response signals, especially the reduced set of response signals obtained through the above embodiments, it is possible to determine whether the response signal will participate in the reverberation calculation based on the signal intensity attribute of the response signal therein. In particular, whether it participates in the convolution calculation for obtaining the binaural impulse response, so as to calculate the sound pressure level for each channel through the absolute psychological auditory threshold to reduce the complexity of the binaural impulse response based on convolution.

[0105] In some embodiments, the absolute response characteristic corresponds to the intensity-related information of the signal, and the signal processing module can be configured to compare the intensity-related information with a specific intensity-related threshold during signal processing. When the intensity-related information is lower than the specific intensity-related threshold (which can also be referred to as the perception intensity threshold or the absolute perception intensity threshold), the corresponding sound signal, especially the corresponding impulse response, can be removed and does not need to be used for audio rendering processing, which can effectively reduce the computational burden of the audio rendering processing. In some embodiments, the intensity-related information can be in various appropriate representation forms, such as a sound intensity signal, a sound pressure signal, a relative value obtained based on a reference intensity signal, a relative value obtained based on a reference sound pressure signal, etc. The intensity-related threshold can be a threshold in the corresponding form. In other embodiments, the intensity-related information can be determined in an appropriate manner, such as being determined for a frequency band, for a channel, etc.

[0106] As an example, for the loudness signal, calculate the relative intensity value related to hearing for each channel

[0107]

[0108] Among them, p represents the sound pressure of the loudness signal, and p ref represents the reference sound pressure, which is defined as the minimum sound pressure that can be heard by young people with normal hearing for a sound signal at room temperature of 25°C, standard atmospheric pressure, and 1000 Hz, which is 20 μPa. Then, it is compared with the standard absolute hearing threshold to determine whether the sound pressure of the current channel is within the audible range of the human ear.

[0109]

[0110] Among them, L audible The corresponding sound signal equal to 1 is a sound that can be effectively perceived and can be used for calculating the impulse response of the binaural room, that is, it can be applied to audio rendering processing. L audible The corresponding sound signal equal to 0 is a sound that cannot be effectively perceived, and the corresponding response signal will be discarded or removed without further involving audio rendering or reverberation calculation. It should be noted that L audible The above values are only exemplary, and it can also be other appropriate values as long as the values can distinguish the above different situations.

[0111] It should be noted that the above calculation is only exemplary, and the intensity-related information can also be determined in other appropriate ways, such as based on frequency bands, based on time blocks, etc. In addition, screening based on the intensity-related information can be performed in various other appropriate ways, such as directly determining the intensity, sound pressure, etc., and then comparing the intensity with the intensity threshold and the sound pressure with the sound pressure threshold for screening.

[0112] In some embodiments, it can be performed on each impulse response in the acquired impulse response set. Among them, the intensity-related information is the sound pressure level of the frequency band corresponding to the impulse response included in the impulse response set. In other embodiments, it can be performed on the impulse response blocks in the acquired impulse response set. Among them, the impulse response block can be an impulse response block obtained by dividing the impulse response set according to time. Among them, the intensity-related information is the sound pressure level of the frequency band corresponding to the impulse response block included in the impulse response set. In particular, each impulse response block can correspond to at least one frequency band, so that the sound pressure level can be obtained for each frequency band corresponding to the impulse response block. Thus, when the sound pressure level of the impulse response is less than a specific threshold, the impulse response will be removed and not used in the calculation for audio rendering. This can effectively reduce the amount of data used in the audio rendering calculation, reduce the computational complexity and computational time, and improve the computational efficiency.

[0113] According to embodiments of the present disclosure, signal processing can also utilize both relative perception characteristics and absolute perception characteristics simultaneously, that is, utilize both intensity-related information and masking condition-related information to screen the impulse response, thereby further reducing the amount of data used for audio rendering processing, reducing the computational complexity and computational workload, and improving the processing efficiency. In some embodiments, preferably, the impulse response is first appropriately processed according to the masking condition-related information, such as merging, retaining, ignoring, removing, etc., and then for the processed impulse response, each impulse response is further screened according to the intensity-related information of the signal, so as to further obtain a reduced set of impulse responses. In other embodiments, for a given set of response signals, each impulse response can be screened according to the intensity-related information of the signal to obtain a reduced set of impulse responses, and then for the reduced set of impulse responses, the impulse response can be appropriately processed according to the masking condition-related information, such as merging, retaining, removing, ignoring, etc., so as to obtain a further reduced set of impulse responses.

[0114] The above mainly describes the signal processing operations performed when the perception characteristics include perception data, including determining the perception status (such as whether it is masked, whether it is not perceptible enough, etc.) and the corresponding processing based on the determination result. It should be noted that when the perception characteristics include perception status-related information, the signal processing operations can also be performed similarly. For example, the perception status-related information can be set by comparing the perception data with a threshold as described above. In particular, the perception status can be determined by judging the value of the perception status-related information, and then the corresponding processing is performed based on the determination result. For example, it can be determined whether the perception status-related information is 1 or 0, and in the case of 0, the above signal processing such as merging, ignoring, removing, etc. is performed.

[0115] According to embodiments of the present disclosure, after optimizing the response signal applicable to audio rendering, further processing can be performed on the response signal, such as dividing the response signal into blocks, especially time blocks, and then performing audio rendering on the divided response signal, such as calculating ARIR, optionally or additionally calculating BRIR. The operations such as block division, ARIR or BRIR calculation, etc. can be performed in various appropriate ways, such as various well-known ways in the art, which will not be described in detail here.

[0116] In particular, the signal processing according to embodiments of the present disclosure can be applied to audio rendering processing in a suitable manner. In particular, it can be applied to audio rendering processing either centrally or distributively. In particular, compared with the conventional signal processing process shown in FIG. 1, the signal processing process is optimized by newly added modules, and the newly added modules can correspond to the signal processing device according to embodiments of the present disclosure, where the response signal is optimized according to relative perceptual characteristics, especially by removing redundant responses with the help of information related to the mutual masking situation, and / or the response signal is optimized according to absolute perceptual characteristics, especially by calculating the perceptual channels as intensity-related information to further process the signal, so that an optimized set of impulse signals can be obtained for audio rendering.

[0117] In some other embodiments, the signal processing according to embodiments of the present disclosure can be applied before chunking. As Figure 4A (a) shows, specifically, after obtaining the impulse response set R, the signal processing according to embodiments of the present disclosure can be applied to the impulse responses in the impulse response set R. In particular, redundant responses can be removed with the help of information related to the mutual masking situation, and / or perceptual channels can be calculated for the impulse responses as intensity-related information to further process the signal. For example, impulse responses with intensity-related information lower than a specific threshold can be removed, and then the optimized set of impulse signals thus obtained is time-chunked, and then audio rendering is performed based on the chunked impulse signals, such as calculating ARIR, optionally or additionally calculating BRIR.

[0118] In some embodiments, the signal processing according to embodiments of the present disclosure can be applied after chunking. As Figure 4A (b) shows, specifically, after obtaining the impulse response set R and chunking it according to time, the signal processing according to embodiments of the present disclosure can be applied to the impulse responses in each time chunk. In particular, redundant responses can be removed with the help of information related to the mutual masking situation, and / or perceptual channels can be calculated for the impulse responses as intensity-related information to further process the signal. For example, impulse responses with intensity-related information lower than a specific threshold can be removed, so that they need to participate in the reverberation calculation for audio rendering, and thus an optimized set of impulse signals can be obtained for audio rendering, such as calculating ARIR, optionally or additionally, BRIR is also calculated.

[0119] In still some other embodiments, the signal processing according to embodiments of the present disclosure can be dispersed before and after chunking. As Figure 4AAs shown in (c), after obtaining the impulse response set R, signal processing according to an embodiment of the present disclosure can be applied to the impulse responses in the impulse response set R. In particular, redundant responses can be removed by means of mutual masking condition-related information, and then the processed impulse responses can be time-blocked. Then, for each impulse response block, perceptual channels are calculated for the impulse responses as intensity-related information to further process the signal. For example, impulse responses with intensity-related information lower than a specific threshold can be removed, and thus audio rendering can be performed based on the further processed signal, such as calculating ARIR, and optionally or additionally calculating BRIR. It should be noted that in this distributed implementation, the operation of removing redundant responses by means of mutual masking condition-related information and the operation of calculating perceptual channels as intensity-related information to further process the signal can be performed interchangeably. For example, perceptual channels can be calculated as intensity-related information to process the signal before blocking, and redundant responses can be removed by means of mutual masking condition-related information after blocking.

[0120] Thus, in the present disclosure, by determining whether the perceptual characteristics of the response signal meet the perceptual requirements, such as whether the perceptual characteristics in the time and / or space dimensions meet the perceptual requirements, and performing at least one process such as removal, ignoring, merging, etc. on the response signals that do not meet the requirements, this is equivalent to performing psychoacoustic masking on the response signals that do not meet the requirements, thereby reducing the number of impulse responses while the performance of the algorithm still maintains high performance and high fidelity.

[0121] According to some embodiments of the present disclosure, an audio rendering device is also proposed, which includes a signal processing module as described herein, configured to process a response signal obtained from a sound signal from a sound source to a listening position, and a rendering module, configured to perform audio rendering based on the processed response signal, as Figure 2C shown. In particular, audio rendering can be implemented using various appropriate known rendering operations in the art. For example, various appropriate rendering signals can be obtained for rendering. As an example, for a more advanced scene information processor, a spatial room reverberation response of the scene may be generated, including but not limited to RIR (Room Impulse Response), ARIR (Ambisonics Room Impulse Response), BRIR (Binaural Room Impulse Response), MO-BRIR (Multi orientation Binaural RoomImpulse Response). For such information, a convolver can be added to this module to obtain the processed signal. Depending on the type of reverberation, the generated result may be an intermediate signal (ARIR), an omnidirectional signal (RIR), or a binaural signal (BRIR, MO-BRIR).

[0122] Specifically, according to an embodiment of the present disclosure, the processing of optimizing a signal based on the absolute perception characteristics of the signal as described above can also be implemented by a rendering module in an audio rendering device. That is to say, in the audio rendering device, for a response signal obtained from a sound signal from a sound source to a listening position, the signal processing module optimizes the response signal based on the relative perception characteristics of the signal to obtain a reduced number of response signals, and then the reduced number of response signals are subjected to a rendering process in the rendering module. Further, for the reduced number of response signals, signal processing based on the absolute perception characteristics of the signal according to an embodiment of the present disclosure is applied, and in particular, only signals whose absolute perception characteristics are higher than a specific threshold are used for reverberation calculation for audio rendering, for example, audio rendering is performed by convolution, which can further reduce the computational complexity, reduce the computational overhead, and improve the computational efficiency.

[0123] It should be noted that each module of the signal processing device and the audio rendering device as described above is only a logical module divided according to the specific functions it implements, rather than being used to limit the specific implementation manner. For example, it can be implemented in a software, hardware, or software-hardware combination manner. In actual implementation, the above-mentioned various units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). For example, an encoder, a decoder, etc. can be implemented using a chip (such as an integrated circuit module including a single wafer), a hardware component, or a complete product. In addition, the elements shown by dashed lines in the drawings indicate that these elements may exist, but do not need to actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.

[0124] In addition, optionally, the signal processing device and the audio rendering device may further include other components not shown, such as an interface, a memory, a communication unit, etc. As an example, the interface and / or the communication unit can be used to receive an input audio signal to be rendered, or a set of response signals, and can also output the finally generated audio signal to a playback device in a playback environment for playback. As an example, the memory can store various data, information, programs, etc. used in audio rendering and / or generated during the audio rendering process. The memory can include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), and a flash memory.

[0125] According to some embodiments of the present disclosure, a signal processing method for audio rendering is also proposed. Figure 2B Flowcharts of some embodiments of the signal processing method for audio rendering according to the present disclosure are shown. As Figure 2BAs shown, in step S210 (acquisition step), a set of response signals is acquired, the set of response signals including response signals obtained from a sound signal, where the sound signal is a signal received at a listening position. In step S220 (processing step), the response signals in the set of response signals are processed based on perceptual characteristics related to the response signals to obtain response signals suitable for audio rendering, where the number of response signals suitable for audio rendering is less than or equal to the number of response signals in the set of response signals.

[0126] According to some embodiments of the present disclosure, an audio rendering method is also proposed, which includes processing response signals obtained from sound signals from a sound source to a listening position by using the signal processing method described herein, and performing audio rendering based on the processed response signals, as Figure 2D shown.

[0127] Although not shown, the signal processing method for audio rendering according to the present disclosure may further include other steps to implement the impulse response sorting, psychoacoustic masking characteristic acquisition, comparison / judgment processing described above, which will not be described in detail here. It should be noted that the signal processing method, audio rendering method, and steps therein according to the present disclosure may be executed by any suitable device, such as a processor, integrated circuit, chip, etc., for example, may be executed by the aforementioned signal processing device and each module therein, and the method may also be implemented in a computer program, instruction, computer program medium, computer program product, etc.

[0128] Exemplary processing operations according to embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 4B A flowchart of an exemplary processing operation according to an embodiment of the present disclosure is shown, in which sound signal processing is performed based on both intensity-related information and signal masking condition information for audio rendering processing.

[0129] 1. For the impulse response set R, sort it according to the time in R to obtain the sorted set R t,s , where the subscript t represents time and s represents space.

[0130] 2. Recursively traverse the neighboring response set of the current response r t,s one by one from the time dimension Each r t,s includes three important pieces of data, such as time, spatial direction, and sound intensity. Here, the neighboring response set may be a set of responses within a specific time range including the current response, l represents the length l of the neighboring response set, which may indicate the time range, or the number of responses that the neighboring response set needs to include, etc.

[0131] 3. According to the adjacent response set Calculate the cosine set of the spatial vectors in the current set as the aforementioned masking condition response information

[0132]

[0133] Where and represent the vector representations of two impulse responses in the adjacent response set. The arrow here indicates the direction because each impulse response has a direction coordinate value in space, which is equivalent to a vector. The |r |, |r i |, |r j | respectively indicate the magnitudes of these two impulse responses, such as the magnitude of a vector in a specific coordinate system. In this way, the cosine set between every two responses in the obtained adjacent response set is obtained.

[0134] 4. According to the set and the spatial cosine threshold ζ T , determine whether to merge the responses and generate a new set R′ t,s

[0135]

[0136] Specifically, for each value in the set , compare it with a specific threshold, and in the case of being less than the threshold, merge the two impulse responses corresponding to this value in the set , such as the mean of the two impulse responses. It should be noted that other merging methods can also be used. For other cases, these two impulse responses can be retained. In this way, through merging, the impulse responses included in the impulse response set can be reduced to obtain a new set.

[0137] 5. According to the new set R′ t,s Calculate the sound pressure level of the corresponding frequency band of the response as the intensity-related information in the psychoacoustic perception characteristics. Here, the sound pressure level can be calculated for a channel, especially for a high-fidelity reverberation channel (Ambisonic channel).

[0138] Preferably, the sound pressure level can be calculated for an impulse response block, which is obtained by partitioning the new set. The block size can be set by various appropriate methods. In some embodiments, the block size can correspond to the size of the head-related transfer function (HRTF) used in audio rendering. The calculation of the sound pressure level is as follows:

[0139]

[0140] Preferably, the sound pressure where z0 represents the acoustic impedance the sum of the sound pressures of each frequency band in each block, P ref represents the reference sound pressure

[0141] 6. Calculate the set R′ t,s of the ARIR, and determine whether to perform the convolution calculation to obtain R according to the SPL calculated in the previous step arir

[0142]

[0143] The convolution operation here can be implemented in various ways known in the art. The selected hrtf function can be various appropriate functions known in the art, and will not be described in detail here. In this way, signals with high sound pressure intensity levels are retained and convolved to obtain the corresponding ARIR, while signals with low sound pressure levels do not need to be convolved, which can reduce the computational operation overhead and improve the computational efficiency

[0144] 7. According to R arir convert it to the corresponding R brir . The conversion operation here can adopt various conversion methods known in the art and will not be described in detail here

[0145] The conversion operation here can be performed by various appropriate conversion methods in the art and will not be described in detail here

[0146] The following will describe the advantageous technical effects achieved by the optimization process according to the embodiments of the present disclosure. Through this method, the number of calculated impulse responses and the computational complexity and time consumption of the binaural impulse response can be effectively reduced

[0147] Here, the spatial scene of sibenik and the ambisonic order of 3 are taken as an example for description. Among them, through spatio-temporal calculation, the ratio of the number of impulse responses that are blocked / filtered to the total number of impulse responses can be obtained. The calculation formula is

[0148]

[0149] where R m is the number of impulse responses that are blocked / filtered, R n is the total number of impulse responses, and p n is the ratio of the number of impulse responses that are blocked / filtered to the total number of impulse responses when the current number of impulse responses is n. Specifically, as the number of impulse responses increases, the number of impulse responses that are blocked / filtered also increases. When the range of impulse responses is [1000, 10000], the proportion of impulse responses that are blocked / filtered is [1%, 17.5%]

[0150] As another example, the ratio of the number of channels perceived to be below the absolute hearing threshold to the total number of channels can be obtained by calculating the absolute auditory threshold, and the calculation formula is

[0151]

[0152] where is the number of channels perceived to be below the absolute hearing threshold, is the total number of channels, and is the ratio of the number of channels perceived to be below the absolute hearing threshold to the total number of channels when the current number of impulse responses is i.

[0153] Specifically, as the number of impulse responses increases, the proportion of channels perceived to be below the absolute hearing threshold also increases. As an example, when the range of impulse responses is [1000, 10000], the proportion of channels perceived to be below the absolute threshold is [50%, 70%].

[0154] By statistically analyzing the time consumption of different high-fidelity reverberation orders with an impulse response of 1000, the performance ratio of the time consumption after optimized calculation to the time consumption of the original method can be obtained, and the calculation formula is as follows:

[0155]

[0156] where is the calculation time consumption of the original method when the order is n, is the calculation time consumption after spatio-temporal absolute threshold perception, and is the ratio of the saved time to the time consumed by the original method.

[0157] As an example, when the order of high-fidelity reverberation is in the range of [3, 7], the calculation time consumption of BRIR in the sibenik scenario can be saved by [30%, 50%].

[0158] Generally speaking, for the process of calculating the binaural room impulse response of late reverberation from the impulse response in the present disclosure, the calculation time consumption will be greatly reduced, thereby reducing the calculation overhead and improving the calculation efficiency.

[0159] Figure 5 The block diagram of some embodiments of the electronic device of the present disclosure is shown.

[0160] As Figure 5 shown, the electronic device 5 of this embodiment includes: a memory 51 and a processor 52 coupled to the memory 51. The processor 52 is configured to execute the reverberation duration estimation method or the audio signal rendering method in any one of the embodiments of the present disclosure based on the instructions stored in the memory 51.

[0161] Among them, the memory 51 can, for example, include a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, a database, and other programs.

[0162] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure can include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0163] Figure 6 A block diagram showing other embodiments of the electronic device of the present disclosure.

[0164] As Figure 6 shown, the electronic device can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0165] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device shown has various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0166] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0167] In some embodiments, a chip is further provided, including: at least one processor and an interface, the interface is used to provide computer execution instructions for at least one processor, and at least one processor is used to execute the computer execution instructions to implement the reverberation time estimation method or the audio signal rendering method of any one of the above embodiments.

[0168] Figure 7 A block diagram showing some embodiments of the chip of the present disclosure.

[0169] As Figure 7 shown, the processor 70 of the chip is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the HostCPU. The core part of the processor 70 is an arithmetic circuit, and the controller 704 controls the arithmetic circuit 703 to extract data from the memory (weight memory or input memory) and perform arithmetic operations.

[0170] In some embodiments, the arithmetic circuit 703 internally includes a plurality of processing units (Process Engine, PE). In some embodiments, the arithmetic circuit 703 is a two-dimensional systolic array. The arithmetic circuit 703 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some embodiments, the arithmetic circuit 703 is a general matrix processor.

[0171] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 702 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 701 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 708.

[0172] The vector calculation unit 707 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc.

[0173] In some embodiments, the vector calculation unit 707 can store the vector of the processed output into the unified buffer 706. For example, the vector calculation unit 707 can apply a non-linear function to the output of the arithmetic circuit 703, such as a vector of accumulated values, to generate activation values. In some embodiments, the vector calculation unit 707 generates normalized values, combined values, or both. In some embodiments, the vector of the processed output can be used as an activation input to the arithmetic circuit 703, for example, for use in subsequent layers in a neural network.

[0174] The unified memory 706 is used to store input data and output data.

[0175] The storage unit access controller 705 (Direct Memory Access Controller, DMAC) transfers the input data in the external memory to the input memory 701 and / or the unified memory 706, stores the weight data in the external memory into the weight memory 702, and stores the data in the unified memory 706 into the external memory.

[0176] The Bus Interface Unit (BIU) 510 is used to enable interaction between the main CPU, the DMAC, and the instruction fetch memory 709 through the bus.

[0177] The instruction fetch buffer 709 connected to the controller 704 is used to store the instructions used by the controller 704;

[0178] The controller 704 is used to call the instructions cached in the instruction fetch memory 709 to control the working process of the arithmetic accelerator.

[0179] Generally, the unified memory 706, the input memory 701, the weight memory 702, and the instruction fetch memory 709 are all on-chip memories, and the external memory is a memory external to the NPU. The external memory can be a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), a High Bandwidth Memory (HBM), or other readable and writable memories.

[0180] In some embodiments, a computer program is further provided, including: instructions that, when executed by a processor, cause the processor to execute the reverberation duration estimation method or the audio signal rendering method of any one of the above embodiments.

[0181] Those skilled in the art should understand that the present disclosure may be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Moreover, the present disclosure may be implemented in the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0182] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A signal processing method for audio rendering, comprising: obtaining a set of response signals, the set of response signals including response signals obtained from sound signals, where the sound signals are signals received at a listening position; and processing the response signals in the set of response signals based on perceptual characteristics related to the response signals to obtain response signals suitable for audio rendering, where the number of response signals suitable for audio rendering is less than or equal to the number of response signals in the set of response signals, wherein processing the response signals in the set of response signals based on perceptual characteristics related to the response signals includes: merging or removing response signals in the set of response signals whose related perceptual characteristics do not meet the perceptual requirements, where the perceptual requirements include information related to the conditions or requirements that need to be met for the sound corresponding to the response signal to be effectively perceived.

2. The signal processing method according to claim 1, wherein, the perceptual characteristics include relative perceptual characteristics between response signals, and processing the response signals in the set of response signals based on perceptual characteristics related to the response signals includes: determining whether the relative perceptual characteristics between the response signals in the set of response signals meet the perceptual requirements, and merging or removing the response signals when it is determined that the relative perceptual characteristics between the response signals in the set of response signals do not meet the perceptual requirements.

3. The signal processing method according to claim 1, wherein, the perceptual characteristics include relative perceptual characteristics between response signals, and processing the response signals in the set of response signals based on perceptual characteristics related to the response signals includes: obtaining a set of adjacent response signals in the set of response signals, determining whether the relative perceptual characteristics between the response signals in the set of adjacent response signals meet the perceptual requirements, and merging or removing the response signals when it is determined that the relative perceptual characteristics between the response signals in the set of adjacent response signals do not meet the perceptual requirements.

4. The signal processing method according to claim 2, wherein, the relative perceptual characteristics and the perceptual requirements are related to the mutual masking situation between response signals, the determining whether the relative perceptual characteristics between the response signals in the set of response signals meet the perceptual requirements includes: obtaining information related to the mutual masking situation between each two response signals in the set of response signals, and determining the magnitude of the mutual masking situation between each two response signals in the set of response signals; the merging or removing the response signals when it is determined that the relative perceptual characteristics between the response signals in the set of response signals do not meet the perceptual requirements includes: merging the two response signals when the mutual masking situation between the two response signals in the set of response signals is large to obtain an updated response signal.

5. The signal processing method according to claim 3, wherein, the relative perceptual characteristics and the perceptual requirements are related to the mutual masking situation between response signals, Determining whether the relative perception characteristics between the response signals in the determined proximity response signal set meet the perception requirements includes: Obtaining relevant information on the mutual masking situation between every two response signals in the proximity response signal set, and determining the magnitude of the mutual masking situation between every two response signals in the proximity response set; When it is determined that the relative perception characteristics between the response signals in the proximity response signal set do not meet the perception requirements, merging or removing the response signals includes: When the mutual masking situation between two response signals in the proximity response set is large, merging the two response signals to obtain an updated response signal.

6. The signal processing method according to claim 4 or 5, wherein, The relevant information on the mutual masking situation between the two response signals includes the spatial interval information between the two response signals, and a spatial interval less than a specific interval threshold between the two response signals indicates a large mutual masking situation between the two response signals.

7. The signal processing method according to claim 6, wherein, The spatial interval information between the two response signals is represented by a statistic of the spatial vector between the two response signals.

8. The signal processing method according to claim 6, wherein, The spatial interval information between the two response signals is determined based on at least one of the time information, spatial information, and intensity information of the two response signals.

9. The signal processing method according to claim 3 or 5, wherein, The proximity response signal set in the response signal set includes response signals whose time interval, spatial interval, or frequency domain interval between each other is less than a second proximity threshold.

10. The signal processing method according to claim 2 or 3, wherein, The relative perception characteristics between the response signals and the perception requirements are related to the proximity between the response signals, Determining whether the relative perception characteristics between the response signals in the response signal set meet the perception requirements includes: For each response signal in the response signal set, determining whether the proximity between the response signal and any other response signal in the response signal set is less than a first proximity threshold, When it is determined that the relative perception characteristics between the response signals in the response signal set do not meet the perception requirements, merging or removing the response signals includes: When it is determined that the proximity between two response signals is less than the first proximity threshold, merging the two response signals.

11. The signal processing method according to claim 10, wherein, The proximity between the response signals includes at least one of temporal proximity, spatial proximity, and frequency domain proximity.

12. The signal processing method according to claim 1, wherein, The method further includes: Before processing the response signals in the response signal set based on the perception characteristics related to the response signals, performing time sorting or spatial sorting on the response signals in the response signal set.

13. The signal processing method according to claim 1, wherein, The merging includes performing mathematical statistics on the attribute information of the response signals as the attribute information of the merged response signals, where the attribute information of the response signals includes at least one of time information, spatial information, and sound intensity information.

14. The signal processing method according to claim 13, wherein, the mathematical statistics includes averaging the attribute information of the response signals.

15. The signal processing method according to claim 1, wherein, the perceptual characteristics related to the response signals include the perceptual intensity characteristics of the response signals themselves, and processing the response signals in the response signal set based on the perceptual characteristics related to the response signals includes: not using the response signal for audio rendering when the perceptual intensity characteristic of the response signal itself is lower than a specific absolute perceptual threshold.

16. The signal processing method according to claim 15, wherein, the perceptual intensity characteristic of the response signal itself includes at least one of the sound pressure level of the sound signal corresponding to the loudness signal and the ratio of the channel-based sound pressure level of the sound signal corresponding to the loudness signal to the reference sound pressure level.

17. The signal processing method according to claim 1, wherein, the response signals include response signals converted from at least one of the direct sound signals and the reflected sound signals received at the listening position.

18. An audio rendering method, including: processing a response signal set obtained from sound signals from a sound source to a listening position by using the signal processing method according to any one of claims 1 - 17; and performing audio rendering based on the processed response signal set.

19. A signal processing apparatus for audio rendering, including: an acquisition module configured to acquire a response signal set, the response signal set including response signals obtained from sound signals, where the sound signals are signals received at the listening position; and a processing module configured to process the response signals in the response signal set based on the perceptual characteristics related to the response signals to obtain response signals suitable for audio rendering, where the number of the response signals suitable for audio rendering is less than or equal to the number of the response signals in the response signal set, wherein the processing module is further configured to: merge or remove the response signals in the response signal set whose related perceptual characteristics do not meet the perceptual requirements, where the perceptual requirements include information related to the conditions or requirements that need to be met for the sound corresponding to the response signal to be effectively perceived.

20. An audio rendering apparatus, including: the signal processing apparatus according to claim 19, configured to process a response signal set obtained from sound signals from a sound source to a listening position; and a rendering module configured to perform audio rendering based on the processed response signal set.

21. A chip, including: At least one processor and an interface, the interface being configured to provide computer-executable instructions for the at least one processor, the at least one processor being configured to execute the computer-executable instructions to implement the signal processing method according to any one of claims 1-17 or the audio rendering method according to claim 18.

22. An electronic device, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute, based on instructions stored in the memory device, the signal processing method according to any one of claims 1-17 or the audio rendering method according to claim 18.

23. A non-transitory computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the signal processing method according to any one of claims 1-17 or the audio rendering method according to claim 18.

24. A computer program product comprising instructions that, when executed by a processor, cause the processor to execute the signal processing method according to any one of claims 1-17 or the audio rendering method according to claim 18.

Citation Information

Patent Citations

  • Parametric wave field coding for real-time sound propagation for dynamic sources

    CN106465037A