Audio processing method and apparatus
By obtaining the HRTF corresponding to the positions of the left and right ears, the audio signal processing method was optimized, solving the problem of low audio signal quality in virtual reality devices and achieving optimal audio signal transmission effect.
Patent Information
- Application Number
- CN202211471838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-08-20
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2038-08-20
AI Technical Summary
Existing binaural playback methods based on multi-channel headphones in virtual reality devices cannot optimally transmit audio signals to the listener's ears, resulting in poor audio signal quality.
By acquiring M first HRTFs centered on the left ear position and N second HRTFs centered on the right ear position, M first audio signals and N second audio signals are processed respectively to optimize the audio signal quality transmitted to the left and right ears.
It improves the quality of the audio signal output from the audio signal receiver, making the signal transmitted to the left and right ears optimal.
Smart Images

Figure CN115866505B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 201810950088.1 and the original application date is August 20, 2018. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to sound processing technology, and more particularly to an audio processing method and apparatus. Background Technology
[0003] With the rapid development of high-performance computing and signal processing technologies, virtual reality (VR) technology has received increasing attention. An immersive VR system requires not only stunning visual effects but also realistic auditory effects; the fusion of audio and video greatly enhances the VR experience. The core of VR audio is 3D audio technology. Currently, there are various playback methods for 3D audio (such as multi-channel and object-based methods), but the most commonly used method in existing VR devices is still binaural playback based on multi-channel headphones.
[0004] Binaural reproduction based on multi-channel headphones is mainly achieved using the Head Related Transfer Function (HRTF). HRTF characterizes the effects of scattering, reflection, and refraction by organs such as the head, trunk, and auricle when sound waves from a sound source propagate to the ear canal. Assuming the sound source is at a certain location, the audio signal receiver selects the HRTF corresponding to the distance from that location to the center of the listener's head and convolves it with the audio signal emitted by that sound source. The sweet spot of the processed audio signal is the center of the listener's head; that is, the sound signal transmitted to the center of the listener's head is the optimal audio signal.
[0005] However, the position of the listener's ears is not the center of the listener's head. Therefore, the sound signal transmitted to the listener's ears after the above processing is not the optimal audio signal, that is, the quality of the audio signal output by the frequency signal receiver is not high. Summary of the Invention
[0006] This application provides an audio processing method and apparatus that improves the quality of the audio signal output from the audio signal receiver.
[0007] In a first aspect, embodiments of this application provide an audio processing method, including:
[0008] The system acquires M first audio signals after the audio signal to be processed has been processed by M first virtual speakers, and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals, and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers.
[0009] Obtain M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0010] A first target audio signal is obtained based on the M first audio signals and the M first HRTFs; a second target audio signal is obtained based on the N second audio signals and the N second HRTFs.
[0011] In the scheme, a first target audio signal is obtained by using M first audio signals and M first HRTFs centered on the left ear position, so that the signal transmitted to the left ear position is optimal. Similarly, a second target audio signal is obtained by using N second audio signals and N second HRTFs centered on the right ear position, so that the signal transmitted to the right ear position is optimal. Therefore, the quality of the audio signal output by the audio signal receiver is improved.
[0012] Optionally, the "obtaining the first target audio signal based on the M first audio signals and the M first HRTFs" in the above scheme includes: performing convolution processing on the M first audio signals with their corresponding first HRTFs to obtain M first convolutional audio signals; and obtaining the first target audio signal based on the M first convolutional audio signals.
[0013] Optionally, the "obtaining the second target audio signal based on the N second audio signals and the N second HRTFs" in the above scheme includes: convolving the N second audio signals with the corresponding second HRTFs to obtain N second convolutional audio signals; and obtaining the second target audio signal based on the N second convolutional audio signals.
[0014] Specifically, "obtaining M first HRTFs" can be implemented in the following ways:
[0015] In one implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; obtaining M first HRTFs includes:
[0016] Obtain the M first positions of the M first virtual speakers relative to the current left ear position;
[0017] Based on the M first positions and the corresponding relationship, the M HRTFs corresponding to the M first positions are determined as the M first HRTFs.
[0018] In this embodiment, the M first HRTFs corresponding to the M virtual speakers are the M HRTFs centered on the left ear position obtained by actual measurement. These M first HRTFs best represent the HRTFs corresponding to the transmission of the M first audio signals to the current left ear position, so that the signal transmitted to the left ear position is optimal.
[0019] In another implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; the acquisition of N second HRTFs includes:
[0020] Obtain the N second positions of the N second virtual speakers relative to the current right ear position;
[0021] Based on the N second positions and the corresponding relationship, the N HRTFs corresponding to the N second positions are determined as the N second HRTFs.
[0022] In this embodiment, the M first HRTFs are obtained by converting the HRTF centered on the head, which is highly efficient in obtaining the first HRTFs.
[0023] In another implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; obtaining the M first HRTFs includes:
[0024] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0025] M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0026] Based on the M fourth positions and the corresponding relationship, the M HRTFs corresponding to the M fourth positions are determined to be the M first HRTFs.
[0027] In this embodiment, the M first HRTFs are obtained by converting the HRTF centered on the head, and when obtaining the fourth position mentioned above, the size of the current listener's head is not considered, which further improves the efficiency of obtaining the first HRTFs.
[0028] Specifically, "obtaining N second HRTFs" can be implemented in the following ways:
[0029] In another implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; obtaining N second HRTFs includes:
[0030] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0031] N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position include the same pitch angle and the same distance. The sum of the azimuth angle included in the sixth position and the second value is the second azimuth angle included in the corresponding fifth position. The second value is the difference between the third included angle and the second included angle. The second included angle is the angle between the second straight line and the first surface. The third included angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0032] Based on the N sixth positions and the corresponding relationship, the N HRTFs corresponding to the N sixth positions are determined as the N second HRTFs.
[0033] In this embodiment, the N second HRTFs are the N HRTFs centered on the right ear position obtained by actual measurement. The N second HRTFs obtained best represent the HRTFs corresponding to the current right ear position of the N second audio signals transmitted to the listener, so that the signal transmitted to the right ear position is optimal.
[0034] In another implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; the acquisition of M first HRTFs includes
[0035] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0036] M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position.
[0037] Based on the M seventh positions and the corresponding relationship, the M HRTFs corresponding to the M seventh positions are determined to be the M first HRTFs.
[0038] In this embodiment, the N second HRTFs are obtained by converting the HRTF centered on the head, which is highly efficient in obtaining the second HRTFs.
[0039] In another implementation, a plurality of preset locations and a plurality of HRTFs are pre-stored; the acquisition of N second HRTFs includes:
[0040] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0041] N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position.
[0042] Based on the N eighth positions and the corresponding relationship, the N HRTFs corresponding to the N eighth positions are determined as the N second HRTFs.
[0043] In this embodiment, the N second HRTFs are obtained by converting the HRTF centered on the head, and when obtaining the eighth position mentioned above, the size of the current listener's head is not considered, which further improves the efficiency of obtaining the second HRTFs.
[0044] In one possible design, before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers, the process further includes:
[0045] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the M first virtual speakers;
[0046] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0047] The acquisition of the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers includes:
[0048] The audio signal to be processed is processed according to the M tenth positions to obtain the M first audio signals.
[0049] In this embodiment, a target virtual speaker group is virtually mapped, and M first virtual speakers corresponding to the left ear are obtained based on the target virtual speaker group. The overall efficiency of mapping virtual speakers is high.
[0050] In one possible design, M = N, and before the N second audio signals processed by the N second virtual speakers, the following is also included:
[0051] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers;
[0052] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual loudspeakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0053] The acquisition of N second audio signals after the audio signal to be processed has been processed by N second virtual speakers includes:
[0054] The audio signal to be processed is processed according to the N eleventh positions to obtain the N second audio signals.
[0055] In this embodiment, a target virtual speaker group is mapped, and N second virtual speakers corresponding to the right ear are obtained based on the target virtual speaker group. The overall efficiency of mapping virtual speakers is high.
[0056] In one possible design, the M first virtual speakers are speakers in a first speaker group, and the N second virtual speakers are speakers in a second speaker group, wherein the first speaker group and the second speaker group are two independent speaker groups; or...
[0057] The M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group. The first speaker group and the second speaker group are the same speaker group, and M = N.
[0058] Secondly, embodiments of this application provide an audio processing apparatus, including:
[0059] The processing module is used to acquire M first audio signals after the audio signal to be processed has been processed by M first virtual speakers and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers;
[0060] The acquisition module is used to acquire M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0061] The acquisition module is further configured to acquire a first target audio signal based on the M first audio signals and the M first HRTFs; and to acquire a second target audio signal based on the N second audio signals and the N second HRTFs.
[0062] In one possible design, the acquisition module is specifically used for:
[0063] The M first audio signals are convolved with their corresponding first HRTFs to obtain M first convolved audio signals.
[0064] The first target audio signal is obtained based on the M first convolutional audio signals.
[0065] In one possible design, the acquisition module is specifically used for:
[0066] The N second audio signals are convolved with their corresponding second HRTF signals to obtain N second convolved audio signals;
[0067] The second target audio signal is obtained based on the N second convolutional audio signals.
[0068] In one possible design, the acquisition module is specifically used for:
[0069] Obtain the M first positions of the M first virtual speakers relative to the current left ear position;
[0070] Based on the M first positions and their corresponding relationships, the M HRTFs corresponding to the M first positions are determined as the M first HRTFs. The corresponding relationships are the pre-stored correspondences between multiple preset positions and multiple HRTFs.
[0071] In one possible design, the acquisition module is specifically used for:
[0072] Obtain the N second positions of the N second virtual speakers relative to the current right ear position;
[0073] Based on the N second positions and their corresponding relationships, the N HRTFs corresponding to the N second positions are determined as the N second HRTFs. The corresponding relationships are multiple preset positions and multiple HRTFs stored in advance.
[0074] In one possible design, the acquisition module is specifically used for:
[0075] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0076] M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0077] Based on the M fourth positions and their corresponding relationships, the M HRTFs corresponding to the M fourth positions are determined as the M first HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
[0078] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module is specifically used for:
[0079] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0080] N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position include the same pitch angle and the same distance. The sum of the azimuth angle included in the sixth position and the second value is the second azimuth angle included in the corresponding fifth position. The second value is the difference between the third included angle and the second included angle. The second included angle is the angle between the second straight line and the first surface. The third included angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0081] Based on the N sixth positions and their corresponding relationships, the N HRTFs corresponding to the N sixth positions are determined as the N second HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
[0082] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module is specifically used for:
[0083] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0084] M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position.
[0085] Based on the M seventh positions and their corresponding relationships, the M HRTFs corresponding to the M seventh positions are determined as the M first HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
[0086] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module is specifically used for:
[0087] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0088] N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position.
[0089] Based on the N eighth positions and their corresponding relationships, the N HRTFs corresponding to the N eighth positions are determined as the N second HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
[0090] In one possible design, the acquisition module is further configured to: before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers,
[0091] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the M first virtual speakers;
[0092] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0093] The processing module is specifically used to: process the audio signal to be processed according to the M tenth positions to obtain the M first audio signals.
[0094] In one possible design, M = N, the acquisition module is further configured to: before the N second audio signals processed by the N second virtual speakers:
[0095] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers;
[0096] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual loudspeakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0097] The processing module is specifically used to: process the audio signal to be processed according to the N eleventh positions to obtain the N second audio signals.
[0098] In one possible design, the M first virtual speakers are speakers in a first speaker group, and the N second virtual speakers are speakers in a second speaker group, wherein the first speaker group and the second speaker group are two independent speaker groups; or...
[0099] The M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group. The first speaker group and the second speaker group are the same speaker group, and M = N.
[0100] Thirdly, embodiments of this application provide an audio processing apparatus, which includes a processor;
[0101] The processor is configured to couple with a memory, read and execute instructions in the memory to implement the method as described in any of the first aspects.
[0102] In one possible design, the memory is also included.
[0103] Fourthly, embodiments of this application provide a readable storage medium on which a computer program is stored; when the computer program is executed, it implements the method described in any of the first aspects.
[0104] Fifthly, embodiments of this application provide a computer program product, which, when executed, implements the method described in any of the first aspects.
[0105] This application obtains a first target audio signal transmitted to the left ear based on M first audio signals and M first HRTFs centered on the left ear position, thus optimizing the signal transmitted to the left ear position; and obtains a second target audio signal transmitted to the right ear based on N second audio signals and N second HRTFs centered on the right ear position, thus optimizing the signal transmitted to the right ear position. Therefore, it improves the quality of the audio signal output by the audio signal receiver. Attached Figure Description
[0106] Figure 1 This is a schematic diagram of the structure of an audio signal system provided in an embodiment of this application;
[0107] Figure 2 System architecture diagram provided for embodiments of this application;
[0108] Figure 3 A structural block diagram of an audio signal receiving device is provided for embodiments of this application;
[0109] Figure 4 Flowchart of the audio processing method provided in the embodiments of this application Figure 1 ;
[0110] Figure 5 A measurement scenario diagram with the head center as the center of HRTF measurement provided for embodiments of this application;
[0111] Figure 6 Flowchart of the audio processing method provided in the embodiments of this application Figure 2 ;
[0112] Figure 7 This application provides a measurement scenario diagram with the center of the left ear position as the center for measuring HRTF;
[0113] Figure 8 Flowchart of the audio processing method provided in the embodiments of this application Figure 3 ;
[0114] Figure 9 Flowchart of the audio processing method provided in the embodiments of this application Figure 4 ;
[0115] Figure 10 Flowchart of the audio processing method provided in the embodiments of this application Figure 5 ;
[0116] Figure 11 This application provides a measurement scenario diagram with the center of the right ear location as the center for measuring HRTF;
[0117] Figure 12 Flowchart of the audio processing method provided in the embodiments of this application Figure 6 ;
[0118] Figure 13 Flowchart of the audio processing method provided in the embodiments of this application Figure 7 ;
[0119] Figure 14 Flowchart of the audio processing method provided in the embodiments of this application Figure 8 ;
[0120] Figure 15 Flowchart of the audio processing method provided in the embodiments of this application Figure 9 ;
[0121] Figure 16 This is a spectrum showing the difference between the rendered spectrum of the signal corresponding to the left ear position in the existing technology and the theoretical spectrum corresponding to the left ear position.
[0122] Figure 17 The difference spectrum between the rendered spectrum of the rendered signal corresponding to the right ear position in the existing technology and the theoretical spectrum corresponding to the right ear position;
[0123] Figure 18 A spectrum showing the difference between the rendered spectrum of the rendered signal corresponding to the left ear position and the theoretical spectrum corresponding to the left ear position in the method provided for implementation of this application;
[0124] Figure 19 A spectrum showing the difference between the rendered spectrum of the rendered signal corresponding to the right ear position and the theoretical spectrum corresponding to the right ear position in the method provided for implementation of this application;
[0125] Figure 20 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application. Detailed Implementation
[0126] First, the relevant technical terms involved in this application will be explained.
[0127] Head Related Transfer Function (HRTF): Sound waves emitted from a sound source are scattered by the head, auricle, and torso before reaching both ears. This physical process can be viewed as a linear, time-invariant acoustic filtering system, and its characteristics can be described by the HRTF. In other words, the HRTF describes the transmission process of sound waves from the sound source to the ears. A more intuitive explanation is: if the audio signal emitted by the sound source is X, and the corresponding audio signal Y after X reaches a predetermined location is Y, then X*Z = Y (X convolved with Z equals Y), where Z is the HRTF.
[0128] In this embodiment, the preset positions in the correspondence between multiple preset positions and multiple HRTFs can be positions relative to the left ear, in which case the multiple HRTFs are multiple HRTFs centered on the left ear; the preset positions in this embodiment can also be positions relative to the right ear, in which case the multiple HRTFs are multiple HRTFs centered on the right ear; the preset positions in this embodiment can also be positions relative to the center of the head, in which case the multiple HRTFs are multiple HRTFs centered on the center of the head.
[0129] Figure 1 This is a schematic diagram of the structure of an audio signal system provided in an embodiment of this application. The audio signal system includes an audio signal transmitter 11 and an audio signal receiver 12.
[0130] The audio signal transmitter 11 is used to collect and encode the signal emitted by the sound source to obtain an audio signal encoded bitstream. After obtaining the audio signal encoded bitstream, the audio signal receiver 12 decodes and renders the audio signal encoded bitstream to obtain a rendered audio signal.
[0131] Optionally, the audio signal transmitter 11 and the audio signal receiver 12 can be connected by wired or wireless means.
[0132] Figure 2 This is a system architecture diagram provided for an embodiment of this application. Figure 2 As shown, the system architecture includes mobile terminal 130 and mobile terminal 140; mobile terminal 130 can be an audio signal transmitter, and mobile terminal 140 can be an audio signal receiver.
[0133] Among them, mobile terminal 130 and mobile terminal 140 can be independent electronic devices with audio signal processing capabilities, such as mobile phones, wearable devices, virtual reality (VR) devices, or augmented reality (AR) devices, etc., and mobile terminal 130 and mobile terminal 140 are connected to each other via wireless or wired networks.
[0134] Optionally, the mobile terminal 130 may include a data acquisition component 131, an encoding component 110, and a channel coding component 132, wherein the data acquisition component 131 is connected to the encoding component 110, and the encoding component 110 is connected to the encoding component 132.
[0135] Optionally, the mobile terminal 140 may include an audio playback component 141, a decoding and rendering component 120, and a channel decoding component 142, wherein the audio playback component 141 is connected to the decoding and rendering component 120, and the decoding and rendering component 120 is connected to the channel decoding component 142.
[0136] After the mobile terminal 130 acquires the audio signal through the acquisition component 131, it encodes the audio signal through the encoding component 110 to obtain the audio signal encoded bitstream; then, it encodes the audio signal encoded bitstream through the channel coding component 132 to obtain the transmission signal.
[0137] Mobile terminal 130 transmits the signal to mobile terminal 140 via a wireless or wired network.
[0138] After receiving the transmission signal, mobile terminal 140 performs channel decoding on the transmission signal through channel decoding component 142 to obtain an audio signal encoded bitstream; it then decodes the audio signal encoded bitstream through decoding and rendering component 120 to obtain an audio signal to be processed, and renders the audio signal to be processed to obtain a rendered audio signal; finally, it plays the rendered audio signal through audio playback component. It is understood that mobile terminal 130 may also include the components included in mobile terminal 140, and vice versa.
[0139] In addition, the mobile terminal 140 may also include an audio playback component, a decoding component, a rendering component, and a channel decoding component, wherein the channel decoding component is connected to the decoding component, the decoding component is connected to the rendering component, and the rendering component is connected to the audio playback component. When the mobile terminal 140 receives the transmission signal, it decodes the transmission signal using the channel decoding component to obtain an audio signal encoded bitstream; it decodes the audio signal encoded bitstream using the decoding component to obtain the audio signal to be processed; the rendering component renders the audio signal to be processed to obtain the rendered audio signal; and the audio playback component plays the rendered audio signal.
[0140] Figure 3 A structural block diagram of an audio signal receiving device is provided for embodiments of this application; see also Figure 3 The audio signal receiving device 20 of this application embodiment may include: at least one processor 21, a memory 22, at least one communication bus 23, a receiver 24, and a transmitter 25. The communication bus 203 is used to establish communication between the processor 21, the memory 22, the receiver 24, and the transmitter 25. The processor 21 may include a signal decoding component 211, a decoding component 212, and a rendering component 213.
[0141] Specifically, the memory 22 can be any one or any combination of the following: solid state drives (SSDs), hard disk drives, magnetic disks, disk arrays, and other storage media that can provide instructions and data to the processor 201.
[0142] The memory 22 is used to store the following data: the correspondence between multiple preset positions and multiple HRTFs: (1) multiple positions relative to the left ear position, and the HRTF centered on the left ear position corresponding to each position relative to the left ear position; (2) multiple positions relative to the right ear position, and the HRTF centered on the right ear position corresponding to each position relative to the right ear position; (3) multiple positions relative to the center of the head, and the HRTF centered on the center of the head corresponding to each position relative to the center of the head.
[0143] Optionally, memory 22 may also be used to store elements such as operating system and application modules.
[0144] The operating system contains various system programs used to implement basic business operations and handle hardware-based tasks. The application module contains various applications used to implement various application business operations.
[0145] Processor 21 may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. A general-purpose processor may be a microprocessor, or it may be any conventional processor.
[0146] Receiver 24 is used to receive audio signals from the audio signal transmitting device.
[0147] The processor can execute the following steps by calling the program or instructions and data stored in memory 22: performing channel decoding on the received audio signal to obtain the audio signal encoded bitstream (this step can be implemented by the processor's channel decoding component), and then further decoding the audio signal encoded bitstream (this step can be implemented by the processor's decoding component) to obtain the audio signal to be processed.
[0148] After receiving the signal to be processed, the processor 21 is configured to: acquire M first audio signals after the audio signal to be processed has been processed by M first virtual speakers and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals, and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers;
[0149] Obtain M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0150] A first target audio signal is obtained based on the M first audio signals and the M first HRTFs; a second target audio signal is obtained based on the N second audio signals and the N second HRTFs.
[0151] Wherein, the M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group, and the first speaker group and the second speaker group are two independent speaker groups; or, the M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group, and the first speaker group and the second speaker group are the same speaker group, where M = N.
[0152] The processor 21 is specifically configured to: perform convolution processing on the M first audio signals with their corresponding first HRTFs to obtain M first convolutional audio signals; and obtain the first target audio signal based on the M first convolutional audio signals.
[0153] The processor 21 is further configured to: convolve the N second audio signals with their corresponding second HRTFs to obtain N second convolutional audio signals; and obtain the second target audio signal based on the N second convolutional audio signals.
[0154] The processor 21 is further specifically configured to: obtain M first positions of the M first virtual speakers relative to the current left ear position; determine the M HRTFs corresponding to the M first positions as the M first HRTFs based on the M first positions and the first correspondence stored in the memory 22; the first correspondence includes: multiple positions relative to the left ear position, and the HRTF centered on the left ear position corresponding to each position relative to the left ear position.
[0155] The processor 21 is further specifically configured to: obtain N second positions of the N second virtual speakers relative to the current right ear position; determine N HRTFs corresponding to the N second positions as the N second HRTFs based on the N second positions and the second correspondence stored in the memory 22; the second correspondence includes: multiple positions relative to the right ear position, and HRTFs centered on the right ear position corresponding to each position relative to the right ear position.
[0156] The processor 21 is further configured to: obtain M third positions of the M first virtual speakers relative to the current head center; the third positions include a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0157] M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0158] Based on the M fourth positions and the third correspondence stored in memory 22, the M HRTFs corresponding to the M fourth positions are determined to be the M first HRTFs; the third correspondence includes: multiple positions relative to the head center, and the HRTF centered on the head center corresponding to each position relative to the head center.
[0159] The processor 21 is further specifically configured to: obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes a second azimuth angle and a second pitch angle of the second virtual speaker relative to the current head center, and a second distance between the current head center and the second virtual speaker;
[0160] N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position include the same pitch angle and the same distance. The sum of the azimuth angle included in the sixth position and the second value is the second azimuth angle included in the corresponding fifth position. The second value is the difference between the third included angle and the second included angle. The second included angle is the angle between the second straight line and the first surface. The third included angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0161] Based on the N sixth positions and the aforementioned third correspondence, the N HRTFs corresponding to the N sixth positions are determined to be the N second HRTFs.
[0162] The processor 21 is further configured to: obtain M third positions of the M first virtual speakers relative to the current head center; the third positions include a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0163] M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position.
[0164] Based on the M seventh positions and the aforementioned third correspondence, the M HRTFs corresponding to the M seventh positions are determined to be the M first HRTFs.
[0165] The processor 21 is further specifically configured to: obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes a second azimuth angle and a second pitch angle of the second virtual speaker relative to the current head center, and a second distance between the current head center and the second virtual speaker;
[0166] N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position.
[0167] Based on the N eighth positions and the aforementioned third correspondence, the N HRTFs corresponding to the N eighth positions are determined to be the N second HRTFs.
[0168] The processor 21 is further configured to: before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers, acquire a target virtual speaker group, wherein the target virtual speaker group includes M target virtual speakers and the M target virtual speakers correspond one-to-one with the M first virtual speakers;
[0169] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0170] The processor 21 is specifically used to: process the audio signal to be processed according to the M tenth positions to obtain the M first audio signals.
[0171] Processor 21 is further configured to: acquire a target virtual speaker group before the N second audio signals processed by the N second virtual speakers to be processed, wherein the target virtual speaker group includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers; M = N.
[0172] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual loudspeakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0173] The processor 21 is specifically used to: process the audio signal to be processed according to the N eleventh positions to obtain the N second audio signals.
[0174] It is understandable that the various methods after the processor 21 receives the signal to be processed can be executed by the rendering component in the processor.
[0175] The audio signal receiving device of this embodiment obtains a first target audio signal to be transmitted to the left ear based on M first audio signals and M first HRTFs centered on the left ear position, thereby optimizing the signal transmitted to the left ear position. It also obtains a second target audio signal to be transmitted to the right ear based on N second audio signals and N second HRTFs centered on the right ear position, thereby optimizing the signal transmitted to the right ear position. Therefore, the quality of the audio signal output by the audio signal receiving end is improved.
[0176] The audio processing method involved in this application will be described below using specific embodiments. The execution subject of each of the following embodiments is an audio signal receiving end, such as... Figure 2 The mobile terminal 140 shown is shown.
[0177] Figure 4 Flowchart of the audio processing method provided in the embodiments of this application Figure 1 See also Figure 4 The method in this embodiment includes:
[0178] Step S101: Obtain M first audio signals after the audio signal to be processed has been processed by M first virtual speakers and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals, and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers;
[0179] Step S102: Obtain M first HRTFs and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0180] Step S103: Obtain the first target audio signal based on M first audio signals and M first HRTFs; obtain the second target audio signal based on N second audio signals and N second HRTFs.
[0181] Specifically, the method in this embodiment can be the method executed by the mobile terminal 140 described above. The encoding end acquires the stereo signal emitted by the sound source. The encoding component of the encoding end encodes the stereo signal emitted by the sound source to obtain an encoded signal. The encoded signal is transmitted wirelessly or via a wired network to the audio signal receiving end. The audio signal receiving end decodes the encoded signal, and the decoded signal is the audio signal to be processed in this embodiment. That is, the audio signal to be processed in this embodiment can be the signal decoded by the decoding component in the processor, or... Figure 2 The signal obtained by the decoding and rendering component 120 or the decoding component in the mobile terminal 140.
[0182] It is understandable that if the standard used when processing audio signals is Ambisonic, then the encoded signal obtained at the encoding end will be a standard Ambisonic signal. Correspondingly, the signal decoded at the audio signal receiving end will also be an Ambisonic signal, such as an Ambisonic B-format signal. Ambisonic signals include first-order Ambisonics (FOA) and high-order Ambisonics.
[0183] The following description uses an Ambisonic B-format audio signal decoded by the audio signal receiver as an example to illustrate this embodiment.
[0184] Specifically, in step S101, M first virtual speakers can form a first virtual speaker group, and N second virtual speakers can form a second virtual speaker group. The first virtual speaker group and the second virtual speaker group can be the same virtual speaker group or different virtual speaker groups. If the first virtual speaker group and the second virtual speaker group are the same virtual speaker group, then M = N, and the first virtual speaker and the second virtual speaker are the same.
[0185] Optionally, M can be any of 4, 8, 16, etc., and N can be any of 4, 8, 16, etc.
[0186] The first virtual speaker can process the audio signal to be processed into a first audio signal using the following formula: M first virtual speakers correspond one-to-one with M first audio signals.
[0187]
[0188] Where 1≤m≤M; P 1mLet W be the m-th first audio signal after the m-th first virtual speaker processes the audio signal to be processed. W represents the components corresponding to all sounds in the environment where the sound source is located, called the environmental component. X represents the components of all sounds in the environment where the sound source is located along the X-axis, called the X-coordinate component. Y represents the components of all sounds in the environment where the sound source is located along the Y-axis, called the Y-coordinate component. Z represents the components of all sounds in the environment where the sound source is located along the Z-axis, called the Z-coordinate component. Here, the X-axis, Y-axis, and Z-axis are the X-axis, Y-axis, and Z-axis of the three-dimensional coordinate system corresponding to the sound source (i.e., the three-dimensional coordinate system corresponding to the audio signal transmitter), respectively. L is the energy adjustment coefficient. φ 1m Let θ be the pitch angle of the m-th virtual loudspeaker relative to the origin of the three-dimensional coordinate system corresponding to the audio signal receiver. 1m The azimuth angle of the m-th first virtual speaker relative to the origin of the coordinate system.
[0189] The first audio signal can be a multi-channel signal or a mono signal.
[0190] The second virtual speaker can process the audio signal to be processed into a second audio signal using the following formula: N second virtual speakers correspond one-to-one with N second audio signals.
[0191]
[0192] Where 1≤n≤N; P 1n Let W be the nth first audio signal after the audio signal to be processed by the nth first virtual speaker, and W be the components corresponding to all sounds in the environment where the sound source is located, called the environmental component. X is the component of all sounds in the environment where the sound source is located along the X-axis, called the X-coordinate component. Y is the component of all sounds in the environment where the sound source is located along the Y-axis, called the Y-coordinate component. Z is the component of all sounds in the environment where the sound source is located along the Z-axis, called the Z-coordinate component. Here, the X-axis, Y-axis, and Z-axis are the X-axis, Y-axis, and Z-axis of the three-dimensional coordinate system of the environment where the sound source is located, respectively. L is the energy adjustment coefficient. 1n Let θ be the pitch angle of the nth virtual loudspeaker relative to the origin of the three-dimensional coordinate system corresponding to the audio signal receiver. 1n The azimuth angle of the nth first virtual speaker relative to the origin of the coordinate system.
[0193] The second audio signal can be a multi-channel signal or a mono signal.
[0194] Specifically, in step S102, the M first HRTFs can be referred to as the M first HRTFs corresponding to the M first virtual speakers, with each first virtual speaker corresponding to one first HRTF, that is, the M first HRTFs correspond one-to-one with the M first virtual speakers; the N second HRTFs can be referred to as the N second HRTFs corresponding to the N second virtual speakers, with each second virtual speaker corresponding to one second HRTF, that is, the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0195] In the prior art, the first HRTF is a head-centered HRTF, and the second HRTF is also a head-centered HRTF.
[0196] In this embodiment, "centering on the head" means taking the head center as the center for measuring HRTF.
[0197] Figure 5 This is a measurement scenario diagram provided for an embodiment of this application, with the head center as the center for measuring HRTF. See also... Figure 5 , Figure 5 The diagram illustrates several positions 61 relative to the head center 62. It can be understood that there are multiple HRTFs centered on the head center, and audio signals from the first sound source at different positions 61 are transmitted to the head center corresponding to different HRTFs centered on the head center. The head center used to measure the HRTF centered on the head center can be the head center of the current listener, the head center of other listeners, or the head center of a virtual listener.
[0198] By setting the first sound source at different preset positions relative to the head center 62, multiple HRTFs corresponding to the preset positions can be obtained. Specifically, if the position of the first sound source 1 relative to the head center 62 is position c, the measured HRTF1 of the first sound source 1 transmitted to the head center 62 is the HRTF1 corresponding to position c and centered on the head center. Similarly, if the position of the first sound source 2 relative to the head center 62 is position d, the measured HRTF2 of the first sound source 2 transmitted to the head center 62 is the HRTF2 corresponding to position d and centered on the head center. The center HRTF2, etc.; where position c includes azimuth angle 1, pitch angle 1 and distance 1, azimuth angle 1 is the azimuth angle of the first sound source 1 relative to the head center 62, pitch angle 1 is the pitch angle of the first sound source 1 relative to the head center 62, and distance 1 is the distance between the first sound source 1 and the head center 62; similarly, position d includes azimuth angle 2, pitch angle 2 and distance 2, azimuth angle 2 is the azimuth angle of the first sound source 2 relative to the head center 62, pitch angle 2 is the pitch angle of the first sound source 2 relative to the head center 62, and distance 2 is the distance between the first sound source 2 and the head center 62.
[0199] Specifically, when setting the position of the first sound source relative to the head center 62, when the distance and pitch angle remain unchanged, the azimuth angle of adjacent first sound sources can be spaced apart by a first preset angle; when the distance and azimuth angle remain unchanged, the pitch angle of adjacent first sound sources can be spaced apart by a second preset angle; when the pitch angle and azimuth angle remain unchanged, the distance between adjacent first sound sources can be spaced apart by a first preset distance. The first preset angle can be any from 3° to 10°, for example, 5°; the second preset angle can be any from 3° to 10°, for example, 5°; and the first distance can be any from 0.05m to 0.2m, for example, 0.1m.
[0200] For example, the process of obtaining the HRTF1 centered on the head center at position c (100°, 50°, 1m) is as follows: a first sound source 1 is set at a distance of 1m with an azimuth angle of 100° and a pitch angle of 50° relative to the head center. The audio signal emitted by the first sound source 1 is measured and transmitted to the HRTF corresponding to the head center 62 to obtain the HRTF1 centered on the head center. The measurement method is the existing method, which will not be described in detail here.
[0201] For example, the process of obtaining the HRTF1 centered on the head center at position d(100°, 45°, 1m) is as follows: a first sound source 2 is set at a distance of 1m with an azimuth angle of 100° and a pitch angle of 45° relative to the head center. The audio signal emitted by the first sound source 2 is measured and transmitted to the HRTF corresponding to the head center 62 to obtain the HRTF2 centered on the head center.
[0202] For example, the process of obtaining the HRTF1 centered on the head center at position e (95°, 45°, 1m) is as follows: a first sound source 3 is set at a distance of 1m with an azimuth angle of 95° and a pitch angle of 45° relative to the head center. The audio signal emitted by the first sound source 3 is measured and transmitted to the HRTF corresponding to the head center 62 to obtain the HRTF3 centered on the head center.
[0203] For example, the process of obtaining the HRTF1 centered on the head center at position f(95°, 50°, 1m) is as follows: a first sound source 4 is set at a distance of 1m with an azimuth angle of 95° and a pitch angle of 50° relative to the head center. The audio signal emitted by the first sound source 4 is measured and transmitted to the HRTF corresponding to the head center 62 to obtain the HRTF4 centered on the head center.
[0204] For example, the process of obtaining the HRTF1 centered on the head center at position g (100°, 50°, 1.1m) is as follows: a first sound source 5 is set at a distance of 1m with an azimuth angle of 95° and a pitch angle of 50° relative to the head center. The audio signal emitted by the first sound source 5 is measured and transmitted to the HRTF corresponding to the head center 62 to obtain the HRTF5 centered on the head center.
[0205] It is worth noting that in the subsequent positions (x, x, x), the first x is the azimuth angle, the second x is the elevation angle, and the third x is the distance.
[0206] Using the above method, the correspondence between multiple positions and multiple head-centered HRTFs can be measured. It is understood that the multiple positions where the first sound source is placed when measuring the head-centered HRTF can be called preset positions. Therefore, using the above method, the correspondence between multiple preset positions and multiple head-centered HRTFs can be measured. This correspondence can be called a second correspondence, and the second correspondence can be stored in the memory 22 shown in Figure 3.
[0207] In practical applications, the aforementioned prior art obtains the position 'a' of the first virtual speaker relative to the current left ear position. The HRTF (Head-to-Head Response Time) measured at position 'a', centered on the head center, is the HRTF corresponding to the first virtual speaker. Similarly, it obtains the position 'b' of the second virtual speaker relative to the current right ear position. The HRTF measured at position 'b', centered on the head center, is the HRTF corresponding to the second virtual speaker. It is clear that position 'a' is not the position of the first virtual speaker relative to the head center, but rather relative to the left ear position. If the HRTF at position 'a', centered on the head center, is still used as the HRTF for the first virtual speaker, the signal transmitted to the left ear will not be optimal; the optimal signal is at the head center position. Likewise, position 'b' is not the position of the second virtual speaker relative to the head center, but rather relative to the right ear position. If the HRTF at position 'b', centered on the head center, is still used as the HRTF for the second virtual speaker, the signal transmitted to the right ear will not be optimal; the optimal signal is at the head center position.
[0208] In this embodiment, the first HRTF corresponding to the first virtual speaker is the HRTF centered on the left ear position; the second HRTF corresponding to the second virtual speaker is the HRTF centered on the right ear position.
[0209] In this embodiment, "centering on the left ear" means using the left ear position as the center for measuring HRTF, and "centering on the right ear" means using the right ear position as the center for measuring HRTF.
[0210] The HRTF centered on the left ear position can be obtained through actual measurement, that is, by collecting the audio signal a emitted by the sound source at position a relative to the left ear position, and collecting the audio signal b transmitted from the audio signal a to the left ear position, and obtaining it based on the audio signal a and the audio signal b; the HRTF centered on the left ear position can also be obtained by converting the HRTF centered on the head center; these two acquisition methods will be described in detail in subsequent embodiments.
[0211] Similarly, the HRTF centered on the right ear position can be obtained through actual measurement, that is, by collecting the audio signal c emitted by the sound source at position b relative to the right ear position, and collecting the audio signal d transmitted from the audio signal c to the right ear position, and obtaining it based on the audio signal c and the audio signal d; the HRTF centered on the right ear position can also be obtained by converting the HRTF centered on the head center; these two acquisition methods will be described in detail in subsequent embodiments.
[0212] For step S103, obtain the first target audio signal based on M first audio signals and M first HRTFs, and obtain the second target audio signal based on N second audio signals and N second HRTFs.
[0213] Specifically, based on M first audio signals and M first HRTFs, the first target audio signal is obtained, including:
[0214] M first audio signals are convolved with their corresponding first HRTFs to obtain M first convolved audio signals.
[0215] The first target audio signal is obtained from M first convolutional audio signals.
[0216] That is, the mth first audio signal output by the mth first virtual speaker is convolved with the first HRTF corresponding to the mth first virtual speaker to obtain the mth convolutional audio signal. When there are M first virtual speakers, M first convolutional audio signals will be obtained.
[0217] The signal obtained by superimposing M first convolutional audio signals is the first target audio signal, which is the audio signal transmitted to the left ear position, or the audio signal corresponding to the left ear position obtained by rendering.
[0218] Since the m-th first audio signal output by the m-th first virtual speaker is convolved with the first HRTF corresponding to the m-th first virtual speaker, and the first HRTF corresponding to the m-th first virtual speaker is the HRTF of the m-th first audio signal centered on the left ear position, the signal of the first target audio signal transmitted to the left ear position is the optimal signal.
[0219] Based on N second audio signals and N second HRTFs, obtain the second target audio signal;
[0220] The N second audio signals are convolved with their corresponding second HRTF signals to obtain N second convolved audio signals.
[0221] The second target audio signal is obtained from N second convolutional audio signals.
[0222] That is, the nth second audio signal output by the nth second virtual speaker is convolved with the second HRTF corresponding to the nth second virtual speaker to obtain the nth convolved audio signal. When there are N first virtual speakers, N second convolved audio signals will be obtained.
[0223] The signal obtained by superimposing N second convolutional audio signals is the second target audio signal, which is the audio signal transmitted to the right ear position, or the audio signal corresponding to the right ear position obtained by rendering.
[0224] Since the nth second audio signal output by the nth second virtual speaker is convolved with the second HRTF corresponding to the nth second virtual speaker, and the second HRTF corresponding to the nth second virtual speaker is an HRTF centered on the right ear position, the signal of the second target audio signal transmitted to the right ear position is the optimal signal.
[0225] It is understandable that the first target audio signal and the second target audio signal here are the rendered audio signals, and the first target audio signal and the second target audio signal constitute the stereo signal finally output by the audio signal receiver.
[0226] In this embodiment, a first target audio signal is obtained by using M first audio signals and M first HRTFs centered on the left ear position, so that the signal transmitted to the left ear position is optimal. Similarly, a second target audio signal is obtained by using N second audio signals and N second HRTFs centered on the right ear position, so that the signal transmitted to the right ear position is optimal. Therefore, the quality of the audio signal output by the audio signal receiver is improved.
[0227] The following uses Figures 6 to 15 The illustrated embodiments are for Figure 4 The embodiments shown are described in detail. Figures 6 to 15 The embodiments shown involve the following: Figure 4 In the embodiments shown, words with the same name have the same meaning.
[0228] First of all, Figure 4 The first method for obtaining the M first HRTFs in step S102 of the illustrated embodiment will be described. Figure 6Flowchart of the audio processing method provided in the embodiments of this application Figure 2 See Figure 6 The method in this embodiment includes:
[0229] Step S201: Obtain the M first positions of the M first virtual speakers relative to the current left ear position;
[0230] Step S202: Based on the M first positions and the first correspondence, determine the M HRTFs corresponding to the M first positions as M first HRTFs; wherein, the first correspondence is the correspondence between multiple preset positions and multiple HRTFs centered on the left ear position that are stored in advance.
[0231] Specifically, for step S201, obtaining the first position of each first virtual speaker relative to the current left ear position, if there are M first virtual speakers, then M first positions will be obtained.
[0232] Each first position includes a third pitch angle and a third azimuth angle of the corresponding first virtual speaker relative to the current left ear position, as well as a third distance between the first virtual speaker and the current left ear position. The current left ear position is the left ear of the current listener.
[0233] For step S202, before step S202, it is necessary to obtain multiple preset positions and multiple HRTFs centered on the left ear position in advance;
[0234] Figure 7 This application provides a measurement scenario diagram with the center of the left ear location as the center for measuring HRTF. See also... Figure 7 , Figure 7 The diagram illustrates several positions 81 relative to the left ear position 82. It can be understood that there are multiple HRTFs centered on the left ear position. Audio signals transmitted from the second sound source at different positions 81 correspond to different HRTFs when transmitted to the left ear position. Therefore, before step S202, it is necessary to measure the HRTFs centered on the left ear position for multiple positions 81. The left ear position when measuring the HRTF centered on the left ear position can be the current left ear position of the current listener, the left ear position of another listener, or the left ear position of a virtual listener.
[0235] Second sound sources are positioned at different locations relative to the left ear position 82 to obtain multiple HRTFs centered on the left ear position corresponding to positions 81. Specifically, if the position of second sound source 1 relative to the left ear position 82 is position c, the measured signal emitted by second sound source 1 is transmitted to the HRTF of the left ear position 82, which is HRTF1 corresponding to position c and centered on the left ear position. Similarly, if the position of second sound source 2 relative to the left ear position 82 is position d, the measured signal emitted by second sound source 2 is transmitted to the HRTF of the left ear position 82, which is HRTF1 corresponding to position d and... HRTF2, centered on the left ear position; wherein, position c includes azimuth angle 1, elevation angle 1, and distance 1, azimuth angle 1 is the azimuth angle of the second sound source 1 relative to the left ear position 82, elevation angle 1 is the elevation angle of the second sound source 1 relative to the left ear position 82, and distance 1 is the distance between the second sound source 1 and the left ear position 82; similarly, position d includes azimuth angle 2, elevation angle 2, and distance 2, azimuth angle 2 is the azimuth angle of the second sound source 2 relative to the left ear position 82, elevation angle 2 is the elevation angle of the second sound source 2 relative to the left ear position 82, and distance 2 is the distance between the second sound source 2 and the left ear position 82.
[0236] Understandably, when setting the position of the second sound source relative to the left ear position 82, with the distance and elevation angle remaining constant, the azimuth angle of adjacent second sound sources can be spaced apart by a first angle; with the distance and azimuth angle remaining constant, the elevation angle of adjacent second sound sources can be spaced apart by a second angle; and with the elevation angle and azimuth angle remaining constant, the distance between adjacent second sound sources can be spaced apart by a first distance. The first angle can be any value between 3° and 10°, for example, 5°; the second angle can be any value between 3° and 10°, for example, 5°; and the first distance can be any value between 0.05m and 0.2m, for example, 0.1m.
[0237] For example, the process of obtaining HRTF1 centered on the left ear position corresponding to position c (100°, 50°, 1m) is as follows: a second sound source 1 is set at a distance of 1m with an azimuth angle of 100° and a pitch angle of 50° relative to the left ear position 82. The HRTF corresponding to the left ear position is measured by measuring the audio signal emitted by the second sound source 1. HRTF1 centered on the left ear position is obtained.
[0238] For example, the process of obtaining the HRTF2 centered on the left ear position corresponding to position d (100°, 45°, 1m) is as follows: a second sound source 2 is set at a distance of 1m with an azimuth angle of 100° and an elevation angle of 45° relative to the left ear position 82. The HRTF corresponding to the left ear position is measured by measuring the audio signal emitted by the second sound source 2. The HRTF2 centered on the left ear position is obtained.
[0239] For example, the process of obtaining the HRTF3 centered on the left ear position corresponding to position e (95°, 50°, 1m) is as follows: a second sound source 3 is set at a distance of 1m with an azimuth angle of 95° and an elevation angle of 50° relative to the left ear position 82. The audio signal emitted by the second sound source 3 is measured to the HRTF corresponding to the left ear position, and the HRTF3 centered on the left ear position is obtained, and so on.
[0240] For example, the process of obtaining the HRTF4 centered on the left ear position corresponding to position f (95°, 45°, 1m) is as follows: a second sound source 4 is set at a distance of 1m with an azimuth angle of 95° and an elevation angle of 40° relative to the left ear position 82. The audio signal emitted by the second sound source 4 is measured to the HRTF corresponding to the left ear position, and the HRTF4 centered on the left ear position is obtained.
[0241] For example, the process of obtaining the HRTF5 centered on the left ear position corresponding to position g (100°, 50°, 1.2m) is as follows: a second sound source 5 is set at a distance of 1.2m with an azimuth angle of 100° and an elevation angle of 50° relative to the left ear position 82. The audio signal emitted by the second sound source 5 is measured and transmitted to the HRTF corresponding to the left ear position to obtain the HRTF5 centered on the left ear position.
[0242] For example, the process of obtaining the HRTF5 centered on the left ear position corresponding to position h (95°, 50°, 1.1m) is as follows: a second sound source 6 is set at a distance of 1.1m with an azimuth angle of 95° and an elevation angle of 50° relative to the left ear position 82. The audio signal emitted by the second sound source 6 is measured and transmitted to the HRTF corresponding to the left ear position to obtain the HRTF6 centered on the left ear position.
[0243] Understandably, since the azimuth range is (-180° to 180°) and the pitch range is (-90° to 90°), if the first angle is 5°, the second angle is 5°, the first distance can be 0.05m, and the total distance is 2m, for example, 0.1m, then 72×36×21 positions corresponding to the HRTF centered on the left ear position can be obtained.
[0244] Using the above method, the correspondence between multiple positions and multiple HRTFs centered on the left ear position can be measured. It is understood that the multiple positions where the second sound source is placed when measuring the HRTF centered on the left ear position can be called preset positions. Therefore, using the above method, the correspondence between multiple preset positions and multiple HRTFs centered on the left ear position can be measured. This correspondence can be called the first correspondence, and the first correspondence can be stored in the memory 22 shown in Figure 3.
[0245] Next, based on the M first positions and the first correspondence, the M HRTFs corresponding to the M first positions are determined as the M first HRTFs. The first correspondence includes the correspondence between multiple preset positions and multiple HRTFs centered on the left ear position, including:
[0246] Determine the M first preset positions associated with the M first positions; the M first preset positions are preset positions in the first correspondence relationship;
[0247] According to the first correspondence, the M HRTFs centered on the left ear position corresponding to the M first preset positions are determined as the M first HRTFs; the M HRTFs centered on the left ear position are actually the M HRTFs centered on the left ear position 82 corresponding to the audio signals emitted by the sound sources at the M first preset positions transmitted to the left ear position 82.
[0248] Wherein, the first preset position associated with the first position may be the first position itself; or...
[0249] The first preset position includes a pitch angle that is the target pitch angle closest to the third pitch angle included in the first position, an azimuth angle that is the target azimuth angle closest to the third azimuth angle included in the first position, and a distance that is the target distance closest to the third distance included in the first position. Specifically, the target azimuth angle is the azimuth angle included in the preset position when measuring the HRTF centered on the left ear position, which is the azimuth angle of the second sound source placed relative to the left ear position when measuring the HRTF centered on the left ear position. The target pitch angle is the pitch angle in the preset position when measuring the HRTF centered on the left ear position, which is the pitch angle of the second sound source placed relative to the left ear position when measuring the HRTF centered on the left ear position. The target distance is the distance in the preset position when measuring the HRTF centered on the left ear position, which is the distance of the second sound source placed relative to the left ear position when measuring the HRTF centered on the left ear position. In other words, the first preset position refers to the position where the second sound source is placed when measuring multiple HRTFs centered on the left ear position, meaning that the HRTF centered on the left ear position for each first preset position has been measured beforehand.
[0250] Understandably, if the third azimuth angle included in the first position is located between the two target azimuth angles, the choice of which of the two target azimuth angles to include in the first preset position can be determined according to preset rules. For example, the preset rule is: if the third azimuth angle included in the first position is located between the two target azimuth angles, then the smaller of the two target azimuth angles is determined as the azimuth angle included in the first preset position. Similarly, if the third elevation angle included in the first position is located between the two target elevation angles, the choice of which of the two target elevation angles to include in the first preset position can be determined according to preset rules. For example, the preset rule is: if the third elevation angle included in the first position is located between the two target elevation angles, then the smaller of the two target elevation angles is determined as the elevation angle included in the first preset position. If the third distance included in the first position is located in the middle of the two target distances, the choice of which of the two target distances is to be included in the first preset position can be determined according to a preset rule. For example, the preset rule is: if the third distance included in the first position is located in the middle of the two target distances, then the smaller of the two target distances is determined as the distance included in the first preset position.
[0251] For example, if the third azimuth angle of the first position of the m-th first virtual speaker relative to the left ear position measured in step S201 is 88°, the third pitch angle is 46°, and the third distance is 1.02m, the correspondence between the multiple preset positions measured in advance and the multiple HRTFs centered on the left ear position includes (90°, 45°, 1m) corresponding to the HRTF centered on the left ear position, (85°, 45°, 1m) corresponding to the HRTF centered on the left ear position, (90°, 50°, 1m) corresponding to the HRTF centered on the left ear position, (85°, 50°, 1m) corresponding to the HRTF centered on the left ear position, (90°, 45 ... The HRTFs centered on the left ear position are (1m), (85°, 45°, 1.1m), (90°, 50°, 1.1m), and (85°, 50°, 1.1m). Since 88° is between 85° and 90° but closer to 90°, 46° is between 45° and 50° but closer to 45°, and 1.02m is between 1m and 1.1m but closer to 1m, (90°, 45°, 1m) is determined as the first preset position m associated with the first position of the m-th first virtual speaker relative to the left ear position.
[0252] After determining the M first preset positions associated with the M first positions, the M HRTFs centered on the left ear position corresponding to the M first preset positions are the M first HRTFs. For example, in the above example, in the first correspondence, the HRTF centered on the left ear position corresponding to the first preset position m (90°, 45°, 1m) is the HRTF corresponding to the first position of the m-th first virtual speaker relative to the current left ear position. That is to say, in the first correspondence, the HRTF centered on the left ear position corresponding to the first preset position m (90°, 45°, 1m) is the m-th first HRTF, or one of the M first HRTFs.
[0253] In this embodiment, the M first HRTFs corresponding to the M virtual speakers are the M HRTFs centered on the left ear position obtained by actual measurement. These M first HRTFs best represent the HRTFs corresponding to the transmission of the M first audio signals to the current left ear position, so that the signal transmitted to the left ear position is optimal.
[0254] Next, regarding Figure 4 The second method for obtaining the M first HRTFs in step S102 of the illustrated embodiment will be described. Figure 8 Flowchart of the audio processing method provided in the embodiments of this application Figure 3 See Figure 8 The method in this embodiment includes:
[0255] Step S301: Obtain M third positions of the first virtual speaker relative to the current head center; the third position includes the first azimuth angle and the first pitch angle of the first virtual speaker relative to the current head center, and the first distance between the current head center and the first virtual speaker;
[0256] Step S302: Determine M fourth positions based on M third positions. Each of the M third positions corresponds one-to-one with the M fourth positions. A fourth position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle of the fourth position and the first value is the first azimuth angle of the corresponding third position. The first value is the difference between the first included angle and the second included angle. The first included angle is the first included angle between the first straight line and the first surface. The second included angle is the angle between the second straight line and the first surface. The first straight line is the straight line passing through the current left ear and the origin of the three-dimensional coordinate system. The second straight line is the straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0257] Step S303: Based on the M fourth positions and the second correspondence, determine the M HRTFs corresponding to the M fourth positions as the M first HRTFs; wherein, the second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center;
[0258] Specifically, in step S301, the third position of each first virtual speaker relative to the current head center is obtained. If there are M first virtual speakers, then M third positions will be obtained. The current head center is the head center of the current listener.
[0259] Each third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker.
[0260] For step S302, for each third position, the second pitch angle included in the third position is used as the pitch angle included in the corresponding fourth position, the second distance included in the third position is used as the distance included in the corresponding fourth position, and the second azimuth angle included in the third position is added to the first value to obtain the azimuth angle included in the corresponding fourth position. For example, if the third position is (52°, 73°, 0.5m) and the first value is 6°, then the fourth position is (58°, 73°, 0.5m).
[0261] The three-dimensional coordinate system in this embodiment is the same as the three-dimensional coordinate system corresponding to the audio receiver mentioned above.
[0262] Before step S303, it is necessary to obtain the correspondence between multiple preset locations and multiple head-centered HRTFs. The method for obtaining the correspondence between multiple preset locations and multiple head-centered HRTFs is described in [reference needed]. Figure 4 The descriptions in the illustrated embodiments will not be repeated in this embodiment.
[0263] Specifically, based on the M fourth positions and the second correspondence, the M HRTFs corresponding to the M fourth positions are determined to be the M first HRTFs. The second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center, including:
[0264] Based on the M fourth positions, determine the M second preset positions associated with the M fourth positions; the M second preset positions are preset positions in the pre-stored second correspondence relationship;
[0265] Based on the second correspondence, determine the M HRTFs centered on the head center corresponding to the second preset positions, which are the M first HRTFs.
[0266] Specifically, the second preset position associated with the fourth position may be the fourth position itself; or,
[0267] The second preset position includes the following: the pitch angle is the target pitch angle closest to the pitch angle included in the fourth position; the azimuth angle included in the second preset position is the target azimuth angle closest to the azimuth angle included in the fourth position; and the distance included in the second preset position is the target distance closest to the distance included in the fourth position. Specifically, the target azimuth angle is the azimuth angle included in the preset position when measuring the HRTF centered on the head center, which is the azimuth angle of the first sound source placed relative to the head center when measuring the HRTF centered on the head center. The target pitch angle is the pitch angle in the preset position when measuring the HRTF centered on the head center, which is the pitch angle of the first sound source placed relative to the head center when measuring the HRTF centered on the head center. The target distance is the distance in the preset position when measuring the HRTF centered on the head center, which is the distance of the first sound source placed relative to the head center when measuring the HRTF centered on the head center. In other words, the second preset position is the position where the first sound source is placed when measuring multiple HRTFs centered on the head center, meaning that the HRTF centered on the head center corresponding to each second preset position has been measured beforehand.
[0268] It is understandable that the methods for determining the azimuth of the second preset position if the azimuth of the fourth position is located between the two target azimuths, the methods for determining the pitch of the second preset position if the pitch of the fourth position is located between the two target pitch angles, and the methods for determining the pitch of the second preset position if the pitch of the fourth position is located between the two target pitch angles are the same as the description of the first preset position associated with the first position, and will not be repeated here.
[0269] In determining M fourth positions associated with M second preset positions, the HRTFs centered on the head for each of the M second preset positions are the M first HRTFs. For example, if a certain fourth position is associated with a second preset position of (30°, 60°, 0.5m), then in the second correspondence, the HRTF corresponding to (30°, 60°, 0.5m) is the HRTF centered on the head for that fourth position. In other words, in the second correspondence, the HRTF centered on the head for (30°, 60°, 0.5m) is one of the M first HRTFs.
[0270] In this embodiment, the M first HRTFs are obtained by converting the HRTF centered on the head, which is highly efficient in obtaining the first HRTFs.
[0271] Next, regarding Figure 4 The third method for obtaining the M first HRTFs in step S102 of the illustrated embodiment will be described. Figure 9Flowchart of the audio processing method provided in the embodiments of this application Figure 4 See Figure 9 The method in this embodiment includes:
[0272] Step S401: Obtain M third positions of the first virtual speaker relative to the current head center; the third position includes the first azimuth angle and the first pitch angle of the first virtual speaker relative to the current head center, and the first distance between the current head center and the first virtual speaker;
[0273] Step S402: Determine M seventh positions based on M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in a seventh position and the first preset value is the first azimuth angle included in the corresponding third position. The correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center.
[0274] Step S403: Based on the M seventh positions and the second correspondence, determine the M HRTFs corresponding to the M seventh positions as the M first HRTFs; wherein, the second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center.
[0275] Specifically, step S401 in this embodiment refers to Figure 8 Step S301 in the illustrated embodiment will not be repeated here.
[0276] For step S402, the three-dimensional coordinate system in this embodiment is the three-dimensional coordinate system corresponding to the audio receiver mentioned above.
[0277] Specifically, for each third position, the second pitch angle included in the third position is used as the pitch angle included in the corresponding seventh position, the second distance included in the third position is used as the distance included in the corresponding seventh position, and the second azimuth angle included in the third position is added to a first preset value to obtain the azimuth angle included in the corresponding seventh position. For example, if the third position is (52°, 73°, 0.5m) and the first preset value is 5°, then the seventh position is (57°, 73°, 0.5m).
[0278] The first preset value is pre-set and does not consider the size of the listener's head. In the previous embodiment, the first value was the difference between the first angle and the second angle, taking into account the size of the current listener's head. Optionally, the first preset value and... Figure 4 The first preset angle described in the illustrated embodiments is the same.
[0279] Before step S403, it is necessary to obtain the correspondence between multiple preset locations and multiple head-centered HRTFs. The method for obtaining the correspondence between multiple preset locations and multiple head-centered HRTFs is described in [reference needed]. Figure 4 The descriptions in the illustrated embodiments will not be repeated in this embodiment.
[0280] Specifically, based on the M seventh positions and the second correspondence, the M HRTFs corresponding to the M seventh positions are determined to be the M first HRTFs. The second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center, including:
[0281] Based on the M seventh positions, determine the M third preset positions associated with the M seventh positions; the M third preset positions are preset positions in the second correspondence relationship;
[0282] Based on the second correspondence, determine the M HRTFs centered on the head center corresponding to the third preset positions, which are the M first HRTFs.
[0283] Specifically, for the third preset position associated with the seventh position, refer to Figure 6 The explanation of the first preset position associated with the first position in the illustrated embodiment will not be repeated here.
[0284] After determining the M third preset positions associated with the M seventh positions, the HRTFs centered on the head for each of the M third preset positions are the M first HRTFs. For example, if a certain seventh position is associated with a third preset position of (35°, 60°, 0.5m), then in the second correspondence, the HRTF centered on the head for (35°, 60°, 0.5m) is the same as the HRTF centered on the head for that seventh position. In other words, in the second correspondence, the HRTF centered on the head for (35°, 60°, 0.5m) is one of the first HRTFs.
[0285] In this embodiment, the M first HRTFs are obtained by converting the HRTF centered on the head, and when obtaining the fourth position mentioned above, the size of the current listener's head is not considered, which further improves the efficiency of obtaining the first HRTFs.
[0286] Secondly, for Figure 4 The first acquisition process of the N second HRTFs in step S102 of the illustrated embodiment will be described. Figure 10 Flowchart of the audio processing method provided in the embodiments of this application Figure 5 See Figure 10 The method in this embodiment includes:
[0287] Step S501: Obtain the N second positions of the N second virtual speakers relative to the current right ear position;
[0288] Step S502: Based on the N second positions and the third correspondence, determine the N HRTFs corresponding to the N second positions as N second HRTFs; wherein, the third correspondence is the correspondence between multiple preset positions and multiple HRTFs centered on the right ear position that are stored in advance.
[0289] Specifically, for step S501, the second position of each second virtual speaker relative to the listener's right ear position is obtained. If there are N second virtual speakers, then N second positions will be obtained.
[0290] Each second position includes a fourth pitch angle and a fourth azimuth angle of the corresponding second virtual speaker relative to the current right ear position, as well as a fourth distance between the second virtual speaker and the current right ear position. The current right ear position is the right ear of the current listener.
[0291] For step S502, before step S502, it is necessary to obtain multiple preset positions and multiple HRTFs centered on the right ear position in advance;
[0292] Figure 11 This application provides a measurement scenario diagram with the center of the right ear location as the center for measuring HRTF. See also... Figure 11 , Figure 11 The diagram illustrates several positions 51 relative to the right ear position 52. It can be understood that there are multiple HRTFs centered on the right ear position. Audio signals sent by the third sound source at different positions 51 are transmitted to the right ear position and correspond to different HRTFs. The right ear position when measuring the HRTF centered on the right ear position can be the current right ear position of the current listener, the right ear position of other listeners, or the right ear position of a virtual listener.
[0293] By setting up third sound sources at different positions relative to the right ear position 52, multiple HRTFs centered on the right ear position are obtained for each position 51. Specifically, if the position of third sound source 1 relative to the right ear position 52 is position c, the measured signal emitted by third sound source 1 is transmitted to the HRTF of the right ear position 52, which is the HRTF1 centered on the right ear position corresponding to position c. Similarly, if the position of third sound source 2 relative to the right ear position 52 is position d, the measured signal emitted by third sound source 2 is transmitted to the HRTF of the right ear position 52, which is the HRTF centered on the right ear position. Correspondingly, HRTF2, centered on the right ear position, etc.; wherein, position c includes azimuth angle 1, elevation angle 1, and distance 1, azimuth angle 1 is the azimuth angle of the third sound source 1 relative to the right ear position 52, elevation angle 1 is the elevation angle of the third sound source 1 relative to the right ear position 52, and distance 1 is the distance between the third sound source 1 and the right ear position 52; similarly, position d includes azimuth angle 2, elevation angle 2, and distance 2, azimuth angle 2 is the azimuth angle of the third sound source 2 relative to the right ear position 52, elevation angle 2 is the elevation angle of the third sound source 2 relative to the right ear position 52, and distance 2 is the distance between the third sound source 2 and the right ear position 52.
[0294] Understandably, when setting the position of the third sound source relative to the right ear position 52, with the distance and pitch angle remaining constant, the azimuth angle of adjacent third sound sources can be spaced apart by a first preset angle; with the distance and azimuth angle remaining constant, the pitch angle of adjacent third sound sources can be spaced apart by a second preset angle; and with the pitch angle and azimuth angle remaining constant, the distance between adjacent third sound sources can be spaced apart by a first preset distance. The first preset angle can be any value between 3° and 10°, for example, 5°; the second preset angle can be any value between 3° and 10°, for example, 5°; and the first preset distance can be any value between 0.05m and 0.2m, for example, 0.1m.
[0295] For example, the process of obtaining the HRTF1 centered on the right ear position corresponding to position c (100°, 50°, 1m) is as follows: a third sound source 1 is set at a distance of 1m with an azimuth angle of 100° and a pitch angle of 50° relative to the right ear position. The HRTF corresponding to the right ear position is measured by measuring the audio signal emitted by the third sound source 1. The HRTF1 centered on the right ear position is obtained.
[0296] For example, the process of obtaining the HRTF2 centered on the right ear position corresponding to position d(100°, 45°, 1m) is as follows: a third sound source 2 is set at a distance of 1m with an azimuth angle of 100° and a pitch angle of 45° relative to the right ear position. The HRTF corresponding to the right ear position is measured by measuring the audio signal emitted by the third sound source 2. The HRTF2 centered on the right ear position is obtained.
[0297] For example, the process of obtaining the HRTF3 centered on the right ear position corresponding to position e (95°, 50°, 1m) is as follows: a third sound source 3 is set at a distance of 1m with an azimuth angle of 95° and a pitch angle of 50° relative to the right ear position. The audio signal emitted by the third sound source 3 is measured to the HRTF corresponding to the right ear position, and the HRTF3 centered on the right ear position is obtained, and so on.
[0298] For example, the process of obtaining the HRTF3 centered on the right ear position corresponding to position f (95°, 45°, 1m) is as follows: a third sound source 4 is set at a distance of 1m with an azimuth angle of 95° and a pitch angle of 40° relative to the right ear position. The audio signal emitted by the third sound source 4 is measured to the HRTF corresponding to the right ear position, and the HRTF4 centered on the right ear position is obtained, and so on.
[0299] For example, the process of obtaining the HRTF5 centered on the right ear position corresponding to position g (100°, 50°, 1.2m) is as follows: a third sound source 5 is set at a distance of 1.2m with an azimuth angle of 100° and a pitch angle of 50° relative to the right ear position. The HRTF corresponding to the right ear position is measured by the audio signal emitted by the third sound source 5. The HRTF5 centered on the right ear position is obtained.
[0300] For example, the process of obtaining the HRTF5 centered on the right ear position corresponding to position h (95°, 50°, 1.1m) is as follows: a third sound source 6 is set at a distance of 1.1m with an azimuth angle of 95° and an elevation angle of 50° relative to the right ear position. The HRTF corresponding to the right ear position is measured by the audio signal emitted by the third sound source 6. The HRTF6 centered on the right ear position is then obtained.
[0301] Understandably, since the azimuth range is (-180° to 180°) and the pitch range is (-90° to 90°), if the first preset angle is 5°, the second preset angle is 5°, the first preset distance can be 0.05m, and the total distance is 2m, for example, 0.1m, then 72×36×21 HRTFs centered on the right ear position can be obtained.
[0302] Using the method described above, the correspondence between multiple locations and multiple HRTFs centered on the right ear position can be measured. It can be understood that the multiple locations where the third sound source is placed when measuring the HRTF centered on the right ear position can be called preset locations. Therefore, using the method described above, the correspondence between multiple preset locations and multiple HRTFs centered on the right ear position can be measured. This correspondence is called the third correspondence, and the third correspondence can be stored... Figure 3 In the memory 22 shown.
[0303] Next, based on the N second positions and the third correspondence, the N HRTFs corresponding to the N second positions are determined to be the N second HRTFs. The third correspondence is the correspondence between multiple pre-stored preset positions and multiple HRTFs centered on the right ear position, including:
[0304] Determine the N fourth preset positions associated with the N second positions;
[0305] Based on the third correspondence, the N HRTFs centered on the right ear position corresponding to the N fourth preset positions are determined as the N second HRTFs.
[0306] Wherein, the fourth preset position associated with the second position may be the second position itself; or,
[0307] The fourth preset position includes a pitch angle that is the target pitch angle closest to the fourth pitch angle included in the second position; an azimuth angle that is the target azimuth angle closest to the fourth azimuth angle included in the second position; and a distance that is the target distance closest to the fourth distance included in the second position. Specifically, the target azimuth angle is the azimuth angle included in the preset position when measuring the HRTF centered on the right ear position, which is the azimuth angle of the third sound source relative to the right ear position when measuring the HRTF centered on the right ear position; the target pitch angle is the pitch angle included in the preset position when measuring the HRTF centered on the right ear position, which is the pitch angle of the third sound source relative to the right ear position when measuring the HRTF centered on the right ear position; and the target distance is the distance included in the preset position when measuring the HRTF centered on the right ear position, which is the distance of the third sound source relative to the right ear position when measuring the HRTF centered on the right ear position. In other words, the fourth preset position is the position where the third sound source is placed when measuring multiple HRTFs, meaning that the HRTF centered on the right ear position for each fourth preset position has been measured beforehand.
[0308] It is understandable that the methods for determining the azimuth of the fourth preset position if the fourth azimuth included in the second position is located between the two target azimuths, the methods for determining the pitch angle included in the fourth preset position if the fourth pitch angle included in the second position is located between the two target pitch angles, and the methods for determining the pitch angle included in the fourth preset position if the fourth pitch angle included in the second position is located between the two target pitch angles are the same as the description of the first preset position associated with the first position, and will not be repeated here.
[0309] For example, if the fourth azimuth angle of the second position of the nth second virtual speaker relative to the right ear position measured in step S401 is 88°, the fourth pitch angle is 46°, and the fourth distance is 1.02m, the correspondence between multiple preset positions and multiple HRTFs centered on the right ear position includes (90°, 45°, 1m) corresponding to the HRTF centered on the right ear position, (85°, 45°, 1m) corresponding to the HRTF centered on the right ear position, (90°, 50°, 1m) corresponding to the HRTF centered on the right ear position, (85°, 50°, 1m) corresponding to the HRTF centered on the right ear position, (90°, 45°, 1.02 ... The HRTFs centered on the right ear position are (85°, 45°, 1.1m) and (90°, 50°, 1.1m). Since 88° is between 85° and 90° but closer to 90°, 46° is between 45° and 50° but closer to 45°, and 1.02m is between 1m and 1.1m but closer to 1m, (90°, 45°, 1m) is determined as the fourth preset position n associated with the second position of the nth second virtual speaker relative to the right ear position.
[0310] After determining the N fourth preset positions associated with the N second positions, the N HRTFs centered on the right ear position corresponding to the N fourth preset positions are the N second HRTFs. For example, in the above example, in the third correspondence, the HRTF centered on the right ear position corresponding to (90°, 45°, 1m) is the HRTF centered on the right ear position corresponding to the second position of the nth second virtual speaker relative to the right ear position. That is to say, in the third correspondence, the HRTF centered on the right ear position corresponding to the fourth preset position n (90°, 45°, 1m) is the nth first HRTF, or the first HRTF corresponding to the nth first virtual speaker.
[0311] In this embodiment, the N second HRTFs are the N HRTFs centered on the right ear position obtained by actual measurement. The N second HRTFs obtained best represent the HRTFs corresponding to the current right ear position of the N second audio signals transmitted to the listener, so that the signal transmitted to the right ear position is optimal.
[0312] Next, regarding Figure 4 The second acquisition process of the N second HRTFs in step S102 of the illustrated embodiment will be described. Figure 12 Flowchart of the audio processing method provided in the embodiments of this application Figure 6 See Figure 12 The method in this embodiment includes:
[0313] Step S601: Obtain N fifth positions of the second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0314] Step S602: Determine N sixth positions based on N fifth positions. Each of the N fifth positions corresponds one-to-one with the N sixth positions. A sixth position and its corresponding fifth position share the same pitch angle and distance. The sum of the azimuth angle and the second value of the sixth position is the second azimuth angle of the corresponding fifth position. The second value is the difference between the third angle and the second angle. The second angle is the angle between the second straight line and the first plane. The third angle is the angle between the third straight line and the first plane. The second straight line is the line passing through the current head center and the origin of the coordinate system. The third straight line is the line passing through the current right ear and the origin of the coordinate system. The first plane is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0315] Step S603: Based on the N sixth positions and the second correspondence, determine the N HRTFs corresponding to the N sixth positions as N second HRTFs; the second correspondence is the correspondence between multiple preset positions and multiple HRTFs centered on the head center, which is stored in advance.
[0316] Specifically, for step S601, obtaining the fifth position of each second virtual speaker relative to the listener's head center, if there are N second virtual speakers, then N fifth positions will be obtained. The current head center is the current listener's head center.
[0317] Each fifth position includes a second pitch angle and a second azimuth angle of the corresponding second virtual speaker relative to the current head center, as well as a second distance between the second virtual speaker and the current head center.
[0318] For step S602, for each fifth position, the second pitch angle included in the fifth position is used as the pitch angle included in the corresponding sixth position, the second distance included in the fifth position is used as the distance included in the corresponding sixth position, and the second azimuth angle included in the fifth position is subtracted from the second value to obtain the azimuth angles included in the M sixth positions corresponding to the fifth position. For example, if the fifth position is (52°, 73°, 0.5m) and the second value is 6°, then the sixth position is (46°, 73°, 0.5m).
[0319] The three-dimensional coordinate system in this embodiment is the same as the three-dimensional coordinate system corresponding to the audio receiver mentioned above.
[0320] For step S603, prior to step S603, it is necessary to obtain the correspondence between multiple preset locations and multiple head-centered HRTFs. The method for obtaining the correspondence between multiple preset locations and multiple head-centered HRTFs is described in [reference needed]. Figure 4 The descriptions in the illustrated embodiments will not be repeated in this embodiment.
[0321] Specifically, based on the N sixth positions and the second correspondence, the N HRTFs corresponding to the N sixth positions are determined to be the N second HRTFs. This second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center, including:
[0322] Based on the N sixth positions, determine the N fifth preset positions; the N fifth preset positions are the preset positions in the second correspondence relationship;
[0323] Based on the second correspondence, N HRTFs centered on the head center are determined for the N fifth preset positions, which are the N second HRTFs.
[0324] The explanation of the fifth preset position associated with the sixth position is the same as that of the second preset position associated with the fourth position, and will not be repeated here.
[0325] After determining the N fifth preset positions associated with the N sixth positions, the N head-centered HRTFs corresponding to the N fifth preset positions are the N second HRTFs. For example, if a certain sixth position is associated with a fifth preset position (40°, 60°, 0.5m), then the head-centered HRTF corresponding to (40°, 60°, 0.5m) in the second correspondence is the head-centered HRTF corresponding to that sixth position. In other words, the head-centered HRTF corresponding to (30°, 60°, 0.5m) in the second correspondence is one of the N second HRTFs.
[0326] In this embodiment, the N second HRTFs are obtained by converting the HRTF centered on the head, which is highly efficient in obtaining the second HRTFs.
[0327] Next, regarding Figure 6 The third acquisition process of N second HRTFs in step S102 of the illustrated embodiment will be described. Figure 13 Flowchart of the audio processing method provided in the embodiments of this application Figure 7 See Figure 13 The method in this embodiment includes:
[0328] Step S701: Obtain N fifth positions of the second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0329] Step S702: Determine N eighth positions based on N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position include the same pitch angle and the same distance. The sum of the azimuth angle included in an eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position.
[0330] Step S703: Based on the N eighth positions and the second correspondence, determine the N HRTFs corresponding to the N eighth positions as N second HRTFs. The second correspondence is the correspondence between multiple preset positions and multiple HRTFs centered on the head center, which is stored in advance.
[0331] Specifically, step S701 in this embodiment refers to Figure 12 Step S601 in the embodiment will not be repeated here.
[0332] For step S702, the three-dimensional coordinate system in this embodiment is the three-dimensional coordinate system corresponding to the audio receiver mentioned above.
[0333] Specifically, for each fifth position, the second pitch angle included in that fifth position is taken as the pitch angle included in the corresponding eighth position, the second distance included in that fifth position is taken as the distance included in the corresponding eighth position, and the second azimuth angle included in that fifth position is subtracted from the first preset value to obtain the azimuth angle included in the corresponding eighth position. For example, if the fifth position is (52°, 73°, 0.5m) and the first preset value is 5°, then the eighth position is (47°, 73°, 0.5m).
[0334] The first preset value is pre-set and does not consider the size of the listener's head. In the previous embodiment, the second value is the difference between the third angle and the second angle, taking into account the size of the current listener's head. Optionally, the first preset value and... Figure 6 The first preset angle described in the illustrated embodiments is the same.
[0335] Before step S703, it is necessary to obtain the correspondence between multiple preset locations and multiple head-centered HRTFs. The method for obtaining the correspondence between multiple preset locations and multiple head-centered HRTFs is described in [reference needed]. Figure 6 The descriptions in the illustrated embodiments will not be repeated in this embodiment.
[0336] Specifically, based on the N eighth positions and the second correspondence, the N HRTFs corresponding to the N eighth positions are determined to be the N second HRTFs. This second correspondence is a pre-stored correspondence between multiple preset positions and multiple HRTFs centered on the head center, including:
[0337] Based on the N eighth positions, determine the N sixth preset positions associated with the N eighth positions; the N sixth preset positions are preset positions in the second correspondence relationship;
[0338] Based on the second correspondence, determine the N HRTFs centered on the head center corresponding to the sixth preset position, which are the N second HRTFs.
[0339] Specifically, the explanation of the sixth preset position associated with the eighth position is the same as that of the second preset position associated with the fourth position, and will not be repeated here.
[0340] After determining the N sixth preset positions associated with the N eighth positions, the HRTFs centered on the head for each of the N sixth preset positions are the N second HRTFs. For example, if the sixth preset position associated with a certain eighth position is (45°, 60°, 0.5m), then in the second correspondence, the HRTF centered on the head for (45°, 60°, 0.5m) is the HRTF centered on the head for that eighth position. In other words, in the second correspondence, the HRTF centered on the head for (45°, 60°, 0.5m) is one of the second HRTFs.
[0341] In this embodiment, the N second HRTFs are obtained by converting the HRTF centered on the head, and when obtaining the eighth position mentioned above, the size of the current listener's head is not considered, which further improves the efficiency of obtaining the second HRTFs.
[0342] pass Figures 6 to 13 The illustrated embodiment describes the process of obtaining M first HRTFs and N second HRTFs. Wherein, Figure 6 , Figure 8 , Figure 9 The method and its embodiment shown in any of the embodiments are related to Figure 10 , Figure 12 , Figure 13 The methods shown in any of the embodiments are used in combination.
[0343] Furthermore, the positions of the M first virtual speakers relative to the aforementioned coordinate origin and the positions of the N second virtual speakers relative to the aforementioned coordinate origin can be obtained in the following manner. It is understood that the positions of the M first virtual speakers relative to the aforementioned coordinate origin and the positions of the N second virtual speakers relative to the aforementioned coordinate origin are obtained before step S101.
[0344] First, the method for obtaining the position of the first virtual speaker relative to the aforementioned coordinate origin will be explained;
[0345] Figure 14 Flowchart of the audio processing method provided in the embodiments of this application Figure 8 See Figure 14 The method in this embodiment includes:
[0346] Step S801: Obtain the target virtual speaker group, which includes M target virtual speakers;
[0347] Step S802: Based on the M ninth positions of the M target virtual speakers relative to the origin, determine the M tenth positions of the M first virtual speakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position.
[0348] Specifically, in step S801, the audio signal receiving end renders and processes to obtain a target virtual speaker group, which includes M target virtual speakers.
[0349] For step S802, determining the M tenth positions of the M first virtual speakers relative to the origin based on the M ninth positions of the M target virtual speakers relative to the origin includes:
[0350] For each ninth position, the pitch angle included in the ninth position is taken as the pitch angle of the corresponding tenth position, the second distance included in the ninth position is taken as the distance included in the corresponding tenth position, and the azimuth angle included in the ninth position is added to the second preset value to take as the azimuth angle included in the corresponding tenth position.
[0351] For example, if the ninth position is (40°, 90°, 0.8m) and the second preset value is 5°, then the tenth position is (45°, 90°, 0.8m).
[0352] It is understandable that after obtaining the tenth position of the first virtual speaker relative to the origin, as shown in Formula 1, M first audio signals can be obtained based on the M tenth positions of the first virtual speaker relative to the origin.
[0353] In other words, the process of obtaining the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers includes: processing the audio signal to be processed according to the M tenth positions of the M first virtual speakers relative to the origin of the coordinate system to obtain the M first audio signals.
[0354] Secondly, the method for obtaining the position of the second virtual speaker relative to the above-mentioned coordinate origin will be explained; Figure 15 Flowchart of the audio processing method provided in the embodiments of this application Figure 9 See Figure 15 The method in this embodiment includes:
[0355] Step S901: Obtain the target virtual speaker group, which includes M target virtual speakers;
[0356] Step S902: Based on the M ninth positions of the M target virtual speakers relative to the origin, determine the N eleventh positions of the N second virtual speakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, and an eleventh position and its corresponding ninth position have the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position.
[0357] Specifically, in step S901, the audio signal receiving end renders and processes to obtain a target virtual speaker group, which includes M or N target virtual speakers, where M = N.
[0358] For step S902, based on the M ninth positions of the M target virtual speakers relative to the origin, determine the N eleventh positions of the N second virtual speakers relative to the origin, including:
[0359] For each ninth position, the pitch angle included in the ninth position is taken as the pitch angle of the corresponding eleventh position, the second distance included in the ninth position is taken as the distance included in the corresponding eleventh position, and the azimuth angle included in the ninth position is subtracted from the second preset value to obtain the azimuth angle included in the corresponding eleventh position.
[0360] For example, if the ninth position is (40°, 90°, 0.8m) and the second preset value is 5°, then the eleventh position is (35°, 90°, 0.8m).
[0361] It is understandable that after obtaining the eleventh position of the second virtual speaker relative to the origin, N second audio signals can be obtained based on the N eleventh positions of the second virtual speaker relative to the origin, as shown in Formula 2.
[0362] In other words, the acquisition of N second audio signals after the audio signal to be processed has been processed by N second virtual speakers includes: processing the audio signal to be processed according to the N eleventh positions of the N second virtual speakers relative to the origin of the coordinate system to obtain N second audio signals.
[0363] The following describes the effect of the audio processing method of this application in practical applications.
[0364] Figure 16 This is a spectrum showing the difference between the rendered spectrum of the signal corresponding to the left ear position in the existing technology and the theoretical spectrum corresponding to the left ear position. Figure 17 This is a spectrum showing the difference between the rendered spectrum of the signal corresponding to the right ear position in the existing technology and the theoretical spectrum corresponding to the right ear position. Figure 18 The difference spectrum between the rendered spectrum of the rendered signal corresponding to the left ear position and the theoretical spectrum corresponding to the left ear position in the method provided for implementation of this application. Figure 19 The difference spectrum between the rendered spectrum of the rendered signal corresponding to the right ear position and the theoretical spectrum corresponding to the right ear position in the method provided for implementation of this application.
[0365] Figures 16-19 In the middle graph, lighter colors indicate that the rendered spectrum is closer to the theoretical spectrum, while darker colors indicate a greater difference between the rendered spectrum and the theoretical spectrum. (Comparison) Figure 16 and Figure 18 Therefore, Figure 18 The area of the lighter-colored region is significantly larger than that of the medium-colored region. Figure 16 The area of the lighter-colored region indicates that the signal rendered by the method in this embodiment of the application is closer to the theoretically obtained signal, meaning the rendered signal effect is better. (Comparison) Figure 17 and Figure 19 It can be seen that, Figure 19 The area of the lighter-colored region is significantly larger than that of the medium-colored region. Figure 17 The area of the lighter-colored region indicates that the closer the signal rendered by the method of this embodiment of the application is to the signal corresponding to the right ear position, the better the rendered signal effect.
[0366] The above description of the functions implemented by the audio signal receiver has introduced the solutions provided in the embodiments of this application. It is understood that, in order to implement the above functions, the audio signal receiver includes corresponding hardware structures and / or software modules for executing each function. Based on the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of this application.
[0367] This application embodiment can divide the audio signal receiver into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0368] Figure 20 This is a schematic diagram of the structure of the audio processing device provided in the embodiments of this application; see also Figure 20 The device in this embodiment includes: a processing module 31 and an acquisition module 32;
[0369] Processing module 31 is used to acquire M first audio signals after the audio signal to be processed has been processed by M first virtual speakers and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers;
[0370] The acquisition module 32 is used to acquire M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers.
[0371] The acquisition module 32 is further configured to acquire a first target audio signal based on the M first audio signals and the M first HRTFs; and to acquire a second target audio signal based on the N second audio signals and the N second HRTFs.
[0372] The apparatus in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be repeated here.
[0373] In one possible design, the acquisition module 32 is specifically used for:
[0374] The M first audio signals are convolved with their corresponding first HRTFs to obtain M first convolved audio signals.
[0375] The first target audio signal is obtained based on the M first convolutional audio signals.
[0376] In one possible design, the acquisition module 32 is specifically used for:
[0377] The N second audio signals are convolved with their corresponding second HRTF signals to obtain N second convolved audio signals;
[0378] The second target audio signal is obtained based on the N second convolutional audio signals.
[0379] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0380] Obtain the M first positions of the M first virtual speakers relative to the current left ear position;
[0381] Based on the M first positions and the corresponding relationship, the M HRTFs corresponding to the M first positions are determined as the M first HRTFs.
[0382] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0383] Obtain the N second positions of the N second virtual speakers relative to the current right ear position;
[0384] Based on the N second positions and the corresponding relationship, the N HRTFs corresponding to the N second positions are determined as the N second HRTFs.
[0385] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0386] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0387] M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0388] Based on the M fourth positions and the corresponding relationship, the M HRTFs corresponding to the M fourth positions are determined to be the M first HRTFs.
[0389] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0390] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0391] N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position include the same pitch angle and the same distance. The sum of the azimuth angle included in the sixth position and the second value is the second azimuth angle included in the corresponding fifth position. The second value is the difference between the third included angle and the second included angle. The second included angle is the angle between the second straight line and the first surface. The third included angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system.
[0392] Based on the N sixth positions and the corresponding relationship, the N HRTFs corresponding to the N sixth positions are determined as the N second HRTFs.
[0393] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0394] Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker;
[0395] M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position.
[0396] Based on the M seventh positions and the corresponding relationship, the M HRTFs corresponding to the M seventh positions are determined to be the M first HRTFs.
[0397] In one possible design, multiple preset locations and multiple HRTFs are pre-stored; the acquisition module 32 is specifically used for:
[0398] Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker;
[0399] N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position.
[0400] Based on the N eighth positions and the corresponding relationship, the N HRTFs corresponding to the N eighth positions are determined as the N second HRTFs.
[0401] In one possible design, the acquisition module 32 is further configured to: before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers,
[0402] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the M first virtual speakers;
[0403] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0404] The processing module 32 is specifically used to: process the audio signal to be processed according to the M tenth positions to obtain the M first audio signals.
[0405] In one possible design, M = N, and the acquisition module 32 is further configured to: before the N second audio signals processed by the N second virtual speakers:
[0406] Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers;
[0407] Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual loudspeakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position;
[0408] The processing module 32 is specifically used to: process the audio signal to be processed according to the N eleventh positions to obtain the N second audio signals.
[0409] In one possible design, the M first virtual speakers are speakers in a first speaker group, and the N second virtual speakers are speakers in a second speaker group, wherein the first speaker group and the second speaker group are two independent speaker groups; or...
[0410] The M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group. The first speaker group and the second speaker group are the same speaker group, and M = N.
[0411] The apparatus in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be repeated here.
[0412] This application provides a computer-readable storage medium storing instructions that, when executed, cause a computer to perform the method described in the above-described method embodiments of this application.
[0413] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0414] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0415] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.
Claims
1. An audio processing method, characterized in that, include: The system acquires M first audio signals after the audio signal to be processed has been processed by M first virtual speakers, and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals, and the N second virtual speakers correspond one-to-one with the N second audio signals. M and N are positive integers, and the audio signal to be processed is an Ambisonic signal; Obtain M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers. Here, "centered on the left ear position" means that the left ear position is the center of the HRTF measurement, and "centered on the right ear position" means that the right ear position is the center of the HRTF measurement. Multiple preset locations and their corresponding HRTFs are pre-stored; The step of obtaining the M first HRTFs is as follows: obtaining the M first positions of the M first virtual speakers relative to the current left ear position; and determining the M HRTFs corresponding to the M first positions as the M first HRTFs based on the M first positions and the correspondence. The step of obtaining N second HRTFs involves: obtaining N second positions of the N second virtual speakers relative to the current right ear position; and determining the N HRTFs corresponding to the N second positions as the N second HRTFs based on the N second positions and the correspondence. A first target audio signal is obtained based on the M first audio signals and the M first HRTFs; a second target audio signal is obtained based on the N second audio signals and the N second HRTFs. The M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group. The first speaker group and the second speaker group are two independent speaker groups.
2. The method according to claim 1, characterized in that, The step of obtaining the first target audio signal based on the M first audio signals and the M first HRTFs includes: The M first audio signals are convolved with their corresponding first HRTFs to obtain M first convolved audio signals. The first target audio signal is obtained based on the M first convolutional audio signals.
3. The method according to claim 1 or 2, characterized in that, The second target audio signal is obtained based on N second audio signals and N second HRTFs; The N second audio signals are convolved with their corresponding second HRTF signals to obtain N second convolved audio signals; The second target audio signal is obtained based on the N second convolutional audio signals.
4. The method according to claim 1 or 2, characterized in that, The process of obtaining M first HRTFs also includes: Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker; M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system. Based on the M fourth positions and the corresponding relationship, the M HRTFs corresponding to the M fourth positions are determined to be the M first HRTFs.
5. The method according to claim 1 or 2, characterized in that, The process of obtaining N second HRTFs also includes: Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker; N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle of the sixth position and the second value is the second azimuth angle of the corresponding fifth position. The second value is the difference between the third angle and the second angle. The second angle is the angle between the second straight line and the first surface. The third angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is a plane formed by the X-axis and Z-axis of the three-dimensional coordinate system. Based on the N sixth positions and the corresponding relationship, the N HRTFs corresponding to the N sixth positions are determined as the N second HRTFs.
6. The method according to claim 1 or 2, characterized in that, The process of obtaining M first HRTFs also includes: Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker; M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position. Based on the M seventh positions and the corresponding relationship, the M HRTFs corresponding to the M seventh positions are determined to be the M first HRTFs.
7. The method according to claim 1 or 2, characterized in that, The process of obtaining N second HRTFs also includes: Obtain the N fifth positions of the N second virtual speakers relative to the current head center; the fifth positions include This includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker; N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position. Based on the N eighth positions and the corresponding relationship, the N HRTFs corresponding to the N eighth positions are determined as the N second HRTFs.
8. The method according to claim 1 or 2, characterized in that, Before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers, the process further includes: Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the M first virtual speakers; Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position; The process of obtaining the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers includes: processing the audio signal to be processed according to the M tenth positions to obtain the M first audio signals.
9. The method according to claim 1 or 2, characterized in that, Before obtaining the N second audio signals after the audio signal to be processed by the N second virtual speakers, M=N, the method further includes: Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers; Based on the M ninth positions of the M target virtual speakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual speakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position; The acquisition of N second audio signals after the audio signal to be processed has been processed by N second virtual speakers includes: The audio signal to be processed is processed according to the N eleventh positions to obtain the N second audio signals.
10. An audio processing apparatus, characterized in that, include: The processing module is used to acquire M first audio signals after the audio signal to be processed has been processed by M first virtual speakers and N second audio signals after the audio signal to be processed has been processed by N second virtual speakers; the M first virtual speakers correspond one-to-one with the M first audio signals and the N second virtual speakers correspond one-to-one with the N second audio signals; M and N are positive integers, and the audio signal to be processed is an Ambisonic signal; The acquisition module is used to acquire M first head-related transfer functions (HRTFs) and N second HRTFs. The M first HRTFs are all HRTFs centered on the left ear position, and the N second HRTFs are all HRTFs centered on the right ear position. The M first HRTFs correspond one-to-one with the M first virtual speakers, and the N second HRTFs correspond one-to-one with the N second virtual speakers. The "centered on the left ear position" means that the left ear position is the center of the HRTF measurement, and the "centered on the right ear position" means that the right ear position is the center of the HRTF measurement. The acquisition module is further configured to acquire a first target audio signal based on the M first audio signals and the M first HRTFs; and acquire a second target audio signal based on the N second audio signals and the N second HRTFs. The acquisition module is specifically used to: acquire M first positions of the M first virtual speakers relative to the current left ear position; and determine the M HRTFs corresponding to the M first positions as the M first HRTFs based on the M first positions and the corresponding relationship, wherein the corresponding relationship is a pre-stored correspondence between multiple preset positions and multiple HRTFs. The acquisition module is specifically used to: acquire N second positions of the N second virtual speakers relative to the current right ear position; and determine the N HRTFs corresponding to the N second positions as the N second HRTFs based on the N second positions and the correspondence. The M first virtual speakers are speakers in the first speaker group, and the N second virtual speakers are speakers in the second speaker group. The first speaker group and the second speaker group are two independent speaker groups.
11. The apparatus according to claim 10, characterized in that, The acquisition module is specifically used for: The M first audio signals are convolved with their corresponding first HRTFs to obtain M first convolved audio signals. The first target audio signal is obtained based on the M first convolutional audio signals.
12. The apparatus according to claim 10 or 11, characterized in that, The acquisition module is specifically used for: The N second audio signals are convolved with their corresponding second HRTF signals to obtain N second convolved audio signals; The second target audio signal is obtained based on the N second convolutional audio signals.
13. The apparatus according to claim 10 or 11, characterized in that, The acquisition module is further used for: Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker; M fourth positions are determined based on the M third positions, with each of the M third positions corresponding one-to-one with the M fourth positions. A fourth position and its corresponding third position share the same pitch angle and distance. The difference between the azimuth angle of the fourth position and a first value is the first azimuth angle of the corresponding third position. The first value is the difference between a first included angle and a second included angle. The first included angle is the angle between a first straight line and a first surface, and the second included angle is the angle between a second straight line and the first surface. The first straight line is a straight line passing through the current left ear and the origin of the three-dimensional coordinate system, and the second straight line is a straight line passing through the current head center and the origin of the coordinate system. The first surface is the plane formed by the X-axis and Z-axis of the three-dimensional coordinate system. Based on the M fourth positions and their corresponding relationships, the M HRTFs corresponding to the M fourth positions are determined as the M first HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
14. The apparatus according to claim 10 or 11, characterized in that, The acquisition module is further used for: Obtain N fifth positions of the N second virtual speakers relative to the current head center; the fifth position includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker; N sixth positions are determined based on the N fifth positions, and the N fifth positions correspond one-to-one with the N sixth positions. A sixth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle of the sixth position and the second value is the second azimuth angle of the corresponding fifth position. The second value is the difference between the third angle and the second angle. The second angle is the angle between the second straight line and the first surface. The third angle is the angle between the third straight line and the first surface. The second straight line is a straight line passing through the current head center and the origin of the coordinate system. The third straight line is a straight line passing through the current right ear and the origin of the coordinate system. The first surface is a plane formed by the X-axis and Z-axis of the three-dimensional coordinate system. Based on the N sixth positions and their corresponding relationships, the N HRTFs corresponding to the N sixth positions are determined as the N second HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
15. The apparatus according to claim 10 or 11, characterized in that, The acquisition module is further used for: Obtain M third positions of the M first virtual speakers relative to the current head center; the third position includes a first azimuth angle and a first pitch angle of the first virtual speaker relative to the current head center, and a first distance between the current head center and the first virtual speaker; M seventh positions are determined based on the M third positions. The M third positions correspond one-to-one with the M seventh positions. A seventh position and its corresponding third position have the same pitch angle and the same distance. The difference between the azimuth angle included in the seventh position and the first preset value is the first azimuth angle included in the corresponding third position. Based on the M seventh positions and their corresponding relationships, the M HRTFs corresponding to the M seventh positions are determined as the M first HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
16. The apparatus according to claim 10 or 11, characterized in that, The system pre-stores the correspondence between multiple preset locations and multiple HRTFs; the acquisition module is further specifically used for: Obtain the N fifth positions of the N second virtual speakers relative to the current head center; the fifth positions include This includes the second azimuth angle and the second pitch angle of the second virtual speaker relative to the current head center, and the second distance between the current head center and the second virtual speaker; N eighth positions are determined based on the N fifth positions. The N fifth positions correspond one-to-one with the N eighth positions. An eighth position and its corresponding fifth position have the same pitch angle and the same distance. The sum of the azimuth angle included in the eighth position and the first preset value is the second azimuth angle included in the corresponding fifth position. Based on the N eighth positions and their corresponding relationships, the N HRTFs corresponding to the N eighth positions are determined as the N second HRTFs. The corresponding relationships are pre-stored relationships between multiple preset positions and multiple HRTFs.
17. The apparatus according to claim 10 or 11, characterized in that, The acquisition module is further configured to: before acquiring the M first audio signals after the audio signal to be processed has been processed by the M first virtual speakers, Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the M first virtual speakers; Based on the M ninth positions of the M target virtual loudspeakers relative to the origin of the three-dimensional coordinate system, the M tenth positions of the M first virtual loudspeakers relative to the origin are determined; wherein, the M ninth positions correspond one-to-one with the M tenth positions, and a tenth position and its corresponding ninth position include the same pitch angle and the same distance, and the difference between the azimuth angle included in the tenth position and the second preset value is the azimuth angle included in the corresponding ninth position; The processing module is specifically used to: process the audio signal to be processed according to the M tenth positions. The M first audio signals are obtained.
18. The apparatus according to claim 10 or 11, characterized in that, M=N, the acquisition module is further configured to: before acquiring the N second audio signals after the audio signal to be processed has been processed by the N second virtual speakers: Obtain a target virtual speaker group, which includes M target virtual speakers, and the M target virtual speakers correspond one-to-one with the N second virtual speakers; Based on the M ninth positions of the M target virtual speakers relative to the origin of the three-dimensional coordinate system, determine the N eleventh positions of the N second virtual speakers relative to the origin; wherein, the M ninth positions correspond one-to-one with the N eleventh positions, an eleventh position and its corresponding ninth position include the same pitch angle and the same distance, and the sum of the azimuth angle included in the eleventh position and the second preset value is the azimuth angle included in the corresponding ninth position; The processing module is specifically used to: process the audio signal to be processed according to the N eleventh positions to obtain the N second audio signals.
19. An audio processing apparatus, characterized in that, Including processor and memory; The memory is used to store computer instructions; The processor is used to read and execute instructions in the memory to implement the method as described in any one of claims 1-9.
20. A readable storage medium, characterized in that, The readable storage medium stores a computer program; when the computer program is executed, it implements the method as described in any one of claims 1-9.
21. A computer program product, characterized in that, The computer program product stores a computer program; when the computer program is executed, it implements the method as described in any one of claims 1-9.
Citation Information
Patent Citations
Audio signal processing method and apparatus, and terminal
CN108156575A