Sound signal processing method, apparatus, device, and storage medium

By decomposing and reconstructing the spatial information of the stereo frequency domain signal and using the head-related transfer function to simulate the transmission of virtual sound sources, the problem of insufficient spatiality in headphone playback is solved, achieving a stronger sense of space and virtual surround sound experience.

CN116261086BActive Publication Date: 2026-01-20SHENZHEN BLUETRUM TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211101303.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2026-01-20
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

When headphones play stereo sound, they cannot effectively create a sound field, resulting in insufficient spatiality of the stereo sound and an inability to produce an immersive experience.

Method used

By decomposing the stereo frequency domain signal into spatial information, multiple virtual sound signals are obtained. The spatial information is then reconstructed using the head-related transfer function to simulate the transmission process of the virtual sound source and generate rich spatial information signals.

Benefits of technology

It enhances the spatial sense and virtual surround sound of stereo playback, bringing users a more immersive listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116261086B_ABST
    Figure CN116261086B_ABST
Patent Text Reader

Abstract

The application provides a sound signal processing method, device and equipment and a storage medium. The method comprises the following steps: obtaining a to-be-processed stereo frequency domain signal; performing spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations; performing spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein the head-related transfer function corresponding to a target virtual sound signal is used to represent the phase and frequency response of the target virtual sound signal from the spatial orientation corresponding to the target virtual sound signal to the head, and the target virtual sound signal is any one of the plurality of virtual sound signals. The technical solution makes the stereo sound have rich spatial information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of signal processing, in particular to a sound signal processing method and device, equipment and a storage medium. BACKGROUND

[0002] Stereo sound refers to sound with a sense of three-dimensionality. In daily life, the sound of nature heard by human ears is stereo sound, and the sound source has a certain spatial position. People can perceive the spatial position distribution of various sound sources by hearing, and thus feel the orientation of the sound source relative to the head.

[0003] In audio processing technology, a processing system composed of recording, transmission and playback systems reflects the spatial positions of different sound sources, so that people can obtain the spatial distribution impression of sound and produce a sense of immediacy and stereo sound when listening to audio. The sound obtained by processing through such a processing system is stereo sound. When playing stereo sound using earphones, the earphones cannot form a sound field like loudspeakers, and there is a problem of positioning the sound image inside the head. This results in insufficient spatiality of the stereo sound perceived by the ear, and cannot produce a sense of being there. SUMMARY

[0004] The present application provides a sound signal processing method, device, equipment and storage medium to solve the technical problem of insufficient spatiality of stereo sound played by earphones.

[0005] In a first aspect, a sound signal processing method is provided, comprising:

[0006] obtaining a stereo sound frequency domain signal to be processed;

[0007] performing spatial information decomposition on the stereo sound frequency domain signal to be processed to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations;

[0008] performing spatial information reconstruction on the stereo sound frequency domain signal to be processed according to the plurality of virtual sound signals and the head-related transfer functions corresponding to the plurality of virtual sound signals, to obtain a frequency domain output signal corresponding to the stereo sound frequency domain signal to be processed, wherein the head-related transfer function corresponding to a target virtual sound signal is used to represent the phase and frequency response of the target virtual sound signal from the spatial orientation corresponding to the target virtual sound signal to the head, and the target virtual sound signal is any one of the plurality of virtual sound signals.

[0009] In the technical solution, after obtaining the to-be-processed stereo frequency domain signal, spatial information decomposition is performed on the to-be-processed stereo frequency domain signal to obtain virtual sound signals at multiple different spatial positions, so that the virtual sound sources at different positions can be simulated; then, spatial information reconstruction is performed on the to-be-processed stereo frequency domain signal according to the virtual sound signals at multiple different spatial positions and the head-related transfer functions corresponding to the virtual sound signals at multiple different spatial positions, so that the transmission process of the virtual sound sources at different positions to the human ear can be simulated and reproduced, and therefore, the frequency domain output signal obtained through reconstruction has rich spatial information and can bring stronger virtual surround and spatial sense to the user when output.

[0010] With reference to the first aspect, in a possible implementation manner, the to-be-processed stereo frequency domain signal includes a left channel frequency domain signal and a right channel frequency domain signal; and the spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain virtual sound signals corresponding to multiple spatial positions respectively includes: performing spectrum analysis on the left channel frequency domain signal and the right channel frequency domain signal to determine a target position angle, the target position angle being a position angle at which the difference between signals is the largest; and performing signal superposition and signal decomposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain the multiple virtual sound signals. The signal superposition and signal decomposition are performed based on the position angle at which the difference between signals is the largest, so that the virtual sound signals obtained through decomposition are independent of each other, and the spatial characteristics of the virtual sound sources that can be simulated are more different, thereby bringing stronger spatial sense.

[0011] With reference to the first aspect, in a possible implementation manner, the multiple virtual sound signals include an illusion sound source signal, a left channel residual signal and a right channel residual signal; and the signal superposition and signal decomposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain the multiple virtual sound signals include: performing signal superposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain the illusion sound source signal; performing signal decomposition according to the target position angle, the illusion sound source signal and the right channel frequency domain signal to obtain the left channel residual signal; and performing signal decomposition according to the target position angle, the illusion sound source signal and the left channel frequency domain signal to obtain the right channel residual signal. The illusion sound source signal is obtained through signal superposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal, and the left channel residual signal and the right channel residual signal are obtained through signal decomposition according to the target position angle, the illusion sound source signal, the left channel frequency domain signal and the right channel frequency domain signal, so that the virtual sound source with higher precision can be obtained.

[0012] With reference to the first aspect, in a possible implementation manner, the plurality of virtual sound signals comprises an illusion sound source signal, a left channel residual signal and a right channel residual signal; and before the spatial information of the to-be-processed stereo frequency domain signal is reconstructed according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, the method further comprises: obtaining a first position angle and a second position angle, the first position angle being a position angle corresponding to a left virtual speaker, and the second position angle being a position angle corresponding to a right virtual speaker; determining a spatial direction corresponding to the left channel residual signal according to the first position angle, and determining a spatial direction corresponding to the right channel residual signal according to the second position angle; and determining a spatial direction corresponding to the illusion sound source signal according to the first position angle, the second position angle and a target position angle. By determining the spatial direction of each virtual sound source according to the target position angle and the position angles of the left and right virtual speakers, the spatial directions of the virtual sound sources can be made to be more different, thereby bringing stronger spatial sense.

[0013] With reference to the first aspect, in a possible implementation manner, the spatial information of the to-be-processed stereo frequency domain signal is reconstructed according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, comprising: determining a target in-ear signal according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, the target in-ear signal being used to indicate a mixed signal received by the head when the plurality of virtual sound signals are transmitted to the head; and mixing the target in-ear signal and the to-be-processed stereo frequency domain signal to obtain the frequency domain output signal. By mixing the signal when each virtual sound source is transmitted to the ear and the to-be-processed stereo frequency domain signal, the frequency domain output signal obtained by mixing can carry the spatial information of each virtual sound source, thereby having stronger spatial sense.

[0014] In a possible implementation manner of the first aspect, the determining the target in-ear signal according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals comprises: determining a virtual in-ear signal corresponding to the target virtual sound signal according to the target virtual sound signal and the head-related transfer function corresponding to the target virtual sound signal, the virtual in-ear signal corresponding to the target virtual sound signal being used to indicate a signal received by a head when the target virtual sound signal is transmitted to the head; and performing signal superposition on the virtual in-ear signals corresponding to the plurality of virtual sound signals to obtain the target in-ear signal. By performing signal superposition on the signals when each virtual sound source is transmitted to the ear, the transmission process of the plurality of virtual sound sources transmitted to the ear can be simulated, so that the target in-ear signal has stronger authenticity.

[0015] In a possible implementation manner of the first aspect, the mixing the target in-ear signal and the to-be-processed stereo frequency domain signal to obtain the frequency domain output signal comprises: obtaining a stereo low frequency signal corresponding to the to-be-processed stereo frequency domain signal, the stereo low frequency signal having a frequency lower than a preset frequency; and performing weighted sum modulation on the target in-ear signal, the to-be-processed stereo frequency domain signal, and the stereo low frequency signal to obtain the frequency domain output signal. By performing weighted sum modulation on the target in-ear signal, the to-be-processed stereo frequency domain signal, and the stereo low frequency signal, the mixing of the to-be-processed stereo frequency domain signal can be implemented, so that the frequency domain output signal has richer levels.

[0016] In a second aspect, a sound signal processing apparatus is provided, comprising:

[0017] The obtaining module is configured to obtain a to-be-processed stereo frequency domain signal.

[0018] The decomposing module is configured to perform spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations.

[0019] The reconstructing module is configured to perform spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein the head-related transfer function corresponding to a target virtual sound signal is used to represent a phase and a frequency response of the target virtual sound signal transmitted from a spatial orientation corresponding to the target virtual sound signal to a head, and the target virtual sound signal is any one of the plurality of virtual sound signals.

[0020] In a third aspect, an audio device is provided, comprising a memory connected to one or more processors, and the one or more processors are configured to execute one or more computer programs stored in the memory, and the one or more processors, when executing the one or more computer programs, cause the audio device to implement the sound signal processing method of the first aspect.

[0021] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions, when executed by a processor, cause the processor to execute the sound signal processing method of the first aspect.

[0022] The present application can achieve the following technical effects: by decomposing the spatial information of the to-be-processed stereo frequency-domain signal to obtain virtual sound signals at multiple different spatial orientations, the simulation of virtual sound sources at different orientations can be achieved; then, according to the virtual sound signals at multiple different spatial orientations and the head-related transfer functions corresponding to the virtual sound signals at multiple different spatial orientations, the spatial information of the to-be-processed stereo frequency-domain signal is reconstructed, which can simulate the transmission process of the virtual sound sources at different orientations to the human ear, so that the frequency-domain output signal obtained by reconstruction has rich spatial information, and can bring stronger virtual surround and spatial sense to the user when output. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 FIG. 1 is a flowchart of a sound signal processing method provided by an embodiment of the present application;

[0024] Figure 2 FIG. 2 is a structural diagram of an audio data processing apparatus provided by an embodiment of the present application;

[0025] Figure 3 FIG. 3 is a structural diagram of an audio device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0026] The technical solutions of the embodiments of the present application will be described below with reference to the drawings.

[0027] The technical solutions of the present application can be applied to an audio playing scene, and can be specifically applied to processing and outputting stereo sound, so that the output stereo sound has stronger spatial sense. The stereo sound can be sound recorded by stereo recording technology. For example, the sound can be recorded by two microphones in a specific arrangement manner. The technical solutions of the present application can be applied to earphones, loudspeakers, terminals (such as mobile phones) with audio players, and other audio devices for playing audio data.

[0028] The technical principle of the present application is as follows: before outputting the stereo sound, the input stereo sound signal is converted into a frequency domain, the input stereo sound is spatial information decomposed in the frequency domain to extract a virtual sound signal with different spatial information characteristics, the simulation of the virtual sound source is realized; then, the spatial information is reconstructed and restored in combination with the virtual sound source and the head related transfer function corresponding to the virtual sound source, the simulation and reproduction of the transmission process of the virtual sound source to the human ear are realized, so that the frequency domain signal with rich spatial information is obtained; finally, the frequency domain signal is converted into a time domain for output, so that the output stereo sound signal has stronger spatial sense and brings the user an immersive listening experience.

[0029] For the convenience of understanding, first, some concepts involved in the technical solutions of the present application are introduced.

[0030] The head related transfer function (HRTF) can also be called an anatomical transfer function (ATF), which describes the transmission process of sound waves from a sound source to the ears, and is used to express the results of the comprehensive filtering of the physiological structure (such as the head, auricle and torso, etc.) of the human body on the sound waves. The physical process of the sound waves emitted by the sound source reaching the ears after scattering through the physiological structure of the human body can be regarded as a linear time-invariant sound wave system, and the HRTF is the frequency domain transfer function of this sound wave system, which is defined as:

[0031]

[0032]

[0033] Wherein, HL and HR are the HRTF corresponding to the left ear and the right ear respectively, PL and PR are the complex sound pressure produced by the point sound source at the left and right ears of the listener respectively; Po is the complex sound pressure at the center of the head when the head does not exist; r is the distance from the sound source to the center of the head; θ is the horizontal azimuth angle of the sound source; is the vertical azimuth angle of the sound source; ω is the angular frequency of the sound wave; α is a set of parameters related to the physiological structure and size. The transmission path of the sound wave to the center of the head is different, the transmission medium of the sound wave to the human ear is different, and the HRTF is different; in the case that the center of the head is fixed, the sound source is in different spatial positions, and the corresponding transmission path is different, so the corresponding HRTF is different.

[0034] Next, the technical solutions of the present application are introduced, please refer to Figure 1 , Figure 1 The flowchart of a sound signal processing method provided by the embodiment of the present application, the method can be applied to various audio devices for playing audio data mentioned above, such asFigure 1 The method comprises the following steps:

[0035] S101, obtaining a stereo audio domain signal to be processed.

[0036] Here, the stereo audio domain signal to be processed refers to a stereo audio domain signal that needs to be processed. The stereo audio domain signal to be processed can be obtained by time-frequency conversion of a stereo time domain signal to be processed. The stereo time domain signal can be an original sound signal recorded by stereo recording technology, or a sound signal obtained by signal processing (such as editing, noise reduction, etc.) on the original sound signal recorded by stereo recording technology. The stereo time domain signal to be processed can refer to a stereo time domain signal that needs to be played in an audio device, that is, a stereo time domain signal input into an audio device; or can refer to a stereo time domain signal that needs to be processed in other scenarios. The stereo time domain signal to be processed can include a left channel time domain signal and a right channel time domain signal; correspondingly, the stereo audio domain signal to be processed can also include a left channel frequency domain signal and a right channel frequency domain signal, wherein the left channel frequency domain signal can be obtained by time-frequency conversion of the left channel time domain signal, and the right channel frequency domain signal can be obtained by time-frequency conversion of the right channel time domain signal.

[0037] In some possible scenarios, the stereo audio domain signal to be processed can be obtained by Fourier transform of the stereo time domain signal to be processed. The above obtaining the stereo audio domain signal to be processed specifically includes: obtaining the stereo time domain signal to be processed, and performing Fourier transform on the stereo time domain signal to be processed to obtain the stereo audio domain signal to be processed.

[0038] In order to facilitate spectral analysis and processing, in some possible embodiments, the stereo time domain signal to be processed can be processed by discrete Fourier transform to obtain the stereo audio domain signal to be processed. The above obtaining the stereo audio domain signal to be processed specifically includes: obtaining the stereo time domain signal to be processed, sampling the stereo time domain signal to be processed according to a preset sampling frequency to obtain a stereo time domain sampling signal to be processed; and performing discrete Fourier transform on the stereo time domain sampling signal to be processed to obtain the stereo audio domain signal to be processed. The preset sampling frequency and the number of points of the discrete Fourier transform can be set according to requirements. For example, the sampling frequency can be set to 44.1 kHz (HZ), the left channel time domain signal and the right channel time domain signal are sampled to obtain the left channel time domain sampling signal x l and the right channel time domain sampling signal x r , and then 512 sampling points are taken per frame to perform 512-point discrete Fourier transform on the left channel time domain sampling signal x l and the right channel time domain sampling signal x r to obtain the left channel time domain sampling signal x lX l a left channel time-domain sample signal x r X r a right channel time-domain sample signal x

[0039] Optionally, before performing the discrete Fourier transform on the to-be-processed stereo time-domain sample signal to obtain the to-be-processed stereo frequency-domain signal, the to-be-processed stereo time-domain sample signal can also be preprocessed by overlapping windowing to obtain a preprocessed stereo time-domain sample signal; and in the process of performing the discrete Fourier transform on the to-be-processed stereo time-domain sample signal to obtain the to-be-processed stereo frequency-domain signal, the preprocessed stereo time-domain sample signal is subjected to the discrete Fourier transform to obtain the to-be-processed stereo frequency-domain signal. For example, after obtaining the left channel time-domain sample signal x l and the right channel time-domain sample signal x r , 512 sample points can be extracted from the left channel time-domain sample signal x l and the right channel time-domain sample signal x r in each frame in a manner that there is 50% overlap between adjacent two frames, and a square root Hanning window is used to process each frame of the extracted left channel time-domain sample signal x l and the right channel time-domain sample signal x r to obtain preprocessed left channel time-domain sample signals and preprocessed right channel time-domain sample signals, and then the preprocessed left channel time-domain sample signals and the preprocessed right channel time-domain sample signals are subjected to a 512-point discrete Fourier transform to obtain the left channel frequency-domain signal and the right channel frequency-domain signal. By performing the overlapping windowing preprocessing on the to-be-processed stereo time-domain sample signal, the amplitude modulation of the to-be-processed stereo time-domain sample signal can be realized, and the spectral leakage can be reduced, so that the loss of the signal can be prevented, and the to-be-processed stereo frequency-domain signal can be better analyzed in the frequency domain.

[0040] Optionally, the time-frequency conversion of the to-be-processed stereo time-domain signal to obtain the to-be-processed stereo frequency-domain signal can also be realized by other manners, which is not limited in the present application.

[0041] S102, spatial information decomposition is performed on the to-be-processed stereo frequency-domain signal to obtain a plurality of virtual sound signals.

[0042] Here, the spatial information decomposition of the to-be-processed stereo frequency-domain signal refers to determining a sound source positioning clue in the stereo sound through a spectrum analysis manner of the to-be-processed stereo frequency-domain signal, and extracting a frequency-domain signal with different spatial information characteristics from the to-be-processed stereo frequency-domain signal based on the sound source positioning clue as a plurality of virtual sound signals. The plurality of virtual sound signals are respectively virtual sound signals corresponding to different spatial directions, and can be used to reflect sound signals in different directions in the stereo recording process. The spectrum analysis of the to-be-processed stereo frequency-domain signal can specifically include spectrum analysis of the left-channel frequency-domain signal and the right-channel frequency-domain signal.

[0043] In a feasible manner, the spatial information decomposition of the to-be-processed stereo frequency-domain signal can be implemented through the following steps A1-A2:

[0044] A1, performing spectrum analysis on the left-channel frequency-domain signal and the right-channel frequency-domain signal to determine a target position angle.

[0045] The target position angle is a position angle that maximizes the difference between signals; the target position angle can be solved under the condition that the plurality of virtual sound signals satisfy the mutual independent energy constraint condition. The solving formula of the target position angle is as follows:

[0046]

[0047]

[0048]

[0049]

[0050]

[0051] wherein X l (k) is the left-channel frequency-domain signal, X r (k) is the right-channel frequency-domain signal; γ b is the target position angle, and the variable ρ b represents the normalized correlation coefficient of the left-channel frequency-domain signal X l (k) and the right-channel frequency-domain signal X r (k) in the processing frequency band b, represents the energy of the left-channel frequency-domain signal X l (k) in the processing frequency band b, represents the energy of the left-channel frequency-domain signal X r (k) in the processing frequency band b, and R represents the real part.

[0052] The bandwidth of the processing frequency band b can be set based on the frequency band bandwidth of the left channel frequency domain signal and the right channel frequency domain signal. In some possible scenarios, after the left channel frequency domain signal and the right channel frequency domain signal are obtained through time-frequency conversion, the frequency band of each frame of the left channel frequency domain signal and the right channel frequency domain signal can be divided into a plurality of Bark domains according to critical frequency bands represented by a "Bark scale" (i.e., Bark domain), and then each Bark domain is taken as a processing frequency band b to determine the target position angle corresponding to each Bark domain. Since the human ear structure will resonate at about 24 frequency points, and the sound signal will also exhibit 24 critical frequency bands on the frequency band, the frequency band of each frame of the left channel frequency domain signal and the right channel frequency domain signal can be divided into 24 processing frequency bands for solving the target position angle. By dividing the frequency band of each frame of the left channel frequency domain signal and the right channel frequency domain signal into critical frequency bands to solve the target position angle, the calculation amount can be saved, and the processing efficiency can be improved. In addition, since the critical frequency band is the frequency spectrum corresponding to the frequency point at which the human ear structure resonates, by dividing the frequency band of each frame of the left channel frequency domain signal and the right channel frequency domain signal into critical frequency bands for spectral analysis, the target position angle obtained through analysis can also conform to the human auditory structure, that is, the target position angle is more accurate.

[0053] Specifically, after the spectrum of the left channel frequency domain signal and the right channel frequency domain signal on the processing frequency band b is obtained, the energy of the left channel frequency domain signal and the energy of the right channel frequency domain signal are calculated according to the above formula (4), respectively. l,b and σ r,b are obtained. b is calculated according to the above formula (3), then v b is calculated according to the above formula (2), and β b is calculated according to the above formula (5). b is calculated according to the above formula (1), and the target position angle γ b is calculated. By calculating the target position angle that maximizes the difference between the plurality of virtual sound signals, the virtual sound signals obtained through decomposition are mutually independent, and thus have stronger azimuth distinguishability.

[0054] A2, signal superposition and decomposition are performed according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal, to obtain a plurality of virtual sound signals.

[0055] Here, the signal superposition and decomposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain the plurality of virtual sound signals refers to extracting the phantom sound source signal representing the main signal and the residual signal representing the environment signal from the left channel frequency domain signal and the right channel frequency domain signal in a manner of signal superposition and / or signal decomposition. The main signal refers to the sound signal emitted by the main sound source in the stereo recording process, and the main sound source refers to the main sound source emitting the sound, which can be understood as the sound source mainly recorded in the stereo recording process. The main signal can be, for example, the sound emitted by an airplane taking off, the sound of a cannonball, and the like. The environment signal refers to the sound signal emitted by other sound sources in the stereo recording process, and the other sound sources refer to the environmental sound sources in the recording environment except the main sound source. The residual signal can specifically include a left channel residual signal and a right channel residual signal, which are respectively used to represent the environmental signals in the left and right directions. Thus, the plurality of virtual sound signals includes the phantom sound source signal, the left channel residual signal and the right channel residual signal.

[0056] Specifically, the left channel frequency domain signal and the right channel frequency domain signal can be subjected to signal superposition and decomposition to obtain the plurality of virtual sound signals through the following steps A21-A23:

[0057] A21, signal superposition is performed according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain the phantom sound source signal.

[0058] Specifically, the phantom sound source signal can be calculated based on the following formula (6):

[0059]

[0060] wherein S(k) is the phantom sound source signal, γ b is the target position angle, X l (k) is the left channel frequency domain signal, and X r (k) is the right channel frequency domain signal.

[0061] A22, signal decomposition is performed according to the target position angle, the phantom sound source signal and the right channel frequency domain signal to obtain the left channel residual signal.

[0062] Specifically, the left channel residual signal can be calculated based on the following formula (7):

[0063] D l (k) = sin(γ b )S(k) - X r (k) formula (7)

[0064] wherein D l (k) is the left channel residual signal, S(k) is the phantom sound source signal, γ b is the target position angle, and X r(k) is a right channel frequency domain signal.

[0065] A23, signal decomposition is performed according to the target position angle, the phantom sound source signal and the left channel frequency domain signal to obtain a right channel residual signal.

[0066] Specifically, the right channel residual signal can be calculated based on the following formula (8):

[0067] D r (k) = sin (γ b ) S (k) - X l (k) formula (8)

[0068] wherein, D r (k) is a right channel residual signal, S(k) is a phantom sound source signal, γ b is a target position angle, and X l (k) is a left channel frequency domain signal.

[0069] By performing signal superposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal to obtain a phantom sound source signal, and performing signal decomposition according to the target position angle, the phantom sound source signal, the left channel frequency domain signal and the right channel frequency domain signal to obtain a left channel residual signal and a right channel residual signal, a virtual sound source with more accurate position can be obtained. By performing signal superposition and decomposition based on the position angle with the largest difference degree among signals, a plurality of virtual sound signals can be obtained, which are independent of each other, and the spatial characteristics of the virtual sound sources simulated are more different, thereby bringing stronger spatial sense.

[0070] Optionally, the spatial information of the to-be-processed stereo frequency domain signal can also be decomposed to obtain a plurality of virtual sound signals by other manners, which is not limited in the present application.

[0071] S103, according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, spatial information of the to-be-processed stereo frequency domain signal is reconstructed to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal.

[0072] Here, according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, spatial information of the to-be-processed stereo frequency domain signal is reconstructed, which means that based on the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, the simulation of the virtual sound source signal entering the human ear and the reconstruction of the sound signal heard by the human ear are completed.

[0073] Specifically, the spatial information reconstruction on the to-be-processed stereo frequency domain signal can be performed through the following steps B1-B2 to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal:

[0074] B1, determining a target in-ear signal according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals.

[0075] The target in-ear signal is used to indicate a mixed signal received by a pair of ears at a head position due to the plurality of virtual sound signals. The mixed signal is obtained by mixing and superimposing signals received by the pair of ears at the head position due to the plurality of virtual sound signals. The target in-ear signal can include a left in-ear signal and a right in-ear signal. The left in-ear signal is used to indicate a mixed signal received by a left ear at a head position due to the plurality of virtual sound signals. The right in-ear signal is used to indicate a mixed signal received by a right ear at a head position due to the plurality of virtual sound signals.

[0076] In an available implementation, the target in-ear signal can be obtained through the following steps B11-B12:

[0077] B11, determining a virtual in-ear signal corresponding to a target virtual sound signal according to the target virtual sound signal and the head-related transfer function corresponding to the target virtual sound signal.

[0078] B12, performing signal superposition on the virtual in-ear signals corresponding to the plurality of virtual sound signals to obtain the target in-ear signal.

[0079] The virtual in-ear signal corresponding to the target virtual sound signal is used to indicate a signal received by a pair of ears at a head position due to the target virtual sound signal. The target virtual sound signal can be the aforementioned phantom sound source signal, left channel residual signal or right channel residual signal.

[0080] The calculation formula of the target in-ear signal can be seen in the following formulas (9)-(16):

[0081] Y l (k) = Y l1 (k) + Y l2 (k) - Y l3 (k) Formula (9)

[0082]

[0083]

[0084]

[0085] Y r (k) = Y r1 (k) + Y r2(k) -Y r3 (k) Formula (13)

[0086]

[0087]

[0088]

[0089] wherein Y l (k) is a left ear-in signal, Y l1 (k) is a left virtual ear-in signal corresponding to the phantom sound source signal, Y l2 (k) is a left virtual ear-in signal corresponding to the left channel residual signal, Y l3 (k) is a left virtual ear-in signal corresponding to the right channel residual signal; p l , b, a s , 0 is a left ear head-related transfer function corresponding to the phantom sound source signal, used to represent the phase and frequency response of the phantom sound source signal being transferred to the left ear from the direction a s , p l , b, a l , 0 is a left ear head-related transfer function corresponding to the left channel residual signal, used to represent the phase and frequency response of the left channel residual signal being transferred to the left ear from the direction a l , p l , b, a r , 0 is a head-related transfer function corresponding to the right channel residual signal, used to represent the phase and frequency response of the right channel residual signal being transferred to the left ear from the direction a r ; Yr(k) is a right ear-in signal, Yr1(k) is a right virtual ear-in signal corresponding to the phantom sound source signal, Yr2(k) is a right virtual ear-in signal corresponding to the right channel residual signal, Yr3(k) is a right virtual ear-in signal corresponding to the right channel residual signal; pr, b, a s , 0 is a right ear head-related transfer function corresponding to the phantom sound source signal, used to represent the phase and frequency response of the phantom sound source signal being transferred to the right ear from the direction a s , pr, b, a l , 0 is a right ear head-related transfer function corresponding to the left channel residual signal, used to represent the phase and frequency response of the left channel residual signal being transferred to the right ear from the direction a l , pr, b, a r , 0 is a head-related transfer function corresponding to the right channel residual signal, used to represent the phase and frequency response of the right channel residual signal being transferred to the right ear from the direction a r ; φb, a s , 0 is a phase compensation corresponding to the phantom sound source signal, φb, a l, 0 is the phase compensation corresponding to the left channel residual signal; φb,a r , 0 is the phase compensation corresponding to the left channel residual signal.

[0090] After the spatial directions corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal are determined, the head-related transfer functions corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal are determined according to the spatial directions corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal, then the left virtual ear-in signals corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal are calculated according to the above formulas (10)-(12), and the left ear-in signal is calculated according to the above formula (9); and the right virtual ear-in signals corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal are calculated according to the above formulas (14)-(15), and the right ear-in signal is calculated according to the above formula (13); so as to obtain the target ear-in signal.

[0091] The spatial directions corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal can be determined through the following steps C1-C3.

[0092] C1, obtain a first position angle and a second position angle.

[0093] The first position angle is a position angle corresponding to a left virtual loudspeaker, and the sound emitted by the left virtual loudspeaker corresponds to the sound emitted by the environmental sound source in the left direction during the stereo recording process. It can be understood that the left virtual loudspeaker is used to simulate the environmental sound source located on the left side of the recording device during the stereo recording process. The second position angle is a position angle corresponding to a right virtual loudspeaker, and the sound emitted by the right virtual loudspeaker corresponds to the sound emitted by the environmental sound source in the right direction during the stereo recording process. It can be understood that the right virtual loudspeaker is used to simulate the environmental sound source located on the right side of the recording device during the stereo recording process.

[0094] The first position angle and the second position angle can be preset, for example, they can be set to-30° and 30° respectively. It should be understood that the position angle corresponding to the virtual loudspeaker can be set based on actual needs, and the present application does not limit it.

[0095] C2, determine the spatial direction corresponding to the left channel residual signal according to the first position angle, and determine the spatial direction corresponding to the right channel residual signal according to the second position angle.

[0096] C3, determine the spatial direction corresponding to the phantom sound source signal according to the first position angle, the second position angle and the target position angle.

[0097] Specifically, the calculation formulas of the spatial orientations corresponding to the left channel residual signal, the right channel residual signal and the phantom sound source signal can refer to the following formulas (17)-(19)

[0098] a s = c1*(a2+(a 1- a2)*γ b / 90)+c0 Formula (17)

[0099] a l = c1*a1+c0 Formula (18)

[0100] a r = c1*a2+c0 Formula (19)

[0101] wherein a l is the first position angle, a r is the second position angle, c1 is the orientation scaling factor, c0 is the offset angle, a s is the spatial orientation corresponding to the phantom sound source signal, a l is the spatial orientation corresponding to the left channel residual signal, and a r is the spatial orientation corresponding to the right channel residual signal.

[0102] By determining the spatial orientations of the virtual sound sources according to the target position angle and the position angles of the left and right virtual loudspeakers, the spatial orientations of the virtual sound sources can be made to be more different, thereby being able to bring stronger spatial sense.

[0103] In the process of determining the head-related transfer functions corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal, the head-related transfer function corresponding to the spatial orientation of the phantom sound source signal can be obtained from the preset head-related transfer function library as the head-related transfer function corresponding to the phantom sound source signal, the head-related transfer function corresponding to the spatial orientation of the left channel residual signal can be obtained as the head-related transfer function corresponding to the left channel residual signal, and the head-related transfer function corresponding to the spatial orientation of the right channel residual signal can be obtained as the head-related transfer function corresponding to the right channel residual signal, according to the spatial orientations corresponding to the phantom sound source signal, the left channel residual signal and the right channel residual signal. The head-related transfer functions in the preset head-related transfer function library can be obtained by pre-measurement.

[0104] In some possible cases, in the case that the head-related transfer function corresponding to the spatial orientation of the target virtual sound signal does not exist in the preset head-related transfer function library, the head-related transfer function corresponding to the spatial orientation of the target virtual sound signal can be determined by linear interpolation, the target virtual sound signal being any one of the phantom sound source signal, the left-channel residual signal or the right-channel residual signal. By linear interpolation, the head-related transfer function can be supplemented, thereby ensuring the diversity of the head-related transfer function.

[0105] B2, mixing the target in-ear signal and the to-be-processed stereo frequency-domain signal to obtain a frequency-domain output signal.

[0106] Here, the mixing of the target in-ear signal and the to-be-processed stereo frequency-domain signal refers to mixing the signal frequencies of the target in-ear signal and the to-be-processed stereo frequency-domain signal to obtain a new signal.

[0107] In a possible implementation, the mixing of the target in-ear signal and the to-be-processed stereo frequency-domain signal can be performed through the following steps B21-B22.

[0108] B21, obtaining a stereo low-frequency signal corresponding to the to-be-processed stereo frequency-domain signal.

[0109] The frequency of the stereo low-frequency signal is lower than a preset frequency. The stereo low-frequency signal can be a stereo low-frequency signal extracted by a low-pass filter. For example, the cutoff frequency of the low-pass filter can be 200 HZ.

[0110] In a specific implementation, the to-be-processed stereo frequency-domain signal can be filtered by a low-pass filter to obtain the stereo low-frequency signal.

[0111] B22, performing weighted sum modulation on the target in-ear signal, the to-be-processed stereo frequency-domain signal and the stereo low-frequency signal to obtain the frequency-domain output signal.

[0112] Specifically, the formula of the weighted sum modulation is as follows:

[0113] OUT l = M1*Yl + M2*Xl + M3*Z

[0114] OUTr = M1*Yr + M2*Xr + M3*Z

[0115] wherein, OUT l and OUTr are left and right ear frequency-domain output signals respectively, M1, M2 and M3 are weighting coefficients respectively, Yl and Yr are left and right ear in-ear signals respectively, and Xl and Xr are left and right channel frequency-domain signals respectively.

[0116] By mixing the signals when each virtual sound source is transmitted to the ear with the stereo frequency domain signal to be processed, the frequency domain output signal obtained by mixing can carry the spatial information of each virtual sound source, thereby having stronger spatial sense.

[0117] Specifically, the target in-ear signal, the stereo frequency domain signal to be processed and the stereo low frequency signal can be weighted and summed to obtain the frequency domain output signal in the form of up-mixing or down-mixing.

[0118] Optionally, after obtaining the frequency domain output signal, the frequency domain output signal can be inversely Fourier transformed to obtain a time domain output signal, that is, out l IFFT (M1*Yl+M2*Xl+M3*Z), outr=IFFT (M1*Yr+M2*Xr+M3*Z).

[0119] In the above Figure 1 In the corresponding technical solution, after obtaining the stereo frequency domain signal to be processed, the stereo frequency domain signal to be processed is spatially information decomposed to obtain virtual sound signals at multiple different spatial orientations, which can simulate virtual sound sources at different orientations; then, the stereo frequency domain signal to be processed is spatially information reconstructed according to the virtual sound signals at multiple different spatial orientations and the head-related transfer functions corresponding to the virtual sound signals at multiple different spatial orientations, which can simulate the transmission process of the virtual sound sources at different orientations to the human ear, so that the frequency domain output signal obtained by reconstruction has rich spatial information and can bring stronger virtual surround sense and spatial sense to the user when output.

[0120] Referring to Figure 2 , Figure 2 is a structural schematic diagram of an audio data processing device provided by an embodiment of the present application. As shown in Figure 2 , the sound signal processing device 20 includes:

[0121] The acquisition module 201 is configured to acquire a stereo frequency domain signal to be processed.

[0122] The decomposition module 202 is configured to spatially information decompose the stereo frequency domain signal to be processed to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations.

[0123] The reconstruction module 203 is configured to perform spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein the head-related transfer function corresponding to a target virtual sound signal is used to represent the phase and frequency response of the target virtual sound signal transmitted to the head from a spatial orientation corresponding to the target virtual sound signal, and the target virtual sound signal is any one of the plurality of virtual sound signals.

[0124] In a possible design, the to-be-processed stereo frequency domain signal includes a left channel frequency domain signal and a right channel frequency domain signal; and the decomposition module 202 is specifically configured to perform spectral analysis on the left channel frequency domain signal and the right channel frequency domain signal, to determine a target position angle, the target position angle being a position angle at which the difference between signals is the largest; and perform signal superposition and signal decomposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal, to obtain the plurality of virtual sound signals.

[0125] In a possible design, the plurality of virtual sound signals include an illusion sound source signal, a left channel residual signal and a right channel residual signal; and the decomposition module 202 is specifically configured to perform signal superposition according to the target position angle, the left channel frequency domain signal and the right channel frequency domain signal, to obtain the illusion sound source signal; perform signal decomposition according to the target position angle, the illusion sound source signal and the right channel frequency domain signal, to obtain the left channel residual signal; and perform signal decomposition according to the target position angle, the illusion sound source signal and the left channel frequency domain signal, to obtain the right channel residual signal.

[0126] In a possible design, the plurality of virtual sound signals include an illusion sound source signal, a left channel residual signal and a right channel residual signal; and the sound signal processing apparatus 20 further includes a position angle acquisition module 204 configured to acquire a first position angle and a second position angle, the first position angle being a position angle corresponding to a left virtual loudspeaker, and the second position angle being a position angle corresponding to a right virtual loudspeaker; a direction determination module 205 configured to determine a spatial orientation corresponding to the left channel residual signal according to the first position angle, and determine a spatial orientation corresponding to the right channel residual signal according to the second position angle; and determine a spatial orientation corresponding to the illusion sound source signal according to the first position angle, the second position angle and the target position angle.

[0127] In a possible design, the reconstruction module 203 is specifically configured to: determine a target in-ear signal according to the plurality of virtual sound signals and the head-related transfer functions corresponding to the plurality of virtual sound signals, the target in-ear signal being used to indicate a mixed signal received by a pair of ears when the plurality of virtual sound signals are transmitted to the head; and perform mixing processing on the target in-ear signal and the to-be-processed stereo frequency-domain signal to obtain the frequency-domain output signal.

[0128] In a possible design, the reconstruction module 203 is specifically configured to: determine a virtual in-ear signal corresponding to the target virtual sound signal according to the target virtual sound signal and the head-related transfer function corresponding to the target virtual sound signal, the virtual in-ear signal corresponding to the target virtual sound signal being used to indicate a signal received by a pair of ears when the target virtual sound signal is transmitted to the head; and perform signal superposition on the virtual in-ear signals corresponding to the plurality of virtual sound signals to obtain the target in-ear signal.

[0129] In a possible design, the reconstruction module 203 is specifically configured to: obtain a stereo low-frequency signal corresponding to the to-be-processed stereo frequency-domain signal, the stereo low-frequency signal being lower than a preset frequency; and perform weighted sum modulation on the target in-ear signal, the to-be-processed stereo frequency-domain signal and the stereo low-frequency signal to obtain the frequency-domain output signal.

[0130] It should be noted that, Figure 2 The contents not mentioned in the corresponding embodiments can be referred to the descriptions of the foregoing method embodiments, which will not be described here.

[0131] The apparatus, after obtaining the to-be-processed stereo frequency-domain signal, performs spatial information decomposition on the to-be-processed stereo frequency-domain signal to obtain a plurality of virtual sound signals at different spatial orientations, which can simulate virtual sound sources at different orientations; then performs spatial information reconstruction on the to-be-processed stereo frequency-domain signal according to the plurality of virtual sound signals at different spatial orientations and the head-related transfer functions corresponding to the plurality of virtual sound signals at different spatial orientations, which can simulate the transmission process of the virtual sound sources at different orientations to the human ear, and thus the frequency-domain output signal obtained through the reconstruction has rich spatial information, which can bring stronger virtual surround and spatial sense to the user when output.

[0132] See Figure 3 , Figure 3 is a structural schematic diagram of an audio device provided by an embodiment of the present application. The audio device 30 includes a processor 301 and a memory 302. The memory 302 is connected to the processor 301, for example, through a bus.

[0133] The processor 301 is configured to support the audio device 30 to perform the corresponding functions in the methods in the above method embodiments. The processor 301 can be a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof. The hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0134] The memory 302 is configured to store program codes and the like. The memory 302 can include a volatile memory (VM) such as a random access memory (RAM), and / or a non-volatile memory (NVM) such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), and / or a combination thereof.

[0135] The processor 301 can invoke the program codes to perform the following operations:

[0136] obtain a to-be-processed stereo frequency-domain signal;

[0137] perform spatial information decomposition on the to-be-processed stereo frequency-domain signal to obtain a plurality of virtual sound signals corresponding to different spatial directions;

[0138] According to the plurality of virtual sound signals and the head-related transfer function corresponding to each of the plurality of virtual sound signals, spatial information of the to-be-processed stereo frequency domain signal is reconstructed to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein the head-related transfer function corresponding to a target virtual sound signal is used to represent a phase and a frequency response of the target virtual sound signal transmitted from a spatial position corresponding to the target virtual sound signal to a head, and the target virtual sound signal is any one of the plurality of virtual sound signals.

[0139] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program comprises program instructions. When the program instructions are executed by a computer, the computer executes the method according to the foregoing embodiments.

[0140] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments can be included. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0141] The above only describes the preferred embodiments of the present application, and of course cannot limit the scope of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. A method of processing a sound signal, characterized by, The method comprises: acquiring a to-be-processed stereo frequency domain signal, the to-be-processed stereo frequency domain signal comprising a left-channel frequency domain signal and a right-channel frequency domain signal; performing spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations, the plurality of virtual sound signals comprising an illusionary sound source signal, a left-channel residual signal and a right-channel residual signal, the illusionary sound source signal being a sound signal emitted by a main sound source in a stereo recording process; performing spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals, to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein a head-related transfer function corresponding to a target virtual sound signal is used to represent a phase and frequency response of the target virtual sound signal from a spatial orientation corresponding to the target virtual sound signal to a head, the target virtual sound signal being any one of the plurality of virtual sound signals; the spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain the plurality of virtual sound signals comprises: performing spectral analysis on the left-channel frequency domain signal and the right-channel frequency domain signal to determine a target position angle, the target position angle being a position angle that maximizes a difference between signals; performing signal superposition according to the target position angle, the left-channel frequency domain signal and the right-channel frequency domain signal to obtain an illusionary sound source signal; performing signal decomposition according to the target position angle, the illusionary sound source signal and the right-channel frequency domain signal to obtain a left-channel residual signal; performing signal decomposition according to the target position angle, the illusionary sound source signal and the left-channel frequency domain signal to obtain a right-channel residual signal.

2. The method of claim 1, wherein, Before the spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals, to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, the method further comprises: acquiring a first position angle and a second position angle, the first position angle being a position angle corresponding to a left virtual loudspeaker, and the second position angle being a position angle corresponding to a right virtual loudspeaker; determining a spatial orientation corresponding to the left-channel residual signal according to the first position angle, and determining a spatial orientation corresponding to the right-channel residual signal according to the second position angle; determining a spatial orientation corresponding to the illusionary sound source signal according to the first position angle, the second position angle and the target position angle.

3. The method according to any of claims 1 or 2, characterized in that, the spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals, to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, comprises: determine a target in-ear signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals, the target in-ear signal being used to indicate a mixed signal received by a pair of ears when the plurality of virtual sound signals are transmitted to the head; mix the target in-ear signal with the to-be-processed stereo frequency domain signal to obtain the frequency domain output signal.

4. The method of claim 3, wherein, The method according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals, the target in-ear signal includes: determine a virtual in-ear signal corresponding to the target virtual sound signal according to the target virtual sound signal and the head-related transfer function corresponding to the target virtual sound signal, the virtual in-ear signal corresponding to the target virtual sound signal being used to indicate a signal received by a pair of ears when the target virtual sound signal is transmitted to the head; superimpose the virtual in-ear signals corresponding to the plurality of virtual sound signals to obtain the target in-ear signal.

5. The method of claim 4, wherein, The method of mixing the target in-ear signal with the to-be-processed stereo frequency domain signal to obtain the frequency domain output signal includes: obtain a stereo low frequency signal corresponding to the to-be-processed stereo frequency domain signal, the stereo low frequency signal having a frequency lower than a preset frequency; weight and sum modulate the target in-ear signal, the to-be-processed stereo frequency domain signal, and the stereo low frequency signal to obtain the frequency domain output signal.

6. A sound signal processing apparatus characterized by comprising: The method includes: an obtaining module configured to obtain a to-be-processed stereo frequency domain signal, the to-be-processed stereo frequency domain signal including a left channel frequency domain signal and a right channel frequency domain signal; a decomposition module configured to perform spatial information decomposition on the to-be-processed stereo frequency domain signal to obtain a plurality of virtual sound signals, the plurality of virtual sound signals being virtual sound signals corresponding to different spatial orientations, the plurality of virtual sound signals including a phantom sound source signal, a left channel residual signal, and a right channel residual signal, the phantom sound source signal being a sound signal emitted by a main sound source in a stereo recording process; a reconstruction module configured to perform spatial information reconstruction on the to-be-processed stereo frequency domain signal according to the plurality of virtual sound signals and head-related transfer functions corresponding to the plurality of virtual sound signals to obtain a frequency domain output signal corresponding to the to-be-processed stereo frequency domain signal, wherein a head-related transfer function corresponding to a target virtual sound signal is used to represent a phase and a frequency response of the target virtual sound signal transmitted from a spatial orientation corresponding to the target virtual sound signal to the head, the target virtual sound signal being any one of the plurality of virtual sound signals. The decomposition module is specifically configured to: perform spectrum analysis on the left-channel frequency-domain signal and the right-channel frequency-domain signal to determine a target position angle, the target position angle being a position angle that makes the difference between signals maximum, perform signal superposition according to the target position angle, the left-channel frequency-domain signal and the right-channel frequency-domain signal to obtain a phantom sound source signal, perform signal decomposition according to the target position angle, the phantom sound source signal and the right-channel frequency-domain signal to obtain a left-channel residual signal, and perform signal decomposition according to the target position angle, the phantom sound source signal and the left-channel frequency-domain signal to obtain a right-channel residual signal.

7. An audio device, comprising: The audio device comprises a memory, a processor, the memory being connected to the processor, the processor being used to execute one or more computer programs stored in the memory, and the processor, when executing the one or more computer programs, causes the audio device to implement the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, the computer program comprising program instructions, the program instructions, when executed by a processor, causing the processor to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Sound signal processing method, device and system, terminal and storage medium

    CN108966110A