Signal generation method, device, readable storage medium, and computer program product
By integrating omnidirectional and figure-eight directional microphones into audio devices, channel separation and sound source parameter rendering are achieved, solving the problem of insufficient universality of stereo recording in small devices and realizing effective reproduction of stereo effects.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GOERTEK INC
- Filing Date
- 2025-07-10
- Publication Date
- 2026-04-23
AI Technical Summary
Traditional stereo recording equipment, especially small devices such as mobile phones, tablets, VR glasses and AR glasses, cannot effectively achieve the universality of stereo recording and lacks a sense of stereo sound.
An omnidirectional microphone and a figure-eight directional microphone are integrated into the audio device. The left and right channel signals are obtained through channel separation processing. The original audio signal is rendered based on the sound source parameters to generate the target audio signal to simulate the sound heard by the human ear.
It improves the versatility of stereo recording, enabling small devices to effectively reproduce the spatial sense of sound, increase the listener's spatial directionality and depth perception, and ensure stereo effect.
Smart Images

Figure CN2025107829_23042026_PF_FP_ABST
Abstract
Description
Signal generation methods, devices, readable storage media, and computer program products Technical Field
[0001] This application relates to the field of signal processing technology, and in particular to a signal generation method, apparatus, readable storage medium, and computer program product. Background Technology
[0002] The basic characteristics of sound include pitch, timbre, and intensity. Stereo (also known as 3D stereo) refers to sound that, in addition to these basic characteristics, also possesses spatial features such as directionality, distance, and layering. For example, in a concert hall, each instrument has its own position, and listeners can discern the direction and distance of the sound source based on what they hear; this is the directionality and distance characteristic of stereo.
[0003] Traditional stereo recording typically requires two or more microphones, and the two microphones need to be spaced relatively far apart (usually 2 to 3 meters) to achieve a clear stereo separation. Stereo recording is generally only possible with larger audio devices. For smaller audio devices that require tightly integrated microphones, such as mobile phones, tablets, VR (Virtual Reality) glasses, and AR (Augmented Reality) glasses, the smaller spacing between the microphones may result in recording quality similar to a single microphone, lacking a sense of depth and limiting the versatility of stereo recording.
[0004] Therefore, how to improve the universality of stereo recording is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] The main objective of this application is to provide a signal generation method, apparatus, readable storage medium, and computer program product, aiming to solve the technical problem of how to improve the universality of stereo recording.
[0006] To achieve the above objectives, this application provides a signal generation method applied to an audio device, the audio device including an omnidirectional microphone and a figure-eight directional microphone, the signal generation method comprising the following steps:
[0007] Acquire the original audio signal, wherein the original audio signal includes a first audio signal and a second audio signal, the first audio signal being the audio signal picked up by the omnidirectional microphone, and the second audio signal being the audio signal picked up by the figure-eight directional microphone;
[0008] The original audio signal is subjected to channel separation processing to obtain a two-channel signal, wherein the two-channel signal includes a left channel signal and a right channel signal;
[0009] Based on the dual-channel signal, the sound source parameters corresponding to the sound source are calculated, and the original audio signal is rendered according to the sound source parameters to obtain the generated target audio signal. The sound source parameters include the sound source orientation and / or the sound source size.
[0010] In one embodiment, the step of performing channel separation processing on the original audio signal to obtain a two-channel signal includes:
[0011] Alignment processing is performed on the first audio signal and the second audio signal to obtain an aligned audio signal, wherein the alignment processing includes time alignment processing and / or frequency response alignment processing, and the aligned audio signal includes a third audio signal and a fourth audio signal, wherein the third audio signal is the first audio signal after alignment processing, and the fourth audio signal is the second audio signal after alignment processing;
[0012] The left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0013] The left channel signal and the right channel signal are determined to be stereo signals.
[0014] In one embodiment, after the step of aligning the first audio signal and the second audio signal to obtain an aligned audio signal, the method further includes:
[0015] For the third audio signal and the fourth audio signal of the initial frame, the left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0016] The signal value of the third audio signal is determined to be a first signal value, the signal value of the left channel signal is determined to be a second signal value, and the signal value of the right channel signal is determined to be a third signal value.
[0017] The target signal value is obtained by calculating the weighted sum of the second signal value and the third signal value.
[0018] A target gain is determined based on the first signal value and the target signal value, wherein the target gain is negatively correlated with the first signal value and positively correlated with the target signal value;
[0019] The third audio signal of the next frame is adjusted with the target gain to obtain an intermediate audio signal. The intermediate audio signal is determined to be the new third audio signal. Based on the new third audio signal and the fourth audio signal of the next frame, the steps of adding the third audio signal to the fourth audio signal to obtain the left channel signal and subtracting the phase-inverted fourth audio signal from the third audio signal to obtain the right channel signal are returned to the execution.
[0020] Once the preset separation termination condition is met, all the obtained left channel signals and all the obtained right channel signals are determined to be stereo signals.
[0021] In one embodiment, the step of calculating the sound source parameters corresponding to the sound source based on the dual-channel signal includes:
[0022] The dual-channel signal is divided into multiple sub-band signals by frequency band division.
[0023] For each sub-band signal, determine the target frequency band to which the sub-band signal belongs, and calculate the sound source parameters corresponding to the sound source in the target frequency band based on the sub-band signal.
[0024] In one embodiment, the step of calculating the sound source parameters corresponding to the target frequency band based on the sub-band signal includes:
[0025] Spatial cue parameters are calculated based on the left and right channel sub-signals in the sub-band signals, and the sound source parameters corresponding to the sound source in the target frequency band are determined based on the spatial cue parameters.
[0026] The left channel sub-signal and the right channel sub-signal belong to the same frequency band.
[0027] In one embodiment, the step of determining the sound source parameters corresponding to the sound source in the target frequency band based on the spatial cue parameters includes:
[0028] If the spatial cue parameters include the binaural time difference and the target frequency band belongs to a preset low-frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural time difference.
[0029] If the spatial cue parameters include the binaural sound level difference and the target frequency band does not belong to the preset low frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural sound level difference.
[0030] If the spatial cue parameters include the binaural cross-correlation coefficient, then the sound source size corresponding to the sound source in the target frequency band is determined based on the binaural cross-correlation coefficient.
[0031] In one embodiment, the step of calculating spatial cue parameters based on the left and right channel sub-signals in the sub-band signals includes:
[0032] If there are multiple left channel sub-signals or multiple right channel sub-signals in the sub-band signal, then select one target left channel sub-signal and one target right channel sub-signal.
[0033] Spatial cue parameters are calculated based on the target left vocal tract sub-signal and the target right vocal tract sub-signal.
[0034] In addition, to achieve the above objectives, this application also provides an audio device, which includes an omnidirectional microphone, a figure-eight directional microphone, and a processor. The omnidirectional microphone and the figure-eight directional microphone are respectively connected to the processor, and the processor is used to execute the steps of the signal generation method described above.
[0035] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a program implementing the signal generation method is stored, and the program implementing the signal generation method is executed by a processor to implement the steps of the signal generation method as described above.
[0036] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the signal generation method described above.
[0037] One or more technical solutions proposed in this application have at least the following technical effects:
[0038] An omnidirectional microphone and a figure-eight directional microphone are integrated into an audio device for recording, acquiring the original audio signals picked up by the omnidirectional and figure-eight directional microphones. The original audio signals are then processed by channel separation to obtain a two-channel signal, which includes a left channel signal and a right channel signal. Based on the two-channel signal, sound source parameters corresponding to the sound source are calculated, and the original audio signal is rendered according to the sound source parameters to obtain the generated target audio signal. The sound source parameters include the sound source orientation and / or sound source size. Thus, this embodiment separates the left and right channel signals from the original audio signals picked up by the omnidirectional and figure-eight directional microphones, simulating the sound heard by human binaural hearing through the left and right channel signals. This allows for the close arrangement of the figure-eight directional and omnidirectional microphones, enabling the construction of a very compact miniature directional recording device. This device occupies little space and can be easily integrated into small audio devices, thereby improving the versatility of stereo recording. Furthermore, the sound source orientation and / or sound source size are calculated based on the two-channel signal. The sound source orientation allows the listener to determine the specific location of the sound source, thereby increasing the listener's spatial direction perception. The sound source size can increase the sense of depth of the sound, thereby increasing the listener's spatial depth perception. Therefore, both the sound source orientation and size can be used to restore the spatial sense of stereo. Thus, the original audio signal is rendered based on the sound source parameters to simulate the sound heard by the listener according to the sound source parameters, so that the target audio signal generated by the rendering can better restore the spatial sense of the sound, that is, ensure the stereo effect of the generated signal. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 is a flowchart illustrating the first embodiment of the signal generation method of this application;
[0042] Figure 2 is a schematic diagram of a computational ITDG scenario involving an embodiment of the signal generation method of this application;
[0043] Figure 3 is a schematic diagram of the ITDG variation curve involved in an embodiment of the signal generation method of this application;
[0044] Figure 4 is a schematic diagram of the device structure of the signal generation apparatus of this application;
[0045] Figure 5 is a schematic diagram of the hardware operating environment involved in the signal generation device in the embodiments of this application.
[0046] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0047] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] With the widespread use of open-back audio products and devices such as AR (Augmented Reality), VR (Virtual Reality), and smart audio glasses, users' demand for immersive near-ear open-back audio devices is increasing. Currently, in stereo recording applications, users typically wear audio devices, which use the device's pickup unit (e.g., a microphone) to capture sound signals from the surrounding environment. The captured audio signal is then saved to obtain a stereo recording signal. However, due to the limitations of spatial perception in hardware devices, the recording signal obtained in this way cannot accurately represent the spatial information of the sound signal, failing to achieve the effect perceived by the human ear. Subsequent playback cannot reproduce the spatial sense, resulting in a poor user experience.
[0049] Traditional stereo recording typically requires two or more microphones, and the two microphones need to be spaced relatively far apart (usually 2-3 meters) to achieve a clear stereo separation. However, when microphones are placed close together (such as on mobile phones, tablets, VR glasses, AR glasses, etc.), the recording effect may be similar to that of a single microphone, lacking a sense of stereo sound.
[0050] Based on this, the main solution of this application is: to integrate an omnidirectional microphone and a figure-eight directional microphone into an audio device for recording, and to obtain the original audio signals picked up by the omnidirectional microphone and the figure-eight directional microphone; to perform channel separation processing on the original audio signals to obtain a two-channel signal, wherein the two-channel signal includes a left channel signal and a right channel signal; to calculate the sound source parameters corresponding to the sound source based on the two-channel signal, and to render the original audio signal according to the sound source parameters to obtain the generated target audio signal, wherein the sound source parameters include the sound source orientation and / or the sound source size.
[0051] This application separates the left and right channel signals from the original audio signals picked up by an omnidirectional microphone and a figure-eight directional microphone. These signals are then used to simulate the sound heard by both ears, allowing for a close arrangement of the figure-eight and omnidirectional microphones. This creates a very compact, miniature directional recording device that occupies little space and can be easily integrated into smaller audio devices, thus improving the versatility of stereo recording. Furthermore, the application calculates the sound source location and / or sound source size based on the two-channel signals. The sound source location allows the listener to determine the specific location of the sound source, increasing their spatial directionality perception. The sound source size increases the sense of sound depth, further enhancing the listener's spatial depth perception. Therefore, both sound source location and size can be used to recreate the spatial sense of stereo. Based on these sound source parameters, the original audio signal is rendered to simulate the sound heard by the listener. This ensures that the generated target audio signal can effectively reproduce the spatial sense of sound, guaranteeing the stereo effect of the generated signal.
[0052] It should be noted that the execution subject of each embodiment of the signal generation method of this application can be an audio device capable of realizing the above functions, such as headphones, AR glasses, AR helmets, VR glasses, VR helmets, smart audio glasses, etc. This embodiment does not impose specific limitations on this.
[0053] Based on this, this application proposes a signal generation method according to a first embodiment, applied to an audio device, the audio device including an omnidirectional microphone and a figure-eight directional microphone. Referring to Figure 1, the signal generation method includes steps S10 to S30:
[0054] Step S10: Acquire the picked-up original audio signal, wherein the original audio signal includes a first audio signal and a second audio signal, the first audio signal is the audio signal picked up by the omnidirectional microphone, and the second audio signal is the audio signal picked up by the figure-eight directional microphone;
[0055] An omnidirectional microphone is a type of microphone that can capture sound from all directions. A figure-eight microphone, also known as a supercardioid or bicardioid microphone, is a type of microphone with a specific directional characteristic. It has high sensitivity on the front and back of the microphone, and lower sensitivity on the sides.
[0056] The omnidirectional microphone and the figure-eight directional microphone are positioned at different locations to pick up direct sound signals and ambient sound signals. Direct sound signals refer to signals that travel directly from the sound source to the pickup unit without any reflection or diffraction. Ambient sound signals refer to signals that have undergone reflection, diffraction, and scattering during propagation; these include multiple reflections of sound waves from surfaces such as walls, ceilings, and floors within the room, as well as acoustic effects caused by the room's geometry and materials.
[0057] Furthermore, the audio device may include one or more omnidirectional microphones and one or more figure-eight directional microphones. This embodiment does not impose a specific limit on the specific number of omnidirectional microphones and figure-eight directional microphones.
[0058] After obtaining the original audio signal, noise reduction can be performed on the original audio signal, and subsequent processing can be carried out based on the noise-reduced original audio signal to improve the quality of the original audio signal.
[0059] Step S20: Perform channel separation processing on the original audio signal to obtain a dual-channel signal, wherein the dual-channel signal includes a left channel signal and a right channel signal;
[0060] It should be noted that the left channel signal and the right channel signal are used to simulate the sound signals heard by human ears.
[0061] As one implementation, the step of performing channel separation processing on the original audio signal to obtain a dual-channel signal includes steps S201 to S203:
[0062] Step S201: Align the first audio signal and the second audio signal to obtain an aligned audio signal. The alignment process includes time alignment and / or frequency response alignment. The aligned audio signal includes a third audio signal and a fourth audio signal. The third audio signal is the first audio signal after alignment, and the fourth audio signal is the second audio signal after alignment.
[0063] Considering that the omnidirectional microphone and the figure-eight directional microphone are positioned at different locations on the audio device, the signals picked up by the omnidirectional microphone and the figure-eight directional microphone may have a certain relative time delay due to their different positions. To avoid attributing the relative time delay caused by the difference in the positions of the omnidirectional microphone and the figure-eight directional microphone to the location of the sound source, this embodiment performs alignment processing on the first audio signal and the second audio signal. Specifically, the relative time delay between the two can be calculated, and the end time of the first audio signal or the second audio signal can be aligned after being delayed by the delay time.
[0064] Specifically, delay = d / c*fs, where d represents the distance between the omnidirectional microphone and the figure-eight microphone, c represents the speed of sound, and fs represents the sampling rate of the omnidirectional microphone and the figure-eight microphone (both sampling rates are set to be the same).
[0065] Frequency response refers to the response characteristics of signals with different frequency components. It describes the relationship between the amplitude and phase of the output signal and the frequency of the input signal. Specifically, frequency response includes amplitude response and phase response. Amplitude response refers to the characteristic that the amplitude of the output signal changes with the frequency of the input signal. Phase response refers to the characteristic that the phase of the output signal changes with the frequency of the input signal.
[0066] Frequency response alignment processing specifically involves aligning the amplitude and phase responses of the first and second audio signals, i.e., aligning the amplitude and phase responses of the omnidirectional microphone and the figure-eight microphone. In one specific embodiment, the output of the figure-eight microphone (i.e., the second audio signal) undergoes time-frequency domain filtering to calibrate the amplitude and phase responses of the figure-eight microphone output, aligning them with those of the omnidirectional microphone. This avoids the influence of directivity and differences between different microphones, ensuring that the amplitude responses of the omnidirectional and figure-eight microphones are consistent within the 0–180 degree range, with only a phase difference. The time-frequency domain filtering processing includes, but is not limited to, NLMS (Normalized Least Mean Square) filtering, convolution, IIR (Infinite Impulse Response) filtering, FIR (Finite Impulse Response) filtering, Fourier transform filtering, wavelet transform filtering, and equalization processing.
[0067] Preferably, the first audio signal and the second audio signal are time-aligned and frequency-response aligned to eliminate signal differences caused by non-sound source factors such as the positional differences between the omnidirectional microphone and the figure-eight microphone and their own frequency response characteristics, thereby improving the quality of subsequent signal separation.
[0068] Step S202: The left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0069] Let the third audio signal be Mid and the fourth audio signal be Side. Then the left channel signal and the right channel signal can be expressed by the formulas: Left channel signal = Mid + Side (Mid signal plus the un-inverted phase Side signal), Right channel signal = Mid - Side (Mid signal minus the inverted phase Side signal).
[0070] Step S203: Determine that the left channel signal and the right channel signal are dual-channel signals.
[0071] Furthermore, after obtaining the two-channel signal, the target amplitude can be obtained by using a weighted average of the amplitudes (i.e., signal values) of the left channel signal and the right channel signal as the target amplitude. Gain control can then be applied to the first audio signal (e.g., if the target amplitude is A and the amplitude of the first audio signal is B, then A / B can be used to adjust the gain of the first audio signal) to obtain an intermediate signal with a suitable relative intensity. This separates the left channel signal, right channel signal, and intermediate signal from the original audio signal, allowing subsequent mixing by adjusting these signals to achieve the best stereo effect.
[0072] As another implementation, as one embodiment, the step of performing channel separation processing on the original audio signal to obtain a dual-channel signal includes steps S204 to S210:
[0073] Step S204: Align the first audio signal and the second audio signal to obtain an aligned audio signal;
[0074] The alignment process is the same as the one described above, and will not be repeated here.
[0075] Step S205: For the third audio signal and the fourth audio signal of the initial frame, the left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0076] For the third and fourth audio signals of the initial frame, that is, for the third and fourth audio signals of the first frame, the left channel signal of the first frame is obtained by adding the third audio signal of this frame to the fourth audio signal of this frame. Similarly, the right channel audio signal of the first frame is obtained by subtracting the phase-inverted fourth audio signal of this frame from the third audio signal of this frame.
[0077] Step S206: Determine the signal value of the third audio signal as the first signal value, determine the signal value of the left channel signal as the second signal value, and determine the signal value of the right channel signal as the third signal value;
[0078] Specifically, the first signal value can be the average of the signal values corresponding to all sampling points of the first audio signal. Similarly, the second signal value can be the average of the signal values corresponding to all sampling points of the left channel signal, and the third signal value can be the average of the signal values corresponding to all sampling points of the right channel signal.
[0079] Step S207: Calculate the weighted sum of the second signal value and the third signal value to obtain the target signal value;
[0080] Step S208: Determine the target gain based on the first signal value and the target signal value, wherein the target gain is negatively correlated with the first signal value and positively correlated with the target signal value;
[0081] Specifically, the target gain can be the ratio between the first signal value and the target signal value, such as the target signal value divided by the first signal value as the target gain.
[0082] Step S209: Adjust the third audio signal of the next frame with the target gain to obtain an intermediate audio signal, determine the intermediate audio signal as a new third audio signal, and return to the steps of adding the third audio signal to the fourth audio signal to obtain the left channel signal and subtracting the phase-inverted fourth audio signal from the third audio signal to obtain the right channel signal based on the new third audio signal and the fourth audio signal of the next frame.
[0083] Adjust the third audio signal of the next frame with the target gain. Specifically, this adjustment can be achieved by multiplying the signal value by the target gain, such as multiplying the signal value corresponding to each sampling point by the target gain.
[0084] The adjusted intermediate audio signal is used as the new third audio signal, that is, the third audio signal of the next frame. It is then used to perform a sum-difference calculation with the fourth audio signal of the next frame to obtain the left channel signal and the right channel signal of the next frame.
[0085] Step S210: After the preset separation end condition is met, determine that all the obtained left channel signals and all the right channel signals are dual channel signals.
[0086] The preset separation termination condition can be set by relevant personnel according to actual needs, such as having obtained a certain number of frames of dual-channel signals, or the third or fourth audio signal of the current frame (the signal frame for which sum and difference calculation is performed this time) being the last frame signal, etc. This embodiment does not impose specific limitations on this.
[0087] It should be noted that a frame of dual-channel signal includes a frame of left channel signal and a frame of right channel signal.
[0088] Step S30: Calculate the sound source parameters corresponding to the sound source based on the dual-channel signal, and render the original audio signal according to the sound source parameters to obtain the generated target audio signal, wherein the sound source parameters include the sound source orientation and / or the sound source size.
[0089] These sound source parameters can be used to simulate sounds heard by the human ear, including but not limited to sound source location and / or sound source size, such as sound source intensity. Sound source location refers to the spatial position of the sound source relative to the listener, including its position in the horizontal plane (left and right, front and back) and the vertical plane (up and down). Sound source size refers to the perceived volume of the sound source. Sound source size is not entirely consistent with the physical size of the sound source, but is related to factors such as the sound's spectral distribution, dynamic range, and reverberation effect. Sound source size can affect the listener's perception of the spatial characteristics of the sound, including its width, depth, and height. Sound source intensity usually refers to the loudness of the sound, that is, the strength or volume of the sound. It is a measure of sound energy and is closely related to the sound pressure level (SPL). Sound source intensity determines the degree of impact of the sound on the listener's senses, commonly referred to as "volume."
[0090] The location of a sound source is fundamental to stereo systems, as it allows listeners to pinpoint its exact location. Through stereo imaging, a sound source can be located in front of, behind, to the left, to the right, above, and below the listener. Sound source size increases the sense of depth. For example, a sound with a long reverberation time is perceived as coming from a distant space, thus increasing the listener's spatial depth perception. Combining sound source location and size allows for the rendering of more realistic and immersive audio signals. For instance, in cinemas or home theaters, precise sound source localization and reverberation effects can simulate the spatial feel of a movie scene, making the audience feel as if they are in the environment where the story takes place. Therefore, the embodiments of this application select sound source size and location as sound source parameters to ensure that the generated signal has a better spatial reproduction effect.
[0091] The sound source parameters can be calculated based on the two signals using known methods. For example, the size of the sound source can be estimated based on the frequency components of the two signals (a wide-spectrum sound is usually perceived as a large sound source). The location of the sound source can be calculated based on the time difference, intensity difference, phase difference, and other characteristics of the two signals. This embodiment will not describe the specific calculation process of the sound source parameters in detail.
[0092] After obtaining the sound source parameters, the original audio signal is rendered based on the sound source parameters to obtain the generated target audio signal.
[0093] Signal rendering typically refers to the process of processing an audio signal to produce the desired auditory effect on a specific playback device. In this embodiment, to restore the spatial sense of sound, a stereo rendering method is adopted. Stereo rendering is an audio processing technique that uses two or more channels to create a sense of spatial sound. It controls the volume, frequency content, and time differences (phase differences) between different channels based on sound source parameters, so that when the generated signal is played, the listener perceives the sound as coming from a sound source with these sound source parameters.
[0094] In this embodiment, an omnidirectional microphone and a figure-eight directional microphone are integrated into the audio device for recording, acquiring the original audio signals picked up by the omnidirectional and figure-eight directional microphones. The original audio signals are then processed by channel separation to obtain a two-channel signal, which includes a left channel signal and a right channel signal. Based on the two-channel signal, the corresponding sound source parameters are calculated, and the original audio signal is rendered according to the sound source parameters to obtain the generated target audio signal. The sound source parameters include the sound source orientation and / or sound source size. Thus, this embodiment separates the left and right channel signals from the original audio signals picked up by the omnidirectional and figure-eight directional microphones, simulating the sound heard by human binaural hearing through the left and right channel signals. This allows for the close arrangement of the figure-eight directional and omnidirectional microphones, enabling the construction of a very compact miniature directional recording device. This device occupies little space and can be easily integrated into small audio devices, thereby improving the versatility of stereo recording. Furthermore, the sound source orientation and / or sound source size are calculated based on the two-channel signal. The sound source orientation allows the listener to determine the specific location of the sound source, thereby increasing the listener's spatial direction perception. The sound source size can increase the sense of depth of the sound, thereby increasing the listener's spatial depth perception. Therefore, both the sound source orientation and size can be used to restore the spatial sense of stereo. Thus, the original audio signal is rendered based on the sound source parameters to simulate the sound heard by the listener according to the sound source parameters, so that the target audio signal generated by the rendering can better restore the spatial sense of the sound, that is, ensure the stereo effect of the generated signal.
[0095] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of calculating the sound source parameters corresponding to the sound source based on the dual-channel signal includes:
[0096] Step A10: Perform frequency band division processing on the dual-channel signal to obtain multiple sub-band signals;
[0097] It should be noted that the audio signal initially picked up by the microphone is usually a time-domain signal. Therefore, if the stereo signal is a time-domain signal, a Fourier transform is performed on the stereo signal to obtain a frequency-domain stereo signal, and a frequency band division process is performed based on the frequency-domain stereo signal.
[0098] One or more specific frequency boundaries can be set in advance, and the stereo signal is subjected to a frequency band division process according to the pre-set frequency boundaries. Exemplarily, assuming that the pre-set frequency boundaries are A, B, C, and D, where A < B < C < D, the stereo signal can be divided into four frequency bands f1, f2, f3, and f4 according to this frequency boundary. Specifically, f1, f2, f3, and f4 can be less than A, [A, B), [B, C), and greater than or equal to C, respectively.
[0099] It should be noted that the stereo signal includes a left-channel signal and a right-channel signal. Performing a frequency band division process on the stereo signal means performing a frequency band division process on both the left-channel signal and the right-channel signal. Each sub-band signal obtained by the frequency band division process includes a left-channel sub-signal and a right-channel signal belonging to the same frequency band.
[0100] Exemplarily, denoting the stereo signal as M, the left-channel signal in M as L, and the right-channel signal in M as R. After performing a frequency band division process on M according to the frequency boundaries in the above example, four sub-band signals S1, S2, S3, and S4 belonging to f1, f2, f3, and f4 are obtained. Then S1 includes [L1, R1], S2 includes [L2, R2], S3 includes [L3, R3], and S4 includes [L4, R4]. Here, L1 represents the left-channel sub-signal in L belonging to the f1 frequency band, L2 represents the left-channel sub-signal in L belonging to the f2 frequency band, L3 represents the left-channel sub-signal in L belonging to the f3 frequency band, L4 represents the left-channel sub-signal in L belonging to the f4 frequency band, R1 represents the right-channel sub-signal in R belonging to the f1 frequency band, R2 represents the right-channel sub-signal in R belonging to the f2 frequency band, R3 represents the right-channel sub-signal in R belonging to the f3 frequency band, and R4 represents the right-channel sub-signal in R belonging to the f4 frequency band.
[0101] Step A20, for each of the sub-band signals, determine the target frequency band to which the sub-band signal belongs, and calculate the sound source parameters corresponding to the sound source in the target frequency band based on the sub-band signal.
[0102] It should be noted that, based on the target frequency band to which each sub-band signal belongs, the sound source parameters corresponding to each target audio are calculated. That is, the sound source parameters corresponding to each frequency band are calculated, and the target audio signal of that frequency band is generated based on the sound source parameters corresponding to each frequency band. For example, let the sound source parameter be I, and the sub-band signals be S1, S2, S3, and S4, belonging to the target frequency bands f1, f2, f3, and f4 respectively. Then, based on S1, the sound source parameter I1 corresponding to the sound source in the f1 frequency band is calculated; based on S2, the sound source parameter I2 corresponding to the sound source in the f2 frequency band is calculated; based on S3, the sound source parameter I3 corresponding to the sound source in the f3 frequency band is calculated; and based on S4, the sound source parameter I4 corresponding to the sound source in the f4 frequency band is calculated. The final sound source parameter I specifically includes I1, I2, I3, and I4. During rendering, I1 is used to generate the target audio signal in the f1 frequency band, I2 is used to generate the target audio signal in the f2 frequency band, I3 is used to generate the target audio signal in the f3 frequency band, and I4 is used to generate the target audio signal in the f4 frequency band.
[0103] Considering that the sound source parameters depend on the frequency response characteristics of the sound source, the calculated sound source parameters may vary with frequency. Based on this, this embodiment performs frequency band division processing on the original audio signal picked up by the pickup unit, calculates the sound source parameters corresponding to each frequency band, and renders the original audio signal based on the sound source parameters corresponding to each frequency band, so that the sound of each frequency band has a good spatial sense reproduction effect, further improving the spatial sense reproduction effect of the sound.
[0104] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of calculating the sound source parameters corresponding to the target frequency band based on the sub-band signal includes:
[0105] Step B10: Calculate spatial cue parameters based on the left and right channel sub-signals in the sub-band signals, and determine the sound source parameters corresponding to the sound source in the target frequency band based on the spatial cue parameters; wherein the left and right channel sub-signals belong to the same frequency band.
[0106] Spatial cue parameters refer to various factors that influence the perception of sound source location and size when sound propagates in three-dimensional space. These may include interaural time difference (ITD), interaural level difference (ILD), and interaural cross-correlation coefficient (IACC), though this embodiment does not impose specific limitations on these parameters.
[0107] The interaural time difference (ITD) refers to the difference in the time it takes for sound waves to reach each of the two ears. Because human ears are located on the sides of the head, when a sound source is located on the left or right side of the head, the sound waves reach the ear closer to the head earlier than the ear farther away. The brain detects this time difference to determine the location of the sound source relative to the head. By measuring the time difference in sound waves reaching both ears, the horizontal orientation of the sound source can be estimated. Specifically, the ITD can be calculated using known methods, such as... Where n represents the index of the sampling point (the index of the i-th sampling point is i-1), N represents the total number of sampling points, x1 and x2 represent the left channel sub-signal and the right channel sub-signal respectively, x1(n) represents the signal value of the sampling point with index n of the left channel sub-signal, x2(n) represents the signal value of the sampling point with index n of the right channel sub-signal, d represents the number of delay samples, and argmax d This indicates the value of d when the maximum value is returned.
[0108] Binaural sound level difference (ILD) refers to the difference in intensity or volume of sound waves reaching the two ears. Because human ears are located on opposite sides of the head, when a sound source is located on the left or right side, the sound waves will exhibit intensity differences during propagation due to factors such as distance, head obstruction, and auricular filtering. The intensity difference of sound waves reaching the two ears provides information about the horizontal orientation and distance of the sound source. When the sound source is biased to one side, the ear farther from the source will receive a reduced sound intensity. Specifically, ILD can be calculated using known methods, such as...
[0109] The binaural cross-correlation coefficient (IACC) is a parameter that indicates the similarity of sound signals received by both ears. It is a dimensionless parameter, ranging from +1 to -1. Specifically, the IACC can be calculated using known methods, such as... Where p1 represents the sound pressure level of the left vocal tract sub-signal, and p2 represents the sound pressure level of the right vocal tract sub-signal. The signal mean of the sound pressure level signal of the left vocal tract sub-signal. This represents the average sound pressure level of the right vocal tract sub-signal.
[0110] As one implementation method for calculating sound source parameters, a parameter calculation model can be pre-trained using machine learning. In this way, the spatial cue parameters can be input into the pre-trained parameter calculation model to output the sound source parameters.
[0111] As another implementation method, a mapping table can be set up based on the correlation between sound source parameters and spatial cue parameters. Sound source parameters can be found based on the mapping table. For example, the binaural time difference and binaural sound level difference are correlated with the sound source location. Based on this correlation, mapping relationships between binaural time difference and sound source location, and between binaural sound level difference and sound source location can be set respectively. The binaural hearing cross-correlation coefficient is correlated with the sound source size. Based on this correlation, a mapping relationship between the binaural hearing cross-correlation coefficient and the sound source size can be set.
[0112] The above are just two feasible implementation methods for calculating sound source parameters based on spatial cue parameters provided in this embodiment. This embodiment does not impose specific limitations on the specific implementation methods for calculating sound source parameters based on spatial cue parameters.
[0113] In one possible implementation, the step of determining the sound source parameters corresponding to the sound source in the target frequency band based on the spatial cue parameters includes:
[0114] Step C10: If the spatial cue parameters include the binaural time difference and the target frequency band belongs to a preset low-frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural time difference.
[0115] Step C20: If the spatial cue parameters include the binaural sound level difference and the target frequency band does not belong to the preset low frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural sound level difference.
[0116] The preset low-frequency band is a frequency band set in advance. Considering that the correlation between the binaural sound level difference and the sound source location is strong in the frequency band above 1.5kHz, and the correlation between the binaural time difference and the sound source location is strong in the frequency band below 1.5kHz, preferably, the preset low-frequency band is a frequency band less than or equal to 1.5kHz.
[0117] It should be noted that, in order to avoid the phenomenon that some frequency bands of a target frequency band belong to the preset low frequency band while other frequency bands do not belong to the preset low frequency band, such as the preset low frequency band being less than or equal to 1.5kHz, while a target frequency band is 0.8kHz to 2kHz, when setting the frequency boundary for frequency band division processing of the original audio signal, the set frequency boundary must at least include the frequency boundary of the preset low frequency band.
[0118] Step C30: If the spatial cue parameters include the binaural cross-correlation coefficient, then the sound source size corresponding to the sound source in the target frequency band is determined based on the binaural cross-correlation coefficient.
[0119] In this embodiment, considering that in the low-frequency band, the binaural time difference is more sensitive to changes in the sound source location and has a strong correlation with it, and in the high-frequency band, the binaural sound level difference is more sensitive to changes in the sound source location and has a strong correlation with it, the sound source location is determined based on the binaural time difference when the target frequency band belongs to the preset low-frequency band, and based on the binaural sound level difference when the target frequency band does not belong to the preset low-frequency band. This improves the accuracy of the determined sound source location. Furthermore, the binaural cross-correlation coefficient is more sensitive to changes in the sound source size and has a certain correlation with it; therefore, the sound source size is determined based on the binaural cross-correlation coefficient.
[0120] In one possible implementation, the step of calculating spatial cue parameters based on the left and right channel sub-signals in the sub-band signals includes:
[0121] Step D10: If there are multiple left channel sub-signals or multiple right channel sub-signals in the sub-band signal, then select one target left channel sub-signal and one target right channel sub-signal.
[0122] Understandably, when there are multiple omnidirectional microphones or figure-eight directional microphones, there will be multiple audio signals picked up, and there will also be multiple sub-audio signals in the sub-band signals. At this time, a target left channel sub-signal and a target right channel sub-signal can be randomly selected.
[0123] Step D20: Calculate spatial cue parameters based on the target left vocal tract sub-signal and the target right vocal tract sub-signal.
[0124] In this embodiment, when there are multiple left channel sub-signals or multiple right channel sub-signals, a target left channel sub-signal and a target right channel sub-signal are selected to calculate the spatial cue parameters. This reduces the number of audio signals involved in the calculation, thereby reducing the computational load.
[0125] In another feasible embodiment, a left channel sub-signal and a right channel sub-signal can be combined sequentially. A spatial cue parameter is calculated based on the currently combined left and right channel sub-signals. Finally, the average parameter of all the spatial cue parameters calculated is determined and used as the final calculated spatial cue parameter. In this way, the occurrence of accidental errors in a certain audio signal, such as accidental failure to pick up a signal, which may lead to incorrect calculated spatial cue parameters, can be reduced, thereby improving the accuracy of the final calculated spatial cue parameter.
[0126] For example, suppose an audio device has two omnidirectional microphones and two figure-eight microphones. The left channel audio signals picked up by the two omnidirectional microphones are L1 and L2, and the right channel audio signals picked up by the two figure-eight microphones are R1 and R2. The original audio signal M includes L1, L2, R1, and R2. After dividing M into two frequency bands f1 and f2, two sub-band signals S1 and S2 are obtained. S1 includes [L1_1, L2_1, R1_1, R2]. Let S1 be a subband signal. S2 includes L1_2, L2_2, R1_2, and R2_2. For S1, spatial cue parameters D1, D2, D3, and D4 can be calculated sequentially based on L1_1 and R1_1, D2, D3, and D4, respectively. The average of D1, D2, D3, and D4 is then used as the final spatial cue parameter for S1. It should be noted that when multiple spatial cue parameters are included, the average value of each parameter is taken separately.
[0127] In a preferred embodiment, the step of determining the sound source parameters corresponding to the sound source in the target frequency band based on the spatial cue parameters includes:
[0128] Step E10: Based on the preset mapping relationship and the spatial cue parameters, determine the sound source parameters corresponding to the sound source in the target frequency band.
[0129] Specifically, mapping relationships can be set according to the correspondence between different spatial cues and different sound source parameters. For example, different mapping relationships can be set between different binaural time differences and sound source locations, different binaural sound level differences and sound source locations, and different binaural hearing cross-correlation coefficients and sound source sizes. In this way, sound source parameters can be found based on the mapping relationships, reducing the amount of calculation and improving the calculation efficiency of sound source parameters.
[0130] Based on the first, second, and / or third embodiments of this application, in the fourth embodiment of this application, the content that is the same as or similar to the above-described embodiments one, two, and three can be referred to the above description and will not be repeated hereafter. In addition, the audio device further includes an intermediate pickup unit. After the step of calculating the sound source parameters corresponding to the sound source based on the original audio signal, the method further includes:
[0131] Step F10: Obtain the distance between the intermediate pickup unit and the sound source, and obtain the distance between the sound source and the reflecting surface;
[0132] Specifically, the reflecting surface can be the plane where the audio emitted by the sound source is first emitted during its propagation. For ease of subsequent explanation, the distance between the sound source and the reflecting surface will be denoted as d. b The distance between the intermediate pickup unit and the sound source is denoted as d. d .
[0133] For example, referring to Figure 2, a user wears an audio device with a speaker as the sound source. The audio played is first emitted onto a wall during its propagation process; this wall is the reflecting surface. b The distance between the speaker and the wall is d. The distance between the central pickup unit and the speaker is close to the distance between the user and the speaker. The distance between the user and the speaker is taken as d. d .
[0134] Step F20: The distance between the intermediate pickup unit and the sound source is taken as the first distance, and the distance between the sound source and the reflective surface is taken as the second distance;
[0135] Step F30: Calculate the first acoustic reflection time difference based on the first distance and the second distance;
[0136] The first acoustic reflection time difference is the rate of change of the time difference between the arrival of sound waves at the two ears. This parameter reflects how the time difference between the arrival of sound waves at the two ears changes with the direction of the sound source as the sound propagates in space, due to factors such as the location and shape of the sound source, as well as the shape of the listener's head and ears.
[0137] Specifically, the initial time-delay gap (ITDG) of acoustic reflection can be calculated based on a preset formula, which can be: Where c represents the speed of sound.
[0138] Step F40: Determine the spatial dimensions of the space where the sound source is currently located based on the first acoustic reflection time difference;
[0139] For example, referring to Figure 3, the vertical axis of Figure 3 represents ITDG in milliseconds, and the horizontal axis represents d. d The unit is meters, and the different values of d are... b (Figure 3 shows ITDG at depths of 0.6m, 0.7m, 0.8m, 0.9m, 1m, 1.1m, and 1.2m) as d d The curve showing the change in d can be seen. d Under certain circumstances, ITDG and d b Positive correlation, while d b It is usually positively correlated with the current spatial size, so the spatial size can be determined based on this correlation according to ITDG.
[0140] The step of rendering the original audio signal based on the sound source parameters to obtain the generated target audio signal includes:
[0141] Step F50: Render the original audio signal according to the spatial dimensions and the sound source parameters to obtain the generated target audio signal.
[0142] The original audio signal is rendered based on the spatial dimensions and sound source parameters to further improve the spatial reproduction effect of the generated target audio signal.
[0143] For example, to aid in understanding the technical concept or principle of signal generation in combination with Embodiments 1, 2, and 3 above, a specific embodiment is provided. In this specific embodiment, the signal generation process is as follows:
[0144] (1) First, the omnidirectional microphone picks up the omnidirectional microphone signal (i.e., the first audio signal).
[0145] (2) The figure-eight microphone pickup module obtains the figure-eight microphone signal (i.e., the second audio signal). Based on the positions of the omnidirectional microphone and the figure-eight microphone, the relative delay between the figure-eight microphone and the omnidirectional microphone is calculated as follows: delay=d / c*fs.
[0146] (3) After delay processing, the output of the figure-eight microphone is subjected to time-frequency domain equalization filtering. The amplitude and phase responses of the figure-eight microphone output are calibrated and aligned with those of the omnidirectional microphone to ensure that the amplitude responses of the omnidirectional and figure-eight microphones are consistent within the 0-180 degree range, with only a phase difference. The time-frequency domain filtering methods include, but are not limited to, NLMS filtering, convolution, IIR filtering, FIR filtering, Fourier transform filtering, wavelet transform filtering, and equalization processing.
[0147] (4) The processed omnidirectional microphone signal and figure-eight microphone signal are summed and subtracted. When their amplitude and phase responses are in the same direction, they are kept consistent. The two signals are then added together to obtain a beam pointing directly forward (i.e., the left channel signal). At 180 degrees, the signals are reversed and subtracted to obtain a beam pointing directly backward (i.e., the right channel signal), thus achieving separation of the left and right channels. Information in the 90-degree direction can be provided by adjusting the amplitude of the omnidirectional microphone.
[0148] The signals from the left and right sides can be obtained by amplitude processing of the omnidirectional microphone signals. Based on the weighted average of the left and right channel amplitudes as the target amplitude, gain control is applied to the omnidirectional microphone output to obtain the center channel output information with appropriate relative intensity.
[0149] (6) Perform time-frequency transformation and frequency band division on the dual-channel signal to obtain sub-band signals belonging to four frequency bands respectively. The four frequency bands are low frequency (less than 200Hz), mid-low frequency (200Hz~1.5kHz), mid-high frequency (1.5kHz~4kHz) and high frequency (greater than 4kHz).
[0150] (7) Extract spatial cue parameters from the sub-band signals of the four frequency bands respectively; the spatial cue parameters include binaural time difference (ITD), binaural sound level difference (ILD), and binaural auditory cross-correlation coefficient (IACC).
[0151] (8) The sound source size is obtained by IACC calculation. The sound source location is calculated in the low and mid-low frequencies according to the correspondence between ITD and sound source location. In the mid-high and high frequencies, the sound source location is calculated according to the correspondence between ILD and sound source location.
[0152] (9) Calculate ITDG and estimate the space size based on ITDG.
[0153] (10) Render the original audio signal according to the sound source location, sound source size and spatial dimensions to generate the target audio signal.
[0154] As can be seen, the final generated target audio signal fully integrates the location information of the sound source, which can well express the spatial information of the sound signal. After playback, it can achieve the effect of human hearing, resulting in a better user experience.
[0155] It should be noted that the above specific embodiments are only used to understand this application and do not constitute a limitation on the signal generation process of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0156] Furthermore, this application embodiment also provides a signal generation device. Referring to FIG4, the signal generation device includes an omnidirectional microphone and a figure-eight directional microphone, and the signal generation device further includes:
[0157] The acquisition module 10 is used to acquire the original audio signal picked up, wherein the original audio signal includes a first audio signal and a second audio signal, the first audio signal is the audio signal picked up by the omnidirectional microphone, and the second audio signal is the audio signal picked up by the figure-eight directional microphone;
[0158] The separation module 20 is used to perform channel separation processing on the original audio signal to obtain a two-channel signal, wherein the two-channel signal includes a left channel signal and a right channel signal;
[0159] The rendering module 30 is used to calculate the sound source parameters corresponding to the sound source based on the dual-channel signal, and render the original audio signal according to the sound source parameters to obtain the generated target audio signal, wherein the sound source parameters include the sound source orientation and / or the sound source size.
[0160] In one embodiment, the separation module 20 is further configured to:
[0161] Alignment processing is performed on the first audio signal and the second audio signal to obtain an aligned audio signal, wherein the alignment processing includes time alignment processing and / or frequency response alignment processing, and the aligned audio signal includes a third audio signal and a fourth audio signal, wherein the third audio signal is the first audio signal after alignment processing, and the fourth audio signal is the second audio signal after alignment processing;
[0162] The left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0163] The left channel signal and the right channel signal are determined to be stereo signals.
[0164] In one embodiment, the separation module 20 is further configured to:
[0165] For the third audio signal and the fourth audio signal of the initial frame, the left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal.
[0166] The signal value of the third audio signal is determined to be a first signal value, the signal value of the left channel signal is determined to be a second signal value, and the signal value of the right channel signal is determined to be a third signal value.
[0167] The target signal value is obtained by calculating the weighted sum of the second signal value and the third signal value.
[0168] A target gain is determined based on the first signal value and the target signal value, wherein the target gain is negatively correlated with the first signal value and positively correlated with the target signal value;
[0169] The third audio signal of the next frame is adjusted with the target gain to obtain an intermediate audio signal. The intermediate audio signal is determined to be the new third audio signal. Based on the new third audio signal and the fourth audio signal of the next frame, the steps of adding the third audio signal to the fourth audio signal to obtain the left channel signal and subtracting the phase-inverted fourth audio signal from the third audio signal to obtain the right channel signal are returned to the execution.
[0170] Once the preset separation termination condition is met, all the obtained left channel signals and all the obtained right channel signals are determined to be stereo signals.
[0171] In one embodiment, the rendering module 30 is further configured to:
[0172] The dual-channel signal is divided into multiple sub-band signals by frequency band division.
[0173] For each sub-band signal, determine the target frequency band to which the sub-band signal belongs, and calculate the sound source parameters corresponding to the sound source in the target frequency band based on the sub-band signal.
[0174] In one embodiment, the rendering module 30 is further configured to:
[0175] Spatial cue parameters are calculated based on the left and right channel sub-signals in the sub-band signals, and the sound source parameters corresponding to the sound source in the target frequency band are determined based on the spatial cue parameters.
[0176] The left channel sub-signal and the right channel sub-signal belong to the same frequency band.
[0177] In one embodiment, the rendering module 30 is further configured to:
[0178] If the spatial cue parameters include the binaural time difference and the target frequency band belongs to a preset low-frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural time difference.
[0179] If the spatial cue parameters include the binaural sound level difference and the target frequency band does not belong to the preset low frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural sound level difference.
[0180] If the spatial cue parameters include the binaural cross-correlation coefficient, then the sound source size corresponding to the sound source in the target frequency band is determined based on the binaural cross-correlation coefficient.
[0181] In one embodiment, the rendering module 30 is further configured to:
[0182] If there are multiple left channel sub-signals or multiple right channel sub-signals in the sub-band signal, then select one target left channel sub-signal and one target right channel sub-signal.
[0183] Spatial cue parameters are calculated based on the target left vocal tract sub-signal and the target right vocal tract sub-signal.
[0184] Furthermore, this application also proposes an audio device, which includes an omnidirectional microphone, a figure-eight directional microphone, and a processor. The omnidirectional microphone and the figure-eight directional microphone are respectively connected to the processor, and the processor is used to execute the steps of the signal generation method described above.
[0185] As shown in Figure 7, the audio device may further include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the audio device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows audio devices to communicate wirelessly or wiredly with other devices to exchange data. While audio devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0186] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0187] The audio device provided in this application, employing the signal generation method described in the above embodiments, can solve the technical problem of how to improve the universality of stereo recording. Compared with the prior art, the beneficial effects of the audio device provided in this application are the same as those of the signal generation method described in the above embodiments, and other technical features of the audio device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0188] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0189] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0190] In addition, to achieve the above objectives, this application also provides a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the signal generation method described in the above embodiments.
[0191] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0192] The aforementioned computer-readable storage medium may be included in an audio device or may exist independently without being assembled into an audio device.
[0193] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an audio device, cause the audio device to: acquire a picked-up raw audio signal, wherein the raw audio signal includes a first audio signal and a second audio signal, the first audio signal being an audio signal picked up by the omnidirectional microphone and the second audio signal being an audio signal picked up by the figure-eight directional microphone; perform channel separation processing on the raw audio signal to obtain a dual-channel signal, wherein the dual-channel signal includes a left channel signal and a right channel signal; calculate the sound source parameters corresponding to the sound source based on the dual-channel signal; render the raw audio signal according to the sound source parameters to obtain a generated target audio signal, wherein the sound source parameters include the sound source orientation and / or sound source size.
[0194] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0195] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0196] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0197] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described signal generation method, thereby solving the technical problem of how to improve the universality of stereo recording. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the signal generation method provided in the above embodiments, and will not be repeated here.
[0198] Furthermore, embodiments of this application also propose a computer program product, including a signal generation program, which, when executed by a processor, implements the steps of the signal generation method described above.
[0199] The specific implementation of the computer program product in this application is basically the same as the various embodiments of the signal generation method described above, and will not be repeated here.
[0200] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0201] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0203] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A signal generation method characterized by, Applied to an audio device, the audio device including an omnidirectional microphone and a figure-eight directional microphone, the signal generation method includes the following steps: Acquire the original audio signal, wherein the original audio signal includes a first audio signal and a second audio signal, the first audio signal being the audio signal picked up by the omnidirectional microphone, and the second audio signal being the audio signal picked up by the figure-eight directional microphone; The original audio signal is subjected to channel separation processing to obtain a two-channel signal, wherein the two-channel signal includes a left channel signal and a right channel signal; Based on the dual-channel signal, the sound source parameters corresponding to the sound source are calculated, and the original audio signal is rendered according to the sound source parameters to obtain the generated target audio signal. The sound source parameters include the sound source orientation and / or the sound source size.
2. The signal generating method of claim 1, wherein, The step of performing channel separation processing on the original audio signal to obtain a two-channel signal includes: Alignment processing is performed on the first audio signal and the second audio signal to obtain an aligned audio signal, wherein the alignment processing includes time alignment processing and / or frequency response alignment processing, and the aligned audio signal includes a third audio signal and a fourth audio signal, wherein the third audio signal is the first audio signal after alignment processing, and the fourth audio signal is the second audio signal after alignment processing; The left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal. The left channel signal and the right channel signal are determined to be stereo signals.
3. The signal generating method of claim 2, wherein, After the step of aligning the first audio signal and the second audio signal to obtain an aligned audio signal, the method further includes: For the third audio signal and the fourth audio signal of the initial frame, the left channel signal is obtained by adding the third audio signal to the fourth audio signal, and the right channel signal is obtained by subtracting the phase-inverted fourth audio signal from the third audio signal. The signal value of the third audio signal is determined to be a first signal value, the signal value of the left channel signal is determined to be a second signal value, and the signal value of the right channel signal is determined to be a third signal value. The target signal value is obtained by calculating the weighted sum of the second signal value and the third signal value. A target gain is determined based on the first signal value and the target signal value, wherein the target gain is negatively correlated with the first signal value and positively correlated with the target signal value; The third audio signal of the next frame is adjusted with the target gain to obtain an intermediate audio signal. The intermediate audio signal is determined to be the new third audio signal. Based on the new third audio signal and the fourth audio signal of the next frame, the steps of adding the third audio signal to the fourth audio signal to obtain the left channel signal and subtracting the phase-inverted fourth audio signal from the third audio signal to obtain the right channel signal are returned to the execution. Once the preset separation termination condition is met, all the obtained left channel signals and all the obtained right channel signals are determined to be stereo signals.
4. The signal generating method of claim 1, wherein, The step of calculating the sound source parameters corresponding to the sound source based on the dual-channel signal includes: The dual-channel signal is divided into multiple sub-band signals by frequency band division. For each sub-band signal, determine the target frequency band to which the sub-band signal belongs, and calculate the sound source parameters corresponding to the sound source in the target frequency band based on the sub-band signal.
5. The signal generating method of claim 4, wherein, The step of calculating the sound source parameters corresponding to the target frequency band based on the sub-band signal includes: Spatial cue parameters are calculated based on the left and right channel sub-signals in the sub-band signals, and the sound source parameters corresponding to the sound source in the target frequency band are determined based on the spatial cue parameters. The left channel sub-signal and the right channel sub-signal belong to the same frequency band.
6. The signal generating method of claim 5, wherein, The step of determining the sound source parameters corresponding to the sound source in the target frequency band based on the spatial cue parameters includes: If the spatial cue parameters include the binaural time difference and the target frequency band belongs to a preset low-frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural time difference. If the spatial cue parameters include the binaural sound level difference and the target frequency band does not belong to the preset low frequency band, then the sound source location corresponding to the target frequency band is determined based on the binaural sound level difference. If the spatial cue parameters include the binaural cross-correlation coefficient, then the sound source size corresponding to the sound source in the target frequency band is determined based on the binaural cross-correlation coefficient.
7. The signal generating method of claim 5, wherein, The step of calculating spatial cue parameters based on the left and right channel sub-signals in the sub-band signal includes: If there are multiple left channel sub-signals or multiple right channel sub-signals in the sub-band signal, then select one target left channel sub-signal and one target right channel sub-signal. Spatial cue parameters are calculated based on the target left vocal tract sub-signal and the target right vocal tract sub-signal.
8. An audio device, comprising: The audio device includes an omnidirectional microphone, a figure-eight directional microphone, and a processor. The omnidirectional microphone and the figure-eight directional microphone are respectively connected to the processor, and the processor is used to execute the steps of the signal generation method as described in any one of claims 1 to 7.
9. A readable storage medium, characterized by, The readable storage medium is a computer-readable storage medium, on which a program implementing the signal generation method is stored, and the program implementing the signal generation method is executed by a processor to implement the steps of the signal generation method as described in any one of claims 1 to 7.
10. A computer program product, characterised in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the signal generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to dirac based spatial audio coding using direct component compensation
CN113424257A
Audio signal processing method, electronic equipment and computer readable storage medium
CN115802274A
Signal generation method and device, readable storage medium and computer program product
CN119421077A
Miniature directional recording device and electronic equipment
CN220043611U
Synthesizing method for stereophonic music
JP1993153699A