A multi-channel audio signal acquisition method, device and system
Through the audio signal processing of the main device and the additional device, combined with multi-channel rendering and ambient sound suppression technology, the problem of poor audio signal recording effect in multiple sound sources or noisy environments is solved, achieving better audio signal recording effect and immersive experience.
Patent Information
- Application Number
- CN202011027264.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-09-25
AI Technical Summary
In the prior art, in multi-sound sources or noisy environments, distributed audio signals have poor recording effects and severe environmental sound impact, resulting in the sounds of interest to users being flooded with background noise.
The audio signal is obtained through the main device and the additional device, and multi-channel rendering and ambient sound suppression processing are performed. Combined with adaptive filtering and spatial filtering technology, ambient sound is suppressed and the target audio signal is enhanced to achieve mixing and rendering of the audio signal.
It improves the recording effect of audio signals, simulates point-shaped auditory goals in the spatial sound field, creates an immersive experience, and enhances the focus and tracking of the sounds that are of interest to users.
Smart Images

Figure CN114255781B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio technology, and particularly to a method, device and system for acquiring multi-channel audio signals. Background Art
[0002] With the progress of technology, people have put forward higher requirements for the photography and recording effects of mobile devices. Currently, with the popularization of true wireless stereo (TWS) Bluetooth headsets, a distributed audio capture scheme has emerged. This scheme uses the microphones on TWS Bluetooth headsets to capture high-quality close-up audio signals far from the user, and mixes and binaurally renders the spatial audio signals collected by the microphone array on the main device, simulating a point-like auditory target in the spatial sound field and creating a more realistic immersive experience. However, this scheme only mixes the distributed audio signals and does not suppress the ambient sound. When using a mobile device to shoot a video in a place with multiple sound sources or a relatively noisy environment, the sound that the user is really interested in will be mixed with various irrelevant sound sources and even submerged in the background noise. Therefore, the existing scheme may result in poor recording effects of audio signals due to the influence of ambient sound. Summary of the Invention
[0003] Embodiments of the present invention provide a method, device and system for acquiring multi-channel audio signals, which can use the relationship between distributed audio signals to suppress ambient sound and improve the recording effect of audio signals.
[0004] To solve the above technical problems, the embodiments of the present invention are implemented as follows:
[0005] In a first aspect, an embodiment of the present invention provides a method for acquiring multi-channel audio signals, including:
[0006] Acquire the main audio signal collected when the main device shoots a video, and perform multi-channel rendering to obtain an ambient multi-channel audio signal;
[0007] Acquire the audio signal collected by the additional device, and determine the first additional audio signal; wherein, the distance between the additional device and the target shooting object is less than a first threshold;
[0008] Perform ambient sound suppression processing on the first additional audio signal and the main audio signal to obtain a target audio signal;
[0009] Perform multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal;
[0010] Mix the ambient multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal.
[0011] Second aspect, there is provided a multi-channel audio signal acquisition device, including:
[0012] An acquisition module, configured to acquire a main audio signal collected when a main device performs video shooting on a target object, and perform first multi-channel rendering to obtain an ambient multi-channel audio signal; acquire an audio signal collected by an additional device, and determine a first additional audio signal, where the distance between the additional device and the target object is less than a first threshold;
[0013] A processing module, configured to perform ambient sound suppression processing on the first additional audio signal and the main audio signal to obtain a target audio signal;
[0014] Perform multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal;
[0015] Mix the ambient multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal.
[0016] Third aspect, there is provided a terminal device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the multi-channel audio signal acquisition method as in the first aspect.
[0017] Fourth aspect, there is provided a terminal device, including: the multi-channel audio signal acquisition device as in the second aspect and a main device,
[0018] The main device is configured to collect a main audio signal when shooting a video, and send the main audio signal to the multi-channel audio signal acquisition device.
[0019] Fifth aspect, there is provided a multi-channel audio signal acquisition system, the system including: the multi-channel audio signal acquisition device as in the second aspect, a main device, and an additional device, where the main device and the additional device are respectively communicatively connected to the multi-channel audio signal;
[0020] The main device is configured to collect a main audio signal when shooting a video, and send the main audio signal to the multi-channel audio signal acquisition device;
[0021] The additional device is configured to collect a second additional audio signal, and send the second additional audio signal to the multi-channel audio signal acquisition device;
[0022] where the distance between the additional device and the target object is less than a first threshold.
[0023] Sixth aspect, there is provided a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, where when the computer program is executed by the processor, it implements the multi-channel audio signal acquisition method as in the first aspect.
[0024] In an embodiment of the present invention, the main audio signal collected when the main device shoots a video can be obtained and multi-channel rendering can be performed to obtain an environmental multi-channel audio signal; and the audio signal collected by an additional device whose distance from the target shooting object is less than a first threshold can be obtained to determine a first additional audio signal; environmental noise suppression processing is performed on the first additional audio signal and the main audio signal to obtain a target audio signal; multi-channel rendering is performed on the target audio signal to obtain a target multi-channel audio signal; and the environmental multi-channel audio signal and the target multi-channel audio signal are mixed to obtain a mixed multi-channel audio signal. Through this solution, distributed audio signals can be obtained from the main device and the additional device, and the relationship between the distributed audio signals can be utilized. Based on the first additional audio signal obtained from the audio signal collected by the additional device and the main audio signal collected by the main device, environmental noise suppression processing is performed to suppress the environmental noise during the recording process to obtain a target multi-channel audio signal. Then, when the environmental multi-channel audio signal (obtained by performing multi-channel rendering on the main audio signal) and the target multi-channel audio signal are mixed, not only is the mixing of the distributed audio signals realized to simulate the point auditory target in the spatial sound field, but also the environmental noise is suppressed, thereby improving the recording effect of the audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments and the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and other drawings can be obtained according to these drawings.
[0026] Figure 1 The figure shows a schematic diagram of a multi-channel audio signal acquisition system provided by an embodiment of the present invention;
[0027] Figure 2A The figure shows a schematic diagram of a multi-channel audio signal acquisition method provided by an embodiment of the present invention Figure 1 ;
[0028] Figure 2B The figure shows a schematic diagram of the interface of a terminal device provided by an embodiment of the present invention;
[0029] Figure 3 The figure shows a second schematic diagram of a multi-channel audio signal acquisition method provided by an embodiment of the present invention;
[0030] Figure 4 The figure shows a schematic diagram of a multi-channel audio signal acquisition device provided by an embodiment of the present invention;
[0031] Figure 5 The figure shows a schematic diagram of the structure of a terminal device provided by an embodiment of the present invention;
[0032] Figure 6 The following is a schematic diagram of the hardware structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner. In addition, in the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality of" refers to two or more.
[0035] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0036] The embodiments of the present invention provide a multi-channel audio signal acquisition method, device and system, which can be applied in a video shooting scenario, especially in a scenario with multiple sound sources or a noisy environment for video shooting. It can mix distributed audio signals, simulate a point auditory target in a spatial sound field, and suppress ambient sound, thereby improving the recording effect of audio signals.
[0037] As Figure 1 shown, it is a schematic diagram of a multi-channel audio signal acquisition system provided by an embodiment of the present invention. The system may include a main device, an additional device and an audio processing device (which can be the multi-channel audio acquisition device in the embodiment of the present invention). Among them, Figure 1 The additional device among them is a TWS Bluetooth headset, which can be used to collect an audio stream (that is, the additional audio signal in the embodiment of the present invention). The main device can be used to collect a video stream and an audio stream (that is, the main audio signal in the embodiment of the present invention). The audio processing device may include the following modules: target tracking, scene sound source classification, delay compensation, adaptive filtering, spatial filtering, binaural rendering and mixer, etc. Among them, the specific function introduction of each module will be described in combination with the multi-channel audio signal acquisition method described in the following embodiments, and will not be elaborated here.
[0038] It should be noted that the master device and the audio processing device in the embodiments of the present invention can be two independent devices. Optionally, the master device and the audio processing device can also be an integrated device, for example, it can be a terminal device integrating the functions of the master device and the audio processing device.
[0039] In the embodiments of the present invention, the additional device and the terminal device, or the additional device and the audio processing device can be connected by a wireless communication method, for example, can be connected by Bluetooth or by WiFi. The connection method is not specifically limited in the embodiments of the present invention.
[0040] The terminal device in the embodiments of the present invention can include: mobile phones, tablet computers, laptop computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, personal digital assistants (PDAs), wearable devices (such as watches, wristbands, glasses, helmets, headbands, etc.) and other terminal devices. The specific form of the terminal device is not particularly limited in the embodiments of the present application.
[0041] In the embodiments of the present invention, the additional device can be a terminal device independent of the master device and the audio processing device. The mobile terminal device can be a portable terminal device, for example, it can be a Bluetooth headset, wearable devices (such as watches, wristbands, glasses, helmets, headbands, etc.) and other terminal devices.
[0042] In a video shooting scenario, the master device can shoot a video, obtain the main audio signal and send it to the audio processing device, while the additional device is closer to a certain target shooting object in the video shooting scenario (for example, the distance between the two is less than the first threshold), and obtain the additional audio device, and then send it to the audio processing device.
[0043] Optionally, the target shooting object can be a person or an instrument in the video shooting scenario, etc.
[0044] Optionally, in a video shooting scenario, there can be multiple shooting objects, and the target shooting object can be one of the multiple shooting objects.
[0045] Figure 2A As shown, it is a schematic diagram of a multi-channel audio signal acquisition method provided in the embodiments of the present invention. Exemplarily, the execution subject of this method can be the audio processing device (i.e., the multi-channel audio acquisition device) as shown above Figure 1 as shown, or can be an integrated above Figure 1An audio processing device and a terminal device with master device functions are shown. At this time, the master device can be a functional module or entity for collecting audio and video in the terminal device. In the following embodiments, the terminal device is used as the execution subject for exemplary description.
[0046] The following is a detailed introduction to this method. As Figure 2A shown, this method includes:
[0047] 201. Obtain the main audio signal collected when the master device performs video shooting on the target shooting object, and perform first multi-channel rendering to obtain an environmental multi-channel audio signal.
[0048] Among them, the distance between the target shooting object and the additional device can be less than the first threshold.
[0049] Optionally, the user can set the additional device on the target shooting object to be tracked, start video shooting on the terminal device, and select the target shooting object in the video content by clicking on the video content displayed on the screen. The sound collection module on the master device in the terminal device and the sound collection module on the additional device can start recording to collect audio signals.
[0050] Optionally, the sound collection module on the master device can be a microphone array, and the main audio signal is collected through this microphone array. The sound collection module on the additional device can be a microphone.
[0051] As Figure 2B shown, it can be a schematic diagram of an interface of the terminal device. The video content can be displayed on the screen of the terminal device. Among them, the user can click on the person 21 displayed in this interface using a mobile phone to determine the person 21 as the target shooting object. A Bluetooth headset (i.e., the above-mentioned additional device) can be carried on the person 21 to collect the audio signal near the person 21 and send it to the terminal device.
[0052] In the embodiments of the present invention, multi-channel can refer to dual-channel, four-channel, 5.1 or more channels.
[0053] When the audio signal obtained in the embodiments of the present invention is a dual-channel audio signal, the main audio signal can be binaurally rendered through a head related transfer function (HRTF) to obtain an environmental binaural audio signal.
[0054] Exemplarily, the main audio signal can be binaurally rendered through the Figure 1 binaural renderer in to obtain an environmental binaural audio signal.
[0055] 202. Obtain the audio signal collected by the additional device and determine the first additional audio signal.
[0056] Optionally, obtaining the audio signal collected by the additional device on the target object and determining the first additional audio signal may include two implementation manners:
[0057] The first implementation manner: obtaining the second additional audio signal collected by the additional device on the target object and determining the second additional audio signal as the first additional audio signal;
[0058] The second implementation manner: obtaining the second additional audio signal collected by the additional device on the target object, aligning the second additional audio signal with the main audio signal in the time domain to obtain the first additional audio signal.
[0059] Since there may be a certain distance between the main device and the additional device, there may be a certain time delay between the obtained main audio signal and the second additional audio signal. The main audio signal and the second additional audio signal can be aligned in the time domain according to the time delay between the main audio signal and the second additional audio signal to obtain the first additional audio signal.
[0060] Generally, in an audio signal acquisition system, for example, Figure 1 in the multi-channel audio signal acquisition system shown, there will also be a certain system time delay (for example, the time delay caused by Bluetooth transmission and the time delay caused by the decoding module for decoding), and this system time delay can be obtained through testing. Optionally, in the embodiments of the present invention, the actual time delay can be obtained by combining the estimated sound wave propagation time delay (that is, the time delay between the above-mentioned main audio signal and the second additional audio signal) with the system time delay, and the main audio signal and the second additional audio signal can be aligned in the time domain according to this actual time delay to obtain the first additional audio signal.
[0061] Figure 1 The delay compensator in can be used to align the additional audio signal with the main audio signal in the time domain according to the time delay between the main audio signal and the second additional audio signal to obtain the first additional audio signal.
[0062] 203. Perform ambient sound suppression processing on the first additional audio signal and the main audio signal to obtain the target audio signal.
[0063] In the embodiments of the present invention, for the case where the target object is within the shooting field of view of the main device and for the case where the target object is outside the shooting field of view of the main device, the manner of obtaining the target audio signal by performing ambient sound suppression processing on the first additional audio signal and the main audio signal is different.
[0064] (1) For the case where the target object is within the shooting field of view of the main device.
[0065] According to the shooting field of view of the master device, perform spatial filtering on the master audio signal in the area outside the shooting field of view of the master device to obtain a reverse focused audio signal; use the reverse focused audio signal as a reference signal to perform adaptive filtering processing on the first additional audio signal to obtain a target audio signal.
[0066] This method first performs spatial filtering on the master audio signal in the area outside the shooting field of view of the master device to obtain a reverse focused audio signal, suppressing the sound components at the position of the target object included in the master audio signal and obtaining a purer ambient audio signal. Then, using the reverse focused audio signal as a reference signal to perform adaptive filtering processing on the first additional audio signal can further suppress the ambient sound in the additional audio signal.
[0067] (2) For the case where the target object is outside the shooting field of view of the master device.
[0068] According to the shooting field of view of the master device, perform spatial filtering on the master audio signal in the area within the shooting field of view to obtain a focused audio signal; use the first additional audio signal as a reference signal to perform adaptive filtering processing on the focused audio signal to obtain a target audio signal.
[0069] This method first performs spatial filtering on the master audio signal in the area within the shooting field of view to obtain a focused audio signal, suppressing some of the ambient sound in the master audio signal. Then, using the first additional audio signal as a reference signal to perform adaptive filtering processing on the focused audio signal can further suppress the ambient sound outside the focused area that has not been fully suppressed in the focused audio signal, especially the sound components at the position of the target object included in the ambient sound.
[0070] Figure 1 The spatial filter in [ID] can be used to perform spatial filtering on the master audio signal to obtain a directionally enhanced audio signal. When the target object is within the shooting field of view of the master device, since a high-quality close-up audio signal has been obtained through the first additional audio signal, the main purpose of spatial filtering is to obtain a purer ambient audio signal, and the target area of spatial filtering is the area outside the shooting field of view, and the obtained signal is called a reverse focused audio signal; while when the target object is outside the shooting field of view of the master device, since it is necessary to obtain a close-up audio signal of the area within the shooting field of view through spatial filtering, the target area of spatial filtering is then the area within the shooting field of view, and the obtained signal is a focused audio signal.
[0071] Among them, the method of spatial filtering can be a beamforming-based method, such as the minimum variance distortionless response (MVDR) method, or the beamforming method using a generalized sidelobe canceller (GSC), etc.
[0072] Figure 1 It includes two groups of adaptive filters, and these two groups of adaptive filters act on the target audio signals obtained in the above two cases respectively. Specifically, according to the change of the target object in the shooting field of view, only one group of adaptive filters can be enabled. When the target object is within the shooting field of view of the main device, the adaptive filter acting on the first additional audio signal is activated, and the reverse focused audio signal is input as a reference signal to further suppress the ambient sound from the first additional audio signal, making the sound near the target object more prominent. When the target object is outside the shooting field of view of the main device, the adaptive filter acting on the focused audio signal is activated, and the first additional audio signal is input as a reference signal to further suppress the sound outside the shooting field of view, especially the sound at the position of the target object.
[0073] Among them, the method of adaptive filtering can be the least mean square (LMS) method, etc.
[0074] 204. Perform second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal.
[0075] Exemplarily, Figure 1 The three groups of binaural renderers in act on the main audio signal, the target audio signal after adaptive filtering in the above case (1), and the target audio signal after adaptive filtering in the above case (2) respectively to obtain three groups of binaural signals: ambient binaural signal, additional binaural signal, and focused binaural signal.
[0076] Among them, since the above case (1) and (2) do not exist simultaneously, the binaural renderer acting on the target audio signal in the above case (1) and the binaural renderer acting on the target audio signal in the above case (2) can be enabled non-simultaneously, and can be selected and enabled according to the change of the target object in the shooting field of view of the main device. The binaural renderer acting on the main audio signal is always enabled.
[0077] Further, when the target object is within the shooting field of view of the main device, the binaural renderer for the target audio signal obtained in the above situation (1) is enabled. When the target object is outside the shooting field of view of the main device, the binaural renderer for the target audio signal obtained in the above situation (2) is enabled.
[0078] Optionally, the binaural renderer may include a decorrelator and a convolver inside, and requires the HRTF corresponding to the target position to simulate the perception of the auditory target in the desired direction and distance.
[0079] Optionally, the scene sound source classification module can be used to determine the rendering rule according to the determined current scene and the sound source type of the target object. The determined rendering rule can be applied to the decorrelator to obtain different rendering styles. The azimuth angle and distance between the accessory device and the main device can be used to control the generation of the HRTF. The HRTF corresponding to a specific position can be obtained by interpolating on a pre-stored set of HRTFs, or can also be obtained using a method based on a deep neural network (DNN).
[0080] 205. Mix the environmental multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal.
[0081] In the embodiment of the present invention, mixing the environmental multi-channel audio signal and the target multi-channel audio signal means adding the environmental multi-channel audio signal and the target multi-channel audio signal according to the gain. Specifically, when adding the environmental multi-channel audio signal and the target multi-channel audio signal according to the gain, it can be adding the signal sampling points in the environmental multi-channel audio signal and adding the signal sampling points in the target multi-channel audio signal.
[0082] Among them, the gain can be a preset fixed value or a variable gain.
[0083] Optionally, the variable gain can be specifically determined according to the shooting field of view.
[0084] Figure 1 The mixer in is used to mix two of the foregoing three groups of binaural signals. When the target object is within the shooting field of view of the main device, the environmental binaural signal and the accessory binaural signal are mixed; when the target object is outside the shooting field of view of the main device, the environmental binaural signal and the focused binaural signal are mixed.
[0085] In an embodiment of the present invention, a main audio signal collected when a main device captures a video can be obtained and subjected to first multi-channel rendering to obtain an environmental multi-channel audio signal; and an audio signal collected by an additional device whose distance from a target object to be captured is less than a first threshold can be obtained, and a first additional audio signal can be determined; environmental noise suppression processing is performed on the first additional audio signal and the main audio signal to obtain a target audio signal; the target audio signal is subjected to second multi-channel rendering to obtain a target multi-channel audio signal; the environmental multi-channel audio signal and the target multi-channel audio signal are mixed to obtain a mixed multi-channel audio signal. Through this solution, distributed audio signals can be obtained from the main device and the additional device, and the relationship between the distributed audio signals can be utilized. Based on the first additional audio signal obtained from the audio signal collected by the additional device and the main audio signal collected by the main device, environmental noise suppression processing is performed to suppress environmental noise during the recording process and obtain a target multi-channel audio signal. Then, when the environmental multi-channel audio signal (obtained by performing multi-channel rendering on the main audio signal) is mixed with the target multi-channel audio signal, not only is the mixing of the distributed audio signals achieved to simulate a point auditory target in the spatial sound field, but also the environmental noise is suppressed, thereby improving the recording effect of the audio signal.
[0086] As Figure 3 shown, an embodiment of the present invention further provides a method for obtaining a multi-channel audio signal, and the method includes:
[0087] 301. Obtain a main audio signal collected by a microphone array on a main device.
[0088] 302. Obtain a second additional audio signal collected by an additional device.
[0089] After the user selects a target object to be captured on the main device and starts capturing a video, the terminal device can execute the above 301 and 302, and the terminal device can continuously respond to changes in the shooting field of view of the main device to track the movement of the target object to be captured in the shooting field of view.
[0090] Optionally, video data (including the main audio signal) captured by the main device and a second additional audio signal collected by the additional device can be obtained.
[0091] Furthermore, the current scene category and the target object category can be determined based on the above video data and / or the second additional audio signal, and through a rendering rule that matches the current scene category and the target object category. And according to the determined rendering rule, multi-channel rendering is performed on subsequent audio signals.
[0092] Optionally, according to the determined rendering rules, perform second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal, and perform first multi-channel rendering on the main audio signal according to the determined rendering rules to obtain an ambient multi-channel audio signal.
[0093] Optionally, according to the determined rendering rules, perform multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal, which may include:
[0094] Obtain the video data captured by the main device and the second additional audio signal collected by the additional device;
[0095] Determine the current scene category and the target shooting object category;
[0096] Perform multi-channel rendering on the target audio signal through a first rendering rule that matches the current scene category and the target shooting object category to obtain a target multi-channel audio signal.
[0097] Optionally, according to the determined rendering rules, perform multi-channel rendering on the main audio signal to obtain an ambient multi-channel audio signal, which may include:
[0098] Obtain the main audio signal collected by the main device when shooting a video of the target shooting object;
[0099] Determine the current scene category;
[0100] Perform first multi-channel rendering on the main audio signal through a second rendering rule that matches the current scene category to obtain an ambient multi-channel audio signal.
[0101] Figure 1 In [description], the scene sound source classification module can include two paths, one using video stream information and the other using audio stream information. Both paths consist of a scene analyzer and a voice / instrument classifier. Among them, the scene analyzer can analyze the space type where the current user is located from the video or audio, such as a small room, a medium-sized room, a large room, a concert hall, a stadium, outdoors, etc. The voice / instrument classifier analyzes the sound source type near the current target shooting object from the video or audio, such as male voice, female voice, children's voice, or accordion, guitar, bass, piano, keyboard, and percussion instruments, etc.
[0102] Optionally, both the scene analyzer and the voice / instrument classifier can be DNN-based methods. The input of the video is the image of each frame, and the input of the audio can be the Mel spectrum of the sound or the Mel-frequency cepstrum coefficient (MFCC).
[0103] Optionally, it is also possible to combine the spatial scene analysis and the results obtained by the voice / instrument classifier with the user's preference settings to determine the rendering rules to be used in the subsequent binaural rendering module.
[0104] 303. Generate a first multi-channel transfer function according to the microphone array configuration on the main device, and perform multi-channel rendering on the main audio signal according to the first multi-channel transfer function to obtain an environmental multi-channel audio signal.
[0105] It should be noted that when the multi-channel in the embodiment of the present invention is a two-channel, the above first multi-channel transfer function can be an HRTF function.
[0106] In the embodiment of the present invention, Figure 1 in the binaural renderer, there can be a set of preset HRTF functions and binaural rendering methods. Determine the preset HRTF function according to the microphone array configuration on the main device, and perform binaural rendering on the main audio signal using the HRTF to obtain an environmental binaural audio signal.
[0107] 304. Determine whether the target object is within the shooting field of view of the main device.
[0108] If it is detected that the target object is within the shooting field of view of the main device, then execute the following 305 to 312, and 320 to 323; if it is detected that the target object is outside the shooting field of view of the main device, then execute the following 313 to 319, and 320 to 323.
[0109] Figure 1 The target tracking module in [[ ]] consists of a visual target tracker and an audio target tracker, and can be used to determine the position of the target object and estimate the azimuth angle and distance between the target object and the main device by using visual data and / or audio signals. When the target object is within the shooting field of view of the main device, at this time, visual data and audio signals can be used together to determine the position of the target object, and at this time, the visual target tracker and the audio target tracker are both enabled. When the target object is outside the shooting field of view of the main device, audio signals can be used to determine the position of the target object, and at this time, only the audio target tracker can be enabled.
[0110] Optionally, when the target object is within the shooting field of view of the main device, one of the visual data and the audio signal can also be used to determine the position of the target object.
[0111] 305. Determine the first azimuth angle between the target object and the main device according to the video information and shooting parameters obtained by the main device, obtain the first active time and the first distance of the second additional audio signal, and determine the second active time of the main audio signal according to the first active time and the first distance.
[0112] Among them, the first distance is the target distance between the target object captured last time and the master device.
[0113] 306. Use the main audio signal within the second active time to estimate the angle of arrival, obtain the second azimuth angle between the target object and the master device, and smooth the first azimuth angle and the second azimuth angle to obtain the target azimuth angle.
[0114] 307. Determine the second distance between the target object and the master device according to the video information obtained by the master device, and calculate the second time delay according to the second distance and the speed of sound.
[0115] 308. Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal, and determine the first time delay between the beamformed signal and the second additional audio signal.
[0116] Figure 1 Among them, the sound source direction finding and the beamformer can be used to perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal, and the delay estimator can further determine the first time delay between the beamformed signal and the second additional audio signal.
[0117] 309. Smooth the second time delay and the first time delay to obtain the target time delay, and calculate the target distance according to the target time delay and the speed of sound.
[0118] When the target object is within the shooting field of view of the master device, the video data obtained at this time includes the target object. At this time, according to the position of the target object captured in the video frame in the video frame, combined with prior information such as camera parameters (for example, focal length) and zoom scale (different shooting fields of view correspond to different zoom scales), the above-mentioned first azimuth angle can be obtained. The azimuth angle and distance between the target object and the master device can also be estimated by using the audio signal to obtain the above-mentioned second azimuth angle. The target azimuth angle is obtained by smoothing the above-mentioned first azimuth angle and the second azimuth angle.
[0119] Furthermore, by comparing the size of the target object captured in the video frame with the typical size of the target object recorded in advance, and combined with prior information such as camera parameters (for example, focal length) and zoom scale (different shooting fields of view correspond to different zoom scales), a rough distance estimate can be made to obtain the above-mentioned second distance. According to the second distance, the speed of sound, and the known system delay, the above-mentioned second time delay can be obtained, and the delay between the second additional audio signal and the main audio signal (i.e., the first time delay) can be calculated. By smoothing the first time delay and the second time delay, the target time delay can be obtained.
[0120] In the embodiments of the present invention, the smoothing process may refer to calculating an average value. For example, after smoothing the first azimuth angle and the second azimuth angle, the target azimuth angle can be obtained by calculating the average value of the first azimuth angle and the second azimuth angle as the target azimuth angle; for the smoothing process of the first time delay and the second time delay, the target time delay can be obtained, which can be calculating the average value of the first time delay and the second time delay as the target time delay.
[0121] When the target object to be photographed is within the shooting field of view of the master device, Figure 1 the visual target tracker in can use the captured video to detect the target azimuth angle and the target distance between the target object to be photographed and the master device. The advantage of using a visual target tracker is that in a noisy environment or when there are a large number of sound sources, its tracking result is more accurate than that of an audio target tracker.
[0122] Furthermore, by simultaneously using a visual target tracker and an audio target tracker to detect the target azimuth angle and the target distance between the target object to be photographed and the master device, the accuracy can be further improved.
[0123] 310. Align the second additional audio signal with the main audio signal in the time domain according to the target time delay to obtain the first additional audio signal.
[0124] 311. Perform spatial filtering on the main audio signal in the area outside the shooting field of view according to the shooting field of view of the master device to obtain a reverse focused audio signal.
[0125] 312. Use the reverse focused audio signal as a reference signal to perform adaptive filtering on the first additional audio signal to obtain the target audio signal.
[0126] 313. Obtain the first active time and the first distance of the second additional audio signal, and determine the second active time of the main audio signal according to the first active time and the first distance.
[0127] Wherein, the first distance is the target distance between the target object to be photographed and the master device determined last time.
[0128] In the embodiments of the present invention, the active time of an audio signal refers to the time period during which there is a valid audio signal in the audio signal. Optionally, the first active time of the second additional audio signal may refer to the time period during which there is a valid audio signal in the second additional audio signal.
[0129] Optionally, the valid audio signal may refer to human voice or musical instrument sound, etc. Exemplarily, it may be the sound of the target object to be photographed.
[0130] In an embodiment of the present invention, the time delay between the second additional audio signal and the main audio signal can be determined according to the first distance and the speed of sound. Then, according to the time delay and the first active time, the audio signal corresponding to the second active time in the second additional audio signal in the main audio signal can be determined.
[0131] 314. Use the main audio signal within the second active time for angle-of-arrival estimation to obtain the target azimuth angle between the target object and the main device.
[0132] 315. Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal, and determine the first time delay between the beamformed signal and the second additional audio signal.
[0133] 316. Calculate the target distance between the target object and the main device according to the first time delay and the speed of sound.
[0134] When the target object is outside the shooting field of view of the main device, the video data obtained at this time does not include the target object. At this time, audio signals can be used to determine the position of the target object.
[0135] Figure 1 In this case, the audio target tracker can use the main audio signal and the additional audio signal to estimate the target azimuth angle and the target distance between the target object and the main device. Specifically, it can include steps such as sound source direction finding, beamforming, and delay estimation.
[0136] Specifically, the target azimuth angle can be obtained by performing direction-of-arrival (DOA) estimation on the main audio signal. To avoid the influence of a noisy environment or multiple sound sources on the DOA estimation, before performing the DOA estimation, first, the second additional audio can be analyzed to obtain the time corresponding to the active part of the effective audio signal (which can refer to the audio signal with the sound of the target object) in the second additional audio, that is, the above-mentioned first active time. Then, according to the previously estimated target distance, the delay (i.e., the first time delay) between the second additional audio signal and the main audio signal is obtained, and the first active time is mapped to the second active time in the main audio signal. Then, a segment of the main audio signal is intercepted at the second active time, and DOA estimation is performed to obtain the azimuth angle between the target object and the main device, and this azimuth angle is used as the above-mentioned target azimuth angle.
[0137] Optionally, when performing DOA estimation, the generalized cross correlation (GCC) method with phase weighting (PHAT) can be first used to perform time delay of arrival (TDOA) estimation, and then combined with the formation information of the microphone array to obtain the DOA. After obtaining the DOA estimation, the multi-channel main audio signal will pass through a beamformer in a fixed direction to obtain a beamformed signal, which is directionally enhanced towards the direction of the above target direction angle to improve the accuracy of the subsequent delay estimation. The beamforming method can be delay-sum or minimum variance distortion response (MVDR). The estimation of the above first delay is also performed using the TDOA method between the main audio beamformed signal and the second additional audio signal. Similarly, the TDOA estimation is also only performed during the active time of the second additional audio signal. Based on the first delay, the speed of sound, and the known system delay, the distance between the target object and the main device can be obtained, that is, the above target distance.
[0138] 317. Align the second additional audio signal with the main audio signal in the time domain according to the first time delay to obtain the first additional audio signal.
[0139] When the target object is outside the shooting field of view of the main device, use the first time delay as the target time delay between the main audio signal and the second additional audio signal, and align the second additional audio signal with the main audio signal in the time domain according to the first time delay to obtain the first additional audio signal.
[0140] Figure 1 The delay compensator in [description] can align the second additional audio signal with the main audio signal in the time domain according to the above first delay to obtain the first additional audio signal.
[0141] 318. Perform spatial filtering on the main audio signal in the area within the shooting field of view of the main device to obtain a focused audio signal.
[0142] 319. Use the first additional audio signal as a reference signal to perform adaptive filtering on the focused audio signal to obtain the target audio signal.
[0143] When the target object is within the shooting field of view, since a high-quality close-up audio signal has been obtained through the additional audio signal, the main purpose of spatial filtering is to obtain a purer ambient audio signal. Therefore, the target area of spatial filtering is outside the shooting field of view, and the obtained signal is hereinafter referred to as the reverse-focus audio signal; when the target object is outside the shooting field of view, since spatial filtering is required to obtain the close-up audio signal within the shooting field of view, the target area of spatial filtering is the shooting field of view, and the obtained signal is hereinafter referred to as the focus audio signal.
[0144] Furthermore, when performing spatial filtering, it combines with the shooting field of view of the main device and can follow the change of the shooting field of view of the main device, so as to perform directional enhancement on the local audio signal.
[0145] Figure 1 Among them, two groups of adaptive filters act on the focus audio signal and the additional audio signal respectively. According to the change of the target in the shooting field of view, only one of the groups of adaptive filters is enabled. When the target is in the shooting field of view, the adaptive filter acting on the additional audio signal is activated, and the reverse-focus audio signal is input as a reference signal to further suppress the ambient sound from the additional audio signal, making the sound near the target object more prominent. When the target is outside the shooting field of view, the adaptive filter acting on the focus audio signal is activated, and the additional audio signal is input as a reference signal to further suppress the sound outside the shooting field of view from the focus audio signal. The method of adaptive filtering can be the least mean square error (LMS, Least Mean Square), etc.
[0146] 320. Generate a second multi-channel transfer function according to the target distance and the target azimuth angle.
[0147] 321. Perform multi-channel rendering on the target audio signal according to the second multi-channel transfer function to obtain a target multi-channel audio signal.
[0148] 322. Determine the first gain of the ambient multi-channel audio signal and the second gain of the target multi-channel audio signal according to the shooting parameters of the main device.
[0149] 323. Mix the ambient multi-channel audio signal and the target multi-channel audio signal according to the first gain and the second gain to obtain a mixed multi-channel audio signal.
[0150] Figure 1In this case, a hybrid gain controller can determine the hybrid gain according to the user's shooting field of view, that is, the proportion of the two groups of signals in the mixed signal. For example, when increasing the zoom level of the camera, that is, reducing the shooting field of view, the gain of the environmental binaural audio signal will decrease, while the gain of the additional binaural audio signal (i.e., the determined target multi-channel audio signal when the target object is within the field of view) or the focused binaural audio signal (i.e., the determined target multi-channel audio signal when the target object is outside the field of view) will increase. In this way, while the shooting field of view of the video is focused on a specified area, the audio will also be focused on the specified area.
[0151] In the embodiments of the present invention, according to the shooting parameters of the main device (such as the zoom level of the camera), the size of the shooting field of view is determined, and based on this, the first gain of the environmental multi-channel audio signal and the second gain of the target multi-channel audio signal are determined, so that while the shooting field of view of the video is focused on a specified area, the audio will also be focused on the specified area, thereby creating an effect of "immersive, sound following the image movement".
[0152] The multi-channel audio signal acquisition method provided by the embodiments of the present invention is a distributed recording and audio focusing method that can create a more realistic sense of presence. This method can simultaneously use the microphone array on the main device in the terminal device and the microphones on the additional device (TWS Bluetooth headset) for distributed audio acquisition and fusion. The microphone array in the terminal device acquires the spatial audio at the location of the main device (i.e., the main audio signal involved in the embodiments of the present invention), and the TWS Bluetooth headset can be set on the target object to be tracked and, as the target object moves, acquire high-quality close-up audio signals in the distance (i.e., the first additional audio signal involved in the embodiments of the present invention). Combining with the FOV change during the video shooting process, corresponding adaptive filtering processing is performed on the two groups of acquired signals to achieve environmental sound suppression, and spatial filtering processing of the spatial audio signal in the specified area is performed to achieve directional enhancement. Then, combining the two positioning methods of vision and sound, the target of interest is tracked and positioned, and the HRTF binaural rendering and upmixing or downmixing are respectively performed on the three groups of signals of the obtained spatial audio, high-quality close-up audio, and directionally enhanced audio to obtain three groups of binaural signals: environmental binaural signals, additional binaural signals, and focused binaural signals. Finally, the mixing ratio of the above three groups of binaural signals is determined according to the size of the FOV and mixed.
[0153] Such a technical solution can produce the following beneficial effects:
[0154] When the finally output binaural audio signal is played on stereo headphones, it can simultaneously simulate the spatial sound field and the point-like auditory target at the specified position.
[0155] By using distributed audio signals, better directional enhancement effects can be obtained, and the suppression of interfering sounds and ambient sounds is more obvious during focusing.
[0156] It can follow the change of FOV, better focus on and track the sounds that users are interested in, so as to create an immersive experience of "immersive, sound following the image".
[0157] Such as Figure 4 As shown, an embodiment of the present invention provides a multi-channel audio signal acquisition device 400, which includes:
[0158] An acquisition module 401, configured to acquire the main audio signal collected when the main device performs video shooting on the target shooting object, and perform first multi-channel rendering to obtain an ambient multi-channel audio signal; acquire the audio signal collected by the additional device, and determine the first additional audio signal; wherein, the distance between the additional device and the target shooting object is less than the first threshold;
[0159] A processing module 402, configured to perform ambient sound suppression processing on the first additional audio signal and the main audio signal to obtain a target audio signal;
[0160] Perform second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal;
[0161] Mix the ambient multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal.
[0162] Optionally, the processing module 402 is specifically configured to determine the first gain of the ambient multi-channel audio signal and the second gain of the target multi-channel audio signal according to the shooting parameters of the main device;
[0163] Mix the ambient multi-channel audio signal and the target multi-channel audio signal according to the first gain and the second gain to obtain a mixed multi-channel audio signal.
[0164] Optionally, the acquisition module 401 is specifically configured to acquire the main audio signal collected by the microphone array on the main device;
[0165] Generate a first multi-channel transfer function according to the microphone array formation on the main device,
[0166] Perform multi-channel rendering on the main audio signal according to the first multi-channel transfer function to obtain an ambient multi-channel audio signal.
[0167] Optionally, the acquisition module 401 is specifically configured to acquire the second additional audio signal collected by the additional device on the target shooting object, and determine the second additional audio signal as the first additional audio signal;
[0168] Or,
[0169] Obtain a second additional audio signal collected by an additional device, align the second additional audio signal with the main audio signal in the time domain to obtain a first additional audio signal.
[0170] Optionally, the processing module 402 is specifically configured to obtain a target azimuth angle between the target object and the main device;
[0171] Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal;
[0172] Determine a target time delay between the main audio signal and the second additional audio signal;
[0173] According to the first time delay, align the second additional audio signal with the main audio signal in the time domain to obtain a first additional audio signal.
[0174] Optionally, the processing module 402 is specifically configured to obtain a target distance and a target azimuth angle between the target object and the main device;
[0175] Generate a second multi-channel transfer function according to the target distance and the target azimuth angle;
[0176] Perform multi-channel rendering on the target audio signal according to the second multi-channel transfer function to obtain a target multi-channel audio signal.
[0177] Optionally, the acquisition module 401 is specifically configured to, when it is detected that the target object is outside the shooting field of view of the main device, obtain a first active time and a first distance of the second additional audio signal, and the first distance is the target distance between the target object and the main device determined last time;
[0178] Determine a second active time of the main audio signal according to the first active time and the first distance;
[0179] Use the main audio signal within the second active time to perform angle-of-arrival estimation to obtain a target azimuth angle between the target object and the main device.
[0180] Optionally, the acquisition module 401 is specifically configured to, when it is detected that the target object is outside the shooting field of view of the main device, perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal;
[0181] Determine a first time delay between the beamformed signal and the second additional audio signal;
[0182] Calculate the target distance between the target object and the main device according to the first time delay and the speed of sound.
[0183] Optionally, the processing module 402 is specifically configured to, when detecting that the target object is outside the shooting field of view of the master device, perform spatial filtering on the main audio signal in the area within the shooting field of view according to the shooting field of view of the master device to obtain a focused audio signal;
[0184] Use the first additional audio signal as a reference signal to perform adaptive filtering processing on the focused audio signal to obtain a target audio signal.
[0185] Optionally, the acquisition module 401 is specifically configured to, when detecting that the target object is within the shooting field of view of the master device, determine a first azimuth angle between the target object and the master device according to the video information and shooting parameters obtained by the master device;
[0186] Obtain a first active time and a first distance of the second additional audio signal, where the first distance is the target distance between the target object and the master device determined last time;
[0187] Determine a second active time of the main audio signal according to the first active time and the first distance;
[0188] Use the main audio signal within the second active time to perform arrival angle estimation to obtain a second azimuth angle between the target object and the master device;
[0189] Perform smoothing processing on the first azimuth angle and the second azimuth angle to obtain a target azimuth angle.
[0190] Optionally, the acquisition module 401 is specifically configured to, when detecting that the target object is within the shooting field of view of the master device, determine a second distance between the target object and the master device according to the video information obtained by the master device;
[0191] Calculate a second time delay according to the second distance and the speed of sound;
[0192] Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal;
[0193] Determine a first time delay between the beamformed signal and the second additional audio signal;
[0194] Perform smoothing processing on the second time delay and the first time delay to obtain a target time delay;
[0195] Calculate the target distance according to the target time delay and the speed of sound.
[0196] Optionally, the processing module 402 is configured to, when detecting that the target object is within the shooting field of view of the master device, perform spatial filtering on the main audio signal in the area outside the shooting field of view according to the shooting field of view of the master device to obtain a reverse focused audio signal;
[0197] Using the reverse-focused audio signal as a reference signal, perform adaptive filtering on the first additional audio signal to obtain a target audio signal.
[0198] Optionally, the processing module 402 is specifically configured to obtain the video data captured by the master device and the second additional audio signal collected by the additional device;
[0199] Determine the current scene category and the target shooting object category;
[0200] Perform multi-channel rendering on the target audio signal through a first rendering rule that matches the current scene category and the target shooting object category to obtain a target multi-channel audio signal.
[0201] Optionally, the processing module 402 is specifically configured to obtain the main audio signal collected by the master device when shooting a video of the target shooting object;
[0202] Determine the current scene category;
[0203] Perform first multi-channel rendering on the main audio signal through a second rendering rule that matches the current scene category to obtain the environmental multi-channel audio signal.
[0204] An embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the multi-channel audio signal acquisition method provided in the above method embodiment.
[0205] As Figure 5 shown, an embodiment of the present invention further provides a terminal device, which includes the above multi-channel audio signal acquisition device 400 and a master device 500.
[0206] Wherein, the master device is configured to collect a main audio signal when shooting a video and send the main audio signal to the multi-channel audio signal acquisition device.
[0207] As Figure 6 shown, an embodiment of the present invention further provides a terminal device, which includes but is not limited to: radio frequency (RF) circuit 601, memory 602, input unit 603, display unit 604, sensor 605, audio circuit 606, wireless fidelity (WiFi) module 607, processor 608, Bluetooth module 609, and camera 610, etc. Among them, the RF circuit 601 includes a receiver 6011 and a transmitter 6012. Those skilled in the art can understand, Figure 6The terminal device structure shown does not limit the terminal device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0208] The RF circuit 601 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 608 for processing; in addition, it transmits the designed uplink data to the base station. Generally, the RF circuit 601 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 601 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0209] The memory 602 can be used to store software programs and modules. The processor 608 executes various functional applications and data processing of the terminal device by running the software programs and modules stored in the memory 602. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, applications required for at least one function (such as the voice playback function, the image playback function, etc.); the data storage area can store data created according to the use of the terminal device (such as audio signals, phone books, etc.). In addition, the memory 602 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0210] The input unit 603 can be used to receive input numerical or character information, and generate key signal inputs related to the user settings and function control of the terminal device. Specifically, the input unit 603 can include a touch panel 6031 and other input devices 6032. The touch panel 6031, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 6031), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 6031 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 608, and can receive and execute the commands sent by the processor 608. In addition, various implementations such as resistive, capacitive, infrared, and surface acoustic wave can be adopted for the touch panel 6031. In addition to the touch panel 6031, the input unit 603 can also include other input devices 6032. Specifically, the other input devices 6032 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0211] The display unit 604 can be used to display the information input by the user or the information provided to the user, as well as various menus of the terminal device. The display unit 604 can include a display panel 6041. Optionally, the display panel 6041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 6031 can cover the display panel 6041. After the touch panel 6031 detects a touch operation thereon or nearby, it is transmitted to the processor 608 to determine the touch event. Subsequently, the processor 608 provides corresponding visual output on the display panel 6041 according to the touch event. Although in Figure 6 the touch panel 6031 and the display panel 6041 are implemented as two independent components to realize the input and input functions of the terminal device, in some embodiments, the touch panel 6031 and the display panel 6041 can be integrated to realize the input and output functions of the terminal device.
[0212] The terminal device may further include at least one sensor 605, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 6041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 6041 and / or the backlight when the terminal device is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the terminal device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors that the terminal device can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here. In the embodiments of the present invention, the terminal device may include an acceleration sensor, a depth sensor, or a distance sensor, etc.
[0213] The audio circuit 606, the speaker 6061, and the microphone 6062 can provide an audio interface between the user and the terminal device. The audio circuit 606 can transmit the electrical signal converted from the received audio signal to the speaker 6061, and the speaker 6061 converts it into a sound signal for output. On the other hand, the microphone 6062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 606 and then converted into an audio signal. After the audio signal is output and processed by the processor 608, it is sent through the RF circuit 601 to, for example, another terminal device, or the audio signal is output to the memory 602 for further processing. Among them, the above-mentioned microphone 6062 can be a microphone array.
[0214] WiFi belongs to short-distance wireless transmission technology. The terminal device can help users send and receive emails, browse the web, and access streaming media through the WiFi module 607. It provides users with wireless broadband Internet access. Although Figure 6 the WiFi module 607 is shown, it can be understood that it does not belong to an essential component of the terminal device and can be completely omitted within the scope of not changing the essence of the invention according to needs.
[0215] The processor 608 is the control center of the terminal device, connecting various parts of the entire terminal device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 602, and by invoking the data stored in the memory 602, it executes various functions of the terminal device and processes data, thereby monitoring the terminal device as a whole. Optionally, the processor 608 may include one or more processing units; preferably, the processor 608 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 608 either.
[0216] The terminal device further includes a Bluetooth module 609, which is used for short-distance wireless communication and is divided into a Bluetooth data module and a Bluetooth voice module according to functions. The Bluetooth module refers to the basic circuit set of a chip integrating Bluetooth functions and is used for wireless network communication. It can be roughly divided into three major types: data transmission module, Bluetooth audio module, Bluetooth audio + data dual-in-one module, etc.
[0217] Although not shown, the terminal device may further include other functional modules, which will not be elaborated here.
[0218] In the embodiment of the present invention, the microphone 6062 can be used to collect the main audio signal. The terminal device can be connected to an additional device through the above-mentioned WiFi module 607 or Bluetooth module 609, and receive the second additional audio signal collected by the additional device.
[0219] The processor 608 is used to obtain the main audio signal, perform multi-channel rendering to obtain an ambient multi-channel audio signal; obtain the audio signal collected by the additional device and determine the first additional audio signal; perform ambient noise suppression processing on the basis of the first additional audio signal and the main audio signal to obtain a target audio signal; perform multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal; mix the ambient multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal. Wherein, the distance between the additional device and the target shooting object is less than a first threshold;
[0220] Optionally, the above-mentioned processor 608 can also be used to implement other processes implemented by the terminal device in the above method embodiments, which will not be elaborated here.
[0221] The embodiment of the present invention further provides a multi-channel audio signal acquisition system, which includes: a multi-channel audio signal acquisition device, a main device, and an additional device. The main device and the additional device are respectively communicatively connected to the multi-channel audio signal;
[0222] The master device is configured to collect a master audio signal during video shooting of a target object and send the master audio signal to the multi-channel audio signal acquisition device;
[0223] The additional device is configured to collect a second additional audio signal and send the second additional audio signal to the multi-channel audio signal acquisition device.
[0224] Exemplarily, the multi-channel audio signal acquisition system may be as shown in the above Figure 1 wherein Figure 1 the audio processing device in the above may be the multi-channel audio signal acquisition device.
[0225] An embodiment of the present invention further provides a computer-readable storage medium, including: a computer program stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the multi-channel audio signal acquisition method in the above method embodiment.
[0226] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all should fall within the scope of protection of the present invention.
[0227] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0228] In several embodiments provided by the present invention, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in electrical, mechanical, or other forms.
[0229] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0230] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0231] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0232] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for obtaining multi-channel audio signals, characterized in that, Including: Obtaining a main audio signal collected when a main device performs video shooting on a target shooting object, and performing first multi-channel rendering to obtain an environmental multi-channel audio signal; Obtaining an audio signal collected by an additional device, and determining a first additional audio signal, where the distance between the additional device and the target shooting object is less than a first threshold; Performing environmental noise suppression processing on the first additional audio signal and the main audio signal to suppress environmental noise and obtain a target audio signal; Performing second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal; and mixing the environmental multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal; When it is detected that the target shooting object is within the shooting field of view of the main device, the performing environmental noise suppression processing on the first additional audio signal and the main audio signal to suppress environmental noise and obtain a target audio signal includes: Performing spatial filtering on the main audio signal in an area outside the shooting field of view according to the shooting field of view of the main device to obtain a reverse focused audio signal; using the reverse focused audio signal as a reference signal, and performing adaptive filtering processing on the first additional audio signal to obtain the target audio signal; or When it is detected that the target shooting object is outside the shooting field of view of the main device, the performing environmental noise suppression processing on the first additional audio signal and the main audio signal to suppress environmental noise and obtain a target audio signal includes: Performing spatial filtering on the main audio signal in an area within the shooting field of view according to the shooting field of view of the main device to obtain a focused audio signal; using the first additional audio signal as a reference signal, and performing adaptive filtering processing on the focused audio signal to obtain the target audio signal.
2. The method according to claim 1, wherein The mixing the environmental multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal includes: Determining a first gain of the environmental multi-channel audio signal and a second gain of the target multi-channel audio signal according to the shooting parameters of the main device; Mixing the environmental multi-channel audio signal and the target multi-channel audio signal according to the first gain and the second gain to obtain a mixed multi-channel audio signal.
3. The method according to claim 1, wherein The obtaining a main audio signal collected when a main device performs video shooting on a target shooting object, and performing first multi-channel rendering to obtain an environmental multi-channel audio signal includes: Obtaining the main audio signal collected by a microphone array on the main device; Generating a first multi-channel transfer function according to the microphone array configuration on the main device; Performing first multi-channel rendering on the main audio signal according to the first multi-channel transfer function to obtain the environmental multi-channel audio signal.
4. The method according to claim 1, wherein The obtaining an audio signal collected by an additional device, and determining a first additional audio signal includes: Obtaining a second additional audio signal collected by the additional device, and determining the second additional audio signal as the first additional audio signal; Or Obtain the second additional audio signal collected by the additional device, and align the second additional audio signal with the main audio signal in the time domain to obtain the first additional audio signal.
5. The method according to claim 4, characterized in that The step of aligning the second additional audio signal with the main audio signal in the time domain to obtain the first additional audio signal includes: Obtain the target azimuth angle between the target object and the main device; Determine the target time delay between the main audio signal and the second additional audio signal; According to the target time delay, align the second additional audio signal with the main audio signal in the time domain to obtain the first additional audio signal.
6. The method according to claim 1, wherein The step of performing second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal includes: Obtain the target distance and the target azimuth angle between the target object and the main device; Generate a second multi-channel transfer function according to the target distance and the target azimuth angle; Perform second multi-channel rendering on the target audio signal according to the second multi-channel transfer function to obtain a target multi-channel audio signal.
7. The method according to claim 6, wherein When it is detected that the target object is within the shooting field of view of the main device, the step of obtaining the target azimuth angle between the target object and the main device includes: Determine the first azimuth angle between the target object and the main device according to the video information and shooting parameters obtained by the main device; Obtain the first active time and the first distance of the second additional audio signal, where the first distance is the target distance between the target object and the main device determined last time, and the second additional audio signal is the signal collected by the additional device; Determine the second active time of the main audio signal according to the first active time and the first distance; Perform arrival angle estimation using the main audio signal within the second active time to obtain the second azimuth angle between the target object and the main device; Perform smoothing processing on the first azimuth angle and the second azimuth angle to obtain the target azimuth angle.
8. The method according to claim 7, wherein The step of obtaining the target distance between the target object and the main device includes: Determine the second distance between the target object and the main device according to the video information obtained by the main device; Calculate the second time delay according to the second distance and the speed of sound; Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal; Determine the first time delay between the beamformed signal and the second additional audio signal; Perform smoothing processing on the second time delay and the first time delay to obtain the target time delay; Calculate the target distance according to the target time delay and the speed of sound.
9. The method according to claim 6, wherein When it is detected that the target object is outside the shooting field of view of the main device, the step of obtaining the target azimuth angle between the target object and the main device includes: Obtain the first active time and the first distance of the second additional audio signal, where the first distance is the target distance between the target object and the main device determined last time, and the second additional audio signal is the signal collected by the additional device; Determine the second active time of the main audio signal according to the first active time and the first distance; Use the main audio signal within the second active time to perform angle-of-arrival estimation to obtain the target azimuth angle between the target object and the main device.
10. The method according to claim 6, wherein When it is detected that the target object is outside the shooting field of view of the main device, the obtaining of the target distance between the target object and the main device includes: Perform beamforming processing on the main audio signal towards the target azimuth angle to obtain a beamformed signal; Determine the first time delay between the beamformed signal and the second additional audio signal, where the second additional audio signal is the signal collected by the additional device; Calculate the target distance between the target object and the main device according to the first time delay and the speed of sound.
11. The method according to claim 1, characterized in that, The performing second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal includes: Obtain the video data captured by the main device and the second additional audio signal collected by the additional device; Determine the current scene category and the target object category; Perform second multi-channel rendering on the target audio signal through a first rendering rule matching the current scene category and the target object category to obtain the target multi-channel audio signal.
12. The method according to claim 1, wherein Obtain the main audio signal collected when the main device performs video shooting on the target object, and perform first multi-channel rendering to obtain an environmental multi-channel audio signal, including: Obtain the main audio signal collected when the main device shoots a video of the target object; Determine the current scene category; Perform first multi-channel rendering on the main audio signal through a second rendering rule matching the current scene category to obtain the environmental multi-channel audio signal.
13. A multi-channel audio signal acquisition device, characterized in that, Including: An acquisition module, configured to obtain the main audio signal collected when the main device performs video shooting on the target object, and perform first multi-channel rendering to obtain an environmental multi-channel audio signal; Obtain the audio signal collected by the additional device and determine the first additional audio signal, where the distance between the additional device and the target object is less than a first threshold; A processing module, configured to perform environmental noise suppression processing through the first additional audio signal and the main audio signal to suppress environmental noise and obtain a target audio signal; Perform second multi-channel rendering on the target audio signal to obtain a target multi-channel audio signal; Mix the environmental multi-channel audio signal and the target multi-channel audio signal to obtain a mixed multi-channel audio signal; The processing module is specifically configured to, when detecting that the target shooting object is within the shooting field of view of the master device, perform spatial filtering on the main audio signal in the area outside the shooting field of view according to the shooting field of view of the master device to obtain a reverse focused audio signal; use the reverse focused audio signal as a reference signal to perform adaptive filtering processing on the first additional audio signal to obtain the target audio signal; when detecting that the target shooting object is outside the shooting field of view of the master device, perform spatial filtering on the main audio signal in the area within the shooting field of view according to the shooting field of view of the master device to obtain a focused audio signal; use the first additional audio signal as a reference signal to perform adaptive filtering processing on the focused audio signal to obtain the target audio signal.
14. A terminal device, characterized in that, Comprising: A processor, a memory, and a computer program stored on the memory and executable on the processor, where the computer program, when executed by the processor, implements the multi-channel audio signal acquisition method according to any one of claims 1 to 12.
15. A terminal device, characterized in that, Comprising: The multi-channel audio signal acquisition device and the master device according to claim 13. The master device is configured to collect a main audio signal during video shooting of a target shooting object and send the main audio signal to the multi-channel audio signal acquisition device.
16. A multi-channel audio signal acquisition system, characterized in that The system comprises: the multi-channel audio signal acquisition device, the master device, and an additional device according to claim 13, where the master device and the additional device are respectively communicatively connected to the multi-channel audio signal. The master device is configured to collect a main audio signal during video shooting of a target shooting object and send the main audio signal to the multi-channel audio signal acquisition device. The additional device is configured to collect a second additional audio signal and send the second additional audio signal to the multi-channel audio signal acquisition device. Wherein, the distance between the additional device and the target shooting object is less than a first threshold.
17. A computer-readable storage medium, characterized in that, Comprising: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the multi-channel audio signal acquisition method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Distributed audio capture and mixing
CN108370471A
Remote sound collection device, monitoring device and remote sound collection method
CN108389586A
Intelligent audio rendering for video recording
US20190222950A1