Dynamic pickup method and system based on microphone array
Through the dynamic sound pickup method of the microphone array, the audio signal characteristics are obtained, the target audio signal and sound source position are determined, the sound pickup beam is formed within the sound pickup area, and the alternative beam is controlled to switch to the sound pickup beam to achieve accurate sound pickup of the target sound source, solving the problems of sound pickup delay and sound quality deterioration in complex environments, and improving the accuracy and stability of sound pickup.
Patent Information
- Application Number
- CN202510880360.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-12
AI Technical Summary
The existing technology has reduced beam switching speed and repositioning accuracy in complex environments, resulting in sound pickup delay or poor sound quality.
Through a dynamic sound pickup method based on a microphone array, audio signal characteristics are obtained, the target audio signal and sound source position are determined, and a sound pickup beam within the pickup area is formed. By setting the pickup area within the pickup area, any beam is controlled to switch to the target sound source and switch to a gathering beam, pointing to the target audio signal to pick up the target audio signal within the pickup area, and when no voice signal is detected within a preset time, the pickup beam is switched to a standby state. Technical measures are taken to achieve accurate pickup of the target sound source.
It improves the focus of sound pickup and the accuracy and stability of sound source tracking, reduces the response delay to an extremely low level, ensures rapid response of sound pickup, significantly improves the accuracy and stability of the sound source, solves the problems of beam switching speed and repositioning accuracy in the existing technology, significantly improves the accuracy and stability of the sound source, solves the problems of beam switching speed and positioning accuracy in the existing technology, solves the problems of sound source positioning accuracy in the existing technology, improves the accuracy and stability of the sound source, solves the problems of beam switching positioning accuracy in the existing technology, and improves the accuracy and stability of sound source positioning.
Smart Images

Figure CN120640175A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of sound pickup technology, and in particular to a dynamic sound pickup method and system based on a microphone array. Background Art
[0002] In the field of microphone array sound pickup technology, existing technologies primarily achieve directional sound pickup of a target sound source by deploying multiple pickup beams and switching and repositioning them based on the sound source's location data. Specifically, existing technologies detect sound activity in the environment, first selecting one of the multiple pickup beams for switching and pickup, and then determining whether the metric of the selected pickup beam is greater than or equal to a metric related to the sound activity. If so, the selected pickup beam is repositioned based on the sound source's location data to track and pick up the target sound source.
[0003] However, when the environment is complex or the detection of sound source location information is not sensitive enough, the speed of beam switching and the accuracy of repositioning will decrease significantly, resulting in sound pickup delay or poor sound quality. Summary of the Invention
[0004] To this end, the embodiments of the present application provide a dynamic sound pickup method and system based on a microphone array, which can dynamically switch the state of the sound pickup beam and improve the accuracy of sound pickup.
[0005] In a first aspect, the present application provides a dynamic sound pickup method based on a microphone array.
[0006] This application is achieved through the following technical solutions:
[0007] A dynamic sound pickup method based on a microphone array, comprising:
[0008] receiving at least one original audio signal from a conference environment, and obtaining audio features of each of the original audio signals;
[0009] Determining a target audio signal containing a speech signal from each of the original audio signals based on the audio features of the original audio signals;
[0010] Determining a target sound source position to be picked up based on the target audio signal;
[0011] According to at least one pickup area set in the conference environment, forming a plurality of pickup beams within each pickup area;
[0012] When the target sound source position is within the sound pickup area, any one of the alternative beams is controlled to switch to a sound pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a sound pickup beam that is switched to a standby state when no voice signal is detected within a preset time length.
[0013] In a preferred example of this application, it can be further set as follows:
[0014] Switch from the pickup beam to an alternative beam, including:
[0015] Acquiring beam pickup signals collected by a plurality of pickup beams;
[0016] Determine the time-frequency point of each beam pickup signal and the signal energy corresponding to each time-frequency point;
[0017] Comparing the signal energies of the various pickup beams at the time-frequency points, determining the pickup beam with the maximum signal energy corresponding to each time-frequency point, and marking the time-frequency point as the effective sound source time-frequency point of the pickup beam with the maximum signal energy;
[0018] The number of effective sound source time-frequency points corresponding to each sound pickup beam is counted, and when the number of effective sound source time-frequency points of the sound pickup beam within a preset time length is less than a preset threshold, the sound pickup beam is switched to an alternative beam.
[0019] In a preferred example of the present application, it can be further configured to include:
[0020] When the difference in signal energy between the two pickup beams at a time-frequency point is less than a preset difference, the speech similarity of the beam pickup signals of the two pickup beams is determined. When the speech similarity of the beam pickup signals is greater than a preset similarity threshold, the time-frequency point is marked as a valid sound source time-frequency point of one of the pickup beams.
[0021] In a preferred example of this application, it can be further set as follows:
[0022] The audio features include signal energy and signal frequency of the original audio signals; and determining a target audio signal containing a speech signal from each of the original audio signals based on the audio features of the original audio signals includes:
[0023] The signal energy and signal frequency of each of the original audio signals are input into a well-trained speech detection model; if the speech detection model detects that a speech signal exists in the original audio signal, the original audio signal is the target audio signal.
[0024] In a preferred example of this application, it can be further set as follows:
[0025] Determining a target sound source position to be picked up based on the target audio signal includes:
[0026] Obtaining a covariance matrix of the target audio signal in the frequency domain;
[0027] Performing eigendecomposition on the covariance matrix to obtain a signal subspace and a noise subspace;
[0028] The correlation between the sound source direction and the noise subspace in different directions is scanned, and the sound source direction with the smallest correlation is determined as the target sound source position to be picked up.
[0029] In a preferred example of the present application, it can be further configured to obtain at least one shielded area set in the conference environment;
[0030] When the target sound source position is within the shielding area, the target sound source position is shielded.
[0031] In a preferred example of the present application, it can be further set that the beam width of the alternative beam is set to 10° to 50°.
[0032] In a preferred example of the present application, it can be further configured that the pickup beam also includes a main beam, and the main beam is used to pick up other effective sound sources with voice signals in the pickup area. When the target sound source position is within the pickup area and the main beam picks up other effective sound sources with voice signals in the pickup area, any one of the alternative beams is controlled to switch to the pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a pickup beam that is switched to a standby state when no voice signal is detected within a preset time period.
[0033] In a preferred example of the present application, it can be further configured to include:
[0034] The interference energy generated in each sound pickup beam by an external interference sound source of each sound pickup beam is calculated and filtered out, so that each sound pickup beam outputs a speech signal of a sound source at the center of the beam.
[0035] In a second aspect, the present application provides a dynamic sound pickup system based on a microphone array, wherein the system includes a microphone array, and the microphone array includes a plurality of microphone elements and a processor connected to the microphone elements.
[0036] This application is achieved through the following technical solutions:
[0037] A dynamic sound pickup system based on a microphone array, the system comprising a microphone array, the microphone array comprising a plurality of microphone elements and a processor connected to the microphone elements, the microphone array being configured to execute the dynamic sound pickup method based on the microphone array described in the first aspect above, comprising:
[0038] The microphone element is used to receive at least one original audio signal of the conference environment;
[0039] The processor is configured to:
[0040] Acquiring audio features of each of the original audio signals;
[0041] Determining, based on the audio features of the original audio signals, a target audio signal containing a valid speech signal from each of the original audio signals;
[0042] Determining a target sound source position to be picked up based on the target audio signal;
[0043] According to at least one sound pickup area set in the conference environment, controlling the plurality of microphone elements to form a plurality of sound pickup beams in each sound pickup area;
[0044] When the target sound source position is within the sound pickup area, any one of the alternative beams is controlled to switch to a sound pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a sound pickup beam that is switched to a standby state when no voice signal is detected within a preset time length.
[0045] In summary, compared with the prior art, the technical solutions provided by the embodiments of the present application have at least the following beneficial effects:
[0046] The present application obtains a set pickup area; receives at least one original audio signal from a conference environment and obtains audio features of each of the original audio signals; based on the audio features of the original audio signals, determines a target audio signal containing a voice signal from each of the original audio signals, thereby determining the target sound source position to be picked up; forms a plurality of pickup beams within each pickup area according to at least one pickup area set in the conference environment; when the target sound source position is within the pickup area, controls any one of the alternative beams to switch to a pickup beam, pointing to the target sound source position to pick up the target pickup signal, wherein the alternative beam refers to a pickup beam of a microphone array that switches to a standby state when no voice signal is detected within a preset time period. First, the pickup area of the conference environment is set so that the microphone array can accurately process the sound source in the area, effectively shielding the background noise outside the pickup area, and improving the focus of the pickup; when no valid voice is detected within the preset time, the relevant beam automatically switches to a low-power standby state. Once the target sound source activity is detected, the standby beam is immediately activated and accurately points to the direction of the sound source. The response delay is extremely low, ensuring a rapid response to the pickup and significantly improving the accuracy and stability of sound source tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic flow chart of a dynamic sound pickup method based on a microphone array provided in one embodiment of the present application;
[0048] Figure 2 A schematic structural diagram of a dynamic sound pickup device based on a microphone array provided in one embodiment of the present application;
[0049] Description of reference numerals:
[0050] Sound pickup area acquisition module 10 , beamforming module 20 , sound source activity detection module 30 , and sound pickup module 40 . DETAILED DESCRIPTION
[0051] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make non-creative modifications to the present embodiment as needed, but as long as they are within the scope of the claims of the present application, they are protected by the patent law.
[0052] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0053] In addition, the term "and / or" in this application is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application, unless otherwise specified, generally indicates that the related objects are in an "or" relationship.
[0054] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.
[0055] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0056] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:
[0057] (1) A microphone array is a system that consists of a number of microphone elements and processors connected to each microphone element. It samples and filters the spatial characteristics of the sound field. It collects sound signals synchronously through multiple microphone elements and uses the time difference, phase difference and intensity difference of the received signals of each microphone element for joint processing. Based on beamforming technology, it can selectively enhance the sound in a specific direction and suppress the noise in other directions.
[0058] (2) The pickup area refers to the sound collection range defined within a specific physical space. In a conference environment, the role of the pickup area is to accurately capture the voice signal in the target area while suppressing noise or interfering sound sources in non-target areas.
[0059] (3) A pickup beam refers to a working beam generated by a number of microphone elements that picks up audio within a pickup area. The pickup beams configured within each pickup area can pick up only valid sound sources within that pickup area, or can be configured to pick up valid sound sources within other pickup areas. A valid sound source refers to a sound source that contains a voice signal.
[0060] (4) A standby beam is a pickup beam generated by a plurality of microphone elements and switches to a standby state when no voice signal is detected within a preset time period. The standby state of a standby beam means that it does not pick up any valid sound source within the pickup area. The standby beam needs to be switched to the pickup state under the control of the processor before it can pick up audio within the pickup area.
[0061] The embodiments of the present application are described in further detail below with reference to the accompanying drawings.
[0062] Reference Figure 1 As shown, the first exemplary embodiment of the present application provides a dynamic sound pickup method based on a microphone array, the specific steps of which include:
[0063] S1: Receive at least one original audio signal in a conference environment, and obtain audio features of each original audio signal.
[0064] The microphone array uses the basic sound pickup capabilities of all microphone elements in the conference environment to collect the original audio signals. Each microphone element corresponds to a channel, and the audio characteristics of the original audio signals of each channel are extracted. All microphone elements in the microphone array are synchronously sampled using a unified clock source to avoid time base deviation between channels.
[0065] Audio features include at least one of time-frequency energy, frequency domain energy, signal frequency, and sound intensity vectors. These audio features can describe the characteristics of sound from different dimensions. Specifically, time-frequency energy represents the energy accumulation of the sound signal in the time dimension, reflecting the intensity change of the signal; frequency domain energy represents the energy distribution of the signal within a specific frequency range, and is analyzed by converting the time-frequency signal into the frequency domain through Fourier transform (FFT); signal frequency refers to the periodic vibration characteristics of the sound; and the sound intensity vector represents the propagation direction and intensity of the sound in space.
[0066] S2: Based on the audio features of the original audio signals, determine a target audio signal containing a speech signal from each original audio signal.
[0067] By performing correlation analysis on the audio features of each original audio signal and determining whether there is a speech signal in each original audio signal based on a preset correlation standard, the original audio signal with a speech signal is marked as the target audio signal, completing the first screening of the original audio signal and improving the accuracy of subsequent audio positioning.
[0068] S3: Determine the target sound source position to be picked up based on the target audio signal.
[0069] The target audio signal is further localized to determine the target sound source position, including azimuth and elevation angles.
[0070] S4: According to at least one sound pickup area set in the conference environment, a plurality of sound pickup beams are formed in each sound pickup area.
[0071] The user sets at least one pickup area in the conference environment where sound needs to be picked up through software or other control devices. The pickup area specifically includes a set spatial coordinate system and the spatial range of the pickup area in the spatial coordinate system. For example, a unified Cartesian coordinate system is established in the conference environment, and the origin and coordinate axis direction of the Cartesian coordinate system are defined; the conference environment is divided into at least one pickup area, and the geometric parameters of each pickup area in the Cartesian coordinate system are determined, including edge coordinates and center coordinates. The geometric parameters of all pickup areas are summarized to form pickup area information in JSON format and transmitted to the microphone array.
[0072] It should be noted that the geometric type of the sound pickup area can be set to a cube, a sphere, or a cylinder according to actual needs, and multiple sound pickup areas do not overlap.
[0073] The execution order of the above step S4 can be placed before S1, or can be executed simultaneously with step S1.
[0074] S5: When the target sound source is within the pickup area, any one of the alternative beams is controlled to switch to a pickup beam, pointing to the target sound source to pick up the target audio signal. The alternative beam refers to a pickup beam that switches to a standby state when no voice signal is detected within a preset time.
[0075] After determining the target sound source position, it is necessary to convert the target sound source position into coordinates in the Cartesian coordinate system through coordinate transformation, perform encirclement detection with the geometric boundary of the pickup area, determine whether the target sound source position is in the pickup area, complete the second screening of the original audio signal, improve the stability and accuracy of the microphone array pickup, and prevent the problem of lobe overlap and pointing to useless directions during the pickup process. When the target sound source position matches one of the pickup areas, control any one of the alternative beams to point to the target sound source position to pick up the target audio signal. If a target sound source position is detected in the pickup area, control any one of the alternative beams to point to the target sound source position to pick up the target audio signal; if two target sound source positions are detected in the pickup area, control any two alternative beams to point to the target sound source position to pick up the target audio signal.
[0076] It should be noted that a pickup beam that has not detected a voice signal within a preset time period is switched to a standby state and marked as a candidate beam. The candidate beam is in a low-power monitoring mode, not completely shut down. However, it will not be able to pick up voice signals within the pickup area and output valid voice signals until it is reselected by the processor and switched to a pickup beam.
[0077] In some embodiments, the microphone array forms a plurality of sound pickup beams, and switching from a sound pickup beam to a standby state as an alternative beam includes:
[0078] Acquiring beam pickup signals collected by a plurality of pickup beams;
[0079] Determine the time-frequency point of each beam pickup signal and the signal energy corresponding to each time-frequency point;
[0080] Compare the signal energy of each pickup beam at the time-frequency point, determine the pickup beam with the maximum signal energy corresponding to each time-frequency point, and mark the time-frequency point as the effective sound source time-frequency point of the pickup beam with the maximum signal energy;
[0081] The number of effective sound source time-frequency points corresponding to each sound pickup beam is counted. When the number of effective sound source time-frequency points of the sound pickup beam within a preset time length is less than a preset threshold, the sound pickup beam is switched to an alternative beam.
[0082] Specifically, beam pickup signals collected by several pickup beams are obtained, all beam pickup signals are expanded on the time scale and frequency scale, and each beam pickup signal is framed, with one frame as a time-frequency point, and the signal energy corresponding to each time-frequency point is calculated; for each time-frequency point, the signal energy of each pickup beam is compared, and the time-frequency point is marked as the effective sound source time-frequency point of the pickup beam with the highest signal energy; the number of effective sound source time-frequency points of each pickup beam is counted and compared with a preset threshold; if the number of effective sound source time-frequency points of a certain pickup beam is less than the preset threshold for a period of time, the pickup beam is switched to a standby state as an alternative beam.
[0083] For example, the microphone array forms eight beams: Beam A, Beam B, Beam C, Beam D, Beam E, Beam F, Beam G, and Beam H. Beams A through H are all active pickup beams. The eight pickup beams collect two seconds of audio. The beam pickup signals collected by all pickup beams are divided into 100 frames, each 20ms long. An FFT is performed on each frame to obtain a two-dimensional time-frequency point. Each frame is treated as a time-frequency point, and the signal energy of each time-frequency point is calculated. For example, for time-frequency point x1, compare the signal energies of all beams at time-frequency point x1, and find that the signal energy of beam A is the highest, then mark the time-frequency point x1 as the effective sound source time-frequency point of beam A; for time-frequency point x2, compare the signal energies of all beams at time-frequency point x2, and find that the signal energy of beam C is the highest, then mark the time-frequency point x2 as the effective sound source time-frequency point of beam C, and so on. Count the number of effective sound source time-frequency points of each beam, and the number of effective sound source time-frequency points of beam A is M1, the number of effective sound source time-frequency points of beam B is M2, the number of effective sound source time-frequency points of beam C is M3, the number of effective sound source time-frequency points of beam D is M4, the number of effective sound source time-frequency points of beam E is M5, the number of effective sound source time-frequency points of beam F is M6, the number of effective sound source time-frequency points of beam G is M7, and the number of effective sound source time-frequency points of beam H is M8. The number of effective sound source time-frequency points in each beam is compared with a preset threshold σ. Beams with a value less than the threshold σ are considered candidate beams. In this embodiment, the threshold σ is set to 10% of the total time-frequency points, or 10. It should be noted that the above threshold is merely an example and does not limit the threshold. The threshold can be adjusted based on actual needs.
[0084] In some embodiments, when comparing the signal energies of the various sound pickup beams at the time-frequency points and determining the sound pickup beam with the maximum signal energy corresponding to each time-frequency point, the method further includes:
[0085] When the difference in signal energy between the two pickup beams at a time-frequency point is less than a preset difference, the speech similarity of the beam pickup signals of the two pickup beams is determined. When the speech similarity of the beam pickup signals is greater than a preset similarity threshold, the time-frequency point is marked as a valid sound source time-frequency point of one of the pickup beams.
[0086] In actual implementation, for each pickup beam's signal energy, if the difference in signal energy between two pickup beams at a time-frequency point is less than a preset difference (that is, the difference between the maximum energy and the second-largest energy is less than a preset difference), speech similarity judgment is triggered. Feature vectors (such as cosine similarity) are used to determine whether the two pickup beams originate from the same sound source. Combined with the beam's spatial directivity, the pickup beam closer to the estimated sound source direction is selected as the valid sound source time-frequency point. When the pickup beams' signal energy is close and the speech is similar, the appropriate allocation to the pickup beam closer to the estimated sound source direction can reduce beam jitter caused by boundary effects.
[0087] In some embodiments, the audio features include signal energy and signal frequency of the original audio signals; and determining the target audio signal containing the speech signal from each original audio signal based on the audio features of the original audio signals includes:
[0088] The signal energy and signal frequency of each original audio signal are input into a well-trained speech detection model. If the speech detection model detects that one of the original audio signals contains a speech signal, the original audio signal is marked as the target audio signal.
[0089] The speech detection model includes a first model.
[0090] Inputting the signal energy and signal frequency of each original audio signal into the fully trained first model;
[0091] The first model is used to determine the probability log-likelihood ratio of the speech signal and the noise signal in the original audio signal;
[0092] Whether the original audio signal is a target audio signal containing a speech signal is determined based on the probability log-likelihood ratio and a preset threshold.
[0093] Exemplarily, the first model can be a pre-trained Webrtc-VAD model. A Gaussian mixture model is constructed using the first model to model the audio features of the original audio signal. The signal energy and signal frequency of the original audio signal are input into the Webrtc-VAD model to determine the probability log-likelihood ratio of the speech signal and the noise signal in the original audio signal. Based on the probability log-likelihood ratio and a preset threshold, it is determined whether the original audio signal is a target audio signal containing a speech signal.
[0094] It should be noted that Webrtc-VAD is trained using labeled speech and noise data from various scenarios. The labels are used to identify speech data. The speech data is framed and the spectrum is divided into multiple subbands, and the energy contribution of each subband is calculated. Gaussian mixture models are built for the speech and noise signals, respectively. The speech Gaussian mixture model (speech GMM) learns the distribution of speech features, while the noise Gaussian mixture model (noise GMM) learns the distribution of noise features. The likelihood difference between the input features in the speech GMM and the noise GMM is calculated, and a preset threshold is used to determine whether the signal is speech or noise.
[0095] In another implementation, the speech detection model includes a second model:
[0096] The signal energy and signal frequency of each original audio signal are input into the fully trained second model, the audio features of the input original audio signal are classified, and whether the original audio signal is a target audio signal containing a speech signal is determined based on the classification result.
[0097] Exemplarily, the second model is obtained by training based on the SileroVAD model.
[0098] In some embodiments, determining the target sound source position to be picked up based on the target audio signal includes:
[0099] Obtain the covariance matrix of the target audio signal within the frequency range;
[0100] Perform eigendecomposition on the covariance matrix to obtain the signal subspace and noise subspace;
[0101] Scan the correlation between the sound source direction and the noise subspace in different directions, and determine the sound source direction with the smallest correlation as the target sound source position to be picked up.
[0102] Specifically, the target audio signal x(t) is subjected to a short-time Fourier transform (STFT) to obtain a frequency domain signal X(f);
[0103] Select the target frequency band, process the frequency points, and calculate the estimated value of the covariance matrix for each frequency f:
[0104] where X H represents conjugate transpose;
[0105] Perform eigendecomposition on the covariance matrix R(f): R(f) = U∑U H ,
[0106] Among them, ∑=diag(λ1,λ2,…,λ m ) is the eigenvalue matrix, U=[u1,u2,…,u m] is the corresponding eigenvector matrix;
[0107] According to the change of eigenvalue, the eigenvalue is divided into signal subspace U s and noise subspace U n ;
[0108] For each sound source direction θ, the MUSIC algorithm is used to calculate the correlation P between different sound source directions and the noise subspace music (θ), and the sound source direction with the smallest correlation is determined as the target sound source direction to be picked up.
[0109] In some embodiments, the user can use software to access at least one shielded area in the conference environment set by other control devices and transmit the shielded area to the microphone array. When the target sound source is located within the shielded area, the target sound source is shielded, and no alternative beam is directed at the target sound source. In this way, when the active location of the sound source is not within the pickup area, it is determined to be an interfering sound source and is not picked up. This ensures that all valid sound sources in the pickup area are covered by the beam, while interfering sound sources in the shielded area are not covered by the beam.
[0110] In some embodiments, the beam width of the alternative beam is set to 10° to 50°. Preferably, the beam width of the alternative beam is set to 30°. Setting the beam width of the alternative beam within the above range can achieve better sound pickup quality while having a better shielding effect against other sound sources.
[0111] In some embodiments, after the candidate beam is directed to the target sound source position to pick up the target audio signal, the method further includes:
[0112] The interference energy generated by the external interference sound source of the pickup beam in the pickup beam is calculated and filtered out, so that each pickup beam outputs the voice signal of the sound source at the center of the beam. In actual implementation, the interference energy generated by the external interference sound source of each pickup beam in each pickup beam can be filtered out by a GSC sidelobe elimination filter and a multi-channel Wiener filter. In an embodiment of the present application, a multi-channel Wiener filter is used to calculate the pickup energy of the external interference sound source in the beam based on the original audio data of the array microphone, and then it is filtered out by the Wiener filter, so that after each beam is processed, only the voice information of the sound source at the center of the beam remains.
[0113] In some embodiments, the sound pickup beam further includes a main beam, which is used to pick up other effective sound sources with voice signals in the sound pickup area. The dynamic sound pickup method further includes:
[0114] When the target sound source is within the pickup area and the main beam is used to pick up other valid sound sources with voice signals within the pickup area, any alternative beam is controlled to switch to a pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a pickup beam that switches to a standby state when no voice signal is detected within a preset time.
[0115] In detail, after the user sets the pickup area, the microphone array will automatically generate several pickup beams in the pickup area. When all the pickup beams do not detect a voice signal within the preset time and switch to the backup beam, one of the pickup beams will be retained and not switched to the backup beam as the main beam. When there is no voice, the main beam will continue to pick up the ambient noise in the pickup area for filtering, so that there is always at least one working beam in the conference environment.
[0116] If a new original audio signal appears in the conference environment, and that original audio signal is a target audio signal containing a speech signal, the target sound source position of the target audio signal is calculated. If the target sound source position is within the pickup area, the main beam continues to pick up the ambient noise within the pickup area, and the processor controls any alternative beam to switch to the pickup beam, pointing to the target sound source position to pick up the target audio signal.
[0117] Furthermore, when any backup beam in a conference environment switches to a sound pickup beam and points to the target sound source to pick up the target audio signal, if the main beam does not detect a voice signal within a preset time, the main beam switches to a backup beam. The active sound pickup beam in a conference scene serves as the main beam for that scene, and a conference scene can contain multiple main beams.
[0118] Reference Figure 2 As shown, another exemplary embodiment of the present application provides a dynamic sound pickup device based on a microphone array, and the dynamic sound pickup device specifically includes:
[0119] A sound pickup area acquisition module 10 is used to acquire at least one sound pickup area set according to a conference environment;
[0120] A beamforming module 20, configured to form a plurality of pickup beams within each pickup area;
[0121] The sound source activity detection module 30 is configured to receive at least one original audio signal from the conference environment and obtain audio features of each original audio signal; based on the audio features of the original audio signal, determine a target audio signal containing a speech signal from each original audio signal; and determine the location of a target sound source to be picked up based on the target audio signal;
[0122] The beam switching pickup module 40 controls any one of the alternative beams to point to the target sound source position to pick up the target audio signal when the target sound source position is within the pickup area. The alternative beam refers to the pickup beam that switches to the standby state when no voice signal is detected within a preset time.
[0123] In some embodiments, the beam switching pickup module 40 switches from a pickup beam to a standby state as an alternative beam, and is specifically used to: obtain beam pickup signals collected by several pickup beams; determine the time-frequency point of each beam pickup signal and the signal energy corresponding to each time-frequency point; compare the signal energy of each pickup beam at the time-frequency point, determine the pickup beam with the largest signal energy corresponding to each time-frequency point, and mark the time-frequency point as the effective sound source time-frequency point of the pickup beam with the largest signal energy; count the number of effective sound source time-frequency points corresponding to each pickup beam, and when the number of effective sound source time-frequency points of the pickup beam within a preset time length is less than a preset threshold, switch the pickup beam to an alternative beam.
[0124] The beam width of the candidate beam is set to 10° to 50°.
[0125] In some embodiments, the beam switching pickup module 40 is further used to determine the speech similarity of the beam pickup signals of the two pickup beams when the difference in signal energy of the two pickup beams at a time-frequency point is less than a preset difference; and when the speech similarity of the beam pickup signals is greater than a preset similarity threshold, mark the time-frequency point as an effective sound source time-frequency point of one of the pickup beams.
[0126] In some embodiments, the audio features include the signal energy and signal frequency of the original picked-up signal. The sound source activity detection module 30 is specifically used to input the signal energy and signal frequency of each original audio signal into a well-trained speech detection model. If the speech detection model detects that there is a speech signal in the original audio signal, the original audio signal is the target audio signal.
[0127] In some embodiments, the sound source activity detection module 30 is also used to obtain the covariance matrix of the target audio signal in the frequency domain; perform eigendecomposition on the covariance matrix to obtain the signal subspace and the noise subspace; scan the correlation between the sound source directions and the noise subspace in different directions, and determine the sound source direction with the smallest correlation as the target sound source position to be picked up.
[0128] In some embodiments, the pickup area acquisition module 10 is also used to obtain at least one shielding area set in the conference environment; the device also includes a sound shielding module, which shields the target sound source position when the target sound source position is within the shielding area.
[0129] In some embodiments, the beam switching pickup module 40 is further configured to calculate and filter out interference energy generated in each pickup beam by an external interference sound source, so that each pickup beam outputs a valid speech signal of the beam center sound source.
[0130] In some embodiments, the pickup beam also includes a main beam, which is used to pick up ambient noise in the pickup area. The beam switching pickup module 40 is also used to: when the target sound source position is within the pickup area and the main beam picks up ambient noise in the pickup area, control any one of the alternative beams to switch to the pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a pickup beam that switches to a standby state when no voice signal is detected within a preset time period.
[0131] The present application provides a dynamic sound pickup system based on a microphone array, which includes a microphone array including a plurality of microphone elements and a processor connected to the microphone elements.
[0132] The microphone element is used to receive at least one original audio signal from the conference environment.
[0133] The processor is used to: obtain audio features of each original audio signal; based on the audio features of the original audio signal, determine the target audio signal containing a valid voice signal from each original audio signal; based on the target audio signal, determine the target sound source position to be picked up; according to at least one pickup area set in the conference environment, control a number of microphone elements to form a number of pickup beams in each pickup area; when the target sound source position is within the pickup area, control any one of the alternative beams to switch to the pickup beam, pointing to the target sound source position to pick up the target audio signal, the alternative beam refers to the pickup beam that switches to the standby state when no voice signal is detected within a preset time period.
[0134] In some embodiments, the processor is further used to: obtain beam pickup signals collected by several pickup beams; determine the time-frequency point of each beam pickup signal and the signal energy corresponding to each time-frequency point; compare the signal energy of each pickup beam at the time-frequency point, determine the pickup beam with the largest signal energy corresponding to each time-frequency point, and mark the time-frequency point as the effective sound source time-frequency point of the pickup beam with the largest signal energy; count the number of effective sound source time-frequency points corresponding to each pickup beam, and when the number of effective sound source time-frequency points of the pickup beam within a preset time length is less than a preset threshold, switch the pickup beam to an alternative beam.
[0135] In some embodiments, the processor is further configured to: when the difference in signal energy between the two pickup beams at a time-frequency point is less than a preset difference, determine the speech similarity of the beam pickup signals of the two pickup beams; and when the speech similarity of the beam pickup signals is greater than a preset similarity threshold, mark the time-frequency point as a valid sound source time-frequency point of one of the pickup beams.
[0136] In some embodiments, the audio features include the signal energy and signal frequency of the original audio signal; the processor is further used to: input the signal energy and signal frequency of each original audio signal into a fully trained speech detection model; if the speech detection model detects that there is a speech signal in the original audio signal, then the original audio signal is the target audio signal.
[0137] In some embodiments, the processor is also used to: obtain the covariance matrix of the target audio signal in the frequency domain; perform eigendecomposition on the covariance matrix to obtain a signal subspace and a noise subspace; scan the correlation between the sound source directions and the noise subspace in different directions, and determine the sound source direction with the smallest correlation as the target sound source position to be picked up.
[0138] In some embodiments, the processor is further configured to: obtain a shielding area set in the conference environment; and shield the target sound source position when the target sound source position is within the shielding area.
[0139] In some embodiments, the beam width of the candidate beam is set to 10° to 50°.
[0140] In some embodiments, the pickup beam also includes a main beam, which is used to pick up other effective sound sources with voice signals in the pickup area. The processor is also used to: when the target sound source position is within the pickup area and the main beam picks up other effective sound sources with voice signals in the pickup area, control any one of the alternative beams to switch to the pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a pickup beam that switches to a standby state when no voice signal is detected within a preset time period.
[0141] In some embodiments, the processor is further configured to calculate and filter out interference energy generated in each sound pickup beam by an external interfering sound source of each sound pickup beam, so that each sound pickup beam outputs a speech signal of a sound source at the center of the beam.
[0142] The working process, working details and technical effects of the processor provided in this embodiment can be found in the embodiment of the dynamic sound pickup method based on the microphone array above, and will not be repeated here.
[0143] The present application provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the dynamic sound pickup method based on a microphone array as described in any of the above embodiments. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a memory stick. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device.
[0144] The working process, working details and technical effects of the computer-readable storage medium provided in this embodiment can be found in the above embodiment of a dynamic sound pickup method based on a microphone array, which will not be described in detail here.
[0145] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0146] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0147] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, the division of the above-mentioned functional units and modules is only used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system described in this application is divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A dynamic sound pickup method based on a microphone array, characterized in that: include: receiving at least one original audio signal from a conference environment, and obtaining audio features of each of the original audio signals; Determining a target audio signal containing a speech signal from each of the original audio signals based on the audio features of the original audio signals; Determining a target sound source position to be picked up based on the target audio signal; According to at least one pickup area set in the conference environment, forming a plurality of pickup beams within each pickup area; When the target sound source position is within the sound pickup area, any one of the alternative beams is controlled to switch to a sound pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a sound pickup beam that is switched to a standby state when no voice signal is detected within a preset time length.
2. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: Switch from the pickup beam to an alternative beam, including: Acquiring beam pickup signals collected by a plurality of pickup beams; Determine the time-frequency point of each beam pickup signal and the signal energy corresponding to each time-frequency point; Comparing the signal energies of the various pickup beams at the time-frequency points, determining the pickup beam with the maximum signal energy corresponding to each time-frequency point, and marking the time-frequency point as the effective sound source time-frequency point of the pickup beam with the maximum signal energy; The number of effective sound source time-frequency points corresponding to each sound pickup beam is counted, and when the number of effective sound source time-frequency points of the sound pickup beam within a preset time length is less than a preset threshold, the sound pickup beam is switched to an alternative beam.
3. The dynamic sound pickup method based on microphone array according to claim 2, characterized in that: Also includes: When the difference in signal energy between the two pickup beams at a time-frequency point is less than a preset difference, the speech similarity of the beam pickup signals of the two pickup beams is determined. When the speech similarity of the beam pickup signals is greater than a preset similarity threshold, the time-frequency point is marked as a valid sound source time-frequency point of one of the pickup beams.
4. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: The audio features include signal energy and signal frequency of the original audio signals; and determining a target audio signal containing a speech signal from each of the original audio signals based on the audio features of the original audio signals includes: The signal energy and signal frequency of each of the original audio signals are input into a well-trained speech detection model; if the speech detection model detects that a speech signal exists in the original audio signal, the original audio signal is the target audio signal.
5. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: Determining a target sound source position to be picked up based on the target audio signal includes: Obtaining a covariance matrix of the target audio signal in the frequency domain; Performing eigendecomposition on the covariance matrix to obtain a signal subspace and a noise subspace; The correlation between the sound source direction and the noise subspace in different directions is scanned, and the sound source direction with the smallest correlation is determined as the target sound source position to be picked up.
6. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: include: Obtain at least one shielded area set in the conference environment; When the target sound source position is within the shielding area, the target sound source position is shielded.
7. The dynamic sound pickup method based on a microphone array according to any one of claims 1 to 6, characterized in that: The beam width of the candidate beam is set to 10° to 50°.
8. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: The sound pickup beam further includes a main beam, which is used to pick up other effective sound sources with voice signals in the sound pickup area, and further includes: When the target sound source position is within the sound pickup area and the main beam picks up other valid sound sources with voice signals in the sound pickup area, any one of the alternative beams is controlled to switch to the sound pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to the sound pickup beam that switches to the standby state when no voice signal is detected within a preset time period.
9. The dynamic sound pickup method based on microphone array according to claim 1, characterized in that: Also includes: The interference energy generated in each sound pickup beam by an external interference sound source of each sound pickup beam is calculated and filtered out, so that each sound pickup beam outputs a speech signal of a sound source at the center of the beam.
10. A dynamic sound pickup system based on a microphone array, characterized in that: The system includes a microphone array, wherein the microphone array includes a plurality of microphone elements and a processor connected to the microphone elements. The microphone element is used to receive at least one original audio signal of the conference environment; The processor is configured to: Acquiring audio features of each of the original audio signals; Determining, based on the audio features of the original audio signals, a target audio signal containing a valid speech signal from each of the original audio signals; Determining a target sound source position to be picked up based on the target audio signal; According to at least one sound pickup area set in the conference environment, controlling the plurality of microphone elements to form a plurality of sound pickup beams in each sound pickup area; When the target sound source position is within the sound pickup area, any one of the alternative beams is controlled to switch to a sound pickup beam, pointing to the target sound source position to pick up the target audio signal. The alternative beam refers to a sound pickup beam that is switched to a standby state when no voice signal is detected within a preset time length.
Citation Information
Cited By
Sound source angle detection method and device, electronic equipment, medium and product
CN120949158A