Adaptive array microphone cascaded mixing method
By employing an adaptive array microphone cascade mixing method, which combines a synchronization module, a channel gain equalization estimation module, and an adaptive mixing module, the problems of inaccurate sound localization and volume imbalance in microphone array cascade systems are solved, resulting in better voice quality and user experience.
Patent Information
- Application Number
- CN202411704878.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-11-26
AI Technical Summary
In existing technologies, microphone array cascade systems ignore the actual distance differences between different microphone arrays and the speaker, as well as the uneven sound pickup quality, when mixing speech, leading to problems such as inaccurate sound localization, uneven volume, noise interference, and echo effects.
An adaptive array microphone cascade mixing method is adopted. Through the collaborative work of a synchronization module, a channel gain equalization estimation module, and an adaptive mixing module, adaptive mixing processing is achieved, including techniques such as relative delay analysis, speech activation analysis, channel gain equalization, and delay addition.
It achieves volume equalization, target speech enhancement, signal quality optimization, and defect repair, ensuring balanced output of the speaker's voice from different positions, thus improving overall speech quality and user experience.
Smart Images

Figure CN119545255B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of communication, and particularly relates to a cascaded mixing method for adaptive array microphones. BACKGROUND
[0002] In a conference scenario, a microphone device cascading scheme is widely used to cover a wider pickup area and ensure that the speech of a participant in any corner of the conference room can be clearly captured. A traditional conference microphone cascading system is usually composed of multiple microphone array devices that are scattered and arranged at key positions in the conference room to achieve comprehensive coverage of the entire conference space. For example, in the prior art, a Chinese patent with publication number CN205666934U discloses a headset and a cascading system thereof, wherein the headset includes a headset speaker and a microphone, and further includes: an audio interface in electrical signal communication with the headset speaker and the microphone, for signal interaction with an external audio device; at least one cascading interface in electrical signal communication with the headset speaker, the microphone, and the audio interface, for cascading with other headsets; wherein the cascading system is composed of multiple headsets in a star connection or a serial connection structure.
[0003] Although the above-mentioned prior art can cascade multiple headsets for use by multiple people and can achieve a director function, it can be particularly applied in rich application scenarios such as chorus and network conference, but in actual application, how to effectively mix the output speech of these microphone array devices has become a technical problem in the field.
[0004] In a traditional mixing method, the output speech of each microphone array device is often simply added or averaged, and this processing method ignores the actual distance difference between different microphone arrays and speakers, as well as the uneven quality of pickup of each array. Therefore, the final output mixed sound result often has the following problems: first, the output speech of the microphone array far from the speaker may occupy too much weight, resulting in inaccurate sound positioning and reducing the clarity of the speech; second, when multiple participants at different positions speak at the same time, the sound volume at each position is difficult to balance, and some sounds may be suppressed or amplified too much; third, the output of a single microphone array may have defects such as noise interference and echo effect, which may be amplified in the mixing process, thereby affecting the overall speech quality.
[0005] The most commonly used method in the prior art is to configure different mixing weights according to the energy of the output signal of each device, and the calculation formula of the mixed signal is:
[0006] ;
[0007] wherein, represents the energy of the output signal of the i-th microphone array device, and time domain signal of the channel; is a weighting coefficient of the mth channel; is a microphone array channel index, taking a value between 0 and the number of microphones of the microphone array, represents a discrete time sequence number, is a positive integer. Assuming that the output energy of the mth array channel is maximum, and the maximum output energy is: , then the maximum output energy The weight configuration of the corresponding mth array is: ; the weight configuration of other arrays except the mth array is: ; the output energy of other arrays is all less than the maximum output energy ; the advantage of this scheme is simple to implement, but there are also the above problems: if there are people speaking at different positions, this will lead to the inability to simultaneously enhance all voices, and when speaking, one voice is enhanced and one voice is weakened, and multiple microphone arrays cannot be used to simultaneously enhance voices of people at a distance.
[0008] Therefore, there is an urgent need for an adaptive array microphone cascading mixing method to obtain better cascading voice quality. SUMMARY
[0009] To solve the problems in the related art, the present application provides an adaptive array microphone cascading mixing method to overcome the above technical problems existing in the prior art. Through the cooperative work of the synchronization module, the channel gain equalization estimation module and the adaptive mixing module, the present application can intelligently process signals from different microphone arrays and realize adaptive mixing processing. The method also enables the present application to achieve multiple beneficial effects such as volume equalization, target voice enhancement, signal quality optimization, adaptive mixing processing, output defect repair and low-quality output exclusion.
[0010] The technical solution of the present application is as follows: an adaptive array microphone cascading mixing method applied to a microphone cascading mixing system, the microphone cascading mixing system comprising a synchronization module, a channel gain equalization estimation module and an adaptive mixing module, the synchronization module and the channel gain equalization estimation module being connected through multiple channel one communication connections, the channel gain equalization estimation module and the adaptive mixing module being connected through multiple channel two communication connections, and the synchronization module and the adaptive mixing module being connected through multiple channel three communication connections.
[0011] Further comprising a first microphone array device, a second microphone array device, …, an Nth microphone array device, N being a positive integer; wherein the first microphone array device is connected with the synchronization module through a Ch0 channel, the second microphone array device is connected with the synchronization module through a Ch1 channel, …, and the Nth microphone array device is connected with the synchronization module through a Ch(N-1) channel;
[0012] The method comprises the following steps:
[0013] Step one: each microphone array device transmits data to the synchronization module, the synchronization module performs relative delay analysis and voice activation analysis on the data of each channel received, and obtains an analysis result; and then performs secondary phase-level synchronization processing on the data of each channel according to the analysis result, and obtains a first signal; the secondary phase-level synchronization processing comprises sequentially performed coarse delay synchronization processing and phase synchronization processing;
[0014] Step two: the channel gain equalization estimation module first receives the first signal after the secondary phase-level synchronization processing, and selects a first signal that needs to be processed by the sound mixing synchronization processing from the first signal, and obtains a second signal after the sound mixing synchronization processing; in the selection process, the channel gain equalization estimation module analyzes the relative distance of a plurality of channels one from the speaker according to the energy of each channel one, and selects the plurality of channels one as candidate channels, while discarding other channels one; wherein the first signal corresponding to the candidate channels needs to be processed by the sound mixing synchronization processing;
[0015] Step three: the adaptive sound mixing module first receives the second signal processed by the channel gain equalization estimation module; the adaptive sound mixing module then performs smoothing processing on the gain obtained by the second signal in the channel gain equalization estimation module, and applies the gain to each aligned channel three signal, and then performs delay addition and sound mixing processing, and finally obtains a sound mixing signal after the adaptive array microphone cascade sound mixing.
[0016] Further, the synchronization module comprises a coarse delay synchronization unit and a phase synchronization unit which are in communication with each other; the coarse delay synchronization unit is used for performing coarse delay synchronization processing, and the phase synchronization unit is used for performing phase synchronization processing;
[0017] Each of the microphone array devices is connected with the coarse delay synchronization unit through different channels, i.e., the first microphone array device is connected with the coarse delay synchronization unit through a Ch0 channel, …, and the Nth microphone array device is connected with the coarse delay synchronization unit through a Ch(N-1) channel; the phase synchronization unit is in communication with the channel gain equalization estimation module and the adaptive sound mixing module;
[0018] Further, the phase synchronization unit is in communication with the channel gain equalization estimation module through a plurality of channels two; and the phase synchronization unit is in communication with the adaptive sound mixing module through a plurality of channels three.
[0019] Further, in the step one, the process of the secondary phase level synchronization processing is as follows:
[0020] The Ch0 channel is selected as the reference channel for the delay analysis;
[0021] The VAD algorithm is used to analyze whether the signal in the Ch0 channel is speech or noise; it should be noted that the VAD algorithm can be any optional existing algorithm, which is not specifically limited in the present application;
[0022] When the VAD algorithm detects that the speech frame of the signal in a certain channel is 1, the delay analysis is performed:
[0023] Taking the Ch0 channel and the Ch1 channel as examples, it is assumed that the real Fourier transform is performed on and to obtain the frequency domain signals of points and , and the calculation formula of the frequency domain signal is:
[0024] ;
[0025] wherein, 0 represents the channel ID, represents the speech frame, z is a positive integer, for example: if a frame of data has 320 sample points, represents all data of 1:320 points, represents all data of 321:640 points, and so on;
[0026] The calculation formula of the frequency domain signal is:
[0027] ;
[0028] wherein, ; K represents the number of points of the fast Fourier transform (FFT); represents the frequency; is the real Fourier transform;
[0029] The cross-correlation calculation is performed on the frequency domain signals and to obtain the coherence coefficient , and the cross-correlation calculation formula is:
[0030] ;
[0031] wherein, is the smoothing coefficient, and * represents the conjugate;
[0032] The coherent coefficient is normalized and weighted The calculation formula is as follows:
[0033] ;
[0034] The inverse real Fourier transform of the length is performed on to obtain R, and the calculation formula is as follows:
[0035] ;
[0036] wherein, is the inverse real Fourier transform;
[0037] Finally, the calculation formula of the relative delay of the two signals is as follows:
[0038] ;
[0039] wherein, indicates the index value when the maximum value is obtained; M is the up-sampling multiple, and the value range is 4-6; indicates the delay sample number of the Ch0 channel and the Ch1 channel.
[0040] Further, after the relative delay of each channel with respect to the Ch0 channel is calculated, secondary phase level synchronization processing is performed on each channel, and a first signal is obtained after the secondary phase level synchronization processing.
[0041] It should be noted that, if only the coarse synchronization is used in the application, the delay synchronization granularity is too large and inaccurate; if only the phase synchronization is used, the delay between signals is too large, and the high-frequency part is spatially mixed to cause signal distortion; therefore, the application adopts secondary synchronization processing, that is, the coarse synchronization and the phase synchronization are combined.
[0042] The process of the coarse synchronization processing is as follows:
[0043] For the signal except the Ch0 channel, the calculation formula of the alignment to the Ch0 channel in time is as follows:
[0044] ;
[0045] wherein, indicates the integer not greater than the value; indicates the sequence number of the discrete time, and is a positive integer.
[0046] Further, the process of the phase synchronization processing is as follows:
[0047] For the signal of removing Ch0 channel, higher precision than coarse synchronization processing is aligned to the time of Ch0 channel, and the calculation formula is:
[0048] ;
[0049] wherein, is the kth frequency point of the current frame, k represents the kth frequency point of the current frame, represents the short-time Fourier transform;
[0050] The calculation formula of fractional delay is:
[0051] ;
[0052] ;
[0053] wherein, is a complex exponential function, is an imaginary unit, is a sampling frequency;
[0054] The calculation formula of the signal after the same synchronization is:
[0055] ;
[0056] wherein, represents the short-time inverse Fourier transform;
[0057] It should be noted that the signal after strict synchronization can be obtained according to the above steps, and the signal and , , and the like can be obtained by the same method as the signal .
[0058] Further, in the step two, for the first signal which needs to be mixed and synchronized, the channel gain equalization estimation module sets the corresponding channel gain to , and the channel gain corresponding to the first signal which does not need to be mixed and synchronized is set to ; the specific steps are as follows:
[0059] The first-order smooth energy of each channel signal after alignment is calculated as , and the calculation formula is:
[0060] ;
[0061] wherein, represents the sequence number of discrete time, and is a positive integer; is the smoothing coefficient, is the smoothing energy of channel m, m is a positive integer; represents the current gain of the current channel m.
[0062] Further, according to the first-order smoothing energy obtained, the channel with the current maximum energy is obtained, and the calculation formula is:
[0063] , is set at the same time ;
[0064] The energy of all other channels is compared with the maximum channel energy obtained, if , is set , otherwise is set
[0065] , wherein, is the channel index number with the maximum energy; represents the index value when the maximum value is obtained; is the vad flag of channel m; is the energy interpolation threshold.
[0066] Further, in order to prevent fluctuations, hangover processing is also performed:
[0067] ;
[0068] ;
[0069] ;
[0070] ;
[0071] , wherein, represents the hangover frame number of the mth channel;
[0072] if the value is greater than 0, is set , otherwise is set .
[0073] Further, in the third step, the adaptive mixing module performs gain smoothing processing according to the channel gain calculation result in the channel gain equalization estimation module, and the calculation formula is:
[0074] ;
[0075] , wherein, is a smoothing constant, and its value ranges from 0 to 1; is the smoothing gain of the channel m; is the current gain of the current channel m.
[0076] Further, the calculation formula of the delay addition and the mixing processing is:
[0077]
[0078] wherein, is the serial number of the discrete time, and is a positive integer; the mixed signal after mixing is that is, the mixed signal after the final mixing of the adaptive array microphone cascade.
[0079] Advantages of the present application:
[0080] (1) Firstly, the present application can intelligently process signals from different microphone arrays and realize adaptive mixing processing through the cooperative work of the synchronization module, the channel gain equalization estimation module and the adaptive mixing module; and the present application can better equalize the volume when multiple target speakers speak at the same time in the cascade array scene, avoid the situation that a certain speaker is suppressed, and ensure that the sound of speakers in different positions can obtain more balanced volume output.
[0081] (2) Moreover, when the target speaker is at a relative distance from multiple microphone arrays, the present application further enhances the target voice by using a method similar to delay addition, and the enhancement effect is particularly obvious when the target speaker is far away from all arrays.
[0082] (3) In addition, when the target speaker is close to one microphone and far from the other microphone, the present application can recognize this situation and preferentially use the signal of the close microphone (because its quality is much better than that of the far microphone), while discarding the input of the far microphone, thereby ensuring the quality of the target voice. BRIEF DESCRIPTION OF DRAWINGS
[0083] Fig. 1 is a flow chart of an adaptive array microphone cascade mixing method of the present application;
[0084] Fig. 2 is a specific implementation flow chart of the synchronization module of the present application. DETAILED DESCRIPTION
[0085] With reference to the accompanying drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below, obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present application.
[0086] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like indicate the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.
[0087] As shown in Figs. 1-2 The embodiment provides a self-adaptive array microphone cascade mixing method, which is applied to a microphone cascade mixing system, the microphone cascade mixing system comprises a synchronization module, a channel gain equalization estimation module and a self-adaptive mixing module, the synchronization module and the channel gain equalization estimation module are connected through a plurality of channel one communication connections, the channel gain equalization estimation module and the self-adaptive mixing module are connected through a plurality of channel two communication connections, and the synchronization module and the self-adaptive mixing module are connected through a plurality of channel three communication connections.
[0088] Further comprising a first microphone array device, a second microphone array device,..., and an Nth microphone array device, N being a positive integer; wherein the first microphone array device is connected with the synchronization module through a Ch0 channel, the second microphone array device is connected with the synchronization module through a Ch1 channel,..., and the Nth microphone array device is connected with the synchronization module through a Ch (N-1) channel.
[0089] The method comprises the following steps:
[0090] Step one: each microphone array device transmits data to the synchronization module, the synchronization module performs relative delay analysis and voice activation analysis on the data of each channel received, obtains an analysis result, and then performs two-level phase-level synchronization processing on the data of each channel according to the analysis result to obtain a first signal; the two-level phase-level synchronization processing comprises sequentially performed coarse delay synchronization processing and phase synchronization processing.
[0091] Step two: the channel gain equalization estimation module first receives the first signal after the secondary phase level synchronization processing, and selects the first signal that needs to be processed by the sound mixing synchronization from the first signal, and obtains the second signal after the sound mixing synchronization processing; in the selection process, the channel gain equalization estimation module analyzes the relative distance of a plurality of channel one from the speaker according to the energy of each channel one, and selects a plurality of channel one as a candidate channel, and discards other channel one; wherein, the first signal corresponding to the candidate channel needs to be processed by the sound mixing synchronization;
[0092] Step three: the adaptive sound mixing module receives the second signal processed by the channel gain equalization estimation module; the adaptive sound mixing module further smoothes the gain obtained by the second signal in the channel gain equalization estimation module, and applies the gain to each aligned channel three signal, and then performs delay addition and sound mixing processing, and finally obtains the sound mixing signal after the adaptive array microphone cascade sound mixing.
[0093] Firstly, through the cooperative work of the synchronization module, the channel gain equalization estimation module and the adaptive sound mixing module, the embodiment can intelligently process signals from different microphone arrays, and realize adaptive sound mixing processing.
[0094] Secondly, the embodiment can better balance the volume of multiple target speakers speaking at the same time in the cascade array scene, avoid the situation that a speaker is suppressed, and ensure that the voices of speakers in different positions can obtain more balanced volume output.
[0095] Moreover, in the case that the target speaker is at a relatively close distance from the multiple microphone arrays, the embodiment further enhances the target voice by using a method similar to delay addition, and the enhancement effect is particularly obvious when the target speaker is far away from all arrays.
[0096] In addition, when the target speaker is close to one microphone and far from the other microphone, the embodiment can identify this situation, and preferentially use the signal of the close microphone (because its quality is much better than that of the far microphone), and discard the input of the far microphone, so as to ensure the quality of the target voice.
[0097] Specifically, the synchronization module comprises a coarse delay synchronization unit and a phase synchronization unit which are in communication with each other; the coarse delay synchronization unit is used for coarse delay synchronization processing, and the phase synchronization unit is used for phase synchronization processing.
[0098] The microphone array devices are connected with the coarse delay synchronization unit through different channels, that is, the first microphone array device is connected with the coarse delay synchronization unit through Ch0 channel, and the Nth microphone array device is connected with the coarse delay synchronization unit through Ch(N-1) channel; the phase synchronization unit is in communication connection with the channel gain equalization estimation module and the adaptive mixing module respectively;
[0099] More specifically, the phase synchronization unit is in communication connection with the channel gain equalization estimation module through multiple channels two; and the phase synchronization unit is in communication connection with the adaptive mixing module through multiple channels three.
[0100] Specifically, in the step one, the process of the secondary phase synchronization is as follows:
[0101] Ch0 channel is selected as the reference channel for delay analysis;
[0102] The VAD algorithm is used to analyze whether the signal in Ch0 channel is speech or noise; it should be noted that the VAD algorithm can be any optional existing algorithm, which is not specifically limited in the embodiment;
[0103] When the VAD algorithm detects that the speech frame of the signal in a certain channel is 1, delay analysis is performed:
[0104] Taking Ch0 channel and Ch1 channel as an example, it is assumed that and are subjected to real Fourier transform, and points of and are obtained, and the calculation formula of the frequency domain signal is:
[0105] ;
[0106] wherein, 0 represents the channel ID, represents the speech frame, z is a positive integer, for example: if a frame of data has 320 sample points, represents all data of 1:320 points, represents all data of 321:640 points, and so on;
[0107] The calculation formula of the frequency domain signal is:
[0108] ;
[0109] wherein, ; K represents the number of points of fast Fourier transform (FFT); represents the frequency; is the real Fourier transform;
[0110] The frequency domain signals And Do each other's coherence calculation, get the coherence coefficient , The mutual coherence calculation formula is:
[0111] ;
[0112] Among them, The smoothing coefficient, * indicates taking the conjugate;
[0113] Again, the coherence coefficient Do the normalized weighting, the calculation formula is:
[0114] ;
[0115] Then, the Do Length of inverse real Fourier transform to get R, the calculation formula is:
[0116] ;
[0117] Among them, Inverse real Fourier transform;
[0118] Finally, the calculation formula of the relative delay of two signals is:
[0119] ;
[0120] Among them, Indicates the index value when the maximum value is obtained; M is the upsampling multiple, the value range is 4-6; Indicates the delay sample number of Ch0 channel and Ch1 channel.
[0121] Specifically, after calculating the relative delay of each channel corresponding to the Ch0 channel, the secondary phase level synchronization processing is started for each channel, and the first signal is obtained after the secondary phase level synchronization processing.
[0122] It should be noted that if only the coarse synchronization is used in the embodiment, the delay synchronization granularity will be too large and inaccurate; if only the phase synchronization is used, the delay between signals will be too large, resulting in spatial aliasing of high frequency part and signal distortion; therefore, the embodiment adopts two-level synchronization processing, that is, coarse synchronization and phase synchronization are combined;
[0123] Among them, the process of coarse synchronization processing is as follows:
[0124] For the signals except Ch0 channel, align to the time of Ch0 channel, the calculation formula is:
[0125] ;
[0126] wherein, represents an integer not greater than the value; represents a serial number of discrete time, and is a positive integer.
[0127] Specifically, the process of phase synchronization processing is as follows:
[0128] For the signal of removing Ch0 channel, higher accuracy than coarse synchronization processing is aligned to the time of Ch0 channel, and the calculation formula is:
[0129] ;
[0130] wherein, is a short-time Fourier transform of the first frame, k represents the kth frequency point of the current frame,
[0131] The calculation formula of fractional delay is:
[0132] ;
[0133] ;
[0134] wherein, is a complex exponential function, is an imaginary unit, is a sampling frequency;
[0135] The calculation formula of the signal after phase synchronization is:
[0136] ;
[0137] wherein, is a short-time inverse Fourier transform;
[0138] It should be noted that the signal after strict synchronization can be obtained according to the above steps, and the signal and , , and the like have the same acquisition method as the signal .
[0139] Specifically, in the step two, for the first signal which needs to be processed by mixing sound synchronization, the channel gain equalization estimation module sets the corresponding channel gain of the first signal to , and the channel gain corresponding to the first signal which does not need to be processed by mixing sound synchronization is set to ; the specific steps are as follows:
[0140] the first-order smoothed energy of each channel signal after alignment , the calculation formula is:
[0141] ;
[0142] wherein, represents the sequence number of discrete time, and is a positive integer; is a smoothing coefficient, is the smoothed energy of channel m, m is a positive integer; represents the current gain of the current channel m.
[0143] Specifically, according to the result of the first-order smoothed energy obtained, the channel with the current maximum energy is obtained , the calculation formula is:
[0144] , and ;
[0145] the energy of all other channels is compared with the maximum channel energy obtained, if , is set , otherwise ;
[0146] wherein, is the index number of the channel with the maximum energy; represents the index value when the maximum value is obtained; is the vad flag of channel m; is the energy interpolation threshold.
[0147] Specifically, in order to prevent fluctuations, hangover processing is also performed:
[0148] ;
[0149] ;
[0150] ;
[0151] ;
[0152] wherein, represents the hangover frame number of the mth channel;
[0153] if the value is greater than 0, is set , otherwise .
[0154] Specifically, in the step three, the adaptive mixing module performs gain smoothing processing according to the channel gain calculation result in the channel gain equalization estimation module, and the calculation formula is:
[0155] ;
[0156] wherein, is a smoothing constant, and the value range is between 0 and 1; is the smoothing gain of the channel m; represents the current gain of the current channel m.
[0157] Specifically, the calculation formula of the delay addition and mixing processing is:
[0158] ;
[0159] wherein, represents the serial number of the discrete time, and is a positive integer; the mixed signal after mixing is , which is the final mixed signal after the adaptive array microphone cascade mixing.
[0160] In the embodiment, the adaptive array microphone cascade mixing method can make the multiple microphone arrays after cascade aggregation close to the speaker array output, enhance the speaker voice, and repair the output defects of a single array using the output of different microphone arrays, thereby improving the overall mixed voice quality. Moreover, while maintaining the above advantages, the embodiment excludes the relatively far array output to prevent mixing the relatively poor quality array output, thereby avoiding the deterioration of the overall mixed voice quality.
[0161] In summary, the embodiment realizes multiple beneficial effects such as volume equalization, target voice enhancement, signal quality optimization, adaptive mixing processing, output defect repair, and low-quality output exclusion in the cascade array microphone scene, significantly improving the overall voice processing performance and user experience.
[0162] According to the disclosure and teaching of the above description, those skilled in the art of the present application can also make changes and modifications to the above embodiments. Therefore, the present application is not limited to the specific embodiments disclosed and described above, and some modifications and changes of the present application should also fall within the protection scope of the claims of the present application. In addition, although some specific terms are used in the present specification, these terms are only for convenience of explanation and do not constitute any limitation on the present application.
Claims
1. An adaptive array microphone cascading mixing method applied to a microphone cascading mixing system, the microphone cascading mixing system comprising a synchronization module, a channel gain equalization estimation module and an adaptive mixing module, the synchronization module being connected with the channel gain equalization estimation module through a plurality of channel one communications, the channel gain equalization estimation module being connected with the adaptive mixing module through a plurality of channel two communications, and the synchronization module being connected with the adaptive mixing module through a plurality of channel three communications; Further comprising a first microphone array device, a second microphone array device,..., and an Nth microphone array device, N being a positive integer; wherein, the first microphone array device is connected with the synchronization module through a Ch0 channel, the second microphone array device is connected with the synchronization module through a Ch1 channel,..., and the Nth microphone array device is connected with the synchronization module through a Ch(N-1) channel; characterized in that the method comprises the following steps: Step one: each microphone array device transmits data to the synchronization module, the synchronization module performs relative delay analysis and voice activation analysis on the data received by each channel to obtain an analysis result, and then performs two-level phase-level synchronization processing on the data of each channel according to the analysis result to obtain a first signal; the two-level phase-level synchronization processing comprises sequentially performing coarse delay synchronization processing and phase synchronization processing; Step two: the channel gain equalization estimation module first receives the first signal after the two-level phase-level synchronization processing, and selects a first signal that needs to be processed by mixing synchronization from the first signal; in the selection process, the channel gain equalization estimation module analyzes the energy in each channel one to determine the relative distance from the speaker of a plurality of channel ones, and selects the channel ones as candidate channels while discarding other channel ones; wherein, the first signal corresponding to the candidate channel needs to be processed by mixing synchronization; Step three: the adaptive mixing module first receives the second signal processed by the channel gain equalization estimation module; the adaptive mixing module then smoothes the gain obtained by the second signal in the channel gain equalization estimation module and applies it to each aligned channel three signal, and then performs delay addition and mixing processing to finally obtain a mixing signal after adaptive array microphone cascading mixing.
2. The adaptive array microphone cascading mixing method according to claim 1, characterized in that, The synchronization module comprises a coarse delay synchronization unit and a phase synchronization unit that are in communication with each other; the coarse delay synchronization unit is used for coarse delay synchronization processing, and the phase synchronization unit is used for phase synchronization processing; Each microphone array device is connected with the coarse delay synchronization unit through a different channel, i.e., the first microphone array device is connected with the coarse delay synchronization unit through a Ch0 channel,..., and the Nth microphone array device is connected with the coarse delay synchronization unit through a Ch(N-1) channel.
3. The adaptive array microphone cascading mixing method according to claim 2, characterized in that, In the step one, the process of two-level phase-level synchronization processing is as follows: select Ch0 channel as a reference channel for delay analysis; analyze whether the signal in Ch0 channel is voice or noise through VAD algorithm; when the VAD algorithm detects that the voice frame of the signal in a certain channel is 1, perform delay analysis: Taking Ch0 channel and Ch1 channel as examples, it is assumed that real Fourier transform is performed on and to obtain frequency domain signals of points of and , and the calculation formula of the frequency domain signals is ; Wherein, 0 represents the channel ID, z represents the speech frame, z is a positive integer; Frequency domain signal The calculation formula is: ; wherein, denotes a time domain signal of the Ch0 channel, denotes a time domain signal of the Ch1 channel; K denotes a number of points of a fast Fourier transform, FFT; denotes a frequency; is a real Fourier transform; On the frequency domain signal And Do mutual coherence calculation, get coherence coefficient ; The coherence coefficient is normalized again The calculation formula is: ; Then R is obtained by taking the inverse real Fourier transform of the length-2N The inverse real Fourier transform of length 2N gives R, with the formula ; wherein is the inverse real Fourier transform; The relative delay calculation formula of the two signals is: ; wherein, represents the index value at the time of obtaining the maximum value; M is the up-sampling multiple, and the value range is 4-6; represents the delay sample number of Ch0 channel and Ch1 channel.
4. The adaptive array microphone cascading mixing method according to claim 3, characterized in that, After calculating the relative delay of each channel corresponding to the Ch0 channel, the secondary phase level synchronization processing is performed on each channel, and the first signal is obtained after the secondary phase level synchronization processing; The coarse synchronization processing process is as follows: For the signals except the Ch0 channel, the calculation formula for aligning to the time of the Ch0 channel is: ; wherein represents an integer not greater than the value; represents a serial number of a discrete time, and is a positive integer.
5. The adaptive array microphone cascading mixing method according to claim 4, characterized in that, The phase synchronization processing process is as follows: For the signals except the Ch0 channel, the calculation formula for aligning to the time of the Ch0 channel with higher accuracy than the coarse synchronization processing is: ; wherein, is indicative of the kth frequency bin of the current frame k, is indicative of the kth frequency bin of the current frame k, is indicative of the kth frequency bin of the current frame k, The fractional delay calculation formula is: ; ; wherein is a complex exponential function, is the imaginary unit, is the sampling frequency; Phase-synchronized signal The calculation formula is: ; wherein denotes the inverse short-time Fourier transform.
6. The adaptive array microphone cascading mixing method according to claim 5, characterized in that, In the second step, for the first signal which needs to be processed by the mixing synchronization, the channel gain equalization estimation module sets the corresponding channel gain as For the first signal which does not need to be processed by the mixing synchronization, the corresponding channel gain is set as The specific steps are as follows: first order smoothed energy of each channel signal after alignment The formula is: ; wherein denotes a discrete-time sequence number, and is a positive integer; is a smoothing coefficient, is the smoothed energy of channel m, m being a positive integer; denotes the current gain of the current channel m.
7. The adaptive array microphone cascading mixing method according to claim 6, characterized in that, According to the obtained first-order smoothed energy , the channel with the current maximum energy is obtained , and the calculation formula is: while setting ; set the energy of the other channels to zero and compare it to the maximum channel energy found if then set , otherwise set ; wherein, is the index number of the channel with the maximum energy; denotes the index value at which the maximum is taken; is the vad flag for channel m; is the energy interpolation threshold.
8. The adaptive array microphone cascading mixing method according to claim 7, characterized in that, In order to prevent fluctuations, hangover processing is also performed: ; ; ; ; wherein, hangover number of frames for the mth channel; If value is greater than 0, set , otherwise set .
9. The adaptive array microphone cascading mixing method according to claim 1, wherein, In the third step, the adaptive mixing module performs gain smoothing processing according to the channel gain calculation result in the channel gain equalization estimation module to obtain the smoothed gain of the current channel m .
10. The adaptive array microphone cascading mixing method according to claim 9, characterized in that, The delay addition and mixing processing calculation formula is: ; wherein, denotes a sequence number of a discrete time, and is a positive integer; the mixed sound after is the mixed sound signal after the final mixing of the adaptive array microphones in cascade.
Citation Information
Patent Citations
Headset and cascade system thereof
CN205666934U
Microphone sound mixing method, device and equipment and storage medium
CN118155639A
Networked automatic mixer system and method
CN118216161A