Far and near field voice separation method and apparatus, and wearable device
Through the combination of beamforming algorithm and bone conduction microphone signal, the voice masking matrix is calculated, which solves the problem that the existing technology cannot completely separate near-field and far-field voice, and achieves more efficient voice separation and improves user experience.
Patent Information
- Application Number
- CN202510304343.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology cannot completely separate near-field voice and far-field voice, resulting in insufficient signal-to-noise ratio in application scenarios such as voice calls, voice assistants and far-field translation, which affects the user experience.
By obtaining the microphone array signal collected by the first microphone in the target scenario, the beamforming algorithm is used to obtain the near-field and far-field beamforming signals, and combining the bone conduction microphone signals, the near-field and far-field speech masking matrix is calculated, and the near-field and far-field speech masking matrix is then separated.
Effectively separate near-field voice and far-field voice, improve signal-to-noise ratio, provide clearer call experience and more accurate voice assistant response.
Smart Images

Figure CN120108413A_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of speech processing technology, and in particular to a far-field and near-field speech separation method, device and wearable device. [Background technology]
[0002] With the gradual popularity of wearable devices such as True Wireless Stereo (TWS) and Augmented Reality (AR) glasses, voice applications based on wearable devices are becoming more and more widespread. In application scenarios such as voice calls, voice assistants, and far-field translation, the technology to distinguish between the wearer's speech and other people's speech is crucial.
[0003] In terms of voice calls, wearable devices need to be able to eliminate the voices of background speakers other than the wearer to provide a clear calling experience; for voice assistant functions, wearable devices need to be able to recognize the wearer's voice commands and provide responses; for far-field translation scenarios, wearable devices need to be able to distinguish between the wearer's voice and the voice of far-field speakers, and then be able to translate only the far-field voice.
[0004] The common solution in the prior art is to use directional beamforming technology to pick up far-field or near-field voice inputs to related applications based on the multi-microphone audio signals picked up by the microphone array. However, although this solution can improve the signal-to-noise ratio of near-field or far-field voice to a certain extent, it still cannot completely separate near-field voice from far-field voice.
[0005] Therefore, it is necessary to provide a far-field and near-field speech separation method, apparatus and wearable device. [Summary of the invention]
[0006] The purpose of the embodiments of the present invention is to provide a far-field and near-field speech separation method, apparatus and wearable device.
[0007] The technical solution of the present invention is as follows:
[0008] In a first aspect, an embodiment of the present invention provides a method for far-field and near-field speech separation, comprising:
[0009] Acquire a near-field beamforming signal and a far-field beamforming signal according to a microphone array signal collected by a first microphone in a target scene and a beamforming algorithm;
[0010] Calculating a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal;
[0011] Outputting near-field speech according to the near-field beamforming signal and the near-field speech masking matrix;
[0012] The far-field speech is outputted according to the far-field beamforming signal and the far-field speech masking matrix.
[0013] Furthermore, the bone conduction microphone signal is obtained by the following steps:
[0014] Acquire a voice signal of the wearer collected by the second microphone in the target scene;
[0015] The wearer's voice signal is used as the bone conduction microphone signal.
[0016] Further, the calculating, according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal, a near-field speech masking matrix and a far-field speech masking matrix comprises:
[0017] The near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are input into a separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
[0018] Further, the separation neural network is a convolutional recurrent neural network, and the convolutional recurrent neural network includes a convolutional neural network, a recurrent neural network, and a deep neural network. The near-field beamforming signal, the far-field beamforming signal, and the bone conduction microphone signal are input into the separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix, including:
[0019] Inputting the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal into the convolutional neural network to extract near-field audio features and far-field audio features;
[0020] Inputting the near-field audio features and the far-field audio features into the recurrent neural network to extract near-field voice information and far-field voice information;
[0021] The near-field speech information and the far-field speech information are input into the deep neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
[0022] Furthermore, the step of inputting the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal into the convolutional neural network to extract near-field audio features and far-field audio features comprises:
[0023] If the effective bandwidth of the bone conduction microphone signal is within a preset frequency range, the number of convolution kernels of the convolutional neural network will be increased and / or the stride of the convolutional neural network will be reduced;
[0024] If the effective bandwidth of the bone conduction microphone signal is outside the preset frequency range, the number of convolution kernels of the convolutional neural network is reduced and / or the stride of the convolutional neural network is increased.
[0025] Furthermore, according to the microphone array signal collected by the first microphone in the target scene and the beamforming algorithm, the near-field beamforming signal and the far-field beamforming signal are obtained, and the calculation formula is as follows:
[0026]
[0027] Among them, N t,f represents the near-field beamforming signal, F t,f represents the far-field beamforming signal, S i,t,f The time domain signal S represents the signal of the i-th microphone array i The frequency domain signal of the fth frequency point of the tth speech frame after SIFT transformation, represents the first filter weight, represents the second filter weight, i, t, and f are all positive integers.
[0028] Further, the near-field speech is output according to the near-field beamforming signal and the near-field speech masking matrix, and the calculation formula is as follows:
[0029]
[0030] Among them, N t ' ,f represents the near-field speech, N t,f represents the near-field beamforming signal, represents the near-field speech masking matrix, and t and f are both positive integers.
[0031] Further, the far-field speech is output according to the far-field beamforming signal and the far-field speech masking matrix, and the calculation formula is as follows:
[0032]
[0033] Among them, F t ' ,f represents the far-field speech, F t,f represents the far-field beamforming signal, represents the far-field speech masking matrix, and t and f are both positive integers.
[0034] In a second aspect, an embodiment of the present invention provides a far-field and near-field speech separation device, comprising:
[0035] An acquisition module, used to acquire a near-field beamforming signal and a far-field beamforming signal according to a microphone array signal acquired by a first microphone in a target scene and a beamforming algorithm;
[0036] A calculation module, configured to calculate a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal;
[0037] A first output module, configured to output near-field speech according to the near-field beamforming signal and the near-field speech masking matrix;
[0038] The second output module is used to output the far-field speech according to the far-field beamforming signal and the far-field speech masking matrix.
[0039] In a third aspect, an embodiment of the present invention provides a wearable device, comprising a far-field and near-field speech separation device as provided in the first aspect.
[0040] The beneficial effects of the present invention are: obtaining the microphone array signal collected by the first microphone in the target scene, and using the beamforming algorithm to obtain the near-field beamforming signal and the far-field beamforming signal; based on the near-field beamforming signal and the far-field beamforming signal, taking the bone conduction microphone signal as the basis, calculating the near-field speech shielding matrix and the far-field speech shielding matrix, and by fusing the bone conduction microphone signal with the traditional microphone array signal, the near-field speech and the far-field speech can be effectively separated, thereby obtaining the wearer's voice and the far-field voice.
Brief Description of the Drawings
[0041] Figure 1 A schematic diagram of an application scenario of a far-field and near-field speech separation method is provided for an embodiment of the present invention;
[0042] Figure 2 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention;
[0043] Figure 3 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention;
[0044] Figure 4 A schematic diagram of the structure of a far-field and near-field speech separation device provided by an embodiment of the present invention;
[0045] Figure 5 A schematic diagram of the structure of a wearable device provided in an embodiment of the present invention. [Specific implementation method]
[0046] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.
[0047] Figure 1 A schematic diagram of an application scenario of a far-field and near-field speech separation method is provided for an embodiment of the present invention. Figure 1 As shown, this method can be applied to wearable devices such as TWS and AR glasses. AR glasses are used as an example for explanation. Generally, AR glasses include multiple first microphones, and the wearer wears AR glasses. When the controller inside the AR glasses receives the separation instruction, the near-field beamforming signal and the far-field beamforming signal are obtained according to the microphone array signal and the beamforming algorithm collected by the first microphone in the target scene; the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are used to calculate the near-field speech masking matrix and the far-field speech masking matrix; the near-field speech is output according to the near-field beamforming signal and the near-field speech masking matrix; the far-field speech is output according to the far-field beamforming signal and the far-field speech masking matrix. Therefore, AR glasses can well separate near-field speech and far-field speech, thereby eliminating far-field speech other than the wearer, providing the wearer with a clear call experience, and improving the wearer's experience of using AR glasses.
[0048] Figure 2 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the microphone array signal collected by the first microphone is subjected to a near-field and far-field beamforming algorithm to obtain a near-field beam signal and a far-field beam signal; the near-field beam signal, the far-field beam signal and the bone conduction microphone signal are subjected to a near-field and far-field separation algorithm to obtain a masking matrix corresponding to the near-field and far-field speech; the near-field beam signal and the far-field beam signal are respectively fused with the corresponding masking matrices, and the near-field and far-field beam signals are multiplied with the corresponding masking matrices to obtain the separated near-field speech and far-field speech.
[0049] Figure 3 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the method includes:
[0050] An embodiment of the present invention provides a near-field and far-field speech separation method. The near-field and far-field speech separation method is applied to AR glasses as an example for explanation. The AR glasses include five first microphones, all of which are traditional microphones. Four of the first microphones are respectively located on two temples of the AR glasses. Figure 1 1 to 5 in the figure; another first microphone is located on the nose pad of the AR glasses; the AR glasses also include a second microphone, which is also located on the nose pad of the AR glasses, see Figure 1 This layout can make better use of the space on the glasses, and there are microphone distances with smaller and larger spacing, providing more flexibility for the beamforming algorithm.
[0051] The target scenario may be that the wearer wears AR glasses and talks with other people. In this target scenario, the wearer's voice is near-field voice, and the voices of other people are far-field voice. By executing the far-field and near-field voice separation method, the voices of the wearer and other people can be effectively separated, thereby improving the wearer's experience of using AR glasses.
[0052] S110, acquiring a near-field beamforming signal and a far-field beamforming signal according to a microphone array signal collected by a first microphone in a target scene and a beamforming algorithm;
[0053] The microphone array signal collected by the first microphone in the target scene is obtained. The microphone array signal includes the audio signals collected by the above five first microphones, including both the audio of the wearer speaking and the audio of other people speaking nearby. After obtaining the microphone array signal, a near-field beamforming signal and a far-field beamforming signal can be obtained through a beamforming algorithm.
[0054] Among them, the beamforming algorithm is a technology used for signal processing. The core concept of the beamforming algorithm is to form a beam with a specific directionality by weighted summing the microphone array signals received by each first microphone in the AR glasses. This beam can enhance the signal strength in the desired direction while suppressing interference signals in other directions.
[0055] For example, when using a beamforming algorithm to obtain near-field beamforming signals and far-field beamforming signals, it is first necessary to establish a signal model to distinguish between the near field and the far field; then, a steering vector is constructed to obtain a near-field steering vector and a far-field steering vector; based on the near-field steering vector and the far-field steering vector, a covariance matrix is estimated, and finally, separation is performed based on a minimum variance distortionless response algorithm or a subspace method to obtain a near-field beamforming signal and a far-field beamforming signal.
[0056] S120, calculating a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal;
[0057] After obtaining the near-field beamforming signal and the far-field beamforming signal in the above steps, the bone conduction microphone signal is combined and the near-field speech masking matrix is calculated according to the relationship between the near-field beamforming signal and the bone conduction microphone signal; the far-field speech masking matrix is calculated according to the relationship between the far-field beamforming signal and the bone conduction microphone signal.
[0058] In near-field speech processing, near-field speech refers to speech signals from speakers that are relatively close to the microphone (usually within tens of centimeters). Since the near-field environment is relatively complex, there may be various background noises and interference signals; the near-field speech masking matrix is used to describe the relationship between speech signals and these interference signals, especially in the process of separating or enhancing speech signals.
[0059] In a far-field speech environment, the speech signal source is relatively far away from the microphone array. At this time, the signal propagation process will be interfered by more environmental factors, such as reverberation, long-distance background noise, etc. The far-field speech shielding matrix is mainly used to deal with this complex far-field speech situation, with the purpose of highlighting the target speech signal and suppressing the interference signal.
[0060] S130, outputting near-field speech according to the near-field beamforming signal and the near-field speech masking matrix;
[0061] S140: Output the far-field speech according to the far-field beamforming signal and the far-field speech masking matrix.
[0062] After the near-field speech masking matrix is calculated according to the above method, the near-field beamforming signal is further refined and denoised according to the near-field beamforming signal and the near-field speech masking matrix to obtain the near-field speech; the far-field beamforming signal is further refined and denoised according to the far-field beamforming signal and the far-field speech masking matrix to obtain the far-field speech.
[0063] For example, the near-field beamforming signal is multiplied by the near-field speech masking matrix to obtain the near-field speech; and the far-field beamforming signal is multiplied by the far-field speech masking matrix to obtain the far-field speech.
[0064] An embodiment of the present invention provides a method for separating near-field and far-field speech. The method first obtains a microphone array signal collected by a first microphone in a target scene, and uses a beamforming algorithm to obtain a near-field beamforming signal and a far-field beamforming signal. Based on the near-field beamforming signal and the far-field beamforming signal, a near-field speech shielding matrix and a far-field speech shielding matrix are calculated based on a bone conduction microphone signal. By fusing the bone conduction microphone signal with a traditional microphone array signal, the near-field speech and the far-field speech can be effectively separated, thereby obtaining the wearer's voice and the far-field voice.
[0065] In one implementation, the bone conduction microphone signal is obtained by the following steps:
[0066] Acquire a voice signal of the wearer collected by the second microphone in the target scene;
[0067] The wearer's voice signal is used as the bone conduction microphone signal.
[0068] In an embodiment of the present invention, a voice signal of the wearer in a target scene is collected by a second microphone on the AR glasses. Generally, the second microphone can be located on the nose pads of the AR glasses to collect the wearer's speaking voice at a close distance.
[0069] In actual implementation, after collecting the voice signal of the wearer, the voice signal of the wearer is usually transformed into the frequency domain, and the voice signal of the wearer after being transformed into the frequency domain is used as the bone conduction microphone signal.
[0070] Among them, the second microphone can be a voice accelerometer (VACC) or a voice pickup sensor (VPU). The voice signals collected by the voice accelerometer and the voice pickup sensor can capture the wearer's voice signals through bones, muscles and other tissues, effectively isolating external environmental noise and interfering human voices, thereby improving the accuracy and clarity of voice pickup.
[0071] In some embodiments, the calculating, according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal, a near-field speech masking matrix and a far-field speech masking matrix comprises:
[0072] The near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are input into a separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
[0073] In an embodiment of the present invention, a near-field beamforming signal, a far-field beamforming signal and a bone conduction microphone signal are input into a separation neural network, and the separation neural network can output a near-field speech masking matrix and a far-field speech masking matrix.
[0074] As an implementation mode, the separation neural network is a convolutional recurrent neural network, and the convolutional recurrent neural network includes a convolutional neural network, a recurrent neural network, and a deep neural network. The near-field beamforming signal, the far-field beamforming signal, and the bone conduction microphone signal are input into the separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix, including:
[0075] Inputting the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal into the convolutional neural network to extract near-field audio features and far-field audio features;
[0076] Inputting the near-field audio features and the far-field audio features into the recurrent neural network to extract near-field voice information and far-field voice information;
[0077] The near-field speech information and the far-field speech information are input into the deep neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
[0078] In the embodiment of the present invention, the separation neural network is a convolutional recurrent neural network, which is composed of three parts: a convolutional neural network (CNN), a recurrent neural network (RNN), and a deep neural network (DNN). The steps for obtaining the near-field voice shielding matrix and the far-field voice shielding matrix are as follows:
[0079] First, the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are combined into input features. The input features can be expressed as Fea∈C T×F×C Indicates, where T represents the number of signal frames, F represents the number of frequency bands of the signal, and C represents the number of input channels. In the embodiment of the present invention, C=3, indicating that the wearable device includes three channels: near field, far field and bone conduction.
[0080] The input features are then input into a convolutional neural network, which performs feature extraction to extract near-field audio features and far-field audio features. The near-field audio features and far-field audio features are then input into a recurrent neural network, which can model the time series of speech signals to extract near-field speech information and far-field speech information. Finally, the near-field speech information and far-field speech information are input into a deep neural network to obtain near-field speech masking matrices and far-field speech masking matrices.
[0081] As an implementation manner, the near-field beamforming signal, the far-field beamforming signal, and the bone conduction microphone signal are input into the convolutional neural network to extract near-field audio features and far-field audio features, and the previous steps include:
[0082] If the effective bandwidth of the bone conduction microphone signal is within a preset frequency range, the number of convolution kernels of the convolutional neural network will be increased and / or the stride of the convolutional neural network will be reduced;
[0083] If the effective bandwidth of the bone conduction microphone signal is outside the preset frequency range, the number of convolution kernels of the convolutional neural network is reduced and / or the stride of the convolutional neural network is increased.
[0084] In an embodiment of the present invention, before the bone conduction microphone signal is input into the convolutional neural network, the relationship between the effective bandwidth of the bone conduction microphone signal and the preset frequency range is compared. If the effective bandwidth of the bone conduction microphone signal is within the preset frequency domain range, the number of convolution kernels of the convolutional neural network is increased, and / or the stride of the convolutional neural network is reduced; if the effective bandwidth of the bone conduction microphone signal is outside the preset frequency domain range, the number of convolution kernels of the convolutional neural network is reduced, and / or the stride of the convolutional neural network is increased.
[0085] For example, the preset frequency range is 0 to 2KHz. Combined with the frequency response characteristics of the VACC / VPU signal, different network structures can be used for different frequency bands to effectively utilize the VACC / VPU signal while reducing the overall complexity. For example, the effective bandwidth of VACC is only about 2KHz. For frequency bands below 2KHz, the convolutional neural network can use more convolution kernels and a smaller stride, and better utilize VACC for far-field and near-field speech separation through refined modeling; for frequency bands above 2KHz, fewer convolution kernels and larger strides can be used, thereby reducing the overall complexity of the convolutional neural network.
[0086] As an example, according to the microphone array signal collected by the first microphone in the target scene and the beamforming algorithm, the near-field beamforming signal and the far-field beamforming signal are obtained, and the calculation formula is as follows:
[0087]
[0088] Among them, N t,f represents the near-field beamforming signal, F t,f represents the far-field beamforming signal, S i,t,f The time domain signal S represents the signal of the i-th microphone array i The frequency domain signal of the fth frequency point of the tth speech frame after STFT transformation is: represents the first filter weight, represents the second filter weight, i, t, and f are all positive integers.
[0089] It should be noted that the first filter weight and the second filter weight are determined according to specific layout information of the microphone.
[0090] As an example, the near-field speech is output according to the near-field beamforming signal and the near-field speech masking matrix, and the calculation formula is as follows:
[0091]
[0092] Among them, N t ' ,f represents the near-field speech, Nt,f represents the near-field beamforming signal, represents the near-field speech masking matrix, and t and f are both positive integers.
[0093] As an example, the far-field speech is output according to the far-field speech masking matrix and the far-field beamforming signal, and the calculation formula is as follows:
[0094]
[0095] Among them, F t ' ,f represents the far-field speech, F t,f represents the far-field beamforming signal, represents the far-field speech masking matrix, and t and f are both positive integers.
[0096] Figure 4 A schematic diagram of the structure of a far-field and near-field speech separation device provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the device comprises:
[0097] The acquisition module 210 is used to acquire a near-field beamforming signal and a far-field beamforming signal according to the microphone array signal and the beamforming algorithm collected by the first microphone in the target scene;
[0098] A calculation module 220, configured to calculate a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal;
[0099] A first output module 230, configured to output near-field speech according to the near-field beamforming signal and the near-field speech masking matrix;
[0100] The second output module 240 outputs the far-field speech according to the far-field beamforming signal and the far-field speech masking matrix.
[0101] This embodiment is a device embodiment corresponding to the above method, and its implementation process is the same as that of the above method embodiment. For details, please refer to the above method embodiment, and this device embodiment will not elaborate on it.
[0102] Figure 5 A structural diagram of a wearable device provided by an embodiment of the present invention is shown in FIG. Figure 5 As shown, the wearable device 300 includes the above-mentioned far-field and near-field speech separation device 200. The wearable device includes TWS, AR glasses, etc.
[0103] The above description is only an implementation mode of the present invention. It should be pointed out that, for ordinary technicians in this field, improvements can be made without departing from the creative concept of the present invention, but these all belong to the protection scope of the present invention.
Claims
1. A method for separating far-field and near-field speech, characterized in that: include: Acquire a near-field beamforming signal and a far-field beamforming signal according to a microphone array signal collected by a first microphone in a target scene and a beamforming algorithm; Calculating a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal; Outputting near-field speech according to the near-field beamforming signal and the near-field speech masking matrix; The far-field speech is outputted according to the far-field beamforming signal and the far-field speech masking matrix.
2. The far-field and near-field speech separation method according to claim 1, characterized in that: The bone conduction microphone signal is obtained by the following steps: Acquire a voice signal of the wearer collected by the second microphone in the target scene; The wearer's voice signal is used as the bone conduction microphone signal.
3. The far-field and near-field speech separation method according to claim 1, characterized in that: The calculating, according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal, a near-field speech masking matrix and a far-field speech masking matrix comprises: The near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are input into a separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
4. The far-field and near-field speech separation method according to claim 3, characterized in that: The separation neural network is a convolutional recurrent neural network, which includes a convolutional neural network, a recurrent neural network and a deep neural network. The near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal are input into the separation neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix, including: Inputting the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal into the convolutional neural network to extract near-field audio features and far-field audio features; Inputting the near-field audio features and the far-field audio features into the recurrent neural network to extract near-field voice information and far-field voice information; The near-field speech information and the far-field speech information are input into the deep neural network to obtain the near-field speech masking matrix and the far-field speech masking matrix.
5. The method for separating far-field and near-field speech according to claim 4, characterized in that: The step of inputting the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal into the convolutional neural network to extract near-field audio features and far-field audio features comprises: If the effective bandwidth of the bone conduction microphone signal is within a preset frequency range, the number of convolution kernels of the convolutional neural network will be increased and / or the stride of the convolutional neural network will be reduced; If the effective bandwidth of the bone conduction microphone signal is outside the preset frequency range, the number of convolution kernels of the convolutional neural network is reduced and / or the stride of the convolutional neural network is increased.
6. The far-field and near-field speech separation method according to claim 1, characterized in that: According to the microphone array signal collected by the first microphone in the target scene and the beamforming algorithm, the near-field beamforming signal and the far-field beamforming signal are obtained, and the calculation formula is as follows: Among them, N t,f represents the near-field beamforming signal, F t,f represents the far-field beamforming signal, S i,t,f The time domain signal S represents the signal of the i-th microphone array i The frequency domain signal of the fth frequency point of the tth speech frame after SIFT transformation, represents the first filter weight, represents the second filter weight, i, t, and f are all positive integers.
7. The far-field and near-field speech separation method according to claim 1, characterized in that: The near-field speech is output according to the near-field beamforming signal and the near-field speech masking matrix, and the calculation formula is as follows: Among them, N t ' ,f represents the near-field speech, N t,f represents the near-field beamforming signal, represents the near-field speech masking matrix, and t and f are both positive integers.
8. The far-field and near-field speech separation method according to claim 1, characterized in that: The far-field speech is output according to the far-field beamforming signal and the far-field speech masking matrix, and the calculation formula is as follows: Among them, F t ' ,f represents the far-field speech, F t,f represents the far-field beamforming signal, represents the far-field speech masking matrix, and t and f are both positive integers.
9. A far-field and near-field speech separation device, characterized in that: include: An acquisition module, used to acquire a near-field beamforming signal and a far-field beamforming signal according to a microphone array signal acquired by a first microphone in a target scene and a beamforming algorithm; A calculation module, configured to calculate a near-field speech masking matrix and a far-field speech masking matrix according to the near-field beamforming signal, the far-field beamforming signal and the bone conduction microphone signal; A first output module, configured to output near-field speech according to the near-field beamforming signal and the near-field speech masking matrix; The second output module is used to output the far-field speech according to the far-field beamforming signal and the far-field speech masking matrix.
10. A wearable device, characterized in that: It comprises the far-field and near-field speech separation device as described in claim 9.