Near-far-field speech separation method and apparatus, and wearable device
Patent Information
- Application Number
- PCT/CN2025/082499
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025082499_17092026_PF_FP_ABST
Abstract
Description
Near-field and far-field speech separation methods, devices and wearable devices [Technical Field]
[0001] This application relates to the field of speech processing technology, and in particular to a method, apparatus and wearable device for far-field and near-field speech separation. [Background Technology]
[0002] With the increasing popularity of wearable devices such as True Wireless Stereo (TWS) and Augmented Reality (AR) glasses, voice applications based on wearable devices are becoming more and more widespread. In application scenarios such as voice calls, voice assistants, and far-field translation, the technology to distinguish the wearer's speech from that of other people is crucial.
[0003] In terms of voice calls, wearable devices need to be able to eliminate background speaker voices other than the wearer's to provide a clear call experience; for voice assistant functions, wearable devices need to be able to recognize the wearer's voice commands and provide responses; for far-field translation scenarios, wearable devices need to be able to distinguish between the wearer's voice and the voices of far-field speakers, so as to translate only far-field speech.
[0004] A common approach in existing technologies is to use multi-microphone arrays to pick up audio signals and employ directional beamforming technology to capture far-field or near-field speech input for relevant applications. However, while this approach can improve the signal-to-noise ratio of near-field or far-field speech to some extent, it still cannot completely separate near-field and far-field speech.
[0005] Therefore, it is necessary to provide a method, apparatus, and wearable device for near-field speech separation. [Summary of the Invention]
[0006] The purpose of this application is to provide a method, apparatus, and wearable device for far-field and near-field speech separation.
[0007] The technical solution of this application is as follows:
[0008] In a first aspect, embodiments of this application provide a near-field and far-field speech separation method, including:
[0009] Based on the microphone array signals and beamforming algorithms collected in the target scene, near-field beamforming signals and far-field beamforming signals are obtained.
[0010] Calculate the distinguishing features between the near-field beamforming signal and the far-field beamforming signal;
[0011] Based on the discriminative features, the near-field speech probability of the microphone array signal is obtained;
[0012] Based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal, the near-field signal and the far-field signal are output.
[0013] Further, the step of obtaining the near-field speech probability of the microphone array signal based on the discriminative features includes:
[0014] Based on the relationship between each discriminative feature and the first and second thresholds, the near-field speech probability of each discriminative feature is determined.
[0015] The near-field speech probability of the microphone array signal is calculated based on the near-field speech probability of each distinguishing feature.
[0016] Furthermore, the near-field speech probability of each discriminative feature is determined based on the relationship between each discriminative feature and the first and second thresholds, and is calculated using the following formula:
[0017] Among them, P j D represents the near-field speech probability of the j-th discriminative feature. j Let thH represent the j-th discriminative feature. j Let thL represent the first threshold corresponding to the j-th discriminative feature. j This represents the second threshold corresponding to the j-th discriminative feature, where j is a positive integer.
[0018] Further, the near-field signal and far-field signal are output based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal, and are calculated using the following formula: S′ near,f1 =S near,f1 *P; S′ far,f1 =S far,f1 *(1-P);
[0019] Where f1 represents the sampling frequency band, S′ near,f1 S′ represents the near-field signal corresponding to the sampling frequency band. far,f1 S represents the far-field signal corresponding to the sampling frequency band. near,f1 S represents the near-field beamforming signal corresponding to the sampling frequency band. far,f1 The frequency band to be sampled represents the far-field beamforming signal, and P represents the near-field speech probability of the microphone array signal.
[0020] Further, the step of calculating the distinguishable features between the near-field beamforming signal and the far-field beamforming signal includes:
[0021] Calculate the energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point;
[0022] Based on the near-field beamforming signal, the far-field beamforming signal, and the capability difference signal, the distinguishable features of each sampling frequency point are obtained.
[0023] Further, the step of obtaining the near-field speech probability of the microphone array signal based on the discriminative features includes:
[0024] The discriminative features of each sampling frequency point are concatenated and input into the neural network to obtain the near-field speech probability of the microphone array signal.
[0025] Further, the near-field signal and far-field signal are output based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal, and are calculated using the following formula: S′ near,f =S near,f *q f ;
[0026] S′ far,f =S far,f *(1-q f );
[0027] Among them, S′ near,f S′ represents the near-field signal corresponding to the f-th sampling frequency point. far,f S represents the far-field signal corresponding to the f-th sampling frequency point. near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f q represents the far-field beamforming signal corresponding to the f-th sampling frequency point. f This represents the near-field speech probability corresponding to the f-th sampling frequency point.
[0028] Furthermore, the microphone array signal is acquired by microphones arranged in a small triangle. The near-field beamforming signal and far-field beamforming signal are obtained based on the microphone array signal acquired in the target scene and the beamforming algorithm, calculated using the following formula:
[0029] Where f represents the sampling frequency, S near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f Let represent the far-field beamforming signal corresponding to the f-th sampling frequency point, i represent the signal of the i-th microphone array, and w near,i,f w represents the weight of the first filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. far,i,fM represents the weight of the second filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. i,f This represents the frequency domain signal at the f-th sampling frequency point after the time domain signal corresponding to the i-th microphone array signal is processed by STFT.
[0030] Secondly, embodiments of this application provide a near-field and far-field speech separation device, comprising:
[0031] The acquisition module is used to acquire near-field beamforming signals and far-field beamforming signals based on the microphone array signals and beamforming algorithms collected in the target scene.
[0032] The feature module is used to calculate the distinguishable features between the near-field beamforming signal and the far-field beamforming signal;
[0033] A probability module is used to obtain the near-field speech probability of the microphone array signal based on the discriminative features.
[0034] The separation module is used to output near-field signals and far-field signals based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal.
[0035] Thirdly, embodiments of this application provide a wearable device, including a near-field voice separation device as provided in the first aspect.
[0036] The beneficial effects of this application are as follows: It acquires microphone array signals collected in a target scene, uses beamforming algorithms to obtain near-field beamforming signals and far-field beamforming signals; calculates the distinguishing features between the near-field and far-field beamforming signals, and calculates the near-field speech probability of the microphone array signals based on these distinguishing features; finally, based on the near-field speech probability, it outputs near-field and far-field speech. This application embodiment can effectively separate near-field and far-field speech using microphone array signals, thereby obtaining the wearer's voice and far-field voice. [Attached Image Description]
[0037] Figure 1 is a schematic diagram illustrating an application scenario of a near-field speech separation method provided in an embodiment of this application.
[0038] Figure 2 is a flowchart of a near-field speech separation method provided in an embodiment of this application;
[0039] Figure 3 is a schematic diagram of a near-field speech separation device provided in an embodiment of this application;
[0040] Figure 4 is a structural schematic diagram of a wearable device provided in an embodiment of this application.
Detailed Implementation Methods
[0041] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0042] Figure 1 illustrates an application scenario of a near-field and far-field speech separation method provided in this application. As shown in Figure 1, this method can be applied to wearable devices such as TWS and AR glasses. Taking AR glasses as an example, AR glasses typically include multiple microphones, and the wearer wears the AR glasses. When the controller inside the AR glasses receives a separation command, it acquires near-field beamforming signals and far-field beamforming signals based on the microphone array signals collected in the target scene and the beamforming algorithm; it calculates the discriminative features between the near-field and far-field beamforming signals; based on the discriminative features, it obtains the near-field speech probability of the microphone array signals; and based on the near-field beamforming signals, far-field beamforming signals, and near-field speech probabilities, it outputs near-field and far-field signals. Thus, the AR glasses can effectively separate near-field and far-field speech, thereby eliminating far-field speech outside the wearer's field of vision, providing the wearer with a clear call experience, and improving the wearer's user experience of the AR glasses.
[0043] Figure 2 is a flowchart of a near-field speech separation method provided in an embodiment of this application. As shown in Figure 2, the method includes:
[0044] This application provides a near-field speech separation method. Taking the application of this near-field speech separation method to AR glasses as an example, the AR glasses include three microphones, all of which are traditional microphones. The three microphones are located on one temple of the AR glasses and form a small triangular layout, as shown in Figure 1. In the figure, labels 1 to 3 all represent traditional microphones.
[0045] The target scenario could be a user wearing AR glasses talking to others. In this scenario, the user's voice is near-field speech, while the voices of others are far-field speech. By implementing this near-field and far-field speech separation method, the voices of the user and others can be effectively separated, thereby improving the user's experience with AR glasses.
[0046] S110: Based on the microphone array signals and beamforming algorithms collected in the target scene, obtain near-field beamforming signals and far-field beamforming signals.
[0047] The microphone array signal is acquired from the microphones in the target scene. This microphone array signal includes audio signals collected by the three microphones mentioned above, including both the audio of the wearer speaking and the audio of other people speaking nearby. After acquiring the microphone array signal, a beamforming algorithm is used to obtain near-field beamforming signals and far-field beamforming signals.
[0048] Beamforming is a signal processing technique. The core concept of beamforming is to form a beam with a specific direction by weighted summation of the microphone array signals received by each microphone in the AR glasses. This beam can enhance the signal strength in the desired direction while suppressing interference signals in other directions.
[0049] For example, when using beamforming algorithms to obtain near-field and far-field beamforming signals, it is first necessary to establish a signal model to distinguish between the near and far fields; then, the steering vector is constructed to obtain the near-field steering vector and the far-field steering vector; based on the near-field and far-field steering vectors, the covariance matrix is estimated; finally, separation is performed based on the minimum variance distortionless response algorithm or the subspace method to obtain the near-field and far-field beamforming signals.
[0050] S120, Calculate the distinguishable features between the near-field beamforming signal and the far-field beamforming signal;
[0051] After obtaining the near-field beamforming signal and the far-field beamforming signal in the above steps, frame-level discriminative features are calculated based on the speech characteristics of the wearer and the non-wearer next to them. These discriminative features can be any features that can distinguish the near-field beamforming signal from the far-field beamforming signal, such as the near-field energy difference and the correlation between the near-field and far-field signals.
[0052] The distinguishing feature can be the energy difference between the near and far fields. Since the microphone is close to the mouth, the voice of the wearer is often more energetic than that of the non-wearer. After processing by the beamforming algorithm, the energy difference between the near and far fields is further amplified. Therefore, the energy difference between the near and far fields between the near-field beamforming signal and the far-field beamforming signal can effectively determine the near and far field voice.
[0053] This discriminative feature can be the correlation between near-field and far-field signals. When the wearer speaks, due to the higher energy of their voice, after beamforming processing, the far-field beamout signal will also contain some of the wearer's voice in addition to the near-field output signal, resulting in a relatively high correlation between the near-field and far-field signals. Conversely, when a non-wearer speaks, the near-field output signal contains very little of the non-wearer's voice, while the far-field output signal does contain it, resulting in a relatively low correlation between the near-field and far-field signals. Therefore, the magnitude of the near-field correlation can be used to distinguish between the wearer's and non-wearer's voices.
[0054] It should be noted that the above discriminative features can be calculated based on the full frequency band of the microphone array signal, or a fixed frequency band (such as 300Hz-4000Hz) can be selected for feature calculation according to the frequency range of the microphone array signal and the discriminative power of the features, thereby reducing the complexity of the algorithm.
[0055] S130, based on the discriminative features, obtain the near-field speech probability of the microphone array signal;
[0056] After calculating the discriminative characteristics, the near-field speech probability of the microphone array signal is determined based on the magnitude of the discriminative characteristics. For example, based on the discriminative characteristics, a fixed threshold can be set. When the discriminative characteristics are greater than the threshold, the near-field speech probability is determined to be P; otherwise, the far-field speech probability is determined to be 1-P.
[0057] It should be noted that the AR glasses in this embodiment include three microphones. Therefore, each microphone collects a microphone array signal. Each microphone array signal is processed by a beamforming algorithm to obtain corresponding near-field beamforming signals and far-field beamforming signals. Thus, each microphone array signal corresponds to a distinguishing feature.
[0058] Based on each distinguishing feature, the near-field speech probability corresponding to each microphone array signal can be multiplied to determine the near-field speech probability of the entire microphone array signal. The near-field speech probability of this microphone array signal is the near-field speech probability of the speech audio collected by the AR glasses.
[0059] S140, based on the near-field beamforming signal, the far-field beamforming signal and the near-field speech probability, output the near-field signal and the far-field signal.
[0060] In this embodiment, the near-field beamforming signal and the near-field speech probability are further refined to remove noise from the near-field beamforming signal to obtain a near-field signal; and the far-field beamforming signal and the far-field speech probability are further refined to remove noise from the far-field beamforming signal to obtain a far-field signal.
[0061] This application provides a method for separating near-field and far-field speech. First, microphone array signals collected in a target scene are acquired. Then, beamforming algorithms are used to obtain near-field and far-field beamforming signals. The discriminative features between the near-field and far-field beamforming signals are calculated. Based on these discriminative features, the near-field speech probability of the microphone array signal is calculated. Finally, based on the near-field speech probability, near-field and far-field speech are output. This application effectively separates near-field and far-field speech from the microphone array signal, thereby obtaining the wearer's voice and far-field voice.
[0062] In one implementation, the step of obtaining the near-field speech probability of the microphone array signal based on the discriminative features includes:
[0063] Based on the relationship between each discriminative feature and the first and second thresholds, the near-field speech probability of each discriminative feature is determined.
[0064] The near-field speech probability of the microphone array signal is calculated based on the near-field speech probability of each distinguishing feature.
[0065] In this embodiment of the application, when the distinguishing feature is the correlation between near and far field signals, a first threshold and a second threshold are set for each distinguishing feature. The first and second thresholds corresponding to different distinguishing features can be the same or different, and can be determined according to the actual situation. This embodiment of the application does not make specific limitations on this.
[0066] Based on the relationship between each discriminative feature and the first and second thresholds, the near-field speech probability of each discriminative feature is determined, and the near-field speech probabilities of each discriminative feature are multiplied together to obtain the near-field speech probability of the speech audio collected by the AR glasses.
[0067] As an example, the near-field speech probability of each discriminative feature is determined based on the relationship between each discriminative feature and the first and second thresholds, and is calculated using the following formula:
[0068] Among them, P j D represents the near-field speech probability of the j-th discriminative feature. j Let thH represent the j-th discriminative feature. j Let thL represent the first threshold corresponding to the j-th discriminative feature. j This represents the second threshold corresponding to the j-th discriminative feature, where j is 1, 2, or 3.
[0069] As an example, the output of near-field and far-field signals based on the near-field speech probabilities of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal is calculated using the following formula:
[0070] S′ near,f1 =S near,f1 *P;
[0071] S′ far,f1 =S far,f1 *(1-P);
[0072] Where f1 represents the sampling frequency band, S′ near,f1 S′ represents the near-field signal corresponding to the sampling frequency band. far,f1 S represents the far-field signal corresponding to the sampling frequency band. near,f1 S represents the near-field beamforming signal corresponding to the sampling frequency band. far,f1 The frequency band to be sampled represents the far-field beamforming signal, and P represents the near-field speech probability of the microphone array signal.
[0073] It should be noted that f1 can be the entire frequency band of the microphone array signal or a fixed frequency band within it. The specific frequency band can be determined according to the actual situation, and this application embodiment does not make specific limitations in this regard.
[0074] In some embodiments, the step of calculating the distinguishable features between the near-field beamforming signal and the far-field beamforming signal includes:
[0075] Calculate the energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point;
[0076] Based on the near-field beamforming signal, the far-field beamforming signal, and the capability difference signal, the distinguishable features of each sampling frequency point are obtained.
[0077] In this embodiment of the application, when the distinguishing feature is the energy difference signal, multiple sampling frequency points are set, and the energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point is calculated; the near-field beamforming signal, the far-field beamforming signal, and the energy difference signal at each sampling frequency point are used as the distinguishing feature at each sampling frequency point.
[0078] As one implementation, the step of obtaining the near-field speech probability of the microphone array signal based on the distinguishing features includes:
[0079] The discriminative features of each sampling frequency point are concatenated and input into the neural network to obtain the near-field speech probability of the microphone array signal.
[0080] After calculating the discriminative features of each sampling frequency point, the discriminative features of each sampling frequency point are concatenated together and then input into the neural network to obtain the near-field speech probability of the entire AR glasses' collected speech audio.
[0081] For example, the discriminative feature of the f-th sampling frequency point can be represented as (S near,f ) 2 -(S far,f ) 2 The near-field beamforming signal at the f-th sampling frequency can be expressed as S near,f The far-field beamforming signal at the f-th sampling frequency can be represented as S far,f The discriminative features of all sampling frequency points are concatenated together and used as the input feature vector of the neural network. After passing through the neural network, the wearer's speech probability can be expressed as: Q=Net(K);
[0082] in, f represents the f-th sampling frequency point, F is the total number of sampling frequency points, and K is the input feature vector, which is the vector obtained by concatenating the discriminative features of all sampling frequency points.
[0083] As one implementation, the near-field signal and far-field signal are output based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal, and are calculated using the following formula: S′ near,f =S near,f *q f ; S′ far,f =S far,f *(1-q f );
[0084] Among them, S′ near,f S′ represents the near-field signal corresponding to the f-th sampling frequency point. far,f S represents the far-field signal corresponding to the f-th sampling frequency point. near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f q represents the far-field beamforming signal corresponding to the f-th sampling frequency point. f This represents the near-field speech probability corresponding to the f-th sampling frequency point.
[0085] In one implementation, the microphone array signal is acquired by microphones arranged in a small triangle. The near-field beamforming signal and far-field beamforming signal are obtained based on the microphone array signal acquired in the target scene and a beamforming algorithm, calculated using the following formula:
[0086] Where f represents the sampling frequency, S near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f Let represent the far-field beamforming signal corresponding to the f-th sampling frequency point, i represent the signal of the i-th microphone array, and w near,i,f w represents the weight of the first filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. far,i,f M represents the weight of the second filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. i,f This represents the frequency domain signal at the f-th sampling frequency point after the time domain signal corresponding to the i-th microphone array signal is processed by STFT.
[0087] Figure 3 is a schematic diagram of a near-field speech separation device provided in an embodiment of this application. As shown in Figure 3, the device includes:
[0088] The acquisition module 210 is used to acquire near-field beamforming signals and far-field beamforming signals based on the microphone array signals and beamforming algorithms acquired in the target scene.
[0089] Feature module 220 is used to calculate the distinguishing features between the near-field beamforming signal and the far-field beamforming signal;
[0090] The probability module 230 is used to obtain the near-field speech probability of the microphone array signal based on the discriminative features.
[0091] The separation module 240 is used to output near-field signals and far-field signals based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal and the microphone array signal.
[0092] This embodiment is a device embodiment corresponding to the above method. Its implementation process is the same as that of the above method embodiment. For details, please refer to the above method embodiment. This device embodiment will not be described in detail here.
[0093] Figure 4 is a schematic diagram of a wearable device provided in an embodiment of this application. As shown in Figure 4, the wearable device 300 includes the aforementioned near-field and far-field voice separation device 200. The wearable device includes TWS, AR glasses, etc.
[0094] The above description is merely an embodiment of this application. It should be noted that those skilled in the art can make improvements without departing from the inventive concept of this application, but these improvements all fall within the protection scope of this application.
Claims
1. A method for separating near and far-field speech, characterized in that, include: Based on the microphone array signals and beamforming algorithms collected in the target scene, near-field beamforming signals and far-field beamforming signals are obtained. Calculate the distinguishing features between the near-field beamforming signal and the far-field beamforming signal; Based on the discriminative features, the near-field speech probability of the microphone array signal is obtained; Based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal, the near-field signal and the far-field signal are output.
2. The near-field and far-field speech separation method according to claim 1, characterized in that, The step of obtaining the near-field speech probability of the microphone array signal based on the discriminative features includes: Based on the relationship between each discriminative feature and the first and second thresholds, the near-field speech probability of each discriminative feature is determined. The near-field speech probability of the microphone array signal is calculated based on the near-field speech probability of each distinguishing feature.
3. The near-field and far-field speech separation method according to claim 2, characterized in that, The near-field speech probability of each discriminative feature is determined based on the relationship between each discriminative feature and the first and second thresholds, and is calculated using the following formula: Among them, P j D represents the near-field speech probability of the j-th discriminative feature. j Let thH represent the j-th discriminative feature. j Let thL represent the first threshold corresponding to the j-th discriminative feature. j This represents the second threshold corresponding to the j-th discriminative feature, where j is a positive integer.
4. The near-field and far-field speech separation method according to claim 2, characterized in that, The near-field signal and far-field signal are output based on the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal, and are calculated using the following formula: S′ near,f1 =S near,f1 *P; S′ far,f1 =S far,f1 *(1-P); Where f1 represents the sampling frequency band, S′ near,f1 S′ represents the near-field signal corresponding to the sampling frequency band. far,f1 S represents the far-field signal corresponding to the sampling frequency band. near,f1 S represents the near-field beamforming signal corresponding to the sampling frequency band. far,f1 The frequency band to be sampled represents the far-field beamforming signal, and P represents the near-field speech probability of the microphone array signal.
5. The near-field and far-field speech separation method according to claim 1, characterized in that, The step of calculating the distinguishing features between the near-field beamforming signal and the far-field beamforming signal includes: Calculate the energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point; Based on the near-field beamforming signal, the far-field beamforming signal, and the energy difference signal, the distinguishable features of each sampling frequency point are obtained.
6. The near-field and far-field speech separation method according to claim 5, characterized in that, The step of obtaining the near-field speech probability of the microphone array signal based on the discriminative features includes: The discriminative features of each sampling frequency point are concatenated and input into the neural network to obtain the near-field speech probability of the microphone array signal.
7. The near-field and far-field speech separation method according to claim 6, characterized in that, The near-field signal and far-field signal are output based on the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal, and are calculated using the following formula: S′ near,f =S near,f *q f ; S′ far,f =S far,f *(1-q f ); Among them, S′ near,f S′ represents the near-field signal corresponding to the f-th sampling frequency point. far,f S represents the far-field signal corresponding to the f-th sampling frequency point. near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f q represents the far-field beamforming signal corresponding to the f-th sampling frequency point. f This represents the near-field speech probability corresponding to the f-th sampling frequency point.
8. The near-field and far-field speech separation method according to claim 1, characterized in that, The microphone array signal is acquired by microphones arranged in a small triangle. Based on the microphone array signal acquired in the target scene and the beamforming algorithm, the near-field beamforming signal and far-field beamforming signal are obtained and calculated using the following formula: Where f represents the sampling frequency, S near,f S represents the near-field beamforming signal corresponding to the f-th sampling frequency point. far,f Let represent the far-field beamforming signal corresponding to the f-th sampling frequency point, i represent the signal of the i-th microphone array, and w near,i,f w represents the weight of the first filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. far,i,f M represents the weight of the second filter corresponding to the f-th sampling frequency point of the i-th microphone array signal. i,f This represents the frequency domain signal at the f-th sampling frequency point after the time domain signal corresponding to the i-th microphone array signal is processed by STFT.
9. A near-field and far-field speech separation device, characterized in that, include: The acquisition module is used to acquire near-field beamforming signals and far-field beamforming signals based on the microphone array signals and beamforming algorithms collected in the target scene. The feature module is used to calculate the distinguishable features between the near-field beamforming signal and the far-field beamforming signal; A probability module is used to obtain the near-field speech probability of the microphone array signal based on the discriminative features. The separation module is used to output near-field signals and far-field signals based on the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal.
10. A wearable device, characterized in that, Includes the near-field and far-field speech separation device as described in claim 9.