Far and near field voice separation method and apparatus, and wearable device
Through the beamforming algorithm and differentiating feature calculation of microphone array signals, the separation of near-field voice and far-field voice is achieved, the problem of poor separation effect in the prior art is solved, and the signal-to-noise ratio and user experience of voice applications are improved.
Patent Information
- Application Number
- CN202510304411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art cannot completely separate near-field voice and far-field voice, resulting in improved signal-to-noise ratio but poor separation effect in application scenarios such as voice calls, voice assistants and far-field translation.
Through the acquisition and application of microphone array signal, the near-field beamforming signal and the application of beamforming algorithms, the near-field beamforming signal and the far-field beamforming signal are obtained, the distinctive characteristics are calculated, the near-field speech probability of the microphone array signal is determined, and the near-field signal and the far-field signal are output based on this probability.
Effective separation of near-field voice and far-field voice is achieved, the clarity and signal-to-noise ratio of the wearer's voice is improved, and the performance of wearable devices in voice applications is enhanced.
Smart Images

Figure CN120108414A_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of speech processing technology, and in particular to a far-field and near-field speech separation method, device and wearable device. [Background technology]
[0002] With the gradual popularity of wearable devices such as True Wireless Stereo (TWS) and Augmented Reality (AR) glasses, voice applications based on wearable devices are becoming more and more widespread. In application scenarios such as voice calls, voice assistants, and far-field translation, the technology to distinguish between the wearer's speech and other people's speech is crucial.
[0003] In terms of voice calls, wearable devices need to be able to eliminate the voices of background speakers other than the wearer to provide a clear calling experience; for voice assistant functions, wearable devices need to be able to recognize the wearer's voice commands and provide responses; for far-field translation scenarios, wearable devices need to be able to distinguish between the wearer's voice and the voice of far-field speakers, and then be able to translate only the far-field voice.
[0004] The common solution in the prior art is to use directional beamforming technology to pick up far-field or near-field voice inputs to related applications based on the multi-microphone audio signals picked up by the microphone array. However, although this solution can improve the signal-to-noise ratio of near-field or far-field voice to a certain extent, it still cannot completely separate near-field voice from far-field voice.
[0005] Therefore, it is necessary to provide a far-field and near-field speech separation method, apparatus and wearable device. [Summary of the invention]
[0006] The purpose of the embodiments of the present invention is to provide a far-field and near-field speech separation method, apparatus and wearable device.
[0007] The technical solution of the present invention is as follows:
[0008] In a first aspect, an embodiment of the present invention provides a method for far-field and near-field speech separation, comprising:
[0009] Acquire near-field beamforming signals and far-field beamforming signals based on microphone array signals and beamforming algorithms collected in the target scene;
[0010] calculating a distinguishing feature between the near-field beamformed signal and the far-field beamformed signal;
[0011] Obtaining a near-field speech probability of the microphone array signal according to the distinguishing feature;
[0012] A near-field signal and a far-field signal are output according to the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal.
[0013] Furthermore, the step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature comprises:
[0014] Determining the near-field speech probability of each distinguishing feature according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold;
[0015] The near-field speech probability of the microphone array signal is calculated according to the near-field speech probability of each distinguishing feature.
[0016] Furthermore, the near-field speech probability of each distinguishing feature is determined according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold, and is calculated by the following formula:
[0017]
[0018] Among them, P j represents the near-field speech probability of the jth discriminative feature, D j represents the jth discriminative feature, thH j represents the first threshold corresponding to the jth distinguishing feature, thL j represents the second threshold corresponding to the j-th distinguishing feature, where j is a positive integer.
[0019] Further, the near-field speech probability according to the near-field beamforming signal, the far-field beamforming signal and the microphone array signal is outputted as a near-field signal and a far-field signal, which are calculated by the following formula:
[0020] S n ' ear,f1 =S near,f1 *P;
[0021] S′ far,f1 =S far,f1 *(1-P);
[0022] Among them, f1 represents the sampling frequency band, S′ near,f1 represents the near-field signal corresponding to the sampling frequency band, S′ far,f1 represents the far-field signal corresponding to the sampling frequency band, S near,f1 represents the near-field beamforming signal corresponding to the sampling frequency band, S far,f1 represents the far-field beamforming signal corresponding to the sampling frequency band, and P represents the near-field speech probability of the microphone array signal.
[0023] Further, the step of calculating the distinguishing feature between the near-field beamforming signal and the far-field beamforming signal comprises:
[0024] Calculating an energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point;
[0025] A distinguishing feature of each sampling frequency point is obtained according to the near-field beamforming signal, the far-field beamforming signal and the capability difference signal.
[0026] Furthermore, the step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature comprises:
[0027] The distinguishing features of each sampling frequency point are spliced and input into the neural network to obtain the near-field speech probability of the microphone array signal.
[0028] Further, the near-field speech probability according to the near-field beamforming signal, the far-field beamforming signal and the microphone array signal is outputted as a near-field signal and a far-field signal, which are calculated by the following formula:
[0029] S n ' ear,f =S near,f *q f ;
[0030] S′ far,f =S far,f *(1-q f );
[0031] Among them, S′ near,f represents the near-field signal corresponding to the f-th sampling frequency, S′ far,f represents the far-field signal corresponding to the fth sampling frequency, S near,f represents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, q f represents the near-field speech probability corresponding to the f-th sampling frequency.
[0032] Furthermore, the microphone array signal is collected by microphones presenting a small triangle layout, and the near-field beamforming signal and the far-field beamforming signal are obtained according to the microphone array signal collected in the target scene and the beamforming algorithm, and are calculated by the following formula:
[0033]
[0034] Where f represents the sampling frequency, S near,frepresents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, i represents the i-th microphone array signal, and w near,i,f represents the first filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, w far,i,f represents the second filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, M i,f It represents the frequency domain signal at the fth sampling frequency after the time domain signal corresponding to the i-th microphone array signal passes through STFT.
[0035] In a second aspect, an embodiment of the present invention provides a far-field and near-field speech separation device, comprising:
[0036] An acquisition module, used to acquire near-field beamforming signals and far-field beamforming signals according to microphone array signals and beamforming algorithms acquired in a target scene;
[0037] A feature module, configured to calculate a distinguishing feature between the near-field beamforming signal and the far-field beamforming signal;
[0038] A probability module, used to obtain the near-field speech probability of the microphone array signal according to the distinguishing feature;
[0039] A separation module is used to output a near-field signal and a far-field signal according to the near-field beamforming signal, the far-field beamforming signal and the near-field speech probability of the microphone array signal.
[0040] In a third aspect, an embodiment of the present invention provides a wearable device, comprising a far-field and near-field speech separation device as provided in the first aspect.
[0041] The beneficial effects of the present invention are: obtaining microphone array signals collected in a target scene, using a beamforming algorithm to obtain near-field beamforming signals and far-field beamforming signals; calculating the distinguishing features between the near-field beamforming signals and the far-field beamforming signals, and calculating the near-field speech probability of the microphone array signal based on the distinguishing features; and finally outputting the near-field speech and far-field speech based on the near-field speech probability. The embodiment of the present invention can effectively separate the near-field speech and the far-field speech through the microphone array signal, thereby obtaining the wearer's voice and the far-field voice.
Brief Description of the Drawings
[0042] Figure 1 A schematic diagram of an application scenario of a far-field and near-field speech separation method provided by an embodiment of the present invention
[0043] Figure 2 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of the structure of a far-field and near-field speech separation device provided by an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of the structure of a wearable device provided in an embodiment of the present invention. [Specific implementation method]
[0046] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.
[0047] Figure 1 A schematic diagram of an application scenario of a far-field and near-field speech separation method is provided for an embodiment of the present invention. Figure 1 As shown, this method can be applied to wearable devices such as TWS and AR glasses. AR glasses are used as an example for explanation. Generally, AR glasses include multiple microphones, and the wearer wears AR glasses. When the controller inside the AR glasses receives the separation instruction, it obtains the near-field beamforming signal and the far-field beamforming signal according to the microphone array signal and the beamforming algorithm collected in the target scene; calculates the distinguishing features between the near-field beamforming signal and the far-field beamforming signal; obtains the near-field speech probability of the microphone array signal according to the distinguishing features; outputs the near-field signal and the far-field signal according to the near-field beamforming signal, the far-field beamforming signal and the near-field speech probability. Therefore, AR glasses can well separate the near-field speech and the far-field speech, thereby eliminating the far-field speech other than the wearer, providing the wearer with a clear call experience, and improving the wearer's experience of using AR glasses.
[0048] Figure 2 A flowchart of a far-field and near-field speech separation method provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the method includes:
[0049] An embodiment of the present invention provides a near-field and far-field speech separation method. The near-field and far-field speech separation method is applied to AR glasses as an example for explanation. The AR glasses include three microphones, all of which are traditional microphones. The three microphones are all located on one of the legs of the AR glasses, and the three microphones form a small triangle layout. Figure 1 As shown, numbers 1 to 3 in the figure all represent traditional microphones.
[0050] The target scenario may be that the wearer wears AR glasses and talks with other people. In this target scenario, the wearer's voice is near-field voice, and the voices of other people are far-field voice. By executing the far-field and near-field voice separation method, the voices of the wearer and other people can be effectively separated, thereby improving the wearer's experience of using AR glasses.
[0051] S110, acquiring a near-field beamforming signal and a far-field beamforming signal according to the microphone array signal and the beamforming algorithm collected in the target scene;
[0052] The microphone array signal collected by the microphone in the target scene is obtained. The microphone array signal includes the audio signals collected by the above three microphones, including both the audio of the wearer speaking and the audio of other people speaking nearby. After obtaining the microphone array signal, the near-field beamforming signal and the far-field beamforming signal can be obtained through the beamforming algorithm.
[0053] Among them, the beamforming algorithm is a technology used for signal processing. The core concept of the beamforming algorithm is to form a beam with a specific directionality by weighted summing the microphone array signals received by each microphone in the AR glasses. This beam can enhance the signal strength in the desired direction while suppressing interference signals in other directions.
[0054] For example, when using a beamforming algorithm to obtain near-field beamforming signals and far-field beamforming signals, it is first necessary to establish a signal model to distinguish between the near field and the far field; then, a steering vector is constructed to obtain a near-field steering vector and a far-field steering vector; based on the near-field steering vector and the far-field steering vector, a covariance matrix is estimated, and finally, separation is performed based on a minimum variance distortionless response algorithm or a subspace method to obtain a near-field beamforming signal and a far-field beamforming signal.
[0055] S120, calculating a distinguishing feature between the near-field beamforming signal and the far-field beamforming signal;
[0056] After obtaining the near-field beamforming signal and the far-field beamforming signal in the above steps, the frame-level distinguishing features are calculated according to the voice characteristics of the wearer and the non-wearer nearby. The distinguishing features can be any features that can distinguish the near-field beamforming signal from the far-field beamforming signal, such as the energy difference between the near-field and far-field and the correlation between the near-field and far-field signals.
[0057] The distinguishing feature can be the energy difference between the far and near fields. Since the microphone is close to the human mouth, the wearer's voice often has greater energy than the non-wearer's voice. After processing based on the above-mentioned beamforming algorithm, the energy difference between the far and near fields is further amplified. Therefore, the far and near field energy difference between the near-field beamforming signal and the far-field beamforming signal can effectively judge the far and near field voice.
[0058] The distinguishing feature can be the correlation between far-field and near-field signals. When the wearer speaks, due to the large energy of the wearer's voice, after being processed by the beamforming algorithm, in addition to the near-field output, there will also be a certain amount of the wearer's voice in the far-field beam output signal. At this time, the correlation between far-field and near-field signals is relatively large; when the non-wearer speaks, there is little non-wearer voice in the near-field output signal, and the far-field output signal contains non-wearer voice. At this time, the correlation between far-field and near-field is relatively small. Therefore, the wearer's and non-wearer's voices can be distinguished based on the size of the far-field and near-field correlations.
[0059] It should be noted that the above discriminative features can be calculated based on the full frequency band of the microphone array signal, or a fixed frequency band (such as 300Hz-4000Hz) can be selected for feature calculation according to the frequency range of the microphone array signal and the degree of discrimination of the features to reduce the complexity of the algorithm.
[0060] S130, obtaining a near-field speech probability of the microphone array signal according to the distinguishing feature;
[0061] After calculating the discriminative characteristics, the near-field speech probability of the microphone array signal is determined according to the size of the discriminative features. For example, based on the discriminative features, a fixed threshold can be set. When the discriminative features are greater than the threshold, the near-field speech probability is determined to be P, otherwise the far-field speech probability is determined to be 1-P.
[0062] It should be noted that the AR glasses in the embodiment of the present invention include three microphones, so each microphone collects a microphone array signal, and each microphone array signal is processed by a beamforming algorithm to obtain corresponding near-field beamforming signals and far-field beamforming signals, so that each microphone array signal corresponds to a distinguishing feature.
[0063] According to each distinguishing feature, the near-field speech probability corresponding to each microphone array signal can be multiplied to determine the near-field speech probability of the entire microphone array signal. The near-field speech probability of the microphone array signal is the near-field speech probability of the speech audio collected by the AR glasses.
[0064] S140: Output a near-field signal and a far-field signal according to the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability.
[0065] The embodiment of the present invention further refines the near-field beamforming signal and the near-field speech probability, removes the noise in the near-field beamforming signal, and obtains the near-field signal; and further refines the far-field beamforming signal and the far-field speech probability, removes the noise in the far-field beamforming signal, and obtains the far-field signal.
[0066] The embodiment of the present invention provides a method for separating near-field and far-field speech. The method first obtains the microphone array signal collected in the target scene, and uses the beamforming algorithm to obtain the near-field beamforming signal and the far-field beamforming signal; calculates the distinguishing features between the near-field beamforming signal and the far-field beamforming signal, and calculates the near-field speech probability of the microphone array signal based on the distinguishing features; finally, based on the near-field speech probability, outputs the near-field speech and the far-field speech. The embodiment of the present invention can effectively separate the near-field speech and the far-field speech through the microphone array signal, thereby obtaining the wearer's voice and the far-field voice.
[0067] In one implementation, the step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature includes:
[0068] Determining the near-field speech probability of each distinguishing feature according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold;
[0069] The near-field speech probability of the microphone array signal is calculated according to the near-field speech probability of each distinguishing feature.
[0070] In an embodiment of the present invention, when the distinguishing feature is the correlation between far-field and near-field signals, a first threshold and a second threshold are set for each distinguishing feature, respectively, wherein the first / second thresholds corresponding to different distinguishing features may be the same or different, and may be specifically determined based on actual conditions, and the embodiment of the present invention does not make any specific limitation on this.
[0071] According to the size relationship between each distinguishing feature and the first threshold and the second threshold, the near-field speech probability of each distinguishing feature is determined, and the near-field speech probability of each distinguishing feature is multiplied to obtain the near-field speech probability of the speech audio collected by the AR glasses.
[0072] As an example, the near-field speech probability of each distinguishing feature is determined according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold, and is calculated by the following formula:
[0073]
[0074] Among them, P j represents the near-field speech probability of the jth discriminative feature, D j represents the jth discriminative feature, thH j represents the first threshold corresponding to the jth distinguishing feature, thL j represents the second threshold corresponding to the j-th discriminative feature, where j is 1, 2, or 3.
[0075] As an example, the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal are outputted as a near-field signal and a far-field signal, which are calculated by the following formula:
[0076] S n ' ear,f1 =S near,f1 *P;
[0077] S′ far,f1 =S far,f1 *(1-P);
[0078] Among them, f1 represents the sampling frequency band, S n ' ear,f1 represents the near-field signal corresponding to the sampling frequency band, S′ far,f1 represents the far-field signal corresponding to the sampling frequency band, S near,f1 represents the near-field beamforming signal corresponding to the sampling frequency band, S far,f1 represents the far-field beamforming signal corresponding to the sampling frequency band, and P represents the near-field speech probability of the microphone array signal.
[0079] It should be noted that f1 may be the entire frequency band of the microphone array signal or a fixed frequency band therein, which may be determined based on actual conditions and is not specifically limited in the embodiment of the present invention.
[0080] In some embodiments, the step of calculating the distinguishing feature between the near-field beamforming signal and the far-field beamforming signal comprises:
[0081] Calculating an energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point;
[0082] A distinguishing feature of each sampling frequency point is obtained according to the near-field beamforming signal, the far-field beamforming signal and the capability difference signal.
[0083] In an embodiment of the present invention, when the distinguishing feature is an energy difference signal, multiple sampling frequencies are set, and the energy difference signal of the near-field beamforming signal and the far-field beamforming signal at each sampling frequency is calculated; the near-field beamforming signal, the far-field beamforming signal and the energy difference signal at each sampling frequency are used as the distinguishing features at each sampling frequency.
[0084] As an implementation manner, the step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature includes:
[0085] The distinguishing features of each sampling frequency point are spliced and input into the neural network to obtain the near-field speech probability of the microphone array signal.
[0086] After calculating the discriminative features of each sampling frequency point, the discriminative features of each sampling frequency point are spliced together again and input into the neural network to obtain the near-field speech probability of the speech audio collected by the entire AR glasses.
[0087] For example, the distinguishing feature of the fth sampling frequency point can be expressed as (S near,f ) 2 -(S far,f ) 2 , the near-field beamforming signal at the fth sampling frequency can be expressed as S near,f , the far-field beamforming signal at the fth sampling frequency can be expressed as S far,f , the distinguishing features of all sampling frequency points are spliced together as the input feature vector of the neural network. After the neural network, the wearer's voice probability can be expressed as:
[0088] Q = Net(K);
[0089] in, f represents the fth sampling frequency point, F is the total number of sampling frequencies, and K is the input feature vector, that is, the vector obtained by concatenating the distinguishing features of all sampling frequencies.
[0090] As an implementation manner, the near-field speech probability of the near-field beamforming signal, the far-field beamforming signal, and the microphone array signal is outputted as a near-field signal and a far-field signal, which are calculated by the following formula:
[0091] S n ' ear,f =S near,f *q f ;
[0092] S′ far,f =S far,f *(1-q f );
[0093] Among them, S n ' ear,f represents the near-field signal corresponding to the f-th sampling frequency, S′ far,f represents the far-field signal corresponding to the fth sampling frequency, S near,f represents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, q f represents the near-field speech probability corresponding to the f-th sampling frequency.
[0094] As an implementation mode, the microphone array signal is collected by a microphone presenting a small triangle layout, and the near-field beamforming signal and the far-field beamforming signal are obtained according to the microphone array signal collected in the target scene and the beamforming algorithm, and are calculated by the following formula:
[0095]
[0096] Where f represents the sampling frequency, S near,f represents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, i represents the i-th microphone array signal, and w near,i,f represents the first filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, w far,i,f represents the second filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, M i,f It represents the frequency domain signal at the fth sampling frequency after the time domain signal corresponding to the i-th microphone array signal passes through STFT.
[0097] Figure 3 A schematic diagram of the structure of a far-field and near-field speech separation device provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the device comprises:
[0098] The acquisition module 210 is used to acquire near-field beamforming signals and far-field beamforming signals according to the microphone array signals and beamforming algorithms collected in the target scene;
[0099] A feature module 220, configured to calculate a distinguishing feature between the near-field beamforming signal and the far-field beamforming signal;
[0100] A probability module 230, configured to obtain a near-field speech probability of the microphone array signal based on the distinguishing feature;
[0101] The separation module 240 is configured to output a near-field signal and a far-field signal according to the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal.
[0102] This embodiment is a device embodiment corresponding to the above method, and its implementation process is the same as that of the above method embodiment. For details, please refer to the above method embodiment, and this device embodiment will not elaborate on it.
[0103] Figure 4 A structural diagram of a wearable device provided by an embodiment of the present invention is shown in FIG. Figure 4As shown, the wearable device 300 includes the above-mentioned far-field and near-field speech separation device 200. The wearable device includes TWS, AR glasses, etc.
[0104] The above description is only an implementation mode of the present invention. It should be pointed out that, for ordinary technicians in this field, improvements can be made without departing from the creative concept of the present invention, but these all belong to the protection scope of the present invention.
Claims
1. A method for separating far-field and near-field speech, characterized in that: include: Acquire near-field beamforming signals and far-field beamforming signals based on microphone array signals and beamforming algorithms collected in the target scene; calculating a distinguishing feature between the near-field beamformed signal and the far-field beamformed signal; Obtaining a near-field speech probability of the microphone array signal according to the distinguishing feature; A near-field signal and a far-field signal are output according to the near-field beamforming signal, the far-field beamforming signal, and the near-field speech probability of the microphone array signal.
2. The far-field and near-field speech separation method according to claim 1, characterized in that: The step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature comprises: Determining the near-field speech probability of each distinguishing feature according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold; The near-field speech probability of the microphone array signal is calculated according to the near-field speech probability of each distinguishing feature.
3. The far-field and near-field speech separation method according to claim 2, characterized in that: The near-field speech probability of each distinguishing feature is determined according to the magnitude relationship between each distinguishing feature and the first threshold and the second threshold, and is calculated by the following formula: Among them, P j represents the near-field speech probability of the jth discriminative feature, D j represents the jth discriminative feature, thH j represents the first threshold corresponding to the jth distinguishing feature, thL j represents the second threshold corresponding to the j-th distinguishing feature, where j is a positive integer.
4. The far-field and near-field speech separation method according to claim 2, characterized in that: The near-field speech probability according to the near-field beamforming signal, the far-field beamforming signal and the microphone array signal is outputted as a near-field signal and a far-field signal, which is calculated by the following formula: S′ near,f1 =S near,f1 *P; S′ far,f1 =S far,f1 *(1-P); Among them, f1 represents the sampling frequency band, S′ near,f1 represents the near-field signal corresponding to the sampling frequency band, S′ far,f1 represents the far-field signal corresponding to the sampling frequency band, S near,f1 represents the near-field beamforming signal corresponding to the sampling frequency band, S far,f1 represents the far-field beamforming signal corresponding to the sampling frequency band, and P represents the near-field speech probability of the microphone array signal.
5. The far-field and near-field speech separation method according to claim 1, characterized in that: The step of calculating the distinguishing feature between the near-field beamforming signal and the far-field beamforming signal comprises: Calculating an energy difference signal between the near-field beamforming signal and the far-field beamforming signal at each sampling frequency point; A distinguishing feature of each sampling frequency point is obtained according to the near-field beamforming signal, the far-field beamforming signal and the energy difference signal.
6. The far-field and near-field speech separation method according to claim 5, characterized in that: The step of obtaining the near-field speech probability of the microphone array signal according to the distinguishing feature comprises: The distinguishing features of each sampling frequency point are spliced and input into the neural network to obtain the near-field speech probability of the microphone array signal.
7. The far-field and near-field speech separation method according to claim 6, characterized in that: The near-field speech probability according to the near-field beamforming signal, the far-field beamforming signal and the microphone array signal is outputted as a near-field signal and a far-field signal, which is calculated by the following formula: S′ near,f =S near,f *q f ; S′ far,f =S far,f *(1-q f ); Among them, S′ near,f represents the near-field signal corresponding to the f-th sampling frequency, S′ far,f represents the far-field signal corresponding to the fth sampling frequency, S near,f represents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, q f represents the near-field speech probability corresponding to the f-th sampling frequency.
8. The far-field and near-field speech separation method according to claim 1, characterized in that: The microphone array signal is collected by microphones presenting a small triangle layout. The near-field beamforming signal and the far-field beamforming signal are obtained according to the microphone array signal collected in the target scene and the beamforming algorithm, and are calculated by the following formula: Where f represents the sampling frequency, S near,f represents the near-field beamforming signal corresponding to the f-th sampling frequency, S far,f represents the far-field beamforming signal corresponding to the f-th sampling frequency, i represents the i-th microphone array signal, and w near,i,f represents the first filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, w far,i,f represents the second filter weight corresponding to the fth sampling frequency of the i-th microphone array signal, M i,f It represents the frequency domain signal at the fth sampling frequency after the time domain signal corresponding to the i-th microphone array signal passes through STFT.
9. A far-field and near-field speech separation device, characterized in that: include: An acquisition module, used to acquire near-field beamforming signals and far-field beamforming signals according to microphone array signals and beamforming algorithms acquired in a target scene; a feature module, configured to calculate a distinguishing feature between the near-field beamforming signal and the far-field beamforming signal; A probability module, configured to obtain a near-field speech probability of the microphone array signal based on the distinguishing feature; A separation module is used to output a near-field signal and a far-field signal according to the near-field beamforming signal, the far-field beamforming signal and the near-field speech probability of the microphone array signal.
10. A wearable device, characterized in that: It comprises the far-field and near-field speech separation device as described in claim 9.
Citation Information
Cited By
Voice separation method, electronic device, storage medium and computer program product
CN120581022A