Voice separation method, electronic device, storage medium and computer program product

Through the combination of microphone array beamforming and separation models, efficient speech separation in resource-constrained devices is achieved, solving the problem of difficulty in deploying DNN models in computing power-constrained devices, and reducing computing and memory requirements.

CN120581022AActive Publication Date: 2025-09-02GOERTEK INC

Patent Information

Application Number
CN202511086198.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-02
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

The existing end-to-end deep neural network (DNN) speech separation model is difficult to effectively deploy in end-side devices with limited computing power, mainly due to the large number of parameters and high computational complexity.

Method used

The beamforming process is performed through the microphone array, and the near-field and far-field speaker signals are initially separated, and the preliminary separation signals are further processed using the preset near-field and far-field separation models to reduce the demand for computing and memory resources.

Benefits of technology

The use of computing and memory resources by the voice separation algorithm is greatly reduced, so that the end-side devices with resource-constrained resources such as wearable devices can effectively deploy the voice separation algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581022A_ABST
    Figure CN120581022A_ABST
Patent Text Reader

Abstract

The invention discloses a voice separation method, electronic equipment, a storage medium and a computer program product, and relates to the technical field of signal processing, and the method comprises the steps: carrying out the beam forming of a microphone array signal through employing a preset near-field speaker direction as a voice enhancement direction, and obtaining a first near-field speaker signal; performing beam forming processing on the microphone array signal by taking a preset far-field speaker direction as a voice enhancement direction to obtain a first far-field speaker signal; inputting the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal into a preset near-field separation model for processing to obtain a second near-field speaker signal; and inputting the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal into a preset far-field separation model for processing to obtain a second far-field speaker signal. According to the invention, the far and near field signals are preliminarily separated, so that the computing power resource consumption of a voice separation algorithm is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of signal processing technology, and in particular to a speech separation method, electronic equipment, storage medium, and computer program product. Background Art

[0002] Speech separation technology can be applied to mixed far-field and near-field audio pickup scenarios in electronic devices. When multiple sound sources (such as far-field and near-field speakers) speak simultaneously, this technology can be used to extract and separate the target speech signal. The current mainstream solution uses end-to-end deep neural networks (DNNs) for speech separation. This technology uses a trained model to separate the far-field and near-field speakers in mixed audio into two independent signals, addressing the problem of mutual interference between sound sources that cannot be addressed by traditional speech noise reduction or enhancement technologies. However, existing DNN models suffer from the drawbacks of large parameter counts and high computational complexity, resulting in significant consumption of computing resources and making them difficult to effectively deploy on edge devices with limited computing power.

[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a speech separation method, electronic device, storage medium and computer program product, aiming to solve the technical problem that the current end-to-end deep neural network (DNN) speech separation consumes a lot of computing resources and is difficult to be effectively deployed in end-side devices with limited computing power.

[0005] To achieve the above objectives, the present application proposes a speech separation method, which includes: Performing beamforming processing on the microphone array signal with a preset near-field speaker direction as the speech enhancement direction to obtain a first near-field speaker signal; Performing beamforming processing on the microphone array signal with a preset far-field speaker direction as a speech enhancement direction to obtain a first far-field speaker signal; Using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing to obtain a second near-field speaker signal; The first far-field speaker signal is used as a main signal, and the first near-field speaker signal is used as a reference signal, and input into a preset far-field separation model for processing to obtain a second far-field speaker signal.

[0006] Optionally, the speech separation method is applied to a head-mounted device, the preset near-field speaker direction includes the direction of the user's mouth when the head-mounted device is worn, and the preset far-field speaker direction includes the direction directly in front of the user when the head-mounted device is worn.

[0007] Optionally, before the step of performing beamforming processing on the microphone array signal with the preset far-field speaker direction as the speech enhancement direction to obtain the first far-field speaker signal, the method further includes: Sound source positioning is performed on the microphone array signal within a 360° range in the horizontal plane to obtain a target sound source direction, and the preset far-field speaker direction is obtained according to the target sound source direction.

[0008] Optionally, the speech separation method further includes: performing a voice activity detection on the first near-field speaker signal to obtain a first detection result, and / or performing a voice activity detection on the first far-field speaker signal to obtain a second detection result; If the first detection result indicates that a human voice is present, performing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal; When the second detection result indicates that a human voice is present, the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signals into a preset far-field separation model for processing, and obtaining a second far-field speaker signal is performed.

[0009] Optionally, the speech separation method further includes: Determining a calling scenario of the microphone array, and triggering a speech separation instruction according to the calling scenario; When the speech separation instruction instructs near-field and far-field speech separation or instructs near-field speech separation, performing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal; When the speech separation instruction indicates near-field and far-field speech separation or indicates far-field speech separation, the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signal into a preset far-field separation model for processing, and obtaining a second far-field speaker signal is performed.

[0010] Optionally, the speech separation method further includes: Obtaining first training data, wherein the first training data includes a first main signal and a first reference signal as model input data, and a first near-field clean signal as a model output label, the first main signal and the first reference signal are both obtained by mixing the first near-field clean signal and a first far-field clean signal, energy of the first near-field clean signal in the first main signal is greater than that of the first far-field clean signal, and energy of the first near-field clean signal in the first reference signal is less than that of the first far-field clean signal; The first training data is used to train a preset first deep neural network model to obtain the near-field separation model.

[0011] Optionally, the speech separation method further includes: Obtaining second training data, wherein the second training data includes a second main signal and a second reference signal as model input data, and a second far-field clean signal as a model output label, the second main signal and the second reference signal are both obtained by mixing a second near-field clean signal and the second far-field clean signal, energy of the second near-field clean signal in the second main signal is less than that of the second far-field clean signal, and energy of the second near-field clean signal in the second reference signal is greater than that of the second far-field clean signal; The second training data is used to train a preset second deep neural network model to obtain the far-field separation model.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a speech separation device, which includes: a near-field signal enhancement module, configured to perform beamforming processing on the microphone array signal with a preset near-field speaker direction as a speech enhancement direction to obtain a first near-field speaker signal; a far-field signal enhancement module, configured to perform beamforming processing on the microphone array signal with a preset far-field speaker direction as a speech enhancement direction to obtain a first far-field speaker signal; a near-field signal separation module, configured to process the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal into a preset near-field separation model to obtain a second near-field speaker signal; The far-field signal separation module is configured to process the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal into a preset far-field separation model to obtain a second far-field speaker signal.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the speech separation method as described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by the processor, the steps of the speech separation method described above are implemented.

[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the speech separation method described above.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: The first near-field speaker signal is obtained by beamforming the microphone array signal with the preset near-field speaker direction as the speech enhancement direction, thereby realizing preliminary near-field speaker signal separation; the first far-field speaker signal is obtained by beamforming the microphone array signal with the preset far-field speaker direction as the speech enhancement direction, thereby realizing preliminary far-field speaker signal separation; the first near-field speaker signal is used as the main signal, and the first far-field speaker signal is used as the reference signal, and is input into the preset far-field separation model for processing to obtain the second near-field speaker signal, thereby realizing preliminary far-field speaker signal separation. The speaker signal is used as a reference to further and more thoroughly remove the residual far-field signal from the initially separated near-field speaker signal, resulting in a final separated near-field speaker signal. The first far-field speaker signal is used as the main signal, and the first near-field speaker signal is used as the reference signal, which are input into a preset near-field separation model for processing to obtain a second far-field speaker signal. This achieves the goal of using the initially separated near-field speaker signal as a reference to further and more thoroughly remove the residual near-field signal from the initially separated far-field speaker signal, resulting in a final separated far-field speaker signal. Beamforming is used to perform preliminary separation of the near-field and far-field speaker signals, respectively. The initially separated near-field and far-field speaker signals serve as the main and reference signals for further near-field and far-field separation, respectively. This significantly improves the entry requirements of the back-end separation model and greatly reduces the scale requirements for the near-field and far-field separation models. This reduces the computational and memory resources consumed by the speech separation algorithm, making it possible to deploy the speech separation algorithm on resource-constrained devices such as wearable devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A flowchart of the first embodiment of the speech separation method of the present application is provided; Figure 2 A schematic diagram of the far-field and near-field speaker directions provided in the first embodiment of the present application; Figure 3 This is a diagram showing the speech separation effect provided in the second embodiment of this application; Figure 4 This is a schematic diagram of the module structure of the speech separation device according to an embodiment of the present application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the speech separation method in the embodiment of the present application.

[0020] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0022] It should be noted that, in the description of this application specification and the appended claims, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0024] The current mainstream solution uses end-to-end deep neural networks (DNNs) for speech separation. This technology uses a trained model to separate the far-field and near-field speakers in mixed audio into two independent signals, addressing the problem of mutual interference between sound sources that cannot be addressed by traditional speech noise reduction or enhancement techniques. However, existing DNN models suffer from large parameter counts and high computational complexity, resulting in significant consumption of computing resources and making them difficult to effectively deploy on edge devices with limited computing power.

[0025] The embodiment of the present application provides a solution, which performs beamforming processing on the microphone array signal with a preset near-field speaker direction as the speech enhancement direction to obtain a first near-field speaker signal, thereby achieving preliminary near-field speaker signal separation; performs beamforming processing on the microphone array signal with a preset far-field speaker direction as the speech enhancement direction to obtain a first far-field speaker signal, thereby achieving preliminary far-field speaker signal separation; uses the first near-field speaker signal as the main signal and the first far-field speaker signal as the reference signal, inputs a preset far-field separation model for processing, obtains a second near-field speaker signal, thereby achieving preliminary far-field speaker signal separation. The initially separated far-field speaker signal is used as a reference to further and more thoroughly remove the remaining far-field signal from the initially separated near-field speaker signal, resulting in a final separated near-field speaker signal. The first far-field speaker signal is used as the main signal, and the first near-field speaker signal is used as the reference signal, both of which are input into a preset near-field separation model for processing to obtain a second far-field speaker signal. This achieves the goal of further and more thoroughly removing the remaining near-field signal from the initially separated far-field speaker signal, using the initially separated near-field speaker signal as a reference, resulting in a final separated far-field speaker signal. Beamforming is used to perform preliminary separation of the near-field and far-field speaker signals, respectively. The initially separated near-field and far-field speaker signals serve as the main and reference signals for further near-field and far-field separation, respectively. This significantly improves the entry requirements of the backend separation model and greatly reduces the scale requirements for the near-field and far-field separation models. This reduces the computational and memory resources consumed by the speech separation algorithm, making it possible to deploy the speech separation algorithm on resource-constrained devices, such as wearable devices.

[0026] It should be noted that the execution entity of this embodiment can be an electronic device that is deployed with a microphone array and has data processing, network communication and program running functions, such as AR (Augmented Reality) / VR (Virtual Reality) glasses, helmets and other head-mounted devices, or tablet computers, personal computers, mobile phones and other devices.

[0027] Reference Figure 1 , Figure 1 This is a flowchart of the first embodiment of the speech separation method of the present application.

[0028] In this embodiment, the speech separation method includes steps S10 to S40: Step S10, performing beamforming processing on the microphone array signal with a preset near-field speaker direction as the speech enhancement direction to obtain a first near-field speaker signal; The near-field speaker direction refers to the position direction of the near-field speaker relative to the microphone array, which is usually represented by a direction vector in a spatial coordinate system and can be pre-set or dynamically calculated.

[0029] A microphone array is a collection of multiple microphones arranged in a certain manner. Each microphone can be regarded as an acoustic sensor that can convert sound signals into electrical signals. The multi-channel voice signals collected by these microphones at the same time will be read separately to form the microphone array signal.

[0030] Beamforming refers to the spatial filtering of microphone array signals in the direction of speech enhancement. By adjusting the phase and amplitude of each channel signal in the microphone array signal, the sound in the direction of speech enhancement can be structurally enhanced.

[0031] The first near-field speaker signal refers to a human voice signal with an improved signal-to-noise ratio obtained by beamforming speech enhancement on a mixed microphone array signal with a preset near-field speaker speaking direction as the speech enhancement direction. However, the first near-field speaker signal may still contain residual far-field speaker signals.

[0032] Optionally, a spatial coordinate system can be established based on a microphone array or an electronic device, and the direction of the near-field speaker can be determined based on the position of the near-field speaker, wherein the position of the near-field speaker can be determined by sound source localization; then, through a beamforming algorithm, the sound signal in the direction of the near-field speaker in the microphone array signal is enhanced, and the noise in other directions is suppressed, thereby obtaining a relatively pure near-field speaker signal after preliminary enhancement.

[0033] For example, a user may directly input the speaker direction (such as "1 meter directly in front") through the device interface. After receiving the input direction, the device converts it into a unit direction vector in the spatial coordinate system, such as (1,0,0), and stores the vector as the preset near-field speaker direction. Then, while collecting voice signals in real time through the microphone array, the collected microphone array signal is beamformed with the near-field speaker direction as the voice enhancement direction to obtain a first near-field speaker signal with a relatively pure audio signal in the near-field speaker direction.

[0034] Step S20, performing beamforming processing on the microphone array signal with a preset far-field speaker direction as a speech enhancement direction to obtain a first far-field speaker signal; A far-field speaker is one that is relatively far from the microphone array. Their sound signals travel through the air for a long time before reaching the microphone array, making them susceptible to interference from ambient noise and reverberation. Correspondingly, the preset far-field speaker direction refers to the direction of the far-field speaker relative to the microphone array. This can be determined by a user-selected orientation or by performing sound source localization on the far-field speaker signal.

[0035] By performing beamforming processing on the microphone array signal with the preset far-field speaker direction as the speech enhancement direction, near-field speaker signals are suppressed and far-field speaker signals are enhanced, thereby obtaining a first far-field speaker signal. However, some near-field speaker signals may still remain in the first far-field speaker signal.

[0036] Optionally, an interactive interface can be output, which illustrates an image of an electronic device equipped with a microphone array and multiple directions around the electronic device. The user can select a direction from the interactive interface as the preset far-field speaker direction based on the location of the far-field speaker.

[0037] The specific implementation of step S20 in this embodiment can refer to the specific implementation of step S10 above, and the near-field speaker in step S10 can be replaced by the far-field speaker, which will not be described in detail here.

[0038] In a feasible implementation manner, before step S20, the method further includes: Step S201 : performing sound source localization on the microphone array signal within a 360° range on a horizontal plane to obtain a target sound source direction, and obtaining the preset far-field speaker direction according to the target sound source direction.

[0039] In one feasible embodiment, audio signals within a 360° range on the horizontal plane are collected by a microphone array, and then the time difference between the audio signals reaching different microphones is analyzed through sound source localization to determine the direction of the sound source of the audio signal in space, and based on this, the preset far-field speaker direction is determined.

[0040] Optionally, in addition to the time difference, the sound source localization may be achieved by analyzing the phase difference or energy distribution of the sound reaching different microphones. This embodiment does not impose any specific limitation on the method for achieving the sound source localization.

[0041] Optionally, the output of sound source localization can usually be expressed in the form of azimuth, elevation, three-dimensional coordinates, unit direction vectors, etc. In practical applications, the sound source direction can be expressed in vector form to facilitate subsequent beamforming processing. For the convenience of description, the sound source direction is expressed in azimuth and elevation. The azimuth indicates the direction of the sound source on the horizontal plane (a two-dimensional plane perpendicular to the direction of gravity), usually with the front of the microphone array as 0°, and the range is generally 0°~360°; the elevation indicates the angle of the sound source in the vertical direction, and the range is generally -90°~90°. A pitch angle of 0° indicates that the sound source position is at the same height as the microphone array, +90° indicates that the sound source position is directly above the microphone array, and -90° indicates that the sound source position is directly below the microphone array.

[0042] For example, each sound source will produce a specific time delay between different microphones. The direction of the sound source can be inferred through multiple delays. In the case of multiple sound sources, the generalized cross-correlation (GCC) algorithm can be used to first estimate multiple time delay differences (TDOAs). Then, all possible TDOA sequence combinations are generated using permutations and combinations, and different TDOA sequence combinations are solved to determine the azimuth angles corresponding to the multiple sound sources. The obtained azimuth angle information is then displayed on the interactive interface of the electronic device, and the azimuth angle selected by the user is determined as the target sound source direction. The azimuth angle (target sound source direction) is then converted into a vector form to obtain the preset far-field speaker direction.

[0043] In this embodiment, by dynamically determining the far-field speaker direction through sound source localization, the speech signal from a specific direction can be enhanced in real time, thereby reducing the interference of background noise and speech signals from other directions during the speech separation process.

[0044] Optionally, a speaker selection option can be provided on the device's interface to help users determine the direction of the far-field speaker. For devices with a display, such as mobile phones, this can be displayed directly on the interface; for devices without a display, such as head-mounted devices, the corresponding speaker selection option can be displayed on the device's associated display.

[0045] For example, an interactive interface can be output, which illustrates the image of a device equipped with a microphone array and the images of selectable speakers around the device, and the relative position between the device and the selectable speakers is determined by sound source localization; furthermore, the user can select a far-field speaker in the interactive interface based on the azimuth relationship between the communication object and the electronic device in the actual scenario; furthermore, the electronic device can determine the direction of the far-field speaker based on the far-field speaker selected by the user.

[0046] For example, an interactive interface can be output, displaying controls for near-field speaker identification and far-field speaker identification. When the user triggers the near-field speaker identification control, the device can issue a prompt message, "Near-field speaker, please speak," and perform sound source localization to determine the direction of the near-field speaker. Similarly, when the user triggers the far-field speaker identification control, the device can issue a prompt message, "Far-field speaker, please speak," and perform sound source localization to determine the direction of the far-field speaker. Alternatively, the control triggering operation can be omitted. When entering the interactive interface for speaker direction setting, the near-field speaker speaks first, and the electronic device performs sound source localization to determine the direction of the near-field speaker. Then, the far-field speaker speaks again, and the electronic device performs sound source localization to determine the direction of the far-field speaker.

[0047] In a feasible implementation, the speech separation method is applied to a head-mounted device, the preset near-field speaker direction includes the direction of the user's mouth when the head-mounted device is worn, and the preset far-field speaker direction includes the direction directly in front of the user when the head-mounted device is worn.

[0048] In one feasible embodiment, the above-mentioned speech separation method is applied to head-mounted devices such as VR / AR glasses. In order to conveniently describe the directions of near-field speakers and far-field speakers, a spatial coordinate system is established based on the microphone array of the head-mounted device. The preset near-field speaker direction can be a spatial cone area pointing from the microphone array to the user's mouth, such as the azimuth angle. , pitch angle ; This includes the direction of the user's mouth when the head-mounted device is worn, such as the azimuth angle is 0, the pitch angle The preset far-field speaker direction can be a fan-shaped area directly in front of the user wearing the head-mounted device, with the front of the head-mounted device as the reference, such as the azimuth angle , pitch angle , this area must include the area directly in front of the user when the head-mounted device is worn, and the azimuth and pitch angles are both 0.

[0049] For example, please refer to Figure 2 , Figure 2 In the figure, a spatial coordinate system is established based on the head-mounted device when it is worn by the user. The xOy plane is a two-dimensional plane perpendicular to the direction of gravity. The positive direction of the x-axis is the front direction of the head-mounted device. Assume that there are users (not shown in the figure), speakers 1 and 2 in the current scene. Figure 2 Region A is the direction of the user's mouth, region B is the direction of speaker 1, and region C is the direction of speaker 2. Speaker 1 is directly in front of the user. Figure 2A slight offset is made to distinguish it from the x-axis, but in fact the direction vector from the origin to speaker 1 should coincide with the x-axis. When separating the surrounding audio signals through the headset, the preset near-field speaker direction is first determined to be the direction of the user's mouth when the headset is worn, that is, the direction of area A. Therefore, the azimuth angle of the preset near-field speaker direction can be determined to be 0, and the pitch angle is , so that it can be called during the subsequent beamforming processing. Since the preset far-field speaker direction includes the front of the user when the head-mounted device is worn, that is, the direction of area C, it can be determined that the azimuth and pitch angles of a far-field speaker direction are both 0; at the same time, the preset far-field speaker direction can also include a fan-shaped area in front of the head-mounted device as the reference. Figure 2 In the scenario, the far-field speaker direction also includes the direction of area B, and the azimuth of the other far-field speaker direction can be determined as , the pitch angle is ; In the case where there are multiple far-field speaker directions, an interactive interface containing these far-field speaker directions can be output on the display screen of the head-mounted device, and the preset far-field speaker direction is determined according to the direction vector selected by the user.

[0050] from Figure 2 As can be seen from the figure, the preset near-field speaker direction (the direction vector pointing from the coordinate origin to region A) is located on the yOz plane, while the preset far-field speaker direction is located on the xOz plane. Therefore, the preset near-field speaker direction and the preset far-field speaker direction do not overlap.

[0051] It can be understood that the audio signals in the speaking directions of the selected near-field speakers and far-field speakers are enhanced through beamforming processing, making the near-field speaker signals and the far-field speaker signals clearer, thereby reducing the workload of subsequent model processing.

[0052] Step S30: using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing to obtain a second near-field speaker signal; In one feasible embodiment, a preset algorithm model utilizes the more significant noise characteristics in the reference signal (the first far-field speaker signal) to further suppress the residual far-field speaker signal in the main signal (the first near-field speaker signal), thereby obtaining a pure near-field speaker signal (hereinafter referred to as the second near-field speaker signal) for speech recognition and other purposes.

[0053] Optionally, the primary signal refers to the target signal to be enhanced, which may contain interference features such as residual noise and reverberation, while the reference signal refers to the non-target signal that carries the interference features. It is important to note that the time and frequency domains of the reference signal are correlated with the primary signal to allow for separation model references and to separate interference signals in the same time or frequency domain as the primary signal.

[0054] Optionally, the near-field separation model is a trained machine learning model used to process the input main signal and reference signal to separate a purer near-field speech signal. It is usually implemented through signal processing algorithms (such as adaptive filtering, blind source separation, etc.) or deep learning algorithms. This embodiment does not limit the specific implementation method of the near-field separation model.

[0055] Optionally, the microphone array receives a time domain signal, and the output after beamforming processing may be a time domain signal or a frequency domain signal. For example, when beamforming is achieved by delaying and weighted summing the signals from different microphones, the corresponding output is a time domain signal, while Fourier domain beamforming outputs a frequency domain signal. The speech separation model can process the time domain signal to learn time domain features for speech separation, or it can process the frequency domain signal to learn spectral features for speech separation.

[0056] In a feasible embodiment, the near-field separation model is obtained by training a preset first deep neural network model through first training data, and the first training data includes a first main signal and a first reference signal as model input data, and a first near-field pure signal as a model output label. The first main signal and the first reference signal are both obtained by mixing the first near-field pure signal and the first far-field pure signal. The energy of the first near-field pure signal in the first main signal is greater than that of the first far-field pure signal, and the energy of the first near-field pure signal in the first reference signal is less than that of the first far-field pure signal.

[0057] Optionally, the first training data is a dataset consisting of input data and output labels, which may include one or more pieces of training data. The input data includes an artificially constructed first main signal (dominated by a first near-field clean signal) and a first reference signal (dominated by a first far-field clean signal). The first near-field clean signal and the first far-field clean signal are both original signals uninterrupted by other sound sources, representing independent audio signals of the target sound source (near field) and the non-target sound source (far field), respectively. The first near-field clean signal serves as the model's output label to supervise the near-field signal separation performance of the near-field separation model (deep neural network model).

[0058] Optionally, during each iterative training process, the first deep neural network model will output its predicted near-field pure signal based on the input data; in order to further improve the signal separation effect of the first deep neural network model, the error loss calculation can be performed between the predicted near-field pure signal output by it and the first near-field pure signal in the first training data, and the first deep neural network model can be reversely optimized according to the error loss function until the loss value converges or reaches a preset number of training times to obtain a near-field separation model.

[0059] Exemplarily, loss functions such as mean square error loss, L1 norm loss, and cross entropy loss can be used to achieve reverse optimization. This embodiment does not impose any specific restrictions on the loss function used in the training process.

[0060] In this embodiment, the first deep neural network is trained by taking the artificially constructed main signal (dominated by the near-field pure signal) and the reference signal (dominated by the far-field pure signal) as the model input and the near-field pure signal as the output label, so that the trained near-field separation model can extract a purer near-field speaker signal from the main signal.

[0061] Step S40 : Using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, the signals are input into a preset far-field separation model for processing to obtain a second far-field speaker signal.

[0062] The specific implementation of step S40 in this embodiment can refer to the specific implementation of step S30 above, and the near-field speaker signal in step S30 can be replaced by the far-field speaker signal, and the far-field speaker signal can be replaced by the near-field speaker signal. No further details are given here.

[0063] Optionally, the far-field separation model, similar to the near-field separation model, is a trained machine learning model that processes the input primary and reference signals. It uses the first near-field speaker signal as a reference to suppress any interference in the primary signal, isolate a purer far-field speech signal, and further improve the quality of the far-field speaker signal. The far-field separation model can be a deep neural network model, a Gaussian mixture model, or other similar model. This embodiment does not limit the specific implementation of the far-field separation model.

[0064] For example, please refer to Figure 3 , Figure 3A general framework diagram for performing speech separation on microphone array signals is provided. First, microphone array information is obtained and input into a beamforming module. Then, the beamforming module performs beamforming processing on the microphone array signal with the preset near-field speaker direction as the speech enhancement direction to obtain a first near-field speaker signal. Simultaneously, the microphone array signal is beamformed with the preset far-field speaker direction as the speech enhancement direction to obtain a first far-field speaker signal, thereby achieving preliminary separation of near- and far-field signals. Furthermore, through a near-field separation model, the first far-field speaker signal is used as a reference signal to further suppress the far-field speaker signal remaining in the first near-field speaker signal, thereby obtaining a purer second near-field speaker signal. Simultaneously, through a far-field separation model, the first near-field speaker signal is used as a reference signal to further suppress the near-field speaker signal remaining in the first far-field speaker signal, thereby obtaining a purer second far-field speaker signal.

[0065] In a feasible embodiment, the far-field separation model can be obtained by training a preset second deep neural network model through second training data, wherein the second training data includes a second main signal and a second reference signal as model input data, and a second far-field clean signal as a model output label, and the second main signal and the second reference signal are both obtained by mixing a second near-field clean signal and a second far-field clean signal, the energy of the second near-field clean signal in the second main signal is less than that of the second far-field clean signal, and the energy of the second near-field clean signal in the second reference signal is greater than that of the second far-field clean signal.

[0066] Optionally, the first deep neural network model and the second deep neural network model can adopt the same model structure; since the two use different model output labels, during the reverse optimization process, the adjustment directions of the model parameters of the two are different. The first deep neural network model uses the near-field pure signal as the model output label, learns how to separate the near-field signal, and obtains the near-field separation model after the training is completed, while the second deep neural network uses the far-field pure signal as the model output label, learns how to separate the far-field signal, and obtains the far-field separation model after the training is completed.

[0067] Optionally, a first set of training data and a second set of training data may be constructed simultaneously based on a section of near-field pure signal and a section of far-field pure signal.

[0068] In this embodiment, beamforming is first performed on the microphone array signal in the preset near-field and far-field speaker directions to obtain a first near-field speaker signal and a first far-field speaker signal, respectively, thereby achieving preliminary separation of the near-field and far-field speaker audio signals. The trained near-field separation model then uses the first far-field speaker signal obtained from the preliminary separation as a reference to further suppress the residual far-field speaker signal in the first near-field speaker signal, thereby obtaining a purer near-field speaker signal. Similarly, the trained far-field separation model obtains a purer far-field speaker signal. This preliminary separation based on beamforming improves the entry requirements of the backend separation model and reduces the scale requirements for the near-field and far-field separation models in the subsequent separation process. This reduces the computational and memory resources used by the speech separation algorithm, making it possible to deploy the speech separation algorithm on resource-constrained devices such as wearable devices.

[0069] Based on the above first embodiment, a second embodiment of the speech separation method of the present application is proposed. In this embodiment, the same or similar contents as those of the above first embodiment can be referred to the above introduction and will not be repeated hereafter. In this embodiment, the speech separation method further includes steps B10 to B40: Step B10: performing voice activity detection on the first near-field speaker signal to obtain a first detection result, and / or performing voice activity detection on the first far-field speaker signal to obtain a second detection result; Voice Activity Detection (VAD) can be used to detect whether there is valid speech in an audio signal. In this embodiment, a voice activity detection method is used to detect whether the first near-field speaker signal and / or the first far-field speaker signal contains a human voice.

[0070] The first detection result may be a result indicating whether a human voice exists in the first near-field speaker signal. The second detection result may be a result indicating whether a human voice exists in the first far-field speaker signal.

[0071] Step B20: If the first detection result indicates the presence of a human voice, performing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal. Step B30: When the second detection result indicates the presence of a human voice, executing the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signals into a preset far-field separation model for processing, and obtaining a second far-field speaker signal.

[0072] For example, assuming that a user is wearing a head-mounted device for two-way real-time translation, the sound signal collected by the microphone array of the head-mounted device is processed by beamforming to obtain a first near-field speaker signal and a first far-field speaker signal; the first near-field speaker signal is then detected, and it is found that the user is speaking (i.e., the first detection result is that there is a human voice); at the same time, the first far-field speaker signal is detected, and it is found that there is no voice activity (i.e., the second detection result is that there is no human voice); based on these detection results, the first near-field speaker signal is used as the main signal and the first far-field speaker signal is used as the reference signal, and are input into the near-field separation model for processing to obtain a purer second near-field speaker signal for subsequent speech recognition and real-time translation, while the first far-field speaker signal where there is no human voice is not processed.

[0073] Optionally, when both the first detection result and the second detection result indicate that there is no human voice, a prompt message of "no one is speaking" can be output, and the prompt message can be presented through voice broadcast, text prompt, vibration, etc.; in addition, other operations can also be taken according to the configuration information of the electronic device.

[0074] For example, when a user hosts a meeting by wearing a head-mounted device, if the head-mounted device detects that there is no human voice in the near field and far field, the next meeting item can be displayed on the corresponding display screen or supporting interactive application to advance the meeting process.

[0075] For example, when a user watches a VR movie by wearing a head-mounted device, the head-mounted device can shield the surrounding near-field speaker signals and only output the far-field speaker signals. When there is no human voice in the far field, the surrounding environmental acoustic information can be further collected to more accurately simulate the propagation effect of sound in the virtual scene, making the user feel as if they are in the movie scene.

[0076] In this embodiment, after performing preliminary separation of the far-field and near-field signals, voice activity detection is performed on them to ensure that the far-field and near-field signals obtained by the preliminary separation contain human voices. Only then is the subsequent processing model started to perform further voice separation to avoid unnecessary processing of invalid signals, thereby avoiding unnecessary separation operations from occupying the computing resources of the voice separation model (including the near-field separation model and the far-field separation model).

[0077] In one feasible implementation, the speech separation method further includes steps E10 to E30: Step E10, determining a calling scenario of the microphone array, and triggering a speech separation instruction according to the calling scenario; Step E20, when the speech separation instruction indicates near-field and far-field speech separation or indicates near-field speech separation, executing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal; Step E30, when the speech separation instruction indicates near-field and far-field speech separation or indicates far-field speech separation, executes the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signal into a preset far-field separation model for processing, and obtaining a second far-field speaker signal.

[0078] Optionally, the calling scenario refers to a scenario in which an upper-layer application or function of an electronic device needs to call a microphone array to pick up sound. For example, in a head-mounted device, the calling scenario may be a translation scenario, a call scenario, etc.

[0079] Optionally, language separation instructions corresponding to different calling scenarios may be pre-set in the electronic device, and after the calling scenario is determined, the corresponding voice separation instruction is triggered.

[0080] For example, when an electronic device invokes the microphone array through call software, and the corresponding speech separation instruction of the call software instructs near-field speech separation, the first near-field speaker signal obtained after preliminary separation is simply used as the main signal, and the first far-field speaker signal is used as the reference signal. These are input into the near-field separation model for processing to obtain a second near-field speaker signal, without performing speech separation on the far-field signal that is invalid in the current call scenario, thereby reducing inefficient resource consumption of the far-field speech separation model. Similarly, when the microphone array is invoked through a translation scenario, assuming the user only needs to translate information from other speakers, the speech separation instruction instructing far-field speech separation is triggered, and no further processing of the near-field speaker signal is performed, thereby reducing inefficient resource consumption of the near-field speech separation model. If the user needs to translate information for both themselves and other speakers, the speech separation instruction instructing speech separation of both near-field and far-field signals is triggered, and further speech separation is performed on the first near-field speaker signal and the first far-field speaker signal, thereby improving the signal-to-noise ratio of the near-field and far-field speaker signals, thereby improving the recognition rate of both the far-field and near-field speech signals.

[0081] In this embodiment, the speech separation instruction is triggered according to the calling scenario of the current microphone array to decide whether to perform further speech separation on the far-field and near-field speaker signals after the initial separation, thereby avoiding processing of speech signals that are invalid for the current scenario and reducing the resource occupation and resource consumption of the speech separation model.

[0082] To further illustrate the effectiveness of the speech separation method obtained by combining the first and second embodiments, Table 1 shows the ASR (Automatic Speech Recognition) recognition rates for different speech signals at different far-field and near-field distances.

[0083] Table 1 ASR recognition rate results

[0084] The near-field and far-field distances refer to the distances between the near-field and far-field human voice sources. The percentages in the "mix-in" column represent the speech recognition rates for the near-field and far-field signals, respectively, when performing speech recognition on the mixed signal. The percentages in the "alg-out" column represent the recognition rates for the far-field and near-field signals, respectively, obtained using the aforementioned speech separation method. As Table 1 shows, at far-field and near-field distances of 1m, 1.5m, and 2m, the recognition rates for speech recognition after speech separation are significantly higher than those for the mixed signal. Furthermore, the ASR recognition rate for simultaneous speech in the near-field and far-field scenarios at a distance of 1m exceeded 96%, meeting commercial standards.

[0085] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the speech separation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0086] The present application also provides a speech separation device, please refer to Figure 4 , the speech separation device includes: A near-field signal enhancement module 10 is configured to perform beamforming processing on the microphone array signal with a preset near-field speaker direction as a speech enhancement direction to obtain a first near-field speaker signal; A far-field signal enhancement module 20 is configured to perform beamforming processing on the microphone array signal with a preset far-field speaker direction as a speech enhancement direction to obtain a first far-field speaker signal; a near-field signal separation module 30 for processing the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal into a preset near-field separation model to obtain a second near-field speaker signal; The far-field signal separation module 40 is configured to process the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal into a preset far-field separation model to obtain a second far-field speaker signal.

[0087] The speech separation device provided in the embodiment of the present application adopts the speech separation method in the above embodiment. Compared with the prior art, the beneficial effects of the speech separation device provided in the present application are the same as the beneficial effects of the speech separation method provided in the above embodiment, and the other technical features in the speech separation device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0088] An embodiment of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the speech separation method in the above-mentioned embodiment 1.

[0089] Reference below Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0090] like Figure 5As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wired to exchange data. Although the figures show electronic devices with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0091] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0092] The electronic device provided in the embodiment of the present application adopts the speech separation method in the above embodiment. Compared with the prior art, the beneficial effects of the electronic device provided in the present application are the same as the beneficial effects of the speech separation method provided in the above embodiment, and the other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0093] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0094] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0095] An embodiment of the present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the speech separation method in the above embodiment.

[0096] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0097] The computer-readable storage medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0098] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device performs the functions defined in the method of the embodiment disclosed in this application.

[0099] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0100] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0101] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0102] The readable storage medium provided in the embodiment of the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned speech separation method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiment of the present application are the same as the beneficial effects of the speech separation method provided in the above-mentioned embodiment, and will not be repeated here.

[0103] An embodiment of the present application further provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned speech separation method when executed by a processor.

[0104] Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of the present application are the same as the beneficial effects of the speech separation method provided in the above embodiment, and will not be repeated here.

[0105] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A speech separation method, characterized in that: The speech separation method comprises: Performing beamforming processing on the microphone array signal with a preset near-field speaker direction as the speech enhancement direction to obtain a first near-field speaker signal; Performing beamforming processing on the microphone array signal with a preset far-field speaker direction as a speech enhancement direction to obtain a first far-field speaker signal; Using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing to obtain a second near-field speaker signal; The first far-field speaker signal is used as a main signal, and the first near-field speaker signal is used as a reference signal, and input into a preset far-field separation model for processing to obtain a second far-field speaker signal.

2. The speech separation method according to claim 1, wherein The speech separation method is applied to a head-mounted device, the preset near-field speaker direction includes the direction of the user's mouth when the head-mounted device is worn, and the preset far-field speaker direction includes the direction directly in front of the user when the head-mounted device is worn.

3. The speech separation method according to claim 1, wherein Before the step of performing beamforming processing on the microphone array signal with the preset far-field speaker direction as the speech enhancement direction to obtain the first far-field speaker signal, the method further includes: Sound source positioning is performed on the microphone array signal within a 360° range in the horizontal plane to obtain a target sound source direction, and the preset far-field speaker direction is obtained according to the target sound source direction.

4. The speech separation method according to claim 1, wherein The speech separation method further comprises: performing a voice activity detection on the first near-field speaker signal to obtain a first detection result, and / or performing a voice activity detection on the first far-field speaker signal to obtain a second detection result; If the first detection result indicates that a human voice is present, performing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal; When the second detection result indicates that a human voice is present, the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signals into a preset far-field separation model for processing, and obtaining a second far-field speaker signal is performed.

5. The speech separation method according to claim 1, wherein The speech separation method further comprises: Determining a calling scenario of the microphone array, and triggering a speech separation instruction according to the calling scenario; When the speech separation instruction instructs near-field and far-field speech separation or instructs near-field speech separation, performing the step of using the first near-field speaker signal as a main signal and the first far-field speaker signal as a reference signal, inputting the signals into a preset near-field separation model for processing, and obtaining a second near-field speaker signal; When the speech separation instruction indicates near-field and far-field speech separation or indicates far-field speech separation, the step of using the first far-field speaker signal as a main signal and the first near-field speaker signal as a reference signal, inputting the signal into a preset far-field separation model for processing, and obtaining a second far-field speaker signal is performed.

6. The speech separation method according to claim 1, wherein: The speech separation method further comprises: Obtaining first training data, wherein the first training data includes a first main signal and a first reference signal as model input data, and a first near-field clean signal as a model output label, the first main signal and the first reference signal are both obtained by mixing the first near-field clean signal and a first far-field clean signal, energy of the first near-field clean signal in the first main signal is greater than that of the first far-field clean signal, and energy of the first near-field clean signal in the first reference signal is less than that of the first far-field clean signal; The first training data is used to train a preset first deep neural network model to obtain the near-field separation model.

7. The speech separation method according to any one of claims 1 to 6, characterized in that: The speech separation method further comprises: Obtaining second training data, wherein the second training data includes a second main signal and a second reference signal as model input data, and a second far-field clean signal as a model output label, the second main signal and the second reference signal are both obtained by mixing a second near-field clean signal and the second far-field clean signal, energy of the second near-field clean signal in the second main signal is less than that of the second far-field clean signal, and energy of the second near-field clean signal in the second reference signal is greater than that of the second far-field clean signal; The second training data is used to train a preset second deep neural network model to obtain the far-field separation model.

8. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the speech separation method according to any one of claims 1 to 7.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the speech separation method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the speech separation method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Speaker speech separation method and device

    CN108520756A

  • Voice consultation device, voice consultation method and storage medium

    CN116132878A

  • Far and near field voice separation method and apparatus, and wearable device

    CN120108414A

  • Signal processing device and program

    JP2010217773A

  • Dual Omnidirectional Microphone Array (DOMA)

    US20090003623A1

Cited By

  • Intelligent glasses double-target speech enhancement method and system based on multi-microphone array, terminal equipment and medium

    CN121938388A

  • A method, system, terminal device, and medium for dual-target speech enhancement in smart glasses based on a multi-microphone array.

    CN121938388B