Voice separation method and device of audio signal, equipment and storage medium

By performing environmental background noise processing and feature fusion on the original microphone signal, combined with Bayesian non-parametric model and voiceprint enhancement network, the problems of low efficiency and poor stability of speech separation in the existing technology are solved, and efficient and stable speech separation effect is achieved, adapting to dynamic conference scenarios and improving user experience.

CN120279931APending Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512524.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The speech separation technology of existing conference scenarios has problems such as preset fixed number of speakers that cannot adapt to dynamic changes, single domain feature processing cannot fully utilize multi-domain information in time and frequency space, lack of effective continuous tracking mechanism for speaker identity, insufficient separation speech quality and spectrum distortion and artifacts, and difficulty in maintaining stable performance in complex acoustic environments.

Method used

By processing the original microphone signal based on environmental background noise, position estimation is used to use preset position estimation technology and confidence weighting optimization algorithm to perform position estimation, combined with preset feature extraction and fusion rules, Bayesian non-parametric model harmony student model, voiceprint enhancement network and psychoacoustic model are used to perform speech separation, so as to achieve identity label determination and speech separation of the target speaker's voice signal.

Benefits of technology

It improves the efficiency of speech separation, improves user experience, enhances the stable performance and separation quality in complex acoustic environments, adapts to dynamically changing conference scenarios, and improves separation accuracy and speech clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279931A_ABST
    Figure CN120279931A_ABST
Patent Text Reader

Abstract

The invention discloses a voice separation method, device and equipment for an audio signal and a storage medium, and relates to the technical field of voice signal processing, and the method comprises the steps: carrying out the position estimation of a to-be-processed microphone signal obtained through the processing of an original microphone signal based on environment background noise through employing a preset position estimation technology and a confidence coefficient weighted optimization algorithm, and obtaining a to-be-processed microphone signal; obtaining a target microphone three-dimensional coordinate set; processing the to-be-processed microphone signal to obtain a feature fusion result, and processing the feature fusion result and the target microphone three-dimensional coordinate set by using a Bayesian nonparametric model and an acoustic imaging model to obtain a target speaker voice signal and a corresponding three-dimensional coordinate; and processing the voice signal of each target speaker by using a preset voiceprint enhancement network to obtain a corresponding identity tag, and determining voice separation information based on the voice signal of each target speaker, the identity tag and the three-dimensional coordinate by using a preset psychological acoustic model and a preset separation strategy. Therefore, the efficiency of separating audio signals from voice can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech signal processing, and particularly relates to a method, device, equipment and storage medium for separating voices of audio signals. Background Art

[0002] Currently, the existing voice separation technology in conference scenarios has the following main problems: First, a fixed number of preset speakers is difficult to adapt to the dynamically changing conference scenarios; second, single-domain feature processing cannot fully utilize multi-domain information in the time-frequency space; third, there is a lack of an effective speaker identity continuous tracking mechanism; fourth, the quality of the separated voices is insufficient, with spectrum distortion and artifacts; fifth, it is difficult to maintain stable performance in complex acoustic environments.

[0003] As can be seen from the above, how to improve the efficiency of separating voices from audio signals during the voice separation process of audio signals is an urgent problem to be solved at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for separating voices of audio signals, which can improve the efficiency of separating voices from audio signals during the voice separation process of audio signals. The specific scheme is as follows:

[0005] In the first aspect, the present application provides a method for separating voices of audio signals, including:

[0006] Processing the original microphone signal based on the environmental background noise to obtain a to-be-processed microphone signal, and using a preset position estimation technology and a confidence-weighted optimization algorithm to perform position estimation on the to-be-processed microphone signal to obtain a set of three-dimensional coordinates of the target microphone;

[0007] Processing the to-be-processed microphone signal using a preset feature extraction and fusion rule to obtain a feature fusion result, and using a Bayesian nonparametric model and an acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain several target speaker voice signals and the three-dimensional coordinates corresponding to each speaker;

[0008] Processing each of the target speaker voice signals using a preset voiceprint enhancement network to obtain an identity label corresponding to each of the target speaker voice signals, and then using a preset psychoacoustic model and a preset separation strategy, and determining voice separation information corresponding to the original microphone signal based on each of the target speaker voice signals, the identity label and the corresponding three-dimensional coordinates.

[0009] Optionally, processing the original microphone signal based on the ambient background noise to obtain a microphone signal to be processed, and using a preset position estimation technique and a confidence-weighted optimization algorithm to perform position estimation on the microphone signal to be processed, to obtain a set of three-dimensional coordinates of the target microphone, including:

[0010] Establishing a noise dictionary based on historical ambient background noise, and performing a noise suppression operation on the original microphone signal based on the current ambient background noise and the noise dictionary to obtain an intermediate microphone signal;

[0011] Using the cross-correlation function corresponding to each microphone pair and the intermediate microphone signal to determine the acoustic wave propagation delay corresponding to each microphone pair, and determining an initial distance matrix based on the product of the acoustic wave propagation delay and the speed of sound;

[0012] Using a preset confidence-weighted mechanism to process each distance value in the initial distance matrix to obtain a target distance matrix, and using a three-dimensional coordinate reconstruction model and a preset position estimation technique to process the target distance matrix to obtain a set of initial three-dimensional coordinates of the microphone;

[0013] Using a preset gradient descent algorithm and based on the noise level corresponding to the microphone signal to be processed to process the set of initial three-dimensional coordinates of the microphone to obtain a set of three-dimensional coordinates of the target microphone.

[0014] Optionally, processing the microphone signal to be processed using a preset feature extraction and fusion rule to obtain a feature fusion result, including:

[0015] Using a preset attention mechanism to establish a deep convolutional network including a progressive receptive field, and using the deep convolutional network to perform an acoustic feature extraction operation on the microphone signal to be processed to obtain corresponding first acoustic features, second acoustic features, and third acoustic features; the time scales corresponding to the first acoustic features, the second acoustic features, and the third acoustic features increase in sequence;

[0016] Using the preset attention mechanism to perform a fusion operation on the first acoustic features, the second acoustic features, and the third acoustic features to obtain an acoustic feature fusion result;

[0017] Using a preset complex-valued neural network and a preset phase continuity constraint to perform processing operations on the frequency-domain amplitude and phase information of the acoustic feature fusion result to obtain time-domain features and frequency-domain features, and establishing a graph attention network based on the microphone array corresponding to the microphone signal to be processed; wherein, the nodes in the graph attention network are the microphones in the microphone array, and the edges in the graph attention network are the correlations between the microphones.

[0018] Determine spatial topological features by using the graph attention network and based on the microphone signal to be processed, and fuse the time-domain features, the frequency-domain features and the spatial topological features to obtain a feature fusion result.

[0019] Optionally, processing the feature fusion result and the set of three-dimensional coordinates of the target microphone by using a Bayesian non-parametric model and an acoustic generation model to obtain a plurality of target speaker voice signals and three-dimensional coordinates corresponding to each speaker, including:

[0020] Cluster the feature space corresponding to the feature fusion result by using a Dirichlet process mixture model to obtain pitch features corresponding to the feature fusion result;

[0021] Determine the number of speakers corresponding to the microphone signal to be processed by using a Bayesian non-parametric model based on the feature fusion result and the pitch features;

[0022] Determine a bidirectional fusion mechanism based on a preset separation strategy and a preset weighting mechanism, and then use the bidirectional fusion mechanism and based on a preset index to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain a plurality of initial speaker voice signals and three-dimensional coordinates corresponding to each speaker; the preset separation strategy includes time-frequency masking and spatial filtering;

[0023] Determine a spectral compensation network based on an acoustic generation model and a speech prior model including a variational autoencoder, so as to use the spectral compensation network to determine the spectral structure characteristics of the initial speaker voice signal to obtain spectral structure characteristics, and then determine the target speaker voice signal based on the spectral structure characteristics and the initial speaker voice signal.

[0024] Optionally, processing each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain identity labels corresponding to each of the target speaker voice signals, including:

[0025] Perform a feature extraction operation on each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain voiceprint features, speaker spatial position features and speech content features corresponding to the target speaker voice signal;

[0026] Establish a modal similarity metric based on the voiceprint features, the speaker spatial position features and the speech content features, and then use an adaptive memory model and based on the modal similarity metric to process the speaker features corresponding to the target speaker voice signal to obtain a processing result;

[0027] Determine the activity level corresponding to the speech signal of the target speaker, and use a speaker state tracking algorithm based on particle filtering to track the identity of the speaker based on the activity level and the processing result, so as to obtain identity tags corresponding to the speech signals of each target speaker.

[0028] Optionally, the determining the speech separation information corresponding to the original microphone signal by using a preset psychoacoustic model and a preset separation strategy, and based on the speech signals of each target speaker, the identity tags, and the corresponding three-dimensional coordinates includes:

[0029] Determine a blind de-reverberation algorithm based on a preset neural network, and use the blind de-reverberation algorithm and a preset dry-wet sound decomposition model to decompose the speech signal of the target speaker to obtain a direct sound signal and a reverberation component signal;

[0030] Perform an inhibition operation on the reverberation component signal by using a preset inhibition rule to obtain an inhibited signal, and determine a signal to be repaired based on the inhibited signal and the direct sound signal;

[0031] Determine a harmonic structure constraint condition based on the sensitivity of the user to several sound frequency bands respectively, and perform a repair operation on the signal to be repaired based on the harmonic structure constraint condition and the spectral repair algorithm in the preset psychoacoustic model to obtain speech information corresponding to each speaker;

[0032] Determine the speech separation information corresponding to the original microphone signal based on the speech information corresponding to each speaker, the identity tags, and the three-dimensional coordinates.

[0033] Optionally, after determining the speech separation information corresponding to the original microphone signal based on the speech signals of each target speaker, the identity tags, and the corresponding three-dimensional coordinates, it further includes:

[0034] Determine the vowels corresponding to each pronunciation part based on the pronunciation part corresponding to the speech separation information, determine the corresponding enhancement strategy based on each vowel, and use each enhancement strategy to perform an enhancement processing operation on the corresponding vowel to obtain enhanced speech information;

[0035] Perform a speech loudness optimization operation on the enhanced speech information by using a preset adaptive compressor to obtain target speech information.

[0036] In a second aspect, the present application provides a speech separation device for an audio signal, including:

[0037] A signal processing module, configured to process the original microphone signal based on the environmental background noise to obtain a to-be-processed microphone signal, and perform position estimation on the to-be-processed microphone signal by using a preset position estimation technique and a confidence-weighted optimization algorithm to obtain a set of three-dimensional coordinates of the target microphone;

[0038] A feature fusion module, configured to process the to-be-processed microphone signal by using a preset feature extraction and fusion rule to obtain a feature fusion result, and process the feature fusion result and the set of three-dimensional coordinates of the target microphone by using a Bayesian nonparametric model and an acoustic generation model to obtain a plurality of target speaker voice signals and three-dimensional coordinates corresponding to each speaker;

[0039] A voice separation information determination module, configured to process each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain an identity label corresponding to each of the target speaker voice signals, and then use a preset psychoacoustic model and a preset separation strategy, and determine voice separation information corresponding to the original microphone signal based on each of the target speaker voice signals, the identity label, and the corresponding three-dimensional coordinates.

[0040] In a third aspect, the present application provides an electronic device, including:

[0041] A memory, configured to store a computer program;

[0042] A processor, configured to execute the computer program to implement the foregoing method for separating voices of audio signals.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, where the computer program, when executed by a processor, implements the foregoing method for separating voices of audio signals.

[0044] As can be seen from the above, before separating the voices of the audio signal in the present application, it is necessary to process the original microphone signal based on the environmental background noise to obtain a to-be-processed microphone signal, and perform position estimation on the to-be-processed microphone signal by using a preset position estimation technique and a confidence-weighted optimization algorithm to obtain a set of three-dimensional coordinates of the target microphone; process the to-be-processed microphone signal by using a preset feature extraction and fusion rule to obtain a feature fusion result, and process the feature fusion result and the set of three-dimensional coordinates of the target microphone by using a Bayesian nonparametric model and an acoustic generation model to obtain a plurality of target speaker voice signals and three-dimensional coordinates corresponding to each speaker; process each target speaker voice signal by using a preset voiceprint enhancement network to obtain an identity label corresponding to each speaker voice signal, and then use a preset psychoacoustic model and a preset separation strategy, and determine voice separation information corresponding to the original microphone signal based on each target speaker voice signal, the identity label, and the corresponding three-dimensional coordinates.

[0045] It can be seen that, first, the present application needs to process the original microphone signal based on the environmental background noise to obtain the microphone signal to be processed, and use the preset position estimation technology and confidence-weighted optimization algorithm to estimate the position of the microphone signal to be processed, so as to obtain the set of three-dimensional coordinates of the target microphone; subsequently, use the preset feature extraction and fusion rules to process the microphone signal to be processed to obtain the feature fusion result, and use the Bayesian non-parametric model and acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain a number of target speaker voice signals and the three-dimensional coordinates corresponding to each speaker; finally, use the preset voiceprint enhancement network to process each target speaker voice signal to obtain the identity label corresponding to each speaker voice signal, and then use the preset psychoacoustic model and preset separation strategy, and determine the voice separation information corresponding to the original microphone signal based on each target speaker voice signal, identity label and the corresponding three-dimensional coordinates. In this way, the efficiency of voice separation of the audio signal is improved during the voice separation process of the audio signal, thereby enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a method for voice separation of an audio signal disclosed in the present application;

[0048] Figure 2 It is a schematic diagram of a specific process for preprocessing and calibrating a signal disclosed in the present application;

[0049] Figure 3 It is a schematic diagram of a specific process for extracting and fusing time-frequency space features disclosed in the present application;

[0050] Figure 4 It is a schematic diagram of a specific voice separation based on deep learning disclosed in the present application;

[0051] Figure 5 It is a schematic diagram of a specific process for tracking and maintaining the identity of a speaker disclosed in the present application;

[0052] Figure 6 It is a schematic diagram of a specific process for post-processing and quality enhancement operations on voice information disclosed in the present application;

[0053] Figure 7 Schematic structural diagram of a voice separation device for an audio signal disclosed in this application;

[0054] Figure 8 Structural diagram of an electronic device disclosed in this application. Specific embodiments

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] The existing voice separation technology in conference scenarios has the following main problems: First, the preset fixed number of speakers is difficult to adapt to the dynamically changing conference scenarios; second, single-domain feature processing cannot make full use of multi-domain information in the time-frequency space; third, there is a lack of an effective speaker identity continuous tracking mechanism; fourth, the quality of the separated voice is insufficient, with spectral distortion and artifacts; fifth, it is difficult to maintain stable performance in complex acoustic environments. For this reason, this application provides a method for separating voices from audio signals, which can improve the efficiency of separating voices from audio signals during the voice separation process of audio signals.

[0057] See Figure 1 As shown, the embodiments of the present invention disclose a method for separating voices from audio signals, including:

[0058] Step S11: Process the original microphone signal based on the environmental background noise to obtain a microphone signal to be processed, and use the preset position estimation technology and confidence-weighted optimization algorithm to perform position estimation on the microphone signal to be processed to obtain a set of three-dimensional coordinates of the target microphone.

[0059] In this embodiment, it is selected to convert the microphone array calibration problem into a position estimation problem based on acoustic fingerprints, and the schematic diagram of the signal preprocessing and calibration process is as Figure 2 shown. That is, use the background noise in the conference room environment as the calibration signal source, without additional calibration procedures, and the specific process is as follows: First, estimate the time delay by analyzing the cross-correlation function between microphone pairs, construct a distance matrix D, and the expression formula for constructing the distance matrix D is as follows:

[0060] ;

[0061] where D[i,j] represents the distance between microphones i and j estimated by signal analysis, represents the speed of sound propagation in air, To obtain the time delay that maximizes the value of function, is the time point, is the microphone at the time point the received signal value, is the microphone at the time point the received signal value.

[0062] Then, the three-dimensional coordinate set of the microphone is reconstructed by solving the optimization problem :

[0063] ;

[0064] wherein, is a regularization parameter, a non-negative scalar, used to balance the importance between the main optimization objective and the regularization term, is a regularization term that depends on the positions of all microphones .

[0065] Furthermore, to improve the calibration stability, the embodiment of the present application introduces a confidence-weighted mechanism to assign different weights to each distance estimate based on the signal-to-noise ratio and the correlation strength, thereby reducing the impact of unreliable estimates.

[0066] Subsequently, the embodiment of the present application needs to design an adaptive noise feature learning algorithm for the specific noise in the conference room, and then identify the non-speech segments through voice activity detection, and update the noise dictionary using the identified non-speech segments:

[0067] ;

[0068] wherein, represents the updated noise pattern or noise dictionary, represents the old noise pattern or noise dictionary before update, is the smoothing coefficient, a decimal between 0 and 1, used to control the influence degree of new information on the update result, represents the input signal frame currently identified as a non-speech segment (i.e., noise), represents an update function used to calculate the update amount of the noise pattern according to the old noise pattern and the current noise frame to represent the process of extracting information from the current noise frame and using it to update the noise model.

[0069] That is, for each input frame, the embodiments of the present application estimate its noise component through sparse representation, and finally suppress the estimated noise through soft spectral subtraction. In this way, through multi-channel noise feature learning, the embodiments of the present application can identify and suppress noise sources with different spatial characteristics.

[0070] Specifically, the original microphone signal is processed based on the environmental background noise to obtain the microphone signal to be processed, and the position estimation technology and confidence weighted optimization algorithm are used to estimate the position of the microphone signal to be processed, and the target microphone three-dimensional coordinate set can be obtained, including: establishing a noise dictionary based on the historical environmental background noise, and performing noise suppression operation on the original microphone signal based on the current environmental background noise and the noise dictionary to obtain an intermediate state microphone signal; determining the acoustic wave propagation delay corresponding to each microphone pair by using the cross-correlation function of each microphone pair and the intermediate state microphone signal, and determining the initial distance matrix based on the product of the acoustic wave propagation delay and the speed of sound; processing each distance value in the initial distance matrix by using a preset confidence weighted mechanism to obtain a target distance matrix, and processing the target distance matrix by using a three-dimensional coordinate reconstruction model and a preset position estimation technology to obtain an initial microphone three-dimensional coordinate set; processing the initial microphone three-dimensional coordinate set by using a preset gradient descent algorithm based on the noise level corresponding to the microphone signal to be processed to obtain the target microphone three-dimensional coordinate set.

[0071] Further, after obtaining the calibrated target microphone three-dimensional coordinate set, the embodiments of the present application need to construct an adaptive beamformer with low computational complexity. That is, a minimum variance distortionless response (MVDR) beamforming algorithm fused with geometric prior is designed, and the beamforming weight vector can be solved by the following optimization formula:

[0072] ;

[0073] Among them, , represents the beamforming weight vector for a specific direction , used to determine the weight vector that can minimize the value of the following expression , is the weight vector of the beamformer, is the vector Hermitian conjugate transpose, is the noise covariance matrix, is the array manifold vector, corresponding to the phase relationship when the sound source signal from the direction arrives at each microphone of the array, which is determined by the geometric structure of the array and the signal frequency.

[0074] It is worth mentioning that, to reduce the computational complexity, the embodiment of the present application uses the block diagonalization method for approximate processing of the inverse operation, so that the computational complexity is reduced from to , enabling the algorithm to run in real time on an embedded platform.

[0075] Step S12: Process the to-be-processed microphone signal by using a preset feature extraction and fusion rule to obtain a feature fusion result, and process the feature fusion result and the set of three-dimensional coordinates of the target microphone by using a Bayesian nonparametric model and an acoustic generation model to obtain a plurality of target speaker voice signals and three-dimensional coordinates corresponding to each speaker.

[0076] In this embodiment, a deep convolutional network with a progressive receptive field is used to capture time-domain features at three scales of short, medium, and long time, and the flow chart of extracting and fusing time-frequency space features is as shown in Figure 3 . Among them, the network architecture adopts a multi-branch design, and each branch is responsible for extracting features at a specific time scale, and then feature fusion is performed through an attention mechanism, and the expression is as follows:

[0077] ;

[0078] Among them, represents the finally fused time-domain feature representation, represents the index of the time scale (short, medium, long), represents the attention weight assigned to the feature from scale , represents the time-domain feature extracted by the network branch responsible for extracting features at a specific time scale .

[0079] In addition, the embodiment of the present application develops a phase-aware frequency-domain feature extraction network to process amplitude and phase information by using the frequency-domain feature extraction network, and the expression of processing amplitude and phase information by using the frequency-domain feature extraction network is as follows:

[0080] ;

[0081] Among them, represents the complex feature output by a complex-valued convolutional neural network (ComplexCNN), and simultaneously contains amplitude and phase information, represents the functional operation of the complex-valued convolutional neural network, represents the input complex-valued time-frequency information, where is the time frame index, is the frequency index.

[0082] In addition, the embodiments of the present application design a phase continuity constraint to ensure that the extracted phase features conform to physical laws. It is worth mentioning that compared with the method that only uses amplitude information, the method adopted in the embodiments of the present application improves the word recognition rate by 15% in the speech clarity test.

[0083] Furthermore, the embodiments of the present application design a spatial feature network based on self-attention to perform extraction operations on spatial features and model the obtained extraction results as a graph learning problem. That is, the microphones are regarded as the nodes of the graph, and the correlation is used as the weight of the edge. Subsequently, the spatial features are extracted through a graph attention network to capture the microphone topology relationship in the non-Euclidean space, and the importance of different microphone channels is adaptively adjusted through the attention mechanism. Moreover, the embodiments of the present application design a cross-domain feature adaptive fusion network to calculate the mutual influence between domains through a domain interaction attention module, and finally fuse the original features and the interaction features through a gating mechanism. It is worth mentioning that the weights can be dynamically adjusted according to the current acoustic environment. Compared with simple splicing, the method adopted in the embodiments of the present application improves the separation accuracy by 18% on average in different acoustic environments.

[0084] Specifically, using a preset feature extraction and fusion rule to process the microphone signal to be processed to obtain a feature fusion result may include: using a preset attention mechanism to establish a deep convolutional network including a progressive receptive field, and using the deep convolutional network to perform acoustic feature extraction operations on the microphone signal to be processed to obtain corresponding first acoustic features, second acoustic features, and third acoustic features; the time scales corresponding to the first acoustic features, second acoustic features, and third acoustic features increase sequentially; using a preset attention mechanism to perform fusion operations on the first acoustic features, second acoustic features, and third acoustic features to obtain an acoustic feature fusion result; using a preset complex-valued neural network and a preset phase continuity constraint to perform processing operations on the frequency domain amplitude and phase information of the acoustic feature fusion result to obtain time domain features and frequency domain features, and establishing a graph attention network based on the microphone array corresponding to the microphone signal to be processed; wherein, the nodes in the graph attention network are the microphones in the microphone array, and the edges in the graph attention network are the correlations between the microphones; using the graph attention network and based on the microphone signal to be processed to determine spatial topological features, and fusing the time domain features, frequency domain features, and spatial topological features to obtain a feature fusion result.

[0085] In this embodiment, after obtaining the feature fusion result, it is necessary to determine the corresponding speech separation information based on the feature fusion result, and the schematic diagram of speech separation based on deep learning is as Figure 4As shown below. First, the embodiment of the present application uses a Bayesian non-parametric model to automatically estimate the number of currently active speakers. That is, the Dirichlet process mixture model is used to cluster the feature space, and pitch is introduced as an auxiliary clustering feature:

[0086] ;

[0087] where, represents the posterior probability that the nth data point (feature vector) belongs to the cluster (speaker) c given the cluster assignment of the previous data points and the pitch information of the current data point of the nth data point; Here, denotes "proportional to", meaning that the left side of the symbol is equal to the right side multiplied by a normalization constant, denotes the cluster assignment of the nth data point / feature vector (i.e., which speaker it belongs to), represents the index of a specific cluster (speaker), denotes the cluster assignments from the 1st to the nth data points, denotes the fundamental frequency associated with the nth data point / feature vector, i.e., pitch, denotes the prior probability of assigning the nth point to the cluster c given the assignments of the previous data points, and denotes the likelihood probability of observing the pitch f0 given that the data point belongs to the cluster (speaker) c. It is worth mentioning that the embodiment of the present application can automatically adapt to the dynamically changing number of speakers, does not require a preset upper limit on the number of speakers, and improves the discrimination ability of speakers with similar voices by introducing a pitch prior. Subsequently, the embodiment of the present application combines two complementary separation strategies, time-frequency masking and spatial filtering, to design a reliability-weighted bidirectional fusion mechanism, and the expression is as follows: where, denotes the time-frequency representation of the speech signal finally reconstructed (separated) for speaker s at the time-frequency point (t, f), denotes the reliability weight assigned to the separation result based on the time-frequency mask for speaker s at the time-frequency point (t, f), and

[0088] ;

[0089] denotes the time-frequency representation of the speech signal finally reconstructed (separated) for speaker s at the time-frequency point (t, f).

[0090] ;

[0091] where, denotes the reliability weight assigned to the separation result based on the time-frequency mask for speaker s at the time-frequency point (t, f), and denotes the time-frequency representation of the speech signal finally reconstructed (separated) for speaker s at the time-frequency point (t, f). The reliability weight for the speaker at the time-frequency point of the speech component. Indicates the speaker estimated using the time-frequency masking method at the time-frequency point of the speech component. Indicates the reliability weight assigned to the result of spatial filtering separation for the speaker at the time-frequency point of the speech component, Indicates the speaker estimated using the spatial filtering method at the time-frequency point of the speech component.

[0092] It is worth mentioning that the embodiments of the present application combine the high speech quality of time-frequency masking and the interference suppression ability of spatial filtering, so as to adaptively adjust the contributions of the two methods through reliability weighting, and can maintain stable separation performance in different acoustic environments. In addition, the embodiments of the present application design a spectral compensation network based on an acoustic generation model to repair the lost speech details during the separation process using the spectral compensation network. That is, a speech prior model based on a variational autoencoder is designed to capture the spectral structure characteristics corresponding to natural speech, so as to infer and compensate for the lost spectral details according to the intrinsic characteristics of the speech and maintain the naturalness and coherence of the speech.

[0093] Specifically, using a Bayesian nonparametric model and an acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone, several target speaker speech signals and the corresponding three-dimensional coordinates of each speaker can be obtained, including: using a Dirichlet process mixture model to cluster the feature space corresponding to the feature fusion result to obtain the pitch features corresponding to the feature fusion result; using a Bayesian nonparametric model to determine the number of speakers based on the feature fusion result and the pitch features to obtain the number of speakers corresponding to the microphone signal to be processed; determining a two-way fusion mechanism based on a preset separation strategy and a preset weighting mechanism, and then using the two-way fusion mechanism and based on a preset index to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain several initial speaker speech signals and the corresponding three-dimensional coordinates of each speaker; the preset separation strategy includes time-frequency masking and spatial filtering; determining a spectral compensation network based on an acoustic generation model and a speech prior model including a variational autoencoder, and using the spectral compensation network to determine the spectral structure characteristics of the initial speaker speech signal to obtain the spectral structure characteristics, and then determining the target speaker speech signal based on the spectral structure characteristics and the initial speaker speech signal.

[0094] Step S13: Process each of the target speaker voice signals using a preset voiceprint enhancement network to obtain identity tags corresponding to the target speaker voice signals, and then use a preset psychoacoustic model and a preset separation strategy, and determine voice separation information corresponding to the original microphone signal based on the target speaker voice signals, the identity tags, and the corresponding three-dimensional coordinates.

[0095] In this embodiment, the flow diagram for tracking and maintaining the identity of the speaker is as Figure 5 shown, and the embodiments of the present application design a lightweight voiceprint feature extractor for short speech segments to improve the discriminability of short speech voiceprints through adversarial training, and the expressions are as shown below respectively:

[0096] ;

[0097] Among them, represents the short-time voiceprint feature vector extracted for the th speaker / speech segment, represents the encoder network (usually a deep neural network) used to extract voiceprint features, is the voice data input to the encoder, which is the feature representation separated for the th speaker.

[0098] ;

[0099] Among them, is the loss function of the discriminator network, represents the expected value (usually approximated by taking the average over a batch of data). is for the output of the discriminator network, is the voiceprint vector sampled from the true, target voiceprint distribution. is the voiceprint vector generated by the encoder based on the speaker .

[0100] ;

[0101] Among them, is the loss function of the encoder network, is used to indicate that this part of the loss encourages the encoder to generate voiceprints that can "fool" the discriminator , that is, to make the discriminator output a high probability (close to 1) for . is the weight coefficient, is used to calculate the triplet loss on the generated voiceprint .

[0102] Subsequently, the embodiments of the present application designed a multi-modal identity tracking algorithm that combines voiceprint, location, and speech content, and fused the above three features with adaptive weights, and the expression is as follows:

[0103] ;

[0104] Among them, is the overall similarity score between the current speech segment and the speaker being tracked ; , and are the adaptive weights corresponding to the three modal features of voiceprint, position, and content respectively, is the similarity between the speaker and the speech segment calculated based on the voiceprint feature, is the similarity between the speaker and the speech segment calculated based on the estimated spatial position information, is the similarity between the speaker and the speech segment calculated based on the speech content (such as language features, topics, etc.).

[0105] In addition, the embodiments of the present application designed a speaker state tracking algorithm based on particle filter, and then used an adaptive memory model to dynamically adjust the feature update rate based on the speaker characteristics, so as to be able to track the slow changes of the speaker features using the adaptive memory model, and effectively handle the situation of intermittent speech of the speaker by balancing stability and adaptability.

[0106] Specifically, using a preset voiceprint enhancement network to process the speech signals of each target speaker to obtain identity labels corresponding to the speech signals of each target speaker may include: using the preset voiceprint enhancement network to perform feature extraction operations on the speech signals of each target speaker to obtain voiceprint features, speaker spatial position features, and speech content features corresponding to the speech signals of the target speaker; establishing a modal similarity metric based on the voiceprint features, speaker spatial position features, and speech content features, and then using the adaptive memory model to process the speaker features corresponding to the speech signals of the target speaker based on the modal similarity metric to obtain a processing result; determining the activity level corresponding to the speech signals of the target speaker, and using the speaker state tracking algorithm based on particle filter to perform identity tracking on the identity of the speaker based on the activity level and the processing result to obtain identity labels corresponding to the speech signals of each target speaker.

[0107] In this embodiment, after obtaining the identity tags corresponding to the voice signals of each target speaker, the embodiments of the present application need to perform post-processing and quality enhancement operations on the obtained voice information, and the schematic flow chart of performing post-processing and quality enhancement operations on the voice information is as shown in Figure 6 the figure. First, a blind de-reverberation algorithm based on a neural network is designed, and the voice is decomposed into direct sound and reverberation components using a dry-wet sound decomposition model, and only the reverberation components are selectively suppressed to retain early reflections. In addition, to avoid the "dry" effect caused by overprocessing, the embodiments of the present application introduce perceptual constraints to ensure natural listening.

[0108] Furthermore, the embodiments of the present application design an inter-spectral consistency restoration network, that is, using the spectral repair algorithm in the psychoacoustic model and considering the sensitivity of the human ear to different frequency bands, introducing harmonic structure constraints to ensure the naturalness of the repair.

[0109] Specifically, using a preset psychoacoustic model and a preset separation strategy, and determining the voice separation information corresponding to the original microphone signal based on the voice signals, identity tags, and corresponding three-dimensional coordinates of each target speaker may include: determining a blind de-reverberation algorithm based on a preset neural network, and using the blind de-reverberation algorithm and a preset dry-wet sound decomposition model to decompose the voice signal of the target speaker to obtain a direct sound signal and a reverberation component signal; performing an inhibition operation on the reverberation component signal using a preset inhibition rule to obtain an inhibited signal, and determining a signal to be repaired based on the inhibited signal and the direct sound signal; determining harmonic structure constraint conditions based on the sensitivity of the user to several sound frequency bands respectively, and performing a repair operation on the signal to be repaired based on the harmonic structure constraint conditions and the spectral repair algorithm in the preset psychoacoustic model to obtain voice information corresponding to each speaker; determining the voice separation information corresponding to the original microphone signal based on the voice information, identity tags, and three-dimensional coordinates corresponding to each speaker.

[0110] It is worth mentioning that the embodiments of the present application design naturalness enhancement and dynamic range processing for conference voice, that is, designing a selective enhancement operation based on the pronunciation part to adopt different enhancement strategies for different vowels, and finally applying a multi-band adaptive compressor to optimize the voice loudness. Specifically, after determining the voice separation information corresponding to the original microphone signal based on the voice signals, identity tags, and corresponding three-dimensional coordinates of each target speaker, it may further include: determining the vowels corresponding to each pronunciation part based on the pronunciation part corresponding to the voice separation information, and determining the corresponding enhancement strategies based on each vowel, and performing an enhancement processing operation on the corresponding vowel using each enhancement strategy to obtain enhanced voice information; performing a voice loudness optimization operation on the enhanced voice information using a preset adaptive compressor to obtain target voice information.

[0111] As can be seen from the above, in the embodiments of the present application, it is first necessary to process the original microphone signal based on the environmental background noise to obtain a microphone signal to be processed, and use a preset position estimation technique and a confidence-weighted optimization algorithm to estimate the position of the microphone signal to be processed, so as to obtain a set of three-dimensional coordinates of the target microphone. Subsequently, use the preset feature extraction and fusion rules to process the microphone signal to be processed to obtain a feature fusion result, and use a Bayesian nonparametric model and an acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain a plurality of target speaker voice signals and the three-dimensional coordinates corresponding to each speaker. Finally, use a preset voiceprint enhancement network to process each target speaker voice signal to obtain an identity label corresponding to each target speaker voice signal, and then use a preset psychoacoustic model and a preset separation strategy, and determine the voice separation information corresponding to the original microphone signal based on each target speaker voice signal, the identity label, and the corresponding three-dimensional coordinates. In this way, the efficiency of voice separation of the audio signal is improved during the voice separation process of the audio signal.

[0112] Correspondingly, as shown in Figure 7 the present application also provides a voice separation device for an audio signal, including:

[0113] A signal processing module 11, configured to process the original microphone signal based on the environmental background noise to obtain a microphone signal to be processed, and use a preset position estimation technique and a confidence-weighted optimization algorithm to estimate the position of the microphone signal to be processed, so as to obtain a set of three-dimensional coordinates of the target microphone;

[0114] A feature fusion module 12, configured to use preset feature extraction and fusion rules to process the microphone signal to be processed to obtain a feature fusion result, and use a Bayesian nonparametric model and an acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone to obtain a plurality of target speaker voice signals and the three-dimensional coordinates corresponding to each speaker;

[0115] A voice separation information determination module 13, configured to use a preset voiceprint enhancement network to process each target speaker voice signal to obtain an identity label corresponding to each target speaker voice signal, and then use a preset psychoacoustic model and a preset separation strategy, and determine the voice separation information corresponding to the original microphone signal based on each target speaker voice signal, the identity label, and the corresponding three-dimensional coordinates.

[0116] As can be seen from the above, before performing voice separation on the audio signal in the embodiment of the present application, it is first necessary to process the original microphone signal based on the environmental background noise to obtain the microphone signal to be processed, and use the preset position estimation technology and the confidence weighted optimization algorithm to perform position estimation on the microphone signal to be processed to obtain the target microphone three-dimensional coordinate set; subsequently, use the preset feature extraction and fusion rules to process the microphone signal to be processed to obtain the feature fusion result, and use the Bayesian nonparametric model and the acoustic generation model to process the feature fusion result and the target microphone three-dimensional coordinate set to obtain a plurality of target speaker voice signals and the three-dimensional coordinates corresponding to each speaker; finally, use the preset voiceprint enhancement network to process each target speaker voice signal to obtain the identity label corresponding to each speaker voice signal, and then use the preset psychoacoustic model and the preset separation strategy, and determine the voice separation information corresponding to the original microphone signal based on each target speaker voice signal, identity label, and the corresponding three-dimensional coordinates. In this way, the efficiency of voice separation of the audio signal is improved during the voice separation process of the audio signal.

[0117] In some specific embodiments, the signal processing module 11 may specifically include:

[0118] The microphone signal suppression unit is configured to establish a noise dictionary based on the historical environmental background noise, and perform a noise suppression operation on the original microphone signal based on the current environmental background noise and the noise dictionary to obtain an intermediate microphone signal;

[0119] The distance matrix determination unit is configured to determine the acoustic wave propagation delay corresponding to each microphone pair by using the cross-correlation function of each microphone pair and the intermediate microphone signal, and determine the initial distance matrix based on the product of the acoustic wave propagation delay and the speed of sound;

[0120] The distance matrix processing unit is configured to process each distance value in the initial distance matrix by using a preset confidence weighting mechanism to obtain a target distance matrix, and process the target distance matrix by using a three-dimensional coordinate reconstruction model and a preset position estimation technology to obtain an initial microphone three-dimensional coordinate set;

[0121] The coordinate set processing unit is configured to process the initial microphone three-dimensional coordinate set by using a preset gradient descent algorithm based on the noise level corresponding to the microphone signal to be processed to obtain a target microphone three-dimensional coordinate set.

[0122] In some specific embodiments, the feature fusion module 12 may specifically include:

[0123] An acoustic feature determination unit, configured to establish a deep convolutional network including a progressive receptive field by using a preset attention mechanism, and perform an acoustic feature extraction operation on the to-be-processed microphone signal by using the deep convolutional network to obtain corresponding first acoustic features, second acoustic features, and third acoustic features; the time scales corresponding to the first acoustic features, the second acoustic features, and the third acoustic features increase in sequence;

[0124] A first feature fusion subunit, configured to perform a fusion operation on the first acoustic features, the second acoustic features, and the third acoustic features by using the preset attention mechanism to obtain an acoustic feature fusion result;

[0125] A graph attention network establishment unit, configured to perform processing operations on the frequency-domain amplitude and phase information of the acoustic feature fusion result by using a preset complex-valued neural network and a preset phase continuity constraint to obtain time-domain features and frequency-domain features, and establish a graph attention network based on the microphone array corresponding to the to-be-processed microphone signal; wherein, the nodes in the graph attention network are the microphones in the microphone array, and the edges in the graph attention network are the correlations between the microphones;

[0126] A second feature fusion subunit, configured to determine a spatial topology feature by using the graph attention network and based on the to-be-processed microphone signal, and fuse the time-domain features, the frequency-domain features, and the spatial topology feature to obtain a feature fusion result.

[0127] In some specific embodiments, the feature fusion module 12 may specifically include:

[0128] A pitch feature determination unit, configured to cluster the feature space corresponding to the feature fusion result by using a Dirichlet process mixture model to obtain a pitch feature corresponding to the feature fusion result;

[0129] A speaker number determination unit, configured to perform a speaker number determination operation based on the feature fusion result and the pitch feature by using a Bayesian nonparametric model to obtain the number of speakers corresponding to the to-be-processed microphone signal;

[0130] A three-dimensional coordinate determination unit, configured to determine a bidirectional fusion mechanism based on a preset separation strategy and a preset weighting mechanism, and then use the bidirectional fusion mechanism and process the feature fusion result and the target microphone three-dimensional coordinate set based on a preset index to obtain a plurality of initial speaker voice signals and three-dimensional coordinates corresponding to each speaker; the preset separation strategy includes time-frequency masking and spatial filtering;

[0131] A voice signal determination unit, configured to determine a spectral compensation network based on an acoustic generation model and a voice prior model including a variational autoencoder, so as to use the spectral compensation network to determine the spectral structure characteristics of the initial speaker voice signal, obtain the spectral structure characteristics, and then determine the target speaker voice signal based on the spectral structure characteristics and the initial speaker voice signal.

[0132] In some specific embodiments, the voice separation information determination module 13 may specifically include:

[0133] A feature extraction unit, configured to perform a feature extraction operation on each of the target speaker voice signals by using a preset voiceprint enhancement network, to obtain a voiceprint feature, a speaker spatial position feature, and a voice content feature corresponding to the target speaker voice signal;

[0134] A feature processing unit, configured to establish a modal similarity metric based on the voiceprint feature, the speaker spatial position feature, and the voice content feature, and then use an adaptive memory model and process the speaker feature corresponding to the target speaker voice signal based on the modal similarity metric, to obtain a processing result;

[0135] An identity label determination unit, configured to determine the activity level corresponding to the target speaker voice signal, and perform identity tracking on the identity of the speaker by using a speaker state tracking algorithm based on particle filtering for the activity level and the processing result, to obtain an identity label corresponding to each of the target speaker voice signals.

[0136] In some specific embodiments, the voice separation information determination module 13 may specifically include:

[0137] A voice signal decomposition unit, configured to determine a blind de-reverberation algorithm based on a preset neural network, and decompose the target speaker voice signal by using the blind de-reverberation algorithm and a preset dry-wet sound decomposition model, to obtain a direct sound signal and a reverberation component signal;

[0138] A signal to be repaired determination unit, configured to perform a suppression operation on the reverberation component signal by using a preset suppression rule, to obtain a suppressed signal, and determine a signal to be repaired based on the suppressed signal and the direct sound signal;

[0139] A signal repair unit, configured to determine a harmonic structure constraint condition based on the sensitivity of the user to several voice frequency bands respectively, and perform a repair operation on the signal to be repaired based on the harmonic structure constraint condition and a spectral repair algorithm in a preset psychoacoustic model, to obtain voice information corresponding to each of the speakers;

[0140] A voice separation information determination unit, configured to determine voice separation information corresponding to the original microphone signal based on the voice information corresponding to each of the speakers, the identity tag, and the three-dimensional coordinates.

[0141] In some specific embodiments, the voice separation device of the audio signal may further include:

[0142] A voice information enhancement unit, configured to determine vowels corresponding to each of the pronunciation parts based on the pronunciation parts corresponding to the voice separation information, determine corresponding enhancement strategies based on each of the vowels, and use each of the enhancement strategies to perform an enhancement processing operation on the corresponding vowels to obtain enhanced voice information;

[0143] A voice information optimization unit, configured to perform a voice loudness optimization operation on the enhanced voice information by using a preset adaptive compressor to obtain target voice information.

[0144] Furthermore, an embodiment of the present application also discloses an electronic device. Figure 8 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation to the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the voice separation method of the audio signal disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0145] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0146] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0147] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of implementing the voice separation method of the audio signal executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0148] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the voice separation method of the audio signal disclosed above. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0149] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method part.

[0150] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0151] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0152] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0153] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for separating speech from an audio signal, characterized in that, Including: Processing the original microphone signal based on the environmental background noise to obtain a microphone signal to be processed, and using a preset position estimation technique and a confidence weighted optimization algorithm to perform position estimation on the microphone signal to be processed, so as to obtain a set of three-dimensional coordinates of the target microphone; Processing the microphone signal to be processed by using a preset feature extraction and fusion rule to obtain a feature fusion result, and using a Bayesian nonparametric model and an acoustic generation model to process the feature fusion result and the set of three-dimensional coordinates of the target microphone, so as to obtain a plurality of target speaker voice signals and three-dimensional coordinates corresponding to each speaker; Processing each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain an identity label corresponding to each of the target speaker voice signals, and then using a preset psychoacoustic model and a preset separation strategy, and determining voice separation information corresponding to the original microphone signal based on each of the target speaker voice signals, the identity label, and the corresponding three-dimensional coordinates.

2. The method for separating speech of an audio signal according to claim 1, wherein The processing the original microphone signal based on the environmental background noise to obtain a microphone signal to be processed, and using a preset position estimation technique and a confidence weighted optimization algorithm to perform position estimation on the microphone signal to be processed, so as to obtain a set of three-dimensional coordinates of the target microphone, includes: Establishing a noise dictionary based on historical environmental background noise, and performing a noise suppression operation on the original microphone signal based on the current environmental background noise and the noise dictionary to obtain an intermediate microphone signal; Using the cross-correlation function corresponding to each microphone pair and the intermediate microphone signal to determine the acoustic wave propagation time delay corresponding to each microphone pair, and determining an initial distance matrix based on the product of the acoustic wave propagation time delay and the speed of sound; Processing each distance value in the initial distance matrix by using a preset confidence weighted mechanism to obtain a target distance matrix, and using a three-dimensional coordinate reconstruction model and a preset position estimation technique to process the target distance matrix to obtain a set of initial three-dimensional coordinates of the microphone; Processing the set of initial three-dimensional coordinates of the microphone by using a preset gradient descent algorithm and based on the noise level corresponding to the microphone signal to be processed to obtain a set of three-dimensional coordinates of the target microphone.

3. The method for separating speech from an audio signal according to claim 1, wherein The processing the microphone signal to be processed by using a preset feature extraction and fusion rule to obtain a feature fusion result, includes: Establishing a deep convolutional network including a progressive receptive field by using a preset attention mechanism, and performing an acoustic feature extraction operation on the microphone signal to be processed by using the deep convolutional network to obtain corresponding first acoustic features, second acoustic features, and third acoustic features; the time scales corresponding to the first acoustic features, the second acoustic features, and the third acoustic features increase in sequence; Performing a fusion operation on the first acoustic features, the second acoustic features, and the third acoustic features by using the preset attention mechanism to obtain an acoustic feature fusion result; Perform processing operations on the frequency-domain amplitude and phase information of the acoustic feature fusion result by using a preset complex-valued neural network and a preset phase continuity constraint to obtain time-domain features and frequency-domain features, and establish a graph attention network based on the microphone array corresponding to the microphone signal to be processed; wherein, the nodes in the graph attention network are the microphones in the microphone array, and the edges in the graph attention network are the correlations between the microphones. Use the graph attention network and determine spatial topology features based on the microphone signal to be processed, and fuse the time-domain features, the frequency-domain features, and the spatial topology features to obtain a feature fusion result.

4. The method for separating speech from an audio signal according to claim 1, wherein, Process the feature fusion result and the set of three-dimensional coordinates of the target microphone by using a Bayesian nonparametric model and an acoustic generation model to obtain a plurality of target speaker voice signals and the three-dimensional coordinates corresponding to each speaker, including: Cluster the feature space corresponding to the feature fusion result by using a Dirichlet process mixture model to obtain pitch features corresponding to the feature fusion result. Perform an operation to determine the number of speakers corresponding to the microphone signal to be processed by using a Bayesian nonparametric model based on the feature fusion result and the pitch features to obtain the number of speakers corresponding to the microphone signal to be processed. Determine a bidirectional fusion mechanism based on a preset separation strategy and a preset weighting mechanism, and then use the bidirectional fusion mechanism and process the feature fusion result and the set of three-dimensional coordinates of the target microphone based on a preset index to obtain a plurality of initial speaker voice signals and the three-dimensional coordinates corresponding to each speaker; the preset separation strategy includes time-frequency masking and spatial filtering. Determine a spectral compensation network based on an acoustic generation model and a voice prior model including a variational autoencoder, so as to use the spectral compensation network to determine the spectral structure characteristics of the initial speaker voice signal to obtain spectral structure characteristics, and then determine the target speaker voice signal based on the spectral structure characteristics and the initial speaker voice signal.

5. The method for separating speech of an audio signal according to claim 1, characterized in that Process each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain an identity label corresponding to each of the target speaker voice signals, including: Perform a feature extraction operation on each of the target speaker voice signals by using a preset voiceprint enhancement network to obtain voiceprint features, speaker spatial position features, and speech content features corresponding to the target speaker voice signals. Establish a modal similarity metric based on the voiceprint features, the speaker spatial position features, and the speech content features, and then use an adaptive memory model and process the speaker features corresponding to the target speaker voice signal based on the modal similarity metric to obtain a processing result. Determine the activity level corresponding to the target speaker voice signal, and use a speaker state tracking algorithm based on particle filtering to perform identity tracking on the identity of the speaker based on the activity level and the processing result to obtain an identity label corresponding to each of the target speaker voice signals.

6. The method for separating voices of an audio signal according to claim 1, characterized in that The method for determining speech separation information corresponding to the original microphone signal by using a preset psychoacoustic model and a preset separation strategy and based on the speech signals of each target speaker, the identity tags, and the corresponding three-dimensional coordinates includes: Determining a blind de-reverberation algorithm based on a preset neural network, and using the blind de-reverberation algorithm and a preset dry-wet sound decomposition model to decompose the speech signal of the target speaker to obtain a direct sound signal and a reverberation component signal; Performing an inhibition operation on the reverberation component signal by using a preset inhibition rule to obtain an inhibited signal, and determining a signal to be repaired based on the inhibited signal and the direct sound signal; Determining a harmonic structure constraint condition based on the sensitivity of the user to several sound frequency bands respectively, and performing a repair operation on the signal to be repaired by using the harmonic structure constraint condition and a spectral repair algorithm in the preset psychoacoustic model to obtain speech information corresponding to each speaker; Determining speech separation information corresponding to the original microphone signal based on the speech information corresponding to each speaker, the identity tags, and the three-dimensional coordinates.

7. The method for separating voices of an audio signal according to any one of claims 1 to 6, characterized in that, After determining the speech separation information corresponding to the original microphone signal based on the speech signals of each target speaker, the identity tags, and the corresponding three-dimensional coordinates, it further includes: Determining vowels corresponding to each pronunciation part based on the pronunciation part corresponding to the speech separation information, determining a corresponding enhancement strategy based on each vowel, and performing an enhancement processing operation on the corresponding vowel by using each enhancement strategy to obtain enhanced speech information; Performing a speech loudness optimization operation on the enhanced speech information by using a preset adaptive compressor to obtain target speech information.

8. A voice separation device for an audio signal, characterized in that, It includes: A signal processing module, configured to process the original microphone signal based on environmental background noise to obtain a microphone signal to be processed, and perform a position estimation on the microphone signal to be processed by using a preset position estimation technique and a confidence weighted optimization algorithm to obtain a set of target microphone three-dimensional coordinates; A feature fusion module, configured to process the microphone signal to be processed by using a preset feature extraction and fusion rule to obtain a feature fusion result, and process the feature fusion result and the set of target microphone three-dimensional coordinates by using a Bayesian nonparametric model and an acoustic generation model to obtain several target speaker speech signals and three-dimensional coordinates corresponding to each speaker; A speech separation information determination module, configured to process the speech signal of each target speaker by using a preset voiceprint enhancement network to obtain an identity tag corresponding to the speech signal of each target speaker, and then use a preset psychoacoustic model and a preset separation strategy, and determine speech separation information corresponding to the original microphone signal based on the speech signals of each target speaker, the identity tags, and the corresponding three-dimensional coordinates.

9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the method for separating speech of an audio signal according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein when the computer program is executed by a processor, it implements the method for separating voices of audio signals according to any one of claims 1 to 7.