A speech separation method, device, apparatus and storage medium
By fusing audio features with voiceprint features, emotional features, and deep features in the speech separation system, and using the GALR network for separation, the problem of low accuracy in real-world scenarios of the speech separation system is solved, achieving higher accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AUTOMOBILE RES INST OF TSINGHUA UNIV IN SUZHOU XIANGCHENG
- Filing Date
- 2022-10-25
- Publication Date
- 2026-04-10
AI Technical Summary
Existing deep learning-based speech separation systems suffer from low accuracy, instability, and poor anti-interference capabilities in real-world scenarios due to the availability of visual features.
By fusing audio features with voiceprint features, emotional features, and deep features, deep learning algorithms are used to extract and adjust multi-dimensional features of speech data, and the GALR network is used for separation, thereby improving the accuracy and stability of speech separation.
While improving the accuracy of speech separation, it also enhances the system's anti-interference capabilities and the stability of the separation effect.
Smart Images

Figure CN115662463B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech data processing, and in particular to a speech separation method, device, equipment and storage medium. BACKGROUND
[0002] The speech separation system based on deep learning usually adopts an end-to-end structure of an encoder, a separator and a decoder. The potential features in the mixed speech are extracted through the encoder, which is difficult to completely and accurately represent the necessary features in the mixed speech. Considering the multi-modal features for speech separation is a good solution.
[0003] At present, the speech separation system extracts the visual features of the speaker's expression and mouth shape and the like to assist in improving the speech separation performance. However, the use of visual features requires the support of a large amount of image data, which is very costly. In addition, in actual scenes, real-time collected images are easily unusable due to reasons such as occlusion and insufficient light, and even affect the quality of speech separation, greatly limiting the application of the speech separation system in real scenes. Therefore, how to add targeted auxiliary features to improve the accuracy of speech separation has become a problem to be solved. SUMMARY
[0004] The present application provides a speech separation method, device, equipment and storage medium to solve the problem of low speech separation accuracy, which can improve the accuracy of speech separation while ensuring the stability and anti-interference of the separation effect.
[0005] According to an aspect of the present application, a speech separation method is provided, the method comprising:
[0006] obtaining speech data to be separated, determining audio features of the speech data to be separated and auxiliary features of the speech data to be separated; wherein the auxiliary features include voiceprint features, emotional features and deep features;
[0007] determining fusion features according to the audio features and the auxiliary features;
[0008] determining a separation result of the speech data to be separated according to the fusion features and the audio features.
[0009] According to another aspect of the present application, a speech separation device is provided, the device comprising:
[0010] a feature determination module configured to obtain speech data to be separated, determine audio features of the speech data to be separated and auxiliary features of the speech data to be separated; wherein the auxiliary features include voiceprint features, emotional features and deep features;
[0011] a fusion feature determination module configured to determine fusion features according to the audio features and the auxiliary features;
[0012] a separation result determination module, configured to determine a separation result of the to-be-separated speech data according to the fusion feature and the audio feature.
[0013] According to another aspect of the present application, an electronic device is provided, which comprises:
[0014] at least one processor; and
[0015] a memory in communication with the at least one processor; wherein
[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the speech separation method according to any one of the embodiments of the present application.
[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the speech separation method according to any one of the embodiments of the present application when executed by the processor.
[0018] The technical solution of the embodiments of the present application fuses auxiliary features such as voiceprint features, emotion features and deep features on the basis of audio features to solve the problem of low speech separation accuracy, and can improve the speech separation accuracy while ensuring the stability and anti-interference of the separation effect.
[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0021] Figure 1 is a flowchart of a speech separation method according to an embodiment of the present application;
[0022] Figure 2 is a flowchart of a speech separation method according to an embodiment of the present application;
[0023] Figure 3 is a structural schematic diagram of a speech separation device according to an embodiment of the present application;
[0024] Figure 4 is a structural schematic diagram of an electronic device for implementing a voice separation method according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should belong to the scope of protection of the present application.
[0026] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing and the like of data in the technical solutions of the present application all comply with the relevant provisions of national laws and regulations.
[0027] Embodiment one
[0028] Figure 1 A flowchart of a voice separation method is provided for the first embodiment of the present application. The present embodiment can be applied to the scene of mixed voice separation. The method can be performed by a voice separation device, which can be realized in the form of hardware and / or software. The device can be configured in an electronic device. As shown in the figure, the method comprises: Figure 1
[0029] S110, obtaining voice data to be separated, determining audio features of the voice data to be separated and auxiliary features of the voice data to be separated.
[0030] The scheme can be executed by a speech separation system, which can include a microphone, a microphone, and other speech collection devices for collecting speech data to be separated in a deployment scene. The speech separation system can also directly read the generated speech data to be separated. The speech data to be separated can be mixed speech data, including speech data of at least two speakers. After obtaining the speech data to be separated, the speech separation system can perform pre-emphasis, framing, windowing, transformation, and filtering on the speech data to be separated, and extract audio features of the speech data to be separated. The speech separation system can also extract audio features based on a deep learning algorithm. In a preferred scheme, the speech separation system can encode the speech data to be separated into a frame-based multi-dimensional embedding sequence through an audio encoder.
[0031] The speech separation system can also extract auxiliary features of the speech data to be separated, which can be other features of the speech data to be separated in addition to the audio features, such as voiceprint features, emotion features, deep features, and timing features. In this scheme, optionally, the voiceprint features are determined by a voiceprint feature extractor; the emotion features are determined based on the spectral information of the speech data to be separated; and the deep features are obtained by domain conversion based on the features of the speech data to be separated output by a deep feature extractor.
[0032] The speech separation system can use the speaker dialization open source toolkit pyannote as a voiceprint feature extractor to extract speaker embedding vectors as voiceprint features of the speech data to be separated. After obtaining the voiceprint features, the speech separation system can encode the voiceprint features through a one-dimensional convolutional neural network to facilitate feature fusion.
[0033] The emotional feature is a feature that is difficult to quantitatively describe. In order to comprehensively describe the emotional feature of the to-be-separated speech data, the speech separation system can perform signal processing operations such as framing, windowing, transformation, and filtering on the to-be-separated speech data, and then determine the spectral information of the to-be-separated speech data. The spectral information can include an average value of spectral flatness, an average value of an audio time series zero-crossing rate, a p-order spectral bandwidth, amplitude statistical data, spectral centroid statistical data, tuning offset statistical data, root mean square statistical data, and mel-frequency cepstral coefficient feature statistical data, and the like. Among them, the average value of spectral flatness, the average value of the audio time series zero-crossing rate, and the p-order spectral bandwidth can be used to represent the spectral characteristics of the to-be-separated speech data. The spectral centroid statistical data and the like can be used to measure the brightness feature of the to-be-separated speech data. The root mean square statistical data and the like can be used to describe the loudness feature of the to-be-separated speech data. In addition to the static features such as spectral characteristics, brightness features, and loudness features, the speech separation system can also describe the dynamic features of the to-be-separated speech data through the first-order difference and second-order difference of the mel-frequency cepstral coefficient features.
[0034] It is easy to understand that the deep feature can be a deeper feature of the to-be-separated speech data compared with the audio feature. The speech separation system can use a pre-trained speech model such as WavLM as a deep feature extractor to extract deep features from the to-be-separated speech data, and then adaptively adjust the deep features extracted by the pre-trained model through a domain conversion network to meet the application of the current speech separation scene. The speech separation system can also extract the time sequence features of the to-be-separated speech data through a recurrent convolutional neural network to obtain information about the timing of different speakers, and then assist in the accurate separation of the to-be-separated speech data.
[0035] S120, determining a fusion feature according to the audio feature and the auxiliary feature.
[0036] The audio feature and the auxiliary feature can have the same dimension, and the speech separation system can directly combine and splice the audio feature and the auxiliary feature to obtain the fusion feature. If the dimensions of the audio feature and the auxiliary feature are different, the speech separation system can adjust the dimensions of the audio feature and the auxiliary feature to be consistent through feature screening, feature replication, or the like, so as to facilitate combination and splicing.
[0037] S130, determining a separation result of the to-be-separated speech data according to the fusion feature and the audio feature.
[0038] After obtaining the fused features, the speech separation system can input the fused features to a separation network, and determine a separation result of the to-be-separated speech data according to an output result of the separation network and the audio features. The separation network can be constructed based on a globally attentive locally recurrent (GALR) network.
[0039] The technical solution fuses the auxiliary features such as the voiceprint feature, the emotion feature and the deep feature on the basis of the audio features, so as to solve the problem of low speech separation accuracy, and can improve the speech separation accuracy while ensuring the stability and anti-interference of the separation effect.
[0040] Embodiment Two
[0041] Figure 2 A flowchart of a speech separation method provided for the second embodiment of the present application is shown in FIG. 2. The embodiment is based on the above-described embodiment and is refined. As shown in FIG. 2, the method comprises the following steps. Figure 2
[0042] S210, obtaining to-be-separated speech data, determining audio features of the to-be-separated speech data and auxiliary features of the to-be-separated speech data.
[0043] In the present solution, the auxiliary features can include a voiceprint feature, an emotion feature and a deep feature. The voiceprint feature is determined by a voiceprint feature extractor performing feature extraction on the to-be-separated speech data; the emotion feature is determined based on spectral information of the to-be-separated speech data; and the deep feature is obtained by domain conversion based on features of the to-be-separated speech data output by a deep feature extractor. In one feasible solution, the emotion feature includes static features and dynamic features; wherein the static features include spectral characteristic features, brightness features and loudness features.
[0044] The spectral characteristic features can include an average value of spectral flatness, an average value of an audio time series zero-crossing rate and a p-order spectral bandwidth. The brightness features can include an average value, a standard deviation and a maximum value of spectral centroid. The loudness features can include an average value, a standard deviation and a maximum value of a root mean square. The dynamic features can include a mel-frequency cepstral coefficient feature, an average value of the mel-frequency cepstral coefficient feature, a standard deviation of the mel-frequency cepstral coefficient feature, a maximum value of the mel-frequency cepstral coefficient feature, a first-order difference of the mel-frequency cepstral coefficient feature and a second-order difference of the mel-frequency cepstral coefficient feature.
[0045] In one specific example, the emotion feature can be a feature composed of 276-dimensional parameters, and the content of each dimension of the parameters can be as shown in Table 1:
[0046] Table 1:
[0047] Feature No. Feature Name 0 Mean of spectral flatness 1 Mean of audio time series zero-crossing rate 2~4 Mean, standard deviation, maximum of amplitude 5~7 Mean, standard deviation, maximum of spectral centroid 8~11 Tuning deviation and its mean, standard deviation, maximum 12~14 Mean, standard deviation, maximum of root mean square (RMS) 15~86 0-24 order MFCC features and their mean, standard deviation, maximum 87~134 First-order difference, second-order difference of 0-24 order MFCC 135~146 Spectrogram 147~274 Mel frequency 275 p-th order spectral bandwidth (default p = 2)
[0048] The above scheme can obtain multi-dimensional features of the to-be-separated speech data, and is beneficial to accurate speech separation according to the multi-dimensional features.
[0049] It should be noted that the audio features and the auxiliary features in the scheme can include three dimensions of frames, time and channels. In order to ensure the correspondence of the features, the audio features and the auxiliary features can have the same frame dimension. Since the extraction methods of the audio features and the auxiliary features are different, the audio features and the auxiliary features can have differences in the time dimension and the channel dimension. Generally, in order to cover more comprehensive time span information, the time dimension of the audio features can be greater than or equal to the time dimension of the auxiliary features.
[0050] S220, adjusting the auxiliary features to have the same time dimension as the audio features.
[0051] If the time dimension of the audio features is greater than that of the auxiliary features, the speech separation system can perform a replication operation on each auxiliary feature to obtain a feature with the same time dimension as the audio features.
[0052] S230, splicing the audio features and the auxiliary features in the time dimension or the channel dimension, and determining the fusion features according to the splicing result.
[0053] After obtaining the audio features and the auxiliary features with the same time dimension, the speech separation system can splice the audio features and the auxiliary features in the time dimension, or can splice them in the channel dimension, and take the splicing result as the fusion features.
[0054] In one feasible scheme, the splicing the audio features and the auxiliary features in the time dimension, and determining the fusion features according to the splicing result, includes:
[0055] The audio features and the auxiliary features are spliced in the time dimension, and the splicing result is reshaped to obtain fusion features matching the time dimension of the audio features.
[0056] It should be noted that if the audio feature and the auxiliary feature are spliced in the time dimension, the speech separation system can reshape the spliced result to maintain the time dimension while fusing the features, so as to obtain a fused feature consistent with the time dimension of the audio feature. For example, the audio feature, the voiceprint feature, the emotion feature and the depth feature after time dimension adjustment are each dimension [B, N, T], wherein the first element represents the frame dimension, the second element represents the channel dimension, and the third dimension represents the time dimension. After splicing the audio feature, the voiceprint feature, the emotion feature and the depth feature in the time dimension, the spliced result is a feature with each dimension [B, N, T+T+T+T]. The speech separation system can reshape the spliced result to obtain a fused feature with dimensions [B, N+N+N+N, T].
[0057] S240, input the fused feature into a speech separation network to determine a separated speech prediction feature.
[0058] The speech separation system can input the fused feature into the speech separation network to obtain the separated speech prediction feature. The speech separation network can use a scale-invariant signal-to-noise ratio as a loss function. The calculation formula of the loss function can be represented as:
[0059]
[0060] wherein, the separated speech prediction feature is represented as y, the audio feature is represented as x.
[0061] S250, determining a separation result of the to-be-detected audio data according to the separated speech prediction feature and the audio feature.
[0062] Optionally, the determining the separation result of the to-be-detected audio data according to the separated speech prediction feature and the audio feature comprises:
[0063] determining a dot product result of the separated speech prediction feature and the audio feature;
[0064] inputting the dot product result as an input of an audio decoder, and determining the separation result of the to-be-detected audio data according to an output of the audio decoder.
[0065] The speech separation system can perform dot product operation on the separated speech prediction feature and the audio feature, and input the dot product result into the audio decoder to obtain the pure speech of each speaker. The audio decoder can have a structure opposite to that of the feature extractor, and is used to restore the features to separated speech data. The audio decoder can include structures such as deconvolution and depooling. It should be noted that in order to ensure consistency of the restoration, the parameter settings in the audio decoder are usually consistent with those in the feature extractor, such as convolution kernel size, convolution step length, etc.
[0066] The technical scheme fuses the auxiliary features such as the voiceprint feature, the emotion feature and the deep feature on the basis of the audio feature, so as to solve the problem of low speech separation accuracy, and can improve the speech separation accuracy while ensuring the stability and anti-interference of the separation effect.
[0067] Embodiment three
[0068] Figure 3 A structural schematic diagram of a speech separation device provided for the third embodiment of the present application is shown in FIG. 3. As shown in the figure, the device comprises: Figure 3
[0069] The feature determination module 310 is configured to acquire the speech data to be separated, determine the audio feature of the speech data to be separated and the auxiliary feature of the speech data to be separated; wherein the auxiliary feature comprises the voiceprint feature, the emotion feature and the deep feature.
[0070] The fusion feature determination module 320 is configured to determine the fusion feature according to the audio feature and the auxiliary feature.
[0071] The separation result determination module 330 is configured to determine the separation result of the speech data to be separated according to the fusion feature and the audio feature.
[0072] In the present scheme, optionally, the voiceprint feature is determined by feature extraction of the speech data to be separated by a voiceprint feature extractor; the emotion feature is determined based on the spectral information of the speech data to be separated; and the deep feature is obtained by domain conversion based on the features of the speech data to be separated output by a deep feature extractor.
[0073] In the above scheme, the emotion feature comprises static features and dynamic features; wherein the static features comprise spectral characteristic features, brightness features and loudness features.
[0074] In a feasible scheme, the time dimension of the audio feature is greater than or equal to the time dimension of the auxiliary feature.
[0075] The fusion feature determination module 320 comprises:
[0076] The dimension adjustment unit is configured to adjust the auxiliary feature to have the same time dimension as the audio feature.
[0077] The fusion feature determination unit is configured to splice the audio feature and the auxiliary feature in the time dimension or the channel dimension, and determine the fusion feature according to the splicing result.
[0078] In the above scheme, the fusion feature determination unit is specifically configured to:
[0079] The audio feature and the auxiliary feature are spliced in a time dimension, and a reshaping operation is performed on the spliced result to obtain a fusion feature matching the time dimension of the audio feature.
[0080] In the scheme, optionally, the separation result determination module 330 includes:
[0081] The prediction feature determination unit is configured to input the fusion feature into a speech separation network to determine a separation speech prediction feature.
[0082] The separation result determination unit is configured to determine a separation result of the to-be-detected audio data according to the separation speech prediction feature and the audio feature.
[0083] On the basis of the above scheme, the separation result determination unit is specifically configured to:
[0084] Determine a dot product result of the separation speech prediction feature and the audio feature.
[0085] Take the dot product result as an input of an audio decoder, and determine the separation result of the to-be-detected audio data according to an output of the audio decoder.
[0086] The speech separation device provided in the embodiments of the present application can execute the speech separation method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0087] Embodiment four
[0088] Figure 4 A structural schematic diagram of an electronic device 410 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0089] As Figure 4As shown, the electronic device 410 includes at least one processor 411, and a memory, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc., connected to the at least one processor 411 in communication. The memory stores computer programs executable by the at least one processor 411, and the processor 411 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 412 or loaded into the random access memory (RAM) 413 from the storage unit 418. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other through a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.
[0090] Various components in the electronic device 410 are connected to the I / O interface 415, including an input unit 416, such as a keyboard, a mouse, etc., an output unit 417, such as various types of displays, a speaker, etc., a storage unit 418, such as a magnetic disk, an optical disk, etc., and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0091] The processor 411 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 411 performs various methods and processes described above, such as the speech separation method.
[0092] In some embodiments, the speech separation method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded onto the RAM 413 and executed by the processor 411, one or more steps of the speech separation method described above can be performed. Alternatively, in other embodiments, the processor 411 can be configured to perform the speech separation method by any other appropriate means, such as by means of firmware.
[0093] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0094] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.
[0095] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0096] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0097] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0098] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0099] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0100] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.
Claims
1. A speech separation method characterized by, The method comprises: acquiring to-be-separated voice data, determining audio features of the to-be-separated voice data, and determining auxiliary features of the to-be-separated voice data; wherein the auxiliary features comprise voiceprint features, emotion features, and deep features; determining fusion features according to the audio features and the auxiliary features; determining a separation result of the to-be-separated voice data according to the fusion features and the audio features; the emotion features comprise static features and dynamic features; wherein the static features comprise spectral characteristic features, brightness features, and loudness features.
2. The method of claim 1, wherein, The voiceprint features are determined by feature extraction of the to-be-separated voice data by a voiceprint feature extractor; the emotion features are determined based on spectral information of the to-be-separated voice data; and the deep features are obtained by domain conversion based on features of the to-be-separated voice data output by a deep feature extractor.
3. The method of claim 1, wherein, The time dimension of the audio features is greater than or equal to the time dimension of the auxiliary features. The determining of the fusion features according to the audio features and the auxiliary features comprises: adjusting the auxiliary features to have the same time dimension as the audio features; splicing the audio features and the auxiliary features in the time dimension or the channel dimension, and determining the fusion features according to the splicing result.
4. The method of claim 3, wherein, The splicing of the audio features and the auxiliary features in the time dimension and the determining of the fusion features according to the splicing result comprise: splicing the audio features and the auxiliary features in the time dimension, and performing a reshaping operation on the splicing result to obtain fusion features matching the time dimension of the audio features.
5. The method of claim 1, wherein, The determining of the separation result of the to-be-detected audio data according to the separation voice prediction features and the audio features comprises: inputting the fusion features into a voice separation network to determine separation voice prediction features; determining the separation result of the to-be-detected audio data according to the separation voice prediction features and the audio features.
6. The method of claim 5, wherein, The determining of the separation result of the to-be-detected audio data according to the separation voice prediction features and the audio features comprises: determining a dot product result of the separation voice prediction features and the audio features; inputting the dot product result into an audio decoder, and determining the separation result of the to-be-detected audio data according to an output of the audio decoder.
7. A speech separation apparatus characterized by comprising: comprise: a feature determination module configured to acquire to-be-separated voice data, determine audio features of the to-be-separated voice data, and determine auxiliary features of the to-be-separated voice data; wherein the auxiliary features comprise voiceprint features, emotion features, and deep features; the emotion features comprise static features and dynamic features; wherein the static features comprise spectral characteristic features, brightness features, and loudness features; a fusion feature determination module configured to determine fusion features according to the audio features and the auxiliary features; a separation result determination module configured to determine a separation result of the to-be-separated voice data according to the fusion features and the audio features.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the voice separation method in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to implement the voice separation method of any one of claims 1-6 when executed.
Citation Information
Patent Citations
Visual voiceprint assisted voice separation method and device
CN113035225A