Speech separation method, electronic device, chip and computer-readable storage medium
By time-domain encoding and visual semantic feature extraction of user audio and video information, combined with preset visual speech separation network, the problem of low speech separation accuracy of unknown speakers in the prior art is solved, and the speech separation effect with high accuracy and low latency is achieved.
Patent Information
- Application Number
- CN202011027680.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-09-25
AI Technical Summary
The existing voice separation technology based on audio and video fusion has poor generalization ability for unknown speakers and low speech separation accuracy, resulting in poor user experience and difficulty in applying it in real-time voice separation scenarios.
By obtaining the user's audio and video information, using the convolutional neural network for time domain encoding, extracting visual semantic features, and modal fusion is performed in the preset visual speech separation network to achieve speech separation.
It improves the accuracy and generalization ability of speech separation for unknown speakers, reduces the delay in speech separation, and improves the user experience.
Smart Images

Figure CN114333896B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a voice separation method, a chip electronic device, a chip, and a computer-readable storage medium. Background Art
[0002] Voice interaction technology has been increasingly applied in actual products, such as mobile phone intelligent voice assistants, voice control of smart speakers, video conferencing devices, etc. However, in the case of being interfered by a noisy environment and surrounding human voices, situations such as low speech recognition accuracy and degraded call quality will occur. To solve the above problems, the industry has proposed a voice separation technology based on audio-visual fusion. This audio-visual fusion voice separation technology performs voice separation based on face representations. Its basic idea is: using a pre-trained face model to extract face representations, and then based on the face representations, mixed speech, and deep learning algorithms, extracting the speech of a specified speaker. However, this technology has poor generalization ability for unknown speakers, that is, when the speech of the target speaker has not appeared in the training dataset, the accuracy of its voice separation is poor, resulting in a poor user experience, and the latency of voice separation is large, making it difficult to be applied in real-time voice separation application scenarios. Summary of the Invention
[0003] In view of this, it is necessary to provide a voice separation method, which can overcome the above problems, has strong generalization ability for unknown speakers, high voice separation accuracy, and improves the user experience.
[0004] In the first aspect of the embodiments of this application, a voice separation method is disclosed, including: obtaining audio information containing user speech and video information containing the user's face during the user's speech; encoding the audio information to obtain mixed acoustic features; extracting the user's visual semantic features from the video information, where the visual semantic features include the facial movement features of the user during the speech; inputting the mixed acoustic features and the visual semantic features into a preset visual voice separation network to obtain the user's acoustic features; decoding the user's acoustic features to obtain the user's voice signal.
[0005] By adopting this technical solution, voice separation of mixed speech containing user speech and environmental noise can be realized based on visual semantic features, and the user's voice can be accurately separated, improving the user experience.
[0006] In a possible implementation manner, the audio information is mixed speech information containing the user's speech and environmental noise, and the encoding of the audio information includes: constructing a time-domain audio encoder based on a convolutional neural network; using the time-domain audio encoder to perform time-domain encoding on the audio information.
[0007] By adopting this technical solution, time-domain encoding is performed on the mixed speech, so that the subsequent time-domain speech signal can be decoded, reducing the loss of speech phase information, improving the speech separation performance, and having the advantage of low speech separation delay.
[0008] In a possible implementation manner, the decoding of the acoustic features of the user to obtain the speech signal of the user includes: constructing a time-domain audio decoder based on the convolutional neural network; using the time-domain audio decoder to decode the acoustic features of the user to obtain the time-domain speech signal of the user.
[0009] By adopting this technical solution, the time-domain speech signal can be decoded, reducing the loss of speech phase information, improving the speech separation performance, and having the advantage of low speech separation delay.
[0010] In a possible implementation manner, the audio information is mixed speech information including the user's speech and environmental noise, and the encoding of the audio information includes: performing time-domain encoding on the audio information by using a preset short-time Fourier transform algorithm.
[0011] By adopting this technical solution, time-domain encoding is performed on the mixed speech, so that the subsequent time-domain speech signal can be decoded, reducing the loss of speech phase information, improving the speech separation performance, and having the advantage of low speech separation delay.
[0012] In a possible implementation manner, the decoding of the acoustic features of the user to obtain the speech signal of the user includes: using a preset inverse short-time Fourier transform algorithm to decode the acoustic features of the user to obtain the time-domain speech signal of the user.
[0013] By adopting this technical solution, the time-domain speech signal can be decoded, reducing the loss of speech phase information, improving the speech separation performance, and having the advantage of low speech separation delay.
[0014] In a possible implementation manner, the extraction of the visual semantic features of the user from the video information includes: converting the video information into image frames arranged in the order of frame playback; processing each image frame to obtain multiple face thumbnails with a preset size and including the user's face; inputting the multiple face thumbnails into a preset decoupling network to extract the visual semantic features of the user.
[0015] By adopting this technical solution, speech separation of the mixed speech including the user's speech and environmental noise is realized based on the visual semantic features, and the user's voice can be accurately separated, improving the user experience.
[0016] In a possible implementation, the processing of each of the image frames to obtain multiple face thumbnails having a preset size and including the user's face includes: locating the image region including the user's face in each of the image frames; performing magnification or reduction processing on the image region to obtain a face thumbnail having the preset size and including the user's face.
[0017] By adopting this technical solution, speech separation of the mixed speech including the user's voice and environmental noise is realized based on visual semantic features, and the user's voice can be accurately separated, improving the user experience.
[0018] In a possible implementation, the inputting of the multiple face thumbnails into a preset decoupling network to extract the visual semantic features of the user includes: inputting the multiple face thumbnails into the preset decoupling network; using the preset decoupling network to map each of the face thumbnails into a visual representation including face identity features and the visual semantic features, and separating the visual semantic features from the visual representation.
[0019] By adopting this technical solution, separating the visual semantic features from the visual representation by using a preset decoupling network is realized, and speech separation of the mixed speech including the user's voice and environmental noise is realized, and the user's voice can be accurately separated, improving the user experience.
[0020] In a possible implementation, the inputting of the mixed acoustic features and the visual semantic features into a preset visual speech separation network to obtain the acoustic features of the user includes: obtaining the time dependence relationship of the mixed acoustic features to obtain deep mixed acoustic features based on the time dependence relationship of the mixed acoustic features; obtaining the time dependence relationship of the visual semantic features to obtain deep visual semantic features based on the time dependence relationship of the visual semantic features; performing modal fusion on the deep mixed acoustic features and the deep visual semantic features to obtain audiovisual features; predicting the acoustic features of the user based on the audiovisual features.
[0021] By adopting this technical solution, speech separation of the mixed speech including the user's voice and environmental noise is realized by using a preset visual speech separation network, and the user's voice can be accurately separated, improving the user experience.
[0022] In a possible implementation, before performing modal fusion on the deep mixed acoustic features and the deep visual semantic features, it further includes: performing time dimension synchronization processing on the deep mixed acoustic features and the deep visual semantics to make the time dimension of the deep mixed acoustic features synchronized with the time dimension of the deep visual semantics.
[0023] By adopting this technical solution, it is possible to perform voice separation on the mixed voice containing the user's voice and environmental noise by using a preset visual voice separation network, accurately separate the voice of the user, and improve the user experience.
[0024] In a possible implementation manner, the obtaining of the acoustic features of the user based on the audiovisual features includes: predicting a masking value of the user's voice based on the audiovisual features; performing output mapping processing on the masking value by using a preset activation function; and performing a matrix dot product operation on the masking value processed by the preset activation function and the mixed acoustic features to obtain the acoustic features of the user.
[0025] By adopting this technical solution, it is possible to perform voice separation on the mixed voice containing the user's voice and environmental noise by using a preset visual voice separation network, accurately separate the voice of the user, and improve the user experience.
[0026] In a possible implementation manner, the performing of output mapping processing on the masking value by using a preset activation function includes: if encoding the audio information based on a convolutional neural network, performing output mapping processing on the masking value by using the sigmoid function; or if encoding the audio information based on the short-time Fourier transform algorithm, performing output mapping processing on the masking value by using the Tanh function.
[0027] By adopting this technical solution, it is possible to perform output mapping processing by using an activation function corresponding to the audio encoding algorithm according to different audio encoding algorithms.
[0028] In a second aspect, an embodiment of the present application provides a computer-readable storage medium, including computer instructions, which when running on an electronic device, cause the electronic device to execute the voice separation method as described in the first aspect or the second aspect.
[0029] In a third aspect, an embodiment of the present application provides an electronic device, in which at least an agent service process is installed. The electronic device includes a processor and a memory. The memory is used to store instructions, and the processor is used to call the instructions in the memory, so that the electronic device executes the voice separation method as described in the first aspect or the second aspect.
[0030] In a fourth aspect, an embodiment of the present application provides a computer program product, which when running on a computer, causes the computer to execute the voice separation method as described in the first aspect or the second aspect.
[0031] Fifth aspect, an embodiment of the present application provides a device, which has the function of implementing the behavior of the first electronic device in the method provided in the above first aspect or second aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0032] It can be understood that the computer-readable storage medium described in the above second aspect, the electronic device described in the third aspect, the computer program product described in the fourth aspect, and the device described in the fifth aspect all correspond to the method in the above first aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Schematic diagram of an application scenario of a voice separation device provided by an embodiment of the present application;
[0034] Figure 2 Schematic flowchart of a voice separation method provided by an embodiment of the present application;
[0035] Figure 3 Schematic diagram of the network structure of a preset decoupling network provided by an embodiment of the present application;
[0036] Figure 4 Schematic diagram of the network structure of a preset visual voice separation network provided by an embodiment of the present application;
[0037] Figure 5 Schematic diagram of the functional modules of a voice separation device provided by an embodiment of the present application;
[0038] Figure 6 Schematic diagram of the structure of a possible electronic device provided by an embodiment of the present application; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] It should be noted that "at least one" in the present application means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0040] For ease of understanding, some explanations of concepts related to the embodiments of the present application are given by way of example for reference.
[0041] The following is combined withFigure 1 Schematic diagram of an application scenario of a voice separation device provided by an exemplary embodiment of the present invention. The voice separation device may be disposed in the electronic device 100.
[0042] When the user uses the electronic device 100 for a call, a video conference, voice interaction control, etc., if the current scene where the user is located contains the voices of other people or the voices of other objects, the user's voice can be separated and enhanced, so as to highlight the user's voice and reduce the interference of ambient noise on the user's voice.
[0043] The electronic device 100 may be a device such as a mobile phone, a computer, a smart home appliance, a car stereo, etc.
[0044] Referring to Figure 2 As shown, a voice separation method provided by an embodiment of the present application is applied to the electronic device 100. In this embodiment, the voice separation method includes:
[0045] 21. Obtain audio information containing the user's voice and video information containing the user's face during the user's speech.
[0046] In some embodiments, the electronic device 100 may include a camera function and a sound pickup function. For example, the electronic device 100 includes a camera and a microphone. The camera is used to collect video information containing the user's face during the user's speech, and the microphone is used to collect audio information containing the user's voice during the user's speech. Furthermore, audio information containing the user's voice and video information containing the user's face during the user's speech can be obtained from the camera and the microphone.
[0047] It can be understood that the video information collected by the camera not only includes the user's face information, but may also include other body part information of the user, the current shooting background information, or the body part information of other users. The audio information collected by the microphone not only includes the user's current speech, but may also include ambient noise. For example, the ambient noise is the voice of other users and / or the voice of other objects.
[0048] 22. Encode the audio information to obtain a mixed acoustic feature.
[0049] In some embodiments, a preset audio encoder may be used to encode the audio information to obtain hybrid acoustic features. The preset audio encoder may be an encoder constructed based on a Convolutional Neural Network (CNN), but is not limited to CNN, and may also be other types of neural networks, such as Long Short-Term Memory (LSTM), Recurrent Neural Network (RNN), etc. The construction method of constructing the preset audio encoder using CNN may be the construction method recorded in the existing solutions, which will not be elaborated here.
[0050] In some embodiments, the processing of audio information generally includes time-domain processing and frequency-domain processing. Compared with frequency-domain processing, time-domain processing can reduce the length of speech frames, facilitate the design of a low-latency speech separation model, reduce the loss of speech phase information, and thus improve speech separation performance. The preset audio encoder is preferably an audio encoder for time-domain encoding constructed based on CNN.
[0051] In some embodiments, the audio information is mixed speech containing user speech, and the hybrid acoustic features may refer to vectors containing hybrid speech features obtained by CNN encoding.
[0052] In some embodiments, the short-time Fourier transform algorithm may also be used to perform time-domain encoding on the audio information to obtain hybrid acoustic features.
[0053] 23. Extract the visual semantic features of the user from the video information.
[0054] In some embodiments, the visual semantic features include the facial movement features of the user during speech, such as lip movement features and cheek movement features. The extraction of the visual semantic features of the user from the video information may be achieved in the following way:
[0055] a. Convert the video information into image frames arranged in the order of frame playback, and process the image frames to obtain a face thumbnail with a preset size and containing the user's face;
[0056] Specifically, existing video decoding techniques can be used to decode the video information to obtain multiple image frames arranged in the order of frame playback. Then, existing face detection techniques can be used to locate the image area containing the user's face in each image frame. Finally, the image area is enlarged or reduced to obtain a face thumbnail with the preset size and containing the user's face. The preset size can be set according to actual needs. For example, the preset size is 256*256, that is, the image area of the user's face is uniformly converted into a face thumbnail of 256*256.
[0057] In some embodiments, since the sizes of the image areas containing the user's face in each image frame may be different, in order to uniformly convert them into face thumbnails of 256*256, some image areas may need to be enlarged and some image areas may need to be reduced. Specifically, it can be determined whether to select enlargement processing or reduction processing according to the size of the actually located image area of the user's face.
[0058] b. Input multiple face thumbnails into a preset decoupling network to extract the visual semantic features of the user.
[0059] Specifically, when a face thumbnail of the preset size is obtained, the face thumbnail can be input into a preset decoupling network that has undergone adversarial training, and the preset decoupling network is used to extract the visual semantic features of the user. The schematic diagram of the network structure of the preset decoupling network is as Figure 3 shown. The preset decoupling network may include a visual encoder E v 、a speech encoder E a 、a classifier D1, a binary classifier D2, and an identity discriminator Dis.
[0060] In some embodiments, N video samples and N audio samples can be used to train the preset decoupling network, where N is a positive integer greater than 1:
[0061] i. Conduct joint audio-visual representation learning to map the face thumbnail into a visual representation containing face identity features and visual semantic features;
[0062] During training, randomly select the m-th audio sample from the audio samples of size N, and randomly select the n-th video sample from the video samples of size N. Define the label as: when the n-th video sample matches the m-th audio sample (that is, the audio sample is the playback sound of the video sample), record it as l mn =1, when the n-th video sample does not match the m-th audio sample, record it as l mn =0. The n-th video sample can be input into the visual encoder E v(Constructed based on CNN), obtaining a visual representation f that contains face identity features and visual semantic features v(n) Input the m-th audio example into the speech encoder E a (Constructed based on CNN), obtaining a speech representation f that contains sound features a(m) ;
[0063] When the visual representation f v(n) and the speech representation f a(m) are obtained, the distance between the visual representation f v(n) and the speech representation f a(m) can be reduced through the following three processing methods:
[0064] a). The visual representation f v(n) and the speech representation f a(m) share the same classifier D1 for word-level audiovisual speech recognition tasks, and the loss is denoted as;
[0065]
[0066] Among them, is the total number of words in the training set, p k is the true class label, and each class label can correspond to a word. k is a positive integer greater than zero.
[0067] b). Use a binary classifier D2 for adversarial training to identify whether the input representation is a visual representation or an audio representation;
[0068] First, freeze the weights of the visual encoder E v and the speech encoder E a (that is, fix the weights of the visual encoder E v and the speech encoder E a so that their weights are not trained), and train the binary classifier D2 so that it can correctly distinguish whether the input representation is a visual representation or an audio representation. Its training loss is denoted as Then, freeze the weights of the binary classifier D2, and train the visual encoder E v and the speech encoder E a so that the binary classifier D2 cannot correctly distinguish whether the input representation is a visual representation or an audio representation. Its training loss is denoted as The loss and the loss are as follows:
[0069]
[0070]
[0071] Among them, p v= 0 represents that the input representation is a visual representation, p a = 1 represents that the input representation is an audio representation.
[0072] c). Minimize the visual representation f c and the speech representation f v(n) by the contrastive loss L a(m) . The loss c is defined as follows:
[0073]
[0074] where d mn is the Euclidean distance between the visual representation f v(n) and the speech representation f a(m) , and d mn = ||f v(n) - f a(m) ||2;
[0075] ii. Use an adversarial approach to separate the visual semantic features from the visual representation.
[0076] First, freeze the weights of the visual encoder E v to train the identity discriminator Dis so that the identity discriminator Dis can correctly identify the identity of each face in the video example. Its training loss is denoted as Then, freeze the weights of the identity discriminator Dis and train the visual encoder E v so that the visual representation encoded by the visual encoder E v completely loses the identity information (i.e., loses the face identity features). Its training loss is denoted as After the training of the visual encoder E v is completed, if the visual representation f v(n) is input into the visual encoder E v , for each class of identity, it tends to output equal probabilities, that is, it completely loses the identity information. The trained visual encoder E v can be used to separate the visual semantic features from the visual representation. Each class of identity corresponds to an identity ID, representing a person. The loss and the loss are as follows:
[0077]
[0078]
[0079] where N p is the number of identity categories, p jIt is a one-hot label, where j is a positive integer greater than zero. For example, if N video samples in total include 10 types of identities (the first type of identity to the tenth type of identity), if the first video sample belongs to the first type of identity, the corresponding one-hot can be expressed as "1000000000", and if the second video sample belongs to the third type of identity, the corresponding one-hot can be expressed as "0010000000".
[0080] 24. Input the mixed acoustic features and the visual semantic features into a preset visual speech separation network to obtain the acoustic features of the user.
[0081] In some embodiments, the preset visual speech separation network may be a network constructed based on a Temporal Convolutional Network (TCN). The schematic diagram of the network structure of the preset visual speech separation network may be as Figure 4 shown. The preset visual speech separation network includes a first TCN unit TCN-1, a second TCN unit TCN-2, a third TCN unit TCN-3, an upsampling unit Upsample, a modal fusion unit Modal_fusion, a normalization-convolution unit LN_convld, an activation-convolution unit PreLU_convld, an activation unit σ / Tanh, and a matrix dot product unit Matrix_dm.
[0082] The regularized-convolution unit LN_convld is used to regularize the input mixed acoustic features and process them through a one-dimensional convolutional layer; the first TCN unit TCN-1 is used to capture the temporal dependencies of the mixed acoustic features to obtain deep mixed acoustic features; the third TCN unit TCN-3 is used to capture the temporal dependencies of the input visual semantic features to obtain deep visual semantic features; the upsampling unit Upsample is used to upsample the deep visual semantic features to synchronize them with the deep mixed acoustic features in the time dimension; the modality fusion unit Modal_fusion is used to concatenate the deep visual semantic features and the deep mixed acoustic features in the channel dimension and perform a dimensional transformation through a linear layer to obtain fused audiovisual features. The fused audiovisual features can be represented by the following formula: =P([a;Upsample(V)]), where f is the fused audiovisual features, i.e., the input of the second TCN unit TCN-2, P is a linear mapping, is the deep mixed acoustic features, and V is the deep visual semantic features; the second TCN unit TCN-2 and the activation-convolution unit PreLU_convld are used to predict the mask value of the user's speech according to the fused audiovisual features f; the activation unit σ / Tanh is used to introduce non-linear characteristics to perform mapping output processing on the mask value; the matrix dot product unit Matrix_dm is used to perform a matrix dot product operation on the mask output by the activation unit σ / Tanh and the mixed acoustic features to obtain the acoustic features of the user.
[0083] In some embodiments, when the mixed acoustic features are obtained by CNN encoding, the activation unit σ / Tanh can optionally use the sigmoid function to introduce non-linear characteristics. When the mixed acoustic features are obtained by short-time Fourier transform, the activation unit σ / Tanh can optionally use the Tanh function to introduce non-linear characteristics.
[0084] 25. Decode the acoustic features of the user to obtain the speech signal of the user.
[0085] In some embodiments, when the acoustic features of the user are obtained through the preset visual speech separation network, a preset audio decoder can be used to decode the acoustic features of the user to obtain the speech signal of the user. The preset audio decoder can be a decoder constructed based on CNN, but is not limited to CNN, and can also be other types of neural networks, such as LSTM, RNN, etc. The construction method of constructing the preset audio decoder using CNN can be the construction method recorded in the existing solutions and will not be elaborated here.
[0086] It can be understood that when the short-time Fourier transform algorithm is used to encode the audio information to obtain the mixed acoustic features, at this time, the inverse short-time Fourier transform algorithm can be used to decode the acoustic features of the user to obtain the speech signal of the user.
[0087] In some embodiments, since the CNN or the short-time Fourier transform algorithm is used to perform time-domain encoding on the audio information, the decoded user speech signal is a time-domain speech signal.
[0088] The above speech separation method can perform speech separation on the mixed speech in the time domain based on the visual semantic features, and can accurately and real-time separate the speech of the target speaker from the environmental noise interference. It has high accuracy and strong generalization for the speech separation of unknown speakers, low speech separation delay, and supports the application scenarios of real-time speech separation.
[0089] Refer to Figure 5 As shown, a speech separation device 110 provided by an embodiment of the present application can be applied to Figure 1 the electronic device 100 shown. The electronic device 100 can include a camera function and a sound pickup function. In this embodiment, the speech separation device 110 can include an acquisition module 101, an encoding module 102, an extraction module 103, a separation module 104, and a decoding module 105.
[0090] The acquisition module 101 is used to acquire audio information containing the user's speech and video information containing the user's face during the user's speech.
[0091] The encoding module 102 is used to encode the audio information to obtain mixed acoustic features.
[0092] The extraction module 103 is used to extract the visual semantic features of the user from the video information. The visual semantic features include the facial movement features of the user during the speech.
[0093] The separation module 104 is used to input the mixed acoustic features and the visual semantic features into a preset visual speech separation network to obtain the acoustic features of the user.
[0094] The decoding module 105 is used to decode the acoustic features of the user to obtain the speech signal of the user.
[0095] It can be understood that the division of each module in the above device 110 is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. For example, the above-mentioned each module can be a separately established processing element, or can be integrated in a certain chip of the terminal. In addition, it can also be stored in the storage element of the controller in the form of program code, and the functions of the above-mentioned each module are called and executed by a certain processing element of the processor. In addition, the above-mentioned each module can be integrated together or independently implemented. The processing element described here can be an integrated circuit chip with signal processing capabilities. The processing element can be a general-purpose processor, such as a central processing unit (CPU), or can also be one or more integrated circuits configured to implement the above function modules, such as: one or more application-specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field-programmable gate arrays (FPGAs), etc.
[0096] Reference Figure 6 , is a schematic diagram of the hardware structure of the electronic device 100 provided by the embodiment of the present application. As Figure 6 shown, the electronic device 100 may include a processor 1001, a memory 1002, a communication bus 1003, a camera assembly 1004, a microphone assembly 1005, and a speaker assembly 1006. The memory 1002 is used to store one or more computer programs 1007. The one or more computer programs 1007 are configured to be executed by the processor 1001. The one or more computer programs 1007 include instructions, and the above instructions can be used to implement the above voice separation method or the above voice separation device 110 in the electronic device 100.
[0097] It can be understood that the structure schematically shown in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements.
[0098] The processor 1001 may include one or more processing units. For example, the processor 1001 may include an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a DSP, a CPU, a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0099] The processor 1001 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 1001 is a cache memory. This memory can save the instructions or data that the processor 1001 has just used or recycled. If the processor 1001 needs to use the instruction or data again, it can be directly called from this memory. This avoids repeated accesses, reduces the waiting time of the processor 1001, and thus improves the efficiency of the system.
[0100] In some embodiments, the processor 1001 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface, etc.
[0101] In some embodiments, the memory 1002 may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0102] The camera component 1004 is used to capture the facial information of the speaker to generate video information including the speaker's face. The camera component 1004 may include a lens, an image sensor, an image signal processor, etc. The microphone component 1005 is used to record the voice of the speaker and the surrounding environmental sounds to obtain audio information including the user's voice. The microphone component 1005 may include a microphone and peripheral circuits or components cooperating with the microphone. The speaker component 1006 is used to play the voice of the speaker obtained through voice separation processing. The speaker component 1006 may include a speaker and peripheral circuits or components cooperating with the speaker.
[0103] This embodiment also provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device is enabled to execute the above related method steps to implement the voice separation method in the above embodiment.
[0104] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above related steps to implement the voice separation method in the above embodiment.
[0105] In addition, an embodiment of the present application also provides a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other. Among them, the memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory to enable the chip to execute the voice separation method in each of the above method embodiments.
[0106] Among them, the first electronic device, the computer storage medium, the computer program product or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved may refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0107] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0108] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0109] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0110] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0111] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0112] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application.
Claims
1. A voice separation method, characterized in that, The method includes: Obtaining audio information containing the user's voice and video information containing the user's face during the user's speech; Encoding the audio information to obtain mixed acoustic features; Extracting the visual semantic features of the user from the video information, where the visual semantic features include the facial motion features of the user during speech; Inputting the mixed acoustic features and the visual semantic features into a preset visual speech separation network to obtain the acoustic features of the user; Decoding the acoustic features of the user to obtain the voice signal of the user; Among them, the extracting the visual semantic features of the user from the video information includes: Converting the video information into image frames arranged in the order of frame playback, where the image frames contain human faces; Inputting the image frames into a preset decoupling network, using the preset decoupling network to map each image frame into a visual representation containing the face identity features and the visual semantic features, and performing identity feature loss processing on the visual representation to separate the visual semantic features from the visual representation.
2. The voice separation method according to claim 1, wherein, The audio information is mixed speech information containing the user's voice and environmental noise, and the encoding the audio information includes: Constructing a time-domain audio encoder based on a convolutional neural network; Using the time-domain audio encoder to perform time-domain encoding on the audio information.
3. The voice separation method according to claim 2, wherein The decoding the acoustic features of the user to obtain the voice signal of the user includes: Constructing a time-domain audio decoder based on the convolutional neural network; Using the time-domain audio decoder to decode the acoustic features of the user to obtain the time-domain voice signal of the user.
4. The voice separation method according to claim 1, wherein The audio information is mixed speech information containing the user's voice and environmental noise, and the encoding the audio information includes: Performing time-domain encoding on the audio information using a preset short-time Fourier transform algorithm.
5. The voice separation method according to claim 4, characterized in that The decoding the acoustic features of the user to obtain the voice signal of the user includes: Performing inverse decoding on the acoustic features of the user using a preset inverse short-time Fourier transform algorithm to obtain the time-domain voice signal of the user.
6. The voice separation method according to claim 1, wherein, The inputting the image frames into a preset decoupling network includes: Processing the image frames to obtain a face thumbnail with a preset size and containing the user's face; Inputting the face thumbnail into the preset decoupling network.
7. The voice separation method according to claim 6, wherein, The processing the image frames to obtain a face thumbnail with a preset size and containing the user's face includes: Locating the image area in the image frame that contains the user's face; Performing magnification or reduction processing on the image area to obtain a face thumbnail with the preset size and containing the user's face.
8. The voice separation method according to claim 1, characterized in that, The inputting the mixed acoustic features and the visual semantic features into a preset visual speech separation network to obtain the acoustic features of the user includes: Obtaining the time dependence of the mixed acoustic features to obtain deep mixed acoustic features based on the time dependence of the mixed acoustic features; Obtaining the time dependence of the visual semantic features to obtain deep visual semantic features based on the time dependence of the visual semantic features; Perform modal fusion on the depth - mixed acoustic features and the depth - visual semantic features to obtain audiovisual features; Predict the acoustic features of the user based on the audiovisual features.
9. The voice separation method according to claim 8, characterized in that Before performing the modal fusion on the depth - mixed acoustic features and the depth - visual semantic features, it further includes: Perform temporal - dimension synchronization processing on the depth - mixed acoustic features and the depth - visual semantics so that the temporal dimension of the depth - mixed acoustic features is synchronized with the temporal dimension of the depth - visual semantics.
10. The voice separation method according to claim 8, wherein The predicting the acoustic features of the user based on the audiovisual features includes: Predict the masking value of the user's speech based on the audiovisual features; Perform output - mapping processing on the masking value using a preset activation function; Perform matrix dot - product operation on the masking value processed by the preset activation function and the mixed acoustic features to obtain the acoustic features of the user.
11. The voice separation method according to claim 10, characterized in that, The performing output - mapping processing on the masking value using a preset activation function includes: If encoding the audio information based on a convolutional neural network, use the sigmoid function to perform output - mapping processing on the masking value; or If encoding the audio information based on the short - time Fourier transform algorithm, use the Tanh function to perform output - mapping processing on the masking value.
12. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the speech separation method according to any one of claims 1 to 11.
13. An electronic device, characterized in that, The electronic device includes a processor and a memory. The memory is used to store instructions, and the processor is used to call the instructions in the memory so that the electronic device executes the speech separation method according to any one of claims 1 to 11.
14. A chip, coupled to a memory in an electronic device, characterized in that, The chip is used to control the electronic device to execute the speech separation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Audio-visual speech separation
EP3607547A1