Voice signal extraction method and device, readable storage medium and electronic equipment

By combining deep learning methods with single-channel audio signals and lip image sequences, the shortcomings of single-channel and multi-channel speech separation methods are overcome, achieving high-accuracy, low-cost, and low-latency speech signal extraction.

CN115910037BActive Publication Date: 2025-12-12BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211179551.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-12-12
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Existing single-channel speech separation methods are ineffective, while multi-channel speech separation methods are costly and have limited application scenarios. Deep learning-based speech separation methods have low accuracy, and traditional methods are computationally complex, resulting in long latency.

Method used

By combining a single-channel mixed audio signal with the target user's lip image sequence, a deep learning method is used for speech separation. The speech signal of the target user is extracted by fusing lip state feature data and audio feature data.

Benefits of technology

It improves the accuracy of speech separation, reduces hardware costs and data processing volume, reduces speech separation latency, and enhances the scalability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910037B_ABST
    Figure CN115910037B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a speech signal extraction method and device, a computer readable storage medium and an electronic device. The method comprises: acquiring a single-channel mixed audio signal and an image sequence collected in a target area; determining a target user in the target area based on the image sequence; determining a lip region image sequence of the target user based on the image sequence; determining lip state feature data based on the lip region image sequence; determining audio feature data based on the single-channel mixed audio signal; fusing the lip state feature data and the audio feature data to obtain fused feature data; and extracting a speech signal of the target user from the single-channel mixed audio signal based on the fused feature data. The embodiments of the present disclosure can effectively improve the accuracy of extracting the speech signal of the target user, reduce the delay time of speech separation, and improve the scalability of the method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular to a speech signal extraction method and device, computer readable storage medium and electronic equipment. BACKGROUND

[0002] Speech separation technology for target speaker voice extraction refers to separating the voice of a target speaker from a mixed speech signal in which multiple people speak at the same time. As a front-end technology of speech recognition, speech separation has always been one of the key technologies in human-computer interaction.

[0003] According to different interference sources, the speech separation task can be divided into three categories: when the interference source is a noise signal, it can be called "speech enhancement"; when the interference source is the voice of other speakers, it can be called "multi-speaker separation"; when the interference is the reflection wave of the target speaker's own voice, it can be called "de-reverberation".

[0004] Since the sound collected by the microphone can include noise, the voices of other people, reverberation and other interferences, if no speech separation is performed and direct recognition is performed, the accuracy of recognition will be affected. Therefore, adding speech separation technology in the front end of speech recognition can separate the voice of the target speaker from other interferences, thereby improving the robustness of the speech recognition system, which has thus become an indispensable part of modern speech recognition systems.

[0005] Existing speech separation methods include traditional signal processing methods and deep learning-based methods, and according to the number of sensors or microphones, they can be divided into single-channel methods (single microphone) and multi-channel methods (multiple microphones). Two traditional methods of single-channel speech separation include speech enhancement and computational auditory scene analysis (CASA). Two traditional methods of multi-channel speech separation include beamforming method and blind source separation (BSS) method.

[0006] The existing single-channel traditional speech separation method needs to be improved in terms of separation effect due to the lack of reference of other channel signals, and the multi-channel speech separation method requires multiple microphones, resulting in high cost, large data processing amount and large use scene limitation.

[0007] The existing deep learning-based speech separation method usually only uses audio feature data, resulting in low accuracy of speech separation. SUMMARY

[0008] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a speech signal extraction method, device, computer readable storage medium and electronic equipment.

[0009] Embodiments of the present disclosure provide a speech signal extraction method, which comprises: acquiring a single-channel mixed audio signal and an image sequence single-channel mixed audio signal collected in a target area; determining a target user in the target area based on the image sequence; determining a lip region image sequence of the target user based on the image sequence; determining lip state feature data based on the lip region image sequence; determining audio feature data based on the single-channel mixed audio signal; fusing the lip state feature data and the audio feature data to obtain fusion feature data; and extracting a speech signal of the target user from the single-channel mixed audio signal based on the fusion feature data.

[0010] According to another aspect of the embodiments of the present disclosure, a speech signal extraction device is provided, which comprises: an acquisition module configured to acquire a single-channel mixed audio signal and an image sequence single-channel mixed audio signal collected in a target area; a first determination module configured to determine a target user in the target area based on the image sequence; a second determination module configured to determine a lip region image sequence of the target user based on the image sequence; a third determination module configured to determine lip state feature data based on the lip region image sequence; a fourth determination module configured to determine audio feature data based on the single-channel mixed audio signal; a fusion module configured to fuse the lip state feature data and the audio feature data to obtain fusion feature data; and an extraction module configured to extract a speech signal of the target user from the single-channel mixed audio signal based on the fusion feature data.

[0011] According to another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program configured to be executed by a processor to implement the speech signal extraction method described above.

[0012] According to another aspect of the embodiments of the present disclosure, an electronic equipment is provided, which comprises: a processor; a memory configured to store processor executable instructions; and the processor configured to read the executable instructions from the memory and execute the instructions to implement the speech signal extraction method described above.

[0013] The speech signal extraction method, device, computer readable storage medium and electronic device provided by the above embodiments of the present disclosure are based on the following. A single-channel mixed audio signal collected in a target area and a lip region image sequence of a target user are obtained, then based on the lip region image sequence, lip state feature data is determined, and based on the single-channel mixed audio signal, audio feature data is determined, then the lip state feature data and the audio feature data are fused to obtain fusion feature data, and finally based on the fusion feature data, the speech signal of the target user is extracted from the single-channel mixed audio signal, thereby realizing multi-modal speech separation combining audio signals and lip images. The feature data used for speech separation is more abundant, and compared with the method of relying only on single-modal audio signals for speech separation, the multi-modal speech separation method provided by the embodiments of the present disclosure has higher accuracy of the extracted speech signal of the target user. In addition, since only a single microphone is used to collect the audio signal, the hardware cost can be reduced, and the data processing amount is also reduced. The traditional speech separation method for single-channel mixed audio signals uses a complex algorithm, and needs a certain convergence time during calculation, causing a long delay time of speech separation. The method provided by the embodiments of the present disclosure combines lip image feature data, and does not need to use the traditional speech separation algorithm, thereby effectively reducing the delay time of speech separation. In addition, in a multi-person scene, only the lip image sequences of different users need to be obtained, and the method provided by the embodiments of the present disclosure is performed on different lip image sequences, so that the speech signals of multiple persons can be extracted, thereby effectively improving the scalability of the method.

[0014] The technical solutions of the present disclosure will be described in further detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0015] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. The drawings provided below for the purpose of explanation only and are constitute a part of the specification. They serve to explain the present disclosure together with the embodiments of the present disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 is a system diagram to which the present disclosure is applicable.

[0017] Figure 2 is a flowchart of a speech signal extraction method provided by an exemplary embodiment of the present disclosure.

[0018] Figure 3 is a flowchart of a speech signal extraction method provided by another exemplary embodiment of the present disclosure.

[0019] Figure 4is a structural schematic diagram of a first neural network model provided by another exemplary embodiment of the present disclosure.

[0020] Figure 5 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0021] Figure 6 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0022] Figure 7 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0023] Figure 8 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0024] Figure 9 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0025] Figure 10 is a flow schematic diagram of a voice signal extraction method provided by another exemplary embodiment of the present disclosure.

[0026] Figure 11 is an exemplary schematic diagram of generating fusion feature data provided by an exemplary embodiment of the present disclosure.

[0027] Figure 12 is a structural schematic diagram of a voice signal extraction device provided by an exemplary embodiment of the present disclosure.

[0028] Figure 13 is a structural schematic diagram of a voice signal extraction device provided by another exemplary embodiment of the present disclosure.

[0029] Figure 14 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] In the following, example embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the example embodiments described herein.

[0031] It should be noted that: unless otherwise specifically stated, the relative arrangement, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0032] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent a logical sequence between them.

[0033] It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more.

[0034] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, it can be understood as one or more in general, without explicit limitation or in the context of the opposite indication.

[0035] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present disclosure generally represents that the front and rear associated objects are in an "or" relationship.

[0036] It should also be understood that the description of the embodiments of the present disclosure focuses on the differences between the embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, they will not be repeated.

[0037] At the same time, it should be understood that in order to facilitate description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship.

[0038] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application or uses.

[0039] The techniques, methods and devices known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods and devices should be considered as part of the specification.

[0040] It should be noted that: similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0041] The disclosed embodiments can be applied to terminal devices, computer systems, servers, and other electronic devices, which can operate with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations that can be suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computers, server computers, thin clients, thick clients, hand-held or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputers, mainframe computers, and distributed cloud computing technology environments that include any of the above systems, and the like.

[0042] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, that perform particular tasks or implement particular abstract data types. Computer systems / servers can be practiced in distributed cloud-computing environments with remote processing devices that are linked through a communications network. In a distributed cloud-computing environment, program modules can reside on local or remote computer system storage media including memory storage devices.

[0043] Summary

[0044] Two traditional approaches to single-channel speech separation are speech enhancement and computational auditory scene analysis. Speech enhancement methods require analysis of the entire data of speech and noise, and then through noise estimation of noisy speech, the clear speech is estimated. The simplest and most widely used enhancement method is the spectral subtraction method, in which the estimated noise power spectrum is subtracted from the noisy speech. In order to estimate the background noise, speech enhancement techniques generally assume that the background noise is stable, that is, the spectral characteristics do not change over time, or at least some stability than speech.

[0045] Computational auditory scene analysis is based on the perceptual theory of auditory scene analysis, and uses clustering cues such as pitch and onset. For example, the tandem algorithm separates speech by exchanging pitch estimation and pitch-based clustering.

[0046] The above two algorithms are based on certain scene constraints, so the speech separation effect is not good.

[0047] Speech separation methods based on an array of multiple microphones, such as beamforming, also known as spatial filter, enhance the signal coming from a specific direction by appropriate array structure, and reduce the interference from other directions. The simplest beamforming is a delay-sum technique that can add the signals from multiple microphones in the target direction with the same phase, and reduce the signals from other directions according to the phase difference. The amount of noise reduction depends on the interval, size and structure of the array, and generally increases with the number of microphones and the length of the array. Obviously, when the target source and the interference source are close to each other, the spatial filter cannot be applied. In addition, in the echo scenario, the effectiveness of beamforming is greatly reduced, and the determination of the sound source direction becomes ambiguous.

[0048] Another conventional multi-channel separation technique is blind signal separation (BSS), which means estimating the source signal only according to the observed mixed signal without knowing the source signal and signal mixing parameters. Independent component analysis (ICA) is a new technology gradually developed to solve the problem of blind signal separation.

[0049] The above-mentioned conventional single-channel speech separation method has poor effect, and the conventional multi-channel speech separation method requires more microphones, has high cost, large data processing amount, and exists underdetermined and overdetermined scenarios, and has large use scenario limitation. The existing deep learning-based speech separation method usually only uses audio feature data, resulting in low accuracy of speech separation.

[0050] Embodiments of the present disclosure aim to solve the above technical problems, and combine single-channel mixed audio signals and target user's lip image sequence to perform speech separation by using a deep learning method, thereby greatly improving the accuracy and efficiency of speech separation.

[0051] Exemplary System

[0052] Figure 1 An exemplary system architecture 100 of a speech signal extraction method or a speech signal extraction apparatus to which embodiments of the present disclosure can be applied is shown.

[0053] As shown in Figure 1 The system architecture 100 can include a terminal device 101, a network 102, a server 103, a microphone 104 and a camera 105. The network 102 is used to provide a communication link medium between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0054] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various applications can be installed on the terminal device 101, such as a speech recognition application, an image recognition application, a search application, etc.

[0055] The microphone 104 and the camera 105 are used to collect a single-channel mixed audio signal and an image of a target user. The microphone 104 and the camera 105 can be directly connected with the terminal device 101, or connected with the terminal device 101 through the network 102, or connected with the server 103 through the network 102.

[0056] The terminal device 101 can be various electronic devices, including but not limited to mobile terminals such as a vehicle-mounted terminal, a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), etc., and fixed terminals such as a digital TV, a desktop computer, a smart home appliance, etc.

[0057] The server 103 can be a server providing various services, such as a background server processing an audio signal, an image, etc. uploaded by the terminal device 101. The background server can perform speech separation using the received single-channel mixed audio signal and image sequence to obtain a speech signal of the target user.

[0058] It should be noted that the speech signal extraction method provided by the embodiments of the present disclosure can be executed by the server 103 or the terminal device 101, and correspondingly, the speech signal extraction apparatus can be arranged in the server 103 or the terminal device 101.

[0059] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system architecture is only illustrative. Any number of terminal devices, networks and servers can be provided according to the implementation needs. In the case where the single-channel mixed audio signal and the image sequence do not need to be obtained remotely, the above system architecture can not include the network, but only include the server or the terminal device.

[0060] Exemplary Method

[0061] Figure 2 FIG. 1 is a schematic diagram of a system architecture according to an example embodiment of the present disclosure. The system architecture can include a terminal device 101, a network 102 and a server 103. The terminal device 101 can be connected with the server 103 through the network 102. Figure 1 The terminal device 101 or the server 103 shown in FIG. 1 can be an electronic device such as a mobile terminal, a fixed terminal, etc. Figure 2 The method shown in FIG. 2 includes the following steps:

[0062] In step 201, a single-channel mixed audio signal and an image sequence collected in a target area are obtained.

[0063] In this embodiment, the target area can be a spatial area provided with a microphone and a camera, and the type of the target area can include, but is not limited to, a vehicle interior, a room interior, etc. The single-channel mixed audio signal can be an audio signal collected by a single microphone, which can include a speech signal of at least one user and a noise signal, etc. The image sequence can be images taken by the camera of the user in the target area. It should be understood that the single-channel mixed audio signal and the image sequence in this embodiment are collected synchronously in the same time length (for example, 1 second).

[0064] In step 202, a target user in the target area is determined based on the image sequence.

[0065] Optionally, the camera can take an image of a single user in a certain area (for example, a driver seat, a front passenger seat, etc. in a vehicle), and if the electronic device identifies the user from the taken image sequence, the user is determined as the target user.

[0066] The camera can also take images of multiple users, identify the multiple users from the taken image sequence, and the electronic device determines one of the users as the target user for which the method is currently performed. For example, a user located in a specified central image area can be determined as the target user from the identified multiple users; or each user can be determined as the target user, and the method is performed once for each target user; or a user matching preset user feature data (for example, face feature data) can be identified from the image sequence based on the user feature data, and the user is determined as the target user.

[0067] In step 203, a lip region image sequence of the target user is determined based on the image sequence.

[0068] Specifically, the images in the image sequence can include the lip region of the target user, and the electronic device can extract the lip region images from the images included in the image sequence based on a lip image detection method (for example, a lip region image is determined based on a face key point detection method), to obtain the lip region image sequence.

[0069] Generally, the size of the lip region images extracted from the image sequence can be adjusted to a fixed size (for example, 96x96), to obtain the lip region image sequence with uniform size.

[0070] In step 204, lip state feature data is determined based on the lip region image sequence.

[0071] The lip state feature data is used to represent the change characteristics of the mouth shape. Generally, the electronic device can identify the lip contour feature data (e.g., including the distance between the corners of the mouth, the distance between the upper and lower lips, etc.) of each lip region image in the sequence of lip region images, and combine the lip contour feature data of each lip region image into the lip state feature data. It should be understood that based on the sequence of lip region images, the method of determining the lip state feature data can determine the lip state feature data by using methods such as lip reading, which will not be described here.

[0072] In step 205, audio feature data is determined based on the single-channel mixed audio signal.

[0073] Optionally, the electronic device can determine the audio feature data based on a neural network method. For example, the neural network can include but is not limited to RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), UNet (U-shaped network), Complex UNet, and Transformer architecture based on self-attention mechanism and cross-domain attention mechanism.

[0074] In step 206, the lip state feature data and the audio feature data are fused to obtain fusion feature data.

[0075] The fusion of the lip state feature data and the audio feature data can be realized by various methods, such as concat feature fusion method, elemwise_add feature fusion method, single gate feature fusion method, attention feature fusion method, etc.

[0076] In step 207, the speech signal of the target user is extracted from the single-channel mixed audio signal based on the fusion feature data.

[0077] Optionally, the neural network can be used to decode the fusion feature data to obtain mask data, multiply the mask data with the frequency domain signal (e.g., obtained by performing short-time Fourier transform on the single-channel mixed audio signal) of the single-channel mixed audio signal to obtain feature data representing the speech signal of the target user, and then perform processing such as inverse Fourier transform on the feature data representing the speech signal of the target user to obtain the time-domain speech signal.

[0078] The method provided by the above embodiments of the present disclosure comprises: acquiring a single-channel mixed audio signal collected in a target region and an image sequence of a lip region of a target user, then determining lip state feature data based on the image sequence of the lip region and determining audio feature data based on the single-channel mixed audio signal, then fusing the lip state feature data and the audio feature data to obtain fused feature data, and finally extracting a speech signal of the target user from the single-channel mixed audio signal based on the fused feature data, so that multi-modal speech separation combining audio signals and lip images is realized, the feature data used for speech separation is more abundant, and compared with a method of performing speech separation only by using a single-mode audio signal, the method of multi-modal speech separation provided by the embodiments of the present disclosure has higher accuracy of the extracted speech signal of the target user. In addition, since only a single microphone is used to collect an audio signal, the hardware cost can be reduced, and the data processing amount is also reduced. A traditional speech separation method for a single-channel mixed audio signal uses a complex algorithm, and a certain convergence time is required during calculation, so that the delay time of speech separation is relatively long. The method provided by the embodiments of the present disclosure combines lip image feature data, and does not need to use a traditional speech separation algorithm, so that the delay time of speech separation is effectively reduced. In addition, in a multi-person scene, only the lip image sequences of different users are obtained, the method provided by the embodiments of the present disclosure is performed on different lip image sequences, and speech signals of multiple persons can be extracted, so that the scalability of the method is effectively improved.

[0079] In some optional implementations, as shown in Figure 3 Step 205 comprises:

[0080] Step 2051, pre-processing the single-channel mixed audio signal to obtain to-be-encoded data.

[0081] The pre-processing method can comprise converting the single-channel mixed audio signal in the time domain to the frequency domain, compressing the frequency domain signal, and the like.

[0082] Step 2052, encoding the to-be-encoded data by using a down-sampling module of a pre-trained first neural network model to obtain audio feature data.

[0083] The first neural network model can be a UNet network, and a structural diagram of the UNet network is as shown in Figure 4 401 is a down-sampling module, that is, the left half of the UNet network. The down-sampling module can perform a series of convolution, pooling and the like on the to-be-encoded data, and convert large-scale data into small-scale data. For example, 403 in Figure 4 is the audio feature data.

[0084] As shown in Figure 3 Step 207 comprises:

[0085] Step 2071: Use the upsampling module of the first neural network model to decode the fused feature data to obtain mask data.

[0086] like Figure 4 As shown, 402 is the upsampling module, which is the right half of the UNet network. After the audio feature data 403 and the lip state feature data 404 are fused, the fused feature data 405 is input into the upsampling module 402. The upsampling module 402 can restore the scale of the small-scale fused feature data to the large-scale mask data.

[0087] Step 2072: Extract the target user's speech signal from the single-channel mixed audio signal based on the mask data.

[0088] The mask data is used to filter the frequency domain signal of the single-channel mixed audio signal (e.g., obtained by performing a short-time Fourier transform on the single-channel mixed audio signal) to obtain the frequency domain signal of the target user's speech signal. Optionally, the mask data can be directly multiplied with the frequency domain signal of the single-channel mixed audio signal to obtain the frequency domain signal of the target user's speech signal; or the mask data can be normalized first using an activation function such as tanh, and then the normalized mask data can be multiplied with the frequency domain signal of the single-channel mixed audio signal to obtain the frequency domain signal of the target user's speech signal. Then, the frequency domain signal of the target user's speech signal is processed by methods such as inverse Fourier transform to obtain the time domain speech signal.

[0089] The first neural network model can be trained using machine learning methods. Specifically, training samples can be pre-acquired, including sample data to be encoded, sample lip state feature data, and labeled mask data. The sample data to be encoded can be used as input to a downsampling module, and the audio feature data output by the downsampling module is fused with the sample lip state feature data to obtain fused feature data. This fused feature data is then input to an upsampling module, and the labeled mask data corresponding to the input sample data to be encoded is used as the expected output of the upsampling module to train the initial first neural network model. For each training input of the sample data to be encoded and the sample lip state feature data, the actual output can be obtained. The actual output is the mask data actually output by the initial first neural network model. Then, gradient descent and backpropagation methods can be used to adjust the parameters of the initial first neural network model based on the actual and expected outputs. The model obtained after each parameter adjustment is used as the initial first neural network model for the next training. Training ends when preset training termination conditions are met (e.g., the loss value calculated based on a preset loss function converges, or the number of training iterations exceeds a preset number). This process completes the training of the first neural network model.

[0090] The embodiment pre-processes the single-channel mixed audio signal to obtain to-be-encoded data, and encodes and decodes the to-be-encoded data by using the downsampling module and the upsampling module of the first neural network model. Since the first neural network model is usually a UNet network, and the UNet network is usually used for semantic segmentation, the first neural network model is used to help accurately segment the feature data of the speech signal of the target user from the audio feature data, thereby effectively improving the accuracy of extracting the speech signal of the target user.

[0091] In some optional implementations, as shown in Figure 5 Step 2051 includes:

[0092] 20511. Perform frequency domain conversion on the single-channel mixed audio signal to obtain frequency domain data.

[0093] Specifically, the method of performing frequency domain conversion on the single-channel mixed audio signal can be implemented by various means, such as STFT (Short-Time Fourier Transform), DFT (Discrete Fourier Transform), etc.

[0094] 20512. Compress the frequency domain data to obtain to-be-encoded data.

[0095] The purpose of compressing the frequency domain data is to reduce the numerical range of the frequency domain data. The method of compressing the frequency domain data can be implemented based on various ways, for example, an exponential compression method can be used, that is, all numerical values included in the frequency domain data are taken to a preset power (for example, 0.3).

[0096] The embodiment performs frequency domain conversion on the single-channel mixed audio signal in the time domain, and then compresses the frequency domain data to obtain to-be-encoded data, which can reduce the numerical range of the frequency domain data, reduce the data processing difficulty of the neural network, and further improve the efficiency of extracting the speech signal of the target user.

[0097] In some optional implementations, as shown in Figure 6 Step 206 includes:

[0098] Step 2061. Merge the audio feature data and the lip state feature data to obtain merged feature data.

[0099] Specifically, the method of merging the audio feature data and the lip state feature data can be implemented by various means. For example, each channel included in the audio feature data and the lip state feature data can be directly merged; or the audio feature data and the lip state feature data can be merged by using a concat feature fusion method or the like.

[0100] At step 2062, the audio feature data and the merged feature data are fused to generate first fused feature data.

[0101] Optionally, the method of fusing the audio feature data and the merged feature data can include but is not limited to an elemwise_add feature fusion method, a single gate feature fusion method, an attention feature fusion method, etc.

[0102] At step 2063, the lip state feature data and the merged feature data are fused to generate second fused feature data.

[0103] Optionally, the method of fusing the lip state feature data and the merged feature data can also include but is not limited to an elemwise_add feature fusion method, a single gate feature fusion method, an attention feature fusion method, etc.

[0104] At step 2064, the first fused feature data and the second fused feature data are merged into fused feature data.

[0105] Specifically, the method of merging the first fused feature data and the second fused feature data can be implemented in various manners. For example, the respective channels included in the first fused feature data and the second fused feature data can be directly merged; or the first fused feature data and the second fused feature data can be merged by using a concat feature fusion method, etc.

[0106] Based on the audio feature data and the lip state feature data, the embodiment realizes more sufficient feature fusion for the audio feature data and the lip state feature data by using the multiple merging and fusing manner, which helps to make the fused feature data express more abundant audio features and visual features, and improves the accuracy of extracting the speech signal of the target user. In the scene of lip occlusion, etc., since the two kinds of feature data are sufficiently fused, the error rate of speech signal extraction can be effectively reduced.

[0107] In some optional implementation manners, as shown in Figure 7 Step 2062 includes:

[0108] At step 20621, the first convolutional layer and the first activation function included in the pre-trained second neural network model are used to perform first convolutional processing on the merged feature data to obtain first feature data.

[0109] The second neural network model can be a neural network model parallel to the first neural network described in the optional embodiments above, or it can be included in the first neural network model, that is, the second neural network model serves as a fusion module of the first neural network model. During training, the first neural network model and the second neural network model can be jointly trained using the same training samples.

[0110] The first activation function described above is used to normalize the data output from the first convolutional layer, ensuring that the values ​​of the first feature data fall within the range of 0 to 1. Optionally, the first activation function can be the tanh activation function.

[0111] Step 20622: Using the second convolutional layer and the second activation function included in the second neural network model, perform a second convolutional process on the merged feature data to obtain the first weight data.

[0112] Optionally, the second activation function can be the sigmoid activation function.

[0113] Step 20623: Generate second feature data based on the first feature data and the first weight data.

[0114] Typically, the first feature data and the first weight data can be multiplied element-wise (elemwise_mul) to obtain the second feature data. Alternatively, the first feature data and the first weight data can be multiplied element-wise, and then a corresponding bias can be added to each multiplied value to obtain the second feature data.

[0115] Step 20624: Generate first fused feature data based on audio feature data and second feature data.

[0116] Optionally, the audio feature data can be directly fused with the second feature data to obtain the first fused feature data by using methods such as elementwise addition (elemwise_add) or concat fusion.

[0117] This embodiment performs two convolutional processes on the merged feature data and generates second feature data based on the results of the two convolutional processes. This allows for the extraction of more common features representing audio feature data and lip state feature data from the merged feature data. When combined with the audio feature data, the resulting first fused feature data can simultaneously express the features of the audio and the common features of the audio and lip state, thereby helping to extract the target user's speech signal more accurately from a single-channel mixed audio signal.

[0118] In some alternative implementations, such as Figure 8 As shown, step 20624 includes:

[0119] Step 206241: Using the third convolutional layer and third activation function included in the second neural network model, the merged feature data is processed by the third convolution to obtain the second weight data.

[0120] Optionally, the third activation function can be the sigmoid activation function.

[0121] Step 206242: Generate third feature data based on audio feature data and second weight data.

[0122] Specifically, the method for generating the third feature data in this step can be the same as the method for generating the second feature data in step 20623 above. For example, the audio feature data is multiplied element-wise with the second weight data to obtain the third feature data.

[0123] Step 206243: Generate first fused feature data based on the third feature data and the second feature data.

[0124] Optionally, the third feature data and the second feature data can be fused together using methods such as elementwise_add or concat to obtain the first fused feature data.

[0125] This embodiment obtains second weight data by convolving the merged feature data, and obtains third feature data based on the audio feature data and the merged feature data. The third feature data and the second feature data are then fused to obtain first fused feature data. Since the third feature data is obtained by combining the audio feature data and the merged feature data, the third feature data can represent common features of audio and lip state in addition to mainly representing audio features. Thus, the obtained first fused feature data fully integrates lip state features on the basis of using audio features as the main feature representation, which helps to make the features represented by the final fused feature data richer and more targeted, and improves the accuracy of extracting the speech signal of the target user.

[0126] In some alternative implementations, such as Figure 9 As shown, step 2063 includes:

[0127] Step 20631: Using the fourth convolutional layer and fourth activation function included in the second neural network model, the merged feature data is processed by the fourth convolution to obtain the fourth feature data.

[0128] Step 20632: Using the fifth convolutional layer and fifth activation function included in the second neural network model, the merged feature data is processed by the fifth convolution to obtain the third weight data.

[0129] Step 20633: Generate the fifth feature data based on the fourth feature data and the third weight data.

[0130] In step 20634, the second fusion feature data is generated based on the lip state feature data and the fifth feature data.

[0131] It should be noted that the steps included in this embodiment are basically the same as the steps described in the corresponding embodiment, the processing process and the network structure used are basically the same, and the difference lies in that the data processed by the two is different. Figure 7 The steps, processing process and network structure used in this embodiment are basically the same as those described in the corresponding embodiment, and the difference lies in that the data processed by the two is different.

[0132] In this embodiment, the fifth feature data is generated based on the merged feature data by twice convolution processing, and the fifth feature data is generated based on the results of the twice convolution processing. The features representing the audio feature data and the lip state feature data with more commonality can be extracted from the merged feature data. Then, the second fusion feature data is obtained by combining the lip state feature data, which can express the features of the lip state and the commonality of the audio and the lip state, thereby helping to more accurately extract the speech signal of the target user from the single-channel mixed audio signal.

[0133] In some optional implementations, as shown in FIG. 20, step 20634 includes: Figure 10

[0134] In step 206341, the sixth convolution processing is performed on the merged feature data by using the sixth convolution layer and the sixth activation function included in the second neural network model, to obtain the fourth weight data.

[0135] In step 206342, the sixth feature data is generated based on the lip state feature data and the fourth weight data.

[0136] In step 206343, the second fusion feature data is generated based on the sixth feature data and the fifth feature data.

[0137] It should be noted that the steps included in this embodiment are basically the same as the steps described in the corresponding embodiment, the processing process and the network structure used are basically the same, and the difference lies in that the data processed by the two is different. Figure 8 The steps, processing process and network structure used in this embodiment are basically the same as those described in the corresponding embodiment, and the difference lies in that the data processed by the two is different.

[0138] In this embodiment, the fourth weight data is obtained by performing convolution on the merged feature data, the sixth feature data is obtained based on the lip state feature data and the merged feature data, and the second fusion feature data is obtained by fusing the sixth feature data and the fifth feature data. Since the sixth feature data is obtained by combining the audio feature data and the merged feature data, the sixth feature data can represent the commonality of the audio and the lip state on the basis of mainly representing the lip state feature data. Therefore, the second fusion feature data obtained on the basis of mainly representing the lip state feature data can fully fuse the audio feature, which helps to make the feature represented by the finally obtained fusion feature data more rich and more targeted, and improves the accuracy of extracting the speech signal of the target user.​

[0139] Referring to Figure 11 , Figure 11 is an example of a method for generating fusion feature data according to the voice signal extraction method of the present embodiment. As shown in Figure 11 , the merged feature data is passed through a first convolutional layer and a first activation function 1101 (e.g., a tanh activation function) to generate first feature data; the merged feature data is passed through a second convolutional layer and a second activation function 1102 (e.g., a sigmoid activation function) to generate first weight data. The first feature data and the first weight data are element-wise multiplied 1103 to generate second feature data. The merged feature data is passed through a third convolutional layer and a third activation function 1104 (e.g., a sigmoid activation function) to generate second weight data; the audio feature data is again element-wise multiplied 1105 with the second weight data to generate third feature data. The third feature data and the second feature data are fused by an elemwise_add method 1106 to generate first fusion feature data.

[0140] The merged feature data is passed through a fourth convolutional layer and a fourth activation function 1107 (e.g., a tanh activation function) to generate fourth feature data; the merged feature data is passed through a fifth convolutional layer and a fifth activation function 1108 (e.g., a sigmoid activation function) to generate third weight data. The fourth feature data and the third weight data are element-wise multiplied 1109 to generate fifth feature data. The merged feature data is passed through a sixth convolutional layer and a sixth activation function 1110 (e.g., a sigmoid activation function) to generate fourth weight data; the lip state feature data is again element-wise multiplied 1111 with the fourth weight data to generate sixth feature data. The sixth feature data and the fifth feature data are fused by an elemwise_add method 1112 to generate second fusion feature data.

[0141] The first fusion feature data and the second fusion feature data are merged to generate fusion feature data.

[0142] Figure 11 The method for generating fusion feature data shown in the figure can also be referred to as a dual gate method. The method performs feature fusion twice according to similar steps and network structures, with the audio feature data as the main feature data and the lip state feature data as the slave feature data, and with the lip state feature data as the main feature data and the audio feature data as the slave feature data. The weight data generated during execution is used as a gate parameter to operate on the corresponding feature data, which can more specifically extract information representing the target user's voice from the audio feature data and the lip state feature data, thereby achieving more accurate extraction of the target user's voice signal.

[0143] Exemplary Apparatus

[0144] Figure 12 is a structural schematic diagram of a speech signal extraction device provided by an example embodiment of the present disclosure. The embodiment can be applied to an electronic device, such as a vehicle, a smart home, etc. Figure 12 As shown in the figure, the speech signal extraction device includes: an acquisition module 1201 configured to acquire a single-channel mixed audio signal and an image sequence collected in a target area; a first determination module 1202 configured to determine a target user in the target area based on the image sequence; a second determination module 1203 configured to determine a lip region image sequence of the target user based on the image sequence; a third determination module 1204 configured to determine lip state feature data based on the lip region image sequence; a fourth determination module 1205 configured to determine audio feature data based on the single-channel mixed audio signal; a fusion module 1206 configured to fuse the lip state feature data and the audio feature data to obtain fusion feature data; and an extraction module 1207 configured to extract a speech signal of the target user from the single-channel mixed audio signal based on the fusion feature data.

[0145] In the embodiment, the acquisition module 1201 can acquire a single-channel mixed audio signal and an image sequence collected in a target area. The target area can be a spatial area provided with a microphone and a camera, and the type of the target area can include but is not limited to a vehicle interior, a room interior, etc. The single-channel mixed audio signal can be an audio signal collected by a single microphone, which can include speech signals of at least one user and noise signals, etc. The image sequence can be an image taken by the camera of a user in the target area. It should be understood that the single-channel mixed audio signal and the image sequence in the embodiment are collected synchronously in the same time length (for example, 1 second).

[0146] In the embodiment, the first determination module 1202 can determine a target user in the target area based on the image sequence. Here, the target user refers to a specific user.

[0147] Optionally, the camera can take an image of a single user in a specific area (for example, a driver seat, a copilot seat, etc. in a vehicle). If the first determination module 1202 identifies the user from the taken image sequence, the user is determined as the target user.

[0148] The camera can also capture multiple users, identify multiple users from the captured image sequence, and the first determining module 1202 determines one of the users as a target user for which the method is currently performed. In this embodiment, the second determining module 1203 can determine a lip region image sequence of the target user based on the image sequence. Specifically, the images in the image sequence can include the lip region of the target user, and the second determining module 1203 can extract lip region images from the images included in the image sequence based on a lip image detection method (such as a lip region image based on a face key point detection method), to obtain a lip region image sequence.

[0149] Generally, the size of the lip region image extracted from the image sequence can be adjusted to a fixed size (such as 96x96), to obtain a lip region image sequence with uniform size.

[0150] In this embodiment, the third determining module 1204 can determine lip state feature data based on the lip region image sequence.

[0151] The lip state feature data is used to represent the change characteristics of the mouth shape. Generally, the third determining module 1204 can identify the lip shape feature data (such as the distance between the corners of the mouth, the distance between the upper and lower lips, etc.) of each lip region image in the lip region image sequence, and combine the lip shape feature data of each lip region image into lip state feature data. It should be understood that the method of determining the lip state feature data based on the lip region image sequence can determine the lip state feature data using methods such as lip reading, which will not be described here.

[0152] In this embodiment, the fourth determining module 1205 can determine audio feature data based on the single-channel mixed audio signal.

[0153] Optionally, the fourth determining module 1205 can also determine the audio feature data based on a neural network method. For example, the neural network can include but is not limited to RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), UNet (U-shaped network), Complex UNet, and Transformer architecture based on self-attention mechanism and cross-domain attention mechanism.

[0154] In this embodiment, the fusion module 1206 can fuse the lip state feature data and the audio feature data to obtain fusion feature data.

[0155] The fusion of the lip state feature data and the audio feature data can be achieved by various methods, such as a concat feature fusion method, an elemwise_add feature fusion method, a single gate feature fusion method, an attention feature fusion method, and the like.

[0156] In this embodiment, the extraction module 1207 can extract the speech signal of the target user from the single-channel mixed audio signal based on the fusion feature data.

[0157] Optionally, the fusion feature data can be decoded by using a neural network to obtain mask data, the mask data is multiplied by the audio feature data to obtain feature data representing the speech signal of the target user, and then the feature data representing the speech signal of the target user is processed, such as inverse Fourier transform, to obtain the time-domain speech signal.

[0158] Referring to Figure 13 , Figure 13 is a structural schematic diagram of a speech signal extraction device provided by another exemplary embodiment of the present disclosure.

[0159] In some optional implementations, the fourth determination module 1205 includes: a preprocessing unit 12051 configured to preprocess the single-channel mixed audio signal to obtain to-be-encoded data; and an encoding unit 12052 configured to encode the to-be-encoded data by using a down-sampling module of a pre-trained first neural network model to obtain the audio feature data; the extraction module 1207 includes: a decoding unit 12071 configured to decode the fusion feature data by using an up-sampling module of the first neural network model to obtain mask data; and an extraction unit 12072 configured to extract the speech signal of the target user from the single-channel mixed audio signal based on the mask data.

[0160] In some optional implementations, the preprocessing unit 12051 includes: a conversion sub-unit 120511 configured to perform frequency domain conversion on the single-channel mixed audio signal to obtain frequency domain data; and a compression sub-unit 120512 configured to compress the frequency domain data to obtain the to-be-encoded data.

[0161] In some optional implementations, the fusion module 1206 includes: a first merging unit 12061 configured to merge the audio feature data and the lip state feature data to obtain merged feature data; a first fusion unit 12062 configured to fuse the audio feature data and the merged feature data to generate first fusion feature data; a second fusion unit 12063 configured to fuse the lip state feature data and the merged feature data to generate second fusion feature data; and a second merging unit 12064 configured to merge the first fusion feature data and the second fusion feature data into the fusion feature data.

[0162] In some optional implementation manners, the first fusion unit 12062 includes: a first processing sub-unit 120621, configured to perform first convolution processing on the merged feature data by using a first convolution layer and a first activation function included in the pre-trained second neural network model, to obtain first feature data; a second processing sub-unit 120622, configured to perform second convolution processing on the merged feature data by using a second convolution layer and a second activation function included in the second neural network model, to obtain first weight data; a first generation sub-unit 120623, configured to generate second feature data based on the first feature data and the first weight data; and a second generation sub-unit 120624, configured to generate the first fusion feature data based on the audio feature data and the second feature data.

[0163] In some optional implementation manners, the second generation sub-unit 120624 is further configured to: perform third convolution processing on the merged feature data by using a third convolution layer and a third activation function included in the second neural network model, to obtain second weight data; generate third feature data based on the audio feature data and the second weight data; and generate the first fusion feature data based on the third feature data and the second feature data.

[0164] In some optional implementation manners, the second fusion unit 12063 includes: a third processing sub-unit 120631, configured to perform fourth convolution processing on the merged feature data by using a fourth convolution layer and a fourth activation function included in the second neural network model, to obtain fourth feature data; a fourth processing sub-unit 120632, configured to perform fifth convolution processing on the merged feature data by using a fifth convolution layer and a fifth activation function included in the second neural network model, to obtain third weight data; a third generation sub-unit 120633, configured to generate fifth feature data based on the fourth feature data and the third weight data; and a fourth generation sub-unit 120634, configured to generate the second fusion feature data based on the lip state feature data and the fifth feature data.

[0165] In some optional implementation manners, the fourth generation sub-unit 120634 is further configured to: perform sixth convolution processing on the merged feature data by using a sixth convolution layer and a sixth activation function included in the second neural network model, to obtain fourth weight data; generate sixth feature data based on the lip state feature data and the fourth weight data; and generate the second fusion feature data based on the sixth feature data and the fifth feature data.

[0166] The speech signal extraction apparatus provided in the above embodiments of this disclosure acquires a single-channel mixed audio signal and a sequence of lip region images of the target user within a target area. Then, based on the lip region image sequence, it determines lip state feature data, and based on the single-channel mixed audio signal, it determines audio feature data. Next, it fuses the lip state feature data and the audio feature data to obtain fused feature data. Finally, based on the fused feature data, it extracts the target user's speech signal from the single-channel mixed audio signal. This achieves multimodal speech separation by combining audio signals and lip images. The speech separation utilizes richer feature data, and compared to methods that rely solely on single-modal audio signals for speech separation, the multimodal speech separation method provided in this disclosure extracts the target user's speech signal with higher accuracy. Furthermore, since only a single microphone is needed to acquire the audio signal, hardware costs and data processing volume can be reduced. Traditional speech separation methods for single-channel mixed audio signals employ complex algorithms that require convergence time, resulting in significant delays in speech separation. The method provided in this disclosure, by incorporating lip image feature data, eliminates the need for traditional speech separation algorithms, effectively reducing the delay. Furthermore, in multi-user scenarios, only lip image sequences from different users need to be obtained; the method provided in this application can be applied to each lip image sequence separately to extract speech signals from multiple users, thereby significantly improving the scalability of the method.

[0167] Exemplary Electronic Device

[0168] Below, for reference Figure 14 To describe an electronic device according to embodiments of the present disclosure. The electronic device may be as follows: Figure 1 The terminal device 101 and server 103 shown, or either one or both, or a standalone device independent of them, can communicate with the terminal device 101 and server 103 to receive the collected input signals from them.

[0169] Figure 14 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0170] like Figure 14 As shown, the electronic device 1400 includes one or more processors 1401 and memory 1402.

[0171] The processor 1401 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1400 to perform desired functions.

[0172] The memory 1402 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk drives, solid-state drives, and / or the like. The computer-readable storage media can store one or more computer program instructions executable by the processor 1401 to implement the method of extracting a voice signal of various embodiments of the present disclosure described above and / or other desired functions. Various contents such as a monaural mixed audio signal, an image sequence, and the like can also be stored in the computer-readable storage media.

[0173] In one example, the electronic device 1400 can further include an input device 1403 and an output device 1404, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0174] For example, when the electronic device is the terminal device 101 or the server 103, the input device 1403 can be a microphone, a camera, a mouse, a keyboard, and the like, which are used to input a monaural mixed audio signal, an image sequence, various commands, and the like. When the electronic device is a stand-alone device, the input device 1403 can be a communication network connector, which is used to receive a monaural mixed audio signal, an image sequence, various commands, and the like, which are input from the terminal device 101 and the server 103.

[0175] The output device 1404 can output various information including a voice signal of a target user to the outside. The output device 1404 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0176] Of course, in order to simplify, Figure 14 Only some of the components of the electronic device 1400 related to the present disclosure are shown in FIG. 14, and components such as a bus, an input / output interface, and the like are omitted. In addition to this, the electronic device 1400 can further include any other appropriate components according to a specific application.

[0177] Exemplary Computer Program Product and Computer-Readable Storage Medium

[0178] In addition to the above-described method and device, an embodiment of the present disclosure can be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the method of extracting a voice signal according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of the specification.

[0179] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0180] Furthermore, embodiments of the present disclosure can also be a computer readable storage medium, having stored thereon computer program instructions which, when executed by a processor, cause the processor to perform the steps described in the above "Exemplary Method" section of the present specification for the method of extracting a voice signal according to various embodiments of the present disclosure.

[0181] The computer readable storage medium can be any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0182] The above describes the basic principles of the present disclosure in combination with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present disclosure are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as the must-have of each embodiment of the present disclosure. In addition, the above specific details are only for the purpose of example and understanding, and the above details do not limit the present disclosure to the must-use specific details.

[0183] Each embodiment in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between each embodiment can be referred to each other. For system embodiments, since they are basically corresponding to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0184] The block diagrams of devices, apparatuses, equipment, systems referred to in this disclosure are merely illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. These devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner as will be appreciated by those skilled in the art. Words such as "include," "contain," "have," and the like are open-ended words that are to be interpreted to mean "including but not limited to," and are not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein. The words "or" and "and" as used herein are to be interpreted as the word "and / or," and are not to be interpreted as requiring both features, elements, and / or steps disclosed herein. The word "such as" as used herein is to be interpreted as the phrase "such as but not limited to," and is not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein.

[0185] The methods and apparatuses of this disclosure can be implemented in a number of ways. For example, the methods and apparatuses of this disclosure can be implemented using software, hardware, firmware, or any combination of these. The above described order of steps for the methods is merely illustrative, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, the disclosure can also be implemented as a program recorded on a recording medium, which includes machine readable instructions for implementing the methods according to the disclosure. Thus, the disclosure also covers a recording medium storing a program for executing the methods according to the disclosure.

[0186] It is also important to note that the devices, equipment, and methods of this disclosure can be embodied in a variety of ways. These variations are contemplated as being within the scope of the present disclosure.

[0187] The above description of the disclosed aspects is given for illustrative purposes and is not intended to limit the scope of the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0188] The above description has been given for illustrative and descriptive purposes. In addition, this description is not intended to limit embodiments of the disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those of skill in the art will recognize certain modifications, permutations, additions, and sub-combinations thereof.

Claims

1. A method for extracting a speech signal from a single-channel mixed audio signal, comprising: obtaining a single-channel mixed audio signal and an image sequence collected in a target region; determining a target user in the target region based on the image sequence; determining a lip region image sequence of the target user based on the image sequence; determining lip state feature data based on the lip region image sequence; determining audio feature data based on the single-channel mixed audio signal; fusing the lip state feature data and the audio feature data to obtain fused feature data; and extracting a speech signal of the target user from the single-channel mixed audio signal based on the fused feature data; wherein the fusing the lip state feature data and the audio feature data to obtain fused feature data comprises: merging the audio feature data and the lip state feature data to obtain merged feature data; fusing the audio feature data and the merged feature data to generate first fused feature data; fusing the lip state feature data and the merged feature data to generate second fused feature data; and merging the first fused feature data and the second fused feature data into the fused feature data. The determining audio feature data based on the single-channel mixed audio signal comprises: pre-processing the single-channel mixed audio signal to obtain to-be-encoded data; and encoding the to-be-encoded data using a down-sampling module of a pre-trained first neural network model to obtain the audio feature data. The extracting a speech signal of the target user from the single-channel mixed audio signal based on the fused feature data comprises: decoding the fused feature data using an up-sampling module of the first neural network model to obtain mask data; and extracting the speech signal of the target user from the single-channel mixed audio signal based on the mask data. The pre-processing the single-channel mixed audio signal to obtain to-be-encoded data comprises: performing frequency domain conversion on the single-channel mixed audio signal to obtain frequency domain data; and compressing the frequency domain data to obtain the to-be-encoded data. The fusing the audio feature data and the merged feature data to generate first fused feature data comprises: performing first convolution processing on the merged feature data using a first convolution layer and a first activation function included in a pre-trained second neural network model to obtain first feature data; performing second convolution processing on the merged feature data using a second convolution layer and a second activation function included in the second neural network model to obtain first weight data; generating second feature data based on the first feature data and the first weight data; and generating the first fused feature data based on the audio feature data and the second feature data. The generating the first fused feature data based on the audio feature data and the second feature data comprises: ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 2. The method of claim 1, wherein, ​ ​ ​ ​ ​ ​ 3. The method of claim 2, wherein, ​ ​ ​ 4. The method of claim 1, wherein, ​ ​ ​ ​ ​ 5. The method of claim 4, wherein, ​ The second weight data is obtained by performing third convolution processing on the merged feature data by using a third convolution layer and a third activation function included in the second neural network model; The third feature data is generated based on the audio feature data and the second weight data; The first fusion feature data is generated based on the third feature data and the second feature data.

6. The method of claim 4, wherein, The second fusion feature data is generated by fusing the lip state feature data and the merged feature data, including: The fourth feature data is obtained by performing fourth convolution processing on the merged feature data by using a fourth convolution layer and a fourth activation function included in the second neural network model; The third weight data is obtained by performing fifth convolution processing on the merged feature data by using a fifth convolution layer and a fifth activation function included in the second neural network model; The fifth feature data is generated based on the fourth feature data and the third weight data; The second fusion feature data is generated based on the lip state feature data and the fifth feature data.

7. The method of claim 6, wherein, The second fusion feature data is generated based on the lip state feature data and the fifth feature data, including: The fourth weight data is obtained by performing sixth convolution processing on the merged feature data by using a sixth convolution layer and a sixth activation function included in the second neural network model; The sixth feature data is generated based on the lip state feature data and the fourth weight data; The second fusion feature data is generated based on the sixth feature data and the fifth feature data.

8. An apparatus for extracting a speech signal from a single-channel mixed audio signal, comprising: an acquisition module configured to acquire a single-channel mixed audio signal and an image sequence collected in a target region; a first determination module configured to determine a target user in the target region based on the image sequence; a second determination module configured to determine a lip region image sequence of the target user based on the image sequence; a third determination module configured to determine lip state feature data based on the lip region image sequence; a fourth determination module configured to determine audio feature data based on the single-channel mixed audio signal; a fusion module configured to fuse the lip state feature data and the audio feature data to obtain fusion feature data; the fusion module includes: a first merging unit configured to merge the audio feature data and the lip state feature data to obtain merged feature data; a first fusion unit configured to fuse the audio feature data and the merged feature data to generate first fusion feature data; a second fusion unit configured to fuse the lip state feature data and the merged feature data to generate second fusion feature data; and a second merging unit configured to merge the first fusion feature data and the second fusion feature data into the fusion feature data; an extraction module configured to extract a speech signal of the target user from the single-channel mixed audio signal based on the fusion feature data. 9.A computer readable storage medium, the storage medium storing a computer program for being executed by a processor to implement the method of any one of claims 1-7. 10.An electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Electronic equipment, voice recognition method thereof and medium

    CN114141230A