Voice signal extraction method and device, readable storage medium and electronic equipment
By combining lip region images and spatial location feature data of the microphone array, the problem of insufficient lip image quality in multimodal speech separation methods is solved, and high accuracy and stability of speech signal extraction are achieved even when lips are occluded or unclear.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
- Filing Date
- 2022-09-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing multimodal speech separation methods have high requirements for the quality of the speaker's lip images. When the lips are occluded or the lip images are unclear, visual information cannot be effectively utilized, which affects the speech separation effect.
By acquiring multi-channel mixed audio signals and image sequences, the spatial location feature data of the lip region image sequence and microphone array are determined. Combined with lip state features and audio feature data, the speech signal of the target user is extracted from the multi-channel mixed audio signals.
It improves the accuracy and stability of speech signal extraction and reduces the impact of image quality degradation when lips are obscured or the lip image quality is poor.
Smart Images

Figure CN115910038B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, computer-readable storage medium, and electronic device for extracting speech signals. Background Technology
[0002] With the continuous development of human-computer interaction methods, efficiency, accuracy, and convenience have become research goals in related fields. Multimodal speech separation, as a method of human-computer interaction, has been widely researched and applied. Multimodal speech separation refers to combining audio and images, using neural networks and other methods to perform multimodal fusion of auditory and visual signals to solve the problem of sound source separation. This method trains a model to simultaneously learn the features of audio and images, using images as an aid to better learn the vocal information of different speakers in the audio.
[0003] Current multimodal speech separation methods typically require high-quality images of the speaker's lips. When lips are obscured or the lip images are unclear, the speech separation effect is significantly affected. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a method, apparatus, computer-readable storage medium, and electronic device for extracting speech signals.
[0005] Embodiments of this disclosure provide a method for extracting speech signals. The method includes: acquiring a multi-channel mixed audio signal and an image sequence collected within a target area; identifying a target user within the target area; determining a lip region image sequence of the target user based on the image sequence; determining lip state feature data based on the lip region image sequence; determining audio feature data based on the multi-channel mixed audio signal; determining spatial position feature data of the target user's lips and a microphone array based on the lip region image sequence; and extracting the target user's speech signal from the multi-channel mixed audio signal based on the lip state feature data, audio feature data, and spatial position feature data.
[0006] According to another aspect of the present disclosure, a speech signal extraction apparatus is provided. The apparatus includes: an acquisition module for acquiring a multi-channel mixed audio signal and an image sequence collected within a target area; a first determination module for determining a target user within the target area; a second determination module for determining a lip region image sequence of the target user based on the image sequence; a third determination module for determining lip state feature data based on the lip region image sequence; a fourth determination module for determining audio feature data based on the multi-channel mixed audio signal; a fifth determination module for determining spatial position feature data of the target user's lips relative to a microphone array based on the lip region image sequence; and an extraction module for extracting the target user's speech signal from the multi-channel mixed audio signal based on the lip state feature data, audio feature data, and spatial position feature data.
[0007] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which stores a computer program for execution by a processor to implement the above-described method for extracting speech signals.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the above-described method for extracting voice signals.
[0009] Based on the speech signal extraction method, apparatus, computer-readable storage medium, and electronic device provided in the above embodiments of this disclosure, the method involves acquiring multi-channel mixed audio signals and image sequences collected within a target area. Then, based on the lip region image sequence, lip state feature data is determined. Based on the lip region image sequence, lip state feature data and spatial position feature data of the target user's lips and microphone array are determined. Audio feature data is determined based on the multi-channel mixed audio signal. Finally, based on the lip state feature data, audio feature data, and spatial position feature data, the speech signal of the target user is extracted from the multi-channel mixed audio signal. This disclosure achieves multimodal speech separation by combining multi-channel mixed audio signals and spatial position feature data. It effectively utilizes the positional relationship between the spatial position of the lips and the positions of multiple microphones as auxiliary information for speech separation, enabling more targeted tracking of the target user's lip position and improving the accuracy of speech signal extraction. In scenarios where lip occlusion or poor lip image quality occurs, the positional relationship between the lips and microphone array can be effectively utilized to reduce the impact of image quality degradation, thereby improving the stability of speech signal extraction.
[0010] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0011] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 This is the system diagram to which this disclosure applies.
[0013] Figure 2 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0014] Figure 3 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0015] Figure 4 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0016] Figure 5 This is a schematic diagram of the angle between the target line where the lip position is located and the reference line of the microphone array, provided in an exemplary embodiment of this disclosure.
[0017] Figure 6 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0018] Figure 7 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0019] Figure 8 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure.
[0020] Figure 9 This is an exemplary schematic diagram of generating fused feature data provided by an exemplary embodiment of the present disclosure.
[0021] Figure 10 This is a schematic diagram of the structure of a speech signal extraction device provided in an exemplary embodiment of the present disclosure.
[0022] Figure 11 This is a schematic diagram of the structure of a speech signal extraction device provided in another exemplary embodiment of this disclosure.
[0023] Figure 12 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0024] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0025] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0026] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0027] It should also be understood that in the embodiments disclosed herein, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0028] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0029] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0030] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0031] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0034] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0035] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0036] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0037] Application Overview
[0038] Current multimodal speech separation methods typically require high-quality images of the speaker's lips. When lips are obscured or the lip images are unclear, visual information cannot be fully utilized, which significantly affects the speech separation effect.
[0039] The embodiments disclosed herein aim to solve this problem by introducing spatial location feature data representing the positional relationship between the lips and the microphone array, based on the determined audio feature data and lip state feature data. Using these feature data for speech separation, the stability and accuracy of extracting the speech signal of the target user are effectively improved.
[0040] Exemplary System
[0041] Figure 1 An exemplary system architecture 100 is shown for a speech signal extraction method or speech signal extraction apparatus to which embodiments of the present disclosure may be applied.
[0042] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, server 103, microphone array 104, and camera 105. Network 102 serves as the medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0043] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various applications can be installed on terminal device 101, such as voice recognition applications, image recognition applications, search applications, etc.
[0044] Microphone array 104 and camera 105 are used to acquire multi-channel mixed audio signals and images of the target user. Microphone array 104 and camera 105 can be directly connected to terminal device 101, or connected to terminal device 101 via network 102. Microphone array 104 and camera 105 can also be connected to server 103 via network 102. Microphone array 104 and camera 105 are positioned within a target area, which can be any type of spatial area, such as inside a vehicle or room.
[0045] The microphone array 104 includes at least two microphones to collect sound within the target area and obtain a multi-channel mixed audio signal.
[0046] Terminal device 101 can be various electronic devices, including but not limited to mobile terminals such as vehicle terminals, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., as well as fixed terminals such as digital TVs, desktop computers, smart home appliances, etc.
[0047] Server 103 can be a server that provides various services, such as a backend server that processes audio signals and images uploaded by terminal device 101. The backend server can use the received multi-channel mixed audio signals and image sequences to perform speech separation to obtain the target user's speech signal.
[0048] It should be noted that the voice signal extraction method provided in the embodiments of this disclosure can be executed by the server 103 or by the terminal device 101. Accordingly, the voice signal extraction device can be set in the server 103 or in the terminal device 101.
[0049] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. In cases where multi-channel mixed audio signals and image sequences do not need to be acquired remotely, the above system architecture may exclude networks and servers, including only microphone arrays, cameras, and terminal devices.
[0050] Exemplary methods
[0051] Figure 2 This is a schematic flowchart of a speech signal extraction method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices (such as...). Figure 1 On the terminal device 101 or server 103 shown, such as Figure 2 As shown, the method includes the following steps:
[0052] Step 201: Acquire the multi-channel mixed audio signal and image sequence collected within the target area.
[0053] In this embodiment, the electronic device can acquire multi-channel mixed audio signals and image sequences collected within a target area. The target area can be defined as follows: Figure 1 The spatial area of the microphone array 104 and camera 105 shown can be, but is not limited to, the interior of a vehicle or room. The multi-channel mixed audio signal can be an audio signal acquired by the microphone array 104, which may include multiple channels, each corresponding to an audio signal acquired by one microphone. The multi-channel mixed audio signal includes at least one user's voice signal and noise signals, etc. The image sequence can be images captured by the camera 105 of the user within the target area. It should be understood that in this embodiment, the multi-channel mixed audio signal and the image sequence are acquired synchronously within the same duration (e.g., 1 second).
[0054] Step 202: Identify the target users within the target area.
[0055] In this embodiment, the electronic device can determine the target user within the target area based on various methods.
[0056] Optionally, the camera can capture images of a single user within a specific area (e.g., the driver's seat, passenger seat, etc. in a vehicle). If the electronic device identifies this user from the captured image sequence, it determines that user as the target user. Alternatively, the camera can capture images of multiple users, identify multiple users from the captured image sequence, and the electronic device identifies one of these users as the target user for which the method is currently being executed. For example, the user located in the central region of a specified image can be identified as the target user from among the identified users; or, each user can be identified as the target user, and the method can be executed once for each target user; or, based on preset user feature data (e.g., facial feature data), the user matching the user feature data can be identified from the image sequence, and that user can be identified as the target user.
[0057] Optionally, the electronic device can also use other methods to determine the target user within the target area. For example, the target area may include multiple sub-areas (e.g., the area where each seat is located is a sub-area), and each sub-area may have a button. When the electronic device detects that a button has been pressed, it determines from the image sequence the user in the sub-area corresponding to the pressed button as the target user. As another example, each sub-area may have a microphone. When the electronic device detects that the microphone in a certain sub-area has picked up a voice signal and the strength of the voice signal picked up by that microphone is the highest, it determines from the image sequence the user in that sub-area as the target user.
[0058] Step 203: Based on the image sequence, determine the image sequence of the target user's lip region.
[0059] In this embodiment, the electronic device can determine the image sequence of the target user's lip region based on the image sequence.
[0060] Specifically, the images in the image sequence may include the lip region of the target user. The electronic device may extract lip region images from the images included in the image sequence based on a lip image detection method (e.g., a facial key point detection method to determine the lip region image) to obtain a lip region image sequence.
[0061] Typically, the size of the lip region images extracted from the image sequence can be adjusted to a fixed size (e.g., 96×96) to obtain a lip region image sequence of uniform size.
[0062] Step 204: Determine lip state feature data based on the lip region image sequence.
[0063] In this embodiment, the electronic device can determine lip state feature data based on a sequence of lip region images. This lip state feature data is used to characterize changes in mouth shape. Typically, the electronic device can identify the lip shape feature data (e.g., the distance between the corners of the mouth, the distance between the upper and lower lips, etc.) of each lip region image in the sequence, and merge the lip shape feature data of each lip region image into lip state feature data. Determining lip state feature data based on a sequence of lip region images can be achieved using methods such as lip reading, which will not be elaborated upon here.
[0064] Step 205: Determine audio feature data based on the multi-channel mixed audio signal.
[0065] In this embodiment, the electronic device can determine audio feature data based on multi-channel mixed audio signals.
[0066] Specifically, the aforementioned audio feature data can be determined for the audio signal of any channel in a multi-channel mixed audio signal; or the feature data of the audio signal of each channel can be determined separately, and then the feature data of each channel can be fused into the aforementioned audio feature data.
[0067] Optionally, the electronic device can determine the audio feature data of a channel based on a neural network method. For example, the neural network can include, but is not limited to, RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory Network), UNet (U-shaped network), Complex UNet, and Transformer architecture based on self-attention mechanism and cross-domain attention mechanism.
[0068] Step 206: Based on the lip region image sequence, determine the spatial location feature data of the target user's lips and microphone array.
[0069] In this embodiment, the electronic device can determine the target user's lips based on a sequence of lip region images. Figure 1 Spatial location feature data of the microphone array 104 shown.
[0070] The spatial location feature data is used to characterize the spatial relationship between the target user's lips and the microphone array 104. This spatial location feature data can be obtained based on the position of the microphone array and the position between the target user's lips. The position of the microphone array can be pre-calibrated, and the position of the target user's lips can be obtained by identifying a sequence of lip region images. For example, based on the intrinsic, extrinsic, and pose information of the camera 104, the two-dimensional coordinates of the lip region in the original image captured by the camera are transformed to three-dimensional coordinates in the camera coordinate system or the world coordinate system.
[0071] As an example, spatial location feature data may include, but is not limited to: the distance between the lips and the reference point of the microphone array, the angle between the line connecting the lip position and the aforementioned reference point and the reference line of the microphone array, and the phase difference between the audio signals of each channel in the multi-channel mixed audio signal (representing the difference between the paths of sound emitted from the lip position to each microphone), etc.
[0072] The methods for determining the above-mentioned reference points, reference lines, and phase differences can be referred to in the following optional embodiments.
[0073] Step 207: Extract the target user's speech signal from the multi-channel mixed audio signal based on lip state feature data, audio feature data, and spatial location feature data.
[0074] In this embodiment, the electronic device can extract the target user's voice signal from a multi-channel mixed audio signal based on lip state feature data, audio feature data, and spatial location feature data.
[0075] Specifically, lip state feature data, audio feature data, and spatial location feature data can be fused first to obtain fused feature data. Then, a method such as a neural network can be used to decode the fused feature data to obtain mask data. The mask data is multiplied with the frequency domain data corresponding to the audio signal of any channel in the multi-channel mixed audio signal (e.g., obtained by performing a short-time Fourier transform on the audio signal of one channel) to obtain the frequency domain data of the target user's speech signal. Then, the frequency domain data of the target user's speech signal is processed by methods such as inverse Fourier transform to obtain the time domain speech signal.
[0076] The method provided in the above embodiments of this disclosure acquires multi-channel mixed audio signals and image sequences collected within a target area. Then, based on the lip region image sequence, it determines lip state feature data, and based on the lip region image sequence, it determines the spatial position feature data of the target user's lips and the microphone array. Based on the multi-channel mixed audio signal, it determines audio feature data. Finally, based on the lip state feature data, audio feature data, and spatial position feature data, it extracts the target user's speech signal from the multi-channel mixed audio signal. This disclosure achieves multimodal speech separation by combining multi-channel mixed audio signals and spatial position feature data. It effectively utilizes the positional relationship between the spatial position of the lips and the positions of multiple microphones as auxiliary information for speech separation, enabling more targeted tracking of the target user's lip position, thereby improving the accuracy of speech signal extraction. In scenarios where lip occlusion or poor lip image quality occurs, the positional relationship between the lips and the microphone array can be effectively utilized to reduce the impact of image quality degradation, thereby improving the stability of speech signal extraction.
[0077] In some alternative implementations, such as Figure 3 As shown, step 205 includes:
[0078] Step 2051: Perform frequency domain conversion on the multi-channel mixed audio signal to obtain frequency domain data.
[0079] Optionally, the audio signal of any channel can be frequency-domain converted to obtain frequency domain data, or the audio signal of each channel can be frequency-domain converted to obtain multi-channel frequency domain data, and then the multi-channel frequency domain data can be fused into single-channel frequency domain data according to a preset fusion method (e.g., averaging the frequency domain data of each frequency point).
[0080] Step 2052: Compress the frequency domain data to obtain compressed frequency domain data.
[0081] Compression of frequency domain data can be achieved in various ways. For example, exponential compression can be used, which involves calculating all the values included in the frequency domain data to a preset power (e.g., 0.3).
[0082] Step 2053: Encode the compressed frequency domain data using the audio coding network of the pre-trained neural network model to obtain audio feature data.
[0083] Audio coding networks can be implemented using various neural network structures, such as RNN, LSTM, Complex UNet, and Transformer architecture based on self-attention and cross-domain attention mechanisms.
[0084] This embodiment performs frequency domain conversion on the multi-channel mixed audio signal in the time domain, and then compresses the frequency domain data to obtain compressed frequency domain data. This reduces the numerical range of the frequency domain data, lowers the data processing difficulty of the neural network, and thus improves the efficiency of extracting the target user's speech signal.
[0085] In some alternative implementations, such as Figure 4 As shown, step 206 includes:
[0086] Step 2061: Based on the lip region image sequence and the preset parameters of the camera used to acquire the image sequence, determine the lip position information representing the spatial position of the target user's lips.
[0087] The camera's preset parameters can include pre-calibrated intrinsic parameters, extrinsic parameters, and pose information. Based on these parameters, the two-dimensional coordinates of the lip region in the captured image can be transformed to three-dimensional coordinates in the camera coordinate system or the world coordinate system, thus obtaining the lip position information. Methods for determining the lip region's position from a two-dimensional image and transforming its position to a three-dimensional coordinate system can be implemented using image-based coordinate transformation techniques, which will not be elaborated upon here.
[0088] Since the lip region image sequence consists of multiple frames acquired over a period of time, the lip position information can be identified from any one of the images. For example, it can be identified from the last frame of the lip region image sequence. Optionally, the lip position information can be identified from each of the multiple frames in the lip region image sequence, and then averaged to obtain the lip position information representing the spatial position of the target user's lips.
[0089] Step 2062: Based on the lip position information and the preset position information of the microphone array, determine the angle between the target line where the target user's lips are located and the baseline of the microphone array.
[0090] The preset position information can be obtained through pre-calibration. For example, the coordinates of the microphone array in the two-dimensional image captured by the camera can be determined in advance, and then the two-dimensional coordinates of the microphone array can be transformed into the camera coordinate system or the world coordinate system using the camera's intrinsic, extrinsic, and pose information, thereby obtaining the preset position information of the microphone array.
[0091] Since a microphone array includes at least two microphones, preset position information can represent a specific point within the range of the microphone array, and the baseline of the microphone array can be a straight line pre-specified based on the position of the microphone array. For example... Figure 5As shown, if the microphone array includes two microphones 501 and 502, the line 503 connecting the two microphones can be a reference line, and the midpoint 504 of the reference line can be used as a reference point. The coordinates of this reference point in the three-dimensional coordinate system are the preset position information.
[0092] The target straight line mentioned above can be the line connecting the aforementioned reference point and the lip position. For example... Figure 5 As shown, the lip position is 505, the target line 506 is the line connecting the lip position 505 and the reference point 504, and the angle α is the angle between the target line 506 and the reference line 503.
[0093] It should be noted that, Figure 5 This is merely an example and does not constitute a limitation on the reference point, reference line, or target line. The reference line and reference point mentioned above can be arbitrarily specified. For example, the reference point can be a point represented by 501 or 502, and the reference line can be a straight line perpendicular to line segment 503.
[0094] Step 2063: Based on the angle, determine the angular feature data between the target user's lip position and the microphone array.
[0095] Optionally, the aforementioned angles can be defined as angular feature data. Alternatively, based on current techniques for calculating steering vectors, the steering vector can be calculated from the angles and then defined as angular feature data.
[0096] Step 2064: Determine spatial location feature data based on angular feature data.
[0097] Optionally, the angular feature data can be determined as spatial feature data, or the spatial location feature data can be determined according to the method provided in the following embodiments.
[0098] This embodiment determines the angular feature data between the target user's lip position and the microphone array. The angular feature data can accurately represent the relative positional relationship between the target user's lips and the microphone array. Using the angular feature data as auxiliary information for extracting the target user's speech signal helps to continuously track the position of the lips in three-dimensional space, thereby improving the stability of extracting the target user's speech signal.
[0099] In some alternative implementations, after step 201, the following steps may also be performed:
[0100] Determine the phase difference characteristic data representing the multi-channel mixed audio signals.
[0101] Specifically, phase difference feature data represents the phase difference between audio signals from multiple channels acquired by a microphone array. Typically, if the microphone array includes two microphones, the phase difference between the audio signals of the two channels can be determined. If the microphone array includes more than two microphones, one channel can be used as a reference channel, and the phase difference between the audio signals of the other channels and the audio signal of the reference channel can be determined separately.
[0102] Since audio signals contain multiple frequency components, the audio signal of each channel can usually be converted to the frequency domain using methods such as Fast Fourier Transform. The phase difference of each frequency component can be determined by calculating the inter-channel phase difference (IPD). The set of phase differences of each frequency component is then defined as phase difference feature data.
[0103] Step 2063 above can be performed as follows:
[0104] Angle feature data is determined based on angle and phase difference feature data.
[0105] Optionally, based on the aforementioned angles, the steering vector can be determined using a method that calculates the steering vector. For each frequency component of the signal, angular characteristic data can be determined based on the steering vector and the phase difference. For example, the difference between the phase difference and the steering vector can be determined, and then the cosine or sine value of the difference can be taken as the angular characteristic data.
[0106] Optionally, angle feature data can also be determined using other methods based on the aforementioned angle and phase difference feature data. For example, the aforementioned angle and phase difference feature data can be merged to obtain angle feature data.
[0107] Since the phase difference between channels can represent the difference in the path taken by the target user's voice to each microphone, this embodiment combines the phase difference feature data with the angle of the lips relative to the microphone array to determine the angle feature data. The angle feature data can be used to accurately represent the difference in the sound transmission path between each microphone of the sound source, and use it as auxiliary information for extracting the target user's voice signal. This helps to utilize the transmission path of the target user's voice and improve the stability of extracting the target user's voice signal.
[0108] In some alternative implementations, step 2064 above can be performed as follows:
[0109] Spatial location feature data is determined based on angular feature data and phase difference feature data.
[0110] Specifically, angular feature data and phase difference feature data can be combined into spatial location feature data.
[0111] The spatial location feature data provided in this embodiment includes angle feature data and phase difference feature data, which enriches the content of the spatial location feature data. This allows the spatial location feature data to more fully represent the positional relationship between the lips and the microphone array and the transmission path of the target user's voice. This helps to include richer spatial location feature data in the fused feature data in the subsequent feature fusion step, thereby more accurately extracting the target user's voice signal from the multi-channel mixed audio signal.
[0112] In some alternative implementations, such as Figure 6 As shown, step 207 includes:
[0113] Step 2071: Using the fusion network of the pre-trained neural network model, the lip state feature data, audio feature data and spatial location feature data are fused to obtain fused feature data.
[0114] The fusion of lip state feature data, audio feature data, and spatial location feature data can be achieved through various methods, such as the concat feature fusion method, the elemwise_add feature fusion method, the single-gated feature fusion method, and the attention feature fusion method.
[0115] Step 2072: Use the decoding network of the neural network model to decode the fused feature data to obtain mask data.
[0116] Typically, decoding networks can have upsampling capabilities. The fused feature data is usually small-scale feature data. Through the decoding network, the small-scale fused feature data can be upsampled to obtain mask data with the same scale as the frequency domain data of each channel (e.g., obtained by performing a short-time Fourier transform on the audio signal of a channel).
[0117] Optionally, this neural network model can be combined with the above. Figure 3 The neural network model described in the corresponding embodiments is the same model, or it can be the same as... Figure 3 Another model that differs from the neural network model described in the corresponding embodiment.
[0118] The aforementioned neural network model can be trained using machine learning methods. The neural network model may include the audio encoding network, fusion network, and decoding network described above. Specifically, training samples can be obtained in advance. These training samples include the sample data to be encoded (i.e., the data processed by the audio encoding network, such as the one described above). Figure 3The data includes frequency domain data (or compressed frequency domain data) in the corresponding embodiment, sample lip state feature data, and sample spatial location feature data, as well as labeled mask data. The sample data to be encoded can be used as input to the audio encoding network, and the audio feature data output by the audio encoding network is fused with the sample lip state feature data and sample spatial location feature data to obtain fused feature data. This fused feature data is then input to the decoding network, and the labeled mask data corresponding to the input sample data to be encoded is used as the expected output of the decoding network to train the initial neural network model. For each training input sample data to be encoded, sample lip state feature data, and sample spatial location feature data, the actual output can be obtained. The actual output is the mask data actually output by the initial neural network model. Then, gradient descent and backpropagation methods can be used to adjust the parameters of the initial neural network model based on the difference between the actual output and the expected output, gradually reducing the difference. The model obtained after each parameter adjustment is used as the initial neural network model for the next training session. Training ends when a preset training termination condition is met (e.g., the loss value calculated based on a preset loss function converges, or the number of training iterations exceeds a preset number, etc.), thereby training the aforementioned neural network model.
[0119] Step 2073: Extract the target user's voice signal from the multi-channel mixed audio signal based on the mask data.
[0120] The mask data is used to filter the frequency domain data of a single-channel audio signal (e.g., obtained by performing a short-time Fourier transform on the single-channel audio signal) to obtain the frequency domain data of the target user's speech signal. Optionally, the mask data can be directly multiplied by the aforementioned frequency domain data of the single-channel audio signal to obtain the frequency domain data of the target user's speech signal. Then, processing such as inverse Fourier transform is performed on the frequency domain data of the target user's speech signal to obtain the time-domain speech signal.
[0121] Optionally, if the audio feature data is as described above Figure 3 According to the method described in the corresponding embodiment, the mask data can be multiplied with the compressed frequency domain data to obtain the compressed frequency domain data of the target user's speech signal. Then, the compressed frequency domain data of the target user's speech signal is decompressed to obtain the frequency domain data of the target user's speech signal. Finally, the frequency domain data of the target user's speech signal is processed such as inverse Fourier transform to obtain the time domain speech signal.
[0122] It should be noted that since multi-channel mixed audio signals include audio signals from multiple channels, and the mask data has the same scale as the frequency domain data of a single-channel audio signal, the mask data can be correlated with the frequency domain data of any channel's audio signal to extract the target user's speech signal from the single-channel audio signal. Optionally, the frequency domain data of the audio signals from each channel can be fused (e.g., by averaging the signals at each frequency point) into single-channel frequency domain data, and then the mask data can be correlated with the fused single-channel frequency domain data.
[0123] This embodiment effectively utilizes the characteristics of fused feature data, which can represent audio features, lip state features, and lip spatial position features. It uses a neural network model trained by machine learning methods to output mask data, and the mask data can be used to extract the target user's speech signal more accurately from the multi-channel mixed audio signal.
[0124] In some alternative implementations, such as Figure 7 As shown, step 2073 includes:
[0125] Step 20731: Compress the mask data using a preset activation function to obtain compressed data.
[0126] The purpose of compressing mask data is to reduce the numerical range of the mask data. As an example, the preset activation function can be the tanh activation function. By inputting each value in the mask data into the tanh activation function, a value between 0 and 1 can be obtained.
[0127] Step 20732: Extract the target user's speech signal from the multi-channel mixed audio signal based on the compressed data.
[0128] Specifically, compressed data can be multiplied by the frequency domain data of a single-channel audio signal to obtain the frequency domain data of the target user's speech signal. Then, processing such as inverse Fourier transform is performed on the frequency domain data of the target user's speech signal to obtain the time domain speech signal.
[0129] This embodiment compresses the mask data, which reduces the numerical range of the mask data, allowing the compressed mask data to be used as the proportion of the target user's voice signal to the original audio signal, thereby extracting the target user's voice signal more accurately.
[0130] In some alternative implementations, such as Figure 8 As shown, step 2071 includes:
[0131] Step 20711: The audio feature data and spatial location feature data are fused using the first fusion sub-network included in the fusion network to obtain fused audio feature data.
[0132] The first fusion process can be implemented in various ways. For example, the concat feature fusion method can be used to fuse audio feature data and spatial location feature data to obtain fused audio feature data.
[0133] Step 20712: Use the second fusion sub-network included in the fusion network to perform a second fusion process on the fused audio feature data and lip state feature data to obtain fused feature data.
[0134] The second fusion process can be implemented in various ways, and it can be implemented in the same or different ways as the first fusion process. For example, the methods used for the second fusion process may include, but are not limited to: the elemwise_add feature fusion method, the single-gated feature fusion method, the attention feature fusion method, and the dual-gated feature fusion method.
[0135] See Figure 9 , Figure 9 This is an exemplary schematic diagram illustrating the generation of fused feature data according to the speech signal extraction method of this embodiment. Figure 9 As shown, the merged feature data is the feature data obtained by merging the fused audio feature data and lip state feature data. 901-912 in the figure represent the functional modules included in the second fusion sub-network. The merged feature data passes through a first convolutional layer and a first activation function 901 (e.g., tanh activation function) to generate first feature data; the merged feature data passes through a second convolutional layer and a second activation function 902 (e.g., sigmoid activation function) to generate first weight data. The first feature data and the first weight data are multiplied element-wise 903 to generate second feature data. The merged feature data passes through a third convolutional layer and a third activation function 904 (e.g., sigmoid activation function) to generate second weight data; the audio feature data is then multiplied element-wise with the second weight data 905 to generate third feature data. The third feature data and the second feature data are fused using the elemwise_add method 906 to generate the first fused feature data.
[0136] The merged feature data is passed through the fourth convolutional layer and the fourth activation function (907, e.g., tanh activation function) to generate the fourth feature data; the merged feature data is then passed through the fifth convolutional layer and the fifth activation function (908, e.g., sigmoid activation function) to generate the third weight data. The fourth feature data and the third weight data are then multiplied element-wise (909) to generate the fifth feature data. The merged feature data is then passed through the sixth convolutional layer and the sixth activation function (910, e.g., sigmoid activation function) to generate the fourth weight data; the lip state feature data is then multiplied element-wise with the fourth weight data (911) to generate the sixth feature data. The sixth feature data and the fifth feature data are then fused using the `elemwise_add` method (912) to generate the second fused feature data.
[0137] The first and second fusion feature data are combined to generate fusion feature data.
[0138] Figure 9 The generated fused feature data shown is the dual-gated feature fusion method. This method uses fused audio feature data as the primary feature data and lip state feature data as the secondary feature data, and also uses lip state feature data as the primary feature data and fused audio feature data as the secondary feature data. It performs two feature fusions following similar steps and network structures to obtain the first and second fused feature data. The weight data generated during the process is used as gating parameters to perform calculations with the corresponding feature data. This allows for more targeted extraction of information representing the target user's speech from the audio feature data and lip state feature data, thereby achieving more accurate extraction of the target user's speech signal.
[0139] Exemplary device
[0140] Figure 10 This is a schematic diagram of the structure of a speech signal extraction device provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 10As shown, the speech signal extraction device includes: an acquisition module 1001 for acquiring multi-channel mixed audio signals and image sequences collected within a target area; a first determination module 1002 for determining a target user within the target area; a second determination module 1003 for determining a lip region image sequence of the target user based on the image sequence; a third determination module 1004 for determining lip state feature data based on the lip region image sequence; a fourth determination module 1005 for determining audio feature data based on the multi-channel mixed audio signal; a fifth determination module 1006 for determining the spatial position feature data of the target user's lips and microphone array based on the lip region image sequence; and an extraction module 1007 for extracting the target user's speech signal from the multi-channel mixed audio signal based on the lip state feature data, audio feature data, and spatial position feature data.
[0141] In this embodiment, the acquisition module 1001 acquires multi-channel mixed audio signals and image sequences collected within a target area. The target area can be defined as follows: Figure 1 The spatial area of the microphone array 104 and camera 105 shown can be, but is not limited to, the interior of a vehicle or room. The multi-channel mixed audio signal can be an audio signal acquired by the microphone array 104, which may include multiple channels, each corresponding to an audio signal acquired by one microphone. The multi-channel mixed audio signal includes at least one user's voice signal and noise signals, etc. The image sequence can be images captured by the camera 105 of the user within the target area. It should be understood that in this embodiment, the multi-channel mixed audio signal and the image sequence are acquired synchronously within the same duration (e.g., 1 second).
[0142] In this embodiment, the first determining module 1002 can determine the target user within the target area.
[0143] Optionally, the camera can capture images of a single user in a specific area (such as the driver's seat, passenger seat, etc. in a vehicle). If the first determining module 1002 identifies the user from the captured image sequence, then the user is determined to be the target user.
[0144] The camera can also capture images of multiple users, identifying multiple users from the captured image sequence. The first determining module 1002 then identifies one of these users as the target user for which the method is currently being executed. For example, the user located in the central region of a specified image can be identified as the target user from among the identified multiple users; alternatively, each user can be identified as the target user, and the method can be executed once for each target user; or, based on preset user feature data (e.g., facial feature data), a user matching the user feature data can be identified from the image sequence, and that user can be identified as the target user.
[0145] In this embodiment, the second determining module 1003 can determine the image sequence of the target user's lip region based on the image sequence.
[0146] Specifically, the images in the image sequence may include the lip region of the target user. The second determining module 1003 may extract lip region images from the images included in the image sequence based on a lip image detection method (e.g., determining lip region images based on a facial key point detection method) to obtain a lip region image sequence.
[0147] Typically, the size of the lip region images extracted from the image sequence can be adjusted to a fixed size (e.g., 96×96) to obtain a lip region image sequence of uniform size.
[0148] In this embodiment, the third determining module 1004 can determine lip state feature data based on the lip region image sequence. The lip state feature data is used to characterize the changes in mouth shape. Typically, the third determining module 1004 can identify the lip shape feature data (e.g., the distance between the corners of the mouth, the distance between the upper and lower lips, etc.) of each lip region image in the lip region image sequence, and merge the lip shape feature data of each lip region image into lip state feature data. Determining lip state feature data based on the lip region image sequence can be achieved using methods such as lip reading recognition, which will not be elaborated here.
[0149] In this embodiment, the fourth determining module 1005 can determine audio feature data based on multi-channel mixed audio signals.
[0150] Specifically, the aforementioned audio feature data can be determined for the audio signal of any channel in a multi-channel mixed audio signal; or the feature data of the audio signal of each channel can be determined separately, and then the feature data of each channel can be fused into the aforementioned audio feature data.
[0151] Optionally, the fourth determining module 1005 can determine the audio feature data of a channel based on a neural network method. For example, the neural network can include, but is not limited to, RNN, LSTM, UNet, Complex UNet, and Transformer architecture based on self-attention mechanism and cross-domain attention mechanism.
[0152] In this embodiment, the fifth determining module 1006 can determine the target user's lips based on the lip region image sequence. Figure 1 Spatial location feature data of the microphone array 104 shown.
[0153] The spatial location feature data is used to characterize the spatial relationship between the target user's lips and the microphone array 104. This spatial location feature data can be obtained based on the position of the microphone array and the position between the target user's lips. The position of the microphone array can be pre-calibrated, and the position of the target user's lips can be obtained by identifying a sequence of lip region images. For example, based on the intrinsic, extrinsic, and pose information of the camera 104, the two-dimensional coordinates of the lip region in the original image captured by the camera are transformed to three-dimensional coordinates in the camera coordinate system or the world coordinate system.
[0154] In this embodiment, the extraction module 1007 can extract the target user's speech signal from the multi-channel mixed audio signal based on lip state feature data, audio feature data, and spatial location feature data.
[0155] Specifically, lip state feature data, audio feature data, and spatial location feature data can be fused first to obtain fused feature data. Then, a method such as a neural network can be used to decode the fused feature data to obtain mask data. The mask data is multiplied with the frequency domain data corresponding to the audio signal of any channel in the multi-channel mixed audio signal (e.g., obtained by performing a short-time Fourier transform on the audio signal of one channel) to obtain the frequency domain data of the target user's speech signal. Then, the frequency domain data of the target user's speech signal is processed by methods such as inverse Fourier transform to obtain the time domain speech signal.
[0156] Reference Figure 11 , Figure 11 This is a schematic diagram of the structure of a speech signal extraction device provided in another exemplary embodiment of this disclosure.
[0157] In some optional implementations, the fifth determining module 1006 includes: a first determining unit 10061, used to determine lip position information representing the spatial position of the target user's lips based on the lip region image sequence and preset parameters of the camera used to acquire the image sequence; a second determining unit 10062, used to determine the angle between the target line where the target user's lips are located and the baseline of the microphone array based on the lip position information and preset position information of the microphone array; a third determining unit 10063, used to determine the angle feature data between the target user's lip position and the microphone array based on the angle; and a fourth determining unit 10064, used to determine the spatial position feature data based on the angle feature data.
[0158] In some alternative implementations, the device further includes: a sixth determining module 1008 for determining phase difference feature data representing the multi-channel mixed audio signals; and a third determining unit 10063 for further determining angle feature data based on the angle and phase difference feature data.
[0159] In some optional implementations, the fourth determining unit 10064 is further used to: determine spatial position feature data based on angle feature data and phase difference feature data.
[0160] In some optional implementations, the extraction module 1007 includes: a fusion unit 10071, used to fuse lip state feature data, audio feature data and spatial location feature data using a fusion network of a pre-trained neural network model to obtain fused feature data; a decoding unit 10072, used to decode the fused feature data using a decoding network of the neural network model to obtain mask data; and an extraction unit 10073, used to extract the speech signal of the target user from the multi-channel mixed audio signal based on the mask data.
[0161] In some optional implementations, the extraction unit 10073 includes: a compression subunit 100731, used to compress the mask data using a preset activation function to obtain compressed data; and an extraction subunit 100732, used to extract the target user's speech signal from the multi-channel mixed audio signal based on the compressed data.
[0162] In some optional implementations, the fusion unit 10071 includes: a first fusion subunit 100711, used to perform a first fusion process on audio feature data and spatial location feature data using the first fusion subnetwork included in the fusion network to obtain fused audio feature data; and a second fusion subunit 100712, used to perform a second fusion process on the fused audio feature data and lip state feature data using the second fusion subnetwork included in the fusion network to obtain fused feature data.
[0163] In some optional implementations, the fourth determining module 1005 includes: a conversion unit 10051, used to perform frequency domain conversion on the multi-channel mixed audio signal to obtain frequency domain data; a compression unit 10052, used to compress the frequency domain data to obtain compressed frequency domain data; and an encoding unit 10053, used to encode the compressed frequency domain data using an audio encoding network of a pre-trained neural network model to obtain audio feature data.
[0164] The speech signal extraction apparatus provided in the above embodiments of this disclosure acquires multi-channel mixed audio signals and image sequences collected within a target area. Then, based on the lip region image sequence, it determines lip state feature data, and based on the lip region image sequence, it determines the spatial position feature data of the target user's lips and the microphone array. Based on the multi-channel mixed audio signal, it determines audio feature data. Finally, based on the lip state feature data, audio feature data, and spatial position feature data, it extracts the target user's speech signal from the multi-channel mixed audio signal. This disclosure achieves multimodal speech separation by combining multi-channel mixed audio signals and spatial position feature data. It effectively utilizes the positional relationship between the spatial position of the lips and the positions of multiple microphones as auxiliary information for speech separation, enabling more targeted tracking of the target user's lip position, thereby improving the accuracy of speech signal extraction. In scenarios where lip occlusion occurs or the lip image quality is poor, the positional relationship between the lips and the microphone array can be effectively utilized to reduce the impact of image quality degradation, thereby improving the stability of speech signal extraction.
[0165] Exemplary electronic devices
[0166] Below, for reference Figure 12 To describe an electronic device according to embodiments of the present disclosure. The electronic device may be as follows: Figure 1 The terminal device 101 and server 103 shown, or either one or both, or a standalone device independent of them, can communicate with the terminal device 101 and server 103 to receive the collected input signals from them.
[0167] Figure 12 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0168] like Figure 12 As shown, the electronic device 1200 includes one or more processors 1201 and memory 1202.
[0169] The processor 1201 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1200 to perform desired functions.
[0170] The memory 1202 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1201 may execute the program instructions to implement the speech signal extraction methods of the various embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0171] In one example, the electronic device 1200 may also include an input device 1203 and an output device 1204, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0172] For example, when the electronic device is a terminal device 101 or a server 103, the input device 1203 can be a microphone, camera, mouse, keyboard, or other devices used to input multi-channel mixed audio signals, image sequences, various commands, etc. When the electronic device is a standalone device, the input device 1203 can be a communication network connector used to receive input multi-channel mixed audio signals, image sequences, various commands, etc. from the terminal device 101 and the server 103.
[0173] The output device 1204 can output various information to the outside, including the voice signal of the target user. The output device 1204 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0174] Of course, for the sake of simplicity, Figure 12 Only some of the components of the electronic device 1200 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1200 may include any other suitable components depending on the specific application.
[0175] Exemplary computer program products and computer-readable storage media
[0176] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the methods for extracting speech signals according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0177] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0178] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the speech signal extraction methods according to various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0179] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0180] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0181] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0182] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0183] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0184] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0185] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0186] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A method for extracting speech signals, comprising: Acquire multi-channel mixed audio signals and image sequences collected within the target area; Identify the target users within the target area; Based on the image sequence, determine the lip region image sequence of the target user; Based on the image sequence of the lip region, determine the lip state feature data; Based on the multi-channel mixed audio signal, audio feature data is determined; Based on the lip region image sequence, determine the spatial location feature data of the target user's lips and microphone array; The lip state feature data, the audio feature data, and the spatial location feature data are fused to obtain fused feature data. Based on the fused feature data, the speech signal of the target user is extracted from the multi-channel mixed audio signal. The step of determining the spatial location feature data of the target user's lips and microphone array based on the lip region image sequence includes: Based on the lip region image sequence and preset parameters of the camera used to acquire the image sequence, lip position information representing the spatial position of the target user's lips is determined; Based on the lip position information and the preset position information of the microphone array, the angle between the target line where the target user's lips are located and the baseline of the microphone array is determined; Based on the angle, determine the angular feature data between the target user's lip position and the microphone array; Based on the angular feature data, the spatial location feature data is determined.
2. The method according to claim 1, wherein, Also includes: Determine the phase difference feature data representing the multi-channel mixed audio signals; The step of determining the angular feature data between the target user's lip position and the microphone array based on the angle includes: The angle feature data is determined based on the angle and the phase difference feature data.
3. The method according to claim 2, wherein, The step of determining the spatial location feature data based on the angular feature data includes: Based on the angular feature data and the phase difference feature data, the spatial position feature data is determined.
4. The method according to claim 1, wherein, The process of fusing the lip state feature data, the audio feature data, and the spatial location feature data to obtain fused feature data, and extracting the target user's speech signal from the multi-channel mixed audio signal based on the fused feature data, includes: The fusion network of a pre-trained neural network model is used to fuse the lip state feature data, the audio feature data, and the spatial location feature data to obtain fused feature data. The fused feature data is decoded using the decoding network of the neural network model to obtain mask data; Based on the mask data, the voice signal of the target user is extracted from the multi-channel mixed audio signal.
5. The method according to claim 4, wherein, Extracting the target user's voice signal from the multi-channel mixed audio signal based on the mask data includes: The mask data is compressed using a preset activation function to obtain compressed data. Based on the compressed data, the voice signal of the target user is extracted from the multi-channel mixed audio signal.
6. The method according to claim 4, wherein, The fusion network, utilizing a pre-trained neural network model, fuses the lip state feature data, the audio feature data, and the spatial location feature data to obtain fused feature data, including: The audio feature data and the spatial location feature data are subjected to a first fusion process using the first fusion sub-network included in the fusion network to obtain fused audio feature data; The fusion feature data is obtained by using the second fusion sub-network included in the fusion network to perform a second fusion process on the fused audio feature data and the lip state feature data.
7. The method according to claim 1, wherein, The process of determining audio feature data based on the multi-channel mixed audio signal includes: The multi-channel mixed audio signal is frequency domain converted to obtain frequency domain data; The frequency domain data is compressed to obtain compressed frequency domain data; The compressed frequency domain data is encoded using a pre-trained neural network model to obtain the audio feature data.
8. A speech signal extraction device, comprising: The acquisition module is used to acquire multi-channel mixed audio signals and image sequences collected within the target area; The first determining module is used to determine the target user within the target area; The second determining module is used to determine the lip region image sequence of the target user based on the image sequence; The third determining module is used to determine lip state feature data based on the lip region image sequence; The fourth determining module is used to determine audio feature data based on the multi-channel mixed audio signal; The fifth determining module is used to determine the spatial position feature data of the target user's lips and microphone array based on the lip region image sequence; The extraction module is used to fuse the lip state feature data, the audio feature data and the spatial location feature data to obtain fused feature data, and extract the speech signal of the target user from the multi-channel mixed audio signal based on the fused feature data; The fifth determining module includes: The first determining unit is used to determine lip position information representing the spatial position of the target user's lips based on the lip region image sequence and preset parameters of the camera used to acquire the image sequence; The second determining unit is used to determine the angle between the target line where the target user's lips are located and the baseline of the microphone array, based on the lip position information and the preset position information of the microphone array. The third determining unit is used to determine the angular feature data between the target user's lip position and the microphone array based on the angle; The fourth determining unit is used to determine the spatial location feature data based on the angle feature data.
9. A computer-readable storage medium storing a computer program for execution by a processor to implement the method of any one of claims 1-7.
10. An electronic device, the electronic device comprising: processor; Memory for storing the executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-7.