Voice data processing method, device, intelligent device and computer storage medium
By combining facial images and voice spectrum data to generate a spectrum mask, the problem of voice recognition accuracy of intelligent voice devices in noisy environments is solved, and normal interaction in noisy environments is achieved.
Patent Information
- Application Number
- CN202110232082.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-02
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-03-02
AI Technical Summary
In noisy environments, smart voice devices find it difficult to accurately recognize user voices, resulting in a poor interaction experience between users and the device.
By acquiring facial image data and voice spectrum data, and combining facial features and voiceprint features, a spectrum mask is generated to filter out noise and enhance the target user's voice.
Accurately identify the target user's voice in noisy environments and improve the user's interactive experience with smart voice devices.
Smart Images

Figure CN114999499B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method, apparatus, intelligent device, and computer storage medium for processing voice data. Background Art
[0002] With the development of AI (Artificial Intelligence) technology, more and more intelligent voice devices based on AI voice interaction are being widely used in people's work and life.
[0003] Existing intelligent voice devices use microphone arrays to pick up user voices, recognize the picked-up speech, and interact with users based on the recognition results. However, in noisy environments such as subway stations, showrooms, and homes streaming TV, user voices are susceptible to significant interference. Intelligent voice devices process and recognize all the picked-up speech around the device, resulting in reduced speech recognition accuracy.
[0004] As a result, when using smart voice devices in these scenarios, users are often unable to interact normally with the smart voice devices, resulting in a poor user experience. Summary of the Invention
[0005] In view of this, an embodiment of the present application provides a voice data processing solution to at least partially solve the above-mentioned problem.
[0006] According to a first aspect of an embodiment of the present application, a speech data processing method is provided, comprising: obtaining facial image data and speech spectrum data containing multiple faces; processing the facial image data and the speech spectrum data to determine a target face; obtaining facial features and voiceprint features corresponding to the target face, and determining a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, the voiceprint features and the speech spectrum data; and performing speech enhancement processing on the speech spectrum data according to the spectrum mask.
[0007] According to a second aspect of an embodiment of the present application, a speech data processing device is provided, comprising: a data acquisition module for acquiring facial image data and speech spectrum data containing multiple faces; a processing and determination module for processing the facial image data and the speech spectrum data to determine a target face; a spectrum mask acquisition module for acquiring facial features and voiceprint features corresponding to the target face, and determining a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, the voiceprint features and the speech spectrum data; and a speech enhancement module for performing speech enhancement processing on the speech spectrum data according to the spectrum mask.
[0008] According to a third aspect of an embodiment of the present application, an electric intelligent device is provided, comprising: a voice acquisition device, an image acquisition device, and a processor; wherein the voice acquisition device is used to acquire voice data; the image acquisition device is used to acquire facial images; the processor is used to receive facial image data containing multiple faces acquired by the image acquisition device and voice data acquired by the voice acquisition device and convert them into voice spectrum data; and, based on the facial image data and the voice spectrum data, perform operations corresponding to the voice data processing method described in the first aspect.
[0009] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the voice data processing method as described in the first aspect is implemented.
[0010] According to the voice data processing solution provided by the embodiment of the present application, when using an intelligent voice device in a crowded and noisy environment, the voice and image are combined. First, the target user who issues voice commands to the intelligent voice device, that is, the user corresponding to the target face, is determined based on the data after the fusion of the facial image data and the voice spectrum data; then, a spectrum mask is obtained based on the facial features, voiceprint features and voice spectrum data corresponding to the target face; and then voice enhancement is performed through the spectrum mask. Because the facial image data will not be affected even in a noisy environment, the target user can still be determined more accurately. On this basis, a spectrum mask of noise data that can be used to indicate voices other than the target user is determined. The spectrum mask is used to filter out the voices of non-target users as much as possible, thereby achieving the effect of enhancing the voice of the target user. As a result, even when using an intelligent voice device in a noisy environment, the user can interact normally with the intelligent voice device, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0012] Figure 1A This is a flowchart of a method for processing voice data according to the first embodiment of the present application;
[0013] Figure 1B for Figure 1A A schematic diagram of an example scenario in the illustrated embodiment;
[0014] Figure 2This is a flowchart of a method for processing voice data according to the second embodiment of the present application;
[0015] Figure 3 This is a structural block diagram of a voice data processing device according to the third embodiment of the present application;
[0016] Figure 4 This is a structural diagram of an intelligent voice device according to the fourth embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0018] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0019] Example 1
[0020] Reference Figure 1A , shows a step flow chart of a voice data processing method according to embodiment 1 of the present application.
[0021] The voice data processing method of this embodiment includes the following steps:
[0022] Step S102: Acquire facial image data and speech spectrum data containing multiple faces.
[0023] In this step, facial image data and voice spectrum data containing multiple faces can be obtained through appropriate equipment such as equipment with image acquisition and sound acquisition functions, such as multimodal intelligent voice equipment.
[0024] Taking multimodal intelligent voice devices as an example, a multimodal intelligent voice device is a type of intelligent voice device that, in addition to the microphone of a conventional intelligent voice device, also has an image acquisition device such as a camera. This allows it to collect multimodal information.
[0025] In this embodiment, the corresponding device, such as a multimodal intelligent voice device, collects at least facial images and voice. The collection method can be implemented as real-time collection, or collection at preset time intervals, or collection based on trigger conditions (such as after being awakened by a wake-up word), etc. The embodiment of this application does not limit this.
[0026] Taking the multimodal intelligent voice device as an example, specifically in this embodiment, in a noisy environment, the images captured by the multimodal intelligent voice device typically contain multiple faces, that is, facial image data containing multiple faces is captured. Furthermore, in this embodiment, the multimodal intelligent voice device can also convert the collected voice data into voice spectrum data. In a specific implementation, those skilled in the art can use any appropriate method to convert the collected raw voice data into voice spectrum data, including but not limited to Fourier transform. Voice spectrum data is a digital representation of the voice spectrum, which characterizes the frequency distribution of speech.
[0027] Step S104: Process the facial image data and the speech spectrum data to determine the target face.
[0028] The processing may be any suitable processing that can determine the target face based on the facial image data and the speech spectrum data, for example, matching the facial image data and the speech spectrum data, or performing matching processing after feature extraction, etc.
[0029] In one feasible way, the processing can be implemented as multimodal fusion. Multimodal fusion can integrate data from multiple (in the embodiments of this application, unless otherwise specified, "multiple", "multiple", etc., and quantities related to "multiple" all mean two or more) modalities to perform more accurate subsequent processing. Because the data of a single modality usually cannot contain all the valid data to produce the required results, multimodal fusion can combine data from multiple modalities to achieve information supplementation and broaden the coverage of the information contained in the data.
[0030] Through multimodal fusion, facial image data and voice spectrum data can be effectively matched and supplemented. Based on this, the target face can be determined, that is, the facial image of the target user who is currently sending voice commands to the multimodal intelligent voice device.
[0031] Step S106: Acquire facial features and voiceprint features corresponding to the target face, and determine a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, voiceprint features and speech spectrum data.
[0032] In one feasible approach, facial features and voiceprint features can be pre-stored in a corresponding device, such as a multimodal intelligent voice device. After a target face is identified, the facial features and voiceprint features corresponding to the target face can be obtained based on one or more pre-defined mapping relationships. By pre-storing these features locally on a corresponding device, such as a multimodal intelligent voice device, facial features and voiceprint features can be quickly obtained, speeding up the overall execution of the solution.
[0033] In another feasible approach, facial features and voiceprint features, as well as one or more mapping relationships between the target face, facial features, and voiceprint features, can be pre-stored on the server side (server or cloud). In this way, when the multimodal intelligent voice device determines the target face, it can obtain the corresponding facial features and voiceprint features from the server side. This approach can greatly reduce the data storage burden on the corresponding device, such as the multimodal intelligent voice device.
[0034] Once facial and voiceprint features are obtained, combined with the previously acquired speech spectrum data, a spectrum mask can be obtained to identify noise data within the speech spectrum data. In noisy environments, speech spectrum data often contains data for non-target users (i.e., users not corresponding to the target face). However, the acquired facial and voiceprint features contain information specific to the target user. Based on this, a spectrum mask can be generated through appropriate methods, such as neural network models, to identify noise data within the speech spectrum data, i.e., speech data that does not belong to the target user.
[0035] Step S108: performing speech enhancement processing on the speech spectrum data according to the spectrum mask.
[0036] Since the spectrum mask can indicate and mark the noise data in the speech spectrum data, the speech spectrum data can be processed accordingly according to the spectrum mask to filter the noise data and achieve enhanced processing of the target user's speech data.
[0037] The following is an example of a specific scenario in which a multimodal intelligent voice device performs the above-mentioned voice data processing to illustrate the above process. Figure 1B shown.
[0038] Figure 1BIn the scenario shown, the multimodal intelligent voice device is specifically implemented as a smart speaker, and the specific scene is set as a living room where a party is being held. Suppose that user A issues a command to the smart speaker to "turn on the air conditioner", but at the same time, user B behind him is behind user A and leans forward to see user A's operation and makes a "let me see" sound, and there are other noisy sounds in the living room. The sounds made by user B and other noisy sounds are noise. In this case, when the traditional method is adopted, because the voice of user A is mixed with all the sounds, it is very likely that the smart speaker will not be able to accurately identify it, and thus the corresponding operation cannot be completed. However, using the solution of the embodiment of the present application, the smart speaker will simultaneously collect the voice of user A containing noise, as well as images containing user A and user B. Furthermore, the smart speaker converts the collected sound into voice spectrum data, and fuses it with the facial image data containing user A and user B to obtain the fused data. Based on this, in a specific implementation method, the facial image data can be detected to determine that the mouth of user A is moving, and the time matches the time corresponding to the sound "turn on the air conditioner". Based on this, the face of user A in the facial image data can be determined as the target face.
[0039] Then, based on the facial features and voiceprint features pre-stored in the database and their corresponding relationship with the target face, the facial features and voiceprint features corresponding to the target face of user A are obtained. In this example, user A's facial features, voiceprint features, and the aforementioned voice spectrum data are input into the neural network model used to generate the spectral mask, resulting in an output spectral mask. This spectral mask is then multiplied by the aforementioned voice spectrum data, and then voice conversion processing is performed to obtain enhanced voice data of user A, such as "Turn on the air conditioner." The smart speaker can then respond to user A's command "Turn on the air conditioner" and execute the operation to turn on the air conditioner.
[0040] Through this embodiment, when using an intelligent voice device in a crowded and noisy environment, voice and image are combined. First, the target user who issues voice commands to the intelligent voice device, that is, the user corresponding to the target face, is determined based on the data obtained by fusing facial image data and voice spectrum data. Then, a spectrum mask is obtained based on the facial features, voiceprint features, and voice spectrum data corresponding to the target face. Then, voice enhancement is performed using the spectrum mask. Because the facial image data will not be affected even in a noisy environment, the target user can still be determined relatively accurately. On this basis, a spectrum mask is determined for noise data that can be used to indicate voices other than the target user. The spectrum mask is used to filter out the voices of non-target users as much as possible, thereby achieving the effect of enhancing the target user's voice. As a result, even when using an intelligent voice device in a noisy environment, the user can still interact with the intelligent voice device normally, improving the user experience.
[0041] Example 2
[0042] Reference Figure 2 , shows a step flow chart of a voice data processing method according to embodiment 2 of the present application.
[0043] In this embodiment, the solution provided by the embodiment of the present application is described by taking a multimodal intelligent voice device for voice data processing as an example. However, it should be clear to those skilled in the art that other intelligent devices capable of collecting multimodal data (including at least image data and voice data) are also applicable. The voice data processing method of this embodiment includes the following steps:
[0044] Step S202: Collect facial images and voice data through a multimodal intelligent voice device.
[0045] In this embodiment, the collected facial image is a facial image containing multiple faces.
[0046] Step S204: Acquire facial image data corresponding to the facial image and speech spectrum data corresponding to the speech data.
[0047] The speech spectrum data can be obtained by converting the speech data through methods such as Fourier transform.
[0048] Step S206: performing multimodal fusion on the facial image data and the speech spectrum data, and determining a target face from the multiple faces based on the multimodal fusion result.
[0049] In one feasible manner, this step can be implemented as follows: performing face detection on multiple facial image data within a preset time period, and intercepting the facial image portion corresponding to each of the multiple faces from the multiple facial image data based on the face detection results; sorting the intercepted facial image portions of each face according to the image acquisition timing to generate multiple facial image sequences; matching the multiple facial image sequences with the voice spectrum data respectively, and determining the target face from the multiple faces based on the matching results. In this way, the voice spectrum data can be accurately matched to the corresponding face, thereby achieving accurate detection of the target face. The preset time period can be appropriately set by those skilled in the art according to actual needs, and the embodiments of the present application are not limited to this.
[0050] In a specific implementation, the method of matching the multiple facial image sequences with the speech spectrum data and determining the target face from the multiple faces based on the matching results can be performed based on the time information corresponding to each facial image sequence and the time information corresponding to the speech spectrum data, and determining the target face from the multiple faces based on the matching results. For example, based on the time information corresponding to each facial image sequence and the time information corresponding to the speech spectrum data, a facial image sequence whose duration of face appearance coincides with the duration of the speech spectrum data can be matched, and the target face can be determined from the multiple faces based on the matching results. Alternatively, feature extraction can be performed on the facial image portions of the multiple faces included in the multiple facial image sequences to obtain the corresponding multiple facial features; feature extraction can also be performed on the speech spectrum data to obtain the corresponding voiceprint features; based on the pre-stored correspondence between facial features and voiceprint features, the facial features among the multiple facial features that have a corresponding relationship with the voiceprint features can be determined as the target facial features; and the face corresponding to the target facial features can be determined as the target face. The corresponding relationship may be any appropriate form of corresponding relationship, including but not limited to a corresponding relationship represented by an ID, or a corresponding relationship represented by a serial number, etc., and the present embodiment does not limit this. The matching method based on time information is efficient and has a high matching degree; while the matching method based on feature matching is more accurate.
[0051] Optionally, when the pre-stored correspondence between facial features and voiceprint features is a correspondence between a facial identifier of a facial feature and a voiceprint identifier of a voiceprint feature, when determining, based on the pre-stored correspondence between facial features and voiceprint features, a facial feature that has a correspondence with the voiceprint feature among multiple facial features as the target facial feature, the method may include determining, from the multiple facial features, a facial feature that has a corresponding facial identifier; and determining the voiceprint identifier corresponding to the voiceprint feature; and, based on the correspondence, determining as the target facial feature the facial feature corresponding to the facial identifier that has a correspondence with the voiceprint identifier. This can reduce the amount of data matching and improve matching efficiency.
[0052] For example, N faces (N is an integer greater than or equal to 2) can be detected from facial image data corresponding to multiple facial images collected continuously; the detected N faces can be cut out from all images; according to the image collection sequence of the multiple facial images, the facial image parts corresponding to the same face can be arranged to form N facial image sequences, where one facial image sequence corresponds to the same face; each facial image sequence can be matched and compared with the speech spectrum data, and the target face can be determined based on the comparison results.
[0053] When performing a matching comparison, one approach is to determine if a facial image sequence appears continuously within the time period corresponding to the speech spectrum data. That is, if the duration of the face appearance is consistent with the duration of the speech, the face corresponding to the facial image sequence can be determined as the target face. If such a facial image sequence includes multiple (two or more) facial image sequences, it is possible to detect whether the faces in the facial image sequence have any mouth movements within the time period corresponding to the speech spectrum data. If so, the face corresponding to the facial image sequence is determined as the target face.
[0054] In another approach, facial feature extraction is performed on each image in the facial image sequence to obtain facial features corresponding to each face, such as N facial features. Simultaneously, feature extraction can also be performed on the speech spectrum data to obtain corresponding voiceprint features. Then, a determination is made as to which of the N facial features corresponds to the voiceprint features, and the face corresponding to the corresponding facial feature is identified as the target face.
[0055] Through the above process, the target user can be accurately identified through the target face.
[0056] Step S208: Acquire facial features and voiceprint features corresponding to the target face, and determine a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, voiceprint features and speech spectrum data.
[0057] In one feasible approach, obtaining the facial features and voiceprint features corresponding to the target face can be achieved by: performing facial recognition on the target face, obtaining a facial identifier for the target face based on the facial recognition result; determining the voiceprint identifier corresponding to the facial identifier, and obtaining the voiceprint features corresponding to the voiceprint identifier. In this case, the system pre-establishes a database for storing facial features and a database for storing voiceprint features, wherein facial features correspond to facial identifiers such as face IDs, and voiceprint features correspond to voiceprint identifiers such as voiceprint IDs. The system also stores the correspondence between facial identifiers and voiceprint identifiers, using which the corresponding facial features and voiceprint features can be determined. This approach improves the efficiency of feature data management and the matching efficiency of corresponding features.
[0058] For example, after obtaining the corresponding facial features through face recognition, you can search the database that stores the facial features to determine the corresponding face ID; then, based on the correspondence between the face ID and the voiceprint ID, find the voiceprint ID corresponding to the face ID; and then, through the database that stores the voiceprint features, find the voiceprint features corresponding to the voiceprint ID.
[0059] However, the present invention is not limited thereto. In practical applications, other methods of obtaining facial features and voiceprint features corresponding to a determined target face may also be applicable to the embodiments of the present application.
[0060] When determining a spectral mask to indicate noise in speech spectrum data based on facial features, voiceprint features, and speech spectrum data, one feasible approach involves performing feature fusion on facial and voiceprint features to obtain a fused voiceprint-face feature. Furthermore, feature extraction is performed on the speech spectrum data to obtain spectral features. Using the fused voiceprint-face feature and spectral features as input, a pre-trained neural network model is used to generate a spectral mask probability map. Each probability value in the spectral mask probability map indicates the probability that the data at the corresponding position in the speech spectrum data is noise. This spectral masking approach enables precise noise suppression and voice enhancement for the target user. Furthermore, after obtaining the fused voiceprint-face feature, further feature extraction can be performed on the facial and voiceprint features separately, and based on the results of this further feature extraction, feature fusion is performed to obtain the fused voiceprint-face feature. This secondary feature extraction approach not only achieves feature enhancement but also reduces the amount of subsequent data computation.
[0061] Among them, feature extraction of corresponding data can be achieved by using an appropriate neural network model.
[0062] For example, facial features are input into the first LSTM model to produce the first output feature vector; voiceprint features are input into the second LSTM model to produce the second output feature vector; and speech spectrum data is input into the third LSTM model to produce the third output feature vector, the spectrum feature vector. The first and second feature vectors are fused using a concat layer to form a fused feature vector, the voiceprint-face fusion feature. This fused voiceprint-face feature and the spectrum feature vector are then input into a backbone network, such as a CNN, to produce a masked feature vector. This masked feature vector is then converted into a probability map through a softmax layer and output. The value of a specific pixel in this probability map represents the probability that the corresponding position in the speech spectrum data is noise data.
[0063] The backbone network takes facial features, voiceprint features, and speech spectrum features as input, and outputs a mask probability map. During training, clean speech spectrum data can be used as the ground truth. This speech spectrum data is then subjected to noise enhancement (for example, including superimposed Gaussian noise, daily environmental noise, and other subtle noise). The enhanced speech spectrum data is then Fourier transformed and feature extracted to convert it into a speech spectrum feature containing noise. The backbone network automatically learns and generates the spectrum mask probability map.
[0064] However, those skilled in the art should understand that the description of the above-mentioned neural network model is only an example description, and the embodiments of the present application do not limit the specific implementation form, structure and training process of the neural network model that outputs the mask probability map based on facial features, voiceprint features, and speech spectrum features.
[0065] Step S210: performing speech enhancement processing on the speech spectrum data according to the spectrum mask.
[0066] For example, matrix multiplication operation is performed on the spectrum mask and the speech spectrum data, and enhanced speech spectrum data is obtained according to the operation result; and inverse Fourier transform is performed on the enhanced speech spectrum data to obtain corresponding enhanced speech data.
[0067] The spectrum mask and speech spectrum data are usually represented in matrix form. Based on this, matrix multiplication operations can be performed on the two. Then, the speech spectrum data with noise at the corresponding position indicated by the spectrum mask will be removed or suppressed, thereby enhancing the non-noise data.
[0068] Because facial features are obtained through images, noise signals will not affect facial features. Therefore, in noisy environments, the introduction of facial features can help improve voiceprint features, thereby improving the noise suppression effect of the entire neural network model.
[0069] Through this embodiment, when using an intelligent voice device in a crowded and noisy environment, voice and image are combined. First, the target user who issues voice commands to the intelligent voice device, that is, the user corresponding to the target face, is determined based on the data obtained by fusing facial image data and voice spectrum data. Then, a spectrum mask is obtained based on the facial features, voiceprint features, and voice spectrum data corresponding to the target face. Then, voice enhancement is performed using the spectrum mask. Because the facial image data will not be affected even in a noisy environment, the target user can still be determined relatively accurately. On this basis, a spectrum mask is determined for noise data that can be used to indicate voices other than the target user. The spectrum mask is used to filter out the voices of non-target users as much as possible, thereby achieving the effect of enhancing the target user's voice. As a result, even when using an intelligent voice device in a noisy environment, the user can still interact with the intelligent voice device normally, improving the user experience.
[0070] Example 3
[0071] Reference Figure 3 , shows a structural block diagram of a voice data processing device according to embodiment three of the present application.
[0072] The speech data processing device of this embodiment includes: a data acquisition module 302 , a multimodal fusion module 304 , a spectrum mask acquisition module 306 , and a speech enhancement module 308 .
[0073] Among them, the data acquisition module 302 is used to obtain facial image data and speech spectrum data containing multiple faces; the processing and determination module 304 is used to process the facial image data and the speech spectrum data to determine the target face; the spectrum mask acquisition module 306 is used to obtain the facial features and voiceprint features corresponding to the target face, and based on the facial features, the voiceprint features and the speech spectrum data, determine the spectrum mask used to indicate the noise data in the speech spectrum data; the speech enhancement module 308 is used to perform speech enhancement processing on the speech spectrum data according to the spectrum mask.
[0074] Optionally, in any embodiment of the present application, the processing and determination module 304 is used to perform face detection on the multiple facial image data within a preset time period, and based on the face detection results, intercept the facial image portion corresponding to each of the multiple faces from the multiple facial image data; sort the intercepted facial image portions of each face according to the image acquisition timing to generate multiple facial image sequences; match the multiple facial image sequences with the speech spectrum data respectively, and determine the target face from the multiple faces based on the matching results.
[0075] Optionally, in any embodiment of the present application, when the processing and determination module 304 matches the multiple facial image sequences with the speech spectrum data respectively and determines the target face from the multiple faces based on the matching results: according to the time information corresponding to each facial image sequence and the time information corresponding to the speech spectrum data, match a facial image sequence whose continuous appearance time of the face is consistent with the duration of the speech spectrum data, and determine the target face from the multiple faces based on the matching results.
[0076] Optionally, in any embodiment of the present application, when the processing and determination module 304 matches the multiple facial image sequences with the voice spectrum data respectively and determines the target face from the multiple faces based on the matching results: feature extraction is performed on the facial image parts of the multiple faces included in the multiple facial image sequences to obtain the corresponding multiple facial features; and feature extraction is performed on the voice spectrum data to obtain the corresponding voiceprint features; based on the pre-stored correspondence between facial features and voiceprint features, the facial features among the multiple facial features that have a corresponding relationship with the voiceprint features are determined as target facial features; and the face corresponding to the target facial features is determined as the target face.
[0077] Optionally, in any embodiment of the present application, the correspondence between the pre-stored facial features and the voiceprint features is the correspondence between the face identifier of the pre-stored facial features and the voiceprint identifier of the voiceprint features; when the processing determination module 304 determines the facial feature that has a correspondence with the voiceprint feature among the multiple facial features as the target facial feature based on the correspondence between the pre-stored facial features and the voiceprint features: determines the facial feature that has a corresponding face identifier from the multiple facial features; and determines the voiceprint identifier corresponding to the voiceprint feature; based on the correspondence, determines the facial feature corresponding to the face identifier that has a correspondence with the voiceprint identifier as the target facial feature.
[0078] Optionally, in any embodiment of the present application, the spectrum mask acquisition module 306 is used to perform feature fusion on the facial features and the voiceprint features to obtain voiceprint-face fusion features; and perform feature extraction on the speech spectrum data to obtain spectrum features; and use the voiceprint-face fusion features and the spectrum features as input to obtain a spectrum mask probability map using a pre-trained neural network model, wherein each probability value in the spectrum mask probability map is used to indicate the probability that the data at the corresponding position in the speech spectrum data is noise data.
[0079] Optionally, in any embodiment of the present application, the speech enhancement module 308 is used to perform matrix multiplication operation on the spectrum mask and the speech spectrum data, and obtain enhanced speech spectrum data according to the operation result; and perform inverse Fourier transform on the enhanced speech spectrum data to obtain corresponding enhanced speech data.
[0080] Optionally, in any embodiment of the present application, when acquiring the facial features and voiceprint features corresponding to the target face, the spectrum mask acquisition module 306: performs face recognition on the target face, and obtains a face identifier of the target face based on the face recognition result; determines a voiceprint identifier corresponding to the face identifier, and acquires the voiceprint features corresponding to the voiceprint identifier.
[0081] Optionally, the data acquisition module 302 is configured to acquire facial image data and speech spectrum data containing multiple faces through a multimodal intelligent voice device.
[0082] The speech data processing device of this embodiment is used to implement the corresponding speech data processing methods of the aforementioned multiple method embodiments and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here. In addition, the functional implementation of each module in the speech data processing device of this embodiment can refer to the description of the corresponding parts in the aforementioned method embodiments, which will not be described in detail here.
[0083] Example 4
[0084] Reference Figure 4, shows a structural diagram of an intelligent voice device according to the fourth embodiment of the present application. The specific embodiment of the present application does not limit the specific implementation of the intelligent voice device.
[0085] like Figure 4 As shown, the intelligent voice device may include: a processor 402 , a voice acquisition device 404 (such as a microphone), an image acquisition device 406 (such as a camera), a memory 408 , and a communication bus 410 .
[0086] in:
[0087] The processor 402 , the voice acquisition device 404 , the image acquisition device 406 , and the memory 408 communicate with each other via a communication bus 410 .
[0088] Optionally, a communication interface 412 may be further included for communicating with other electronic devices or servers.
[0089] The voice collecting device 404 is used to collect voice data.
[0090] The image acquisition device 406 is used to acquire facial images.
[0091] The processor 402 is configured to execute the program 414 , and specifically may execute the relevant steps in the above-mentioned embodiment of the voice data processing method.
[0092] Specifically, the program 414 may include program codes, which include computer operating instructions.
[0093] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0094] The memory 408 is used to store the program 414. The memory 408 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0095] Program 414 can specifically be used to enable the processor 402 to perform the following operations: receive facial image data containing multiple faces collected by the image acquisition device and voice data collected by the voice acquisition device and convert them into voice spectrum data; and, based on the facial image data and the voice spectrum data, process the facial image data and the voice spectrum data to determine the target face; obtain facial features and voiceprint features corresponding to the target face, and based on the facial features, the voiceprint features and the voice spectrum data, determine a spectrum mask for indicating noise data in the voice spectrum data; and perform voice enhancement processing on the voice spectrum data according to the spectrum mask.
[0096] In an optional embodiment, the program 414 is also used to enable the processor 402 to perform the following operations when processing the facial image data and the speech spectrum data to determine the target face: perform face detection on the multiple facial image data within a preset time period, and based on the face detection results, intercept the facial image parts corresponding to each of the multiple faces from the multiple facial image data; sort the intercepted facial image parts of each face according to the image acquisition timing to generate multiple facial image sequences; match the multiple facial image sequences with the speech spectrum data respectively, and determine the target face from the multiple faces based on the matching results.
[0097] In an optional embodiment, the program 414 is also used to enable the processor 402 to perform the following operations when matching the multiple facial image sequences with the speech spectrum data respectively and determining the target face from the multiple faces based on the matching results: matching a facial image sequence whose continuous appearance time is consistent with the duration of the speech spectrum data based on the time information corresponding to each facial image sequence and the time information corresponding to the speech spectrum data, and determining the target face from the multiple faces based on the matching results.
[0098] In an optional embodiment, the program 414 is further configured to cause the processor 402 to perform the following operations when matching the multiple facial image sequences with the speech spectrum data respectively and determining a target face from the multiple faces based on the matching results: performing feature extraction on the facial image portions of the multiple faces included in the multiple facial image sequences to obtain a corresponding multiple facial features; and performing feature extraction on the speech spectrum data to obtain corresponding voiceprint features; determining, based on a pre-stored correspondence between facial features and voiceprint features, a facial feature among the multiple facial features that has a corresponding relationship with the voiceprint feature as a target facial feature; and determining a face corresponding to the target facial feature as a target face.
[0099] In an optional embodiment, the correspondence between the pre-stored facial features and the voiceprint features is a correspondence between the face identifier of the pre-stored facial features and the voiceprint identifier of the voiceprint features; the program 414 is also used to enable the processor 402 to perform the following operations when determining, based on the pre-stored correspondence between the facial features and the voiceprint features, a facial feature that has a correspondence with the voiceprint feature among the multiple facial features as a target facial feature: determining, from the multiple facial features, a facial feature that has a corresponding face identifier; and determining the voiceprint identifier corresponding to the voiceprint feature; and determining, based on the correspondence, the facial feature corresponding to the face identifier that has a correspondence with the voiceprint identifier as the target facial feature.
[0100] In an optional embodiment, the program 414 is also used to enable the processor 402 to perform the following operations when determining the spectrum mask for indicating noise data in the speech spectrum data based on the facial features, the voiceprint features and the speech spectrum data: perform feature fusion on the facial features and the voiceprint features to obtain voiceprint-face fusion features; and perform feature extraction on the speech spectrum data to obtain spectrum features; and use the voiceprint-face fusion features and the spectrum features as input to obtain a spectrum mask probability map using a pre-trained neural network model, wherein each probability value in the spectrum mask probability map is used to indicate the probability that the data at the corresponding position in the speech spectrum data is noise data.
[0101] In an optional embodiment, the program 414 is also used to enable the processor 402 to perform the following operations when performing speech enhancement processing on the speech spectrum data according to the spectrum mask: perform matrix multiplication operation on the spectrum mask and the speech spectrum data, and obtain enhanced speech spectrum data according to the operation result; perform inverse Fourier transform on the enhanced speech spectrum data to obtain corresponding enhanced speech data.
[0102] In an optional embodiment, the program 414 is also used to enable the processor 402 to perform the following operations when obtaining the facial features and voiceprint features corresponding to the target face: perform facial recognition on the target face, and obtain the facial identifier of the target face based on the facial recognition result; determine the voiceprint identifier corresponding to the facial identifier, and obtain the voiceprint features corresponding to the voiceprint identifier.
[0103] In an optional embodiment, program 414 is also used to enable processor 402 to perform the following operations when acquiring facial image data and voice spectrum data containing multiple faces: acquiring facial image data and voice spectrum data containing multiple faces through a multimodal intelligent voice device.
[0104] The specific implementation of each step in program 414 can be found in the corresponding descriptions of the corresponding steps and units in the above-mentioned voice data processing method embodiment, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the above-mentioned method embodiment, and will not be repeated here.
[0105] Through the intelligent voice device of this embodiment, when the intelligent voice device is used in a crowded and noisy environment, the voice and image are combined. First, based on the data obtained by fusing the facial image data and the voice spectrum data, the target user who issues voice commands to the intelligent voice device, that is, the user corresponding to the target face, is determined; then, based on the facial features, voiceprint features, and voice spectrum data corresponding to the target face, a spectrum mask is obtained; and then voice enhancement is performed through the spectrum mask. Because the facial image data will not be affected even in a noisy environment, the target user can still be determined relatively accurately. On this basis, a spectrum mask is determined for noise data that can be used to indicate voices other than the target user. The spectrum mask is used to filter out the voices of non-target users as much as possible, thereby achieving the effect of enhancing the target user's voice. As a result, even when using the intelligent voice device in a noisy environment, the user can still interact with the intelligent voice device normally, improving the user experience.
[0106] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0107] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the voice data processing method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the voice data processing method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the voice data processing method shown here.
[0108] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0109] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.
Claims
1. A method for processing speech data, comprising: Acquire facial image data and speech spectrum data containing multiple faces; Processing the facial image data and the speech spectrum data to determine a target face; Obtaining facial features and voiceprint features corresponding to the target face, and determining a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, the voiceprint features, and the speech spectrum data, wherein the voiceprint features are determined based on a pre-stored correspondence between facial features and voiceprint features; Performing speech enhancement processing on the speech spectrum data according to the spectrum mask.
2. The method according to claim 1, wherein The processing of the facial image data and the voice spectrum data to determine a target face includes: Performing face detection on the plurality of face image data within a preset time period, and intercepting a face image portion corresponding to each of the plurality of faces from the plurality of face image data according to the face detection result; sorting the facial image parts of each face captured according to the image acquisition time sequence to generate multiple facial image sequences; The multiple facial image sequences are matched with the speech spectrum data respectively, and a target face is determined from the multiple faces according to the matching results.
3. The method according to claim 2, wherein: The matching of the plurality of facial image sequences with the speech spectrum data respectively, and determining a target face from the plurality of faces according to the matching results, comprises: Based on the time information corresponding to each facial image sequence and the time information corresponding to the speech spectrum data, a facial image sequence in which the continuous appearance time of the face is consistent with the duration of the speech spectrum data is matched, and the target face is determined from the multiple faces based on the matching results.
4. The method according to claim 2, wherein: The matching of the plurality of facial image sequences with the speech spectrum data respectively, and determining a target face from the plurality of faces according to the matching results, comprises: performing feature extraction on facial image portions of multiple faces included in the multiple facial image sequences to obtain corresponding multiple facial features; and performing feature extraction on the speech spectrum data to obtain corresponding voiceprint features; According to the pre-stored correspondence between facial features and voiceprint features, determining a facial feature among the multiple facial features that has a correspondence with the voiceprint feature as a target facial feature; The face corresponding to the target facial feature is determined as the target face.
5. The method according to claim 4, wherein The pre-stored correspondence between facial features and voiceprint features is a correspondence between a facial identifier of a facial feature and a voiceprint identifier of a voiceprint feature; The step of determining, based on the pre-stored correspondence between facial features and voiceprint features, a facial feature among the plurality of facial features that has a correspondence with the voiceprint feature as a target facial feature includes: Determining, from the plurality of facial features, a facial feature having a corresponding facial identifier; and determining a voiceprint identifier corresponding to the voiceprint feature; According to the corresponding relationship, the facial features corresponding to the facial identifier having a corresponding relationship with the voiceprint identifier are determined as target facial features.
6. The method according to claim 1, wherein The determining, based on the facial features, the voiceprint features, and the speech spectrum data, a spectrum mask for indicating noise data in the speech spectrum data includes: Performing feature fusion on the facial features and the voiceprint features to obtain voiceprint-face fusion features; and performing feature extraction on the speech spectrum data to obtain spectrum features; Taking the voiceprint face fusion feature and the spectral feature as input, a pre-trained neural network model is used to obtain a spectral mask probability map, wherein each probability value in the spectral mask probability map is used to indicate the probability that the data at the corresponding position in the speech spectrum data is noise data.
7. The method according to claim 6, wherein: The performing speech enhancement processing on the speech spectrum data according to the spectrum mask includes: Performing a matrix multiplication operation on the spectrum mask and the speech spectrum data, and obtaining enhanced speech spectrum data according to the operation result; Perform an inverse Fourier transform on the enhanced speech spectrum data to obtain corresponding enhanced speech data.
8. The method according to claim 1, wherein The acquiring of facial features and voiceprint features corresponding to the target face includes: Performing facial recognition on the target face, and obtaining a facial identifier of the target face according to the facial recognition result; Determine a voiceprint identifier corresponding to the face identifier, and obtain a voiceprint feature corresponding to the voiceprint identifier.
9. The method according to claim 1, wherein The acquiring of facial image data and speech spectrum data containing multiple faces includes: Acquire facial image data and speech spectrum data containing multiple faces through a multimodal intelligent voice device.
10. A speech data processing device, comprising: A data acquisition module is used to acquire facial image data and speech spectrum data containing multiple faces; a processing and determining module, configured to process the facial image data and the speech spectrum data to determine a target face; a spectrum mask acquisition module, configured to acquire facial features and voiceprint features corresponding to the target face, and determine a spectrum mask for indicating noise data in the speech spectrum data based on the facial features, the voiceprint features, and the speech spectrum data, wherein the voiceprint features are determined based on a pre-stored correspondence between facial features and voiceprint features; The speech enhancement module is used to perform speech enhancement processing on the speech spectrum data according to the spectrum mask.
11. A smart device comprising: Voice acquisition device, image acquisition device, processor; in, The voice collection device is used to collect voice data; An image acquisition device, for acquiring facial images; A processor, configured to receive facial image data containing multiple faces acquired by the image acquisition device and voice data acquired by the voice acquisition device and convert them into voice spectrum data; and, based on the facial image data and the voice spectrum data, perform operations corresponding to the voice data processing method according to any one of claims 1 to 9.
12. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for processing speech data according to any one of claims 1 to 9 is implemented.