Multimodal speech recognition method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202211150783.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-09-21
AI Technical Summary
但是,本案发明人研究发现,唇语视频虽然能够在一定程度上补充语音识别的可用信息量,但是并未考虑全面,语音识别的准确度还有很大的提升空间
[0018] Using the above technical solution, this application acquires the speech of a speaker and a video of the speaker, and processes the speech and video using a pre-configured multimodal speech recognition model to obtain the output recognized text. The video input to the multimodal speech recognition model includes at least one of facial video and body video. The inventors of this application have discovered that the facial expressions, head posture, and body movements of a person during speech convey certain signals and are highly correlated with the content of the speech. Therefore, existing technologies that rely solely on lip video to utilize visual signals are insufficient. This application expands the video information from traditional lip video to a wider range of facial and body videos, thereby utilizing richer visual cues to provide more auxiliary information and effectively improve the accuracy of speech recognition.
Smart Images

Figure CN115565534B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and more specifically, to a multimodal speech recognition method, apparatus, device, and storage medium. Background Technology
[0002] When a person speaks, the lips are the vocal organs; therefore, video footage of lip movements is highly correlated with the speaker's spoken text content. Lip reading and multimodal speech recognition rely on visual signal input.
[0003] Existing lip-reading and multimodal speech recognition technologies utilize videos of human lip movements to extract useful representations related to speech content, thereby enabling automatic recognition of spoken text. Because lip-reading videos are unaffected by environmental noise, fusing visual signals can effectively improve the robustness of speech recognition systems in noisy environments. However, the inventors in this study found that while lip-reading videos can supplement the available information in speech recognition to some extent, they are not comprehensive, and there is still significant room for improvement in the accuracy of speech recognition.
[0004] In addition, existing technologies typically require processing facial videos such as key point detection, affine transformation, and lip cropping to extract lip features during speech recognition, which places high demands on the performance of the equipment. Summary of the Invention
[0005] In view of the above problems, this application is made to provide a multimodal speech recognition method, apparatus, device, and storage medium to further improve the accuracy of speech recognition and reduce the performance requirements of speech recognition processing equipment. The specific solution is as follows:
[0006] Firstly, a multimodal speech recognition method is provided, including:
[0007] Acquire the speaker's voice and the captured video during the speaking process, wherein the video includes at least one of facial video and human body video;
[0008] The speech and video are processed using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model;
[0009] The multimodal speech recognition model is configured to: extract visual features related to the speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features.
[0010] Secondly, a multimodal speech recognition device is provided, comprising:
[0011] The voice and video acquisition unit is used to acquire the voice of a speaker and the video captured during the speaker's speech process, wherein the video includes at least one of face video and body video;
[0012] The model processing unit is used to process the speech and the video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model.
[0013] The multimodal speech recognition model is configured to: extract visual features related to the speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features.
[0014] Thirdly, a multimodal speech recognition device is provided, including: a memory and a processor;
[0015] The memory is used to store programs;
[0016] The processor is used to execute the program to implement the various steps of the multimodal speech recognition method described above.
[0017] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the multimodal speech recognition method as described above.
[0018] Using the above technical solution, this application acquires the speech of a speaker and a video of the speaker, and processes the speech and video using a pre-configured multimodal speech recognition model to obtain the output recognized text. The video input to the multimodal speech recognition model includes at least one of facial video and body video. The inventors of this application have discovered that the facial expressions, head posture, and body movements of a person during speech convey certain signals and are highly correlated with the content of the speech. Therefore, existing technologies that rely solely on lip video to utilize visual signals are insufficient. This application expands the video information from traditional lip video to a wider range of facial and body videos, thereby utilizing richer visual cues to provide more auxiliary information and effectively improve the accuracy of speech recognition.
[0019] Furthermore, considering that facial videos and body videos contain more visual cues than lip videos, they may also contain more interfering information, the multimodal speech recognition model in this application is configured to extract facial visual features related to the speech content when the input video includes facial videos, and to extract human visual features related to the speech content when the input video includes body videos. This enables the model to extract information related to speech recognition from the video, thereby better assisting speech recognition and improving its accuracy.
[0020] Furthermore, since the video used in the speech recognition process of this application is a face video or a human body video, it does not require further processing to extract lip video as in existing technologies. That is, it does not require key point detection, affine transformation, lip cropping, etc., which greatly reduces the performance requirements of speech recognition processing equipment and can further expand the application scenarios of speech recognition systems. Attached Figure Description
[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0022] Figure 1 A flowchart illustrating the multimodal speech recognition method provided in this application embodiment;
[0023] Figure 2 A schematic diagram illustrating the processing procedure of a multimodal speech recognition model is provided.
[0024] Figure 3 This example illustrates the pre-training process of a multimodal speech recognition model.
[0025] Figures 4-7 The diagrams illustrate the processing steps of the multimodal speech recognition model under several different input conditions.
[0026] Figure 8 A schematic diagram of a multimodal speech recognition device provided in an embodiment of this application;
[0027] Figure 9 This is a schematic diagram of the structure of a multimodal speech recognition device provided in an embodiment of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] This application provides a multimodal speech recognition scheme that improves the accuracy of speech recognition results by acquiring the speaker's speech and video of the speaker, and combining the speech and video.
[0030] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.
[0031] Next, combined Figure 1 The multimodal speech recognition method of this application may include the following steps:
[0032] Step S100: Obtain the voice recording of the speaker's speech and the captured video, wherein the video includes at least one of a face video and a body video.
[0033] Specifically, in order to recognize the speaker's speech, this step may acquire the recorded audio of the speaker speaking, as well as a video of the speaker speaking. The video acquired in this step may include at least one of facial video and body video.
[0034] Among them, facial video refers to video containing the facial expressions of the speaker, and human body video refers to video containing the head, body posture and movements of the human body.
[0035] This application takes into account that facial expressions, head and body postures, and movements during speech convey certain signals and are highly correlated with the content of the speech. For example, when a speaker's facial expression is confident and firm, the speech content usually matches the emotion. When a speaker is emotionally agitated, the corresponding speech content also conveys the corresponding emotional information. In addition, a person's head, body posture, and movements will also vary depending on the content of speech, which is consistent with human communication habits. Therefore, people naturally use body language to help communicate and convey meaning in communication. For example, a person may shake their head to indicate negation or wave goodbye during speech.
[0036] Step S110: Process the speech and the video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model.
[0037] Specifically, this application can pre-train a multimodal speech recognition model that can support simultaneous input of speech and video. The video can include only human body video or face video, or both human body video and face video.
[0038] In one feasible approach, the multimodal speech recognition model can be configured to: extract visual features related to speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. Wherein, when the input video includes a face video, the extracted visual features include face visual features, and when the input video includes a human body video, the extracted visual features include human body visual features.
[0039] It should be noted that in this embodiment, the multimodal speech recognition model is configured to extract facial visual features from input facial videos and human visual features from input human videos. However, it is not limited to requiring both facial and human videos as input. When the input to the multimodal speech recognition model consists only of speech and facial videos, the model can fuse the extracted facial visual features and speech features, and then encode and decode the fused features to obtain the predicted recognition text. Similarly, when the input to the multimodal speech recognition model consists only of speech and human videos, the model can fuse the extracted human visual features and speech features, and then encode and decode the fused features to obtain the predicted recognition text.
[0040] Understandably, compared to lip-sync videos, face videos and body videos, while containing more useful visual cues, also contain more interfering information, meaning their relevance to the speech content is weaker. Therefore, in this embodiment, pre-training can enable the multimodal speech recognition model to effectively capture the connection between the video and abstract semantics, extract visual features related to the speech content from the video, and then fuse them with the speech features. This fused feature contains more information, and the accuracy of the recognized text obtained by encoding and decoding based on this is also higher.
[0041] The multimodal speech recognition method provided in this application acquires the speech of a speaker and a video of the speaker, and processes the speech and video using a pre-configured multimodal speech recognition model to obtain the output recognized text. The video input to the multimodal speech recognition model includes at least one of facial video and body video. The inventors of this application have discovered that the facial expressions, head posture, and body movements of a person during speech convey certain signals that are highly correlated with the content of the speech. Therefore, existing technologies that rely solely on lip-based video do not fully utilize visual signals. This application expands the video information beyond traditional lip-based video to include a wider range of facial and body videos, thereby utilizing richer visual cues to provide more auxiliary information and effectively improve the accuracy of speech recognition.
[0042] Furthermore, considering that facial videos and body videos contain more visual cues than lip videos, they may also contain more interfering information, the multimodal speech recognition model in this application is configured to extract facial visual features related to the speech content when the input video includes facial videos, and to extract human visual features related to the speech content when the input video includes body videos. This enables the model to extract information related to speech recognition from the video, thereby better assisting speech recognition and improving its accuracy.
[0043] Furthermore, since the video used in the speech recognition process of this application is a face video or a human body video, it does not require further processing to extract lip video as in existing technologies. That is, it does not require key point detection, affine transformation, lip cropping, etc., which greatly reduces the performance requirements of speech recognition processing equipment and can further expand the application scenarios of speech recognition systems.
[0044] In some embodiments of this application, a multimodal speech recognition model is described.
[0045] Combination Figure 2 The paper discloses a schematic diagram of the processing procedure of a multimodal speech recognition model. The multimodal speech recognition model may include a first visual feature extractor, a second visual feature extractor, a speech feature extractor, an encoder, and a decoder.
[0046] The first visual feature extractor is used to extract facial visual features related to the speech content from the input facial video.
[0047] The second visual feature extractor is used to extract human visual features related to the speech content from the input human video.
[0048] A speech feature extractor is used to extract speech features from input speech.
[0049] The encoder is used to encode features based on the fusion of facial visual features, human visual features, and speech features, and the decoder predicts the corresponding recognition text based on the encoding results.
[0050] It should be noted that when the input to the multimodal speech recognition model consists only of speech and facial video, the second visual feature extractor does not operate; instead, the fused feature is obtained by fusing speech features and facial visual features. Similarly, when the input to the multimodal speech recognition model consists only of speech and human video, the first visual feature extractor does not operate; instead, the fused feature is obtained by fusing speech features and human visual features.
[0051] Among them, the first visual feature extractor, the second visual feature extractor, and the speech feature extractor adopt neural network structures, such as CNN.
[0052] Furthermore, in order to enable the first visual feature extractor and the second visual feature extractor to extract visual features related to the speech content, and the speech feature extractor to extract speech features helpful for speech recognition, this embodiment can employ a pre-training method to pre-train the structure of the multimodal speech recognition model excluding the decoder. The pre-training process can use a mask prediction method for training. Next, combined with... Figure 3 As shown, the pre-training process is described, including the following steps:
[0053] S1. Obtain a training sample set, which contains multiple sets of training samples. Each set of training samples includes face training videos, human body training videos, and training audio taken from at least one viewpoint.
[0054] Specifically, considering that in real-world scenarios, facial and human body videos may be captured from different perspectives, or may contain videos captured from multiple different perspectives simultaneously, this application, during the model pre-training process, includes a training sample set where each training sample set comprises facial training videos, human body training videos, and training audio captured from at least one perspective.
[0055] S2. Based on the acoustic features of each frame in each training speech, determine the pseudo-label for each frame.
[0056] Specifically, this application employs a masked training method when pre-training the multimodal speech recognition model. The extracted features are masked, and training is performed with the encoder predicting the pseudo-labels corresponding to the mask as the target. For this purpose, it is necessary to pre-determine the pseudo-labels for each frame in each training speech. To enable the first and second visual feature extractors to extract visual features related to the speech content, this embodiment determines the pseudo-labels for each frame based on the acoustic features of each frame.
[0057] Among them, acoustic features can be MFCC (Mel-frequency cepstral coefficients) or other types of acoustic features.
[0058] This embodiment describes a specific implementation of step S2, including the following steps:
[0059] S21. Extract the acoustic features of all training speech in the training sample set, and cluster each acoustic feature to obtain the class centers of several clusters.
[0060] Specifically, when clustering various acoustic features, clustering algorithms such as k-means can be used.
[0061] S22. Calculate the distance between the acoustic features of each frame in each training speech and the center of each class, and use the label of the nearest class center as the pseudo label of the corresponding frame.
[0062] S3. For each set of training samples: randomly select a face training video from a viewpoint and input it into the first visual feature extractor to obtain face visual training features; randomly select a human body training video from a viewpoint and input it into the second visual feature extractor to obtain human body visual training features; input the training speech into the speech feature extractor to obtain speech training features.
[0063] S4. For each of the facial visual training features, the human visual training features, and the speech training features, randomly select several consecutive frames of features and perform masking processing to obtain masked facial visual training features, masked human visual training features, and masked speech training features.
[0064] Specifically, any one of the following features—face visual training features, human visual training features, and speech training features—can be represented as F = {f1, f2, ..., f...} T In the formula}, where T represents the length, the starting position of the mask segment is randomly selected from the position 1-T using a Bernoulli distribution with probability p. Then, the features of the consecutive s frames starting from the starting position are masked. During masking, a learnable embedding vector can be used to replace the original features, thereby obtaining the mask training features.
[0065] For ease of description, the visual training features of the masked face are defined as F. v1 The masked human visual training features are represented as F. v2 The masked speech training features are represented as F a .
[0066] S5. The masked face visual training features, masked human body visual training features, and masked voice training features are fused to obtain fused training features.
[0067] Specifically, during fusion, channel-dimensional splicing can be used for fusion, thus fusing the training features F. multi Represented as:
[0068] F multi =Concat(F v1 ,F v2 ,F a )
[0069] Here, `Concat()` represents concatenation along the channel dimension. For example, F... v1 F v2 ,F a The dimensions are all T*D, where T represents the time dimension with T frames and D represents the channel dimension of the feature. The dimension of the fused training feature after the three features are spliced together in the channel dimension is T*3D.
[0070] Optionally, to further improve the model's robustness to feature loss and to make fuller use of each feature, this embodiment may also introduce a random feature discarding strategy during fusion, specifically:
[0071] According to the preset first discard ratio p1 corresponding to the masked face visual training features, the second discard ratio p2 corresponding to the masked human body visual training features, and the third discard ratio p3 corresponding to the masked speech training features, the masked face visual training features, the masked human body visual training features, and the masked speech training features are spliced and fused using a random feature discarding strategy to obtain fused training features.
[0072] Then the fusion training features F multi Represented as:
[0073] F multi =Concat(Dropout(F v1 ,p1),Dropout(F v2 ,p2),Dropout(F a ,p3))
[0074] Dropout refers to the random dropping operation.
[0075] It should be noted that the above random discard operation refers to randomly setting all zeros.
[0076] S6. Input the fused training features into the encoder, and the encoder predicts the pseudo-label prediction result corresponding to the mask position in the fused training features.
[0077] The encoder can be a transformer-structured encoder network. The fused training features are input into the encoder, and the encoder's output corresponding to the feature mask position passes through a fully connected layer and a softmax activation function, outputting the pseudo-label prediction result.
[0078] S7. Based on the pseudo-label prediction results of the mask position predicted by the encoder, and the pseudo-labels corresponding to each frame of the mask position, calculate the loss, and train the network parameters of each structure in the multimodal speech recognition model except for the decoder based on the loss, until the training termination condition is met.
[0079] According to the pre-training method provided in this embodiment, the network parameters of the first and second visual feature extractors, the speech feature extractor, and the encoder in the multimodal speech recognition model can be trained, enabling the first and second visual feature extractors to extract visual features related to the speech content. Based on this, a downstream multimodal speech recognition model can be effectively constructed.
[0080] Optionally, the process for determining pseudo-labels for each frame of the training speech in the pre-training phase described above can be implemented. In the early stages of pre-training, pseudo-labels can be set using acoustic features, as described in step S2 above. After training to a certain stage, the model itself gains the ability to extract abstract features. Therefore, hidden features can be extracted from the encoder output of the model, and these hidden features can replace the acoustic features. Clustering and pseudo-label determination are then performed again, and the model is trained further based on the redefined pseudo-labels until training is complete. Specifically:
[0081] Before reaching the training termination conditions, the pre-training process, after several rounds of pre-training, also includes:
[0082] Extract the hidden layer features from the last hidden layer output of the encoder, and cluster each hidden layer feature to obtain the cluster centers of several clusters. Calculate the distance between the hidden layer features of each frame in each training speech and each cluster center, and use the label of the nearest cluster center as the updated pseudo-label for the corresponding frame. Replace the original pseudo-label with the updated pseudo-labels of each frame, and continue training the multimodal speech recognition model using training samples until the training termination condition is met.
[0083] Furthermore, after completing the pre-training of the multimodal speech recognition model, the model can be fine-tuned using downstream speech recognition task data. During the fine-tuning stage, a decoder module is introduced. This decoder is responsible for predicting and recognizing text labels from the encoder's output, thus completing the speech recognition task. Specifically:
[0084] After pre-training, the following is also included:
[0085] Obtain a training dataset, which includes multiple sets of training samples and the recognition text corresponding to each set of training samples. Each set of training samples includes face training videos, human body training videos, and training audio taken from at least one viewpoint.
[0086] The network parameters of each structure in the multimodal speech recognition model are fine-tuned using the training dataset. In this model, the network parameters of each structure except the decoder are reused from the pre-trained network parameters.
[0087] Optionally, during the fine-tuning phase, similar to the pre-training phase, the operation of randomly selecting a face training video and a human body training video taken from a single viewpoint from each training sample group can still be retained. Simultaneously, the random discarding operation during multimodal feature fusion can also be retained, ensuring that the fine-tuned multimodal speech recognition model does not depend on any single feature and can still effectively perform speech recognition even when a certain feature is missing.
[0088] The above embodiments described the structure and training process of the multimodal speech recognition model. This application embodiment further explains the speech recognition process based on the aforementioned trained multimodal speech recognition model.
[0089] The foregoing Figure 1 In the corresponding speech recognition process, step S100 has already been described, which involves acquiring the speaker's speech and the captured video, wherein the video includes at least one of facial video and body video. In different scenarios, it may be possible to acquire only the speech and facial video, or only the speech and body video. The acquired facial video may include facial videos from one or more viewpoints, and the body video may also include body videos from one or more viewpoints. Therefore, the input data for the multimodal speech recognition model can contain various different combinations.
[0090] This embodiment exemplifies the specific implementation process of using a multimodal speech recognition model to process speech and video to obtain recognized text under several different combinations.
[0091] In the first case:
[0092] The acquired videos include any one or both of face videos and body videos, and for both face videos and body videos, only videos shot from one perspective are included. Of course, when a video contains both face videos and body videos, the shooting perspectives for the two can be the same or different.
[0093] In this case, the process of processing the speech and the video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model may include:
[0094] The multimodal speech recognition model takes a face video from one perspective and / or a human body video from one perspective, along with the acquired speech input, to obtain the recognized text output by the model.
[0095] like Figure 4 , Figure 5 As shown:
[0096] in, Figure 4 This example illustrates the processing flow of a multimodal speech recognition model when the input consists only of a face video and speech from shooting angle i.
[0097] Figure 5 The example illustrates the processing flow of a multimodal speech recognition model when the input consists only of a face video from shooting angle i, and a human body video and speech from shooting angle j.
[0098] Among them, angle i and angle j can be the same or different.
[0099] In the second case:
[0100] The acquired videos include any one or both of face videos and body videos, and for each type of video, at least one type is captured from two or more perspectives. That is, the acquired videos contain face videos from at least two perspectives, or body videos from at least two perspectives.
[0101] In this case, the process of using a pre-configured multimodal speech recognition model to process the speech and the video to obtain the recognized text output by the model can include two processing methods: 1)
[0103] Select a face video and / or a human body video from any viewpoint from the acquired video, and input them together with the acquired speech into a multimodal speech recognition model to obtain the recognized text output by the model.
[0104] Specifically, because the multimodal speech recognition model incorporates a random feature discarding strategy during both the pre-training and fine-tuning phases, the trained model can perform speech recognition based on any one or two types of video and audio. Therefore, in this processing method, a face video and / or a human body video from a random perspective can be selected from the acquired video and input along with the acquired audio into the multimedia speech recognition model to obtain the recognized text output by the model. Details are as follows... Figure 6 As shown. 2)
[0106] The acquired speech is combined with facial video and / or human body video from any viewpoint to obtain at least one set of inputs.
[0107] Understandably, depending on the shooting angle of the face video and the shooting angle of the body video, there can be multiple combinations of the above, meaning that there can be multiple sets of inputs.
[0108] Each set of inputs is fed into a multimodal speech recognition model. The posterior probabilities of the different multimodal speech recognition models are averaged, and the final recognized text is predicted based on the average posterior probability. The network parameters of the multimodal speech recognition models corresponding to each set of inputs are consistent.
[0109] Specifically, this application can set the same number of multimodal speech recognition models according to the number of input groups, and the network parameters of each multimodal speech recognition model are completely identical.
[0110] Then, multiple multimodal speech recognition models are integrated. Each multimodal speech recognition model processes a set of input data accordingly. Finally, the posterior probabilities of the decoder outputs of each multimodal speech recognition model are averaged, and the final recognized text is predicted based on the average posterior probability.
[0111] By using multiple sets of inputs corresponding to multiple multimodal speech recognition models and averaging the posterior probabilities, the final recognized text can be predicted. This approach can fully utilize different video data to achieve optimal recognition results.
[0112] See Figure 7 , Figure 7 The example illustrates the processing flow of a multimodal speech recognition model when it includes face videos from shooting angles i and j, as well as human body videos and audio from shooting angle j.
[0113] Depend on Figure 7 It can be seen that this application constructs three sets of inputs, namely:
[0114] First input: Face video and voice at angle i
[0115] Second input: Face video and audio at angle j
[0116] The third input group: human body video and audio at angle j.
[0117] The multimodal speech recognition device provided in the embodiments of this application will be described below. The multimodal speech recognition device described below can be referred to in correspondence with the multimodal speech recognition method described above.
[0118] See Figure 8 , Figure 8This is a schematic diagram of the structure of a multimodal speech recognition device disclosed in an embodiment of this application.
[0119] like Figure 8 As shown, the device may include:
[0120] The voice and video acquisition unit 11 is used to acquire the voice of the speaker and the video captured during the speaker's speech process, wherein the video includes at least one of face video and body video;
[0121] The model processing unit 12 is used to process the speech and the video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model.
[0122] The multimodal speech recognition model is configured to: extract visual features related to the speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features.
[0123] Optionally, the above multimodal speech recognition model may include: a first visual feature extractor, a second visual feature extractor, a speech feature extractor, an encoder, and a decoder;
[0124] The first visual feature extractor is used to extract facial visual features related to the speech content from the input facial video;
[0125] The second visual feature extractor is used to extract human visual features related to the speech content from the input human video.
[0126] The speech feature extractor is used to extract speech features from the input speech;
[0127] The encoder is used to encode based on the fusion features of facial visual features, human visual features and speech features, and the decoder predicts the corresponding recognition text based on the encoding result.
[0128] Optionally, the apparatus of this application may further include a model pre-training unit for training the structure of the multimodal speech recognition model, excluding the decoder, using a sampling pre-training method. The pre-training process includes:
[0129] Obtain a training sample set, which contains multiple sets of training samples. Each set of training samples includes at least one face training video, human body training video, and training audio taken from at least one perspective.
[0130] Based on the acoustic features of each frame in each training speech, determine the pseudo-label for each frame.
[0131] For each set of training samples: randomly select a face training video from a viewpoint and input it into the first visual feature extractor to obtain face visual training features; randomly select a human body training video from a viewpoint and input it into the second visual feature extractor to obtain human body visual training features; input the training speech into the speech feature extractor to obtain speech training features.
[0132] For each of the facial visual training features, the human visual training features, and the speech training features, several consecutive frames of features are randomly selected and masked to obtain masked facial visual training features, masked human visual training features, and masked speech training features.
[0133] The masked face visual training features, masked human body visual training features, and masked speech training features are fused to obtain fused training features;
[0134] The fused training features are input into the encoder, which then predicts the pseudo-label prediction results corresponding to the mask positions in the fused training features.
[0135] The loss is calculated based on the pseudo-label prediction results of the mask position predicted by the encoder, and the pseudo-labels corresponding to each frame at the mask position.
[0136] The network parameters of each structure in the multimodal speech recognition model, excluding the decoder, are trained based on the loss until the training termination condition is met.
[0137] Optionally, the process by which the model pre-training unit determines the pseudo-label for each frame based on the acoustic features of each frame in each training speech may include:
[0138] The acoustic features of all training speech in the training sample set are extracted, and each acoustic feature is clustered to obtain the class centers of several clusters.
[0139] Calculate the distance between the acoustic features of each frame in each training speech and the center of each class, and use the label of the nearest class center as the pseudo-label of the corresponding frame.
[0140] Optionally, after several rounds of pre-training during the pre-training process, the aforementioned model pre-training unit is also used before reaching the training termination condition to:
[0141] Extract the hidden layer features of the last hidden layer output of the encoder, and cluster each hidden layer feature to obtain the class centers of several clusters;
[0142] Calculate the distance between the hidden layer features of each frame in each training speech and the class center of each class. Use the label of the nearest class center as the updated pseudo-label of the corresponding frame. Replace the original pseudo-label with the updated pseudo-label of each frame. Continue to train the multimodal speech recognition model using training samples until the training termination condition is met.
[0143] Optionally, the apparatus of this application may further include a model fine-tuning unit, used to acquire a training dataset after pre-training, the training dataset including multiple sets of training samples and recognition text corresponding to each set of training samples, each set of training samples including face training video, human body training video and training speech taken from at least one viewpoint; and to fine-tune the network parameters of each structure in the multimodal speech recognition model using the training dataset, wherein the other structures in the multimodal speech recognition model, except for the decoder, reuse the pre-trained network parameters.
[0144] Optionally, the process by which the above-mentioned model pre-training unit fuses the masked face visual training features, masked human body visual training features, and masked speech training features to obtain fused training features may include:
[0145] According to a preset first discard ratio corresponding to the masked face visual training features, a second discard ratio corresponding to the masked human visual training features, and a third discard ratio corresponding to the masked speech training features, a random feature discarding strategy is used to splice and fuse the masked face visual training features, the masked human visual training features, and the masked speech training features to obtain fused training features.
[0146] Optionally, the video acquired by the voice and video acquisition unit includes a face video from one viewpoint and / or a human body video from one viewpoint. Based on this, the model processing unit processes the voice and video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model. This process may include:
[0147] The facial video from one perspective, and / or the human body video from one perspective, along with the voice input multimodal speech recognition model, are used to obtain the recognized text output by the model.
[0148] Optionally, the video acquired by the speech and video acquisition unit includes facial videos from at least two perspectives, or human body videos from at least two perspectives. Based on this, the process by which the model processing unit processes the speech and video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model may include:
[0149] Select a face video and / or a human body video from any viewpoint from the acquired video, and input them together with the acquired speech into a multimodal speech recognition model to obtain the recognized text output by the model.
[0150] or,
[0151] Combine the acquired speech with facial video and / or human body video from any viewpoint to obtain at least one set of inputs;
[0152] Each set of inputs is fed into a multimodal speech recognition model. The posterior probabilities of the different multimodal speech recognition models are averaged, and the final recognized text is predicted based on the average posterior probability. The network parameters of the multimodal speech recognition models corresponding to each set of inputs are consistent.
[0153] The multimodal speech recognition device provided in this application embodiment can be applied to multimodal speech recognition devices, such as terminals: mobile phones, computers, etc. Optionally, Figure 9 The hardware structure block diagram of the multimodal speech recognition device is shown below. Figure 9 The hardware structure of a multimodal speech recognition device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0154] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;
[0155] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0156] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0157] The memory stores a program, which the processor can call. The program is used for:
[0158] Acquire the speaker's voice and the captured video during the speaking process, wherein the video includes at least one of facial video and human body video;
[0159] The speech and video are processed using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model;
[0160] The multimodal speech recognition model is configured to: extract visual features related to the speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features.
[0161] Optionally, the refined and extended functions of the program can be found in the description above.
[0162] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:
[0163] Acquire the speaker's voice and the captured video during the speaking process, wherein the video includes at least one of facial video and human body video;
[0164] The speech and video are processed using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model;
[0165] The multimodal speech recognition model is configured to: extract visual features related to the speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features.
[0166] Optionally, the refined and extended functions of the program can be found in the description above.
[0167] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0168] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0169] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal speech recognition method, characterized in that, include: Acquire the speaker's voice and the captured video during the speaking process, wherein the video includes at least one of facial video and human body video; The speech and video are processed using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model; The multimodal speech recognition model is configured to: extract visual features related to speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. The multimodal speech recognition model can mine information related to speech recognition from the video to assist speech recognition. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features. It also includes: enabling the multimodal speech recognition model to effectively capture the connection between video and abstract semantics through pre-training, extracting visual features related to speech content from the video, and fusing them with speech features; The multimodal speech recognition model includes: a first visual feature extractor, a second visual feature extractor, a speech feature extractor, an encoder, and a decoder; The first visual feature extractor is used to extract facial visual features related to the speech content from the input facial video; The second visual feature extractor is used to extract human visual features related to the speech content from the input human video; The speech feature extractor is used to extract speech features from the input speech; The encoder is used to encode based on the fusion features of facial visual features, human visual features and speech features, and the decoder predicts the corresponding recognition text based on the encoding results.
2. The method according to claim 1, characterized in that, The multimodal speech recognition model, excluding the decoder, is trained using a pre-training method. The pre-training process includes: Obtain a training sample set, which contains multiple sets of training samples. Each set of training samples includes at least one face training video, human body training video, and training audio taken from at least one perspective. Based on the acoustic features of each frame in each training speech, determine the pseudo-label for each frame. For each set of training samples: randomly select a face training video from a viewpoint and input it into the first visual feature extractor to obtain face visual training features; randomly select a human body training video from a viewpoint and input it into the second visual feature extractor to obtain human body visual training features; input the training speech into the speech feature extractor to obtain speech training features. For each of the facial visual training features, the human visual training features, and the speech training features, several consecutive frames of features are randomly selected and masked to obtain masked facial visual training features, masked human visual training features, and masked speech training features. The masked face visual training features, masked human body visual training features, and masked speech training features are fused to obtain fused training features; The fused training features are input into the encoder, which then predicts the pseudo-label prediction results corresponding to the mask positions in the fused training features. The loss is calculated based on the pseudo-label prediction results of the mask position predicted by the encoder, and the pseudo-labels corresponding to each frame at the mask position. The network parameters of each structure in the multimodal speech recognition model, excluding the decoder, are trained based on the loss until the training termination condition is met.
3. The method according to claim 2, characterized in that, The process of determining the pseudo-label for each frame based on the acoustic features of each frame in each training speech includes: The acoustic features of all training speech in the training sample set are extracted, and each acoustic feature is clustered to obtain the class centers of several clusters. Calculate the distance between the acoustic features of each frame in each training speech and the center of each class, and use the label of the nearest class center as the pseudo-label of the corresponding frame.
4. The method according to claim 2, characterized in that, The pre-training process, after several rounds of pre-training but before reaching the end-of-training conditions, also includes: Extract the hidden layer features of the last hidden layer output of the encoder, and cluster each hidden layer feature to obtain the class centers of several clusters; Calculate the distance between the hidden layer features of each frame in each training speech and the class center of each class. Use the label of the nearest class center as the updated pseudo-label of the corresponding frame. Replace the original pseudo-label with the updated pseudo-label of each frame. Continue to train the multimodal speech recognition model using training samples until the training termination condition is met.
5. The method according to claim 2, characterized in that, After pre-training, the following is also included: Obtain a training dataset, which includes multiple sets of training samples and the recognition text corresponding to each set of training samples. Each set of training samples includes at least one face training video, human body training video and training speech taken from at least one perspective. The network parameters of each structure in the multimodal speech recognition model are fine-tuned using the training dataset. In this model, the network parameters of each structure except the decoder are reused from the pre-trained network parameters.
6. The method according to claim 2, characterized in that, The process of fusing the masked face visual training features, masked human body visual training features, and masked speech training features to obtain fused training features includes: According to a preset first discard ratio corresponding to the masked face visual training features, a second discard ratio corresponding to the masked human visual training features, and a third discard ratio corresponding to the masked speech training features, a random feature discarding strategy is used to splice and fuse the masked face visual training features, the masked human visual training features, and the masked speech training features to obtain fused training features.
7. The method according to any one of claims 1-6, characterized in that, The video includes a face video from one perspective and / or a human body video from one perspective; The process of using a pre-configured multimodal speech recognition model to process the speech and video to obtain the recognized text output by the model includes: The facial video from one perspective, and / or the human body video from one perspective, along with the voice input multimodal speech recognition model, are used to obtain the recognized text output by the model.
8. The method according to any one of claims 1-6, characterized in that, The video includes facial videos from at least two perspectives, or human body videos from at least two perspectives; The process of using a pre-configured multimodal speech recognition model to process the speech and video to obtain the recognized text output by the model includes: Select a face video and / or a human body video from any viewpoint from the acquired video, and input them together with the acquired speech into a multimodal speech recognition model to obtain the recognized text output by the model. or, Combine the acquired speech with facial video and / or human body video from any viewpoint to obtain at least one set of inputs; Each set of inputs is fed into a multimodal speech recognition model. The posterior probabilities of the different multimodal speech recognition models are averaged, and the final recognized text is predicted based on the average posterior probability. The network parameters of the multimodal speech recognition models corresponding to each set of inputs are consistent.
9. A multimodal speech recognition device, characterized in that, include: The voice and video acquisition unit is used to acquire the voice of a speaker and the video captured during the speaker's speech process, wherein the video includes at least one of face video and body video; The model processing unit is used to process the speech and the video using a pre-configured multimodal speech recognition model to obtain the recognized text output by the model. The multimodal speech recognition model is configured to: extract visual features related to speech content from the input video, extract speech features from the input speech, fuse the extracted visual features and speech features, and encode and decode the fused features to obtain the predicted recognition text. The multimodal speech recognition model can mine information related to speech recognition from the video to assist speech recognition. When the input video includes a face video, the extracted visual features include face visual features; when the input video includes a human body video, the extracted visual features include human body visual features. The multimodal speech recognition device is also used to: enable the multimodal speech recognition model to effectively capture the connection between video and abstract semantics through pre-training, extract visual features related to speech content from the video, and fuse them with speech features; The multimodal speech recognition model includes: a first visual feature extractor, a second visual feature extractor, a speech feature extractor, an encoder, and a decoder; The first visual feature extractor is used to extract facial visual features related to the speech content from the input facial video; The second visual feature extractor is used to extract human visual features related to the speech content from the input human video; The speech feature extractor is used to extract speech features from the input speech; The encoder is used to encode based on the fusion features of facial visual features, human visual features and speech features, and the decoder predicts the corresponding recognition text based on the encoding results.
10. A multimodal speech recognition device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the multimodal speech recognition method as described in any one of claims 1 to 8.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the multimodal speech recognition method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Voice recognition method and device, medium and computing equipment
CN113889081A
Multi-modal speech recognition model training method, speech recognition method and equipment
CN114724548A
Multi-modal fusion natural interaction method and system of intelligent robot and medium
CN114995657A