Chinese lip language recognition method based on credible visual position element acquisition
Through deep clustering and deep learning methods, a trusted visual head element library was established, and the visual head category was directly extracted from the lip motion video data, solving the generalization and accuracy of the Chinese lip recognition model, and achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510302505.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the Chinese lip recognition model has insufficient generalization and accuracy, especially when facing unseen data and unseen speakers, the recognition effect is poor, and relying on audio data assistance leads to inaccurate mapping relationships, resulting in bottlenecks in accuracy.
Through deep clustering and deep learning methods, a trusted visual-point element library is established, and the visual-point categories are directly extracted from the lip motion video data, frame-by-frame image data annotation, and feature extraction and sequence decoding are used using 3D convolutional neural network and Transformer encoder to realize the mapping of visual-point to Chinese characters and reduce recognition errors.
It improves the generalization and accuracy of lip word recognition, can be applicable to unknown data and unknown speakers, reduces the cumulative error of recognition prediction, and breaks the accuracy bottleneck of optote lip word recognition.
Smart Images

Figure CN120260118A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual speech recognition, and in particular to a Chinese lip-reading recognition method based on obtaining credible visual phonemes. Background Art
[0002] Voice interaction is the most commonly used and natural way in interpersonal communication and is also a reliable and effective carrier in information transmission.
[0003] In the implementation of voice interaction, the most obvious and mature method is to encode and decode the audio signal generated by the speaker and conducted through the air to achieve speech recognition. However, in scenarios such as field operations, emergency rescue, military applications, and medical rehabilitation, the traditional air-conducted audio path cannot fully or even completely lose the representation of speech information due to reasons such as high noise interference, strong concealment requirements, and the health status of the speaker. Therefore, it is difficult to implement speech information encoding in the above complex scenarios on the transmitting side. We need to seek other information modalities released by the speaker when performing a speaking behavior or having a speaking intention to assist or replace the role of air-conducted audio. This type of processing of speech information based on non-air-conducted audio information modalities is usually referred to as silent speech recognition (or covert communication).
[0004] In visual speech recognition, or more directly lip-reading recognition, it has a strong correlation with speech recognition, and lip-reading images are also the information modality closest to actual applications for implementing silent speech recognition. Since the language content of the speaker is recognized based on the visual information of lip movement and mouth shape changes, it is not affected by the acoustic environment. Therefore, this is a very valuable research topic in practice. In recent years, with the rapid development of deep learning technology, the effect of lip-reading recognition has also been significantly improved.
[0005] The lip-reading recognition model mainly processes the lip images of speakers and is mainly oriented towards the recognition of limited datasets and a limited number of speakers. However, in actual scenarios, the recognition targets are likely to be unknown data or the speakers to be recognized are likely to be unknown (i.e., not present in the training set). There are significant differences in the corresponding features of different data, and there are often large differences in the pronunciation habits and lip region image features of different speakers. Moreover, compared with the research on English lip-reading recognition, English lip-reading recognition has made great progress, while relatively less attention has been paid to Chinese lip-reading recognition. English words are mainly composed of 26 letters, while Chinese characters are pictographic characters with tone changes. There are more than 3,000 common Chinese characters alone, and the number space at the word level and sentence level for prediction is larger. Moreover, most Chinese characters are homophones, that is, the same pronunciation mouth shape can correspond to multiple Chinese characters. Representing speech information through lip shape changes inevitably leads to the problem of homophony (same lip shape but different meanings), especially as the length of the target speech content increases, which affects comprehensibility. Therefore, compared with English, Chinese lip-reading recognition is more challenging. Therefore, how to establish a general lip-reading recognition model with good generalization ability that can be applied to unseen data and unseen speakers is a major challenge in the Chinese lip-reading recognition task.
[0006] Visemes are used as an intermediate bridge to boost the performance of lip-reading recognition. A viseme is the visual description of a phoneme in spoken language. A phoneme is the smallest sound unit in human language that can distinguish meaning, while a viseme is the smallest visual primitive that can distinguish the meaning of pronunciation, defining the instantaneous / short-term position (state) of a person's face and mouth when speaking. It can be considered that all lip-readings are composed of the sequential splicing of a limited number of viseme categories. Therefore, lip-reading recognition using viseme intermediate representation, that is, the task of fine-grained visual speech recognition relying on viseme primitives, can be simplified as follows: first, identify the viseme category of each frame, and then regard the lip-reading video as the splicing of different visemes. In this way, based on the high-accuracy recognition of a limited number of specific visual visemes, the decoding of long-sequence seen (in the training) / unseen lip-reading videos can be improved. Through deep learning algorithms, the mapping from video segments to single-frame or multi-frame specific viseme categories is completed, and then through the corresponding mapping relationships between visemes and pronunciation phonemes (pinyin in Chinese) and between pronunciation phonemes and text, lip-reading recognition is achieved. Generally speaking, lip-reading recognition based on visemes deconstructs all unseen words and phrases into linear combinations of the seen primitive library by recognizing the primitive library, thereby enhancing the model's ability to transfer and generalize to unseen text words and phrases.
[0007] The accuracy of lip reading based on visemes lies in an accurate viseme primitive library. The viseme primitive library can cover all pronunciation units of general spoken sentences, can fit the lip shapes of different people as much as possible, and has good inter-class discrimination, which ensures better performance of lip reading based on this. However, in current work, the method of determining visemes is simple and crude. Using the traditional artificial correspondence method between phonemes and visemes has a large error and cannot completely find an ideal viseme primitive library from the perspective of data distribution. At the same time, such a method relies on the assistance of audio data to specify visual symbol categories by phoneme categories, which requires high-quality audio-visual data and external toolkits. However, these resources are relatively scarce. At the same time, over-reliance on audio data, especially the assistance of phonemes, will naturally expand the problem of the one-to-many mapping relationship between vision and phonemes. Therefore, compared with phonemes that have more explicit language standards and good discrimination, the research on visemes has not been standardized, and there is no standard, well-generalized, and highly discriminative viseme primitive library, which directly leads to the technical problem of low accuracy of lip reading based on visemes in the existing technology. Summary of the Invention
[0008] Aiming at the problems existing in the prior art, the purpose of the present invention is to establish a method mechanism for viseme-based lip reading driven by extracting credible visemes, and establish a corresponding recognition prediction system. By obtaining a credible viseme primitive library with higher discrimination, better generalization, and more accurate and comprehensive characterization of pronunciation forms, as the prediction space of the intermediate representation of lip reading, the cumulative error of recognition prediction is reduced, and the performance of viseme-based lip reading is further improved, breaking through the accuracy bottleneck of viseme-based lip reading.
[0009] To achieve the above purpose, the present invention provides a Chinese lip reading method based on obtaining credible visemes, and the method includes the following steps:
[0010] S1. Data collection and preprocessing: To obtain video data depicting lip movements;
[0011] S2. Deep clustering: Perform deep clustering on the video data depicting lip movements to obtain the number of viseme categories in the clustering distribution, the corresponding viseme categories and viseme library, so as to obtain frame-by-frame image data with viseme category annotations corresponding to the video data depicting lip movements;
[0012] S3. Recognition of cascaded Chinese character sequences based on viseme intermediate representation: Extract features based on the frame-by-frame image data with viseme category annotations to realize the recognition of cascaded Chinese character sequences with visemes as the intermediate representation.
[0013] Further, step S1 further includes: after reading and frame-dividing the collected video data into a picture sequence, first performing preprocessing to focus the picture on the key information of the lips; the preprocessing includes lip region detection, cropping, and image size adjustment.
[0014] Further, in step S2, the data is sent into a deep clustering network for unsupervised clustering based on a deep neural network to obtain credible viseme clusters, and at the same time, the corresponding data distribution and annotation are formed.
[0015] Further, the deep clustering network includes an autoencoder and a deep clustering module based on mutual information, and the loss L of the deep clustering network dc includes the reconstruction loss L of the autoencoder n and the loss L of the mutual information clustering module c .
[0016] Further, the deep clustering network includes a variational autoencoder, a generative adversarial network, a siamese network, and a graph neural network model; it also includes a deep clustering module based on a K-means module, a spectral clustering module, a subspace clustering module, a KL divergence deep clustering module, and / or a Gaussian mixture model. Different neural network models and deep clustering methods can be used in combination respectively.
[0017] Further, step S3 includes:
[0018] S3.1 Based on the frame-by-frame image data with annotations after deep clustering, perform preprocessing;
[0019] S3.2 Further input the preprocessed frame-by-frame image data into a spatio-temporal feature extractor to extract the reduced spatio-temporal features through a 3D convolutional neural network;
[0020] S3.3 Input the spatio-temporal features into a viseme sub-encoder-decoder unit to output a viseme sequence, which is constrained by the ctc loss; at the same time, the intermediate encoded representation output by the encoder in the viseme sub-encoder-decoder unit is reserved for the next step;
[0021] S3.4 Based on the intermediate encoded representation output by the viseme sub-encoder-decoder unit in S3.2, input it into a Chinese character sub-encoder-decoder unit to further output a Chinese character sequence, which is constrained by the cross-entropy loss.
[0022] Further, in step S3.2, the spatio-temporal feature extractor uses a spatio-temporal convolutional neural network to extract features in the time and space dimensions for the obtained lip image sequence. The spatio-temporal feature extractor consists of a 3D convolutional network layer, a batch normalization layer, a ReLU activation function, a regularization dropout layer, and a max pooling layer.
[0023] Further, in step S3, in the visual phoneme sequence prediction and Chinese character sequence prediction, a Transformer encoder is used. The Transformer encoder includes 6 layers of Transformer Encoder structures. Each layer is composed of a multi-head self-attention module and a feed-forward network module, and also includes a residual connection and layer normalization. The feed-forward network module is composed of two linear layers and a non-linear activation function in the middle.
[0024] Further, in step S3.3, based on the output of the visual phoneme sequence prediction, character sequence decoding of an attention-based sequence-to-sequence architecture is performed. In the character sequence prediction, the encoder is a 6-layer Transformer encoder. Each layer includes a masked multi-head attention mechanism, a multi-head attention mechanism, and a feed-forward neural network. In addition, it also includes a residual connection and layer normalization.
[0025] Further, the input of the encoder is the vector representation of the visual phoneme, and the output of the decoder is the character sequence. During the training process, an autoregressive method is used for parallel optimization. The prediction result at each time step will be used as the input for the decoding of the next time step. During the decoding process, the historical prediction output results are utilized, and at the same time, meaningful parts in the visual phoneme sequence are selectively extracted through the attention mechanism.
[0026] The beneficial effects of the present invention are as follows:
[0027] 1. The present invention proposes a Chinese sign language recognition method based on obtaining credible visual phonemes. Through a deep neural network and the natural distribution of sign language data, a credible visual phoneme primitive library that can more accurately and comprehensively depict the pronunciation form and is suitable for subsequent sign language recognition can be found. This method does not rely on audio resource support and the supplement of phoneme (pinyin) information, is more flexible, and has a low dependence on data resources, additional tools, and manual operations.
[0028] 2. The present invention relies on visual depth clustering and visual Chinese character modeling to directly establish a mapping from visual phonemes to Chinese characters, endowing the video with meaning encoded by the visual phoneme sequence for the first time without the need for an intermediate interpretation of pinyin information, reducing the cumulative error, and further improving the sign language recognition performance based on the intermediate representation of visual phonemes. Description of the Drawings
[0029] Figure 1 It is a schematic diagram of the overall algorithm flow of the Chinese sign language recognition method based on obtaining credible visual phonemes according to the present invention;
[0030] Figure 2 It is a schematic diagram of the deep clustering network architecture according to the present invention;
[0031] Figure 3 It is a schematic diagram of the structure of the spatio-temporal feature extractor according to the present invention;
[0032] Figure 4 Schematic diagram of the Transformer encoder structure (for visual isotope sequence prediction) according to the present invention;
[0033] Figure 5 Schematic diagram of the Transformer decoder structure (for character sequence prediction) according to the present invention. Detailed implementation manners
[0034] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0036] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "install", "connect", "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0037] The following is combined with Figures 1 - 5 The specific implementation manners of the present invention will be described in detail. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0038] The inventive concept of this application lies in establishing a method mechanism for viseme-based lip-reading recognition driven by extracting credible visemes, and establishing a corresponding recognition prediction system. By obtaining a credible viseme primitive library (i.e., viseme categories and viseme library) with higher discrimination, better generalization, and more accurate and comprehensive characterization of pronunciation forms, as the prediction space for the intermediate representation of lip-reading recognition, the cumulative error of recognition prediction is reduced, and the performance of viseme-based lip-reading recognition is further improved. The present invention can obtain a standard, highly generalizable, and highly discriminative viseme primitive library, and based on this, establish a generalizable lip-reading recognition model that can be applied to unseen data. The unseen data refers to data of new categories, sentence categories that have not appeared in model training. The general lip-reading recognition model for unseen speakers breaks through the accuracy bottleneck of viseme-based lip-reading recognition. It can be applied to the communication and rehabilitation of patients with speech disorders in medical and rehabilitation engineering; communication in scenarios such as high-noise or concealed and private military operations and emergency rescue.
[0039] The improvement of the present invention mainly lies in the assistance of viseme prediction. Through fine-grained prediction of frame-by-frame representation, the accuracy of sequence prediction is further improved. A viseme can be understood as having a specific category attribute in each frame of an image in a lip movement video (similar to the pinyin in the sense of visual representation). In deep learning-based lip-reading recognition, the characteristics of each frame of the video having a specific category (viseme category) are utilized to improve the lip-reading recognition effect. To solve the above problems and achieve the above goals, the present invention proposes a Chinese lip-reading recognition method based on obtaining credible visemes. The specific algorithm process is as Figure 1 shown, including the following steps:
[0040] S1. Data collection and preprocessing: To obtain lip movement video data for characterization. After reading and frame-dividing the collected video data into a picture sequence, preprocessing is first performed, including lip region detection, cropping, size adjustment, etc., so that the pictures focus on the key lip information.
[0041] S2. Deep clustering: Perform deep clustering on the lip movement video data to obtain the number of viseme categories with a reasonable clustering distribution, the corresponding viseme categories and viseme library, and thus also obtain the original video's corresponding frame-by-frame image data with viseme category annotations. The viseme deep clustering module determines reasonable viseme categories and performs clustering based on the video data, deep neural network, and clustering algorithm, and realizes viseme information annotation for video data without viseme annotations.
[0042] S3. Cascade Chinese character recognition based on viseme intermediate representation.
[0043] In step S2, the preprocessed data is fed into a deep clustering network for unsupervised clustering based on a deep neural network to obtain reliable viseme clusters (i.e., viseme categories), and at the same time, the corresponding data distribution and annotation are formed. Here, the present invention does not preset the categories that visemes should have, but fully respects the distribution of the data itself and the drive of the deep neural network, so that the clustering results in solutions with high inter-cluster discrimination and high intra-cluster similarity, thus providing better data annotation and primitive information support for subsequent lip reading based on visemes. The present invention assumes that the results of the deep clustering network give K types of visemes Viseme, and each type of viseme (each picture) is labeled as V1, V2, V3, …, V K . Each video is correspondingly labeled as V i1 V i2 V i3 ……V in . Where n represents the total number of frames of a data instance. Each subscript is one of the categories from 1 to K. The purpose of viseme extraction is to obtain a set of image data with viseme category annotations, so that the data can be used to train the viseme encoding and decoding unit.
[0044] Specifically, for the deep clustering network, the structure is as Figure 2 shown.
[0045] The deep clustering network includes two parts: an autoencoder and a deep clustering module based on mutual information. First, the model is prompted to learn feature representations that are conducive to discrimination, improving the performance of subsequent clustering algorithms. At the same time, the deep clustering module based on mutual information can, in turn, guide the front-end neural network to learn better features. Both the autoencoder and the deep clustering module based on mutual information adopt classical algorithm structures. The loss L dc of the deep clustering network includes the reconstruction loss L n of the autoencoder and the loss L c of the mutual information clustering module. α and β are respectively set to 0.5.
[0046] L dc = αL n + βL c
[0047] Among them, the loss L dc of the deep clustering network, the reconstruction loss L n of the autoencoder, and the loss L c of the mutual information clustering module. The loss calculated at each place is backpropagated into the deep learning network in the subsequent process to constrain the network to update parameters.
[0048] Step S3 further includes:
[0049] S3.1 Preprocess the frame-by-frame image data with annotations after deep clustering. The annotated data is grayscaled and the image size is adjusted.
[0050] S3.2 Further input the preprocessed frame-by-frame image data into a spatio-temporal feature extractor, mainly a 3D convolutional neural network, to extract the spatio-temporal features after dimensionality reduction. The spatio-temporal features are the compact feature representations obtained by reducing the dimensionality of the image data through a spatio-temporal feature extractor (mainly 3D convolution). The spatio-temporal feature extractor can efficiently extract the deep representations in the lip video information, which plays an important role in subsequent modeling.
[0051] S3.3 Input the spatio-temporal features into the viseme sub-encoder-decoder unit to output the viseme sequence, which is constrained by the ctc loss. At the same time, the intermediate encoded representation output by the encoder in the viseme sub-encoder-decoder unit is reserved for the next step.
[0052] S3.4 Based on the intermediate encoded representation output by the viseme sub-encoder-decoder unit in S3.2, input it into the Chinese character sub-encoder-decoder unit to further output the Chinese character sequence, which is constrained by the cross-entropy loss.
[0053] The overall encoder-decoder consists of the above two encoder-decoder units. The viseme sequence prediction task is used as an auxiliary task, and the Chinese character prediction task is used as the main task. The two tasks are jointly constrained by the ctc (continuous time classification) loss and the ce (cross-entropy) loss.
[0054] The spatio-temporal features pass through the character sequence prediction module to obtain the character sequence. It can be seen that the processing of the annotated data goes through a cascaded encoder-decoder structure. First, the data is encoded and decoded into the viseme intermediate representation, and then the viseme intermediate representation is encoded and decoded into the character (Chinese character) sequence for output. The network structure at this stage is trained jointly in two levels.
[0055] Specifically, the spatio-temporal feature extractor uses a spatio-temporal convolutional neural network to extract the features in the temporal and spatial dimensions of the obtained lip image sequence. The structure is as Figure 3 shown. It consists of a 3D convolutional network layer (3DCNN), a batch normalization layer (BatchNormalization, BN), a ReLU activation function, a regularization dropout layer, and a max pooling layer.
[0056] Specifically, in order to further perform temporal modeling on the front-end visual features, in visual pixel sequence prediction and Chinese character sequence prediction, a Transformer encoder is used to better capture global information interaction. The Transformer encoder consists of 6 layers of classical Transformer Encoder structures. Each layer consists of a multi-head self-attention module and a feed-forward network module. In addition, residual connections and layer normalization (LayerNorm) are also used. The overall structure Figure 4 is shown as follows. The feed-forward network module mainly consists of two linear layers and a non-linear activation function in the middle.
[0057] The Transformer architecture can efficiently handle long-distance video text dependence problems such as lip reading recognition tasks.
[0058] The multi-head self-attention mechanism simultaneously pays attention to information from different representation subspaces at different positions. Through multiple independent attention calculations acting as an integration, it can prevent overfitting and more comprehensively capture various potential semantic associations in the sequence. For the lip reading recognition task, which is a video-semantic sequence prediction task, it can effectively conduct in-depth mining of the relevance between video and semantic information.
[0059] Among them, the multi-head self-attention module projects the query vector Q, key vector K, and value vector V through h different linear transformations, and finally splices the modeling results of different attention mechanisms with the softmax function:
[0060] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0061]
[0062] Therefore, the Encoder contains 6 identical attention layers, the multi-head attention sub-layer and the fully connected feed-forward sub-layer. Each sub-layer is added with a residual connection and normalization. Therefore, the output of the sub-layer can be expressed as:
[0063] sublayer output =LayerNorm(x+(SubLayer(x)))
[0064] The output of each time step of the Transformer encoder in the lip pose prediction is processed through a linear layer and a Softmax layer to obtain the recognition probability result of the lip pose. The recognition probability result of the lip pose is input into a Connectionist Temporal Classification (CTC) to obtain the classification result of the lip pose sequence. The CTC loss function enables the model to learn the automatic alignment between the lip image sequence and the lip pose sequence without maintaining a one-to-one mapping relationship in time.
[0065] Among them, the CTC loss is based on the conditional independence assumption, and the specific calculation method is as follows:
[0066] L ctc = -ln P(y|x)
[0067]
[0068] n is the length of the input image sequence x, y is the label of the actual lip pose sequence, and P πt represents the probability of the predicted lip pose label output after Softmax, and π t represents the lip pose predicted at the t-th frame. The mapping function F represents deleting adjacent duplicate lip pose labels (such as the sequence F(V1 V1 V3) = F(V1 V1 V3V3) = V1 V3), and F -1 (y) represents the set of all possible lip pose output sequences, that is, all possible CTC paths that map to y. During inference, the GreedySearch algorithm is used to decode the CTC path, and the lip pose with the highest probability is selected at each time step, laying a foundation for the subsequent prediction of the lip pose sequence to the character sequence.
[0069] Based on the lip pose sequence prediction output, character sequence decoding of an attention-based Sequence-to-Sequence (Seq2Seq) architecture is performed. In the character sequence prediction, the encoder is a 6-layer Transformer encoder, as Figure 5 shown. The decoder uses a classic 6-layer stacked Transformer decoder layer. Each layer includes a masked multi-head attention mechanism, a multi-head attention mechanism, and a feed-forward neural network. In addition, residual connections and layer normalization (LayerNorm) are also used.
[0070] The feedforward network module mainly consists of two linear layers and a non-linear activation function in the middle. The input of the encoder is the vector representation of the visual phonemes, and the output of the decoder is the character sequence. During the training process, it is optimized in parallel in an auto-regressive manner. The prediction result at each time step will be used as the input for the decoding of the next time step. During the decoding process, not only the results of historical predictions are utilized, but also the meaningful parts of the visual phoneme sequence selectively extracted through the attention mechanism are used. After passing through a linear layer and a Softmax layer, the hidden layer output representation of the decoder yields the probability distribution of the Chinese character sequence, and the cross-entropy loss is used as the model optimization objective, which is defined as follows:
[0071]
[0072] Among them, L represents the length of the predicted text, x v represents the input previous representation, c t represents the Chinese character predicted at time step t, and ĉ represents the predicted Chinese characters before the current time step. y t represents the true label of the Chinese characters before time step t.
[0073] Since it is fitted through a deep learning network, the actual process is to use the backpropagation of the loss function to constrain the update of each parameter of the network, so as to achieve the accurate effect of the model on the input data and finally output the probability distribution. Therefore, the network structure, the loss function, and the decoding method in the inference stage can cover the actual application process)
[0074] Based on the probability distribution of different output classes, according to the decoding algorithm in the inference stage, index the Chinese text units in the existing dictionary, and finally combine and splice them to form the most likely text sequence.
[0075] In the inference stage, the beam search method is used to decode the prediction results of the Chinese character sequence, that is, based on the conditional probability, several possible Chinese character sequences with the highest probability are selected at each time step for the input visual phoneme sequence as candidate results, and the best prediction result of the Chinese character sequence is obtained from them.
[0076] Among them, the conditional probability refers to the probability of event A occurring under the condition that another event B has already occurred. The conditional probability is expressed as: P(A|B).
[0077] In this task, it refers to the possible probability distribution of the next character under the existing character / string results. In order to improve the accuracy of both pinyin and Chinese character sequence prediction simultaneously, first optimize the deep clustering network model in the visual phoneme deep clustering extraction stage alone (based on L dcLoss function). Then in stage 2, the visual element prediction and Chinese character prediction are jointly optimized (based on L ctc and L ce The total loss L of the two can also be preliminarily defined as the weighted sum based on α=0.5.
[0078] L=αL ctc +(1-α)L ce
[0079] In terms of specific technical implementation:
[0080] In the deep clustering part, the autoencoder can be replaced by other neural network models such as variational autoencoder, generative adversarial network, twin network and graph neural network; the clustering based on mutual information can be replaced by deep clustering methods based on K-means, spectral clustering, subspace clustering, KL divergence deep clustering and Gaussian mixture model. Different neural network models and deep clustering methods can be used in combination.
[0081] For the clustering and use of visemes, a viseme is not limited to only one frame of image, but the concept of visemes of dynamic (multi-frame) images can be used to regard a video as a combination of multiple multi-frame visemes. The basis of the method mentioned above is that a video is a combination of n (maximum number of frames) single-frame visemes. Adjust the temporal granularity of the viseme primitive, and the other implementation methods are the same.
[0082] The encoders for viseme prediction and character prediction can use a Conformer structure based on a combination of CNN and Transformer to further extract local spatiotemporal features while paying attention to global information; or other encoder structures to optimize the effect.
[0083] The present invention establishes a method for extracting credible visemes driven by visemes for lip reading recognition based on visemes, which is independent of audio data support and phoneme and pinyin information supplementation. The method can determine the relatively accurate number of viseme primitive categories suitable for lip reading recognition, obtain a credible viseme primitive library with higher discrimination, better generalization, and more accurate and comprehensive description of pronunciation morphology, and can automatically annotate data. At the same time, lip reading recognition based on the credible data and credible viseme intermediate representation can reduce the cumulative error of recognition prediction, further improve the viseme-based lip reading recognition performance, and thus provide overall benefits for improving the generalization, robustness and accuracy of the lip reading recognition model.
[0084] Any process or method description depicted in the flowchart of the present invention or described otherwise herein may be construed as representing a module, segment, or portion of code that includes one or more executable instructions for implementing a specific logical function or process, which can be implemented in any computer-readable medium for an instruction execution system, apparatus, or device. The computer-readable medium can be any medium that contains, communicates, propagates, or transports a program for use by or in connection with an instruction execution system, apparatus, or device. This includes read-only memories, magnetic disks, or optical disks, etc.
[0085] In the description of this specification, the descriptions referring to the terms "embodiment", "example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Additionally, those skilled in the art can combine or combine the different embodiments or examples described in this specification and the features therein without contradiction.
[0086] Although the above has shown and described embodiments of the present invention, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can perform update operations such as changes, modifications, substitutions, and variations on the above embodiments within the scope of the present invention.
Claims
1. A Chinese lip-reading recognition method based on obtaining trusted visual elements, characterized in that The method includes the following steps: S1. Data acquisition and preprocessing: To obtain video data depicting lip movements; S2. Deep clustering: Perform deep clustering on the video data depicting lip movements to obtain the number of viseme categories in the clustering distribution, the corresponding viseme categories and viseme library, so as to obtain frame-by-frame image data with viseme category annotations corresponding to the video data depicting lip movements; S3. Cascade Chinese character sequence recognition based on viseme intermediate representation: Perform feature extraction based on the frame-by-frame image data with viseme category annotations to achieve cascade Chinese character sequence recognition with visemes as the intermediate representation.
2. The Chinese lip-reading recognition method based on obtaining a trusted visual isotope according to claim 1, wherein, Step S1 further includes: After reading and frame-dividing the collected video data into a picture sequence, first perform preprocessing to focus the picture on the key lip information; the preprocessing includes lip region detection, cropping, and image size adjustment.
3. The Chinese lip-reading recognition method based on obtaining a trusted visual element according to claim 2, wherein In step S2, the data is sent into a deep clustering network for unsupervised clustering based on a deep neural network to obtain credible viseme clusters, and at the same time form the corresponding data distribution and annotation.
4. The Chinese lip-reading recognition method based on obtaining credible visual elements according to claim 3, characterized in that, The deep clustering network includes an autoencoder and a deep clustering module based on mutual information, and the loss L of the deep clustering network dc includes the reconstruction loss L of the autoencoder n and the loss L of the mutual information clustering module c .
5. The Chinese lip-reading recognition method based on obtaining trusted visual elements according to claim 3, characterized in that, The deep clustering network includes a variational autoencoder, a generative adversarial network, a siamese network, and a graph neural network model; it also includes a deep clustering module based on a K-means module, a spectral clustering module, a subspace clustering module, a KL divergence deep clustering module, and / or a Gaussian mixture model. Different neural network models and deep clustering methods are used in combination.
6. The Chinese lip-reading recognition method based on obtaining a trusted visual isotope according to claim 1, characterized in that, Step S3 further includes: S3.1 Based on the frame-by-frame image data with annotations after deep clustering, perform preprocessing; S3.2 Further input the preprocessed frame-by-frame image data into a spatio-temporal feature extractor to extract the spatio-temporal features after dimensionality reduction through a 3D convolutional neural network; S3.3 Input the spatio-temporal features into a viseme sub-encoder-decoder unit to output a viseme sequence, which is constrained by the ctc loss; at the same time, the intermediate encoded representation output by the encoder in the viseme sub-encoder-decoder unit is reserved for the next step; S3.4 Based on the intermediate encoded representation output by the viseme sub-encoder-decoder unit in S3.2, input it into a Chinese character sub-encoder-decoder unit to further output a Chinese character sequence, which is constrained by the cross-entropy loss.
7. The Chinese lip-reading recognition method based on obtaining trusted visual elements according to claim 6, wherein In step S3.2, the spatio-temporal feature extractor uses a spatio-temporal convolutional neural network to extract features in the time and space dimensions for the obtained lip image sequence. The spatio-temporal feature extractor consists of a 3D convolutional network layer, a batch normalization layer, a ReLU activation function, a regularization dropout layer, and a max pooling layer.
8. The Chinese lip-reading recognition method based on obtaining a trusted visual isotope according to claim 6, characterized in that, In step S3, in the viseme sequence prediction and Chinese character sequence prediction, a Transformer encoder is used. The Transformer encoder contains 6 layers of Transformer Encoder structures, each layer consists of a multi-head self-attention module and a feed-forward network module, and also includes residual connections and layer normalization. The feed-forward network module consists of two linear layers and a non-linear activation function in the middle.
9. The Chinese lip-reading recognition method based on obtaining a trusted visual isotope according to claim 6, wherein, In step S3.3, based on the predicted output of the viseme sequence, character sequence decoding is performed using an attention-based sequence-to-sequence architecture. In the character sequence prediction, the encoder is a 6-layer Transformer encoder, and each layer includes a masked multi-head attention mechanism, a multi-head attention mechanism, and a feed-forward neural network. In addition, it also includes residual connections and layer normalization.
10. The Chinese lip-reading recognition method based on obtaining trusted visual elements according to claim 9, wherein, The input of the encoder is the vector representation of visemes, and the output of the decoder is the character sequence. During the training process, parallel optimization is performed in an autoregressive manner, and the prediction result at each time step will be used as the input for the next time step decoding. During the decoding process, the historical prediction output results are utilized, and at the same time, the attention mechanism is used to selectively extract the meaningful parts in the viseme sequence.
Citation Information
Cited By
GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion
CN120526486A