Gnn-lstm method for chinese lip speech classification based on node multi-association graph information fusion
Through the GNN-LSTM method based on node multi-association graph information fusion, the problems of spatiotemporal correlation and the collaborative relationship between initials and finals in Chinese lip reading recognition are solved, and the accuracy and robustness of lip reading recognition are improved.
Patent Information
- Application Number
- CN202511014893.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing Chinese lip reading recognition methods cannot effectively capture the spatiotemporal correlation between lip key points and the synergistic relationship between initials and finals, resulting in low recognition efficiency.
The GNN-LSTM method based on node multi-association graph information fusion is adopted. The lip key points are represented by adjacency graph, symmetry graph and upper and lower lip relationship graph. Combined with the long short-term memory network, high-dimensional spatiotemporal features are extracted, and the lip shape library is constructed by dividing the initials and finals into lip shape categories.
The accuracy and robustness of lip reading recognition are improved, and it can adapt to the complex changes and many-to-one mapping phenomena in the actual pronunciation process, reducing the impact of lip shape similarity and visual confusion on model performance.
Smart Images

Figure CN120526486B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lip reading and mouth shape analysis, and particularly relates to a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion. Background Art
[0002] Lip reading is a key research area in multimodal perception. By integrating visual analysis with semantic understanding, it develops a method for identifying speech content from lip movements. By extracting and analyzing the spatiotemporal characteristics of lip movements, this technology enables intelligent recognition of silent speech. This technology has significant application value in scenarios such as assisting communication for the hearing-impaired, security monitoring, intelligent human-computer interaction, and analyzing silent video content.
[0003] With the introduction of new architectures such as graph convolutional neural networks, lip reading models have evolved from a single image classification task to a complex spatiotemporal feature modeling task. Graph convolutional networks have the ability to efficiently model local topological structures and can accurately capture the dynamic changes of lip key points. Their application in lip reading recognition has significantly improved the recognition accuracy. However, existing Chinese lip reading recognition methods mostly extract lip key points from videos and combine them with speech or pinyin models for decoding. This has certain shortcomings. Traditional convolutional neural networks are difficult to effectively capture the spatiotemporal global correlation between lip key points, resulting in a disconnect between local motion features and overall semantics. In addition, existing classification strategies fail to fully consider the multi-level collaborative relationship between initials and finals. Traditional static lip shape libraries cannot dynamically adapt to the complex changes and many-to-one mapping phenomena in the actual pronunciation process. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that the Chinese lip reading recognition method in the existing technology cannot capture the spatiotemporal correlation between lip key points and does not consider the collaborative relationship between initials and finals, resulting in low recognition efficiency. Therefore, a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion is provided.
[0005] The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion includes the following steps: obtaining a face speaking dataset, wherein the face speaking dataset includes a face speaking video and a corresponding text label with a timestamp; extracting video frames from the face speaking video, cropping the video frames to obtain a lip area map, and extracting the coordinates of key points from the lip area map; the key points are used as nodes of the graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left and right symmetric key points of the symmetry graph are connected as edges to form another closed loop. Points are connected to form edges; the upper lip key points at corresponding positions of the upper and lower lip relationship graph are connected to the lower lip key points to form edges; a graph convolutional neural network is formed through the adjacency graph, the symmetry graph and the upper and lower lip relationship graph, the coordinates of the key points are input to output the global features of the three graph structures, and the global features are fused to form fusion features; each lip area graph is formed into a fusion feature through graph convolution, and multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features; the text labels are converted into pinyin, and are divided into lip shape categories based on initials and finals, and a lip shape library is established based on the lip shape categories and high-dimensional spatiotemporal features; the lip shape library inputs the face speaking video and outputs the lip reading text content.
[0006] Furthermore, the face speaking dataset is a CMLR dataset, which is converted into npy format and the mouth area is obtained through the dlib68-point face key point detector. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.
[0007] Furthermore, when extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are located at the inner midpoint of the upper lip, the right side of the inner upper lip, the inner right corner of the mouth, the right side of the inner lower lip, the middle point of the inner lower lip, the left side of the inner lower lip, the inner left corner of the mouth, and the left side of the inner upper lip; the key points of the outer lip are located at the outer left corner of the mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the middle point of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer right corner of the mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the middle point of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.
[0008] Furthermore, the adjacency graph, symmetry graph, and upper and lower lip relationship graph are expressed as:
[0009] The single-layer convolution operation of the adjacency graph is expressed as:
[0010] ;
[0011] The single-layer convolution operation of the symmetric graph is expressed as:
[0012] ;
[0013] The single-layer convolution operation of the upper and lower lip relationship graph is expressed as:
[0014] ;
[0015] in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of represents the trainable weight matrix, Represents the nonlinear activation function ReLU.
[0016] Furthermore, the global features are fused to form fusion features including:
[0017] Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score:
[0018] ;
[0019] in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector used to map global features to an attention score;
[0020] The attention scores are normalized into attention weights through the softmax function:
[0021] ;
[0022] ;
[0023] ;
[0024] in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively;
[0025] Weighted fusion of global features based on attention weights:
[0026] ;
[0027] in, Indicates fusion features.
[0028] Furthermore, each lip region map is subjected to graph convolution to form fused features. Multiple fused features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as:
[0029] ;
[0030] ;
[0031] in, represents the fusion feature per unit time, Represents high-dimensional spatiotemporal features.
[0032] Furthermore, the mouth shape categories are divided based on the initials and finals, including: classifying the initials b, p, m and f as labiodental sounds, classifying the initials z, c, s as tip-front sounds, classifying the initials d, t, n, l as tip-middle sounds, classifying the initials zh, ch, sh, r as tip-back sounds, classifying the initials j, q, x as palatal sounds, classifying the initials g, k, h as root sounds, and classifying the initials y, w as Y / X type sounds; classifying the finals a, o, e, i, u, v as simple finals, classifying the finals an, en, in, un, vn as front nasal finals, classifying the finals ang, eng, ing, ong as back nasal finals, classifying the finals ai, ei, ui, ao, ou, iu as front-sounding complex finals, and classifying the finals ie, ve, er as back-sounding complex finals.
[0033] Furthermore, the lip shape class and high-dimensional spatiotemporal features are used to establish a lip shape library, including: combining the initial consonant class and the final vowel class to form a lip shape class, and constructing a lip shape library based on the pinyin corresponding to the text label corresponding to the high-dimensional spatiotemporal features.
[0034] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned lip reading classification method when executing the computer program.
[0035] A computer-readable storage medium stores a computer program, which implements the steps of the above-mentioned lip reading classification method when executed by a processor.
[0036] Beneficial effect: The present invention discloses a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, which represents lip key points through three structures: adjacency graph, symmetry graph and upper and lower lip relationship graph, extracts high-dimensional spatiotemporal features under the synergistic effect of GNN and LSTM, effectively captures the spatiotemporal global correlation between lip key points, and divides them into lip shape categories based on initials and finals. The multi-level collaborative relationship between initials and finals is considered to reduce the influence of lip shape similarity and visual confusion on model performance, and uses the above-mentioned extracted high-dimensional spatiotemporal features to establish a lip shape library, thereby performing a more discriminative mapping and induction between lip morphology and corresponding pinyin, reducing the influence of lip shape similarity and visual confusion on model performance, and being able to adapt to the complex changes and many-to-one mapping phenomenon in the actual pronunciation process, effectively enhancing the accuracy and robustness of subsequent lip shape classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0038] Figure 1 The figure is a schematic block diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0039] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0040] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0041] Reference Figure 1 As shown, this embodiment provides a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, including the following steps:
[0042] Step S1: Obtain a face speech dataset, which includes a face speech video and corresponding time-stamped text labels. Extract video frames from the face speech video, crop the video frames to obtain a lip region map, and extract the coordinates of key points from the lip region map. In this embodiment, the time-stamped text label is represented by the time corresponding to a Chinese character segment in the video, and the text label is the Chinese character. In steps S2 to S4 of this embodiment, the video segment corresponding to each text label is processed separately.
[0043] Step S2: The key points are used as nodes of a graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip in the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left-right symmetrical key points in the symmetry graph are connected to form edges; the upper lip key points and the lower lip key points at corresponding positions in the upper and lower lip relationship graph are connected to form edges; a graph convolutional neural network is formed through the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, and the coordinates of the key points are input to output the global features of the three graph structures;
[0044] Step S3: Global features are fused to form fused features; each lip region map is subjected to graph convolution to form fused features, and multiple fused features are processed through a long short-term memory network to output high-dimensional spatiotemporal features;
[0045] Step S4: convert the text labels into pinyin, divide them into lip shape categories based on initials and finals, and build a lip shape library based on the lip shape categories and high-dimensional spatiotemporal features;
[0046] The lip shape library inputs the video of the human face speaking and outputs the lip reading text content.
[0047] Specifically, in step S1, the face speaking dataset is a CMLR dataset. The mp4 format in the dataset is converted into npy format and the mouth area is obtained through the dlib68-point face key point detector. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.
[0048] The Chinese Mandarin Lip Reading (CMLR) dataset is derived from Chinese News Broadcast videos and is highly standardized. It provides a relatively unified and standardized sample source for lip reading recognition research. It contains 102,076 sentences uttered by 11 anchors, each containing up to 29 Chinese characters and excluding English letters, Arabic numerals, and rare punctuation marks. This ensures the data's linguistic purity and facilitates the extraction and analysis of Chinese lip reading features.
[0049] As a further improvement to this embodiment, the video frame sequence is time-aligned according to the timestamps provided by the dataset to ensure that the mouth movement in each time period corresponds to the correct pronunciation of Chinese characters. At the same time, linear interpolation is used to fill or truncate the number of frames to meet the fixed length requirement.
[0050] In this embodiment, when extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are respectively located at the midpoint of the inner side of the upper lip, the right side of the inner side of the upper lip, the inner side of the right corner of the mouth, the right side of the inner side of the lower lip, the midpoint of the inner side of the lower lip, the left side of the inner side of the lower lip, the inner side of the left corner of the mouth, and the left side of the inner side of the upper lip; the key points of the outer lip are respectively located at the outer side of the left corner of the mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the midpoint of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer side of the right corner of the mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the midpoint of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.
[0051] In step S2, a lip landmark graph (LLG) is constructed, and an adjacency matrix is set, which is expressed as:
[0052] ;
[0053] in, Represents the adjacency matrix.
[0054] In this embodiment, the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop. Relying on the natural topological relationship of the lip key points, the local topological characteristics of the lip contour can be reflected; the left and right symmetrical key points of the symmetry graph are connected to form edges. Considering the physiological structure characteristics of the left and right symmetry of the lips, the spatial symmetry relationship between the left and right sides of the lips can be highlighted, helping the network to capture the symmetry characteristics of the lips; the upper lip key points and the lower lip key points at the corresponding positions of the upper and lower lip relationship graph are connected to form edges. According to the physiological characteristics of the opening and closing of the lips, the dynamic relationship between the upper and lower lips reflected by the opening and closing of the lips and the changes in the movement trajectory can be captured.
[0055] Specifically, the adjacency graph, symmetry graph, and upper and lower lip relationship graph are represented as:
[0056] The single-layer convolution operation of the adjacency graph is expressed as:
[0057] ;
[0058] The single-layer convolution operation of the symmetric graph is expressed as:
[0059] ;
[0060] The single-layer convolution operation of the upper and lower lip relationship graph is expressed as:
[0061] ;
[0062] in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of represents the trainable weight matrix, Represents the nonlinear activation function ReLU.
[0063] In step S3 of this embodiment, fusing global features to form fused features includes:
[0064] Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score:
[0065] ;
[0066] in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector used to map the global feature into an attention score. In this embodiment, it takes the form of an inner product.
[0067] The attention scores are normalized into attention weights through the softmax function:
[0068] ;
[0069] ;
[0070] ;
[0071] in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively;
[0072] Weighted fusion of global features based on attention weights:
[0073] ;
[0074] in, Indicates fusion features.
[0075] In this embodiment, each lip region map is fused through graph convolution to form a fusion feature. Multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as:
[0076] ;
[0077] ;
[0078] in, represents the fusion feature per unit time, specifically represents the fusion feature of one frame in this embodiment, Represents high-dimensional spatiotemporal features, which provide multi-dimensional spatiotemporal feature support for the construction of the lip shape library.
[0079] In step S4 of this embodiment, the mouth shape categories are divided based on the initials and finals, including: classifying the initials b, p, m, and f as labiodental sounds, classifying the initials z, c, and s as tip-front sounds, classifying the initials d, t, n, and l as tip-medial sounds, classifying the initials zh, ch, sh, and r as tip-back sounds, classifying the initials j, q, and x as palatal sounds, classifying the initials g, k, and h as root sounds, and classifying the initials y and w as Y / X sounds; classifying the finals a, o, e, i, u, and v as simple finals, classifying the finals an, en, in, un, and vn as front nasal finals, classifying the finals ang, eng, ing, and ong as back nasal finals, classifying the finals ai, ei, ui, ao, ou, and iu as front-sounding complex finals, and classifying the finals ie, ve, and er as back-sounding complex finals.
[0080] Refer to Table 1 and Table 2:
[0081] Table 1: Classification of initial consonants
[0082] Initial consonant type List of initial consonants Bilabial b,p,m labiodental f frontal consonants z,c,s Tip of tongue mid-range d,t,n,l back tongue sound zh,ch,sh,r tongue-tip sounds j,q,x root of tongue g,k,h Y / X type sound y,w
[0083] Table 2 Classification of finals
[0084] Final type Finals list Single vowel a,o,e,i,u,v front nasal finals an,en,in,un,vn back nasal finals ang,eng,ing,ong Front-sounding complex vowels ai,ei,ui,ao,ou,iu Back-sounding complex vowels ie,ve,er
[0085] As a further improvement of the present embodiment, considering the situation of compound vowels, in order to avoid the misclassification caused by this situation, an alignment method based on character splitting is adopted. For example, "uan" can be decomposed into "u" and "an", and "an" is a complete vowel. When aligning, the "an" part of "uan" and "an" is classified into the same category, and the "u" and "a" parts are respectively classified into different categories, ensuring that different combinations of parts of the same vowel will not be mistakenly classified into the same category. Through this alignment method, it is guaranteed that pinyin with similar pronunciation mouth shapes can be accurately distinguished, eliminating the confusion that compound vowels may cause. This classification and alignment method can accurately describe the mouth shapes of different pinyin.
[0086] In this embodiment, a lip shape library is established based on lip shape classes and high-dimensional spatiotemporal features, including: combining initial consonant classes and final vowel classes to form lip shape classes, and constructing a lip shape library based on the lip shape classes corresponding to the pinyin of the text tags corresponding to the high-dimensional spatiotemporal features.
[0087] Specifically, the lip shape embedding vectors extracted through a graph convolutional neural network and a long short-term memory network are grouped and stored according to the "initial consonant-final vowel" category to form a multi-label lip shape library. Each category contains: high-dimensional spatiotemporal features; related Chinese characters: a list of corresponding Chinese characters; and a data structure that records the category name, feature center (mean vector), and feature variance.
[0088] As a further improvement to this embodiment, for each category of high-dimensional spatiotemporal features, the mean and variance are calculated to optimize lip matching in the subsequent model, helping to adjust the classification boundaries and improve recognition accuracy:
[0089] ;
[0090] ;
[0091] in, represents the mean, represents the variance, Represents high-dimensional spatiotemporal features.
[0092] In this embodiment, the lip shape library inputs a video of a human face speaking and outputs lip reading text content, which specifically includes the following method steps:
[0093] Step S5: Video feature extraction and time alignment.
[0094] In this embodiment, for the video data to be recognized, the start and end timestamps corresponding to each Chinese character are first parsed from the video's corresponding tag file (.txt). This is then used to delineate the precise time period corresponding to each Chinese character in the video. Then, for each delineated time period, the coordinates of key mouth points are extracted according to the methods described in steps S1 through S4. Three topological structures are constructed: an adjacency graph, a symmetry graph, and an upper-lower-lip relationship graph. A graph convolutional network (GNN) and a long-short-term memory network (LSTM) are then combined with an attention mechanism to extract the corresponding high-dimensional spatiotemporal features. These high-dimensional spatiotemporal features effectively capture the spatial structural and temporal dynamic characteristics of the current character, providing a stable and robust basic feature representation for subsequent lip shape library matching.
[0095] Step S6: matching the lip shape library and extracting candidate pinyin.
[0096] The high-dimensional spatiotemporal features of each Chinese character period extracted are matched with the features in the constructed lip shape library. In this embodiment, the cosine similarity is used to calculate the similarity between the features of the lip shape to be recognized and the features in the library, and the three closest lip shape categories are selected as the candidate results, so as to determine the candidate pinyin set corresponding to the current Chinese character. Specifically, for the high-dimensional spatiotemporal features of each Chinese character period to be recognized, the multiple feature samples pre-stored in the lip shape library are sequentially subjected to similarity calculation, and the category set with the highest similarity is selected, so as to obtain several candidate pinyins for each Chinese character.
[0097] Step 7: Pinyin sequence combination based on MoE expert mixture model.
[0098] Since a single lip shape category often corresponds to multiple possible pinyin combinations, in order to accurately recognize the Chinese character sequence, the expert mixture model (Mixture-of-Experts, MoE) is used in this embodiment to optimize and combine the candidate pinyin sequence. The MoE model includes several expert networks, each of which focuses on processing a specific pronunciation mode or context relationship, and dynamically weights and fuses the outputs of these expert networks through a gating network. Taking the candidate pinyin sequence as input, the MoE model evaluates the rationality of each candidate pinyin combination according to the pinyin context relationship and language characteristics, and generates a series of score values of the pinyin sequence, so as to select the optimal pinyin sequence.
[0099] Step 8: Translation of pinyin sequence to Chinese text.
[0100] After obtaining the optimal pinyin sequence, a pre-trained pinyin-Chinese character translation language model (sequence-to-sequence model based on Transformer) is further used in this embodiment to convert the pinyin sequence into the final Chinese text output. The model has fully learned the mapping relationship between pinyin and Chinese characters in the training stage, and has captured the syntactic and semantic rules inherent in Chinese language, so as to accurately and efficiently convert the pinyin sequence to the Chinese character sequence. The final output Chinese text is automatically checked and post-processed, which can effectively reduce the recognition error and improve the overall accuracy and robustness of Chinese lip reading recognition.
[0101] In some other embodiments of this embodiment, the Python library Pinyin2Hanzi is used to convert the optimal pinyin sequence into the final Chinese text output.
[0102] This embodiment also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the lip language classification method when executing the computer program.
[0103] This embodiment further provides a computer-readable storage medium storing a computer program, which implements the steps of the above-mentioned lip reading classification method when executed by a processor.
[0104] This embodiment provides a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, which represents lip key points through three structures: adjacency graph, symmetry graph and upper and lower lip relationship graph. Under the synergistic effect of GNN and LSTM, high-dimensional spatiotemporal features are extracted to effectively capture the spatiotemporal global correlation between lip key points. The method is divided into lip shape categories based on initials and finals, and the multi-level collaborative relationship between initials and finals is considered to reduce the impact of lip shape similarity and visual confusion on model performance. The high-dimensional spatiotemporal features extracted above are used to establish a lip shape library, thereby performing a more discriminative mapping and induction between lip morphology and corresponding pinyin, reducing the impact of lip shape similarity and visual confusion on model performance, and being able to adapt to the complex changes and many-to-one mapping phenomenon in the actual pronunciation process, effectively enhancing the accuracy and robustness of subsequent lip shape classification.
[0105] In this invention, GNN stands for Graph Neural Networks, and Graph Convolutional Networks (GCN) is an implementation of Graph Neural Networks.
[0106] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0107] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion is characterized by: The following steps are involved: Obtain a face speaking dataset, which includes a face speaking video and corresponding text labels with timestamps; extract video frames from the face speaking video, crop the video frames to obtain a lip area map, and extract the coordinates of key points from the lip area map; the key points are used as nodes of the graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left-right symmetrical key points of the symmetry graph are connected to form edges; the upper lip key points and the lower lip key points at corresponding positions in the upper and lower lip relationship graph are connected to form edges; a graph convolutional neural network is formed by the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, the coordinates of the key points are input to output the global features of the three graph structures, and the global features are fused to form fusion features; each lip area map is formed into a fusion feature through graph convolution, and multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features; the text labels are converted into pinyin, and the initials and finals are divided into lip shape categories, and a lip shape library is established based on the lip shape categories and high-dimensional spatiotemporal features; The lip shape library inputs the face speaking video and outputs the lip reading text content; The mouth shape class includes: classifying the initial consonants b, p, m and f as labiodental sounds, classifying the initial consonants z, c, s as tip-front sounds, classifying the initial consonants d, t, n, l as tip-median sounds, classifying the initial consonants zh, ch, sh, r as tip-back sounds, classifying the initial consonants j, q, x as palatal sounds, classifying the initial consonants g, k, h as root sounds, and classifying the initial consonants y, w as Y / X class sounds; classifying the finals a, o, e, i, u, v as simple finals, classifying the finals an, en, in, un, vn as front nasal finals, classifying the finals ang, eng, ing, ong as back nasal finals, classifying the finals ai, ei, ui, ao, ou, iu as front-sounding complex finals, and classifying the finals ie, ve, er as back-sounding complex finals.
2. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The face speech dataset is a CMLR dataset, which is converted into npy format and uses the dlib 68-point face key point detector to obtain the mouth area. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.
3. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: When extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are located at the inner midpoint of the upper lip, the right side of the inner upper lip, the inner right corner of the right mouth, the right side of the inner lower lip, the middle point of the inner lower lip, the left side of the inner lower lip, the inner left corner of the left mouth, and the left side of the inner upper lip; the key points of the outer lip are located at the outer left corner of the left mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the middle point of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer right corner of the right mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the middle point of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.
4. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The single-layer convolution operation of the adjacency graph is expressed as: ; The single-layer convolution operation of the symmetric graph is expressed as: ; The single-layer convolution operation of the upper and lower lip relationship graph is expressed as: ; in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of Represents the nonlinear activation function ReLU.
5. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 4 is characterized in that: The fusion features formed by global feature fusion include: Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score: ; in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector for mapping global features into an attention score; The attention scores are normalized into attention weights through the softmax function: ; ; ; in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively; Weighted fusion of global features based on attention weights: ; in, Indicates fusion features.
6. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 5 is characterized in that: Each lip region map is convolved to form a fusion feature. Multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as: ; 。 7. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The lip shape class and high-dimensional spatiotemporal features are used to establish a lip shape library, including: combining the initial consonant class and the final vowel class to form a lip shape class, and constructing a lip shape library based on the pinyin corresponding to the lip shape class of the text label corresponding to the high-dimensional spatiotemporal features.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the lip reading classification method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the lip reading classification method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Lip reading method based on adaptive semantic space-time diagram convolutional network
CN111259875A
Lip language recognition method and system based on multi-scale space-time convolution
CN115830688A