GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion

Through the GNN-LSTM method based on node multi-correlation graph information fusion, the problem of spatiotemporal correlation and initial vowel synergy in Chinese lip recognition is solved, high-dimensional spatiotemporal features are extracted and lip library is constructed, which improves the accuracy and robustness of lip recognition.

CN120526486AActive Publication Date: 2025-08-22XIANGJIANG LAB
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511014893.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-08-22
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

The existing Chinese lip recognition method cannot effectively capture the spatial and temporal correlation between lip key points and the synergistic relationship between initial consonants and final vowels, resulting in low recognition efficiency.

Method used

The GNN-LSTM method based on node multi-correlation graph information fusion is adopted to represent the key points of the lip through the adjacency graph, symmetric graph and upper and lower lip relationship graph. Combined with long and short-term memory networks, high-dimensional spatiotemporal features are extracted, and the lip library is constructed based on the initial consonant and final vowel to be divided into lip type categories.

Benefits of technology

It improves the accuracy and robustness of lip recognition, can adapt to complex pronunciation changes and many-to-one mapping phenomena, and reduces the impact of lip similarity and visual confusion on model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526486A_ABST
    Figure CN120526486A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of lip language mouth shape analysis, and discloses a GNN-LSTM Chinese lip language classification method based on node multi-association graph information fusion, lip key points are represented through three structures of an adjacent graph, a symmetric graph and an upper and lower lip relation graph, high-dimensional spatial-temporal features are extracted under the synergistic effect of a graph convolutional neural network and a long-short-term memory network, and the lip language mouth shape is classified. The method comprises the following steps: effectively capturing space-time global relevance between lip key points, dividing initial consonants and vowels into mouth shape classes, considering a multi-level collaborative relationship between the initial consonants and the vowels, reducing the influence of mouth shape similarity and visual confusion on model performance, and establishing a mouth shape library by utilizing extracted high-dimensional space-time characteristics, so as to improve the mouth shape similarity and visual confusion. Therefore, mapping induction with higher discriminative ability is carried out on the lip shape and the corresponding pinyin, the influence of mouth shape similarity and visual confusion on the model performance is reduced, the method can adapt to complex changes and many-to-one mapping phenomena in the actual pronunciation process, and the accuracy and robustness of subsequent mouth shape classification are effectively enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lip reading and mouth shape analysis, and particularly relates to a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion. Background Art

[0002] Lip reading is a key research area in multimodal perception. By integrating visual analysis with semantic understanding, it develops a method for identifying speech content from lip movements. By extracting and analyzing the spatiotemporal characteristics of lip movements, this technology enables intelligent recognition of silent speech. This technology has significant application value in scenarios such as assisting communication for the hearing-impaired, security monitoring, intelligent human-computer interaction, and analyzing silent video content.

[0003] With the introduction of new architectures such as graph convolutional neural networks, lip reading models have evolved from a single image classification task to a complex spatiotemporal feature modeling task. Graph convolutional networks have the ability to efficiently model local topological structures and can accurately capture the dynamic changes of lip key points. Their application in lip reading recognition has significantly improved the recognition accuracy. However, existing Chinese lip reading recognition methods mostly extract lip key points from videos and combine them with speech or pinyin models for decoding. This has certain shortcomings. Traditional convolutional neural networks are difficult to effectively capture the spatiotemporal global correlation between lip key points, resulting in a disconnect between local motion features and overall semantics. In addition, existing classification strategies fail to fully consider the multi-level collaborative relationship between initials and finals. Traditional static lip shape libraries cannot dynamically adapt to the complex changes and many-to-one mapping phenomena in the actual pronunciation process. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that the Chinese lip reading recognition method in the existing technology cannot capture the spatiotemporal correlation between lip key points and does not consider the collaborative relationship between initials and finals, resulting in low recognition efficiency. Therefore, a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion is provided.

[0005] The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion includes the following steps: obtaining a face speaking dataset, wherein the face speaking dataset includes a face speaking video and a corresponding text label with a timestamp; extracting video frames from the face speaking video, cropping the video frames to obtain a lip area map, and extracting the coordinates of key points from the lip area map; the key points are used as nodes of the graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left and right symmetric key points of the symmetry graph are connected as edges to form another closed loop. Points are connected to form edges; the upper lip key points at corresponding positions of the upper and lower lip relationship graph are connected to the lower lip key points to form edges; a graph convolutional neural network is formed through the adjacency graph, the symmetry graph and the upper and lower lip relationship graph, the coordinates of the key points are input to output the global features of the three graph structures, and the global features are fused to form fusion features; each lip area graph is formed into a fusion feature through graph convolution, and multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features; the text labels are converted into pinyin, and are divided into lip shape categories based on initials and finals, and a lip shape library is established based on the lip shape categories and high-dimensional spatiotemporal features; the lip shape library inputs the face speaking video and outputs the lip reading text content.

[0006] Furthermore, the face speaking dataset is a CMLR dataset, which is converted into npy format and the mouth area is obtained through the dlib68-point face key point detector. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.

[0007] Furthermore, when extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are located at the inner midpoint of the upper lip, the right side of the inner upper lip, the inner right corner of the mouth, the right side of the inner lower lip, the middle point of the inner lower lip, the left side of the inner lower lip, the inner left corner of the mouth, and the left side of the inner upper lip; the key points of the outer lip are located at the outer left corner of the mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the middle point of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer right corner of the mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the middle point of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.

[0008] Furthermore, the adjacency graph, symmetry graph, and upper and lower lip relationship graph are expressed as: The single-layer convolution operation of the adjacency graph is expressed as: ; The single-layer convolution operation of the symmetric graph is expressed as: ; The single-layer convolution operation of the upper and lower lip relationship graph is expressed as: ; in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of represents the trainable weight matrix, Represents the nonlinear activation function ReLU.

[0009] Furthermore, the fusion features formed by global feature fusion include: Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score: ; in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector for mapping global features into an attention score; The attention scores are normalized into attention weights through the softmax function: ; ; ; in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively; Weighted fusion of global features based on attention weights: ; in, Indicates fusion features.

[0010] Furthermore, each lip region map is subjected to graph convolution to form fused features. Multiple fused features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as: ; ; in, represents the fusion feature per unit time, Represents high-dimensional spatiotemporal features.

[0011] Furthermore, the mouth shape categories are divided based on the initials and finals, including: classifying the initials b, p, m and f as labiodental sounds, classifying the initials z, c, s as tip-front sounds, classifying the initials d, t, n, l as tip-middle sounds, classifying the initials zh, ch, sh, r as tip-back sounds, classifying the initials j, q, x as palatal sounds, classifying the initials g, k, h as root sounds, and classifying the initials y, w as Y / X type sounds; classifying the finals a, o, e, i, u, v as simple finals, classifying the finals an, en, in, un, vn as front nasal finals, classifying the finals ang, eng, ing, ong as back nasal finals, classifying the finals ai, ei, ui, ao, ou, iu as front-sounding complex finals, and classifying the finals ie, ve, er as back-sounding complex finals.

[0012] Furthermore, the lip shape class and high-dimensional spatiotemporal features are used to establish a lip shape library, including: combining the initial consonant class and the final vowel class to form a lip shape class, and constructing a lip shape library based on the pinyin corresponding to the text label corresponding to the high-dimensional spatiotemporal features.

[0013] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned lip reading classification method when executing the computer program.

[0014] A computer-readable storage medium stores a computer program, which implements the steps of the above-mentioned lip reading classification method when executed by a processor.

[0015] Beneficial effect: The present invention discloses a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, which represents lip key points through three structures: adjacency graph, symmetry graph and upper and lower lip relationship graph, extracts high-dimensional spatiotemporal features under the synergistic effect of GNN and LSTM, effectively captures the spatiotemporal global correlation between lip key points, and divides them into lip shape categories based on initials and finals. The multi-level collaborative relationship between initials and finals is considered to reduce the influence of lip shape similarity and visual confusion on model performance, and uses the above-mentioned extracted high-dimensional spatiotemporal features to establish a lip shape library, thereby performing a more discriminative mapping and induction between lip morphology and corresponding pinyin, reducing the influence of lip shape similarity and visual confusion on model performance, and being able to adapt to the complex changes and many-to-one mapping phenomenon in the actual pronunciation process, effectively enhancing the accuracy and robustness of subsequent lip shape classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 The figure is a schematic block diagram of the method flow of the present invention. DETAILED DESCRIPTION

[0018] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0019] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0020] Reference Figure 1 As shown, this embodiment provides a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, including the following steps: Step S1: Obtain a face speech dataset, which includes a face speech video and corresponding time-stamped text labels. Extract video frames from the face speech video, crop the video frames to obtain a lip region map, and extract the coordinates of key points from the lip region map. In this embodiment, the time-stamped text label is represented by the time corresponding to a Chinese character segment in the video, and the text label is the Chinese character. In steps S2 to S4 of this embodiment, the video segment corresponding to each text label is processed separately.

[0021] Step S2: The key points are used as nodes of a graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip in the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left-right symmetrical key points in the symmetry graph are connected to form edges; the upper lip key points and the lower lip key points at corresponding positions in the upper and lower lip relationship graph are connected to form edges; a graph convolutional neural network is formed through the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, and the coordinates of the key points are input to output the global features of the three graph structures; Step S3: Global features are fused to form fused features; each lip region map is subjected to graph convolution to form fused features, and multiple fused features are processed through a long short-term memory network to output high-dimensional spatiotemporal features; Step S4: convert the text labels into pinyin, divide them into lip shape categories based on initials and finals, and build a lip shape library based on the lip shape categories and high-dimensional spatiotemporal features; The lip shape library inputs the video of the human face speaking and outputs the lip reading text content.

[0022] Specifically, in step S1, the face speaking dataset is a CMLR dataset. The mp4 format in the dataset is converted into npy format and the mouth area is obtained through the dlib68-point face key point detector. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.

[0023] CMLR (Chinese Mandarin Lip Reading) is highly standardized and normative, providing a relatively unified and standardized sample source for lip reading recognition research. It contains a total of 102,076 sentences uttered by 11 hosts. Each sentence contains a maximum of 29 Chinese characters and does not contain English letters, Arabic numerals, and rare punctuation marks. This ensures the purity of the data in terms of language type and is conducive to focusing on the extraction and analysis of Chinese lip reading features.

[0024] As a further improvement to this embodiment, the video frame sequence is time-aligned according to the timestamps provided by the dataset to ensure that the mouth movement in each time period corresponds to the correct pronunciation of Chinese characters. At the same time, linear interpolation is used to fill or truncate the number of frames to meet the fixed length requirement.

[0025] In this embodiment, when extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are respectively located at the midpoint of the inner side of the upper lip, the right side of the inner side of the upper lip, the inner side of the right corner of the mouth, the right side of the inner side of the lower lip, the midpoint of the inner side of the lower lip, the left side of the inner side of the lower lip, the inner side of the left corner of the mouth, and the left side of the inner side of the upper lip; the key points of the outer lip are respectively located at the outer side of the left corner of the mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the midpoint of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer side of the right corner of the mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the midpoint of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.

[0026] In step S2, a lip landmark graph (LLG) is constructed, and an adjacency matrix is ​​set, which is expressed as: ; in, Represents the adjacency matrix.

[0027] In this embodiment, the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop. Relying on the natural topological relationship of the lip key points, the local topological characteristics of the lip contour can be reflected; the left and right symmetrical key points of the symmetry graph are connected to form edges. Considering the physiological structure characteristics of the left and right symmetry of the lips, the spatial symmetry relationship between the left and right sides of the lips can be highlighted, helping the network to capture the symmetry characteristics of the lips; the upper lip key points and the lower lip key points at the corresponding positions of the upper and lower lip relationship graph are connected to form edges. According to the physiological characteristics of the opening and closing of the lips, the dynamic relationship between the upper and lower lips reflected by the opening and closing of the lips and the changes in the movement trajectory can be captured.

[0028] Specifically, the adjacency graph, symmetry graph, and upper and lower lip relationship graph are represented as: The single-layer convolution operation of the adjacency graph is expressed as: ; The single-layer convolution operation of the symmetric graph is expressed as: ; The single-layer convolution operation of the upper and lower lip relationship graph is expressed as: ; in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of represents the trainable weight matrix, Represents the nonlinear activation function ReLU.

[0029] In step S3 of this embodiment, fusing global features to form fused features includes: Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score: ; in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector used to map the global feature into an attention score. In this embodiment, it takes the form of an inner product. The attention scores are normalized into attention weights through the softmax function: ; ; ; in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively; Weighted fusion of global features based on attention weights: ; in, Indicates fusion features.

[0030] In this embodiment, each lip region map is fused through graph convolution to form a fusion feature. Multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as: ; ; in, represents the fusion feature per unit time, specifically represents the fusion feature of one frame in this embodiment, Represents high-dimensional spatiotemporal features, which provide multi-dimensional spatiotemporal feature support for the construction of the lip shape library.

[0031] In step S4 of this embodiment, the mouth shape categories are divided based on the initials and finals, including: classifying the initials b, p, m, and f as labiodental sounds, classifying the initials z, c, and s as tip-front sounds, classifying the initials d, t, n, and l as tip-medial sounds, classifying the initials zh, ch, sh, and r as tip-back sounds, classifying the initials j, q, and x as palatal sounds, classifying the initials g, k, and h as root sounds, and classifying the initials y and w as Y / X sounds; classifying the finals a, o, e, i, u, and v as simple finals, classifying the finals an, en, in, un, and vn as front nasal finals, classifying the finals ang, eng, ing, and ong as back nasal finals, classifying the finals ai, ei, ui, ao, ou, and iu as front-sounding complex finals, and classifying the finals ie, ve, and er as back-sounding complex finals.

[0032] Refer to Table 1 and Table 2: Table 1: Classification of initial consonants Initial consonant type List of initial consonants Bilabial b,p,m labiodental f frontal consonants z,c,s Tip of tongue mid-range d,t,n,l back tongue sound zh,ch,sh,r tongue-tip sounds j,q,x root of tongue g,k,h Y / X type sound y,w Table 2 Classification of finals Final type Finals list Single vowel a,o,e,i,u,v front nasal finals an,en,in,un,vn back nasal finals ang,eng,ing,ong Front-sounding complex vowels ai,ei,ui,ao,ou,iu Back-sounding complex vowels ie,ve,er As a further improvement of the present embodiment, considering the situation of compound vowels, in order to avoid the misclassification caused by this situation, an alignment method based on character splitting is adopted. For example, "uan" can be decomposed into "u" and "an", and "an" is a complete vowel. When aligning, the "an" part of "uan" and "an" is classified into the same category, and the "u" and "a" parts are respectively classified into different categories, ensuring that different combinations of parts of the same vowel will not be mistakenly classified into the same category. Through this alignment method, it is guaranteed that pinyin with similar pronunciation mouth shapes can be accurately distinguished, eliminating the confusion that compound vowels may cause. This classification and alignment method can accurately describe the mouth shapes of different pinyin.

[0033] In this embodiment, a lip shape library is established based on lip shape classes and high-dimensional spatiotemporal features, including: combining initial consonant classes and final vowel classes to form lip shape classes, and constructing a lip shape library based on the lip shape classes corresponding to the pinyin of the text tags corresponding to the high-dimensional spatiotemporal features.

[0034] Specifically, the lip shape embedding vectors extracted through a graph convolutional neural network and a long short-term memory network are grouped and stored according to the "initial consonant-final vowel" category to form a multi-label lip shape library. Each category contains: high-dimensional spatiotemporal features; related Chinese characters: a list of corresponding Chinese characters; and a data structure that records the category name, feature center (mean vector), and feature variance.

[0035] As a further improvement to this embodiment, for each category of high-dimensional spatiotemporal features, the mean and variance are calculated to optimize lip matching in the subsequent model, helping to adjust the classification boundaries and improve recognition accuracy: ; ; in, represents the mean, represents the variance, Represents high-dimensional spatiotemporal features.

[0036] In this embodiment, the lip shape library inputs a video of a human face speaking and outputs lip reading text content, which specifically includes the following method steps: Step S5: Video feature extraction and time alignment.

[0037] In this embodiment, for the video data to be recognized, the start and end timestamps corresponding to each Chinese character are first parsed from the video's corresponding tag file (.txt). This is then used to delineate the precise time period corresponding to each Chinese character in the video. Then, for each delineated time period, the coordinates of key mouth points are extracted according to the methods described in steps S1 through S4. Three topological structures are constructed: an adjacency graph, a symmetry graph, and an upper-lower-lip relationship graph. A graph convolutional network (GNN) and a long-short-term memory network (LSTM) are then combined with an attention mechanism to extract the corresponding high-dimensional spatiotemporal features. These high-dimensional spatiotemporal features effectively capture the spatial structural and temporal dynamic characteristics of the current character, providing a stable and robust basic feature representation for subsequent lip shape library matching.

[0038] Step S6: matching the lip shape library and extracting candidate pinyin.

[0039] The extracted high-dimensional spatiotemporal features of each Chinese character time period are used to perform similarity matching with the features in the constructed lip shape library. In this embodiment, cosine similarity is used to calculate the similarity between the lip shape features to be identified and the features in the library, and the three closest lip shape categories are selected as candidate results to determine the candidate pinyin set corresponding to the current Chinese character. Specifically, for the high-dimensional spatiotemporal features of each Chinese character time period to be identified, the multiple feature samples pre-stored in the lip shape library will be calculated for similarity one by one, and the category set with the highest similarity will be selected to obtain several candidate pinyins for each Chinese character.

[0040] Step 7: Pinyin sequence combination based on MoE expert mixture model.

[0041] Since a single lip shape category often corresponds to multiple possible pinyin combinations, this embodiment uses a Mixture of Experts (MoE) model to optimize the combination of candidate pinyin sequences to accurately identify Chinese character sequences. The MoE model consists of several expert networks, each of which focuses on processing specific pronunciation patterns or contextual relationships. The outputs of these expert networks are dynamically weighted and fused through a gating network. Taking the candidate pinyin sequence as input, the MoE model evaluates the rationality of each candidate pinyin combination based on the pinyin contextual relationship and language characteristics, generates a series of pinyin sequence scoring values, and selects the optimal pinyin sequence.

[0042] Step 8: Translation of Pinyin sequence to Chinese text.

[0043] After obtaining the optimal pinyin sequence, this embodiment further uses a pre-trained pinyin-to-Chinese character translation language model (a Transformer-based sequence-to-sequence model) to convert the pinyin sequence into the final Chinese text output. During the training phase, this model fully learns the mapping relationship between pinyin and Chinese characters, capturing the inherent syntactic and semantic regularities of the Chinese language. This enables precise and efficient conversion of pinyin sequences into Chinese character sequences. The final Chinese text output undergoes automatic verification and post-processing, effectively reducing recognition errors and improving the overall accuracy and robustness of Chinese lip reading recognition.

[0044] In some other implementations of this embodiment, the Python library Pinyin2Hanzi is used to convert the optimal pinyin sequence into the final Chinese text output.

[0045] This embodiment further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned lip reading classification method when executing the computer program.

[0046] This embodiment further provides a computer-readable storage medium storing a computer program, which implements the steps of the above-mentioned lip reading classification method when executed by a processor.

[0047] This embodiment provides a GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion, which represents lip key points through three structures: adjacency graph, symmetry graph and upper and lower lip relationship graph. Under the synergistic effect of GNN and LSTM, high-dimensional spatiotemporal features are extracted to effectively capture the spatiotemporal global correlation between lip key points. The method is divided into lip shape categories based on initials and finals, and the multi-level collaborative relationship between initials and finals is considered to reduce the impact of lip shape similarity and visual confusion on model performance. The high-dimensional spatiotemporal features extracted above are used to establish a lip shape library, thereby performing a more discriminative mapping and induction between lip morphology and corresponding pinyin, reducing the impact of lip shape similarity and visual confusion on model performance, and being able to adapt to the complex changes and many-to-one mapping phenomenon in the actual pronunciation process, effectively enhancing the accuracy and robustness of subsequent lip shape classification.

[0048] In this invention, GNN stands for Graph Neural Networks, and Graph Convolutional Networks (GCN) is an implementation of Graph Neural Networks.

[0049] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0050] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion is characterized by: The following steps are involved: Obtain a face speech dataset, the face speech dataset including face speech videos and corresponding text labels with timestamps; extract video frames from the face speech videos, crop the video frames to obtain a lip area map, and extract the coordinates of key points from the lip area map; the key points are used as nodes of the graph to form an adjacency graph, a symmetry graph, and an upper and lower lip relationship graph; the adjacent key points of the outer lip of the adjacency graph are connected as edges to form a closed loop, and the adjacent key points of the inner lip are connected as edges to form another closed loop; the left and right symmetrical key points of the symmetry graph are connected to form edges; the corresponding positions of the upper and lower lip relationship graph are The upper lip key points are connected to the lower lip key points to form edges; a graph convolutional neural network is formed through the adjacency graph, symmetry graph and upper and lower lip relationship graph, and the coordinates of the key points are input to output the global features of the three graph structures, and the global features are fused to form fusion features; each lip area graph is formed into a fusion feature through graph convolution, and multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features; the text labels are converted into pinyin, and the initials and finals are divided into lip shape categories, and a lip shape library is established based on the lip shape categories and high-dimensional spatiotemporal features; the lip shape library inputs the face speaking video and outputs the lip reading text content.

2. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The face speech dataset is a CMLR dataset, which is converted into npy format and uses the dlib 68-point face key point detector to obtain the mouth area. The mouth video frame is cropped and grayscale processed to form a 96×96 pixel image with lip key points.

3. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: When extracting key points from the lip area map, 8 key points are formed on the inner lip of the lip and 12 key points are formed on the outer lip; the key points of the inner lip are located at the inner midpoint of the upper lip, the right side of the inner upper lip, the inner right corner of the right mouth, the right side of the inner lower lip, the middle point of the inner lower lip, the left side of the inner lower lip, the inner left corner of the left mouth, and the left side of the inner upper lip; the key points of the outer lip are located at the outer left corner of the left mouth, the left outer side of the upper lip, the left side of the top of the upper lip, the middle point of the top of the upper lip, the right side of the top of the upper lip, the right outer side of the upper lip, the outer right corner of the right mouth, the right outer side of the lower lip, the right side of the bottom of the lower lip, the middle point of the bottom of the lower lip, the left side of the bottom of the lower lip, and the left outer side of the lower lip.

4. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The adjacency graph, symmetry graph, and upper and lower lip relationship graph are expressed as: The single-layer convolution operation of the adjacency graph is expressed as: ; The single-layer convolution operation of the symmetric graph is expressed as: ; The single-layer convolution operation of the upper and lower lip relationship graph is expressed as: ; in, represents the symmetric normalized adjacency matrix, Indicates the current time step The node characteristics of represents the trainable weight matrix, Represents the nonlinear activation function ReLU.

5. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 4 is characterized in that: The fusion features formed by global feature fusion include: Score the global feature attention formed by the single-layer convolution operation of the adjacency graph, symmetry graph, and upper and lower lip relationship graph to form an attention score: ; in, Represent the global features formed from the single-layer convolution operation of the adjacency graph, the symmetry graph, and the upper and lower lip relationship graph, respectively. represents a learnable parameter vector for mapping global features into an attention score; The attention scores are normalized into attention weights through the softmax function: ; ; ; in, Represents the attention weights of the adjacency graph, symmetry graph, and upper and lower lip relationship graph features respectively; Weighted fusion of global features based on attention weights: ; in, Indicates fusion features.

6. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 5 is characterized in that: Each lip region map is convolved to form a fusion feature. Multiple fusion features are processed through a long short-term memory network to output high-dimensional spatiotemporal features, which can be expressed as: ; ; in, represents the fusion feature per unit time, Represents high-dimensional spatiotemporal features.

7. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 1 is characterized in that: The mouth shape categories are divided based on the initials and finals, including: initials b, p, m and f are classified as labiodental sounds, initials z, c, s are classified as tip-front sounds, initials d, t, n, l are classified as tip-medial sounds, initials zh, ch, sh, r are classified as tip-back sounds, initials j, q, x are classified as palatal sounds, initials g, k, h are classified as root sounds, and initials y, w are classified as Y / X type sounds; finals a, o, e, i, u, v are classified as simple finals, finals an, en, in, un, vn are classified as front nasal finals, finals ang, eng, ing, ong are classified as back nasal finals, finals ai, ei, ui, ao, ou, iu are classified as front-sounding complex finals, and finals ie, ve, er are classified as back-sounding complex finals.

8. The GNN-LSTM Chinese lip reading classification method based on node multi-association graph information fusion according to claim 7 is characterized in that: The lip shape class and high-dimensional spatiotemporal features are used to establish a lip shape library, including: combining the initial consonant class and the final vowel class to form a lip shape class, and constructing a lip shape library based on the pinyin corresponding to the lip shape class of the text label corresponding to the high-dimensional spatiotemporal features.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the lip reading classification method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the lip reading classification method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Lip reading method based on adaptive semantic space-time diagram convolutional network

    CN111259875A

  • Method for constructing Chinese lip language recognition modeling unit set

    CN112766101A

  • Lip language recognition method combining graph neural network and multi-feature fusion

    CN112861791A

  • Lip language recognition method and device based on adaptive matrix feature fusion network, and electronic equipment

    CN114359785A

  • Lip language recognition method and system based on multi-scale space-time convolution

    CN115830688A