A real-time lip-reading method for Chinese mute people based on headphone inertial sensors
By acquiring and processing inertial data on COTS headsets, combining DTW method and Chinese context information, high-precision real-time lip translation of Chinese voiceless people is achieved, solving the problems of high cost and low accuracy of existing equipment, and providing cost-effective barrier-free communication solutions.
Patent Information
- Application Number
- CN202510194575.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing ear-wearing devices based on inertial sensing are mainly carried out on homemade prototypes and are costly, so high-precision real-time lip translation cannot be achieved on commercial off-the-shelf equipment (COTS).
By acquiring and expanding the inertial data set, denoising technology and lightweight methods are used to segment the syllable inertial data, the dynamic time regular distance (DTW) method is used to identify the syllable data, and the context information of Chinese is combined with the character selection to realize real-time reading of Chinese voiceless crowd lip language based on COTS headsets.
It realizes high-precision real-time lip translation on COTS headsets, reduces costs, provides barrier-free communication and interactive services, and has high recognition accuracy and accuracy.
Smart Images

Figure CN119691468B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lip language translation, and particularly relates to a real-time lip language reading method for Chinese mute people based on inertial sensors of earphones. Background Art
[0002] In cases of organic aphasia, phonation disorders, after laryngectomy, and voice disorders, these users have complete pronunciation skills (such as lip and jaw movements), but they cannot enjoy the convenience of voice communication and voice interaction. Lip language provides a friendly communication method without vocalization, using the existing pronunciation process without additional learning costs. However, lip language recognition is still very difficult, and solutions based on computer vision and acoustic sensing perform poorly in non-line-of-sight (NLOS) scenarios.
[0003] Ear-worn sensors provide a new solution for lip language translation. Users naturally wear wireless earphones, which are equipped with various sensors, including inertial sensors (accelerometers and gyroscopes), capable of measuring minute vibrations. Therefore, the inertial sensors on the earphones can detect and identify the movements of the jaw and temporomandibular joint caused by lip language.
[0004] Existing ear-worn devices based on inertial sensing are mainly carried out on self-made prototypes rather than on commercial off-the-shelf (COTS) devices. Users need to pay additional costs for these dedicated peripherals or hardware modifications, with a high cost and low customer willingness to purchase.
[0005] Therefore, how to use inertial sensors on COTS earphones to help Chinese mute people achieve high-precision real-time lip language translation is the technical problem that the present invention wants to solve. Summary of the Invention
[0006] The purpose of the present invention is to provide a real-time lip language reading method for Chinese mute people based on inertial sensors of earphones to solve the problems raised in the above background art.
[0007] The purpose of the present invention is achieved as follows: A real-time lip language reading method for Chinese mute people based on inertial sensors of earphones, characterized in that: the method includes the following steps:
[0008] Step S1: Obtain the inertial data set in the registration stage and expand the inertial data set;
[0009] Step S2: Use denoising technology to obtain error-free inertial data;
[0010] Step S3: Use a lightweight method to segment the syllable inertial data;
[0011] The syllable inertial data includes accelerometer syllable inertial data and gyroscope syllable inertial data;
[0012] Step S4: Use the consistency method to fuse the accelerometer syllable inertial data and the gyroscope syllable inertial data;
[0013] Step S5: Identify the syllable inertial data according to the reference data pre-registered by the user using the dynamic time warping distance DTW method;
[0014] Step S6: Use the context information of Chinese to improve and correct the selection of characters.
[0015] Preferably, in step S1, an inertial data set in the registration stage is obtained and the inertial data set is expanded. The specific operation is as follows:
[0016] Step S1-1: Collect inertial data using the headphone inertial sensor;
[0017] Step S1-2: Construct a converter model, and use the converter model to modify the inertial data of the source user and the silent pronunciation user to achieve the expansion of the inertial data;
[0018] The converter model includes an ID encoder for identifying individual characteristics , a content encoder for extracting content vectors and a decoder D for combining the identity and content of the inertial data. The ID encoder includes two long short-term memory network units LSTM with a feature dimension size of 768 and a fully connected layer. The fully connected layer converts the output result of the long short-term memory network unit LSTM into a final prediction result; The ID encoder The input data is the variables collected from the target user with speech disorders , where represents the user identification of the lip-reading user group is the lip-reading content vector collected from the lip-reading user group represents the inertial readings related to vocalization collected from the lip-reading user group;
[0019] The content encoder includes three convolutional layers with a convolutional kernel size of 5*1 and two bidirectional long short-term memory networks BLSTM with a feature dimension of 32. The content encoder The input is the variables collected from the source user , where represents the user identification of the group of people with normal pronunciation functions is the speech content vector collected from the normal population represents the inertial readings related to pronunciation collected from the normal population;
[0020] The decoder D includes a concatenate layer, three first convolutional layers with a convolutional kernel size of 5*1, three long short-term memory network units (LSTM) with a feature dimension size of 512, four second convolutional layers with a convolutional kernel size of 5*1, and a third convolutional layer with a convolutional kernel size of 5*1. The concatenate layer connects the outputs of the ID encoder and the content encoder, merging different IMU data sources;
[0021] The three first convolutional layers with a convolutional kernel size of 5*1 and the three long short-term memory network units (LSTM) with a feature dimension size of 512 are used to adjust the number of features and combine the features;
[0022] The four second convolutional layers with a convolutional kernel size of 5*1 are used to extract features, and the third convolutional layer with a convolutional kernel size of 5*1 is used for the final output adjustment.
[0023] Preferably, in step S1-2, a transducer model is used to modify the inertial data of the source user and the silent pronunciation to achieve the expansion of inertial data. The specific operation is as follows:
[0024] The inertial data with variable characteristics collected by the source user is represented as , and the inertial data with variable characteristics collected by the silent pronunciation user is represented as ;
[0025] The source user and the source user dictate the same content , the ID encoder pre-trains to extract features , and the content encoder constructs an information bottleneck from the concatenated features of the spectrogram and the features ;
[0026] The decoder D follows the Auto VC principle under the loss function, that is:
[0027] min E CV ,D L=& E [∥I MU ^ S1→S2 -IM U S2 ∥ 2 2 ] &-λ E [|| E CV (I M ^ U S1→S2 )- E CV (IM U S2 )| | 1 ] ;
[0028] Among them, is the generated spectrogram; is the real spectrogram; is the expected value, indicating that the loss function is the average loss over all training samples; is the regularization parameter, used to balance the weights of the two parts in the loss function, controlling the contribution degree of the regularization term to the total loss;
[0029] The real spectrogram and the generated spectrogram There is a high similarity between them, and the Griffin-Lim algorithm is used to estimate the signal of inertial data from the spectrogram.
[0030] Preferably, in step S2, a denoising technique is adopted to obtain error-free inertial data, specifically:
[0031] Step S2-1: Wiener filtering is used to eliminate the inherent noise; the inherent noise includes DC bias and additive white noise;
[0032] Step S2-2: A deep regression model is used to remove the inertial data interfered by the sensor during user movement; user movement includes the movement of the user's body and head, mainly including walking, cycling, driving, and head shaking;
[0033] Step S2-3: A deep regression model is used to remove the influence of user lip movement on inertial data; the inertial data corresponding to the collected user lip movement includes signals containing semantic information and signals not containing semantic information. The signals containing semantic information are the required signals, and the signals not containing semantic information are noise signals that need to be removed, mainly divided into two categories: continuous movement and sudden movement. The representative action of continuous movement is chewing, and the representative actions of sudden movement are yawning and coughing;
[0034] Step S2-4: Downsampling and an adaptive noise filter based on the normalized least mean square (NLMS) algorithm are used to remove the influence of the audio played by the headphone speaker on inertial data;
[0035] According to the Shannon-Nyquist sampling theorem, downsampling operation is performed on the headphone audio with a sampling rate much higher than that of the inertial sensor:
[0036] ;
[0037] Among them, represents the downsampling noise of the axis of the inertial sensor, ; is the transfer function, , and represent the amplitude, frequency, and phase of the audio respectively;
[0038] An adaptive noise filter based on the normalized least mean square (NLMS) algorithm is used to separate the pronunciation-related inertial signal without interference from the pronunciation-related inertial signal affected by interference in real time, so as to eliminate the vibration influence generated by the headphone speaker during audio playback.
[0039] Preferably, in step S3, a lightweight method is used to segment the syllable inertial data, and the specific operation is as follows:
[0040] Step S3-1: Use a double-threshold scheme to remove irrelevant inertial data and retain the inertial data corresponding to lip movement; use the OTSU algorithm to calculate the thresholds of the inertial data corresponding to the required lip movement on each sensor respectively. The part of the inertial data that exceeds the respective sensor thresholds on both sensors is considered the inertial data related to lip pronunciation;
[0041] Step S3-2: Expand the starting point and ending point of each recognition area forward and backward by 5 samples respectively to ensure that the pronunciation actions related to the entire syllable are covered;
[0042] Step S3-3: Segment the pronunciation actions by syllable to segment out the syllable inertial data.
[0043] Preferably, in step S4, the accelerometer syllable inertial data and the gyroscope syllable inertial data are fused using the consistency method, specifically as follows:
[0044] Step S4-1: Reduce the dimension of the accelerometer syllable inertial data. The formula is as follows:
[0045] ;
[0046] Among them, represents the inertial reading along the accelerometer axis, represents the inertial reading along the accelerometer axis, represents the inertial reading along the accelerometer axis; is the sign function, indicating to return the sign; represents the maximum value function;
[0047] Reduce the dimension of the gyroscope syllable inertial data. The formula is as follows:
[0048] ;
[0049] Among them, represents the inertial reading along the gyroscope axis, represents the inertial reading along the gyroscope axis, represents the inertial reading along the gyroscope axis;
[0050] Step S4-2: Use a point-by-point multiplier to fuse the two sources into a one-dimensional feature vector. The formula is as follows:
[0051] ;
[0052] Among them, is the one-dimensional accelerometer syllable inertial data; is the syllable inertial data of a one-dimensional gyroscope, is a point-by-point multiplier.
[0053] Preferably, in step S5, the syllable inertial data is identified by using the dynamic time warping distance Fast-DTW method according to the reference data pre-registered by the user, specifically:
[0054] The Fast-DTW method aligns the fused data in step S4 with the pre-registered reference database and ranks the potential pronunciation gestures according to the DTW distance. Among them, a low distance indicates a high similarity between the reference and the gesture;
[0055] Step S5-1: Detect the pinyin situation of single-vowel letters:
[0056] Pick out the single vowels , , and ; including the following situations:
[0057] Only appear after the initial consonant ;
[0058] Only appear after ;
[0059] The single vowel is only used alone;
[0060] The single vowel only appears in ;
[0061] Step S5-2: Identify the finals containing vowels;
[0062] Step S5-3: Use the prior knowledge of the identified finals to identify the entire syllable inertial data and achieve independent identification of the initial consonants.
[0063] Preferably, in step S6, the context information of Chinese is used to improve and correct the selection of characters, specifically:
[0064] Step S6-1: Use the context information to correct the mis-identified syllables;
[0065] The Chinese pinyin initial consonants are divided into seven categories, represented by ; Define the transition probability , where is the initial consonant and is the final; Use the Bayesian equation to calculate the transmission probability:
[0066] ;
[0067] Among them, represents when under the vowel condition it holds the number of initial consonants;
[0068] Step S6-2: Consider the constraints of Chinese language and context analysis;
[0069] Step S6-3: Continuous translation based on the minimum DTW distance and the highest probability;
[0070] By calculating the minimum DTW distance and the highest transition probability, continuously translate the lip-reading sentences; for each recognized syllable, select the syllable with the minimum DTW distance and the highest transition probability as the final result.
[0071] Preferably, considering the constraints of Chinese language and context analysis in step S6-2 specifically includes:
[0072] The constraints include: the feasibility constraint of syllable combination, the non-uniformity constraint of syllable usage frequency, and the frequent usage constraint of phrases;
[0073] The feasibility constraint of syllable combination: The specificity of the initial syllable and vowel combination in Chinese is reflected in:
[0074] only appears after , , , and and other specific initial consonants, generating the conditional probability ;
[0075] The non-uniformity constraint of syllable usage frequency: The distribution of syllable usage frequency in Chinese is non-uniform, which is reflected in:
[0076] The syllable is more common than , and the syllables and have four similar pronunciation gestures but are not used in Chinese;
[0077] The frequent usage constraint of phrases: Phrases are often used in Chinese instead of single characters, and the context will affect the transition probability. It is necessary to consider the context information to adjust the transition probability. Improve the Bayesian equation in step S6-1 as follows:
[0078] ;
[0079] Among them, represents the number of letters before or after an initial consonant , For the preceding letter, For the following letter.
[0080] Compared with the prior art, the present invention has the following improvements and advantages:
[0081] 1. Expand the inertial data through a converter, and perform further processing on the data to eliminate noise data; use the normalized least mean square algorithm NLMS to separate the clean clarity-related signal from the clarity-related signal affected by the speaker interference in real time, so as to eliminate the vibration influence generated by the headphone speaker during audio playback; at the same time, suppress the influence of the diversity between the user and the device, and ensure the correctness and accuracy of real-time lip language translation.
[0082] 2. Provide real-time barrier-free communication and interaction services for users who cannot use voice through the method of the present invention. Using COTS headphones, no other peripherals are required, and the limited capabilities of mobile devices are fully utilized to achieve sensing based on inertial data. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 Is the overall flowchart of the method of the present invention.
[0084] Figure 2 Is the structural diagram of the converter model.
[0085] Figure 3 Is the structural diagram of the energy distribution of the inherent noise.
[0086] Figure 4 Is the change diagram of the x-axis of the accelerometer after removing the inherent noise.
[0087] Figure 5 Is the frequency response diagram of the experimental equipment.
[0088] Figure 6 Is the comparison diagram of the correct rate of phrase recognition.
[0089] Figure 7 Is the comparison diagram of the accuracy rate of syllable recognition. DETAILED DESCRIPTION OF THE INVENTION
[0090] The following further outlines the present invention in conjunction with the accompanying drawings.
[0091] As Figure 1 shown, a real-time lip language reading method for Chinese mute people based on headphone inertial sensors, the method includes the following steps:
[0092] Step S1: Obtain the inertial data set in the registration stage, and expand the inertial data set;
[0093] Step S1-1: Use a COTS headset with an inertial sensor and a corresponding Apple phone as data collection devices;
[0094] Step S1-2: Construct a converter model, and use the converter model to modify the inertial data of the source user and the silent pronunciation user to achieve the expansion of inertial data;
[0095] The converter is expressed as follows:
[0096] ;
[0097] As Figure 2 shown, the converter model includes an ID encoder for identifying individual characteristics , a content encoder for extracting content vectors and a decoder D for combining the identity and content of inertial data. The ID encoder includes two long short-term memory network units LSTM with a feature dimension size of 768 and a fully connected layer. The fully connected layer converts the output result of the long short-term memory network unit LSTM into the final prediction result; The input data of the ID encoder is a variable collected from the target user with speech disorders , where represents the user identification of the lip-reading user group, is the lip-reading content vector collected from the lip-reading user group, represents the inertial readings related to vocalization collected from the lip-reading user group;
[0098] The content encoder includes three convolutional layers with a convolutional kernel size of 5*1 and two bidirectional long short-term memory networks BLSTM with a feature dimension of 32. The input of the content encoder is a variable collected from the source user , where represents the user identification of the group with normal pronunciation function, is the speech content vector collected from the normal population, represents the inertial readings related to pronunciation collected from the normal population;
[0099] The decoder D includes a concatenate layer, three first convolutional layers with a convolutional kernel size of 5*1, three long short-term memory network units LSTM with a feature dimension size of 512, four second convolutional layers with a convolutional kernel size of 5*1, and a third convolutional layer with a convolutional kernel size of 5*1. The concatenate layer connects the outputs of the ID encoder and the content encoder to merge different IMU data sources;
[0100] Three first convolutional layers with a convolutional kernel size of 5*1 and three long short-term memory network units (LSTM) with a feature dimension size of 512 are used to adjust the number of features and combine features;
[0101] Four second convolutional layers with a convolutional kernel size of 5*1 are used to extract features, and one third convolutional layer with a convolutional kernel size of 5*1 is used for the final output adjustment.
[0102] The transformer model is used to modify the inertial data of the source user and the silent pronunciation to achieve the expansion of inertial data. The specific operation is as follows:
[0103] The inertial data with variable characteristics collected by the source user is represented as The inertial data with variable characteristics collected by the silent pronunciation user is represented as ;
[0104] Source user And source user Oralize the same content ID encoder Pre-train to extract features Content encoder From the spectrogram And features Construct an information bottleneck from the concatenated features;
[0105] The decoder D follows the Auto VC principle under the loss function, that is:
[0106] min E CV ,D L=& E [∥I MU ^ S1→S2 -IM U S2 ∥ 2 2 ] &-λ E [|| E CV (I M ^ U S1→S2 )- E CV (IM U S2 )| | 1 ] ;
[0107] Among them, Is the generated spectrogram; Is the real spectrogram; Is the expected value, indicating that the loss function Is the average loss over all training samples; Is the regularization parameter, which is used to balance the weights of the two parts in the loss function and controls the contribution degree of the regularization term to the total loss;
[0108] Real spectrogram And generated spectrogram Have a high similarity, and the Griffin-Lim algorithm is used to estimate the signal of the inertial data from the spectrogram;
[0109] The Griffin-Lim algorithm is used to estimate the signal of the inertial data from the spectrogram. Specifically:
[0110] Initialization: Read the magnitude spectrum calculated from the IMU data. Since the Griffin-Lim algorithm does not require the original phase information, then randomly initialize a phase spectrum;
[0111] Iterative process: Use the current phase spectrum and magnitude spectrum to reconstruct the audio signal through the inverse short-time Fourier transform (ISTFT) of the short-time Fourier transform (STFT);
[0112] Perform STFT on the reconstructed audio signal again to obtain a new magnitude spectrum and phase spectrum;
[0113] Keep the original magnitude spectrum consistent with the new magnitude spectrum, and update the current phase spectrum with the new phase spectrum;
[0114] Repeat iteration: Repeat the iterative process until the preset number of iterations is reached or the difference between the spectrogram of the reconstructed audio signal and the original spectrogram is small enough;
[0115] Termination condition: When the difference between the magnitude spectrum of the reconstructed signal and the target magnitude spectrum is less than a certain threshold, or the maximum number of iterations is reached, the algorithm terminates;
[0116] Reconstructed signal: Finally, use the final phase spectrum and the target magnitude spectrum to perform the inverse short-time Fourier transform (ISTFT) to obtain the reconstructed inertial data signal.
[0117] Step S2: Adopt a denoising technique to obtain error-free inertial data. The specific operation is as follows:
[0118] Step S2-1: Use Wiener filtering to eliminate the inherent noise;
[0119] The inherent noise includes a DC bias and additive white noise; The energy distribution of the inherent noise is as Figure 3 shown. In the figure, the abscissa represents the power spectral density, with the unit of 10 -6 m 2 / s 4 Hz, and the ordinate represents the frequency, ranging from 0 Hz to 12.5 Hz;
[0120] The DC bias is expressed as:
[0121] DC bias ;
[0122] where is the sampled value of the signal, is the total number of sampling points;
[0123] Subtract the calculated DC bias value from the original signal, expressed as:
[0124] ;
[0125] Among them, is the signal sampling value after debiasing;
[0126] In the frequency domain, multiply the transfer function of the Wiener filter by the Fourier transform of the debiased signal, and then perform the inverse Fourier transform to obtain the filtered signal. The transfer function of the Wiener filter is expressed as:
[0127] ;
[0128] The filtered signal is expressed as:
[0129] ;
[0130] Among them, is the Fourier transform of the filtered signal, is the Fourier transform of the debiased signal. The change after removing noise is as Figure 4 shown.
[0131] Step S2-2: Use a deep regression model to remove the inertial data interfered by the user's movement; the user's movement includes the movement of the user's body and head, mainly including walking, cycling, driving, and head shaking;
[0132] Step S2-3: Use a deep regression model to remove the influence of the user's lip movement on the inertial data; the inertial data corresponding to the user's lip movement collected includes signals containing semantic information and signals not containing semantic information. The signals containing semantic information are the required signals, and the signals not containing semantic information are the noise signals to be removed, mainly divided into two categories: continuous movement and sudden movement. The representative action of continuous movement is chewing, and the representative actions of sudden movement are yawning and coughing;
[0133] The deep regression model includes a flattening layer for converting the input raw data into a one-dimensional array; two fully connected layers for learning complex features in the input data; two dropout layers for randomly "dropping out" a part of the neurons to reduce the risk of overfitting; a regression layer for outputting the result using a linear activation function;
[0134] Step S2-4: Use downsampling and an adaptive noise filter based on the normalized least mean square (NLMS) algorithm to remove the influence of the audio played by the headphone speaker on the inertial data;
[0135] According to the Shannon-Nyquist sampling theorem, perform downsampling on the headphone audio with a sampling rate much higher than that of the inertial sensor:
[0136] ;
[0137] Among them, Indicates an inertial sensor The downsampling noise of the axis, ; is the transfer function, 、 and respectively represent the amplitude, frequency and phase of the audio; for a given sensor, the transfer function is consistent, as shown in Figure 5 , which is the frequency response diagram of the experimental equipment; using a chirp signal to advance and deal with , and using an adaptive noise filter, an adaptive noise filter based on the normalized least mean square algorithm NLMS, to separate the pronunciation-related inertial signal without interference from the pronunciation-related inertial signal affected by interference in real time, so as to eliminate the vibration impact generated by the headphone speaker during audio playback.
[0138] Step S3: Use a lightweight method to segment the syllable inertial data. The specific operation is as follows:
[0139] Step S3-1: Adopt a double-threshold scheme to remove irrelevant inertial data and retain the inertial data corresponding to the lip-reading action; use the OTSU algorithm to calculate the threshold of the inertial data corresponding to the required lip-reading action on each sensor respectively. When the inertial data parts on both sensors exceed their respective sensor thresholds, they are considered as the inertial data related to lip-reading pronunciation;
[0140] Calculate the histogram of the signal strength of each sensor, calculate the cumulative histogram of each signal strength value, and represent it as , where is the signal strength level;
[0141] For each possible threshold , the signal strength is divided into two classes, including class 1 and class 2. Class 1 is the signal strength less than or equal to the threshold , and class 2 is the signal strength greater than the threshold . Calculate the between-class variance :
[0142] σ between 2 (t)= P 0 (t)⋅ P 1 (t) P(t ) 2 ⋅[ μ 0 (t)- μ 1 (t) ] 2 ;
[0143] Among them: is the probability of class 1; is the probability of class 2; is the total probability, that is ; is the average signal strength of class 1; is the average signal strength of class 2.
[0144] Step S3-2: Expand the start and end points of each recognition region forward and backward by 5 samples respectively to ensure that the pronunciation actions related to the entire syllable are covered;
[0145] Step S3-3: Segment the pronunciation actions by syllable to segment out the syllable inertia data.
[0146] According to the language structure of Chinese, the initial consonant is only composed of consonants, while the final must start with a vowel. The vowel is generally the loudest sound and is most easily detected. Therefore, the first occurrence of the vowel is used as the entry point for syllable segmentation; at the same time, the vowel pronunciation will cause a significant change in the inertial sensor readings; after finding the first peak exceeding the threshold on the axis with the highest rotational change, the zero crossing point before this peak is designated as the cutting point.
[0147] Step S4: Use the consistency method to fuse the accelerometer syllable inertia data and the gyroscope syllable inertia data. Specifically:
[0148] Step S4-1: Reduce the dimension of the accelerometer syllable inertia data. The formula is as follows:
[0149] ;
[0150] Among them, represents the inertial reading along the accelerometer axis, represents the inertial reading along the accelerometer axis, represents the inertial reading along the accelerometer axis; is the sign function, indicating to return the sign; represents the maximum value function;
[0151] Reduce the dimension of the gyroscope syllable inertia data. The formula is as follows:
[0152] ;
[0153] Among them, represents the inertial reading along the gyroscope axis, represents the inertial reading along the gyroscope axis, represents the inertial reading along the gyroscope axis;
[0154] Step S4-2: Use a point-by-point multiplier to fuse two sources into a one-dimensional feature vector, with the formula as follows:
[0155] ;
[0156] where is the one-dimensional accelerometer syllable inertial data; is the one-dimensional gyroscope syllable inertial data, is the point-by-point multiplier.
[0157] Step S5: Use the Fast-DTW method of dynamic time warping distance to identify syllable inertial data based on the reference data pre-registered by the user. Specifically:
[0158] The Fast-DTW method aligns the fused data in Step S4 with the pre-registered reference database and ranks potential pronunciation gestures according to the DTW distance. Among them, a low distance indicates a high similarity between the reference and the gesture;
[0159] Step S5-1: Detect the spelling of single-vowel letters:
[0160] Pick out the single vowels , , and ; including the following situations:
[0161] Only appear after the initial consonant ;
[0162] Only appear after ;
[0163] The single vowel is only used alone;
[0164] The single vowel only appears in ;
[0165] Step S5-2: Identify the finals containing vowels;
[0166] It is observed that each final has different clarity. Compared with the initial consonants, the finals have the characteristics of long duration and high intensity. At the same time, the end of a pronunciation gesture may be affected by subsequent syllables, resulting in distortion and overlap; match the final segmentation result with the reference final data, compare the DTW distance between the pronunciation action and the reference signal, and select the three finals with the shortest DTW distance for the subsequent steps;
[0167] Step S5-3: Use the prior knowledge of the identified finals to identify the entire syllable inertial data and achieve independent identification of the initial consonants.
[0168] The pronunciation of consonants is easily affected by the subsequent finals, and it cannot be distinguished based on inertial readings alone. The prior knowledge of the recognized finals is used to recognize the entire syllable, thereby achieving independent recognition of initials; these indistinguishable initials are named visual elements.
[0169] Step S6: Use the context information of Chinese to improve and correct the selection of characters, specifically:
[0170] Step S6-1: Use context information to correct misrecognized syllables;
[0171] The initials of Chinese pinyin are divided into seven categories, represented by ; Define the transition probability , where is the initial, is the final; The Bayesian equation is used to calculate the transmission probability:
[0172] ;
[0173] Among them, represents the number of when the condition of the final is established ;
[0174] Step S6-2: Consider the constraints of Chinese language and context analysis, specifically:
[0175] The constraints include: the feasibility constraint of syllable combination, the non-uniformity constraint of syllable usage frequency, and the frequent usage constraint of phrases;
[0176] The feasibility constraint of syllable combination: The specificity of the initial syllable and final combination in Chinese is reflected in:
[0177] only appears after , , , and and other specific initials, generating the conditional probability ;
[0178] The non-uniformity constraint of syllable usage frequency: The distribution of syllable usage frequency in Chinese is non-uniform, reflected in:
[0179] The syllable is more common than , and the four pronunciation gestures of the syllables and are similar, but they are not used in Chinese;
[0180] Frequency constraint of phrases: In Chinese, phrases are often used instead of single characters, and the context affects the transition probability. Context information needs to be considered to adjust the transition probability. The Bayesian equation in step S6-1 is improved as follows:
[0181] ;
[0182] where, represents the number of preceding or following letters of an initial consonant ; is the preceding letter, is the following letter; for example: although , but the probability of the following will exceed .
[0183] Step S6-3: Continuous translation based on the minimum DTW distance and the highest probability;
[0184] By calculating the minimum DTW distance and the highest transition probability, continuously translate lip-reading sentences; for each recognized syllable, select the syllable with the minimum DTW distance and the highest transition probability as the final result.
[0185] To verify the feasibility and effectiveness of the method of the present invention, the experimental results are presented through relevant performance index evaluations and usage effects; the experimental settings are as follows:
[0186] Experiments are carried out on two COTS headphones. By calling the application programming interface API, the headphones collect inertial data at 25 Hz and send it to the paired mobile device in real time via Bluetooth and run the written Mute - aid application on it;
[0187] The converter model is trained on a server equipped with an Intel (R) Xeon (R) Silver 4210R CPU @ 2.40 GHz and 2 Nvidia RTX 3090s; the lip-reading is played in the audio or converted to text through ASR;
[0188] The data collected in this experiment is mainly divided into two cases: The data collection of 7 voiceless subjects is carried out on Apple airpods3 and iPad 9 (iPadOS 17.5.1); the data collection of 26 vocal volunteers is carried out on Apple Wireless Headphones Pro 2 and iPhone 13 mini (iOS17.6.1), including a total of 33 participants' inertial data, with a total of more than 80,000 items;
[0189] As Figure 6As shown, since in Chinese, the daily usage frequency of phrases is higher than that of single words, the experimental task to be recognized is set to 20 phrases, and the average recognition accuracy rate is 90.32%; the performance of the present invention on COTS earplugs is only 4.18% lower than that of the state-of-the-art (SOTA) lip-reading translator (94.5%), and the latter relies on dedicated sensors / devices. In the present invention, the performance of each phrase is presented along the diagonal. The correct rate of 30% of the phrases exceeds 95%, the correct rate of 45% of the phrases remains at 90%, the correct rate of 85% of the phrases ≥ 80%, and the correct rate of none of the phrases is lower than 70%.
[0190] To evaluate the ability of the method of the present invention to recognize syllables, without context information, it is more challenging to recognize single syllables than to recognize disyllabic phrases; as Figure 7 shown, the method of the present invention achieved a top-1 accuracy rate of 41.28% and a top-3 accuracy rate of 92.67% in 2500 characters containing 50 types of syllables. In addition, it maintained a top-1 accuracy rate of over 70% and a top-3 accuracy rate of over 95% in the recognition of finals and visemes, which indicates that the inventive method can recognize both phrases and characters, and practically demonstrates the potential of the present invention in continuous lip-reading.
[0191] The above is only the implementation manner of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A real-time lip reading method for Chinese voiceless people based on earphone inertial sensor, characterized by: The method comprises the following steps: Step S1: Obtain the inertial data set in the registration phase and expand the inertial data set; Step S1-1: using the inertial data collected by the headset inertial sensor; Step S1-2: construct a converter model, and use the converter model to modify the inertial data of the source user and the silent pronunciation user to achieve the expansion of the inertial data; The transformer model includes an ID encoder E for identifying individual features ID , content encoder E for extracting content vector CV and a decoder D for combining the identity and content of inertial data, an ID encoder E ID It includes two layers of long short-term memory network units LSTM with a feature dimension size of 768 and a fully connected layer. The fully connected layer converts the output results of the long short-term memory network unit LSTM into the final prediction results; ID encoder E ID The input data is the variables (ID T ,CV T ,IMU T ), where ID T User ID indicating the lip reading user population, CV T is the lip reading content vector collected from lip reading users, IMU T represents the inertial readings related to phonation collected from a population of lip readers; The content encoder E CV It includes three convolutional layers with a convolution kernel size of 5*1 and two layers of bidirectional long short-term memory network BLSTM with a feature dimension of 32, content encoder E CV The input is the variable collected from the source user (ID S ,CV S ,IMU S ), where ID S User ID representing people with normal pronunciation function, CV S is the speech content vector collected from normal people, IMU S represents the inertial readings related to articulation collected from a normal population; The decoder D includes a concatenate layer, three first convolution layers with a convolution kernel size of 5*1, three layers of long short-term memory network units LSTM with a feature dimension size of 512, four second convolution layers with a convolution kernel size of 5*1, and a third convolution layer with a convolution kernel size of 5*1. The concatenate layer connects the outputs of the ID encoder and the content encoder together to merge different IMU data sources; The first convolutional layer with three convolution kernels of size 5*1 and three layers of long short-term memory network units LSTM with feature dimension size of 512 are used to adjust the number of features and combine features; The second convolution layer with four convolution kernels of size 5*1 is used to extract features, and the third convolution layer with a convolution kernel of size 5*1 is used for the final output adjustment; In step S1-2, the inertial data of the source user and the silent pronunciation are modified by using the converter model to achieve the expansion of the inertial data. The specific operation is: The inertial data with variable characteristics collected by the source user is represented as (ID S ,CV S ,IMU S ), the inertial data with variable characteristics collected by silent pronunciation users is represented as (ID T ,CV T ,IMU T ); Source User ID S1 and source user ID S2 Describe the same CV S , ID encoder E ID Pre-training to extract feature ID S1 , content encoder E CV From spectrogram IMU S1 and feature ID S1 Constructing information bottlenecks in the concatenated features; The decoder D follows the Auto VC principle under the loss function, that is: in, To generate spectrogram; IMU S2 is the real spectrum; is the expected value, indicating that the loss function L is the average loss over all training samples; λ is the regularization parameter, which is used to balance the weights of the two parts in the loss function and controls the contribution of the regularization term to the total loss; Real Spectrum IMU S2 and generate spectra There is a high similarity between them, and the signal of inertial data is estimated from the spectrum using the Griffin-Lim algorithm; Step S2: using denoising technology to obtain error-free inertial data; Step S3: segmenting syllable inertia data using a lightweight method; The syllable inertia data includes accelerometer syllable inertia data and gyroscope syllable inertia data; Step S4: fusing the accelerometer syllable inertial data and the gyroscope syllable inertial data using a consistency method; Step S5: using the dynamic time warping distance DTW method to identify syllable inertial data according to the reference data pre-registered by the user; Step S6: Using the context information of Chinese language to improve and correct the character selection.
2. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 1, characterized in that: In step S2, a denoising technique is used to obtain error-free inertial data, specifically: Step S2-1: using Wiener filtering to eliminate inherent noise; the inherent noise includes DC bias and additive white noise; Step S2-2: using a deep regression model to remove inertial data interfered by sensors when the user moves; user movement includes movement of the user's body and head, mainly including walking, riding a bicycle, driving, and shaking the head; Step S2-3: using a deep regression model to remove the influence of the user's lip movement on the inertial data; the collected inertial data corresponding to the user's lip movement includes a signal containing semantic information and a signal not containing semantic information, the signal containing semantic information is the required signal, and the signal not containing semantic information is a noise signal that needs to be removed, which is mainly divided into two categories: continuous movement and sudden movement, the representative action of continuous movement is chewing, and the representative action of sudden movement is yawning and coughing; Step S2-4: using downsampling and an adaptive noise filter based on the normalized least mean square algorithm NLMS to remove the influence of the audio played by the earphone speaker on the inertial data; According to the Shannon-Nyquist sampling theorem, the headphone audio with a higher sampling rate than the inertial sensor is downsampled: Among them, IES i (t) represents the down-sampling noise of the i-axis of the inertial sensor, i = x, y, z; H i (f) is the transfer function, where A0, f0 and φ0 represent the amplitude, frequency and phase of the audio frequency respectively; An adaptive noise filter based on the normalized least mean square algorithm (NLMS) is used to separate the pronunciation-related inertial signals without interference from the pronunciation-related inertial signals affected by interference in real time, thereby eliminating the vibration effect generated by the headphone speaker during audio playback.
3. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 1, characterized in that: In step S3, the syllable inertia data is segmented using a lightweight method, and the specific operation is as follows: Step S3-1: Use a double threshold scheme to remove irrelevant inertial data and retain the inertial data corresponding to the lip reading action; use the OTSU algorithm to calculate the threshold of the inertial data corresponding to the required lip reading action on each sensor respectively, and when the inertial data part on both sensors exceeds the threshold of each sensor, it is considered to be the inertial data related to the lip reading pronunciation; Step S3-2: Extend the starting point and the end point of each recognition area forward and backward by 5 samples respectively to ensure that the pronunciation actions related to the entire syllable are covered; Step S3-3: Segment the pronunciation action into syllables to obtain syllable inertia data.
4. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 1, characterized in that: In step S4, the accelerometer syllable inertia data and the gyroscope syllable inertia data are fused using a consistency method, specifically: Step S4-1: Reduce the dimension of the accelerometer syllable inertial data, the formula is as follows: Among them, a x represents the inertial reading along the x-axis of the accelerometer, a y represents the inertial reading along the y-axis of the accelerometer, a z Represents the inertial reading along the z-axis of the accelerometer; sign(·) is a sign function, indicating the return sign; max(·) represents the maximum value function; To reduce the dimension of the gyroscope syllable inertial data, the formula is as follows: Among them, g x represents the inertial reading along the gyroscope x-axis, g y Represents the inertial reading along the y-axis of the gyroscope, g z Represents the inertial reading along the gyroscope's z-axis; Step S4-2: Use a point-by-point multiplier to fuse the two sources into a one-dimensional feature vector, the formula is as follows: in, is the one-dimensional accelerometer syllable inertial data; is the one-dimensional gyroscope syllable inertial data, and ⊙ is the point-by-point multiplier.
5. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 1, characterized in that: In step S5, the syllable inertia data is recognized using the dynamic time warping distance Fast-DTW method according to the reference data pre-registered by the user, specifically: The Fast-DTW method aligns the fused data in step S4 with the pre-registered reference database and ranks potential articulatory gestures according to the DTW distance, where a low distance indicates a high similarity between the reference and the gesture; Step S5-1: Detect the pinyin of monosyllabic letters: Single vowel as well as Selected; including the following: Only appears in the initials z / ts / ,c / ts h / ,s / s / after; Only appears in after; Single vowel Only used alone; Single vowel It only appears in ′ye / iε / ′; Step S5-2: Identify the finals containing vowels; Step S5-3: using the prior knowledge of the identified finals to identify the inertial data of the entire syllable, and realizing independent recognition of the initial consonant.
6. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 1, characterized in that: The step S6 uses the context information of Chinese to improve and correct the selection of characters, specifically: Step S6-1: Correcting the misrecognized syllables using context information; The Chinese Pinyin initials are divided into seven categories, with G i (i=1,2,...,7) represents; define the transition probability P(I,G i |F), where I is the initial consonant and F is the final vowel; the Bayesian equation is used to calculate the transmission probability: Among them, N(I,G i |F) means that under the condition of F final, G i The number of I initials at the time of establishment; Step S6-2: Consider the constraints of Chinese language and context analysis; Step S6-3: Continuous translation based on minimum DTW distance and highest probability; The lip sentences are translated continuously by calculating the minimum DTW distance and the highest transition probability; for each recognized syllable, the syllable with the minimum DTW distance and the highest transition probability is selected as the final result.
7. The method for real-time lip reading of Chinese voiceless people based on earphone inertial sensor according to claim 6, characterized in that: The constraints of Chinese language and context analysis are considered in step S6-2, specifically: Constraints include: The feasibility constraints of syllable combinations, the uneven frequency constraints of syllable usage, and the frequent usage constraints of phrases; Feasibility constraints of syllable combination: The specificity of the combination of initial syllable and final vowel in Chinese is reflected in: 'ia' only appears after the initial consonants of 'd', 'j', 'l', 'q' and 'x', resulting in the conditional probability ∑ I∈{d,j,l,q,x} P(I,G i |ia)=1; The uneven frequency of syllables: The frequency of syllables in Chinese is uneven, as shown in the following: Syllable ′dui / duei / ′ Ratio ′tui / d h uei / ′ is more common, syllable ′nui / nuei / ′ and ′lui / luei / ′ The four pronunciation gestures are similar but not used in Chinese; Frequent usage constraint of phrases: Phrases are often used in Chinese instead of single characters. The context will affect the transition probability. It is necessary to consider the context information to adjust the transition probability. The Bayesian equation in step S6-1 is improved as follows: Among them, N(I,G i |F,C forward,back ) indicates the number of letters before or after an initial consonant I, C forward The first letter is C backward Followed by letters.