Multi-modal lip language recognition method and device fusing respiratory airflow data

Through a multimodal lip recognition method that fuses respiratory airflow data, combined with lip movement characteristics and airflow characteristics, the problem of low accuracy and applicability of lip recognition in the prior art in the larynx patients is solved, and higher recognition accuracy and applicability are achieved.

CN120217274AActive Publication Date: 2025-06-27THE FIRST AFFILIATED HOSPITAL OF SUN YAT SEN UNIV +1

Patent Information

Application Number
CN202510107622.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-27
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing lip recognition model has low recognition accuracy and applicability in laryngeal patients, mainly because the model training dataset has only a single modal face image, and it is impossible to effectively capture the unique lip movement and airflow characteristics of laryngeal patients.

Method used

A multimodal lip recognition method that integrates breathing air flow data is adopted. By obtaining pronunciation action video signals and breathing air flow signals, pre-processing, visual feature extraction, airflow feature extraction and multimodal fusion processing are performed, fusion features are generated and lip recognition model is input to realize multimodal lip recognition.

Benefits of technology

The accuracy and applicability of lip recognition in laryngeal-free patients was significantly improved, and by combining lip movement characteristics and airflow characteristics, the feature expression ability and recognition accuracy of the model were enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217274A_ABST
    Figure CN120217274A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode lip language recognition method and device fusing respiratory airflow data. The method comprises the steps that a pronunciation action video signal and a respiratory airflow signal are acquired; performing first preprocessing on the pronunciation action video signal to obtain a pronunciation action video; performing second preprocessing on the respiratory airflow signal to obtain respiratory airflow data; performing visual feature extraction processing on the pronunciation action video to obtain lip movement features; performing component analysis processing on the respiratory airflow data to obtain a frequency component and an intensity component; according to the frequency component and the intensity component, airflow feature extraction processing is conducted on the respiratory airflow data, and airflow features are obtained; performing multi-modal fusion processing on the lip movement features and the airflow features to obtain fusion features; and inputting the fusion features into a lip language recognition model to obtain a lip language recognition result. According to the invention, multi-modal lip language recognition is realized, and the accuracy and applicability are improved. The method can be widely applied to the technical field of artificial intelligence visual speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence visual speech recognition, and in particular to a multimodal lip-reading recognition method and device that integrates respiratory airflow data. Background Art

[0002] Laryngectomy is an effective way to cure laryngeal cancer, but it will cause the patient to lose the ability to speak normally. Traditional solutions (such as writing boards, sign languages, electronic larynxes, or esophageal speech, etc.) are still not convenient and intuitive enough in daily communication and have limitations. Visual speech recognition programs (VSRPs) can convert silent pronunciation actions into expected speech texts and then broadcast them through speakers. However, most of the existing model training datasets are from healthy people with standard pronunciation. The lip movements and vocalizations of healthy people are synchronized. While the lip movements and expressions of laryngectomized patients without a larynx are often more exaggerated when speaking, and the model training datasets usually only have a single modality of facial images, resulting in low recognition accuracy and low applicability of the models.

[0003] In summary, the technical problems existing in the related art need to be improved. Summary of the Invention

[0004] Embodiments of the present invention provide a multimodal lip-reading recognition method and device that integrates respiratory airflow data, effectively improving the accuracy and applicability.

[0005] On the one hand, embodiments of the present invention provide a multimodal lip-reading recognition method that integrates respiratory airflow data, including the following steps:

[0006] Obtain a pronunciation action video signal and a respiratory airflow signal;

[0007] Perform a first preprocessing on the pronunciation action video signal to obtain a pronunciation action video;

[0008] Perform a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data;

[0009] Perform visual feature extraction processing on the pronunciation action video to obtain lip movement features;

[0010] Perform component analysis processing on the respiratory airflow data to obtain frequency components and intensity components;

[0011] According to the frequency components and the intensity components, perform airflow feature extraction processing on the respiratory airflow data to obtain airflow features, where the airflow features include frequency features, intensity features, or time features;

[0012] Perform multimodal fusion processing on the lip movement features and the airflow features to obtain fusion features;

[0013] Input the fusion feature into the lip-reading recognition model to obtain the lip-reading recognition result.

[0014] In some embodiments, the first preprocessing of the pronunciation action video signal to obtain the pronunciation action video includes:

[0015] Perform digital signal conversion processing on the pronunciation action video signal to obtain a first video;

[0016] Perform frame extraction processing on the first video to obtain a second video;

[0017] Perform time series alignment processing on the second video to obtain a third video;

[0018] Perform grayscale processing on the third video to obtain a fourth video;

[0019] Use a Gaussian background model to perform background noise elimination processing on the fourth video to obtain the pronunciation action video.

[0020] In some embodiments, the second preprocessing of the respiratory airflow signal to obtain the respiratory airflow data includes:

[0021] Perform digital signal conversion processing on the respiratory airflow signal to obtain first airflow data;

[0022] Use a preset denoising method to perform denoising processing on the first airflow data to obtain second airflow data, where the preset denoising method includes wavelet transform or Fourier transform;

[0023] Use a preset filter to perform filtering processing on the second airflow data to obtain third airflow data, where the preset filter includes a low-pass filter or a band-pass filter;

[0024] Perform amplification processing on the third airflow data to obtain the respiratory airflow data.

[0025] In some embodiments, the visual feature extraction processing of the pronunciation action video to obtain the lip movement feature includes:

[0026] Use a face detection method to perform lip region localization on the pronunciation action video to obtain a lip video;

[0027] Use an optical flow method to perform motion trajectory extraction on the lip video to obtain a motion trajectory;

[0028] Use a preset deep learning network to perform motion feature extraction on the motion trajectory to obtain the lip movement feature.

[0029] In some embodiments, the component analysis processing of the respiratory airflow data to obtain frequency components and intensity components includes:

[0030] Performing frequency analysis on the respiratory airflow data using frequency domain analysis to obtain the frequency components;

[0031] Performing intensity analysis on the respiratory airflow data using time domain analysis to obtain the intensity components.

[0032] In some embodiments, the airflow feature extraction processing of the respiratory airflow data according to the frequency components and the intensity components to obtain airflow features includes:

[0033] Performing frequency feature extraction on the respiratory airflow data using the autocorrelation function or power spectral density according to the frequency components to obtain the frequency features, where the frequency features include fundamental frequency, frequency change range or frequency stability;

[0034] Performing intensity feature extraction on the respiratory airflow data using statistical analysis according to the intensity components to obtain the intensity features, where the intensity features include airflow peak value, average airflow intensity or airflow change rate, and the statistical analysis methods include mean method, variance method or peak method;

[0035] Performing time feature extraction on the respiratory airflow data using time series analysis to obtain the time features, where the time features include airflow duration or airflow change time sequence, and the time series analysis methods include difference method or integration method.

[0036] In some embodiments, the multi-modal fusion processing of the lip movement features and the airflow features to obtain fusion features includes:

[0037] Performing time alignment on the lip movement features and the airflow features according to the time stamp using time alignment method, where the time alignment method includes dynamic time warping;

[0038] Generating a query matrix according to the airflow features after time alignment;

[0039] Generating a key matrix and a value matrix according to the lip movement features after time alignment;

[0040] Calculating attention weights according to the query matrix, the key matrix, the value matrix, the dimension of the key and a preset activation function;

[0041] Performing feature fusion on the lip movement features and the airflow features after time alignment according to the attention weights to obtain the fusion features.

[0042] In some embodiments, the lip language recognition model is obtained through the following steps:

[0043] Obtain a pronunciation action video dataset of laryngectomee patients;

[0044] Annotate the pronunciation action video dataset of laryngectomee patients to obtain a training set;

[0045] According to the performance evaluation index, input the training set into the initial deep learning model so that the initial deep learning model is trained to obtain the lip-reading recognition model.

[0046] In some embodiments, the construction process of the initial deep learning model includes:

[0047] Construct an image feature encoding module, which is used to generate a feature sequence according to the video frame sequence by using a deep residual network;

[0048] After the image feature encoding module, construct a temporal feature extraction module, which is used to generate temporal features according to the feature sequence by using a multi-head attention mechanism.

[0049] On the other hand, an embodiment of the present invention provides a multi-modal lip-reading recognition device integrating respiratory airflow data, including:

[0050] A first module, used to obtain a pronunciation action video signal and a respiratory airflow signal;

[0051] A second module, used to perform a first preprocessing on the pronunciation action video signal to obtain a pronunciation action video;

[0052] A third module, used to perform a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data;

[0053] A fourth module, used to perform visual feature extraction processing on the pronunciation action video to obtain lip movement features;

[0054] A fifth module, used to perform component analysis processing on the respiratory airflow data to obtain frequency components and intensity components;

[0055] A sixth module, used to perform airflow feature extraction processing on the respiratory airflow data according to the frequency components and the intensity components to obtain airflow features, where the airflow features include frequency features, intensity features or time features;

[0056] A seventh module, used to perform multi-modal fusion processing on the lip movement features and the airflow features to obtain fusion features;

[0057] An eighth module, used to input the fusion features into a lip-reading recognition model to obtain a lip-reading recognition result.

[0058] The beneficial effects of the present invention are as follows:

[0059] In the embodiment of the present invention, first, a pronunciation action video signal and a respiratory airflow signal are acquired. The pronunciation action video signal is subjected to a first preprocessing to obtain a pronunciation action video, and the respiratory airflow signal is subjected to a second preprocessing to obtain respiratory airflow data. Then, visual feature extraction processing is performed on the pronunciation action video to obtain lip movement features, component analysis processing is performed on the respiratory airflow data to obtain frequency components and intensity components, and based on the frequency components and intensity components, airflow feature extraction processing is performed on the respiratory airflow data to obtain airflow features. Next, multimodal fusion processing is performed on the lip movement features and the airflow features to obtain fusion features. Finally, the fusion features are input into a lip-reading recognition model to obtain a lip-reading recognition result, so that multimodal lip-reading recognition can be realized through the lip movement features and the airflow features, thereby improving the accuracy and applicability.

[0060] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the drawings. Brief Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is a flowchart of a multimodal lip-reading recognition method integrating respiratory airflow data according to an embodiment of the present invention;

[0063] Figure 2 It is a schematic diagram of multimodal data cross-fusion combining an attention mechanism according to an embodiment of the present invention;

[0064] Figure 3 It is a schematic diagram of an initial deep learning model architecture according to an embodiment of the present invention;

[0065] Figure 4 It is a schematic diagram of the structure of a multimodal lip-reading recognition device integrating respiratory airflow data according to an embodiment of the present invention. Detailed Embodiments

[0066] In order to make the objectives, technical solutions, and advantages of the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0067] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0068] The terms "at least one", "a plurality of", "each", "any one", etc. used in the present application, at least one includes one, two, or more than two, a plurality of includes two or more than two, each refers to each one of the corresponding plurality, and any one refers to any one of the plurality.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0070] Before the embodiments of the present application are described in detail, some nouns and terms involved in the embodiments of the present application will be described first. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0071] Lipreading: Also known as lip reading or visual speech recognition, it refers to the process of understanding and identifying the speech content by observing the lip movements, facial expressions, and movements of the speech organs of the speaker. Artificial intelligence visual speech recognition technology is a technology that combines computer vision and speech recognition. Computer vision enables a computer to understand visual information from images or videos, while speech recognition technology enables a computer to recognize and process human speech. This technology can simulate human perception capabilities, enabling a computer to simultaneously understand and analyze visual and auditory data, thereby improving the accuracy and efficiency of information processing. This technology can be used to assist patients after laryngectomy in speech reconstruction. Lipreading is a computer vision technology that can generate text by identifying the words spoken by patients after laryngectomy through their lip movements and output sound through speech synthesis technology to achieve daily communication for laryngectomees. Through software applications installed and running on smartphones. In this case, the mobile application may integrate lipreading technology, making it a portable assistive tool to help patients after laryngectomy with their daily communication.

[0072] Total Laryngectomy: It is a surgical procedure used to treat laryngeal cancer or other laryngeal diseases by removing the entire laryngeal structure. After the surgery, the patient loses the ability to speak through traditional means (such as vocal cord vibration) and becomes a laryngectomee.

[0073] In the related art, laryngectomy is an effective way to treat laryngeal cancer radically, but it will cause the patient to lose the ability of normal vocalization, and the daily communication is severely restricted. Losing the voice will lead to a significant decline in the patient's quality of life, because the inability to speak and communicate normally will cause great pain, anxiety, fear, distress and frustration. In a survey of patients undergoing mechanical ventilation after laryngectomy, 82% of the patients reported moderate to extreme distress due to the inability to speak. Therefore, for these patients, the need to restore their voice is urgent. Traditional solutions such as writing boards, sign languages, electronic larynxes, and esophageal speech are not convenient and intuitive enough for daily communication, and they all have their limitations. Although tracheoesophageal speech is the "gold standard", it has the most complications; although esophageal speech and electronic larynx speech are more natural than tracheoesophageal speech, there are still problems such as speech distortion, weakening, and recognition difficulties in noisy environments or when using the phone. In addition, these methods cannot be used immediately after surgery and need to wait for the surgical sutures to heal sufficiently. To overcome these limitations, researchers have made efforts to develop vision-based speech recognition programs, namely visual speech recognition programs (VSRP). Such programs can convert silent pronunciation movements into the expected speech text and then broadcast it through a speaker. VSRP is essentially a computer-aided lip-reading program that relies on hardware to capture pronunciation movement data and then convert this data into speech text. More importantly, existing lip-reading recognition datasets mostly come from healthy people with standard pronunciation. The lip movements of healthy people are synchronized with vocalization. However, for laryngectomized patients without a larynx, since the larynx is removed, their lip movements and expressions are often more exaggerated when speaking, and the movement amplitudes of their lips and mandibular facial muscles during pronunciation may be different from those of normal people. The accuracy and applicability of the lip-reading recognition model trained with the lip language dataset of healthy people may be greatly reduced. First, laryngectomized patients usually communicate through a neck anterior artificial larynx or esophageal voice. This vocalization method may cause the lip movements to be less clear than those of normal people, or the lip movement amplitude or facial expression to be larger than that of normal people. Second, the lip shape characteristics of laryngectomized patients may change due to the absence of the vocal cords, and the lip shapes of some syllables may be unclear or atypical. Third, due to the absence of the larynx, visual cues such as laryngeal vibration are missing, and lip-reading recognition can only rely on lip movements and facial expressions, which may require stronger data acquisition and feature extraction capabilities. All of the above factors increase the difficulty of lip-reading recognition and make lip-reading recognition for laryngectomized patients relatively more challenging. In related research, the peak expiratory airflow and peak air pressure of laryngectomized patients during breathing are higher than those of normal people. Three different breathing patterns have been observed in laryngectomized patients when speaking with an electronic larynx: breath-holding, exhalation, and inhalation. Among 12 laryngectomized patients, 4 long-term users of electronic larynxes hold their breath when speaking, 7 exhale continuously when speaking, and only 1 maintains breathing when speaking. Therefore, the pneumatic characteristics such as breathing frequency and intensity of laryngectomized patients when speaking are very different from those of normal people, and the changes in breathing frequency and intensity when speaking will both affect the accuracy of lip-reading recognition.Integrate the movements of the lip and facial muscles and the airflow frequency and intensity when a laryngectomee speaks, which are different from those of healthy people, and input them into the trained lip-reading recognition model to improve the specificity and applicability of the model in the laryngectomee group.

[0074] In view of this, in this embodiment, the frequency and intensity of the breathing airflow at the tracheotomy of a laryngectomee and the lip-reading recognition video are subjected to feature extraction, and multi-modal fusion is performed, and then lip-reading recognition is performed, which can effectively improve the convenience and intuitiveness of laryngectomee patients in daily communication, and improve the recognition accuracy and applicability.

[0075] A multi-modal lip-reading recognition method integrating breathing airflow data provided by an embodiment of the present application relates to the technical field of artificial intelligence visual speech recognition. The multi-modal lip-reading recognition method integrating breathing airflow data provided by an embodiment of the present application can be applied to a terminal, can also be applied to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing a multi-modal lip-reading recognition method integrating breathing airflow data, etc., but is not limited to the above forms.

[0076] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0077] The following specifically explains the embodiments of the present application with reference to the accompanying drawings:

[0078] Figure 1 It is an optional flowchart of a multi-modal lip-reading recognition method for fusing respiratory airflow data provided by an embodiment of the present application. Figure 1 The method in it may include but is not limited to steps S101 to S108.

[0079] Step S101: Obtain a pronunciation action video signal and a respiratory airflow signal;

[0080] Step S102: Perform first preprocessing on the pronunciation action video signal to obtain a pronunciation action video;

[0081] Step S103: Perform second preprocessing on the respiratory airflow signal to obtain respiratory airflow data;

[0082] Step S104: Perform visual feature extraction processing on the pronunciation action video to obtain lip movement features;

[0083] Step S105: Perform component analysis processing on the respiratory airflow data to obtain frequency components and intensity components;

[0084] Step S106: According to the frequency components and intensity components, perform airflow feature extraction processing on the respiratory airflow data to obtain airflow features, and the airflow features include frequency features, intensity features or time features;

[0085] Step S107: Perform multi-modal fusion processing on the lip movement features and the airflow features to obtain fusion features;

[0086] Step S108: Input the fusion features into a lip-reading recognition model to obtain a lip-reading recognition result.

[0087] Steps S101 to S108 illustrated in the embodiments of the present application achieve multi-modal lip-reading recognition and improve the accuracy and applicability.

[0088] In step S101 of some embodiments, the pronunciation action video signal can be obtained through a high-definition camera device, and the respiratory airflow signal can be obtained through an airflow sensing device. The pronunciation action video signal and the respiratory airflow signal can also be obtained by other means, not limited to this. Exemplarily, in the acquisition of the pronunciation action video signal, a high-definition camera with a frame rate of at least 30fps can be used to capture the amplitude of subtle lip and jaw movements and the changes in facial expressions. The best installation position of the high-definition camera is facing the face of the laryngectomee, at a distance of about 40 cm - 50 cm, to ensure that the details of the lips, jaw, and facial expressions can be clearly captured. In the acquisition of the respiratory airflow signal, a highly sensitive micro-airflow sensor can be used to accurately capture the airflow changes at the tracheostomy site, and sensors such as thermistor or differential pressure sensors that respond quickly and accurately to airflow changes can be used. The sensor is embedded in the tracheostomy mask or neck guard worn by the laryngectomee to ensure close contact between the sensor and the tracheostomy site and to accurately measure the airflow data.

[0089] In some embodiments, in step S102, the first preprocessing of the pronunciation action video signal to obtain the pronunciation action video may include but is not limited to the following steps:

[0090] Perform digital signal conversion processing on the pronunciation action video signal to obtain the first video;

[0091] Perform frame extraction processing on the first video to obtain the second video;

[0092] Perform time series alignment processing on the second video to obtain the third video;

[0093] Perform grayscale processing on the third video to obtain the fourth video;

[0094] Use the Gaussian background model to perform background noise elimination processing on the fourth video to obtain the pronunciation action video.

[0095] In some embodiments, the pronunciation action video signal can first be subjected to digital signal conversion processing by an image processing unit to obtain a first video, so that the video signal is converted into a digital signal. Then, frame extraction processing is performed on the first video to obtain a second video. Exemplarily, since the number of frames of the video is large and there is a large amount of redundant information between frames (several adjacent frames of a certain frame are actually the same as the content expressed by that frame), a large number of redundant frames will only increase the duration of model training. A frame extraction technique can be used to process the video frames, that is, one frame is extracted every few frames as data, and the redundant frames are deleted. Then, time series alignment processing is performed on the second video to obtain a third video. Exemplarily, since the frame extraction technique is used, the obtained video frames are not in the original time sequence. The time sequence can be adjusted according to the frame extraction ratio, and time alignment processing is performed on each video frame so that each aligned frame has the same time length and resolution. Then, grayscale processing is performed on the third video to obtain a fourth video. It can be understood that converting a color image into a grayscale image can reduce the amount of data and highlight facial features. Finally, a Gaussian background model is used to perform background noise elimination processing on the fourth video to obtain a pronunciation action video, so as to eliminate background noise and only retain the pronunciation actions of the laryngectomee. Additionally, a deep learning algorithm can also be used for background noise elimination processing.

[0096] In some embodiments, in step S103, the second preprocessing of the respiratory airflow signal to obtain respiratory airflow data may include, but is not limited to, the following steps:

[0097] Perform digital signal conversion processing on the respiratory airflow signal to obtain first airflow data;

[0098] Use a preset denoising method to perform denoising processing on the first airflow data to obtain second airflow data. The preset denoising method includes wavelet transform or Fourier transform;

[0099] Use a preset filter to perform filtering processing on the second airflow data to obtain third airflow data. The preset filter includes a low-pass filter or a band-pass filter;

[0100] Perform amplification processing on the third airflow data to obtain respiratory airflow data.

[0101] In some embodiments, the respiratory airflow signal can first be processed by a signal processing unit for digital signal conversion to obtain first airflow data, converting the airflow signal into a digital signal. Then, a preset denoising method is used to denoise the first airflow data to obtain second airflow data, where the preset denoising method includes wavelet transform or Fourier transform. Exemplarily, signal processing techniques such as wavelet transform or Fourier transform can be used to remove the noise components in the airflow signal (first airflow data) and retain the effective signal to obtain the second airflow data. Then, a preset filter is used to filter the second airflow data to obtain third airflow data, where the preset filter includes a low-pass filter or a band-pass filter. Exemplarily, a low-pass filter or a band-pass filter can be used to filter out the high-frequency noise and low-frequency drift in the second airflow data, and only retain the signal components related to the airflow frequency to obtain the third airflow data. Finally, the third airflow data is amplified to obtain the respiratory airflow data, so that the amplitude of the airflow signal in the third airflow data has a sufficient dynamic range in subsequent processing.

[0102] In some embodiments, in step S104, performing visual feature extraction processing on the pronunciation action video to obtain lip movement features may include, but is not limited to, the following steps:

[0103] Using a face detection method to locate the lip region of the pronunciation action video to obtain a lip video;

[0104] Using an optical flow method to extract the motion trajectory of the lip video to obtain a motion trajectory;

[0105] Using a preset deep learning network to extract motion features from the motion trajectory to obtain lip movement features.

[0106] In some embodiments, a face detection method can first be used to locate the lip region of the pronunciation action video to obtain a lip video. Exemplarily, the lip region part can be located through the face detection algorithm (Haar Cascades) of OpenCV, and this part can be cropped to obtain the lip video. Then, an optical flow method is used to extract the motion trajectory of the lip video, and the motion trajectory of the pixels in the image is calculated by analyzing the pixel changes in the image sequence. Finally, a preset deep learning network (such as a convolutional neural network, a recurrent neural network, a Transformer, etc.) is used to extract motion features from the motion trajectory, and the motion features of the lips, mandible, and facial expressions are extracted to obtain lip movement features.

[0107] In some embodiments, in step S105, performing component analysis processing on the respiratory airflow data to obtain frequency components and intensity components may include, but is not limited to, the following steps:

[0108] Using a frequency domain analysis method to perform frequency analysis on the respiratory airflow data to obtain frequency components;

[0109] The intensity analysis of the respiratory airflow data is carried out by using time-domain analysis method to obtain the intensity components.

[0110] In some embodiments, the frequency analysis of the respiratory airflow data can be first carried out by using frequency-domain analysis method to obtain the frequency components. Exemplarily, the frequency components in the respiratory airflow data can be extracted by using frequency-domain analysis method (such as fast Fourier transform FFT) to analyze the fundamental frequency and harmonic components of the airflow, and to identify the fundamental frequency of the airflow and the frequency change during pronunciation. Then, the intensity analysis of the respiratory airflow data is carried out by using time-domain analysis method to obtain the intensity components. Exemplarily, the intensity components in the respiratory airflow data can be extracted by using time-domain analysis method (such as peak detection, average value calculation, etc.) to analyze the amplitude change of the signal in the airflow and to identify the intensity change of the airflow.

[0111] In some embodiments, in step S106, according to the frequency components and intensity components, the airflow feature extraction process is carried out on the respiratory airflow data to obtain the airflow features, which may include but are not limited to the following steps:

[0112] According to the frequency components, the frequency features of the respiratory airflow data are extracted by using the autocorrelation function or power spectral density to obtain the frequency features, and the frequency features include the fundamental frequency, frequency change range or frequency stability;

[0113] According to the intensity components, the intensity features of the respiratory airflow data are extracted by using statistical analysis method to obtain the intensity features, and the intensity features include the airflow peak value, average airflow intensity or airflow change rate, and the statistical analysis method includes the mean method, variance method or peak method;

[0114] The time features of the respiratory airflow data are extracted by using time series analysis method to obtain the time features, and the time features include the airflow duration or airflow change time sequence, and the time series analysis method includes the difference method or integral method.

[0115] In some embodiments, the respiratory airflow data can be processed to extract airflow features based on frequency components and intensity components, obtaining airflow features, where the airflow features include frequency features, intensity features, or time features. In the extraction of frequency features, the frequency features can be obtained by extracting the frequency features of the respiratory airflow data based on the frequency components using the autocorrelation function or the power spectral density. The frequency features can include the fundamental frequency, the frequency change range, or the frequency stability. In the extraction of intensity features, the intensity features can be obtained by extracting the intensity features of the respiratory airflow data based on the intensity components using statistical analysis methods. The intensity features can include the airflow peak value, the average airflow intensity, or the airflow change rate. The statistical analysis methods can include the mean method, the variance method, or the peak method. In the extraction of time features, the time features can be obtained by extracting the time features of the respiratory airflow data using time series analysis methods. The time features can include the airflow duration or the airflow change time sequence. The time series analysis methods can include the difference method or the integration method.

[0116] In some embodiments, in step S107, the multi-modal fusion processing of the lip movement features and the airflow features to obtain the fusion features may include, but is not limited to, the following steps:

[0117] According to the time stamp, the lip movement features and the airflow features are time-aligned using the time alignment method, and the time alignment method includes dynamic time warping;

[0118] Generate a query matrix according to the airflow features after time alignment;

[0119] Generate a key matrix and a value matrix according to the lip movement features after time alignment;

[0120] Calculate the attention weights according to the query matrix, the key matrix, the value matrix, the dimension of the key, and the preset activation function;

[0121] Feature fusion is performed on the lip movement features and the airflow features after time alignment according to the attention weights to obtain the fusion features.

[0122] In some embodiments, the lip movement features and the airflow features can be time-aligned according to the time stamp using the time alignment method first, where the time alignment method can include dynamic time warping to synchronize the lip movement features and the airflow features on the time axis. The cross-modal data fusion combining the attention mechanism is as Figure 2 shown. A query matrix Q can be generated according to the airflow features after time alignment, and a key matrix K and a value matrix V can be generated according to the lip movement features after time alignment. Then, the attention weights are calculated according to the query matrix, the key matrix, the value matrix, the dimension of the key, and the preset activation function. The calculation formula of the attention weights is: Where Attention is the attention weight, Q is the query matrix, K is the key matrix, V is the value matrix, Softmax is the activation function, and d k is the dimension of the key. Finally, according to the attention weight, the lip movement features and airflow features after time alignment are fused to obtain the fused features. It can be understood that the airflow features can be used as the query matrix Q, and the lip movement features can be used as the key matrix K and the value matrix V, so as to model the relationship between the two modalities and fully fuse the airflow features and lip movement features.

[0123] In some embodiments, in step S108, the fused features can be input into the lip reading recognition model to obtain the lip reading recognition result. Exemplarily, the fused features can be input into the trained lip reading recognition model, and after passing through the neural network, the model front-end encoder and decoder, the final lip reading recognition result is output.

[0124] In some embodiments, the lip reading recognition model is obtained through the following steps:

[0125] Obtain the pronunciation action video dataset of laryngectomee patients;

[0126] Annotate the pronunciation action video dataset of laryngectomee patients to obtain the training set;

[0127] According to the performance evaluation index, input the training set into the initial deep learning model to train the initial deep learning model to obtain the lip reading recognition model.

[0128] In some embodiments, first, the pronunciation action video dataset of laryngectomee patients can be obtained. The pronunciation action video dataset of laryngectomee patients contains video samples under different pronunciation actions, different lighting conditions, and different background environments. Then, the pronunciation action video dataset of laryngectomee patients is annotated to obtain the training set. Then, according to the performance evaluation index, the training set is input into the initial deep learning model to train the initial deep learning model to obtain the lip reading recognition model. During the training process, data augmentation techniques (such as image rotation, scaling, color transformation, etc.) can be used to increase the diversity of the dataset and improve the generalization ability of the model. After the training is completed, an independent test dataset can be used to evaluate the performance of the trained model, and calculate indicators such as the accuracy rate, recall rate, and F1 score of the model to ensure that the performance of the model meets the requirements. Moreover, this embodiment can efficiently process video data in a real-time environment, capture and analyze the real-time pronunciation actions of laryngectomee patients, and has the characteristics of low latency and high throughput. It can complete action capture and feature extraction within milliseconds to achieve real-time processing. This embodiment has a real-time feedback mechanism and can output the speech recognition result in real time according to the pronunciation actions of laryngectomee patients. The feedback mechanism can be realized through speech synthesis, text display, etc. to ensure that laryngectomee patients can obtain feedback in a timely manner.

[0129] In some embodiments, the construction process of the initial deep learning model includes:

[0130] Construct an image feature encoding module, which is used to generate a feature sequence according to the video frame sequence by using a deep residual network;

[0131] After the image feature encoding module, construct a temporal feature extraction module, which is used to generate temporal features according to the feature sequence by using a multi-head attention mechanism.

[0132] In some embodiments, the architecture of the initial deep learning model is as Figure 3 shown, mainly including an image feature encoding module and a temporal feature extraction module. The image feature encoding module can be constructed first, and then the temporal feature extraction module can be constructed. Among them, the image feature encoding module is used to generate a feature sequence according to the video frame sequence by using a deep residual network. It can be understood that in the image feature encoding module, in order to enhance the ability to extract lip features, a deep ResNet (deep residual network) model that has been pre-trained on a large-scale lip dataset can be used as the image feature encoder, which combines a multi-level convolutional feature extraction mechanism, not only can capture the fine-grained features of the lips, but also effectively suppresses the problems of information loss and gradient disappearance while retaining important details by introducing a residual learning structure. Further, the temporal feature extraction module is used to generate temporal features according to the feature sequence by using a multi-head attention mechanism. It can be understood that since lip movement is a dynamic process with a time dimension, a temporal feature extraction module can be introduced. The traditional LSTM structure processes time series in a relatively fixed manner, while in this embodiment, a multi-head attention mechanism (Multi-Head Attention) is adopted, which enables the window of temporal modeling to be adaptively adjusted according to the lip movement features at each moment, dynamically captures key temporal information, and obtains temporal features. Moreover, in order to further improve the adaptability to complex lip movements, an image style conversion technology can be combined, so that the model can automatically enhance the contrast and edge sharpness of the lip region, so as to still maintain efficient feature extraction under complex backgrounds and lighting conditions, in order to achieve more robust and accurate lip feature encoding, effectively improving the performance of lip reading recognition.

[0133] In some embodiments, this embodiment captures the lip movement data and airflow data of laryngectomees through a high-definition camera and a micro airflow sensor respectively, and uses advanced image processing and signal processing technologies to accurately analyze lip movement, amplitude, mandibular movement, facial expressions, as well as airflow frequency and intensity. Through multimodal data fusion technology, the lip movement data and airflow data are fused, significantly improving the accuracy of speech recognition for laryngectomees and enabling the system to more accurately recognize the speech information of laryngectomees. This embodiment uses a highly sensitive micro airflow sensor to real-time monitor the changes in airflow frequency and intensity at the tracheostomy site, and through signal processing technology, it accurately analyzes and extracts the time features, frequency features, and intensity features of the airflow. By accurately capturing the airflow features, it is possible to more comprehensively understand the pronunciation method of laryngectomees and enhance the robustness and accuracy of the speech recognition system. This embodiment uses multimodal data fusion technology to perform time alignment and weighted fusion on the lip movement data and airflow data to generate comprehensive multimodal fusion features as important inputs for speech recognition. Multimodal data fusion technology can comprehensively utilize visual and airflow information, enhance the feature expression ability of the system, improve the accuracy and reliability of speech recognition, and reduce the misrecognition rate. This embodiment uses efficient image processing and signal processing technologies to ensure that the system can efficiently process video and airflow data in a real-time environment and achieve low latency in milliseconds. The real-time processing and low latency characteristics enable the system to promptly feedback the pronunciation results of laryngectomees and improve the naturalness and fluency of interaction. Therefore, through high-precision data capture, advanced image processing and signal processing technologies, multimodal data fusion, and the application of deep learning models, this embodiment significantly improves the accuracy of speech recognition for laryngectomees and the robustness of the system, has real-time and low latency characteristics, and is highly adaptable and universal. It not only enhances the convenience and intuitiveness of laryngectomees in daily communication but also improves their quality of life, having important social significance and application value.

[0134] The beneficial effects of implementing the embodiments of the present invention include: The embodiments of the present invention first obtain the pronunciation action video signal and the respiratory airflow signal, perform a first preprocessing on the pronunciation action video signal to obtain the pronunciation action video, and perform a second preprocessing on the respiratory airflow signal to obtain the respiratory airflow data. Then, perform visual feature extraction processing on the pronunciation action video to obtain lip movement features, perform component analysis processing on the respiratory airflow data to obtain frequency components and intensity components, and based on the frequency components and intensity components, perform airflow feature extraction processing on the respiratory airflow data to obtain airflow features. Then, perform multimodal fusion processing on the lip movement features and the airflow features to obtain fusion features. Finally, input the fusion features into the lip movement recognition model to obtain the lip movement recognition result, thereby enabling multimodal lip movement recognition through lip movement features and airflow features, and further improving the accuracy and applicability.

[0135] Such as Figure 4As shown in the figure, an embodiment of the present invention further provides a multi-modal lip-reading recognition device that fuses respiratory airflow data, including:

[0136] A first module 801, configured to obtain a pronunciation action video signal and a respiratory airflow signal;

[0137] A second module 802, configured to perform a first preprocessing on the pronunciation action video signal to obtain a pronunciation action video;

[0138] A third module 803, configured to perform a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data;

[0139] A fourth module 804, configured to perform visual feature extraction processing on the pronunciation action video to obtain lip movement features;

[0140] A fifth module 805, configured to perform component analysis processing on the respiratory airflow data to obtain frequency components and intensity components;

[0141] A sixth module 806, configured to perform airflow feature extraction processing on the respiratory airflow data according to the frequency components and intensity components to obtain airflow features, where the airflow features include frequency features, intensity features, or time features;

[0142] A seventh module 807, configured to perform multi-modal fusion processing on the lip movement features and the airflow features to obtain fusion features;

[0143] An eighth module 808, configured to input the fusion features into a lip-reading recognition model to obtain a lip-reading recognition result.

[0144] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0145] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A multimodal lip reading recognition method integrating respiratory airflow data, characterized in that: The following steps are involved: Acquire pronunciation action video signals and respiratory airflow signals; Performing a first preprocessing on the pronunciation action video signal to obtain a pronunciation action video; Performing a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data; Performing visual feature extraction processing on the pronunciation action video to obtain lip movement features; Performing component analysis on the respiratory airflow data to obtain frequency components and intensity components; performing airflow feature extraction processing on the respiratory airflow data according to the frequency component and the intensity component to obtain airflow features, wherein the airflow features include frequency features, intensity features or time features; Performing multimodal fusion processing on the lip movement feature and the airflow feature to obtain a fusion feature; The fusion feature is input into the lip reading recognition model to obtain the lip reading recognition result.

2. The method according to claim 1, characterized in that The first preprocessing of the pronunciation action video signal to obtain the pronunciation action video includes: Performing digital signal conversion processing on the pronunciation action video signal to obtain a first video; Performing frame extraction processing on the first video to obtain a second video; Performing time series alignment processing on the second video to obtain a third video; Performing grayscale processing on the third video to obtain a fourth video; The background noise of the fourth video is eliminated by using a Gaussian background model to obtain the pronunciation action video.

3. The method according to claim 1, characterized in that The performing a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data includes: Performing digital signal conversion processing on the respiratory airflow signal to obtain first airflow data; De-noising the first airflow data using a preset de-noising method to obtain second airflow data, wherein the preset de-noising method includes a wavelet transform or a Fourier transform; Using a preset filter to filter the second airflow data to obtain third airflow data, wherein the preset filter includes a low-pass filter or a band-pass filter; The third airflow data is amplified to obtain the respiratory airflow data.

4. The method according to claim 1, characterized in that: The performing of visual feature extraction processing on the pronunciation action video to obtain lip movement features includes: Using a face detection method to locate the lip area of ​​the pronunciation action video to obtain a lip video; Extracting the motion trajectory of the lip video using an optical flow method to obtain a motion trajectory; The motion feature of the motion trajectory is extracted using a preset deep learning network to obtain the lip movement feature.

5. The method according to claim 1, characterized in that The component analysis processing of the respiratory airflow data to obtain frequency components and intensity components includes: Performing frequency analysis on the respiratory airflow data using a frequency domain analysis method to obtain the frequency components; The intensity component is obtained by performing intensity analysis on the respiratory airflow data using a time domain analysis method.

6. The method according to claim 1, characterized in that The step of performing airflow feature extraction processing on the respiratory airflow data according to the frequency component and the intensity component to obtain the airflow feature includes: According to the frequency components, frequency features of the respiratory airflow data are extracted using an autocorrelation function or a power spectrum density to obtain the frequency features, wherein the frequency features include a fundamental frequency, a frequency variation range, or a frequency stability; According to the intensity component, extracting the intensity feature of the respiratory airflow data by using a statistical analysis method to obtain the intensity feature, wherein the intensity feature includes an airflow peak value, an average airflow intensity or an airflow change rate, and the statistical analysis method includes a mean method, a variance method or a peak method; The time characteristics of the respiratory airflow data are extracted by using a time series analysis method to obtain the time characteristics, wherein the time characteristics include airflow duration or airflow change sequence, and the time series analysis method includes a difference method or an integration method.

7. The method according to claim 1, characterized in that The performing multimodal fusion processing on the lip movement feature and the airflow feature to obtain a fusion feature includes: According to the timestamp, the lip movement feature and the airflow feature are time-aligned using a time alignment method, wherein the time alignment method includes dynamic time warping; generating a query matrix according to the airflow characteristics after time alignment; generating a key matrix and a value matrix according to the time-aligned lip movement features; Calculating attention weights according to the query matrix, the key matrix, the value matrix, the dimension of the key, and a preset activation function; According to the attention weight, feature fusion is performed on the time-aligned lip movement features and the airflow features to obtain the fused features.

8. The method according to claim 1, characterized in that The lip reading recognition model is obtained by the following steps: Obtain a video dataset of pronunciation movements of patients without a larynx; Annotating the video dataset of pronunciation movements of the laryngeal patient to obtain a training set; According to the performance evaluation index, the training set is input into the initial deep learning model so that the initial deep learning model is trained to obtain the lip reading recognition model.

9. The method according to claim 8, characterized in that The construction process of the initial deep learning model includes: Constructing an image feature encoding module, wherein the image feature encoding module is used to generate a feature sequence using a deep residual network according to a video frame sequence; After the image feature encoding module, a temporal feature extraction module is constructed, and the temporal feature extraction module is used to generate temporal features according to the feature sequence using a multi-head attention mechanism.

10. A multimodal lip reading recognition device integrating respiratory airflow data, characterized in that: include: The first module is used to obtain the pronunciation action video signal and the breathing airflow signal; The second module is used to perform a first preprocessing on the pronunciation action video signal to obtain a pronunciation action video; A third module is used to perform a second preprocessing on the respiratory airflow signal to obtain respiratory airflow data; The fourth module is used to perform visual feature extraction processing on the pronunciation action video to obtain lip movement features; A fifth module is used to perform component analysis on the respiratory airflow data to obtain frequency components and intensity components; A sixth module is used to perform airflow feature extraction processing on the respiratory airflow data according to the frequency component and the intensity component to obtain airflow features, where the airflow features include frequency features, intensity features or time features; A seventh module is used to perform multimodal fusion processing on the lip movement feature and the airflow feature to obtain a fusion feature; The eighth module is used to input the fusion features into the lip reading recognition model to obtain the lip reading recognition results.

Citation Information

Patent Citations

  • Multimodal speech recognition system and method based on millimeter wave radar

    CN116416996A

  • Multi-mode speech recognition method, device and equipment and computer readable medium

    CN118748008A

  • Magnetic resonance breath-holding imaging artificial intelligence auxiliary device based on voice recognition

    CN119046614A

  • Methods and apparatus for detection of disordered breathing

    US20220007965A1

  • Internet calling method and apparatus, computer device, and storage medium

    US20220044693A1

Cited By

  • Voice input system and method based on friction nanometer power generation

    CN121237131A