A data processing method and system for multimodal facial motion point data and vocal cord motion data
By collecting and processing facial and throat data of normal people, a face and neck movement model is established to help deaf-mute people imitate vocalization, solving the problem of autonomous vocalization of deaf-mute people in existing technologies and achieving instant feedback and efficient communication.
Patent Information
- Application Number
- CN202411114777.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-08-14
AI Technical Summary
Existing communication methods and devices for the deaf and mute usually require additional devices and logic, and sign language has learning barriers and geographical restrictions, making it difficult to achieve autonomous Chinese pronunciation.
By collecting continuous facial vocalization images and laryngeal vibration data of normal people, extracting time and space features after preprocessing, and establishing a facial and neck movement model for Chinese vocalization, it helps deaf-mute people imitate normal vocalization and provides instant feedback.
It makes it possible for deaf-mute people to speak Chinese independently, lowers the learning threshold, provides instant feedback, and improves vocal skills and communication efficiency.
Smart Images

Figure CN119025825B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biological signal processing, and in particular relates to a data processing method and system for multimodal facial motion point data and vocal cord motion data. Background Art
[0002] Speech is the primary means of communication between humans and a fundamental survival skill. Speech recognition, a voice technology used in everyday life, helps computers understand the speaker's intent. Its development has greatly improved the relationship between humans and computers, enabling natural and efficient human-computer interaction. However, for the deaf and mute, who cannot or will not speak, speech recognition technology alone is of little use.
[0003] It's often said that "nine out of ten deaf people are mute." The "mute" here doesn't mean "unable to speak," but "unable to speak." Normal people are constantly monitored by their ears when they speak, receiving timely feedback as they learn to speak. However, people with bilateral deafness cannot hear their own pronunciation, cannot judge the effectiveness of their speech, and without timely auditory feedback, they are unable to express their intentions verbally. Over time, simple deafness progresses to deaf-muteness.
[0004] Sign language uses gestures to represent movements, and simulates images or syllables based on changes in gestures to form certain meanings or words. It is a hand language that allows people with hearing impairments or who cannot speak to communicate and exchange ideas with each other.
[0005] However, in actual application scenarios, sign language itself still has problems and limitations. As a language with low universality, sign language itself requires special training and learning, and there are certain barriers, which limits the universality of sign language communication. In addition, there are many local sign languages around the world, which makes sign language communication even more inconvenient; in addition, the language logic of sign language itself is different from that of normal language, and the written language of hearing-impaired people will also be restricted by sign language and cannot form correct written language;
[0006] Existing methods, devices, and systems for assisting deaf-mute people in communicating typically collect other physiological signals, such as eye movement signals or physical button presses, requiring the addition of additional devices and logic; or they are extensions of traditional sign language translation and still fall into the sign language stereotype.
[0007] Therefore, it is necessary to provide a device for assisting the deaf-mute to learn and train normal vocalization to solve the above problems.
[0008] Invention patent content
[0009] The present invention provides a data processing method and system for multimodal facial motion point data and vocal cord motion data to solve the technical problems existing in the known technology.
[0010] A data processing method for multimodal facial motion point data and vocal cord motion data comprises the following steps:
[0011] Step 1: Provide a single-word text and collect continuous images or videos of normal people's facial vocalization and laryngeal vibration data;
[0012] Step 2: preprocessing the collected offline facial continuous images or videos and laryngeal vibration data to obtain preprocessed facial motion point data and vocal cord movement data;
[0013] Step 3: Input the pre-processed vocal cord motion data into the vocal cord motion data recognition model to extract the temporal and spatial features of the vocal cord motion data. Input the pre-processed facial motion point data into the facial motion point data recognition model to extract the temporal and spatial features of the facial motion point data.
[0014] Step 4: Process the temporal and spatial features of the facial motion point data and the vocal cord motion data to obtain the facial feature point motion trajectory and the vocal cord motion trajectory of a normal person speaking Chinese, and finally establish a face and neck motion model for Chinese speech;
[0015] Step 5: The deaf-mute person imitates the Chinese pronunciation model of facial and neck movements, while simultaneously collecting the sound signal during the pronunciation. The deaf-mute person's facial motion data, vocal cord movement data, and sound signal are input into the Chinese pronunciation model of facial and neck movements, and the evaluation results are fed back to the deaf-mute person.
[0016] In the present invention, compared with auxiliary communication devices for the deaf-mute based on other physiological signals, the data collected in step 1 is generated during normal speech, which is simple and easy to obtain. The decoding method in steps 3 and 4 and the constructed face and neck movement model of Chinese speech follow the physiological changes of Chinese speech, conform to the normal logic of written language, and have no additional mapping relationship. In step 5, the deaf-mute person imitates speech based on the face and neck movement model of Chinese speech, which is expected to get rid of the assistance of additional systems and achieve autonomous Chinese speech.
[0017] In step 1, the laryngeal vibration signal includes the vibration generated by the friction between the vocal cords and the air, the vibration generated by the opening and closing movement of the vocal cords, and the vibration signal transmitted through the neck muscles;
[0018] The facial motion point data includes: peripheral facial motion data of the face, lips, and mandible during speech production.
[0019] In step 2, the preprocessing includes signal filtering, amplification, dimensionality reduction and Fourier transform.
[0020] In step 3, the method for extracting spatiotemporal features is Conformer;
[0021] In step 5, the evaluation result is calculated using the Multimodal DBM method.
[0022] A system for assisting deaf-mute voice training using multimodal facial motion point data and vocal cord motion data, comprising:
[0023] A depth camera for capturing continuous images or video, a sensor for recording throat vibrations, and a microphone for recording sound;
[0024] A depth camera for capturing continuous images or video, a sensor for recording throat vibrations, and a microphone for recording sound;
[0025] The temporal and spatial feature extraction module is used to pre-process the continuous images or videos captured by the depth camera to obtain facial motion point data of the vocalization, and extract the temporal and spatial features from the facial motion point data; it is used to pre-process the laryngeal vibration signal collected by the laryngeal vibration sensor to obtain vocal cord movement data, and extract the temporal and spatial features from the vocal cord movement data;
[0026] The feature reconstruction module is used to use the temporal and spatial characteristics of normal facial motion point data and vocal cord motion data to form the facial feature point motion trajectory and vocal cord motion trajectory of Chinese speech, and construct a face and neck motion model for Chinese speech;
[0027] The correction module is used to read the prompts and guidance of the facial and neck movement model of deaf-mute people based on Chinese pronunciation, imitate the facial and neck movements of normal people when pronouncing Chinese, perform multimodal recognition on offline pronunciation data, facial motion point data, vocal cord movement data and sound signals of deaf-mute people, and output correction results.
[0028] The deaf-mute person, guided by the facial and neck movement model for Chinese pronunciation, mimics the facial and neck movements of a hearing person and produces Chinese pronunciation. Simultaneously, continuous images or video, laryngeal vibration signals, and actual sounds are recorded and processed again. The model then evaluates the results and provides feedback to the deaf-mute person.
[0029] The advantages and positive effects of the present invention are:
[0030] Promoting Speech Vocalization Training: This invention combines facial motion data with vocal cord movement data to provide a novel speech vocalization training method for the deaf and mute. By comprehensively utilizing multimodal data, deaf and mute individuals can more intuitively understand the physiological process of vocalization, thereby helping to improve vocalization skills.
[0031] Provide instant feedback: Traditional voice training methods often lack instant auditory feedback. However, this invention can provide deaf-mute people with more intuitive and accurate instant feedback by collecting facial motion point data and vocal cord movement data, helping them to adjust and improve their voice effects in a timely manner.
[0032] Lowering the learning threshold: Compared with traditional sign language training, the method provided by the present invention is more intuitive and convenient, and does not require additional language learning, thereby lowering the learning threshold and enabling more deaf-mute people to easily learn and master vocal skills.
[0033] Improve communication efficiency: By using the method provided by the present invention for voice training, deaf-mute people can express their intentions more accurately, thereby improving the efficiency and quality of communication and promoting their integration and communication with the surrounding society.
[0034] Wide applicability: The present invention is not restricted by region or language and is applicable to deaf-mute groups around the world, providing them with a universal speech vocalization training method with broad application prospects and social significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a structural schematic diagram of a system for assisting deaf-mute voice training using multimodal facial motion point data and vocal cord movement data of the present invention.
[0036] Figure 2 It is a schematic diagram of the workflow of vocal cord movement data extraction, analysis and modeling of the present invention.
[0037] Figure 3 This is a schematic diagram of the workflow for facial motion point data extraction, analysis and modeling of the present invention.
[0038] Figure 4 This is a workflow diagram of a data processing method for multimodal facial motion point data and vocal cord motion data of the present invention. DETAILED DESCRIPTION
[0039] To further understand the content, features and effects of the present invention, the following embodiments are listed and described in detail with reference to the accompanying drawings:
[0040] The Chinese meanings of the English words and abbreviations in this article are as follows:
[0041] Conformer: The Conformer algorithm is a speech recognition model that combines the advantages of the Transformer and Convolutional Neural Networks (CNNs). It is designed to simultaneously capture both global and local features of speech signals, thereby improving speech recognition performance and efficiency. This paper uses it to extract temporal and spatial features from facial motion data and vocal cord movement data.
[0042] Multimodal DBM: Multimodal Deep Boltzmann Machine (MDBM) is a deep generative model for processing multiple input modes. The model can capture the complex relationship between different modes by jointly learning the potential representations of multiple input modes. Specifically, MDBM fuses data from different modalities through a multi-layer neural network structure to generate a unified potential representation. This paper uses Multimodal DBM to calculate the multimodal data input to the face and neck movement model of Chinese pronunciation and obtain the final evaluation results.
[0043] Figure 1 The system, which uses multimodal facial motion data and vocal cord movement data to assist deaf-mute voice training, includes a camera for capturing continuous facial images or video during vocalization, a microphone for capturing laryngeal vibration signals, and actual sound. Deaf-mute individuals train by imitating the Chinese pronunciation of a hearing person using a facial and neck movement model of Chinese pronunciation displayed on a screen.
[0044] Figure 2 It demonstrates the extraction, analysis and modeling process of vocal cord movement data. After the laryngeal vibration signal collected by the sensor is input, it is first preprocessed and then the time and space features are extracted. Based on the extracted time and space features, the vocal cord movement trajectory is reconstructed and together with other information, a face and neck movement model for Chinese speech is constructed.
[0045] Figure 3 The demonstration shows the extraction, analysis and modeling process of facial motion point data. Continuous images or videos taken by a depth camera are used as input. They are first preprocessed and then the spatiotemporal features are extracted. Based on the extracted spatiotemporal features, the motion trajectory of the facial feature points during speech is reconstructed, and together with other information, a face and neck motion model for Chinese speech is constructed.
[0046] Figure 4 What is shown is the implementation process of a preferred embodiment of the present invention.
[0047] Example 1
[0048] The present invention provides a data processing method for multimodal facial motion point data and vocal cord motion data, and the specific implementation steps are as follows:
[0049] The following is a preferred embodiment of the present invention to further illustrate the workflow and working principle of the present invention:
[0050] Step 1: Data Collection
[0051] The system provides text containing a single Chinese character, uses a depth camera to capture continuous facial images or videos of a normal person speaking, and uses a laryngeal vibration sensor to record laryngeal vibration data. The facial images captured by the depth camera include movement data of the face, lips, and jaw. The laryngeal vibration sensor records vibrations caused by friction between the vocal cords and air, vibrations caused by the opening and closing of the vocal cords, and vibration signals transmitted through the neck muscles.
[0052] Step 2: Data Preprocessing
[0053] The collected facial images or videos and laryngeal vibration data are preprocessed. This preprocessing step includes signal filtering, amplification, dimensionality reduction, and Fourier transform. The goal is to remove noise, enhance the effective signal components, simplify the data dimension, and retain the main feature information.
[0054] Step 3: Feature Extraction
[0055] The pre-processed vocal cord motion data is fed into the vocal cord motion data recognition model, and the temporal and spatial features of the vocal cord motion data are extracted using the Conformer method. The pre-processed facial motion point data is fed into the facial motion point data recognition model, and the temporal and spatial features of the facial motion point data are also extracted using the Conformer method.
[0056] Step 4: Model Building
[0057] The extracted spatiotemporal features of facial feature point data and vocal cord motion data were processed to obtain the motion trajectories of facial feature points and vocal cords based on normal Chinese speech. This ultimately led to the establishment of a face and neck motion model for Chinese speech. This model accurately reflects the facial and vocal cord motion states of normal individuals during speech.
[0058] Step 5: Voice training for the deaf
[0059] The deaf-mute person performs imitation vocalization training based on the established Chinese vocalization face and neck movement model. At the same time, a depth camera and a laryngeal vibration sensor are used to collect facial motion point data and vocal cord movement data of the deaf-mute person when speaking, and the actual sound signals emitted are recorded. These data are input into the Chinese vocalization face and neck movement model, and the evaluation results are calculated using the Multimodal DBM method. The feedback results are provided to the deaf-mute person to help him adjust and improve his vocalization. Through the above embodiment, the deaf-mute person can perform more accurate Chinese vocalization training and improve his vocalization ability under the guidance of visual and auditory multimodality.
[0060] The embodiments described above are only used to illustrate the technical ideas and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of the patent of the present invention cannot be limited by these embodiments alone. That is, any equivalent changes or modifications made to the spirit disclosed by the present invention still fall within the scope of the patent of the present invention.
Claims
1. A data processing method based on multimodal facial motion point data and vocal cord motion data, characterized in that: The steps include: Step 1: Provide a single-word text and collect continuous facial vocalization images and laryngeal vibration data of a normal person; Step 2: preprocessing the collected facial vocalization continuous images and laryngeal vibration data to obtain preprocessed facial motion point data and vocal cord movement data; Step 3: input the pre-processed vocal cord motion data into a vocal cord motion data recognition model to extract the temporal and spatial features of the vocal cord motion data; The pre-processed facial motion point data is input into the facial motion point data recognition model, and the temporal and spatial features of the facial motion point data points are extracted using Conformer. Step 4, processing the temporal and spatial features of the vocal cord motion data obtained in step 3 to obtain a vocal cord motion trajectory based on a normal person's Chinese pronunciation; Processing the temporal and spatial features of the facial moving point data points obtained in step 3 to obtain a motion trajectory of facial feature points based on the Chinese pronunciation of a normal person; A face and neck motion model for Chinese speech is established based on the vocal cord motion trajectory of normal Chinese speech and the facial feature point motion trajectory of normal Chinese speech; In step 5, the data of the deaf-mute person imitating the Chinese pronunciation is input into the face and neck movement model of Chinese pronunciation obtained in step 4. The model extracts the facial motion point data, vocal cord movement data and sound signals of the deaf-mute person and uses the Multimodal DBM method for multimodal recognition to output the correction result.
2. The data processing method of multimodal facial motion point data and vocal cord motion data according to claim 1, characterized in that: The laryngeal vibration data includes: vibration data generated by the friction between the vocal cords and the air, vibration data generated by the opening and closing movement of the vocal cords, and vibration signals transmitted through the neck muscles; The facial vocalization continuous image includes: peripheral facial movement data of the face, lips, and mandible during vocalization.
3. The data processing method of multimodal facial motion point data and vocal cord motion data according to claim 1, characterized in that: In step 2, the preprocessing includes: signal filtering, amplification, dimensionality reduction and Fourier transform.
4. A system for assisting deaf-mute people in Chinese pronunciation training using multimodal facial motion point data and vocal cord motion data, which implements the data processing method according to any one of claims 1 to 3, characterized in that: include: A depth camera for capturing continuous images or video, a sensor for recording throat vibrations, and a microphone for recording sound; The temporal and spatial feature extraction module is used to pre-process the continuous images or videos taken by the depth camera to obtain the facial motion point data of the voice, and extract the temporal and spatial features of the facial motion point data; for preprocessing the laryngeal vibration signal collected by the laryngeal vibration sensor to obtain vocal cord motion data, and extracting time features and spatial features from the vocal cord motion data; The feature reconstruction module is used to use the temporal and spatial characteristics of normal facial motion point data and vocal cord motion data to form the facial feature point motion trajectory and vocal cord motion trajectory of Chinese speech, and construct a face and neck motion model for Chinese speech; The correction module is used to read the prompts and guidance of the facial and neck movement model of deaf-mute people based on Chinese pronunciation, imitate the facial and neck movements of normal people when pronouncing Chinese, perform multimodal recognition on offline pronunciation data, facial motion point data, vocal cord movement data and sound signals of deaf-mute people, and output correction results.
Citation Information
Patent Citations
AI auxiliary correction method for children functional articulation phonological disorders
CN112168147A
Visual face contour motion-based dysarthria speech recognition method and system
CN113241065A