Face and neck surface electromyographic electrode position selection method for silent speech recognition
By quantifying the spatial activation mode of key pronunciation muscle groups on the face and neck, the electrode position is optimized, and the problem of unreasonable electrode layout in silent speech recognition is solved, signal capture efficiency and wear comfort are improved, and efficient silent speech recognition is achieved.
Patent Information
- Application Number
- CN202510576198.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-22
AI Technical Summary
In the existing silent speech recognition technology, the electrode layout design is unreasonable, resulting in low signal capture efficiency and redundant calculations, affecting wear comfort and recognition accuracy.
By quantifying the spatial activation modes of multiple key pronunciation muscle groups on the face and neck, combined with the surface electromyography signal acquisition array, accurately obtain nerve discharge and muscle movement information, optimize electrode position selection, reduce redundant electrodes, and improve signal processing efficiency and wear comfort.
The accuracy and efficiency of electrode layout are achieved, unnecessary computational complexity is reduced, signal acquisition efficiency and wearer comfort are improved, while maintaining data quality.
Smart Images

Figure CN120514401A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of silent speech recognition, and specifically relates to a method for selecting the position of surface electromyography electrodes of the face and neck for silent speech recognition, which is specifically applied to human-computer interaction systems and silent speech recognition systems based on surface electromyography signals. Background Art
[0002] Speech is the most natural and efficient form of human communication, and automatic speech recognition technology has been widely used in smart devices. However, this technology relies heavily on acoustic signals, resulting in reduced recognition performance in noisy environments. Furthermore, it struggles to meet the needs of private communication and individuals with speech impairments. Consequently, silent speech recognition technology based on non-acoustic signals has garnered significant attention.
[0003] Existing research on silent speech recognition mainly includes two types of solutions based on imaging technology and electrophysiological signals. Imaging technologies such as ultrasound or optical images can capture the movement information of the vocal organs, but they have high requirements for acquisition conditions and are easily affected by occlusion and ambient light, making them difficult to promote and use in natural scenes and portable devices. In contrast, surface electromyography signals record the electrophysiological activities of muscles in areas such as the face and neck related to vocalization by attaching electrodes to the surface of the skin. They have the advantages of being non-invasive, convenient, with relatively simple acquisition equipment and moderate signal processing difficulty. Surface electromyography signals can effectively reflect the activation patterns of the vocal muscles and can still capture meaningful muscle movement characteristics in silent pronunciation scenarios. They are convenient and reliable.
[0004] Current surface electromyography (EMG) acquisition methods include a small-lead arrangement and a high-density electrode arrangement. The small-lead electrodes are placed based on experience without prior quantitative analysis, which may lead to inaccurate speech recognition results. High-density electrodes can cause data redundancy and discomfort for the experimenter. Therefore, it is necessary to perform a secondary fine division of the EMG signal area based on the distribution of facial and neck muscles to screen for the optimal electrode distribution pattern suitable for silent speech. Summary of the Invention
[0005] To address the aforementioned problems of the prior art, the present invention aims to provide a method for selecting surface EMG electrode locations for silent speech recognition. This method addresses the existing issues of irrational electrode layout design, low signal capture efficiency, and computational redundancy. By quantitatively analyzing the spatial activation patterns of multiple key articulatory muscles in the face and neck, and integrating this with a surface EMG signal acquisition array, the present invention accurately captures neural discharge and muscle movement information.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for selecting face and neck surface electromyography electrode positions for silent speech recognition comprises the following steps:
[0008] Step 1: Deploy surface electromyography electrode arrays in symmetrical areas of the face and throat to cover the key articulatory muscles.
[0009] Step 2: Surface electromyography (SEM) signals under different speech tasks are obtained through silent articulation experiments to construct a silent articulation SEM dataset.
[0010] Step 3: Extract the surface electromyography root mean square (RMS) characteristic value of each channel in each electrode array, construct a two-dimensional surface electromyography root mean square (RMS) distribution map, and smooth and normalize the surface electromyography root mean square (RMS) abnormal channels in the distribution map;
[0011] Step 4: Evaluate the symmetry of the spatial position of muscle activation and the coordination of surface EMG signals in the temporal dynamic process. Calculate the surface EMG cosine similarity of the symmetrical channels and the average surface EMG dynamic similarity under the sliding window.
[0012] Step 5: Based on the results of activation level and symmetry analysis, redundant electrodes are streamlined to retain only key channels in highly active or asymmetric muscle areas.
[0013] Preferably, in the step three, when the difference between the surface electromyography root mean square (RMS) value in a channel and the surface electromyography root mean square (RMS) mean of the adjacent channels exceeds three standard deviations, it is determined to be an RMS abnormal channel, and is replaced by the average value of the adjacent channels to achieve local signal smoothing and noise suppression. The advantage of selecting a channel to be determined as an RMS abnormal channel when the difference exceeds three standard deviations is that, based on the assumption of approximate normal distribution, values exceeding ±3 standard deviations from the mean are considered abnormal, which can cover approximately 99.7% of the normal sample range. Therefore, setting the difference between the channel RMS value and the adjacent channel RMS mean exceeding three standard deviations as the abnormality criterion can effectively identify local extreme values caused by factors such as poor electrode contact, motion artifacts or environmental noise during electromyography acquisition, and avoid misleading the overall distribution map and electrode activity assessment.
[0014] Preferably, after normalizing the root mean square (RMS) value, if the RMS value of a channel is ≥0.7, the channel is defined as a high muscle activity area and is retained preferentially. The advantage of defining channels with RMS values ≥0.7 as high muscle activity areas is that after normalizing the root mean square (RMS) eigenvalues of the surface electromyography signal, the activation levels of all channels are mapped to the [0,1] interval. Setting 0.7 as the judgment threshold for high muscle activity areas can effectively identify key muscle group areas whose activation levels are significantly higher than the average level during the pronunciation process. This value is in the high percentile of the normalized distribution, corresponding to muscle activity channels with significantly enhanced signals, which helps to eliminate interference from background low-amplitude signals or non-pronunciation-related muscles.
[0015] Preferably, if the spatial distribution similarity and the average dynamic similarity of a symmetrical channel pair are both ≥ 0.8, only the electrodes on either side of the channel pair are retained, thereby streamlining and optimizing the number of electrodes. The advantage of choosing 0.8 is that the value range of cosine similarity and dynamic collaborative similarity is [-1, 1], and 0.8-1.0 is classified as "extremely strong correlation" in statistics. Using 0.8 as the judgment criterion can ensure the accuracy of symmetry recognition of facial and neck signals while retaining the physiological information of active areas during silent pronunciation to the greatest extent. For channel pairs with a similarity greater than 0.8, only one side of the electrode is retained, thereby streamlining and optimizing the number of electrodes and avoiding interference from redundant information. When the similarity is lower than 0.8, it means that the facial and neck signals lack sufficient symmetry. At this time, electrodes are arranged on both the left and right sides to ensure the comprehensiveness of data collection.
[0016] The present invention visualizes muscle activation areas under different speech tasks, generates heat maps using root mean square (RMS) values, and quantifies the spatial distribution of muscle activation. Symmetry and activation differences are evaluated using cosine similarity and variance analysis. This optimization scheme avoids electrode redundancy while determining key channel locations and reducing coverage of inactive areas. This optimized design improves the system's computational efficiency and reduces unnecessary signal interference. Compared to existing technologies, the present invention has the following advantages:
[0017] 1) By analyzing the spatial distribution and symmetry of muscle activation areas, the electrode placement is precisely determined, avoiding redundant electrode placement. In traditional methods, excessive redundant channels not only waste resources but also increase computational complexity. By reducing redundant channels, this invention significantly reduces the collection of invalid data and improves signal processing efficiency.
[0018] 2) By optimizing the electrode layout, the number of electrodes is reduced, avoiding the discomfort caused by excessive electrodes in traditional methods. Fewer electrodes not only improves signal acquisition efficiency but also effectively increases wearer comfort, reducing facial and neck pressure or friction caused by excessive electrodes, making the wearer more comfortable during silent speech recognition experiments.
[0019] 3) The optimization of electrode layout reduces the computational burden of subsequent signal processing and improves computational efficiency by streamlining the number of electrodes and improving the rationality of electrode positions. This reduces unnecessary computational complexity while maintaining data quality and improves the system's operating speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Flowchart of the method for selecting surface EMG electrode locations for the face and neck.
[0021] Figure 2 Schematic diagram of electrode patching position.
[0022] Figure 3 Acquire paradigm for database.
[0023] Figure 4 The RMS mapping results of different phoneme categories, where (a) is the electrode array position, (b) is a vowel, (c) is a diphthong, and (d) is a consonant. DETAILED DESCRIPTION
[0024] The present invention is further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments described are only a portion of the present invention, not all of the embodiments. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.
[0025] like Figure 1 As shown, the present invention provides a method for selecting the position of surface myoelectric electrodes of the face and neck for silent speech recognition, comprising the following steps:
[0026] Step 1: Surface EMG electrode placement
[0027] like Figure 2As shown, four 3×3 surface electromyography (SEM) electrode arrays are used, with a center-to-center distance of 20 mm between adjacent electrode channels. These electrodes are located at marked positions A1, A2, B1, and B2, respectively. Arrays A1 and B1 cover the facial region, recording activity associated with oral movements; arrays A2 and B2 cover the laryngeal region, recording activity in the laryngeal articulatory muscles. The facial electrodes are symmetrically placed on both cheeks, covering key articulatory muscles such as the orbicularis oris and buccinator, recording EMG activity during mouth shape changes. The laryngeal electrodes are placed on both sides of the thyroid cartilage, covering neck muscles involved in vocal control, such as the cricothyroid and sternocleidomastoid muscles, recording EMG signals during articulation.
[0028] Step 2: Database acquisition and construction of a multi-task silent articulation surface electromyography dataset
[0029] A total of N healthy adult volunteers were recruited to participate in the data collection experiment, including N / 2 males and N / 2 females. During the data collection process, the subjects sat in a comfortable chair and faced the computer screen to perform the experimental tasks. Figure 3 As shown in the figure, the speech material contains a total of five phonemes. The vowels include the single vowel / i / and the diphthong / ai / , and the consonants include the plosive / p / , the fricative / f / , and the nasal / m / . In the silent articulation mode, the subjects performed articulation movements according to the prompts but did not actually produce any sound. Each phoneme was repeated 10 times. A silent articulation surface electromyography dataset was constructed.
[0030] Step 3: Quantify the spatial activation area of the voiceless articulatory muscles
[0031] For each trial, the surface electromyography root mean square (RMS) eigenvalue of each channel in each electrode array was extracted, and a two-dimensional surface electromyography root mean square (RMS) distribution map was constructed based on the eigenvalue. The formula for the surface electromyography root mean square (RMS) eigenvalue is:
[0032]
[0033] Among them, N is the total number of sampling points in the window, x n Indicates the surface electromyography signal amplitude at the nth sampling point in the window.
[0034] For each electrode array, if the difference between the surface electromyography root mean square (RMS) value of the channel and the average surface electromyography RMS value of its adjacent channels (the vertex channel includes 3 adjacent channels, and the edge channel includes 5 adjacent channels) exceeds three standard deviations, the channel is marked as an RMS abnormal channel. For the detected RMS abnormal channel, its RMS value will be replaced with the RMS mean of its adjacent channels to achieve local smoothing and noise suppression. After the outlier correction is completed, the RMS distribution map corresponding to each electrode array is normalized, and the RMS values of all channels are scaled to the [0,1] interval. Channels with RMS values ≥ 0.7 are defined as high-activity areas.
[0035] like Figure 4 As shown in the RMS results, in the neck region, surface EMG activation is primarily concentrated around the sternohyoid and omohyoid muscles, while facial EMG activation is primarily concentrated around the levator labii superioris, depressor anguli oris, buccinator, and masseter muscles. These muscles are involved in key articulatory movements, including upper lip elevation, mouth angle retraction, cheek control, and jaw closure.
[0036] Step 4: Static symmetry analysis and dynamic synergy analysis
[0037] To evaluate the symmetry of the spatial position of muscle activation, cosine similarity is used as a measure of spatial distribution consistency. The RMS eigenvectors of the two electrode array regions are set to be a = [a1, a2, ..., a n ] and b=[b1,b2,…,b n ], then the similarity Sim between them is calculated as:
[0038]
[0039] Across five experimental tasks across four subjects, similarity calculated based on RMS feature maps extracted from symmetry channels showed that the symmetry similarity for the facial region ranged from a low of 0.84 to a high of 0.99; for the neck region, it ranged from a low of 0.81 to a high of 0.89, both exceeding 0.8, indicating high consistency in overall spatial distribution. This analysis focuses on symmetry assessment at the static distribution level.
[0040] A sliding window was used to reduce the data dimension and analyze the coordination of surface electromyographic signals in a temporal dynamic process. The signal was segmented using a sliding window with a window length of 250ms and a step size of 100ms, and the RMS eigenvalue of each segment was extracted. The RMS eigenvalue within each window was calculated as follows:
[0041]
[0042] Among them, x t+irepresents the surface electromyography value of the i-th sampling point in the t-th sliding window, and N is the total number of sampling points in the window.
[0043] Then, the RMS time series of each channel is normalized, and the RMS values of all channels are scaled to the range [0,1]. On this basis, for each pair of left and right symmetrical channels, the similarity between their RMS waveforms is calculated in each sliding window to measure the synergy of myoelectric activation during that period of the pronunciation process. Finally, the similarity results of all sliding windows are averaged to obtain the average dynamic similarity AvgSim of the channel pair during the entire pronunciation process:
[0044]
[0045] Where T is the total number of sliding windows.
[0046] A threshold of 0.8 was set as the symmetry threshold: when both the spatial distribution similarity and temporal dynamic synergy of muscle activation were greater than 0.8, the region was judged to have high symmetry during silent articulation; if any dimension was below the threshold, it was considered to have low symmetry.
[0047] The analysis results show that in the five types of experiments on four subjects, the NRMS similarity of the facial channel was generally high, with a minimum of 0.94 and a maximum of 0.99, reflecting the high degree of coordination and synchronization of the activation of the central muscles during pronunciation; the similarity of the cervical channel was between 0.90 and 0.98, slightly lower than that of the face, but still showed strong temporal coordination overall.
[0048] Step 5: Select the electrode placement location on the face and neck
[0049] Based on the symmetry analysis results of the spatial activation areas of the quantified silent articulatory muscles and the spatial distribution of the electromyographic signals, the electrode layout design was optimized by placing electrodes in the active areas and retaining only the electrodes on one side of the symmetrical channels. The electrode positions were finally selected as follows:
[0050] Electrode 1: 1 cm to the right of the nose
[0051] Electrode 2: 4 cm to the right of the nose
[0052] Electrode 3: 1 cm to the right of the corner of the mouth
[0053] Electrode 4: Left side of cheek, 3 cm in front of the ear Electrode 5: 1 cm to the lower right corner of the mouth Electrode 6: 5 cm below the midline of the mandible, 3 cm to the right Electrode 7: 5 cm below the midline of the chin Electrode 8: 5 cm below the midline of the mandible, 3 cm to the left Electrode 9: In the middle of the thyroid cartilage Electrode 10: 2 cm to the right of the thyroid cartilage Electrode 11: 4 cm to the right of the thyroid cartilage Electrode 12: 2 cm below the thyroid cartilage, 4 cm to the right Electrode 13: 2 cm below the thyroid cartilage, 2 cm to the right Electrode 14: 2 cm below the thyroid cartilage.
Claims
1. A method for selecting face and neck surface electromyography electrode positions for silent speech recognition, characterized in that: The following steps are involved: Step 1: Deploy surface electromyography electrode arrays in symmetrical areas of the face and throat to cover the key articulatory muscles. Step 2: Surface electromyography (SEM) signals under different speech tasks are obtained through silent articulation experiments to construct a silent articulation SEM dataset. Step 3: Extract the surface electromyography root mean square (RMS) characteristic value of each channel in each electrode array, construct a two-dimensional surface electromyography root mean square (RMS) distribution map, and smooth and normalize the surface electromyography root mean square (RMS) abnormal channels in the distribution map; Step 4: Evaluate the symmetry of the spatial position of muscle activation and the coordination of surface EMG signals in the temporal dynamic process. Calculate the surface EMG cosine similarity of the symmetrical channels and the average surface EMG dynamic similarity under the sliding window. Step 5: Based on the results of activation level and symmetry analysis, redundant electrodes are streamlined to retain only key channels in highly active or asymmetric muscle areas.
2. The method according to claim 1, characterized in that In step three, when the difference between the surface electromyography root mean square (RMS) value in a channel and the surface electromyography root mean square (RMS) mean value of an adjacent channel exceeds three standard deviations, the channel is determined to be an RMS abnormal channel and replaced with the average value of the adjacent channels to achieve local signal smoothing and noise suppression.
3. The method according to claim 1, characterized in that After normalizing the RMS value, if the RMS value of a channel is ≥0.7, the channel is defined as a high-activity muscle area and is retained first.
4. The method according to claim 1, wherein If the spatial distribution similarity and average dynamic similarity of a symmetrical channel pair are both ≥ 0.8, only the electrodes on either side of the channel pair are retained, thereby streamlining and optimizing the number of electrodes.