Throat motion monitoring method based on stereoscopic perception and hybrid coding Transform model
Through the stereoscopic perception and hybrid coded Transformer model, combined with mechanical motion and electrical characteristic measurement units, the comprehensive tracking and feature fusion of the dynamic characteristics of the throat is achieved, solving the portability and accuracy of traditional monitoring methods, and improving the continuity and accuracy of throat monitoring.
Patent Information
- Application Number
- CN202510514192.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art is difficult to achieve portable continuous monitoring of the physiological state of the throat. The imaging method has the risk of invasiveness and radiation exposure, and it is impossible to achieve high-precision real-time monitoring of the throat function.
The stereoscopic perception and hybrid encoding Transformer model are adopted to obtain the epidermal motion characteristics of the laryngeal epidermal motion sensor unit, and the electrical characteristic measurement unit obtains the deep muscle group characteristics, and performs feature analysis through the HE-Transformer model to achieve comprehensive tracking of the changes in the dynamic characteristics of the laryngeal and the extraction and fusion of local timing characteristics and global timing characteristics.
It realizes high-precision and all-round monitoring of throat movements, breaks through the limitations of traditional monitoring methods, improves the portability and continuity of monitoring, and improves the accuracy and information utilization of monitoring.
Smart Images

Figure CN120419946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical sensing technology, and in particular to a laryngeal movement monitoring method based on stereo perception and a hybrid coding Transformer model. Background Art
[0002] The larynx is the core hub of human breathing, vocalization, and swallowing functions. High-precision real-time monitoring of its physiological state has important clinical significance for early warning of respiratory diseases, functional rehabilitation evaluation after laryngeal surgery, vocal cord health management, and airway protection in critically ill patients.
[0003] Currently, laryngeal function assessment relies primarily on imaging techniques such as video laryngoscopes and fiberoptic endoscopes (FEES) and swallowing scintigraphy (VFSS). However, while FEES can provide intuitive anatomical information, its invasive procedure can easily cause patient discomfort. While VFSS can capture laryngeal motion during swallowing, it carries the risk of radiation exposure and is highly device-dependent, making portable and continuous monitoring difficult. Therefore, it is crucial to design a laryngeal motion monitoring method based on stereo perception and a hybrid encoding Transformer model. Summary of the Invention
[0004] The purpose of the present invention is to provide a laryngeal motion monitoring method based on stereo perception and hybrid coding Transformer model, so as to achieve all-round tracking of the dynamic characteristics of the larynx through the surface-deep stereo perception technology, and to realize the extraction and fusion of local and global temporal features through the short-term component extraction and patched Transformer framework.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A laryngeal movement monitoring method based on stereo perception and a hybrid coding Transformer model includes the following steps:
[0007] Acquiring the motion characteristics of the laryngeal epidermis through a mechanical motion sensing unit;
[0008] The characteristics of deep laryngeal muscle groups are obtained through an electrical characteristic measurement unit;
[0009] Data preprocessing was performed on both the laryngeal surface movement characteristics and the laryngeal deep muscle group characteristics to obtain a data set;
[0010] The dataset is input into the HE-Transformer model for feature analysis to obtain the classification results; the HE-Transformer model consists of an embedding layer, a patch layer, a position encoding layer, a Transformer layer, and a linear layer connected in sequence.
[0011] Optionally, data preprocessing is performed on both the laryngeal skin movement features and the laryngeal deep muscle group features to obtain a data set, specifically: the laryngeal skin movement features and the laryngeal deep muscle group features are integrated into a feature set, and the feature set is downsampled to a preset specific length to obtain a data set.
[0012] Optionally, the dataset is input into the HE-Transformer model for feature analysis to obtain classification results, including:
[0013] Perform feature enhancement on the dataset through the embedding layer to obtain an enhanced sequence;
[0014] The enhanced sequence is divided into several patches through the patch layer, and the patches are spliced into the original sequence;
[0015] The position coding layer adds position codes to the patches to obtain coding patches, and concatenates the coding patches and the original sequence into coding sequences.
[0016] The attention head output is obtained by performing an attention enhancement operation on the encoded sequence through the multi-head attention mechanism built into the Transformer layer;
[0017] The attention head output is classified through the feedforward neural network of the linear layer and the linear head to obtain the classification result.
[0018] Optionally, feature enhancement is performed on the dataset through an embedding layer to obtain an enhanced sequence, including:
[0019] Extract the time window of the dataset according to the preset time step;
[0020] Calculate the mean and standard deviation of the time window respectively;
[0021] Perform one-dimensional convolution operations on the mean and standard deviation respectively to obtain their respective short-term features;
[0022] The dataset and short-term features are concatenated in the channel dimension to obtain an enhanced sequence.
[0023] Optionally, the mean and standard deviation are calculated as follows:
[0024]
[0025] Where μ(t) is the mean, σ(t) is the standard deviation, l is the length of the time window, and x local (t,i) is the i-th time window, t is the time step, and ∈ is an additional term.
[0026] Optionally, the positional encoding expression is:
[0027]
[0028] Among them, pos is the position index, i is the dimension index, W (·) Encode the position.
[0029] Optionally, the expression of the multi-head attention mechanism is:
[0030]
[0031] in, is the attention head output, is the query for the i-th encoding patch, is the key of the i-th encoded patch, is the value of the i-th encoded patch, d k is the feature dimension, and Attention(·) is the attention enhancement operation.
[0032] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: The present invention provides a laryngeal motion monitoring method based on stereo perception and a hybrid coding Transformer model, the method comprising: obtaining laryngeal epidermal motion characteristics through a mechanical motion sensing unit; obtaining laryngeal deep muscle group characteristics through an electrical characteristic measurement unit; performing data preprocessing on both the epidermal motion characteristics and the laryngeal deep muscle group characteristics to obtain a data set; inputting the data set into a HE-Transformer model for feature analysis to obtain a classification result; the HE-Transformer model is composed of an embedding layer, a patch layer, a position encoding layer, a Transformer layer, and a linear layer connected in sequence. This method achieves all-round tracking of changes in the dynamic characteristics of the larynx through epidermal-deep stereo perception technology, and achieves the extraction and fusion of local and global temporal features through short-term component extraction and a patched Transformer framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 is a flow chart of the multimodal laryngeal activity monitoring method of the present invention;
[0035] Figure 2 This is the workflow diagram of the HE-Transformer model of the present invention. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0037] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] like Figure 1 As shown, the present invention provides a laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model, comprising the following steps:
[0039] Step 100: Acquire laryngeal epidermal motion characteristics through a mechanical motion sensing unit;
[0040] The mechanical motion sensing unit of this embodiment is a strain sensor (including but not limited to piezoresistive, capacitive, piezoelectric or triboelectric) sensor, a three-axis acceleration sensor or a surface electromyography sensor.
[0041] It should be noted that the mechanical motion sensing unit acquires laryngeal motion characteristics through mechanical deformation of the body surface. When the pharynx moves, the epidermis and the sensing interface undergo compression deformation, generating a changing electrical signal through the electromechanical conversion effect. The characteristics of the electrical signal directly reflect the surface displacement amplitude and motion trajectory of actions such as swallowing and phonation, thereby achieving high-precision laryngeal epidermal motion monitoring.
[0042] Step 200: Acquire characteristics of deep laryngeal muscle groups through an electrical characteristic measurement unit;
[0043] The electrical characteristic measurement unit of this embodiment is an active sensor (including but not limited to electromagnetic field, ultrasonic or radiation) or a bioelectric sensor. The electrical characteristic measurement unit monitors the internal tissue characteristics of the throat through a penetrating sensor. An excitation signal is applied to the throat. When the laryngeal cartilage displaces and the muscles contract, the spatial orientation of the tissue changes with the anatomical deformation. The received signal is then compared with the excitation signal to obtain changes in the deep biological characteristics of the throat muscles, cartilage, etc.
[0044] Example 1 is a combination of a triboelectric sensor and a bioelectrical impedance sensor. When the larynx moves, the triboelectric sensor generates dynamic charges due to skin contact deformation, and the resulting electrical signal reflects the characteristics of the swallowing action. The bioelectrical impedance sensor applies a swept frequency AC signal (<10mA) to the excitation electrode, and the measuring electrode collects the impedance spectrum of the deep tissue of the larynx, and then uses a phase-locked amplifier to extract the real part (Re) and imaginary part (Im) characteristics of the impedance spectrum. Finally, the two signals of the triboelectric sensor and the bioelectrical impedance sensor are synchronized in time and space through the same clock signal to form a fused data stream containing the mechanical displacement-tissue impedance correlation characteristics.
[0045] Example 2 combines a triaxial accelerometer and a surface electromyography (SEM) sensor. The triaxial accelerometer uses vibration spectrum analysis to capture the peak acceleration and direction vector of swallowing movements in real time. The SEM sensor then collects EMG signals at a sampling rate >1kHz to extract muscle and skin electrical characteristics. Finally, the signals from the triaxial accelerometer and SEM sensor are synchronized using a unified clock source.
[0046] It should be noted that the complementary sensing system formed by the fusion of a mechanical motion sensing unit and an electrical property measurement unit transcends the limitations of single-dimensional monitoring and achieves a coordinated, three-dimensional perception of the laryngeal mechanical motion trajectory and internal tissue properties. The electrical property measurement unit employs an active stimulus-response detection mode, detecting changes in dielectric and mechanical properties caused by laryngeal anatomical deformation. This fills the gap in the ability of traditional surface monitoring to capture deep tissue characteristics.
[0047] Step 300: Preprocessing the laryngeal skin movement features and the laryngeal deep muscle group features to obtain a data set;
[0048] Specifically, the laryngeal skin movement features and the laryngeal deep muscle group features are integrated into a feature set, and the feature set is downsampled to a preset specific length to obtain a dataset.
[0049] Step 400: Input the dataset into the HE-Transformer model for feature analysis to obtain the classification result; the HE-Transformer model consists of an embedding layer, a patch layer, a position encoding layer, a transformer layer, and a linear layer connected in sequence. The specific steps are as follows Figure 2 Shown, including:
[0050] Step 401: Perform feature enhancement on the dataset through an embedding layer to obtain an enhanced sequence;
[0051] Specifically, perform feature enhancement operations on the dataset through local component encoding. First, extract the local components of the initial sequence of the dataset. For each time step t (t ≥ l), extract a time window x local (t) with a length of l in the forward direction. The expression is:
[0052] x local (t) = [x(t - l + 1), x(t - l + 2), …, x(t)];
[0053] For time step t (t < l), repeat the time window extraction operation and merge all time windows to obtain x local ∈R C×D×l , where x(·) is the initial sequence of the dataset, C is the number of channels, D is the dimension of the hidden space, and R is the set of real numbers. Then calculate the mean μ(t) and standard deviation σ(t) of each time window. The calculation formulas are:
[0054]
[0055] where x local (t, i) is the i-th time window, and ∈ is an additional term to prevent the variance from being 0. Then perform a one-dimensional convolution operation on the mean and standard deviation to extract more abstract short-term features. The expression is:
[0056] μ conv = Conv1D(μ); σ conv = Conv1D(σ);
[0057] where μ conv is the short-term feature of the mean, σ conv is the short-term feature of the standard deviation, and Conv1D(·) is the one-dimensional convolution operation. Finally, concatenate the initial sequence of the dataset, the convolved mean, and the convolved standard deviation in the channel dimension to obtain the enhanced sequence x concat . The expression is:
[0058] It should be noted that local encoding significantly improves the expressive ability of time series features through local trend modeling. By juxtaposing and connecting the features after local encoding with the initial sequence, an organic integration of multi-scale features is achieved. While maintaining the linear computational complexity, it captures both the microscopic fluctuation patterns and retains the global evolution trend.
[0059] Step 402: Divide the enhanced sequence into several patches through the Patch layer and splice the patches into the original sequence;
[0060] Specifically, the enhanced sequence is divided into multiple patches, each of which has a length of P and a step size of S (i.e., the non-overlapping length between two adjacent patches), and then the patches are spliced into a new sequence representation x p =R {P×N} (original sequence), N is the sum of the number of split patches, and the calculation formula is: Where L is the sequence length.
[0061] Step 403: adding position codes to the patches through the position coding layer to obtain coded patches, and concatenating the coded patches and the original sequence into a coded sequence;
[0062] Specifically, a position code is added to each patch, and the expression of the position code is:
[0063]
[0064] Among them, pos is the position index, i is the dimension index, W (·) is the positional encoding. After adding the positional encoding, the encoded patch and the original sequence are concatenated to obtain the new representation of the input sequence (encoded sequence):
[0065]
[0066] Among them, W pos is the position code, W p is a learnable parameter.
[0067] Step 404: Perform an attention enhancement operation on the encoded sequence using the multi-head attention mechanism built into the Transformer layer to obtain an attention head output;
[0068] Specifically, for each patch sequence x i , generate Q, K and V vectors through linear transformation, the expressions are: Q = x i W Q ; K = x i W K ; V = x i W V ;W Q 、W K and W V Both are learnable weight matrices. Then the dot product of Q and K is calculated and scaled to complete the attention enhancement operation. The process expression is:
[0069]
[0070] in, is the attention head output, is the query for the i-th encoding patch, is the key of the i-th encoded patch, is the value of the i-th encoded patch, d k is the feature dimension, and Attention(·) is the attention enhancement operation.
[0071] Step 405: Classify the attention head output through the feedforward neural network of the linear layer and the linear head to obtain the classification result.
[0072] Specifically, the output of the attention head is sequentially classified by a feedforward neural network (FFN) and a linear head to output the classification results of the physiological movements of pronunciation, drinking water and eating.
[0073] The beneficial effects of the present invention are as follows:
[0074] 1) The HE-Transformer model of the present invention introduces short-term component coding, which can directly encode the short-term fluctuations of the sequence and improve local stability;
[0075] 2) The non-uniform patch partitioning strategy enables the HE-Transformer model to better adapt to the key phases of laryngeal motion and improve the resolution of sequence features;
[0076] 3) By combining position encoding with the multi-head self-attention mechanism, it not only retains the global information but also enhances the expression of short-term component features;
[0077] 4) The HE-Transformer model fuses data from different sensors. This multimodal data fusion strategy can fully utilize the complementary information of different sensors and improve the accuracy of classification.
[0078] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0079] The present invention uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model, characterized in that: The steps include: Acquiring the motion characteristics of the laryngeal epidermis through a mechanical motion sensing unit; The characteristics of deep laryngeal muscle groups are obtained through an electrical characteristic measurement unit; performing data preprocessing on both the laryngeal skin movement characteristics and the laryngeal deep muscle group characteristics to obtain a data set; The data set is input into the HE-Transformer model for feature analysis to obtain a classification result; the HE-Transformer model consists of an embedding layer, a patch layer, a position encoding layer, a Transformer layer and a linear layer connected in sequence.
2. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 1, characterized in that: Data preprocessing is performed on both the laryngeal epidermal movement features and the laryngeal deep muscle group features to obtain a data set. Specifically, the laryngeal epidermal movement features and the laryngeal deep muscle group features are integrated into a feature set, and the feature set is downsampled to a preset specific length to obtain the data set.
3. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 1, characterized in that: The dataset is input into the HE-Transformer model for feature analysis to obtain classification results, including: Performing feature enhancement operation on the data set through the embedding layer to obtain an enhanced sequence; Dividing the enhanced sequence into a plurality of patches through the patching layer, and splicing the patches into the original sequence; Adding position codes to the patches through the position coding layer to obtain coding patches, and concatenating the coding patches and the original sequence into a coding sequence; Performing an attention enhancement operation on the encoding sequence through the multi-head attention mechanism built into the Transformer layer to obtain an attention head output; The attention head output is classified by the feedforward neural network and the linear head of the linear layer to obtain the classification result.
4. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 3, characterized in that: The feature enhancement operation is performed on the dataset through the embedding layer to obtain an enhanced sequence, including: Extracting a time window of the data set according to a preset time step; Calculating the mean and standard deviation of the time window respectively; Performing a one-dimensional convolution operation on the mean and the standard deviation respectively to obtain respective short-term features; The data set and the short-term features are concatenated in the channel dimension to obtain the enhanced sequence.
5. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 4, characterized in that: The calculation formulas for the mean and the standard deviation are respectively: Where μ(t) is the mean, σ(t) is the standard deviation, l is the length of the time window, and x local (t,i) is the i-th time window, t is the time step, and ∈ is an additional term.
6. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 3, characterized in that: The expression of the position encoding is: Among them, pos is the position index, i is the dimension index, W (·) Encode the position.
7. The laryngeal movement monitoring method based on stereo perception and hybrid coding Transformer model according to claim 3, characterized in that: The expression of the multi-head attention mechanism is: in, is the attention head output, is the query for the i-th encoding patch, is the key of the i-th encoded patch, is the value of the i-th encoded patch, d k is the feature dimension, and Attention(·) is the attention enhancement operation.
Citation Information
Cited By
Mechanical drilling speed prediction method based on improved Transform
CN122366532A