Dysphagia detection and automatic scoring system based on high-resolution acoustic sensor

By adopting high-resolution acoustic sensors and multi-level signal processing technology in the swallowing dysphagia detection system, the shortcomings of signal processing and scoring models in the prior art are solved, and high-precision swallowing dysphagia detection and personalized scoring are achieved.

CN120048494APending Publication Date: 2025-05-27THE FIRST AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV (GUANGZHOU RESPIRATORY CENT)
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510356266.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing swallowing dysphagia detection methods have significant problems in signal processing and scoring models, including insufficient resistance to environmental noise and lack of dynamic adjustment capabilities, resulting in low diagnostic accuracy.

Method used

The system based on high-resolution acoustic sensors is adopted, including a three-dimensional curved acoustic sensing array module, an adaptive environmental noise compensation module, a dynamic segmentation module of swallowing stage, a multi-dimensional feature extraction engine, a swallowing abnormal mode classifier and a dynamic weight scoring model. Through the combination of multi-level signal decomposition, noise reduction processing and deep learning models, the precise capture and personalized scoring of the swallowing process can be achieved.

Benefits of technology

It improves the accuracy and reliability of swallowing dysphagia detection, realizes accurate scoring of swallowing function and personalized treatment recommendations, and meets multiple clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048494A_ABST
    Figure CN120048494A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical diagnosis, in particular to a dysphagia detection and automatic scoring system based on a high-resolution acoustic sensor. Comprising a three-dimensional curved surface type acoustic sensing array module, a self-adaptive environment noise compensation module, a swallowing stage dynamic segmentation module, a multi-dimensional feature extraction engine, a swallowing abnormal mode classifier and a dynamic weight scoring model. A self-adaptive environment noise compensation module, a time-frequency domain feature extraction module and a dynamic weight scoring model are combined, the swallowing process is monitored and scored in real time, the system automatically identifies and analyzes acoustic features of an oral cavity preparation period, a pharyngeal period and an esophageal period, and through multi-dimensional feature fusion, dynamic weighting and neural network optimization, the swallowing process is monitored and scored in real time. The method has the advantages of being non-invasive, easy and convenient to operate, high in real-time performance and the like, and can be widely applied to early screening, evaluation and follow-up visit of clinical dysphagia.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical diagnosis, and particularly to a dysphagia detection and automatic scoring system based on a high-resolution acoustic sensor. Background Art

[0002] Dysphagia refers to the symptoms of difficulty or incomplete swallowing during eating or drinking, which is common in the elderly, patients with neurological diseases, and patients during the recovery period after head and neck surgery. Dysphagia not only affects the patient's nutrient absorption, but may also lead to complications such as aspiration and pneumonia, and even threatens life. With the aging of the population and the increase in related diseases, the early screening and accurate diagnosis of dysphagia have become increasingly important.

[0003] Existing dysphagia detection methods mostly rely on traditional imaging examinations, such as video fluoroscopy or fiber endoscopy. Although these methods can provide detailed imaging data, they are complex to operate, costly, and rely on professionals, and cannot achieve instant detection and large-scale screening. In addition, although the dysphagia detection system based on an acoustic sensor has the advantages of non-invasiveness and easy operation, there are still significant problems in signal processing and scoring models in the prior art. First, the existing acoustic sensor system has insufficient resistance to environmental noise, resulting in the inability to effectively distinguish the signals during swallowing and background noise in practical applications. Second, the existing scoring systems usually use fixed rules or single features for determination, lacking the ability of dynamic adjustment and unable to be optimized in real time according to individual differences, resulting in low diagnostic accuracy and unable to meet the clinical personalized needs. Summary of the Invention

[0004] The present invention provides a dysphagia detection and automatic scoring system based on a high-resolution acoustic sensor.

[0005] The dysphagia detection and automatic scoring system based on a high-resolution acoustic sensor includes a three-dimensional curved surface acoustic sensing array module, an adaptive environmental noise compensation module, a swallowing stage dynamic segmentation module, a multi-dimensional feature extraction engine, a swallowing abnormal pattern classifier, and a dynamic weight scoring model, wherein;

[0006] The three-dimensional curved surface acoustic sensing array module is configured with a 16-channel MEMS microphone matrix, which is arranged in a ring at a spacing of 2 mm in the thyroid cartilage-cricoid cartilage complex area to output the original acoustic signal;

[0007] The adaptive environmental noise compensation module receives the original acoustic signal of the acoustic sensing array, performs hybrid noise reduction processing of wavelet packet decomposition and empirical mode decomposition, and outputs an acoustic signal with a high signal-to-noise ratio;

[0008] The swallowing stage dynamic segmentation module receives high signal-to-noise ratio acoustic signals and generates staged signal data based on the sliding window energy entropy change detection and dynamic threshold segmentation algorithm;

[0009] The multi-dimensional feature extraction engine processes the staged signal data, synchronously extracts time-frequency domain features, non-linear dynamics features and phase coupling features, and generates a multi-dimensional feature set;

[0010] The swallowing abnormality pattern classifier receives the multi-dimensional feature set and outputs the probability of swallowing function abnormality through a bidirectional long short-term memory network with an attention mechanism;

[0011] The dynamic weight scoring model generates a standardized swallowing function score based on the abnormality probability and feature deviation degree.

[0012] Optionally, the three-dimensional curved surface acoustic sensing array module includes:

[0013] Bionic laryngeal curved surface base construction: A three-dimensional curved surface base adapted to the laryngeal anatomical structure is constructed using a flexible material, with 16 precision positioning grooves on the surface and a dynamic fitting compensation mechanism integrated;

[0014] Multi-channel MEMS microphone array arrangement: 16 MEMS microphone units are distributed in a spiral ring, and each unit automatically adjusts to be perpendicular to the skin surface through a micro universal joint;

[0015] Temperature-pressure coupling compensation mechanism: Each microphone unit is equipped with a micro piezoresistive sensor to monitor pressure changes and perform dynamic gain compensation when it exceeds the preset range. A temperature sensor array is provided in the center of the base, and the signal baseline drift is corrected using a temperature compensation algorithm;

[0016] Multi-modal signal synchronous acquisition: Each microphone is equipped with a pre-amplification circuit to provide programmable gain. The original acoustic signals are synchronously acquired using an analog-to-digital converter, and the spatial coordinate data of the marking points are acquired through an external infrared camera, and a three-dimensional space mapping relationship between the sensor array and the laryngeal anatomical landmark points is established;

[0017] Dynamic fitting calibration: Start the automatic calibration program, and the base automatically adjusts to fit according to the user's laryngeal contour. A pressure distribution map is generated using pressure feedback, and the corresponding microphone channel is closed when the set pressure is exceeded. After completion, the channel availability status code is output;

[0018] Signal preprocessing and transmission: The original acoustic signals are subjected to high-pass filtering, differential amplification and digital processing, and the processed signals are transmitted through a flexible cable using the I2S protocol, and a meta data packet composed of temperature, pressure and space data is carried.

[0019] Optionally, the adaptive environmental noise compensation module includes:

[0020] Multi-channel signal spatio-temporal alignment: Receive the original acoustic signal and the spatial coordinate data of infrared marker points, calculate the signal time offset based on the time delay model, and use the interpolation algorithm for synchronous alignment to generate spatio-temporal calibration signals;

[0021] Hybrid multi-scale signal decomposition: Perform wavelet packet decomposition and improved empirical mode decomposition on the calibration signal to generate wavelet packet coefficients and intrinsic mode components, and construct a time-frequency representation space.

[0022] Dynamic noise feature library construction: Combine temperature, pressure, and environmental noise baselines to construct a noise template library, calculate the band energy entropy value based on a sliding time window, identify the noise dominant region, and establish a noise feature fingerprint.

[0023] Hierarchical adaptive noise reduction processing: Perform differential noise reduction processing on signals in different frequency bands. The high-frequency signal retains the swallowing characteristics, the intermediate-frequency signal eliminates noise-related components, and the low-frequency signal eliminates breathing noise through an LMS adaptive filter.

[0024] Multi-modal signal reconstruction and optimization: Reconstruct the noise-reduced signal. The high-frequency part is restored through inverse wavelet transform, the intermediate-frequency part is restored by superposition, and the low-frequency noise is removed through a zero-phase FIR filter to generate an acoustic signal with high signal-to-noise ratio.

[0025] Optionally, the swallowing stage dynamic segmentation module includes:

[0026] Multi-channel signal consistency verification: Receive the acoustic signal with high signal-to-noise ratio and the channel availability status code, calculate the peak value of the cross-correlation function of the signals between effective channels, and eliminate the abnormal channel data with a correlation coefficient less than 0.85. Normalize the remaining channels, and generate a standardized signal based on the channel with the maximum energy;

[0027] Dynamic energy entropy change feature extraction: Use a 50ms Hamming window to slide the standardized signal at a step of 10ms, calculate the time-frequency distribution matrix of the Teager energy operator of the signal within each time window, construct a time-varying Shannon entropy curve based on the energy probability density function, and extract the second derivative extreme points of the entropy value curve as candidate segmentation markers through cubic spline interpolation;

[0028] Adaptive threshold model construction: Extract the energy entropy statistical features of the signal during the pre-swallowing resting period, calculate the mean and standard deviation, establish the initial segmentation benchmark for the dynamic threshold T, and update the threshold using the exponential moving average algorithm in combination with the real-time signal characteristics;

[0029] Multi-stage joint segmentation decision: Input the candidate segmentation markers into a multi-channel voting decision mechanism. If the energy entropy transition amplitude exceeds 120% of the threshold T in ≥3 adjacent channels within a 50 ms time window, it is determined as the starting point of the oral preparation phase. The pharyngeal phase trigger requires that at least 5 channel energy peaks reach 80% of the maximum value and the duration is greater than 150 ms. The detection of the esophageal phase termination point uses dual-condition constraints;

[0030] Physiological rationality verification: Conduct biomechanical verification on the preliminary segmentation results, limit the durations of the oral phase, pharyngeal phase, and esophageal phase, as well as the consistency of the energy gradient direction at the phase transition points, and exclude false segmentation points through the multi-channel phase synchronization index;

[0031] Stage-based data encapsulation and output: Generate a structured data packet including accurate timestamps, energy transition amplitudes, the number of activated channels, and physiological verification markers, and output the stage-based signal data to a multi-dimensional feature extraction engine.

[0032] Optionally, the multi-dimensional feature extraction engine includes:

[0033] Dynamic optimization of effective channels: Based on the number of activated channels and physiological verification markers in the stage-based signal data, calculate the correlation coefficient of each channel with the reference template, and select the top 8 channels with a correlation coefficient greater than 0.9 as effective channels for signal enhancement processing;

[0034] Time-frequency domain feature analysis: For each swallowing phase signal, use Mel cepstral coefficients for time-frequency feature extraction, and simultaneously calculate the wavelet packet energy spectrum entropy to generate time-frequency domain features;

[0035] Nonlinear dynamics feature mining: Conduct recursive quantitative analysis on the pharyngeal phase signal segment, calculate deterministic, laminarity, and recurrence entropy parameters, and simultaneously calculate the maximum Lyapunov exponent to track chaotic characteristics and generate nonlinear dynamics features.

[0036] Cross-channel phase coupling analysis: Calculate the bispectral coherence of the pharyngeal phase trigger period, analyze the phase synchronization strength between channels, and extract phase coupling features.

[0037] Optionally, the multi-dimensional feature extraction engine further includes:

[0038] Multi-modal feature fusion: Normalize the time-frequency domain features, nonlinear dynamics features, and phase coupling features, use principal component analysis for dimensionality reduction, retain the principal components with a cumulative contribution rate greater than 95%, and perform weighted fusion according to the feature importance of different stages to generate a 128-dimensional fusion feature vector.

[0039] Dynamic Feature Optimization Mechanism: Based on the feedback results of the swallowing abnormality pattern classifier, the random forest algorithm is used to rank the feature importance, and redundant features with importance scores lower than 0.05 are removed;

[0040] Feature Data Encapsulation and Output: Generate a multi-dimensional feature set containing timestamp alignment information, feature type markers, and data quality indicators.

[0041] Optionally, the swallowing abnormality pattern classifier includes:

[0042] Multi-dimensional Feature Preprocessing: Receive the multi-dimensional feature set output by the multi-dimensional feature extraction engine, perform timestamp alignment verification, fill in the missing feature values caused by channel failures, and reorganize the feature sequences into an oral preparatory phase feature matrix, a pharyngeal phase feature tensor, and an esophageal phase feature vector according to the swallowing stage division. After completing the feature space standardization, input it into the neural network architecture;

[0043] Bidirectional Long Short-Term Memory Network - Attention Network Construction: Deploy a core architecture containing a three-layer bidirectional long short-term memory network, add a temporal convolutional layer to the pharyngeal phase feature processing branch, and integrate a multi-head attention mechanism at the top of the network;

[0044] Temporal Feature Enhancement Processing: Extract local dependence features through the temporal convolutional layer, combine with the hidden state of the bidirectional long short-term memory network to generate an enhanced temporal feature vector, and use a gating mechanism to dynamically adjust the fusion weight of convolutional features and recurrent features;

[0045] Dynamic Allocation of Attention Weights: Perform cross-stage feature correlation analysis in the multi-head attention layer, calculate the cross-attention scores between the oral preparatory phase and pharyngeal phase features, set the time attention weight benchmark value for the pharyngeal phase trigger stage to 0.6, and automatically increase the attention weight of the 200ms period after the pharyngeal phase to above 0.8 when esophageal phase feature abnormalities are detected, generating an attention heat map with clinical interpretability;

[0046] Implementation of the Multi-task Learning Framework: Parallelly output three types of abnormality probabilities, including aspiration risk probability, pharyngeal phase delay probability, and esophageal dysfunction probability, and use dynamic weighted cross-entropy as the loss function;

[0047] Abnormality Probability Fusion Decision: Perform weighted fusion on the multi-task outputs according to clinical rules to calculate the final swallowing function abnormality probability.

[0048] Optionally, the dynamic weight scoring model includes:

[0049] Data Preparation: Receive the swallowing function abnormality probability output from the swallowing abnormality pattern classifier and the feature deviation data obtained from the multi-dimensional feature extraction engine, where;

[0050] The probabilities of abnormal swallowing functions include the probabilities of aspiration risk, pharyngeal phase delay, and esophageal dysfunction;

[0051] The feature deviation degree data for each swallowing stage includes the cepstrum deviation value of the Mel cepstral coefficients in the oral preparatory phase, the non-linear dynamics deviation rate in the pharyngeal phase, and the phase synchronization index decay slope in the esophageal phase;

[0052] Calculation and standardization of feature deviation degree: According to the feature deviation degree calculation formulas for each swallowing stage, calculate the deviation degree coefficients for each stage and perform standardization processing;

[0053] Intelligent allocation of clinical weights: Dynamically allocate the clinical weights for each stage according to the clinical requirements of different swallowing stages;

[0054] Calculation of multi-dimensional score fusion: Generate the final swallowing function score using a weighted fusion algorithm based on the abnormal probability, feature deviation degree, and real-time weights.

[0055] Optionally, the dynamic weight scoring model further includes.

[0056] Physiological constraint calibration mechanism: Introduce the physiological boundary conditions of the swallowing safety period to calibrate the score.

[0057] Real-time feedback optimization: Conduct correlation analysis between the current score result and the attention heat map output by the swallowing abnormality pattern classifier.

[0058] Score grading and risk assessment: According to the final score, classify the swallowing function into normal swallowing, mild disorder, moderate disorder, and high-risk aspiration.

[0059] Advantages of the present invention:

[0060] Based on high-resolution acoustic sensors, combined with multi-channel signal synchronization, noise compensation, and multi-dimensional feature extraction technologies, the present invention can accurately capture the subtle changes during swallowing. Through multi-level signal decomposition and adaptive noise reduction processing, it can not only remove noise interference but also enhance the ability to capture swallowing transient features. Especially through the extraction of improved Mel cepstral coefficients and non-linear dynamics features, combined with deep learning models, the system can effectively identify swallowing disorder patterns in various stages such as the oral phase, pharyngeal phase, and esophageal phase, improving the accuracy and reliability of swallowing disorder detection.

[0061] Based on the traditional detection of dysphagia, the present invention adds a swallowing function scoring system based on multi-dimensional features. Through a dynamic weight scoring model, by combining the probability of swallowing abnormality and the degree of feature deviation, a standardized swallowing function score is generated. This score not only takes into account multiple clinical factors such as aspiration risk, pharyngeal phase delay, and esophageal dysfunction, but also can be optimized in real time according to clinical feedback, automatically adjusting the weight allocation of the scoring model to make the scoring result more in line with the individual's swallowing function status. It can automatically output a detailed scoring report, including stage-by-stage scoring, feature deviation maps, and risk hot zone markings, providing doctors with precise clinical intervention plan suggestions to help implement personalized treatment and rehabilitation training plans.

[0062] In the present invention, the dynamic weight scoring model can dynamically adjust the weight allocation at each stage by receiving the abnormality probability and feature deviation data output by the swallowing abnormality pattern classifier in real time and combining clinical feedback information, ensuring high precision and consistency of the scoring under different clinical conditions, and providing real-time and precise swallowing disorder scoring and personalized treatment suggestions for clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a schematic diagram of the system flow of the embodiment of the present invention;

[0065] Figure 2 It is a schematic diagram of the flow of the swallowing stage dynamic segmentation module of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The present invention will be described in detail below in combination with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0067] It should be noted that in the specification, references to "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes such specific features, structures, or characteristics. Additionally, when describing a specific feature, structure, or characteristic in connection with an embodiment, implementing such feature, structure, or characteristic in connection with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0068] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0069] As Figure 1 - Figure 2 shown, a dysphagia detection and automatic scoring system based on a high-resolution acoustic sensor includes a three-dimensional curved surface acoustic sensing array module, an adaptive environmental noise compensation module, a swallowing stage dynamic segmentation module, a multi-dimensional feature extraction engine, a swallowing abnormality pattern classifier, and a dynamic weight scoring model, where;

[0070] The three-dimensional curved surface acoustic sensing array module is configured with a 16-channel MEMS microphone matrix, which is arranged in a ring at a 2-mm spacing in the thyroid cartilage-cricoid cartilage complex area to output the original acoustic signal;

[0071] The adaptive environmental noise compensation module receives the original acoustic signal from the acoustic sensing array, performs a hybrid noise reduction process of wavelet packet decomposition and empirical mode decomposition, and outputs a high signal-to-noise ratio acoustic signal;

[0072] The swallowing stage dynamic segmentation module receives the high signal-to-noise ratio acoustic signal and generates staged signal data based on the sliding window energy entropy change detection and dynamic threshold segmentation algorithm;

[0073] The multi-dimensional feature extraction engine processes the staged signal data, synchronously extracts time-frequency domain features, non-linear dynamics features, and phase coupling features, and generates a multi-dimensional feature set;

[0074] The swallowing abnormality pattern classifier receives the multi-dimensional feature set and outputs the probability of swallowing function abnormality through a bidirectional long short-term memory network with an attention mechanism;

[0075] The dynamic weight scoring model generates a standardized swallowing function score based on the abnormality probability and the feature deviation degree.

[0076] The three-dimensional curved surface acoustic sensing array module includes:

[0077] Construction of bionic laryngeal curved surface substrate: A three-dimensional curved surface substrate adapted to the laryngeal anatomical structure is constructed using flexible materials. 16 precision positioning grooves are set on the surface, and a dynamic fitting compensation mechanism is integrated, specifically as follows:

[0078] A three-dimensional curved surface substrate matching the anatomical structure of the human thyroid cartilage-cricoid cartilage composite area is prepared using flexible polyimide materials. The radius of curvature of the substrate is dynamically adjusted from 15 to 25 mm through an internal shape memory alloy wire. 16 equally spaced positioning grooves with a depth of 0.5 mm are set on the surface of the substrate. Micro spring contacts are integrated at the bottom of the grooves to provide a displacement compensation amount of ±0.3 mm when the pressure on the laryngeal skin surface changes, ensuring the dynamic close fitting of the sensor array with the skin surface;

[0079] Arrangement of multi-channel MEMS microphone array: 16 MEMS microphone units are distributed in a spiral ring shape, and each unit automatically adjusts to be perpendicular to the skin surface through a micro universal joint, specifically as follows:

[0080] The 16 MEMS microphone units are embedded in the substrate positioning grooves in a clockwise spiral ring array. The center distance between adjacent units is strictly controlled within a tolerance range of 2 mm ± 0.1 mm through laser micro positioning technology. Among them, the 1st - 8th channels form a semicircular distribution along the upper edge of the thyroid cartilage, covering the area 3 - 5 mm above the laryngeal prominence. The 9th - 16th channels form a closed ring array along the lower edge of the cricoid cartilage. The normal direction of the receiving surface of each unit is automatically adjusted to be perpendicular to the laryngeal surface through a micro universal joint mechanism, and the tilt angle compensation range is ±15°;

[0081] Temperature-pressure coupling compensation mechanism: Each microphone unit is equipped with a micro piezoresistive sensor to monitor pressure changes and perform dynamic gain compensation when it exceeds the preset range. A temperature sensor array is provided in the center of the substrate, and a temperature compensation algorithm is used to correct the signal baseline drift, specifically as follows:

[0082] A micro piezoresistive sensor is embedded at the bottom of each MEMS microphone unit to monitor the pressure fluctuation of the laryngeal skin contact in real time. When the detected pressure value exceeds the preset range of 2 - 5 N / cm 2 a dynamic gain compensation circuit is triggered to perform an amplitude correction of 0.5 - 3 dB on the corresponding channel signal. A 4×4 high-precision temperature sensor array is integrated in the center of the substrate to collect local epidermal temperature data with a resolution of 0.1 °C. A compensation algorithm with a temperature coefficient of -0.15% / °C is used to eliminate the signal baseline drift caused by thermal expansion, and the compensation accuracy reaches ±0.05 mV / °C;

[0083] Multi-modal signal synchronous acquisition: Each microphone is equipped with a pre-amplification circuit that provides programmable gain. The original acoustic signals are synchronously acquired using an analog-to-digital converter. The spatial coordinate data of the marker points are acquired by an external infrared camera, and a three-dimensional spatial mapping relationship between the sensor array and the laryngeal anatomical landmark points is established as follows:

[0084] Each MEMS microphone unit is connected to an independent pre-amplification circuit, which provides two levels of programmable gain modes of 40 dB and 60 dB, automatically switches the amplification factor according to the signal strength, and uses a 24-bit analog-to-digital converter to synchronously acquire 16-channel acoustic signals at a sampling rate of 48 kHz. The synchronous clock deviation is less than 10 ns. Three groups of infrared reflective marker points are set at the base edge, and the spatial coordinate data of the marker points are acquired by an external infrared camera at a frequency of 120 Hz. A three-dimensional spatial mapping relationship between the sensor array and the laryngeal anatomical landmark points is established, and the positioning accuracy reaches 0.1 mm;

[0085] Dynamic fitting calibration: Start the automatic calibration program. The base automatically adjusts to fit according to the user's laryngeal contour, generates a pressure distribution map using pressure feedback, closes the corresponding microphone channel when the set pressure is exceeded, and outputs the channel availability status code after completion as follows:

[0086] Before each detection, start the automatic calibration program. Inject controllable air pressure into the base sandwich through a micro air pump to make the curved surface adaptively match the current user's laryngeal contour within 5 seconds. Generate a thermal map of the contact pressure distribution using the data feedback from the piezoresistive sensor. When the local pressure is detected to exceed 8 N / cm 2 automatically close the microphone channels in the corresponding area, reconstruct the three-dimensional sound field acquisition matrix according to the spatial distribution of the remaining effective channels, and output the channel availability status code when the calibration is completed;

[0087] Signal preprocessing and transmission: Perform high-pass filtering, differential amplification, and digital processing on the original acoustic signals, and transmit the processed signals through a flexible cable using the I2S protocol, and carry temperature, pressure, and spatial data to form a meta-data packet as follows:

[0088] Integrate a multi-stage signal conditioning circuit inside the base. Perform high-pass filtering with a cut-off frequency of 300 Hz on the original acoustic signals to eliminate vascular pulsation interference, perform differential amplification processing with a common-mode rejection ratio greater than 120 dB, and finally perform digital data encapsulation based on the I2S protocol. The processed multi-channel signals are transmitted to the adaptive ambient noise compensation module through a flexible flat cable, and at the same time carry temperature compensation parameters, pressure correction coefficients, and spatial coordinate mapping data to form a meta-data packet.

[0089] The adaptive ambient noise compensation module includes:

[0090] Multi-channel signal spatio-temporal alignment: Receive the original acoustic signal and the spatial coordinate data of infrared marker points, calculate the signal time offset based on the time delay model, and use the interpolation algorithm for synchronous alignment to generate the spatio-temporal calibration signal, as follows:

[0091] Receive the 16-channel original acoustic signal and the spatial coordinate data of infrared marker points output by the three-dimensional curved surface acoustic sensing array module, calculate the time offset of each channel signal based on the inter-channel transmission time delay model, use the interpolation resampling algorithm to achieve microsecond-level synchronous alignment, with a synchronous error less than 50 μs, and generate the spatio-temporal calibration signal;

[0092] Hybrid multi-scale signal decomposition: Perform wavelet packet decomposition (db8 wavelet basis, 8 frequency bands) and improved empirical mode decomposition on the calibration signal to generate wavelet packet coefficients and intrinsic mode components, and construct the time-frequency representation space, as follows:

[0093] Parallelly perform three-level wavelet packet decomposition and improved empirical mode decomposition on the calibrated signals of each channel. Among them, the wavelet packet decomposition uses the db8 wavelet basis function to divide the signal into 8 frequency subspaces, and the empirical mode decomposition generates 5 intrinsic mode components by injecting white noise with 0.2 times the standard deviation, and constructs a hybrid time-frequency representation space containing the wavelet packet coefficient matrix and the mode component tensor.

[0094] Dynamic noise feature library construction: Combine temperature, pressure, and environmental noise baseline to construct a noise template library, calculate the frequency band energy entropy value based on a sliding time window, identify the noise dominant area and establish a noise feature fingerprint, as follows:

[0095] Call the temperature and pressure compensation parameters of the three-dimensional curved surface acoustic sensing array module, combine the current environmental noise baseline to construct a dynamic noise template library, calculate the frequency band energy entropy value through a sliding time window (200 ms), identify the frequency bands with an energy entropy mutation exceeding the threshold of ±15% as the noise dominant area, and establish a noise feature fingerprint containing the respiration interference spectrum and the environmental vibration spectrum.

[0096] Hierarchical adaptive noise reduction processing: Perform differential noise reduction processing on signals in different frequency bands. Retain the swallowing characteristics for high-frequency signals, remove noise-related components for medium-frequency signals, and eliminate respiration noise for low-frequency signals through an LMS adaptive filter, as follows:

[0097] Perform differential processing on the signal components after hybrid decomposition. For high-frequency components (>8 kHz), use the wavelet coefficient shrinkage method with an improved threshold function to retain the swallowing transient characteristics. For medium-frequency components (1 - 8 kHz), screen and remove the mode components with a correlation coefficient >0.7 based on the noise feature library. For low-frequency components (<1 kHz), activate a 16th-order LMS adaptive filter bank and perform active noise cancellation using the respiration reference signal generated from the pressure sensor data.

[0098] Multimodal signal reconstruction optimization: The denoised signal is reconstructed, the high-frequency part is restored by inverse wavelet transform, the intermediate-frequency part is restored by superposition, and the low-frequency noise is removed by a zero-phase FIR filter to generate a high signal-to-noise ratio acoustic signal, as follows:

[0099] The components of each frequency band after denoising are fused and reconstructed, and the inverse wavelet packet transform reconstruction retains the high-frequency characteristics. The filtered intrinsic modes are superimposed to restore the intermediate frequency components. The residual noise is filtered out twice through a zero-phase FIR filter (passband 300Hz-12kHz) to generate a high signal-to-noise ratio acoustic signal with complete time-frequency characteristics.

[0100] The swallowing phase dynamic segmentation module includes:

[0101] Multi-channel signal consistency check: Receive high signal-to-noise ratio acoustic signals and channel availability status codes, calculate the peak value of the cross-correlation function of the signals between valid channels, and remove abnormal channel data with a correlation coefficient less than 0.85. Normalize the retained channels and generate standardized signals based on the maximum energy channel;

[0102] Dynamic energy entropy change feature extraction: A 50ms Hamming window is used to slide the standardized signal with a step size of 10ms. The time-frequency distribution matrix of the Teager energy operator of the signal in each time window is calculated. The time-varying Shannon entropy curve is constructed based on the energy probability density function. The second-order derivative extreme points of the entropy value curve are extracted as candidate segmentation marks through the cubic spline interpolation method.

[0103] Adaptive threshold model construction: extract the energy entropy statistical characteristics of the signal in the pre-swallowing resting period (5 seconds before the start of the test), calculate the mean and standard deviation, establish the initial segmentation benchmark of the dynamic threshold T, and use the exponential moving average algorithm to update the threshold in combination with the real-time signal characteristics, constrain the fluctuation range to ±0.3σ, and correct the abnormal threshold jump by making the thyroid cartilage displacement less than 8mm / s;

[0104] The dynamic threshold is expressed as:

[0105] T = μ + 1.5σ;

[0106] Among them, μ and σ are the mean and standard deviation of the signal respectively, and T is the dynamic threshold;

[0107] The exponential moving average update is expressed as: T new =α·T new +(1-α)·T old

[0108] Among them, α is the exponential smoothing coefficient, T new is the updated dynamic threshold, T old is the dynamic threshold calculated last time;

[0109] Multi-stage joint segmentation decision: Input the candidate segmentation markers into a multi-channel voting decision mechanism. If the energy entropy transition amplitude exceeds 120% of the threshold T in ≥3 adjacent channels within a 50-ms time window, it is determined as the starting point of the oral preparatory phase. The pharyngeal phase trigger requires that at least 5 channels reach 80% of the maximum energy peak and the duration is greater than 150 ms. The detection of the esophageal phase termination point uses a dual-condition constraint (the energy drops back to ±10% of the baseline and remains stable for more than 300 ms);

[0110] Energy transition amplitude determination:

[0111]

[0112] Among them, E max is the signal peak energy, and E min is the minimum energy.

[0113] Energy peak condition: E peak ≥0.8·E max ;

[0114] Among them, E peak is the peak energy of the current channel, and R max is the maximum energy of all channels;

[0115] Physiological rationality verification: Conduct a biomechanical verification on the preliminary segmentation results, limit the durations of the oral phase, pharyngeal phase, and esophageal phase, as well as the consistency of the energy gradient direction at the phase transition points. Exclude false segmentation points through the multi-channel phase synchronization index, as follows:

[0116] Limit the duration of the oral phase to 0.3 - 1.2 s and the pharyngeal phase to 0.5 - 1.5 s. When out of range, initiate backtracking and re-segmentation; Check the consistency of the energy gradient direction at the phase transition points (positive gradient is required from the oral phase to the pharyngeal phase, and negative gradient is required from the pharyngeal phase to the esophageal phase); Exclude false segmentation points through the multi-channel phase synchronization index (PSI > 0.7).

[0117] The phase synchronization index (PSI) is expressed as:

[0118] Among them, φ i and φ j are the phases of the i-th and j-th channels, and N is the number of channels;

[0119] Stage-based data encapsulation and output: Generate a structured data packet including accurate timestamps (±5-ms accuracy), energy transition amplitudes, the number of activated channels, and physiological verification markers, and output the stage-based signal data to a multi-dimensional feature extraction engine. The data format includes a three-segment time domain marker for the oral preparatory phase (0 - 1), pharyngeal phase (1 - 2), and esophageal phase (2 - 3) and the corresponding frequency domain energy distribution map.

[0120] The multi-dimensional feature extraction engine includes:

[0121] Effective channel dynamic optimization: Based on the number of channel activations and physiological verification markers in the phased signal data, calculate the correlation coefficient of each channel with the reference template, and select the top 8 channels with a correlation coefficient greater than 0.9 as effective channels for signal enhancement processing;

[0122] Time-frequency domain feature analysis: For each swallowing stage signal, use Mel cepstral coefficients to extract time-frequency features, and calculate the wavelet packet energy spectrum entropy at the same time to generate time-frequency domain features, specifically as follows:

[0123] Parallelly perform Mel cepstral coefficient extraction on each swallowing stage signal, set a 26-channel triangular filter bank to cover the 0 - 12 kHz frequency band, and extract the first 15-order cepstral coefficients; synchronously calculate the wavelet packet energy spectrum entropy value, perform 5-layer decomposition using the db10 wavelet basis, calculate the energy proportion in the γ frequency band (30 - 60 Hz) and the θ frequency band (4 - 8 Hz) respectively, and generate time-frequency domain features including time-domain envelope features, frequency-domain energy distribution features, and cross-frequency band correlation features;

[0124] Mel cepstral coefficients:

[0125]

[0126] Among them, MFCC k is the k-th order Mel cepstral coefficient, X n is the frequency domain value of the signal, N is the number of Mel filters, and k is the order of the cepstral coefficient (such as the first 15 orders);

[0127] Nonlinear dynamics feature mining: Perform recurrence quantification analysis on the pharyngeal phase signal segment, calculate deterministic, laminarity, and recurrence entropy parameters, and calculate the maximum Lyapunov exponent at the same time to track chaotic characteristics and generate nonlinear dynamics features, specifically as follows:

[0128] Perform recurrence quantification analysis on the key signal segment of the pharyngeal phase (phase markers 1 - 2), set the time delay τ = 8 ms and the embedding dimension m = 5, and calculate the deterministic (DET > 85%), laminarity, and recurrence entropy parameters; synchronously calculate the maximum Lyapunov exponent, and use the small data method to track the evolution of chaotic characteristics with a time resolution of 0.5 ms to generate nonlinear dynamics features.

[0129] Cross-channel phase coupling analysis: Calculate the bispectral coherence of the pharyngeal phase trigger period, analyze the phase synchronization strength between channels, and extract phase coupling features, specifically as follows:

[0130] Calculate the bispectral coherence during the pharyngeal trigger period (phase markers 1.2 - 1.8), select the laryngeal motion characteristic frequency band (800 - 1200 Hz) as the main frequency band, and analyze the inter-channel phase difference distribution; adopt the improved phase locking value (PLV) algorithm, slide the 20 ms time window to calculate the phase synchronization strength between channels 3 - 7 and channels 9 - 12, and extract the pharyngeal coupling strength index and the phase modulation depth parameter;

[0131] The calculation of bispectral coherence is as follows:

[0132] where Γ(ω) is the bispectral coherence, S xy (ω) is the cross-spectral density between signal x and signal y, is the auto-spectral density of signal x, S yy (ω) is the auto-spectral density of signal y,

[0133] The calculation of phase synchronization strength (PLV) is as follows:

[0134]

[0135] where PLV is the phase synchronization strength, φ n is the phase of signal n, N 1 is the number of signal channels;

[0136] The calculation of the pharyngeal coupling strength index (PCSI) is as follows:

[0137]

[0138] where PCSI is the pharyngeal coupling strength index, PLV ij is the phase synchronization strength between the i-th and j-th channels.

[0139] The multi-dimensional feature extraction engine also includes:

[0140] Multi-modal feature fusion: Normalize the time-frequency domain features, non-linear dynamics features, and phase coupling features, use principal component analysis for dimensionality reduction, retain the principal components with a cumulative contribution rate greater than 95%, and perform weighted fusion according to the feature importance of different stages to generate a 128-dimensional fusion feature vector, specifically as follows:

[0141] Normalize the time-frequency domain features, non-linear dynamics features, and phase coupling features, use principal component analysis to reduce the feature dimensions, and retain the principal components with a cumulative contribution rate > 95%; perform dynamic weighted fusion on the cross-stage features, assign 60% weight to the time-frequency features in the oral preparation period, strengthen the non-linear features in the pharyngeal period (70% weight), and focus on the phase coupling features in the esophageal period (65% weight) to generate a 128-dimensional fusion feature vector.

[0142] Dynamic feature optimization mechanism: Based on the feedback results of the swallowing abnormality pattern classifier, the random forest algorithm is used to rank the feature importance, and redundant features with an importance score lower than 0.05 are removed;

[0143] Calculation of random forest feature importance:

[0144]

[0145] where I j is the importance score of the jth feature, M is the number of decision trees, Impurity beforesplit m is the impurity before splitting, and Impurity after split m is the impurity after splitting; Removal of redundant features, if I j <0.05, then this feature is removed;

[0146] Feature data encapsulation and output: Generate a multi-dimensional feature set containing timestamp alignment information, feature type markers, and data quality indicators. The output format is a four-layer nested structure, including stage markers, time-frequency features, non-linear features, phase features, and fusion weights. The feature numerical precision is retained to four decimal places and transmitted to the swallowing abnormality pattern classifier through a high-speed data bus.

[0147] The swallowing abnormality pattern classifier includes:

[0148] Multi-dimensional feature preprocessing: Receive the multi-dimensional feature set output by the multi-dimensional feature extraction engine, perform timestamp alignment verification, fill in the missing feature values caused by channel failures (using the adjacent time window feature interpolation method), and reorganize the feature sequence into an oral preparatory phase feature matrix (T×32 dimensions), a pharyngeal phase feature tensor (T×64 dimensions), and an esophageal phase feature vector (T×32 dimensions) according to the swallowing stage division. After completing the feature space normalization, input it into the neural network architecture;

[0149] Construction of a bidirectional long short-term memory network-attention network: Deploy a core architecture containing three layers of bidirectional long short-term memory networks (Bi-LSTM), with the number of hidden layer nodes in each layer being 128 / 64 / 32 respectively. Add a temporal convolutional layer (TCN, convolution kernel size 5, stride 2) to the pharyngeal phase feature processing branch, and integrate a multi-head attention mechanism (4 attention heads) at the top of the network. The attention weight calculation uses the scaled dot product method, focusing on the feature mutation patterns within the 300ms time window before and after pharyngeal phase triggering;

[0150] Temporal feature enhancement processing: Extract local dependence features through the temporal convolutional layer, combine with the hidden state of the bidirectional long short-term memory network to generate an enhanced temporal feature vector, and use a gating mechanism to dynamically adjust the fusion weight of convolutional features and recurrent features, specifically as follows:

[0151] Extract the local dependence features of the pharyngeal phase signal through a temporal convolutional layer. The convolutional kernel covers a time span of 50 ms. The output feature map is concatenated with the hidden state of the Bi-LSTM at the channel level to generate an enhanced temporal feature vector. A gating mechanism is used to dynamically adjust the fusion weight of the convolutional features and the recurrent features (range 0.3 - 0.7) to form a hybrid feature representation with multi-scale time perception ability;

[0152] Dynamic allocation of attention weights: Perform cross-stage feature correlation analysis in the multi-head attention layer, calculate the cross-attention scores between the oral preparation phase and pharyngeal phase features, set the time attention weight benchmark value for the pharyngeal phase trigger stage to 0.6, and automatically increase the attention weight of the 200 ms period after the pharyngeal phase to above 0.8 when abnormal esophageal phase features are detected to generate an attention heat map with clinical interpretability;

[0153] Implementation of the multi-task learning framework: Output three types of anomaly probabilities in parallel, including the aspiration risk probability, pharyngeal phase delay probability, and esophageal dysfunction probability. Each sub-task is configured with an independent fully connected layer (256→128→1 node), and the dynamic weighted cross-entropy is used as the loss function (aspiration risk weight 1.5, other weights 1.0), and cross-task knowledge transfer is achieved through the feature sharing layer;

[0154] Fusion decision of anomaly probabilities: Perform weighted fusion on the multi-task outputs according to clinical rules, calculate the final swallowing function anomaly probability. When multi-stage concurrent anomalies occur (such as high aspiration risk and pharyngeal phase delay), activate the emergency warning coefficient to generate a swallowing function anomaly probability value in the range of 0 - 1, specifically as follows:

[0155] Perform weighted fusion on the multi-task outputs based on clinical rules. The aspiration risk probability accounts for 60% of the final score, the pharyngeal phase delay accounts for 30%, and the esophageal dysfunction accounts for 10%. When multi-stage concurrent anomalies are detected, activate the emergency warning coefficient to generate a standardized swallowing function anomaly probability value in the range of 0 - 1, with the precision reserved to three decimal places.

[0156] The dynamic weight scoring model includes:

[0157] Data preparation: Receive the swallowing function anomaly probabilities output from the swallowing anomaly pattern classifier and the feature deviation data obtained from the multi-dimensional feature extraction engine, where;

[0158] The swallowing function anomaly probabilities include the aspiration risk probability, pharyngeal phase delay probability, and esophageal dysfunction probability;

[0159] The feature deviation data is the feature deviation of each swallowing stage, including the cepstrum deviation value of the Mel cepstral coefficients in the oral preparation phase, the non-linear dynamics deviation rate in the pharyngeal phase, and the decay slope of the phase synchronization index in the esophageal phase;

[0160] Feature deviation calculation and standardization: According to the feature deviation calculation formula for each stage of swallowing, calculate the deviation coefficient for each stage respectively, and perform standardization processing to normalize the deviation value to the range of 0-1. The deviation calculation formula is as follows:

[0161]

[0162] Among them, D oral , D pharyngeal , D esophageal are the feature deviations of each stage, MFCC is the MFCC cepstral feature vector, MFCC template is the standard template, λ max (t) is the maximum Lyapunov exponent, representing the change of nonlinear dynamics, and PCSI(t) is the phase coupling synchronization index;

[0163] Intelligent allocation of clinical weights: According to the clinical needs of different swallowing stages, dynamically allocate the clinical weights for each stage. The initial weights are set as follows:

[0164] The initial value of the time-domain feature weight in the oral preparatory phase is 0.6. If the swallowing start delay exceeds 500 ms, the weight is adjusted to 0.8;

[0165] The basic weight of the nonlinear feature in the pharyngeal phase is 0.7. If laryngeal closure insufficiency (phase coupling index lower than 0.5) occurs, the weight is adjusted to 1.2;

[0166] The weight of the esophageal phase is dynamically adjusted based on the predicted value of the residual food volume, ranging from 0.4 to 0.9;

[0167] Multi-dimensional scoring fusion calculation: According to the abnormal probability, feature deviation, and real-time weight, use the weighted fusion algorithm to generate the final swallowing function score. The score calculation formula is as follows:

[0168]

[0169] Among them, P final is the final abnormal probability of swallowing function, taken from the output of the swallowing abnormal pattern classifier, D final is the standardized feature deviation, after standardization processing, W final is the real-time weight calculated according to the intelligent allocation of clinical weights, and S is the final generated swallowing function score. The score range is 0 to 100, and the precision is reserved to one decimal place.

[0170] The dynamic weight scoring model also includes.

[0171] Physiological constraint calibration mechanism: Introduce the physiological boundary conditions of the swallowing safety period to calibrate the score. The specific rules are as follows:

[0172] If the duration of the pharyngeal phase is less than 300 ms, 15 points will be deducted.

[0173] If the degree of complete laryngeal closure is less than 70%, the upper limit of 20 points deduction will be triggered.

[0174] If the esophageal efficiency is greater than 85%, a reward of 5 - 10 points will be given.

[0175] The final scoring result is ensured to be within the valid range of 0 - 100.

[0176] Real - time feedback optimization: Correlate the current scoring result with the attention heat map output by the swallowing abnormality pattern classifier.

[0177] When the influence degree of a specific feature dimension (such as the phase coupling index) on the score exceeds 30%, automatically increase the allocation ratio of this feature in the weight matrix by up to 50%. This process will be dynamically adjusted after every 50 samples are processed.

[0178] Scoring classification and risk assessment: According to the final score, the swallowing function is divided into normal swallowing, mild disorder, moderate disorder, and high - risk aspiration, as follows:

[0179] Grade I (80 - 100 points): Normal swallowing;

[0180] Grade II (60 - 79 points): Mild disorder;

[0181] Grade III (40 - 59 points): Moderate disorder;

[0182] Grade IV (<40 points): High - risk aspiration.

[0183] Each grade corresponds to different clinical intervention plan suggestions, including the recommended urgency of VFSS examination and the type of rehabilitation training.

[0184] This invention covers any substitutions, modifications, equivalent methods, and schemes made within the essence and scope of this invention. For the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of this invention, well - known methods, processes, procedures, components, and circuits are not described in detail.

[0185] The above are only the preferred embodiments of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.

Claims

1. A swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensors, characterized in that: It includes a three-dimensional curved acoustic sensor array module, an adaptive environmental noise compensation module, a swallowing stage dynamic segmentation module, a multi-dimensional feature extraction engine, a swallowing abnormality pattern classifier and a dynamic weight scoring model, among which; The three-dimensional curved acoustic sensor array module is configured with a 16-channel MEMS microphone matrix, which is arranged in a ring with a spacing of 2 mm in the thyroid cartilage-cricoid cartilage complex area to output the original acoustic signal; The adaptive environmental noise compensation module receives the original acoustic signal of the acoustic sensor array, performs a hybrid noise reduction process of wavelet packet decomposition and empirical mode decomposition, and outputs a high signal-to-noise ratio acoustic signal; The swallowing stage dynamic segmentation module receives high signal-to-noise ratio acoustic signals and generates staged signal data based on sliding window energy entropy change detection and dynamic threshold segmentation algorithm; The multi-dimensional feature extraction engine processes the phased signal data, synchronously extracts time-frequency domain features, nonlinear dynamic features, and phase coupling features, and generates a multi-dimensional feature set; The swallowing abnormality pattern classifier receives a multi-dimensional feature set and outputs the probability of swallowing abnormality through a bidirectional long short-term memory network with an attention mechanism; The dynamic weight scoring model generates a standardized swallowing function score according to the abnormal probability and the characteristic deviation.

2. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 1 is characterized in that: The three-dimensional curved surface acoustic sensor array module comprises: Bionic laryngeal curved surface base construction: Flexible materials are used to construct a three-dimensional curved surface base that adapts to the anatomical structure of the larynx. 16 positioning grooves are set on the surface, and a dynamic fitting compensation mechanism is integrated; Multi-channel MEMS microphone array arrangement: 16 MEMS microphone units are arranged in a spiral ring, and each unit is automatically adjusted to be perpendicular to the skin surface through a micro universal joint; Temperature-pressure coupling compensation mechanism: Each microphone unit is equipped with a micro piezoresistive sensor to monitor pressure changes and perform dynamic gain compensation when it exceeds the preset range. A temperature sensor array is set in the center of the substrate to correct signal baseline drift using a temperature compensation algorithm. Multimodal signal synchronous acquisition: Each microphone is equipped with a preamplifier circuit, which provides programmable gain, uses an analog-to-digital converter to synchronously acquire the original acoustic signal, collects the spatial coordinate data of the marker points through an external infrared camera, and establishes a three-dimensional spatial mapping relationship between the sensor array and the laryngeal anatomical landmarks; Dynamic fit calibration: Start the automatic calibration procedure, the base automatically adjusts the fit according to the user's throat contour, uses pressure feedback to generate a pressure distribution map, and closes the corresponding microphone channel when the set pressure is exceeded. After completion, the channel availability status code is output; Signal preprocessing and transmission: The original acoustic signal is subjected to high-pass filtering, differential amplification and digital processing. The processed signal is transmitted through a flexible cable using the I2S protocol, and carries temperature, pressure and spatial data to form a metadata package.

3. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 2 is characterized in that: The adaptive environmental noise compensation module comprises: Multi-channel signal time-space alignment: Receive the original acoustic signal and the spatial coordinate data of the infrared marker point, calculate the signal time offset based on the time delay model, use the interpolation algorithm to synchronize and align, and generate the time-space calibration signal; Hybrid multi-scale signal decomposition: perform wavelet packet decomposition and improved empirical mode decomposition on the calibration signal to generate wavelet packet coefficients and intrinsic mode components, and construct a time-frequency representation space; Dynamic noise feature library construction: Combine temperature, pressure and environmental noise baseline to build a noise template library, calculate the frequency band energy entropy value based on the sliding time window, identify the noise-dominated area and establish the noise feature fingerprint; Hierarchical adaptive noise reduction processing: Differentiated noise reduction processing is performed on signals in different frequency bands. High-frequency signals retain swallowing characteristics, medium-frequency signals remove noise-related components, and low-frequency signals use LMS adaptive filters to eliminate breathing noise. Multimodal signal reconstruction optimization: The denoised signal is reconstructed, the high-frequency part is restored by inverse wavelet transform, the intermediate-frequency part is restored by superposition, and the low-frequency noise is removed by a zero-phase FIR filter to generate a high signal-to-noise ratio acoustic signal.

4. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 3 is characterized in that: The swallowing stage dynamic segmentation module includes: Multi-channel signal consistency check: Receive high signal-to-noise ratio acoustic signals and channel availability status codes, calculate the peak value of the cross-correlation function of the signals between valid channels, remove abnormal channel data with a correlation coefficient less than 0.85, normalize the retained channels, and generate standardized signals based on the maximum energy channel; Dynamic energy entropy change feature extraction: A 50ms Hamming window is used to slide the standardized signal with a step size of 10ms. The time-frequency distribution matrix of the Teager energy operator of the signal in each time window is calculated. The time-varying Shannon entropy curve is constructed based on the energy probability density function. The second-order derivative extreme points of the entropy value curve are extracted as candidate segmentation marks through the cubic spline interpolation method. Adaptive threshold model construction: extract the energy entropy statistical characteristics of the pre-swallowing resting period signal, calculate the mean and standard deviation, establish the initial segmentation benchmark of the dynamic threshold T, and use the exponential moving average algorithm to update the threshold in combination with the real-time signal characteristics; Multi-stage joint segmentation decision: The candidate segmentation marks are input into the multi-channel voting decision mechanism. If ≥3 adjacent channels simultaneously detect that the energy entropy transition amplitude exceeds 120% of the threshold T within the 50ms time window, it is determined to be the start point of the oral preparation phase. The pharyngeal phase triggering needs to meet the requirement that at least 5 channels have energy peaks reaching 80% of the maximum value and the maintenance time is greater than 150ms. The esophageal phase termination point detection adopts double condition constraints. Physiological plausibility verification: biomechanical verification of the preliminary segmentation results, limiting the duration of the oral, pharyngeal and esophageal phases, as well as the consistency of the energy gradient direction at the phase transition points, and excluding pseudo-segmentation points through the multi-channel phase synchronization index; Phased data packaging output: Generates structured data packets including timestamp, energy transition amplitude, channel activation number, and physiological verification mark, and outputs phased signal data to the multi-dimensional feature extraction engine.

5. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 4 is characterized in that: The multi-dimensional feature extraction engine comprises: Dynamic optimization of effective channels: Based on the number of channel activations and physiological verification markers in the phased signal data, the correlation coefficient between each channel and the reference template is calculated, and the first 8 channels with a correlation coefficient greater than 0.9 are selected as effective channels for signal enhancement processing; Time-frequency domain feature analysis: For each swallowing stage signal, the Mel-frequency cepstral coefficients are used to extract the time-frequency features, and the wavelet packet energy spectrum entropy is calculated to generate the time-frequency domain features; Mining of nonlinear dynamic features: recursive quantitative analysis of pharyngeal signal segments, calculation of deterministic, laminar and recursive entropy parameters, and calculation of the maximum Lyapunov exponent, tracking chaotic characteristics, and generating nonlinear dynamic features; Cross-channel phase coupling analysis: Calculate the bispectral coherence of the pharyngeal trigger period, analyze the phase synchronization strength between channels, and extract phase coupling characteristics.

6. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 5 is characterized in that: The multi-dimensional feature extraction engine also includes: Multimodal feature fusion: normalize the time-frequency domain features, nonlinear dynamic features, and phase coupling features, use principal component analysis to reduce the dimension, retain the principal components with a cumulative contribution rate greater than 95%, and perform weighted fusion according to the importance of features at different stages to generate a 128-dimensional fusion feature vector; Dynamic feature selection mechanism: Based on the feedback results of the swallowing abnormality pattern classifier, the random forest algorithm is used to sort the feature importance and eliminate redundant features with importance scores lower than 0.05; Feature data encapsulation output: Generates a multi-dimensional feature set containing timestamp alignment information, feature type tags, and data quality indicators.

7. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 6 is characterized in that: The abnormal swallowing pattern classifier comprises: Multi-dimensional feature preprocessing: Receive the multi-dimensional feature set output by the multi-dimensional feature extraction engine, perform timestamp alignment verification, fill in the missing feature values ​​caused by channel failure, and reorganize the feature sequence into the oral preparation phase feature matrix, pharyngeal phase feature tensor, and esophageal phase feature vector according to the swallowing stage. After completing the feature space standardization, input it into the neural network architecture; Bidirectional LSTM-Attention Network Construction: Deploy a core architecture consisting of a three-layer bidirectional LSTM network, add a temporal convolution layer to the pharyngeal feature processing branch, and integrate a multi-head attention mechanism at the top of the network; Temporal feature enhancement processing: extract local dependency features through the temporal convolution layer, combine the hidden state of the bidirectional long short-term memory network to generate an enhanced temporal feature vector, and use the gating mechanism to dynamically adjust the fusion weight of the convolutional feature and the cyclic feature; Dynamic allocation of attention weights: Perform cross-stage feature correlation analysis in the multi-head attention layer, calculate the cross-attention scores of the oral preparation phase and the pharyngeal phase features, set the time attention weight baseline value of the pharyngeal phase trigger phase to 0.6, and automatically increase the attention weight of the 200ms period after the pharyngeal phase to above 0.8 when abnormal esophageal phase features are detected, and generate a clinically interpretable attention heat map; Implementation of the multi-task learning framework: parallel output of three types of abnormal probabilities, including aspiration risk probability, pharyngeal delay probability, and esophageal dysfunction probability, using dynamic weighted cross entropy as the loss function; Abnormal probability fusion decision: Weighted fusion of multi-task outputs is performed according to clinical rules to calculate the final probability of abnormal swallowing function.

8. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 7, characterized in that: The dynamic weight scoring model includes: Data preparation: receiving the swallowing function abnormality probability output from the swallowing abnormality pattern classifier and the feature deviation data obtained from the multi-dimensional feature extraction engine, wherein; The probability of abnormal swallowing function includes the probability of aspiration risk, the probability of delayed pharyngeal phase and the probability of esophageal dysfunction; The characteristic deviation data of each swallowing stage includes the Mel-frequency cepstral coefficient cepstral deviation value of the oral preparation period, the nonlinear dynamic deviation rate of the pharyngeal period, and the phase synchronization exponential decay slope of the esophageal period. Characteristic deviation calculation and standardization: According to the characteristic deviation calculation formula of each stage of swallowing, the deviation coefficient of each stage is calculated and standardized; Intelligent allocation of clinical weights: Dynamically allocate clinical weights for each stage based on clinical needs at different swallowing stages; Multi-dimensional score fusion calculation: Based on the abnormal probability, feature deviation and real-time weight, a weighted fusion algorithm is used to generate the final swallowing function score.

9. The swallowing disorder detection and automatic scoring system based on high-resolution acoustic sensor according to claim 8, characterized in that: The dynamic weight scoring model also includes: Physiological constraint calibration mechanism: introduce physiological boundary conditions of the swallowing safety period and calibrate the score; Real-time feedback optimization: Correlation analysis is performed between the current scoring result and the attention heat map output by the swallowing abnormality pattern classifier; Scoring and risk assessment: Based on the final score, swallowing function is divided into normal swallowing, mild disorder, moderate disorder and high-risk aspiration.

Citation Information

Cited By

  • Automatic braised chicken packaging and detecting method based on machine vision

    CN120445318A

  • Multi-modal sensing fusion swallowing rehabilitation evaluation system and method

    CN120899193A

  • Preservation degree monitoring system for picked green bamboo shoots

    CN121347513A

  • Medical data multi-dimensional perception and integrated processing system

    CN121601128A

  • Network anomaly monitoring method and device based on flow fingerprint learning, and program product

    CN121907616A