Portable swallowing function screening system based on multiple modes

Through multimodal sensing fusion and deep learning technology, accurate and rapid identification of portable swallowing function screening system in intensive care units was achieved, solving the portability, real-time and accuracy problems of swallowing function screening in existing technologies and improving screening efficiency and accuracy.

CN120678384APending Publication Date: 2025-09-23广州医科大学附属清远医院(清远市人民医院)
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510620870.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies are unable to achieve portable, real-time monitoring and early warning swallowing function screening in intensive care units. Traditional methods are highly invasive, complex to operate, have a high misjudgment rate, cannot accurately reflect the complex biomechanical changes during swallowing, and lack real-time interaction with electronic medical record systems.

Method used

Using multimodal sensing fusion and deep learning technology, the signal acquisition device collects laryngeal pressure, laryngeal acoustic vibration and swallowing sound wave signals. Combined with feature extraction, feature fusion and deep learning models, the risk probability of swallowing disorders is calculated, and real-time display and voice feedback are provided through an interactive display device.

Benefits of technology

It has achieved accurate and rapid identification of swallowing disorders, significantly improved the efficiency and accuracy of swallowing function screening, and filled the gap in intelligent screening equipment for clinical swallowing assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120678384A_ABST
    Figure CN120678384A_ABST
Patent Text Reader

Abstract

The invention discloses a portable swallowing function screening system based on multiple modes. The portable swallowing function screening system comprises a signal acquisition device, a data processing device and an interactive display device which are connected in sequence, the signal acquisition device is used for acquiring a laryngeal pressure signal, a laryngeal acoustic vibration signal and a swallowing acoustic wave signal of a patient; the data processing device is used for performing feature extraction and feature fusion according to the throat pressure signal, the throat sound vibration signal and the swallowing sound wave signal to obtain a multi-modal fusion feature, and calculating a swallowing disorder risk probability through a preset deep learning model according to the multi-modal fusion feature; the interactive display device is used for judging the swallowing function according to the swallowing disorder risk probability and displaying the screening result of the swallowing function according to the judgment result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical detection instruments, and in particular to a portable swallowing function screening system based on multimodality. Background Art

[0002] In the intensive care unit (ICU), rapid screening for PED (Postextubation Dysphagia) is crucial for preventing aspiration and reducing complications such as lung infections. However, existing technologies have significant shortcomings. Traditional testing methods, such as VFSS (Videofluoroscopic Swallowing Study), rely on radiology equipment and cannot be performed at the bedside, increasing patient transfer risks and examination costs. Although FEES (Fiberoptic Endoscopic Evaluation of Swallowing) can directly observe the throat, it is highly invasive, has poor patient tolerance, and is complex to perform, making it unsuitable for emergency screening scenarios in the ICU. Commonly used bedside screening tools in clinical practice, such as the Swallowing Function Screening Scale, rely on the nurse's subjective judgment and have a sensitivity of only 62-75%, which is difficult to meet the needs of accurate screening. In addition, existing swallowing assessment equipment is often bulky. For example, surface electromyography requires the connection of multiple leads, which limits its flexible application in the ICU. In terms of technical implementation, existing devices often use single-modality detection, such as monitoring only laryngeal movement or sound. This results in a high rate of misjudgment and fails to accurately reflect the complex biomechanical changes during swallowing. Furthermore, the lack of real-time interaction with electronic medical record systems hinders data integration and analysis, limiting the efficiency and accuracy of clinical decision-making.

[0003] Therefore, how to realize portable swallowing function screening for real-time clinical monitoring and early warning is a problem that needs to be solved urgently. Summary of the Invention

[0004] The main purpose of this application is to provide a portable swallowing function screening system based on multimodality, aiming to solve the technical problem of how to realize portable swallowing function screening for clinical real-time monitoring and early warning.

[0005] To achieve the above objectives, the present application proposes a portable swallowing function screening system based on multimodality, wherein the portable swallowing function screening system based on multimodality comprises: a signal acquisition device, a data processing device, and an interactive display device connected in sequence;

[0006] The signal acquisition device is used to collect the patient's laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal;

[0007] The data processing device is used to extract and fuse features based on the laryngeal pressure signal, the laryngeal acoustic vibration signal, and the swallowing sound wave signal to obtain a multimodal fusion feature, and calculate the dysphagia risk probability through a preset deep learning model based on the multimodal fusion feature;

[0008] The interactive display device is used to judge the swallowing function according to the swallowing disorder risk probability, and to display the swallowing function screening result according to the judgment result.

[0009] In one embodiment, the data processing device includes: a feature extraction module and a feature fusion module;

[0010] The feature extraction module is used to detect the rising edge of the laryngeal pressure signal and generate a pressure waveform feature;

[0011] The feature extraction module is further used to perform short-time energy analysis on the acoustic vibration signal to extract vibration energy features;

[0012] The feature extraction module is further used to perform spectrum analysis on the swallowing sound wave signal to extract sound wave spectrum features;

[0013] The feature fusion module is used to perform time sequence alignment and feature fusion on the pressure waveform features, vibration energy features and sound wave spectrum features to generate multimodal fusion features.

[0014] In one embodiment, the feature extraction module is further configured to divide the acoustic vibration signal into frames according to a first preset time window, and perform weighted processing on each frame signal to obtain a weighted frame signal;

[0015] The feature extraction module is further configured to calculate the short-time energy value of each frame signal based on the weighted frame signal, and extract the energy peak value and rise time in the short-time energy value as vibration energy features.

[0016] In one embodiment, the feature extraction module is further configured to extract Mel-frequency cepstral coefficients from the swallowing sound wave signal to obtain a spectrum envelope feature;

[0017] The feature extraction module is further configured to calculate an energy ratio of a preset frequency band based on the spectrum envelope feature, and use the energy ratio as a sound wave spectrum feature.

[0018] In one embodiment, the data processing device further comprises: a deep learning module;

[0019] The deep learning module is used to divide the multimodal fusion features into three-dimensional input data according to a second preset time window, and calculate the dysphagia risk probability through a preset convolutional neural network model and the three-dimensional input data.

[0020] In one embodiment, the data processing device further includes: a pre-processing module;

[0021] The preprocessing module is configured to obtain a preset filtering strategy and an adaptive noise reduction strategy, and perform filtering and noise reduction processing on the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal based on the preset filtering strategy and the adaptive noise reduction strategy to obtain an initial feature signal, and send the initial feature signal to the feature extraction module in the data processing device, wherein the initial feature signal includes the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal after filtering and noise reduction processing;

[0022] The feature extraction module is used to perform feature extraction and feature fusion on the initial feature signal to obtain multimodal fusion features.

[0023] In one embodiment, the signal acquisition device includes: a data calibration module;

[0024] The data calibration module is used to obtain the current ambient humidity and the current static signal, and when the current ambient humidity is greater than a preset humidity threshold, automatically calibrate the sensors that collect various signals according to the current static signal and a preset standard baseline.

[0025] In one embodiment, the interactive display device includes: a display module;

[0026] The display module is used to display patient identification information, a dynamic waveform of a laryngeal pressure signal, a frequency spectrum distribution diagram of a swallowing sound wave signal, and dysphagia risk level prompt information in real time.

[0027] In one embodiment, the interactive display device further includes: a voice interaction module;

[0028] The voice interaction module is used to receive voice instructions through a preset voice recognition model and generate corresponding multilingual voice feedback based on the swallowing disorder risk level prompt information. The voice feedback includes risk prompt content and operation suggestions.

[0029] The present application provides a portable swallowing function screening system based on multimodality, and the portable swallowing function screening system based on multimodality includes: a signal acquisition device, a data processing device and an interactive display device connected in sequence; the signal acquisition device is used to collect the patient's laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal; the data processing device is used to extract and fuse features based on the laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal to obtain multimodal fusion features, and calculate the swallowing disorder risk probability through a preset deep learning model based on the multimodal fusion features; the interactive display device is used to judge the swallowing function based on the swallowing disorder risk probability, and display the swallowing function screening results based on the judgment results. In summary, the present application realizes the accurate and rapid identification of swallowing disorders through multimodal sensing fusion and deep learning technology, thereby significantly improving the screening efficiency and accuracy of swallowing function, and filling the gap in domestic clinical swallowing assessment intelligent screening equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0032] Figure 1 This is a structural block diagram of the first embodiment of the portable multimodal swallowing function screening system of the present application;

[0033] Figure 2 This is a structural block diagram of the second embodiment of the multimodal portable swallowing function screening system of the present application;

[0034] Figure 3 This is a multimodal signal processing flow chart in one embodiment of the multimodal portable swallowing function screening system of the present application;

[0035] Figure 4 This is a structural block diagram of the third embodiment of the multimodal portable swallowing function screening system of this application.

[0036] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings.

[0037] Description of Figure Numbers:

[0038] Label name Label name 10 Signal acquisition device 202 Feature fusion module 20 Data processing device 203 Deep Learning Module 30 Interactive display device 204 Preprocessing module 101 Data calibration module 301 Display Module 201 Feature extraction module 302 Voice interaction module DETAILED DESCRIPTION

[0039] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0040] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0041] In the intensive care unit (ICU), rapid screening for PEDs is crucial for preventing aspiration and reducing complications such as lung infections. However, existing technologies have significant shortcomings. Traditional testing methods, such as VFSS, rely on radiology equipment and cannot be performed at the bedside, increasing patient transfer risks and examination costs. While FEES allows direct observation of the throat, it is highly invasive, poorly tolerated by patients, and complex to perform, making it unsuitable for emergency screening in the ICU. Commonly used bedside screening tools, such as the swallowing function screening scale, rely on subjective judgment by nurses and have a sensitivity of only 62-75%, failing to meet the requirements for accurate screening. Furthermore, existing swallowing assessment devices are often bulky; surface electromyography, for example, requires multiple leads, limiting their flexible application in the ICU. Technically, existing devices often utilize single-modality testing, such as monitoring only laryngeal movement or sound, resulting in high rates of false positives and an inability to accurately reflect the complex biomechanical changes during swallowing. Furthermore, the lack of real-time interaction with electronic medical record systems hinders data integration and analysis, limiting the efficiency and accuracy of clinical decision-making. Therefore, how to realize portable swallowing function screening for real-time clinical monitoring and early warning is a problem that needs to be solved urgently.

[0042] This application achieves accurate and rapid identification of swallowing disorders through multimodal sensor fusion and deep learning technology, thereby significantly improving the screening efficiency and accuracy of swallowing function, and filling the gap in domestic clinical swallowing assessment intelligent screening equipment.

[0043] Based on this, the embodiment of the present application provides a portable swallowing function screening system based on multimodality, referring to Figure 1 , Figure 1 This is a structural block diagram of the first embodiment of the multimodal portable swallowing function screening system of this application.

[0044] The portable swallowing function screening system based on multimodality in this embodiment includes: a signal acquisition device 10, a data processing device 20 and an interactive display device 30, wherein the signal acquisition device 10 is connected to the data processing device 20 and the interactive display device 30 in sequence; the signal acquisition device 10 is used to collect the patient's laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal.

[0045] It should be noted that the signal acquisition device 10 is composed of a wearable sensor module, including a laryngeal pressure sensor array, an acoustic vibration sensor and a directional MEMS (micro-electromechanical system) microphone. Among them, the laryngeal pressure sensor array adopts a flexible PCB substrate, integrates 8 miniature piezoresistive sensors, and is arranged in the area from the upper edge of the thyroid cartilage to the lower edge of the cricoid cartilage, for real-time acquisition of laryngeal pressure changes during the patient's swallowing process, that is, laryngeal pressure signals. The acoustic vibration sensor is a piezoelectric ceramic type with a sensitivity of 0.5mV / μm. It is attached to the throat and is used to capture the tiny vibrations of the throat during swallowing, that is, the throat acoustic vibration signal. The directional MEMS microphone has a frequency response range of 100-5000Hz and is used to accurately collect the sound wave signals generated during swallowing, that is, the swallowing sound wave signals.

[0046] It can be understood that the laryngeal pressure signal refers to the pressure changes caused by the contraction of the laryngeal muscles during swallowing, reflecting the strength and coordination of the swallowing action. The laryngeal acoustic vibration signal refers to the weak mechanical waves generated by the vibration of the laryngeal tissue during swallowing, which can reflect the patency of the swallowing passage. Swallowing acoustic wave signal: refers to the sound signal generated during swallowing, including swallowing sounds, airflow sounds, etc., which can be used to assess the integrity of the swallowing function.

[0047] In this embodiment, the data processing device 20 is used to perform feature extraction and feature fusion based on the laryngeal pressure signal, the laryngeal acoustic vibration signal, and the swallowing sound wave signal to obtain a multimodal fusion feature, and calculate the swallowing disorder risk probability through a preset deep learning model based on the multimodal fusion feature.

[0048] It should be noted that the data processing device 20 includes a multimodal data fusion chip (FPGA implementation) and intelligent analysis software. The multimodal data fusion chip will perform time alignment on the processed signals (accuracy ±5ms), and extract characteristic quantities such as the slope of the laryngeal pressure rise and the energy ratio of the main frequency band of the swallowing sound to form multimodal fusion features. The intelligent analysis software has a built-in preset deep learning model (1D-CNN architecture). The model uses 300 pre-collected swallowing event annotated data as a training set, can receive multimodal fusion features as input, and output the risk probability of swallowing disorders.

[0049] Additionally, it should be noted that multimodal fusion features combine features from different modalities (laryngeal pressure, pharyngeal acoustic vibrations, and swallowing sound waves) to form a feature vector that comprehensively reflects swallowing function. The pre-set deep learning model is based on a convolutional neural network (CNN) architecture, trained on a large amount of labeled swallowing event data. It can automatically learn complex patterns of swallowing function and accurately predict the risk of dysphagia.

[0050] In this embodiment, the interactive display device 30 is used to judge the swallowing function according to the swallowing disorder risk probability, and to display the swallowing function screening result according to the judgment result.

[0051] It should be noted that the interactive display device 30 includes a 7-inch anti-glare touch screen (IP65 protection level) and a Cantonese / Mandarin bilingual voice prompt module. After receiving the swallowing disorder risk probability transmitted by the data processing device 20, the interactive display device will make a swallowing function judgment based on a preset threshold value (such as 70%). If the risk probability exceeds the threshold, it is determined to be a swallowing dysfunction, and a red highlighted prompt is given on the screen. For example, a text prompt such as "High Risk-It is recommended to start the anti-aspiration protocol." If the risk probability is lower than the threshold, a green normal prompt bar is displayed. At the same time, the bilingual voice prompt module plays the corresponding voice prompt based on the judgment result.

[0052] The present embodiment provides a portable swallowing function screening system based on multimodality, and the portable swallowing function screening system based on multimodality includes: a signal acquisition device, a data processing device and an interactive display device connected in sequence; the signal acquisition device is used to collect the patient's laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal; the data processing device is used to extract and fuse features based on the laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal to obtain multimodal fusion features, and calculate the swallowing disorder risk probability through a preset deep learning model based on the multimodal fusion features; the interactive display device is used to judge the swallowing function based on the swallowing disorder risk probability, and display the swallowing function screening results based on the judgment results. In summary, this embodiment realizes the accurate and rapid identification of swallowing disorders through multimodal sensing fusion and deep learning technology, thereby significantly improving the screening efficiency and accuracy of swallowing function, and filling the gap in domestic clinical swallowing assessment intelligent screening equipment.

[0053] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , Figure 2 This is a structural block diagram of the second embodiment of the multimodal portable swallowing function screening system of this application.

[0054] Based on the above-mentioned first embodiment, the portable swallowing function screening system based on multimodality in this embodiment includes a data processing device 20, and the data processing device 20 includes: a feature extraction module 201, which is used to perform rising edge detection on the laryngeal pressure signal to generate pressure waveform characteristics; perform short-time energy analysis on the acoustic vibration signal to extract vibration energy characteristics; perform spectrum analysis on the swallowing sound wave signal to extract sound wave spectrum characteristics; a feature fusion module 202, which is used to perform time alignment and feature fusion on the pressure waveform characteristics, vibration energy characteristics and sound wave spectrum characteristics to generate multimodal fusion characteristics.

[0055] It should be noted that in this step, the feature extraction module 201 is a core component of the data processing device 20. It is responsible for extracting features from the laryngeal pressure signal, laryngeal acoustic vibration signal, and swallowing sound wave signal transmitted by the signal acquisition device 10. Laryngeal pressure signal processing: The feature extraction module 201 first detects the rising edge of the laryngeal pressure signal. Using an algorithm, it identifies the moment when the pressure waveform begins to rise, calculates the rising slope, and generates pressure waveform features. For the laryngeal acoustic vibration signal, the feature extraction module 201 uses a short-time energy analysis method to calculate the energy value of the signal within a short period of time (e.g., 10ms) and extract the vibration energy feature. This feature can reflect the vibration amplitude and frequency of the laryngeal tissue during swallowing, helping to determine the patency of the swallowing passage. Swallowing sound wave signal processing: The feature extraction module 201 performs spectral analysis on the swallowing sound wave signal, converting the time domain signal into a frequency domain signal using methods such as Fourier transform, and extracting the sound wave spectral features. This feature includes the main frequency, bandwidth, energy distribution, etc. of the sound wave, and can comprehensively reflect the sound characteristics produced during swallowing.

[0056] It can be understood that the pressure waveform characteristics refer to the rising slope and peak value of the laryngeal pressure signal during swallowing, reflecting the activity of the laryngeal muscles during swallowing. The vibration energy characteristics refer to the energy value of the laryngeal acoustic vibration signal in a short period of time, reflecting the vibration amplitude of the laryngeal tissue during swallowing. The acoustic spectrum characteristics refer to the distribution characteristics of the swallowing acoustic wave signal in the frequency domain, including the main frequency, bandwidth, energy distribution, etc., which can comprehensively reflect the sound characteristics produced during swallowing.

[0057] In addition, it should be noted that if Figure 3 As shown, because there may be some time difference between the laryngeal pressure signal, the laryngeal acoustic vibration signal, and the swallowing sound wave signal, the feature fusion module 202 first performs time alignment to ensure the consistency of each feature on the time axis. After time alignment, the feature fusion module 202 uses a preset fusion algorithm (such as decision-level fusion) to fuse the pressure waveform features, vibration energy features, and sound wave spectrum features to generate a multimodal fusion feature. This feature integrates information from multiple modalities and can more comprehensively reflect the patient's swallowing function status.

[0058] In a feasible embodiment, the feature extraction module 201 is also used to frame the acoustic vibration signal according to a first preset time window, and perform weighted processing on each frame signal to obtain a weighted frame signal; calculate the short-time energy value of each frame signal based on the weighted frame signal, and extract the energy peak and rise time in the short-time energy value as vibration energy features.

[0059] It should be noted that in this step, when processing the acoustic vibration signal, feature extraction module 201 first uses a first preset time window to segment the continuous acoustic vibration signal into multiple frames, each of which represents a short signal segment. This step facilitates subsequent frame-by-frame analysis of the signal and captures the temporal characteristics of the signal.

[0060] It is understandable that the length of the first preset time window can be adjusted according to actual needs, for example, it can be set to 200ms to ensure that each frame signal can contain sufficient swallowing-related feature information. After framing, each frame signal is regarded as an independent processing unit. In order to enhance the key information in the signal, the feature extraction module 201 will perform weighted processing on each frame signal to obtain a weighted frame signal. The weighted processing can be performed based on the energy, frequency or other characteristics of the signal to highlight the vibration characteristics related to swallowing. Short-time energy calculation refers to the weighted frame signal, and the feature extraction module 201 will calculate the short-time energy value of each frame signal. The short-time energy value reflects the energy of the signal in a short time (i.e. within one frame) and is used to evaluate the vibration intensity of the throat tissue during swallowing.

[0061] In a feasible embodiment, the feature extraction module 201 is further used to extract Mel-frequency cepstral coefficients of the swallowing sound wave signal to obtain a spectrum envelope feature; calculate the energy ratio of a preset frequency band based on the spectrum envelope feature, and use the energy ratio as the sound wave spectrum feature.

[0062] It should be noted that if Figure 3 As shown, in this step, feature extraction module 201 uses Mel-Frequency Cepstral Coefficient (MFCC) extraction technology to obtain the spectral envelope characteristics of the swallowing sound wave signal when processing the swallowing sound wave signal. MFCC is a feature extraction method widely used in speech recognition and audio processing. By simulating the human ear's perception of sound, it converts the sound wave signal into an MFCC feature vector, which can effectively describe the characteristics of the sound wave signal in the frequency domain.

[0063] It is understood that energy ratio calculation refers to the calculation of energy ratios within preset frequency bands (e.g., 2000-3500 Hz vs. 500-1500 Hz) based on the extracted MFCC spectral envelope features. This ratio reflects the energy distribution differences of swallowing sound waves in different frequency bands and helps distinguish between normal swallowing and swallowing disorders.

[0064] In a feasible embodiment, the data processing device 20 also includes: a deep learning module 203, which is used to divide the multimodal fusion features into three-dimensional input data according to a second preset time window, and calculate the swallowing disorder risk probability through a preset convolutional neural network model and the three-dimensional input data.

[0065] It should be noted that in this step, the deep learning module 203 first receives the multimodal fusion features from the data processing device 20. These features are obtained by synchronously collecting laryngeal pressure, laryngeal acoustic vibration and swallowing sound wave signals, and undergoing preliminary processing (such as filtering and noise reduction). The deep learning module 203 divides these multimodal fusion features into a series of three-dimensional input data blocks according to a predefined second preset time window (for example, 300 milliseconds). Each three-dimensional input data block contains the feature information of all modalities within the time window. Then, the deep learning module 203 processes these three-dimensional input data blocks using a preset convolutional neural network model. The convolutional neural network automatically extracts the corresponding features in the data through its convolutional layer, pooling layer, and fully connected layer structures, and ultimately outputs a risk probability of swallowing disorder.

[0066] It is understandable that the second preset time window refers to the time length used to segment the multimodal fusion features in the deep learning module, which is 300 milliseconds in this example. Figure 3 As shown in the figure, the 3D input data refers to the 3D fusion data formed by organizing multimodal fusion features (laryngeal pressure, pharyngeal acoustic vibrations, and swallowing sound waves) within each time window for processing by the convolutional neural network model. The pre-trained convolutional neural network model refers to a pre-trained convolutional neural network model for processing time series data. In this example, the 1D-CNN architecture is used.

[0067] In a feasible embodiment, the data processing device 20 also includes: a preprocessing module 204, which is used to obtain a preset filtering strategy and an adaptive noise reduction strategy, and filter and noise-reduce the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal based on the preset filtering strategy and the adaptive noise reduction strategy to obtain an initial feature signal, and send the initial feature signal to the feature extraction module in the data processing device. The initial feature signal includes the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal after filtering and noise reduction.

[0068] It should be noted that in this step, the pre-processing module 204 filters and performs noise reduction on the collected laryngeal pressure signal, laryngeal acoustic vibration signal, and swallowing sound wave signal according to a preset filtering strategy (e.g., 50Hz power frequency filtering) and an adaptive noise reduction strategy. The filtering process can remove noise of specific frequencies (e.g., power frequency interference), while the adaptive noise reduction strategy can dynamically adjust the noise reduction parameters based on the real-time characteristics of the signal to achieve the best noise reduction effect.

[0069] It is understood that the preset filtering strategy refers to a pre-set filtering method for removing noise of a specific frequency, which in this case is 50Hz power frequency filtering. The adaptive noise reduction strategy refers to a strategy that can dynamically adjust the noise reduction parameters according to the real-time characteristics of the signal. Figure 3 As shown, motion artifacts are eliminated using the ICA (Independent Component Analysis) blind source separation algorithm to achieve optimal noise reduction. The signals after filtering and noise reduction are called initial feature signals. These signals retain key information about the swallowing process while reducing the effects of noise and interference. The preprocessing module 204 sends these initial feature signals to the feature extraction module in the data processing device for subsequent analysis.

[0070] In this embodiment, an adaptive noise reduction strategy is used to cope with noise interference in different environments, improve the signal quality, enhance the robustness of the system, and provide a reliable data basis for subsequent feature extraction and swallowing disorder identification; at the same time, through deep learning technology, accurate prediction of swallowing disorders is achieved, and the accuracy and efficiency of screening are improved; using multimodal data fusion, a more comprehensive swallowing function assessment is provided, reducing the misjudgment that may be caused by single modality detection.

[0071] Based on the first and second embodiments of the present application, in the third embodiment of the present application, the same or similar contents as those in the first and second embodiments can be referred to above and will not be described in detail. Figure 4 , Figure 4This is a structural block diagram of the third embodiment of the multimodal portable swallowing function screening system of this application.

[0072] Based on the above-mentioned first and second embodiments, the multimodal portable swallowing function screening system in this embodiment includes an interactive display device, and the interactive display device 30 includes: a display module 301, which is used to display patient identification information, the dynamic waveform of the laryngeal pressure signal, the spectrum distribution diagram of the swallowing sound wave signal, and swallowing disorder risk level prompt information in real time.

[0073] It should be noted that in this step, when the device is started and connected to the patient, the display module 301 will first obtain the patient identification information, which is obtained synchronously by the patient information management system built into the device or the external electronic medical record system. Subsequently, the device begins to collect the patient's laryngeal pressure signal and swallowing sound wave signal, and converts the pressure signal into a dynamic waveform graph, and converts the sound wave signal into a spectrum distribution graph. At the same time, the device uses the built-in swallowing disorder recognition model to analyze the collected signals to derive the risk level of swallowing disorder. Finally, the display module 301 displays the patient identification information, the dynamic waveform of the laryngeal pressure signal, the spectrum distribution graph of the swallowing sound wave signal, and the swallowing disorder risk level prompt information on the screen in real time.

[0074] It is understood that patient identification information refers to the patient's unique identifying information, such as name, hospitalization number, and bed number, which is used to accurately identify the patient. The dynamic waveform of the pharyngeal pressure signal refers to the swallowing sound wave signal collected by the laryngeal pressure sensor array. The processed spectral distribution diagram shows the energy distribution of the swallowing sound wave at different frequencies and is used to determine whether the swallowing function is normal.

[0075] In a feasible embodiment, the interactive display device 30 also includes: a voice interaction module 302, which is used to receive voice instructions through a preset voice recognition model, and generate corresponding multilingual voice feedback according to the swallowing disorder risk level prompt information, and the voice feedback includes risk prompt content and operation suggestions.

[0076] It should be noted that in this step, the voice interaction module 302 has a built-in preset voice recognition model that can recognize the medical staff's voice commands, such as "Start screening" and "Display results." After the device completes the swallowing function screening and generates risk level prompt information, the voice interaction module 302 generates corresponding multilingual voice feedback based on this information. The voice feedback content includes risk warning content, such as "The patient has a high risk of swallowing difficulties, please take immediate action," and operational suggestions, such as "It is recommended to initiate the anti-aspiration protocol."

[0077] It is understandable that the preset voice recognition model refers to a pre-trained voice recognition model that can recognize specific voice instructions and convert them into commands executable by the device. Multilingual voice feedback means that the device supports voice feedback in multiple languages ​​(such as Cantonese / Mandarin) to meet the language needs of different medical staff. Risk warning content and operation suggestions refer to the voice feedback content generated based on the swallowing disorder risk level prompt information, including risk warnings and specific operation suggestions, which can help medical staff make quick judgments and handle them.

[0078] In this embodiment, the signal acquisition device 10 includes: a data calibration module 101, which is used to obtain the current ambient humidity and the current static signal, and when the current ambient humidity is greater than a preset humidity threshold, automatically calibrate the sensors that collect various types of signals according to the current static signal and a preset standard baseline.

[0079] It should be noted that in this step, the data calibration module 101 first obtains the current ambient humidity value through the built-in humidity sensor. At the same time, the module will collect the current static signal, that is, the signal value output by each sensor when there is no swallowing action. The data calibration module 101 will compare the current ambient humidity with the preset humidity threshold (such as 80% RH). If the current ambient humidity is greater than the preset humidity threshold, the data calibration module 101 will automatically calibrate the sensors that collect various signals (such as laryngeal pressure signals and sound wave signals) based on the difference between the current static signal and the preset standard baseline to eliminate the impact of humidity changes on sensor output.

[0080] It's understood that the current ambient humidity refers to the real-time humidity value of the device's environment, measured by the built-in humidity sensor. The current static signal refers to the signal value output by each sensor when there is no swallowing action, which is used as a calibration benchmark.

[0081] In a feasible implementation, the data calibration module 101 is further configured to fit the current ambient humidity according to a preset fitting model to obtain a dynamic gain coefficient, and automatically calibrate sensors that collect various signals using the dynamic gain coefficient and the signal offset.

[0082] It should be noted that in this step, the data calibration module 101 has a built-in preset fitting model that describes the relationship between ambient humidity and the sensor output signal. The data calibration module 101 will perform a fitting calculation based on the current ambient humidity value using the preset fitting model to obtain a dynamic gain coefficient. At the same time, the offset between the current signal and the standard signal (i.e., signal offset) will also be calculated. Finally, the data calibration module 101 will use the dynamic gain coefficient and signal offset to automatically calibrate the sensors that collect various signals and adjust the sensor output to eliminate the impact of humidity changes on the signal.

[0083] The dynamic gain coefficient is calculated based on the current ambient humidity using a preset humidity fitting model and is used to adjust the sensor output signal gain. The signal offset is the difference between the current signal and the reference signal and is used to adjust the sensor output offset to eliminate system errors.

[0084] In this embodiment, real-time visualization of patient information, pressure waveform, acoustic spectrum, and risk level is achieved through an interactive display device, and multilingual voice feedback is supported. At the same time, the data calibration module dynamically adjusts the gain and offset of the sensor according to the ambient humidity, thereby improving the output accuracy of the sensor under different ambient humidity conditions, effectively improving the efficiency and accuracy of swallowing function screening, and enhancing the convenience and reliability of clinical applications.

[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the system, device and possible architecture, function and operation of various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code includes one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0086] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0087] In addition, for technical details not fully described in this embodiment, please refer to the multimodal portable swallowing function screening system provided in any embodiment of the present application, and will not be repeated here.

[0088] In addition, it should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0089] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A portable swallowing function screening system based on multimodality, characterized in that: The portable swallowing function screening system based on multimodality includes: a signal acquisition device, a data processing device and an interactive display device connected in sequence; The signal acquisition device is used to collect the patient's laryngeal pressure signal, laryngeal acoustic vibration signal and swallowing sound wave signal; The data processing device is used to extract and fuse features based on the laryngeal pressure signal, the laryngeal acoustic vibration signal, and the swallowing sound wave signal to obtain a multimodal fusion feature, and calculate the dysphagia risk probability through a preset deep learning model based on the multimodal fusion feature; The interactive display device is used to judge the swallowing function according to the swallowing disorder risk probability, and to display the swallowing function screening result according to the judgment result.

2. The multimodal portable swallowing function screening system according to claim 1, characterized in that: The data processing device includes: a feature extraction module and a feature fusion module; The feature extraction module is used to detect the rising edge of the laryngeal pressure signal and generate a pressure waveform feature; The feature extraction module is further used to perform short-time energy analysis on the acoustic vibration signal to extract vibration energy features; The feature extraction module is further used to perform spectrum analysis on the swallowing sound wave signal to extract sound wave spectrum features; The feature fusion module is used to perform time sequence alignment and feature fusion on the pressure waveform features, vibration energy features and sound wave spectrum features to generate multimodal fusion features.

3. The multimodal portable swallowing function screening system according to claim 2, characterized in that: The feature extraction module is further configured to divide the acoustic vibration signal into frames according to a first preset time window, and perform weighted processing on each frame signal to obtain a weighted frame signal; The feature extraction module is further configured to calculate the short-time energy value of each frame signal based on the weighted frame signal, and extract the energy peak value and rise time in the short-time energy value as vibration energy features.

4. The multimodal portable swallowing function screening system according to claim 2, characterized in that: The feature extraction module is further configured to extract Mel-frequency cepstral coefficients from the swallowing sound wave signal to obtain a spectrum envelope feature; The feature extraction module is further configured to calculate an energy ratio of a preset frequency band based on the spectrum envelope feature, and use the energy ratio as a sound wave spectrum feature.

5. The multimodal portable swallowing function screening system according to claim 1, characterized in that: The data processing device further includes: a deep learning module; The deep learning module is used to divide the multimodal fusion features into three-dimensional input data according to a second preset time window, and calculate the dysphagia risk probability through a preset convolutional neural network model and the three-dimensional input data.

6. The multimodal portable swallowing function screening system according to claim 2, characterized in that: The data processing device further includes: a pre-processing module; The preprocessing module is configured to obtain a preset filtering strategy and an adaptive noise reduction strategy, and perform filtering and noise reduction processing on the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal based on the preset filtering strategy and the adaptive noise reduction strategy to obtain an initial feature signal, and send the initial feature signal to the feature extraction module in the data processing device, wherein the initial feature signal includes the laryngeal pressure signal, the pharyngeal acoustic vibration signal, and the swallowing sound wave signal after filtering and noise reduction processing; The feature extraction module is used to perform feature extraction and feature fusion on the initial feature signal to obtain multimodal fusion features.

7. The multimodal portable swallowing function screening system according to claim 1, characterized in that: The signal acquisition device includes: a data calibration module; The data calibration module is used to obtain the current ambient humidity and the current static signal, and when the current ambient humidity is greater than a preset humidity threshold, automatically calibrate the sensors that collect various signals according to the current static signal and a preset standard baseline.

8. The multimodal portable swallowing function screening system according to claim 7, characterized in that: The data calibration module is further configured to calculate a signal offset based on the current static signal and a preset standard baseline; The data calibration module is further used to fit the current ambient humidity according to a preset fitting model to obtain a dynamic gain coefficient, and automatically calibrate the sensors that collect various signals through the dynamic gain coefficient and the signal offset.

9. The multimodal portable swallowing function screening system according to claim 1, characterized in that: The interactive display device includes: a display module; The display module is used to display patient identification information, a dynamic waveform of a laryngeal pressure signal, a frequency spectrum distribution diagram of a swallowing sound wave signal, and dysphagia risk level prompt information in real time.

10. The multimodal portable swallowing function screening system according to claim 1, characterized in that: The interactive display device further includes: a voice interaction module; The voice interaction module is used to receive voice instructions through a preset voice recognition model and generate corresponding multilingual voice feedback based on the swallowing disorder risk level prompt information. The voice feedback includes risk prompt content and operation suggestions.

Citation Information

Cited By

  • Swallowing disorder risk prediction and personalized intervention recommendation system for stroke patient

    CN121709224A

  • Multi-signal analysis-based dysphagia risk early warning method and system

    CN121867704A