Poped field water drinking experiment method for evaluating dysphagia of patient
By using 4K high-frame rate cameras and infrared cameras to capture throat movements, and combining video preprocessing and deep learning models to identify swallowing movements, the problem of misclassification of swallowing monitoring in noisy environments is solved, and high-precision swallowing movement detection is achieved.
Patent Information
- Application Number
- CN202510068094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art uses a high-sensitivity throat microphone for swallow-related events monitoring in noisy environments, and is susceptible to interference from surrounding environment noise, resulting in misclassification.
The 4K high-frame rate camera and infrared high-frame rate camera are used to capture throat movements, and the swallowing action is recognized through video preprocessing and deep learning models. Combined with optical flow method and sliding window method, the precise detection of swallowing action is achieved.
It effectively improves image quality and detection accuracy, reduces the impact of artifacts, improves the recognition accuracy of swallowing actions and the generalization ability of deep learning models.
Smart Images

Figure FT_1 
Figure QLYQS_1 
Figure QLYQS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patient dysphagia assessment, and particularly relates to a Kubota water drinking test method for patient dysphagia assessment. Background Art
[0002] The Kubota water swallowing test (Fukuda Water Swallowing Test) is a simple, rapid, and effective clinical assessment tool for the preliminary screening and assessment of swallowing dysfunction.
[0003] The patent with the application number CN202410789658.9 records in its specification that "the present invention discloses a wearable laryngeal microphone swallowing ability intelligent screening system based on a smartphone for realizing the automatic classification and quantitative assessment of swallowing-related signals. The working principle of the system is divided into three steps: First, the swallowing sound signals of the subject are collected in real time through the laryngeal microphone, and the collected signals are displayed, saved, and analyzed in real time by using the developed application program. Second, through the transfer learning of the pre-trained deep convolutional neural network YAMNet on the AudioSet-Youtube corpus, a model S-YAMNet for detecting swallowing-related events is constructed and deployed in the mobile device to realize the real-time classification of swallowing sound signals. Finally, the subject conducts the Kubota water drinking test, and the swallowing ability is scored by counting the swallowing and coughing events. The test results show that the swallowing ability score evaluated by this system is highly consistent with the expert score. The present invention can be used for the intelligent screening of swallowing ability in daily life, which helps to more effectively and timely detect the potential risk of swallowing difficulties." Although the above technology combines the convenience of the smartphone and the high sensitivity of the laryngeal microphone, and can monitor and analyze swallowing-related events in real time in various environments, providing an effective auxiliary diagnostic tool for medical professionals, even though a high-sensitivity laryngeal microphone is selected, the above technology is still affected by ambient noise, especially in noisy environments such as hospitals, streets, and home environments, and the noise will affect the distinguishability between swallowing, coughing, and noise, resulting in misclassification.
[0004] In summary, developing a Kubota water drinking test method for patient dysphagia assessment is still a key problem that urgently needs to be solved in the technical field of patient dysphagia assessment. Summary of the Invention
[0005] The object of the present invention is to solve the problem that although the above-mentioned technology combines the convenience of a smart phone and the high sensitivity of a laryngeal microphone, it can monitor and analyze swallowing-related events in real time in various environments and provide an effective auxiliary diagnostic tool for medical professionals. However, despite the selection of a high-sensitivity laryngeal microphone, the above-mentioned technology is still affected by ambient noise interference, especially in noisy environments such as hospitals, streets, and home environments, and the noise will affect the discrimination between swallowing, coughing, and noise, resulting in misclassification situations.
[0006] To achieve the above object, the present invention provides the following technical solutions: The present invention provides a Kubota drinking water test method for evaluating patients' swallowing disorders, including the following steps: S1. Use a 4K high-frame-rate camera and an infrared high-frame-rate camera to capture subtle changes during the laryngeal movement to generate video images; S2. Preprocess the video images to obtain high-quality images; S3. Perform swallowing action detection on the high-quality images to obtain feature images; S4. After labeling the feature images, input them into a deep learning model to train the deep learning model to recognize swallowing actions; S5. Evaluate and optimize the trained deep learning model.
[0007] Further, in step S1, the method of using a 4K high-frame-rate camera and an infrared high-frame-rate camera to capture subtle changes during the laryngeal movement to generate video images is as follows: Select a 4K resolution camera and an infrared camera with a frame rate of more than 60 frames per second, install them in front of and on the side of the subject to capture the vertical and front-back movements of the larynx, transmit the video images to the cloud storage for archiving in the H.264 video encoding format, and regularly back up the data. Use software synchronization technology to synchronize the time stamps of the video images.
[0008] Further, in step S2, the method of preprocessing the video images to obtain high-quality images is as follows: Use a video processing library to decode the video images into continuous image frames, remove noise through information of the front and back frame images based on mean filtering and Kalman filtering in a time window, and correct the images using an image alignment algorithm. Mean filtering formula: , where is the image value after mean filtering in the time window, representing the filtered image value at the pixel position on the current frame, is the size of the filtering window, is the range of the time window, is at the moment The image value after being processed by the denoising algorithm at time is the sum of the denoised image values of all frames (from to ) within the time window. The Kalman filter formula: , where is the state vector at the current moment , is the state transition matrix, is the state vector at the previous moment, is the control matrix, is the control input, is the process noise, is the observation value, is the observation matrix, is the observation noise. The image alignment algorithm: , where is the image after alignment, is the transformation matrix, is the original image, is the image transformation type.
[0009] Furthermore, in step S2, the method for preprocessing the video image to obtain a high-quality image is as follows: Improve the brightness distribution of the image through adaptive histogram equalization, adjust the brightness and contrast of the image to make the details of the larynx clearer, use a sharpening filter to enhance the details of the laryngeal muscles and movements, use Canny edge detection and Sobel operator to extract the edge information of the image, identify the laryngeal contour, use an artifact removal algorithm to remove some artifacts that appear during swallowing movements. The adaptive histogram equalization formula: , where is the image after being processed by adaptive histogram equalization, is the local neighborhood of the image centered at point , is the number of pixel points in the local neighborhood, is the pixel value at position in the image, The sharpening filter formula: , where is the pixel value of the sharpened image at this position , is the pixel value of the denoised image at this position , is the sharpening coefficient, is the blurred version of the image . The Canny edge detection formula: , where is the image after edge detection, It is the result after applying the Canny edge detection algorithm. The formula for the artifact removal algorithm is: , where is the pixel value of the image after artifact removal at this position , is the pixel value of the original image at position , is the artifact part.
[0010] Furthermore, in step S3, the method for obtaining the feature image by performing swallowing action detection based on the high-quality image is as follows: Use the Harris corner detection and SIFT (Scale-Invariant Feature Transform) algorithms to extract key points in the laryngeal region to support subsequent motion analysis. Use the optical flow method to estimate the motion of pixels between different frames and track the subtle motion of the larynx to determine whether it is a swallowing action. The Harris corner detection formula is: , where is the Harris response value at the position, is the gradient matrix of the image, is the gradient of the image in the horizontal change direction, is the gradient of the image in the vertical change direction, is the determinant of the matrix , is the trace of the matrix , is an empirical constant. The SIFT feature extraction formula is: , where is the SIFT feature set extracted from the image , is the th coordinate of the feature point, is the th scale of the feature point, is the th direction of the feature point, is the descriptor. The optical flow method formula is: , where respectively represent the speeds of a certain pixel point in the horizontal and vertical directions in the image, is the gradient of the image in the horizontal direction, is the gradient of the image in the vertical direction, is the change of the image over time .
[0011] Furthermore, in step S3, the method for obtaining the feature image by performing swallowing action detection based on the high-quality image is as follows: The sliding window method is used to process 30 consecutive frames as a window, calculate the changes in the laryngeal region within this time window, analyze whether it conforms to the characteristics of swallowing actions, and perform end-to-end detection of the characteristics of swallowing actions through an end-to-end deep learning network, so as to directly output the time points and durations of each swallowing action. The formula for the sliding window method: , where is the change amount of the laryngeal region features within the sliding window, is the size of the sliding window, is the time index within the sliding window, is at time the calculated laryngeal region features, is at time the calculated laryngeal region features. The formula for the end-to-end deep learning network: , where is the probability of a swallowing action occurring after a given image , are the features extracted by the deep neural network, is the feature weight, is the bias term, is the activation function.
[0012] Furthermore, in step S4, after annotating the feature images and inputting them into the deep learning model, the method for training the deep learning model to recognize swallowing actions is as follows: Using the open-source image and video annotation tool ComputerVisionAnnotationTool, according to the time points and durations of swallowing actions, label either "swallowing" or "non-swallowing" in each frame of the image. Before training the deep learning model, divide the labeled image set into a training set, a validation set, and a test set. The training set accounts for 70%-80% of the image set, the validation set accounts for 10%-15% of the image set, and the test set accounts for 10%-15% of the image set.
[0013] Furthermore, in step S4, after annotating the feature images and inputting them into the deep learning model, the method for training the deep learning model to recognize swallowing actions is as follows: Using generative adversarial network technology to balance the number of positive and negative samples (swallowing and non-swallowing), and then sequentially input the training set, validation set, and test set as input data into the deep learning model (3D convolutional neural network). For the binary classification task (swallowing and non-swallowing), use the binary cross-entropy loss function. The formula for the 3D convolutional neural network: , where is the feature map extracted by the th convolutional layer, is the number of channels of the input image, is the number of frames in the video sequence, is the height and width of the input image, is the weight parameter of the convolutional kernel, are the indices of different dimensions of the convolutional kernel in space and time respectively, is the middle position of the input image sequence of the pixel values, is the bias term, the binary cross-entropy loss function formula: where is the number of samples, is the true label value for the binary classification task, 1 for swallowing action and 0 for non-swallowing action, is the predicted value of the model on the sample , is the logarithm of the predicted probability when the true label is 1, is the logarithm of the predicted probability when the true label is 0, is the value of the loss function.
[0014] Furthermore, in step S5, the method for evaluating and optimizing the trained deep learning model is as follows: Use the validation set and the test set to evaluate the trained deep learning model. The evaluation metrics include: accuracy, recall rate, and precision. Use the cross-validation technique to divide the training data into k subsets. Each time, use one of the subsets as the validation set, and the remaining k - 1 subsets as the training set. Through k times of validation, obtain the average performance of the deep learning model. Identify the misclassifications that occur in the deep learning model in a specific scenario through the confusion matrix, and at the same time analyze whether the misclassifications are caused by any one of the reasons such as image quality problems, the complexity of specific actions, and environmental noise. The cross-validation formula: where represents the average evaluation metric value of all times of validation, represents the number of validations during cross-validation, represents the evaluation metric calculated in the i-th validation.
[0015] Furthermore, in step S5, the method for evaluating and optimizing the trained deep learning model is as follows: Introduce a spatial attention mechanism to make the deep learning model focus on the key regions related to swallowing actions in the image. Use methods such as grid search, random search, and Bayesian optimization to optimize the learning rate. After the deep learning model, use the validation set for re-evaluation to prevent overfitting during the optimization process. At the same time, in practical applications, through online learning, the deep learning model continuously learns and adjusts from newly collected data to adapt to the swallowing action changes of different people.
[0016] Beneficial effects Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects: When the present invention is in use, it is beneficial to capture the subtle movements of the throat, improve the image quality and accuracy, and at the same time improve the integrity and accuracy of the data, which is beneficial to accurately identify swallowing movements. The use of the artifact removal algorithm effectively removes the artifacts generated during swallowing movements, ensuring the authenticity and accuracy of the images, which is beneficial to improving the detection accuracy of swallowing movements, enhancing the generalization ability of the deep learning model, improving the recognition accuracy of swallowing movements, optimizing the training process of the deep learning model, being beneficial to effectively evaluating the model performance, discovering and improving the deficiencies of the model, enabling the present invention to continuously adapt to new data, and improving the accuracy and adaptability of swallowing movement recognition. Brief description of the drawings
[0017] Figure 1 It is a flowchart of a Kubota drinking water test method for evaluating swallowing disorders in patients according to the present invention. Detailed implementation manners
[0018] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but includes other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0020] The present invention will be further described in detail below in conjunction with the drawings: Embodiment: As Figure 1As shown in the figure, the present invention provides a Kubota drinking water test method for evaluating patients' swallowing disorders, including the following steps: S1. Use a 4K high-frame-rate camera and an infrared high-frame-rate camera to capture the subtle changes during the laryngeal movement to generate video images; Further, in step S1, the method of using a 4K high-frame-rate camera and an infrared high-frame-rate camera to capture the subtle changes during the laryngeal movement to generate video images is as follows: Select a 4K resolution camera and an infrared camera with a frame rate of 60 frames per second or more, install them in front of and on the side of the subject to capture the vertical and front-back movements of the larynx, transmit the video images to the cloud storage for archiving in the H.264 video coding format, and regularly back up the data. Use software synchronization technology to synchronize the time stamps of the video images; In this embodiment, a 4K high-frame-rate camera and an infrared high-frame-rate camera are used, which are respectively installed in front of and on the side of the subject to capture the vertical and front-back movements of the larynx, and then software synchronization technology is used to ensure the time stamp synchronization of the video images, which is beneficial to improving the integrity and accuracy of the data.
[0021] S2. Preprocess the video images to obtain high-quality images; Further, in step S2, the method of preprocessing the video images to obtain high-quality images is as follows: Use a video processing library to decode the video images into continuous image frames, remove noise through the information of the front and back frame images based on mean filtering and Kalman filtering in the time window, and correct the images using an image alignment algorithm. The mean filtering formula: , where is the image value after mean filtering in the time window, representing the filtered image value at the pixel position on the current frame, is the size of the filtering window, is the range of the time window, is the image value processed by the denoising algorithm at time , is the sum of the denoised image values of all frames (from to ) in the time window. The Kalman filtering formula: , where is the state vector at the current time , is the state transition matrix, is the state vector at the previous time, is the control matrix, is the control input, is the process noise, is the observed value, is the observation matrix, is the observation noise, and the image alignment algorithm: , where is the aligned image, is the transformation matrix, is the original image, is the image transformation type.
[0022] Furthermore, in step S2, the method for preprocessing the video image to obtain a high-quality image is as follows: Improve the brightness distribution of the image through adaptive histogram equalization, adjust the brightness and contrast of the image to make the details of the larynx clearer, use a sharpening filter to enhance the details of the laryngeal muscles and movements, use Canny edge detection and Sobel operator to extract the edge information of the image, identify the laryngeal contour, use an artifact removal algorithm to remove some artifacts that appear during swallowing movements, and the adaptive histogram equalization formula: , where is the image after adaptive histogram equalization processing, is the local neighborhood in the image centered at the point , is the number of pixel points in the local neighborhood, is the position in the image at which the pixel value is located, and the sharpening filter formula: , where is the pixel value of the sharpened image at this position , is the pixel value of the denoised image at this position , is the sharpening coefficient, is the image in its blurred version, and the Canny edge detection formula: , where is the image after edge detection, is the result after applying the Canny edge detection algorithm, and the artifact removal algorithm formula: , where is the pixel value of the image after artifact removal at this position , is the pixel value of the original image at the position , is the artifact part; In this embodiment, noise is reduced through mean filtering, Kalman filtering, and image alignment algorithms to improve image quality and consistency. Adaptive histogram equalization, sharpening filtering, and edge detection enhance the visibility of laryngeal muscles and movements, facilitating the accurate identification of swallowing actions. The use of an artifact removal algorithm effectively removes artifacts generated during swallowing actions, ensuring the authenticity and accuracy of the images.
[0023] S3. Perform swallowing action detection on the high-quality image to obtain a feature image; Further, in step S3, the method for performing swallowing action detection on the high-quality image to obtain a feature image is as follows: Use the Harris corner detection and SIFT (Scale-Invariant Feature Transform) algorithms to extract key points in the laryngeal region to support subsequent motion analysis. Utilize the optical flow method to estimate the motion of pixels between different frames and track the subtle motions of the larynx to determine whether it is a swallowing action. The Harris corner detection formula: , where is the Harris response value at the position, is the gradient matrix of the image, is the gradient of the image in the horizontal change direction, is the gradient of the image in the vertical change direction, is the matrix 's determinant, is the matrix 's trace, is an empirical constant. The SIFT feature extraction formula: , where is the SIFT feature set extracted from the image , is the coordinate of the th feature point, is the scale of the th feature point, is the direction of the th feature point, is the descriptor. The optical flow method formula: , where respectively represent the speeds of a certain pixel point in the horizontal and vertical directions in the image, is the gradient of the image in the horizontal direction, is the gradient of the image in the vertical direction, is the change of the image at time .
[0024] Further, in step S3, the method for performing swallowing action detection on the high-quality image to obtain a feature image is as follows: The sliding window method is used to process 30 consecutive frames as a window, calculate the changes in the laryngeal region within this time window, analyze whether it conforms to the characteristics of swallowing actions, and perform end-to-end detection of the characteristics of swallowing actions through an end-to-end deep learning network, so as to directly output the time points and durations of each swallowing action. The formula for the sliding window method is: , where is the change amount of the laryngeal region characteristics within the sliding window, is the size of the sliding window, is the time index within the sliding window, is at time the laryngeal region characteristics calculated, is at time the laryngeal region characteristics calculated. The formula for the end-to-end deep learning network is: , where is the probability of a swallowing action occurring after a given image , are the characteristics extracted by the deep neural network, is the feature weight, is the bias term, is the activation function; In this embodiment, the Harris corner detection and SIFT algorithm are used to extract the key points of the laryngeal region, providing support for subsequent motion analysis. The optical flow method is used to estimate the motion of pixels between different frames, track the subtle motions of the larynx, and determine whether it is a swallowing action, which is beneficial to improving the detection accuracy of swallowing actions. The sliding window method can analyze consecutive image frames in a short time, monitor the changes in the larynx in real time, and provide accurate data for the detection of swallowing actions. The end-to-end deep learning network provides efficient detection of swallowing actions by learning image features and accurately outputs the occurrence time and duration of the actions.
[0025] S4. After annotating the feature images, input them into the deep learning model to train the deep learning model to recognize swallowing actions; Further, in step S4, the method of annotating the feature images and then inputting them into the deep learning model to train the deep learning model to recognize swallowing actions is as follows: Using the open-source image and video annotation tool ComputerVisionAnnotationTool, according to the time points and durations of swallowing actions, label either "swallowing" or "non-swallowing" in each frame of the image. Before training the deep learning model, divide the labeled image set into a training set, a validation set, and a test set. The training set accounts for 70%-80% of the image set, the validation set accounts for 10%-15% of the image set, and the test set accounts for 10%-15% of the image set.
[0026] Further, in step S4, after the feature image is labeled, it is input into the deep learning model. The method for training the deep learning model to recognize swallowing actions is as follows: Use the generative adversarial network technology to balance the number of positive and negative samples (swallowing and non-swallowing). Subsequently, the training set, validation set, and test set are sequentially used as input data and input into the deep learning model (3D convolutional neural network). For the binary classification task (swallowing and non-swallowing), the binary cross-entropy loss function is used. The 3D convolutional neural network formula: , where is the feature map extracted by the th convolutional layer, is the number of channels of the input image, is the number of frames in the video sequence, is the height and width of the input image, is the weight parameter of the convolutional kernel, are the indices of different dimensions of the convolutional kernel in space and time respectively, is the middle position of the input image sequence pixel value, is the bias term. The binary cross-entropy loss function formula: , where is the number of samples, is the true label value. For the binary classification task, the swallowing action is 1 and the non-swallowing action is 0, is the predicted value of the model on the sample , is the logarithm of the predicted probability when the true label is 1, is the logarithm of the predicted probability when the true label is 0, is the loss function value; In this embodiment, by accurately labeling and reasonably partitioning the data set, the representativeness and diversity of the training data are ensured, which helps the accurate training and verification of the model. The generative adversarial network technology solves the problem of unbalanced positive and negative samples, improves the generalization ability of the deep learning model, improves the recognition accuracy of swallowing actions, and optimizes the training process of the deep learning model.
[0027] S5. Evaluate and optimize the trained deep learning model; Further, in step S5, the method for evaluating and optimizing the trained deep learning model is as follows: Evaluate the trained deep learning model using the validation set and the test set. The evaluation metrics include: accuracy, recall, and precision. Use the cross-validation technique to divide the training data into k subsets. Each time, use one of the subsets as the validation set, and the remaining k - 1 subsets as the training set. Through k validations, obtain the average performance of the deep learning model. Identify the misclassifications that occur in the deep learning model in a specific scenario through the confusion matrix. At the same time, analyze whether the misclassifications are caused by any one of the reasons such as image quality problems, the complexity of specific actions, and environmental noise. Cross-validation formula: , where represents the average evaluation metric value of all k validations, represents the number of validations during cross-validation, represents the evaluation metric calculated in the i-th validation.
[0028] Furthermore, in step S5, the method for evaluating and optimizing the trained deep learning model is as follows: Introduce a spatial attention mechanism to make the deep learning model focus on the key regions related to the swallowing action in the image. Use methods such as grid search, random search, and Bayesian optimization to optimize the learning rate. After the deep learning model, use the validation set for re-evaluation to prevent overfitting during the optimization process. At the same time, in practical applications, through online learning, the deep learning model continuously learns and adjusts from newly collected data to adapt to the swallowing action changes of different people; In this embodiment, use the validation set and the test set to evaluate the trained model. The evaluation metrics include accuracy, recall, and precision. Adopt the cross-validation technique to divide the training data into k subsets, conduct multiple validations, calculate the average value of the evaluation metrics, analyze the reasons for misclassifications through the confusion matrix, such as image quality problems, action complexity, or environmental noise, etc. At the same time, introduce a spatial attention mechanism to make the model focus on the key regions related to the swallowing action, use grid search, random search, and Bayesian optimization to learn the rate, avoid overfitting. In practical applications, make the model adapt to the swallowing action changes of different people through online learning and adjust the parameters, which is beneficial to effectively evaluate the model performance, discover and improve the deficiencies of the model, make the present invention continuously adapt to new data, and improve the accuracy and adaptability of swallowing action recognition.
[0029] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A Kubota drinking water test method for assessing swallowing disorders in patients, characterized in that: The following steps are involved: S1. Use a 4K high frame rate camera and an infrared high frame rate camera to capture subtle changes in the laryngeal movement and generate video images; S2, preprocessing the video image to obtain a high-quality image; S3, performing swallowing action detection according to the high-quality image to obtain a feature image; S4, annotating the feature image and inputting it into a deep learning model to train the deep learning model to recognize swallowing movements; S5. Evaluate and optimize the deep learning model after training.
2. The Kubota drinking water test method for assessing swallowing disorders according to claim 1, characterized in that: In step S1, the method of using a 4K high frame rate camera and an infrared high frame rate camera to capture subtle changes in the laryngeal movement process to generate a video image is as follows: Select 4K resolution cameras and infrared cameras with a frame rate of more than 60 frames per second, install them in front of and on the side of the subject to capture the vertical movement and front-to-back movement of the throat, transfer the video images to cloud storage in H.264 video encoding format for archiving, back up the data regularly, and use software synchronization technology to synchronize the video images with timestamps.
3. The Kubota drinking water test method for assessing swallowing disorders according to claim 2, characterized in that: In step S2, the method of preprocessing the video image to obtain a high-quality image is: The video image is decoded into continuous image frames using a video processing library. The mean filter and Kalman filter based on the time window are used to remove noise through the information of the previous and next frame images. The image is corrected using an image alignment algorithm. The mean filter formula is: ,in It is the image value after time window mean filtering, represented at the pixel position The image value after filtering on the current frame, The size of the filter window, is the range of the time window, It is at the moment is the image value after being processed by the denoising algorithm, is the sum of all frames (from arrive ) to sum the denoised image values, the Kalman filter formula is: ,in It is the current moment The state vector of is the state transition matrix, is the state vector at the previous moment, is the control matrix, is the control input, is the process noise, is the observed value, is the observation matrix, is the observation noise, image alignment algorithm: ,in is the aligned image, is the transformation matrix, is the original image, is the image transformation type.
4. The Kubota drinking water test method for assessing swallowing disorders according to claim 3, characterized in that: In step S2, the method of preprocessing the video image to obtain a high-quality image is: Adaptive histogram equalization is used to improve the brightness distribution of the image, adjust the brightness and contrast of the image, make the details of the larynx clearer, use a sharpening filter to enhance the details of the laryngeal muscles and movements, use Canny edge detection and Sobel operator to extract the edge information of the image, identify the laryngeal contour, and use the artifact removal algorithm to remove some artifacts that appear during swallowing. The formula for adaptive histogram equalization is: ,in is the image processed by adaptive histogram equalization. is a point in the image The local neighborhood centered on is the number of pixels in the local neighborhood, is the position in the image The pixel value at Sharpening filter formula: ,in is the sharpened image at this position The pixel value on is the denoised image at this position The pixel value on is the sharpening factor, is an image Blurred version of Canny edge detection formula: ,in is the image after edge detection, It is the result after applying the Canny edge detection algorithm. The formula for removing artifacts is: ,in The image after removing artifacts is at this position The pixel value on is the original image at position The pixel value on It is the artifact part.
5. The Kubota drinking water test method for assessing swallowing disorders according to claim 4, characterized in that: In step S3, the method for obtaining a characteristic image by performing swallowing action detection based on the high-quality image is as follows: Harris corner detection and SIFT (Scale Invariant Feature Transform) algorithms are used to extract key points in the throat area to provide support for subsequent motion analysis. The optical flow method is used to estimate the movement of pixels between different frames and track the subtle movement of the throat to determine whether it is a swallowing action. The Harris corner detection formula is: ,in is the Harris response value at location, is the gradient matrix of the image, is the gradient of the image in the horizontal direction of change, is the gradient of the image in the vertical direction of change, is a matrix The determinant of is a matrix traces, is an empirical constant, SIFT feature extraction formula: ,in is an image The SIFT feature set extracted from It is The coordinates of the feature points, It is The scale of the feature points, It is The direction of the feature points, is the descriptor, the optical flow formula is: ,in Respectively represent the speed of a pixel in the image in the horizontal and vertical directions, is the horizontal gradient of the image, is the vertical gradient of the image, is the image at time changes in .
6. The Kubota drinking water test method for assessing swallowing disorders according to claim 5, characterized in that: In step S3, the method for obtaining a characteristic image by performing swallowing action detection based on the high-quality image is as follows: The sliding window method is used to process 30 consecutive frames as a window, calculate the changes in the laryngeal area within the time window, and analyze whether it meets the characteristics of the swallowing action. The characteristics of the swallowing action are detected end-to-end through an end-to-end deep learning network, so as to directly output the time point and duration of each swallowing action. The sliding window method formula is: ,in is the variation of the throat region characteristics within the sliding window, is the size of the sliding window, is the time index within the sliding window, It is at the moment Calculated throat region characteristics, It is at the moment Calculated laryngeal region features, end-to-end deep learning network formula: ,in Is a given image The probability of swallowing after are features extracted by deep neural networks, is the feature weight, is the bias term, is the activation function.
7. The Kubota drinking water test method for assessing swallowing disorders according to claim 6, characterized in that: In step S4, the feature image is annotated and then put into a deep learning model. The method for training the deep learning model to recognize swallowing movements is as follows: Using the open source image and video annotation tool Computer Vision Annotation Tool, each frame of the image was labeled as "swallowing" or "non-swallowing" according to the time point of the swallowing action and its duration. Before training the deep learning model, the annotated image set was divided into training set, validation set and test set. The training set accounted for 70%-80% of the image set, the validation set accounted for 10%-15% of the image set, and the test set accounted for 10%-15% of the image set.
8. The Kubota drinking water test method for assessing swallowing disorders according to claim 7, characterized in that: In step S4, the feature image is annotated and then put into a deep learning model. The method for training the deep learning model to recognize swallowing movements is as follows: Generative adversarial network technology is used to balance the number of positive and negative samples (swallowing and non-swallowing). Then the training set, validation set, and test set are input into the deep learning model (3D convolutional neural network) as input data. For the binary classification task (swallowing and non-swallowing), the binary cross entropy loss function is used. The 3D convolutional neural network formula is: ,in It is The feature map extracted by the convolutional layer, is the number of channels of the input image, is the number of frames in the video sequence, are the height and width of the input image, is the weight parameter of the convolution kernel, are the indexes of the convolution kernel in different dimensions in space and time, is the middle position of the input image sequence The pixel value of is the bias term, and the binary cross entropy loss function formula is: ,in is the number of samples, is the true label value. For the binary classification task, swallowing action is 1 and non-swallowing action is 0. The model is in the sample The predicted value on is the logarithm of the predicted probability when the true label is 1, is the logarithm of the predicted probability when the true label is 0, is the loss function value.
9. The Kubota drinking water test method for assessing swallowing disorders according to claim 8, characterized in that: In step S5, the method for evaluating and optimizing the trained deep learning model is: The trained deep learning model is evaluated using the validation set and the test set. The evaluation indicators include accuracy, recall, and precision. The cross-validation technique is used to divide the training data into k subsets, and one of the subsets is used as the validation set each time. The remaining k-1 subsets are used as the training set. After k validations, the average performance of the deep learning model is obtained. The confusion matrix is used to identify the misclassification of the deep learning model in specific scenarios. At the same time, it is analyzed whether the misclassification is caused by any of the following reasons: image quality problems, the complexity of specific actions, and environmental noise. The cross-validation formula is: ,in Indicates all The average evaluation index value of the validation, Indicates the number of validations during cross validation. Represents the evaluation index calculated in the th validation.
10. The Kubota drinking water test method for assessing swallowing disorders according to claim 8, characterized in that: In step S5, the method for evaluating and optimizing the trained deep learning model is: The spatial attention mechanism is introduced to make the deep learning model focus on the key areas related to swallowing movements in the image. The learning rate is optimized using grid search, random search and Bayesian optimization methods. After the deep learning model is completed, the validation set is used for re-evaluation to prevent the deep learning model from overfitting during the optimization process. At the same time, in actual applications, through online learning, the deep learning model can continuously learn and adjust from the newly collected data to adapt to the changes in swallowing movements of different populations.
Citation Information
Patent Citations
Intelligent screening system for swallowing ability of wearable throat microphone based on smart phone
CN118750021A