levator muscle force evaluation method and system based on slow blinking video analysis
By using a dual-path recognition network based on slow blinking video analysis, combined with text and visual features, the problem of assessing levator palpebrae superioris muscle strength in children has been solved, enabling rapid and accurate muscle strength assessment and treatment plan reference, which is suitable for telemedicine.
Patent Information
- Application Number
- CN202510027784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-01-08
AI Technical Summary
Existing techniques are insufficient for effectively assessing the strength of the levator palpebrae superioris muscle in children, especially during preoperative assessment, where differences in children's understanding and cooperation levels make standardized clinical assessment difficult.
A method based on slow-speed blinking video analysis is adopted. By constructing a dual-path recognition network and combining textual and visual features, the levator palpebrae superioris muscle strength is assessed. This includes a text feature extraction module, a slow-speed channel module, a fast-speed channel module, a fusion module, and a prediction module. A deep learning model is used to predict the muscle strength level.
It enables rapid and accurate assessment of levator palpebrae superioris muscle strength, provides a more comprehensive reference for treatment plans, is applicable to the field of telemedicine, and optimizes the allocation of regional medical resources.
Smart Images

Figure CN119700093B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of medical image auxiliary analysis, and particularly relates to a levator muscle force evaluation method and system based on slow blinking video analysis. BACKGROUND
[0002] Ptosis is a common and frequently-occurring disease in clinical ophthalmology, which refers to the dysfunction or loss of the levator muscle and Müller's smooth muscle, so that the upper eyelid is partially or completely drooping, and the pupil is partially or completely covered. Severe congenital ptosis can cause form sensation deprivation, leading to amblyopia and astigmatism. However, such complications can be alleviated or avoided through early diagnosis and surgical correction.
[0003] Ptosis surgery can be briefly divided into three categories: percutaneous approach, transconjunctival approach and frontal muscle suspension. The choice of surgical approach depends largely on the function of the levator muscle, i.e. the offset distance between the lower and upper gaze of the upper eyelid margin with the eyebrow fixed. However, due to the differences in understanding and cooperation levels of children, especially infants, standard preoperative clinical assessment becomes extremely difficult.
[0004] Patent document CN117523633A discloses an automatic identification method for eye ptosis based on image deep learning, which includes: collecting a front face image and extracting an eye image; inputting the eye image into an eye segmentation model to obtain a segmented eye image; performing Hough transform detection on the segmented eye image to obtain an iris contour and perform circle fitting to obtain an iris circle, and then determining the center of the iris; performing image correction; performing contour extraction detection on the segmented eye image to determine the upper eyelid margin contour, lower eyelid margin contour, intersection point of the iris margin and upper eyelid margin; using an iris diameter ruler to measure MRD1 and MRD2; determining the iris-lid margin intersection angle, inner intersection angle and outer intersection angle formed by the intersection point of the iris and the upper eyelid margin; and identifying the degree of eye ptosis.
[0005] Patent document CN116386114A discloses a face feature recognition system based on visual detection, an image storage module for storing preoperative facial images, surgical plan images and postoperative facial images of a plurality of historical subjects, a neuron network module for storing a neural network model, an image acquisition module for acquiring preoperative facial images of a subject to be operated, an image input module connected with the neuron network module and the image acquisition module respectively for inputting the preoperative facial images of the subject to be operated acquired by the image acquisition module into the neural network model, a data processing module connected with the neuron network module for analyzing and processing mask data output by the neural network model, a feature calculation module connected with the data processing module for calculating a double eyelid surgery plan of the subject to be operated according to the analysis and processing result of the data processing module, and a plan adjustment module connected with the image storage module and the feature calculation module respectively for comparing the double eyelid surgery plan of the subject to be operated with the surgical plan of the historical subject to adjust the double eyelid surgery plan of the subject to be operated. SUMMARY
[0006] The purpose of the present application is to provide a levator muscle strength evaluation method and system based on slow blinking video analysis, which can quickly evaluate the levator muscle strength to provide a more perfect reference for subsequent treatment plans.
[0007] To achieve the first purpose of the present application, the following technical solution is provided: a child levator muscle strength evaluation method based on slow blinking video analysis, comprising the following steps:
[0008] Input the slow blinking video and physiological information of the subject, divide the slow blinking video into a plurality of video segments along the time frame, and label the plurality of video segments based on eye and levator muscle strength grade;
[0009] Divide the video segments into eye regions of interest using a face marker detection method to obtain corresponding single-eye blinking video segments, and form a data set by combining the physiological information, the slow blinking video, the corresponding single-eye blinking video segments and the labels;
[0010] A dual-path recognition network based on deep learning is constructed, which includes a text feature extraction module, parallel slow and fast channel modules, a fusion module and a prediction module for aggregating the slow and fast channel modules;
[0011] The text feature extraction module is used to extract text features of the physiological information and expand the text features in the visual feature dimension to obtain a corresponding text matrix;
[0012] The slow channel module is used for extracting semantic information of each frame of video segment in the slow blinking video to obtain a corresponding visual feature sequence, and performing element-by-element multiplication operation on the visual feature sequence and the text matrix to obtain a space-time feature sequence.
[0013] The fast channel module is used for extracting eye rapid change information of each frame of video segment in the slow blinking video to obtain a time domain feature sequence.
[0014] The fusion module is used for transversely splicing the space-time feature sequence and the time domain feature sequence to obtain a fusion feature sequence.
[0015] The prediction module is used for predicting according to the input fusion feature sequence to output a prediction result, and the prediction result includes eye and levator muscle strength grade.
[0016] The dual-path recognition network is trained by using a data set to obtain a video classification model for recognizing levator muscle strength state.
[0017] The slow blinking video to be recognized and physiological data are input into the video classification model to output a corresponding prediction result.
[0018] The present application uses a multi-modal data fusion mode to enhance the data features in the final prediction, so that a more accurate and perfect reference is obtained.
[0019] Specifically, the slow blinking video is a spontaneous blinking video of a subject shot by a 240FPS slow shooting mode.
[0020] Specifically, the levator muscle strength grade is divided into three levels, including muscle strength good, muscle strength good and muscle strength poor.
[0021] When the levator muscle function is greater than 10mm, it is defined as muscle strength good.
[0022] When the levator muscle function is greater than or equal to 5mm and less than or equal to 10mm, it is defined as muscle strength good.
[0023] When the levator muscle function is less than or equal to 4mm, it is defined as muscle strength poor.
[0024] Specifically, the eye region of interest includes the eye inner canthus, outer canthus, upper and lower eyelids, eyelid fissure and eyebrow.
[0025] Specifically, the slow channel module adopts a feature extraction mode with low sampling rate and high channel number to obtain space-time features rich in semantic information in the slow blinking video.
[0026] Specifically, the fast channel module adopts a feature extraction mode with high sampling rate and low channel number to obtain time domain features of rapid changes in the slow blinking video.
[0027] Specifically, the dual-path recognition network adopts a 3D ResNet101 convolutional neural network model as a network framework for construction.
[0028] Specifically, the physiological data includes gender, age, palpebral fissure height, and upper eyelid margin corneal reflection distance of the subject.
[0029] To achieve the second object of the application, the following technical solution is provided: a levator muscle strength evaluation system, which is realized by the above-mentioned levator muscle strength evaluation method based on slow blinking video analysis for children, and includes a data acquisition unit, a video analysis unit, and an inference evaluation unit.
[0030] The data acquisition unit is configured to acquire a slow blinking video of a subject.
[0031] The video analysis unit is configured to analyze the input slow blinking video to output a prediction result.
[0032] The inference evaluation unit is configured to perform similarity matching on the prediction result and historical case data to output historical case data with the highest similarity value as an inference evaluation result.
[0033] The historical case data includes patient's eye, levator muscle strength grade, and corresponding medical record information.
[0034] Specifically, the inference evaluation result includes:
[0035] When the levator muscle strength is good, regular follow-up is recommended.
[0036] When the levator muscle strength is good, comprehensive evaluation of the operation opportunity is recommended, and the operation method can be selected as a transcutaneous approach or a transconjunctival approach.
[0037] When the levator muscle strength is poor, surgery is recommended as soon as possible, and the operation method can be selected as a frontal muscle suspension operation.
[0038] Compared with the prior art, the present application has the following advantages:
[0039] Based on the dual-path video recognition network, visual features are extracted from the video data, and the text information and the visual features in the slow channel are fused and spliced with the visual features in the fast channel, thereby providing a more comprehensive fusion feature sequence for subsequent prediction work.
[0040] Compared with the clinical standardized levator muscle function test, the blinking video is easy to obtain and convenient to transmit, and the intelligent analysis system has low technical requirements for the operator and the cooperation degree of the photographed children, which is expected to be applied in the field of telemedicine, optimize the allocation of regional medical resources, and greatly improve the clinical management efficiency of children with ptosis. Attached Figure Description
[0041] Figure 1 A flowchart of the levator palpebrae superioris muscle strength assessment method based on slow blink video analysis provided in this embodiment;
[0042] Figure 2 This is a schematic diagram of the framework of the dual-path recognition network provided in this embodiment;
[0043] Figure 3 This is a framework diagram of the levator palpebrae superioris muscle strength assessment system provided in this embodiment;
[0044] Figure 4 The inference logic diagram of the inference evaluation unit provided in this embodiment. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0046] like Figure 1 As shown in this embodiment, a method for assessing levator palpebrae superioris muscle strength based on slow blink video analysis is provided, which includes the following steps:
[0047] Input the subject's slow blink video and physiological information, divide the slow blink video into multiple video segments along the time frame, and label the multiple video segments based on eye type and levator muscle strength level;
[0048] The facial tag detection method is used to divide the video clips into regions of interest for the eyes to obtain the corresponding monocular blink video clips. The physiological information, slow blink video, corresponding monocular blink video clips and labels are combined to form a dataset.
[0049] A deep learning-based dual-path recognition network is constructed, which includes a text feature extraction module, a parallel slow channel module and a fast channel module, as well as a fusion module and a prediction module for summarizing the slow channel module and the fast channel module.
[0050] The text feature extraction module is configured to extract text features of the physiological information, and expand the text features in a visual feature dimension to obtain a corresponding text matrix.
[0051] The slow channel module is configured to extract semantic information of each frame of video segment in the slow blinking video to obtain a corresponding visual feature sequence, and perform an element-by-element multiplication operation on the visual feature sequence and the text matrix to obtain a spatiotemporal feature sequence.
[0052] The fast channel module is configured to extract eye rapid change information of each frame of video segment in the slow blinking video to obtain a time series feature sequence.
[0053] The fusion module is configured to horizontally splice the spatiotemporal feature sequence and the time domain feature sequence to obtain a fusion feature sequence.
[0054] The prediction module is configured to perform prediction according to the input fusion feature sequence to output a prediction result, the prediction result including eye and levator muscle strength grade.
[0055] The dual-path recognition network is trained by using the data set to obtain a video classification model for recognizing the levator muscle strength state.
[0056] The slow blinking video to be recognized and the physiological data are input into the video classification model to output a corresponding prediction result.
[0057] Further, in the embodiment, the slow blinking video of a child is used to illustrate the specific process.
[0058] 1606 blinking video segments of children treated in an ophthalmology center of a hospital are collected, the subjects including children with ptosis and children without ptosis, and the levator muscle strength including three levels of good, good and poor.
[0059] In the embodiment, the ordinary smart phone is used to shoot the blinking video segments in a 240 FPS slow motion mode.
[0060] The physiological data of the subjects include the gender, age, palpebral fissure height and upper eyelid margin corneal reflection distance of the subjects.
[0061] The shot blinking video segments need to be quality screened, which includes: (1) poor quality video, which means defective or image with focal length or illumination;
[0062] (2) poor cooperation video, which means that the whole face is not in the shot or the subject does not look straight during the video acquisition process.
[0063] In the embodiment, 38 poor quality blinking videos are filtered, and the remaining 1568 qualified blinking videos are used as the data set.
[0064] Based on the results of the clinical standardized levator function test, manual label annotation was performed on the data set, including annotating the levator muscle strength grade, which includes three levels: good muscle strength, good muscle strength, and poor muscle strength. Specifically, levator function > 10 mm is defined as good muscle strength, levator function 5-10 mm is defined as good muscle strength, and levator function < 4 mm is defined as poor muscle strength.
[0065] Finally, in the data set, 704 video clips were classified as good muscle strength, 458 video clips were classified as good muscle strength, and the remaining 406 video clips were classified as poor muscle strength.
[0066] In addition, in order to highlight the eye region in the data set, based on the characteristics of the eye region of interest position, the eye region of interest is located by a human face marking detection method to frame the eye region of interest in the video frame, including the inner and outer canthi of the eye, the upper and lower eyelids, the palpebral fissure and the eyebrows.
[0067] After locating the eye region of interest, the video clips of single eye blinking are obtained by frame-by-frame cropping and reorganization. In this embodiment, according to the Pareto principle, the preprocessed data set is randomly split in a ratio of 8:2, of which 1254 (80%) video clips are used as the training set, and the remaining 314 (20%) video clips are used as the test set.
[0068] A dual-path recognition network based on deep learning is constructed, and the dual-path recognition network is trained using the preprocessed training set data. The slow and fast channels extract the space-time features and time domain features in the slow blinking video, respectively, to identify the abnormal muscle strength state of the levator muscle, and obtain the trained video classification model.
[0069] As shown in Figure 2 , the dual-path recognition network proposed in this embodiment, both paths of which use a 3D ResNet-101 convolutional neural network model as the backbone, include a slow path (Slow pathway) with low sampling rate and high channel number, which mainly extracts space-time features in the video for analyzing semantic rich static content; a fast path (Fast pathway) with high sampling rate and low channel number, which mainly extracts time domain features in the video for analyzing fast changing dynamic content. The slow path parameters are set as: time step , sampling frame ; The fast path parameters are set as: time step , sampling frame ; the channel ratio between the fast and slow paths ; the learning rate is 0.01; the batch size is set to 16, and the model is trained for 256 iterations.
[0070] In addition, in order to obtain more rich features, text information is also added to the visual features, and the process is as follows:
[0071] First, the input physiological information is subjected to text feature extraction, and then the text features are transformed through a linear layer to map the text features to a dimension more compatible with the visual features, and an activation function Sigmoid is applied to introduce nonlinearity into the text features; finally, the text features are reshaped into a 7x7 spatial dimension to construct a text matrix with the same dimension as the visual features.
[0072] In the slow channel module, the visual feature sequence of the input video frame is extracted through standard operations such as convolution, batch normalization, ReLU, max pooling and residual block, the dimension of the text matrix is expanded, and the shape is changed to (-1, 1, 1, 7, 7), which is aligned with the visual features in the first visual feature sequence, so as to perform element-wise multiplication. Then the softmax function is applied on the channel dimension (`dim=1`) to convert the value on each 7x7 spatial grid of the text features into a probability distribution. Finally, the visual features in the visual feature sequence are multiplied with the text matrix element by element, which enables the model to weight different spatial positions of the visual features according to the text information, thereby obtaining the final spatio-temporal feature sequence.
[0073] The fast channel module directly extracts the time sequence feature sequence in the input video frame.
[0074] In the fusion module, the final spatio-temporal feature sequence and the time domain feature sequence are horizontally concatenated to output the fusion feature sequence.
[0075] In order to evaluate the classification performance of the video classification model, the precision, recall, accuracy and F1 score of the children's upper eyelid levator muscle strength grading evaluation in the test set are calculated, and the results are shown in Table 1: .
[0076] As shown in Figure 3 , the present embodiment also provides a levator muscle strength evaluation system, which is realized by the levator muscle strength evaluation method based on slow blinking video analysis provided in the above embodiments, and includes a data acquisition unit, a video analysis unit and an inference evaluation unit.
[0077] The data acquisition unit is used to acquire the slow blinking video and physiological data of the subject.
[0078] The video analysis unit analyzes the input slow blinking video and physiological data to output a prediction result.
[0079] The inference evaluation unit performs similarity matching on the prediction result and historical case data to output historical case data with the highest similarity value as an inference evaluation result, wherein the historical case data includes the patient's eye side, levator muscle strength grade, and corresponding medical record information.
[0080] As shown in the embodiment, the inference evaluation result includes: Figure 4
[0081] When the levator muscle strength is good, regular follow-up is recommended.
[0082] When the levator muscle strength is good, comprehensive evaluation of the operation opportunity is recommended, and the operation method can be selected as a percutaneous approach or a conjunctival approach.
[0083] When the levator muscle strength is poor, surgery is recommended as soon as possible, and the operation method can be selected as a frontal muscle suspension operation.
[0084] In addition, the terms "upper", "lower", "inner", "outer", "front", "back" are only for the purpose of description, and cannot be understood as indicating or implying relative importance. Unless otherwise specified, the relative steps, numerical expressions and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0085] Of course, the above only describes specific embodiments of the present application, and is not intended to limit the scope of the present application. Any equivalent changes or modifications made in accordance with the principles and features described in the patent application of the present application shall be included in the patent application of the present application.
[0086] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacements to some technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for evaluating the levator muscle force based on slow blinking video analysis, characterized in that, The method comprises the following steps: inputting slow blinking video and physiological information of a subject, dividing the slow blinking video into multiple video segments along a time frame, and labeling the multiple video segments based on eye and levator muscle force grade; dividing the video segments into eye regions of interest using a facial landmark detection method to obtain corresponding single-eye blinking video segments, and forming a data set by combining physiological information, slow blinking video, corresponding single-eye blinking video segments, and labels; the physiological information includes gender, age, palpebral fissure height, and upper eyelid margin corneal reflection distance of the subject; a dual-path recognition network based on deep learning is constructed, which includes a text feature extraction module, parallel slow and fast channel modules, a fusion module for aggregating the slow and fast channel modules, and a prediction module; the text feature extraction module is used to extract text features of the physiological information and expand the text features in the visual feature dimension to obtain a corresponding text matrix; the expansion is reshaping the text features into a 7x7 spatial dimension to construct a text matrix with the same dimension as the visual features; the slow channel module is used to extract semantic information of each frame of video segment in the slow blinking video to obtain a corresponding visual feature sequence, and perform element-wise multiplication operation on the visual feature sequence and the text matrix to obtain a spatiotemporal feature sequence; the element-wise multiplication operation is to apply a softmax function on the channel dimension `dim=1` to convert the values on each 7x7 spatial grid of the text features into a probability distribution, and then perform element-wise multiplication operation on the visual features in the visual feature sequence and the text matrix; the fast channel module is used to extract eye rapid change information of each frame of video segment in the slow blinking video to obtain a time series feature sequence; the fusion module is used to horizontally concatenate the spatiotemporal feature sequence and the time domain feature sequence to obtain a fusion feature sequence; the prediction module predicts according to the input fusion feature sequence to output a prediction result, which includes eye and levator muscle force grade; the dual-path recognition network is trained using the data set to obtain a video classification model for identifying levator muscle force state; the slow blinking video and physiological data to be identified are input into the video classification model to output the corresponding prediction result.
2. The slow blink video analysis based levator muscle force evaluation method according to claim 1, characterized in that, The slow blinking video is a spontaneous blinking video of the subject shot by a 240FPS slow shooting method.
3. The slow blink video analysis based levator muscle force evaluation method according to claim 1, wherein, The levator muscle force grade is divided into three levels, including good muscle force, good muscle force, and poor muscle force; when the levator muscle function is greater than 10mm, it is defined as good muscle force; when the levator muscle function is greater than or equal to 5mm and less than or equal to 10mm, it is defined as good muscle force; when the levator muscle function is less than or equal to 4mm, it is defined as poor muscle force.
4. The slow blink video analysis based levator muscle force evaluation method according to claim 1, wherein, The eye region of interest includes the inner and outer canthi of the eye, the upper and lower eyelids, the palpebral fissure, and the eyebrow region around the eye.
5. The slow blink video analysis based levator muscle force evaluation method according to claim 1, wherein, The slow channel module adopts a feature extraction method with low sampling rate and high channel number.
6. The slow blink video analysis based levator muscle force evaluation method according to claim 1, wherein, The fast channel module adopts a feature extraction method with high sampling rate and low channel number.
7. The slow blink video analysis based levator muscle force evaluation method according to claim 1, wherein, The dual-path identification network is constructed by using a 3D ResNet101 convolutional neural network model as a network framework.
8. A levator muscle strength assessment system, comprising: The levator muscle force evaluation method based on slow blinking video analysis according to any one of claims 1-7 is implemented, which comprises a data acquisition unit, a video analysis unit and an inference evaluation unit. The data acquisition unit is used to acquire slow blinking video and physiological data of a subject. The video analysis unit analyzes the input slow blinking video and physiological data to output a prediction result. The inference evaluation unit performs similarity matching on the prediction result and historical case data to output the historical case data with the highest similarity value as the inference evaluation result. The historical case data includes patient's eye, levator muscle force grade and corresponding medical record information.
Citation Information
Patent Citations
Facial feature recognition system based on visual detection
CN116386114A
Eye upper blepharoptosis automatic identification method based on image deep learning
CN117523633A
Eyeball motion segmentation positioning method based on cyclic residual convolutional neural network
CN114694236A
Image classification method and device, computer equipment and storage medium
CN115690509A