Method, apparatus, and storage medium for measuring palpebral fissure width

JP2025524078A5Active Publication Date: 2025-09-30SHANGHAI INST FOR ENDOCRINE & METABOLIC DISEASES +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504147
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-18
Filing Date
2023-08-17
Publication Date
2025-09-30
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

Current methods for measuring palpebral fissure width are inaccurate due to operator dependence and fail to distinguish eye features accurately, leading to potential misdiagnosis in conditions like thyroid-associated ophthalmopathy.

Method used

A method using near-infrared imaging and a combined UNet and DenseNet neural network model to segment the eye image, calculating palpebral fissure width based on the intersections of sclera, iris, and pupil, with a composite loss function for training accuracy.

Benefits of technology

Accurately measures palpebral fissure width by distinguishing eye features, reducing measurement errors and improving diagnostic precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, and storage medium for measuring the palpebral fissure width. The method includes the steps of: capturing a first eye position image from the front view of a user at a forward position of the eyeball in a near-infrared light field of 700 to 1200 nm; using a neural network training method to segment the background, iris, sclera, and pupil from the first eye position image; extracting the pupil center from the segmented pupil and obtaining a vertical pupil center line; obtaining the distance between intersections of the sclera, iris, or pupil and the background on the pupil center line, and calculating the palpebral fissure width using the distance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of eye detection, and specifically to a method, apparatus, and storage medium for measuring palpebral fissure width.

Background Art

[0002] The palpebral fissure, also called the eyelid fissure, and the palpebral fissure width refer to the distance between the upper and lower eyelids passing through the pupil. Clinically, eyelid retraction includes retraction of the upper and lower eyelids, is a clinical condition commonly seen in thyroid-associated ophthalmopathy (TAO), and is also an important indicator for clinical diagnosis. Eyelid retraction not only causes a disfigured face that is unacceptable to patients, but may also cause exposure keratopathy that threatens vision, such as corneal ulcers. Therefore, the measurement of eyelid retraction is important for clinical diagnosis, and the degree of eyelid retraction can be reflected by measuring the palpebral fissure width.

[0003] Currently, clinically, the palpebral fissure width is usually detected using a millimeter scale. However, in this method, during measurement, the reading value is not accurate and is affected by the level and habits of the operator, so errors are likely to occur in the measurement data. In the diagnosis and evaluation of TAO patients, a fluctuation of 2 mm often indicates a change in the patient's condition, so the requirement for measurement accuracy is often high. However, the above measurement method may cause missed diagnosis and misdiagnosis of patients, delaying the progression of the patient's condition.

[0004] In the prior art, there are also technical means for determining whether the eyelids have retracted using an automatic measurement method. As described in the Chinese patent "Method and Apparatus for Identifying Ocular Symptoms of Thyroid-Associated Ophthalmopathy" (Application No.: 202010803761.6, Publication Date: October 30, 2020), the cornea and sclera are identified through image identification and neural network training, and it is determined whether the upper edge of the upper eyelid and the corneal region is exposed in the scleral region based on the images of the cornea and sclera, thereby determining whether eyelid retraction exists. However, this method gives false judgments for people with exophthalmos or people with protruding eyeballs due to myopia, resulting in the exposure of the white of the eye. As described in the Chinese patent "Method and System for Analyzing Blink Frequency Based on Image Processing" (Application No.: 201910939612.X, Publication Date: February 04, 2020), the iris contour and scleral contour are determined from the captured human eye image, the eyelid fissure boundary is determined based on the iris contour and scleral contour, and the height of the eyelid fissure is calculated based on the difference between the coordinates of the upper boundary point and the lower boundary point of the human eyelid fissure. However, the height of the eyelid fissure is defined as the distance between the upper and lower eyelids passing through the pupil center line, and the height of the eyelid fissure determined by the above method may be the oblique distance of the eyelid fissure.

[0005] There is also prior art for segmenting eye images using a neural network model training method. However, its training model is relatively old, the training process is time-consuming, and the occupied memory space is large. Therefore, it is necessary to further improve the prior art.

Summary of the Invention

Problems to be Solved by the Invention

[0006] In order to eliminate the deficiency in the measurement accuracy of the eyelid fissure width in the prior art, the present application provides a method for measuring the eyelid fissure width.

Means for Solving the Problems

[0007] To achieve the above object, the present application uses the following technical means.

[0008] A method for measuring the eyelid fissure width according to one aspect is In the near-infrared light field of 700~1200nm, obtaining a first eye position image from the front view of the user at the front position of the eyeball by photographing; Segmenting the background, iris, sclera, and pupil from the first eye position image using a neural network training method; Extracting the pupil center from the segmented pupil and obtaining a vertical pupil center line; Calculating the distance between the intersections of the sclera, iris, or pupil and the background on the pupil center line, and calculating the palpebral fissure width using the distance.

[0009] Furthermore, the pupil center is the average value of the X and Y coordinates of all pupil pixels extracted from the first eye position image, and a vertical line is drawn using the average value of the X coordinates to obtain the pupil center line.

[0010] Furthermore, the step of segmenting the background, iris, sclera, and pupil from the first eye position image using a neural network training method uses a combination of a UNet neural network model and a DenseNet neural network model as the neural network model. In the neural network model, the first eye position image is used as the input, and dimensionality reduction and dimensionality expansion are performed on the input using the UNet neural network model. The output of the dimensionality reduction block is transmitted to the corresponding dimensionality expansion block through skip connections. In each dimensionality reduction block, feature extraction is performed using the DenseNet neural network model, and in each dimensionality expansion module, upsampling is performed using the DenseNet neural network model.

[0011] Furthermore, the method for measuring the palpebral fissure width further includes the step of verifying the training effect of the neural network using a loss function. The loss function is a composite loss and is composed of a focal loss, a generalized dice loss, a surface loss, and a boundary recognition loss.

[0012] The loss function is a composite loss, and the calculation method is as follows: Composite loss = a1 × surface loss + a2 × focus loss + a3 × generalized die loss + a4 × boundary recognition loss × focus loss, where a1, a2, a3, and a4 are hyperparameters. a1 is related to the number of time steps during the training process, a2 = 1, a3 = 1 - a1, and a4 = 20.

[0013] Furthermore, the palpebral fissure width is defined as: palpebral fissure width (B) = pixel distance of the palpebral fissure (A) × length of a single pixel in the palpebral fissure direction × distance from the palpebral fissure to the camera lens (D) ÷ distance from the camera photosensitive sensor to the camera lens (C). Here, the pixel distance of the palpebral fissure (A) is the distance between the intersections of the sclera, iris, or pupil and the background on the pupil center line.

[0014] Specifically, the distance from the palpebral fissure to the camera lens is: distance from the palpebral fissure to the camera lens (D) = distance from the canthus locking point to the camera lens (1) - exophthalmos (2).

[0015] Furthermore, the exophthalmos (2) uses the average value of the exophthalmos of healthy individuals.

[0016] In one aspect, the measuring device for palpebral fissure width according to the present application includes an image acquisition module that acquires a first eye position image from the front view of the user at the front position of the eyeball in a near-infrared light field of 700 - 1200 nm, an image segmentation module that segments the background, iris, sclera, and pupil from the first eye position image using a neural network training method, a feature extraction module that extracts the pupil center from the segmented pupil and obtains the pupil center line in the vertical direction, a calculation module that calculates the distance between the intersections of the sclera, iris, or pupil and the background on the pupil center line and calculates the palpebral fissure width using this distance.

[0017] In one aspect, the computer-readable storage medium according to the present application stores at least one program code that is loaded and executed by a processor to implement the method for measuring the palpebral fissure width.

Advantages of the Invention

[0018] The technical means of the present application has at least the following beneficial effects compared with the prior art.

[0019] 1. By capturing an eye image in a near-infrared light field, the iris, pupil, sclera, background, and lacrimal caruncle can be effectively distinguished, and an intuitive and easy-to-discriminate image can be provided through subsequent neural network training.

[0020] 2. The present application trains a neural network using a combined form of a UNet neural network model and a DenseNet neural network model. By doing so, not only does it utilize the advantages that the UNet neural network model has a simple and stable structure and is widely applicable to medical image processing, but also it applies the DenseNet neural network model to each dimensionality reduction block and dimensionality expansion block of the UNet neural network model, and performs feature reuse through skip connection operations in each dimensionality reduction block and dimensionality expansion block, thereby improving the feature extraction effect. Moreover, the DenseNet neural network model used in the present application simplifies the conventional DenseNet neural network model and avoids the problem that the conventional DenseNet neural network model needs to be associated with all previous layers in each layer, resulting in overly slow calculations.

[0021] 3. By accurately discriminating the iris, pupil, sclera, background, and lacrimal caruncle with a neural network model, the intersection of the sclera, iris, or pupil on the pupil center line and the background can be discriminated, and the distance between the two intersections is the pixel distance of the palpebral fissure width, and the distance of the palpebral fissure width can be obtained through geometric calculation.

Brief Description of the Drawings

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0023] Hereinafter, with reference to the drawings and examples, specific embodiments of the present application will be described in more detail. The following examples are for explaining the present application and do not limit the scope of the present application.

[0024] Based on the idea of neural network model training, the present application performs feature extraction on a captured eye image to find the pupil center, further obtains the pupil center line, positions the pixel distance of the palpebral fissure width based on the intersection of the pupil center line and the eye socket, and further calculates the palpebral fissure width according to the imaging principle. Hereinafter, the concept of the present application will be further explained with specific examples.

[0025] In one aspect, the method for measuring the palpebral fissure width shown in FIG. 1 includes the following steps S1 to S4.

[0026] In step S1, in the near-infrared light field of 700 to 1200 nm, a first eye position image from the front view of the user is captured and obtained at the front position of the eyeball.

[0027] In step S1, in order to capture an eye image in the near-infrared light field of 700 - 1200 nm, in the wavelength band of normal visible light of 400 - 700 nm, the colors of different parts of the eye, namely the pupil, iris, and sclera, have little effect on imaging, and due to the gradation structure of the corneal limbus at the contact part between the iris and the sclera, the center of the eyeball cannot be accurately identified. Human melanin pigment has an absorption peak at approximately 335 nm and is hardly absorbed at all in the wavelength band exceeding 700 nm. Since the reflectance of the iris is quite stable within the near-infrared band with a wavelength exceeding 700 nm, by using the near-infrared light field, the sclera, iris, and pupil boundaries can be well distinguished, thereby enabling better improvement of the accuracy and stability of the algorithm.

[0028] In step S2, the background, iris, sclera, and pupil are segmented from the first eye position image using the training method of the neural network.

[0029] In step S2 above, a combination of the UNet neural network model and the DenseNet neural network model is used as the neural network model. In the neural network model, the first eye position image is used as the input. The UNet neural network model is used to perform dimensionality reduction and dimensionality expansion on the input. The output of the dimensionality reduction block is transmitted to the corresponding dimensionality expansion block through skip connection. In each dimensionality reduction block, feature extraction is performed using the DenseNet neural network model, and in each dimensionality expansion module, upsampling is performed using the DenseNet neural network model.

[0030] The architecture of the neural network used in this application is a combination of UNet and DenseNet, and is realized based on the Pytorch framework. The overall structure of the neural network is in a U shape. As shown in Figure 2, there are a total of five dimensionality reduction blocks, namely, dimensionality reduction block 1 to dimensionality reduction block 5, and four dimensionality expansion blocks, namely, dimensionality expansion block 1 to dimensionality expansion block 4.

[0031] The neural network receives the input image and transmits it from dimensionality reduction block 1 to dimensionality reduction block 5. Except for dimensionality reduction block 1, each dimensionality reduction block is connected by one max pooling layer, which halves the dimensions. Then, dimensionality reduction block 5 transmits the output from dimensionality expansion block 1 to dimensionality expansion block 4. The output from dimensionality expansion block 4 is transmitted to the final convolution. Examples of skip connections include connecting dimensionality reduction block 1 to dimensionality expansion block 4, dimensionality reduction region 2 to dimensionality expansion block 3, dimensionality reduction block 3 to dimensionality expansion block 2, and dimensionality reduction block 4 to dimensionality expansion block 1.

[0032] UNet provides a single deep convolutional structure. First, it reduces the input dimensions to a certain extent, then restores the input to its original size, and through skip connections, realizes the transmission of information from the dimensionality reduction blocks to the corresponding dimensionality expansion blocks, and the output of the initial layer should be transmitted as input to the subsequent layers, thereby preventing information loss inside the deep neural network and enabling maximum learning. On the other hand, DenseNet is a concept rather than a fixed structure. The focus on DenseNet in this application is the specific skip connection technology. In this application, an optimized reduced version is proposed. When each dimensionality expansion block and dimensionality reduction block perform the skip connection operation, it is not necessary to use the output of all previous layers as input like each layer of the conventional DenseNet network model. By performing 1 to 2 skip connection operations, it is possible to prevent the network from being too slow.

[0033] To explain the training process of the neural network, the following concepts are introduced.

[0034] Regarding max pooling, max pooling is a pooling operation that calculates the maximum value in each patch of each feature map. As a result, instead of the average presence of features when averaged, a downsampled or aggregated feature map is obtained in which the most prominent features in the patch are emphasized. This operation corresponds to the torch.nn.MaxPool2d function, the kernel size is defined as 2×2, and this operation reduces the input dimension by half.

[0035] Regarding convolution, convolution simply applies one filter (also called a kernel) to one input to generate one activation. Repeatedly applying the same filter to the input generates an activation map called a feature map, which indicates the location and intensity of the features detected from the input. This operation corresponds to the torch.nn.Conv2d function.

[0036] Regarding same convolution, same convolution is a type of convolution where the output matrix has the same dimensions as the input matrix. This operation corresponds to the torch.nn.Conv2d function. In the example of this application, the padding is 1×1 and the kernel size is 3×3.

[0037] Regarding valid convolution, valid convolution is a convolution operation that does not use any padding on the input. Such an operation corresponds to the torch.nn.Conv2d function. In the example of this application, there is no padding and the kernel size is 1×1.

[0038] Regarding transposed convolution, the transposed convolution layer attempts to reconstruct the spatial dimensions of the convolution layer and apply its downsampling and upsampling techniques in reverse. This operation corresponds to torch.nn.ConvTranspose2d.

[0039] Regarding batch normalization, batch normalization (also called batch normalization) is a method that makes artificial neural networks faster and more stable by re-centering and re-scaling the input of a layer. This operation corresponds to BatchNorm2d, and num_features is 5, which is the number of labels of the network.

[0040] Regarding the leaky RELU (Leaky RELU), the leaky rectified linear activation function or simply the leaky RELU is a piecewise linear function that directly outputs the input if it is positive, and otherwise outputs the product of a small factor (0.1 in the example of this application) and the input.

[0041] The operation processes of the specific dimension expansion block and dimension reduction block are as follows.

[0042] In the dimension reduction block, several convolutional operations, as well as Leaky RELU and skip connection operations, are applied. In this application, the size of the original image is n×n×1 dimension is taken as an example for explanation. Here, the 1 dimension indicates that the original image has only grayscale values. The specific structure description is as shown in Figure 3.

[0043] First, the original input of the dimensionality reduction block passes through the same convolutional layer (Layer-1) with 32 filters (channels). Then, Leaky RELU activation (RELU-Layer-1) is performed on the output of Layer-1. Next, RELU-Layer-1 and the original input are connected in series along the channel dimension (Skip-Connection-Input-1). As a result, the dimension of the output is the sum of 32 and the dimension of the original input. This is the first time the skip connection operation is applied, and the original input is directly transmitted to the later convolutional stage. Third, Skip-Connection-Input-1 is transmitted to an effective convolutional layer (Layer-2) with 32 filters, reducing the number of dimensions to 32 dimensions. This is generally called the "bottleneck layer", which shrinks the channel dimension of the neural network to ensure that the network does not become too slow. Fourth, the output of Layer-2 passes through the same convolutional layer (Layer-3) with 32 filters (channels) again. Then, leaky RELU activation (RELU-Layer-3) is performed on the output of Layer-3. Fifth, RELU-Layer-3 is connected to RELU-Layer-1 and the original input (Skip-Connection-Input-2). This is the second time the skip connection operation is applied. Sixth, Skip-Connection-Input-2 passes through an effective convolutional layer (Layer-4) with 32 filters again, and this layer reduces the dimension to 32 dimensions again. Seventh, the output of Layer-4 passes through the same convolutional layer (Layer-5) with 32 filters (channels) again. Then, one leaky RELU activation (RELU-Layer-5) is added to the output of Layer-5. Finally, the result of RELU-Layer-5 is transmitted to the batch normalization layer.

[0044] Also, before applying any convolution, each dimensionality reduction block applies a max pooling layer to the input, except for the first block, since the input of the first dimensionality reduction block is the original image itself. This is where "dimensionality reduction" actually occurs, as the layer reduces the dimensionality of the input by half. The dimensionality expansion block restores the dimensions to their original values.

[0045] The dimensionality expansion block is almost the same as the dimensionality reduction block, except that a transposed convolution is applied instead of max pooling, and the dimensionality expansion block also receives a skip connection from the dimensionality reduction block. The description of the specific structure is also as shown in FIG. 4.

[0046] First, the dimensionality expansion block applies a transposed convolution (Layer-1) to the input, doubling the dimensionality of the input, where "dimensionality expansion" is performed. Second, the output of Layer-1 is connected along the channel dimension with the skip connection from the dimensionality reduction block (Skip-Connection-Input-1), which is an example of the structural skip connection technique. Third, Skip-Connection-Input-1 passes through a convolutional layer (Layer-2) with 32 filters to reduce the dimensionality to 32 dimensions. Fourth, the output of Layer-2 passes through the same convolutional layer (Layer-3) with 32 filters (channels) again. Then, one leaky ReLU activation (RELU-Layer-3) is added to the output of Layer-3. Fifth, RELU-Layer-3 and Skip-Connection-Input-1 perform a skip connection operation to make the dimensionality of the output 96 (Skip-Connection-Input-2). Sixth, Skip-Connection-Input-2 passes through a convolutional layer (Layer-4) with 32 filters again to reduce the dimensionality to 32 dimensions again. Seventh, the output of Layer-4 passes through the same convolutional layer (Layer-5) with 32 filters (channels) again. Then, one leaky ReLU activation (RELU-Layer-5) is added to the output of Layer-5. Finally, the result of RELU-Layer-5 is transmitted to the batch normalization layer.

[0047] Through the dimension expansion block and the dimension reduction block, the output is a matrix of dimension N×N×32, where N and N are the dimensions of the image. The matrix of dimension N×N×32 passes through the final effective convolution with 5 filters (channels), and the final output is a matrix of 1×N×N×5. Intuitively, this matrix assigns 5 values to each pixel, and these 5 values represent the probabilities of the labels that each pixel should have. Here, 0 is the label of the background, 1 is the label of the sclera, 2 is the label of the iris, 3 is the label of the pupil, and 4 is the label of the lacrimal caruncle. The final output is calculated by finding the maximum value among the 5 values of each pixel, and the index of these maximum values is assigned to a matrix of 1×N×N, which is the final output mask.

[0048] In step S3, the pupil center is extracted from the segmented pupil, and a vertical pupil center line is obtained.

[0049] In step S3 above, the pupil center is the average value of the X and Y coordinates of all pupil pixels extracted from the first eye position image, and a vertical line can be drawn using the average value of the X coordinates to obtain the pupil center line.

[0050] In step S4, the distance between the intersections of the sclera, iris or pupil and the background on the pupil center line is obtained, and the palpebral fissure width is calculated using this distance.

[0051] As shown in FIG. 5, in step S4, the palpebral fissure width is calculated based on the principle of similar triangles, and the calculation formula is Palpebral fissure width (B) = Pixel distance of palpebral fissure width (A) × Length of a single pixel in the direction of palpebral fissure width × Distance from the palpebral fissure to the camera lens (D) ÷ Distance from the camera photosensitive sensor to the camera lens (C), where the pixel distance of the palpebral fissure width (A) is the distance between the intersections of the sclera, iris or pupil and the background on the pupil center line.

[0052] Furthermore, in step S4, the distance D from the palpebral fissure to the camera lens = Distance 1 from the canthal fixation point to the camera lens - Exophthalmos 2.

[0053] Furthermore, the exophthalmos degree (2) uses the average value of the exophthalmos degree of healthy individuals. From the perspective of statistical average theory, the average value of the exophthalmos degree of healthy individuals may be 12 - 14 mm. In actual operation, the distance 1 from the canthal fixation point to the camera lens is known, the distance C from the camera photosensitive sensor to the camera lens is also known, the pixel distance A of the palpebral fissure width may be obtained from the captured first eye position image, and since the length of a single pixel in the palpebral fissure width direction is also a known fixed value, the value of the palpebral fissure width can be calculated.

[0054] In the training process of the neural network, in order to verify whether the result finally output by the training of the neural network is accurate, it is necessary to further introduce the definition of the loss function. The loss function is a function that calculates the distance between the current output and the predicted output in the algorithm and provides a digital indicator regarding the model representation. The loss function is an important component of the neural network. In this application, a composite loss function composed of focal loss, generalized dice loss, surface loss, and boundary recognition loss is designed.

[0055] First, it is necessary to explain the meanings of the parameters TP, FP, TN, and FN in the loss function calculation.

[0056] TP, that is, True Positive, indicates that the prediction result is a positive example and the actual result is also a positive example, that is, the positive example is accurately predicted. FP, that is, False Positive, indicates that the prediction result is a positive example and the actual result is a negative example, that is, the positive example is wrongly predicted and the negative example is not predicted. TN, that is, True Negative, indicates that the prediction result is a negative example and the actual result is also a negative example, that is, the negative example is accurately predicted. FN, that is, False Negative, indicates that the prediction result is a negative example and the actual result is a positive example, that is, the negative example is wrongly predicted and the positive example is not predicted.

[0057] The Generalized Dice Loss (GDL) is derived from the F1 score. Specifically, as a supplement to general accuracy, refer to Equation 2. General accuracy can be a good indicator of model performance, but it does not meet the requirements in the case of unbalanced data. The loss function commonly used in semantic segmentation schemes is the IoU (Intersection over Union) loss, that is, the intersection of unions, which is expressed as the ratio of the common set and the union set of the "predicted bounding box" and the "ground truth". IoU is defined by Equation 1. The Dice Factor defined by Equation 3 can be obtained from Equations 1 and 2, and the Dice Loss can be obtained from the Dice Factor in Equation 4. The Generalized Dice Loss is only one extension of the Dice Loss. The simple Dice Loss is only applicable to two segmentation classes, so the Generalized Dice Loss processes multiple segmentation classes. Here, it processes the five segmentation classes mentioned in the label part. The related equations are as follows.

[0058] [Number] [Number] Here, precision refers to the probability in the classification model index, Recall refers to the recall rate, Intersection refers to the common set, and Union refers to the union set.

[0059] The Focal Loss is a variant of the standard cross-entropy loss in order to focus on strict negative examples. The focal loss function is derived from the following. p T is defined in Equation 5, where p represents the ground truth. By the definition of the composite p T , the focal loss can be described by Equation 6, where a TBoth and γ are hyperparameters. When γ is zero, the focal loss shrinks to the cross - entropy loss. γ is 2 in the example, and γ provides an adjustment effect. When the ground truth is very small, the function obtains losses from easy examples. When the model learns easy examples, the function transitions to more difficult examples. The related formula is as follows.

[0060]

Number

Number

[0061] The above y specifies the ground truth class label of the training set for classification in supervised learning, and p is the estimated probability of the label y = 1.

[0062] The semantic boundary in the Boundary Aware Loss (BAL) is a separation region based on category labels. The loss is weighted based on the distance between each pixel and the two closest fragments, introducing edge awareness. In this application, the Canny edge detector in OpenCV is used to generate boundary pixels, and then two more pixels are expanded to reduce confusion at the boundary. After the numerical values of these boundary pixels are enlarged by a certain multiple, it is added to the conventional standard cross - entropy loss to enhance the model's attention to the boundary.

[0063] Surface Loss (SL) is a distance metric based on the image contour space that preserves small and rare structures with high semantic value. BAL attempts to maximize the accurate pixel probability near the boundary, and GDL provides a stable gradient due to the unbalanced conditions. In contrast, SL measures the loss of each pixel based on the distance between each pixel and the ground truth boundary of each category, and can effectively restore small regions ignored by the region-based loss. Surface loss calculates the distance from a single pixel to the boundary of each label group and normalizes this distance according to the size of the image. These calculation results are combined with the results of the model prediction to obtain an average value.

[0064] The final composite loss function is given by Equation 7 below.

[0065] Composite Loss = a1 × Surface Loss + a2 × Focal Loss + a3 × Generalized Dice Loss + a4 × Boundary Recognition Loss × Focal Loss (7) Here, variables a1, a2, a3, and a4 are hyperparameters. In the specific case of this application, a1 is related to the number of elapsed time periods during training, a2 is 1, a3 is 1 - a1, and a4 is 20.

[0066] In one aspect, as shown in FIG. 6, the eyelid fissure width measuring device according to this application includes an image acquisition module 601, an image segmentation module 602, a feature extraction module 603, and a calculation module 604.

[0067] The image acquisition module 601 captures and acquires a first eye position image from the front view of the user at the front position of the eyeball in the near-infrared light field of 700 - 1200 nm. The image segmentation module 602 segments the background, iris, sclera, and pupil from the first eye position image using the training method of a neural network. The feature extraction module 603 extracts the pupil center from the segmented pupil and obtains the vertical pupil center line. The calculation module 604 calculates the distance between the sclera, iris, or the intersection of the pupil and the background on the pupil center line, and calculates the palpebral fissure width using the distance.

[0068] In one aspect, the computer-readable storage medium according to the present application stores at least one program code that is loaded and executed by a processor to implement the method for measuring the palpebral fissure width in the above embodiments.

[0069] In an exemplary embodiment, there is further provided a computer-readable storage medium including a memory storing at least one program code that is loaded and executed by a processor to implement the method for measuring the palpebral fissure width in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0070] Those skilled in the art will understand that all or part of the steps in the above embodiments may be implemented by hardware, or may be implemented by hardware related to at least one program code, the program may be stored in a computer-readable storage medium, and the above-described storage medium may be a read-only memory, a magnetic disk, an optical disk, or the like.

[0071] The above description is only a preferred embodiment of the present application and does not limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included within the protection scope of the present application.

Claims

1. A computer program comprising: By causing a computer to execute the computer program, capturing a first eye position image from a front view of the user at a position in front of the eyeball in a near-infrared light field of 700 to 1200 nm; Segmenting the background, the iris, the sclera, and the pupil from the first eye position image using a neural network training method; extracting a pupil center from the divided pupil to obtain a pupil center line in the vertical direction; determining a distance between an intersection point of the sclera, the iris, or the pupil on the pupil center line and the background, and calculating a palpebral fissure width using the distance; A computer program characterized by:

2. the pupil center is an average value of X and Y coordinates of all pupil pixels extracted from the first eye position image, Draw a vertical line using the average value of the X coordinate to obtain the pupil centerline.

2. The computer program of claim 1.

3. Segmenting the background, the iris, the sclera, and the pupil from the first eye position image using a neural network training method includes: The method includes the steps of using a combination of a UNet neural network model and a DenseNet neural network model as a neural network model, in which the first eye position image is input to the neural network model, performing dimensional reduction and dimensional expansion on the input using the UNet neural network model, transmitting an output of the dimensional reduction block to a corresponding dimensional expansion block by a skip connection, performing feature extraction in each dimensional reduction block using the DenseNet neural network model, and performing upsampling in each dimensional expansion module using the DenseNet neural network model; 2. The computer program of claim 1.

4. further comprising the step of verifying the training effect of the neural network using the loss function; The loss function is a composite loss, which consists of a focal loss, a generalized Dice loss, a surface loss, and a boundary recognition loss.

4. A computer program according to claim 3.

5. The loss function is a composite loss, The calculation method is: Composite loss = a 1 × surface loss +a 2 ×Focal loss +a 3 × Generalized Dice Loss + a 4 × boundary recognition loss × focus loss, Here, the a 1 , a 2 , a 3 , a 4 is a hyperparameter, a 1 is related to the number of time lapses during the training process, a 2 = 1, a 3 = 1 - a 1 and a 4 = 20, 5. A computer program according to claim 4.

6. The palpebral fissure width is Palpebral fissure width (B) = pixel distance of palpebral fissure width (A) × length of a single pixel in the palpebral fissure width direction × distance from palpebral fissure to camera lens (D) ÷ distance from camera photosensitive sensor to camera lens (C), Here, the pixel distance (A) of the palpebral fissure width is the distance between the intersection points of the sclera, iris, or pupil on the pupil center line and the background.

6. A computer program according to claim 1, wherein the computer program is a program for executing ... a computer.

7. The distance from the palpebral fissure to the camera lens is calculated as follows: Distance from the palpebral fissure to the camera lens (D) = Distance from the canthus anchor point to the camera lens (1) - Exophthalmos (2).

7. A computer program according to claim 6.

8. The exophthalmos degree (2) uses the average value of exophthalmos degree of healthy subjects.

8. A computer program according to claim 7.

9. an image acquisition module that captures and acquires a first eye position image from a front view of the user at a position in front of the eyeball in a near-infrared light field of 700 to 1200 nm; an image segmentation module for segmenting a background, an iris, a sclera, and a pupil from the first eye position image using a neural network training method; a feature extraction module that extracts a pupil center from the divided pupil and obtains a pupil center line in the vertical direction; a calculation module that determines a distance between an intersection point of the sclera, the iris, or the pupil on the pupil center line and a background, and calculates a palpebral fissure width using the distance; A device for measuring palpebral fissure width.