An eye video dynamic fusion dual-branch complementary feature-oriented high eye pressure discrimination method
By using DB-CFFNet, a dual-branch convolutional neural network based on rPPG, and acquiring eye videos using a regular camera to extract eye pulse wave signals, the problem of high equipment cost and reliance on professional personnel in traditional intraocular pressure measurement methods is solved, realizing low-cost, non-contact intraocular pressure status monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional methods of intraocular pressure measurement are difficult to achieve continuous dynamic monitoring, require professional personnel to operate, have high equipment costs, are prone to causing discomfort and cross-infection risks, and are not suitable for routine, large-scale intraocular pressure screening.
A dual-branch convolutional neural network, DB-CFFNet, based on remote optical volumetric plethysmography (rPPG), is used to collect eye videos through a regular camera, extract pulse wave signals from the region of interest in the eye, and generate spatiotemporal feature maps in segments using a sliding window to achieve non-invasive monitoring of intraocular pressure.
It enables low-cost, non-contact continuous dynamic identification of intraocular pressure, improving the convenience and comfort of monitoring and reducing equipment costs.
Smart Images

Figure CN122116453A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of physiological signal detection technology and relates to a method for identifying high intraocular pressure based on dynamic fusion of dual-branch complementary features of eye video. Background Technology
[0002] Intraocular pressure (IOP) recognition holds significant research value and application prospects in the fields of smart health and smart education. Sustained increases in IOP can lead to changes in the structure of the optic disc, resulting in progressive degeneration of optic nerve fibers and ultimately irreversible visual impairment. Furthermore, IOP levels can indirectly reflect an individual's eye strain and visual fatigue, particularly among students, providing objective evidence for assessing eye behavior and conducting health education.
[0003] Traditional methods of intraocular pressure (IOP) measurement, such as applanation tonometers, rebound tonometers, and indentation tonometers, achieve accurate IOP measurement through physical contact devices. The Goldmann applanation tonometer (GAT) is widely considered the gold standard for IOP measurement due to its high accuracy. However, these methods struggle to provide continuous dynamic monitoring of IOP, and the procedures require specialized medical personnel. The devices must directly or indirectly contact the cornea, potentially causing discomfort and posing a risk of cross-infection. Furthermore, the high cost of these devices makes them unsuitable for routine, large-scale IOP screening scenarios. Summary of the Invention
[0004] To address the aforementioned problems in existing methods and considering the significant achievements in remote physiological parameter detection based on remote photoplethysmography (rPPG) in recent years, this invention provides a method for identifying high intraocular pressure (IOP) by dynamically fusing dual-branch complementary features from eye videos. A dual-branch convolutional neural network (DB-CFFNet) based on rPPG and using eye videos for IOP identification is established. During the detection process, eye videos are acquired using a regular camera, and the blood volume pulse (BVP) signal of the region of interest (ROI) is extracted. The signal is then segmented using a sliding window and a spatiotemporal feature map is generated and fed into the DB-CFFNet IOP identification model, achieving imperceptible monitoring of IOP status.
[0005] The technical solution adopted in this invention is as follows:
[0006] First, a close-up video dataset of the eye was constructed for non-contact intraocular pressure (IOP) assessment. The eye region was captured at close range using a regular mobile phone camera, and real IOP measurements were performed using an IOP meter to provide the actual IOP status. To extract the BVP signal from the Region of Interest (ROI), a trained eye ROI segmentation model was used to segment the iris-pupil region and the sclera region. The BVP signals of the segmented ROIs were extracted and then normalized and denoised. Next, a sliding window was used to segment the long BVP sequences of the extracted eye ROIs to augment the data. Finally, the obtained BVP signal segments were converted into spatiotemporal feature maps and classified using the DB-CFFNet network to achieve IOP status identification.
[0007] The high intraocular pressure (IOP) discrimination method proposed in this invention, which dynamically fuses dual-branch complementary features from eye videos, mainly comprises several modules: an eye structure segmentation and RGB signal extraction module based on a semantic segmentation framework; a BVP signal extraction module based on the chromaticity model CHROM; a data segmentation and spatiotemporal feature map generation module; and a high IOP discrimination module based on deep learning. The overall framework structure is as follows: Figure 1 As shown, the following is a description of each module:
[0008] Module 1: Eye Structure Segmentation and RGB Signal Extraction Module Based on Semantic Segmentation Framework. This module trains an eye ROI segmentation model on an eye segmentation dataset using an existing semantic segmentation framework. It identifies the ROI to which each pixel in an eye video frame belongs, namely the iris-pupil region and the sclera region. The entire iris and pupil region is considered as one ROI. Then, the RGB three-channel signals of both regions are extracted separately.
[0009] Module 2: BVP signal extraction module based on the CHROM colorimetric model. CHROM eliminates non-pulse components in skin reflection by linearly combining RGB channels, mainly removing motion artifacts and normalizing color differences, thereby extracting a one-dimensional time pulse signal related to blood volume changes from the RGB three-channel signal.
[0010] Module 3: Data Segmentation and Spatiotemporal Feature Map Generation Module. This module uses overlapping sliding windows to segment the original BVP signal, thereby increasing the data volume. To accommodate the input of the DB-CFFNet network, the BVP signal segments are converted into spatiotemporal feature maps. This method can retain one-dimensional temporal information of the BVP signal while adding two-dimensional spatial information.
[0011] Module 4: High intraocular pressure discrimination module based on deep learning. A dual-branch convolutional neural network is trained using a pre-defined training set to obtain the final high intraocular pressure discrimination model, DB-CFFNet.
[0012] The specific steps are as follows:
[0013] Step 1: Construct a close-up video dataset of the eye for non-contact high intraocular pressure detection.
[0014] Step two, train the segmentation model for eye ROI extraction, as follows:
[0015] (1) Construct a dataset of close-up eye images for training iris-pupil and scleral region segmentation masks;
[0016] (2) The eye ROI segmentation model is obtained by training on the dataset based on the existing semantic segmentation framework.
[0017] Step 3, BVP signal extraction and spatiotemporal feature map generation for the iris-pupil and scleral regions, is performed as follows:
[0018] (1) The trained eye segmentation model is used to segment the iris-pupil and sclera regions of the original eye video frame image;
[0019] (2) The segmentation mask is mapped to the original eye video frame, and the RGB three-channel signals of the two regions are extracted respectively;
[0020] (3) Use the CHROM algorithm to eliminate motion artifacts and separate the BVP signal from the original RGB signal;
[0021] (4) Use a sliding window to segment the separated BVP signal and convert it into a spatiotemporal feature map.
[0022] Step four: The spatiotemporal feature map is fed into the constructed dual-branch convolutional neural network for training to obtain the high intraocular pressure discrimination model DB-CFFNet.
[0023] The beneficial effects of this invention are: it proposes a high intraocular pressure (IOP) discrimination method based on dynamic fusion of dual-branch complementary features from eye videos. By modeling the physiological signal features of the iris-pupil and sclera regions using a dual-branch network, and combining rPPG and deep learning techniques, it achieves a low-cost, non-contact IOP monitoring method based on eye videos. This method can directly perform continuous dynamic IOP recognition using ordinary cameras, making it more convenient, comfortable, and cost-effective for IOP monitoring. Attached Figure Description
[0024] Figure 1 These are schematic diagrams of the various modules of the present invention;
[0025] Figure 2 Example video frames of the eye close-up video dataset constructed in this invention;
[0026] Figure 3Example diagrams of eye images and corresponding segmentation masks used in this invention for training the eye segmentation model.
[0027] Figure 4 This is a diagram showing the results of eye region segmentation by the eye segmentation model trained in this invention.
[0028] Figure 5 This invention describes the process of processing eye videos into spatiotemporal feature maps that are adapted to the input of a high intraocular pressure discrimination model.
[0029] Figure 6 This is a schematic diagram illustrating the process by which the DB-CFFNet model described in this invention performs high intraocular pressure discrimination on the input spatiotemporal feature map; Detailed Implementation
[0030] The present invention will now be described in further detail with reference to the accompanying drawings.
[0031] A method for identifying high intraocular pressure based on dynamic fusion of dual-branch complementary features from eye videos includes the following steps:
[0032] Step 1: Construct a dataset of close-up eye videos for training a high intraocular pressure discrimination model.
[0033] This invention constructs a close-up eye video dataset for non-contact diagnosis of high intraocular pressure. This dataset uses a regular mobile phone camera to capture images of the eye area, reflecting everyday usage scenarios and improving the data's general applicability. The device parameters are as follows: video resolution of 1920×1080, frame rate of 30 frames per second, ensuring image quality while effectively capturing detailed physiological signals from the eye. Eye videos were collected from 60 subjects. During filming, no glasses were allowed, and the distance between the eyes and the camera was approximately 5 cm. Simultaneously, eye videos were collected from both eyes of each subject, including 28 men and 32 women, aged 22 to 70 years. For intraocular pressure (IOP) measurement, three measurements were repeated using an ET-2 type eyelid contact electronic tonometer to obtain the true IOP values. The collected IOP values ranged from 6 to 65 mmHg. According to the international standard, 10 to 21 mmHg is considered the normal IOP range, and the true IOP status was labeled. Samples with IOP values within the 10–21 mmHg range were labeled as normal IOP, and samples with IOP values greater than 21 mmHg were labeled as high IOP. The task was modeled as a binary classification problem, with normal IOP corresponding to label 1 and high IOP corresponding to label 0. The acquisition environment was indoor normal lighting (100–550 lx) to minimize the impact of lighting conditions on the quality of the eye video. Example images of the acquired eye video frames and true labels are shown below. Figure 2 As shown.
[0034] Step 2: Construct an eye region segmentation model based on a semantic segmentation network.
[0035] (1) Construct a dataset of masked close-up eye images for training eye segmentation;
[0036] First, image data containing close-up shots of the eyes needs to be collected from publicly available image databases and medical image materials. The collected close-up eye images require corresponding mask images for eye region segmentation. These mask images are binary images, where white areas represent the target ROI (Region of Interest) and black areas represent the background. The mask images are created by segmenting the region of interest (ROI) of the eye as needed. For example, the corresponding mask image generated from the original close-up eye image is shown below. Figure 3 As shown, the target ROI regions are the iris-pupil region and the sclera region.
[0037] (2) Train a model for segmenting the ROI;
[0038] The experiment used collected close-up images of the eye as data, and corresponding iris-pupil and sclera mask images as labels. Eye segmentation models were trained using semantic segmentation frameworks such as YOLO, U2Net, and DeepLabV3+ to segment the iris-pupil region and the sclera region, respectively. The segmentation models were tested on a test set, and the selected evaluation metrics included precision. ), recall ) and F1 score (F1 score, ), as shown in (1) ~ (3).
[0039]
[0040]
[0041]
[0042] in This indicates the number of pixels that correctly segment the target ROI region. This indicates the number of pixels that were incorrectly segmented into the target ROI region. This indicates the number of pixels that are correctly segmented and belong to the background region. This indicates the number of pixels that were incorrectly segmented into the background region.
[0043] The training set contained 400 eye images, and the test set contained 200 eye images. During training, the cross-entropy loss function was used to calculate the error between the segmentation result and the ground truth mask image labels, and the model's weight parameters were updated using backpropagation. Specifically, the training parameters included 100 training epochs, a batch size of 16, a learning rate of 0.001, and the Adam optimizer. The model trained in the last epoch was used as the final model. Comparison of results on the test set showed that the YOLOv11 model performed best for iris-pupil region segmentation. The score reached 0.989. For scleral region segmentation, the U2Net model performed best. The value reached 0.946. Therefore, for different ROI segmentations, the experiment used the optimal segmentation model to divide the eye region.
[0044] Step 3: Extract the RGB three-channel signals from the ROI region in the eye video.
[0045] A trained eye segmentation model is used to perform semantic segmentation on each frame of the eye video, extracting binary masks corresponding to the iris-pupil region and the sclera region respectively, forming two independent ROI video frames, such as... Figure 4 As shown. For each video frame, the two ROI regions are processed as follows:
[0046] For all pixels within the iris-pupil (IP) mask, the average pixel intensity of each R, G, and B channel is calculated to generate the RGB representation signal of the region in the frame, as shown in (4) ~ (6).
[0047]
[0048]
[0049]
[0050] in , , They represent the first Average R, G, and B signal values of the iris-pupil region in a frame; This represents the total number of valid pixels within the iris-pupil region. , as well as For the first in this region The intensity value of each pixel in the corresponding channel.
[0051] Similarly, the same processing is performed on the sclera (SC) mask region to calculate the average intensity of its RGB channels and obtain the temporal color signal of the region, as shown in (7) ~ (9).
[0052]
[0053]
[0054]
[0055] in, , as well as They represent the first Average R, G, and B signal values in the scleral region of the frame; This represents the total number of effective pixels within the scleral region. , as well as For the first in this region The intensity value of each pixel in the corresponding channel.
[0056] The above calculation process is repeated frame by frame to finally obtain the RGB signal sequences of the iris-pupil region and the sclera region, as shown in (10) and (11).
[0057]
[0058]
[0059] in This represents the RGB signal sequence of the iris-pupil region. Represents the RGB signal sequence of the scleral region. This represents the total number of video frames. These two sets of signals will serve as inputs to the subsequent BVP signal extraction module, used to extract physiological pulse information that may reflect changes in intraocular pressure from different eye structures.
[0060] Step 4: Eliminate motion artifacts using the chromaticity-based CHROM method.
[0061] After obtaining two sets of RGB three-channel time-series signals from the iris-pupil region and the sclera region, a chromaticity-based CHROM algorithm was applied to each set of signals to suppress motion artifacts and extract physiological signals related to blood volume and pulse. Skin color changes caused by BVP have specific chromaticity characteristics in the RGB channels, while brightness changes caused by motion artifacts are approximately consistent across all channels.
[0062] First, the RGB signal of each ROI is normalized to reduce the impact of lighting changes in each frame. Calculate the normalized signal, as shown in (12) ~ (14).
[0063]
[0064]
[0065]
[0066] in , and These correspond to the normalized signals of channels R, G, and B, respectively. , and These correspond to the original signals of channels R, G, and B before normalization, respectively. , , These are the average values of the three RGB channels for that region.
[0067] Next, two orthogonal chromaticity signals are constructed using the normalized RGB three-channel signals. and As shown in (15) and (16).
[0068]
[0069]
[0070] The combination coefficients are designed based on a skin color reflection model to enhance pulse signal suppression of brightness variations. It is assumed that motion artifacts occur... and The two have similar energy distributions, but the BVP component is distributed differently in the two, so the preliminary BVP signal can be extracted by equation (17).
[0071]
[0072] in This represents the initially extracted BVP signal. Weighting coefficients are signals and The ratio of standard deviations , correspond standard deviation correspond The standard deviation. Further substituting... and The expression can be obtained The final expression is shown in (18).
[0073]
[0074] For the processed results The final BVP signal can be obtained further by bandpass filtering, as shown in (19).
[0075]
[0076] in The final extracted BVP signal, The bandpass filter function is defined, with a bandpass frequency of 0.75 ~ 4 Hz. The BVP signals for the RGB three-channel signals of the iris-pupil region and sclera region are processed according to (12) ~ (19) to extract the BVP signals for the iris-pupil region. and BVP signal in the scleral region .
[0077] Step 5: Convert the processed BVP signal into a spatiotemporal feature map.
[0078] To further construct samples suitable for deep learning model input and increase data scale, this step performs sliding window segmentation on the extracted long-sequence BVP signal and uses a structured mapping strategy to transform it into a two-dimensional feature map with spatiotemporal expressive capabilities. The specific processing flow is as follows:
[0079] First, the BVP signal is segmented. The BVP signal extracted by the CHROM algorithm (including the iris-pupil and sclera regions) is segmented using overlapping sliding windows. Let the length of the original BVP signal be... Using a fixed length The window is in steps Slide along the time axis to divide it into multiple subsequence segments, as shown in (20).
[0080]
[0081] in Indicates the first BVP subsequence segments, window length Set to 150 (corresponding to a 5s signal, 30fps frame rate), step size .
[0082] Subsequently, for each BVP signal segment Periodic expansion: The original signal is repeatedly spliced 6 times to obtain the expanded sequence. Then initialize one. Two-dimensional matrix ,in The extended sequence is mapped to a matrix using diagonal padding. As shown in (21).
[0083]
[0084] Finally, to adapt to the input of the convolutional neural network, the matrix is... Perform linear normalization to The interval is shown in (22).
[0085]
[0086] This represents the normalized matrix. and Represent matrices respectively The minimum and maximum values in the matrix. The normalized matrix. The image is saved as a grayscale image and used as the input feature map for the DB-CFFNet network. This transformation process preserves the temporal dynamic characteristics of the BVP signal while introducing potential spatial patterns through a two-dimensional structured representation, facilitating the extraction of discriminative features by the convolutional neural network. The generated spatiotemporal feature map is shown below. Figure 5 As shown.
[0087] Step 6: Train the high intraocular pressure discrimination model DB-CFFNet by fusing complementary features from two branches using spatiotemporal feature maps.
[0088] To construct a high intraocular pressure discrimination model, the proposed dual-branch convolutional neural network DB-CFFNet was trained using a pre-generated spatiotemporal feature map dataset. Its network structure is as follows: Figure 6 As shown, it consists of two parallel feature extraction branches and a feature fusion module.
[0089] The feature extraction module is built on a DenseNet161 densely connected network, consisting of an initial convolutional layer, four dense blocks, and corresponding transition layers. The initial convolutional layer uses a 7×7 kernel with a stride of 2 and padding of 3, resulting in 96 output channels. It is followed by a 3×3 max-pooling layer with a stride of 2 for feature map downsampling. Each dense block consists of multiple convolutional units, with the four blocks containing 6, 12, 36, and 24 units respectively. Each convolutional unit includes a batch normalization layer, a ReLU activation function, and a 3×3 convolutional layer with a stride of 1 and padding of 1. The output features from all preceding layers are concatenated through dense connections to enhance feature reuse. The network growth rate is set to k=48. Transition layers are placed between the dense blocks, consisting of 1×1 convolutional layers and average pooling layers. The network employs 1×1 convolutions to reduce channel dimensionality, and 2×2 windows with a stride of 2 are used in the average pooling layers to reduce the spatial resolution of the feature maps. In branch 1, a channel attention module is introduced after the first dense block (Dense Block 1); in branch 2, it is introduced after the fourth dense block (Dense Block 4). The channel attention module uses a Squeeze-and-Excitation structure. First, global average pooling is used to compress the feature maps into channels. Then, two fully connected layers form a bottleneck structure: the first layer reduces dimensionality, and the second layer increases it. The compression ratio r is set to 8, and the activation functions are ReLU and Sigmoid, respectively, to achieve adaptive learning of channel weights and feature recalibration. At the end of the network, global average pooling is performed on the extracted feature maps, and fully connected layers map the features to the class space. Finally, the Softmax function outputs the corresponding class probabilities.
[0090] The fusion module is located after the classification outputs of the two branches. It is used to perform weighted fusion of the class probabilities of the outputs of the two branches. The fusion weights are determined by the performance of the validation set to achieve adaptive combination of the discrimination results of different branches.
[0091] The two-dimensional spatiotemporal feature maps of the input are set to come from the iris-pupil region (branch 1) and the sclera region (branch 2), as shown in (23).
[0092]
[0093] The input represents the spatiotemporal feature map of the iris-pupil region. The input represents the spatiotemporal feature map of the scleral region, where , and These represent the height, width, and number of channels of the input feature map, respectively.
[0094] For branch 1, input feature map First, intermediate feature representations are obtained through a feature extraction network. Then, a channel attention mechanism is introduced in the shallow layer to enhance key features, ultimately yielding high-level semantic features. The process can be represented by equations (24) to (26).
[0095]
[0096]
[0097]
[0098] in This represents the feature extraction operation for the first dense block. This represents the channel attention weighting process. For the weighted features, This indicates the subsequent feature extraction process.
[0099] For branch 2, the input feature map First, intermediate feature representations are obtained through preliminary feature extraction. Then, a channel attention mechanism is introduced at a deeper level to enhance the features, ultimately yielding high-level semantic features. The process can be represented by equations (27) to (29).
[0100]
[0101]
[0102]
[0103] in, This indicates the initial feature extraction process. This represents the feature extraction operation for the fourth dense block. This indicates the subsequent feature extraction process.
[0104] In the channel attention module, the input features are first compressed into channels using global average pooling to obtain channel descriptors. As shown in equation (30).
[0105]
[0106] in, Indicates the first The channel descriptors obtained after global average pooling of each channel. Indicates the spatial location of the input feature map First Feature values on each channel; and These represent the height and width of the input feature map, respectively. This represents the total number of channels in the feature map. For channel indexing.
[0107] Subsequently, the channel weights are learned through the bottleneck structure constructed by two fully connected layers, as shown in Equation (31).
[0108]
[0109] in, and These are the weight matrices for dimensionality reduction and dimensionality increase, respectively. and These represent the ReLU and Sigmoid activation functions, respectively. This is the channel weight vector. The channel weights are multiplied channel-by-channel by the original features to achieve feature recalibration, resulting in the recalibrated features. As shown in equation (32).
[0110]
[0111] in This indicates that the feature map after channel attention weighting is at the th... Output on each channel Indicates the input feature map at the th Characteristic responses on each channel Indicates the corresponding number Attention weight coefficients for each channel.
[0112] Obtaining high-level features of both branches and Then, global average pooling is performed and the class probability distribution is obtained through a fully connected layer, as shown in equations (33) to (34).
[0113]
[0114]
[0115] in, and The feature vector after pooling; , and , These represent the weights and biases of the two branch classification layers, respectively. and These represent the class prediction probabilities of the two branches, respectively.
[0116] The feature fusion module is located after the classification outputs of the two branches and is used to perform weighted fusion of the class probabilities of the outputs of the two branches. This fusion process does not introduce additional trainable parameters.
[0117] To achieve adaptive fusion of the results from the two branches, a weight parameter is introduced. The prediction results of the two branches are weighted and combined, and the optimal weight is determined by the performance of the validation set, as shown in equations (35) to (36).
[0118]
[0119]
[0120] in, , This represents the classification accuracy on the validation set. This means finding the option that maximizes the objective function. , For optimal fusion weights, This refers to the final predicted probability result under the optimal fusion weights. In actual implementation, this is achieved by traversing the validation set with a fixed step size of 0.05. Take values and calculate different The corresponding fusion classification accuracy is selected, and the one that maximizes the validation set accuracy is chosen. As the final fusion weights, they are used to calculate the final output. .
[0121] The dataset is divided into training, validation, and test sets in a 3:1:1 ratio, and the image sizes are uniformly adjusted to the size required for network input. During training, the Adam optimizer was used with a batch size of 16 and a training cycle of 40 epochs. A segmented learning rate strategy was employed during training: the learning rate for the first 20 epochs was set to a specific value. The learning rate decayed to [a certain value] in the last 20 rounds. This promotes stable convergence of the model and improves generalization performance. The cross-entropy loss function is selected, and a Dropout layer with a ratio of 0.1 is introduced before the fully connected layer to prevent overfitting.
[0122] The high intraocular pressure (IOP) discrimination model DB-CFFNet was obtained through the above training. DB-CFFNet achieved the highest classification accuracy on the validation set during the 40 training epochs, and the high IOP recognition accuracy on the test set reached 95.65%. In the testing phase, two BVP signals were extracted from the input eye video according to the aforementioned steps, and corresponding spatiotemporal feature maps were generated. These maps were then input into two branches of the model for feature extraction and fusion. Finally, the classification output layer provided the high IOP discrimination result, thus obtaining the IOP discrimination result for the test samples.
Claims
1. A method for identifying high intraocular pressure by dynamically fusing dual-branch complementary features from eye videos, characterized in that, Includes the following steps: Step 1: Construct and process the eye video dataset; Close-up videos of the subjects' eyes were captured using a regular mobile phone camera at a resolution of 1920×1080 and a frame rate of 30fps. The actual intraocular pressure (IOP) was measured simultaneously using an ET-2 type eyelid contact electronic tonometer. IOP was defined as 10-21 mmHg as normal and greater than 21 mmHg as high. The IOP status was modeled as a binary classification task, with normal IOP corresponding to category label 1 and high IOP corresponding to category label 0. All samples were labeled according to the above criteria to construct a labeled IOP video dataset. Step 2: Segment the eye region and extract the raw signal; (1) Training of the eye segmentation model: Construct a training dataset containing close-up images of the eyes and corresponding segmentation masks for the iris-pupil joint region and sclera region; using this dataset, train an eye segmentation model based on a semantic segmentation network framework to output the corresponding binary masks for the iris-pupil region and the sclera region for the input eye video frame images. (2) Segment the iris-pupil and sclera regions of the eye video frame and extract the RGB three-channel signals: Using a trained eye segmentation model, each frame of the acquired eye video is segmented to obtain binary masks for the iris-pupil joint region (IP) and the sclera region (SC). Then, for each video frame, the average R, G, and B channel intensities of all pixels within the iris-pupil region mask are calculated to form the RGB temporal signal of that region. Simultaneously, the same operation is performed on the scleral region to generate an RGB timing signal for the scleral region. This results in two independent RGB three-channel sequences; Step 3: Extract pulse wave signals based on the chromaticity model CHROM; The CHROM algorithm is applied to the RGB signal sequences of the two branches obtained in step two to extract the BVP signal. The specific process includes: (1) Normalize the RGB signal as shown in (1) ~ (3); ; ; ; in , and These correspond to the normalized channel signals of the three RGB channels, respectively. , and These correspond to the three RGB channels of the signal before normalization; , , These are the average values of the three RGB channels in this region; (2) Construct orthogonal chromaticity signals. Based on the normalized signals, construct two orthogonal chromaticity signals to enhance the pulse components and suppress brightness variations. and As shown in (4) and (5) respectively; ; ; (3) Calculate the preliminary signal by combining the chroma signals and filtering out non-physiological frequency noise. The weighting coefficient For signal and The ratio of the standard deviations, and Corresponding signals and standard deviation, substitute and The expression is obtained The final expression is shown in (6); ; Next, the initially obtained BVP signal is filtered, with the filtering parameters set to 0.7 ~ 4.0 Hz, to further extract a purer BVP signal, as shown in (7); ; in This represents the bandpass filter function; following the steps of formulas (1) to (7), the two BVP signals for the iris-pupil region and the sclera region are obtained respectively. and ; Step 4: Convert the BVP signal into a spatiotemporal feature map; For the BVP signals of the two regions obtained in step three, perform the following transformations independently: (1) Using a length of Step size is The overlapping sliding window divides the long sequence BVP signal into multiple segments, as shown in (8); ; in Indicates the first A BVP subsequence fragment, For a 5-second video with a frame rate of 30fps, the step size is 150. ; (2) Periodically extend each BVP signal segment by repeatedly splicing each BVP signal segment of length L 6 times to obtain the extended signal. ; (3) The expanded signal is mapped to a [missing information] using a diagonal padding method. Two-dimensional matrix Among them The filling rules are as follows: ; (4) Transform the matrix Normalized to the [0, 255] interval and saved as a single-channel grayscale image, which is used as the spatiotemporal feature map input to the DB-CFFNet network; Step 5: Construct and train the dual-branch discriminant model DB-CFFNet; A dual-branch convolutional neural network, DB-CFFNet, is constructed to classify spatiotemporal feature maps. The model comprises two parallel branches. Each branch starts with a 7×7 convolutional kernel with a stride of 2, followed by a 3×3 max-pooling layer with a stride of 2. The network consists of four dense blocks and corresponding transition layers. Each dense block comprises multiple convolutional units, with the number of convolutional units in the four dense blocks being 6, 12, 36, and 24, respectively. Each convolutional unit contains a batch. The network consists of a normalization layer, a ReLU activation function, and a 3×3 convolutional layer with a stride of 1 and padding of 1. The output features of all preceding layers are concatenated using dense connections to enhance feature reusability. Transition layers, composed of 1×1 convolutional layers and average pooling layers, are placed between the dense blocks. The 1×1 convolutions reduce channel dimensions, while the average pooling uses a 2×2 window with a stride of 2 to reduce the spatial resolution of the feature maps. Furthermore, a channel attention module is introduced after the first dense block in the first branch, and after the fourth dense block in the second branch. This channel attention module uses a Squeeze-and-Excitation structure with a compression ratio of 8. A global average pooling layer is placed at the end of the network, and the classification result is output through a fully connected layer. The two branches output class probabilities respectively, and these probabilities are fused at the decision layer using a weighted fusion method. The dataset was divided into training, validation, and test sets. The training set was used for supervised training of the model. During training, the cross-entropy loss function was used to calculate the error between the predicted values and the true labels, and the Adam optimizer was used to update the model parameters. The batch size was set to 16. The training consisted of 40 epochs, and a segmented descent strategy was used for the learning rate, with the learning rate set to 1×10⁻⁶ for the first 20 epochs. -4 The learning rate for the last 20 rounds was set to 1×10. -5 A Dropout layer with a ratio of 0.1 is introduced before the fully connected layer to prevent overfitting. During training, the model obtained in each round of training is evaluated on the validation set, the corresponding classification accuracy is calculated, and the model parameters with the highest classification accuracy on the validation set are selected as the final model weights; simultaneously, the weight parameters in the decision layer fusion are... The search is performed within the interval [0,1] with a fixed step size of 0.
05. The goal is to optimize the classification accuracy of the fused data. The weights that maximize the accuracy of the validation set are selected as the optimal fusion weights. ; Step six: Perform high intraocular pressure detection; For a new eye video to be judged, steps two through four are executed sequentially to generate spatiotemporal feature maps of the corresponding two regions. These maps are then input into the high intraocular pressure discrimination model DB-CFFNet trained in step five. Through the model's dual-branch feature extraction, fusion, and classification, the discrimination result of high intraocular pressure or normal intraocular pressure is output. The discrimination results of all spatiotemporal feature maps generated for a video are statistically analyzed, and the category that appears most frequently is taken as the final high intraocular pressure discrimination result for that eye video segment.