Method and device for recognizing pathological features based on binocular fundus medical images
By aligning binocular fundus images and calculating pseudo-depth information, a pathological feature recognition model with shared weighted dual-path branches is constructed. This solves the problems of reliance on experience and data adaptability in traditional ophthalmic diagnosis, improves the accuracy and stability of fundus disease recognition, and is suitable for scenarios with insufficient primary healthcare resources.
Patent Information
- Application Number
- CN202511525921.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Traditional diagnosis of fundus diseases in ophthalmology relies on doctors' experience, is subject to individual differences, suffers from insufficient primary healthcare resources, and current technologies are unable to adapt to different data quality, fail to fully utilize binocular fundus image information, and lack the accuracy of multi-category disease classification, thus failing to meet clinical needs.
Fundus images are acquired using binocular cameras and the centers of the left and right optic discs are aligned. A global disparity map is calculated to obtain pseudo-depth information. A pathological feature recognition model is constructed, and a shared-weight dual-path branch is used for feature extraction and attention calculation to improve the model's stability and accuracy.
It reduces diagnostic errors, enhances the model's ability to perceive three-dimensional lesions in the fundus, improves the recognition accuracy and stability of multiple types of fundus diseases, adapts to different data quality, and meets the diagnostic needs of primary healthcare.
Smart Images

Figure CN120997895B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image analysis, in particular to a pathological feature recognition method and device based on binocular fundus medical images. BACKGROUND
[0002] As the core basis for the diagnosis of ophthalmic diseases, fundus images can directly present the pathological features of the retina, optic nerve and blood vessel system, and are the key tool for early screening of high-blinding diseases such as diabetic retinopathy, glaucoma, cataract and age-related macular degeneration. Its basic role in ophthalmic clinical diagnosis has been widely recognized.
[0003] However, the traditional ophthalmic fundus disease diagnosis method highly depends on the personal clinical experience of doctors, and the diagnosis result is easily affected by subjective factors such as the professional level and energy state of doctors, and there is significant individual difference. Especially in the primary areas where medical resources are scarce, the number of professional ophthalmologists is scarce, and it is difficult to meet the large-scale fundus disease screening demand, resulting in frequent misdiagnosis and missed diagnosis, and a large number of patients miss the best opportunity for early intervention and treatment, further aggravating the risk of blindness of ophthalmic diseases.
[0004] In recent years, with the development of deep learning technology, computer-aided diagnosis technology based on ResNet, EfficientNet and other models has made certain progress in the field of monocular fundus image classification, providing a new idea for fundus disease diagnosis. However, the existing technology still faces three major challenges and is difficult to meet the needs of clinical practical application: first, the model performance is highly sensitive to data quality, and the imaging parameter difference of different medical institutions and different types of collection equipment easily leads to uneven brightness and obvious color difference of fundus images, seriously affecting the generalization ability of the model, and making the model accuracy greatly decrease when applied across institutions; second, the existing technology focuses on monocular fundus image analysis, and fails to fully utilize the spatial-channel multi-dimensional correlation information contained in binocular fundus images, and cannot simulate the binocular collaborative judgment mechanism of the human visual system, missing the key path of improving lesion recognition accuracy through binocular comparison; third, the multi-class disease classification accuracy is insufficient, and in clinical practice, 7 or more fundus diseases (including the normal category) need to be accurately identified, but the existing model is prone to class confusion when dealing with complex scenes with uneven class distribution and similar lesion characteristics, and it is difficult to meet the strict accuracy requirements of clinical diagnosis.
[0005] In summary, it is urgent to develop an intelligent fundus disease diagnosis scheme that can adapt to different data quality, fully utilize binocular information and improve the recognition accuracy of multi-class pathological features, in order to make up for the shortcomings of traditional diagnosis methods and relieve the pressure of primary medical resources, and provide strong technical support for early screening and accurate diagnosis of fundus diseases. SUMMARY
[0006] The embodiment of the application provides a pathological feature recognition method and device based on binocular fundus medical images, pseudo-depth information is obtained by aligning images to calculate a global disparity map and fusion, a pathological feature recognition model is constructed, and the stability and effectiveness of feature extraction are improved by using a shared weight double-path branch, so that the diagnostic error is reduced and the perception ability of the model to fundus stereoscopic lesions is improved.
[0007] In a first aspect, the embodiment of the application provides a pathological feature recognition method based on binocular fundus medical images, which comprises:
[0008] The binocular camera is used to obtain the to-be-tested fundus medical images including the first left eye image and the first right eye image, and the first left eye image and the first right eye image are aligned to obtain the second left eye image and the second right eye image.
[0009] The global disparity map is calculated based on the second left eye image and the second right eye image, the pseudo-depth information is obtained based on the global disparity map, the pseudo-depth information is fused with the second left eye image to obtain the third left eye image, and the pseudo-depth information is fused with the second right eye image to obtain the third right eye image.
[0010] The third left eye image and the third right eye image are input into the pathological feature recognition model trained by using binocular images of known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0011] In a second aspect, the embodiment of the application provides an intelligent diagnosis device for ophthalmic diseases based on fundus medical images, which comprises:
[0012] The alignment module is configured to obtain the to-be-tested fundus medical images including the first left eye image and the first right eye image by using the binocular camera, and align the first left eye image and the first right eye image to obtain the second left eye image and the second right eye image.
[0013] The depth calculation module is configured to calculate the global disparity map based on the second left eye image and the second right eye image, obtain the pseudo-depth information based on the global disparity map, fuse the pseudo-depth information with the second left eye image to obtain the third left eye image, and fuse the pseudo-depth information with the second right eye image to obtain the third right eye image.
[0014] The recognition module is configured to input the third left-eye image and the third right-eye image into a pathological feature recognition model trained by using binocular images with known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module, and a classification module, the third left-eye image and the third right-eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left-eye feature map and a right-eye feature map, the left-eye feature map is input into the attention calculation module to perform attention calculation to obtain a left-eye attention result, the right-eye feature map is input into the attention calculation module to perform attention calculation to obtain a right-eye attention result, the left-eye attention result and the right-eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0015] In a third aspect, an electronic device is provided, which includes a memory and a processor, the memory has stored therein a computer program, and the processor is configured to run the computer program to perform a pathological feature recognition method based on binocular fundus medical images.
[0016] In a fourth aspect, a readable storage medium is provided, which has stored therein a computer program, the computer program includes program codes for controlling a process to perform a process, and the process includes a pathological feature recognition method based on binocular fundus medical images.
[0017] The main contributions and innovations of the present application are as follows:
[0018] The present application obtains fundus images by using a binocular camera and aligns the left and right optic disc centers, thereby eliminating the initial position deviation of the binocular images, providing accurate image basis for subsequent global disparity map calculation and pseudo-depth information acquisition, and reducing the diagnostic errors caused by image misplacement; the present application calculates a global disparity map based on the aligned binocular images to obtain pseudo-depth information, and fuses the pseudo-depth information with the binocular images, thereby supplementing the stereo structure information of the fundus images, increasing the information dimension of the model, improving the perception ability of the model to the stereo lesions of the fundus, and enabling the model to more comprehensively capture the lesion features; the present application constructs a pathological feature recognition model comprising a feature extraction module, an attention calculation module, and a classification module, the feature extraction module adopts a double-path branch with shared weights to efficiently extract respective features of the binocular images, ensure the consistency of the binocular feature extraction standards, reduce the feature difference interference, and improve the stability and effectiveness of the feature extraction.
[0019] Details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects, and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:
[0021] Figure 1 is a flow chart of a pathological feature recognition method based on binocular fundus medical images according to an embodiment of the application;
[0022] Figure 2 is a schematic diagram of the overall structure of a pathological feature recognition model according to an embodiment of the application;
[0023] Figure 3 is a schematic diagram of the structure of a second feature extraction branch and a third feature extraction branch according to an embodiment of the application;
[0024] Figure 4 is a schematic diagram of the structure of an attention module according to an embodiment of the application;
[0025] Figure 5 is a schematic diagram of the structure of a classification module according to an embodiment of the application;
[0026] Figure 6 is a structural block diagram of a pathological feature recognition device based on binocular fundus medical images according to an embodiment of the application;
[0027] Figure 7 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION
[0028] The exemplary embodiments will be described in detail herein below with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of the description. Instead, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of the description as detailed in the appended claims.
[0029] It should be noted that the steps of the methods in other embodiments are not necessarily performed in the order shown and described in this description. In some other embodiments, the steps of the methods can include more or fewer steps than those described in this description. Furthermore, a single step described in this description can be split into multiple steps in other embodiments; and multiple steps described in this description can be combined into a single step in other embodiments.
[0030] Embodiment One
[0031] The embodiment of the application provides a pathological feature recognition method based on binocular fundus medical images, pseudo-depth information is obtained by aligning images to calculate a global disparity map and fusion, a pathological feature recognition model is constructed, shared weight double-path branches are used to improve feature extraction stability and effectiveness, so that diagnosis error is reduced and model perception ability for fundus stereoscopic lesions is improved, and specifically, reference is made to Figure 1 , the method comprises:
[0032] Obtaining a to-be-tested fundus medical image comprising a first left eye image and a first right eye image through a binocular camera, and performing left-right optic disc center alignment on the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image;
[0033] Calculating a global disparity map based on the second left eye image and the second right eye image, and obtaining pseudo-depth information based on the global disparity map, fusing the pseudo-depth information with the second left eye image to obtain a third left eye image, and fusing the pseudo-depth information with the second right eye image to obtain a third right eye image;
[0034] Inputting the third left eye image and the third right eye image into a pathological feature recognition model trained by using binocular images of known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using double-path branches with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0035] In some specific embodiments, since the to-be-tested fundus medical image directly obtained by the binocular camera has problems of shadow, size difference and brightness and color difference imbalance, so that the recognition and classification accuracy of eye diseases is affected, therefore, the first left eye image and the first right eye image are preprocessed to eliminate image difference before the left-right optic disc center alignment is performed on the first left eye image and the first right eye image.
[0036] Further, the first left eye image and the first right eye image are preprocessed through graying, filtering denoising, contrast enhancement and the like. Specifically, the first left eye image and the first right eye image are subjected to contrast enhancement through a CLAHE algorithm, so as to increase the contrast of local blood vessels and make the local feature enhancement better. When the CLAHE algorithm is used to enhance the contrast of the first left eye image, the first left eye image is first divided into non-overlapping rectangular sub-regions, and a histogram of each rectangular sub-region is calculated. A contrast threshold is set, all pixel points with pixel values exceeding the contrast threshold in the histogram of each rectangular sub-region are obtained, and these pixel points are evenly distributed to all gray levels to obtain a clipping histogram corresponding to each rectangular sub-region. The cumulative distribution function of each clipping histogram is calculated as a corresponding mapping function, and each pixel point in the first left eye image is subjected to bilinear interpolation based on the mapping function, so as to obtain the first left eye image after contrast enhancement. When the CLAHE algorithm is used to enhance the contrast of the first right eye image, the first right eye image is first divided into non-overlapping rectangular sub-regions, and a histogram of each rectangular sub-region is calculated. A contrast threshold is set, all pixel points with pixel values exceeding the contrast threshold in the histogram of each rectangular sub-region are obtained, and these pixel points are evenly distributed to all gray levels to obtain a clipping histogram corresponding to each rectangular sub-region. The cumulative distribution function of each clipping histogram is calculated as a corresponding mapping function, and each pixel point in the first right eye image is subjected to bilinear interpolation based on the mapping function, so as to obtain the first right eye image after contrast enhancement.
[0037] Exemplarily, the size of the first left eye image is MxN, the first left eye image is divided into KxL rectangular sub-regions, and the size of each rectangular sub-region is mxn, where m=M / K and n=N / L. Then, the histogram of each rectangular sub-region is calculated. In the present scheme, the contrast threshold is wherein, is a limiting coefficient, and is usually 1-4; =256, indicating 256 gray levels, and the clipping histogram is obtained by the formula:
[0038]
[0039] wherein, cxcess is all pixel points exceeding the contrast threshold, represents the histogram of the rectangular sub-region with index k, i is a pixel point therein, and C is the contrast threshold.
[0040] The formula for evenly distributing these pixel points to all gray levels is:
[0041]
[0042] wherein, cutting histogram, represents the histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, = 256 represents 256 gray levels.
[0043] Specifically, the calculation formula of the cumulative distribution function is represented as:
[0044]
[0045] wherein, is the cumulative distribution function of the kth cutting histogram, i and j are pixel indexes in the cutting histogram, is the total number of pixels in the corresponding rectangular sub-region, = 256 represents 256 gray levels.
[0046] Since the blocking processing can cause artifacts at the boundaries of the rectangular sub-regions, the bilinear interpolation is adopted to process each pixel point when the mapping of the pixel points is performed based on the mapping function. Specifically, for any pixel in the first left eye image, the four adjacent rectangular sub-regions of the pixel are determined as , and the distance weight of the pixel to the center points of the four adjacent rectangular sub-regions is calculated as , and the updating mode of the pixel value is:
[0047]
[0048] is the original pixel value of the (x, y) pixel point in the first left eye image, is the pixel value of the (x, y) pixel point in the first left eye image after the contrast enhancement, is the distance weight of the pixel to the center points of the four adjacent rectangular sub-regions, is the mapping function corresponding to the pixel point.
[0049] Specifically, the present scheme better distinguishes the details in the image by evenly distributing the pixel points with high contrast to all pixel levels.
[0050] Specifically, the contrast enhancement mode of the first right eye image is the same as that of the first left eye, which is not repeated here.
[0051] In some embodiments, the feature point matching of the first left eye image and the first right eye image obtains a plurality of pairs of matching feature points, two pairs of matching feature points are randomly selected, an initial transformation matrix between the matching feature points is calculated, and the number of inliers in the initial transformation matrix is recorded, so as to obtain an initial transformation matrix with the largest number of inliers as a target transformation matrix, and a final transformation matrix is obtained by least square estimation based on all inliers in the target transformation matrix, and the first left eye image and the first right eye image are aligned based on the left and right optic disc centers based on the final transformation matrix, wherein the inliers are the number of matching feature points meeting the corresponding transformation matrix.
[0052] Further, the re-projection error of each feature point in the first left eye image mapped to the corresponding feature point in the first right eye image via the initial transformation matrix is calculated, and the matching feature point with a re-projection error less than a preset re-projection threshold is taken as an inlier of the corresponding initial transformation matrix.
[0053] Specifically, the SIFT feature point matching algorithm is used to match the feature points of the first left eye image and the first right eye image, and the final transformation matrix acquisition method of the present scheme can better handle complex image deformation and perspective difference.
[0054] In some embodiments, the pixel coordinate deviation value of the same key point of the fundus structure in the second left eye image and the second right eye image is taken as the disparity value, the disparity values of all key points are counted to obtain a global disparity map, and the global disparity map is input into a pre-trained CNN model for feature extraction to obtain pseudo-depth information, wherein the pseudo-depth information is the distance from each pixel point in the second left eye image and the second right eye image to the binocular camera.
[0055] Specifically, a common fundus image generally contains three-channel data of RGB, but the fundus image contains some stereoscopic structures, so all details of the fundus image cannot be displayed by only three-channel images, so the present scheme provides more information dimension options for the model by calculating and adding pseudo-depth information, which helps to improve the perception ability of the model to stereoscopic structures.
[0056] In some specific embodiments, the pathological feature recognition model is a model pre-trained based on the ODIR-5K dataset, and in the training process, random rotation, flipping, elastic transformation and other techniques covering the spatial level are used, combined with a generative adversarial network to enhance the training data, and based on a conditional GAN (cGAN) architecture, synthetic microaneurysms and other subtle lesions are realized to supplement the small sample dataset.
[0057] In some specific embodiments, the third left eye image and the third right eye image are normalized to the range [0, 1] before being input into the pathological feature recognition model.
[0058] In some embodiments, the overall structure of the pathological feature recognition model is as shown in Figure 2 The feature extraction module is composed of a first feature extraction unit, a second feature extraction unit, and a third feature extraction unit. The first feature extraction unit includes a left eye feature extraction branch and a right eye feature extraction branch that are structurally identical and share weights. The left eye feature extraction branch extracts features from the third left eye image to obtain first left eye features. The right eye feature extraction branch extracts features from the third right eye image to obtain first right eye features. The second feature extraction unit uses a feature pyramid to extract features from the first left eye features and the first right eye features to obtain second left eye features and second right eye features. The third feature extraction unit uses a CNN structure to extract features from the second left eye features and the second right eye features to obtain third left eye features and third right eye features. The third left eye features are taken as a left eye feature map, and the third right eye features are taken as a right eye feature map.
[0059] Specifically, the left eye feature extraction branch and the right eye feature extraction branch in the present scheme use CrossViT structure, and the left eye feature extraction branch and the right eye feature extraction branch drive the extraction of single-eye features by sharing weights.
[0060] Specifically, the second feature extraction branch uses a feature pyramid to perform fine upsampling and dynamic weighted fusion on the first left eye features and the first right eye features, thereby reasonably allocating the contribution of each level to the lesion area, thereby significantly improving the detection ability of small lesions such as microaneurysms.
[0061] Specifically, the third feature extraction branch using a CNN structure after the second feature extraction branch performs further feature extraction, which can integrate the global modeling capability of visual Transformer and the local detail capturing advantage of CNN. The structure of the second feature extraction branch and the third feature extraction branch is as shown in Figure 3
[0062] In some embodiments, the schematic diagram of the attention module is as shown in Figure 4 As shown, the attention calculation module includes a joint attention unit and a binocular collaborative attention mechanism unit, the joint attention unit includes channel attention, spatial attention and size attention connected in sequence, and the left eye feature map / right eye feature map sequentially passes through the channel attention, the spatial attention and the size attention to obtain the left eye joint attention result / right eye joint attention result, the binocular collaborative attention mechanism unit includes a left eye attention mechanism branch and a right eye attention mechanism branch which are structurally identical and weight-shared, the binocular collaborative attention mechanism calculates the feature correlation between the left eye joint attention result and the right eye joint attention result through the left eye attention mechanism branch and the right eye attention mechanism branch to obtain a cost volume, which includes the key features of the left eye and the right eye and the collaborative relationship between the left eye key features and the right eye key features.
[0063] Specifically, the present scheme first strengthens the weight of the channel sensitive to the pathological features of ophthalmic diseases through channel attention calculation, and reduces the influence of redundant channels on model judgment, then fuses the maximum / average pooling double-path weight through the calculation of the spatial attention mechanism, and strengthens the features of key regions such as optic disc and macula, and the scale attention mechanism uses the hollow convolution to capture the multi-scale features such as the cup-disc ratio of glaucoma, and through the re-labeling and enhancement of features of different sizes, the recognition ability of the model to multi-scale objects is improved.
[0064] Further, the left eye attention mechanism branch includes a left eye backbone network layer, a left eye error calculation layer, a left eye adversarial network layer and a left eye convolution layer, the right eye attention mechanism branch includes a right eye backbone network layer, a right eye error calculation layer, a right eye adversarial network layer and a right eye convolution layer, the left eye joint attention result is feature extracted by the left eye backbone network layer to obtain the left eye joint attention feature, the right eye joint attention result is feature extracted by the right eye backbone network layer to obtain the right eye joint attention feature, the average absolute error of the left eye joint attention feature and the right eye joint attention feature is calculated in the left eye error calculation layer to obtain the left eye feature difference, the average absolute error of the right eye joint attention feature and the left eye joint attention feature is calculated in the right eye error calculation layer to obtain the right eye feature difference, the feature distribution of the left eye feature difference is optimized in the left eye adversarial network layer to obtain the left eye adversarial feature, the feature distribution of the right eye feature difference is optimized in the right eye adversarial network layer to obtain the right eye adversarial feature, the left eye adversarial feature is feature integrated in the left eye convolution layer to obtain the left eye collaborative attention result, the right eye adversarial feature is feature integrated in the right eye convolution layer to obtain the right eye collaborative attention result, and the cost volume is obtained by fusing the left eye collaborative attention result and the right eye collaborative attention result.
[0065] Specifically, the present scheme adopts Crossvit as the left and right eye backbone network layer, the error calculation layer calculates the difference between the features by mean absolute error to highlight important feature differences, the adversarial network is used to enhance the generalization ability of the model to better resist adversarial attacks and noise interference, and the cost volume generation can make the model more accurately understand and process visual information from two different perspectives, thereby improving the performance of binocular vision tasks.
[0066] In some embodiments, in order to verify the effectiveness of the attention module of the present scheme, 3000 cases of data are carefully selected from the ODIR-5K dataset, which are randomly divided into a training set (2100 cases) and a test set (900 cases) according to a 7:3 ratio. The first group is the original data group (Original) without any processing; the second group uses the joint attention unit and the binocular collaborative attention mechanism unit for training, and then the seven groups of data are input into the classification module constructed in the present study for training, and the final experimental results are shown in Table 1. From Table 1, it can be seen that the accuracy, precision and recall have been improved.
[0067] Table 1 Superiority proof of attention module
[0068]
[0069] In some embodiments, the structure of the classification module is shown in Figure 5 The classification module is composed of multiple independent classification heads, and each independent classification head corresponds to the pathological features of an ophthalmic disease. The cost volume is sequentially input into each independent classification head, and the classification results of each independent classification head are output through a fully connected layer to obtain the pathological feature recognition results.
[0070] Specifically, the present scheme uses multiple independent classification heads to classify the pathological features of different ophthalmic diseases, thereby achieving the effect of reducing feature coupling. In addition, the present scheme uses a fully connected layer and a Sigmoid activation structure to calculate the probability of each classification result and output.
[0071] Specifically, the present scheme dynamically adjusts the decision threshold of each class based on the ROC curve of the validation set, and ensures the coexistence of high sensitivity and specificity through the precision-recall balance strategy.
[0072] In some embodiments, two groups of controls are set: 1) Original group: keep the original data distribution, directly train with the conventional 8-classification head; 2) HDCH group: adopt 8 independent classification heads for parallel training, the existing mainstream method generally adopts the "single forward-8 class" paradigm, that is, one forward inference gives the probability distribution of 8 classes at the same time. In contrast, the hierarchical decoupling strategy proposed in this paper divides the overall task into 8 independent binary classification sub-tasks, each of which is independently completed by a dedicated classification head for training and inference, thereby completely blocking the semantic coupling between pathological features. Finally, the 8 groups of binary classification results are integrated according to the pre-set fusion rule, and the final experimental results are shown in Table 2. As can be seen from Table 2, the Decoupled group has significantly improved in accuracy, precision and recall: the overall accuracy has increased by 7.02% compared with the Original group, and the average precision and recall of each pathological feature have increased by 4.21% and 5.01% respectively, verifying the effectiveness of the hierarchical decoupling strategy in reducing class interference and improving fine-grained diagnostic performance:
[0073] Table 2 Superiority of classification head module
[0074]
[0075] In some embodiments, in order to verify the superiority of the overall pathological feature recognition model in the present scheme, 3000 cases of data are also selected from the ODIR-5K dataset. The optimal method of experiments 1 and 2 is adopted, and the data is input into different models for comparison. After the model training is completed, the two groups of trained model parameters are extracted and applied to the test set in the test link. The labels output by the model training are paired with the true labels of the test set. Finally, by calculating the three key indicators of Accuracy (accuracy), precision and recall, the performance of the model under different processing methods is comprehensively evaluated. From the data in Table 3, it can be seen that the model compared with other models has improved in the three evaluation indicators, proving the superiority of the model.
[0076] Table 3 Superiority of overall pathological feature recognition model
[0077]
[0078] Embodiment Two
[0079] Based on the same idea, referring to Figure 6 The application also proposes a pathological feature recognition device based on binocular fundus medical images, comprising:
[0080] An alignment module is configured to acquire, by the binocular camera, the to-be-tested fundus medical image including the first left eye image and the first right eye image, and perform left-right optic disc center alignment on the first left eye image and the first right eye image to obtain the second left eye image and the second right eye image.
[0081] A depth calculation module is configured to calculate a global disparity map based on the second left eye image and the second right eye image, acquire pseudo-depth information based on the global disparity map, fuse the pseudo-depth information with the second left eye image to obtain a third left eye image, and fuse the pseudo-depth information with the second right eye image to obtain a third right eye image.
[0082] An identification module is configured to input the third left eye image and the third right eye image into a pathological feature identification model trained by using binocular images with known pathological features, wherein the pathological feature identification model includes a feature extraction module, an attention calculation module, and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature identification result.
[0083] Embodiment three
[0084] This embodiment also provides an electronic device, referring to Figure 7 including a memory 404 and a processor 402, the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0085] Specifically, the above processor 402 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0086] The memory 404 can include a mass storage that stores data or instructions. For example, and without limitation, the memory 404 can include a Hard Disk Drive (HDD), a floppy disk drive, a Solid State Drive (SSD), a flash drive, a Compact Disc Read Only Memory (CD-ROM), a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. The memory 404 can be removable and / or non-removable (or fixed) as appropriate. The memory 404 can be internal or external as appropriate. In particular embodiments, the memory 404 is a Non-Volatile memory. In particular embodiments, the memory 404 includes a Read-Only Memory (ROM) and a Random Access Memory (RAM). The ROM can be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), an Electrically Alterable ROM (EAROM), or a FLASH memory, or a combination of two or more of these, as appropriate. The RAM can be a Static Random-Access Memory (SRAM) or a Dynamic Random Access Memory (DRAM), which can be a Fast Page Mode Dynamic Random Access Memory (FPMDRAM), an Extended Data Output Dynamic Random Access Memory (EDODRAM), a Synchronous Dynamic Random-Access Memory (SDRAM), or the like, as appropriate.
[0087] The memory 404 can be used to store or buffer various data files needed for processing and / or communication, and possible computer program instructions executed by the processor 402.
[0088] The processor 402 can implement any of the above-mentioned pathological feature recognition methods based on binocular fundus medical images by reading and executing the computer program instructions stored in the memory 404.
[0089] Optionally, the above-mentioned electronic device can further include a transmission device 406 connected with the processor 402 and an input / output device 408 connected with the processor 402.
[0090] The transmission device 406 can be used to receive or send data via a network. The above-mentioned network can include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (NIC) which can be connected with other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module which is used to communicate with the Internet in a wireless manner.
[0091] The input / output device 408 is used to input or output information. In the present embodiment, the input information can be a fundus medical image to be tested, and the output information can be a pathological feature recognition result.
[0092] Optionally, in the present embodiment, the processor 402 can be configured to execute the following steps by means of a computer program:
[0093] Obtaining a fundus medical image to be tested including a first left eye image and a first right eye image by means of a binocular camera, and performing left-right optic disc center alignment on the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image;
[0094] Calculating a global disparity map based on the second left eye image and the second right eye image, obtaining pseudo-depth information based on the global disparity map, fusing the pseudo-depth information with the second left eye image to obtain a third left eye image, and fusing the pseudo-depth information with the second right eye image to obtain a third right eye image;
[0095] The third left eye image and the third right eye image are input into a pathological feature recognition model trained by binocular images with known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0096] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and this embodiment will not be described here.
[0097] Generally, various embodiments can be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the application can be implemented in hardware, while other aspects can be implemented by firmware or software running on a controller, microprocessor or other computing device, but the application is not limited thereto. While various aspects of the application can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controlers or other computing devices, or some combination thereof.
[0098] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, can be stored in any apparatus-readable data storage medium and they include program instructions to implement specific tasks. The program product can include one or more computer-executable components. The one or more computer-executable components can be at least one software code or portions thereof. Further, in this regard, it should be noted that any of the Figure 7 Any block in the logical flow of the method described herein, and combinations of blocks in the logical flow, can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software can be stored in any apparatus-readable data storage medium including floppy diskettes, optical storage media, CD-ROM, magnetic cassettes, RAM, ROM, tapes, hard disk drives, or any other storage medium which can be used with the computer software product. The computer software or program is configured to cause the functions / operations described herein to be performed when the software is run by the data processor. The software or program can be written in any of a number of suitable programming languages and can be
[0099] Those skilled in the art should understand that each technical feature of the above embodiments can be combined arbitrarily, and for the sake of brevity, each technical feature in the above embodiments is not described in all possible combinations, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the description.
[0100] The above embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it should not be understood as the limitation of the scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for identifying pathological features based on binocular fundus medical images, characterized in that, Includes the following steps: The medical images of the fundus of the eye under test, including the first left eye image and the first right eye image, are acquired by a binocular camera, and the left and right optic disc centers of the first left eye image and the first right eye image are aligned to obtain the second left eye image and the second right eye image. A global disparity map is calculated based on the second left-eye image and the second right-eye image, and pseudo-depth information is obtained based on the global disparity map. The pseudo-depth information is fused with the second left-eye image to obtain a third left-eye image, and the pseudo-depth information is fused with the second right-eye image to obtain a third right-eye image. The third left-eye image and the third right-eye image are input into a pathological feature recognition model trained using binocular images with known pathological features. The pathological feature recognition model includes a feature extraction module, an attention calculation module, and a classification module. The third left-eye image and the third right-eye image are input into the feature extraction module, and feature extraction is performed using a dual-path branch with shared weights to obtain left-eye feature maps and right-eye feature maps respectively. The left-eye feature map is input into the attention calculation module to calculate the left-eye attention result, and the right-eye feature map is input into the attention calculation module to calculate the right-eye attention result. The left-eye attention result and the right-eye attention result are fused to obtain the cost volume, which is then input into the classification module to obtain the pathological feature recognition result.
2. The method for identifying pathological features based on binocular fundus medical images according to claim 1, characterized in that, Feature point matching is performed on the first left-eye image and the first right-eye image to obtain multiple pairs of matching feature points. Two pairs of matching feature points are randomly selected, and the initial transformation matrix between the matching feature points is calculated and the number of inliers in the initial transformation matrix is recorded. The initial transformation matrix with the largest number of inliers is taken as the target transformation matrix. The final transformation matrix is obtained by least squares estimation using all inliers in the target transformation matrix. Based on the final transformation matrix, the left and right optic disc centers of the first left-eye image and the first right-eye image are aligned. Here, the number of inliers is the number of matching feature points that conform to the corresponding transformation matrix.
3. The method for identifying pathological features based on binocular fundus medical images according to claim 2, characterized in that, Calculate the reprojection error of each feature point in the first left-eye image to the corresponding feature point in the first right-eye image via the initial transformation matrix, and take the matching feature points with reprojection errors less than a preset reprojection threshold as the inliers of the corresponding initial transformation matrix.
4. The method for identifying pathological features based on binocular fundus medical images according to claim 1, characterized in that, The pixel coordinate deviation of the same key point of the fundus structure in the second left eye image and the second right eye image is obtained as the disparity value. The disparity values of all key points are counted to obtain a global disparity map. The global disparity map is input into a pre-trained CNN model for feature extraction to obtain pseudo-depth information, where the pseudo-depth information is the distance from each pixel in the second left eye image and the second right eye image to the binocular camera.
5. The method for identifying pathological features based on binocular fundus medical images according to claim 1, characterized in that, The feature extraction module consists of a first feature extraction unit, a second feature extraction unit, and a third feature extraction unit. The first feature extraction unit comprises a left-eye feature extraction branch and a right-eye feature extraction branch with identical structures and shared weights. The left-eye feature extraction branch extracts features from the third left-eye image to obtain the first left-eye feature, and the right-eye feature extraction branch extracts features from the third right-eye image to obtain the first right-eye feature. The second feature extraction unit uses a feature pyramid to extract features from the first left-eye feature and the first right-eye feature to obtain the second left-eye feature and the second right-eye feature, respectively. The third feature extraction unit uses a CNN structure to extract features from the second left-eye feature and the second right-eye feature to obtain the third left-eye feature and the third right-eye feature, respectively. The third left-eye feature is used as the left-eye feature map, and the third right-eye feature is used as the right-eye feature map.
6. The method for identifying pathological features based on binocular fundus medical images according to claim 1, characterized in that, The attention calculation module includes a joint attention unit and a binocular collaborative attention mechanism unit. The joint attention unit includes channel attention, spatial attention, and size attention in sequence. The left-eye feature map and the right-eye feature map are sequentially processed by channel attention, spatial attention, and size attention to obtain the left-eye joint attention result and the right-eye joint attention result. The binocular collaborative attention mechanism unit includes a left-eye attention mechanism branch and a right-eye attention mechanism branch with the same structure and shared weights. The binocular collaborative attention mechanism calculates the feature correlation between the left-eye joint attention result and the right-eye joint attention result through the left-eye attention mechanism branch and the right-eye attention mechanism branch to obtain the cost body. The cost body includes the key features of the left and right eyes and the collaborative relationship between the key features of the left and right eyes.
7. The method for identifying pathological features based on binocular fundus medical images according to claim 1, characterized in that, The classification module consists of multiple independent classification heads, and each independent classification head corresponds to a pathological feature. The cost body is sequentially input into each independent classification head, and the classification result of each independent classification head is output through a fully connected layer to obtain the pathological feature recognition result.
8. A pathological feature recognition device based on binocular fundus medical imaging, characterized in that, include: The alignment module is used to acquire a medical image of the fundus of the eye to be tested, including a first left-eye image and a first right-eye image, through a binocular camera, and to align the left and right optic disc centers of the first left-eye image and the first right-eye image to obtain a second left-eye image and a second right-eye image. The depth calculation module calculates a global disparity map based on the second left-eye image and the second right-eye image, obtains pseudo-depth information based on the global disparity map, fuses the pseudo-depth information with the second left-eye image to obtain a third left-eye image, and fuses the pseudo-depth information with the second right-eye image to obtain a third right-eye image. The recognition module is used to input the third left-eye image and the third right-eye image into a pathological feature recognition model trained using binocular images with known pathological features. The pathological feature recognition model includes a feature extraction module, an attention calculation module, and a classification module. The third left-eye image and the third right-eye image are input into the feature extraction module, and feature extraction is performed using a dual-path branch with shared weights to obtain left-eye feature maps and right-eye feature maps respectively. The left-eye feature map is input into the attention calculation module to perform attention calculation to obtain the left-eye attention result, and the right-eye feature map is input into the attention calculation module to perform attention calculation to obtain the right-eye attention result. The left-eye attention result and the right-eye attention result are fused to obtain the cost volume, and the cost volume is input into the classification module to obtain the pathological feature recognition result.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform a method for identifying pathological features based on binocular fundus medical images as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, the computer program including program code for controlling a process to execute the process, the process including a method for identifying pathological features based on binocular fundus medical images according to any one of claims 1-7.
Citation Information
Patent Citations
Eye fundus image analysis and ophthalmic disease automatic prediction method and device based on deep learning and readable storage medium thereof
CN120563530A
Image detection and recognition method and device and computer readable storage medium
CN120672756A