Pathological feature recognition method and device based on binocular fundus medical image
By aligning binocular fundus images and constructing a shared-weight pathological feature recognition model, the problems of declining accuracy in traditional ophthalmological diagnosis and insufficient classification accuracy for multiple disease categories have been solved, enabling efficient fundus disease identification in primary healthcare settings.
Patent Information
- Application Number
- CN202511525921.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Traditional diagnosis of fundus diseases in ophthalmology relies on doctors' experience, and the accuracy decreases when applied across institutions. It does not make full use of binocular fundus image information, and the classification accuracy of multiple disease categories is insufficient, making it difficult to meet the needs of primary healthcare.
Fundus images are acquired using a binocular camera and aligned with the optic disc center. A global disparity map is calculated to obtain pseudo-depth information. A pathological feature recognition model is constructed, and a shared-weight dual-path branch is used to improve the stability of feature extraction. An attention calculation module is used to enhance pathological feature recognition.
It improves the accuracy and stability of retinal disease diagnosis, reduces diagnostic errors, enhances the ability to perceive three-dimensional lesions, adapts to different data quality levels, and is suitable for environments with limited primary healthcare resources.
Smart Images

Figure CN120997895A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image analysis, in particular to a pathological feature recognition method and device based on binocular fundus medical images. BACKGROUND
[0002] As the core basis for the diagnosis of ophthalmic diseases, fundus images can directly present the pathological features of the retina, optic nerve and blood vessel system, and are the key tool for early screening of high-blinding diseases such as diabetic retinopathy, glaucoma, cataract and age-related macular degeneration. Its basic role in ophthalmic clinical diagnosis has been widely recognized.
[0003] However, the traditional ophthalmic fundus disease diagnosis method highly depends on the personal clinical experience of doctors, and the diagnosis result is easily affected by subjective factors such as the professional level and energy state of doctors, and there is significant individual difference. Especially in the primary areas where medical resources are scarce, the number of professional ophthalmologists is scarce, and it is difficult to meet the large-scale fundus disease screening demand, resulting in frequent misdiagnosis and missed diagnosis, and a large number of patients miss the best opportunity for early intervention and treatment, further aggravating the risk of blindness of ophthalmic diseases.
[0004] In recent years, with the development of deep learning technology, computer-aided diagnosis technology based on ResNet, EfficientNet and other models has made certain progress in the field of monocular fundus image classification, providing a new idea for fundus disease diagnosis. However, the existing technology still faces three major challenges and is difficult to meet the needs of clinical practical application: first, the model performance is highly sensitive to data quality, and the imaging parameter difference of different medical institutions and different types of collection equipment easily leads to uneven brightness and obvious color difference of fundus images, seriously affecting the generalization ability of the model, and making the model accuracy greatly decrease when applied across institutions; second, the existing technology focuses on monocular fundus image analysis, and fails to fully utilize the spatial-channel multi-dimensional correlation information contained in binocular fundus images, and cannot simulate the binocular collaborative judgment mechanism of the human visual system, missing the key path of improving lesion recognition accuracy through binocular comparison; third, the multi-class disease classification accuracy is insufficient, and in clinical practice, 7 or more fundus diseases (including the normal category) need to be accurately identified, but the existing model is prone to class confusion when dealing with complex scenes with uneven class distribution and similar lesion characteristics, and it is difficult to meet the strict accuracy requirements of clinical diagnosis.
[0005] In summary, it is urgent to develop an intelligent fundus disease diagnosis scheme that can adapt to different data quality, fully utilize binocular information and improve the recognition accuracy of multi-class pathological features, in order to make up for the shortcomings of traditional diagnosis methods and relieve the pressure of primary medical resources, and provide strong technical support for early screening and accurate diagnosis of fundus diseases. SUMMARY
[0006] The embodiment of the application provides a pathological feature recognition method and device based on binocular fundus medical images, pseudo-depth information is obtained by aligning images to calculate a global disparity map and fusion, a pathological feature recognition model is constructed, and the stability and effectiveness of feature extraction are improved by using a shared weight double-path branch, so that the diagnostic error is reduced and the perception ability of the model to fundus stereoscopic lesions is improved.
[0007] In a first aspect, the embodiment of the application provides a pathological feature recognition method based on binocular fundus medical images, which comprises: Obtaining the to-be-tested fundus medical images including the first left eye image and the first right eye image by using the binocular camera, and obtaining the second left eye image and the second right eye image by aligning the centers of the left and right optic discs of the first left eye image and the first right eye image; Calculating the global disparity map based on the second left eye image and the second right eye image, obtaining the pseudo-depth information based on the global disparity map, fusing the pseudo-depth information with the second left eye image to obtain the third left eye image, and fusing the pseudo-depth information with the second right eye image to obtain the third right eye image; Inputting the third left eye image and the third right eye image into the pathological feature recognition model trained by using the binocular images of known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module, and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using the double-path branch with shared weights to obtain the left eye feature map and the right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain the left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain the right eye attention result, the left eye attention result and the right eye attention result are fused to obtain the cost volume, and the cost volume is input into the classification module to obtain the pathological feature recognition result.
[0008] In a second aspect, the embodiment of the application provides an intelligent ophthalmic disease diagnosis device based on fundus medical images, which comprises: The alignment module is configured to obtain the to-be-tested fundus medical images including the first left eye image and the first right eye image by using the binocular camera, and obtain the second left eye image and the second right eye image by aligning the centers of the left and right optic discs of the first left eye image and the first right eye image; The depth calculation module is configured to calculate the global disparity map based on the second left eye image and the second right eye image, obtain the pseudo-depth information based on the global disparity map, fuse the pseudo-depth information with the second left eye image to obtain the third left eye image, and fuse the pseudo-depth information with the second right eye image to obtain the third right eye image; The identification module is configured to input the third left-eye image and the third right-eye image into a pathological feature identification model trained by using binocular images with known pathological features, wherein the pathological feature identification model comprises a feature extraction module, an attention calculation module, and a classification module, the third left-eye image and the third right-eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left-eye feature map and a right-eye feature map, the left-eye feature map is input into the attention calculation module to perform attention calculation to obtain a left-eye attention result, the right-eye feature map is input into the attention calculation module to perform attention calculation to obtain a right-eye attention result, the left-eye attention result and the right-eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature identification result.
[0009] In a third aspect, an electronic device is provided, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform a pathological feature identification method based on binocular fundus medical images.
[0010] In a fourth aspect, a readable storage medium is provided, the readable storage medium stores a computer program, and the computer program includes program code for controlling a process to perform a process, and the process includes a pathological feature identification method based on binocular fundus medical images.
[0011] The main contributions and innovations of the present application are as follows: The present application obtains fundus images by using a binocular camera and aligns the left and right optic disc centers, thereby eliminating the initial position deviation of the binocular images, providing accurate image basis for subsequent global disparity map calculation and pseudo-depth information acquisition, and reducing the diagnostic errors caused by image misplacement; the present application calculates a global disparity map based on the aligned binocular images to obtain pseudo-depth information, and fuses the pseudo-depth information with the binocular images, thereby supplementing the stereo structure information of the fundus images, increasing the information dimension of the model, improving the perception ability of the model to the stereo lesions of the fundus, and enabling the model to capture the lesion features more comprehensively; the present application constructs a pathological feature identification model comprising a feature extraction module, an attention calculation module, and a classification module, the feature extraction module adopts a double-path branch with shared weights to efficiently extract the respective features of the binocular images, ensures the consistency of the binocular feature extraction standards, reduces the feature difference interference, and improves the stability and effectiveness of the feature extraction.
[0012] Details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects, and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings: Figure 1 is a flow chart of a pathological feature recognition method based on binocular fundus medical images according to an embodiment of the application; Figure 2 is a schematic diagram of the overall structure of a pathological feature recognition model according to an embodiment of the application; Figure 3 is a schematic diagram of the structure of a second feature extraction branch and a third feature extraction branch according to an embodiment of the application; Figure 4 is a schematic diagram of the structure of an attention module according to an embodiment of the application; Figure 5 is a schematic diagram of the structure of a classification module according to an embodiment of the application; Figure 6 is a schematic diagram of the structure of a pathological feature recognition device based on binocular fundus medical images according to an embodiment of the application; Figure 7 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION
[0014] The exemplary embodiments will be described in detail herein below with reference to the drawings. The following description is with reference to the drawings, wherein like numerals refer to like elements throughout. The embodiments described in the following exemplary embodiments are not representative of all embodiments consistent with one or more aspects of the present description. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of the present description as detailed in the appended claims.
[0015] It should be noted that the steps of the corresponding method in other embodiments are not necessarily performed in the order shown and described in the present description. In some other embodiments, the steps included in the method can be more or less than described in the present description. In addition, a single step described in the present description can be divided into multiple steps for description in other embodiments, and multiple steps described in the present description can be combined into a single step for description in other embodiments.
[0016] Embodiment One The embodiment of the application provides a pathological feature recognition method based on binocular fundus medical images, pseudo-depth information is obtained by aligning images to calculate a global disparity map and fusion, a pathological feature recognition model is constructed, shared weight double-path branches are used to improve feature extraction stability and effectiveness, so that diagnosis errors are reduced and the perception ability of the model to fundus stereoscopic lesions is improved. Specifically, referring to Figure 1 , the method comprises: obtaining a to-be-tested fundus medical image comprising a first left eye image and a first right eye image through a binocular camera, and performing left-right optic disc center alignment on the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image; calculating a global disparity map based on the second left eye image and the second right eye image, and obtaining pseudo-depth information based on the global disparity map, fusing the pseudo-depth information with the second left eye image to obtain a third left eye image, and fusing the pseudo-depth information with the second right eye image to obtain a third right eye image; inputting the third left eye image and the third right eye image into a pathological feature recognition model trained by using binocular images of known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using double-path branches with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0017] In some specific embodiments, since the to-be-tested fundus medical image directly obtained by the binocular camera has problems of shadow, size difference and brightness and color difference imbalance, the recognition and classification accuracy of eye diseases is affected, so the first left eye image and the first right eye image are preprocessed to eliminate image differences before the left-right optic disc center alignment is performed on the first left eye image and the first right eye image.
[0018] Further, the first left eye image and the first right eye image are preprocessed through gray-scale, filter denoising, contrast enhancement and the like. Specifically, the first left eye image and the first right eye image are subjected to contrast enhancement through a CLAHE algorithm, thereby increasing the contrast of local blood vessels and making the local feature enhancement better. When the CLAHE algorithm is used to enhance the contrast of the first left eye image, the first left eye image is first divided into non-overlapping rectangular sub-regions, and a histogram of each rectangular sub-region is calculated. A contrast threshold is set, all pixel points with pixel values exceeding the contrast threshold in the histogram of each rectangular sub-region are obtained, and these pixel points are evenly distributed to all gray levels to obtain a clipping histogram corresponding to each rectangular sub-region. The cumulative distribution function of each clipping histogram is calculated as a corresponding mapping function, and each pixel point in the first left eye image is subjected to bilinear interpolation based on the mapping function, thereby obtaining the first left eye image after contrast enhancement. When the CLAHE algorithm is used to enhance the contrast of the first right eye image, the first right eye image is first divided into non-overlapping rectangular sub-regions, and a histogram of each rectangular sub-region is calculated. A contrast threshold is set, all pixel points with pixel values exceeding the contrast threshold in the histogram of each rectangular sub-region are obtained, and these pixel points are evenly distributed to all gray levels to obtain a clipping histogram corresponding to each rectangular sub-region. The cumulative distribution function of each clipping histogram is calculated as a corresponding mapping function, and each pixel point in the first right eye image is subjected to bilinear interpolation based on the mapping function, thereby obtaining the first right eye image after contrast enhancement.
[0019] Exemplarily, the size of the first left eye image is MxN, the first left eye image is divided into KxL rectangular sub-regions, and the size of each rectangular sub-region is mxn, where m=M / K and n=N / L. Then, the histogram of each rectangular sub-region is calculated. In the present scheme, the contrast threshold is wherein, is a limiting coefficient, and is usually 1-4; =256, indicating 256 gray levels, and the clipping histogram is obtained by the formula:
[0020] wherein, cxcess is all pixel points exceeding the contrast threshold, represents the histogram of the rectangular sub-region with index k, i is a pixel point therein, and C is the contrast threshold.
[0021] The formula for evenly distributing these pixel points to all gray levels is:
[0022] wherein, is the clipping histogram, Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, =256 represents 256 gray levels.
[0023] Specifically, the calculation formula of the cumulative distribution function is represented as:
[0024] wherein, Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, Nk represents the total number of pixels in the corresponding rectangular sub-region, =256 represents 256 gray levels.
[0025] Since the block processing can cause artifacts at the boundary of the rectangular sub-region, the bilinear interpolation is adopted to process each pixel point when the pixel point is mapped based on the mapping function. Specifically, for any pixel in the first left eye image, the four adjacent rectangular sub-regions of the pixel are determined as and the distance weight of the pixel to the center points of the four adjacent rectangular sub-regions is calculated as The updating method of the pixel value is:
[0026] Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold, Hk(i) represents a histogram of the rectangular sub-region with index k, i is a pixel point therein, C is a contrast threshold, and cxcess is all pixel points exceeding the contrast threshold,
[0027] Specifically, the present scheme better distinguishes the details in the image by evenly distributing the pixel points with high contrast to all pixel levels.
[0028] Specifically, the contrast enhancement method of the first right eye image is the same as that of the first left eye, which is not repeated here.
[0029] In some embodiments, the feature point matching of the first left eye image and the first right eye image obtains a plurality of pairs of matching feature points, two pairs of matching feature points are randomly selected, an initial transformation matrix between the matching feature points is calculated, and the number of inliers in the initial transformation matrix is recorded, so as to obtain an initial transformation matrix with the largest number of inliers as a target transformation matrix, and a final transformation matrix is obtained by least square estimation based on all inliers in the target transformation matrix, and the first left eye image and the first right eye image are aligned based on the left and right optic disc centers based on the final transformation matrix, wherein the inliers are the number of matching feature points meeting the corresponding transformation matrix.
[0030] Further, the re-projection error of each feature point in the first left eye image mapped to the corresponding feature point in the first right eye image via the initial transformation matrix is calculated, and the matching feature point with a re-projection error less than a preset re-projection threshold is taken as an inlier of the corresponding initial transformation matrix.
[0031] Specifically, the SIFT feature point matching algorithm is used to match the feature points of the first left eye image and the first right eye image, and the final transformation matrix acquisition method of the present scheme can better handle complex image deformation and perspective difference.
[0032] In some embodiments, the pixel coordinate deviation value of the same key point of the fundus structure in the second left eye image and the second right eye image is taken as the disparity value, the disparity values of all key points are counted to obtain a global disparity map, and the global disparity map is input into a pre-trained CNN model for feature extraction to obtain pseudo-depth information, wherein the pseudo-depth information is the distance from each pixel point in the second left eye image and the second right eye image to the binocular camera.
[0033] Specifically, a common fundus image generally contains three-channel data of RGB, but the fundus image contains some stereoscopic structures, so all details of the fundus image cannot be displayed by only three-channel images, so the present scheme provides more information dimension options for the model by calculating and adding pseudo-depth information, which helps to improve the perception ability of the model to the stereoscopic structure.
[0034] In some specific embodiments, the pathological feature recognition model is a model pre-trained based on the ODIR-5K dataset, and in the training process, random rotation, flipping, elastic transformation and other techniques covering the spatial level are used, combined with a generative adversarial network to enhance the training data, and based on a conditional GAN (cGAN) architecture, microaneurysms and other subtle lesions are synthesized to supplement the small sample dataset.
[0035] In some specific embodiments, the third left eye image and the third right eye image are normalized to the range [0, 1] before being input into the pathological feature recognition model.
[0036] In some embodiments, the overall structure of the pathological feature recognition model is as shown in Figure 2 The feature extraction module is composed of a first feature extraction unit, a second feature extraction unit, and a third feature extraction unit. The first feature extraction unit includes a left eye feature extraction branch and a right eye feature extraction branch that are structurally identical and share weights. The left eye feature extraction branch extracts features from the third left eye image to obtain first left eye features. The right eye feature extraction branch extracts features from the third right eye image to obtain first right eye features. The second feature extraction unit uses a feature pyramid to extract features from the first left eye features and the first right eye features to obtain second left eye features and second right eye features. The third feature extraction unit uses a CNN structure to extract features from the second left eye features and the second right eye features to obtain third left eye features and third right eye features. The third left eye features are taken as a left eye feature map, and the third right eye features are taken as a right eye feature map.
[0037] Specifically, the left eye feature extraction branch and the right eye feature extraction branch in the present scheme use CrossViT structure, and the left eye feature extraction branch and the right eye feature extraction branch drive the extraction of single-eye features by sharing weights.
[0038] Specifically, the second feature extraction branch uses a feature pyramid to perform fine upsampling and dynamic weighted fusion on the first left eye features and the first right eye features, thereby reasonably allocating the contribution of each level to the lesion area, thereby significantly improving the detection ability of small lesions such as microaneurysms.
[0039] Specifically, the third feature extraction branch using CNN structure after the second feature extraction branch performs further feature extraction, which can integrate the global modeling capability of visual Transformer and the local detail capturing advantage of CNN. The structure of the second feature extraction branch and the third feature extraction branch is as shown in Figure 3
[0040] In some embodiments, the schematic diagram of the attention module is as shown in Figure 4 As shown, the attention calculation module comprises a joint attention unit and a binocular collaborative attention mechanism unit, the joint attention unit comprises channel attention, spatial attention and size attention connected in sequence, and the left eye feature map / right eye feature map sequentially passes through the channel attention, the spatial attention and the size attention to obtain the left eye joint attention result / right eye joint attention result, the binocular collaborative attention mechanism unit comprises a left eye attention mechanism branch and a right eye attention mechanism branch which are structurally identical and share weights, the binocular collaborative attention mechanism calculates the feature correlation between the left eye joint attention result and the right eye joint attention result through the left eye attention mechanism branch and the right eye attention mechanism branch to obtain a cost volume, which comprises the key features of the left eye and the right eye and the collaborative relationship between the left eye key features and the right eye key features.
[0041] Specifically, the present scheme first strengthens the weight of the channel sensitive to the pathological features of ophthalmic diseases through channel attention calculation, and reduces the influence of redundant channels on model judgment, then fuses the maximum / average pooling double-path weight through the calculation of the spatial attention mechanism, and strengthens the features of key regions such as optic disc and macula, and the scale attention mechanism uses the hollow convolution to capture the multi-scale features such as the cup-disc ratio of glaucoma, and through the re-labeling and enhancement of features of different sizes, the recognition ability of the model to multi-scale objects is improved.
[0042] Further, the left eye attention mechanism branch comprises a left eye backbone network layer, a left eye error calculation layer, a left eye adversarial network layer and a left eye convolution layer, the right eye attention mechanism branch comprises a right eye backbone network layer, a right eye error calculation layer, a right eye adversarial network layer and a right eye convolution layer, the left eye joint attention result is subjected to feature extraction in the left eye backbone network layer to obtain left eye joint attention features, the right eye joint attention result is subjected to feature extraction in the right eye backbone network layer to obtain right eye joint attention features, the average absolute error of the left eye joint attention features and the right eye joint attention features is calculated in the left eye error calculation layer to obtain left eye feature difference, the average absolute error of the right eye joint attention features and the left eye joint attention features is calculated in the right eye error calculation layer to obtain right eye feature difference, the feature distribution of the left eye feature difference is optimized in the left eye adversarial network layer to obtain left eye adversarial features, the feature distribution of the right eye feature difference is optimized in the right eye adversarial network layer to obtain right eye adversarial features, the left eye adversarial features are subjected to feature integration in the left eye convolution layer to obtain left eye collaborative attention result, the right eye adversarial features are subjected to feature integration in the right eye convolution layer to obtain right eye collaborative attention result, and the cost volume is obtained by fusing the left eye collaborative attention result and the right eye collaborative attention result.
[0043] Specifically, the present scheme adopts Crossvit as the left and right eye backbone network layer, the error calculation layer calculates the difference between the features by mean absolute error to highlight the important feature difference, the adversarial network is used to enhance the generalization ability of the model, so that it can better resist adversarial attacks and noise interference, and the cost volume generation can make the model more accurately understand and process visual information from two different perspectives, thereby improving the performance of binocular vision tasks.
[0044] In some embodiments, in order to verify the effectiveness of the attention module of the present scheme, 3000 cases of data are carefully selected from the ODIR-5K dataset, which are randomly divided into a training set (2100 cases) and a test set (900 cases) according to a 7:3 ratio. The first group is the original data group (Original) without any processing; the second group uses the joint attention unit and the binocular collaborative attention mechanism unit for training, and then the seven groups of data are input into the classification module constructed in the present study for training, and the final experimental results are shown in Table 1. From Table 1, it can be seen that the accuracy, precision and recall have been improved.
[0045] Table 1 Superiority proof of attention module
[0046] In some embodiments, the structure of the classification module is shown in Figure 5 The classification module is composed of multiple independent classification heads, and each independent classification head corresponds to the pathological features of an ophthalmic disease. The cost volume is sequentially input into each independent classification head, and the classification results of each independent classification head are output through a fully connected layer to obtain the pathological feature recognition results.
[0047] Specifically, the present scheme uses multiple independent classification heads to classify the pathological features of different ophthalmic diseases, thereby achieving the effect of reducing feature coupling. In addition, the present scheme uses a fully connected layer and a Sigmoid activation structure to calculate the probability of each classification result and output.
[0048] Specifically, the present scheme dynamically adjusts the decision threshold of each class based on the ROC curve of the validation set, and ensures the coexistence of high sensitivity and specificity through the precision-recall balance strategy.
[0049] In some embodiments, two groups of controls are set: 1) Original group: keep the original data distribution, directly train with the conventional 8-classification head; 2) HDCH group: adopt 8 independent classification heads for parallel training, the existing mainstream method generally adopts the "single forward-8 class" paradigm, that is, once the forward inference is given, the probability distribution of 8 classes is given. In contrast, the hierarchical decoupling strategy proposed in this paper divides the overall task into 8 independent binary classification sub-tasks, each of which is independently completed by a dedicated classification head for training and inference, thereby completely blocking the semantic coupling between pathological features. Finally, the 8 groups of binary classification results are integrated according to the pre-set fusion rule, and the final experimental results are shown in Table 2. As can be seen from Table 2, the Decoupled group has significantly improved in accuracy, precision and recall: the overall accuracy has increased by 7.02% compared with the Original group, and the average precision and recall of each pathological feature have increased by 4.21% and 5.01% respectively, verifying the effectiveness of the hierarchical decoupling strategy in reducing class interference and improving fine-grained diagnostic performance: Table 2 Superiority of classification head module
[0050] In some embodiments, in order to verify the superiority of the overall pathological feature recognition model in the present scheme, 3000 cases of data are also selected from the ODIR-5K dataset. The optimal method of experiments 1 and 2 is adopted, and the data is input into different models for comparison. After the model training is completed, the two groups of trained model parameters are extracted and applied to the test set in the test link. The labels output by the model training are paired with the true labels of the test set. Finally, by calculating the Accuracy (accuracy), precision, and recall three key indicators, the performance of the model under different processing methods is comprehensively evaluated. From the data in Table 3, it can be seen that the model compared with other models has improved in the three evaluation indicators, proving the superiority of the model.
[0051] Table 3 Superiority of overall pathological feature recognition model
[0052] Embodiment two Based on the same idea, reference Figure 6 The application also proposes a pathological feature recognition device based on binocular fundus medical images, comprising: An alignment module is configured to acquire a to-be-tested fundus medical image comprising a first left eye image and a first right eye image through a binocular camera, and align the centers of the left and right optic discs of the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image; The depth calculation module calculates a global disparity map based on the second left-eye image and the second right-eye image, and acquires pseudo-depth information based on the global disparity map, fuses the pseudo-depth information with the second left-eye image to obtain a third left-eye image, and fuses the pseudo-depth information with the second right-eye image to obtain a third right-eye image. The recognition module is configured to input the third left-eye image and the third right-eye image into a pathological feature recognition model trained by using binocular images with known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module, and a classification module, the third left-eye image and the third right-eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left-eye feature map and a right-eye feature map, the left-eye feature map is input into the attention calculation module to perform attention calculation to obtain a left-eye attention result, the right-eye feature map is input into the attention calculation module to perform attention calculation to obtain a right-eye attention result, the left-eye attention result and the right-eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0053] Embodiment Three This embodiment further provides an electronic device, referring to Figure 7 The electronic device comprises a memory 404 and a processor 402, the memory 404 stores a computer program, and the processor 402 is configured to execute the computer program to perform the steps in any of the above method embodiments.
[0054] Specifically, the processor 402 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0055] The memory 404 can include a mass storage that stores data or instructions. For example, and without limitation, the memory 404 can include a Hard Disk Drive (HDD), a floppy disk drive, a Solid State Drive (SSD), a flash drive, a Compact Disc Read Only Memory (CD-ROM), a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. The memory 404 can be removable and / or non-removable (or fixed) as appropriate. The memory 404 can be internal or external as appropriate. In particular embodiments, the memory 404 is a Non-Volatile memory. In particular embodiments, the memory 404 includes a Read-Only Memory (ROM) and a Random Access Memory (RAM). The ROM can be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), an Electrically Alterable ROM (EAROM), or a FLASH memory, or a combination of two or more of these, as appropriate. The RAM can be a Static Random-Access Memory (SRAM) or a Dynamic Random Access Memory (DRAM), which can be a Fast Page Mode Dynamic Random Access Memory (FPMDRAM), an Extended Data Output Dynamic Random Access Memory (EDODRAM), a Synchronous Dynamic Random-Access Memory (SDRAM), or the like, as appropriate.
[0056] The memory 404 can be used to store or buffer various data files needed for processing and / or communication, and possible computer program instructions executed by the processor 402.
[0057] The processor 402 can implement any of the pathological feature recognition methods based on binocular fundus medical images in the above embodiments by reading and executing the computer program instructions stored in the memory 404.
[0058] Optionally, the electronic device can further include a transmission device 406 connected to the processor 402 and an input / output device 408 connected to the processor 402.
[0059] The transmission device 406 can be used to receive or send data via a network. The network can include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (NIC) which can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module which is used to communicate with the Internet in a wireless manner.
[0060] The input / output device 408 is used to input or output information. In the present embodiment, the input information can be a fundus medical image to be tested, and the output information can be a pathological feature recognition result.
[0061] Optionally, in the present embodiment, the processor 402 can be configured to perform the following steps by computer program: acquire a fundus medical image to be tested including a first left eye image and a first right eye image by using a binocular camera, and perform left-right optic disc center alignment on the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image; calculate a global disparity map based on the second left eye image and the second right eye image, obtain pseudo-depth information based on the global disparity map, fuse the pseudo-depth information with the second left eye image to obtain a third left eye image, and fuse the pseudo-depth information with the second right eye image to obtain a third right eye image; The third left eye image and the third right eye image are input into a pathological feature recognition model trained by binocular images with known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
[0062] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and this embodiment will not be described here.
[0063] Generally, various embodiments can be implemented in hardware or special-purpose circuitry, software, logic or any combination thereof. Some aspects of the application can be implemented in hardware, while other aspects can be implemented using firmware or software executed by a controller, microprocessor or other computing device, but the application is not limited thereto. While various aspects of the application can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controlers or other computing devices, or some combination thereof.
[0064] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, can be stored in any apparatus-readable data storage medium and they include program instructions to implement specific tasks. The program product can include one or more computer-executable components that can be configured to carry out the embodiments. The one or more computer-executable components can be one or more software codes or portions thereof. Further, in this regard, it should be noted that any of the Figure 7 Any block in the logical flow of the method described herein, and combinations of blocks in the logical flow, can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software can be stored in any apparatus-readable data storage medium including floppy diskettes, optical storage media, CD-ROM, magnetic cassettes, RAMs, and ROMs, and can be read by the data processor of the mobile device such as in the processor entity. The computer software in the physical media can cause the processor to carry out a method of the application. The application can be implemented by means of computer software running on the processor entity and the computer software can be stored in a computer program product. The software can cause the processor to carry out the steps of the method of the application.
[0065] Those skilled in the art should understand that each technical feature of the above embodiments can be combined arbitrarily, and for the sake of brevity, each technical feature in the above embodiments is not described in all possible combinations, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the description.
[0066] The above embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it should not be understood as the limitation of the scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for recognizing pathological features based on binocular fundus medical images, characterized in that, The method comprises the following steps: Obtaining the to-be-tested fundus medical image comprising a first left eye image and a first right eye image through a binocular camera, and performing left-right optic disc center alignment on the first left eye image and the first right eye image to obtain a second left eye image and a second right eye image; Calculating a global disparity map based on the second left eye image and the second right eye image, and obtaining pseudo-depth information based on the global disparity map, fusing the pseudo-depth information with the second left eye image to obtain a third left eye image, and fusing the pseudo-depth information with the second right eye image to obtain a third right eye image; Inputting the third left eye image and the third right eye image into a pathological feature recognition model trained by using binocular images with known pathological features, wherein the pathological feature recognition model comprises a feature extraction module, an attention calculation module, and a classification module, the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map, the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result, the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result, the left eye attention result and the right eye attention result are fused to obtain a cost volume, and the cost volume is input into the classification module to obtain a pathological feature recognition result.
2. The pathological feature recognition method based on binocular fundus medical images according to claim 1, characterized in that, Performing feature point matching on the first left eye image and the first right eye image to obtain a plurality of pairs of matching feature points, randomly selecting two pairs of matching feature points, calculating an initial transformation matrix between the matching feature points and recording the number of inliers in the initial transformation matrix to obtain an initial transformation matrix with the most inliers as a target transformation matrix, and performing least squares estimation on all inliers in the target transformation matrix to obtain a final transformation matrix, and performing left-right optic disc center alignment on the first left eye image and the first right eye image based on the final transformation matrix, wherein the inliers are the number of matching feature points that meet the corresponding transformation matrix.
3. The pathological feature recognition method based on binocular fundus medical images according to claim 2, characterized in that, Calculating the re-projection error of each feature point in the first left eye image mapped to the corresponding feature point in the first right eye image through the initial transformation matrix, and taking the matching feature points with a re-projection error less than a preset re-projection threshold as the inliers of the corresponding initial transformation matrix.
4. The pathological feature recognition method based on binocular fundus medical images according to claim 1, characterized in that, Obtaining the pixel coordinate deviation value of the same key point of the fundus structure in the second left eye image and the second right eye image as a disparity value, and obtaining a global disparity map by counting the disparity values of all key points, and inputting the global disparity map into a pre-trained CNN model for feature extraction to obtain pseudo-depth information, wherein the pseudo-depth information is the distance from each pixel point in the second left eye image and the second right eye image to the binocular camera.
5. The pathological feature recognition method based on binocular fundus medical images according to claim 1, characterized in that, The feature extraction module is composed of a first feature extraction unit, a second feature extraction unit and a third feature extraction unit, the first feature extraction unit includes left eye feature extraction branches and right eye feature extraction branches which are the same in structure and share weights, the left eye feature extraction branches extract features from the third left eye image to obtain first left eye features, the right eye feature extraction branches extract features from the third right eye image to obtain first right eye features, the second feature extraction unit uses a feature pyramid to extract features from the first left eye features and the first right eye features respectively to obtain second left eye features and second right eye features, and the third feature extraction unit extracts features from the second left eye features and the second right eye features respectively by a CNN structure to obtain third left eye features and third right eye features, and the third left eye features are taken as a left eye feature map and the third right eye features are taken as a right eye feature map.
6. The pathological feature recognition method based on binocular fundus medical images according to claim 1, characterized in that, The attention calculation module includes a joint attention unit and a binocular cooperative attention mechanism unit, the joint attention unit includes channel attention, spatial attention and size attention connected in sequence, and the left eye feature map / right eye feature map sequentially passes through the channel attention, the spatial attention and the size attention to obtain left eye joint attention results / right eye joint attention results, and the binocular cooperative attention mechanism unit includes left eye attention mechanism branches and right eye attention mechanism branches which are the same in structure and share weights, the binocular cooperative attention mechanism calculates the feature correlation between the left eye joint attention results and the right eye joint attention results by the left eye attention mechanism branches and the right eye attention mechanism branches to obtain a cost volume, and the cost volume includes the key features of the left eye and the right eye and the cooperative relationship between the left eye key features and the right eye key features.
7. The pathological feature recognition method based on binocular fundus medical images according to claim 1, characterized in that, The classification module is composed of a plurality of independent classification heads, each independent classification head corresponds to a pathological feature, the cost volume is sequentially input into each independent classification head, and the classification results of each independent classification head are output by a full connection layer to obtain pathological feature recognition results.
8. A pathological feature recognition device based on binocular fundus medical images, characterized by, It includes: An alignment module is configured to acquire, by a binocular camera, a to-be-tested fundus medical image including a first left eye image and a first right eye image, and align the first left eye image and the first right eye image with respect to left and right optic disc centers to obtain a second left eye image and a second right eye image; A depth calculation module is configured to calculate a global disparity map based on the second left eye image and the second right eye image, and acquire pseudo-depth information based on the global disparity map, fuse the pseudo-depth information with the second left eye image to obtain a third left eye image, and fuse the pseudo-depth information with the second right eye image to obtain a third right eye image; and A feature extraction module is configured to extract features from the third left eye image and the third right eye image by a feature extraction network to obtain left eye features and right eye features. The identification module is configured to input the third left eye image and the third right eye image into a pathological feature identification model trained by binocular images with known pathological features, wherein the pathological feature identification model comprises a feature extraction module, an attention calculation module, and a classification module; the third left eye image and the third right eye image are input into the feature extraction module to perform feature extraction by using a double-path branch with shared weights to obtain a left eye feature map and a right eye feature map; the left eye feature map is input into the attention calculation module to perform attention calculation to obtain a left eye attention result; the right eye feature map is input into the attention calculation module to perform attention calculation to obtain a right eye attention result; the left eye attention result and the right eye attention result are fused to obtain a cost volume; and the cost volume is input into the classification module to obtain a pathological feature identification result. 9.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to run the computer program to perform the pathological feature identification method based on binocular fundus medical images according to any one of claims 1-7.
10. A readable storage medium, characterized by, The readable storage medium stores a computer program, and the computer program comprises program code for controlling a process to perform the process, and the process comprises the pathological feature identification method based on binocular fundus medical images according to any one of claims 1-7.
Citation Information
Patent Citations
Eye fundus image analysis and ophthalmic disease automatic prediction method and device based on deep learning and readable storage medium thereof
CN120563530A
Image detection and recognition method and device and computer readable storage medium
CN120672756A
Ophthalmologic apparatus, and method of controlling the same
US20200237213A1
Method, Device, Electronic Equipment and Storage Medium for Positioning Macular Center in Fundus Images
US20220415087A1
Method for training model for recognizing medical image, method for recognizing medical image, and electronic device
US20250124694A1