Ophthalmology department clinical nursing data preprocessing method and system
By combining K-Medoids clustering and deformable image registration, the holistic and process-oriented processing mechanism of lesion feature extraction and time-driven alignment of multimodal data in the existing technology is solved, and the problems of pathological stage separation, feature extraction and time-driven pathological stage separation, lesion feature extraction and cross-modal alignment and fusion in the existing technology are solved, achieving high-quality multimodal data processing and improving lesion identification and diagnosis effects.
Patent Information
- Application Number
- CN202511001757.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-26
AI Technical Summary
Existing ophthalmic data preprocessing methods lack a holistic and streamlined processing mechanism for multimodal data, especially in terms of pathological stage separation, lesion feature extraction, cross-modal alignment and fusion, resulting in poor image recognition and intelligent diagnosis effects.
The K-Medoids clustering algorithm is used to automatically divide the pathological stages, and deformable image registration and group difference analysis are combined to identify lesion characteristics. Local image enhancement and time-driven alignment of multimodal data are used to form structured multimodal information fusion.
It improves the accuracy and stability of lesion feature extraction, enhances the visibility of lesion areas, and improves the semantic consistency and comparability of multimodal data, providing a high-quality data foundation for subsequent intelligent diagnosis.
Smart Images

Figure CN120708832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, in particular to an ophthalmology clinical nursing data preprocessing method and system. Background Art
[0002] With the continued rise in diabetes prevalence, diabetic retinopathy (DR) has become one of the leading causes of vision loss in ophthalmological diseases. To achieve early detection and targeted intervention, clinical practice involves collecting and analyzing multimodal data from DR patients, including retinal images, nursing records, and biochemical markers. However, raw clinical data often suffer from mixed pathological stages, variable image quality, heterogeneous data modalities, and extensive time spans, severely impacting subsequent image recognition, intelligent diagnosis, and personalized care decisions.
[0003] Currently, mainstream ophthalmic data preprocessing methods focus on techniques such as noise removal, image enhancement, and structural segmentation for single-modality images. These methods lack holistic, streamlined mechanisms for processing multimodal data, particularly in the areas of pathological stage separation, lesion feature extraction, and cross-modal alignment and fusion. This fragmented data processing approach makes it difficult to output high-quality structured data, limiting the application of artificial intelligence technologies such as deep learning in automated DR detection, risk stratification, and clinical decision support.
[0004] Therefore, a method and system for preprocessing ophthalmic clinical nursing data are proposed. Summary of the Invention
[0005] This invention provides a method and system for preprocessing ophthalmic clinical nursing data. Targeting multimodal clinical data from patients with diabetic retinopathy, the system sequentially performs pathological stage classification, lesion feature extraction, local image enhancement, non-image data encoding, and multimodal fusion processing. This method addresses issues such as stage confusion in raw clinical data, ambiguous lesion information, and difficulty in modal alignment, providing a structured, high-quality, unified data foundation for subsequent intelligent diagnosis and personalized nursing care.
[0006] To achieve the above object, the present invention provides the following technical solutions: A method for preprocessing ophthalmology clinical nursing data, comprising: Acquire retinal images containing multiple pathological stages, use the predefined representative images of each stage as the initial cluster centers, and automatically classify the mixed images into the corresponding pathological stages using the K-Medoide algorithm; Deformable image registration is performed on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, regions with spatial significance deviations from the cluster center image are identified, and structural pattern recognition technology is combined to extract lesion features. Perform local image enhancement processing on the corresponding area in the original image based on the extracted lesion features to improve the visibility of the lesion; Acquire non-image modality data corresponding to the image data and perform structured coding processing; The time axis is constructed based on the image acquisition time, the image modality and non-image modality data are time-driven aligned, and the multimodal information within the same time window is fused.
[0007] Furthermore, the implementation steps of the K-Medoide algorithm are: After uniformly preprocessing the obtained diabetic retinopathy mixed image samples, each image is mapped into a feature vector of fixed dimension; According to the clinical staging criteria, typical image samples of each pathological stage were manually selected, and their corresponding feature vectors were extracted as the initial cluster centers in the K-Medoids algorithm; The K-Medoids clustering algorithm was used to iteratively update the central image of each cluster based on the criterion of minimizing the total distance from the sample to the cluster center until the clusters were stable. Finally, all mixed images were divided into multiple categories consistent with the clinical stage number.
[0008] Furthermore, the steps for implementing deformable image registration are: The cluster center selected by the K-Medoide algorithm in each stage of clustering is used as the registration reference benchmark for the images in the cluster; Perform vascular trunk extraction and automatic detection of anatomical feature areas on all images to be registered to form structure-guided registration anchor points; Based on the geometric correspondence between the registration anchor points, a deformable mapping function is established using the B-Spline deformable grid, and the image to be registered is elastically transformed and aligned to the coordinate system of the reference benchmark.
[0009] Furthermore, the steps for implementing group difference analysis and identification are as follows: For all registered images in each cluster, regional segmentation and preliminary screening of candidate lesion areas are performed to ensure that the analysis area is representative and contains potential lesion areas; Perform spatial statistical analysis on the corresponding regions of each image, using regional-level statistical methods to compare the spatial differences between the registered image and the cluster center image; By using the above statistical method, image regions that have significant differences from the cluster center image in a specific area are identified and marked as potential lesion regions according to the statistical significance threshold.
[0010] Furthermore, the steps for extracting lesion features are as follows: In the significant difference area, the edge detection algorithm is used to extract the potential lesion area and compare it with the lesion area in the cluster center image to ensure the consistency between the candidate area and the lesion area; In the potential lesion area, structural pattern recognition technology, including template matching based on morphological features, shape analysis or local texture analysis, is used to identify and extract lesion features.
[0011] Furthermore, the steps for implementing the local image enhancement process are as follows: forming a weighted mask based on the extracted lesion features; The generated weighted mask is fused with the original image pixel by pixel to enhance the intensity of the corresponding pixel value in the lesion area and enhance the contrast and significance of the lesion area.
[0012] Furthermore, the steps for fusing multimodal information within the same time window are as follows: Based on the image acquisition time, a time axis is constructed, and time-driven alignment of multimodal image data and non-image data is performed according to the time axis; In the same time window, image modalities are spatially aligned to ensure consistency of different modality data on the same patient and at the same time point; non-image modality data are standardized to eliminate differences between different data sources; The image modality and non-image modality data in the same time window are spliced together to obtain multimodal comprehensive information.
[0013] The present invention also provides an ophthalmology clinical nursing data preprocessing system, comprising: The image clustering module is used to obtain retinal images containing multiple pathological stages. Based on the predefined representative images of each stage as the initial cluster centers, the mixed images are automatically classified into the corresponding pathological stages using the K-Medoide algorithm; The feature extraction module is used to perform deformable image registration on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, it identifies areas with spatial significance deviations from the cluster center image and extracts lesion features in combination with structural pattern recognition technology. An image enhancement module is used to perform local image enhancement processing on the corresponding area in the original image according to the extracted lesion features to improve the visibility of the lesion; A non-image data encoding module is used to obtain non-image modality data corresponding to the image data and perform structured encoding processing; The multimodal fusion module is used to construct a time axis based on the image acquisition time, perform time-driven alignment of image modality and non-image modality data, and fuse multimodal information within the same time window.
[0014] The beneficial effects of the present invention are: 1. This method uses a semi-supervised K-Medoide clustering algorithm. Given the well-defined clinical stages of diabetic retinopathy, representative images from each stage are manually selected as cluster centers, ensuring that the clustering results closely correspond to the actual pathological stage. Compared to traditional unsupervised clustering methods, this significantly improves the accuracy and stability of image classification, providing a more reliable data foundation for subsequent lesion feature extraction and modality fusion.
[0015] 2. This invention uses registration and population difference analysis methods to automatically identify regions within multiple images from the same phase that differ significantly from typical lesions. It then combines this with structural pattern recognition techniques to extract lesion features. This extracted lesion region is then used to construct a weighted mask, increasing the pixel values of the corresponding region in the original image. This achieves local enhancement of the lesion region, helping to improve subsequent models' ability to identify subtle lesions and their focus. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a method for preprocessing ophthalmic clinical nursing data provided by the present invention; Figure 2 This is a structural diagram of an ophthalmology clinical nursing data preprocessing system provided by the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention. Example 1
[0018] A method for preprocessing ophthalmic clinical nursing data, such as Figure 1 Shown, including: S100: Acquire retinal images containing multiple pathological stages, use the predefined representative images of each stage as the initial cluster centers, and automatically classify the mixed images into the corresponding pathological stages using the K-Medoide algorithm; Furthermore, the implementation steps of the K-Medoide algorithm are: After uniformly preprocessing the obtained diabetic retinopathy mixed image samples, each image is mapped into a feature vector of fixed dimension; According to the clinical staging criteria, typical image samples of each pathological stage were manually selected, and their corresponding feature vectors were extracted as the initial cluster centers in the K-Medoids algorithm; The K-Medoids clustering algorithm was used to iteratively update the central image of each cluster based on the criterion of minimizing the total distance from the sample to the cluster center until the clusters were stable. Finally, all mixed images were divided into multiple categories consistent with the clinical stage number.
[0019] Specifically, the acquired mixed diabetic retinopathy (DR) image samples were uniformly preprocessed. This preprocessing process included denoising, normalization, and contrast enhancement to ensure that images from different sources and of varying quality had a similar feature base. After preprocessing, each image was mapped into a feature vector of fixed dimension. The dimension of the feature vector is determined by the primary features of the image content, including grayscale statistics, morphological features, and texture features.
[0020] Based on the clinical staging criteria for diabetic retinopathy, representative image samples were manually selected for each pathological stage. These representative image samples represent the key features of different pathological stages and ensure that they cover the typical characteristics of each stage. The clinical staging criteria are based on the Chinese diabetic retinopathy staging criteria established by the Fundus Disease Group of the Chinese Ophthalmological Society in 1984, which is divided into two types and six stages.
[0021] The K-Medoids algorithm is used for clustering. The goal of this algorithm is to minimize the total distance from the sample to the cluster center. The steps of the K-Medoids algorithm are as follows: Calculate the distance between each image sample and the current cluster center. The distance metric is usually Euclidean distance or Manhattan distance, calculated based on the image's feature vector.
[0022] Each image sample is assigned to the cluster center closest to it. Each image is classified into a cluster.
[0023] The cluster center is recalculated for the images in each cluster, and the new cluster center is the Medoid of the samples in the cluster (that is, the sample with the smallest total distance from all other samples in the cluster).
[0024] Repeat the above process until the cluster center no longer changes, that is, the algorithm converges. Finally, all mixed images are divided into multiple clusters, each cluster represents a pathological stage, and the number of clusters is consistent with the number of clinical stages.
[0025] By manually selecting representative images as the initial cluster centers for the K-Medoids clustering algorithm, the clinical relevance and interpretability of the clustering results are effectively enhanced. Compared to traditional algorithms such as K-means that rely on random initialization or mean centers, K-Medoids is more robust, reducing the interference of abnormal images on clustering results while preserving the complete expression of key pathological features at each stage.
[0026] S200: Perform deformable image registration on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, identify areas with spatial significance deviations from the cluster center image, and combine structural pattern recognition technology to extract lesion features. Furthermore, the steps for implementing deformable image registration are: The cluster center selected by the K-Medoide algorithm in each stage of clustering is used as the registration reference benchmark for the images in the cluster; Perform vascular trunk extraction and automatic detection of anatomical feature areas on all images to be registered to form structure-guided registration anchor points; Based on the geometric correspondence between the registration anchor points, a deformable mapping function is established using the B-Spline deformable grid, and the image to be registered is elastically transformed and aligned to the coordinate system of the reference benchmark.
[0027] Specifically, for each image cluster generated by the K-Medoids clustering algorithm, the central image within that cluster (i.e., the K-Medoids cluster center) is selected as the registration reference image for that stage. This image typically has good imaging quality and clear vascular anatomy, effectively representing the typical anatomical features of that stage.
[0028] For each image to be registered, image preprocessing operations (such as grayscale normalization and noise filtering) are first applied. Then, a vascular trunk extraction algorithm (such as the Frangi filter) is used to extract the backbone structure of the retinal vascular network. At the same time, key anatomical features such as the optic disc and macula are detected to form a set of structure-guided registration anchor points.
[0029] Based on the set of registration anchor points between the reference image and the image to be registered, the geometric correspondence between the anchor points is calculated, and a deformation control mesh is constructed using the B-Spline free-form deformation model. In this mesh, each control point has a local displacement weight, which allows for flexible adjustment of the deformation range and direction.
[0030] Using the constructed B-Spline deformable model, an elastic transformation function is calculated from the image to be registered to the reference image. By minimizing the structural differences between the images (such as mutual information or SSD), the image to be registered is elastically deformed and mapped and aligned to the coordinate system of the reference image, ensuring good consistency of the retinal structure of the registered image.
[0031] As an optional supplement, this solution supports registration result verification and correction, including structural consistency testing of the registration results, such as calculation of vessel overlap and the structural similarity index (SSIM). For images with substandard registration, a localized re-registration mechanism is automatically triggered to ensure that the registration accuracy meets the requirements for subsequent lesion extraction and analysis.
[0032] A deformable image registration method using cluster center images as a reference effectively addresses the structural inconsistencies in diabetic retinopathy images within the same pathological stage, caused by individual differences. By forming structural anchors based on vascular trunks and anatomical regions and integrating the B-Spline elastic deformation model for precise image alignment, this method not only improves spatial consistency between images but also significantly enhances the comparability and statistical stability of lesion regions within a population.
[0033] Furthermore, the steps for implementing group difference analysis and identification are as follows: For all registered images in each cluster, regional segmentation and preliminary screening of candidate lesion areas are performed to ensure that the analysis area is representative and contains potential lesion areas; Perform spatial statistical analysis on the corresponding regions of each image, using regional-level statistical methods to compare the spatial differences between the registered image and the cluster center image; By using the above statistical method, image regions that have significant differences from the cluster center image in a specific area are identified and marked as potential lesion regions according to the statistical significance threshold.
[0034] Specifically, within each stage of deformable image registration, a structure-guided region segmentation algorithm is first applied to the image set, delineating candidate regions from the retinal image. These regions are segmented based on image features such as vascular density, color histogram distribution, and brightness variations. Background areas of no clinical value are initially excluded, ensuring that the analysis encompasses areas commonly associated with retinal pathology, such as the macula, perioptic disc, and areas with dense vascularity.
[0035] For these candidate regions, a regional-level spatial statistical analysis framework is constructed. A pixel-by-pixel or local block comparison is performed between the local region of each registered image and the corresponding location of the cluster center image within the cluster. Two-sample t-tests, z-tests, or nonparametric methods (such as the Mann-Whitney U test) are used to evaluate the statistical differences in multiple dimensions within each candidate region, such as image intensity, texture, and edge gradient.
[0036] Based on the p-value or difference score map obtained from spatial statistical analysis, a significance threshold (e.g., p < 0.01) is set to screen image regions that exhibit significant deviations from the cluster center image. These regions exhibit collective differences across multiple registered images and are highly likely to be lesions. These identified regions are then labeled as potential lesion regions in the image, and a spatial mask is generated for subsequent feature extraction and local image enhancement.
[0037] By comparing the spatial statistics of registered images and cluster center images within the same pathological stage, it is possible to accurately identify potential lesion regions with significant differences at the collective image level. This method leverages the consistency and deviation information between multiple images to effectively reduce the risk of misidentification in single-image lesion detection due to noise, individual differences, or image quality issues. By screening out areas of significant deviation at the group level, the sensitivity and stability of lesion identification are improved, providing precise spatial positioning support for subsequent structural feature extraction and image enhancement.
[0038] Furthermore, the steps for extracting lesion features are as follows: In the significant difference area, the edge detection algorithm is used to extract the potential lesion area and compare it with the lesion area in the cluster center image to ensure the consistency between the candidate area and the lesion area; In the potential lesion area, structural pattern recognition technology, including template matching based on morphological features, shape analysis or local texture analysis, is used to identify and extract lesion features.
[0039] Specifically, after completing the group difference analysis and obtaining regions of significant difference, these regions are first subjected to edge detection, preferably using classic methods such as the Canny algorithm or the Sobel operator, to extract the contour boundaries of potential lesion regions. Subsequently, the extracted edge regions are matched with the known lesion regions in the cluster center image at that stage in terms of spatial position and morphological features to ensure that the candidate regions are consistent with the typical lesion regions in terms of spatial position, shape characteristics, or brightness distribution, thereby enhancing the accuracy of screening.
[0040] In the confirmed consistent areas, structural pattern recognition technology is used to further extract lesion features. This process includes but is not limited to the following three methods: Morphological template matching: Constructs a two-dimensional template shape of typical lesions such as microaneurysms and bleeding spots, performs sliding matching within the candidate area, and selects highly matching areas as lesions using similarity evaluation indicators; Geometric shape analysis: Analyze shape factors such as area, roundness, and boundary curvature of edge regions to identify targets with typical pathological morphological features; Local texture analysis: Based on texture descriptors such as gray-level co-occurrence matrix and local binary pattern, the texture distribution of the lesion area is encoded, and texture features with significant differences are extracted to further confirm the nature of the lesion.
[0041] Finally, by combining the above structural and texture features, a structured lesion feature set is formed in each image, which serves as the key basis for subsequent image enhancement and diagnostic model training.
[0042] By incorporating a dual strategy of edge detection and structural pattern recognition into regions of significant difference, we can further enhance the structural stability and pathological interpretability of lesion identification while maintaining spatial positioning accuracy. Edge detection ensures the geometric clarity of potential lesion areas and effectively suppresses background noise interference. Structural pattern recognition combines multidimensional feature extraction mechanisms such as template matching, shape analysis, and texture mining to improve the sensitivity and specificity of lesion identification.
[0043] S300: performing local image enhancement processing on the corresponding area in the original image according to the extracted lesion features to improve the visibility of the lesion; Furthermore, the steps for implementing the local image enhancement process are as follows: forming a weighted mask based on the extracted lesion features; The generated weighted mask is fused with the original image pixel by pixel to enhance the intensity of the corresponding pixel value in the lesion area and enhance the contrast and significance of the lesion area.
[0044] Specifically, after lesion features are extracted, a lesion region mask with the same size as the original image is constructed based on its spatial location information in the image. This mask is a binary image, where the pixel value of the lesion area is set to 1 and the pixel value of the non-lesion area is set to 0, thus clearly calibrating the boundary between the enhanced and non-enhanced areas.
[0045] Subsequently, the mask image is used to enhance the original image. The enhancement method adopts a pixel-by-pixel weighted superposition strategy: ; in, is the original image pixel value, To enhance the image pixel value, is the enhancement coefficient (taken as 0.05 in this embodiment), The mask image is at position The pixel value of .
[0046] By constructing a binary mask of the lesion area and combining it with a unified pixel weighting strategy, the image enhancement process is stable and controllable. This method significantly improves the brightness and contrast of the lesion area without introducing artifacts or noise into the overall image, making small lesions more visually prominent. This helps doctors quickly identify key lesions during film reading and provides a clearer basis for subsequent image analysis models.
[0047] S400: Acquire non-image modality data corresponding to the image data and perform structured coding processing; Specifically, while collecting image data, non-image modality clinical data corresponding to each retinal image is obtained. The non-image modality data includes but is not limited to the patient's basic information (age, gender), diabetes course, glycated hemoglobin (HbA1c) level, blood pressure, visual acuity, intraocular pressure, fundus examination report, past medical history, medication records, and nursing observation records.
[0048] To achieve unified processing and analysis of multimodal data, the above non-image modality data is structured and coded according to its type. The specific steps include: Data cleaning and screening: Eliminate fields with a high proportion of missing values, repeated information, or no clinical relevance, unify units and naming, and fill in some missing fields (such as mean filling and regression filling).
[0049] Normalization of numerical data: All continuous variables (such as HbA1c and blood pressure) are normalized and converted into standard values in the interval [0,1] to avoid the impact of different dimensions on subsequent processing.
[0050] Categorical data encoding: Discrete variables (such as gender, smoking status, and types of previous diseases) are converted into vector form using one-hot encoding so that they can be consistently integrated with the image feature vector.
[0051] Text data processing: For text fields such as nursing records and fundus examination reports, keywords are extracted using a medical vocabulary and converted into fixed-length vector representations using a bag-of-words model or word embedding model (such as Word2Vec and BERT).
[0052] Unified format output: Merge all types of encoded data into a structured vector format that corresponds one-to-one with the image data, forming the basic input for multimodal data alignment.
[0053] S500: Construct a time axis based on the image acquisition time, perform time-driven alignment on the image modality and non-image modality data, and fuse the multimodal information within the same time window.
[0054] Furthermore, the steps for fusing multimodal information within the same time window are as follows: Based on the image acquisition time, a time axis is constructed, and time-driven alignment of multimodal image data and non-image data is performed according to the time axis; In the same time window, image modalities are spatially aligned to ensure consistency of different modality data on the same patient and at the same time point; non-image modality data are standardized to eliminate differences between different data sources; The image modality and non-image modality data in the same time window are spliced together to obtain multimodal comprehensive information.
[0055] Specifically, a patient-level timeline is constructed, using each patient's image acquisition time as a unified time reference. The acquisition timestamps of all image and non-image modal data are calibrated to form a time-based multimodal data index structure. For the same patient, data within a time window (e.g., ±1 day or ±3 days) is selected as the matching range. Within this window, data from different sources are aligned to ensure temporal consistency across image data, laboratory tests, electronic medical records, nursing records, and other information.
[0056] Image-level registration is performed on multiple image modalities (such as color photography, OCT, and fluorescence angiography) within the same time window, unifying them to a spatial reference frame consistent with the retinal structure. This ensures spatial consistency of lesion areas across images and improves the accuracy of subsequent joint analysis. Matched non-image modality structured data (such as clinical indicators and nursing record vectors) is standardized to unify dimensions and scales, remove statistical bias between different data sources, and improve the balance and comparability of the fused data.
[0057] The aligned image modality features are concatenated with the normalized non-image modality vectors to generate a unified multimodal comprehensive feature representation.
[0058] By constructing a timeline based on image acquisition time and aligning and fusing image and non-image modality data within a unified time window, this not only improves semantic consistency and temporal correlation between different data sources, but also effectively addresses the asynchrony issues caused by differences in multimodal data acquisition time. Image modalities ensure structural consistency through spatial registration, while non-image modalities improve data fusion compatibility through standardization, ultimately achieving the organic fusion of cross-modal and cross-source data. Example 2
[0059] The present invention also provides an ophthalmology clinical nursing data preprocessing system, such as Figure 2 Shown, including: The image clustering module is used to obtain retinal images containing multiple pathological stages. Based on the predefined representative images of each stage as the initial cluster centers, the mixed images are automatically classified into the corresponding pathological stages using the K-Medoide algorithm; The feature extraction module is used to perform deformable image registration on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, it identifies areas with spatial significance deviations from the cluster center image and extracts lesion features in combination with structural pattern recognition technology. An image enhancement module is used to perform local image enhancement processing on the corresponding area in the original image according to the extracted lesion features to improve the visibility of the lesion; A non-image data encoding module is used to obtain non-image modality data corresponding to the image data and perform structured encoding processing; The multimodal fusion module is used to construct a time axis based on the image acquisition time, perform time-driven alignment of image modality and non-image modality data, and fuse multimodal information within the same time window.
[0060] When constructing datasets for clinically assisted diagnosis of diabetic retinopathy, hospitals often face challenges such as heterogeneous data sources, extensive labeling workloads, and inconsistent data quality, hindering the efficiency and usability of high-quality dataset construction. The ophthalmology clinical nursing data preprocessing system described in this invention is deployed at the front end of the data collection and organization process, effectively improving the automation and accuracy of dataset construction.
[0061] Specifically, the system first uses an image clustering module to automatically classify a large number of unlabeled mixed retinal images into pathological stages, replacing manual coarse classification and improving image grading efficiency. This step also automatically selects representative images for each stage for subsequent expert annotation and verification, reducing annotation redundancy.
[0062] On this basis, the feature extraction module and image enhancement module automatically extract and highlight tiny lesion areas (such as microaneurysms, bleeding points, etc.), significantly improving the visibility of key lesion areas, and assisting manual labelers to frame and generate lesions more quickly and accurately, thereby improving labeling accuracy and reducing the risk of missed and mislabeled labels.
[0063] In addition, the system automatically aligns and fuses the patient's contemporaneous examination data (such as blood sugar, HbA1c, and nursing records) with image data through the non-image data encoding module and the multimodal fusion module, forming a multimodal sample with unified structure and consistent semantics, which helps to build complete data entries and meet the needs of multi-task learning.
[0064] After testing, it was found that when constructing a dataset containing tens of thousands of retinal images and corresponding multimodal labels, the application of this system can save more than 60% of the image preprocessing and stage division time, and improve the average accuracy of the downstream DR lesion detection model by more than 5%, especially in the identification of small lesions in the early DR stage.
[0065] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for preprocessing ophthalmic clinical nursing data, characterized in that: include: Acquire retinal images containing multiple pathological stages, use the predefined representative images of each stage as the initial cluster centers, and automatically classify the mixed images into the corresponding pathological stages using the K-Medoide algorithm; Deformable image registration is performed on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, regions with spatial significance deviations from the cluster center image are identified, and structural pattern recognition technology is combined to extract lesion features. Perform local image enhancement processing on the corresponding area in the original image based on the extracted lesion features to improve the visibility of the lesion; Acquire non-image modality data corresponding to the image data and perform structured coding processing; The time axis is constructed based on the image acquisition time, the image modality and non-image modality data are time-driven aligned, and the multimodal information within the same time window is fused.
2. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The implementation steps of the K-Medoide algorithm are: After uniformly preprocessing the obtained diabetic retinopathy mixed image samples, each image is mapped into a feature vector of fixed dimension; According to the clinical staging criteria, typical image samples of each pathological stage were manually selected, and their corresponding feature vectors were extracted as the initial cluster centers in the K-Medoids algorithm; The K-Medoids clustering algorithm was used to iteratively update the central image of each cluster based on the criterion of minimizing the total distance from the sample to the cluster center until the clusters were stable. Finally, all mixed images were divided into multiple categories consistent with the clinical stage number.
3. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The implementation steps of deformable image registration are: The cluster center selected by the K-Medoide algorithm in each stage of clustering is used as the registration reference benchmark for the images in the cluster; Perform vascular trunk extraction and automatic detection of anatomical feature areas on all images to be registered to form structure-guided registration anchor points; Based on the geometric correspondence between the registration anchor points, a deformable mapping function is established using the B-Spline deformable grid, and the image to be registered is elastically transformed and aligned to the coordinate system of the reference benchmark.
4. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The implementation steps of group difference analysis and identification are as follows: For all registered images in each cluster, regional segmentation and preliminary screening of candidate lesion areas are performed to ensure that the analysis area is representative and contains potential lesion areas; Perform spatial statistical analysis on the corresponding regions of each image, using regional-level statistical methods to compare the spatial differences between the registered image and the cluster center image; By using the above statistical method, image regions that have significant differences from the cluster center image in a specific area are identified and marked as potential lesion regions according to the statistical significance threshold.
5. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The steps to extract lesion features are as follows: In the significant difference area, the edge detection algorithm is used to extract the potential lesion area and compare it with the lesion area in the cluster center image to ensure the consistency between the candidate area and the lesion area; In the potential lesion area, structural pattern recognition technology, including template matching based on morphological features, shape analysis or local texture analysis, is used to identify and extract lesion features.
6. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The implementation steps of local image enhancement processing are: forming a weighted mask based on the extracted lesion features; The generated weighted mask is fused with the original image pixel by pixel to enhance the intensity of the corresponding pixel value in the lesion area and enhance the contrast and significance of the lesion area.
7. The ophthalmology clinical nursing data preprocessing method according to claim 1, characterized in that: The steps to implement the fusion of multimodal information in the same time window are as follows: Based on the image acquisition time, a time axis is constructed, and time-driven alignment of multimodal image data and non-image data is performed according to the time axis; In the same time window, image modalities are spatially aligned to ensure consistency of different modality data on the same patient and at the same time point; non-image modality data are standardized to eliminate differences between different data sources; The image modality and non-image modality data in the same time window are spliced together to obtain multimodal comprehensive information.
8. An ophthalmology clinical nursing data preprocessing system, characterized in that: include: The image clustering module is used to obtain retinal images containing multiple pathological stages. Based on the predefined representative images of each stage as the initial cluster centers, the mixed images are automatically classified into the corresponding pathological stages using the K-Medoide algorithm; The feature extraction module is used to perform deformable image registration on the images within each cluster to calibrate the retinal structural differences between the images. Based on group difference analysis, it identifies areas with spatial significance deviations from the cluster center image and extracts lesion features in combination with structural pattern recognition technology. An image enhancement module is used to perform local image enhancement processing on the corresponding area in the original image according to the extracted lesion features to improve the visibility of the lesion; A non-image data encoding module is used to obtain non-image modality data corresponding to the image data and perform structured encoding processing; The multimodal fusion module is used to construct a time axis based on the image acquisition time, perform time-driven alignment of image modality and non-image modality data, and fuse multimodal information within the same time window.
Citation Information
Cited By
Medical image data analysis system based on pixel-level multi-modal fusion
CN121304582A