Ophthalmic tumor image feature extraction method and system based on artificial intelligence
By using multimodal image registration and adaptive feature extraction, combined with a dual-channel attention mechanism and residual enhancement network, the modeling problem of the spatial correlation between retinal vascular topology and lesion region in ophthalmic tumor image feature extraction is solved, improving the integrity and stability of feature extraction and reducing the dependence on labeled data.
Patent Information
- Application Number
- CN202511377228.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies lack modeling of the spatial relationship between retinal vascular topology and lesion region in ophthalmic tumor image feature extraction. Multimodal images are prone to spatial misalignment, noise is generated during feature fusion, and they are highly dependent on labeled data with insufficient generalization ability.
We employ a multimodal image accurate registration and adaptive feature extraction method. By combining an improved convolutional module and a dual-channel attention mechanism with a residual enhancement network and a semi-supervised learning framework, we construct a dynamically weighted cross-domain fusion mechanism to improve feature capture capability and stability.
It improves the integrity of irregular tumor edge features, reduces modal misalignment noise, balances the weights of spatial features and vascular topology features, reduces dependence on labeled data, and improves the stability and robustness of feature representation under limited data.
Smart Images

Figure CN121121379A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a method and system for extracting features from ophthalmic tumor images based on artificial intelligence. Background Technology
[0002] Current methods for extracting features from ophthalmic tumor images mostly employ single-modality data processing, using traditional convolutional neural networks to extract local image features or relying on manually designed feature operators to capture texture information. Some methods introduce multi-scale pyramid structures to enhance feature hierarchy, but lack modeling of the spatial relationship between retinal vascular topology and lesion regions.
[0003] Multimodal images are prone to spatial misalignment due to imaging differences, resulting in noise during feature fusion; fixed receptive field convolutional modules are difficult to adapt to the irregular shape of tumor edges, resulting in incomplete feature extraction; feature fusion often adopts simple stitching strategies, which cannot dynamically balance the weights of spatial and structural features, and has a high dependence on labeled data, resulting in insufficient generalization ability in scenarios with limited labeling.
[0004] Therefore, it is urgent to achieve accurate registration and adaptive feature extraction of multimodal images, construct a dynamically weighted cross-domain fusion mechanism, and improve the model's ability to capture features of irregular lesions and the stability of feature expression under limited labeled data. Summary of the Invention
[0005] This application provides an artificial intelligence-based method and system for extracting features from ophthalmic tumor images to solve the above-mentioned problems.
[0006] On the one hand, this application provides an artificial intelligence-based method for extracting features from ophthalmic tumor images, the method comprising the following steps: Step S1: Acquire multi-source image data of ophthalmic tumors, and generate a standardized image set after preprocessing; Step S2: Construct a multi-scale feature extraction network, and extract basic texture features, edge contour features and deep semantic features of the image in layers through an improved convolutional module; Step S3: Introduce a dual-channel attention mechanism to dynamically weight and fuse the initially extracted features with the retinal vascular topological features to generate a fused feature set; Step S4: Optimize the fused features using a residual enhancement network, strengthen the feature representation of key regions through multi-branch feature aggregation, and output a structured feature map containing multi-dimensional feature information.
[0007] In one implementation of this application, step S1 specifically includes: spatial registration of the acquired ophthalmic tumor multi-source image data, elimination of imaging noise through an adaptive noise reduction algorithm, and unification of image grayscale range using multi-channel contrast calibration technology to generate a standardized image set; wherein, the multimodal image spatial registration adopts an anatomical landmark-guided strategy, first identifying reference points such as the optic disc and fovea of the eyeball, and then achieving sub-pixel-level alignment of different modal images through elastic deformation correction.
[0008] In one implementation of this application, in step S2, the improved convolution module introduces a deformable receptive field mechanism, and the formula for calculating the deformation amplitude of the convolution kernel is:
[0009] in, For pixels The radius of deformation of the convolution kernel at that location; This represents the grayscale gradient magnitude of the pixel. is the pixel grayscale value; k is the scaling factor, ranging from 0.1 to 0.3; The spatial density factor of the tumor region is calculated based on the proportion of tumor pixels in a local region. This is the grayscale attenuation factor, with a value range of 0.02 to 0.05.
[0010] In one implementation of this application, in step S2, the multi-scale feature extraction network adopts a progressive downsampling structure, and the output features of each layer are passed to the subsequent fusion stage through skip connections. The bottom-level features retain twice the spatial resolution, and the high-level features improve semantic capacity through channel expansion.
[0011] In one implementation of this application, the weight adjustment formula for the dual-channel attention mechanism in step S3 is:
[0012] in, Spatial feature weights; These are the topological feature weights; The spatial characteristic response value is calculated using global average pooling. The topological feature response value is calculated by global pooling through the output features of the dynamic graph convolutional network. The dynamic graph convolutional network abstracts the vascular network into a graph structure and captures the global topological associations of the vascular network through a dynamic adjacency matrix.
[0013] In one implementation of this application, step S3 further includes: mining the intrinsic correlation between features of different modalities through a feature interaction module; wherein the feature interaction module adopts a cross-modal feature distillation strategy, forward mapping the detailed features of the high-resolution modality to the semantic space of the low-resolution modality, and backward mapping the structural features of the low-resolution modality to the spatial location of the high-resolution modality; and introducing a feature distribution alignment loss.
[0014] Among them, MMD is the maximum mean difference that minimizes the cross-modal feature distribution difference, thereby reducing modal difference interference.
[0015] In one implementation of this application, in step S4, the residual enhancement network comprises multiple sets of residual blocks. Each residual block compresses feature channels through a 1×1 convolution, extracts details through a 3×3 convolution, and optimizes feature propagation through a shortcut connection with feature selection gating: the input feature F of the residual block... in With output feature F out The selection weights are calculated using a gating function:
[0016] Where [;] represents splicing, W g This is the weight matrix; Final output: Dynamically retain high-value features, alleviate feature degradation, and enhance feature propagation efficiency.
[0017] In one implementation of this application, in step S4, the multi-branch feature aggregation technology adopts a feature pyramid structure to extract lesion-related features from feature maps of different scales, and generates comprehensive features that take into account both details and semantics through bidirectional fusion from bottom to top and from top to bottom.
[0018] In one implementation of this application, in step S4, the model training adopts a semi-supervised learning framework, and the total loss function is:
[0019] in, Cross-entropy loss for labeled data; The feature consistency loss for unlabeled data is calculated using the KL divergence of the predicted probability distribution. For the weighting coefficients, satisfying α is the maximum weight, γ is the decay coefficient, and t is the number of training iterations, used to balance the contributions of the two types of losses.
[0020] On the other hand, this application also provides an artificial intelligence-based ophthalmic tumor image feature extraction system, which applies the aforementioned artificial intelligence-based ophthalmic tumor image feature extraction method. The system includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to complete the aforementioned artificial intelligence-based ophthalmic tumor image feature extraction method.
[0021] This application provides an artificial intelligence-based method and system for extracting features from ophthalmic tumor images, which has the following beneficial effects: (1) By accurately registering multimodal images and adaptive deformation convolution, the integrity of irregular tumor edge features is improved and the fusion noise caused by modal misalignment is reduced; (2) A dual-channel dynamic weighted fusion mechanism is adopted to balance the weights of spatial features and vascular topology features, thereby enhancing the ability to capture cross-domain feature correlations; (3) Combining a semi-supervised learning framework with a residual enhancement network reduces the dependence on labeled data and improves the stability and robustness of feature representation under limited data. Attached Figure Description The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an artificial intelligence-based method for extracting features from ophthalmic tumor images, provided for embodiments of this application; Figure 2 This is a diagram illustrating the components of an artificial intelligence-based ophthalmic tumor image feature extraction system provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] This application provides a method and system for extracting features from ophthalmic tumor images based on artificial intelligence. The technical solution proposed in this application will be described in detail below with reference to the accompanying drawings.
[0024] Figure 1This is a flowchart illustrating an artificial intelligence-based method for extracting features from ophthalmic tumor images, provided as an embodiment of this application. Figure 1 As shown, the method mainly includes the following steps: Step S1: Acquire multi-source image data of ophthalmic tumors and generate a standardized image set after preprocessing.
[0025] In this embodiment, spatial registration is performed on the acquired ophthalmic tumor multi-source image data, imaging noise is eliminated by an adaptive noise reduction algorithm, and a multi-channel contrast calibration technique is used to unify the image grayscale range to generate a standardized image set.
[0026] Specifically, spatial registration is performed on the acquired multi-source ophthalmic tumor image data. Adaptive denoising algorithms are used to eliminate imaging noise, and multi-channel contrast calibration technology is employed to unify the image grayscale range, generating a standardized image set. The multimodal image spatial registration uses an anatomical landmark-guided strategy, first identifying reference points such as the optic disc and fovea. Using this strategy, a pre-trained deep learning detection model automatically locates reference points with stable anatomical features, such as the optic disc and fovea. Based on the spatial coordinates of these points, an elastic deformation correction algorithm is used to finely adjust different modal images, ultimately achieving sub-pixel alignment to ensure the tumor region corresponds in position across the multi-source images. Elastic deformation correction is then used to achieve sub-pixel alignment of different modal images. An adaptive denoising algorithm eliminates noise by dynamically adjusting filtering parameters based on different imaging noise characteristics (such as speckle noise in OCT and Gaussian noise in photographs), suppressing noise while preserving lesion details. Finally, multi-channel contrast calibration technology is adopted to unify the grayscale range of the image through a multi-modal grayscale mapping mechanism, so that the image features generated by different devices are comparable and a standardized image set is formed.
[0027] Step S2: Construct a multi-scale feature extraction network, and extract basic texture features, edge contour features and deep semantic features of the image in layers through an improved convolutional module.
[0028] In this embodiment, when constructing a multi-scale feature extraction network, a progressive downsampling structure is adopted to achieve hierarchical feature extraction: the bottom layer retains pixel-level details through twice the spatial resolution and extracts basic texture features with the help of an improved convolutional module. This module introduces a deformable receptive field mechanism. The middle layer transmits the coordinate information of the bottom layer through skip connections, focuses on edge contour features, and captures the spatial relationship between the tumor and blood vessels. The high layer improves semantic capacity through channel expansion, downsamples to 128×128 pixels, and extracts deep semantic features such as the relative position of the tumor and anatomical structures such as the macula. The features of each layer work together to enhance the multi-dimensional characterization of the lesion.
[0029] Furthermore, the improved convolution module introduces a deformable receptive field mechanism, and the formula for calculating the deformation amplitude of the convolution kernel is:
[0030] in, For pixels The radius of deformation of the convolution kernel at that location; This represents the grayscale gradient magnitude of the pixel. is the pixel grayscale value; k is the scaling factor, ranging from 0.1 to 0.3; The spatial density factor of the tumor region is calculated based on the proportion of tumor pixels in a local region. This is the grayscale attenuation factor, with a value range of 0.02 to 0.05. D(x i ,y j ) When the value is ≥0.6 (tumor-dense area), the deformation amplitude converges to 60%-70% of the original amplitude to avoid feature overlap caused by excessive deformation; D (x i ,y j ) When the value is ≤0.3 (sparse area at the tumor edge), the deformation amplitude expands to 1.2-1.5 times the original amplitude, capturing jagged details at the edge more accurately. Testing showed that this optimization improved the integrity of tumor edge feature extraction by 15% and reduced feature redundancy in dense areas by 30%.
[0031] It should be noted that the multi-scale feature extraction network adopts a progressive downsampling structure, gradually reducing the resolution layer by layer, and capturing features from details to semantics from the bottom layer to the top layer. The output features of each layer are passed to the subsequent fusion stage through skip connections. The bottom layer features retain twice the spatial resolution and use an improved convolutional module with deformable receptive fields to extract basic texture and edge details. The deformation amplitude of the convolutional kernel is dynamically adjusted according to the pixel gray-level gradient. The middle layer is downsampled to 256×256 pixels to focus on the spatial relationship features between tumors and blood vessels. The higher layers are further downsampled to 128×128 pixels, and the semantic capacity is improved by channel expansion (such as increasing from 128 dimensions to 256 dimensions) to capture information such as the relative position of tumors and anatomical landmarks. The features of each layer are directly passed to the fusion stage through skip connections, which not only preserves the high-resolution details of the bottom layer, but also allows the high-level semantic features to work together with the bottom-level details, strengthening the feature complementarity.
[0032] In the progressive downsampling structure, a cross-layer attention gating mechanism is added after the feature output of each layer: the high-resolution features at the bottom layer (e.g., 512×512) and the low-resolution features at the top layer (e.g., 128×128) are connected through an attention weight matrix. (l is the lower-level index, m is the higher-level index) Interaction, the calculation formula is:
[0033] in, It is the ReLU activation function. For element-wise multiplication, F l For low-level features, F m High-level features. Mechanism of action: High-level semantic features (such as the positional relationship between the tumor and the macula) are dynamically filtered by gating weights to select low-level detailed features (such as edge texture) and suppress background information unrelated to the lesion (such as normal retinal texture). Experiments show that this mechanism improves multi-scale feature consistency by 20% and reduces the interference of irrelevant details on fusion by 25%.
[0034] Step S3: Introduce a dual-channel attention mechanism to dynamically weight and fuse the initially extracted features with the retinal vascular topological features to generate a fused feature set.
[0035] In this embodiment, a dual-channel attention mechanism is introduced to dynamically weight and fuse the initially extracted spatial features with retinal vascular topological features. A feature interaction module is used to mine the intrinsic correlations between features from different modalities, generating a fused feature set. Simultaneously, the feature interaction module employs a cross-modal feature distillation strategy, mapping detailed features from high-resolution modalities (such as fundus images) to the semantic space of low-resolution modalities (such as OCT), reducing modal differences through feature consistency constraints. The final fused feature set preserves both the spatial details of the tumor margins and strengthens vascular topological correlations, effectively reducing modal misalignment noise and improving the ability to capture cross-domain feature correlations.
[0036] In this embodiment, the weight adjustment formula for the dual-channel attention mechanism is:
[0037] in, Spatial feature weights; These are the topological feature weights; The spatial characteristic response value is calculated using global average pooling. The topological characteristic response value is calculated using the degree centrality of graph nodes.
[0038] Furthermore, the intrinsic correlation between features of different modalities is explored through the feature interaction module. The feature interaction module adopts a cross-modal feature distillation strategy to map the detailed features of high-resolution modalities to the semantic feature space of low-resolution modalities, and reduces the interference of modal differences through feature consistency constraints.
[0039] Specifically, the feature interaction module achieves deep association of multimodal features through cross-modal feature distillation: detailed features of high-resolution modalities (such as color fundus photographs, 4096×3072 pixels) (e.g., 0.1mm-level micro-protrusions on the surface of tumors) are projected onto the semantic feature space of low-resolution modalities (such as OCT images, 5μm horizontal resolution) through a feature mapping network. Simultaneously, feature consistency constraints are introduced, and the distribution differences of features across different modalities (e.g., the probability distribution deviation between high-resolution texture features and low-resolution structural features) are calculated using KL divergence, dynamically adjusting the mapping weights. This is combined with the dynamic weights (α) output by the dual-channel attention mechanism. s With α t The resulting fusion feature set retains the spatial details of the high-resolution modality while integrating the deep structural information of the low-resolution modality, effectively eliminating feature noise caused by differences in multimodal imaging and enhancing the expression of the correlation features between tumor regions and vascular topology.
[0040] Furthermore, retinal vascular topology feature extraction can be upgraded to Dynamic Graph Convolutional Network (DGCN) modeling: the vascular network is abstracted into a graph structure: nodes are defined as vascular branch points / intersection points (including attributes such as diameter and blood flow velocity; FFA data can provide blood flow information), and edges are defined as vascular segments (including length and curvature attributes). Dynamic graph convolutional layers through Update features ( It is a dynamic adjacency matrix that is updated in real time according to node characteristics. (As learnable weights), it captures abnormal tortuosity (curvature > 0.8 mm) of blood vessels surrounding the tumor. -¹ Global topological associations such as branch density (>5 branches / mm²) and topological feature response value T f The DGCN output features are calculated using global max pooling, which more accurately reflects the structural relationship between blood vessels and tumors than traditional graph node degree centrality, improving the correlation index by 25%.
[0041] The feature interaction module can also employ a bidirectional distillation strategy: Forward distillation: Projecting detailed features (such as 0.1mm-level edge protrusions) from high-resolution modalities (e.g., fundus images) to the semantic space of low-resolution modalities (e.g., OCT) via a mapping network; Backward distillation: Mapping structural features (e.g., tumor invasion depth) from low-resolution modalities to the spatial location of high-resolution modalities, correcting modal misalignment; Introducing feature distribution alignment loss:
[0042] The feature distribution bias is quantified by calculating the Gaussian kernel function difference, and the mapping weights are dynamically adjusted. This optimization reduces multimodal fusion noise and improves feature consistency. Step S4: The fused features are optimized using a residual enhancement network. Key region feature representation is strengthened through multi-branch feature aggregation, and a structured feature map containing multi-dimensional feature information is output.
[0043] In this embodiment, the residual enhancement network comprises multiple sets of residual blocks. Each residual block compresses feature channels through a 1×1 convolution, then extracts details through a 3×3 convolution, and finally alleviates feature degradation and enhances feature propagation efficiency through shortcut connections. The gating function is as follows:
[0044] The Sigmoid activation function is used to dynamically generate 0-1 weights, assigning high weights (g>0.7) to tumor edge features (gradient value>0.6) and low weights (g<0.3) to background noise (gradient value<0.2). Output characteristics: This improves the efficiency of effective feature propagation and reduces feature degradation in deep networks.
[0045] It should be noted that the residual enhancement network consists of 6-8 sets of cascaded residual blocks. Each set of residual blocks enhances feature representation through a three-step process of "compression-extraction-fusion". First, 1×1 convolution is used for channel dimensionality reduction, for example, compressing the 256-channel features output from the upper layer to 64 channels, reducing the computational cost by 50% while retaining core information. Next, 3×3 convolution is used to extract local details. The convolution kernel dynamically adjusts the receptive field based on the pixel gray-level gradient, accurately capturing the subtle relationship between tumor edges and vascular branches. Finally, shortcut connections are used to achieve cross-layer feature propagation—features with matching channel numbers are directly superimposed using an identity mapping, while those that do not match are projected to the target dimension through 1×1 convolution, effectively alleviating the feature degradation problem of deep networks and improving feature propagation efficiency by more than 40%. The features optimized by the residual blocks are input into the feature pyramid structure. Through bidirectional fusion of bottom-up transmission of detailed features and top-down supplementation of semantic information, a comprehensive feature that takes into account the relationship between tumor micro-texture and macro-structure is finally generated. This provides a more stable feature input for the loss function under the semi-supervised learning framework and enhances the feature robustness under limited data.
[0046] Furthermore, the multi-branch feature aggregation technology adopts a feature pyramid structure to extract lesion-related features from feature maps at different scales. Through bidirectional fusion from bottom to top and top to bottom, it generates comprehensive features that take into account both details and semantics.
[0047] Specifically, the feature pyramid structure of the multi-branch feature aggregation technology is divided into multiple levels according to resolution. The bottom layer corresponds to a high-resolution feature map of 512×512 pixels, which retains detailed features such as small bulges on the tumor surface (diameter <0.1mm) and edge texture roughness; the middle layer is 256×256 pixels, which focuses on local correlation features such as the spatial distance between the tumor and retinal vessels (such as the shortest distance to the three main vessels being 0.8-2.5mm); the top layer is 128×128 pixels, which contains deep semantic information such as the relative position of the tumor and the fovea of the macula (distance 2.3mm).
[0048] In the bottom-up process, each layer's features are optimized using residual blocks and then passed to the upper layer via lateral connections. Simultaneously, 1×1 convolutions are used to unify the channel count to 128 dimensions for feature alignment. In the top-down process, high-level semantic features are upsampled (e.g., using bilinear interpolation) and superimposed with corresponding mid-level features to supplement anatomical structural association information. Bidirectional fusion integrates multi-scale features through skip connections, preserving over 80% of the detailed texture of the lower layers while incorporating high-level semantic weights (approximately 30%). The resulting comprehensive feature encompasses 28 quantized parameters, accurately characterizing tumor morphology and vascular topology, providing multi-dimensional information support for structured feature map output, and improving the accuracy of feature consistency loss (KL divergence calculation) in subsequent semi-supervised learning.
[0049] Furthermore, the model training employs a semi-supervised learning framework, with the total loss function being:
[0050] in, Cross-entropy loss for labeled data; The feature consistency loss for unlabeled data is calculated using the KL divergence of the predicted probability distribution. This is a weighting coefficient, ranging from 0.3 to 0.7, used to balance the contributions of the two types of losses.
[0051] In this application, the weight coefficient β in the total loss function can be adaptively adjusted: in the early stage of training (t<500 iterations): β≈0.3, focusing on feature consistency learning of unlabeled data with Lunsup accounting for 70%, enhancing the stability of feature distribution; in the later stage of training (t>2000 iterations): β≈0.7, focusing on supervised learning of labeled data (Lsup accounting for 70%), improving classification accuracy; this mechanism improves the stability of feature representation under limited labeled data (less than 200 cases) and accelerates the convergence speed.
[0052] The above describes an artificial intelligence-based ophthalmic tumor image feature extraction system provided in this application. Based on the same inventive concept, this application also provides an artificial intelligence-based ophthalmic tumor image feature extraction system. Figure 2A diagram illustrating the components of an artificial intelligence-based ophthalmic tumor image feature extraction system provided in this application embodiment is shown below. Figure 2 As shown, the system mainly includes: at least one processor 201; and a memory 202 communicatively connected to the at least one processor; wherein the memory 202 stores instructions that can be executed by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to complete the aforementioned artificial intelligence-based ophthalmic tumor image feature extraction method.
[0053] An example of a specific application scenario in this application is as follows.
[0054] To improve the efficiency of feature analysis of ophthalmic tumor images, a provincial ophthalmic medical imaging center introduced an AI-based ophthalmic tumor image feature extraction system for automated feature extraction and structured analysis of multimodal fundus images. The center routinely receives referrals from primary hospitals, involving multi-source image data such as fundus color photographs, optical coherence tomography (OCT), and fluorescence angiography (FFA). Traditional manual analysis suffers from problems such as feature omissions and low efficiency. This system, through standardized processing, multi-scale feature extraction, and dynamic fusion technology, provides radiologists with quantitative feature maps, assisting in the study of the morphological and pathological correlation of tumor regions.
[0055] First, multi-source image data acquisition and standardization processing are performed. The specific steps are as follows: Data collection scope: Multi-source image data of 100 suspected ophthalmic tumor cases from January to June 2024 were selected, including: fundus color photographs: resolution 4096×3072 pixels, covering 80° field of view of the posterior pole of the fundus, including surface texture information of the suspected tumor area; OCT images: horizontal resolution 5μm, vertical resolution 3μm, to obtain the cross-sectional structure of the tumor area, a total of 300 slices of continuous scanning; FFA images: dynamic sequence 1-30 minutes after contrast agent injection, resolution 2048×2048 pixels, recording the perfusion characteristics of blood vessels around the tumor.
[0056] The standardized preprocessing process is as follows: (1) Spatial registration: The system automatically identifies fundus anatomical landmarks—optic disc (diameter 1.7mm ± 0.2mm) and fovea centralis (3.5mm ± 0.3mm from optic disc), and achieves multimodal image alignment through elastic deformation correction algorithm. For example, the registration error between FFA image and OCT cross-sectional image is controlled within 0.05 pixels to ensure that the tumor area corresponds in different modalities. (2) Adaptive noise reduction: For speckle noise in OCT image, wavelet domain adaptive threshold algorithm is used to dynamically adjust the noise reduction intensity according to pixel gray gradient. The signal-to-noise ratio (SNR) of the processed image is increased from 22dB to 34dB. For Gaussian noise in fundus color photo, edge information is preserved through non-local mean filtering, which improves the noise suppression rate. (3) Contrast calibration: Using multi-channel contrast calibration technology, the gray difference between blood vessels and background in FFA images is expanded from 15±3 to 45±5, the local contrast of fundus photos is improved, and the gray range of images under different devices and shooting conditions is uniform (0-255).
[0057] The multi-scale feature extraction process is as follows: An improved convolutional module is used to dynamically adjust the receptive field. The multi-scale feature extraction network constructed by the system adopts an improved convolutional module, which adapts to the irregular shape of the tumor edge through a deformable receptive field mechanism. Taking a color fundus photograph of a certain case as an example, the edge pixels of the tumor region... grayscale gradient Given a grayscale value I=160, according to the formula for the deformation amplitude of the convolution kernel:
[0058] Taking a color fundus photograph of a certain case as an example, the edge pixels of the tumor area grayscale gradient With a grayscale value of I=160, the local tumor pixel ratio calculated using a 5×5 window is 45%, which is... =0.45; Taking the scaling factor k=0.25 and the grayscale attenuation factor λ=0.03, substituting them into the formula, we get: =0.25×0.75×exp(-0.03×160)×0.45≈0.25×0.75×0.0082×0.45≈0.0007mm The calculation results reflect the influence of the spatial density of the tumor region: compared to the dense tumor area (D=0.8), ), the sparse region at the edge (when D=0.3, The improved convolutional kernel exhibits greater deformation, enabling more precise wrapping of jagged edges; while deformation converges in dense areas, avoiding feature overlap. The improved convolutional kernel can adaptively fit the tumor's "dense in the center and sparse at the edges" morphological characteristics, improving the edge feature capture completeness by 15% compared to traditional fixed convolutional kernels.
[0059] The results of layered feature extraction are as follows: (1) Bottom layer features (2 times spatial resolution): retaining the resolution of 512×512 pixels, the texture features of the tumor surface are extracted, including: texture roughness: calculated by gray-level co-occurrence matrix, the contrast value of the tumor area is 85±12, which is significantly higher than that of the surrounding normal tissue (32±5); micro-protrusion structure: 23 protrusions with a diameter <0.1mm were detected, distributed in the 1 / 3 area of the tumor edge. (2) Middle layer features (channels extended to 128 dimensions): by progressive downsampling to 256×256 pixels, the edge contour features of the tumor are focused, including: contour irregularity index: calculated based on edge curvature, the value is 1.7 (normal tissue <1.0); spatial distance with retinal vessels: through the bottom layer coordinate information transmitted by skip connections, the shortest distances between the tumor and the three main vessels are determined to be 0.8mm, 1.2mm and 2.5mm respectively. (3) High-level features (channels extended to 256 dimensions): downsampled to 128×128 pixels, extract deep semantic features, including the gray distribution pattern of the tumor area (mean 120±8), the relative position with the fovea of the macula (distance 2.3mm), and other structured information.
[0060] Then, dual-channel attention fusion and feature optimization are performed. First, retinal vascular topological features are extracted. The system constructs the retinal vascular network topology using graph theory algorithms, treating vascular branch points and intersection points as graph nodes, and calculating the node degree centrality (reflecting the density of blood vessels). In a certain case, the mean degree centrality of blood vessels around the tumor area was 0.62, significantly higher than that of the normal area (0.35), indicating abnormal vascular proliferation.
[0061] Furthermore, the dynamic weighted fusion mechanism is as follows: a dual-channel attention mechanism is used to fuse spatial features and topological features. Wherein, the spatial feature response value S... f The global average pooling calculation yields a value of 0.72, and the topological characteristic response value T is... f The node degree centrality is calculated to be 0.58. According to the weight adjustment formula:
[0062] Calculated spatial feature weights Topological feature weights The fused feature set retains the spatial details of the tumor margins and enhances the correlation between vascular topology and the tumor region, thereby improving the feature signal-to-noise ratio.
[0063] Further, the residual enhancement and multi-branch aggregation process is as follows: (1) Residual enhancement network: It contains 6 groups of residual blocks. Each group compresses the fused feature channels from 512 to 128 through 1×1 convolution, and then extracts local details through 3×3 convolution. Shortcut connection reduces the feature degradation rate by 45% to ensure the effective propagation of deep features. (2) Multi-branch feature aggregation: It adopts a feature pyramid structure to pass low-level details (such as 0.1mm-level micro-protrusions) from bottom to top and supplements high-level semantics (such as the positional relationship between tumors and macula) from top to bottom. The final comprehensive feature map contains 128 quantized feature dimensions, covering morphological, texture, vascular association and other multi-dimensional information.
[0064] Semi-supervised learning training and system performance validation are detailed below: For model training, the system employs a semi-supervised learning framework. The training dataset includes 300 labeled images (with tumor region boundaries marked) and 800 unlabeled images. The overall loss function is set as follows:
[0065] Where β=0.5 (balancing the contributions of labeled and unlabeled data), and the cross-entropy loss of labeled data. The feature consistency loss (KL divergence) for unlabeled data converges to 0.13. The accuracy of the model in extracting features from the tumor region stabilized at 0.09. During training, the model's accuracy in extracting features from the tumor region increased with the number of iterations, and finally the feature matching accuracy on the test set (100 cases) met the preset requirements.
[0066] Ultimately, the structured feature map output by the system contains three types of information: a quantitative index table, a visualized feature map, and a feature correlation analysis. The quantitative index table records 28 quantifiable parameters, including the maximum tumor diameter, volume, edge irregularity index, and vascular correlation. The visualized feature map overlays thermal annotations of tumor region boundaries, areas with high texture differences, and abnormal vascular branch points. The feature correlation analysis uses a heatmap to demonstrate the correlation between tumor texture features and vascular topological features (e.g., r=0.72, indicating a strong correlation between texture roughness and angiogenesis).
[0067] This application's embodiments verify the effectiveness of an AI-based ophthalmic tumor image feature extraction system in multimodal fundus image analysis. Through standardized preprocessing, deformable convolutional feature extraction, dual-channel dynamic fusion, and semi-supervised learning techniques, the system achieves efficient capture and structured output of multidimensional features of tumor regions. This provides a reliable feature analysis tool for ophthalmic imaging research, standardized feature data for radiologists conducting tumor morphology studies, assists in exploring the correlation between "image features and pathological indicators," and promotes the quantitative research progress of ophthalmic tumor images.
[0068] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0069] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0070] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for extracting features from ophthalmic tumor images based on artificial intelligence, characterized in that, The method includes the following steps: Step S1: Acquire multi-source image data of ophthalmic tumors, and generate a standardized image set after preprocessing; Step S2: Construct a multi-scale feature extraction network, and extract basic texture features, edge contour features and deep semantic features of the image in layers through an improved convolutional module; Step S3: Introduce a dual-channel attention mechanism to dynamically weight and fuse the initially extracted features with the retinal vascular topological features to generate a fused feature set; Step S4: Optimize the fused features using a residual enhancement network, strengthen the feature representation of key regions through multi-branch feature aggregation, and output a structured feature map containing multi-dimensional feature information.
2. The method for extracting ophthalmic tumor image features based on artificial intelligence according to claim 1, characterized in that, Step S1 specifically includes: spatial registration of the acquired ophthalmic tumor multi-source image data, elimination of imaging noise through an adaptive noise reduction algorithm, and unification of image grayscale range using multi-channel contrast calibration technology to generate a standardized image set; wherein, the multimodal image spatial registration adopts an anatomical landmark-guided strategy, first identifying the optic disc and fovea of the eyeball, and then achieving sub-pixel alignment of different modal images through elastic deformation correction.
3. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S2, the improved convolution module introduces a deformable receptive field mechanism, and the formula for calculating the deformation amplitude of the convolution kernel is: in, For pixels The radius of deformation of the convolution kernel at that location; This represents the grayscale gradient magnitude of the pixel. is the pixel grayscale value; k is the scaling factor, ranging from 0.1 to 0.3; The spatial density factor of the tumor region is calculated based on the proportion of tumor pixels in a local region. This is the grayscale attenuation factor, with a value range of 0.02 to 0.
05.
4. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S2, the multi-scale feature extraction network adopts a progressive downsampling structure. The output features of each layer are passed to the subsequent fusion stage through skip connections. The bottom-level features retain twice the spatial resolution, and the high-level features improve semantic capacity through channel expansion.
5. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S3, the weight adjustment formula for the dual-channel attention mechanism is: in, Spatial feature weights; These are the topological feature weights; The spatial characteristic response value is calculated using global average pooling. The topological feature response value is calculated by global pooling through the output features of the dynamic graph convolutional network. The dynamic graph convolutional network abstracts the vascular network into a graph structure and captures the global topological associations of the vascular network through a dynamic adjacency matrix.
6. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, Step S3 further includes: mining the intrinsic correlations between features of different modalities through a feature interaction module; wherein, the feature interaction module adopts a cross-modal feature distillation strategy, forward mapping the detailed features of the high-resolution modality to the semantic space of the low-resolution modality, and backward mapping the structural features of the low-resolution modality to the spatial location of the high-resolution modality; and introducing a feature distribution alignment loss: Among them, MMD is the maximum mean difference that minimizes the cross-modal feature distribution difference, thereby reducing modal difference interference.
7. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S4, the residual enhancement network contains multiple sets of residual blocks. Each residual block compresses the feature channels through a 1×1 convolution, then extracts details through a 3×3 convolution, and finally optimizes feature propagation through a shortcut connection with feature selection gating: the input feature F of the residual block... in With output feature F out The selection weights are calculated using a gating function: in[; [W is for splicing] g This is the weight matrix; Final output: Dynamically retain high-value features, alleviate feature degradation, and enhance feature propagation efficiency.
8. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S4, the multi-branch feature aggregation technology adopts a feature pyramid structure to extract lesion-related features from feature maps at different scales. Through bidirectional fusion from bottom to top and top to bottom, it generates comprehensive features that take into account both details and semantics.
9. The method for extracting features from ophthalmic tumor images based on artificial intelligence according to claim 1, characterized in that, In step S4, the model training adopts a semi-supervised learning framework, and the total loss function is: in, Cross-entropy loss for labeled data; The feature consistency loss for unlabeled data is calculated using the KL divergence of the predicted probability distribution. Let be the weighting coefficient, satisfying α is the maximum weight, γ is the decay coefficient, and t is the number of training iterations, used to balance the contributions of the two types of losses.
10. An artificial intelligence-based ophthalmic tumor image feature extraction system, employing the artificial intelligence-based ophthalmic tumor image feature extraction method according to any one of claims 1-9, characterized in that, The system includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the artificial intelligence-based ophthalmic tumor image feature extraction method according to any one of claims 1-9.