Medical image processing method, apparatus, device, storage medium, and program product
By using an image generation model that extracts global and local features, the problem of complex operation of OCT devices is solved, enabling fast and convenient acquisition of retinal thickness images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-09-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing OCT equipment is complex to operate, makes it difficult to quickly acquire retinal thickness images, and has poor acquisition efficiency.
By acquiring global eye images, performing global and local feature extraction, and using an image generation model to generate a retinal thickness map, the use of complex OCT equipment is avoided.
It improves the efficiency of acquiring retinal thickness maps, simplifies the operation process, and reduces acquisition costs.
Smart Images

Figure CN121169869B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a medical image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the continuous development of image processing technology and the increasing health awareness of users, users expect to acquire clearer and more accurate ophthalmic medical images through better image processing technology. Members of medical institutions also expect to quickly obtain accurate and clear ophthalmic medical images through better image processing technology.
[0003] In existing technologies, retinal thickness images are typically acquired using specific OCT (Optical Coherence Tomography) equipment. However, OCT equipment is complex to operate and has difficulty acquiring retinal thickness images quickly, resulting in poor acquisition efficiency. Summary of the Invention
[0004] Therefore, it is necessary to provide a medical image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the efficiency of acquiring retinal images, in order to address the aforementioned technical problems.
[0005] In a first aspect, this application provides a medical image processing method, comprising: acquiring a global eye image and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0006] In one embodiment, global feature extraction is performed on the global eye image to obtain a first feature map, and local feature extraction is performed on the local eye image to obtain a second feature map. This includes: inputting the global eye image into the global encoder in the image generation model to obtain the first feature map, and inputting the local eye image into the local encoder in the image generation model to obtain the second feature map.
[0007] In one embodiment, the global eye image is input into the global encoder in the image generation model to obtain a first feature map, including: performing feature extraction and feature enhancement processing on the global eye image through the global encoder to obtain the first feature map output by the global encoder; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and retinal global feature map of the global eye image.
[0008] In one embodiment, the local eye image is input into the local encoder in the image generation model to obtain a second feature map, including: cropping the local eye image and the target image corresponding to the target region through the local encoder, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; and fusing the features of each feature map through the local encoder to obtain a second feature map; wherein the second feature map includes the macular region feature map of the local eye image.
[0009] In one embodiment, generating a retinal thickness map based on a first feature map and a second feature map includes: inputting the first feature map and the second feature map into an image generation layer included in an image generation model, and generating a retinal thickness map based on the first feature map, the second feature map, and a preset noise map through the image generation layer.
[0010] In one embodiment, the training process of the image generation layer of the image generation model includes: obtaining a training sample set, which includes sample images and labeled retinal thickness maps. The sample images include a first sample feature map, a second sample feature map, and a noisy retinal thickness map, wherein the noisy retinal thickness map is obtained by adding noise to the labeled retinal thickness map; and training the image generation layer of the initial image generation model based on the sample images and the labeled retinal thickness map to obtain the image generation layer of the image generation model.
[0011] Secondly, this application also provides a medical image processing apparatus, comprising: an image acquisition module for acquiring a global eye image and cropping a local eye image from the global eye image; a feature processing module for performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; and an image generation module for generating a retinal thickness map based on the first feature map and the second feature map.
[0012] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in the first aspect.
[0013] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0014] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.
[0015] The aforementioned medical image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a global eye image, extract a local eye image from the global eye image, perform global feature extraction on the global eye image to obtain a first feature map of all regions in the global eye image, and perform local feature extraction on the local eye image to obtain a second feature map of the target image in the local eye image. By performing global and local image feature extraction in parallel, image processing efficiency is improved. A retinal thickness map is generated based on the first and second feature maps, avoiding the need to operate complex OCT equipment to obtain a retinal thickness map. Instead, a retinal thickness map is generated through image processing based on an easily accessible eye image, thus improving the efficiency of retinal thickness map acquisition. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is an application environment diagram of a medical image processing method in one embodiment;
[0018] Figure 2 This is a flowchart illustrating a medical image processing method in one embodiment;
[0019] Figure 3 This is a flowchart illustrating step 202 in one embodiment;
[0020] Figure 4 This is a schematic diagram of the process of obtaining the first feature map in one embodiment;
[0021] Figure 5 This is a schematic diagram of the process for obtaining the second feature map in one embodiment;
[0022] Figure 6 This is a flowchart illustrating step 203 in one embodiment;
[0023] Figure 7 This is a schematic diagram of the training process of the image generation layer of an image generation model in one embodiment.
[0024] Figure 8 This is a flowchart illustrating a medical image processing method in another embodiment;
[0025] Figure 9 This is a structural block diagram of a medical image processing device in one embodiment;
[0026] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0028] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0029] The medical image processing method provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown includes at least terminal 101 and may also include server 102.
[0030] The terminal 101 acquires a global eye image and a local eye image input by the user. It performs global feature extraction on the global eye image to obtain a first feature map, and performs local feature extraction on the local eye image to obtain a second feature map. Based on the first and second feature maps, it generates a retinal thickness map and displays the retinal thickness map to the user. The terminal 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses.
[0031] It should be noted that the process of feature extraction and retinal image generation can be performed by either terminal 101 or server 102. Terminal 101 can acquire the global and local eye images input by the user, and send the global and local eye images to server 102 via its medical image processing interface. Server 102 performs global feature extraction on the global eye image to obtain a first feature map, and performs local feature extraction on the local eye image to obtain a second feature map. Based on the first and second feature maps, it generates a retinal thickness map and returns it to terminal 101, which then displays the retinal thickness map to the user. Server 102 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. Terminal 101 communicates with server 102 via a network.
[0032] In one exemplary embodiment, such as Figure 2 As shown, a medical image processing method is provided, which can be applied to... Figure 1 The following steps are used as an example of the terminal in the example, including steps 201 to 203.
[0033] Step 201: Obtain the global eye image and extract a local eye image from the global eye image.
[0034] In this application, an eye image refers to a color fundus image, which is an image of the posterior segment of the human eye captured by an image acquisition device, including structures such as the retina, optic disc (optic nerve head), macula, and choroid. Furthermore, a global eye image refers to a color fundus image of the entire area of the eye, such as a color fundus image of the area including the posterior pole, middle periphery, and distal periphery; a local eye image refers to a color fundus image of a part of the eye, such as a color fundus image including the macula.
[0035] During implementation, the terminal acquires a global eye image and a local eye image based on the interactive page. The local eye image can be cropped from the global eye image. During execution, the user can pre-crop a local eye image from the global eye image and simultaneously input both the global and local eye images through the interactive page; that is, the terminal acquires the global and local eye images input by the user based on the interactive page.
[0036] In addition, users can also input a global eye image simply through the interactive page. After receiving the global eye image, the terminal will extract a local eye image corresponding to the central region of the eye. For example, the terminal can extract a local eye image corresponding to the macular region from the global eye image.
[0037] Furthermore, the size of the global eye image and / or local eye image may not conform to the standard. The terminal can perform image segmentation on the global eye image and / or local eye image to crop out the redundant image. The terminal can detect the central region of the global eye image and / or local eye image through an edge detection algorithm, and crop the edge region based on the preset size and the central region. The central region can be the pupil region.
[0038] The specific standard size can be obtained in advance by registering the eye image and the corresponding infrared fundus image (IR-FP image). During the registration process, the retinal vessel binary map is extracted from the color fundus image (denoted as C-FPglobal) and the corresponding IR-FP image using a pre-trained Unet as a common feature. Next, the AKAZE feature point detection operator is used to extract key points on the vessel map, and the preliminary homography transformation matrix is estimated using the RANSAC algorithm to eliminate outliers in the matching. On this basis, the k-nearest neighbor feature matching strategy is further used to finely match the vessel mesh control points. After rigid registration is completed, non-rigid deformation correction is performed on the C-FP image to fit the IR-FP. Finally, a local color image of the macula corresponding to the OCT scan area is cropped, and the size of this local color image of the macula is used as the standard size.
[0039] In addition, other feature extraction and matching algorithms can be used, as long as they can achieve approximate alignment between C-FP and RTM. If IR-FP images are unavailable, a coordinate mapping between the color image and the template thickness map can be established based on fundus anatomical landmarks (such as the optic disc and macula position) to approximate registration and obtain the standard size.
[0040] In real-world scenarios, color fundus images can be obtained with simple devices, such as basic ophthalmic medical equipment, or even through the image acquisition function of a smart terminal. The acquisition is relatively easy. By performing medical image processing on easily obtained color fundus images, retinal thickness images can be obtained, greatly improving the efficiency of retinal thickness image acquisition.
[0041] Step 202: Perform global feature extraction on the global eye image to obtain the first feature map, and perform local feature extraction on the local eye image to obtain the second feature map.
[0042] During implementation, the terminal can perform feature extraction on both the global eye image and the local eye image in parallel. Global feature extraction is performed on the global eye image to obtain a first feature map, and local feature extraction is performed on the local eye image to obtain a second feature map. The first feature map includes image features from all regions of the global eye image, while the second feature map includes image features from the central region of the eye in the local eye image.
[0043] During execution, the terminal can use a feature extraction algorithm to extract features from the global eye image and the local eye image to obtain a first feature map and a second feature map. Furthermore, the feature extraction results can be enhanced, and the enhanced feature map can be used as the first feature map and the second feature map.
[0044] In addition, feature extraction can also be performed using a pre-trained feature processing model. The global eye image and the local eye image can be input into the pre-trained feature processing model for feature processing to obtain the first feature map and the second feature map output by the feature processing model.
[0045] It should be noted that the feature processing methods for the first feature map and the second feature map can be the same or different, which will not be elaborated here.
[0046] It should also be noted that, due to limitations in the computing power of the terminal, feature extraction operations for global and local eye images can also be performed through the backend server. The terminal can call the server's interface to input global and local eye images, and the server will extract features from the global and local eye images to obtain the first feature map and the second feature map.
[0047] Step 203: Generate a retinal thickness map based on the first feature map and the second feature map.
[0048] In this application, the retinal thickness map is used to display a visual image showing the thickness distribution of the retina at different locations. For this retinal thickness map, warm colors (such as red, orange, and yellow) can be used to represent thicker areas, and cool colors (such as green and blue) can be used to represent thinner or normal-thickness areas. Alternatively, the thickness distribution can be displayed in a circular or square area centered on the macula, which can be divided into multiple quadrants (such as upper, lower, nasal, and temporal) and annular areas (such as inner and outer rings) for quantitative analysis.
[0049] During implementation, the terminal can generate a retinal thickness map based on the feature information of the first feature map and the second feature map, or the terminal can convert a preset image into a retinal thickness map based on the feature information of the first feature map and the second feature map; during execution, a retinal thickness map can be generated based on the feature information of the first feature map and the second feature map and a preset noise map.
[0050] The operation of generating a retinal thickness map based on the first feature map and the second feature map can also be performed by the server. The server can generate a retinal thickness map based on the first feature map and the second feature map and send it to the terminal. The terminal obtains the retinal thickness map returned by the interface call.
[0051] Furthermore, the terminal can also display a retinal thickness map on the page.
[0052] It should be noted that steps 201 to 203 can also be completed on the server. The terminal and the server can cooperate to execute the steps. The terminal can obtain the global eye image and / or local eye image input on the interactive page and transmit the global eye image and / or local eye image to the server. The server can obtain the global eye image and / or local eye image. If the server only obtains the global eye image, it can extract the local eye image corresponding to the central area of the eye from the global eye image. The server can perform global feature extraction on the global eye image to obtain a first feature map and perform local feature extraction on the local eye image to obtain a second feature map. A retinal thickness map is generated based on the first feature map and the second feature map. The server sends the retinal thickness map to the terminal, and the terminal displays the retinal thickness map.
[0053] In the aforementioned medical image processing method, a global eye image is acquired, a local eye image is extracted from the global eye image, global features are extracted from the global eye image to obtain a first feature map of all regions in the global eye image, and local features are extracted from the local eye image to obtain a second feature map of the target image in the local eye image. By performing global and local image feature extraction in parallel, the image processing efficiency is improved. A retinal thickness map is generated based on the first and second feature maps, avoiding the need to operate complex OCT equipment to obtain a retinal thickness map. Instead, a retinal thickness map is generated through image processing based on an easily acquired eye image, thus improving the efficiency of retinal thickness map acquisition.
[0054] Based on the above exemplary embodiment, the following provides a medical image processing method in one or more exemplary embodiments, which is applied to... Figure 1 Taking the terminal in the example, the explanation includes the following content.
[0055] In real-world scenarios, to further improve the efficiency and accuracy of image processing, an overall image generation model can be used for image processing, feature extraction, and image generation. The terminal can obtain global and / or local eye images based on the interactive page, input the global and / or local eye images into the image generation model, and the image generation model performs image segmentation, feature extraction, and image generation operations, outputting a retinal thickness map.
[0056] During feature extraction, the image generation model can perform feature extraction operations on the global eye image and the local eye image in parallel through two encoders; in one optional implementation provided in this application, such as Figure 3 As shown, step 202 includes step 301:
[0057] Step 301: Input the global eye image into the global encoder of the image generation model to obtain the first feature map, and input the local eye image into the local encoder of the image generation model to obtain the second feature map.
[0058] During implementation, the terminal inputs the global eye image and the local eye image into the image generation model. The global eye image is input into the global encoder, and feature processing is performed by the global encoder to obtain the first feature map. The local eye image is input into the local encoder to obtain the second feature map.
[0059] During execution, the terminal can label the global eye image and the local eye image respectively, so that the image generation model can transmit the global eye image and the local eye image to their respective encoders according to the labels, and perform feature processing according to the above operation to obtain the first feature map and the second feature map.
[0060] In addition, the image generation model can also recognize global eye images and local eye images. Based on the recognition results, the global eye image is input into the global encoder, and features are extracted through the global encoder to obtain the first feature map. Based on the recognition results, the local eye image is input into the local encoder, and features are extracted through the local encoder to obtain the second feature map.
[0061] One optional implementation provided in this application improves feature processing efficiency by using a global encoder and a local encoder included in the image generation model to perform feature processing on global and local eye images in parallel. At the same time, the accuracy and reliability of the acquired first and second feature maps are improved by using a pre-trained encoder.
[0062] Furthermore, during the processing of the first feature map, the first feature map can be obtained through feature extraction and feature enhancement to improve its usability; in one optional implementation provided by this application, such as Figure 4 As shown, the process of obtaining the first feature map includes step 401:
[0063] Step 401: Perform feature extraction and feature enhancement processing on the global eye image using a global encoder to obtain the first feature map output by the global encoder.
[0064] In this application, the first feature map is a feature map used to characterize global features in a global eye image, specifically it can be used to characterize image features of the posterior pole, mid-periphery and distal periphery regions; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and global retinal feature map of the global eye image.
[0065] During feature processing, the global encoder can perform multiple rounds of feature extraction on the global eye image and perform feature enhancement processing based on the results of multiple rounds of feature extraction to obtain the first feature map output by the global encoder.
[0066] During execution, the global encoder can use the Visual Transformer backbone (ViT) of the RETFound pre-trained model as the feature extraction network. RETFound is a retinal baseline model trained for fundus diseases, and its ViT encoder has learned rich anatomical and pathological patterns on large-scale fundus data. The model parameters of this ViT backbone can be frozen, and after inputting C-FPglobal, multi-scale high-dimensional feature tokens are extracted from different levels of ViT. A designed adapter module then reorganizes / projects the tokens from the four stages to form a global feature map Fg (approximately H / 4×W / 4×64) with dimensions comparable to local features. This global feature map carries a wide-area retinal anatomical background and extensive lesion context information, such as the optic disc, the course and density of blood vessels in large areas, and the overall state of the peripheral retina. This helps the model understand whether local changes in the macula are part of the global pathology, improving the biological consistency of the predictions.
[0067] In addition, the global encoder can also employ other convolutional neural networks (such as UNet, ResNet+FPN, etc.) or visual Transformer models capable of extracting multi-scale features. Besides the RETFound pre-trained ViT, the global encoder can also use other backbone networks trained on large-scale image data. Even the global branch can take in images of other modalities (such as wide-angle fundus photography) to provide more comprehensive retinal background information.
[0068] One optional implementation provided in this application generates a first feature map through multi-round feature extraction and feature enhancement processing by a global encoder, thereby improving the usability and reliability of the first feature map.
[0069] In the process of acquiring the second feature map, the local encoder can first crop the target image corresponding to the target region, perform multiple rounds of feature extraction on the target image, and then fuse them to obtain the second feature map; in one optional implementation provided by this application, such as Figure 5 As shown, the process of obtaining the second feature map includes steps 501 to 502:
[0070] Step 501: The local eye image is cropped and the target image corresponding to the target region is extracted by the local encoder. Multiple feature extractions are performed on the target image to obtain multiple feature maps.
[0071] In this application, the second feature map is a feature map used to characterize the local features of the target region in a local eye image, specifically it can be used to characterize the image features of the macular region; wherein, the second feature map includes the macular region feature map of the local eye image.
[0072] During implementation, the local encoder identifies and extracts the target image corresponding to the target region from the local eye image, performs multiple rounds of feature extraction on the target image, and obtains multiple feature maps; among them, the target region can be the macula region.
[0073] During execution, the local encoder uses edge detection algorithms to detect regions whose features match those of the target region and extracts the image of that region as the target image. In feature processing, the local encoder can employ a deep network based on SwinTransformer to extract high-resolution local features. SwinTransformer uses a sliding window self-attention mechanism to acquire local details.
[0074] Step 502: The feature maps are fused using a local encoder to obtain a second feature map.
[0075] During implementation, the local encoder performs feature overlay / feature fusion processing on each feature map to obtain a second feature map representing the features of the macular region.
[0076] During execution, the local encoder constructs a hierarchical representation by merging patches layer by layer, and then combines it with the Feature Pyramid Network (FPN) to fuse feature maps of multiple scales. The resulting local feature tensor Fm has a size of about 1 / 4 of the original image (H / 4×W / 4×256), which can characterize the subtle structural changes in the macular region, such as the slight thickness undulations in the fovea and the texture changes caused by local exudation.
[0077] In addition, local encoders can also employ other convolutional neural networks (such as UNet, ResNet+FPN, etc.) or visual Transformer models capable of extracting multi-scale features. Global encoders, besides RETFound pre-trained ViT, can also use other backbone networks trained on large-scale image data. Even the global branch can take in images of other modalities (such as wide-angle fundus photography) to provide more comprehensive retinal background information.
[0078] One optional implementation provided in this application obtains a feature map representing the macular region by performing image segmentation, multi-round feature processing, and feature fusion using a local encoder, thereby improving the usability and accuracy of the generated second feature map and thus enhancing the effectiveness of the generated retinal thickness map.
[0079] In the process of generating a retinal thickness map, the retinal thickness map is also a type of "feature map" used to characterize the thickness of the retina. It can be generated by generating a first feature map and a second feature map. During the generation of the first and second feature maps, some noise may be introduced. The first and second feature maps can be denoised first and second feature maps, and the retinal thickness map can be generated based on the denoised first and second feature maps. In one optional implementation provided in this application, such as... Figure 6 As shown, step 203 includes step 601:
[0080] Step 601: Input the first feature map and the second feature map into the image generation layer of the image generation model, and generate a retinal thickness map based on the first feature map, the second feature map and the preset noise map through the image generation layer.
[0081] During implementation, the terminal can input the first feature map and the second feature map into the image generation layer of the image generation model. The image generation layer uses the feature information of the first feature map and the second feature map as inference guidance information, and performs denoising processing on the noise map based on the feature information to obtain the retinal thickness map. The noise map can be a Gaussian noise map; the preset noise map can be fixed data pre-stored in the image generation layer.
[0082] During execution, retinal thickness maps can also be generated based on an image generation model. The image generation model can input the first feature map generated by the global encoder and the second feature map generated by the local encoder into the image generation layer, and the image generation layer generates a retinal thickness map based on the first feature map, the second feature map and a preset noise map.
[0083] During the retinal thickness map generation process, the prediction result can be generated through a reverse diffusion process, that is, gradually recovering the retinal thickness map from the noise map, using the first feature map and the second feature map as guiding data. Furthermore, optimization can be performed by minimizing the mean square error, thereby ensuring that the difference between the predicted result and the target result is minimized. In the denoising process, multiple denoising steps can be used. After each denoising step, the image is compared with the first and second feature maps, and the next denoising step is performed based on the comparison results, until the feature information of the retinal thickness map matches that of the first and second feature maps.
[0084] For example, in the denoising process, we can start from the random noise map, combine it with the features of the input photo, and iteratively apply a diffusion decoder to generate an RTM. To improve inference efficiency, we introduce the DDIM (Denoising Diffusion Implicit Model) sampling strategy. DDIM is a non-Markovian deterministic sampling method that can generate high-fidelity images with fewer diffusion steps, thereby accelerating the generation speed and making the model potentially meet the needs of real-time clinical applications. Furthermore, the conditional diffusion framework naturally possesses diversity and noise resistance, performing robustly when training on data containing image noise or quality variations, thus improving the robustness of prediction results to changes in photographic quality.
[0085] During the generation of the retinal thickness map, the noise map after denoising can be colored and rendered to obtain a retinal thickness map containing color. Among them, the parts with more noise can be rendered as dark areas, that is, areas with greater retinal thickness, while the parts with less noise can be rendered as light areas, that is, areas with less retinal thickness.
[0086] One optional implementation provided in this application generates a retinal thickness map using a first feature map and a second feature map of a color fundus image, thereby improving the generation efficiency of the retinal thickness map and reducing the acquisition cost of the retinal thickness map.
[0087] In real-world scenarios, the image generation layer of the image generation model needs to be pre-trained. This can be achieved by training the initial image generation model's image generation layer based on a pre-acquired training sample set. In one optional implementation provided in this application, such as... Figure 7 As shown, the training process of the image generation layer of the image generation model includes steps 701 to 702:
[0088] Step 701: Obtain the training sample set.
[0089] During implementation, the model training device acquires a training sample set. This training sample set includes sample images and labeled retinal thickness maps. The sample images include a first sample feature map, a second sample feature map, and a noisy retinal thickness map. The noisy retinal thickness map is obtained by adding noise to the labeled retinal thickness map, and the first and second sample feature maps are obtained by performing feature processing on the labeled retinal thickness map.
[0090] In the process of acquiring noisy retinal thickness maps, a noisy retinal thickness map can be generated by progressively adding noise to the labeled retinal thickness map using conditional diffusion. Optionally, multiple noisy retinal thickness maps can be generated, and the degree of noise addition can vary among the maps.
[0091] During execution, Gaussian noise can be gradually added to the tag retina thickness map according to the preset noise scheduling strategy to obtain thickness maps Mt with different degrees of degradation (t represents the number of time steps / diffusion steps).
[0092] Step 702: Train the image generation layer of the initial image generation model based on the sample image and the labeled retinal thickness map to obtain the image generation layer of the image generation model.
[0093] During implementation, the model training device can input sample images into the image generation layer of the initial image generation model, generate retinal thickness maps through the image generation layer, calculate the model training loss based on the generated retinal thickness map and the labeled retinal thickness map, adjust the model parameters based on the loss value until the loss value converges, and obtain the image generation layer of the image generation model.
[0094] During model parameter tuning, the features (Fm and Fg) extracted by the dual encoders can be input into the diffusion decoder Ed along with the noisy image Mt. The diffusion decoder consists of multiple layers of deformable attention modules, which can align global and local features with the currently generated image content to achieve more refined cross-modal alignment. The decoder learns to predict the thickness image Mt of the previous time step from the noisy thickness image Mt under given conditional features, and gradually denoises to approximate the labeled retinal thickness image. The loss function can be referred to as formula (1):
[0095] Formula (1).
[0096] Furthermore, this model can be extended to other related tasks. For example, through transfer learning adjustments, it can be used to predict changes in retinal layer thickness and segment retinal anatomy using color fundus photographs.
[0097] In one optional implementation provided by this application, during the training process, Gaussian noise is gradually superimposed on the labeled retinal thickness data according to a preset noise scheduler to form noisy retinal thickness data at different time steps. Subsequently, the feature vectors from the dual encoders are concatenated as conditional inputs to guide the diffusion process. During the training phase, the decoder learns how to recover the original mapping information based on the noisy target data and conditional features. By using diffusion training, the image generation model, including the image generation layer, is trained, which improves the effectiveness and accuracy of the trained image generation layer, thereby enhancing the usability of the generated retinal thickness map.
[0098] In one embodiment, see Figure 8 The document illustrates a flowchart of a medical image processing method provided in an embodiment of this application, which can be applied to... Figure 1 In the terminal shown. For example... Figure 8As shown, the medical image processing method may include the following steps:
[0099] Step 801: Obtain the global eye image and extract a local eye image from the global eye image.
[0100] Step 802: Input the global eye image and the local eye image into the image generation model. Perform feature extraction and feature enhancement processing on the global eye image through the global encoder in the image generation model to obtain the first feature map output by the global encoder.
[0101] Optionally, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map, and global retinal feature map of the global eye image.
[0102] Step 803: The local eye image and the target image corresponding to the target region are cropped by the local encoder in the image generation model. Multiple rounds of feature extraction are performed on the target image to obtain multiple feature maps. The feature maps are then fused to obtain the second feature map.
[0103] Optionally, the second feature map includes a macular region feature map of a local eye image.
[0104] Step 804: The image generation model inputs the first feature map and the second feature map into the image generation layer of the image generation model, performs denoising processing on the first feature map and the second feature map through the image generation layer, and generates a retinal thickness map based on the denoised first feature map and the denoised second feature map.
[0105] Step 805: Obtain the retinal thickness map and display it on the screen.
[0106] It should be noted that any one or more of steps 801 to 805, or any combination of steps 201 to 203 provided in the above embodiments, can be selected to form a new implementation method according to the needs of implementation and deployment. Furthermore, any one or more technical features in the technical solution composed of steps 801 to 805 can also be selected to form a new implementation method according to the actual deployment needs, or technical features in one or more optional implementations provided in one or more of the above embodiments can be selected to form a new implementation method. These will not be elaborated on here.
[0107] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0108] This application also provides one or more experimental embodiments. It employs a conditional diffusion model (GLD-RT) combined with multi-scale feature fusion and deformable attention to achieve high-precision thickness reconstruction in various regions, including the fovea (G1), inner ring (G2), and outer ring (G3). Compared with existing methods, the mean thickness error (MAE) is reduced by more than 15%, and the peak signal-to-noise ratio (PSNR) is improved by more than 1.5 dB. It supports near-real-time thickness map generation. Utilizing large-scale pre-trained models such as RETFound, it demonstrates robust performance on different camera devices and population data, exhibiting good clinical applicability. Furthermore, the output format is compatible with OCT thickness maps, allowing doctors to directly observe the thickness distribution. The model only requires a standard fundus camera for deployment, facilitating embedding into existing screening equipment or mobile diagnostic systems, thereby improving the efficiency and confidence of diagnostic decisions.
[0109] Based on the same inventive concept, this application also provides a medical image processing apparatus for implementing the medical image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more medical image processing apparatus embodiments provided below can be found in the limitations of the medical image processing method described above, and will not be repeated here.
[0110] In one exemplary embodiment, such as Figure 9As shown, a medical image processing device is provided, including: an image acquisition module 901, a feature processing module 902, and an image generation module 903, wherein: the image acquisition module 901 is used to acquire a global eye image and extract a local eye image from the global eye image; the feature processing module 902 is used to perform global feature extraction on the global eye image to obtain a first feature map, and to perform local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; the image generation module 903 is used to generate a retinal thickness map based on the first feature map and the second feature map.
[0111] In one embodiment, the feature processing module 902 includes an encoder processing unit, wherein the encoder processing unit is used to input a global eye image into a global encoder in an image generation model to obtain a first feature map, and to input a local eye image into a local encoder in an image generation model to obtain a second feature map.
[0112] In one embodiment, the encoder processing unit includes a first encoder processing unit, wherein: the first encoder processing unit is used to perform feature extraction and feature enhancement processing on the global eye image through a global encoder to obtain a first feature map output by the global encoder; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and retinal global feature map of the global eye image.
[0113] In one embodiment, the encoder processing unit includes a second encoder processing unit and a feature fusion unit, wherein: the second encoder processing unit is used to extract a target image corresponding to a local eye image and a target region through a local encoder, and to perform multiple rounds of feature extraction on the target image to obtain multiple feature maps; the feature fusion unit is used to perform feature fusion on each feature map through a local encoder to obtain a second feature map; wherein the second feature map includes a macular region feature map of the local eye image.
[0114] In one embodiment, the image generation module 903 includes an image generation unit, wherein the image generation unit is used to input a first feature map and a second feature map into an image generation layer included in the image generation model, and generate a retinal thickness map based on the first feature map, the second feature map and a preset noise map through the image generation layer.
[0115] In one embodiment, the apparatus further includes a training sample set acquisition module and a model training module, wherein: the sample set acquisition module is used to acquire a training sample set, the training sample set including sample images and labeled retinal thickness maps, the sample images including a first sample feature map, a second sample feature map and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the labeled retinal thickness map; the model training module is used to train the image generation layer included in the initial image generation model based on the sample images and the labeled retinal thickness map, to obtain the image generation layer included in the image generation model.
[0116] Each module in the aforementioned medical image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0117] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a medical image processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0118] Those skilled in the art will understand that Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0119] In one exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: acquiring a global eye image and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0120] In one embodiment, when the processor executes the computer program, it specifically implements the following steps: inputting the global eye image into the global encoder of the image generation model to obtain a first feature map, and inputting the local eye image into the local encoder of the image generation model to obtain a second feature map.
[0121] In one embodiment, when the processor executes the computer program, it specifically implements the following steps: performing feature extraction and feature enhancement processing on the global eye image through a global encoder to obtain a first feature map output by the global encoder; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and retinal global feature map of the global eye image.
[0122] In one embodiment, when the processor executes the computer program, it specifically implements the following steps: using a local encoder to capture a local eye image and a target image corresponding to the target region, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; using a local encoder to fuse the features of each feature map to obtain a second feature map; wherein, the second feature map includes a macular region feature map of the local eye image.
[0123] In one embodiment, when the processor executes the computer program, it specifically implements the following steps: inputting the first feature map and the second feature map into the image generation layer of the image generation model, and generating a retinal thickness map based on the first feature map, the second feature map and a preset noise map through the image generation layer.
[0124] In one embodiment, when the processor executes the computer program, it further performs the following steps: acquiring a training sample set, the training sample set including sample images and labeled retinal thickness maps, the sample images including a first sample feature map, a second sample feature map and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the labeled retinal thickness map; training the image generation layer included in the initial image generation model based on the sample images and the labeled retinal thickness map, to obtain the image generation layer included in the image generation model.
[0125] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps: acquiring a global eye image and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0126] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: inputting a global eye image into a global encoder in an image generation model to obtain a first feature map, and inputting a local eye image into a local encoder in an image generation model to obtain a second feature map.
[0127] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: performing feature extraction and feature enhancement processing on the global eye image through a global encoder to obtain a first feature map output by the global encoder; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and global retinal feature map of the global eye image.
[0128] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: using a local encoder to extract a local eye image and a target image corresponding to the target region, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; using a local encoder to fuse the features of each feature map to obtain a second feature map; wherein, the second feature map includes a macular region feature map of the local eye image.
[0129] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: inputting the first feature map and the second feature map into the image generation layer of the image generation model, and generating a retinal thickness map based on the first feature map, the second feature map and a preset noise map through the image generation layer.
[0130] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: acquiring a training sample set, the training sample set including sample images and labeled retinal thickness maps, the sample images including a first sample feature map, a second sample feature map and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the labeled retinal thickness map; training the image generation layer included in the initial image generation model based on the sample images and the labeled retinal thickness map, to obtain the image generation layer included in the image generation model.
[0131] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps: acquiring a global eye image and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0132] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: inputting a global eye image into a global encoder in an image generation model to obtain a first feature map, and inputting a local eye image into a local encoder in an image generation model to obtain a second feature map.
[0133] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: performing feature extraction and feature enhancement processing on the global eye image through a global encoder to obtain a first feature map output by the global encoder; wherein, the first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map and global retinal feature map of the global eye image.
[0134] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: using a local encoder to extract a local eye image and a target image corresponding to the target region, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; using a local encoder to fuse the features of each feature map to obtain a second feature map; wherein, the second feature map includes a macular region feature map of the local eye image.
[0135] In one embodiment, when the computer program is executed by the processor, it specifically implements the following steps: inputting the first feature map and the second feature map into the image generation layer of the image generation model, and generating a retinal thickness map based on the first feature map, the second feature map and a preset noise map through the image generation layer.
[0136] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: acquiring a training sample set, the training sample set including sample images and labeled retinal thickness maps, the sample images including a first sample feature map, a second sample feature map and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the labeled retinal thickness map; training the image generation layer included in the initial image generation model based on the sample images and the labeled retinal thickness map, to obtain the image generation layer included in the image generation model.
[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0139] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0140] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A medical image processing method, characterized in that, The method includes: Obtain a global eye image, and then crop a local eye image from the global eye image; Global feature extraction is performed on the global eye image to obtain a first feature map, and local feature extraction is performed on the local eye image to obtain a second feature map. The first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of the target region in the local eye image. The first feature map and the second feature map are input into the image generation layer of the image generation model, and the image generation layer generates a retinal thickness map based on the first feature map, the second feature map and a preset noise map.
2. The method according to claim 1, characterized in that, The step of performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, includes: The global eye image is input into the global encoder of the image generation model to obtain the first feature map, and the local eye image is input into the local encoder of the image generation model to obtain the second feature map.
3. The method according to claim 2, characterized in that, The step of inputting the global eye image into the global encoder of the image generation model to obtain the first feature map includes: The global eye image is processed by the global encoder to extract and enhance features, thereby obtaining the first feature map output by the global encoder. The first feature map includes the optic disc feature map, blood vessel course feature map, blood vessel density feature map, and global retinal feature map of the global eye image.
4. The method according to claim 2, characterized in that, The step of inputting the local eye image into the local encoder of the image generation model to obtain the second feature map includes: The local encoder extracts the local eye image and the target image corresponding to the target region, and performs multiple rounds of feature extraction on the target image to obtain multiple feature maps; The second feature map is obtained by fusing features from each feature map using the local encoder. The second feature map includes the macular region feature map of the local eye image.
5. The method according to any one of claims 1-4, characterized in that, The training process for the image generation layer of the image generation model includes: A training sample set is obtained, which includes sample images and labeled retinal thickness maps. The sample images include a first sample feature map, a second sample feature map, and a noisy retinal thickness map. The noisy retinal thickness map is obtained by adding noise to the labeled retinal thickness map. The initial image generation model is trained based on the sample image and the labeled retinal thickness map to obtain the image generation layer of the image generation model.
6. A medical image processing device, characterized in that, The device includes: An image acquisition module is used to acquire a global eye image and extract a local eye image from the global eye image; The feature processing module is used to perform global feature extraction on the global eye image to obtain a first feature map, and to perform local feature extraction on the local eye image to obtain a second feature map. The first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of the target region in the local eye image. An image generation unit is used to input the first feature map and the second feature map into the image generation layer of the image generation model, and generate a retinal thickness map based on the first feature map, the second feature map and a preset noise map through the image generation layer.
7. The medical image processing apparatus according to claim 6, characterized in that, The feature processing module includes an encoder processing unit; The encoder processing unit is used to input the global eye image into the global encoder in the image generation model to obtain the first feature map, and to input the local eye image into the local encoder in the image generation model to obtain the second feature map.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.