Medical image processing method and device, equipment, storage medium and program product
By extracting features in parallel using global and local encoders, a retinal thickness map is generated, which solves the problem of low acquisition efficiency caused by the complexity of OCT device operation and enables rapid acquisition of retinal thickness maps.
Patent Information
- Application Number
- CN202511316335.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing OCT equipment is complex to operate, makes it difficult to quickly acquire retinal thickness images, and has poor acquisition efficiency.
By acquiring a global eye image and cropping a local eye image from it, parallel feature extraction is performed using global and local encoders in the image generation model to generate a retinal thickness map.
It improves the efficiency of acquiring retinal thickness images, avoids the need for complex OCT equipment, and enables rapid acquisition of retinal thickness images.
Smart Images

Figure CN121169869A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a medical image processing method and device, computer equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] With the continuous development of image processing technology and the continuous improvement of user health awareness, users expect to collect eye medical images through better image processing technology to obtain clearer and more accurate eye medical images, and members of medical institutions also expect to quickly obtain accurate and clear eye medical images through better image processing technology.
[0003] In the prior art, a retinal thickness image is usually collected through a specific OCT (Optical Coherence Tomography) device, but the operation of the OCT device is complex, it is difficult to quickly obtain a retinal thickness image, and there is a problem of poor acquisition efficiency. SUMMARY
[0004] Therefore, it is necessary to provide a medical image processing method, device, computer equipment, computer readable storage medium and computer program product capable of improving the acquisition efficiency of a retinal image to solve the above technical problems.
[0005] In a first aspect, the present application provides a medical image processing method, comprising: acquiring a global eye image, and cutting a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, the first feature map comprising image features of all regions in the global eye image, and the second feature map comprising image features of a target region in the local eye image; and generating a retinal thickness map according to the first feature map and the second feature map.
[0006] In one embodiment, the global feature extraction on the global eye image to obtain the first feature map and the local feature extraction on the local eye image to obtain the second feature map comprise: inputting the global eye image into a global encoder in an image generation model to obtain the first feature map, and inputting the local eye image into a local encoder in the image generation model to obtain the second feature map.
[0007] In one of the embodiments, the global eye image is input into a global encoder in the image generation model to obtain a first feature map, including: performing feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; wherein the first feature map includes a optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map and a retinal global feature map of the global eye image.
[0008] In one of the embodiments, the local eye image is input into a local encoder in the image generation model to obtain a second feature map, including: intercepting a target image corresponding to the target region of the local eye image by the local encoder, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; and performing feature fusion on each feature map by the local encoder to obtain the second feature map; wherein the second feature map includes a macular area feature map of the local eye image.
[0009] In one of the embodiments, the retinal thickness map is generated according to the first feature map and the second feature map, including: inputting the first feature map and the second feature map into an image generation layer included in the image generation model, and generating the retinal thickness map based on the first feature map, the second feature map and a preset noise map by the image generation layer.
[0010] In one of the embodiments, the training process of the image generation layer included in the image generation model includes: obtaining a training sample set, the training sample set including a sample image and a label retinal thickness map, the sample image including a first sample feature map, a second sample feature map and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the label retinal thickness map; training the image generation layer included in the initial image generation model based on the sample image and the label retinal thickness map to obtain the image generation layer included in the image generation model.
[0011] In a second aspect, the present application further provides a medical image processing device, including: an image acquisition module, configured to acquire a global eye image and intercept a local eye image from the global eye image; a feature processing module, configured to perform global feature extraction on the global eye image to obtain a first feature map, and perform local feature extraction on the local eye image to obtain a second feature map, the first feature map including image features of all regions in the global eye image, and the second feature map including image features of a target region in the local eye image; and an image generation module, configured to generate a retinal thickness map according to the first feature map and the second feature map.
[0012] In a third aspect, the present application further provides a computer device, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps of the method of the first aspect when executing the computer program.
[0013] In a fourth aspect, the present application also provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the method of the first aspect.
[0014] In a fifth aspect, the present application also provides a computer program product, comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0015] The medical image processing method, device, computer device, computer readable storage medium and computer program product described above, by acquiring a global eye image, cutting a local eye image from the global eye image, performing global feature extraction on the global eye image to obtain a first feature map of all regions in the global eye image, performing local feature extraction on the local eye image to obtain a second feature map of a target image in the local eye image, performing global and local image feature extraction in parallel, improves the image processing efficiency, and generates a retinal thickness map according to the first feature map and the second feature map, avoiding obtaining the retinal thickness map by operating a complex OCT device, but generating the retinal thickness map by image processing based on an easily acquired eye image, improving the acquisition efficiency of the retinal thickness map. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0017] Figure 1 An application environment diagram of the medical image processing method in an embodiment;
[0018] Figure 2 A flowchart of the medical image processing method in an embodiment;
[0019] Figure 3 A flowchart of step 202 in an embodiment;
[0020] Figure 4 A flowchart of the acquisition of the first feature map in an embodiment;
[0021] Figure 5 A flowchart of the acquisition of the second feature map in an embodiment;
[0022] Figure 6 A flowchart of step 203 in an embodiment;
[0023] Figure 7 a training flowchart of an image generation layer of an image generation model in an embodiment;
[0024] Figure 8 a flowchart of a medical image processing method in another embodiment;
[0025] Figure 9 a structural block diagram of a medical image processing device in an embodiment;
[0026] Figure 10 an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION
[0027] In order to make the purposes, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0028] It should be noted that the terms "first", "second" and the like used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "multiple" used in the present application refers to two and more than two. The term "and / or" used in the present application refers to one of the solutions, or any combination of multiple solutions.
[0029] The medical image processing method provided by the embodiments of the present application can be applied in an application environment as shown in the following. Figure 1 The application environment at least includes a terminal 101, and the application environment can further include a server 102.
[0030] The terminal 101 is configured to acquire a global eye image and a local eye image input by a user, perform global feature extraction on the global eye image to obtain a first feature map, perform local feature extraction on the local eye image to obtain a second feature map, generate a retinal thickness map according to the first feature map and the second feature map, and display the retinal thickness map to the user. The terminal 101 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, unmanned aerial vehicles, low-altitude flying vehicles, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle-mounted device, a projection device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc.
[0031] It should be noted that the process of feature extraction and retinal image generation can be performed by the terminal 101 or the server 102. The terminal 101 can obtain the global eye image and the local eye image input by the user, call the medical image processing interface of the server 102 to send the global eye image and the local eye image, the server 102 performs global feature extraction on the global eye image to obtain the first feature map, performs local feature extraction on the local eye image to obtain the second feature map, generates the retinal thickness map according to the first feature map and the second feature map, and returns the retinal thickness map to the terminal 101. The terminal 101 displays the retinal thickness map to the user. The server 102 can be a physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal 101 communicates with the server 102 through a network.
[0032] In an exemplary embodiment, as shown in Figure 2 , a medical image processing method is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps 201 to 203.
[0033] Step 201, obtaining a global eye image and a local eye image from the global eye image.
[0034] In this application, the eye image refers to a color fundus image, which is an image of the posterior segment of the human eye including the retina, optic disc (optic nerve head), macular area and choroid structure, etc. taken by an image acquisition device. Further, the global eye image refers to a color fundus image of the entire eye region, such as a color fundus image including the posterior polar region, the middle peripheral region, and the far peripheral region. The local eye image refers to a color fundus image of a partial region of the eye, such as a color fundus image including the macular region.
[0035] In the implementation process, the terminal obtains the global eye image and the local eye image based on the interactive page, wherein the local eye image can be obtained by cutting from the global eye image. In the execution process, the user can pre-cut the local eye image from the global eye image and input the global eye image and the local eye image through the interactive page at the same time; that is, the terminal obtains the global eye image and the local eye image input by the user based on the interactive page.
[0036] In addition, the user can also input only the global eye image based on the interactive page, and the terminal obtains the local eye image corresponding to the central region of the eye based on the global eye image after receiving the global eye image. For example, the terminal obtains the local eye image corresponding to the macular region in the global eye image.
[0037] Further, the size of the global eye image and / or the local eye image can not conform to the standard, and the terminal can perform image segmentation on the global eye image and / or the local eye image, cut off the excess image, and the terminal can detect the center region of the global eye image and / or the local eye image through an edge detection algorithm, and cut off the edge region based on the preset size and the center region, wherein the center region can be the pupil region.
[0038] In which, the specific standard size can be obtained by registration according to the eye image and the corresponding infrared fundus photography image (IR-FP image) in advance. In the registration process, the pre-trained Unet is used to extract the retinal blood vessel binary graph as the common feature for the color fundus image (denoted as C-FPglobal) and the corresponding IR-FP image. Then, the AKAZE feature point detection operator is used to extract the key points on the blood vessel graph, and the RANSAC algorithm is used to estimate the preliminary homography transformation matrix to eliminate outliers in the matching. On this basis, the k nearest neighbor feature matching strategy is further used to finely match the blood vessel grid control points. After completing the rigid registration, the non-rigid deformation correction is performed on the C-FP image to fit the IR-FP, and finally the macular local color image corresponding to the OCT scanning region is intercepted, and the size of the macular local color image is taken as the standard size.
[0039] In addition, other feature extraction and matching algorithms can also be used as long as they can achieve the general alignment of C-FP and RTM. If there is no IR-FP image, the coordinate mapping between the color image and the template thickness map can be established according to the fundus anatomical landmarks (such as the optic disc, macular position, etc.), and the registration is approximately completed to obtain the standard size.
[0040] In the actual scene, the color fundus image can be obtained through simple equipment, such as through simple ophthalmic medical equipment, or even through the image acquisition function of the smart terminal. The color fundus image is easy to obtain, and the retinal thickness image is obtained through the medical image processing of the color fundus image, which greatly improves the acquisition efficiency of the retinal thickness image.
[0041] In step 202, global feature extraction is performed on the global eye image to obtain a first feature map, and local feature extraction is performed on the local eye image to obtain a second feature map.
[0042] In the implementation process, the terminal can perform feature extraction on the global eye image and the local eye image in parallel. Global feature extraction is performed on the global eye image to obtain a first feature map, and local feature extraction is performed on the local eye image to obtain a second feature map. The first feature map includes image features of all regions in the global eye image, and the second feature map includes image features of the eye center region in the local eye image.
[0043] In the execution process, the terminal can perform feature extraction on the global eye image and the local eye image through a feature extraction algorithm to obtain a first feature map and a second feature map; further, the terminal can perform feature enhancement on the feature extraction result, and take the enhanced feature map as the first feature map and the second feature map.
[0044] In addition, the feature extraction processing can also be performed through a pre-trained feature processing model. The global eye image and the local eye image can be input into the pre-trained feature processing model for feature processing to obtain a first feature map and a second feature map output by the feature processing model.
[0045] It should be noted that the same feature processing mode can be used for the first feature map and the second feature map, or different feature processing modes can be used, which will not be described here.
[0046] It should be further noted that, due to the computing power of the terminal, the feature extraction operation of the global eye image and the local eye image can also be performed by a backend server. The terminal can call the interface of the server to input the global eye image and the local eye image. The server performs feature extraction on the global eye image and the local eye image to obtain a first feature map and a second feature map.
[0047] In step 203, a retinal thickness map is generated according to the first feature map and the second feature map.
[0048] In this application, the retinal thickness map is used to display a visual image of the thickness distribution of the retina at different positions. For the retinal thickness map, warm colors (such as red, orange, and yellow) can be used to represent thicker areas, and cold colors (such as green and blue) can be used to represent thinner or normal thickness areas. Alternatively, a circular or square area can be displayed with the macular region as the center, and the thickness distribution in the area can be divided into multiple quadrants (such as upper, lower, nasal, and temporal) and ring areas (such as inner and outer rings) for quantitative analysis.
[0049] In the implementation process, the terminal can generate a retinal thickness map according to the feature information of the first feature map and the second feature map, or the terminal can convert a preset image into a retinal thickness map based on the feature information of the first feature map and the second feature map. In the execution process, the retinal thickness map can be generated according to the feature information of the first feature map and the second feature map and a preset noise map.
[0050] The operation of generating a retinal thickness map according to the first feature map and the second feature map can also be performed by a server. The server can generate a retinal thickness map based on the first feature map and the second feature map and send it to the terminal. The terminal obtains the retinal thickness map returned by the interface call.
[0051] Further, the terminal can also display the retinal thickness map in the page.
[0052] It should be noted that steps 201 to 203 can also be completed based on a server, and the terminal and the server can cooperate to execute, the terminal can obtain the global eye image and / or the local eye image input by the interactive page, and transmit the global eye image and / or the local eye image to the server, the server can obtain the global eye image and / or the local eye image, in the case where the server only obtains the global eye image, the local eye image corresponding to the eye center region can be intercepted from the global eye image, the server can perform global feature extraction on the global eye image to obtain a first feature map, and perform local feature extraction on the local eye image to obtain a second feature map, generate a retinal thickness map according to the first feature map and the second feature map, and send the retinal thickness map to the terminal, and the terminal displays the retinal thickness map.
[0053] In the above medical image processing method, the global eye image is obtained, the local eye image is intercepted from the global eye image, the global feature extraction is performed on the global eye image to obtain the first feature map of all regions in the global eye image, the local feature extraction is performed on the local eye image to obtain the second feature map of the target image in the local eye image, the global and local image feature extraction are performed in parallel, the image processing efficiency is improved, the retinal thickness map is generated according to the first feature map and the second feature map, the retinal thickness map is avoided to be obtained by operating the complex OCT device, but the retinal thickness map is generated by image processing based on the easily obtained eye image, and the acquisition efficiency of the retinal thickness map is improved.
[0054] Based on the above one exemplary embodiment, the following is provided in one or more exemplary embodiments, a medical image processing method, which is applied to Figure 1 the terminal in the above embodiment, and specifically includes the following contents.
[0055] In actual scenarios, in order to further improve the efficiency and accuracy of image processing, an overall image generation model can be used for image processing, feature extraction and image generation operations, the terminal can obtain the global eye image and / or the local eye image based on the interactive page, input the global eye image and / or the local eye image into the image generation model, the image generation model performs image segmentation, feature extraction and image generation operations, and outputs the retinal thickness map.
[0056] In the feature extraction process, the image generation model can perform feature extraction operations on the global eye image and the local eye image in parallel through two encoders; in an optional embodiment provided by the present application, as shown in Figure 3 step 202 includes step 301:
[0057] Step 301, input the global eye image into the global encoder in the image generation model to obtain a first feature map, and input the local eye image into the local encoder in the image generation model to obtain a second feature map.
[0058] In the implementation process, the terminal inputs the global eye image and the local eye image into the image generation model, inputs the global eye image into the global encoder, performs feature processing through the global encoder to obtain the first feature map, and inputs the local eye image into the local encoder to obtain the second feature map.
[0059] In the execution process, the terminal can mark the global eye image and the local eye image respectively, so that the image generation model transmits the global eye image and the local eye image to the respective corresponding encoders according to the marks, and performs feature processing according to the above operations to obtain the first feature map and the second feature map.
[0060] In addition, the image generation model can also identify the global eye image and the local eye image, input the global eye image into the global encoder according to the identification result, perform feature extraction through the global encoder to obtain the first feature map, and input the local eye image into the local encoder according to the identification result, perform feature extraction through the local encoder to obtain the second feature map.
[0061] An optional embodiment provided in the present application performs feature processing on the global eye image and the local eye image in parallel through the global encoder and the local encoder included in the image generation model, which improves the feature processing efficiency, and at the same time, through the pre-trained encoder, the accuracy and reliability of the obtained first feature map and second feature map are improved.
[0062] Further, in the processing process of the first feature map, the first feature map can be obtained through feature extraction and feature enhancement to improve the usability of the first feature map; as shown in an optional embodiment provided in the present application, Figure 4 the acquisition process of the first feature map includes step 401:
[0063] Step 401, perform feature extraction processing and feature enhancement processing on the global eye image through the global encoder to obtain the first feature map output by the global encoder.
[0064] In the present application, the first feature map is a feature map used to represent the global features in the global eye image, and can be specifically used to represent the image features of the regions of the posterior pole, the mid-peripheral part and the far-peripheral part; wherein the first feature map includes a optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map and a retinal global feature map of the global eye image.
[0065] In the feature processing process, the global encoder can perform multi-round feature extraction on the global eye image, and perform feature enhancement processing based on the results of the multi-round feature extraction to obtain a first feature map output by the global encoder.
[0066] In the execution process, the global encoder can use a visual Transformer stem (ViT) of a RETFound pre-training model as a feature extraction network. The RETFound is a retinal basic model trained for fundus diseases, and the ViT encoder has learned rich anatomical and pathological patterns on a large-scale fundus data. The model parameters of the ViT stem can be frozen, and after inputting the C-FPglobal, multi-scale high-dimensional feature tokens are extracted from different levels of the ViT. The tokens of the four stages are reorganized / projected through a designed adapter module to form a global feature map Fg (about H / 4×W / 4×64 in size) comparable in size to the local feature. The global feature carries wide-area retinal anatomical background and large-range lesion context information, such as the optic disc, the shape and density of the large-area blood vessels, and the overall state of the peripheral retina, which helps the model to understand whether the changes in the macular region are part of the global pathology, and improves the biological consistency of the prediction.
[0067] In addition, the global encoder can also use other convolutional neural networks (such as Unet, ResNet+FPN, etc.) or visual Transformer models capable of extracting multi-scale features. In addition to the RETFound pre-training ViT, the global encoder can also select other backbone networks trained on large-scale image data. Even the global branch can input images of other modalities (such as wide-angle fundus photography) to provide more comprehensive retinal background information.
[0068] In an optional implementation provided in the present application, the first feature map is generated through the multi-round feature extraction and feature enhancement processing of the global encoder, which improves the availability and reliability of the first feature map.
[0069] In the process of obtaining the second feature map, the local encoder can first intercept a target image corresponding to the target region from the local eye image, perform multi-round feature extraction on the target image, and fuse the target image to obtain the second feature map. In an optional implementation provided in the present application, as shown in Figure 5 the process of obtaining the second feature map includes steps 501 to 502:
[0070] In step 501, the local encoder intercepts a target image corresponding to the target region from the local eye image, performs multi-round feature extraction on the target image, and obtains a plurality of feature maps.
[0071] The second feature map in the application is a feature map for representing local features of a target region in a local eye image, and can be used to represent image features of a macular region; wherein the second feature map includes a macular region feature map of the local eye image.
[0072] In the implementation process, the local encoder identifies and intercepts a target image corresponding to the target region in the local eye image, performs multi-round feature extraction on the target image, and obtains a plurality of feature maps; wherein the target region can be a macular region.
[0073] In the execution process, the local encoder can detect a region matching the features of the target region through an edge detection algorithm, and intercept an image of the region as the target image. In the feature processing process, the local encoder can extract high-resolution local features using a deep network based on SwinTransformer. Swin Transformer obtains local details through a sliding window self-attention mechanism.
[0074] Step 502: performing feature fusion on each feature map by the local encoder to obtain a second feature map.
[0075] In the implementation process, the local encoder performs feature superposition / feature fusion processing on each feature map to obtain a second feature map representing the features of the macular region.
[0076] In the execution process, the local encoder constructs a hierarchical representation by layer-by-layer patch merging, and then combines a feature pyramid network (FPN) to fuse feature maps of multiple scales. The local feature tensor Fm obtained in this way has a size of about 1 / 4 (H / 4×W / 4×256) of the original image, and can represent the fine structure changes of the macular region, such as the small thickness fluctuations in the fovea, texture changes caused by local exudation, etc.
[0077] In addition, the local encoder can also use other convolutional neural networks (such as Unet, ResNet+FPN, etc.) or visual Transformer models that can extract multi-scale features. In addition to the RETFound pre-trained ViT, the global encoder can also use other backbone networks trained on large-scale image data. Even the global branch can input images of other modalities (such as wide-angle fundus photography) to provide more comprehensive retinal background information.
[0078] An optional implementation provided by the application obtains a feature map representing the macular region by using the local encoder to perform image segmentation, multi-round feature processing, and feature fusion, thereby improving the usability and accuracy of the generated second feature map, and further improving the effectiveness of the generated retinal thickness map.
[0079] In the retinal thickness map generation process, the retinal thickness map is also a kind of "feature map" for representing the thickness of the retina. The retinal thickness map can be generated by the first feature map and the second feature map. In the generation process of the first feature map and the second feature map, part of the noise can be introduced. The first feature map and the second feature map can be denoised first, and the retinal thickness map can be generated based on the denoised first feature map and the second feature map. In an optional embodiment of the present application, as shown in Figure 6 Step 203 includes step 601 as shown in
[0080] Step 601 inputs the first feature map and the second feature map into the image generation layer included in the image generation model, and generates the retinal thickness map based on the first feature map, the second feature map and the preset noise map through the image generation layer.
[0081] In the implementation process, the terminal can input the first feature map and the second feature map into the image generation layer included in the image generation model, and use the feature information of the first feature map and the second feature map as inference guide information to denoise the noise map through the image generation layer, to obtain the retinal thickness map. The noise map can be a Gaussian noise map. The preset noise map can be fixed data pre-stored in the image generation layer.
[0082] In the execution process, the retinal thickness map can also be generated based on the image generation model. The image generation model can input the first feature map generated by the global encoder and the second feature map generated by the local encoder into the image generation layer, and generate the retinal thickness map based on the first feature map, the second feature map and the preset noise map through the image generation layer.
[0083] In the retinal thickness map generation process, the generation of the prediction result can be realized through the reverse diffusion process, that is, the noise map is gradually recovered to the retinal thickness map, and the first feature map and the second feature map are used as guide data. Further, the prediction result can be optimized by minimizing the mean square error, so as to minimize the difference between the prediction result and the target result. In the denoising process, multiple denoising methods can be used. After each denoising, the first feature map and the second feature map can be compared, and the next denoising process can be continued according to the comparison result, until the feature information of the retinal thickness map matches the first feature map and the second feature map.
[0084] For example, in the denoising process, the RTM can be generated reversely by iteratively applying the diffusion decoder starting from the random noise map combined with the input photo features. To improve the inference efficiency, we introduce the DDIM (Denoising Diffusion Implicit Model) sampling strategy. DDIM is a deterministic sampling method that is not a Markov chain, which can generate high-fidelity images with fewer diffusion steps, thereby speeding up the generation and making the model meet the needs of real-time clinical applications. In addition, the conditional diffusion framework naturally has diversity and noise resistance, and is robust when training data containing image noise or quality differences, which can improve the robustness of the prediction results to changes in photography quality.
[0085] In the retinal thickness map generation process, the denoised noise map can be rendered with color to obtain a retinal thickness map containing color; wherein the part with more noise points can be rendered as a dark area, i.e. the area with larger retinal thickness, and the part with fewer noise points can be rendered as a light area, i.e. the area with smaller retinal thickness.
[0086] An optional embodiment provided in the present application generates a retinal thickness map through the first feature map and the second feature map of the color fundus image, improves the generation efficiency of the retinal thickness map, and at the same time, reduces the acquisition cost of the retinal thickness map, and improves the acquisition efficiency of the retinal thickness map.
[0087] In actual scenarios, the image generation layer of the image generation model also needs to be pre-trained. The image generation layer of the initial image generation model can be trained based on a pre-acquired training sample set to obtain the image generation layer of the image generation model. In an optional embodiment provided in the present application, as shown in Figure 7 The training process of the image generation layer of the image generation model includes steps 701 to 702:
[0088] Step 701, acquiring a training sample set.
[0089] In the implementation process, the model training device acquires the training sample set. The training sample set includes a sample image and a label retinal thickness map. The sample image includes a first sample feature map, a second sample feature map, and a noisy retinal thickness map. The noisy retinal thickness map is obtained by adding noise to the label retinal thickness map. The first sample feature map and the second sample feature map are obtained by processing the label retinal thickness map.
[0090] In the acquisition process of the noisy retinal thickness map, the label retinal thickness map can be gradually added noise by the conditional diffusion method to generate the noisy retinal thickness map. Optionally, there can be multiple noisy retinal thickness maps, and the noise levels of the noisy retinal thickness maps can be different.
[0091] In the execution process, Gaussian noise can be gradually added to the label retinal thickness map according to a preset noise scheduling strategy to obtain thickness maps of different degradation degrees.
[0092] At step 702, the image generation layer included in the initial image generation model is trained based on the sample image and the label retinal thickness map to obtain the image generation layer included in the image generation model.
[0093] In the implementation process, the model training device can input the sample image into the image generation layer included in the initial image generation model, generate a retinal thickness map through the image generation layer, calculate a model training loss according to the generated retinal thickness map and the label retinal thickness map, adjust the model parameters according to the loss value, and stop until the loss value converges to obtain the image generation layer included in the image generation model.
[0094] In the model parameter adjustment process, the features (Fm and Fg) extracted by the double encoder can be input into the diffusion decoder Ed together with the noise map Mt. The diffusion decoder is composed of multiple deformable attention modules and can realize more fine cross-modal alignment of global and local features and the current generated image content. The decoder learns to predict the thickness map Mt at the previous time from the noisy thickness map Mt under the given conditional features, and gradually denoises to approximate the label retinal thickness map. The loss function can refer to formula (1):
[0095] Formula (1).
[0096] In addition, the model can also be applied to other related tasks. For example, through transfer learning adjustment, the change amount of retinal layer thickness and retinal anatomical structure segmentation can be predicted from color fundus photos.
[0097] In an optional embodiment provided in the present application, in the training process, Gaussian noise is gradually superimposed on the label retinal thickness data according to a preset noise scheduler to form noisy retinal thickness data at different time steps, then the feature vectors from the double encoder are spliced as conditional input to guide the diffusion process, and the decoder learns how to recover the original mapping information based on the noisy target data and the conditional features in the training stage. The training of the image generation layer included in the image generation model is performed through diffusion training, which improves the effectiveness and accuracy of the trained image generation layer, and further improves the usability of the generated retinal thickness map.
[0098] In one embodiment, referring to Figure 8 which shows a flowchart of a medical image processing method provided in an embodiment of the present application. The medical image processing method can be applied to Figure 1 the terminal shown in the figure. As Figure 8As shown, the medical image processing method can include the following steps:
[0099] Step 801, a global eye image is acquired, and a local eye image is cropped from the global eye image.
[0100] Step 802, the global eye image and the local eye image are input into an image generation model, feature extraction processing and feature enhancement processing are performed on the global eye image by a global encoder in the image generation model, and a first feature map output by the global encoder is obtained.
[0101] Optionally, the first feature map includes a disc feature map, a blood vessel shape feature map, a blood vessel density feature map and a global retinal feature map of the global eye image.
[0102] Step 803, a target image corresponding to a target region is cropped from the local eye image by a local encoder in the image generation model, a plurality of feature maps are obtained by performing multi-round feature extraction on the target image, and a second feature map is obtained by performing feature fusion on each feature map.
[0103] Optionally, the second feature map includes a macular area feature map of the local eye image.
[0104] Step 804, the image generation model inputs the first feature map and the second feature map into an image generation layer included in the image generation model, performs denoising processing on the first feature map and the second feature map by the image generation layer, and generates a retinal thickness map based on the denoised first feature map and the denoised second feature map.
[0105] Step 805, the retinal thickness map is acquired and displayed on a screen.
[0106] It should be noted that any one step or combination of any multiple steps of steps 801 to 805 can be selected as any one step or combination of any multiple steps of steps 201 to 203 provided in the above embodiments to form a new implementation manner according to the needs of implementation deployment; and any one or any multiple technical features in the technical solution composed of steps 801 to 805 can also be selected as any one or multiple technical features in the technical solution composed of steps 201 to 203 according to the needs of actual deployment to form a new implementation manner, or the technical features in one or more optional implementation manners provided in one or more embodiments are combined to form a new implementation manner, which will not be described here.
[0107] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by the combination are within the scope of protection of the present application.
[0108] The present application also provides one or more experimental embodiments. The present application uses a conditional diffusion model (GLD-RT) combined with multi-scale feature fusion and deformable attention to achieve high-precision thickness reconstruction in each region such as fovea (G1), inner ring (G2), and outer ring (G3). Compared with existing methods, the average thickness error (MAE) is reduced by more than 15%, and the peak signal-to-noise ratio (PSNR) is improved by more than 1.5 dB; supports near real-time thickness map generation; using RETFound and other large-scale pre-training models, it performs stably on different camera devices and crowd data, and has good clinical promotion adaptability; in addition, the output format is compatible with the OCT thickness map, and doctors can directly observe the thickness distribution; the model only needs a general fundus camera to be deployed, which is convenient for embedding in existing screening equipment or mobile diagnosis and treatment systems, and improves the efficiency and confidence of diagnosis and treatment decisions.
[0109] Based on the same inventive concept, the embodiments of the present application also provide a medical image processing device for implementing the medical image processing method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more medical image processing device embodiments provided below can refer to the limitations of the medical image processing method described above, which will not be repeated here.
[0110] In one exemplary embodiment, as Figure 9As shown, a medical image processing apparatus is provided, comprising: an image acquisition module 901, a feature processing module 902, and an image generation module 903, wherein: the image acquisition module 901 is configured to acquire a global eye image and crop a local eye image from the global eye image; the feature processing module 902 is configured to perform global feature extraction on the global eye image to obtain a first feature map, and perform local feature extraction on the local eye image to obtain a second feature map, the first feature map comprising image features of all regions in the global eye image, and the second feature map comprising image features of a target region in the local eye image; and the image generation module 903 is configured to generate a retinal thickness map based on the first feature map and the second feature map.
[0111] In one of the embodiments, the feature processing module 902 comprises an encoder processing unit, wherein: the encoder processing unit is configured to input the global eye image into a global encoder in an image generation model to obtain the first feature map, and input the local eye image into a local encoder in the image generation model to obtain the second feature map.
[0112] In one of the embodiments, the encoder processing unit comprises a first encoder processing unit, wherein: the first encoder processing unit is configured to perform feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; and the first feature map comprises an optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map, and a retinal global feature map of the global eye image.
[0113] In one of the embodiments, the encoder processing unit comprises a second encoder processing unit and a feature fusion unit, wherein: the second encoder processing unit is configured to perform multi-round feature extraction on a target image corresponding to the target region of the local eye image by the local encoder to obtain a plurality of feature maps; and the feature fusion unit is configured to perform feature fusion on the feature maps by the local encoder to obtain the second feature map; and the second feature map comprises a macular area feature map of the local eye image.
[0114] In one of the embodiments, the image generation module 903 comprises an image generation unit, wherein: the image generation unit is configured to input the first feature map and the second feature map into an image generation layer included in the image generation model, and generate the retinal thickness map based on the first feature map, the second feature map, and a preset noise map by the image generation layer.
[0115] In one of the embodiments, the device further comprises a training sample set obtaining module and a model training module, wherein: the sample set obtaining module is configured to obtain a training sample set, the training sample set comprising a sample image and a label retinal thickness map, the sample image comprising a first sample feature map, a second sample feature map, and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the label retinal thickness map; and the model training module is configured to train an image generation layer included in an initial image generation model based on the sample image and the label retinal thickness map, to obtain the image generation layer included in the image generation model.
[0116] The various modules in the medical image processing device can be implemented wholly or partially by software, hardware, and combinations thereof. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the various modules.
[0117] In one exemplary embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 10 The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to perform wired or wireless communication with external terminals, and the wireless communication can be achieved through WIFI, mobile cellular network, near field communication (NFC), or other technologies. The computer program is executed by the processor to implement a medical image processing method. The display unit of the computer device is configured to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0118] Those skilled in the art can understand that Figure 10The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0119] In one exemplary embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program: obtaining a global eye image, and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, the first feature map including image features of all regions in the global eye image, and the second feature map including image features of a target region in the local eye image; and generating a retinal thickness map according to the first feature map and the second feature map.
[0120] In one embodiment, the processor specifically implements the following steps when executing the computer program: inputting the global eye image into a global encoder in an image generation model to obtain the first feature map, and inputting the local eye image into a local encoder in the image generation model to obtain the second feature map.
[0121] In one embodiment, the processor specifically implements the following steps when executing the computer program: performing feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; and wherein the first feature map includes an optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map, and a retinal global feature map of the global eye image.
[0122] In one embodiment, the processor specifically implements the following steps when executing the computer program: performing multi-round feature extraction on a target image corresponding to the target region of the local eye image by the local encoder to obtain a plurality of feature maps; performing feature fusion on the feature maps by the local encoder to obtain the second feature map; and wherein the second feature map includes a macular area feature map of the local eye image.
[0123] In one embodiment, the processor specifically implements the following steps when executing the computer program: inputting the first feature map and the second feature map into an image generation layer included in the image generation model, and generating the retinal thickness map based on the first feature map, the second feature map, and a preset noise map by the image generation layer.
[0124] In one embodiment, the processor further implements the following steps when executing the computer program: obtaining a training sample set, the training sample set comprising a sample image and a label retinal thickness map, the sample image comprising a first sample feature map, a second sample feature map, and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the label retinal thickness map; and training an image generation layer included in the initial image generation model based on the sample image and the label retinal thickness map to obtain the image generation layer included in the image generation model.
[0125] In one embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps: obtaining a global eye image, and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, the first feature map comprising image features of all regions in the global eye image, and the second feature map comprising image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0126] In one embodiment, the computer program is executed by the processor to specifically implement the following steps: inputting the global eye image into a global encoder in the image generation model to obtain the first feature map, and inputting the local eye image into a local encoder in the image generation model to obtain the second feature map.
[0127] In one embodiment, the computer program is executed by the processor to specifically implement the following steps: performing feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; and the first feature map comprises an optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map, and a retinal global feature map of the global eye image.
[0128] In one embodiment, the computer program is executed by the processor to specifically implement the following steps: performing multi-round feature extraction on a target image corresponding to the target region of the local eye image by the local encoder to obtain a plurality of feature maps; and performing feature fusion on the feature maps by the local encoder to obtain the second feature map; and the second feature map comprises a macular area feature map of the local eye image.
[0129] In one embodiment, the computer program is executed by the processor to specifically implement the following steps: inputting the first feature map and the second feature map into an image generation layer included in the image generation model, and generating the retinal thickness map based on the first feature map, the second feature map, and a preset noise map by the image generation layer.
[0130] In an embodiment, the computer program, when executed by the processor, further implements the following steps: obtaining a training sample set, the training sample set comprising a sample image and a label retinal thickness map, the sample image comprising a first sample feature map, a second sample feature map, and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the label retinal thickness map; and training an image generation layer included in the initial image generation model based on the sample image and the label retinal thickness map to obtain the image generation layer included in the image generation model.
[0131] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the following steps: obtaining a global eye image and cropping a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, the first feature map comprising image features of all regions in the global eye image, and the second feature map comprising image features of a target region in the local eye image; and generating a retinal thickness map based on the first feature map and the second feature map.
[0132] In an embodiment, the computer program, when executed by the processor, specifically implements the following steps: inputting the global eye image into a global encoder in the image generation model to obtain the first feature map, and inputting the local eye image into a local encoder in the image generation model to obtain the second feature map.
[0133] In an embodiment, the computer program, when executed by the processor, specifically implements the following steps: performing feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; and wherein the first feature map comprises an optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map, and a retinal global feature map of the global eye image.
[0134] In an embodiment, the computer program, when executed by the processor, specifically implements the following steps: performing multi-round feature extraction on a target image corresponding to the target region of the local eye image by the local encoder to obtain a plurality of feature maps; and performing feature fusion on the feature maps by the local encoder to obtain the second feature map; and wherein the second feature map comprises a macular area feature map of the local eye image.
[0135] In an embodiment, the computer program, when executed by the processor, specifically implements the following steps: inputting the first feature map and the second feature map into an image generation layer included in the image generation model, and generating the retinal thickness map based on the first feature map, the second feature map, and a preset noise map by the image generation layer.
[0136] In one embodiment, the computer program, when executed by the processor, further implements the following steps: obtaining a training sample set, the training sample set comprising a sample image and a label retinal thickness map, the sample image comprising a first sample feature map, a second sample feature map, and a noisy retinal thickness map, the noisy retinal thickness map being obtained by adding noise to the label retinal thickness map; and training an image generation layer included in the initial image generation model based on the sample image and the label retinal thickness map, to obtain the image generation layer included in the image generation model.
[0137] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0138] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0139] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist, they should be considered as the scope of the present application.
[0140] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A medical image processing method characterized by, The method comprises: obtaining a global eye image and cutting a local eye image from the global eye image; performing global feature extraction on the global eye image to obtain a first feature map, and performing local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map comprises image features of all regions in the global eye image, and the second feature map comprises image features of a target region in the local eye image; generating a retinal thickness map according to the first feature map and the second feature map.
2. The method of claim 1, wherein, The global feature extraction on the global eye image to obtain the first feature map and the local feature extraction on the local eye image to obtain the second feature map comprise: inputting the global eye image into a global encoder in an image generation model to obtain the first feature map, and inputting the local eye image into a local encoder in the image generation model to obtain the second feature map.
3. The method of claim 2, wherein, The inputting the global eye image into the global encoder in the image generation model to obtain the first feature map comprises: performing feature extraction processing and feature enhancement processing on the global eye image by the global encoder to obtain the first feature map output by the global encoder; wherein the first feature map comprises an optic disc feature map, a blood vessel shape feature map, a blood vessel density feature map and a retinal global feature map of the global eye image.
4. The method of claim 2, wherein, The inputting the local eye image into the local encoder in the image generation model to obtain the second feature map comprises: cutting a target image corresponding to the target region from the local eye image by the local encoder, performing multiple rounds of feature extraction on the target image to obtain multiple feature maps; performing feature fusion on each of the feature maps by the local encoder to obtain the second feature map; wherein the second feature map comprises a macular area feature map of the local eye image.
5. The method according to any one of claims 1 to 4, characterized in that, The generating a retinal thickness map according to the first feature map and the second feature map comprises: inputting the first feature map and the second feature map into an image generation layer included in an image generation model, and generating the retinal thickness map based on the first feature map, the second feature map and a preset noise map by the image generation layer.
6. The method of claim 5, wherein, The training process of the image generation layer included in the image generation model comprises: obtaining a training sample set, wherein the training sample set comprises a sample image and a label retinal thickness map, the sample image comprises a first sample feature map, a second sample feature map and a noise-added retinal thickness map, and the noise-added retinal thickness map is obtained by adding noise to the label retinal thickness map; training an image generation layer included in an initial image generation model based on the sample image and the label retinal thickness map to obtain an image generation layer included in the image generation model.
7. A medical image processing apparatus characterized by comprising: The device comprises: an image acquisition module configured to obtain a global eye image and cut a local eye image from the global eye image; The feature processing module is configured to perform global feature extraction on the global eye image to obtain a first feature map, and perform local feature extraction on the local eye image to obtain a second feature map, wherein the first feature map comprises image features of all regions in the global eye image, and the second feature map comprises image features of a target region in the local eye image. The image generation module is configured to generate a retinal thickness map according to the first feature map and the second feature map.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6. The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for glaucoma image detection and storage medium
CN114332002A
Retinal map construction method and device, computer equipment and storage medium
CN115439900A
Method and device for optimizing eye medical image and storage medium
CN117314911A
Global and local information fused retinal vessel segmentation method and system
CN117746037A
Fundus image intelligent enhancement method and system based on multi-modal image fusion
CN120563387A