Medical image enhancement method based on human in-loop and visual basis model

By combining the variational autoencoder and the human-in-loop interaction mechanism, the instability and adaptability problems of existing medical image enhancement methods are solved, efficient and accurate medical image enhancement is achieved, and image quality and diagnostic accuracy are improved.

CN120298246APending Publication Date: 2025-07-11EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444452.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing medical image enhancement methods are difficult to accurately capture expert domain knowledge, the enhancement effect is unstable, the generalization performance is insufficient, and the need for enhancing details of complex medical images is unable to meet the needs of enhancing complex medical images, and lack the adaptability to different lesion areas, tissue types and imaging modalities.

Method used

Combining the variational autoencoder and human-in-loop interaction mechanism, efficient and accurate medical image enhancement is achieved through data preprocessing, image quality evaluation, visual basic model and expert feedback optimization.

Benefits of technology

It improves the quality of medical images, enhances the results to meet clinical diagnosis needs, improves the accuracy and efficiency of diagnosis, and reduces processing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298246A_ABST
    Figure CN120298246A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image enhancement method based on a human in-loop and visual basis model. The method is characterized by comprising the following steps: generating an enhanced image candidate set; selecting an optimal enhanced image from the candidate set; detail enhancement is carried out on the variational auto-encoder; performing automatic preliminary enhancement on the SAM; performing feedback correction on a preliminary enhancement result in a man-machine interaction mode, and optimizing details of a key area; inputting an expert feedback enhancement prompt into the variational auto-encoder and the visual basic model, and iteratively optimizing an image enhancement result; and repeatedly executing the expert feedback step until the quality of the enhanced image reaches a preset threshold value or meets clinical requirements, and the like. Compared with the prior art, the problems of insufficient medical image contrast, serious noise interference, fuzzy details and the like are effectively solved, expert domain knowledge and an automatic vision enhancement model are efficiently combined, the medical image quality is greatly improved, the clinical image processing cost is reduced, the accuracy of medical diagnosis is improved, and the method has remarkable clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and specifically, to a medical image enhancement method based on the combination of human-in-the-loop technology and vision foundation models. Background Art

[0002] Medical image enhancement is a key technology in medical image analysis and clinical diagnosis, and is widely used in medical practice fields such as disease diagnosis, lesion identification, treatment planning, and efficacy evaluation. However, existing medical image enhancement methods generally have problems such as unstable enhancement effects, insufficient generalization performance, and insufficient enhancement of details in complex medical images. At the same time, traditional methods do not effectively utilize expert domain knowledge, and it is difficult to accurately capture experts' understanding of specific clinical needs and image enhancement details, seriously restricting the accuracy and efficiency of clinical diagnosis. Therefore, how to efficiently integrate the generalization ability of automatic enhancement models and expert domain knowledge to form an efficient and accurate medical image enhancement method is a key technical problem urgently needed to be solved in the current medical image processing field.

[0003] Medical image enhancement is a core technology in medical image analysis and clinical diagnosis, and is widely used in various medical practice fields such as disease screening, lesion identification, surgical planning, radiotherapy, tissue anatomy analysis, and efficacy evaluation. High-quality medical image enhancement can effectively improve doctors' interpretation ability, help medical AI systems analyze image data more accurately, and reduce the risk of misdiagnosis and missed diagnosis. However, current medical image enhancement technologies still face many challenges, mainly including problems such as unstable enhancement effects, insufficient generalization performance, and insufficient enhancement of details in complex medical images.

[0004] Early medical image enhancement technologies were mainly based on traditional image processing methods, such as histogram equalization, non-local mean filtering, Laplace transform, etc. These methods can improve the contrast, sharpness, and signal-to-noise ratio of medical images to a certain extent, but they have some inherent limitations: First, using fixed transformation rules and applying the same enhancement strategy to all images, lacking the adaptive ability to different lesion regions, tissue types, and different imaging modalities, resulting in poor enhancement effects in some scenarios. At the same time, over-enhancement may lead to the loss of information in key lesion regions of medical images. For example, although histogram equalization can improve the overall contrast, it may over-enhance the highlight region, making small lesions become blurred. In medical images with low signal-to-noise ratio (such as low-dose CT or low-field MRI), traditional enhancement methods may enhance the noise in the original image, reducing the image quality and affecting doctors' diagnostic judgment.

[0005] With the development of artificial intelligence, especially the introduction of deep learning technology, medical image enhancement has entered the era of intelligence. Image enhancement methods based on convolutional neural networks have made breakthrough progress in recent years. By using deep learning models to directly learn the mapping relationship from low-quality images to high-quality images, data-driven automated image enhancement can be achieved. Another common solution is through generative adversarial networks, which can perform image conversion between different modalities and synthesize high-quality medical images to make up for the information loss caused by low-dose imaging. Using variational autoencoders can achieve detail enhancement of medical images, especially suitable for tissue structures with complex textures (such as lung CT, brain MRI, etc.).

[0006] Although deep learning has greatly improved the automation and enhancement effect of medical image enhancement, these methods still have some problems: deep learning methods require a large amount of high-quality labeled data for training, but it is difficult to obtain medical data, and there is a problem of uneven data distribution. Different imaging devices, imaging parameters, patient individual differences, etc. will all affect the performance of deep learning models, resulting in poor adaptability of the models to new data. Existing deep learning methods mainly rely on end-to-end training and lack direct integration of medical expert domain knowledge, and the enhancement results cannot fully meet the needs of clinicians.

[0007] In summary, the medical image enhancement of the prior art has the following problems: 1) It is difficult to accurately capture the understanding of experts for specific clinical needs and image enhancement details, which severely restricts the accuracy and efficiency of clinical diagnosis; 2) Problems such as unstable enhancement effect, insufficient generalization performance, and insufficient enhancement of details of complex medical images; 3) Lack of adaptability to different lesion regions, tissue types, and different imaging modalities, resulting in poor enhancement performance in some scenarios; 4) Enhancing the highlighted area makes small lesions blurred, enhancing the noise in the original image, reducing the image quality, and affecting the doctor's diagnostic judgment; 5) Lack of direct integration of medical expert domain knowledge, and the enhancement results cannot fully meet the needs of clinicians. Summary of the Invention

[0008] The object of the present invention is to provide a medical image enhancement method based on human-in-the-loop and vision foundation models in view of the deficiencies of the prior art. The medical image enhancement method based on variational autoencoders and human-in-the-loop interaction mechanisms combines automated algorithms with clinical expert knowledge, optimizes the medical image enhancement results through human-computer interaction feedback, and achieves efficient and accurate medical image enhancement. In the data preprocessing stage, the method performs affine transformation, histogram equalization, contrast adjustment, and noise removal on medical images to expand the data distribution and improve the generalization ability of the model. At the same time, indicators such as structural similarity and peak signal-to-noise ratio are used to screen the optimal enhanced images. Subsequently, variational autoencoders are used to enhance the details of medical images, and by minimizing the reconstruction loss and KL divergence loss, it is ensured that the enhanced images reduce noise interference while retaining the details of key lesions. Then, the vision foundation model SAM is combined to analyze medical images, and using the preliminary enhanced results generated by variational autoencoders, combined with edge detection Canny to generate mask prompts to improve the accuracy of lesion segmentation. Further, a "human-in-the-loop" expert feedback mechanism is introduced, allowing experts to manually adjust the enhanced areas on the interactive interface, optimize key details, and input the expert corrections into variational autoencoders and SAM to improve subsequent enhancement effects. Finally, enhanced prompt signals are generated through expert feedback to drive variational autoencoders and SAM to perform multiple rounds of optimization, forming high-quality medical image enhancement results. The present invention effectively solves problems such as insufficient contrast, severe noise interference, and blurred details in medical images through the "human-in-the-loop" expert interaction mechanism, efficiently combines expert domain knowledge with automated vision enhancement models, greatly improves the quality of medical images, reduces the cost of clinical image processing, improves the accuracy of medical diagnosis, and has significant clinical application value.

[0009] The object of the present invention is achieved as follows: A medical image enhancement method based on human-in-the-loop and vision foundation models, which is characterized by performing multi-dimensional data enhancement on the input medical image, including preprocessing steps such as affine transformation (rotation, scaling, translation, shearing, etc.), contrast adjustment (histogram equalization, adaptive enhancement), noise removal (Gaussian filtering, non-local means filtering), etc., to improve the image quality and enrich the training data set; Secondly, based on image quality evaluation metrics (SSIM, PSNR, contrast sensitivity, etc.), automatically select the enhanced image with the optimal quality as the input prompt signal for the model; Then, use a variational autoencoder for detail enhancement of medical images. The variational autoencoder learns high-order medical image features through an encoder-latent variable-decoder structure, thus achieving high-quality reconstruction; Subsequently, use the vision foundation model SAM to automatically perform preliminary enhancement on the images generated by the variational autoencoder. SAM adopts a structure based on Image Encoder + Prompt Encoder + MaskDecoder to segment and enhance key regions of medical images (such as lesions, blood vessels, organ boundaries), and combines the edge detection algorithm Canny to generate mask prompts to improve the accuracy of detail enhancement; Next, through the "human-in-the-loop" mechanism, introduce clinical experts for feedback adjustment. The experts can select specific regions on the interactive interface for optimization. The system automatically records the feedback information and inputs it into the variational autoencoder and SAM for further iterative optimization, finally forming the optimal enhancement result.

[0010] The present invention improves the accuracy and clinical applicability of the enhanced image through expert interactive optimization, making the enhancement effect more in line with the needs of medical diagnosis. Through the latent variable learning ability of the variational autoencoder, capture the deep structural information of medical images and achieve high-quality image reconstruction, which is particularly suitable for detail enhancement of medical images with low signal-to-noise ratio. Combine the automatic preliminary enhancement of the vision foundation model SAM, utilize the segmentation and feature extraction capabilities of SAM to improve the enhancement accuracy of key lesion regions, making the enhanced image more clinically usable. Introduce the "human-in-the-loop" mechanism for optimization. Experts can make fine adjustments to the enhancement results through the interactive interface. The system dynamically adjusts the enhancement strategies of the variational autoencoder and SAM according to the feedback information to achieve personalized optimization. Iteratively optimize the enhancement results, combine expert feedback to generate enhancement prompt signals, and drive the variational autoencoder and SAM for further optimization until the image quality meets the requirements of clinical diagnosis.

[0011] The medical image enhancement of the present invention is carried out according to the following steps: (1) Perform data enhancement processing on the medical image to be enhanced to generate an extended set of enhanced image candidates, specifically including: 1.1: Perform affine transformation (rotation, scaling, translation, shearing, etc.) on a single medical image to expand the data distribution and improve the generalization ability of the model; 1.2: Perform preprocessing such as contrast adjustment (histogram equalization, adaptive enhancement) and noise removal (Gaussian filtering, non-local means filtering) on the image to improve the image quality.

[0012] (II) Based on the image quality evaluation algorithm, automatically select the best enhanced image from the candidate set as the initial input, specifically including: 2.1: Evaluate the image quality using indicators such as structural similarity, peak signal-to-noise ratio, and contrast sensitivity; 2.2: Select the image with the highest comprehensive score as the enhanced prompt image.

[0013] (III) Use a variational autoencoder for detail enhancement, specifically including: 3.1: The variational autoencoder model architecture includes an encoder, a latent variable space, and a decoder. The encoder converts the input medical image into a latent variable distribution; the decoder reconstructs a high-quality enhanced image; 3.2: The training data includes a public medical image database covering BraTS (brain MRI), LIDC-IDRI (lung nodule CT), ISIC (skin cancer), and clinical data. The variational autoencoder model is optimized using adversarial loss and KL divergence loss to enable it to learn the key details of medical images.

[0014] (IV) SAM performs automatic preliminary enhancement, specifically including: 4.1: Use a vision foundation model to further refine the image enhanced by the variational autoencoder to ensure global structural consistency; 4.2: Generate a preliminary enhanced image by fusing the enhancement results of the variational autoencoder and SAM through an attention mechanism.

[0015] (V) The expert makes feedback corrections to the preliminary enhancement results through a human-computer interaction method to optimize the details of key regions, specifically including: 5.1: The expert marks the regions with insufficient details or over-enhancement (such as blurred lesion boundaries, lost texture details, etc.) on the interactive interface; 5.2: The system automatically records the expert feedback and generates an expert feedback enhancement prompt for subsequent optimization.

[0016] (VI) Input the expert feedback enhancement prompt into the variational autoencoder and the vision foundation model to iteratively optimize the image enhancement results, specifically including: 6.1: Use the feedback signal to adjust the parameters of the variational autoencoder decoder to better adapt to the medical image features; 6.2: Update the weights of the vision foundation model in combination with the expert feedback to strengthen the enhancement ability for specific medical features.

[0017] (7) Repeat steps 5) - 6) until the enhanced image quality reaches a preset threshold or meets clinical requirements.

[0018] Compared with the prior art, the present invention has the following beneficial technical effects and significant technological progress: 1) Medical image detail enhancement based on variational autoencoders, leveraging the latent variable feature learning ability of variational autoencoders to effectively remove noise while restoring fine anatomical structure information, thereby improving the readability of medical images. Variational autoencoders can automatically learn the deep features of medical images, enhance the self - adaptability of image enhancement, and reduce artifacts in high - noise environments.

[0019] 2) Through the segmentation and feature extraction capabilities of SAM, accurately enhance key lesion areas, and optimize the enhancement strategy through expert feedback, enabling the finally generated medical images to meet the actual clinical diagnosis requirements, effectively avoiding the opacity problem caused by the pure end - to - end training of deep learning models. Introducing an expert correction mechanism makes the enhancement results more reliable and has the ability of personalized adjustment.

[0020] 3) The iterative optimization mechanism of enhancement results, combining automatic enhancement and interactive feedback optimization, enables the model to dynamically adapt to different types of medical images, improves the stability and generalization ability of enhancement results, greatly reduces the burden of manual annotation, and ensures that the enhanced medical images can meet high - quality standards in various clinical application scenarios.

[0021] 4) Through the "human - in - the - loop" expert interaction mechanism, effectively solve problems such as insufficient contrast, severe noise interference, and blurred details in medical images, efficiently combine expert domain knowledge with automated visual enhancement models, significantly improve the quality of medical images, reduce the cost of clinical image processing, and enhance the accuracy of medical diagnosis, having significant clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the present invention; Figure 2 is a schematic diagram of the specific operation of Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present invention expands the distribution of medical image data through methods such as affine transformation, histogram equalization, and contrast adjustment, and selects the optimal enhancement result using SSIM and PSNR. Subsequently, the VAE trains on datasets such as BraTS, LIDC-IDRI, and ISIC through an encoder-latent variable-decoder structure, minimizing the reconstruction loss and KL divergence to enhance lesion details and reduce noise. Then, SAM combines the VAE results for automatic optimization, generating mask prompts using edge detection Canny to improve the segmentation accuracy. The system also introduces a "human-in-the-loop" mechanism where experts can manually adjust the enhancement area, optimize key details, and incorporate the feedback into the VAE and SAM to improve the model. Finally, through multiple rounds of optimization iterations, high-quality medical images are generated to enhance the diagnostic reliability.

[0024] Refer to Figure 1 , the medical image enhancement of the present invention is carried out according to the following steps: (1) Perform data enhancement processing on the medical image to be enhanced to generate an extended set of enhanced image candidates, specifically including: 1.1: Perform affine transformation (rotation, scaling, translation, shearing, etc.) on a single medical image to expand the data distribution and improve the generalization ability of the model; 1.2: Perform preprocessing such as contrast adjustment (histogram equalization, adaptive enhancement) and noise removal (Gaussian filtering, non-local means filtering) on the image to improve the image quality.

[0025] (2) Automatically select the best enhanced image from the candidate set as the initial input based on the image quality evaluation algorithm, specifically including: 2.1: Evaluate the image quality using indicators such as structural similarity, peak signal-to-noise ratio, and contrast sensitivity; 2.2: Select the image with the highest comprehensive score as the enhanced prompt image.

[0026] (3) Use the variational autoencoder for detail enhancement, specifically including: 3.1: The variational autoencoder model architecture includes an encoder, a latent variable space, and a decoder. The encoder converts the input medical image into a latent variable distribution, and the decoder reconstructs a high-quality enhanced image; 3.2: The training data includes public medical image databases covering BraTS (brain MRI), LIDC-IDRI (lung nodule CT), ISIC (skin cancer), and clinical data. The variational autoencoder model is optimized using adversarial loss and KL divergence loss so that it can learn the key details of medical images.

[0027] (4) SAM performs automatic preliminary enhancement, specifically including: 4.1: Further refine the image enhanced by the variational autoencoder using the vision foundation model to ensure global structural consistency; 4.2: Generate a preliminary enhanced image by fusing the enhanced results of the variational autoencoder and SAM through the attention mechanism.

[0028] (5) Experts perform feedback correction on the preliminary enhanced results through human-computer interaction to optimize the details of key regions, specifically including: 5.1: Experts mark areas with insufficient details or over-enhancement (such as blurred lesion boundaries, lost texture details, etc.) on the interaction interface; 5.2: The system automatically records the expert feedback and generates an expert feedback enhancement prompt for subsequent optimization.

[0029] (6) Input the expert feedback enhancement prompt into the variational autoencoder and the vision foundation model to iteratively optimize the image enhancement results, specifically including: 6.1: Use the feedback signal to adjust the parameters of the variational autoencoder decoder to better adapt to the medical image features; 6.2: Update the weights of the vision foundation model in combination with the expert feedback to strengthen the enhancement ability for specific medical features.

[0030] (7) Repeatedly execute steps 5) and 6) until the quality of the enhanced image reaches the preset threshold or meets the clinical requirements.

[0031] The present invention will be further described below in conjunction with specific examples and accompanying drawings. Example 1

[0032] Refer to Figure 2 , this example performs medical image enhancement according to the following steps: 1) Input the original medical image and perform data preprocessing and enhancement, including affine transformation, contrast adjustment, and noise removal, to optimize the usability of the image and improve the generalization ability of the model.

[0033] 2) The system selects the optimal enhanced image from the preprocessed images as the input for subsequent processing, and uses the variational autoencoder for detail enhancement. The VAE learns the deep features of medical images through the encoder-latent variable-decoder structure, while retaining key lesion information during denoising.

[0034] 3) The enhanced image is input into the vision foundation model SAM, and the Canny edge detection algorithm is combined to optimize the segmentation of the lesion area to further improve the clarity and recognizability of the target area.

[0035] 4) Clinical experts manually adjust the enhanced results through the interactive interface, select specific areas for optimization, and the system automatically records the feedback information corrected by the experts and re-enters it into VAE and SAM for the next round of enhancement iteration.

[0036] 5) When the image quality reaches the preset threshold or meets the clinical requirements, the enhancement process terminates, and the final optimized medical image is output. This image has better contrast, detail clarity, and lesion area recognizability, thus improving the diagnostic value of medical images.

[0037] The above embodiments are only for further illustration of the present invention and are not intended to limit the patent of the present invention. All equivalent implementations of the present invention should be included within the scope of the claims of the patent of the present invention.

Claims

1. A medical image enhancement method based on human-in-the-loop and vision foundation models, characterized in that, The method specifically includes the following steps: 1) Perform data augmentation on the medical images to generate an extended set of enhanced image candidates; 2) Automatically select the best enhanced image from the set of enhanced image candidates as the initial input based on an image quality evaluation algorithm; 3) Use a variational autoencoder for detail enhancement; 4) Use SA and Canny to parse medical images for automatic preliminary enhancement; 5) Feedback correction is performed on the preliminary enhancement results through the human-computer interaction of experts to optimize the details of key regions; 6) Input the expert feedback enhancement prompts into the variational autoencoder and the vision foundation model to iteratively optimize the image enhancement results; 7) Repeatedly execute steps 5) - 6) until the quality of the enhanced image reaches a preset threshold or meets clinical requirements.

2. The medical image enhancement method based on human-in-the-loop and vision foundation model according to claim 1, wherein The specific content of step 1) includes: 1.1: Perform affine transformations including rotation, scaling, translation, and shearing on a single medical image to expand the data distribution and improve the generalization ability of the model; 1.2: Perform preprocessing on the image for contrast adjustment and noise removal. The contrast adjustment includes histogram equalization and adaptive enhancement; the noise removal includes Gaussian filtering and non-local means filtering.

3. The medical image enhancement method based on human-in-the-loop and vision foundation model according to claim 1, wherein, The specific content of step 2) includes: 2.1: Use structural similarity, peak signal-to-noise ratio, and contrast sensitivity for image quality assessment; 2.2: Select the image with the highest comprehensive score as the enhanced prompt image.

4. The medical image enhancement method based on human-in-the-loop and vision foundation model according to claim 1, wherein The specific content of step 3) includes: 3.1: The variational autoencoder model architecture includes an encoder, a latent variable space, and a decoder. The encoder converts the input medical image into a latent variable distribution; the decoder reconstructs a high-quality enhanced image; 3.2: The training includes a public medical image database covering brain MRI, lung nodule CT, skin cancer, and clinical data. The variational autoencoder model is optimized using adversarial loss and KL divergence loss to enable it to learn the key details of medical images.

5. The medical image enhancement method based on the human-in-the-loop and vision-based foundation model according to claim 1, wherein The specific content of step 4) includes: 4.1: Use the vision foundation model to further refine the image enhanced by the variational autoencoder to ensure global structural consistency; 4.2: Generate a preliminary enhanced image by fusing the enhanced results of the variational autoencoder and SAM through an attention mechanism.

6. The medical image enhancement method based on the human-in-the-loop and vision foundation model according to claim 1, wherein The specific content of step 5) includes: 5.1: Experts mark areas with insufficient details such as blurred lesion boundaries and lost texture details or over-enhanced areas on the interactive interface; 5.2: The system automatically records the expert feedback and generates expert feedback enhancement prompts for subsequent optimization.

7. The medical image enhancement method based on human-in-the-loop and vision foundation model according to claim 1 or claim 4, characterized in that The specific content of step 3) includes: 6.1: Use the feedback signal to adjust the parameters of the variational autoencoder decoder to better adapt to the characteristics of medical images; 6.2: Update the weights of the vision foundation model in combination with expert feedback to strengthen the enhancement ability for specific medical features.