An image processing method, system, device and readable storage medium

By collecting multimodal image data through IoT devices and evaluating it using multi-stage progressive enhancement networks and bi-branch deep neural networks, the problems of uneven quality of medical image data and insufficient multimodal fusion are solved, generating high-quality images and providing comprehensive evaluation, thereby improving diagnostic accuracy and efficiency.

CN122158087APending Publication Date: 2026-06-05RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-03-18
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Medical imaging data varies greatly in quality, with serious noise and artifact issues affecting the accuracy and reliability of diagnosis. Existing technologies have failed to effectively address the problems of data scarcity and insufficient multimodal fusion.

Method used

Multimodal image data is collected using IoT devices. Training samples are expanded through synthesis and preprocessing. Quality enhancement is performed using a multi-stage progressive enhancement network and evaluated using a dual-branch deep neural network. Generative adversarial networks and conditional diffusion models are introduced to generate high-quality images. The training strategy is optimized by combining knowledge distillation and federated learning.

Benefits of technology

Generate high-quality images under scarce data conditions, provide comprehensive quality assessment, improve the diagnostic usability and accuracy of images, reduce dependence on labeled data, adapt to different equipment characteristics, and improve diagnostic efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122158087A_ABST
    Figure CN122158087A_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, system, device and readable storage medium. The method comprises the following steps: acquiring multi-modal medical image data, performing synthesis processing on the medical image data, obtaining synthesized medical image data, expanding training samples, and performing optimization training on a multi-stage progressive enhancement network by using the expanded training samples, then performing quality enhancement processing on the medical image data, and performing quality evaluation to obtain an evaluation result; if the evaluation result does not reach a standard, iterative optimization is continuously performed; if the evaluation result reaches the standard, the medical image data after quality enhancement is output. The technical effect of the application is that the training samples of the multi-stage progressive enhancement network are expanded, so that the multi-stage progressive enhancement network is optimized; the quality of the medical image data after quality enhancement is evaluated, so that iterative optimization is performed until the medical image data that reaches the standard is obtained, and the quality enhancement effect of the medical image data can be effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to an image processing method, system, device and readable storage medium. Background Technology

[0002] Medical imaging plays a crucial role in medical diagnosis. Multimodal imaging technologies such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), and PET (Positron Emission Tomography) provide doctors with a wealth of diagnostic information.

[0003] However, in practical applications, the quality of medical imaging data varies greatly, seriously affecting the accuracy and reliability of diagnosis. Noise and artifacts are particularly prominent issues. Images are susceptible to various types of noise interference, such as Gaussian noise and Poisson noise, while motion artifacts and metallic artifacts also frequently occur. These noises and artifacts blur image details, interfering with doctors' observation and judgment of lesions, leading to an increased risk of misdiagnosis or missed diagnosis.

[0004] In conclusion, how to effectively improve the quality of medical imaging data is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide an image processing method, system, device, and readable storage medium to enhance the quality of medical image data.

[0006] To solve the above-mentioned technical problems, this application provides the following technical solution: An image processing method, comprising: Multimodal medical image data is collected using Internet of Things (IoT) devices, and the medical image data is synthesized to obtain synthetic medical image data. The training samples are expanded using the synthetic medical image data, and the expanded training samples are used to optimize the training of the multi-stage progressive enhancement network. The medical image data is enhanced using an optimized multi-stage progressive enhancement network, and the enhanced medical image data is then evaluated to obtain the evaluation results. If the evaluation results do not meet the standards, the current multi-stage progressive augmentation network will be optimized and trained to improve the quality augmentation effect. If the evaluation results meet the standards, the enhanced medical imaging data will be output.

[0007] Preferably, after collecting multimodal medical image data using IoT devices, the method further includes: The medical image data is preprocessed by at least one of the following data processing operations: data cleaning, standardization, outlier detection and removal. Accordingly, the medical image data is processed by synthesis, including: The preprocessed medical image data is then synthesized.

[0008] Preferably, the medical image data is subjected to synthesis processing to obtain synthesized medical image data, including: The medical image data is subjected to geometric transformation and / or grayscale transformation to obtain the synthetic medical image data; Alternatively, the medical image data can be synthesized using a conditional generative adversarial network or a conditional diffusion model to obtain the synthesized medical image data.

[0009] Preferably, the medical image data is enhanced using an optimized multi-stage progressive enhancement network, including: Multi-scale image features of the medical image data are extracted using the feature pyramid network in the optimized multi-stage progressive enhancement network. The offset is learned from the feature map using the deformable convolutional layer in the optimized multi-stage progressive enhancement network. The anatomical structure in the optimized multi-stage progressive augmentation network is used to maintain constraints, and the augmentation process is guided by an organ mask generated by the segmentation network to ensure the accuracy of the anatomical structure.

[0010] Preferably, the enhanced medical imaging data undergoes a quality assessment to obtain assessment results, including: A dual-branch deep neural network architecture was used to assess the quality of enhanced medical image data, and the assessment results were obtained. Alternatively, a trained language model can be used to assess the quality of the enhanced medical image data to obtain the assessment results.

[0011] Preferably, a dual-branch deep neural network architecture is used to perform quality assessment on the enhanced medical image data to obtain the assessment results, including: The local feature extraction branch and the global feature extraction branch in the dual-branch deep neural network architecture are used to extract the local features and global features of the enhanced medical image data, respectively. The local and global features are fused using a cross-attention mechanism to obtain noise level scores, artifact severity ratings, and diagnostic usability probabilities. The noise level score, the artifact severity rating, and the diagnostic availability probability are determined as the evaluation results.

[0012] Preferably, it further includes: Using interpretive visualization components, a heatmap corresponding to the evaluation results is generated to mark areas of quality defects.

[0013] An image processing system, comprising: The data acquisition and preprocessing module is used to acquire multimodal medical image data using IoT devices and to perform synthetic processing on the medical image data to obtain synthetic medical image data. An iterative optimization module is used to expand the training samples using the synthetic medical image data, and to optimize the training of the multi-stage progressive enhancement network using the expanded training samples. The quality enhancement and evaluation module is used to perform quality enhancement processing on the medical image data using an optimized multi-stage progressive enhancement network, and to evaluate the quality of the enhanced medical image data to obtain the evaluation results. An iterative enhancement module is used to optimize and train the current multi-stage progressive enhancement network to improve the quality enhancement effect if the evaluation result does not meet the standard. The results output module is used to output enhanced medical image data if the evaluation results meet the standards.

[0014] An electronic device, comprising: Memory, used to store computer programs; A processor is used to implement the steps of the above-described image processing method when executing the computer program.

[0015] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described image processing method.

[0016] The method provided in this application involves using an IoT device to collect multimodal medical image data and performing synthetic processing on the medical image data to obtain synthetic medical image data. The synthetic medical image data is then used to expand the training samples, and the expanded training samples are used to optimize and train a multi-stage progressive augmentation network. The optimized multi-stage progressive augmentation network is then used to perform quality enhancement processing on the medical image data, and the quality of the enhanced medical image data is evaluated to obtain an evaluation result. If the evaluation result does not meet the standard, the current multi-stage progressive augmentation network is optimized and trained to improve the quality enhancement effect. If the evaluation result meets the standard, the enhanced medical image data is output.

[0017] In this application, multimodal medical image data can be collected based on IoT devices. Considering the scarcity of training samples for medical image data, this application performs synthetic processing on the medical image data to obtain synthetic medical image data, thus expanding the training samples. Then, based on the expanded training samples, a multi-stage progressive augmentation network is optimized and trained, making the multi-stage progressive augmentation network more effective. The optimized multi-stage progressive augmentation network is then used to perform quality enhancement processing on the medical image data, improving its quality. To ensure that the final output medical image data has high quality, a quality assessment can be performed after augmentation to obtain the assessment results. If the assessment results do not meet the standard, the current multi-stage progressive augmentation network is further optimized and trained to improve the quality enhancement effect until the quality enhancement meets the standard; if the assessment results meet the standard, the quality-enhanced medical image data is output.

[0018] The technical effect of this application is that by expanding the training samples of the multi-stage progressive enhancement network, the multi-stage progressive enhancement network is optimized; by evaluating the quality of the enhanced medical image data, iterative optimization is carried out until qualified medical image data is obtained, which can effectively ensure the quality enhancement effect of medical image data.

[0019] Accordingly, embodiments of this application also provide an image processing system, device, and readable storage medium corresponding to the above-described image processing method, which have the aforementioned technical effects, and will not be repeated here. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an implementation of an image processing method in this application. Figure 2 This is a schematic diagram of the structure of an image processing system according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application; Figure 4 This is a schematic diagram of the specific structure of an electronic device in an embodiment of this application. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Please refer to Figure 1 , Figure 1 This is a flowchart of an image processing method according to an embodiment of this application, which includes the following steps.

[0024] S101. Collect multimodal medical image data using IoT devices, and perform synthesis processing on the medical image data to obtain synthetic medical image data.

[0025] In this embodiment, IoT devices can be used to collect medical image data of different modalities. Medical image data can be a single image or multiple images. Multimodal medical image data includes, but is not limited to, the following types: CT (Computed Tomography): Covers plain and enhanced scan images of different parts (such as the head, chest, and abdomen), with a slice thickness ranging from 0.5 to 5 mm; MRI (Magnetic Resonance Imaging): Includes T1-weighted, T2-weighted, FLAIR, DWI sequences, etc., supporting field strengths of 1.5T and 3.0T; PET-CT (Positron Emission Tomography / CT Fusion Imaging): Used for the fusion analysis of metabolic function and anatomical structure; Ultrasound imaging: such as B-mode ultrasound, color Doppler ultrasound, and other dynamic sequences; X-ray images: including digital radiography (DR) and computed radiography (CR); Digital subtraction angiography (DSA): used for high-resolution imaging of vascular structures; Specialized modalities such as ophthalmic OCT and dermoscopy imaging.

[0026] Given the scarcity of training samples for medical image data, in order to increase sample diversity, medical image data can be synthesized to obtain synthetic medical image data, thereby expanding the training samples.

[0027] In one specific embodiment of this application, medical image data is synthesized to obtain synthetic medical image data, including: Perform geometric transformations and / or grayscale transformations on medical image data to obtain synthetic medical image data; Alternatively, conditional generative adversarial networks or conditional diffusion models can be used to synthesize medical image data to obtain synthetic medical image data.

[0028] In other words, at least one of the following operations can be performed on the acquired medical image data: geometric transformation (such as rotation ±15°, scaling 0.9-1.1 times) and grayscale transformation (histogram equalization, contrast adjustment), thereby improving data diversity.

[0029] Conditional Generative Adversarial Networks (GANs) can also be used to synthesize diverse samples from medical image data. The generator uses a U-Net architecture, and the discriminator uses PatchGAN, both trained stably using spectral normalization techniques.

[0030] Diffusion models can also be used for medical image enhancement. Diffusion models generate high-quality images through a progressive denoising process, potentially producing more diverse results under conditions of scarce data. In practice, conditional diffusion models can be used to generate high-quality images from low-quality images. During training, denoising score matching techniques are employed to optimize the model's performance with limited data.

[0031] In one specific embodiment of this application, after acquiring multimodal medical image data using an Internet of Things (IoT) device, the method further includes: Preprocessing of medical imaging data involves at least one of the following data processing operations: data cleaning, standardization, outlier detection and removal. Accordingly, the medical image data undergoes synthetic processing, including: The preprocessed medical image data is then synthesized.

[0032] In other words, to avoid inherent defects in medical image data collected by IoT devices, the collected medical image data can be preprocessed. This preprocessing includes, but is not limited to: Data cleaning, such as using the DICOM protocol to parse metadata, automatically identify and correct differences in device parameters (such as window width / window level, layer thickness, and pixel spacing). Standardization involves normalizing the Hounsfield units of CT images (e.g., truncating the range [-1000, 3000]), while MRI images are corrected using an N4 bias field to eliminate magnetic field inhomogeneities. Outlier detection and removal, such as detecting motion artifacts, metal artifacts and other serious quality problems based on SVM or isolated forest algorithms, and automatically marking or removing low-quality data.

[0033] Accordingly, after preprocessing the medical image data, the preprocessed medical image data can be synthesized.

[0034] Specifically, multimodal medical image data is collected under conditions of data scarcity and intelligent preprocessing is performed. Sensor calibration algorithms can be used to generate real-time data streams, and time-series alignment algorithms can be used to synchronize sensor data. To address the scarcity of data, generative data augmentation techniques can be employed, using Conditional Generative Adversarial Networks (GANs) to generate synthetic image data and expand the training sample set. Preprocessing steps include histogram equalization for brightness adjustment, median filtering for noise reduction, and Laplacian sharpening for image detail enhancement.

[0035] The sensor calibration algorithm includes the following steps: Step 1, Time Series Alignment: Synchronize asynchronous data acquired by multiple sensors based on the Dynamic Time Warping (DTW) algorithm to ensure the temporal consistency of the image sequence.

[0036] Step 2, Spatial Registration: Align the spatial coordinates of the multimodal image using mutual information maximization or feature-based registration methods (such as SIFT, SURF).

[0037] Step 3, Parameter Calibration: Establish a mapping model between the device sensor output and the real physical quantities (such as CT value, MR signal intensity) through linear regression or nonlinear optimization (such as the Levenberg-Marquardt algorithm), and update the calibration parameters regularly.

[0038] S102. Expand the training samples using synthetic medical image data, and optimize the training of the multi-stage progressive augmentation network using the expanded training samples.

[0039] After acquiring synthetic medical image data, the training samples can be augmented based on this data. The augmented training samples can then be used to optimize the training of the multi-stage progressive augmentation network, thereby improving its quality enhancement performance.

[0040] Specifically, in environments with scarce data, knowledge distillation techniques can be introduced, using pre-trained large models as teacher models to guide small student models in learning multimodal image feature representations under limited data conditions. Furthermore, device-adaptive enhancement strategies can be established, such as creating a mapping database between CT, MRI, and PET device parameters and enhancement strategies, allowing for the selection of appropriate enhancement strategies based on the characteristics of different devices.

[0041] In coordinating the overall training process, the optimization of the quality assessment loss function and the image enhancement loss function can be coordinated through dynamic weight allocation. In the early stages of training, the focus is on quality assessment accuracy; in the mid-stage, a balanced optimization is implemented; and in the later stages, the focus shifts to enhancement effectiveness. Online correction achieves a real-time feedback mechanism, iteratively optimizing the enhancement parameters based on the quality assessment results of the enhanced image, establishing a parameter mapping matrix, and setting iteration termination conditions (such as a quality score Q ≥ 0.85 or the number of iterations ≥ 5).

[0042] In this application, a distributed training scheme under federated learning can also be adopted. That is, to address data scarcity while protecting patient privacy, a federated learning framework can be used. Each medical institution trains its model locally, exchanging only model parameters, not raw data. A central server aggregates the local models to generate a global model. This method is particularly suitable for medical research across multiple institutions, fully utilizing dispersed medical data resources while protecting data privacy.

[0043] Furthermore, for scenarios with limited computing resources, a hybrid approach combining traditional image processing techniques with deep learning can be adopted. For example, wavelet denoising can be used for initial enhancement, followed by a lightweight deep learning network for fine-tuning. This significantly reduces computing resource requirements while maintaining performance, making it suitable for edge computing device deployments.

[0044] S103. Use the optimized multi-stage progressive enhancement network to perform quality enhancement processing on medical image data, and evaluate the quality of the enhanced medical image data to obtain the evaluation results.

[0045] To improve the quality optimization of medical image data, this embodiment employs a multi-stage progressive enhancement network to perform quality enhancement processing on the medical image data. Its core objectives are to optimize noise suppression, detail enhancement, and structure preservation.

[0046] In one specific embodiment of this application, medical image data is enhanced using an optimized multi-stage progressive enhancement network, including: Multi-scale image features of medical image data are extracted using the feature pyramid network in the optimized multi-stage progressive enhancement network. The offset is learned from the feature map by utilizing the deformable convolutional layer in the optimized multi-stage progressive enhancement network. The anatomical structure in the optimized multi-stage progressive augmentation network is kept constrained, and the augmentation process is guided by an organ mask generated by the segmentation network to ensure the accuracy of the anatomical structure.

[0047] In other words, a multi-stage progressive enhancement architecture allows each stage to include dense residual blocks and channel attention modules. For environments with scarce data, few-shot learning techniques can be integrated, enabling the model to quickly adapt to limited data scenarios through meta-learning strategies. Specifically, the optimized multi-stage progressive enhancement network is used to enhance the quality of medical image data, including: utilizing a feature pyramid network to extract multi-scale image features and fully utilize hierarchical information in limited data; deformable convolutional layers to learn offsets based on feature maps, adaptively adjusting the receptive field to better capture lesion morphological changes; and anatomical structure preservation constraints to guide the enhancement process through organ masks generated by a segmentation network, ensuring the accuracy of the anatomical structure in the enhanced image.

[0048] After the medical image data undergoes quality enhancement processing, it can be evaluated to obtain the evaluation results.

[0049] In one specific embodiment of this application, the quality of enhanced medical image data is assessed to obtain assessment results, including: A dual-branch deep neural network architecture was used to assess the quality of enhanced medical image data, and the assessment results were obtained. Alternatively, a trained language model can be used to assess the quality of the enhanced medical imaging data and obtain the assessment results.

[0050] That is, large language models can be integrated to generate image assessment reports and describe image quality. The language model can generate descriptive quality reports based on image features and provide improvement suggestions. A multimodal learning framework can be used to map image features and text descriptions to the same semantic space, enabling interpretable image quality assessment.

[0051] Alternatively, a two-branch deep neural network architecture can be directly used to assess the quality of enhanced medical image data, yielding evaluation results. Specifically, the evaluation results obtained by using a two-branch deep neural network architecture to assess the quality of enhanced medical image data include: The local feature extraction branch and the global feature extraction branch in the dual-branch deep neural network architecture are used to extract the local and global features of the enhanced medical image data, respectively. By using a cross-attention mechanism to fuse local and global features, noise level scores, artifact severity ratings, and diagnostic usability probabilities are obtained. The noise level score, artifact severity rating, and diagnostic availability probability were used as the evaluation results.

[0052] In this embodiment, a dual-branch deep neural network architecture, including a quality assessment branch and an image enhancement branch, can be used for comprehensive quality assessment of the generated images. The quality assessment branch fuses global and local features, extracting global and local features respectively through a 3DVision Transformer and a convolutional neural network, and then fusing them through a channel attention mechanism. The output includes a three-dimensional assessment result comprising a noise level score (0-1), an artifact severity rating (I-IV), and a diagnostic usability probability.

[0053] In one specific embodiment of this application, the method further includes: using an interpretable visualization component to generate a heatmap corresponding to the evaluation results to mark areas of quality defects. That is, interpretable visualization results, such as quality heatmaps and defect annotation maps, can be provided to assist doctors in understanding the image quality status. For example, by integrating an interpretable visualization component, a heatmap can be generated to mark major quality defect areas, and an automatic alarm can be triggered when severe artifacts are detected or the diagnostic usability probability is lower than a threshold.

[0054] S104. If the evaluation results do not meet the standards, the current multi-stage progressive augmentation network will be optimized and trained to improve the quality augmentation effect.

[0055] In this embodiment, the evaluation results of the output quality-enhanced medical image data can be pre-set. If the evaluation results do not meet the standards, the current multi-stage progressive enhancement network can be optimized and trained to improve the quality enhancement effect. Then, the medical image data can be iteratively enhanced until the evaluation results meet the standards.

[0056] S105. If the evaluation results meet the standards, output the enhanced quality medical imaging data.

[0057] Once the evaluation results of medical imaging data meet the standards, enhanced medical imaging data can be directly output. This ensures the quality of medical imaging data and avoids radiologists making diagnoses based on poor-quality medical imaging data.

[0058] The method provided in this application involves using an IoT device to collect multimodal medical image data and performing synthetic processing on the medical image data to obtain synthetic medical image data. The synthetic medical image data is then used to expand the training samples, and the expanded training samples are used to optimize and train a multi-stage progressive augmentation network. The optimized multi-stage progressive augmentation network is then used to perform quality enhancement processing on the medical image data, and the quality of the enhanced medical image data is evaluated to obtain an evaluation result. If the evaluation result does not meet the standard, the current multi-stage progressive augmentation network is optimized and trained to improve the quality enhancement effect. If the evaluation result meets the standard, the enhanced medical image data is output.

[0059] In this application, multimodal medical image data can be collected based on IoT devices. Considering the scarcity of training samples for medical image data, this application performs synthetic processing on the medical image data to obtain synthetic medical image data, thus expanding the training samples. Then, based on the expanded training samples, a multi-stage progressive augmentation network is optimized and trained, making the multi-stage progressive augmentation network more effective. The optimized multi-stage progressive augmentation network is then used to perform quality enhancement processing on the medical image data, improving its quality. To ensure that the final output medical image data has high quality, a quality assessment can be performed after augmentation to obtain the assessment results. If the assessment results do not meet the standard, the current multi-stage progressive augmentation network is further optimized and trained to improve the quality enhancement effect until the quality enhancement meets the standard; if the assessment results meet the standard, the quality-enhanced medical image data is output.

[0060] The technical effect of this application is that by expanding the training samples of the multi-stage progressive enhancement network, the multi-stage progressive enhancement network is optimized; by evaluating the quality of the enhanced medical image data, iterative optimization is carried out until qualified medical image data is obtained, which can effectively ensure the quality enhancement effect of medical image data.

[0061] Corresponding to the above method embodiments, this application also provides an image processing system, and the image processing system described below can be referred to in correspondence with the image processing method described above.

[0062] See Figure 2 As shown, the system includes the following modules: The data acquisition and preprocessing module 101 is used to acquire multimodal medical image data using IoT devices and to perform synthetic processing on the medical image data to obtain synthetic medical image data. The iterative optimization module 102 is used to expand the training samples using synthetic medical image data and to optimize the training of the multi-stage progressive enhancement network using the expanded training samples. The quality enhancement and evaluation module 103 is used to perform quality enhancement processing on medical image data using an optimized multi-stage progressive enhancement network, and to evaluate the quality of the enhanced medical image data to obtain the evaluation results. The iterative enhancement module 104 is used to optimize the current multi-stage progressive enhancement network to improve the quality enhancement effect if the evaluation result does not meet the standard. The results output module 105 is used to output enhanced medical image data if the evaluation results meet the standards.

[0063] The system provided in this application uses IoT devices to collect multimodal medical image data and performs synthetic processing on the medical image data to obtain synthetic medical image data. The synthetic medical image data is used to expand the training samples, and the expanded training samples are used to optimize and train a multi-stage progressive augmentation network. The optimized multi-stage progressive augmentation network is used to perform quality enhancement processing on the medical image data, and the quality of the enhanced medical image data is evaluated to obtain an evaluation result. If the evaluation result does not meet the standard, the current multi-stage progressive augmentation network is optimized and trained to improve the quality enhancement effect. If the evaluation result meets the standard, the enhanced medical image data is output.

[0064] In this application, multimodal medical image data can be collected based on IoT devices. Considering the scarcity of training samples for medical image data, this application performs synthetic processing on the medical image data to obtain synthetic medical image data, thus expanding the training samples. Then, based on the expanded training samples, a multi-stage progressive augmentation network is optimized and trained, making the multi-stage progressive augmentation network more effective. The optimized multi-stage progressive augmentation network is then used to perform quality enhancement processing on the medical image data, improving its quality. To ensure that the final output medical image data has high quality, a quality assessment can be performed after augmentation to obtain the assessment results. If the assessment results do not meet the standard, the current multi-stage progressive augmentation network is further optimized and trained to improve the quality enhancement effect until the quality enhancement meets the standard; if the assessment results meet the standard, the quality-enhanced medical image data is output.

[0065] The technical effect of this application is that by expanding the training samples of the multi-stage progressive enhancement network, the multi-stage progressive enhancement network is optimized; by evaluating the quality of the enhanced medical image data, iterative optimization is carried out until qualified medical image data is obtained, which can effectively ensure the quality enhancement effect of medical image data.

[0066] In one specific embodiment of this application, the data acquisition and preprocessing module is further configured to preprocess the medical image data after acquiring multimodal medical image data using IoT devices by performing at least one of the following data processing operations: data cleaning, standardization, outlier detection and removal; correspondingly, the medical image data is synthesized, including: synthesizing the preprocessed medical image data.

[0067] In one specific embodiment of this application, the data acquisition and preprocessing module is specifically used to perform geometric transformation and / or grayscale transformation on medical image data to obtain synthetic medical image data. Alternatively, conditional generative adversarial networks or conditional diffusion models can be used to synthesize medical image data to obtain synthetic medical image data.

[0068] In one specific embodiment of this application, the quality enhancement and evaluation module is specifically used to extract multi-scale image features of medical image data using the feature pyramid network in the optimized multi-stage progressive enhancement network. The offset is learned from the feature map by utilizing the deformable convolutional layer in the optimized multi-stage progressive enhancement network. The anatomical structure in the optimized multi-stage progressive augmentation network is kept constrained, and the augmentation process is guided by an organ mask generated by the segmentation network to ensure the accuracy of the anatomical structure.

[0069] In one specific embodiment of this application, the quality enhancement and evaluation module is specifically used to perform quality evaluation on the enhanced medical image data using a dual-branch deep neural network architecture to obtain evaluation results; Alternatively, a trained language model can be used to assess the quality of the enhanced medical imaging data and obtain the assessment results.

[0070] In one specific embodiment of this application, the quality enhancement and evaluation module is specifically used to extract local features and global features of the enhanced medical image data by utilizing the local feature extraction branch and the global feature extraction branch in the dual-branch deep neural network architecture, respectively. By using a cross-attention mechanism to fuse local and global features, noise level scores, artifact severity ratings, and diagnostic usability probabilities are obtained. The noise level score, artifact severity rating, and diagnostic availability probability were used as the evaluation results.

[0071] In one specific embodiment of this application, the quality enhancement and evaluation module is further used to generate a heat map corresponding to the evaluation results using an interpretive visualization component to mark areas of quality defects.

[0072] To facilitate those skilled in the art to better understand and implement the image processing method provided in the embodiments of this application, the image processing method will be described in detail below with specific scenario examples.

[0073] Artificial intelligence-based medical image quality enhancement still faces significant challenges in environments with scarce data, with the following specific shortcomings: 1. Strong data dependence: Deep learning methods require a large amount of labeled data for training, while medical image data labeling relies on professional radiologists, which is time-consuming and labor-intensive, resulting in a scarcity of high-quality labeled data. This data scarcity problem is particularly prominent in rare diseases and special imaging modalities.

[0074] 2. High risk of overfitting: In situations where data is scarce, complex generative models (such as generative adversarial networks) are prone to overfitting, resulting in insufficient diversity of generated images and severe loss of details, which affects the clinical diagnostic value.

[0075] 3. Inadequate quality assessment: Medical image quality assessment relies heavily on manual evaluation, which is inefficient and highly susceptible to subjective factors. It is difficult to accurately quantify image quality, especially when dealing with complex image features.

[0076] 4. Insufficient multimodal fusion: The multimodal medical image fusion is relatively simple and traditional, and the diversity and authenticity of the data are not fully considered, resulting in the fused data losing specific information of the original modality to some extent.

[0077] Some studies have used generative adversarial networks (GANs) for medical image enhancement, but their performance has significantly declined under conditions of scarce data; other studies have used convolutional neural networks (CNNs) for image quality assessment, but lack specific evaluation metrics for the generated images.

[0078] Inspired Medical's RaDyn series of products (such as RaDynCT and RaDynMR) can improve the quality of CT and MR images through artificial intelligence technology, but their core algorithms still rely on a large amount of clinical data, and their effectiveness is limited in scenarios with scarce data. Another method for assessing and enhancing radiological image quality based on deep learning adopts a dual-branch network structure, but it requires a large amount of data and is not suitable for environments with scarce data.

[0079] This application finds that the shortcomings of these technologies lie in their failure to fully consider the real-world challenges of scarce medical imaging data, the lack of dedicated enhancement and evaluation schemes for scarce data, and the inability to effectively utilize limited data resources while ensuring diagnostic reliability.

[0080] To address the aforementioned problems, this application provides an image processing method and system, aiming to overcome the shortcomings of AIGC quality enhancement and evaluation of medical images in environments with scarce data, and to provide a system capable of generating high-quality medical images and assessing reliability under limited data conditions. This method includes, but is not limited to, the following technical effects: Improving the quality of image generation under scarce data: Through innovative generative models and training strategies, high-quality, high-fidelity medical images are generated under conditions of limited data. Establish a comprehensive quality assessment system: provide quantifiable quality assessment indicators to evaluate the noise level, artifact severity, and diagnostic usability of generated images from multiple dimensions; Achieve multimodal image fusion and enhancement: effectively integrate multimodal image information such as CT, MRI, and PET to enhance the diagnostic value of images; Low dependence on labeled data: Develop algorithms that are adapted to small sample learning, reduce dependence on large amounts of labeled data, and improve the applicability of the method.

[0081] In practical applications, different methods and steps can be divided into different functional modules. For example, in practical applications, compared to... Figure 2 The module division shown can also be further divided into the following modules: This module is designed for medical image acquisition and preprocessing in data-scarce environments. It collects multimodal medical image data and performs intelligent preprocessing under these conditions. The module integrates IoT technology, employs sensor calibration algorithms to generate real-time data streams, and synchronizes sensor data using time-series alignment algorithms. Addressing the characteristics of scarce data, this module utilizes generative data augmentation techniques, generating synthetic image data through Conditional Generative Adversarial Networks (GANs) to expand the training sample set. Preprocessing steps include histogram equalization for brightness adjustment, median filtering for noise reduction, and Laplacian sharpening for image detail enhancement.

[0082] The generative image enhancement module for scarce data is the core enhancement component of the system. It adopts a multi-stage progressive enhancement architecture, with each stage containing dense residual blocks and channel attention modules. For scarce data environments, the module integrates few-shot learning techniques, using a meta-learning strategy to enable the model to quickly adapt to limited data scenarios.

[0083] The quality assessment and reliability analysis module provides a comprehensive quality assessment of the generated images. It employs a dual-branch deep neural network architecture, including a quality assessment branch and an image enhancement branch. The quality assessment branch fuses global and local features, extracting global and local features separately through a 3D Vision Transformer and a convolutional neural network, and then fusing them using a channel attention mechanism. The output includes a three-dimensional assessment result comprising a noise level score (0-1), artifact severity rating (I-IV), and diagnostic usability probability. The module also integrates an interpretable visualization component, generating heatmaps to annotate key quality defect areas and triggering automatic alarms when severe artifacts are detected or the diagnostic usability probability falls below a threshold.

[0084] The multimodal image fusion and knowledge distillation module employs a generative adversarial network (GAN) algorithm for multimodal medical image data fusion. In environments with scarce data, knowledge distillation technology is introduced, using a pre-trained large model as a teacher model to guide smaller student models in learning multimodal image feature representations under limited data conditions. The module also establishes an equipment-adaptive enhancement strategy, creating a mapping database between CT, MRI, and PET equipment parameters and enhancement strategies, selecting appropriate enhancement strategies based on the characteristics of different equipment.

[0085] The adaptive joint training and online correction module coordinates the overall training process of the system. It uses a dynamic weight allocation module to coordinate the optimization of the quality assessment loss function and the image enhancement loss function. In the early stages of training, the focus is on quality assessment accuracy; in the mid-stage, a balanced optimization is implemented; and in the later stages, the focus shifts to enhancement effectiveness. The online correction module implements a real-time feedback mechanism, iteratively optimizing the enhancement parameters based on the quality assessment results of the enhanced image, establishing a parameter mapping matrix, and setting iteration termination conditions (such as a quality score Q ≥ 0.85 or ≥ 5 iterations).

[0086] As can be seen, this application can improve the quality of image generation under scarce data. By employing generative data augmentation and few-shot learning techniques, the problem of data scarcity is effectively alleviated. A multi-stage progressive augmentation architecture is adopted, combining dense residual blocks and channel attention mechanisms to generate high-quality medical images under limited data conditions. Experimental results show that the system achieves a structural similarity index (SSIM) of over 0.93 for generated images using only 10% of the training data, significantly higher than the 0.75-0.85 of traditional methods. The generated images exhibit excellent performance in spatial resolution, noise suppression, and detail preservation, meeting the needs of clinical diagnosis.

[0087] This system provides comprehensive and reliable quality assessment, introducing a multi-dimensional quality assessment system that quantitatively evaluates generated image quality from multiple perspectives, including noise level, artifact severity, and diagnostic usability. The quality assessment branch employs a fusion architecture of 3D VisionTransformer and convolutional neural networks, capable of simultaneously capturing global contextual features and local texture features of the image. The assessment results achieve a 92% consistency with radiologist evaluations, significantly higher than the 70%-80% of traditional methods. The system also provides interpretable heatmaps, intuitively marking areas with quality issues to assist physicians in diagnostic decision-making.

[0088] To enhance clinical applicability and diagnostic value, the system is designed to closely align with clinical needs, featuring adaptive enhancement and multi-protocol output capabilities. It can adapt to the characteristics of different imaging devices, simultaneously generating standard window images and lesion-focused enhanced images. For low-dose CT images, the system employs noise distribution estimation-guided targeted enhancement, ensuring diagnostic image quality while reducing radiation dose. Clinical trials show that the enhanced images produced by the system can improve radiologists' diagnostic accuracy by 15% and shorten diagnostic time by 30%.

[0089] By reducing reliance on labeled data, the system significantly reduces its dependence on large amounts of labeled data through advanced techniques such as meta-learning and knowledge distillation. Using only 100 labeled samples, the system's performance can reach the level of a traditional model trained with 1000 samples. This characteristic makes the invention particularly suitable for data-scarce scenarios such as rare diseases and special imaging modalities, demonstrating broad applicability and significant clinical value.

[0090] For example, in a real-world application environment, the system's composition and connection relationships are as follows: In practical implementation, hardware platforms and software algorithms can work together. The hardware platform includes medical image acquisition equipment (such as CT, MRI, PET, etc.), high-performance computing servers (equipped with GPU accelerator cards), and medical-grade display devices. The software algorithm is implemented based on the Python and PyTorch deep learning frameworks and includes the following core modules: The data acquisition and preprocessing module communicates with medical imaging equipment via the DICOM protocol to acquire multimodal image data. To address the issue of scarce data, the module integrates a Conditional Generative Adversarial Network (GAN) for data augmentation. The generator network uses a U-Net architecture, and the discriminator uses a PatchGAN structure. During training, Spectral Normalization is used to stabilize the training process and prevent modality collapse.

[0091] For medical images of different modalities, the module performs grayscale normalization, mapping pixel values ​​to the [0,1] range. To address the non-uniformity issue in MR images, the N4 bias field correction algorithm is used for preprocessing. For CT images, the module automatically identifies different window widths and levels and standardizes them into a unified format.

[0092] The generative image enhancement module employs a multi-stage progressive enhancement architecture, consisting of three sequentially connected enhancement stages. Each stage contains four dense residual blocks with a growth rate of 32. Each residual block uses a 3×3×3 convolutional kernel and integrates a channel attention mechanism to adaptively adjust the weights of each feature channel.

[0093] For training on scarce data, the module employs a meta-learning approach, specifically based on the MAML (Model-Agnostic Meta-Learning) algorithm. During training, the model is trained alternately across different clinical scenarios (such as different disease types and different imaging devices), enabling the model to quickly adapt to new scarce data scenarios.

[0094] The quality assessment module employs a dual-branch network architecture. The global feature extraction branch uses a 3D VisionTransformer with a patch size of 16×16×16. The local feature extraction branch uses a densely connected CNN with 7×7×7 convolutional kernels. Features from both branches are fused through a cross-attention mechanism, ultimately outputting a noise level score, artifact severity rating, and diagnostic usability probability. The module training uses a multi-task loss function, including cross-entropy loss for quality assessment, structural similarity loss for enhanced images, and anatomical distortion constraints. A dynamic weight adjustment strategy is used, prioritizing quality assessment accuracy in the early stages of training (α=0.9) and optimizing enhancement effects in later stages (β=0.7).

[0095] The following steps can be implemented in this system: Step 1: Data acquisition and preprocessing.

[0096] Multimodal medical image data is collected via IoT devices, and a sensor calibration algorithm is used to generate a real-time data stream. The collected data is then time-series aligned to generate synchronized image data. An SVM anomaly detection algorithm is used to filter out outliers, resulting in a high-quality medical image dataset.

[0097] Step 2: Generative data augmentation.

[0098] For scenarios with scarce data, a conditional generative adversarial network (GAN) is used to generate synthetic image data, expanding the training samples. Anatomical constraints are introduced during the generation process to ensure the anatomical plausibility of the generated images. Data augmentation includes operations such as rotation, scaling, and elastic deformation to improve the model's generalization ability.

[0099] Step 3: Multi-stage image enhancement.

[0100] The preprocessed image is input into a multi-stage progressive enhancement network to gradually improve image quality. Each enhancement stage focuses on a different quality improvement task, such as noise suppression, resolution enhancement, and contrast enhancement. The network is trained end-to-end to optimize the overall enhancement effect.

[0101] Step 4: Quality assessment and reliability analysis.

[0102] The enhanced images are evaluated for quality, generating a structured quality report. Evaluation metrics include noise level, artifact severity, and diagnostic usability. When the evaluation results fail to meet preset standards, an online correction mechanism is triggered to iteratively optimize the enhancement parameters.

[0103] Step 5: Results output and visualization.

[0104] The system integrates enhanced images with quality assessment results, supporting the DICOM standard format. It provides interpretable visualizations, such as quality heatmaps and defect annotation maps, to help physicians understand the image quality status.

[0105] Furthermore, to verify the technical effectiveness of the solution provided in this application, an experiment was conducted using a set of scarce data. The experimental data consisted of 300 low-dose CT images, of which only 100 were annotated by experts. After comparison, it was found that in terms of image quality, the peak signal-to-noise ratio (PSNR) of the CT images generated using the image processing method provided in this application reached 38.2 dB, which is 4.5 dB higher than related deep learning methods. The structural similarity index (SSIM) was 0.94, significantly higher than the 0.82 of related methods.

[0106] In a clinical evaluation, three radiologists blind-reviewed the enhanced images. The results showed that the diagnostic usability score of the enhanced images obtained using this invention reached 4.5 out of 5, while the original low-dose images scored only 2.8. Physicians achieved a diagnostic accuracy of 92% for the enhanced images based on this invention, approaching 95% for standard-dose CT images.

[0107] The superior performance of this application in a scarce data environment makes it of great application value in primary healthcare institutions and the field of rare disease diagnosis, and can significantly improve the diagnostic quality and efficiency of medical images.

[0108] Corresponding to the above method embodiments, this application also provides an electronic device. The electronic device described below and the image processing method described above can be referred to in correspondence.

[0109] See Figure 3 As shown, the electronic device includes: Memory 332 is used to store computer programs; The processor 322 is used to implement the steps of the image processing method of the above method embodiments when executing a computer program.

[0110] For details, please refer to Figure 4 , Figure 4This is a schematic diagram of the specific structure of an electronic device provided in this embodiment. The electronic device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 can be temporary or permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332 and execute the series of instruction operations stored in the memory 332 on the electronic device 301.

[0111] Electronic device 301 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.

[0112] The steps in the image processing method described above can be implemented by the structure of an electronic device.

[0113] Corresponding to the above method embodiments, this application also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the image processing method described above.

[0114] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the image processing method described in the above method embodiments.

[0115] The readable storage medium can specifically be a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or any other readable storage medium capable of storing program code.

[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0117] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0119] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0120] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An image processing method, characterized in that, include: Multimodal medical image data is collected using Internet of Things (IoT) devices, and the medical image data is synthesized to obtain synthetic medical image data. The training samples are expanded using the synthetic medical image data, and the expanded training samples are used to optimize the training of the multi-stage progressive enhancement network. The medical image data is enhanced using an optimized multi-stage progressive enhancement network, and the enhanced medical image data is then evaluated to obtain the evaluation results. If the evaluation results do not meet the standards, the current multi-stage progressive augmentation network will be optimized and trained to improve the quality augmentation effect. If the evaluation results meet the standards, the enhanced medical imaging data will be output.

2. The method according to claim 1, characterized in that, After collecting multimodal medical image data using IoT devices, the process also includes: The medical image data is preprocessed by at least one of the following data processing operations: data cleaning, standardization, outlier detection and removal. Accordingly, the medical image data is processed by synthesis, including: The preprocessed medical image data is then synthesized.

3. The method according to claim 1, characterized in that, The medical image data is processed to obtain synthetic medical image data, including: The medical image data is subjected to geometric transformation and / or grayscale transformation to obtain the synthetic medical image data; Alternatively, the medical image data can be synthesized using a conditional generative adversarial network or a conditional diffusion model to obtain the synthesized medical image data.

4. The method according to claim 1, characterized in that, The medical image data is enhanced using an optimized multi-stage progressive enhancement network, including: Multi-scale image features of the medical image data are extracted using the feature pyramid network in the optimized multi-stage progressive enhancement network. The offset is learned from the feature map using the deformable convolutional layer in the optimized multi-stage progressive enhancement network. The anatomical structure in the optimized multi-stage progressive augmentation network is used to maintain constraints, and the augmentation process is guided by an organ mask generated by the segmentation network to ensure the accuracy of the anatomical structure.

5. The method according to claim 1, characterized in that, The quality of the enhanced medical imaging data was assessed, and the assessment results were obtained, including: A dual-branch deep neural network architecture was used to assess the quality of enhanced medical image data, and the assessment results were obtained. Alternatively, a trained language model can be used to assess the quality of the enhanced medical image data to obtain the assessment results.

6. The method according to claim 5, characterized in that, A dual-branch deep neural network architecture is used to assess the quality of enhanced medical image data, and the assessment results are obtained, including: The local feature extraction branch and the global feature extraction branch in the dual-branch deep neural network architecture are used to extract the local features and global features of the enhanced medical image data, respectively. The local and global features are fused using a cross-attention mechanism to obtain noise level scores, artifact severity ratings, and diagnostic usability probabilities. The noise level score, the artifact severity rating, and the diagnostic availability probability are determined as the evaluation results.

7. The method according to claim 6, characterized in that, Also includes: Using interpretive visualization components, a heatmap corresponding to the evaluation results is generated to mark areas of quality defects.

8. An image processing system, characterized in that, include: The data acquisition and preprocessing module is used to acquire multimodal medical image data using Internet of Things devices, and to perform synthetic processing on the medical image data to obtain synthetic medical image data. An iterative optimization module is used to expand the training samples using the synthetic medical image data, and to optimize the training of the multi-stage progressive enhancement network using the expanded training samples. The quality enhancement and evaluation module is used to perform quality enhancement processing on the medical image data using an optimized multi-stage progressive enhancement network, and to evaluate the quality of the enhanced medical image data to obtain the evaluation results. An iterative enhancement module is used to optimize and train the current multi-stage progressive enhancement network to improve the quality enhancement effect if the evaluation result does not meet the standard. The results output module is used to output enhanced medical image data if the evaluation results meet the standards.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the image processing method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image processing method as described in any one of claims 1 to 7.