Deep Learning-Based Computer-Aided Analysis Methods for Medical Images

By using a lightweight multi-scale feature extraction network and a multimodal deep fusion mechanism, the problems of high computational complexity, low recognition accuracy, and insufficient multimodal fusion in existing technologies are solved, achieving efficient and accurate medical image analysis and improving the model's generalization ability and the level of automation in image analysis.

CN120807509BActive Publication Date: 2025-12-02BEIJING KEPTON PHARM TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511300311.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-02
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing deep learning-based medical image analysis methods suffer from high computational complexity, large memory consumption, and long training cycles when processing high-resolution or 3D medical images, making it difficult to achieve efficient real-time processing. Furthermore, they lack accuracy and recall in identifying small lesions with diverse shapes and blurred boundaries, have limited model generalization ability, and suffer from insufficient fusion of multimodal image information, all of which affect the accuracy and robustness of the analysis.

Method used

A lightweight multi-scale feature extraction network is adopted, combined with a fine-grained lesion recognition module, and a multimodal deep fusion mechanism is designed. An adaptive regularization strategy is introduced to improve computational efficiency and recognition accuracy, and to make full use of the complementary information of multimodal images.

Benefits of technology

It effectively improves the automation level and accuracy of medical image analysis, enhances the generalization ability of the model, realizes efficient and accurate multimodal information fusion, and promotes the clinical application of deep learning in the field of medical imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807509B_ABST
    Figure CN120807509B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence, specifically to a deep learning-based computer-aided analysis method for medical images. It aims to address the shortcomings of existing medical image analysis methods, such as low efficiency in processing high-resolution images, insufficient accuracy in identifying minute lesions, weak model generalization ability, and a lack of multimodal image fusion. This method improves high-resolution image processing efficiency by constructing a lightweight multi-scale feature extraction network, introduces a fine-grained lesion identification module to enhance the detection accuracy of minute lesions, employs an adaptive regularization strategy to enhance model generalization ability, and designs a multimodal deep fusion mechanism to fully utilize the complementary information from different modalities of images. This application enables more efficient, accurate, and generalized medical image analysis that effectively integrates multimodal information, thereby promoting the clinical application of deep learning in the field of medical imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, specifically relating to a computer-aided analysis method for medical images based on deep learning. Background Technology

[0002] Medical image analysis plays an indispensable role in disease diagnosis, treatment planning, and prognostic assessment. With the development of modern medical technology, the application of various imaging modalities such as CT, MRI, and ultrasound has become increasingly widespread, generating massive amounts of medical image data. Efficient and accurate analysis of this complex and information-rich image data is key to improving the quality and efficiency of medical services. Against this backdrop, computer-aided analysis methods have been introduced into the field of medical imaging, aiming to assist physicians in making more objective and accurate diagnoses.

[0003] Among them, the deep learning-based computer-aided analysis method for medical images, by constructing a multi-layer neural network, can automatically learn and extract high-dimensional features from raw image data, thereby achieving the detection, segmentation, classification, and quantitative analysis of lesions. This method effectively avoids the tediousness and subjectivity of traditional manual feature extraction, significantly improving the automation level and objectivity of medical image analysis, and showing great potential in the auxiliary diagnosis of various diseases.

[0004] However, despite significant progress in medical image analysis, deep learning still faces numerous challenges. Traditional deep learning models often suffer from high computational complexity, large memory consumption, and long training cycles when processing high-resolution or 3D medical images, making efficient real-time processing difficult. For small lesions with diverse morphologies and blurred boundaries, the model's recognition accuracy and recall still need improvement. Furthermore, with limited labeled data, the generalization ability of existing models is often limited, prone to overfitting, leading to performance degradation when dealing with data from different devices and patient groups. In addition, existing methods lack effective cross-modal feature alignment and deep fusion mechanisms when fusing multimodal medical image information, failing to fully utilize complementary information between different image modalities, thus affecting the final analysis accuracy and robustness. These problems severely restrict the further clinical application and promotion of deep learning in medical image computer-aided analysis. Therefore, there is an urgent need for a deep learning-based medical image computer-aided analysis method that can overcome the aforementioned shortcomings of existing technologies, providing a more efficient, accurate, and generalizable method that can effectively fuse multimodal information. Summary of the Invention

[0005] To address the challenges of existing computer-aided medical image analysis methods in processing high-resolution or 3D medical images, such as high computational complexity, large memory consumption, long training cycles, difficulty in achieving efficient real-time processing, the need to improve the accuracy and recall of identifying small lesions with diverse shapes and blurred boundaries, the limited generalization ability of models under limited labeled data leading to overfitting, and the lack of effective cross-modal feature alignment and deep fusion mechanisms when fusing multimodal medical image information, this application proposes a deep learning-based computer-aided medical image analysis method.

[0006] This application provides a deep learning-based computer-aided analysis method for medical images. It improves the efficiency of high-resolution image processing by constructing a lightweight multi-scale feature extraction network and introduces a fine-grained lesion identification module to enhance the detection accuracy of small lesions. The method further employs an adaptive regularization strategy to enhance the model's generalization ability and designs a multimodal deep fusion mechanism to fully utilize the complementary information of different modalities of images. This effectively overcomes the limitations of existing technologies in terms of computational efficiency, recognition accuracy, model generalization ability, and multimodal fusion, providing a more efficient, accurate, and generalizable medical image analysis solution that effectively integrates multimodal information, thereby promoting the clinical application and widespread adoption of deep learning in the field of medical imaging.

[0007] According to one aspect of this application, a deep learning-based computer-aided analysis method for medical images is provided, comprising the following steps:

[0008] Receive medical image data to be analyzed, which may include single-modal or multimodal images;

[0009] The medical image data is preprocessed and feature extracted to obtain multi-scale image feature representations;

[0010] Based on the multi-scale image feature representation and pre-trained deep learning model, lesion identification or segmentation is performed to generate preliminary analysis results;

[0011] The preliminary analysis results are then post-processed and evaluated to output a final computer-aided analysis report.

[0012] Compared with existing technologies, this application provides a deep learning-based computer-aided analysis method for medical images. It improves the efficiency of high-resolution image processing by constructing a lightweight multi-scale feature extraction network, while introducing a fine-grained lesion recognition module to enhance the detection accuracy of small lesions. Furthermore, this method employs an adaptive regularization strategy to enhance the model's generalization ability and designs a multimodal deep fusion mechanism to fully utilize the complementary information from different modalities of images. This effectively overcomes the limitations of existing technologies in terms of computational efficiency, recognition accuracy, model generalization ability, and multimodal fusion, providing a more efficient, accurate, and generalizable medical image analysis solution that effectively integrates multimodal information, thereby promoting the clinical application and widespread adoption of deep learning in the field of medical imaging.

[0013] The various embodiments of this application will now be described in detail.

[0014] The process of receiving medical image data to be analyzed, which may include single-modal or multi-modal images, specifically includes acquiring raw medical images from medical imaging equipment or data storage systems. The medical image data can cover multiple imaging modalities, such as computed tomography (CT) images, magnetic resonance imaging (MRI) images, positron emission tomography (PET) images, ultrasound images, digital subtraction angiography (DSA) images, and pathological slide images. Each modality of image data has its unique information characteristics and clinical application focus. For example, CT images are good at displaying bone structures and calcified lesions, while MRI images have advantages in soft tissue contrast. PET images can provide functional metabolic information, while ultrasound images are real-time and radiation-free. During the receiving process, the system performs necessary metadata parsing to obtain auxiliary information such as image acquisition parameters, patient information, and scanning protocols. This auxiliary information is crucial for subsequent preprocessing and analysis. For multi-modal images, the system also performs image sequence matching and association to ensure that image data from different modalities can be correctly mapped to the same lesion area of ​​the same patient, laying the foundation for subsequent deep fusion. The accuracy of this step directly affects the effectiveness and reliability of subsequent analyses.

[0015] The steps of preprocessing and feature extraction of the medical image data to obtain multi-scale image feature representations specifically include a series of data normalization and information enhancement operations. The preprocessing operations aim to eliminate noise and artifacts in the original data and to unify the images, including but not limited to the following steps:

[0016] Image format conversion and data normalization: This involves converting received image data of different formats into a standard format that the system can process. Pixel values ​​are normalized, for example, mapped to a range of zero to one, to eliminate intensity differences caused by different acquisition devices or protocols, thereby improving the training stability and convergence speed of the model.

[0017] Image registration: For multimodal image data, precise registration is performed to spatially align images of different modalities, ensuring that corresponding pixels at the same anatomical location accurately overlap in different modal images. This registration can employ intensity-based or feature-based algorithms, combined with various geometric correction methods such as affine transformation and elastic transformation, to compensate for spatial deviations caused by differences in patient position, respiratory motion, or scanning field of view.

[0018] Noise suppression and artifact removal: Filtering algorithms such as Gaussian filtering, median filtering, or nonlocal mean filtering are used to effectively suppress random noise in the images. Simultaneously, artifacts such as metallic artifacts and motion artifacts caused by equipment, patient movement, or the reconstruction process are identified and removed to improve image quality and ensure the accuracy of subsequent feature extraction.

[0019] Region of Interest (ROI) Cropping and Enhancement: Based on prior knowledge or automatic detection algorithms, the RIO containing potential lesions is cropped to reduce interference from irrelevant background information and lower computational burden. Local contrast enhancement is then applied to the RIO using techniques such as histogram equalization and adaptive histogram equalization to highlight lesion details.

[0020] After preprocessing, the system will perform core feature extraction to construct a lightweight multi-scale feature extraction network, thereby efficiently acquiring deep semantic information from the image. The lightweight multi-scale feature extraction network is designed to balance computational efficiency and feature representation capability, making it particularly suitable for high-resolution or 3D medical images. The network may include the following structure:

[0021] Encoder-decoder architecture: Classic encoder-decoder structures such as U-shaped networks and fully convolutional networks are used as the basic framework. The encoder extracts high-level semantic features and reduces the feature map size step by step through a series of convolutional layers, pooling layers, or stride convolutional layers, while increasing the receptive field. The decoder fuses the low-level spatial information and high-level semantic information extracted by the encoder through upsampling layers and skip connections, gradually restoring the spatial resolution of the feature map, and finally generating a fine segmentation map or feature map.

[0022] Multi-scale feature fusion module: Embedding a multi-scale feature fusion mechanism in the encoder or decoder. This can include a spatial pyramid pooling module, which captures different contextual information through pooling operations at different scales; or a feature pyramid network structure, constructing a feature pyramid with rich semantic information from encoder feature maps at different levels, effectively handling lesions of different sizes. Dilated convolution can expand the receptive field without increasing the number of parameters and computational cost, capturing broader contextual information, especially suitable for recognizing small lesions with blurred boundaries and diverse morphologies. Dilated convolution has multi-branch parallel processing capabilities, capturing contextual information at different scales.

[0023] Lightweight Convolution Module: To reduce computational complexity and memory consumption, lightweight designs such as depthwise separable convolution, group convolution, or channel attention mechanisms are employed. Depthwise separable convolution decomposes standard convolution into two steps: depthwise convolution and pointwise convolution, significantly reducing the number of parameters and computational cost. Group convolution groups input feature channels for convolution, further improving efficiency. Channel attention mechanisms adaptively adjust the weights of different channels based on the importance of input features, allowing the network to focus more on features related to lesions.

[0024] Adaptive attention mechanisms: Introducing channel attention modules or spatial attention modules allows the network to dynamically focus on areas in the image that are critical to diagnosis, such as potential lesions. Channel attention mechanisms enhance or suppress features of different channels by learning the weights of each channel, while spatial attention mechanisms emphasize features of important regions by learning the weights of spatial locations. This mechanism is particularly effective in improving the detection accuracy of fine-grained lesions.

[0025] Feature Dimension Compression and Enhancement: Feature dimensions are compressed and enhanced at different levels using methods such as bottleneck structures and residual connections to reduce computational resource consumption while maintaining information content and mitigating the gradient vanishing problem.

[0026] The step of identifying or segmenting lesions based on the multi-scale image feature representation and pre-trained deep learning model to generate preliminary analysis results is the core diagnostic part of this method. This step utilizes the trained deep learning model to infer the extracted multi-scale image features, achieving accurate identification or segmentation of lesions. Specific implementation schemes may include:

[0027] Lesion detection and localization: For tasks requiring the identification of specific lesions in an image and their location, a target detection model is employed. This model can be based on a single-stage detector such as the YOLO or SSD series, or a two-stage detector such as Faster R-CNN. These models utilize feature maps output by lightweight multi-scale feature extraction networks to predict the lesion category and bounding box coordinates. Introducing a fine-grained lesion recognition module enhances the detection capability for small, irregular lesions. This module can employ a multi-scale feature fusion strategy, combining high-level semantic information with low-level spatial information, or design specific anchor box mechanisms to adapt to small-sized lesions.

[0028] Lesion region segmentation: For tasks requiring precise delineation of lesion boundaries, semantic segmentation or instance segmentation models are employed. These models can be based on U-Net, Deeplab series, etc. These models utilize a decoder to progressively restore spatial resolution and incorporate skip connections to pass detailed information captured by the encoder to the decoder, thereby generating pixel-level lesion masks. The fine-grained lesion recognition module here demonstrates its ability to accurately distinguish lesions from normal tissue, especially when lesion boundaries are blurred or have low contrast with the background. For example, a boundary-aware loss function or attention gating mechanism can be introduced to improve boundary segmentation accuracy.

[0029] Lesion Classification and Differentiation: For tasks requiring classification of detected or segmented lesions, such as benign / malignant differentiation, a classification model is employed. This model can be an independent classification network or serve as the backend classification head of a detection / segmentation model. Deep learning models learn the morphological, textural, and functional features of lesions to achieve accurate lesion classification.

[0030] Multimodal Deep Fusion Mechanism: When processing multimodal medical images, this method designs an efficient multimodal deep fusion mechanism to fully utilize the complementary information between different modalities. This fusion mechanism can be implemented at different stages:

[0031] Early fusion: In the feature extraction stage, the original images or their preliminary features from different modalities are stitched together or fused at the element level, and then input into a unified deep learning model for processing.

[0032] Intermediate-level fusion: In the middle layers of the network, features extracted from different modalities are interactively fused. For example, cross-modal attention mechanisms can be used, allowing features from one modality to guide the enhancement of features from another modality, or shared representations of different modalities can be learned through a common embedding space. This helps capture more complex cross-modal associations, especially in fine-grained lesion identification.

[0033] Late-stage fusion: Images of different modalities are analyzed independently, and then the recognition or segmentation results are fused together. For example, voting mechanisms, weighted averaging, or meta-learning models can be used to integrate prediction results from different modalities.

[0034] Pre-trained deep learning models: The deep learning models used can be pre-trained on large-scale public datasets to learn general medical image features, and then fine-tuned on task-specific datasets. Pre-training helps the model achieve better performance with limited labeled data.

[0035] Adaptive Regularization Strategies: To enhance the model's generalization ability with limited labeled data, this method employs a series of adaptive regularization strategies, including but not limited to the following techniques:

[0036] Data augmentation: Performing operations such as random flipping, rotation, scaling, cropping, brightness adjustment, and elastic deformation on training data to increase the diversity of training samples and reduce the model's dependence on specific data patterns.

[0037] Batch normalization: Adding a batch normalization layer after each layer of the network to accelerate network training and play a certain role in regularization, reducing the internal covariate shift.

[0038] Discarding: Randomly discarding a portion of the neuron outputs during training forces the network to learn more robust feature representations and avoids overfitting.

[0039] Weight decay: Adding an L2 norm penalty term for the model weights to the loss function to limit the size of the model parameters and prevent the model from becoming too complex.

[0040] Label smoothing: In classification tasks, hard labels are softened, for example, the label values ​​of one-hot encoding are adjusted from zero and one to 0.1 and 0.9 to reduce model overconfidence and improve generalization ability.

[0041] Contrastive learning: Learning discriminative feature representations of images using unsupervised or self-supervised methods can effectively improve model performance even when labeled data is scarce.

[0042] The steps of post-processing and evaluating the preliminary analysis results to output a final computer-aided analysis report aim to refine the model's output and facilitate clinical translation, ensuring the accuracy, reliability, and interpretability of the final results. Specific steps include:

[0043] Post-processing operations:

[0044] Connectivity analysis: Perform connectivity analysis on the segmentation results to remove excessively small, isolated false positive regions, or merge adjacent fragments of real lesions.

[0045] Morphological operations: Using morphological operations such as opening, closing, dilation, and erosion to smooth lesion boundaries, fill cavities, or remove burrs, making the segmentation results more consistent with medical anatomical shapes.

[0046] False positive reduction: By combining contextual information or statistical methods, false positives in detection or segmentation can be further reduced. For example, by analyzing the location, size, shape, and other characteristics of lesions, atypical false positives can be filtered out.

[0047] Quantitative analysis and evaluation:

[0048] Lesion measurement: Precise quantitative measurement of identified or segmented lesions, including volume, maximum diameter, average density or intensity, shape index, etc. These quantitative indicators are of significant clinical importance for disease diagnosis, staging, treatment efficacy evaluation, and prognosis.

[0049] Uncertainty Quantification: By combining techniques such as Bayesian deep learning and Monte Carlo dropout, the uncertainty of model predictions is quantified, providing doctors with an assessment of model confidence, which helps with risk management and decision support.

[0050] Report generation and visualization:

[0051] Generate structured reports: The final analysis results are integrated into a structured report template, including detailed information on lesions, quantitative measurement results, descriptions of imaging features, and possible diagnostic suggestions. The report can automatically populate predefined fields and allows physicians to revise and supplement it.

[0052] Results visualization: Segmentation masks or bounding boxes of lesions are overlaid on the original medical images, and the lesion areas are highlighted in the form of pseudo-color images. A 3D reconstructed view is provided to help doctors intuitively understand the spatial location and morphology of the lesions, enhancing the interpretability of the diagnosis.

[0053] Trend analysis: For follow-up images, the system can compare quantitative indicators of the same lesion at different time points and generate a trend chart of lesion changes to help doctors assess treatment effectiveness or disease progression.

[0054] Clinical decision support: The system integrates the analysis results with other information such as the patient's clinical history and laboratory test results to provide physicians with more comprehensive clinical decision support. It can also provide recommendations for relevant medical literature or guidelines to assist physicians in making more accurate diagnostic and treatment choices.

[0055] The deep learning-based computer-aided analysis method for medical images presented in this application provides a comprehensive, efficient, accurate, and robust solution through the synergistic effect of the aforementioned steps. It effectively improves the automation and objectivity of medical image analysis, reduces the workload of physicians, and lowers the risk of missed or misdiagnosed diagnoses, thus having profound significance for accelerating medical research progress and improving patient healthcare services. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the overall technical architecture of the deep learning-based computer-aided analysis method for medical images proposed in this invention.

[0057] Figure 2 This is a schematic diagram of the core principle framework of the lightweight multi-scale feature extraction network in this invention. Detailed Implementation

[0058] In the following description, specific terminology and structures are used to illustrate embodiments of the invention for ease of understanding. However, these descriptions are not intended to limit the invention, and those skilled in the art will understand that various modifications and substitutions can be made without departing from the spirit and scope of the invention. The accompanying drawings are for illustrative purposes only, are not drawn to scale, and should not be construed as limiting the invention.

[0059] This application proposes a deep learning-based computer-aided analysis method for medical images. This method leverages deep learning technology, particularly by introducing a lightweight multi-scale feature extraction network, to enhance the identification of complex lesions and minute structures in medical images. Specifically, this method efficiently extracts multi-level, multi-scale semantic features from raw medical images and performs precise medical image analysis based on these features, such as lesion segmentation, disease classification, or anomaly detection. Through the adaptive learning capability of the deep learning model, this method effectively handles the heterogeneity of image data, reduces reliance on manual interpretation, and improves the automation and consistency of the analysis. Ultimately, this method aims to provide a high-precision, high-efficiency, and high-generalization-capability computer-aided analysis tool for medical images, thereby providing clinicians with objective and reliable diagnostic evidence and assisting in the early detection and precise treatment of diseases.

[0060] like Figure 1 and Figure 2 As shown, the deep learning-based computer-aided analysis method for medical images according to an embodiment of this application includes: S110, acquiring medical image data and preprocessing it; S120, extracting features from the preprocessed medical image data using a lightweight multi-scale feature extraction network; S130, performing medical image analysis based on the extracted features; and S140, outputting the analysis results and generating a report.

[0061] In the aforementioned deep learning-based computer-aided analysis method for medical images, step S110 involves acquiring and preprocessing medical image data. It should be understood that acquiring and preprocessing medical image data is fundamental to ensuring that the subsequent deep learning model can perform efficient and accurate analysis. Medical image data typically originates from various modalities, including but not limited to computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), ultrasound examination, and pathological slide images. These raw data vary significantly in format, resolution, contrast, and noise levels; directly inputting them into a deep learning model would severely impact the model's training efficiency and generalization performance. Specifically, the data acquisition stage involves collecting image files from image archiving and communication systems (IARS), digital imaging and communication medical standards (DICM) compatible devices, or dedicated databases. These files are typically stored in DICM formats and contain image pixel information, patient information, scan parameters, and image metadata. The acquired data first needs to undergo format conversion and standardization to ensure that all images can be recognized and loaded by subsequent processing modules. For example, various digital images and medical communication standard files generated by devices from different manufacturers can be uniformly parsed into an internal standard data structure, which facilitates unified management and access.

[0062] Furthermore, after acquiring medical imaging data, the preprocessing process aims to eliminate redundant information in the original data, correct noise, unify data scale, and enhance the contrast of target lesions. Preprocessing specifically includes the following operations: First, data standardization and normalization. The range of image intensity values ​​generated by different imaging devices can vary greatly, leading to difficulties in model convergence during direct training. Standardization adjusts pixel intensity values ​​to a zero-mean, unit-variance distribution, while normalization scales pixel intensity values ​​to a specific numerical range, such as zero to one or negative one to one. This process helps stabilize model training and improves the model's adaptability to different data sources. Second, image registration. For multimodal images or images of the same organ acquired at different time points, image registration techniques are needed to align them to the same coordinate system to ensure spatial consistency of corresponding anatomical structures. Image registration employs feature-based or mutual information-based methods, optimizing similarity metrics to achieve geometric transformation of the images. This process is crucial for fusing multimodal information or tracking disease progression. Third, noise suppression and artifact removal. Medical images often contain noise and artifacts caused by scanning equipment, patient movement, or environmental factors. Examples include ray hardening artifacts in computed tomography (CT) images and motion artifacts in magnetic resonance imaging (MRI). The preprocessing stage employs various filtering techniques, such as Gaussian filtering, median filtering, or nonlocal mean filtering, as well as deep learning-based denoising algorithms, to improve image quality and lesion visibility. Fourth, contrast enhancement. Histogram equalization, gamma correction, or domain-knowledge-based contrast stretching techniques effectively enhance the contrast between the target lesion and surrounding tissues, making it more prominent in the image and facilitating subsequent feature extraction. Fifth, data augmentation. To increase the diversity and generalization ability of the training dataset, the preprocessing process also includes data augmentation operations. This operation generates new training samples by performing geometric and photometric transformations on existing images, such as random rotation, flipping, scaling, cropping, brightness adjustment, and elastic deformation. This effectively reduces the risk of model overfitting and improves the model's robustness to various variations. The refined implementation of these preprocessing steps provides high-quality input for the lightweight multi-scale feature extraction network, ensuring the accuracy and reliability of subsequent analysis.

[0063] In the aforementioned deep learning-based computer-aided analysis method for medical images, step S120 involves extracting features from the preprocessed medical image data using a lightweight multi-scale feature extraction network. It should be understood that lesions in medical images vary greatly in size, shape, and location, and their boundaries are often blurred. Traditional feature extraction methods, such as those based on hand-designed texture or edge features, struggle to effectively capture this multi-scale, multi-level semantic information, resulting in insufficient performance when identifying minute lesions or complex anatomical structures. Using single-scale convolutional kernels for feature extraction may lead to the loss of local details or a lack of global contextual information. Simultaneously, to meet the requirements of computational efficiency and deployment flexibility in clinical applications, the feature extraction network must be lightweight, avoiding excessively large model parameters and computational load. Therefore, this method introduces a lightweight multi-scale feature extraction network, aiming to simultaneously balance the comprehensiveness of feature extraction with the economy of computational resources.

[0064] Specifically, the design principles of this lightweight multi-scale feature extraction network are as follows: First, it employs depthwise separable convolution. Depthwise separable convolution decomposes the standard convolution operation into two independent steps: depthwise convolution and pointwise convolution. This significantly reduces the number of model parameters and computational cost while maintaining similar feature extraction capabilities. Depthwise convolution performs convolution independently on each input channel, effectively extracting spatial features; pointwise convolution integrates information across different channels through one-to-one convolution. This design allows the network to effectively capture local texture and structural information of images while maintaining a lightweight design. Second, it fuses multi-scale features. The network contains multiple parallel or serial branches, each using convolution kernels with different receptive fields or different downsampling rates to process the input image, thereby capturing feature information at different scales. For example, by setting dilated convolution kernels with different dilation rates, the receptive field can be effectively expanded without increasing parameters and computational cost, thus obtaining broader contextual information. Multi-scale feature fusion is achieved through a feature pyramid network (pyramid feature network) structure, which integrates high-resolution detail features at low levels with semantically rich features at high levels to form a feature representation containing rich details and global contextual information.

[0065] Third, an attention mechanism is introduced. To enable the network to adaptively focus on the most diagnostically significant regions (e.g., lesion areas) in medical images and suppress irrelevant background noise, this lightweight multi-scale feature extraction network integrates channel attention and spatial attention mechanisms. The channel attention mechanism learns the importance weights of each channel feature to weight different semantic information; the spatial attention mechanism learns the importance weights of each spatial location to enhance the feature representation of key regions. This mechanism allows the network to dynamically adjust its attention to different regions and features, further improving the accuracy of lesion identification. Fourth, efficient activation functions and normalization layers are employed. To maintain the network's lightweight nature, computationally efficient and high-performance activation functions are selected, such as leaky rectified linear units (LER) or hybrid attention activation functions. Simultaneously, batch normalization layers (batch regularization) or instance normalization layers (instance regularization) are used to accelerate model training, improve model stability, and reduce internal covariate bias.

[0066] Through the above design, the lightweight multi-scale feature extraction network effectively addresses the shortcomings of traditional methods in multi-scale lesion identification and computational efficiency. Its output feature representation is a high-dimensional semantic vector containing rich anatomical structural information and pathological features from medical images. These features not only encode local information such as the shape, size, and texture of the lesion, but also integrate global contextual information such as its spatial relationship with surrounding tissues. These feature vectors will be input into subsequent medical image analysis modules as the basis for accurate diagnosis.

[0067] In the embodiments of this application, step S120 involves feature extraction from the preprocessed medical image data. During the training phase, the lightweight multi-scale feature extraction network optimizes model parameters by minimizing the difference between the predicted results and the true labels. This process involves a loss function used to quantify model performance. Common loss functions include a combination of cross-entropy loss and Jaccard coefficient loss, which can be expressed as:

[0068]

[0069] in, Represents the total loss function. Represents cross-entropy loss, Indicates the loss of the Jaccard coefficient. and To balance the weighting coefficients of the two loss terms, cross-entropy loss is used. The specific calculation formula is as follows:

[0070]

[0071] in, This represents the total number of pixels. Represents pixels The true labels (e.g., zero represents the background, one represents the lesion). Represents the pixel points predicted by the model. The probability of belonging to a lesion. Jaccard coefficient loss. The calculation formula is as follows:

[0072]

[0073] The combination of the above loss functions allows the model to focus on pixel-level classification accuracy while also taking into account the overall structure and boundary matching of the lesion region during the optimization process, thereby achieving more accurate lesion identification and segmentation in complex medical image analysis tasks.

[0074] In the aforementioned deep learning-based computer-aided medical image analysis method, step S130 involves performing medical image analysis based on the extracted features. It should be understood that, based on the rich semantic features obtained from the lightweight multi-scale feature extraction network in step S120, this step aims to perform specific medical image analysis tasks, such as lesion segmentation, disease classification, or anomaly detection. These tasks are crucial aspects of clinical diagnosis, directly impacting early disease detection, precise localization, and treatment planning. Specifically, the extracted features encode various information in medical images, ranging from low-level texture and edges to high-level semantics (e.g., organ structure, lesion morphology). The core of this step is how to effectively utilize these features for accurate downstream task analysis.

[0075] For lesion segmentation tasks, this method inputs the extracted features into a dedicated segmentation head. This segmentation head typically contains a series of upsampling layers, convolutional layers, and activation layers, designed to progressively restore the low-resolution semantic feature map to the resolution of the original image and predict its category (e.g., normal tissue, tumor, inflammation) at each pixel level. The upsampling operation is implemented through transposed convolution (deconvolution) or bilinear interpolation to increase the spatial dimension of the feature map. Convolutional layers are used to further refine the feature representation and fuse information from different scales. Finally, an output layer (e.g., a convolutional layer with a non-linear activation function) generates a probability map for each pixel to belong to a specific category. For example, in tumor segmentation tasks, the model outputs a probability map of the same size as the original image, where each pixel value represents the probability that the pixel belongs to the tumor region. Then, by setting a threshold, a binarized segmentation mask is obtained, clearly indicating the precise location and boundaries of the lesion. The key to this process is that the segmentation head must be able to fully utilize the multi-scale feature information provided by the lightweight multi-scale feature extraction network to ensure the accuracy and detail integrity of the segmentation results.

[0076] For disease classification or anomaly detection tasks, this method extracts features and passes them through a classification head. This classification head typically consists of one or more fully connected layers, used to map high-dimensional feature vectors to the final class probabilities. Batch normalization layers and activation functions can be inserted between the fully connected layers to enhance the model's non-linear expressiveness and accelerate training. The input to the classification head is the global feature representation output by a lightweight multi-scale feature extraction network, such as a vector obtained by reducing the dimensionality of the feature map through global average pooling or global max pooling. Finally, an output layer with a non-linear activation function (e.g., a non-linear activation function for binary classification or a soft max function for multi-class classification) generates the probabilities that an image belongs to each disease category or normal / abnormal category. By thresholding these probabilities, the final diagnostic result can be obtained.

[0077] It is understandable that the effectiveness of this step lies in its ability to leverage the powerful pattern recognition capabilities of deep learning models to automatically learn the complex mapping relationship between medical image features and clinical diagnostic results. This avoids the tediousness and subjectivity of manual feature extraction in traditional methods, significantly improving the objectivity and accuracy of the analysis. Furthermore, by designing corresponding analysis modules for different tasks (segmentation, classification), the method's versatility and flexibility are ensured, enabling it to adapt to diverse clinical needs.

[0078] In the aforementioned deep learning-based computer-aided analysis method for medical images, step S140 involves outputting the analysis results and generating a report. It should be understood that presenting the analysis results of the deep learning model in an intuitive, easy-to-understand, and clinically compliant manner is a crucial step in realizing the value of computer-aided analysis. Simple model predictions, such as probability values ​​or segmentation masks, require further interpretation and transformation by clinicians to aid decision-making. Therefore, this step aims to transform the model's raw output into a structured report with clinical guidance significance and provide a visual presentation of the results.

[0079] Specifically, the output analysis results include, but are not limited to, the following forms: First, quantitative indicators. For lesion segmentation tasks, the output results include the lesion's volume, dimensions (length, width, height), average density or intensity value, and shape characteristics (e.g., sphericity, flatness). These quantitative indicators provide doctors with an objective description of the lesions, assisting in assessing the severity and progression of the disease. For classification tasks, the output results are probability values ​​for each category, such as the probability that a tumor is benign or malignant. These probability values ​​can serve as a reference for doctors in diagnosis.

[0080] Second, visualization results. By overlaying segmentation masks onto the original medical images, the precise location and extent of lesions can be visually displayed, for example, highlighting tumor areas, inflammatory areas, or organ boundaries with different colors. For classification tasks, heatmaps can be annotated on key areas of the image, such as using gradient-weighted activation mapping (GEM) technology to visualize the key areas where the model makes judgments, enhancing the interpretability of the model results. These visualizations help doctors quickly understand the basis of the model's analysis and perform manual review.

[0081] Third, structured report generation. This method automatically generates a structured report containing all key information based on a predefined report template. The report typically includes: basic patient information, type and date of imaging examinations, model analysis results (quantitative indicators, classification probabilities), visualization images (images overlaid with segmentation results, heatmaps), model confidence scores, and potential clinical recommendations. The report format is compatible with hospital information systems (HAS) or radiology information systems (RIS), facilitating seamless integration into existing clinical workflows. For example, the report automatically generates a table containing parameters such as tumor volume, boundary clarity, and internal homogeneity, along with a multi-view 3D reconstruction of the lesion. This structured report not only improves report generation efficiency but also ensures the consistency and completeness of information presentation.

[0082] Fourth, anomaly handling and quality control. During the results output phase, the system also integrates an anomaly handling mechanism and a quality control module. This module can verify the reasonableness of the model's prediction results, such as checking whether the segmentation results are connected and whether the classification probabilities are within a reasonable range. If an anomaly is detected, the system will issue an alarm and may trigger a manual review process to ensure the accuracy and reliability of the final output report. For example, if the model predicts a lesion volume that is too large or too small, exceeding the normal range, the system will prompt the doctor to conduct a secondary confirmation. This mechanism effectively improves the security and robustness of the entire computer-aided analysis system.

[0083] In summary, the deep learning-based computer-aided medical image analysis method based on the embodiments of this application has been clarified. By utilizing deep learning technology and introducing a lightweight multi-scale feature extraction network, it can efficiently extract multi-level and multi-scale semantic features from the original medical images, and perform accurate medical image analysis on this basis, such as lesion segmentation, disease classification or abnormality detection. Finally, it outputs the results in the form of structured reports and visualizations, thereby providing clinicians with objective and reliable diagnostic evidence and assisting in the early detection and precise treatment of diseases.

Claims

1. A computer-aided analysis method for medical images based on deep learning, characterized in that, Includes the following steps: Receive medical image data to be analyzed, including single-modal or multimodal images, and perform metadata parsing and image sequence matching; The medical image data is preprocessed to remove noise and artifacts from the original data and to unify the image format and pixel values. By constructing a lightweight multi-scale feature extraction network, the deep semantic information and multi-scale image feature representation of the medical image data are obtained; Based on the multi-scale image feature representation and pre-trained deep learning model, lesion identification or segmentation is performed through a fine-grained lesion identification module. During the identification process, an adaptive regularization strategy is adopted to enhance the model's generalization ability, and a multi-modal deep fusion mechanism is applied to fully utilize the complementary information of different modal images, thereby generating preliminary analysis results. The preliminary analysis results are post-processed and evaluated to output a final computer-aided analysis report containing quantitative indicators, visualizations, and a structured report.

2. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, Receiving the medical image data to be analyzed specifically includes: Obtain raw medical images from medical imaging equipment or data storage systems; The medical imaging data includes at least one of the following: computed tomography (CT) images, magnetic resonance imaging (MRI) images, positron emission tomography (PET) images, ultrasound images, digital subtraction angiography (DSA) images, and pathological slide images. Metadata parsing is performed to obtain image acquisition parameters, patient information, and scanning protocols; For multimodal imaging, image sequence matching and association are performed to ensure that image data from different modalities correspond to the same lesion area of ​​the same patient.

3. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, The preprocessing of the medical image data specifically includes: Image format conversion and data normalization converts received image data of different formats into a standard format and normalizes the pixel values ​​of the images. Image registration, for multimodal image data, performs precise registration operations to align images of different modalities in space; Noise suppression and artifact removal: The filtering algorithm is used to suppress random noise in the image and to identify and remove artifacts caused by equipment, patient movement or reconstruction process. Region of Interest (ROI) cropping and enhancement: Based on prior knowledge or automatic detection algorithms, the ROI containing potential lesions is cropped, and local contrast enhancement is performed on the ROI.

4. The deep learning-based computer-aided analysis method for medical images according to claim 3, characterized in that, The image registration employs intensity-based or feature-based algorithms, combined with affine or elastic transformations for geometric correction; the noise suppression and artifact removal utilize Gaussian filtering, median filtering, or nonlocal mean filtering; and the region of interest cropping and enhancement employ histogram equalization or adaptive histogram equalization.

5. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, The lightweight multi-scale feature extraction network includes: The encoder-decoder architecture consists of an encoder part and a decoder part. The encoder part extracts high-level semantic features through convolutional layers and pooling layers, while the decoder part restores the spatial resolution of the feature map through upsampling layers and skip connections. The multi-scale feature fusion module captures contextual information at different scales through spatial pyramid pooling, feature pyramid network structures, or dilated convolutions. The lightweight convolution module employs depthwise separable convolution, group convolution, or channel attention mechanisms to reduce computational complexity and memory consumption. An adaptive attention mechanism, which introduces a channel attention module or a spatial attention module to dynamically focus on key diagnostic areas in the image; The feature dimension compression and enhancement module uses a bottleneck structure or residual connections to compress and enhance the feature dimension at different levels.

6. The deep learning-based computer-aided analysis method for medical images according to claim 5, characterized in that, During the training phase, the lightweight multi-scale feature extraction network optimizes model parameters by minimizing the difference between the predicted results and the true labels. The process of minimizing the predicted results uses a combination of cross-entropy loss and Jaccard coefficient loss as the loss function.

7. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, Based on the multi-scale image feature representation and pre-trained deep learning model, lesion identification or segmentation is performed through a fine-grained lesion identification module, specifically including: The multi-scale image feature representation is input into the target detection model, which predicts the lesion category and bounding box coordinates based on a single-stage detector or a two-stage detector for lesion detection and localization. The multi-scale image feature representation is input into a semantic segmentation model or an instance segmentation model. The model restores the spatial resolution through a decoder and generates a pixel-level lesion mask by combining skip connections, so as to segment the lesion region. The multi-scale image feature representation is input into a classification model, which classifies and identifies lesions by learning their morphological, textural, and functional features.

8. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, The multimodal deep fusion mechanism is performed in at least one of the following stages: Early fusion involves stitching together or fusing the original images or their preliminary features from different modalities during the feature extraction stage, and then inputting them into a unified deep learning model for processing. Mid-level fusion involves interactively fusing features extracted from different modalities at the intermediate layers of the network, employing cross-modal attention mechanisms or learning shared representations through a common embedding space. In late-stage fusion, images of different modalities are analyzed independently, and then the recognition or segmentation results of each modality are fused into a decision through a voting mechanism, weighted average, or meta-learning model.

9. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, The adaptive regularization strategy includes at least one of the following techniques: Data augmentation involves performing random flipping, rotation, scaling, cropping, brightness adjustment, or elastic deformation operations on the training data. Batch normalization: Add a batch normalization layer after the network layer. Discarding: Randomly discarding a portion of the neuron's output during training; Weight decay involves adding a penalty term for the model weights to the loss function. Label smoothing softens hard labels in classification tasks; Contrastive learning utilizes unsupervised or self-supervised methods to learn discriminative feature representations of images.

10. The deep learning-based computer-aided analysis method for medical images according to claim 1, characterized in that, Post-processing and result evaluation of the preliminary analysis results to output the final computer-aided analysis report specifically include: The preliminary analysis results were subjected to connected component analysis, morphological manipulation, and false positive reduction. Precise quantitative measurements of identified or segmented lesions, including volume, maximum diameter, average density or intensity, and shape index; By combining Bayesian deep learning or Monte Carlo dropout techniques, the uncertainty of model predictions can be quantified. Generate structured reports containing detailed information about lesions, quantitative measurement results, descriptions of imaging features, and diagnostic recommendations; Results visualization involves overlaying a segmentation mask or bounding box of the lesion onto the original medical image and highlighting the lesion area in pseudo-color format to provide a three-dimensional reconstruction view. Trend analysis is performed to compare quantitative indicators of the same lesion at different time points, generating a trend chart of lesion changes.

Citation Information

Patent Citations

  • Multi-modal fine-grained thesis classification method and system based on regularization ensemble learning

    CN116956214A

  • Medical image identification method and medical image identification device

    CN118537648A