Identification and classification method for lesions in medical images

Through the hierarchical attention feature fusion network and dynamic architecture division algorithm based on image complexity, the problems of insufficient fusion accuracy and intricate lesion area analysis in multimodal medical imaging are solved, and higher accuracy and robustness of lesion recognition are achieved. The iterative classification algorithm for Bayesian uncertainty estimation is ensured to ensure the accuracy and reliability of classification results.

CN120219799AInactive Publication Date: 2025-06-27NANJING MAITUO MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510208404.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, multimodal medical imaging fusion accuracy is insufficient, lesion area analysis is not fine, and traditional blocking strategies cannot effectively identify the different allocation and complexity of lesions in the image.

Method used

Using a hierarchical attention feature fusion network and a dynamic architecture division algorithm based on image complexity, a multimodal medical image data set including CT, MRI and PET images is constructed, the feature fusion weight is dynamically adjusted, and the features that have obvious distinction to the lesions are automatically focused, and finely divided according to the image complexity.

Benefits of technology

It improves the accuracy and robustness of lesion recognition, can extract valuable features in multimodal medical images more accurately, avoids the problems of local details loss or overfusion, and ensures the accuracy and reliability of lesion classification results through the iterative classification algorithm of Bayesian uncertainty estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219799A_ABST
    Figure CN120219799A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing and analysis, and discloses a method for identifying and classifying lesions in medical images, which comprises the following steps: acquiring a plurality of medical image data, and constructing a multimode medical image data set containing CT, MRI and PET images; training the multi-mode medical image data set by using a hierarchical attention feature fusion network to generate a focus recognition and classification model; a target medical image is partitioned by a dynamic architecture partitioning algorithm based on image complexity, and analysis is performed by using the focus recognition and classification model to obtain focus features; and classifying the lesion features by using an iterative classification algorithm of Bayesian uncertainty estimation, calculating a classification threshold in combination with a medical expert knowledge base, and outputting a final classification result about the lesion. According to the invention, the accuracy and robustness of identification and analysis are improved, and the accuracy and reliability of a focus classification result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing and analysis, and in particular to a method for identifying and classifying lesions in medical images. Background Art

[0002] In order to overcome the deficiencies of single-modal images, in recent years, multi-modal medical image data fusion technology has been proposed and used for disease detection and classification. By simultaneously using different modal images such as CT, MRI, and PET, more-dimensional structural, tissue, and damage information can be obtained, further improving the detection accuracy of diseases. However, multi-modal data fusion faces a series of challenges. First, different modal images may have registration deviations due to factors such as acquisition equipment, shooting angles, and time differences, affecting the fusion effect. Second, the intensity distribution and restoration differences of images, especially in areas where tissue features are not clear, increase the complexity of image processing, resulting in the fact that the final fused image information cannot effectively improve the diagnostic accuracy in some cases. In addition, due to the heterogeneity of the size and shape of the lesion area in medical image fusion, traditional difference techniques cannot flexibly and dynamically adjust according to the specific characteristics of the lesion, resulting in significant deficiencies in the robustness, goodness, and accuracy of multi-modal fusion technology in some complex cases. Moreover, although the existing block strategies can process local information to a certain extent in programming, the fixed block strategies often cannot effectively identify the different distributions and complexities of lesions in the image, easily resulting in problems such as inaccurate segmentation or loss of details, and cannot achieve ideal lesion detection and classification.

[0003] The present invention proposes a method for multi-modal medical image data fusion and lesion identification and classification. This method improves the accuracy of lesion identification through a hierarchical hot-spot feature fusion network and a dynamic partitioning architecture algorithm based on image complexity. The present invention constructs a multi-modal medical image data set including CT, MRI, and PET images, making full use of the complementarity of multi-modal images. Since traditional multi-modal data fusion methods usually ignore the feature differences of different modal images, and the present invention introduces a hierarchical feature fusion network, which can dynamically adjust the feature fusion weights at multiple levels and automatically focus on the features with obvious distinguishability for lesions, thus effectively solving the problems of local detail loss or over-fusion that occur in traditional methods during image fusion.

[0004] In terms of the chunking strategy, traditional methods mostly adopt the image chunking method with a fixed scale, while the present invention proposes a dynamic architecture division algorithm based on image complexity, which finely divides different regions according to the complexity of the image, enabling the accurate classification of lesion regions. This dynamic chunking strategy can not only ensure the retention of details in local regions of the image but also avoid over-segmentation or loss of key information; in addition, the present invention adopts an iterative classification algorithm based on Bayesian uncertainty estimation to accurately classify lesion features and automatically adjusts the classification threshold in combination with the medical expert knowledge base, thereby making the classification result more reliable and meeting clinical requirements. This process enhances the stability of the classification result and can effectively avoid misjudgment caused by uncertainty in practical applications. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.

[0006] In view of the above existing problems, the present invention is proposed.

[0007] Therefore, the technical problems solved by the present invention are: solving the problems of insufficient accuracy in multi-modal medical image fusion and inaccurate analysis of lesion regions in the prior art.

[0008] To solve the above technical problems, the present invention provides the following technical solutions: obtaining multiple medical image data and constructing a multi-modal medical image dataset including CT, MRI, and PET images;

[0009] Training the multi-modal medical image dataset using a hierarchical attention feature fusion network to generate a lesion recognition and classification model;

[0010] Chunking the target medical image using a dynamic architecture division algorithm based on image complexity and analyzing it using the lesion recognition and classification model to obtain lesion features;

[0011] Classifying the lesion features using an iterative classification algorithm based on Bayesian uncertainty estimation and calculating the classification threshold in combination with the medical expert knowledge base to output the final classification result of the lesion.

[0012] As a preferred solution of the method for lesion recognition and classification in medical images according to the present invention, the construction of the multi-modal medical image dataset includes:

[0013] Collecting multi-modal medical image data of patients to obtain medical image sequences under different devices, including CT images, MRI images, and PET images;

[0014] Perform denoising processing based on wavelet transform, and eliminate noise interference by setting three-layer wavelet reconstruction and soft threshold denoising function;

[0015] Perform anisotropic diffusion activation, and use the diffusion coefficient function based on edge detection to achieve image smoothing while maintaining edge information;

[0016] Adopt a multi-modal registration method based on mutual information to register CT images, MRI images and PET images, align different modal images in the same coordinate system, that is, perform multi-modal image spatial registration, and optimize the spatial transformation parameters by minimizing the registration objective function;

[0017] Among them, the registration adopts the B-spline interpolation method to construct the deformation field, and the control point spacing is set to 10mm;

[0018] Adopt the N4 algorithm for bias field correction to eliminate the scanner signal inhomogeneity in MRI images;

[0019] Perform Z-score normalization processing on the corrected images to make the intensity distributions between different images comparable;

[0020] Construct a multi-modal feature descriptor to extract complementary feature information from different modal images;

[0021] Organize the extracted multi-dimensional features into a data set. Among them, each sample contains the corresponding label information of the features from three modalities of CT, MRI and PET. The dimension provided by the features is the sum of the feature dimensions of each modality, and a complete sample training set is constructed, that is, a multi-modal medical image data set.

[0022] As a preferred scheme of the method for identifying and classifying lesions in medical images according to the present invention, the extracting complementary feature information from different modal images includes:

[0023] Extract Hounsfield units and texture features based on the co-occurrence image matrix from CT images;

[0024] Extract T1, T2 signal intensities and lesion volume, shape and boundary curvature features from MRI;

[0025] Extract SUV value, metabolic volume and total glycolysis value features from PET images.

[0026] As a preferred scheme of the method for identifying and classifying lesions in medical images according to the present invention, the mathematical expression formula of the multi-modal medical image data set is:

[0027]

[0028] Among them, D is a medical image multi-modal dataset, and X i is the feature vector of the i-th sample, and Y i is the corresponding label of the i-th sample, N is the total number of dataset samples, and f CT is the feature extracted from the CT image, and f MRI is the feature extracted from the MRI image, and f PET is the feature extracted from the PET image.

[0029] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, the hierarchical attention feature fusion network includes a feature extraction layer, a multi-dimensional feature fusion layer, and a classification prediction layer, where:

[0030] The feature extraction layer is used for multi-modal feature extraction based on the improved DenseNet;

[0031] The multi-dimensional feature fusion layer is used for dual attention mechanism and cross-modal fusion;

[0032] The classification prediction layer is used for multi-level feature integration and lesion classification.

[0033] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, obtaining the lesion features includes:

[0034] Adopt a complexity evaluation method based on image gradient and local entropy to analyze the local region features of the image;

[0035] According to the calculated local complexity, perform dynamic grid division and optimize it using an energy function;

[0036] Apply the lesion identification and classification model to each divided grid block to extract a multi-dimensional feature map;

[0037] Adopt an attention mechanism to adaptively weight and fuse the features;

[0038] Based on the fused features, construct a standardized lesion feature framework;

[0039] Adopt a feature selection strategy based on L1 regularization, and remove redundant and related features by optimizing the objective function including the loss function and the regularization term, so as to obtain a concise lesion feature description.

[0040] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, constructing the standardized lesion feature framework includes:

[0041] Spatial position information: the center coordinates and bounding box parameters of the lesion;

[0042] Morphological features: geometric parameters such as area, perimeter, circularity, etc.;

[0043] Texture features: statistics based on the gray-level co-occurrence matrix;

[0044] Intensity features: average gray value, standard deviation, skewness, kurtosis;

[0045] Edge features: edge intensity, direction consistency;

[0046] Context features: describe the relationship between the lesion and the surrounding tissues.

[0047] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, a knowledge base is established in combination with medical expert knowledge for classifying and enhancing lesion features, and the knowledge base includes a lesion morphology feature rule base, an imaging feature rule base, and a clinical diagnosis experience base.

[0048] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, the iterative classification algorithm using Bayesian uncertainty estimation is used to classify the lesion features, and the classification threshold is calculated in combination with the medical expert knowledge base, including:

[0049] The output of the Bayesian neural network model and the rules of the expert knowledge base are fused through an adaptive weight, and the weight coefficient α1 of the fusion is dynamically calculated according to the uncertainty of the model:

[0050]

[0051] where U model is the output uncertainty predicted by the model, is the temperature parameter used to control the variability of the weight, and σ is the sigmoid function used to convert the uncertainty into a weight coefficient between 0 and 1;

[0052] The comprehensive classification score is:

[0053] S(x) = α1·P model (x) + (1 - α1)·P expert (x)

[0054] where P model (x) is the prediction probability of the Bayesian neural network model, and P expert (x) is the score based on expert knowledge;

[0055] The method based on ROC curve analysis is used to determine the optimal classification threshold, and the optimal threshold t * is selected by calculating the true positive rate and false positive rate at different thresholds, and the threshold that makes the classification performance optimal is selected as the decision boundary:

[0056] t * = argmax t (TPR(t) - FPR(t))

[0057]

[0058]

[0059] where TP, TN, FP, and FN are true positive, true negative, false positive, and false negative, respectively.

[0060] As a preferred solution of the method for identifying and classifying lesions in medical images according to the present invention, the output of the final classification result of the lesion includes the disease category, classification confidence, and uncertainty estimate value;

[0061] where the disease category includes benign and malignant.

[0062] Advantages of the present invention: In the processing of complex medical image data and lesion classification, the present invention improves the accuracy and robustness of recognition and analysis. Through the hierarchical attention feature fusion network and the dynamic architecture division algorithm based on image complexity, it can more accurately extract valuable features in multi-modal medical images and flexibly adjust the image analysis method, avoiding the deficiencies brought by the fixed final block strategy in the prior art. Combining the iterative classification algorithm of Bayesian uncertainty estimation ensures the accuracy and reliability of the lesion classification result. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0064] Figure 1 is a schematic flowchart of the method for identifying and classifying lesions in medical images shown in the present invention;

[0065] Figure 2 is a schematic diagram of the contrast curve of lesion recognition complexity shown in the present invention;

[0066] Figure 3 is a schematic diagram of the contrast curve of lesion classification accuracy shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.

[0068] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0069] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0070] According to an embodiment of the present invention, in combination with Figure 1 the flowchart shown, a method for identifying and classifying lesions in medical images specifically includes the following steps:

[0071] S1. Obtain multiple medical image data and construct a multi-modal medical image dataset including CT, MRI, and PET images. Herein, it should be noted that:

[0072] Collect multi-modal medical image data of patients to obtain medical image sequences under different devices, including CT images (obtaining Hounsfield gray value information), MRI images (obtaining T1 and T2 signals under different weighting modes), and PET images (obtaining metabolic features such as standardized uptake value (SUV));

[0073] Furthermore, reconstruct the obtained medical image data, including:

[0074] Perform denoising processing based on wavelet transform, and eliminate noise interference by setting three-layer wavelet reconstruction and soft-threshold denoising function. The denoising function is:

[0075] D(x) = W -1 (T λ (W(x)))

[0076]

[0077] where x is the input original image signal, W(x) is the coefficient representation after wavelet transform of the input image x, T λ (·) is the soft-threshold denoising function, λ is the adaptive threshold for denoising, σ is the standard deviation of noise estimation, n is the total number of pixels of the image, and W -1 is the wavelet inverse transform operation to restore the denoised image;

[0078] Exemplarily, the soft threshold function T λ is defined as:

[0079]

[0080] When the absolute value of w is large (e.g., |w| > λ):

[0081] If w > λ, then T λ (w) = w - λ (shrinking λ towards the 0 direction);

[0082] If w < -λ, then T λ (w) = w + λ (shrinking λ towards the 0 direction);

[0083] Through the above limiting operations, small-amplitude noise interference can be effectively reduced while retaining larger signal components;

[0084] When the absolute value of w is small (e.g., |w| ≤ λ), it is directly set to 0, that is, T λ (w) = 0, to remove noise components smaller than the threshold and make the signal cleaner;

[0085] It should be noted that the soft threshold function is used to attenuate or even completely zero out signals close to 0, while signals with larger amplitudes are only slightly shrunk. It is a smoothing denoising method. Compared with the hard threshold (direct truncation), the soft threshold is smoother at the signal edges and reduces possible breakage effects.

[0086] In a preferred embodiment, anisotropic diffusion activation is performed. While achieving image smoothing using a diffusion coefficient function based on edge detection, the edge information is retained. The diffusion equation is:

[0087]

[0088] where t is time, is the rate of change of the image intensity I with respect to time t, is the gradient of the image, representing the direction and magnitude of the pixel intensity change, is the modulus of the gradient, representing the intensity of the edge, is the divergence operator, representing the diffusion process, is the diffusion coefficient control function, and its mathematical expression formula is:

[0089]

[0090] where K is the edge control parameter to prevent excessive edge smoothing.

[0091] In an alternative embodiment, a multi-modal registration method based on mutual information (MI) is used for registration between CT, MRI, and PET, aligning different-modal images in the same coordinate system, i.e., performing multi-modal image spatial registration, and optimizing the spatial transformation parameters by minimizing the registration objective function. The registration optimization objective function is:

[0092]

[0093] where E(T) is the objective function of registration, T is the spatial transformation to be solved (such as affine transformation, non-rigid transformation), I f is the fixed reference image, I m is the moving image to be registered, is the joint probability distribution of the fixed image I f and the transformed image T(I m ), is the marginal probability distribution of the fixed image I f , is the marginal probability distribution of the transformed image T(I m ), ∑ i,j is the summation over all possible gray levels i, j, α is the regularization weight factor, and R(T) is the regularization term, which constrains the smoothness of the transformation T;

[0094] Exemplarily, the registration is implemented using the B-spline interpolation method to construct the deformation field, and the control point spacing is set to 10 mm.

[0095] Furthermore, intensity normalization processing is performed on the registered images, including:

[0096] The N4 algorithm is used for bias field correction to eliminate the scanner signal non-uniformity in the MRI image. The bias field correction formula is:

[0097]

[0098] where V(x) is the image after bias field correction, U(x) is the original image, and F(x) is the estimated bias field;

[0099] Z-score normalization processing is performed on the corrected image to make the intensity distributions between different images comparable;

[0100] A multi-modal feature descriptor is constructed to extract complementary feature information from different-modal images, including:

[0101] Extract the Hounsfield unit and texture features based on the co-occurrence matrix from the CT image;

[0102] Extract T1 and T2 signal intensities, as well as the volume, shape, and boundary curvature features of the lesion from MRI;

[0103] Extract SUV value (standard uptake value), metabolic volume, and total lesion glycolysis (TLG) features from PET images;

[0104] Organize the extracted multi-dimensional features into a data set. Each sample contains the corresponding label information of the features from three modalities: CT, MRI, and PET. The dimension provided by the features is the sum of the feature dimensions of each modality, and a complete sample training set is constructed to provide standardized input data for the training of subsequent deep learning models. The mathematical expression formula of the data set is:

[0105]

[0106] where D is the multi-modal medical image data set, X i is the feature vector of the i-th sample, Y i is the corresponding label of the i-th sample (such as lesion category, diagnosis result), N is the total number of samples in the data set, f CT is the feature extracted from CT images (such as Hounsfield unit histogram statistics, co-occurrence matrix texture features), f MRI is the feature extracted from MRI images (such as T1 and T2 signal intensities, lesion volume, shape, boundary features), f PET is the feature extracted from PET images (such as standard uptake value, metabolic volume, total lesion glycolysis).

[0107] S2. Use a hierarchical attention feature fusion network to train the multi-modal medical image data set to generate a lesion recognition and classification model. It should be noted in this step that:

[0108] Construct a hierarchical attention feature fusion network, including a feature extraction layer, a multi-dimensional feature fusion layer, and a classification prediction layer. The feature extraction layer is used for multi-modal feature extraction based on an improved DenseNet (dense connection network structure). The multi-dimensional feature fusion layer is used for dual attention mechanism and cross-modal fusion. The classification prediction layer is used for multi-level feature integration and lesion classification;

[0109] Exemplarily, for the feature extraction layer, an improved DenseNet structure is adopted. The multi-module extraction of features is realized through dense blocks. Each dense block contains 6 convolutional layers, and the feature is set to 32, realizing the progressive extraction of shallow features to fine features. At the same time, the feature expression ability of the network is improved through feature reuse. The output feature of each dense block is calculated as follows:

[0110] F l =H l ([x0,x1,,xl-1 )

[0111] Among them, F l represents the output feature map of the l-th layer, and H l is the transformation function of this layer, such as ReLU activation and convolution operation, [x0, x1,, x l-1 represents concatenating the feature maps of the previous l layers and inputting them into the l-th layer for processing;

[0112] Construct a multi-dimensional feature fusion layer, enhance feature expression by introducing a dual attention mechanism, and weight the output feature map. The dual attention mechanism includes channel attention and spatial attention. Capture the dependencies between feature channels through the channel attention mechanism and enhance important feature channels. Among them, the channel attention weight is calculated through global pooling and a gradient perception machine;

[0113] As an example, the calculation formula for the channel attention weight is:

[0114] M c =σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))

[0115] Among them, M c represents the channel attention weight, which is used to adjust the importance of each channel. σ is the Sigmoid activation function, which is used to compress the output to the range of [0, 1] to control the range of the weight. MLP is an enhancer that learns the weighted representation of the pooled features. AvgPool(F) performs global average pooling on the feature map F to extract the global features of the image. MaxPool(F) performs global max pooling on the feature map F to extract the local features of the image. F represents the input feature map, which contains different spatial and channel information;

[0116] Obtain the key regions of the attention image through the spatial attention mechanism. The spatial focus weight is extracted from the average pooling and max pooling features. The calculation formula for the spatial focus weight is:

[0117] M s =σ(Conv([AvgPool(F); MaxPool(F)]))

[0118] Among them, M s represents the spatial attention weight, which highlights the key regions in the image. Conv is a convolution operation used to process the concatenated pooled features and generate a spatial attention map. [AvgPool(F); MaxPool(F)] represents concatenating the average pooled and max pooled feature maps in the channel dimension;

[0119] Furthermore, a cross-modal adaptive weighted fusion strategy is adopted to fuse the feature maps of CT, MRI, and PET. The features from different modalities are integrated using an adaptive weighting method, and the weight fusion is re-obtained through automatic learning by the network, ensuring that the features of different modalities can be dynamically fused according to their distributions. At the same time, the key information in each modality is highlighted through each focus map:

[0120]

[0121] Among them, F fused represents the fused feature map, w i represents the fusion weight of the i-th modality, A i represents the attention map of the i-th modality, F i represents the feature map of the i-th modality, and M is the number of modalities;

[0122] In an alternative embodiment, an optimization objective function is defined, which includes a classification loss and a feature consistency loss. Among them, the classification loss uses a cross-disciplinary form to guide the network to learn correct classification, and the feature consistency loss is used to constrain the semantic clustering of features of different modalities. The relative importance of the two losses is adjusted by setting a balance factor of 0.1:

[0123] L total = L ce + τ·L consistency

[0124] Among them, L total represents the total loss function, which includes the classification loss and the feature consistency loss, L ce represents the cross-entropy loss, which is used to supervise the network to learn classification during the training process, L consistency represents the feature consistency loss, which is used to constrain the semantic consistency of features of different modalities, and τ is the balance factor, which is used to adjust the weights of the classification loss and the consistency loss, and is set to 0.1;

[0125] An optimized iterative training strategy is adopted, and the Adam optimizer is used for parameter update. The final learning rate is set to 0.001, and the cosine annealing strategy is used to dynamically adjust the learning rate, which is reduced once every 50 training epochs. When the performance of the validation set has not improved for 10 consecutive epochs, the training stops; specifically:

[0126] Use the Adam optimizer, set the initial learning rate to 0.001, and decay it by 10 times every 50 epochs;

[0127] Adopt the cosine annealing learning rate adjustment strategy:

[0128]

[0129] Among them, lr represents the current learning rate, lrmin is the minimum learning rate, lr max is the maximum learning rate, t is the current training epoch, and T is the total number of training epochs;

[0130] Use the early stopping strategy to stop training when the performance on the validation set has not improved for 10 consecutive epochs;

[0131] Integrate the trained models:

[0132] Execute the model integration process, train 5 models with different random initializations, and the final prediction result is obtained by weighted average:

[0133]

[0134] Among them, P ensemble represents the prediction result of the integrated model, v j represents the weight of the j-th model, P j represents the prediction result of the j-th model, and N1 is the number of models participating in the integration;

[0135] Specifically, the model training process is carried out under the conditions of a batch size of 16 and 200 training epochs. Mixed-precision training is used to accelerate the calculation process, and at the same time, gradient fine-tuning is adopted to prevent gradient explosion. The gradient threshold is set to 5.0 to ensure the stability of the training process.

[0136] The finally generated lesion recognition and classification model is saved in a standardized format, including the complete network structure information and trained weight parameters, enabling the model to directly recognize new medical image analysis tasks.

[0137] It should be noted that in the network structure, channel attention and spatial attention mechanisms are introduced to highlight the key lesion region features in each modality and suppress the interference features. At the same time, through the hierarchical feature extraction module, the lesions can be characterized and fused at the local and global scales respectively; in this way, the network can not only retain the perception of tissue morphology, density, and activation sensitivity, but also take into account the coherence of the overall spatial structure, ultimately improving the recognition accuracy and generalization ability of disease types.

[0138] S3. The target medical image is segmented using the dynamic architecture partitioning algorithm based on image complexity, and the lesion recognition and classification model is used for analysis to obtain lesion features. Among them, it should be noted in this step:

[0139] Adopt a complexity evaluation method based on image gradient and local entropy to analyze the local region features of the image; including:

[0140] Calculate the entropy value distribution of the local region: measure the richness of image information. The more complex the region, the larger the entropy value;

[0141] Calculate the gradient change of the image: It reflects the change of image intensity, especially edge features;

[0142] Calculate the local pixel values: The pixel distribution in the local area reflects the level of detail in the image;

[0143] By setting different weight coefficients, balance these three feature dimensions to obtain an accurate local complexity evaluation result;

[0144] As an example, its mathematical expression formula is:

[0145]

[0146] Among them, C(x,y) is the local complexity at position (x,y), α, β, γ are weight coefficients, which respectively control the influence of local entropy, image gradient, and pixel values on the complexity evaluation. is the L2 norm of the image gradient, representing the edge intensity of the image, V(x,y) is the pixel value of the local area, and H(x,y) is the local entropy, which is calculated through the probability distribution of local pixel values:

[0147] H(x,y) = -∑p(i)log(p(i))

[0148] Among them, p(i) is the probability distribution of the pixel value at position (x,y);

[0149] According to the calculated local complexity, perform dynamic grid division; including:

[0150] In areas with higher complexity, use a smaller grid size to ensure the extraction of detailed features;

[0151] While in areas with lower complexity, use a larger grid size to improve computational efficiency;

[0152] The variation range of the grid size is determined by the preset minimum size S min and the maximum size S max ;

[0153] Exemplarily, the grid size calculation formula:

[0154] S(x,y) = S min +(S max -S min )·exp(-k·C(x,y))

[0155] Among them, S(x,y) is the grid size at position (x,y), S min and S max are respectively the minimum and maximum values of the grid size, and k is an adjustment coefficient that controls the influence of complexity on the grid size.

[0156] To make the grid boundary smoother and avoid discontinuity of image features caused by grid division, an energy function is used for optimization in this embodiment. The energy function consists of a data term and a smoothing term. The data term ensures a good match between the grid boundary and image features, and the smoothing term constrains the size change between adjacent grids, which is adjusted by a balance parameter:

[0157] E(B) = E d (B) + ∈·E s (B)

[0158] where E d (B) is the data term, representing the matching degree between the grid boundary and image features, and E s (B) is the smoothing term, ensuring a stable size change between adjacent grids, and ∈ is the balance parameter, controlling the weights of the data term and the smoothing term.

[0159] Apply the lesion recognition and classification model to each grid block obtained by division to extract a multi-dimensional feature map;

[0160] The feature map includes morphological features (such as geometric parameters like area, perimeter, circularity, etc.), texture features (such as statistics like energy and entropy based on the co-occurrence matrix), intensity features (such as statistical features like average gray value, standard deviation, skewness, kurtosis, etc.), and edge features (such as edge intensity and direction consistency);

[0161] Exemplarily, the feature vector set is:

[0162] F = {f i |i = 1, 2,, N}

[0163] where f i represents the feature map of the i-th grid block, including the above four types of features.

[0164] In an alternative embodiment, adaptive weighted fusion of features is performed through an attention mechanism, including:

[0165] Calculate the feature correlation between the central grid block and other grid blocks, so as to assign appropriate weights to each grid block to ensure effective integration of feature information;

[0166] The fused feature is expressed as:

[0167] F fused (p) = ∑w i ·F i (p)

[0168] where w i is the fusion weight of the grid block i, and F i (p) is the feature of the i-th grid block; the weight coefficient wi Calculated by the following formula:

[0169]

[0170] Where F c is the feature of the central grid block, representing the relatively important area in the image, and θ(F c ) is the feature transformation of the central grid block F c to obtain the transformed central grid block feature, is the feature transformation of the i-th grid block F i to obtain the transformed grid block feature.

[0171] Based on the fused features, a standardized lesion feature framework is constructed, including:

[0172] Spatial location information: the central coordinates and bounding box parameters of the lesion;

[0173] Morphological features: geometric parameters such as area, perimeter, circularity, etc.;

[0174] Texture features: statistics based on the gray-level co-occurrence matrix;

[0175] Intensity features: average gray value, standard deviation, skewness, kurtosis;

[0176] Edge features: edge intensity, direction consistency;

[0177] Context features: describe the relationship between the lesion and the surrounding tissues.

[0178] Adopt a feature selection strategy based on L1 regularization. By optimizing the objective function containing the loss function and the regularization term, redundant and correlated features are removed, so as to obtain a concise description of the lesion features;

[0179] The objective function containing the loss function and the regularization term is:

[0180] J(w) = L(w) + ε·∥w∥1

[0181] Where L(w) is the loss function (such as mean squared error, cross-entropy loss, contrast loss), measuring the error of feature selection, ∥w∥1 is the L1 regularization term, representing the sparsity of the feature weights, and ε is the regularization coefficient, controlling the strength of the regularization term.

[0182] Preferably, in this step, the problem of uneven complexity in the image is effectively processed through an adaptive grid division strategy, the accuracy of feature extraction is improved, and a comprehensive lesion feature representation is constructed through multi-topology fusion features and feature selection, providing a reliable feature basis for subsequent classification tasks.

[0183] Preferably, in this step, by applying a dynamic architecture partitioning algorithm to the target medical image, for the parts with structural abnormalities or signal changes around the lesion, the algorithm can automatically improve the discrimination ability, reduce possible edge blurring or loss of important feature details; while in the areas with simple background and insignificant texture changes, the computational burden is reduced and the overall efficiency is improved; by inputting these segmented results into the trained lesion recognition and classification model, more accurate local information can be extracted.

[0184] S4. Use an iterative classification algorithm based on Bayesian uncertainty estimation to classify the lesion features, and combine with the medical expert knowledge base to calculate the classification threshold, and output the final classification result of the lesion. It should be noted in this step that:

[0185] Construct a Bayesian neural network classifier, which uses variational inference method to estimate the model parameters, and realizes the optimization of the model parameters by minimizing the KL divergence between the model distribution and the true posterior distribution;

[0186] Exemplarily, the Bayesian neural network estimates the posterior distribution of the model parameters through the variational inference method and uses this distribution to generate uncertainty evaluation. In variational inference, the estimation of the posterior distribution is optimized using KL divergence minimization. The improved expression and calculation process are as follows:

[0187] p(w|D)∝p(D|w)·p(w)

[0188] Among them, p(w|D) is the posterior distribution given the training data D, p(D|w) is the likelihood function of the model under the given parameters w, and p(w) is the prior distribution;

[0189] As an example, the mathematical expression formula for KL divergence optimization is as follows:

[0190] L(θ)=-E[logp(D|w)]+KL(q(w|θ)∥p(w))

[0191] Among them, E[logp(D|w)] is the expectation of the data log-likelihood, representing a measure of the model's prediction ability under the given data, KL(q(w|θ)∥p(w)) is the KL divergence between the variational distribution q(w|θ) and the prior p(w), and q(w|θ) is the variational distribution with parameters θ, that is, variational parameters;

[0192] During the model training process, the embodiment of the present invention uses the Adam optimizer to dynamically adjust the learning rate, and realizes the adaptive update of the parameters by setting the initial learning rate and momentum parameters. During the training process, the model parameters are continuously updated and optimized according to the instantaneous information of the loss function until they converge to the optimal solution;

[0193] In each iteration, Adam updates the parameters by calculating the gradient information, and the formula is as follows:

[0194]

[0195] where θ (t) is the variational parameter at the t-th iteration, η is the adaptive learning rate, and θ (t+1) is the variational parameter at the (t + 1)-th iteration, is the gradient of the loss function with respect to the variational parameter θ, representing the rate of change of the loss function;

[0196] This formula dynamically adjusts the learning rate, enabling better optimization of the model parameters and convergence to the optimal solution during the training process.

[0197] In an alternative embodiment, in combination with medical expert knowledge, a knowledge base is established for enhancing the classification of lesion features; the knowledge base includes:

[0198] (1) Lesion morphological feature rule base:

[0199] R m ={(f i , w i , t i ) | i = 1, 2,, N m}

[0200] where f i is the i-th morphological feature (such as area, perimeter), w i is the weight of this feature, and t i is the threshold of this feature;

[0201] (2) Imaging feature rule base:

[0202] R i ={(g j , v j , s j ) | j = 1, 2,, N i}

[0203] where g j is the j-th imaging feature (such as texture, edge intensity), v j is the weight of this feature, and s j is the reference standard of this feature;

[0204] (3) Clinical diagnosis experience base:

[0205] R c ={(h k , u k , c k ) | k = 1, 2,, N c}

[0206] Among them, h k is the k-th clinical feature (such as the patient's age, symptoms), u k is the credibility of this feature, c k is the diagnosis suggestion based on this feature;

[0207] Preferably, in order to improve the accuracy and reliability of the classification results, in this embodiment, the output of the Bayesian neural network model and the rules of the expert knowledge base are fused through adaptive weights, and the weight coefficient α1 of the fusion is dynamically calculated according to the uncertainty of the model:

[0208]

[0209] Among them, U model is the output uncertainty predicted by the model, is the temperature parameter, which is used to control the variability of the weight, and σ is the sigmoid function, which is used to convert the uncertainty into a weight coefficient between 0 and 1;

[0210] The comprehensive classification score is:

[0211] S(x) = α1·P model (x) + (1 - α1)·P expert (x)

[0212] Among them, P model (x) is the prediction probability of the Bayesian neural network model, and P expert (x) is the score based on expert knowledge.

[0213] In an alternative implementation, a method based on ROC curve analysis is used to determine the optimal classification threshold, and the optimal threshold t * is selected by calculating the true positive rate (TPR) and false positive rate (FPR) at different thresholds, and the threshold that makes the classification performance optimal is selected as the decision boundary:

[0214] t * = argmax t (TPR(t) - FPR(t))

[0215]

[0216]

[0217] Among them, TP, TN, FP, and FN are true positive, true negative, false positive, and false negative respectively.

[0218] The finally output classification report includes the following content:

[0219] Basic information of the lesion: description of spatial location and size, morphological features (such as area, perimeter, etc.), and imaging features (such as texture features, edge intensity, etc.);

[0220] Classification results: disease category (such as benign or malignant), classification confidence (confidence based on model prediction), uncertainty estimation (uncertainty output by the model);

[0221] Quantitative indicators: probability distribution of each category, ROC curve and AUC value, classification decision threshold;

[0222] Clinical suggestions: diagnostic suggestions based on expert rules.

[0223] It should be noted that by applying the Bayesian uncertainty estimation method to the lesion features obtained by the above partitioning for iterative classification, the prediction of the posterior distribution of the model parameters is realized. On the one hand, the Bayesian neural network can give the probability distribution of the disease category and its uncertainty, enabling doctors to carefully understand the confidence of the model in the disease category when viewing the classification results;

[0224] In addition, samples with high uncertainty can be deeply integrated with the expert knowledge base, and then combined with ROC analysis or clinical experience to dynamically adjust the threshold, which helps to reduce the misclassification of "boundary samples" or "difficult lesions", improves the stability and safety of classification prediction, and enhances the diagnostic support ability for complex cases.

[0225] Preferably, in order to better verify the good effect of the method of the present invention in the recognition and classification of multi-modal medical image lesions, in this embodiment, a comparison is made with the traditional method adopted by the prior art. The traditional method has the following defects:

[0226] Modal or simple fusion: only using single-modal images such as CT or MRI, lacking the fusion of tumors with information such as single PET, resulting in insufficient characterization of features;

[0227] Fixed partitioning strategy or manual segmentation: dividing the image into grids of fixed size, unable to adjust the partitioning scale according to the image complexity, resulting in insufficient recognition accuracy for complex lesion areas;

[0228] Lack of uncertainty: the traditional method uses a deterministic classifier (such as SVM or random forest), which is difficult to quickly classify the confidence or credibility of the results, and the misjudgment rate for boundary samples and difficult cases is relatively high;

[0229] In summary, the traditional method cannot make full use of multi-modal information and lacks support for complex lesion areas and uncertainty assessment. After the method of the present invention introduces multi-modal fusion, adaptive partitioning of image complexity, and Bayesian uncertainty estimation, it is expected to significantly improve the accuracy and stability of disease classification.

[0230]

Subjects of the Experiment

[0231] There were a total of 6 clinical imaging cases with typical or suspected lesions, numbered from Case 1 to Case 6 respectively. Each case included CT, MRI (T1-weighted, T2-weighted or enhanced sequences), and PET scan images in three modalities;

[0232] The types of lesions included benign, malignant, and some difficult-to-diagnose lesions, distributed in the lung, liver, and brain regions;

[0233]

Experimental Environmental Conditions

[0234] Intel Xeon CPU (16 cores, 2.3 GHz), 128 GB of memory, NVIDIA GPU (CUDA supported);

[0235] Operating system: Ubuntu 20.04;

[0236] Deep learning framework: PyTorch (≥1.8) or TensorFlow (≥2.4);

[0237] MATLAB 2021b: For post-analysis of results and imaging accuracy;

[0238] Data annotation: Professional radiologists annotated the location of the lesions and their pathological nature (benign / malignant, etc.) according to the clinical examination results to verify the classification accuracy;

[0239] Multi-modal imaging sequence: The CT slice thickness was about 1 mm, the MRI included T1 / T2-weighted sequences, and the PET acquisition used 18F-FDG tracer;

[0240] Ground truth of lesion segmentation: The lesion sites of each case were given by experts in the form of binary masks;

[0241] Lesion category information: Clinical diagnosis results such as lesion morphology, benign / malignant, or subtype;

[0242] Expert knowledge base: A rule set formed by morphological indicators (such as diameter, roundness), interpretation values (such as SUVmax), and other clinical reference values;

[0243]

Experimental Process

[0244] Implementation process of traditional methods:

[0245] The CT images or simply the information of MRI / PET were concentrated on the CT without deep registration or unified normalization;

[0246] The images were segmented into blocks with a fixed size (such as 16×16) without considering the differences in local complexity;

[0247] Extract texture features such as the Gray-Level Co-Occurrence Matrix (GLCM);

[0248] Use a Support Vector Machine (SVM) or Random Forest for classification training to obtain predictions for benign or other types of lesions;

[0249] Do not provide uncertainty, a single output class label;

[0250] The implementation process of the method of the present invention:

[0251] Denoise, perform intensity normalization, and perform high-precision registration based on mutual information on CT, MRI, and PET;

[0252] Input the multi-modal image fusion into the network, and use a dual attention mechanism to mine the key features of lesions in CT, MRI, and PET;

[0253] Through the Adam optimizer combined with an early stopping mechanism, finally generate a lesion recognition and classification model;

[0254] Judge the image complexity according to indicators such as local gradient, entropy, and moment, and perform analysis on regions with high complexity to improve the recognition accuracy of lesion edges;

[0255] After partitioning, input the local image patches into the trained model to obtain a high-dimensional lesion feature representation;

[0256] Use a Bayesian neural network to classify the lesions, iterate and update, and output confidence and uncertainty;

[0257] For samples corresponding to the uncertainty, introduce an expert rule base to dynamically adjust the threshold or perform another iteration;

[0258] Finally, give an interpretable disease classification result and give an uncertainty assessment.

[0259] In this experiment, 6 case samples (Case1 - Case6) were used to perform lesion recognition and classification using traditional methods and the method of the present invention respectively, and the following indicators were recorded respectively:

[0260] Classification accuracy (%): Indicates whether the lesion is correctly classified, with the expert diagnosis result as the control;

[0261] Lesion recognition complexity (unitless relative value): Refers to the adaptability of the model to complex lesion regions in the image; the value gain indicates better processing ability for images with high clutter or lesions with blurred boundaries.

[0262] Refer to Figure 2, it can be intuitively seen that the complexity of lesion recognition reflects the adaptability to local high-variability or low-recovery regions of the image. The traditional method is about 0.41, while the method of the present invention reaches above 0.67. The recognition complexity values of the method of the present invention in Case 3 and Case 5 are significantly higher, indicating that it has a stronger ability to distinguish and process lesions with high clutter and irregular edges, reducing the probability of missed detection or false detection.

[0263] Referring to Figure 3 , it can be intuitively seen that the accuracy of the traditional method among 6 cases is between 65% and 78%, while the method of the present invention is between 82% and 90%. Among them, the difference in Case 4 is particularly obvious: the accuracy rate of the traditional method in this case reaches 65%, while the method of the present invention reaches 82%. This shows that in the scenario where the lesion boundary is blurred and the morphology is complex, the method of the present invention can effectively utilize multi-modal information and dynamic block strategy to improve the recognition rate of difficult-to-detect lesions.

[0264] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for identifying and classifying lesions in medical images, characterized in that: include: Acquire multiple medical imaging data and construct a multimodal medical imaging dataset including CT, MRI and PET images; Using a hierarchical attention feature fusion network to train the multimodal medical imaging dataset to generate a lesion recognition and classification model; The target medical image is divided into blocks using a dynamic architecture partitioning algorithm based on image complexity, and the lesion recognition and classification model is used for analysis to obtain lesion features; The lesion features are classified using an iterative classification algorithm based on Bayesian uncertainty estimation, and the classification threshold is calculated in combination with a medical expert knowledge base to output a final classification result on the lesion.

2. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: The constructing of a multimodal medical imaging dataset comprises: Collect multimodal medical imaging data of patients and obtain medical imaging sequences under different devices, including CT images, MRI images and PET images; Perform denoising based on wavelet transform to eliminate noise interference by setting three-layer wavelet reconstruction and soft threshold denoising function; Perform anisotropic diffusion activation, using a diffusion coefficient function based on boundary detection to achieve image smoothing while maintaining edge information; A multimodal registration method based on mutual information is used to register CT images, MRI images and PET images, so that images of different modalities are aligned in the same coordinate system, that is, multimodal image spatial registration is performed, and spatial transformation parameters are optimized by minimizing the registration objective function. Among them, the registration uses the B-spline interpolation method to realize the construction of the deformation field, and the control point spacing is set to 10mm; The N4 algorithm is used for bias field correction to eliminate scanner signal inhomogeneity in MRI images; Perform Z-score normalization on the corrected images to make the intensity distributions between different images comparable; Construct a multimodal feature descriptor to extract complementary feature information from images of different modalities; The extracted multidimensional features are organized into a data set, where each sample contains the corresponding label information from three modality features: CT, MRI, and PET. The dimension provided by the feature is the sum of the dimensions of each modality feature, and a complete sample training set, namely, a multimodal medical imaging data set, is constructed.

3. The method for identifying and classifying lesions in medical images according to claim 2, characterized in that: The step of extracting complementary feature information from images of different modalities includes: Extract Hounsfield units and texture features based on co-occurrence image matrix from CT images; Extract T1 and T2 signal intensity as well as the volume, shape and boundary curvature characteristics of the lesion from MRI; SUV value, metabolic volume and total glycolysis value features were extracted from PET images.

4. The method for identifying and classifying lesions in medical images according to claim 1 or 2, characterized in that: The mathematical expression formula of the multimodal medical image dataset is: Among them, D is a multimodal dataset of medical images, X i is the feature vector of the i-th sample, Y i is the corresponding label of the i-th sample, N is the total number of samples in the data set, and f CT is the feature extracted from the CT image, f MRI is the feature extracted from the MRI image, f PET are features extracted from PET images.

5. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: The hierarchical attention feature fusion network includes a feature extraction layer, a multi-dimensional feature fusion layer and a classification prediction layer, wherein: The feature extraction layer is used for multimodal feature extraction based on the improved DenseNet; The multi-dimensional feature fusion layer is used for dual attention mechanism and cross-modal fusion; The classification prediction layer is used for multi-level feature integration and lesion classification.

6. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: The obtaining of lesion characteristics comprises: The complexity evaluation method based on image gradient and local entropy is used to analyze the local area features of the image; According to the calculated local complexity, dynamic meshing is performed and optimized using energy functions; Apply the lesion recognition and classification model to each grid block obtained by division to extract a multi-dimensional feature map; Adaptively weight features through attention mechanism; Based on the fused features, a standardized lesion feature framework is constructed; A feature selection strategy based on L1 regularization is adopted to remove redundant and relevant features by optimizing the objective function including the loss function and the regularization term, thereby obtaining a concise lesion feature description.

7. The method for identifying and classifying lesions in medical images according to claim 6, characterized in that: The construction of a standardized lesion feature framework includes: Spatial location information: center coordinates and bounding box parameters of the lesion; Morphological characteristics: geometric parameters such as area, perimeter, and circularity; Texture features: statistics based on gray-level co-occurrence matrix; Intensity characteristics: mean gray value, standard deviation, skewness, kurtosis; Edge features: edge strength, direction consistency; Contextual features: describe the relationship of the lesion to the surrounding tissues.

8. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: In combination with the knowledge of medical experts, a knowledge base is established for classifying and enhancing lesion features, and the knowledge base includes a lesion morphology feature rule base, an imaging feature rule base and a clinical diagnosis experience base.

9. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: The iterative classification algorithm using Bayesian uncertainty estimation is used to classify the lesion features, and the classification threshold is calculated in combination with the medical expert knowledge base, including: The output of the Bayesian neural network model and the rules of the expert knowledge base are fused through adaptive weights, and the fusion weight coefficient α1 is dynamically calculated according to the uncertainty of the model: Among them, U model is the output uncertainty predicted by the model, is the temperature parameter used to control the variability of the weights, and σ is the sigmoid function used to convert uncertainty into a weight coefficient between 0 and 1; The comprehensive classification scores are: S(x)=α1·P model (x)+(1-α1)·P expert (x) Among them, P model (x) is the predicted probability of the Bayesian neural network model, P expert (x) is a score based on expert knowledge; The optimal classification threshold is determined by using a method based on ROC curve analysis, and the optimal threshold t is selected by calculating the true positive rate and false positive rate under different thresholds. * , select the threshold that makes the classification performance optimal as the decision boundary: t * =argmax t (TPR(t)-FPR(t)) Among them, TP, TN, FP and FN are true positive, true negative, false positive and false negative, respectively.

10. The method for identifying and classifying lesions in medical images according to claim 1, characterized in that: Outputting the final classification results of the lesions includes disease category, classification confidence and uncertainty estimation; Among them, disease categories include benign and malignant.

Citation Information

Patent Citations

  • Information extraction system based on domain expert knowledge system and information extraction method thereof

    CN108804408A

  • Automobile fault diagnosis method based on hybrid Bayesian network

    CN112085202A

  • Model training method and device, medical image fusion method and device, equipment and medium

    CN114897756A

  • Multi-modal image fusion method and system for skin disease diagnosis

    CN119445303A

Cited By

  • CT image intelligent analysis system for pneumonia auxiliary screening

    CN120953426A

  • Low-rank DCT compressed sensing algorithm realized based on FPGA

    CN120956909A

  • A Low-Rank DCT Compressed Sensing Method Based on FPGA

    CN120956909B

  • Analysis method and system for multi-tracer image, storage medium and computer equipment

    CN120997133A

  • Focus image determination method and device, electronic equipment and storage medium

    CN121120640A