An artificial intelligence-based medical image recognition method

CN122223325BActive Publication Date: 2026-09-25WUHAN AIYANBANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610318147.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-09-25
Estimated Expiration
2046-03-16

AI Technical Summary

Technical Problem

[0005]本发明的目的在于解决上述背景技术中提到的现有技术在处理复杂病灶、模糊边界和多尺度病灶时分割精度低的问题,而提出一种基于人工智能的医疗影像识别方法

Benefits of technology

[0041]通过实施该技术方案,结合二元交叉熵损失与Dice损失,并引入多层级监督机制,提供全面的训练信号,增强模型对正负样本的平衡学习能力,提升分割结果的稳定性与泛化性能,降低假阳性与假阴性率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223325B_ABST
    Figure CN122223325B_ABST
Patent Text Reader

Abstract

The application discloses a medical image recognition method based on artificial intelligence, and relates to the technical field of machine vision. The method uses a medical image segmentation model containing a three-branch encoder and a multi-level prediction network for lesion recognition. The three-branch encoder includes a Transformer branch for extracting global features, a residual network branch for extracting local features, and a feature fusion main branch for fusing the global features and the local features of the corresponding level to obtain multi-level fusion features. The multi-level prediction network refines and gradually aggregates the fusion features through a feature enhancement module and a hierarchical prediction structure to generate a segmentation prediction result. By fusing global context information and local detail features, the method fully utilizes multi-scale and multi-level lesion information, improves the segmentation accuracy and robustness of complex lesions, fuzzy boundaries and multi-scale lesion areas, and realizes automatic and high-precision medical image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine vision technology, and more specifically to a medical image recognition method based on artificial intelligence. Background Technology

[0002] Medical image segmentation is a key technology in computer-aided diagnostic (CAD) systems. Its core objective is to accurately locate and segment lesion regions within images at the pixel level, providing objective quantitative evidence for early disease screening, treatment planning, and efficacy evaluation. Traditional medical image segmentation methods often rely on manually designed features (such as texture and edges) combined with algorithms like region growing and active contour models. These methods are often sensitive to image quality, exhibit poor robustness in the presence of strong artifacts, weak boundaries, or complex backgrounds, and heavily depend on expert experience, making it difficult to achieve efficient and consistent automated analysis.

[0003] In recent years, semantic segmentation methods based on deep learning, especially convolutional neural networks (CNNs), have made groundbreaking progress in medical image analysis. Encoder-decoder structures, represented by U-Net and its variants (such as U-Net++, Attention U-Net, ResUNet, etc.), significantly improve segmentation accuracy by fusing multi-level features through skip connections. CNNs possess powerful local feature extraction capabilities, effectively capturing image details. However, due to the limitations of their local receptive field in convolutional operations, CNNs are insufficient in modeling long-range dependencies and global contextual information, potentially leading to incomplete segmentation or misclassification when processing lesions with complex structures or extensive contextual associations.

[0004] To overcome the limitations of CNNs, the Transformer architecture has been introduced into the field of computer vision and has shown potential in medical image segmentation. Vision Transformer (ViT) and Swin Transformer, among others, can capture global dependencies in images through self-attention mechanisms, improving the understanding of the overall structure. Existing research has attempted to combine Transformers with CNNs, such as hybrid architectures like TransUNet and Swin-UNet, aiming to simultaneously leverage the local detail extraction capabilities of CNNs and the global context modeling advantages of Transformers. However, existing hybrid methods often employ simple feature concatenation or cascading, failing to fully explore and fuse complementary information between features at different levels and scales, resulting in an ineffective combination of detailed information and global semantics in the image. In multi-scale lesions or images with complex morphologies, this simple fusion method often fails to maintain segmentation accuracy, potentially leading to the loss of local information or missegmentation. Summary of the Invention

[0005] The purpose of this invention is to solve the problem of low segmentation accuracy in the prior art when dealing with complex lesions, blurred boundaries and multi-scale lesions, as mentioned in the background art above, and to propose a medical image recognition method based on artificial intelligence.

[0006] A first aspect of this invention provides an artificial intelligence-based medical image recognition method, which is implemented through a medical image segmentation model; the medical image segmentation model includes a three-branch encoder and a multi-level prediction network; wherein:

[0007] The three-branch encoder includes a global feature extraction branch, a local feature extraction branch, and a feature fusion main branch. The global feature extraction branch is used to extract global features from the input image using a Transformer network to obtain multi-level global feature maps. The local feature extraction branch is used to extract local features from the input image using a residual network to obtain multi-level local feature maps. The feature fusion main branch is used to fuse the global and local feature maps corresponding to each level to obtain multi-level fused features.

[0008] The multi-level prediction network includes multiple feature enhancement modules and a hierarchical prediction network. Each feature enhancement module is used to aggregate and enhance multi-level fused features to obtain a refined feature map with the corresponding level spatial resolution. The hierarchical prediction network adopts a progressive feature aggregation strategy from deep to shallow layers, performing convolution operations, upsampling, and feature fusion on each level of refined feature map in sequence to generate a multi-level segmentation prediction map. After the final level segmentation prediction map is subjected to probability discrimination, the lesion identification and segmentation result is output.

[0009] By implementing this technical solution, combining the global modeling capabilities of Transformer with the local detail extraction advantages of residual networks, a three-branch encoder and a multi-level prediction network are constructed. This fully integrates multi-scale and multi-level lesion information, significantly improving the segmentation accuracy and robustness of complex lesion structures, fuzzy boundaries, and multi-scale lesion regions, thereby achieving automated and high-precision medical image recognition.

[0010] Optionally, before inputting the medical images into the medical image segmentation model, the medical images are preprocessed, and the processed images are then input into the model; the preprocessing includes:

[0011] Medical images are processed using contrast-limited adaptive histogram equalization to enhance low-contrast lesions and obtain the first target image.

[0012] The first target image is scaled up to fit the model input size to obtain the second target image;

[0013] The pixel values ​​of the second target image are normalized to obtain the final image data input to the model.

[0014] By implementing this technical solution, the contrast of the input image is enhanced, reducing segmentation bias caused by image differences and improving the model's generalization ability in real clinical images.

[0015] Optionally, the global feature extraction branch is divided into four stages, each stage including a global feature extraction module GFEM; each global feature extraction module contains multiple SwingTransformer blocks, each numbering 2.

[0016] By implementing this technical solution, which uses a phased SwinTransformer module for global feature encoding, it is possible to efficiently capture long-distance dependencies and global contextual information in images, thereby enhancing the model's ability to understand the overall structure and large-scale semantics of lesions. This is especially suitable for lesion regions with complex structures or discrete distributions.

[0017] Optionally, the local feature extraction branch is divided into four stages, each stage including a local feature extraction module (LFEM); the operation process of any local feature extraction module (LFEM) includes:

[0018] ;

[0019] ;

[0020] Where x represents the input of the feature extraction module; Conv 1×1 Conv 3×3 These represent convolutional layers with kernel sizes of 1×1 and 3×3, respectively; X1 represents the intermediate result produced by the operation; RCM is a preset residual convolutional module; the operator ⊕ represents element-wise addition; X2 represents the output of the feature extraction module; y represents the input of the residual convolutional module; BN represents batch normalization. Y represents the activation function; Y1 and Y2 represent the intermediate results produced by the operation; Y3 represents the output of the residual convolution module.

[0021] By implementing this technical solution, the local feature extraction module based on residual stacking, combined with the initial convolution and cascaded residual structure, enhances the ability to extract and preserve local details, effectively captures subtle features such as lesion edges and textures, improves the segmentation accuracy of small targets and complex boundaries, and alleviates the gradient vanishing problem of deep networks.

[0022] Optionally, the main feature fusion branch includes four feature fusion modules (FFMs), which are denoted as FFM1, FFM2, FFM3, and FFM4 in sequence. The input of FFM1 includes the global feature map and local feature map of the corresponding level. The input of FFM2 to FFM4 also includes the output of the previous level feature fusion module. The operation process of any feature fusion module among FFM2 to FFM4 includes:

[0023] ;

[0024] Among them, G i L i M i-1 These are the global feature map, local feature map, and output feature map of the i-th feature fusion module, respectively; Conv 3×3 This indicates a convolutional layer with a kernel size of 3×3; BN indicates batch normalization; the operator ⊙ indicates element-wise multiplication; concat indicates channel concatenation; F1 to F5 represent intermediate results of the operation; Bottleneck indicates a bottleneck layer; M i This represents the output feature map of the i-th feature fusion module.

[0025] By implementing this technical solution, a multimodal fusion mechanism and a Hadamard product interaction strategy are adopted to deeply integrate global semantic features and local detailed features, avoiding the limitations of a single feature source, improving the completeness and discriminative power of feature expression, and enhancing the segmentation accuracy of the model in complex backgrounds and multi-scale lesions.

[0026] Optionally, for any feature enhancement module, it is denoted as the target feature enhancement module FAM. j Its input includes four fusion features, and the corresponding level of fusion features M i As the primary feature, the remaining three fused features are used as auxiliary features, and are denoted as Z1, Z2, and Z3 respectively according to their spatial resolution from largest to smallest; the operation process of the target feature enhancement module includes:

[0027] ;

[0028] Where AGG is a pre-defined multi-scale aggregation module; σ is the sigmoid activation function; T a For self-attention maps; T r For inverse attention graphs; Conv 3×3 This indicates a convolutional layer with a kernel size of 3×3; reshape indicates dimensionality reshaping; K a V a K r V r K i V i Sa S r It is an intermediate result produced during the calculation process; softmax is a normalization function; the operator ⊙ represents element-wise multiplication; concat represents channel concatenation; D j It is the output feature map of the target feature enhancement module.

[0029] By implementing this technical solution, a dual attention mechanism is introduced, combined with a multi-scale aggregation module, to achieve precise enhancement and noise suppression of the core area and boundary of the lesion, improve the quality and discriminativeness of the feature map, and further optimize the segmentation details and boundary integrity.

[0030] Optionally, the operation process of the multi-scale aggregation module includes:

[0031] ;

[0032] Where Down indicates downsampling; Conv indicates convolution, and the superscript 3×3 is the kernel size; the operator ⊙ indicates element-wise multiplication; the operator ⊕ indicates element-wise addition; and resize indicates resizing. , , , It is an intermediate result generated during the calculation process; Z 123 It is the output feature map of the multi-scale aggregation module.

[0033] By implementing this technical solution, multi-scale auxiliary features are unified and fused into the main features through downsampling, convolution, and multiply-accumulate fusion operations, thereby achieving complementary enhancement of cross-level information and improving the model's adaptability and anti-interference ability to multi-scale lesions.

[0034] Optionally, the hierarchical prediction network includes four prediction modules, each of which performs segmentation prediction based on the refined feature map of the corresponding level. The computation process of the hierarchical prediction network includes:

[0035] ;

[0036] Among them, D j It is the refined feature map output by the j-th feature enhancement module; Conv 1×1 This represents a convolutional layer with a kernel size of 1×1; σ is the sigmoid activation function; P j It is the segmentation prediction map output by the j-th prediction module; Up indicates upsampling; the operator ⊕ indicates element-wise addition.

[0037] By implementing this technical solution, a progressive feature aggregation and prediction strategy from deep to shallow layers is adopted to gradually integrate global semantics and local details, thereby achieving gradual optimization from coarse localization to fine segmentation. This effectively improves the accuracy and completeness of segmentation boundaries and avoids the limitations of single-level prediction.

[0038] Optionally, the loss function used during training of the medical image segmentation model includes binary cross-entropy loss L. BCE And Dice lost L Dice Specifically, the total loss function L total for:

[0039] ;

[0040] Among them, P i It is the segmentation prediction map output by the i-th prediction module, GT i λ1 and λ2 are the labeled true values ​​for the corresponding scale; λ1 and λ2 are the weighting coefficients.

[0041] By implementing this technical solution, combining binary cross-entropy loss and Dice loss, and introducing a multi-level supervision mechanism, comprehensive training signals are provided, enhancing the model's ability to learn a balance between positive and negative samples, improving the stability and generalization performance of the segmentation results, and reducing the false positive and false negative rates. Attached Figure Description

[0042] Figure 1 A network architecture diagram of a medical image segmentation model provided in an embodiment of the present invention;

[0043] Figure 2 This is a schematic diagram of the structure of a local feature extraction module provided in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of the structure of a feature fusion module provided in an embodiment of the present invention;

[0045] Figure 4 This is a schematic diagram of the structure of a feature enhancement module provided in an embodiment of the present invention;

[0046] Figure 5 This is a schematic diagram of the structure of a multi-scale aggregation module provided in an embodiment of the present invention. Detailed Implementation

[0047] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention is provided in conjunction with the accompanying drawings and preferred embodiments.

[0048] This invention provides an artificial intelligence-based medical image recognition method, implemented using a pre-trained medical image segmentation model (TH-Net). After preprocessing the medical image, it is input into the model to obtain the lesion recognition and segmentation results. See also... Figure 1 , Figure 1 This is a network architecture diagram of a medical image segmentation model provided in an embodiment of the present invention. The medical image segmentation model includes a three-branch encoder and a multi-level prediction network; wherein:

[0049] The three-branch encoder consists of a global feature extraction branch, a local feature extraction branch, and a feature fusion main branch. The global feature extraction branch uses a Transformer network to extract global features from the input image, resulting in multi-level global feature maps. The local feature extraction branch uses a residual network to extract local features from the input image, resulting in multi-level local feature maps. The feature fusion main branch fuses the global and local feature maps corresponding to each level to obtain multi-level fused features.

[0050] The multi-level prediction network includes multiple feature enhancement modules and a hierarchical prediction network. Each feature enhancement module is used to aggregate and enhance the multi-level fused features to obtain a refined feature map of the corresponding level spatial resolution. The hierarchical prediction network adopts a progressive feature aggregation strategy from deep to shallow layers, performing convolution operations, upsampling, and feature fusion on the refined feature maps of each level in sequence to generate multi-level segmentation prediction maps. After the final level segmentation prediction map is subjected to probability discrimination, the lesion identification and segmentation results are output.

[0051] The medical image recognition method based on artificial intelligence provided in this invention uses a three-branch encoder structure and a multi-level prediction network. It combines global feature extraction of the Transformer network with local feature extraction of the residual network and integrates multi-scale information to effectively improve the segmentation accuracy of complex lesions, blurred boundaries and multi-scale lesions. It overcomes the problems of sensitivity to image quality and reliance on expert experience of traditional methods and has strong automation and robustness.

[0052] In one embodiment, the preprocessing process includes:

[0053] Step 1: The medical image is processed using contrast-limited adaptive histogram equalization (CLAHE) to enhance low-contrast lesions and obtain the first target image.

[0054] Step two: Scale the first target image to fit the model input size to obtain the second target image.

[0055] Step 3: Normalize the pixel values ​​of the second target image to obtain the final image data input to the model.

[0056] This embodiment uses CLAHE to enhance the lesion details in low-contrast areas of the image, making blurry or difficult-to-identify lesions more prominent, helping the model to more accurately locate lesion areas, and improving robustness, especially in cases of poor image quality.

[0057] In one implementation, the global feature extraction branch is divided into four stages, each stage including a Global Feature Extraction Module (GFEM) for feature encoding. Each GFEM contains multiple SwingTransformer blocks, with a total of two blocks. SwingTransformer is a common variant of Transformer, which will not be elaborated upon here. Figure 1 In this context, "Down" indicates downsampling, which reduces the width and height of the feature map by half. This can be achieved using max pooling or interpolation algorithms with a stride of 2.

[0058] In one implementation, the local feature extraction branch is divided into four stages, each stage including a Local Feature Extraction Module (LFEM) for feature encoding. Compared to the dual-layer convolutional structure of the original U-Net encoder, this invention proposes a residual stacked convolutional module for local feature extraction. See also Figure 2 , Figure 2 This is a schematic diagram of the structure of a local feature extraction module provided in an embodiment of the present invention. Figure 2 (a) shows the overall structure of the LFEM. Figure 2 Figure (b) illustrates a specific structure of a Residual Convolutional Module (RCM). In the figure, 1×1Conv represents a 1×1 convolutional layer, Add represents element-wise addition, and CBR represents Conv (convolution) + BN (batch normalization) + ReLU (activation function). The operation process of the Local Feature Extraction Module (LFEM) includes:

[0059] In the formula, Conv 1×1 This indicates a convolutional layer with a kernel size of 1×1; the operator ⊕ indicates element-wise addition. The operation process of RCM includes:

[0060] In the formula, Conv 3×3 These represent convolutional layers with a kernel size of 3×3; This represents the activation function. Unless otherwise specified, the activation function is ReLU.

[0061] The LFEM proposed in this embodiment first employs a 1×1 convolution to expand the channel dimension and enhance the interaction between channels, capturing higher-dimensional feature associations. Then, it uses a cascaded RCM, through a recursive residual structure, to iteratively refine and accumulate features. This allows the network to not only extract richer local information but also better preserve the fine details of the image, improving the segmentation ability of small objects and the refinement of complex boundaries. The internal residuals of RCM and the global residuals of LFEM can effectively alleviate the gradient vanishing and network degradation problems in deep networks, making the model easier to train and optimize.

[0062] In one implementation, the main feature fusion branch includes four feature fusion modules (FFMs). The input to the FFM includes the global feature map G of the corresponding level. i Local feature map L i and the output M of the previous level feature fusion module i-1 .

[0063] See Figure 3 , Figure 3 This is a schematic diagram of the structure of a feature fusion module provided in an embodiment of the present invention. Figure 3 (a) shows the overall structure of FFM; Figure 3 Figure (b) illustrates a specific structure of a Bottleneck. In the figure, Mul represents element-wise multiplication, and Concat represents channel concatenation.

[0064] The computation process of any feature fusion module in FFM2 to FFM4 includes:

[0065] In the formula, the operator ⊙ represents element-wise multiplication; concat represents channel concatenation.

[0066] In one implementation, FFM1 does not have a higher-level feature fusion module; during fusion, That is, in Figure 3 In the middle, the dotted line portion is not executed.

[0067] The FFM proposed in this embodiment first performs a 3×3 convolution operation on all input features to unify the channel dimension and refine the features. Then, it performs a linear Hadamard product on the refined features at the same position to achieve feature interaction, and concatenates the interactive feature F4 with the original bi-branch features along the channel dimension. Finally, the concatenated features are fed into the residual block to obtain the fused features. FFM, through a multimodal mechanism and Hadamard product, absorbs both the local spatial details of the convolutional network (such as lesion boundaries) and the global contextual information of the Swin-Transformer (such as discrete lesion associations), solving the problem of one-sided single-branch features, thereby improving the segmentation accuracy of large-scale lesions and complex backgrounds.

[0068] In one implementation, this invention proposes a Feature Enhancement Module (FAM) to replace the decoding layer in U-Net. For any given Feature Enhancement Module (FAM)... j Its input includes four fused features output from the main branch of feature fusion, and the corresponding level of fused features M i The remaining three fused features are used as primary features and are designated as auxiliary features, denoted as Z1, Z2, and Z3 in descending order of spatial resolution. For example, when FAM3 processes M2, Z1, Z2, and Z3 are M1, M3, and M4, respectively.

[0069] See Figure 4 , Figure 4 This is a schematic diagram of a feature enhancement module provided in an embodiment of the present invention. In the diagram, AGG is a preset multi-scale aggregation module, Inv represents the negation operation, sigmoid and softmax are two common activation functions, reshape represents dimensionality reshaping (converting a layer into a sequence or a sequence into a layer), and Mul represents element-wise multiplication and / or matrix multiplication. The operation process of the target feature enhancement module includes:

[0070] In the formula, σ is the sigmoid activation function; T a For self-attention maps; T r This is an anti-attention graph.

[0071] In one implementation, for the multi-scale aggregation module AGG, see [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of a multi-scale aggregation module provided in an embodiment of the present invention. In the diagram, "resize" represents size adjustment, implemented through an interpolation algorithm. The AGG calculation process includes:

[0072] .in, and Consistent spatial resolution; , , and Consistent spatial resolution; With the corresponding M i The spatial resolution is consistent.

[0073] AGG unifies three auxiliary feature maps to the same resolution through downsampling and convolution operations, avoiding feature conflicts caused by scale differences. It also fuses information from different levels of features using a multiply-add fusion logic, allowing shallow high-resolution details and deep high-semantic global information to mutually enhance each other, avoiding the limitations of single-scale features. The fused feature map Z... 123It contains richer contextual information, which can effectively resist speckle noise and artifact interference in ultrasound images, allowing subsequent attention maps to more accurately locate lesion areas.

[0074] FAM first generates a self-attention map T using AGG and sigmoid. a Then, an anti-attention map T is generated by inverting (1 - pixel value) and using sigmoid. r Then, the dual attention map and the main feature are mapped to keys K and values ​​V, respectively, and cross-attention operations are performed to obtain the enhanced, refined feature map D. j FAM (Focus-Awareness Analysis) uses self-attention to locate the core of the lesion and suppress non-lesion noise, while using reverse attention to capture the boundary and enhance key features. This not only accurately locates the entire lesion but also captures boundary textures, thus significantly improving segmentation accuracy.

[0075] In one implementation, this invention proposes a hierarchical prediction network. The hierarchical prediction network includes four prediction modules (PRMs), each performing segmentation prediction based on the refined feature map of its corresponding level. Specifically, the computation process of the hierarchical prediction network includes:

[0076] In the formula, P j It is the segmentation prediction map output by the j-th prediction module; Up represents upsampling, which can be achieved through transpose convolution or interpolation algorithms.

[0077] The hierarchical prediction proposed in this embodiment follows the logic of deep semantic guidance and shallow detail supplementation:

[0078] Deep prediction map P1: generated by the deepest refined feature D1, containing global lesion semantic information, responsible for coarsely locating the overall lesion region, and avoiding segmentation results deviating from the core of the lesion.

[0079] Mid-layer prediction maps P2 / P3: By upsampling and fusing the upper-layer prediction maps with the current layer's refined features D2 / D3, mid-layer details such as the shape and outline of the lesion are gradually supplemented.

[0080] Shallow prediction map P4: By fusing P3 with shallow high-resolution refinement feature D4, it captures subtle textures at lesion boundaries, achieving fine segmentation. The entire process progressively transfers deep global semantics to the shallow layer, solving the problem of poor prediction accuracy at a single scale.

[0081] In one implementation, during model training, all P1-P4 components participate in the calculation of the segmentation loss, providing multi-level supervision signals to the model. The loss function of the medical image segmentation model includes the binary cross-entropy loss L. BCE And Dice lost L Dice Specifically, the total loss function L total for:

[0082] ;

[0083] Among them, P i It is the segmentation prediction map output by the i-th prediction module, GT i λ1 and λ2 are the labeled true values ​​for the corresponding scale; λ1 and λ2 are weighting coefficients, which can both be set to 1.

[0084] The multi-level supervision proposed in this embodiment filters out the noise that is easily introduced by single-scale prediction, reducing the probability of false positives (noise misidentified as lesions) and false negatives (lesions missed).

[0085] To verify the effectiveness of the proposed medical image segmentation model TH-Net, qualitative and quantitative evaluations were conducted on a self-built ultrasound liver lesion dataset, and comparative experiments were performed with current mainstream medical image segmentation methods. The experiments used commonly used evaluation metrics in the field of medical image segmentation to quantitatively analyze the model performance, including the Dice similarity coefficient (Dice), accuracy (ACC), and sensitivity (Se). The experimental results for each model are shown in Table 1.

[0086] Table 1

[0087] U-Net 78.5 94.22 76.2 SA-UNet 80.1 95.05 77.9 TransUNet 80.6 95.28 78.6 TH-Net 83.8 97.42 82.6

[0088] Compared to existing U-Net, SA-UNet, and TransUNet, the TH-Net proposed in this embodiment achieves the best performance in evaluation metrics such as Dice, accuracy, and sensitivity, indicating that the model can more accurately and completely segment ultrasound liver lesion regions. Experimental results fully verify that TH-Net has stronger feature representation capabilities and segmentation robustness under complex backgrounds, blurred boundaries, and multi-scale lesion conditions, demonstrating its superior performance in automatic medical image segmentation tasks.

[0089] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention are within the scope of the claims of the present invention.

Claims

1. A medical image recognition method based on artificial intelligence, characterized in that, The method is implemented using a medical image segmentation model; the medical image segmentation model includes a three-branch encoder and a multi-level prediction network; wherein: The three-branch encoder includes a global feature extraction branch, a local feature extraction branch, and a feature fusion main branch. The global feature extraction branch is used to extract global features from the input image using a Transformer network to obtain multi-level global feature maps. The local feature extraction branch is used to extract local features from the input image using a residual network to obtain multi-level local feature maps. The feature fusion main branch is used to fuse the global and local feature maps corresponding to each level to obtain multi-level fused features. The multi-level prediction network includes multiple feature enhancement modules and a hierarchical prediction network. Each feature enhancement module is used to aggregate and enhance multi-level fused features to obtain a refined feature map with the corresponding level spatial resolution. The hierarchical prediction network adopts a progressive feature aggregation strategy from deep to shallow layers, performing convolution operations, upsampling, and feature fusion on each level of refined feature map in sequence to generate a multi-level segmentation prediction map. After probability discrimination, the final level segmentation prediction map is output as the lesion identification and segmentation result. The local feature extraction branch is divided into four stages, each stage including a local feature extraction module (LFEM); the operation process of any local feature extraction module (LFEM) includes: ; ; in, x Indicates the input to the feature extraction module; Conv 1×1 Conv 3×3 These represent convolutional layers with kernel sizes of 1×1 and 3×3, respectively; X1 represents the intermediate result produced by the operation; RCM is the preset residual convolution module; the operator ⊕ indicates element-wise addition; X2 represents the output of the feature extraction module. y This represents the input to the residual convolution module; BN represents batch normalization. Y1 and Y2 represent the intermediate results of the operation; Y3 represents the output of the residual convolution module. The main feature fusion branch includes four feature fusion modules (FFMs), which are denoted as FFM1, FFM2, FFM3, and FFM4 in sequence. The input to FFM1 includes the global and local feature maps of the corresponding level. The inputs to FFM2 through FFM4 also include the output of the previous level feature fusion module. The computation process of any one of the feature fusion modules (FFM2 through FFM4) includes: ; Among them, G i L i M i-1 These are the global feature map, local feature map, and output feature map of the previous-level feature fusion module, respectively, which are the inputs of the i-th feature fusion module; Conv 3×3 This indicates a convolutional layer with a kernel size of 3×3; BN indicates batch normalization; the operator ⊙ indicates element-wise multiplication; concat indicates channel concatenation; F1 to F5 represent intermediate results of the operation; Bottleneck indicates a bottleneck layer; M i This represents the output feature map of the i-th feature fusion module; For any feature enhancement module, let it be denoted as the target feature enhancement module FAM. j Its input includes four fusion features, and the corresponding level of fusion features M i As the primary feature, the remaining three fused features are used as auxiliary features, and are denoted as Z1, Z2, and Z3 respectively according to their spatial resolution from largest to smallest; the operation process of the target feature enhancement module includes: ; Where AGG is a pre-defined multi-scale aggregation module; σ is the sigmoid activation function; T a For self-attention maps; T r This is an inattention map; Conv 3×3 This indicates a convolutional layer with a kernel size of 3×3; reshape indicates dimensionality reshaping; K a V a K r V r K i V i V f S a S r It is an intermediate result produced during the calculation process; softmax is a normalization function; the operator ⊙ represents element-wise multiplication; concat represents channel concatenation; D j It is the output feature map of the target feature enhancement module; The operation process of the multi-scale aggregation module includes: ; Where Down indicates downsampling; Conv indicates convolution, and the superscript 3×3 is the kernel size; the operator ⊙ indicates element-wise multiplication; the operator ⊕ indicates element-wise addition; and resize indicates resizing. , , , It is an intermediate result produced during the calculation process; Z 123 It is the output feature map of the multi-scale aggregation module; The hierarchical prediction network includes four prediction modules. Each prediction module performs segmentation prediction based on the refined feature map of the corresponding level. The operation process of the hierarchical prediction network includes: ; Among them, D j It is the refined feature map output by the j-th feature enhancement module; Conv 1×1 This represents a convolutional layer with a kernel size of 1×1; σ is the sigmoid activation function; P j It is the segmentation prediction map output by the j-th prediction module; Up indicates upsampling; the operator ⊕ indicates element-wise addition.

2. The medical image recognition method based on artificial intelligence according to claim 1, characterized in that, Before inputting the medical images into the medical image segmentation model, the medical images are preprocessed, and the processed images are then input into the model; the preprocessing includes: Medical images are processed using contrast-limited adaptive histogram equalization to enhance low-contrast lesions and obtain the first target image. The first target image is scaled up to fit the model input size to obtain the second target image; The pixel values ​​of the second target image are normalized to obtain the final image data input to the model.

3. The medical image recognition method based on artificial intelligence according to claim 1, characterized in that, The global feature extraction branch is divided into four stages, each stage including a global feature extraction module GFEM; each global feature extraction module contains multiple SwinTransformer blocks, with the number of blocks being 2, 2, 6, and 2 respectively.

4. The medical image recognition method based on artificial intelligence according to claim 1, characterized in that, The medical image segmentation model uses a loss function during training, including binary cross-entropy loss L. BCE And Dice lost L Dice Specifically, the total loss function L total for: ; Among them, P j It is the segmentation prediction map output by the j-th prediction module, GT j λ1 and λ2 are the labeled true values ​​for the corresponding scale; λ1 and λ2 are the weighting coefficients.

Citation Information

Patent Citations

  • Collaborative optimization method for image fusion and semantic segmentation

    CN120182786A

  • Transform-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement

    CN120318256A