Stroke non-enhanced ct lesion combined segmentation model construction method and segmentation method

CN122820748APending Publication Date: 2026-09-25HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611293737.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

由于半暗带在非增强CT上通常不具备显性的生理对比度,现有的分割框架往往难以在无CT或MR灌注影像辅助的情况下准确识别半暗带区域

Benefits of technology

(1)本发明在构建缺血性脑卒中非增强CT分割模型时,以非增强CT图像为输入,针对急性缺血性卒中梗死核心与缺血区域难以同时准确分割且缺乏生理约束的问题,构建一种融合基础模型特征表示、区域结构约束以及生理感知表征学习的协同优化方法。利用视觉基础模型协同临床错配比例约束与扩散模型,进行非增强CT图像中梗死核心与缺血区域联合分割的算法模型。本发明将基于图像强度的分割问题扩展为融合结构约束与生理信息的联合建模问题,使得在构建模型过程中利用更多的有效信息,从而准确地分割出缺血性脑卒中非增强CT梗死核心区域和整体缺血区域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820748A_ABST
    Figure CN122820748A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of medical image processing, and relates to a cerebral apoplexy non-enhanced CT lesion joint segmentation model construction method and a segmentation method, which comprises the following steps: constructing a training set, inputting non-enhanced CT images in the training set into a visual basic model for feature extraction to obtain multi-scale features of the non-enhanced CT images; inputting the multi-scale features into a segmentation decoder, wherein the segmentation decoder comprises a branch structure for outputting an infarction core region and an overall ischemic region, and the segmentation decoder is used to obtain a predicted infarction core region and a predicted overall ischemic region of the non-enhanced CT images; and performing perfusion physiological constraints and region-level structure constraints on the predicted infarction core region and the predicted overall ischemic region to complete model construction. The model constructed by the application realizes high-consistency and high-robustness segmentation of the infarction core and the ischemic region under the condition of using only non-enhanced CT data during reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing technology, and more specifically, relates to a method for constructing a combined segmentation model and a segmentation method for non-enhanced CT lesions of stroke. Background Technology

[0002] Computed tomography (CT) is a core technology for the clinical diagnosis of acute ischemic stroke. As the most readily available and time-efficient imaging modality in emergency settings, non-contrast CT is crucial for assessing tissue damage and making decisions regarding reperfusion therapy. Clinically, ischemic injury is typically divided into the infarct core (irreversibly damaged tissue) and the ischemic penumbra (salvageable tissue). The combined segmentation of the infarct core and penumbra is of extremely high clinical value in identifying patients who can benefit from the injury and determining indications for thrombolysis or thrombectomy.

[0003] Deep learning-based automated segmentation algorithms are currently the mainstream approach for quantifying lesions in CT images. These methods utilize convolutional neural networks or visual Transformers to learn the visual features of lesions, establishing a connection between histological manifestations, such as weak low-density regions on non-contrast CT, and anatomical spatial distribution. They minimize the discrepancy between the predicted mask and expert annotations by optimizing model parameters. Compared to traditional manual feature extraction methods, deep learning-based methods can capture deeper and more complex image representations, demonstrating better performance in infarct core segmentation tasks.

[0004] While existing deep learning segmentation algorithms can provide lesion identification results with a certain degree of accuracy, they still face significant challenges in practical applications on non-contrast CT images. Because the penumbra typically lacks prominent physiological contrast on non-contrast CT, existing segmentation frameworks often struggle to accurately identify penumbra regions without the aid of CT or MR perfusion imaging. This results in most current methods focusing only on the infarct core, neglecting the assessment of the penumbra, which is more clinically relevant.

[0005] Meanwhile, from a medical perspective, joint segmentation is a complex problem. Existing algorithms often ignore the inherent anatomical constraints between the infarct core and the penumbra, resulting in numerical improvements in the segmented lesion regions, but often exhibiting overlaps or breaks in spatial distribution that do not conform to physiological logic. Furthermore, models that rely solely on pre-training on natural images lack guidance from medical knowledge, and the feature space they learn lacks consistency with clinically meaningful pathological evolution. This not only severely limits the robustness of the model on different clinical datasets but also hinders the further popularization and application of non-contrast CT (NCCT)-assisted assessment technology in stroke emergency green channels.

[0006] Therefore, it is urgent to accurately and jointly segment the infarct core and ischemic area of ​​stroke patients based on non-contrast CT scans of ischemic stroke. Summary of the Invention

[0007] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a method for constructing and segmenting a combined segmentation model of ischemic stroke lesions on non-contrast CT based on clinical structural constraints and perfusion information guidance. The purpose is to accurately and jointly segment the infarct core and ischemic area of ​​stroke patients based on non-contrast CT of ischemic stroke.

[0008] According to the present invention, a method for constructing a combined segmentation model of non-contrast CT lesions in ischemic stroke is provided, comprising the following steps: (1) Constructing a training set; in the training set, each training data point is a non-contrast CT image of ischemic stroke with labeled infarct core region and overall ischemic region, and the time when the tissue residual function of the perfusion CT corresponding to the non-contrast CT image reaches its maximum value. T max Parameter diagram; (2) Input the non-enhanced CT images in the training set into the visual basic model for feature extraction to obtain the multi-scale features of the non-enhanced CT images; (3) The multi-scale features are input into the segmentation decoder, which includes a branch structure that outputs the infarct core region and the overall ischemic region, to obtain the predicted infarct core region of the non-enhanced CT image. and predicting the overall ischemic area ; (4) For the predicted infarct core region and predicting the overall ischemic area Weights are assigned to each image to obtain an attention heatmap. This attention heatmap is then used as a condition for a conditional diffusion model, where the input to the conditional diffusion model is the non-enhanced CT image, and the output is the prediction corresponding to the non-enhanced CT image. The parameter map and supervision information are from the training set. T max The parameter map enables the prediction of the infarct core region in the non-enhanced CT image. and predicting the overall ischemic area Physiological constraints on perfusion of the segmentation results; (5) Label the whole ischemic area in the non-enhanced CT images of the labeled infarct core region and the overall ischemic region. Compared to the marked infarct core area The ratio is defined as the standard mismatch ratio; using the multi-scale features, a multilayer perceptron is used as a classification decoder to obtain the predicted mismatch ratio category probability; the standard mismatch ratio is compared with a threshold to obtain the standard mismatch ratio binary classification label, and the binary cross-entropy loss is used to constrain the predicted mismatch ratio category probability to be consistent with the standard mismatch ratio binary classification label, thereby achieving regional-level structural constraints; (6) Optimize the parameters of the visual base model, segmentation decoder, classification decoder and diffusion model by means of a loss function, wherein the loss function is a combined loss function, including a segmentation loss function, a classification loss function and a generation loss function, until the visual base model, segmentation decoder, classification decoder and diffusion model converge.

[0009] Preferably, the combined loss function Represented as:

[0010] in: For the segmentation loss function; The classification loss function; To generate the loss function; , and These are the weighting coefficients; The segmentation loss function includes the DICE loss function and the CE loss function. Represented as: ; in: S c This indicates the actual annotation, including the actual infarct core region. S inf and the actual overall ischemic area S isch , This indicates the prediction results, including the predicted infarct core area. and predicting the overall ischemic area , and These are Dice loss and cross-entropy loss, respectively. The classification loss function is the BCE loss function. Represented as:

[0011] Where y is the standard mismatch ratio. The predicted mismatch ratio; The generation loss function is the L2 loss function. Represented as:

[0012] in: This is a time parameter graph showing the time when the predicted tissue residual function of perfusion CT corresponding to the non-contrast CT image reaches its maximum value. T max The time parameter graph shows the actual tissue residual function of the perfusion CT corresponding to the non-enhanced CT image reaching its maximum value.

[0013] Preferably, The value is 1. and As the training process changes, The initial value was set to 0.5, and then gradually reduced to 0. The initial value is set to 0, and then gradually increased to 0.5.

[0014] Preferably, the infarct core region and the overall ischemic region are marked by using a thresholding method on the CT perfusion image after registration of the non-enhanced CT image.

[0015] Preferably, the threshold used in the threshold method is: the relative blood flow rCBF in the infarct core region is less than or equal to 0.3, and the overall ischemic region... T max Greater than 6 seconds.

[0016] Preferably, in the obtained attention heatmap, the infarct core region is predicted. The weight assigned is 1, which is used to predict the overall ischemic area. The assigned weight is 0.5.

[0017] Preferably, before inputting the non-enhanced CT image into the visual base model for feature extraction, the non-enhanced CT image is further preprocessed, including outlier removal, skull removal, and intensity normalization.

[0018] According to another aspect of the present invention, a method for combined segmentation of ischemic stroke lesions on non-contrast CT is provided, wherein the non-contrast CT image of ischemic stroke to be segmented is input into a combined segmentation model of ischemic stroke lesions on non-contrast CT to obtain segmentation results of the predicted infarct core region and the predicted overall ischemic region. The combined segmentation model of ischemic stroke lesions on non-contrast CT is established by the construction method of the combined segmentation model of ischemic stroke lesions on non-contrast CT.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to perform the method for constructing a combined segmentation model of ischemic stroke non-contrast CT lesions, and / or to perform the method for combined segmentation of ischemic stroke non-contrast CT lesions.

[0020] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages: (1) In constructing a non-enhanced CT segmentation model for ischemic stroke, this invention uses non-enhanced CT images as input. Addressing the difficulty in simultaneously and accurately segmenting the infarct core and ischemic region of acute ischemic stroke, and the lack of physiological constraints, a collaborative optimization method integrating basic model feature representation, regional structural constraints, and physiological perception representation learning is constructed. An algorithm model for joint segmentation of the infarct core and ischemic region in non-enhanced CT images is developed using a visual basic model in conjunction with clinical mismatch ratio constraints and a diffusion model. This invention extends the image intensity-based segmentation problem into a joint modeling problem integrating structural constraints and physiological information, enabling the use of more effective information during model construction, thereby accurately segmenting the infarct core region and the overall ischemic region of ischemic stroke on non-enhanced CT.

[0021] (2) This invention overcomes the shortcomings of existing deep learning methods, such as lack of anatomical constraints and inconsistency between feature space and clinical pathological evolution. By constructing a multi-task collaborative optimization network, it organically combines segmentation, classification (mismatch ratio prediction), and generation tasks (pseudo-parameter graph generation) to achieve pixel-level, region-level, and physiological-level correlation modeling of the distribution of real brain tissue in acute ischemic stroke in NCCT, realizing multi-task anatomical and tissue modeling constraints. Furthermore, it provides an innovative cross-modal physiological prior crossing, i.e., from avascular information to vascular information, which can derive and generate pseudo-parameter graphs with rich vascular perfusion information from NCCT modalities that do not contain explicit vascular hydrodynamic information. The parameter map allows for the introduction of physiological-level prior guidance for segmentation during the inference phase without the need for any perfusion imaging assistance.

[0022] (3) This invention achieves highly consistent and robust segmentation of the infarct core and ischemic region using only non-contrast CT data of ischemic stroke. Non-contrast CT images of ischemic stroke are easily acquired, and this segmentation method can generate corresponding perfusion images simultaneously with segmentation. T max Parameter diagram. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the model method proposed in this invention.

[0024] Figure 2 This is a flowchart of the CT-based ischemic stroke lesion joint segmentation algorithm proposed in this invention, where (a) is the training phase of the algorithm and (b) is the inference phase of the algorithm.

[0025] Figure 3 This is a comparison chart of the segmentation results for two examples. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0027] Example 1 The specific method for constructing a non-contrast CT lesion combined segmentation model for ischemic stroke is shown in Table 1 below: Table 1

[0028] Specifically, it includes the following steps: (1) Constructing the training set. Each training data set used for complete joint training includes a non-enhanced CT image of an ischemic stroke patient, the corresponding infarct core region labeling, the overall ischemic region labeling, a brain tissue mask, and a Tmax parameter map registered to the non-enhanced CT space. The non-enhanced CT image can be a two-dimensional slice image or a two-dimensional image obtained layer by layer from volume data; the infarct core region labeling is used to represent irreversibly damaged tissue, the overall ischemic region labeling is used to represent the total area of ​​the infarct core and the surrounding low-perfusion ischemic tissue, and the ischemic penumbra region can be obtained by subtracting the infarct core region from the overall ischemic region. In addition to the complete training data mentioned above, auxiliary samples with only segmentation labels but no Tmax parameter map can be added. These samples only participate in the training of the segmentation branch and the region-level mismatch ratio constraint branch, and do not participate in the generation loss calculation of the conditional diffusion branch.

[0029] The training set is divided into training / validation subsets and test subsets based on cases or slices. Internal data includes NCCT and CT perfusion data from patients with acute ischemic stroke, with the training, validation, and test sets divided by cases (7:1:2). For external validation, publicly available NCCT datasets containing only expert annotations can be used to test the model's generalization ability to different scanning protocols, patient populations, and image qualities. The number of cases, division ratios, and data sources mentioned above are merely examples and can be adjusted based on the scale of clinical data in actual implementation.

[0030] (2) Preprocessing of training data. For each non-contrast CT image, outlier removal, skull removal, and intensity normalization are performed first. In one specific implementation, the CT values ​​of the non-contrast CT images are truncated to the range of 0 HU to 100 HU before normalization to reduce the intensity distribution differences caused by different scanning devices and scanning parameters. For the Tmax parameter map used in the training phase, its effective value range is truncated to 0 s to 30 s, and the Tmax parameter map is registered to the non-contrast CT image space through affine registration or other medical image registration methods.

[0031] The non-enhanced CT images, infarct core region annotations, overall ischemic region annotations, brain tissue masks, and Tmax parameter maps were all uniformly adjusted to 224×224 pixels. Specifically, bilinear interpolation was used for the non-enhanced CT images and Tmax parameter maps, while nearest-neighbor interpolation was used for the infarct core region annotations, overall ischemic region annotations, and brain tissue masks to avoid generating unlabeled values ​​for the mask labels during scaling.

[0032] If no brain tissue mask is provided in the samples, a mask with the same size as the infarct core region and all pixels having a value of 1 is used as the brain tissue mask. For data used for full joint training, the registered Tmax parameter map is used to supervise the conditional diffusion branch; for additional samples without Tmax auxiliary data, only pixel-level segmentation loss and region-level mismatch ratio loss are calculated. This maintains consistency with the definition of full joint training data while allowing the use of additional labeled data without changing the core technical solution.

[0033] (3) Obtain supervisory labels based on the registered perfusion parameter map. Initial labels are generated on the CT perfusion parameter map using commonly used clinical thresholds: regions with relative cerebral blood flow (rCBF) less than or equal to 0.3 are defined as infarct core regions, and regions with Tmax greater than 6 s are defined as overall ischemic regions. The automatically obtained regions are then verified by a physician. The corresponding ischemic penumbra region is obtained by subtracting the infarct core region from the overall ischemic region. It should be noted that the above-mentioned rCBF threshold, Tmax threshold, and physician verification method are preferred embodiments and should not be construed as the sole limitation on the scope of protection of this invention.

[0034] (4) Constructing the visual base model encoder. The preprocessed non-enhanced CT images are input into the visual base model encoder for feature extraction. One approach used in the visual base model encoder is to pre-train the DINOv3-ViT-B model with an input size of 224×224 and a patch size of 16. Since non-enhanced CT images are usually single-channel grayscale images, the single-channel images are copied into three-channel images before being input into the encoder; if the number of input channels is greater than three, the first three channels are used. Subsequently, per-image minimum-maximum normalization is performed on each input image, and ImageNet mean and variance are used for standardization to match the input distribution of the pre-trained visual base model.

[0035] The visual base model encoder outputs a multi-layered hidden state containing category tokens, register tokens, and patch tokens. Patch tokens are extracted and rearranged into a 2D feature map according to the image grid. For a 224×224 input and a 16×16 patch, the spatial size of the patch feature map is 14×14. Multi-scale features are extracted from the hidden states of layers 1, 3, 6, and 9, respectively, forming shallow texture features, mid-level structural features, and deep semantic features. These layer numbers can be adjusted according to the depth of the visual base model used.

[0036] To improve the adaptability of the visual base model to medical images, adapter modules are inserted into layers 3, 6, 9, and 11 of the visual base model. These adapter modules can be bottleneck structures, comprising a lower projection layer, a GELU activation layer, and an upper projection layer. The lower projection layer compresses the feature dimension by a factor of 16, and the upper projection layer restores the original feature dimension, adding it to the input feature residual. Alternatively, the adapter module can employ a low-rank adaptation structure, achieving efficient parameter fine-tuning through low-rank matrix mapping and scaling factors. During training, the visual base model backbone can be frozen, training only the adapter and decoder; alternatively, the visual base model backbone can be fine-tuned with a small learning rate. This feature extraction and parameter adaptation process corresponds to... Figure 1 The encoder of the basic vision model and its parameter update path.

[0037] (5) Constructing a segmentation decoder. The multi-scale features output by the visual base model encoder are input into the ResUNet-style segmentation decoder. Specifically, the features of layers 1, 3, 6 and 9 are first projected to 64 channels, 128 channels, 256 channels and 256 channels respectively through 1×1 convolution, as skip connection features; then the patch feature map of the last layer of the encoder is input into the residual convolution bottleneck module to convert its channel number to 512 channels.

[0038] The segmentation decoder comprises four levels of upsampling blocks. The first level upsampling block upsamples the 512-channel bottleneck features and concatenates them with the 9th layer projection features, then outputs 256-channel features via a residual convolution block. The second level upsampling block fuses the 256-channel features with the 6th layer projection features to output 128-channel features. The third level upsampling block fuses the 128-channel features with the 3rd layer projection features to output 64-channel features. The fourth level upsampling block fuses the 64-channel features with the 1st layer projection features to output 64-channel features. Finally, a 1×1 convolution outputs a two-channel logit map, with the first channel corresponding to the infarct core region and the second channel corresponding to the overall ischemic region.

[0039] Each upsampling block includes transposed convolutional upsampling, skip connection concatenation, and a residual convolutional block. The residual convolutional block includes two 3×3 convolutions, batch normalization, and ReLU activation, and constructs a shortcut branch using 1×1 convolutions when the number of input and output channels differs. If the spatial dimensions of the skip connection features and the upsampling features are inconsistent, the skip connection features are first adjusted to the same spatial dimensions using bilinear interpolation before concatenation.

[0040] The segmentation decoder also includes a deep supervision branch. In one specific implementation, a 1×1 convolutional auxiliary output head is set after the first and second level upsampling blocks, and the auxiliary output is interpolated to restore the original input size. During training, the auxiliary output and the main output jointly receive supervision from labels of the infarct core region and the overall ischemic region, thereby enabling the intermediate layers of the decoder to learn recognizable lesion structural features in advance and reducing detail loss during the upsampling process. The above-mentioned multi-scale feature fusion, dual-channel output, and segmentation supervision process correspond to... Figure 1 The joint segmentation decoder and segmentation output section.

[0041] (6) Obtain pixel-level segmentation results. The two channel logit maps output by the segmentation decoder are denoted as the infarct core prediction logit and the overall ischemic region prediction logit, respectively. Sigmoid operations are performed on both to obtain the infarct core probability map and the overall ischemic region probability map. Thresholding is applied to the probability maps to obtain the predicted infarct core region and the predicted overall ischemic region. Further, subtracting the predicted infarct core region from the predicted overall ischemic region yields the predicted ischemic penumbra region.

[0042] (7) Construct regional mismatch ratio constraint branches. Calculate the standard mismatch ratio based on the actual overall ischemic region and the actual infarct core region.

[0043] This represents the actual total area of ​​ischemic region. This represents the actual area of ​​the infarct core region. To prevent smoothing terms with a denominator of zero, further... With preset threshold The comparison yields the standard mismatch ratio label for binary classification supervision. :when hour, ;when hour, In one specific implementation, Therefore, the "standard mismatch ratio" "" indicates a continuous area ratio, while "standard mismatch ratio label" indicates a continuous area ratio. "" indicates a binary label obtained by discretizing the continuous ratio, and the two should not be used interchangeably.

[0044] The multi-scale features output by the visual baseline model are aggregated, for example, by performing global average pooling and concatenating them into a one-dimensional feature vector, which is then input into a multilayer perceptron classifier. The multilayer perceptron classifier sequentially includes a fully connected layer, a ReLU activation layer, a Dropout layer, and a fully connected layer and a Sigmoid activation layer with an output dimension of 1, outputting the predicted mismatch ratio class probability p. This classifier is used to determine whether the area relationship between the overall ischemic region and the infarct core region exceeds a preset clinical threshold.

[0045] During training, the region-level mismatch ratio classification loss uses binary cross-entropy loss:

[0046] Optionally, to further constrain the area relationship between the segmented outputs, the continuous prediction ratio can also be calculated based on the prediction probability map.

[0047] And on Ratio to true continuous Set a consistency loss. It should be noted that the binary cross-entropy loss directly supervises the classification labels. With predicted probability Instead of directly using continuous ratios As a label for binary cross-entropy. The above-mentioned regional structural constraint process corresponds to Figure 1 Mismatch ratio classification MLP decoder and region constraint path.

[0048] (8) Constructing a lesion attention heatmap. Sigmoid operations are performed on the predicted logit of the infarct core and the predicted logit of the overall ischemic region to obtain the infarct core probability map. Probability map of overall ischemic area and in accordance with

[0049] By performing weighted combination, an attention heatmap of the lesions was obtained. If a brain tissue mask exists Then, the brain tissue mask is interpolated to the size of the attention heatmap and then compared with... Pixel-by-pixel multiplication, i.e.

[0050] It helps to focus attention on specific areas of brain tissue.

[0051] In this model, the weight of the infarct core region is set to 1, and the weight of the overall ischemic region is set to 0.5, consistent with the preferred weights described in the claims. This attention heatmap allows subsequent conditional diffusion branches to focus on the spatial locations of the infarct core, the overall ischemic region, and the ischemic penumbra corresponding to their difference. This step corresponds to... Figure 1 The attention map is generated from the segmented output and the connection path of the conditional diffusion branch is input.

[0052] (9) Constructing the conditional diffusion branch. The conditional diffusion branch includes a diffusion UNet and a Gaussian diffusion process, using non-enhanced CT images and lesion attention heatmaps as conditional information, and the real Tmax parameter map as the training supervision target. Within the internal process of diffusion training, the real Tmax parameter map is denoted as... Random sampling time step and Gaussian noise and in accordance with

[0053] Obtain the noise-added state ;Diffusion UNet reception Time step Non-contrast CT features and lesion attention heatmap Output noise prediction Or the initial image prediction During the inference phase, random noise is gradually denoised to generate a pseudo Tmax parameter map corresponding to the input non-enhanced CT.

[0054] The downsampling path of the diffusion UNet includes multi-level residual convolutional blocks, a linear attention module, and a downsampling module; the upsampling path includes skip connection concatenation, residual convolutional blocks, a linear attention module, and an upsampling module; the intermediate layer includes residual convolutional blocks, an attention module, and a segmentation-guided attention module. Each level of feature can receive a lesion attention heatmap through the segmentation-guided attention module. The segmentation-guided attention module first encodes the attention heatmap into segmentation conditional features with the same number of channels as the current feature, then generates spatial attention and channel attention, and modulates the current diffusion feature with learnable coefficients, so that the diffusion branch focuses on learning the perfusion spatial distribution of the lesion-related region during the denoising process.

[0055] The diffusion UNet also includes a cross-attention module and a FiLM modulation layer for fusing non-contrast CT image features, lesion attention heatmaps, and other optional clinical condition features; when no other clinical conditions are used, the diffusion branch relies on non-contrast CT image features and lesion attention heatmaps for conditional modulation.

[0056] The Gaussian diffusion process can employ 1000 training time steps and 100 sampling time steps, and the noise scheduling can use cosine scheduling. In one specific implementation, the training objective is to predict noise, and the basic diffusion loss is...

[0057] It can also be stacked Loss or boundary-aware loss based on Sobel gradients can be used to enhance the spatial variation trend of pseudo-perfusion representation.

[0058] When training the conditional diffusion branch, non-enhanced CT images and lesion attention heatmaps are used as conditional inputs, and noise prediction is performed under the supervision of the real Tmax parameter map. If a brain tissue mask exists, it is applied to the real Tmax parameter map, the noise-added state, and the loss term; region weighting coefficients can also be set based on the attention heatmap. This gives higher weight to the generation error of lesion-related areas. This training process corresponds to... Figure 1 Tmax supervision, conditional diffusion model and physiological constraint backpropagation path.

[0059] It should be noted that the core function of the conditional diffusion branch is to utilize the Tmax supervised learning during the training phase to learn physiological perceptual representations consistent with the spatial distribution of perfusion anomalies, and to constrain the visual base model encoder and segmentation decoder through joint backpropagation; the generated pseudo Tmax parameter map can be used as an auxiliary output, but should not be presented as a substitute for real CT perfusion examination, nor is it required to accurately reconstruct the real Tmax parameter map pixel by pixel.

[0060] (10) Construct a joint loss function and train the model. Pixel-level segmentation loss includes Dice loss for the infarct core region and the overall ischemic region, and binary cross-entropy loss. Let

[0061] Let these represent the infarct core and the overall ischemic area, respectively.

[0062] If deep supervision output is set, then the loss of each auxiliary output is multiplied by its corresponding weight and added. .

[0063] The regional mismatch ratio loss is denoted as At least including the output of a multilayer perceptron. Mismatch with standard label The binary cross-entropy loss between them; the physiological level generation loss is denoted as The final joint loss function is expressed as:

[0064] These are the weighting coefficients for segmentation loss, region-level mismatch ratio loss, and diffusion generation loss, respectively.

[0065] In a specific embodiment consistent with the claims, ; and It changes dynamically during the training process, especially in the early stages of training. , Subsequently, gradually Reduce to 0, Increase to 0.5. This strategy allows the model to first establish a stable lesion representation under segmentation and regional structural constraints, and then gradually enhance perfusion physiological constraints. Weight changes can be performed using linear, piecewise linear, or other continuous scheduling methods.

[0066] Optionally, a progressive training strategy is employed to control the time step range of the conditional diffusion branch. In the first 30% of training rounds, the diffusion time step sampling range is the latter half of the time steps, allowing the diffusion branch to preferentially learn the global perfusion anomaly distribution under stronger noise conditions; in the 30% to 70% of training rounds, the time step sampling range is expanded to the latter three-quarters of the time steps; in the later stages of training, the time step sampling range is expanded to all time steps, allowing the diffusion branch to simultaneously learn coarse-grained and fine-grained denoising representations.

[0067] The AdamW optimizer is used to update the model parameters. The learning rate for the backbone parameters of the visual base model is set to 1×10⁻⁶. -5 The learning rates for the adapter, segment decoder, conditional diffusion branch, and mismatch ratio classifier were set to 1×10. -4 The learning rate for the attention module can be set to 1.5 times the base learning rate, and the learning rate for the mismatch classifier can be set to 1.2 times the base learning rate. A linear warm-up of 10 epochs is set at the beginning of training, and the learning rate gradually decays according to a preset decay coefficient after the warm-up. After each backpropagation, gradient pruning with a maximum norm of 1.0 is performed on the gradients of the visual base model, attention module, conditional diffusion model, and mismatch classifier to improve training stability.

[0068] During training, each batch of non-enhanced CT images, infarct core labels, overall ischemic region labels, brain tissue masks, and Tmax parameter maps are input into the model. First, the visual baseline model encoder and segmentation decoder obtain segmentation predictions and multi-scale features for both channels. Then, the mismatch classifier outputs the mismatch category probabilities. Next, lesion attention heatmaps are generated based on the segmentation predictions, and the conditional diffusion branch is trained. The model performs backpropagation based on the joint loss function and updates parameters until the training loss converges or the validation set performance meets the preset requirements, thus completing the establishment of the joint segmentation model for ischemic stroke lesions on non-enhanced CT. For additional samples without Tmax auxiliary data, no calculation is performed. This joint optimization process corresponds to Figure 2 The training phase is shown in (a) of the diagram.

[0069] (11) Model Inference. After model training is completed, for the non-enhanced CT images of ischemic stroke to be segmented, only the non-enhanced CT images need to be input; real CT perfusion images or MR perfusion images do not need to be input. After performing the same preprocessing as in the training phase on the images to be segmented, they are input into the visual base model encoder and segmentation decoder to obtain the infarct core probability map. Probability map of overall ischemic area After thresholding, the predicted infarct core region and the predicted overall ischemic region are obtained, and then the aggregation difference is used to obtain the predicted infarct core region and the predicted overall ischemic region.

[0070] The predicted ischemic penumbra region was obtained.

[0071] Simultaneously, the aggregated features output by the visual baseline model can be input into the mismatch ratio classifier to obtain mismatch ratio classification hints; furthermore, the lesion attention heatmap generated from non-enhanced CT images and segmentation results can be input into the trained conditional diffusion branch to generate a pseudo-Tmax parameter map. Both the mismatch ratio classification hints and the pseudo-Tmax parameter map are auxiliary information; the main segmentation output does not depend on real perfusion imaging. This process corresponds to... Figure 2 The reasoning stage shown in (b) is as follows.

[0072] Through the above implementation methods, this invention unifies the pixel-level segmentation task of non-enhanced CT images, the regional structural constraint task between the infarct core and the overall ischemic region, and the physiological diffusion representation learning task based on Tmax supervision into an end-to-end training framework. Figure 1As can be seen, during the training phase, NCCT, after being encoded by the visual baseline model, enters the joint segmentation decoder and mismatch classification branch respectively. The segmentation results further form a lesion attention heatmap and modulate the conditional diffusion branch. The three types of losses jointly update the relevant network parameters in reverse. During the inference phase, only NCCT input can obtain the infarct core region, the overall ischemic region, and the ischemic penumbra region, with optional output of mismatch ratio hints and pseudo-Tmax parameter maps. Compared with segmentation models that rely solely on pixel-level labels for training, this embodiment can reduce situations that do not conform to clinical norms, such as missed segmentation of the infarct core region, fragmentation of the overall ischemic region, boundary overflow, and the infarct core extending beyond the overall ischemic region, thereby improving the stability, spatial consistency, and clinical interpretability of joint segmentation.

[0073] It should be noted that the descriptions of input size, number of network layers, number of channels, learning rate, loss weights, mismatch ratio threshold, number of diffusion time steps, and perfusion threshold in this embodiment are all preferred embodiments for the purpose of illustrating the technical solution of the present invention, and should not be construed as the sole limitation of the present invention. Without departing from the spirit and principles of the present invention, those skilled in the art can make equivalent substitutions or adjustments to the above parameters, module structures, or training strategies according to the actual data scale, equipment conditions, and clinical application needs.

[0074] Example 2 The non-contrast CT images of ischemic stroke to be segmented are input into the non-contrast CT lesion joint segmentation model of ischemic stroke to obtain the segmentation results of the predicted infarct core area and the predicted overall ischemic area. The non-enhanced CT lesion joint segmentation model for ischemic stroke is the non-enhanced CT lesion joint segmentation model for ischemic stroke constructed in this invention.

[0075] Example 3 Figure 3 The results of processing non-enhanced CT images of two patients with acute ischemic stroke using the method of the present invention are shown. Figure 3 Each row corresponds to one patient. The columns, from left to right, are: the original non-enhanced CT image, the superimposed label of the infarct core and the overall ischemic region, the true Tmax perfusion parameter map, the pseudo Tmax perfusion parameter map generated using the conditional diffusion model of this invention, the segmentation result obtained by the method of this invention, and the segmentation results obtained by comparative methods such as AttnUNet, UNet++, nnUNet-2D, UNETR, SwinUNETR, and I²PC-Net. The infarct core region and the overall ischemic region are displayed using different colors superimposed, with the smaller inner region corresponding to the infarct core and the larger outer region corresponding to the overall ischemic region. The difference between the two corresponds to the ischemic penumbra region.

[0076] against Figure 3 For the two patients shown, two-dimensional non-contrast CT slices and their paired CT perfusion images were first acquired. Outlier removal, skull removal, and intensity normalization were performed on the non-contrast CT images. The CT values ​​of the non-contrast CT images were truncated to 0–100 HU, and the image size was uniformly adjusted to 224×224 pixels. For the Tmax parameter map used in the training phase, its effective value range was set to 0–30 s, and the Tmax parameter map was registered to the corresponding non-contrast CT image space using affine registration.

[0077] Training labels are generated based on the registered CT perfusion parameter maps. Regions with relative cerebral blood flow (rCBF) less than 0.3 are defined as the infarct core region, and regions with Tmax greater than 6 s are defined as the overall ischemic region. The automatically obtained regions are then verified by a physician. The corresponding ischemic penumbra region is obtained by subtracting the infarct core region from the overall ischemic region.

[0078] Preprocessed non-enhanced CT images are input into a pre-trained visual baseline model, which extracts image features at different scales. These multi-scale features are then input into a dual-head segmentation decoder, which outputs an infarct core probability map and a global ischemic region probability map, respectively. After thresholding the two probability maps, the predicted infarct core region and the predicted global ischemic region are obtained.

[0079] During training, Dice loss and cross-entropy loss were used to provide pixel-level supervision for the two segmentation results. Simultaneously, a standard mismatch ratio was calculated based on the area ratio between the true overall ischemic region and the true infarct core region. The multi-scale features output from the visual baseline model were then input into a multilayer perceptron to predict the patient's mismatch ratio category. Binary cross-entropy loss was used to constrain the predicted mismatch ratio category probability to be consistent with the binary classification label obtained from the standard mismatch ratio, enabling the model to learn the regional relationship between the overall ischemic region and the infarct core region.

[0080] Furthermore, the predicted infarct core region is assigned a weight of 1, and the predicted overall ischemic region is assigned a weight of 0.5. These two weighted combinations yield a lesion attention heatmap. This attention heatmap is used as conditional information and input into the conditional diffusion model along with the original non-enhanced CT image. Under the supervision of the true Tmax parameter map, the conditional diffusion model generates a pseudo Tmax parameter map. Through the L2 loss between the generated result and the true Tmax parameter map, the perfusion spatial distribution information is back-propagated to the segmentation decoder and the visual baseline model, thus enabling the features extracted by the model to possess both image representation capabilities and perfusion physiological perception capabilities.

[0081] In the initial training phase, segmentation loss is the primary optimization objective, enabling the model to initially achieve stable lesion localization capabilities. With each training iteration, the effects of mismatch ratio structural constraints and Tmax perfusion information constraints are gradually enhanced. Joint backpropagation is performed using pixel-level segmentation loss, region-level classification loss, and physiological-level generation loss until the model converges.

[0082] like Figure 3 As shown in the first row, the lesion in the first patient appears as a relatively localized and low-contrast abnormal area in the non-enhanced CT image. Conventional segmentation methods are prone to omissions of the infarct core, incomplete boundaries of the overall ischemic area, or discrepancies between the predicted area and the actual lesion location. The pseudo-Tmax parameter map generated by the method of this invention exhibits a spatial distribution trend similar to the true Tmax parameter map in the lesion location and its surrounding area; the overall ischemic area obtained based on this physiological information constraint can relatively completely cover the infarct core, and a reasonable inclusion relationship is maintained between the infarct core and the surrounding ischemic tissue.

[0083] like Figure 3 As shown in the second row, the lesion morphology of the second patient is more irregular. Some comparative methods exhibit fragmentation of the predicted region, boundary overflow, underestimation of the lesion range, or unreasonable spatial relationship between the infarct core and the overall ischemic region. The method of this invention implements regional constraints on the two segmentation targets through a mismatch ratio auxiliary task, and guides the model to focus on regions with abnormal perfusion characteristics through a pseudo-Tmax generation task. Therefore, the obtained overall ischemic region has high consistency with the true label in terms of spatial range and morphology. The predicted infarct core is located inside the overall ischemic region, and the difference between the two can form a continuous ischemic penumbra region that conforms to the clinicopathological relationship.

[0084] In the aforementioned practical application phase, it is no longer necessary to input actual CT perfusion images. For the patient to be examined, only the preprocessed non-enhanced CT image is input into the trained model, which can output the infarct core region, the overall ischemic region, and the ischemic penumbra region obtained by the difference between the two; at the same time, it can also output the mismatch ratio classification results and pseudo-Tmax parameter map as reference information to assist clinical judgment.

[0085] Depend on Figure 3 As the results show, compared with the comparison method that only uses pixel-level segmentation supervision, the present invention, by jointly introducing the structural relationship between the infarct core and the overall ischemic area as well as Tmax perfusion physiological information, can reduce situations that do not conform to clinical rules, such as missed segmentation of lesion area, boundary overflow, regional fragmentation, and core area exceeding the overall ischemic area. This improves the spatial consistency, stability, and clinical rationality of the joint segmentation results of the infarct core and the overall ischemic area under non-enhanced CT conditions.

[0086] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a combined segmentation model of non-contrast CT lesions in ischemic stroke, characterized in that, Includes the following steps: (1) Constructing a training set; in the training set, each training data point is a non-contrast CT image of ischemic stroke with labeled infarct core region and overall ischemic region, and the time when the tissue residual function of the perfusion CT corresponding to the non-contrast CT image reaches its maximum value. T max Parameter diagram; (2) Input the non-enhanced CT images in the training set into the visual basic model for feature extraction to obtain the multi-scale features of the non-enhanced CT images; (3) The multi-scale features are input into the segmentation decoder, which includes a branch structure that outputs the infarct core region and the overall ischemic region, to obtain the predicted infarct core region of the non-enhanced CT image. and predicting the overall ischemic area ; (4) For the predicted infarct core region and predicting the overall ischemic area Weights are assigned to each, resulting in an attention heatmap; The attention heatmap is used as a condition for a conditional diffusion model, where the input to the conditional diffusion model is the non-enhanced CT image, and the output is the prediction corresponding to the non-enhanced CT image. The parameter map and supervision information are from the training set. T max The parameter map enables the prediction of the infarct core region in the non-enhanced CT image. and predicting the overall ischemic area Physiological constraints on perfusion of the segmentation results; (5) Label the whole ischemic area in the non-enhanced CT images of the labeled infarct core region and the overall ischemic region. Compared to the marked infarct core area The ratio is defined as the standard mismatch ratio; using the multi-scale features, a multilayer perceptron is used as a classification decoder to obtain the predicted mismatch ratio category probability; the standard mismatch ratio is compared with a threshold to obtain the standard mismatch ratio binary classification label, and the binary cross-entropy loss is used to constrain the predicted mismatch ratio category probability to be consistent with the standard mismatch ratio binary classification label, thereby achieving regional-level structural constraints; (6) Optimize the parameters of the visual base model, segmentation decoder, classification decoder and diffusion model by means of a loss function, wherein the loss function is a combined loss function, including a segmentation loss function, a classification loss function and a generation loss function, until the visual base model, segmentation decoder, classification decoder and diffusion model converge.

2. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 1, characterized in that, The combined loss function Represented as: in: For the segmentation loss function; The classification loss function; To generate the loss function; , and These are the weighting coefficients; The segmentation loss function includes the DICE loss function and the CE loss function. Represented as: ; in: S c This indicates the actual annotation, including the actual infarct core region. S inf and the actual overall ischemic area S isch , This indicates the prediction results, including the predicted infarct core area. and predicting the overall ischemic area , and These are Dice loss and cross-entropy loss, respectively. The classification loss function is the BCE loss function. Represented as: Where y is the standard mismatch ratio. The predicted mismatch ratio; The generation loss function is the L2 loss function. Represented as: in: This is a time parameter graph showing the time when the predicted tissue residual function of perfusion CT corresponding to the non-contrast CT image reaches its maximum value. T max The time parameter graph shows the actual tissue residual function of the perfusion CT corresponding to the non-enhanced CT image reaching its maximum value.

3. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 2, characterized in that, The value is 1. and As the training process changes, Initially set to 0.5, gradually reduced to 0. The initial value is set to 0, and then gradually increased to 0.

5.

4. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 1, characterized in that, The infarct core region and the overall ischemic region were marked by using a thresholding method on the CT perfusion image after registration of the non-enhanced CT image.

5. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 4, characterized in that, The threshold used in the threshold method is: the relative blood flow rCBF in the infarct core area is less than or equal to 0.3, and the overall ischemic area... T max Greater than 6 seconds.

6. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 1, characterized in that, In the attention heatmap obtained, the infarct core region is predicted. The weight assigned is 1, which is used to predict the overall ischemic area. The assigned weight is 0.

5.

7. The method for constructing a combined segmentation model of ischemic stroke lesions on non-enhanced CT as described in claim 1, characterized in that, Before inputting the non-enhanced CT image into the visual base model for feature extraction, the non-enhanced CT image is preprocessed, including outlier removal, skull removal, and intensity normalization.

8. A method for combined segmentation of non-contrast CT lesions in ischemic stroke, characterized in that, The non-contrast CT images of ischemic stroke to be segmented are input into the non-contrast CT lesion joint segmentation model of ischemic stroke to obtain the segmentation results of the predicted infarct core area and the predicted overall ischemic area. The combined segmentation model of ischemic stroke lesions on non-contrast CT is established by the construction method of the combined segmentation model of ischemic stroke lesions on non-contrast CT as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer program includes a stored computer program; when executed by a processor, the computer program controls the device containing the computer-readable storage medium to perform the method for constructing a combined segmentation model of ischemic stroke non-contrast CT lesions as described in any one of claims 1-7, and / or to perform the method for combined segmentation of ischemic stroke non-contrast CT lesions as described in claim 8.