TLS structure sketching system and method based on artificial intelligence

By using an AI-based TLS structure delineation system that combines morphological and spatial distribution characteristics, the problem of difficulty in identifying the differentiation stage of TLSs in traditional HE staining interpretation has been solved, enabling precise analysis of TLSs and accurate prediction of the efficacy of immunotherapy.

CN120852309AInactive Publication Date: 2025-10-28FUJIAN PROVINCIAL HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510915756.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional HE staining interpretation suffers from difficulty in distinguishing the differentiation stages of TLSs and insufficient handling of morphological heterogeneity. Existing AI technologies cannot accurately identify the differentiation stages and spatial distribution characteristics of TLSs, resulting in insufficient recognition accuracy.

Method used

An AI-based TLS structure delineation system is adopted, which achieves accurate analysis of TLSs by combining multi-level feature modeling and spatial semantic understanding through image preprocessing, morphological feature extraction, clustering region identification, morphological analysis and spatial distribution localization.

Benefits of technology

It improves the accuracy of identifying the differentiation stage of TLSs, can distinguish between early, primary and secondary TLSs, enhances the accuracy of predicting the efficacy of immunotherapy, and adapts to the morphological differences of different tumor types, reducing detection time and interpretation differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852309A_ABST
    Figure CN120852309A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical image processing, and particularly discloses a TLS structure sketching system and method based on artificial intelligence, and the method comprises the steps: S1, image preprocessing: carrying out the standardization processing of an input HE staining section image, firstly separating cell nucleus and cytoplasm staining components through a color deconvolution algorithm, and highlighting the nucleoplasm contrast of lymphocytes; then strengthening the cell contour boundary by adopting an edge detection algorithm, and connecting the fracture edge through morphological operation to form a continuous and clear cell boundary mask; s2, morphological feature extraction: performing single cell segmentation based on the preprocessed image, and extracting geometric features and texture features of each cell; through machine learning model training, distinguishing lymphocytes and non-lymphocytes according to the characteristics, and generating a lymphocyte distribution probability graph; and S3, preliminarily identifying the aggregated area. By adopting the technical scheme of the invention, the TLSs differentiation stage can be identified, and the identification accuracy can be improved by combining morphological characteristics and spatial distribution characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing, and in particular to an artificial intelligence-based TLS structure delineation system and method. Background Technology

[0002] Accurate detection of tertiary lymphoid structures (TLSs) in the solid tumor microenvironment is a core element in predicting the efficacy of immunotherapy. TLSs, follicle-like structures formed by the aggregation of immune cells such as B cells and T cells, have morphological characteristics (e.g., lymphocyte density, follicle integrity, and presence or absence of germinal centers) in hematoxylin and eosin (HE) stained sections that form the basis of pathological assessment. However, traditional HE staining interpretation relies on manual microscopic examination, which presents the following technical bottlenecks:

[0003] Inconsistent subjective interpretation standards: There is a lack of standardized thresholds for determining TLSs positivity, and different studies have significantly different definitions of "positive," such as "aggregated lymphocytes > 50," "area > 1 high-power field (HPF)," or "lymphocyte density > 0.0128 μm." 2 Standards such as "..." coexist. Pathologists rely on experience to identify follicular structures and germinal centers (GCs), which can easily lead to missed or incorrect diagnoses. For example, early TLSs only show loose lymphocyte aggregation without obvious follicular structures, and the manual identification rate is less than 60%. Secondary TLSs require confirmation of CD21+ follicular dendritic cell networks and CD23+ germinal centers (requiring IHC staining assistance), and the accuracy of HE morphological interpretation alone is only about 75%.

[0004] Quantitative analysis is inefficient: Traditional whole-field digital slide (WSI) assessment requires pathologists to count the number of TLSs within every 10 HPFs, with a single test taking more than 30 minutes, and fatigue can easily lead to counting errors. In addition, manual annotation of the spatial distribution of TLSs (within the tumor, at the invasive margin, and adjacent to the tumor) requires multiple field switching, making it difficult to quickly complete three-dimensional spatial feature analysis.

[0005] Inadequate handling of morphological heterogeneity: The morphology of lymphocyte tracts (TLSs) varies significantly across different tumor types. For example, intratumoral TLSs in intrahepatic cholangiocarcinoma appear as dense masses, while in colorectal cancer, TLSs at the invasive margins are mostly scattered small follicles. Existing digital pathology tools are based solely on fixed thresholds (e.g., the minimum lymphocyte area is 6.245 μm). 2 The screening process is not adaptive and cannot learn the morphological characteristics of different tumors, resulting in poor generalization ability for cross-disease detection.

[0006] Artificial intelligence (AI) technology has provided a breakthrough direction for the automated analysis of HE images. Existing research attempts to identify lymphocyte aggregation regions in HE slices using convolutional neural networks (CNNs). For example, the model developed by Barmpoutis et al. has achieved preliminary labeling of TLS-positive regions based on lymphocyte number (≥45), area, and density characteristics.

[0007] However, this type of technology still has key drawbacks: ① It can only distinguish between "positive / negative" and cannot further determine the differentiation stage of TLSs (early / primary / secondary); ② It has not designed a dedicated algorithm for HE morphological features (such as follicular boundary ambiguity and cell arrangement regularity), resulting in insufficient accuracy in identifying atypical TLSs (such as aggregates containing necrotic components); ③ It lacks the ability to understand the spatial distribution characteristics of TLSs (such as invasive edge regions with a distance of <1mm from tumor cells), further leading to insufficient identification accuracy.

[0008] Therefore, there is an urgent need to develop a deep learning-based TLS structure delineation system that focuses on HE images, capable of recognizing the differentiation stages of TLSs, and combining morphological features and spatial distribution features to improve recognition accuracy. Summary of the Invention

[0009] This invention provides an AI-based TLS structure delineation system that can identify the differentiation stages of TLSs and can combine morphological features and spatial distribution features to improve the recognition accuracy.

[0010] To solve the above-mentioned technical problems, this application provides the following technical solution:

[0011] The AI-based TLS structure delineation method includes the following:

[0012] S1 image preprocessing standardizes the input HE-stained slice image. First, a color deconvolution algorithm is used to separate the staining components of the cell nucleus and cytoplasm to highlight the nucleocytoplasmic contrast of lymphocytes. Then, an edge detection algorithm is used to enhance the cell outline boundary, and morphological operations are used to connect the broken edges to form a continuous and clear cell boundary mask.

[0013] S2 morphological feature extraction is performed on preprocessed images to segment single cells and extract the geometric and texture features of each cell. Through machine learning model training, lymphocytes and non-lymphocytes are distinguished based on the above features, and a lymphocyte distribution probability map is generated.

[0014] S3 clustering regions were initially identified. Based on the lymphocyte distribution probability map, a density clustering algorithm was used to identify cell clustering foci. Minimum number of clustered cells and intercellular distance thresholds were set to filter out a small number of scattered cells and select candidate regions that meet the density characteristics of TLSs.

[0015] S4 morphological analysis performs semantic segmentation of follicular structures in candidate regions to distinguish between the core region composed of dense lymphocytes and the edge region composed of sparse cells or matrix. Follicular maturity is determined by calculating the regularity of the follicular edge contour: regular edge contours suggest the possibility of a mature follicle, while irregular edge contours suggest an early or primary follicular state.

[0016] S5 TLSs differentiation and grading: low-staining intensity regions were located in the follicular core area, and the TLSs were classified into three levels based on the results of edge regularity analysis.

[0017] S6 spatial distribution localization: Based on the spatial definition of tumor invasion margin, the distribution locations of TLSs within the tumor, at the invasion margin, and adjacent to the tumor are marked.

[0018] The basic principle and beneficial effects of the solution are as follows: This invention addresses the core problems of difficulty in distinguishing the differentiation stages of TLSs and insufficient handling of morphological heterogeneity in traditional HE staining interpretation. It achieves accurate analysis through multi-level feature modeling and spatial semantic understanding.

[0019] In image preprocessing (S1), color deconvolution and edge enhancement algorithms are used to separate the nucleoplasmic features of lymphocytes and enhance their contour clarity, providing a high-quality data foundation for subsequent follicular structure analysis. For example, local entropy analysis of hematoxylin channels can accurately locate germinal centers with low staining intensity, and combined with continuous cell boundary masks generated by edge detection, the single-cell segmentation error is reduced.

[0020] In the morphological feature extraction (S2-S4) stage, geometric features (such as roundness and nucleocytoplasmic ratio) and texture features (such as LBP pattern and GLCM matrix) extracted by single-cell segmentation are combined with deep learning models (such as U-TransNet) for semantic segmentation of the follicular core and edge regions, enabling a quantitative assessment of the maturity of TLSs. For example, by calculating the fractal dimension (FDim) and normalized root mean square error (NRMSE) of the follicular edge contour, early TLSs (FDim > 1.5, irregular edges), primary TLSs (FDim = 1.3-1.5, moderately regular edges), and secondary TLSs (FDim < 1.3, regular edges) can be distinguished, solving the subjective problem of traditional HE staining relying solely on experience to judge follicular structure.

[0021] By spatially defining the tumor invasion margin (fibrous tissue <1mm from tumor cells, S6), and combining distance transformation algorithms with 3D spatial interpolation techniques, the distribution of secondary TLSs is divided into three categories: intratumoral, invasion margin, and peritumoral. Intratumoral secondary TLSs are positively correlated with the efficacy of immunotherapy, while peritumoral TLSs have lower predictive value. This invention achieves differentiated analysis of TLSs in different regions through spatial localization, avoiding assessment bias caused by conflation.

[0022] In the clustering region identification (S3) and differentiation grading (S5), spatial density clustering algorithms (such as adaptive DBSCAN) and multi-parameter decision tree models are introduced. These are combined with lymphocyte clustering density, follicular edge regularity, germinal center characteristics, and spatial location to construct a "morphological-spatial" joint discrimination system. For example, because TLSs in invasive peripheral regions are closer to tumor cells, their edge regularity and maturity have higher weights in predicting treatment efficacy. The model dynamically adjusts feature weights, improving the identification accuracy of TLSs in this region compared to traditional methods.

[0023] From single-cell segmentation (S2) to follicular semantic segmentation (S4), and then to spatial distribution annotation (S6), the outputs of each step form a progressive feature chain: single-cell features provide the foundation for identifying aggregation follicles, follicular structural features provide the basis for differentiation grading, and spatial location features provide dimensions for clinical significance assessment. Through multi-stage feature fusion, the model can capture the dynamic process of TLSs from "loose aggregation → follicle formation → mature differentiation," as well as their spatial distribution patterns in the tumor microenvironment, achieving a comprehensive analysis of TLSs.

[0024] This invention, through a deep integration of morphological and spatial distribution characteristics, accurately identifies the differentiation stages of TLSs, overcoming the limitations of traditional assessment methods. Traditional HE staining can only roughly determine the "presence" of TLSs and cannot distinguish maturity levels. This invention, however, achieves a three-tiered classification of early, primary, and secondary TLSs through edge regularity analysis and germinal center localization, significantly improving accuracy. This provides crucial evidence for the precise stratification of immunotherapy patients. It can provide more effective diagnostic and treatment guidance; for example, patients with secondary TLSs accounting for >40% have an objective response rate (ORR) of up to 65% using PD-1 inhibitors, significantly higher than patients with predominantly early TLSs (ORR = 28%).

[0025] By clearly defining the tumor invasion margin (<1 mm from tumor cells), the model can distinguish the clinical significance of TLSs in different spatial regions. For example, in intrahepatic cholangiocarcinoma, intratumoral TLSs are associated with a good prognosis, while peritumoral TLSs may indicate an inflammatory response rather than antitumor immunity. This invention, through spatial localization, enhances the correlation coefficient (Spearman's ρ) between TLSs and the efficacy of immunotherapy, significantly improving the predictive power of biomarkers.

[0026] To address the morphological differences of TLSs in different tumor types (e.g., TLSs are mostly diffusely aggregated in gastric cancer and mostly follicular structures in breast cancer), the model uses multi-scale feature extraction (e.g., pyramid pooling network) and dynamic thresholding algorithms (e.g., neighborhood radius adjustment of adaptive DBSCAN). The model has high sensitivity in 10 solid tumors, including liver cancer, colorectal cancer, and lung cancer, which is an improvement over the traditional fixed threshold method and solves the limitation of existing tools that require a different method for each tumor.

[0027] Traditional manual assessment of TLSs takes over 30 minutes, and there are significant differences in interpretation consistency among different physicians. This invention achieves fully automated analysis of WSI within 5 minutes using an AI model, and significantly improves interpretation consistency. Combined with standardized report templates (such as differentiation grading and spatial distribution parameters), it can greatly reduce the workload of pathology departments and promote the widespread application of TLS testing in large-scale clinical cohorts.

[0028] In summary, this invention, through the dual innovation of "morphological feature analysis of differentiation stages + spatial distribution feature optimization of evaluation dimensions," can identify the differentiation stages of TLSs and improve identification accuracy by combining morphological and spatial distribution features. Furthermore, it constructs a TLS detection system that better meets clinical needs, not only solving bottlenecks but also providing a new paradigm for the development of immunotherapy biomarkers, demonstrating significant clinical translational value and promoting industry standardization.

[0029] Furthermore, in the S1 image preprocessing step, the color deconvolution algorithm employs an improved Ruifrok algorithm, separating the hematoxylin and eosin staining components using the following formula:

[0030]

[0031] Wherein, H is the hematoxylin staining component, and E is the eosin staining component; I RGB S is the RGB color space vector of the input image. H and S E The normalized staining concentration vectors for hematoxylin and eosin staining, respectively, A H and A E This represents the absorbance matrix of the corresponding staining component. An adaptive weighting factor is introduced. and Wherein, ∈ is a local constant that dynamically adjusts the contrast between the cell nucleus and cytoplasm, thereby enhancing the nucleocytoplasmic ratio characteristic of lymphocytes;

[0032] The edge detection algorithm adopts an improved Canny edge detection model based on dual thresholds. It preprocesses the image using a Gaussian convolution kernel G(σ) to suppress noise larger than σ = 1.5 pixels. The ratio of high to low thresholds is set to 3:1. The broken edges are connected by non-maximum suppression and hysteresis thresholding. Combined with morphological closing operation, the structuring element is a 3×3 rhombic matrix to fill the intercellular gaps and finally generate a cell boundary mask.

[0033] Furthermore, in the S2 morphological feature extraction step, the single-cell segmentation adopts an improved U-TransNet architecture, which integrates the local feature extraction capabilities of CNNs with the global modeling advantages of Transformers, including:

[0034] The multi-scale feature extraction module extracts feature maps F1, F2, F3, and F4 at four different scales through the ResNet50 backbone network, with 64, 256, 512, and 1024 channels respectively, and the spatial resolution is halved at each scale.

[0035] The Transformer augmented encoder performs patch segmentation on the highest resolution feature map F1 (patch size = 4×4), generates a token sequence through linear projection and positional encoding, and inputs it into a Transformer module containing 6 encoder layers. Each layer contains a multi-head self-attention mechanism and a feedforward neural network, and outputs the augmented feature representation T; where the number of heads in the multi-head self-attention mechanism is 8.

[0036] The feature fusion decoder fuses T with CNN feature maps F2, F3, and F4 through skip connections, uses deformable convolution to capture geometric changes in cell morphology, and finally generates a pixel-level segmentation mask M through upsampling.

[0037] The geometric feature extraction is calculated using the following formula:

[0038]

[0039] Among them, the cell nuclear area was determined by hematoxylin channel threshold segmentation, and the total cell area was calculated by mask M;

[0040] The texture feature extraction employs an improved local binary mode combined with a gray-level co-occurrence matrix algorithm:

[0041]

[0042] Among them, g c g represents the grayscale value of the center pixel. p Let be the grayscale value of the neighboring pixels, and s(x) be the sign function; based on this, a rotation invariant factor is introduced. Only retain patterns with U≤2 to reduce feature dimensions;

[0043] The machine learning model employs a multimodal feature fusion classifier to transform the 12-dimensional geometric feature vector V... g Texture feature vector V with dimension = 36 t Integrate with Transformer encoded features T of dimension = 768 through a gated fusion mechanism:

[0044] V fusion =σ(W g ·V g +W t ·V t )⊙T+(1-σ(W g ·V g +W t ·V t ))⊙T

[0045] Among them, W g W t The learning weight matrix is ​​σ, which is the sigmoid activation function, and ⊙ represents element-wise multiplication. Finally, the lymphocyte probability map P is output through the Softmax classifier.

[0046] Furthermore, in the initial identification step of the S3 clustering region, the density clustering algorithm adopts an improved adaptive DBSCAN model, which dynamically determines the neighborhood radius ∈ and the minimum number of samples MinPts using the following formula:

[0047]

[0048] Where p represents a pixel in the lymphocyte probability map P. α = 0.5 is the initial global neighborhood radius, α = 0.5 is the density adjustment factor, LocalDensity(p) is the average lymphocyte probability value within a 30×30 window centered at p, GlobalDensity is the average probability value across the entire image; β = 3 and γ = 10 are empirical parameters, and CellCount(N(p)) is the estimated number of cells within the neighborhood N(p).

[0049] The intercellular spacing threshold is calculated using an adaptive method.

[0050] Threshold(x,y)=μ·MedianDist(x,y)+σ·MAD(x,y)

[0051] Wherein, MedianDist(x,y) is the median distance of the 10 nearest neighbor cells around point (x,y), MAD(x,y) is the corresponding median absolute deviation, and μ=1.5 and σ=2.0 are robustness parameters; this method automatically shrinks the threshold in high-density regions and appropriately widens it in low-density regions, thereby improving its adaptability to different tumor microenvironments;

[0052] The step of filtering scattered cells was evaluated using multi-dimensional features:

[0053]

[0054] Where C represents the cluster, |C| represents the number of cells within the cluster, MinSize = 30 is the minimum cell count threshold, and MeanDensity(C) is the average lymphocyte probability within the cluster. The compactness index is defined by ω1 = 0.4, ω2 = 0.4, and ω3 = 0.2, which are weighting coefficients. When FilterScore(C) < 0.7, it is judged as a scattered cell cluster and filtered.

[0055] The final selected candidate regions are connected by morphological closing operations, with the structuring element being a 5×5 ellipse. Non-continuous regions are then eliminated through convex hull detection, generating a TLSs candidate region mask R.

[0056] Furthermore, the semantic segmentation of the follicle structure employs an improved DeepLabv3+ network architecture, combining dilated convolutions and attention mechanisms, including:

[0057] In the encoder section, a cascaded hollow spatial pyramid pooling module is introduced, and multi-scale contextual features are extracted by hollow convolution with dilation rates of 3, 6 and 9, respectively, to capture follicle structures of different sizes, where small follicles have a diameter of 100-200μm and large follicles have a diameter of >500μm.

[0058] A spatial attention module is added to the decoder, and the attention weight matrix is ​​generated using the following formula:

[0059] S=σ(W1·MaxPool(F)+W2·AvgPool(F))

[0060] Where F is the feature map output by the encoder, W1 and W2 are learnable weights, and σ is the sigmoid activation function, which enhances the feature response of the lymphocyte-dense region and suppresses the interference of the matrix and necrotic tissue.

[0061] The output segmentation results contain three types of labels: core region, dense lymphocytes, probability threshold > 0.8; edge region, sparse cells, probability threshold 0.3-0.8; background, probability threshold < 0.3.

[0062] Furthermore, the regularity of the follicle edge contour is calculated using a composite evaluation model combining Fourier descriptors and fractal dimension:

[0063] First, a Fourier transform is performed on the follicle profile, and the top 10 low-frequency coefficients are extracted to reconstruct the profile shape. Then, the normalized root mean square error is calculated.

[0064]

[0065] Among them, C i The original contour coordinates, To reconstruct the contour coordinates, <NRMSE<0.1 indicates high contour regularity;

[0066] Simultaneously calculate the fractal dimension of the profile:

[0067]

[0068] Where r is the measurement scale, and N(r) is the minimum number of scales required to cover the profile; the fractal dimension of mature follicles is close to 1.1-1.3, so they are regular curves; the fractal dimension of early follicles is >1.5, so they are irregular curves.

[0069] A maturity discrimination function is established by combining the two indicators:

[0070] MaturityScore=0.6·(1-NRMSE)-0.4·(FDim-1.0)

[0071] When MaturityScore > 0.5, it is judged as a mature follicle with regular edges; when MaturityScore < 0.2 < MaturityScore ≤ 0.5, it is judged as a primary follicle with moderately regular edges; when MaturityScore ≤ 0.2, it is judged as an early follicle with irregular edges.

[0072] Furthermore, in the S5 TLSs differentiation and grading step, the localization of the low-staining intensity region employs an adaptive threshold segmentation algorithm based on multimodal feature fusion, including:

[0073] Hematoxylin-eosin dual-channel feature extraction: For the preprocessed HE image, the local entropy E of the hematoxylin channel, i.e., H, is calculated separately. H The local contrast C of the eosin channel, i.e., E. E :

[0074]

[0075] Where, p i Let std be the probability of gray value i within the neighborhood N(x,y) (size 15×15 pixels), and mean be the standard deviation and mean, respectively.

[0076] Adaptive threshold calculation: Combining local and global statistical properties, the threshold T for low-staining regions is dynamically determined using the following formula. low :

[0077] T low (x,y)=α·Otsu(H)+β·LocalMean(H)-γ·C E (x,y)

[0078] Where Otsu(H) is the global Otsu threshold of the hematoxylin channel, LocalMean(H) is the local mean, and α = 0.6, β = 0.3, and γ = 0.1 are weighting coefficients;

[0079] Candidate regions for hair regrowth centers: Regions satisfying H(x,y) are selected. <T low (x,y) and E(x,y)>T high The pixel region (x, y) is used to generate a candidate mask G for the germinal center through morphological closing operation, a 7×7 structuring element, and area filtering; where T high The threshold for the eosin channel is calculated using the same method as T. low same;

[0080] The three-tier differentiation judgment adopts a multi-parameter fusion decision tree model, and the input parameters include:

[0081] Edge regularity score: EdgeScore, from MaturityScore;

[0082] Germ center feature intensity: Where R core For the core area mask;

[0083] Lymphocyte density gradient: Where R edge For edge region mask;

[0084] Nucleolar characteristic index: Calculated based on cell nucleus segmentation results;

[0085] The classification rules for decision trees are as follows:

[0086]

[0087] For uncertain categories, the following composite scoring function is used to further refine the judgment:

[0088] FinalScore=0.4·EdgeScore+0.3·GCIntensity+0.2·DensityGradient+0.1·NucleoliIndex

[0089] When FinalScore > 0.6, it is determined as secondary TLS; when 0.3 < FinalScore ≤ 0.6, it is determined as primary TLS; otherwise, it is determined as early TLS.

[0090] Furthermore, the spatial definition of the tumor invasion margin is realized by the tumor cell boundary recognition and distance transformation algorithm, including:

[0091] Tumor cell boundary extraction: An improved Mask R-CNN model is used to identify the tumor cell region by introducing a boundary-aware loss function:

[0092]

[0093] where, is the boundary pixel of the tumor cell region, S is the internal pixel, λ = 0.3 is the weight coefficient, CE is the cross-entropy loss, Dice is the Dice loss, which improves the boundary localization accuracy;

[0094] Distance transformation and region division: The binary map of the tumor cell boundary is subjected to Euclidean distance transformation to generate a distance field D(x, y), where each pixel value represents the distance from that point to the nearest tumor cell; the spatial region is divided according to the distance threshold:

[0095]

[0096] where, 1000μm corresponds to the physical distance of 1 high-power field of view under an optical microscope.

[0097] Furthermore, the spatial distribution annotation of the TLSs adopts a three-dimensional spatial interpolation and neighborhood correlation analysis algorithm, including:

[0098] Two-dimensional slice spatial mapping: Rigid registration is performed on the continuous slice sequence, and a three-dimensional coordinate mapping relationship is constructed through SIFT feature matching and thin plate spline interpolation (TPS). The two-dimensional TLSs candidate region mask R is projected into the three-dimensional space to generate a volume mask R3;

[0099] Neighborhood correlation probability calculation: For each TLSs cluster C i , calculate its volume proportion in the tumor V intra , invasion margin V marginal , peritumoral V paracancer , and generate a spatial distribution probability vector through the following formula:

[0100]

[0101] where, V total = V intra + V marginal + V paracancer .

[0102] Fuzzy classification decision: Introducing the fuzzy C-means clustering (FCM) algorithm for P space Soft classification was performed, with a membership threshold τ = 0.6. Regions with a membership degree ≥ τ were identified as the main distribution area. For TLSs clusters distributed across regions (e.g., simultaneously located at the invasion margin and adjacent to the tumor), the main region was determined using centroid localization.

[0103]

[0104] Where d is the Euclidean distance. Attached Figure Description

[0105] Figure 1 A flowchart illustrating the method for outlining the TLS structure based on artificial intelligence;

[0106] Figure 2 This is a schematic diagram of the result after S1 image preprocessing in the AI-based TLS structure delineation method.

[0107] Figure 3 This is a schematic diagram showing the results of the initial identification of the S3 clustering region in the AI-based TLS structure delineation method.

[0108] Figure 4 This is a schematic diagram of the result after S4 morphological analysis processing in the AI-based TLS structure delineation method. Detailed Implementation

[0109] The following detailed description illustrates the specific implementation method:

[0110] AI-based TLS structure delineation methods (such as...) Figure 1-4 As shown), it includes the following:

[0111] S1 image preprocessing standardizes the input HE-stained slice image. First, a color deconvolution algorithm is used to separate the staining components of the cell nucleus and cytoplasm to highlight the nucleocytoplasmic contrast of lymphocytes. Then, an edge detection algorithm is used to enhance the cell outline boundary, and morphological operations are used to connect the broken edges to form a continuous and clear cell boundary mask.

[0112] S2 morphological feature extraction is performed on preprocessed images to segment single cells and extract the geometric and texture features of each cell. Through machine learning model training, lymphocytes and non-lymphocytes are distinguished based on the above features, and a lymphocyte distribution probability map is generated.

[0113] S3 clustering regions were initially identified. Based on the lymphocyte distribution probability map, a density clustering algorithm was used to identify cell clustering foci. Minimum number of clustered cells and intercellular distance thresholds were set to filter out a small number of scattered cells and select candidate regions that meet the density characteristics of TLSs.

[0114] S4 morphological analysis performs semantic segmentation of follicular structures in candidate regions to distinguish between the core region composed of dense lymphocytes and the edge region composed of sparse cells or matrix. Follicular maturity is determined by calculating the regularity of the follicular edge contour: regular edge contours suggest the possibility of a mature follicle, while irregular edge contours suggest an early or primary follicular state.

[0115] S5 TLSs differentiation and grading: low-staining intensity regions were located in the follicular core area, and the TLSs were classified into three levels based on the results of edge regularity analysis.

[0116] S6 spatial distribution localization: Based on the spatial definition of tumor invasion margin, the distribution locations of TLSs within the tumor, at the invasion margin, and adjacent to the tumor are marked.

[0117] In practical application, during the S1 image preprocessing step, the color deconvolution algorithm employs a modified Ruifrok algorithm, separating the hematoxylin and eosin staining components using the following formula:

[0118]

[0119] Wherein, H is the hematoxylin staining component, and E is the eosin staining component; I RGB S is the RGB color space vector of the input image. H and S E The normalized staining concentration vectors for hematoxylin and eosin staining, respectively, A H and A E This represents the absorbance matrix of the corresponding staining component. An adaptive weighting factor is introduced. and Wherein, ∈ is a local constant that dynamically adjusts the contrast between the cell nucleus and cytoplasm, thereby enhancing the nucleocytoplasmic ratio characteristic of lymphocytes;

[0120] The edge detection algorithm adopts an improved Canny edge detection model based on dual thresholds. It preprocesses the image using a Gaussian convolution kernel G(σ) to suppress noise larger than σ = 1.5 pixels. The ratio of high to low thresholds is set to 3:1. The broken edges are connected by non-maximum suppression and hysteresis thresholding. Combined with morphological closing operation, the structuring element is a 3×3 rhombic matrix to fill the intercellular gaps and finally generate a cell boundary mask.

[0121] In the S2 morphological feature extraction step, the single-cell segmentation adopts an improved U-TransNet architecture, which integrates the local feature extraction capabilities of CNNs with the global modeling advantages of Transformers, including:

[0122] The multi-scale feature extraction module extracts feature maps F1, F2, F3, and F4 at four different scales through the ResNet50 backbone network, with 64, 256, 512, and 1024 channels respectively, and the spatial resolution is halved at each scale.

[0123] The Transformer augmented encoder performs patch segmentation on the highest resolution feature map F1 (patch size = 4×4), generates a token sequence through linear projection and positional encoding, and inputs it into a Transformer module containing 6 encoder layers. Each layer contains a multi-head self-attention mechanism and a feedforward neural network, and outputs the augmented feature representation T; where the number of heads in the multi-head self-attention mechanism is 8.

[0124] The feature fusion decoder fuses T with CNN feature maps F2, F3, and F4 through skip connections, uses deformable convolution to capture geometric changes in cell morphology, and finally generates a pixel-level segmentation mask M through upsampling.

[0125] The geometric feature extraction is calculated using the following formula:

[0126]

[0127] Among them, the cell nuclear area was determined by hematoxylin channel threshold segmentation, and the total cell area was calculated by mask M;

[0128] The texture feature extraction employs an improved local binary mode combined with a gray-level co-occurrence matrix algorithm:

[0129]

[0130] Among them, g c g represents the grayscale value of the center pixel. p Let be the grayscale value of the neighboring pixels, and s(x) be the sign function; based on this, a rotation invariant factor is introduced. Only retain patterns with U≤2 to reduce feature dimensions;

[0131] The machine learning model employs a multimodal feature fusion classifier to transform the 12-dimensional geometric feature vector V... g Texture feature vector V with dimension = 36 t Integrate with Transformer encoded features T of dimension = 768 through a gated fusion mechanism:

[0132] Vfusion =σ(W g ·V g +W t ·V t )⊙T+(1-σ(W g ·V g +W t ·V t ))⊙T

[0133] Among them, W g W t The learning weight matrix is ​​σ, which is the sigmoid activation function, and ⊙ represents element-wise multiplication. Finally, the lymphocyte probability map P is output through the Softmax classifier.

[0134] In the initial identification step of S3 clustering regions, the density clustering algorithm adopts an improved adaptive DBSCAN model, which dynamically determines the neighborhood radius ∈ and the minimum number of samples MinPts using the following formula:

[0135]

[0136] Where p represents a pixel in the lymphocyte probability map P. α = 0.5 is the initial global neighborhood radius, α = 0.5 is the density adjustment factor, LocalDensity(p) is the average lymphocyte probability value within a 30×30 window centered at p, GlobalDensity is the average probability value across the entire image; β = 3 and γ = 10 are empirical parameters, and CellCount(N(p)) is the estimated number of cells within the neighborhood N(p).

[0137] The intercellular spacing threshold is calculated using an adaptive method.

[0138] Threshold(x,y)=μ·MedianDist(x,y)+σ·MAD(x,y)

[0139] Wherein, MedianDist(x,y) is the median distance of the 10 nearest neighbor cells around point (x,y), MAD(x,y) is the corresponding median absolute deviation, and μ=1.5 and σ=2.0 are robustness parameters; this method automatically shrinks the threshold in high-density regions and appropriately widens it in low-density regions, thereby improving its adaptability to different tumor microenvironments;

[0140] The step of filtering scattered cells was evaluated using multi-dimensional features:

[0141]

[0142] Where C represents the cluster, |C| represents the number of cells within the cluster, MinSize = 30 is the minimum cell count threshold, and MeanDensity(C) is the average lymphocyte probability within the cluster. The compactness index is defined by ω1 = 0.4, ω2 = 0.4, and ω3 = 0.2, which are weighting coefficients. When FilterScore(C) < 0.7, it is judged as a scattered cell cluster and filtered.

[0143] The final selected candidate regions are connected by morphological closing operations, with the structuring element being a 5×5 ellipse. Non-continuous regions are then eliminated through convex hull detection, generating a TLSs candidate region mask R.

[0144] The follicle structure semantic segmentation adopts an improved DeepLabv3+ network architecture, combining dilated convolutions and attention mechanisms, including:

[0145] In the encoder section, a cascaded hollow spatial pyramid pooling module is introduced, and multi-scale contextual features are extracted by hollow convolution with dilation rates of 3, 6 and 9, respectively, to capture follicle structures of different sizes, where small follicles have a diameter of 100-200μm and large follicles have a diameter of >500μm.

[0146] A spatial attention module is added to the decoder, and the attention weight matrix is ​​generated using the following formula:

[0147] S=σ(W1·MaxPool(F)+W2·AvgPool(F))

[0148] Where F is the feature map output by the encoder, W1 and W2 are learnable weights, and σ is the sigmoid activation function, which enhances the feature response of the lymphocyte-dense region and suppresses the interference of the matrix and necrotic tissue.

[0149] The output segmentation results contain three types of labels: core region, dense lymphocytes, probability threshold > 0.8; edge region, sparse cells, probability threshold 0.3-0.8; background, probability threshold < 0.3.

[0150] The regularity of the follicle edge contour was calculated using a composite evaluation model combining Fourier descriptors and fractal dimension.

[0151] First, a Fourier transform is performed on the follicle profile, and the top 10 low-frequency coefficients are extracted to reconstruct the profile shape. Then, the normalized root mean square error is calculated.

[0152]

[0153] Among them, C i The original contour coordinates, To reconstruct the contour coordinates, <NRMSE<0.1 indicates high contour regularity;

[0154] Simultaneously calculate the fractal dimension of the profile:

[0155]

[0156] Where r is the measurement scale, and N(r) is the minimum number of scales required to cover the profile; the fractal dimension of mature follicles is close to 1.1-1.3, so they are regular curves; the fractal dimension of early follicles is >1.5, so they are irregular curves.

[0157] A maturity discrimination function is established by combining the two indicators:

[0158] MaturityScore=0.6·(1-NRMSE)-0.4·(FDim-1.0)

[0159] When MaturityScore > 0.5, it is judged as a mature follicle with regular edges; when MaturityScore < 0.2 < MaturityScore ≤ 0.5, it is judged as a primary follicle with moderately regular edges; when MaturityScore ≤ 0.2, it is judged as an early follicle with irregular edges.

[0160] In the S5 TLSs differentiation and grading step, the localization of low-staining intensity regions employs an adaptive threshold segmentation algorithm based on multimodal feature fusion, including:

[0161] Hematoxylin-eosin dual-channel feature extraction: For the preprocessed HE image, the local entropy E of the hematoxylin channel, i.e., H, is calculated separately. H The local contrast C of the eosin channel, i.e., E. E :

[0162]

[0163] Among them, p i Let std be the probability of gray value i within the neighborhood N(x,y) (size 15×15 pixels), and mean be the standard deviation and mean, respectively.

[0164] Adaptive threshold calculation: Combining local and global statistical properties, the threshold T for low-staining regions is dynamically determined using the following formula. low :

[0165] T low (x,y)=α·Otsu(H)+β·LocalMean(H)-γ·C E (x,y)

[0166] Where Otsu(H) is the global Otsu threshold of the hematoxylin channel, LocalMean(H) is the local mean, and α = 0.6, β = 0.3, and γ = 0.1 are weighting coefficients;

[0167] Screening of germinal center candidate regions: For pixels satisfying H(x,y) < T low (x,y) and E(x,y) > T high (x,y), a germinal center candidate mask G is generated through morphological closing operation, with a structuring element of 7×7, and area filtering; where T high is the eosin channel threshold, and the calculation method is the same as T low ;

[0168] The three-level differentiation judgment adopts a decision tree model with multi-parameter fusion, and the input parameters include:

[0169] Edge regularity score: EdgeScore, from MaturityScore;

[0170] Germinal center feature intensity: where R core is the core region mask;

[0171] Lymphocyte density gradient: where R edge is the marginal zone mask;

[0172] Nucleolus feature index: Calculated from the nucleus segmentation result;

[0173] The decision tree classification rules are as follows:

[0174]

[0175] For uncertain categories, further refined judgment is made through the following composite scoring function:

[0176] FinalScore = 0.4·EdgeScore + 0.3·GCIntensity + 0.2·DensityGradient + 0.1·NucleoliIndex

[0177] When FinalScore > 0.6, it is determined as secondary TLS; when 0.3 < FinalScore ≤ 0.6, it is determined as primary TLS; otherwise, it is determined as early TLS.

[0178] The spatial definition of the tumor invasion margin is realized through tumor cell boundary recognition and distance transformation algorithm, including:

[0179] Extraction of tumor cell boundary: An improved Mask R-CNN model is used to identify the tumor cell region by introducing a boundary-aware loss function:

[0180]

[0181] in, λ represents the boundary pixel of the tumor cell region, S represents the internal pixel, λ = 0.3 is the weight coefficient, CE is the cross-entropy loss, and Dice is the Dice loss, which improves the boundary localization accuracy.

[0182] Distance Transformation and Region Division: A Euclidean distance transformation is performed on the binary map of tumor cell boundaries to generate a distance field D(x,y), where each pixel value represents the distance from that point to the nearest tumor cell; spatial regions are then divided based on distance thresholds.

[0183]

[0184] Here, 1000μm corresponds to the physical distance of one high-power field of view under an optical microscope.

[0185] The spatial distribution labeling of TLSs employs a three-dimensional spatial interpolation and neighborhood association analysis algorithm, including:

[0186] Two-dimensional slice space mapping: rigid registration is performed on the continuous slice sequence, and a three-dimensional coordinate mapping relationship is constructed by SIFT feature matching and thin plate spline interpolation (TPS). The two-dimensional TLSs candidate region mask R is projected onto the three-dimensional space to generate a volume mask R3.

[0187] Neighborhood association probability calculation: For each TLSs cluster C i Calculate its V within the tumor intra 、Edge of Invasion V marginal , tumor-adjacent V paracancer The volume proportion is determined, and the spatial distribution probability vector is generated using the following formula:

[0188]

[0189] Among them, V total =V intra +V marginal +V paracancer .

[0190] Fuzzy classification decision: Introducing the fuzzy C-means clustering (FCM) algorithm for P space Soft classification was performed, with a membership threshold τ = 0.6. Regions with a membership degree ≥ τ were identified as the main distribution area. For TLSs clusters distributed across regions (e.g., simultaneously located at the invasion margin and adjacent to the tumor), the main region was determined using centroid localization.

[0191]

[0192] Where d is the Euclidean distance.

[0193] The dataset of HE-stained slide images of various solid tumors (including breast cancer, lung cancer, colorectal cancer, etc.) provided by a hospital was used as the test sample. The dataset contains TLSs samples with different differentiation stages and different spatial distributions, totaling 500 slide images.

[0194] One hundred slice images were randomly selected from the dataset for image preprocessing. First, the improved Ruifrok algorithm was used to separate hematoxylin and eosin staining components. When processing a breast cancer slice image, the input image's RGB color space vector is I. RGB Through formula Hematoxylin (H) and eosin (E) staining components were calculated. An adaptive weighting factor was introduced during the calculation. and Where ∈ takes the value 1×10 -6 Calculations showed that lymphocytes, which previously exhibited poor nuclear-cytoplasmic contrast under standard staining, showed a significantly enhanced nuclear-cytoplasmic ratio, with a clearer boundary between the nucleus and cytoplasm (e.g., ...). Figure 2 (As shown), to facilitate subsequent observation and analysis.

[0195] Next, an improved Canny edge detection model based on dual thresholds was employed to enhance cell contour boundaries. Image preprocessing using a Gaussian convolution kernel G(σ) (σ = 1.5) effectively suppressed noise larger than 1.5 pixels. A high-to-low threshold ratio of 3:1 was set, and after non-maximum suppression and hysteresis thresholding to connect broken edges, morphological closing operations (using a 3×3 rhombic matrix as the structuring element) were combined to fill cell gaps. The processed image generated a continuous and clear cell boundary mask, laying a solid foundation for subsequent single-cell segmentation and feature extraction. By comparing the images before and after preprocessing, it is clearly visible that the cell contours become clearer, and the boundaries between cells are more distinct, which greatly improves the accuracy of subsequent analysis.

[0196] For the preprocessed image, single-cell segmentation and feature extraction are performed. An improved U-TransNet architecture is used for single-cell segmentation, employing a ResNet50 backbone network to extract four feature maps F1, F2, F3, and F4 at different scales, with 64, 256, 512, and 1024 channels respectively, halving the spatial resolution at each scale. The highest resolution feature map F1 is segmented using a patch (patch size = 4×4), and a token sequence is generated through linear projection and positional encoding. This token sequence is input into a Transformer module containing six encoder layers (the multi-head self-attention mechanism has eight heads), outputting an enhanced feature representation T. T is then fused with the CNN feature maps F2, F3, and F4 via skip connections. Deformable convolutions are used to capture the geometric changes in cell morphology, and finally, upsampling is used to generate a pixel-level segmentation mask M.

[0197] For geometric feature extraction, the cell nucleus area is determined by threshold segmentation using hematoxylin channels, and the total cell area is calculated using a mask M. Then, geometric features such as roundness, nucleocytoplasmic ratio, and eccentricity are calculated according to formulas. When calculating texture features, an improved local binary mode combined with a gray-level co-occurrence matrix algorithm is used. A rotation invariant factor U is introduced, retaining only modes with U ≤ 2, reducing the feature dimension. The geometric feature vector (dimension = 12), texture feature vector (dimension = 36), and Transformer encoded features T (dimension = 768) are integrated through a gated fusion mechanism, and finally, a lymphocyte probability map P is output through a Softmax classifier. From the generated lymphocyte probability map, the distribution of lymphocytes in the image can be clearly seen; regions with higher probability values ​​indicate a greater likelihood of lymphocyte presence.

[0198] Based on the generated lymphocyte probability map P, an improved adaptive DBSCAN model is used to identify cell aggregation foci. For each pixel p in the image, the formula is used... and The neighborhood radius ∈ and the minimum number of samples MinPts are dynamically determined. Wherein, Let α = 0.5, β = 3, and γ = 10. In a colorectal cancer slice image, this model can accurately identify areas of cell aggregation, avoiding misclassification of scattered small numbers of cells as clusters.

[0199] The intercellular distance threshold is calculated using an adaptive method: Threshold(x,y) = μ·MedianDist(x,y) + σ·MAD(x,y), where μ = 1.5 and σ = 2.0. This method automatically narrows the threshold in high-density regions and appropriately widens it in low-density regions. For example, in densely populated tumor cell areas, the intercellular distance threshold is smaller, effectively filtering out scattered cells that do not meet the density characteristics of TLSs (Transient Transient Distances); while in relatively sparse regions, the threshold is widened to avoid missing potential TLS regions.

[0200] Scattered cells are filtered through multi-dimensional feature evaluation. When FilterScore(C) < 0.7, they are identified as scattered cell clusters and filtered. For the selected candidate regions, morphological closing operations (structural element is an ellipse of 5×5) are performed to connect broken parts, and non-contiguous regions are excluded by convex hull detection, generating a TLS candidate region mask R. From the generated mask R, the location and extent of the TLS candidate regions can be visually observed (e.g., ...). Figure 3 As shown in the figure, this provides accurate regional information for subsequent morphological analysis.

[0201] For the region corresponding to the generated TLSs candidate region mask R, an improved DeepLabv3+ network architecture is used for follicular structure semantic segmentation. A cascaded dilated spatial pyramid pooling module is introduced in the encoder part, and dilated convolutions with dilation rates of 3, 6, and 9 are used to extract multi-scale context features respectively to capture follicular structures of different sizes. A spatial attention module is added to the decoder, and an attention weight matrix is generated through the formula S = σ(W1·MaxPool(F)+W2·AvgPool(F)) to enhance the feature response in the lymphocyte-dense region and suppress the interference of stroma and necrotic tissues (as Figure 4 shown). The output segmentation results include three types of labels: the core region (dense lymphocytes, probability threshold > 0.8), the marginal region (sparse cells, probability threshold 0.3 - 0.8), and the background (probability threshold < 0.3).

[0202] When judging the follicular maturity, first perform a Fourier transform on the follicular contour, extract the first 10 low-frequency coefficients to reconstruct the contour shape, and calculate the normalized root mean square error NRMSE. At the same time, calculate the fractal dimension FDim of the contour, and establish a maturity discrimination function MaturityScore = 0.6·(1 - NRMSE) - 0.4·(FDim - 1.0) by combining the two indicators. When MaturityScore > 0.5, it is determined as a mature follicle with regular margins; when 0.2 < MaturityScore ≤ 0.5, it is determined as a primary follicle with moderately regular margins; when MaturityScore ≤ 0.2, it is determined as an early follicle with irregular margins. Through this method, the follicular maturity can be accurately judged, providing an important basis for the analysis of TLSs.

[0203] In the follicular core region, an adaptive threshold segmentation algorithm based on multi-modal feature fusion is used to locate the low-staining intensity region. For the preprocessed HE image, calculate the local entropy E H of the hematoxylin channel and the local contrast C E of the eosin channel respectively. Combining local and global statistical characteristics, the threshold T low (x, y) = α·Otsu(H)+β·LocalMean(H)-γ·C E for the low-staining region is dynamically determined, where α = 0.6, β = 0.3, and γ = 0.1. For the pixel region that satisfies H(x, y) < T low (x, y) and E(x, y) > T low (x, y), a germinal center candidate mask G is generated through morphological closing operation (structural element 7×7) and area filtering.

[0204] ​​A decision tree model with multi-parameter fusion is used to conduct a three-level differentiation judgment on TLSs. The input parameters include the edge regularity score EdgeScore (from MaturityScore), the germinal center feature intensity GCIntensity, the lymphoid cell density gradient DensityGradient, and the nucleolus feature index NucleoliIndex. The decision tree classification rules are as follows: when EdgeScore ≤ 0.2 and GCIntensity < 0.15, it is determined as early TLS; when 0.2 < EdgeScore ≤ 0.5 and GCIntensity < 0.3, it is determined as primary TLS; when EdgeScore > 0.5 and GCIntensity ≥ 0.3, it is determined as secondary TLS; otherwise, it is uncertain. For the uncertain category, a composite scoring function FinalScore = 0.4·EdgeScore + 0.3·GCIntensity + 0.2·DensityGradient + 0.1·NucleoliIndex is used for further refined judgment.

[0205] In terms of spatial distribution positioning, a tumor cell boundary recognition and distance transformation algorithm is used to define the tumor invasion edge. An improved Mask R-CNN model is used to identify the tumor cell region, and a boundary-aware loss function (λ = 0.3) is introduced to improve the boundary positioning accuracy. An Euclidean distance transformation is performed on the binary map of the tumor cell boundary, and the spatial regions are divided according to the distance threshold: D(x, y) = 0 is inside the tumor, 0 < D(x, y) ≤ 1000μm is the invasion edge, and D(x, y) > 1000μm is adjacent to the tumor.

[0206] A three-dimensional spatial interpolation and neighborhood correlation analysis algorithm is used to label the spatial distribution of TLSs. Rigid registration is performed on the continuous slice sequence, and a three-dimensional coordinate mapping relationship is constructed through SIFT feature matching and thin plate spline interpolation (TPS). The two-dimensional TLSs candidate region mask R is projected into the three-dimensional space to generate the volume mask R3. For each TLSs cluster C i , its volume proportions in the tumor V intra , the invasion edge V marginal , and adjacent to the tumor V paracancer are calculated to generate the spatial distribution probability vector P space . The fuzzy C-means clustering (FCM) algorithm is introduced for P spaceSoft classification was performed, with a membership threshold τ = 0.6. Regions with a membership degree ≥ τ were identified as major distribution areas. For TLS clusters distributed across regions, the centroid localization method was used to determine the main region. These steps accurately determined the distribution locations of TLSs within tumors, at invasive margins, and adjacent to tumors, providing strong support for in-depth research on the relationship between TLSs and tumors.

[0207] As can be seen from the above embodiments, the AI-based TLS structure delineation system and method of the present invention can accurately perform image preprocessing, morphological feature extraction, aggregation region identification, follicular structure analysis, TLS differentiation grading, and spatial distribution localization when processing HE-stained slide images of different types of tumors. It effectively solves the problems of inconsistent subjective interpretation standards, low efficiency of quantitative analysis, and insufficient handling of morphological heterogeneity inherent in traditional methods, providing a reliable technical means for the accurate detection of TLSs in the solid tumor microenvironment and the prediction of immunotherapy efficacy. In practical applications, the system parameters and algorithms can be further optimized and adjusted according to different needs and scenarios to improve the system's performance and adaptability.

[0208] The above are merely embodiments of the present invention. The invention is not limited to the fields covered by these embodiments. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are able to access all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A TLS structure delineation method based on artificial intelligence, characterized in that, Includes the following: S1 image preprocessing standardizes the input HE-stained slice image. First, a color deconvolution algorithm is used to separate the staining components of the cell nucleus and cytoplasm to highlight the nucleocytoplasmic contrast of lymphocytes. Then, an edge detection algorithm is used to enhance the cell outline boundary, and morphological operations are used to connect the broken edges to form a continuous and clear cell boundary mask. S2 morphological feature extraction is based on preprocessed images to perform single-cell segmentation and extract the geometric and texture features of each cell. By training a machine learning model, lymphocytes and non-lymphocytes are distinguished based on the above features, and a lymphocyte distribution probability map is generated. S3 clustering regions were initially identified. Based on the lymphocyte distribution probability map, a density clustering algorithm was used to identify cell clustering foci. Minimum number of clustered cells and intercellular distance thresholds were set to filter out a small number of scattered cells and select candidate regions that meet the density characteristics of TLSs. S4 morphological analysis performs semantic segmentation of follicular structures in candidate regions to distinguish between the core region composed of dense lymphocytes and the edge region composed of sparse cells or matrix. Follicular maturity is determined by calculating the regularity of the follicular edge contour: regular edge contours suggest the possibility of a mature follicle, while irregular edge contours suggest an early or primary follicular state. S5 TLSs differentiation and grading: low-staining intensity regions were located in the follicular core area, and the TLSs were classified into three levels based on the results of edge regularity analysis. S6 spatial distribution localization: Based on the spatial definition of tumor invasion margin, the distribution locations of TLSs within the tumor, at the invasion margin, and adjacent to the tumor are marked.

2. The AI-based TLS structure delineation method according to claim 1, characterized in that, In the S1 image preprocessing step, the color deconvolution algorithm uses a modified Ruifrok algorithm to separate hematoxylin and eosin staining components using the following formula: Wherein, H is the hematoxylin staining component, and E is the eosin staining component; I RGB S is the RGB color space vector of the input image. H and S E The normalized staining concentration vectors for hematoxylin and eosin staining, respectively, A H and A E This represents the absorbance matrix of the corresponding staining component. An adaptive weighting factor is introduced. and Wherein, ∈ is a local constant that dynamically adjusts the contrast between the cell nucleus and cytoplasm, thereby enhancing the nucleocytoplasmic ratio characteristic of lymphocytes; The edge detection algorithm adopts an improved Canny edge detection model based on dual thresholds. It preprocesses the image using a Gaussian convolution kernel G(σ) to suppress noise larger than σ = 1.5 pixels. The ratio of high to low thresholds is set to 3:

1. The broken edges are connected by non-maximum suppression and hysteresis thresholding. Combined with morphological closing operation, the structuring element is a 3×3 rhombic matrix to fill the intercellular gaps and finally generate a cell boundary mask.

3. The TLS structure delineation method based on artificial intelligence according to claim 2, characterized in that, In the S2 morphological feature extraction step, the single-cell segmentation adopts an improved U-TransNet architecture, which integrates the local feature extraction capabilities of CNNs with the global modeling advantages of Transformers, including: The multi-scale feature extraction module extracts feature maps F1, F2, F3, and F4 at four different scales through the ResNet50 backbone network, with 64, 256, 512, and 1024 channels respectively, and the spatial resolution is halved at each scale. The Transformer augmented encoder performs patch segmentation on the highest resolution feature map F1 (patch size = 4×4), generates a token sequence through linear projection and positional encoding, and inputs it into a Transformer module containing 6 encoder layers. Each layer contains a multi-head self-attention mechanism and a feedforward neural network, and outputs the augmented feature representation T; where the number of heads in the multi-head self-attention mechanism is 8. The feature fusion decoder fuses T with CNN feature maps F2, F3, and F4 through skip connections, uses deformable convolution to capture geometric changes in cell morphology, and finally generates a pixel-level segmentation mask M through upsampling. The geometric feature extraction is calculated using the following formula: Among them, the cell nuclear area was determined by hematoxylin channel threshold segmentation, and the total cell area was calculated by mask M; The texture feature extraction employs an improved local binary mode combined with a gray-level co-occurrence matrix algorithm: Among them, g c g represents the grayscale value of the center pixel. p Let be the grayscale value of the neighboring pixels, and s(x) be the sign function; based on this, a rotation invariant factor is introduced. Only retain patterns with U≤2 to reduce feature dimensions; The machine learning model employs a multimodal feature fusion classifier to transform the 12-dimensional geometric feature vector V... g Texture feature vector V with dimension = 36 t Integrate with Transformer encoded features T of dimension = 768 through a gated fusion mechanism: V fusion =σ(W g ·V g +W t ·V t )⊙T+(1-σ(W g ·V g +W t ·V t ))⊙T Among them, W g W t The learning weight matrix is ​​σ, which is the sigmoid activation function, and ⊙ represents element-wise multiplication. Finally, the lymphocyte probability map P is output through the Softmax classifier.

4. The TLS structure delineation method based on artificial intelligence according to claim 3, characterized in that, In the initial identification step of the S3 clustering region, the density clustering algorithm adopts an improved adaptive DBSCAN model, which dynamically determines the neighborhood radius ∈ and the minimum number of samples MinPts using the following formula: Where p represents a pixel in the lymphocyte probability map P. α = 0.5 is the initial global neighborhood radius, α = 0.5 is the density adjustment factor, LocalDensity(p) is the average lymphocyte probability value within a 30×30 window centered at p, GlobalDensity is the average probability value across the entire image; β = 3 and γ = 10 are empirical parameters, and CellCount(N(p)) is the estimated number of cells within the neighborhood N(p). The intercellular spacing threshold is calculated using an adaptive method. Threshold(x,y)=μ·MedianDist(x,y)+σ·MAD(x,y) Wherein, MedianDist(x,y) is the median distance of the 10 nearest neighbor cells around point (x,y), MAD(x,y) is the corresponding median absolute deviation, and μ=1.5 and σ=2.0 are robustness parameters; this method automatically shrinks the threshold in high-density regions and appropriately widens it in low-density regions, thereby improving its adaptability to different tumor microenvironments; The step of filtering scattered cells was evaluated using multi-dimensional features: Where C represents the cluster, |C| represents the number of cells within the cluster, MinSize = 30 is the minimum cell count threshold, and MeanDensity(C) is the average lymphocyte probability within the cluster. The compactness index is defined by ω1 = 0.4, ω2 = 0.4, and ω3 = 0.2, which are weighting coefficients. When FilterScore(C) < 0.7, it is judged as a scattered cell cluster and filtered. The final selected candidate regions are connected by morphological closing operations, with the structuring element being a 5×5 ellipse. Non-continuous regions are then eliminated through convex hull detection, generating a TLSs candidate region mask R.

5. The AI-based TLS structure delineation method according to claim 4, characterized in that, The semantic segmentation of the follicle structure employs an improved DeepLabv3+ network architecture, combining dilated convolutions and attention mechanisms, including: In the encoder section, a cascaded hollow spatial pyramid pooling module is introduced, and multi-scale contextual features are extracted by hollow convolution with dilation rates of 3, 6 and 9, respectively, to capture follicle structures of different sizes, where small follicles have a diameter of 100-200μm and large follicles have a diameter of >500μm. A spatial attention module is added to the decoder, and the attention weight matrix is ​​generated using the following formula: S=σ(W1·MaxPool(F)+W2·AvgPool(F)) Where F is the feature map output by the encoder, W1 and W2 are learnable weights, and σ is the sigmoid activation function, which enhances the feature response of the lymphocyte-dense region and suppresses the interference of the matrix and necrotic tissue. The output segmentation results contain three types of labels: core region, dense lymphocytes, probability threshold > 0.8; edge region, sparse cells, probability threshold 0.3-0.8; background, probability threshold < 0.

3.

6. The TLS structure delineation method based on artificial intelligence according to claim 5, characterized in that, The calculation of the degree of regularity of the follicular edge contour uses a composite evaluation model that combines Fourier descriptors and fractal dimension: First, perform a Fourier transform on the follicular contour, extract the first 10 low-frequency coefficients to reconstruct the contour shape, and calculate the normalized root mean square error: Among them, C i The original contour coordinates, To reconstruct the contour coordinates, <NRMSE<0.1 indicates high contour regularity; At the same time, calculate the fractal dimension of the contour: where r is the measurement scale and N(r) is the minimum number of scales required to cover the contour; if the fractal dimension of a mature follicle is close to 1.1 - 1.3, it is a regular curve; if the fractal dimension of an early follicle > 1.5, it is an irregular curve; Establish a maturity discrimination function by integrating the two indicators: MaturityScore = 0.6·(1 - NRMSE) - 0.4·(FDim - 1.0) When > MaturityScore > 0.5, it is determined to be a mature follicle with a regular edge; when < 0.2 < MaturityScore ≤ 0.5, it is determined to be a primary follicle with a moderately regular edge; when MaturityScore ≤ 0.2, it is determined to be an early follicle with an irregular edge.

7. The AI-based TLS structure delineation method according to claim 6, characterized in that, In the S5 TLSs differentiation grading step, the low staining intensity region is located using an adaptive threshold segmentation algorithm that combines multi-modal features, including: Hematoxylin-eosin dual-channel feature extraction: For the preprocessed HE image, the local entropy E of the hematoxylin channel, i.e., H, is calculated separately. H The local contrast C of the eosin channel, i.e., E. E : Where, p i Let std be the probability of gray value i within the neighborhood N(x,y) (size 15×15 pixels), and mean be the standard deviation and mean, respectively. Adaptive threshold calculation: Combining local and global statistical properties, the threshold T for low-staining regions is dynamically determined using the following formula. low : T low (x,y)=α·Otsu(H)+β·LocalMean(H)-γ·C E (x,y) where Otsu(H) is the global Otsu threshold of the hematoxylin channel, LocalMean(H) is the local mean, and α = 0.6, β = 0.3, γ = 0.1 are weight coefficients; Candidate regions for hair regrowth centers: Regions satisfying H(x,y) are selected. <T low (x,y) and E(x,y)>T high The pixel region (x, y) is used to generate a candidate mask G for the germinal center through morphological closing operation, a 7×7 structuring element, and area filtering; where T high The threshold for the eosin channel is calculated using the same method as T. low same; The three-level differentiation judgment uses a decision tree model that combines multiple parameters, and the input parameters include: Edge regularity score: EdgeScore, from MaturityScore; Germ center feature intensity: Where R core For the core area mask; Lymphocyte density gradient: Where R edge For edge region mask; Nucleolar characteristic index: Calculated based on cell nucleus segmentation results; The decision tree classification rules are as follows: For uncertain categories, further refine the judgment through the following composite scoring function: FinalScore = 0.4·EdgeScore + 0.3·GCIntensity + 0.2·DensityGradient + 0.1·NucleoliIndex When FinalScore > 0.6, it is determined to be a secondary TLS; when 0.3 < FinalScore ≤ 0.6, it is determined to be a primary TLS; otherwise, it is determined to be an early TLS.

8. The TLS structure delineation method based on artificial intelligence according to claim 7, characterized in that, The spatial definition of the tumor invasion edge is achieved through a tumor cell boundary recognition and distance transformation algorithm, including: Tumor cell boundary extraction: Use an improved Mask R-CNN model to identify the tumor cell region by introducing a boundary-aware loss function: in, λ represents the boundary pixel of the tumor cell region, S represents the internal pixel, λ = 0.3 is the weight coefficient, CE is the cross-entropy loss, and Dice is the Dice loss, which improves the boundary localization accuracy. Distance transformation and region division: Perform an Euclidean distance transformation on the binary map of the tumor cell boundary to generate a distance field D(x, y), where each pixel value represents the distance from that point to the nearest tumor cell; divide the spatial region according to the distance threshold: where 1000μm corresponds to the physical distance of 1 high-power microscope field of view under an optical microscope.

9. The AI-based TLS structure delineation method according to claim 8, characterized in that, The TLSs spatial distribution annotation uses a three-dimensional spatial interpolation and neighborhood correlation analysis algorithm, including: Two-dimensional slice space mapping: Perform rigid registration on a continuous slice sequence, construct a three-dimensional coordinate mapping relationship through SIFT feature matching and thin plate spline interpolation (TPS), project the two-dimensional TLSs candidate region mask R into three-dimensional space, and generate a volume mask R3; Neighborhood association probability calculation: For each TLSs cluster C i Calculate its V within the tumor intra 、Edge of Invasion V marginal , tumor-adjacent V paracancer The volume proportion is determined, and the spatial distribution probability vector is generated using the following formula: Among them, V total =V intra +V marginal +V paracancer . Fuzzy classification decision: Introducing the fuzzy C-means clustering (FCM) algorithm for P space Soft classification was performed, with a membership threshold τ = 0.

6. Regions with a membership degree ≥ τ were identified as the main distribution area. For TLSs clusters distributed across regions (e.g., simultaneously located at the invasion margin and adjacent to the tumor), the main region was determined using centroid localization. Where d is the Euclidean distance.

10. A TLS structure delineation system based on artificial intelligence, characterized in that, The method described in any one of claims 1-9 was adopted.

Citation Information

Cited By

  • Method and system for detecting semantic change of remote sensing image with high spatial resolution

    CN121305372A