An inflammation hotspot area identification method, device, equipment and storage medium
By calculating the spatial inflammation score coefficient of cells in spatial transcriptomics and constructing an abnormal baseline model, the problems of accuracy and boundary ambiguity in the identification of inflammatory hotspots in existing technologies are solved, and the accurate identification and clear definition of inflammatory hotspots in complex tissues are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KANGMEIHUA GENE TECH CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-05-29
AI Technical Summary
Existing spatial transcriptomics techniques are not adaptable to complex biological tissues when identifying inflammatory hotspots in tissue sections. They cannot deeply integrate gene expression characteristics with cellular spatial relationships and lack objective and standardized health baseline references, resulting in inaccurate inflammation assessments and blurred boundaries of hotspot regions.
By acquiring spatial transcriptome data of the tissue under test, calculating the spatial inflammation score coefficient of cells, constructing an abnormal baseline model by combining data from a healthy control group, screening and clustering abnormal cells, and quantifying the spatial enrichment characteristics of cell clusters to identify inflammatory hotspots.
It enables precise assessment of the inflammatory microenvironment, eliminates tissue noise interference, clearly defines the spatial boundaries of high-incidence inflammatory areas, and improves the accuracy and reliability of identifying inflammatory hotspots in complex tissues.
Smart Images

Figure CN122117010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial transcriptomics technology, and in particular to a method, apparatus, device and storage medium for identifying inflammatory hotspots. Background Technology
[0002] In the field of biomedical research, spatial transcriptomics is gradually becoming an important tool for disease research and biological development analysis. This technology can acquire gene expression information of cells in tissue sections and their corresponding physical spatial coordinates, providing unprecedented spatial resolution and molecular-level insights for pathological analysis at the microscopic level. Using this technology, researchers can deeply analyze the spatial distribution and enrichment patterns of gene expression profiles, which is of paramount importance for revealing the underlying mechanisms of disease development.
[0003] In studies of various tissue injuries or diseases, identifying inflammatory hotspots (i.e., areas with highly concentrated and active inflammatory responses) in tissue sections is a crucial step in pathological diagnosis. Currently, in the field of spatial transcriptome data analysis, existing identification methods typically involve directly summing or averaging the expression levels of specific gene clusters associated with inflammation to roughly estimate and delineate potential inflammatory risk areas within the tissue.
[0004] However, existing identification schemes have several limitations in terms of universality and adaptability to complex biological tissues. First, conventional scoring methods often view gene expression levels in isolation, failing to deeply couple gene expression characteristics with the complex spatial-physical neighborhood relationships between cells, resulting in inaccurate assessments of the true inflammatory microenvironment of cells. Second, in distinguishing between normal cells and abnormal inflammatory cells, current technologies lack an objective and standardized healthy baseline as a reference, making it difficult to accurately screen for truly abnormal cells in complex tissue heterogeneity. Furthermore, in determining the final hotspot regions, existing methods often lack effective quantification and cluster analysis of cell spatial aggregation density, making it difficult to accurately define the spatial boundaries of locally enriched regions.
[0005] In summary, existing spatial transcriptome analysis techniques have significant technical limitations in terms of deep integration of spatial and expression characteristics, establishment of objective benchmarks for abnormal cell screening, and accurate assessment of spatial region enrichment features. These limitations result in low accuracy and ambiguous boundary delineation when identifying inflammatory hotspots in complex tissues (such as alcoholic liver injury), greatly restricting their widespread application in disease mechanism research. A more accurate identification scheme is urgently needed. Summary of the Invention
[0006] In a first aspect, the present invention provides a method for identifying inflammatory hotspots, comprising: Acquire spatial transcriptome data of the tissue to be tested; wherein, the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue to be tested and the expression levels of inflammation-related genes; Based on the spatial coordinates and the expression level, the spatial inflammation score coefficient of each of the tested cells is calculated; An abnormal baseline model was constructed based on data from a healthy control group, and abnormal cells were screened from the test cells by combining the abnormal baseline model with the spatial inflammation score coefficient. The abnormal cells are clustered to obtain at least one cell cluster; Based on the spatial enrichment characteristics of the cell clusters, inflammatory hotspots in the tissue under test are identified.
[0007] In an optional implementation, the step of constructing an abnormal baseline model based on healthy control group data, and combining the abnormal baseline model with the spatial inflammation score coefficient to screen for abnormal cells from the test cells, includes: The spatial inflammation score coefficient of cells in the healthy control group was obtained, and extreme value sample data were extracted according to cell type; The generalized extreme value distribution model was fitted using the extreme value sample data; Substitute the spatial inflammation score coefficient of the cell under test into the generalized extreme value distribution model to calculate the weighted marginal p-value of the cell under test. Cells whose weighted marginal p-value is less than a preset significance threshold are identified as abnormal cells.
[0008] In an optional implementation, substituting the spatial inflammation score coefficient of the test cells into the generalized extreme value distribution model to calculate the weighted marginal p-value of the test cells includes: Determine the proportional weight of each cell type within the spatial neighborhood of the cell to be tested; For each cell type, the spatial inflammation score coefficient of the cell under test is substituted into the cumulative distribution function of its corresponding generalized extreme value distribution model to calculate the conditional p-value; Based on the proportional weights corresponding to each cell type, all the condition p-values are weighted and summed to obtain the weighted marginal p-value of the cell under test.
[0009] In an optional implementation, the weighted marginal p-value of the cell under test is calculated as follows: ; Where, p j The weighted marginal p-value represents the cell j being tested; G represents the set of cell clusters or types; g represents a specific cell cluster or type within set G; π jgThe probability weight representing the cell group or type g to which the test cell j belongs; IID j G represents the spatial inflammation score coefficient of cell j being tested; g (IID j This represents the cumulative distribution function value calculated by substituting the spatial inflammation score coefficient of the cell j under test into the generalized extreme value distribution model corresponding to the cell cluster or type g. The expression for calculating the cumulative distribution function value is as follows: ; Where Gg(x) represents the cumulative distribution function value of the generalized extreme value distribution model corresponding to cell cluster or type g; x represents the input spatial inflammation score coefficient; μ g ξ represents the location parameter of the generalized extreme value distribution model corresponding to cell cluster or type g; g The shape parameter representing the cell cluster or type g; σ g The scale parameter represents the cell cluster or type g; exp represents an exponential function with the natural constant as the base.
[0010] In an optional implementation, the extraction of extreme value sample data by cell type includes: For each cell type in the healthy control group, the spatial inflammation score coefficients of each cell contained therein were ranked. Extract the spatial inflammation score coefficient of the cells that rank in the top preset percentage, and use it as the extreme value sample data corresponding to that cell type.
[0011] In an optional implementation, the healthy control group data includes n c The record information of each cell; the record information of the i-th cell is represented as: ; in, Represents the sample name to which the i-th cell belongs; The cell name representing the i-th cell; The spatial inflammation score coefficient representing the i-th cell; Represents the group or type to which the i-th cell belongs, and 1≤i≤n c ; Furthermore, for each cell type g in the healthy control group, the expression for extracting the extreme value sample data set corresponding to that cell type is: ; Here, Ctrlg represents the initial candidate set of extreme value sample data corresponding to cell type g.
[0012] In an optional implementation, calculating the spatial inflammation score coefficient for each cell under test based on the spatial coordinates and the expression level includes: For each of the cells to be tested, a local neighborhood network is determined based on its spatial coordinates; The spatial-expression coupling potential is calculated by combining the expression levels of inflammation-related genes in the test cells and the local neighborhood network. A spatial smoothing transformation is performed on the spatial-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the tested cells.
[0013] In an optional implementation, calculating the spatial-expression coupling potential value by combining the expression levels of inflammation-related genes in the test cells with the local neighborhood network includes: The basal inflammatory activity of the test cells was calculated using local functional activity functions. Based on the spatial distance and gene expression similarity between cells in the local neighborhood network, the network decay weights are determined; The spatial-expression coupling potential value is calculated based on the inflammatory baseline activity and the network decay weights.
[0014] In an optional implementation, determining the network decay weights based on the spatial distance and gene expression similarity between cells in the local neighborhood network includes: The distance decay factor is calculated based on the spatial Euclidean distance between the cell under test and its neighboring cells in its local neighborhood network. Based on the correlation between the gene expression vectors of the cell under test and the neighboring cells, an expression similarity factor is calculated; The product of the distance decay factor and the expression similarity factor is determined as the network decay weight.
[0015] In an optional implementation, the formula for calculating the basal inflammatory activity of the test cells is: ; Wherein, F(e) represents the basic inflammatory activity of the test cell e; τ represents the set of pathways related to the inflammatory response; T represents any pathway in the set of pathways; ω T The weights representing pathway T; GSVA T (e) represents the evaluation value for assessing the activity of pathway T.
[0016] In an optional implementation, the expression for calculating the space-expression coupling potential is: ; Among them, SECP iF(e) represents the spatial-expression coupling potential of cell i under test; i ) represents the basic inflammatory activity of test cell i; k represents the number of surrounding cells in the local neighborhood network of test cell i; w ij The network decay weights represent the relationship between cell i and its surrounding cells j. α represents the expression similarity factor, and the following conditions are met: The expression correlation between cell i and surrounding cell j is represented by ).
[0017] In an optional implementation, the network attenuation weight w ij The calculation expression is: ; Where (x,y) i The spatial coordinates (x, y) represent the coordinates of cell i to be tested. j The spatial coordinates of the surrounding cell j; d represents the square of the Euclidean distance between the two coordinates; d represents the minimum distance between all cells in the spatial transcriptome data; exp represents an exponential function with the natural constant as the base.
[0018] In an optional implementation, the step of performing a spatial smoothing transformation on the spatial-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the tested cells includes: Calculate the vector of representations after spatial smoothing transformation: ; in, The vector represents the expression level of a specific inflammation-associated gene g after spatial smoothing transformation; SECP represents the vector obtained after normalizing the spatial-expression coupling potential values of each cell; e g This represents the initial expression vector of the gene g; Based on the expression vector, the spatial inflammation score coefficient is calculated: ; Among them, IID i The spatial inflammation score coefficient represents cell i; m represents the total number of inflammation-related genes; k represents the kth inflammation-related gene. The expression level of the k-th inflammation-associated gene in cell i after spatial smoothing transformation is represented by the expression level vector. Extract from.
[0019] In an optional implementation, determining the inflammatory hotspots in the tissue under test based on the spatial enrichment characteristics of the cell clusters includes: For each of the cell clusters, calculate its local background density and the density of outliers within the cluster. The local enrichment score is calculated based on the local background density, the density of outliers within the cluster, and a preset smoothing constant. The cell clusters with local enrichment scores greater than a preset enrichment threshold are identified as the inflammatory hotspots.
[0020] In an optional implementation, the formula for calculating the density of outliers within a cluster is: ; ; Where, λ sig.k Representing cell cluster C k The density of outliers within the cluster; |C k | Represents cell cluster C k The number of abnormal cells contained within; A k Representing cell cluster C k The area of the hull; Area(ConcaveHull(·)) represents the function to solve for the area of the hull of a point set in space.
[0021] In an optional implementation, the method for calculating the local background density includes: determining the cell cluster C k The distance from the center of the cell to its farthest abnormal cell is used as the actual radius R. k With the actual radius R k The defined circular region is designated as the local background region; the expression for calculating the local background density is as follows: ; Where, λ bg.k Representing cell cluster C k The local background density; N abnormal,k The value represents the total number of abnormal cells within the local background region; π represents pi.
[0022] In an optional implementation, the actual radius R k The calculation expression is: ; Among them, R k Represents cell cluster C k The actual radius; The spatial coordinates of any abnormal cell in the array, where x is the abscissa and y is the ordinate; The center x-coordinate; The central ordinate.
[0023] In an optional implementation, the expression for calculating the center abscissa is: ; and / or, The expression for calculating the central ordinate is: ; in, Representing cell cluster C k The calculated x-coordinate of the center is output. Representing cell cluster C k The calculated center ordinate is output. Representing cell cluster C k The total number of abnormal cells contained within; Represents cell cluster C k The summation of the x or y coordinates of all abnormal cells is performed.
[0024] In an optional implementation, the expression for calculating the local enrichment score is: ; Among them, LDES k λ represents the local enrichment fraction of cell cluster k; sig.k The density of outliers within cell cluster k; λ bg.k γ represents the local background density of cell cluster k; γ represents the preset smoothing constant, and γ is a constant greater than 0.
[0025] In an optional implementation, the clustering of the abnormal cells to obtain at least one cell cluster includes: The abnormal cells are clustered using a space-attribute weighted community clustering algorithm, and the expression for the resulting clustering set is as follows: ; Wherein, P represents the set of points of the abnormal cells; C k The cluster represents the k-th cell cluster obtained by clustering; K represents the total number of cell clusters containing abnormal cells that are greater than or equal to the preset minimum number of points; N represents the set of noise points that have not formed cell clusters.
[0026] In a second aspect, the present invention provides an inflammatory hotspot identification device, comprising: The acquisition module is used to acquire spatial transcriptome data of the tissue to be tested; wherein, the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue to be tested and the expression levels of inflammation-related genes; The calculation module is used to calculate the spatial inflammation score coefficient of each of the test cells based on the spatial coordinates and the expression level; The screening module is used to construct an abnormal baseline model based on data from a healthy control group, and to screen out abnormal cells from the cells to be tested by combining the abnormal baseline model and the spatial inflammation score coefficient. A clustering module is used to cluster the abnormal cells to obtain at least one cell cluster; The identification module is used to determine the inflammatory hotspots in the tissue to be tested based on the spatial enrichment characteristics of the cell clusters.
[0027] Thirdly, the present invention provides a computer device including a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the inflammatory hotspot identification method described in any of the foregoing embodiments.
[0028] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed on a processor, implements the inflammatory hotspot identification method according to any one of the foregoing embodiments.
[0029] This application provides a method, apparatus, device, and storage medium for identifying inflammatory hotspots, relating to the field of spatial transcriptomics. The method for identifying inflammatory hotspots acquires spatial transcriptomic data containing spatial coordinates and the expression levels of inflammation-related genes, and calculates the spatial inflammation score coefficient of the cells under test based on these two data points, overcoming the limitations of traditional analysis methods that isolate gene expression levels. This approach deeply couples the molecular expression characteristics of cells with their physical spatial neighborhood, enabling a more realistic and accurate reflection of the inflammatory microenvironment state of cells in tissue sections, effectively solving the problem of inaccurate assessment of the actual degree of cellular inflammation.
[0030] Based on the precise quantification of the inflammatory state of individual cells, an abnormal baseline model was constructed by introducing data from a healthy control group, establishing an objective and standardized reference benchmark for distinguishing normal cells from abnormally inflammatory cells. Combining this abnormal baseline model with spatial inflammation score coefficients for screening effectively eliminates background noise and interference from the heterogeneity of normal cells in complex biological tissues. This allows for high-precision screening of truly abnormal cells in complex tissue environments, significantly improving the reliability of the identification results.
[0031] Furthermore, the screened abnormal cells were clustered to obtain cell clusters, and the final inflammatory hotspots were determined based on the spatial enrichment characteristics of the cell clusters, enabling a quantitative assessment of the density of local abnormal cell aggregation. This process, from precise screening at the single-cell level to enrichment and demarcation at the macroscopic spatial level, overcomes the shortcomings of traditional methods in defining the vague boundaries when identifying hotspot areas. It can clearly and accurately define the spatial boundaries of high-incidence areas of inflammation, providing a more reliable analytical basis for the study of the pathological mechanisms and clinical diagnosis of complex tissues (such as alcoholic liver injury). Attached Figure Description
[0032] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the hardware operating environment involved in an embodiment of the inflammation hotspot identification method of the present invention; Figure 2 This is a flowchart illustrating Embodiment 1 of the inflammation hotspot identification method of the present invention; Figure 3 This is a detailed flowchart of step S300 in Embodiment 2 of the inflammation hotspot identification method of the present invention; Figure 4 This is a detailed flowchart of step S330 in Embodiment 2 of the inflammation hotspot identification method of the present invention; Figure 5 This is a detailed flowchart of step S310 in Embodiment 2 of the inflammation hotspot identification method of the present invention; Figure 6 This is a detailed flowchart of step S200 in Embodiment 3 of the inflammation hotspot identification method of the present invention; Figure 7 This is a detailed flowchart of step S220 in Embodiment 3 of the inflammation hotspot identification method of the present invention; Figure 8 This is a detailed flowchart of step S222 in Embodiment 3 of the inflammation hotspot identification method of the present invention; Figure 9 This is a detailed flowchart of step S500 in Embodiment 4 of the inflammation hotspot identification method of the present invention; Figure 10This is a flowchart of the identification method for agglomeration-type samples in Experiment Example 1 of this application (A is the basic scoring result, B is the spatial inflammation score coefficient distribution result, C is the abnormal cell screening result based on GEV abnormality detection, and D is the final inflammation hotspot delineation result). Figure 11 This is a reference image of the original H&E pathological staining and artificial annotation of the gold standard for the aggregated sample in Experiment Example 1 of this application; Figure 12 This is a diagram showing the results of identifying inflammatory hotspots using the method of this invention in Experiment Example 1 of this application. Figure 13 This is a diagram showing the results of identifying inflammatory hotspots using the comparative scheme in Experiment Example 1 of this application. Figure 14 This is a hotspot region identification map based on the original expression level of inflammation-related genes in Example 1 of this application; Figure 15 This is a flowchart of the identification method for diffusely distributed samples in Experiment Example 2 of this application (A is the basic scoring result, B is the spatial inflammation score coefficient distribution result, C is the abnormal cell screening result based on GEV abnormality detection, and D is the final inflammatory hotspot region delineation result). Figure 16 This is the gold standard reference image in Experimental Example 2 of this application (showing the original H&E (hematoxylin-eosin) pathological staining image of the diffuse sample, which also includes magnified indications of local areas. This image objectively reflects the harsh section state of "blurred boundaries and diffuse lesions" caused by extensive tissue damage, serving as the "gold standard" reference for evaluating the target area localization ability of various algorithms in complex diffuse backgrounds); Figure 17 The image shows the identification results of the present invention in Experiment Example 2 of this application (demonstrating the hotspot delineation results output by the inflammation hotspot area identification method provided in the embodiments of the present invention). Figure 15 It can be seen that even in complex tissue contexts with extremely high overall inflammatory baselines and diffuse lesions, the proposed solution can still resist interference from globally diffuse signals and accurately locate and purify several localized core areas of severe inflammatory infiltration. This fully verifies the extremely high recognition accuracy and spatial noise robustness of the proposed solution in complex microenvironments. Figure 18 This is the comparison recognition result from Experiment Example 2 of this application (showing the recognition result output after parallel processing using the comparison algorithm (i.e., the existing Hotspot+Seurat+OPTICS integrated algorithm)). Figure 15 and Figure 16In contrast, this comparative model, lacking adaptive background density calibration and extreme anomaly screening mechanisms, resulted in a large area of red coverage across the entire image (i.e., a severe "starry sky" false positive phenomenon). This result failed to effectively distinguish between core severe lesions and minor background fluctuations, completely losing its clinical significance in guiding local pathological diagnosis. Figure 19 This is a schematic diagram of the module connections of the inflammation hotspot identification device of the present invention. Detailed Implementation
[0034] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0035] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0036] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0037] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0038] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0039] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0040] like Figure 1 The diagram shown is a structural schematic of the hardware operating environment of the terminal involved in an embodiment of the present invention.
[0041] The inflammation hotspot identification system of this invention can be a PC, or a mobile terminal device such as a smartphone, tablet, or laptop. The system may include: a processor 1001 (e.g., a CPU), a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to establish communication between these components. The user interface 1003 may include a display screen, an input unit such as a keyboard, or a remote control; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (e.g., a Wi-Fi interface). The memory 1005 may be a high-speed RAM or a stable memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the processor 1001. Optionally, the inflammation hotspot identification system may also include RF (Radio Frequency) circuitry, audio circuitry, a Wi-Fi module, etc. In addition, the inflammation hotspot identification system can also be equipped with other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, which will not be elaborated here.
[0042] Those skilled in the art will understand that Figure 1 The inflammation hotspot identification system shown is not intended to limit it and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a data interface control program, a network connection program, and an inflammation hotspot identification program.
[0043] Example 1 refer to Figure 2 This application provides a method for identifying inflammatory hotspots in embodiment 1. This method is mainly used to accurately identify inflammatory lesions in biological tissue sections. The execution subject of this method can be an electronic device or a computer system with data processing capabilities. The inflammatory hotspot identification method includes: Step S100: Obtain spatial transcriptome data of the tissue to be tested; wherein, the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue to be tested and the expression levels of inflammation-related genes.
[0044] In practice, this step involves acquiring sequencing data from the tissue sections to be tested using a spatiotemporal omics sequencing platform. This data not only captures the gene expression characteristics related to inflammatory responses within cells but also simultaneously records the two-dimensional or three-dimensional coordinates of each cell on the physical section.
[0045] As an optional preprocessing method, after acquiring the data, the expression matrix in the spatial transcriptome data can be subjected to central logarithmic ratio transformation and standardization in sequence to effectively eliminate sequencing errors and batch effects between different samples or batches.
[0046] Step S200: Calculate the spatial inflammation score coefficient for each of the tested cells based on the spatial coordinates and the expression level.
[0047] It should be noted that inflammatory responses in organisms typically spread within the local microenvironment. To objectively reflect the true degree of cellular inflammation, this step deeply couples the molecular expression of cells with their spatial physical neighborhood.
[0048] In a preferred embodiment, a local neighborhood network can be constructed based on the spatial coordinates of each cell to be tested. Taking into account both the decay pattern of physical distance between cells and the similarity of gene expression, neighboring cells are assigned corresponding network decay weights. The basal inflammatory activity of the cell to be tested is mathematically coupled with the network decay weights and spatially smoothed to calculate a spatial inflammation score coefficient that can smooth sequencing noise and reflect the true state of the tissue microenvironment.
[0049] Step S300: Construct an abnormal baseline model based on data from the healthy control group, and combine the abnormal baseline model with the spatial inflammation score coefficient to screen out abnormal cells from the cells to be tested.
[0050] To accurately locate diseased cells within the complex natural heterogeneity of tissues, this step introduces an objective statistical benchmark.
[0051] Specifically, sequencing data from healthy or non-inflammatory tissue samples are extracted, and an abnormal baseline model is constructed using the score distribution characteristics of healthy cells (e.g., by fitting a generalized extreme value distribution model by extracting extreme value samples). The spatial inflammation score coefficients of each test cell obtained above are substituted into this baseline model for statistical significance testing. By comparing the significance test results with a preset threshold, cells belonging to normal physiological fluctuations are objectively eliminated, thereby accurately screening test cells with significant abnormal inflammatory characteristics.
[0052] Step S400: Cluster the abnormal cells to obtain at least one cell cluster.
[0053] In the slice space, isolated abnormal cells with discrete distribution often do not have macroscopic pathological significance. Therefore, unsupervised clustering algorithms (preferably community clustering algorithms based on both spatial and attribute weighting, such as the Louvain algorithm) are used to cluster and partition the selected abnormal cells.
[0054] This process can group physically adjacent and similar abnormal cells into the same set, forming cell clusters with spatially continuous morphology, effectively filtering out spatially random noise points.
[0055] Step S500: Based on the spatial enrichment characteristics of the cell clusters, determine the inflammatory hotspots in the tissue to be tested.
[0056] For each cell cluster obtained from clustering, its density and significance need to be further quantified. In one embodiment, the density of intra-cluster anomalies within each cell cluster (e.g., density calculation based on the spatial concave area of the cell cluster) and the local background density of the surrounding area of the cell cluster are calculated. A local enrichment score is calculated based on the comparison between the intra-cluster anomaly density and the local background density. Cell clusters with local enrichment scores greater than a preset enrichment threshold are identified as real lesion regions with extremely significant spatial aggregation, thereby accurately identifying and outputting the inflammatory hotspot areas in the tissue under test. This step achieves high-precision delineation of the physical boundaries of the hotspot areas.
[0057] In summary, this method, by deeply coupling the spatial coordinates of cells with the expression levels of inflammation-related genes to calculate the spatial inflammation score coefficient, accurately reflects the true inflammatory microenvironment of cells, overcoming the limitations of traditional methods that rely on isolated gene expression analysis. Simultaneously, by introducing an abnormal baseline model based on healthy control group data, an objective and standardized reference scale is established, effectively eliminating interference from tissue background noise and normal cell heterogeneity, achieving high-precision screening of truly abnormal cells. Furthermore, by clustering abnormal cells and quantifying their spatial enrichment characteristics, the physical boundaries of inflammatory hotspot areas are precisely defined. Overall, this method achieves comprehensive optimization from microscopic single-cell state assessment to macroscopic hotspot region delineation, effectively addressing the shortcomings of existing technologies such as inaccurate assessment and ambiguous boundaries, and significantly improving the accuracy and reliability of identifying inflammatory hotspot areas in complex tissues.
[0058] Example 2 refer to Figure 3Based on the above embodiments, to accurately isolate truly diseased inflammatory cells from complex tissue heterogeneity and natural background noise in gene expression, this embodiment further provides a specific implementation method for constructing an abnormal baseline model based on healthy control group data, and combining the abnormal baseline model with the spatial inflammation score coefficient to screen abnormal cells from the test cells. Step S300, constructing an abnormal baseline model based on healthy control group data, and combining the abnormal baseline model with the spatial inflammation score coefficient to screen abnormal cells from the test cells, includes: Step S310: Obtain the spatial inflammation score coefficient of cells in the healthy control group and extract extreme value sample data according to cell type.
[0059] It should be noted that in healthy biological tissues, the gene expression of cells naturally fluctuates to a certain extent. In order to establish an accurate judgment benchmark, this invention introduces the extreme value theory.
[0060] Specifically, instead of modeling the massive amount of low-expression background data in the healthy control group, the data is grouped according to cell clusters or types. For each specific cell type, a subset of extreme values with high spatial inflammation scores is extracted (e.g., the top percentage of samples after sorting the scores in descending order) as extreme value sample data.
[0061] In biological data such as spatial transcriptomics, gene expression or inflammation scores in healthy cells typically do not follow a standard normal distribution (Gaussian distribution), but rather exhibit a significant long-tail characteristic in the right tail (i.e., the high-expression side). If a conventional normal distribution model is fitted using full-scale healthy data, underfitting of the tail is highly likely, leading to an extremely high false positive rate. Therefore, this embodiment introduces extreme value theory. The "extreme value sample data" refers to the set of data extracted from the healthy control group whose scores fall within the top distribution range, representing the upper limit of extreme inflammatory fluctuations that can be tolerated under healthy physiological conditions.
[0062] This approach not only significantly reduces redundant computations, but also precisely focuses the model's attention on the maximum tolerable upper limit of inflammatory fluctuations under healthy physiological conditions.
[0063] Step S320: Fit a generalized extreme value distribution model using the extreme value sample data.
[0064] In this step, the extreme value sample data is used to fit a Generalized Extreme Value (GEV) model. The "Generalized Extreme Value (GEV) model" is a continuous probability distribution model specifically fitted to the aforementioned extreme value sample data. It can capture the probability density characteristics of the long-tail region of biological data with extremely high accuracy, providing a probabilistic benchmark that best fits biological facts for distinguishing between normal fluctuations and pathological mutations.
[0065] In practical implementation, for the extreme value sample data of each extracted cell type, this embodiment preferably uses the maximum likelihood estimation (MLE) method to fit and solve the location parameter μ of the corresponding generalized extreme value distribution model. g Scale parameter σ g and shape parameter ξ g The MLE algorithm can guarantee the asymptotic unbiasedness and effectiveness of parameter estimation in statistics, thereby constructing a distribution model that can accurately describe the fluctuation pattern of high health scores.
[0066] Because high-expression score data in biological transcriptomics often exhibit significant long-tail characteristics, traditional normal distributions are insufficient for accurate description. Therefore, this embodiment fits and estimates the location, scale, and shape parameters of the corresponding model for each extracted cell type's extreme value sample data, thereby constructing a generalized extreme value distribution model that can accurately describe the fluctuation patterns of high-expression scores in healthy cells. This model provides a rigorous and biologically relevant probability density benchmark for subsequent anomaly detection.
[0067] Step S330: Substitute the spatial inflammation score coefficient of the cell under test into the generalized extreme value distribution model to calculate the weighted marginal p-value of the cell under test.
[0068] Considering the potential cell type heterogeneity or sequencing spot mixing in spatial transcriptome sequencing slices, the microenvironment composition must be comprehensively considered when evaluating a single cell.
[0069] It should be noted that in spatial transcriptome sequencing slices, due to physical limitations in sequencing resolution or the density of the tissue itself, the signal of a single cell or sequencing spot is often a mixture of signals from multiple cell types within the local microenvironment. Conventional single-type p-value tests cannot handle this spatial heterogeneity. The "weighted marginal p-value" mentioned in this embodiment refers to a comprehensive statistical index obtained by considering the mixing ratio (i.e., spatial weight) of various cell types within the spatial neighborhood of the cell being tested, and integrating the conditional significance probabilities under various cell type assumptions. This index effectively eliminates the uncertainty in cell type identification and the technical noise caused by spatial mixing and overlap.
[0070] In the specific calculation, the probability distribution or proportional weight of each cell type within the spatial neighborhood of the cell to be tested is first determined. Next, the spatial inflammation score coefficient of the cell to be tested is substituted into the cumulative distribution function of the generalized extreme value distribution model specific to each cell type, which is related to its neighborhood, to calculate the conditional p-values under different cell type assumptions. Finally, based on the aforementioned proportional weights, all conditional p-values are weighted and summed to obtain a weighted marginal p-value that takes into account spatial heterogeneity.
[0071] The weighted marginal p-value objectively reflects the probability that the tested cell would have a very high inflammation score if it were in a healthy state.
[0072] Step S340: The test cells whose weighted marginal p-value is less than the preset significance threshold are identified as abnormal cells.
[0073] The aforementioned preset significance threshold can serve as a critical probability standard for rejecting the null hypothesis (i.e., assuming the cell is healthy) in statistical hypothesis testing. In practical applications, this threshold can be flexibly configured according to the specific tissue / organ type, sequencing depth, and tolerance for false positives / false negatives. For example, commonly used significance thresholds in this field can be set to 0.05, 0.01, or 0.001, etc. This embodiment does not impose an absolute numerical limit on this.
[0074] In this step, cells whose weighted marginal p-value is less than a preset significance threshold are identified as abnormal cells. Based on the principle of statistical hypothesis testing, a preset significance threshold (e.g., 0.05 or 0.01) can be set.
[0075] When the weighted marginal p-value of a test cell is less than the preset significance threshold, the null hypothesis that the cell is within the range of healthy physiological fluctuations is statistically rejected. This allows the test cell to be accurately identified and extracted as an abnormal cell with significant pathological characteristics. Through this dynamic adaptive screening based on rigorous statistical distribution, misjudgments caused by using a "one-size-fits-all" fixed numerical threshold are effectively avoided, greatly improving the accuracy and robustness of abnormal inflammatory cell identification.
[0076] Furthermore, considering that spatial transcriptome data involves multiple hypothesis testing of a massive number of cells, in order to strictly control the false positive rate, after obtaining the weighted marginal p-values of each test cell, the Benjamini-Hochberg (BH) method can preferably be used to perform multiple hypothesis testing correction on the weighted marginal p-values. The corrected p-values (p-adj) are then compared with the preset significance threshold (e.g., controlling FDR to 5%, i.e., the threshold is 0.05), and test cells with corrected p-values less than this threshold are identified as high-confidence abnormal cells.
[0077] Building upon the abnormal baseline model based on a healthy control group constructed in the above embodiments, and considering the inherent spatial heterogeneity in spatial transcriptome sequencing slices (e.g., a single sequencing site may cover multiple cell types, or overlapping cell physical boundaries may lead to signal mixing), this embodiment further provides a specific implementation method for calculating the weighted marginal p-value of the cells under test, to achieve a highly robust and significant test for abnormally inflammatory cells. Specifically, refer to... Figure 4 Step S330 involves substituting the spatial inflammation score coefficient of the test cell into the generalized extreme value distribution model to calculate the weighted marginal p-value of the test cell, including: Step S331: Determine the proportional weight of each cell type within the spatial neighborhood of the cell to be tested.
[0078] It should be noted that in the real biological tissue microenvironment, cells do not exist in isolation. The molecular signals of the cell being tested are often influenced by the combined effects of the cell population within its local spatial neighborhood. Therefore, to accurately assess the cell being tested, it is necessary to first analyze the cellular distribution within its spatial neighborhood.
[0079] Specifically, by analyzing the composition of the cell under test and the cell population within its defined neighborhood, the mixing ratio or probability of each known cell type (e.g., epithelial cells, macrophages, T cells, etc.) in the current microenvironment can be calculated. This yields a set of proportional weights corresponding to each cell type, which quantifies the heterogeneous structure of the cell's microenvironment.
[0080] As a specific engineering implementation method, the analysis of the distribution status of each cell type within the aforementioned spatial neighborhood and the calculation of their proportion weights can be achieved using deconvolution algorithms in spatial transcriptomics. For example, those skilled in the art can use open-source bioinformatics analysis tools such as the RCTD algorithm or the SPOTlight algorithm to decompose the mixed gene expression profile of spatial sequencing sites (Spots) into combinations of single cell types, thereby accurately obtaining the mixing ratio or probability of occurrence of each cell type under the current spatial coordinates.
[0081] Step S332: For each cell type, substitute the spatial inflammation score coefficient of the cell under test into the cumulative distribution function of its corresponding generalized extreme value distribution model to calculate the conditional p-value.
[0082] To eliminate baseline bias caused by differences in natural expression levels between different cell types, this step employs a "genotyping independence test" strategy.
[0083] For each specific cell type within the neighborhood identified in the previous steps, a generalized extreme value distribution (GEV) model fitted using extreme value samples of that cell type from a healthy control group is retrieved. The spatial inflammation score coefficient of the test cells is substituted into the cumulative distribution function (CDF) of this specific GEV model. By calculating the probability value of this score in the right-tailed distribution of the corresponding model (e.g., calculating 1 minus the CDF value corresponding to this score), a conditional p-value is obtained under the assumption that the test cells belong entirely to that specific cell type. This conditional p-value rigorously reflects the statistical rarity of this high inflammation score under a single-type benchmark.
[0084] Step S333: Based on the proportional weights corresponding to each cell type, all the condition p-values are weighted and summed to obtain the weighted marginal p-value of the cell to be tested.
[0085] In this step, based on the principle of the law of total probability, the conditional p-values of various cell types calculated in the previous step are multiplied by their corresponding spatial proportion weights, and then the product results of all types are summed.
[0086] This weighted summation method deeply mathematically integrates the anomaly probabilities under various cell type assumptions with the actual local spatial physical composition. The resulting weighted marginal p-value not only accurately measures the degree to which the tested cells deviate from the healthy baseline, but also effectively absorbs and smooths out the technical noise caused by cell type mixing and identification uncertainties commonly found in spatial transcriptome data. This provides an extremely robust statistical basis for high-precision, low-false-positive screening of abnormal cells.
[0087] Based on the abnormal cell screening scheme provided in the above embodiments, in order to enable those skilled in the art to more clearly understand the deep mathematical logic of statistical testing and abnormal baseline construction in this invention, this embodiment further provides specific mathematical expressions for calculating the weighted marginal p-value and cumulative distribution function, and provides detailed explanations of the core statistical concepts and parameter mechanisms involved.
[0088] It should be noted beforehand that in high-throughput biological data analysis such as spatial transcriptomics, gene expression or inflammation scores in healthy cells typically do not follow a standard normal distribution (Gaussian distribution), but rather exhibit a significant "long tail" characteristic on the high-expression side (right tail). If a conventional normal distribution is used for testing, it is highly susceptible to inadequate fitting of the tail probability, leading to an extremely high false positive rate. Therefore, this embodiment creatively introduces the Generalized Extreme Value Distribution (GEV) model from extreme value theory, specifically for accurate probabilistic modeling of long-tail fluctuations.
[0089] Specifically, the cumulative distribution function value calculated by substituting the spatial inflammation score coefficient of the test cells into the generalized extreme value distribution model corresponding to the cell cluster or type has the following mathematical expression: (Formula 1); Where Gg(x) represents the cumulative distribution function value of the generalized extreme value distribution model corresponding to cell cluster or type g, which represents the probability that the inflammation score of cell type g is less than or equal to the given input value under healthy physiological conditions; x represents the spatial inflammation score coefficient of the input; μ g ξ represents the location parameter of the generalized extreme value distribution model corresponding to cell cluster or type g, which determines the central anchor point of the distribution extreme value population; g The shape parameter representing the cell cluster or type g; σ g The scale parameter represents the cell cluster or type g, which quantifies the inherent fluctuation and dispersion of extreme scores in a healthy state; exp represents an exponential function with the natural constant as its base.
[0090] Furthermore, after calculating the cumulative distribution function value for a single type, considering the potential physical overlap or cell identity mixing (i.e., spatial heterogeneity) of spatial transcriptome sequencing sites, this embodiment mathematically fuses the weighted marginal p-values of the test cells based on the law of total probability. The expression for calculating the weighted marginal p-values of the test cells is as follows: (Formula 2); Where, p j π represents the weighted marginal p-value of cell j, which serves as the statistical basis for determining whether the cell significantly deviates from a healthy state; G represents the set of cell clusters or types, covering all known cell categories present in the sequencing data; g represents a specific cell cluster or type within set G; π jg The probability weight representing the cell group or type g to which the test cell j belongs objectively quantifies the heterogeneity mixing ratio of the cell microenvironment, ensuring the model's tolerance for ambiguous regions in cell identification; IID j G represents the spatial inflammation score coefficient of cell j being tested; g (IID j The value represents the cumulative distribution function calculated by substituting the spatial inflammation score coefficient of the cell j under test into the generalized extreme value distribution model corresponding to the cell cluster or type g.
[0091] Through the synergistic operation of the two sets of rigorous mathematical formulas mentioned above, this embodiment effectively avoids the crudeness of the traditional fixed threshold method, and perfectly integrates the physical weight of spatial composition with the probabilistic deduction of extreme value theory, thereby providing an extremely solid and interference-resistant underlying computational support for accurately identifying real lesion cells.
[0092] In some implementations, reference Figure 5 Step S310, extracting extreme value sample data according to cell type, includes: Step S311: For each cell type in the healthy control group, sort the spatial inflammation score coefficients of each cell contained therein.
[0093] Given the complex cellular heterogeneity within biological tissues, different cell types (such as structural cells and various immune cells) exhibit naturally significant differences in the background expression activity of inflammation-related genes under healthy physiological conditions. To eliminate this systematic bias caused by cell type differences, this embodiment does not perform mixed processing of global health data, but instead adopts a "subtype-based independent processing" strategy.
[0094] Specifically, based on cell type labels in the spatial transcriptome data, cells from the healthy control group were divided into corresponding cell type subsets. Then, within each specific cell type subset, the spatial inflammation score coefficients of all normal cells within it were sequentially sorted according to their numerical values, preferably in descending order of score. This grouping and sorting operation achieved standardized alignment of baseline features for cells of the same type.
[0095] Step S312: Extract the spatial inflammation score coefficient of the cells ranked in the top preset percentage, and use it as the extreme value sample data corresponding to that cell type.
[0096] After completing the descending order within each cell type, based on the tail analysis principle of extreme value theory, high-scoring data is directly extracted from the top of each sequence.
[0097] The aforementioned preset percentage refers to the relative threshold used to determine the extreme value cutoff range after the scores are sorted. This percentage can be flexibly set according to the overall scale of the sequencing data, the richness of cell types, and the desired sensitivity of the statistical test. In essence, it is a quantitative expression of the threshold selection in extreme value theory in engineering.
[0098] Specifically, a preset percentage is set (for example, in practice, this can be flexibly configured as the top 1%, top 5%, or top 10%, etc., depending on the sample size; this embodiment does not impose an absolute limitation on its specific value), and the product of the total number of cells included in the specific cell type and the preset percentage is calculated to determine the number of extreme value samples to be extracted. Then, the spatial inflammation score coefficients of cells ranked within this number range in the descending sequence of the cell type are extracted and compiled into an independent dataset, which is used as the extreme value sample data specific to that cell type.
[0099] By using the above-mentioned processing mechanism of grouping and sorting by cell type and truncating the tail extreme values proportionally, not only is the interference of inherent cell heterogeneity effectively eliminated, but also redundant low-scoring background data of health is significantly filtered out, providing the highest quality and most boundary-representation-capable training samples for subsequent independent fitting of generalized extreme value distribution models for each cell type.
[0100] Based on the extreme value sample extraction scheme provided in the above embodiments, in order to enable those skilled in the art to more clearly understand the underlying data structure and logical filtering mechanism of this application when processing large-scale spatial transcriptome data in a computer system, this embodiment further provides a standardized recording format of healthy control group data and extraction expressions for specific data sets.
[0101] The healthy control group data includes n c The record information of each cell; the record information of the i-th cell is represented as: (Formula 3); It is important to note that in spatial transcriptome sequencing analysis, a single cell carries multiple heterogeneous attributes, including source batch, physical identity, computational status, and biological classification. Encapsulating these heterogeneous attributes within a single mathematical tuple ensures that the core information of the cell remains intact and uninterrupted during subsequent complex dimensionality reduction, filtering, and statistical tests. This forms the data foundation for building a stable and efficient bioinformatics analysis pipeline.
[0102] In the above tuple record format, the specific meanings of each parameter are as follows: Wherein, Represents the sample name (or batch identifier) to which the i-th cell belongs, used for tracing sample information; The cell name representing the i-th cell serves as a unique identifier; The spatial inflammation score coefficient representing the i-th cell is the core numerical basis for subsequent anomaly significance testing; The i-th cell belongs to a specific group or type, serving as the classification criterion for heterogeneity analysis; where i is a positive integer, and 1 ≤ i ≤ n. c n c This represents the total number of normal cells included in the healthy control group data.
[0103] Furthermore, to eliminate the natural differences in background expression among different cell types, it is necessary to construct independent assessment baselines for specific cell types. Therefore, for each cell type g in the healthy control group, the system needs to perform a conditional traversal query on the above structured data to extract the expression for the extreme value sample data set corresponding to that cell type: (Formula 4); Here, Ctrlg represents the initial candidate set of extreme value sample data corresponding to cell type g.
[0104] Regarding the initial candidate set mentioned above, it should be noted that before performing extreme value theory modeling, it is necessary to clean and purify a homogeneous subset of data that meets specific conditions from a globally heterogeneous data pool. This initial candidate set is the original set of scores of the same type that has been filtered by cell classification labels but has not yet been numerically sorted or truncated according to extreme value proportions.
[0105] The physical execution logic of the formula is as follows: The computer system traverses all cell records in the healthy control group, selects cell records whose cluster or type label is strictly equal to the target type g through condition matching, and extracts the spatial inflammation score coefficient of these cells that meet the conditions separately, and combines them into a homogeneous numerical set.
[0106] Through the structured tuple definition and rigorous set extraction expression described above, this embodiment transforms the complex biological information classification process into a data retrieval and matrix slicing logic that can be efficiently executed by a computer system. This not only ensures the consistency of massive cell data during the transfer process, but also provides a pure underlying computational data source for subsequent high-precision extreme value sample sorting and statistical distribution model fitting.
[0107] Example 3 refer to Figure 6 Based on the above embodiments, to overcome the limitations of traditional gene scoring methods in isolated analysis of single-cell data, this embodiment further provides a specific implementation method for calculating the spatial inflammation score coefficient of each of the tested cells. Step S200, based on the spatial coordinates and the expression level, calculates the spatial inflammation score coefficient of each of the tested cells, including: Step S210: For each of the cells to be tested, determine a local neighborhood network based on its spatial coordinates.
[0108] It should be noted that in the biological microenvironment, the transmission of inflammatory signals between cells usually depends on physical proximity. The local neighborhood network in this embodiment refers to a graph structure constructed in computer memory, which takes a cell under test as the central node and connects several cells that are physically close to it as neighbor nodes through edges, thereby mathematically mapping the real intercellular physical communication microenvironment.
[0109] In this step, a local neighborhood network is determined for each test cell based on its spatial coordinates. After acquiring spatial transcriptome data, the system extracts the physical spatial coordinates (two-dimensional or three-dimensional coordinates) of each test cell on the tissue slice. Taking any test cell as the central target node, the system uses a spatial distance calculation algorithm (e.g., calculating the Euclidean distance matrix of a fully connected component) to find its spatially adjacent cells.
[0110] In practice, a neighborhood search algorithm based on a fixed Euclidean distance radius can be used, taking all cells within that distance threshold as its neighbors; or the K nearest neighbor (KNN) algorithm can be used, selecting the K nearest cells as neighbors. Through these methods, the connection topology between the central cell to be tested and its neighboring cells is established, thereby constructing a unique local neighborhood network for each cell on the slice.
[0111] Step S220: Calculate the spatial-expression coupling potential value by combining the expression levels of inflammation-related genes in the test cells and the local neighborhood network.
[0112] It should be noted that single sequencing expression levels are easily affected by technical noise (such as the dropout effect). This embodiment introduces the concept of "potential energy" from physics to construct a comprehensive metric that integrates biomolecular expression and spatial topology. The Space-Expression Coupling Potential (SECP) value depends not only on the activity of inflammatory genes in the cell being tested but also on the synergistic modulation of the states of other cells in its neighborhood network, objectively reflecting the comprehensive potential of the cell to cause inflammatory lesions in the current microenvironment.
[0113] In this step, the spatial-expression coupling potential is calculated by combining the expression levels of inflammation-related genes in the test cell and the local neighborhood network. After constructing the spatial topology, the algorithm system further integrates transcriptome expression data. The system extracts the expression matrix of inflammation-related genes contained in the test cell itself and calculates its basic inflammatory activity indicators. Simultaneously, the system traverses the aforementioned constructed local neighborhood network and extracts the attribute features of each surrounding cell within the network.
[0114] In this process, the algorithm comprehensively evaluates the basic inflammatory activity of the central test cell and introduces a network topology weighting mechanism (e.g., assigning different network attenuation weights based on the physical distance attenuation effect between the central cell and its neighboring cells, as well as the correlation and similarity of gene expression profiles). The algorithm mathematically couples and aggregates the expression activity of the test cell with the synergistic weights transmitted through its neighborhood network to derive and calculate the spatial-expression coupling potential value of the test cell. This potential value effectively compensates for signal loss caused by undetected gene expression in single-cell sequencing, achieving efficient fusion of single-point features and neighborhood microenvironment features.
[0115] Step S230: Perform a spatial smoothing transformation on the spatial-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the test cells.
[0116] It should be noted that, because the diffusion of cytokines and chemokines in biological tissues exhibits a continuous concentration gradient, the biological states of adjacent cells should have a certain continuity. Spatial smoothing transformation is a signal filtering technique performed on graph networks, designed to eliminate isolated sequencing noise and restore the true spatial diffusion gradient of inflammatory signals.
[0117] In this step, the spatial-expression coupling potential energy value is spatially smoothed to obtain the spatial inflammation score coefficient for each of the tested cells. To further simulate the real smooth concentration gradient of inflammatory molecules within biological tissues and suppress high-frequency random noise during sequencing, the algorithm performs spatial dimension smoothing filtering on the initially calculated spatial-expression coupling potential energy value.
[0118] In practical implementation, graph signal processing techniques can be used to perform weighted moving average processing, random walk information transfer, or graph convolution operations on the graph structure composed of local neighborhood networks. By performing multiple rounds of smoothing feature transfer and numerical fusion between the potential energy value of the central cell and the potential energy values of its surrounding neighboring cells, isolated extreme spikes that violate the continuous physiological laws on the slice are effectively weakened. After spatial smoothing transformation processing, the system finally outputs a stable, robust value that highly conforms to the real spatial pathological distribution law, which is the spatial inflammation score coefficient used as the benchmark for subsequent anomaly detection.
[0119] Building upon the local neighborhood network constructed in the above embodiments, this embodiment further provides a specific implementation method for calculating the space-expression coupling potential value in order to objectively and robustly quantify the true inflammatory state of each cell. (Reference) Figure 7 Step S220, combining the expression levels of inflammation-related genes in the test cells and the local neighborhood network, calculates the spatial-expression coupling potential value, including: Step S221: Calculate the basic inflammatory activity of the test cells using a local functional activity function.
[0120] Inflammatory responses in organisms are typically driven by complex pathways composed of multiple genes, rather than being determined by a single gene. The local functional activity function (RFI) is a multivariate mapping algorithm that comprehensively evaluates and reduces the dimensionality of the high-dimensional expression profiles of multiple inflammation-related genes within a cell to a scalar value reflecting the intrinsic pathological activity of the cell.
[0121] In this step, the basal inflammatory activity of the test cells is calculated using a local functional activity function. For each test cell, the system extracts the expression level matrix of genes (or gene sets) related to the inflammatory response. Subsequently, a preset local functional activity function is used to comprehensively score this expression level matrix.
[0122] In practical implementation, the local functional activity function can preferably be a Gene Set Variation Analysis (GSVA) algorithm, a Single Sample Gene Set Enrichment Analysis (ssGSEA) algorithm, or similar algorithms. Through the operation of this function, discrete gene expression data is converted into a continuous scalar value, namely the basal inflammatory activity, which objectively characterizes the inherent inflammatory level of the test cells without considering the influence of the external environment.
[0123] Step S222: Determine the network decay weights based on the spatial distance and gene expression similarity between cells in the local neighborhood network.
[0124] It should be noted that the intensity of signal interactions between cells is non-uniform within the tissue microenvironment. Network decay weight is a modulatory factor used to quantify the influence of peripheral cells on the central target cell in a local neighborhood network. This invention creatively points out that this weight should not be determined solely by physical distance, but must also be modulated by the similarity of biological states between cells.
[0125] In this step, network decay weights are determined based on the spatial distance and gene expression similarity between cells in the local neighborhood network. The system traverses the local neighborhood network of the cell under test and performs a dual-feature similarity assessment for each neighboring cell in the network: on the one hand, it extracts the physical spatial coordinates of the cell under test and the surrounding cells, and calculates the Euclidean distance between them; the closer the distance, the stronger the potential for spatial interaction. On the other hand, it extracts the global or local gene expression profiles of the two cells and calculates the gene expression similarity between them (e.g., using the Pearson correlation coefficient for assessment); the higher the similarity, the greater the probability that the two cells will respond cooperatively in their biological states.
[0126] By combining the attenuation effect of spatial distance with the enhancement effect of gene expression similarity, a specific network attenuation weight is assigned to the surrounding cells. This weighting mechanism perfectly aligns with the real physical laws of signal targeted transmission and proximity diffusion in the biological microenvironment.
[0127] Step S223: Calculate the spatial-expression coupling potential value based on the inflammatory baseline activity and the network decay weights.
[0128] The aforementioned space-expression coupling potential value is a comprehensive evaluation index that integrates the intrinsic activity of single cells with the influence of the external microenvironment. It can effectively compensate for the false negative noise of gene expression technology commonly found in single-cell sequencing and truly restore the comprehensive pathological potential of cells in the microenvironment of the tissue section.
[0129] In this step, the spatial-expression coupling potential value is calculated based on the basal inflammatory activity and the network decay weights. After obtaining the weight parameters of all surrounding cells, the algorithm performs network aggregation operations.
[0130] Specifically, the basic inflammatory activity of each surrounding cell is multiplied by its corresponding network decay weight to obtain the "potential energy contribution value" of the surrounding cells to the test cell. Furthermore, the potential energy contribution values of all surrounding cells are aggregated and summarized, and then mathematically fused with the basic inflammatory activity of the test cell itself (e.g., through a weighted summation operation).
[0131] Through this energy coupling mechanism based on network topology, the spatial-expression coupling potential value of the tested cells is finally calculated. This coupling potential value not only reflects the individual pathological characteristics of the cells, but also describes their community interaction state in the real tissue microenvironment at a deeper level, providing a highly robust quantitative basis for subsequent high-precision screening of abnormal cells.
[0132] Building upon the spatial-expression coupling potential calculation framework provided in the above embodiments, this embodiment further provides a specific calculation method for determining network attenuation weights in order to simulate the complex intercellular signal interaction patterns within biological tissues with extremely high accuracy. (Reference) Figure 8 Step S222, based on the spatial distance and gene expression similarity between cells in the local neighborhood network, determines the network decay weights, including: Step S2221: Calculate the distance decay factor based on the spatial Euclidean distance between the cell under test and its neighboring cells in its local neighborhood network.
[0133] It should be noted that in real biological tissue sections, the paracrine diffusion of cytokines strictly follows the concentration decay law in physics. The distance decay factor is a mathematical parameter used to quantify the signal loss with increasing physical distance in a spatial microenvironment. The introduction of this factor ensures that the algorithm does not assign unrealistically high influence to distant cells, making signal transmission conform to the real physical diffusion gradient.
[0134] In this step, within the constructed local neighborhood network, for each connection edge of the cell under test, the system extracts the geometric coordinates of the cell under test and its corresponding neighboring cells in two-dimensional or three-dimensional physical space. The Euclidean distance formula is used to calculate the linear physical distance between the two coordinate points. Subsequently, this spatial Euclidean distance is input into a preset decay function (preferably an exponential decay function with a natural constant base or a Gaussian kernel function) for transformation and calculation. This yields a value that smoothly decreases with increasing spatial distance, namely the distance decay factor, which truly maps the diffusion limitations of the microenvironment in physical space.
[0135] Step S2222: Calculate the expression similarity factor based on the correlation between the gene expression vectors of the cell to be tested and the neighboring cells.
[0136] Tissue heterogeneity means that adjacent cells may not have the same biological functions. For example, immune cells at the edge of a lesion and adjacent normal structural cells respond to inflammation with drastically different mechanisms. Expression similarity factors quantify the biological "affinity" between cells by comparing the correlation of high-dimensional gene expression profiles, effectively preventing erroneous propagation of computational signals across unrelated cell types or tissue boundaries.
[0137] In addition to physical distance, the system simultaneously extracts high-dimensional gene expression profile data of the cell under test and its neighboring cells, constructing high-dimensional gene expression vectors from the gene expression characteristics of each cell. Subsequently, statistical correlation algorithms (such as Pearson correlation coefficient, Spearman correlation coefficient, or cosine similarity) are used to calculate the correlation score between the two gene expression vectors.
[0138] The correlation score objectively reflects the degree of convergence of the physiological state and transcriptome characteristics of the two cells in the current microenvironment. The system uses this score directly or after normalization as an expression similarity factor.
[0139] Step S2223: The product of the distance decay factor and the expression similarity factor is determined as the network decay weight.
[0140] After obtaining the distance parameter representing physical constraints and the similarity parameter representing biological synergy, the algorithm performs a fusion operation. Specifically, it directly multiplies the distance decay factor calculated for the same neighboring cells with the expression similarity factor. This multiplication operation constitutes a rigorous "dual-feature gating" mechanism in mathematical logic.
[0141] This means that only when neighboring cells are physically close enough and highly similar in their biological transcriptional state will their product (i.e., the final network decay weight) maintain a high numerical level. If the conditions in either dimension (too far apart or with vastly different expressions) are not met, the product effect will cause the final weight of that neighboring cell to decay rapidly.
[0142] Through this product coupling of dual-modal features, this embodiment effectively filters out false interference signals in tissue slices caused by physical proximity but functional irrelevance, providing the most reliable weight matrix for subsequent high-fidelity potential energy aggregation calculations.
[0143] In the space-expression coupling potential calculation framework provided in the above embodiments, in order to quantify the intrinsic inflammation degree of a single test cell with extremely high accuracy and objectivity, this embodiment further provides a specific calculation expression for the basic inflammatory activity of the test cell. The calculation expression for the basic inflammatory activity of the test cell is: (Formula 5); Wherein, F(e) represents the overall inflammatory basal activity of the test cell e, which serves as the core input parameter for subsequent spatial energy coupling with the surrounding microenvironment; τ represents the pre-acquired set of pathways related to the inflammatory response, for example, it can be the set of GO pathways related to the inflammatory response in the GSEA-MSigDB library or the Gene Ontology library (GO names appear in the INFLAMMATORY_RESPONSE or inflammatory response field, such as GO:0006954).
[0144] T represents any pathway with a specific function in the pathway set τ. ω T The weight of pathway T can be dynamically configured and assigned based on different tissue types or prior disease pathological characteristics to highlight the influence of the core driving pathway.
[0145] Furthermore, in order to achieve absolute shielding of irrelevant biological background noise at the mathematical computation level, in this embodiment, the weight ω of the pathway T is... T The preferred approach is a binary conditional weight.
[0146] Specifically, this weight ω T The parameter conventions can be as follows: (Formula 6); The logical meaning of the above-mentioned conditions is: when a certain pathway T belongs to a predefined set of pathways τ related to inflammatory responses, its corresponding weight ω TA value of 1 is assigned, meaning the activity evaluation value of that pathway is fully retained; conversely, for other irrelevant functional pathways that do not belong to the pathway set τ, their corresponding weights ω are changed. T The value is forcibly assigned to 0. Through the above-mentioned binary weighted hard filtering mechanism, it is ensured that the final calculated basal inflammatory activity F(e) is absolutely focused on the core inflammatory signal.
[0147] GSVA T (e) represents the evaluation value after assessing the activity of a specific pathway T in the test cell e using functional activity algorithms such as gene set variation analysis.
[0148] Regarding the GSVA value used to evaluate pathway T activity in this embodiment... T (e) The specific method for obtaining this information can be achieved by those skilled in the art using standard gene set variation analysis algorithms. As a preferred implementation method with extremely high engineering feasibility, this GSVA... T The value of (e) can be automatically evaluated and calculated by calling a third-party open-source bioinformatics analysis toolkit (i.e., the R package GSVA) in the computer R language environment. Using this mature algorithm toolkit, discrete single-cell gene expression matrices can be efficiently transformed into continuous pathway enrichment score matrices, thereby providing stable and reliable underlying activity data input for the spatial-expression coupling potential calculation in the embodiments of this application.
[0149] It should be noted that pathological inflammatory responses in organisms are complex cascade amplification processes, involving the coordinated action of hundreds or even thousands of related genes. The expression levels of a single gene are highly susceptible to interference from sequencing depth, resulting in significant random fluctuations (such as technical false negatives). Therefore, this embodiment abandons the single-gene counting method and instead uses "signaling pathways" as the basic assessment unit. A pathway set refers to a combination of multiple functional gene sets highly correlated with a specific inflammatory response, pre-screened from biological databases.
[0150] The GSVA evaluation value mentioned above is the gene set variation analysis evaluation score. Odor is an algorithmic tool that can assess the enrichment of specific biological pathways at the single-cell level. It can transform discrete gene expression data into stable and continuous pathway activity expression levels, effectively overcoming the sparsity problem of single-cell sequencing data.
[0151] In specific tissue microenvironments or disease models, the contributions of different inflammatory pathways to the overall pathological state vary significantly. Pathway weights are used to quantify the importance or dominance of a particular pathway in the overall inflammatory response.
[0152] In the specific computer processing execution, the system first extracts the original gene expression vector of the cell under test; then, by referring to each pathway T defined in the pathway set τ, it calculates the functional activity assessment value (GSVA) of the cell under test in each pathway. T (e) After obtaining the evaluation values of all individual pathways, the system calls the preset weight matrix to calculate the GSVA of each pathway's evaluation value. T (e) and its corresponding weight ω T Perform product operations.
[0153] Finally, all product results are summed using a weighted average. Through the aforementioned dimensionality reduction and feature fusion mechanism, which moves from "multi-gene" to "multi-pathway" and then to "weighted aggregation," this embodiment filters out sequencing noise at the single-molecule level with extreme precision, and the output F(e) can reflect the true physiological pathological potential of the tested cells with high fidelity.
[0154] Based on the spatial and expressive feature fusion framework provided in the above embodiments, in order to enable those skilled in the art to more clearly understand the underlying mathematical execution logic of the algorithm, this embodiment further provides a specific mathematical expression for calculating the spatial-expressive coupling potential (SECP) of the cell under test.
[0155] The expression for calculating the space-expression coupling potential energy value is as follows: (Formula 7); Among them, SECP i F(e) represents the space-expression coupling potential of cell i under test. This value not only includes its own state but also incorporates the synergistic influence of the local microenvironment. i ) represents the basic inflammatory activity of cell i, serving as the basis scalar for potential energy calculation; k represents the number of surrounding cells in the local neighborhood network of cell i, that is, the total number of surrounding cells (i.e., neighboring nodes) contained in the local neighborhood network centered on cell i. This value defines the local topological boundary of signal interaction. j represents the j-th neighboring cell in the local neighborhood network, and 1≤j≤k.
[0156] w ij The weight represents the network attenuation weight between the cell i under test and the surrounding cell j. This weight is mainly modulated by the Euclidean distance between the two in physical space. The greater the distance, the lower the weight, thus simulating the physical diffusion attenuation of the signal.
[0157] α represents the expression similarity factor, which participates as a gating multiplier in the accumulation term of each potential energy transfer, and α meets the following conditions: (Formula 8); Where, sim(e i ,e j The expression correlation between cell i and surrounding cell j is represented by ).
[0158] In specific computer algorithm implementations, if sim(e i ,e j If the expression patterns of the two cells are highly positively correlated, then α is assigned a high coefficient value, allowing the inflammatory potential of the surrounding cell j to be smoothly transferred to the cell i being tested; conversely, if the expression patterns of the two cells are uncorrelated or even negatively correlated (indicating that they belong to different microenvironments or functionally isolated cell categories), then α approaches or is truncated to 0, thereby mathematically blocking the potential transfer of this pathway.
[0159] It should be noted that in the physical microenvironment of biological tissues, adjacent cells do not always belong to the same functional type. Expression correlation refers to quantifying the similarity between two cells in biological function and response state by calculating the statistical correlation between their high-dimensional gene expression vectors (such as Pearson correlation coefficient or Cosine similarity).
[0160] The aforementioned expression similarity factor is a mathematically gated modulation coefficient derived based on the aforementioned expression correlation. The fundamental purpose of introducing this factor is to prevent erroneous mathematical propagation of inflammatory potential signals across unrelated cell types in physical space (e.g., across normal stromal cell layers), thereby ensuring the biological fidelity of signal aggregation.
[0161] By incorporating the aforementioned comprehensive potential energy equation, which includes the active substrate, spatial distance attenuation weights, and expression similarity gating, the algorithm of this invention exquisitely recreates the real, complex, and constrained microenvironmental communication network in pathological sections, thereby outputting a potential energy score with extremely high signal-to-noise ratio and biological interpretability for each cell.
[0162] In the network attenuation weight determination mechanism provided in the above embodiments, in order to simulate the real molecular diffusion law within biological tissues with extremely high accuracy and to ensure that the algorithm has adaptive generalization ability to the resolution differences of different sequencing platforms, this embodiment further provides the network attenuation weight w. ij The specific mathematical expression for calculating the network attenuation weight w. ij The calculation expression is: (Formula 9); Among them, w ij This represents the final network decay weight between cell i and its surrounding cells j. The value range is limited to [0,1]. The larger the value, the higher the efficiency of spatial communication.
[0163] (x,y) iRepresents the spatial geometric coordinates (x, y) of cell i in a two-dimensional slice or three-dimensional tissue space. j Represents the spatial geometric coordinates of the surrounding cell j in the local neighborhood network; d represents the square of the linear Euclidean distance between the coordinates of cell i and its neighboring cell j, serving as the absolute degree vector of spatial physical loss; d represents the minimum distance between all known cells or sequencing sites in the currently input spatial transcriptome dataset; this value serves as a global adaptive smoothing parameter, constituting the denominator constraint term d. 2 This is used to eliminate rigid scale interference caused by differences in the spacing between sequencing sites on different platforms. exp represents an exponential function with the natural constant as its base.
[0164] It should be noted that the efficiency of paracrine communication of intercellular inflammatory signals is strictly limited by physical spatial barriers. The square of the Euclidean distance mentioned above is used to quantify the nonlinear loss caused by this spatial barrier. Using a square term instead of a first-order term can better simulate the real physical diffusion gradient of molecular concentration in three-dimensional / two-dimensional tissue matrix as a function of distance.
[0165] It should be further noted that there are many types of spatial transcriptomics sequencing platforms, and the physical distance between the chip arrays on different platforms varies greatly. If a fixed physical scale is used for attenuation constraints, it will lead to severe generalization bias in the model on different platforms. In this embodiment, the global minimum distance between all cells in the dataset is extracted as an adaptive scale normalization factor, transforming the absolute physical distance into a "relative topological distance," thereby giving the model universality for heterogeneous sequencing data.
[0166] In the specific computer processing flow, the system can first scan the coordinate matrix of the entire spatial transcriptome data, calculate and cache the global minimum distance parameter d. Then, for any test cell i and its neighboring cell j, the system extracts their spatial coordinates and calculates the squared Euclidean distance. This squared value is then divided by d. 2 After normalization to the relative scale, the negative value is taken and substituted into the natural exponential function for continuous mapping transformation. Through this smooth and non-linear exponential transformation, the system finally outputs network decay weights that strictly follow the spatial diffusion decay law of biomolecules. This weighting operation not only filters out invalid noise from distant cells but also provides solid geometric and mathematical support for subsequent high-fidelity aggregation of inflammatory potential.
[0167] Based on the spatial-expression coupling potential value calculated in the above embodiments, in order to eliminate the inherent technical noise in single-cell spatial transcriptome sequencing (such as false negatives caused by the drop-out effect or isolated false positives caused by non-specific amplification), this embodiment further provides a specific implementation method for spatially smoothing the spatial-expression coupling potential value to obtain the final spatial inflammation score coefficient.
[0168] Step S230 involves performing a spatial smoothing transformation on the spatial-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the tested cells, including: Step S231, calculate the expression vector after spatial smoothing transformation: (Formula 10); in, represents the expression vector of a certain inflammation-associated gene g after spatial smoothing transformation, where each element has been filtered out from isolated sequencing spatial noise; SECP represents the confidence weight vector obtained after normalizing the spatial-expression coupling potential values of each cell; e g This represents the initial, unspaced expression vector of the gene g.
[0169] It should be noted that, since the spatial-representation coupling potential values calculated from different slices or batches may have huge differences in dimensions and absolute values, directly using them for product calculation will lead to gradient explosion. Normalization refers to using mathematical mapping (such as Min-Max normalization or Z-score standardization) to scale the potential vector to a uniform, dimensionless weight interval (e.g., the [0,1] interval), thus transforming it into a pure spatial confidence weight factor.
[0170] The aforementioned spatial smoothing transformation is essentially a signal filtering mechanism based on microenvironment topology. It utilizes the normalized potential vector as spatial prior weights to perform inverse modulation and weighted scaling on the original gene expression data. This transformation can intelligently suppress spatially isolated random high-expression noise and enhance real inflammatory signals with localized clustering effects.
[0171] In this step, the computer system first extracts the spatial-expression coupling potential values of all cells in the global dataset, forming a global potential vector, and then processes it using a normalization algorithm to obtain a normalized weight vector. Subsequently, for each preset inflammation-associated gene g, its initial expression level in all cells is extracted and a new expression vector is constructed. The normalized potential vector is then combined with the initial expression vector of gene g (preferably using element-wise multiplication) to obtain a spatially topologically corrected expression vector.
[0172] Step S232: Calculate the spatial inflammation score coefficient based on the expression vector. (Formula 11); Among them, IID i represents the spatial inflammation score coefficient of the test cell i after comprehensive dimensionality reduction. This coefficient is the only numerical basis for subsequent generalized extreme value distribution anomaly testing; m represents the total number of inflammation-related genes; k represents the k-th inflammation-related gene, and satisfies 1≤k≤m; The expression level of the k-th inflammation-associated gene in cell i after spatial smoothing transformation is represented by the expression level vector. Extract the corresponding index position.
[0173] Through the rigorous signal filtering and feature aggregation calculations described above, this embodiment stably integrates high-dimensional single-molecule expression profiles with macroscopic physical spatial networks. The output score coefficients can withstand extremely high sequencing noise interference, providing solid underlying data support for the accurate characterization of real diseased cells.
[0174] Example 4 refer to Figure 9 Building upon the abnormal cell spatial clustering achieved in the above embodiments, this embodiment further provides a specific implementation method for determining inflammatory hotspots in the tissue under test based on the spatial enrichment characteristics of cell clusters in order to accurately isolate core lesions with true pathological significance from a macroscopic tissue morphology perspective. Based on the aforementioned spatial geometry and statistical mechanisms, the determination of inflammatory hotspots in the tissue under test based on the spatial enrichment characteristics of cell clusters in this embodiment specifically includes the following processing steps: Step S500, based on the spatial enrichment characteristics of the cell clusters, determines the inflammatory hotspots in the tissue to be tested, including: Step S510: For each cell cluster, calculate its local background density and the density of abnormal points within the cluster.
[0175] The density of abnormal points within the clusters described above is used to quantify the compactness and density of abnormally inflammatory cells within a specific cell cluster. True inflammatory lesions are typically characterized by extremely dense aggregation of cells, rather than sparse dispersion.
[0176] The aforementioned local background density is used to quantify the average abnormal inflammatory baseline level of the microenvironment surrounding the cell cluster. Determining whether a region is a "hotspot" requires considering its contrast relative to its surrounding environment.
[0177] After clustering abnormal cells, the computer system traverses at least one generated cell cluster. For any given cell cluster, the system extracts the spatial coordinates of all abnormal cells within that cluster. On one hand, by calculating the actual geometric area or volume covered by the cell cluster in two-dimensional or three-dimensional space, and combining this with the total number of abnormal cells it contains, the system calculates the density of abnormal points within the cluster. On the other hand, the system expands outward from this cell cluster to determine a local background geometric space encompassing the cluster and its surrounding neighboring areas. By statistically analyzing the distribution of abnormal cells within this local background geometric space, the corresponding local background density is calculated. Through these dual density calculations, the internal and external features of the cell cluster's spatial morphology are extracted.
[0178] Step S520: Calculate the local enrichment score based on the local background density, the density of outliers within the cluster, and a preset smoothing constant.
[0179] In order to prevent the calculation crash caused by the "denominator being zero" due to the zero number of abnormal cells in the background region, or to prevent the mathematical amplification effect in the extremely low density region, this embodiment introduces a small positive constant (similar to the pseudo-count in statistics), a preset smoothing constant, to ensure the smoothness and robustness of the enrichment score calculation.
[0180] The aforementioned local enrichment score is a quantitative scoring index that combines internal compactness and external contrast. The higher the score, the more significant the abnormal aggregation effect of the cell cluster in space, and the greater the likelihood that it is a core inflammatory lesion.
[0181] In this step, the local enrichment score is calculated based on the local background density, the density of outliers within the cluster, and a preset smoothing constant. After obtaining the aforementioned dual density indices, the system performs a spatial enrichment assessment.
[0182] Specifically, the system substitutes the density of outliers within the cluster and the local background density into a preset enrichment comparison model (e.g., calculating the ratio or logarithmic relative difference between the two). During the calculation, a preset smoothing constant is simultaneously introduced into the calculation model (e.g., adding the smoothing constant to both the numerator and denominator of the ratio formula). Through this aggregation operation with the introduction of a smoothing constant, the mathematical oscillation interference of background noise under extremely low density conditions is effectively suppressed, and a stable and objective value for measuring the local spatial clustering intensity, i.e., the local enrichment score, is finally calculated and output.
[0183] Step S530: The cell clusters with local enrichment scores greater than a preset enrichment threshold are identified as the inflammatory hotspots.
[0184] After calculating the local enrichment scores of all cell clusters, the system calls a preset enrichment threshold for final binary classification. This preset enrichment threshold represents the critical criterion for statistically or pathologically confirming true aggregated lesions. The system compares the local enrichment scores of each cell cluster sequentially with this threshold. For cell clusters with scores less than or equal to the threshold, the system classifies them as random inflammatory fluctuations without significant aggregation and filters them out; while for cell clusters with local enrichment scores significantly greater than the preset enrichment threshold, the system identifies them as having extremely significant high-density aggregation characteristics in space, and then precisely demarcates the physical space area they cover and outputs it as the final inflammatory hotspot area in the tested tissue.
[0185] Through the above-mentioned automated and standardized density comparison and threshold screening, this embodiment effectively overcomes the defects of manual selection, which is highly subjective and has unclear boundaries, and provides a highly reliable target area division for the spatial pathological diagnosis of complex diseases.
[0186] Based on the spatial enrichment feature quantification framework provided in the above embodiments, in order to evaluate the internal density of each anomalous cell aggregation region with extremely high accuracy, this embodiment further provides a specific mathematical expression for calculating the density of anomalous points within the cluster and its implementation method. Based on the above geometric mechanism, the calculation expression for the density of anomalous points within the cluster in this embodiment specifically includes the following set of equations: The formula for calculating the density of outliers within a cluster is: (Formula 12); (Formula 13); Where, λ sig.k Representing cell cluster C k The density of abnormal points within the cluster; the higher this value, the denser the arrangement of abnormal inflammatory cells within the lesion area; |C k | Represents the cell cluster C after clustering. k The number of abnormal cells contained within is the absolute total number (i.e., the cardinality of the point set); A k Representing cell cluster C k The area of the hull is used as the geometric denominator for density calculation; Area(ConcaveHull(·)) represents the function for solving the area of the hull of a point set in space.
[0187] It should be noted that in spatial geometry and graphics algorithms, polygon envelope algorithms used to define the physical boundaries of discrete point sets are mainly divided into convex hull and concave hull algorithms. Because lesions or abnormal cell clusters in biological tissue slices often exhibit highly irregular shapes (such as star-shaped, elongated branching, or porous shapes), using traditional convex hull algorithms or circumscribed rectangles to define the boundaries will incorrectly include healthy tissue gaps or concave areas outside the point set, resulting in an artificially inflated calculated area. Concave hull algorithms (such as the Alpha-shape algorithm) can adaptively shrink inwards from the point set, closely conforming to the true edges of irregular branches, thereby obtaining the geometric polygon that most closely approximates the actual physical space occupied by the point set.
[0188] The density of abnormal points within the aforementioned cluster refers to the number of abnormal cells contained within a unit of actual aggregated area. This indicator eliminates the interference of the overall absolute size of the cell cluster and is a core parameter that objectively reflects the "concentration degree" of local lesions.
[0189] In the specific execution process of a computer system, for any cell cluster C generated by clustering... k The system first extracts the physical geometric coordinates of all abnormal cells within the cluster on the slice. Then, the system calls a preset concave hull generation algorithm to reconstruct the boundary of this coordinate point set. In a preferred implementation, the system can use an Alpha-shape algorithm based on Delaunay triangulation. By setting an appropriate rolling sphere radius parameter (Alpha value), excessively large external gap triangles are eliminated, thereby generating an irregular concave polygon boundary that closely adheres to the outermost periphery of the abnormal cells. The system then uses a polygon area integration algorithm to calculate the actual closed area of this concave polygon, thus obtaining the concave hull area A. k Finally, the system counts the total number of cells within the cluster |C k |, and A k Perform division to find the quotient.
[0190] By using this "bodysuit"-like concave area constraint, this embodiment fundamentally eliminates the problem of overestimation of area and underestimation of density caused by irregular lesion shape, and the output density of abnormal points within the cluster can perfectly restore the dense morphological characteristics of real biological tissue.
[0191] Based on the cluster anomaly density calculation scheme provided in the above embodiments, in order to objectively and adaptively measure the inflammatory baseline level of the macroscopic environment in which the cell cluster is located, this embodiment further provides a specific spatial definition method and mathematical calculation expression for the local background density.
[0192] Based on the above adaptive geometric mechanism, in this embodiment, step S510, the method for calculating the local background density specifically includes: Step S511, determine the cell cluster C kThe distance from the center of the cell to its farthest abnormal cell is used as the actual radius R. k With the actual radius R k The defined circular region is designated as the local background region; the expression for calculating the local background density is as follows: (Formula 14); Where, λ bg.k Representing cell cluster C k The local background density, as stated above, objectively quantifies the severity of inflammation in the macroscopic microenvironment surrounding the lesion; N abnormal,k This represents the total number of abnormal cells within the local background region (i.e., the circular region constructed above). It should be noted that this total number includes not only the cells constituting the cell cluster C. k The abnormal cells themselves also include other free or sporadically distributed abnormal cells that have fallen into this circular buffer area; R k The actual radius is defined by the maximum distance from the center; π represents pi. That is, the circular geometric area representing the local background region.
[0193] It should be noted that inflammatory lesions (cell clusters) in biological tissue sections vary greatly in area. Using a fixed-size geometric window for background assessment would lead to severe spatial scale distortion. This invention creatively introduces an adaptive mechanism: using the distance from the geometric center of the cell cluster to its farthest abnormal cell as the actual radius, a circular region is constructed that perfectly encloses the irregular cell cluster as a local background region. This circular region not only encompasses the lesion itself, but the buffer space naturally formed at the lesion's edge also objectively reflects the true baseline state of the microenvironment surrounding the lesion, achieving proportional dynamic scaling of the background field of view to the lesion scale.
[0194] The aforementioned local background density is used to quantify the overall distribution level of abnormal inflammatory cells within the aforementioned circular macroscopic field of view, serving as an important reference benchmark for subsequent calculation of spatial enrichment contrast.
[0195] In this step, the cell cluster C is determined. k The distance from the center of the cell to its farthest abnormal cell is used as the actual radius R. k With the actual radius R k The defined circular region is used as the local background region. In the specific computer graphics execution flow, for any cell cluster C obtained from clustering... kThe system first calculates and locates the geometric center point of the cell cluster using spatial coordinates. Then, it iterates through all abnormal cells within the cluster, calculating the straight-line physical distance between each abnormal cell and the center point. The maximum value among all distance calculations is extracted and set as the actual radius R specific to the cell cluster. k Furthermore, in the spatial coordinate system, with the center point as the center and the actual radius R, the system... k A standard circular envelope region is constructed for the radius, and this circular region is defined as the local background region for macroscopic evaluation.
[0196] Secondly, based on the geometric region defined above, the local background density is numerically calculated. The specific expression for calculating the local background density is shown in Formula 14.
[0197] By employing the aforementioned adaptive maximum radius-based circular envelope and density calculation mechanism, this embodiment abandons the rigid fixed window calculation method, endowing the algorithm with extremely strong spatial scale invariance. The output local background density can provide a uniform, objective, and impartial macroscopic environmental reference benchmark for all anomalous clusters of different sizes and shapes, thereby ensuring the accuracy of the final inflammatory hotspot delineation.
[0198] Based on the local background density calculation scheme provided in the above embodiments, in order to determine the boundary of the circular background region used to measure the macroscopic microenvironment with extremely high accuracy and adaptability, this embodiment further provides the actual radius R. k The specific algebraic calculation expression and implementation process.
[0199] The actual radius R k The calculation expression is: (Formula 15); Among them, R k Represents cell cluster C k The actual radius of the circular region used to evaluate the background environment is determined by this parameter. The function representing the extreme value extraction is logically based on the following: taking cell cluster C... k The maximum value of the calculated distances from all abnormal cells to the center; (x,y) represents cell cluster C. k The two-dimensional physical space coordinates of any abnormal cell on the slice, where x is the abscissa and y is the ordinate; Representing cell cluster C k The center x-coordinate; Representing cell cluster C k The central ordinate.
[0200] In a two-dimensional physical slice space, the center coordinates represent the geometric centroid or the average anchor point of the spatial distribution of an irregular, abnormal cell aggregation region (cell cluster).
[0201] The edges of cell clusters are typically highly irregular (with synapses or elongated branches). To construct an adaptive background viewing window that can completely encompass this irregular shape, the system must find the outlier cell within the cluster that is furthest from the center. The straight-line physical distance from this furthest cell to the geometric center, mathematically, constitutes the radius of the smallest circumscribed circle that can cover the entire cell cluster, i.e., the maximum center distance (actual radius).
[0202] In the specific computer processing and execution flow, the system first obtains the current cell cluster C through the coordinate resolution module. k The central x-coordinate and the center ordinate Subsequently, the system initiates a loop traversal mechanism, extracting the spatial coordinates (x, y) of each abnormal cell within the cell cluster. In each traversal, the system uses the planar Euclidean distance formula (i.e., the square root of the sum of the squares of the differences between the x and y coordinates) to calculate the straight-line distance from the specific abnormal cell to the aforementioned geometric center. After completing the distance calculation for all abnormal cells within the cluster, the system generates a distance result set and calls a maximum value search algorithm (max operation) to extract the distance item with the largest value from this result set. This largest distance item is then precisely identified by the system as the cell cluster C. k The actual radius R k .
[0203] Through the rigorous algebraic extremum calculations described above, this embodiment, without the need for complex image edge recognition, utilizes dimensionality reduction calculations of the underlying coordinate matrix to determine the dynamic spatial boundary that can strictly enclose the entire irregular inflammatory lesion, providing accurate scale parameters for the objective quantification of the local macroscopic environment.
[0204] In the adaptive local background region definition scheme provided in the above embodiments, in order to objectively and stably anchor a spatial reference point for any irregularly shaped cell cluster, this embodiment further provides the specific algebraic calculation expressions and implementation process for the central abscissa and central ordinate.
[0205] The expression for calculating the central abscissa is: (Formula 16).
[0206] The expression for calculating the central ordinate is: (Equation 17).
[0207] in, Representing cell cluster C kThe calculated x-coordinate of the center is output. Representing cell cluster C k The calculated center ordinate is output. After clustering, cell cluster C represents... k The absolute total number of abnormal cells contained within, which is used as the denominator in the arithmetic mean calculation; Represents cell cluster C k The summation is performed on the x (horizontal coordinate) or y (vertical coordinate) values of all abnormal cells.
[0208] In the specific computer processing and execution flow, for any cell cluster C generated by clustering... k The system first counts and caches the total number of abnormal cells contained in the cluster. The system then iterates through all abnormal cells within the cell cluster, extracting the two-dimensional spatial coordinates (x, y) of each cell. During the iteration, the system uses an accumulator to calculate the sum of all x-coordinates and the sum of all y-coordinates. Finally, the system divides the sum of the x-coordinates by the total number of cells. Obtain the center x-coordinate Divide the sum of the vertical axes by the total number. Obtain the center ordinate .
[0209] Through the linear arithmetic mean operation with extremely low time complexity described above, this embodiment efficiently and robustly outputs absolutely objective two-dimensional geometric center coordinates for various complex lesions, laying a solid coordinate origin foundation for subsequent high-precision spatial radius deduction.
[0210] Based on the calculation of intra-cluster outlier density and local background density provided in the above embodiments, this embodiment further provides the Local Enrichment Score (LDES) to ultimately determine the spatial enrichment significance of each cell cluster. k The specific mathematical expression for calculation.
[0211] The expression for calculating the local enrichment fraction is: (Formula 18); Among them, LDES k λ represents the local enrichment fraction of cell cluster k; sig.k The density of outliers within cell cluster k; λ bg.k γ represents the local background density of cell cluster k; γ represents the preset smoothing constant, and γ is a constant greater than 0, which is a pre-defined, strictly positive real number constant. In specific engineering implementations, the value of γ can be fine-tuned according to the global average density of the spatial transcriptome data to achieve the best noise reduction effect.
[0212] In the specific execution process of the computer system, the system obtains the internal compactness λ of the target cell cluster. sig.k and its adaptive background density λ within the background window bg.k Subsequently, the system activates the enrichment scoring module, adding the two density values to a preset smoothing constant γ to form the numerator and denominator of the proportional formula. Finally, by calculating the quotient of the numerator and denominator, the local enrichment score (LDES) corresponding to the cell cluster is calculated and output. k .
[0213] By incorporating a smoothing term into the density ratio calculation, this embodiment mathematically constructs a highly robust signal-to-noise ratio detector. This formula not only sensitively captures clustered inflammatory lesions emerging in low-background environments but also automatically filters out false high-resolution signals caused by background sparsity by utilizing the constraint effect of the γ parameter. This quantization strategy based on relative density comparison provides an absolutely objective and interference-resistant numerical foundation for the automated and standardized delineation of inflammatory hotspots.
[0214] Example 5 Based on the abnormal cells identified in the above embodiments, in order to identify pathological regions with structural integrity at the macroscopic tissue level, this embodiment further provides a specific implementation method for clustering the abnormal cells.
[0215] Step S400 involves clustering the abnormal cells to obtain at least one cell cluster, including: Step S410: A community clustering algorithm based on spatial-attribute weighting is used to cluster the abnormal cells. The expression for the resulting clustering result set is: ; Wherein, P represents the set of points of the abnormal cells; C k represents the k-th cell cluster obtained by clustering, where k is a positive integer; K represents the total number of valid cell clusters containing abnormal cells that are greater than or equal to a preset minimum point threshold, which is used to exclude scattered signals that do not have a pathological clustering scale; N represents the set of noise points that have not formed cell clusters, that is, those isolated abnormal cells that are extremely discrete in spatial distribution and fail to meet the community division criteria.
[0216] It should be noted that in the complex tissue microenvironment, the aggregation of abnormal cells is not random. In this embodiment, not only is the physical spatial distance between cells considered, but the spatial inflammation score coefficient calculated above is also introduced as an attribute weight into the clustering process. This dual constraint mechanism ensures that the identified cell clusters are highly connected in space and highly homogeneous in terms of inflammation severity.
[0217] The community clustering algorithm described above is an unsupervised clustering method based on graph theory. It divides closely related nodes into "communities" by maximizing the modularity of the network. Compared to traditional centroid-based clustering, community clustering does not require pre-specifying the number of clusters and can adaptively identify various irregularly shaped inflammatory lesions in tissue slices.
[0218] In this embodiment, the abnormal cells are clustered to obtain at least one cell cluster. The specific implementation steps are as follows: First, the system extracts a set P of all points identified as abnormal cells. For each abnormal cell in this set, its spatial coordinate information and corresponding spatial inflammation score coefficient are extracted. Subsequently, the system performs clustering operations using a community clustering algorithm based on spatial-attribute weighting. In the specific execution process, the system first constructs an association graph with abnormal cells as nodes. For node pairs with spatial distances less than a preset range and similar score coefficients, higher connection weights are assigned. Next, a community detection model (preferably the Leiden algorithm or the Louvain algorithm) is used to perform topological partitioning of the weighted graph. After iterative optimization of the algorithm, the system finally outputs a set of clustering results.
[0219] Furthermore, after performing topological partitioning using the community detection model, in order to accurately identify the core clustering regions of each community cluster and remove marginal free points, this embodiment preferably introduces a kernel density estimation method to calculate the core point density of each initial cell cluster, and then extracts the core point set with a density greater than a preset threshold as the final cell cluster. In addition, it should be understood that, besides the aforementioned preferred Louvain or Leiden community clustering algorithm based on spatial-attribute weighting and KDE estimation, those skilled in the art, guided by the technical concept of this application, can also equivalently use the K-means algorithm, hierarchical clustering, spectral clustering to achieve spatial clustering partitioning, or use the DBSCAN algorithm or OPTICS algorithm to replace density estimation. These equivalent algorithm substitutions do not depart from the core protection scope of this invention.
[0220] Through the aforementioned community clustering process based on both spatial and attribute weighting, this embodiment not only automatically identifies the precise number and spatial distribution of each inflammatory lesion in the slice, but also significantly improves the signal-to-noise ratio of hotspot identification by constraining the minimum number of points and stripping the noise point set N. This method can accurately reproduce the real inflammatory cell infiltration patterns in biological tissues, providing a high-purity candidate set for subsequent evaluation of lesion enrichment characteristics.
[0221] The present invention will be further illustrated below with specific experimental examples. However, it should be understood that these experimental examples are only for more detailed illustration and should not be construed as limiting the present invention in any way.
[0222] To verify the technical reliability and significant advancements of the inflammatory hotspot identification method provided in this application, this section provides two clinical validation examples with different tissue morphologies and introduces the conventional spatial analysis method of "Hotspot+Seurat+OPTICS" as a comparative example.
[0223] It should be noted that in the following experimental examples, the "healthy control group data" used to construct the abnormal baseline model all have independent data sources. Specifically, the healthy control group data and their corresponding disease group (target group) data in each experimental example are all taken from different individuals in their respective projects, and not from different tissue regions of the same individual, thus ensuring the objectivity and independence of the constructed health baseline standard.
[0224] Experiment Example 1: Identification and Comparative Evaluation of Aggregation-Type Inflammatory Hotspots In this experimental example, a method for identifying "agglomerated" inflammatory hotspot regions was investigated to verify the accuracy of the algorithm of this invention in delineating the boundaries of densely aggregated lesions, and to compare and verify it with existing technologies.
[0225] 1. Experimental Data and Methods: Experimental data: Stereo-seq (stereo-semiconductor sequencing) data from the Inflammatory Bowel Disease (IBD) project were used. Specifically, tissue slice data from individuals with IBD were used as the target group (disease group), with sample number A02883B5. Simultaneously, gut Stereo-seq data from independent healthy individuals within the IBD project were extracted as a healthy control group for fitting and constructing the aforementioned generalized extreme value distribution (GEV) abnormal baseline model.
[0226] Comparative method: As a control, the existing "Hotspot+Seurat+OPTICS" integrated spatial clustering and hotspot identification algorithm was used to process the same batch of data in parallel.
[0227] 2. Experimental Results and Analysis: To objectively evaluate the accuracy of the identification results, this experimental case uses the inflammatory area manually selected by the pathologist in the H&E map as the gold standard. Spatial registration technology is used to overlap the expression matrix scatter plot with the H&E map, and the coordinate information of the corresponding area is obtained. The following evaluation indicators are defined: 1) Precision = Number of hotspot cells identified within the box / Total number of all identified hotspot cells; 2) False positive rate = Number of hotspot cells identified outside the box / Total number of all identified hotspot cells.
[0228] (1) Verification of the internal process of the algorithm of this invention: refer to Figure 10(Flowchart of chip analysis for A02883B5) shows that the diagram, from left to right (A to D), fully illustrates the progressive signal processing and delimitation process of the algorithm of this invention: Figure 10 The middle A (basic score results) shows the basic inflammatory activity scores of each cell calculated based on the inflammatory gene set in the initial stage. At this time, the signal distribution is relatively discrete and contains a lot of technical noise. Figure 10 Image B (IID Distribution Results): This image shows the spatial inflammation score coefficient (IID) distribution results obtained after spatial-expression coupling and spatial smoothing transformation. It can be seen that the smoothing process effectively suppressed isolated false positive high points, resulting in an initial enhancement of signals with aggregation tendencies. Figure 10 C (GEV Anomaly Detection Results): This section presents the results after significance testing of the healthy baseline using a generalized extreme value distribution (GEV) model. This step successfully isolated normal physiological background fluctuations and accurately screened out highly abnormal cells with substantial pathological significance. Figure 10 The middle D (hotspot delineation results) shows the inflammatory hotspot enrichment region (i.e. the core target area with dense enrichment of highly abnormal cells) that was finally accurately identified and delineated after clustering based on the above abnormal cells and calculating the Local Enrichment Score (LDES).
[0229] This experimental example sequentially presents the results based on the inflammatory gene set scoring, the IID (spatial inflammation score coefficient) distribution, the results based on GEV (generalized extreme value distribution) anomaly detection, and the finally defined inflammatory hotspot enrichment region.
[0230] The results showed that by comparing with the control group and screening out the high-abnormal cells that meet the requirements, and then enriching them in local areas, it is possible to accurately peel off and locate the real inflammatory hotspots where high-abnormal cells are densely enriched.
[0231] (2) Horizontal comparison results with the comparative example: combined with Figure 11 , Figure 12 , Figure 13 (Comparison of results from multiple approaches). Table 1 shows the quantitative comparison of recognition performance between this approach and the comparative approach for the A02883B5 chip sample: Table 1. Comparison of recognition performance between this scheme and the comparative example.
[0232] Combination Figure 11 , Figure 12 , Figure 13(Comparison of results from multiple approaches). This figure shows the original HE pathological staining image (with major inflammatory areas manually marked), the identification results of the proposed method, and the identification results of the comparative approach (Hotspot + Seurat + OPTICS). Specifically: Figure 11 (Gold Standard Reference Image) shows the original H&E (hematoxylin-eosin) pathological staining image of the sample. In this image, the main inflammatory aggregation areas have been clearly marked manually based on pathological characteristics. These manually marked areas serve as the "gold standard" reference boundary for evaluating the recognition accuracy of various algorithms. Figure 12 (Identification Results of the Invention): This section displays the hotspot delineation results output by the inflammation hotspot identification method provided in this embodiment of the invention. Comparison Figure 11 As can be seen, the present invention can clearly and accurately identify key cohesive local inflammatory areas that highly match the H&E pathological images, with very sharp boundary delineation, effectively filtering out irrelevant background signals. Figure 13 (Comparative Recognition Results): This section displays the recognition results after parallel processing using a comparative algorithm (i.e., the existing Hotspot+Seurat+OPTICS integrated algorithm). Compared to... Figure 11 and Figure 12 In contrast, while the comparative method also identified some inflamed areas, its results included a large number of scattered dots diffused in non-core pathological areas (i.e., healthy or mildly inflamed background tissue). This indicates that the comparative method has poor specificity, produces significant false-positive technical noise, and cannot provide precise local lesion target areas for clinical practice.
[0233] Comparative analysis shows that the solution of the present invention can clearly and accurately identify key cohesive local inflammatory areas that highly match the HE pathological images.
[0234] In contrast, the comparative sample (Hotspot+Seurat+OPTICS) identified a large number of scattered points (i.e., technical noise points) in non-core pathological areas, resulting in poor specificity. Quantitative results showed that the accuracy of this method was as high as 97.07%, while the accuracy of the comparative sample was only 51.84%. Nearly half (48.16%) of the identification points in the comparative sample were false positive signals falling outside the pathological frame. This fully demonstrates that this method, through spatial-expression coupling potential and GEV anomaly detection, can significantly suppress background noise, greatly improving the specificity of identification while maintaining high sensitivity. The purity and accuracy of its hotspot delimitation are far superior to existing technologies, providing high-confidence pathological target areas for clinical diagnosis. In addition, to further verify the specific contribution of the spatial-expression coupling potential (SECP) module in this method to the identification accuracy, this experimental example also compared it with the reference method of 'directly using the original expression level of inflammation-related genes for identification'. Figure 14 For chip sample A02883B5, using the area outlined by the pathologist as the gold standard, the performance indicators of the two methods are compared in Table 2: Table 2. Comparison of SECP Module Validity Verification Table
[0235] Note: Precision rate = number of hotspots falling within the box / total number of all hotspots; False positive rate = number of hotspots falling outside the box / total number of all hotspots.
[0236] The results showed that if the original expression level was used directly for identification without the spatial-expression coupling processing of SECP, the false positive rate would increase dramatically from 2.93% to 78.13%, while the accuracy would drop sharply to 21.87%. This strongly demonstrates that the SECP module, by deeply coupling the cell's physical spatial coordinates with gene expression levels, can extremely effectively suppress sequencing random noise, which is a core technical prerequisite for achieving high-precision lesion demarcation.
[0237] Experiment Example 2: Identification and Comparative Evaluation of Inflammatory Hotspots with Diffuse Distribution In this experimental example, a method for identifying inflammatory hotspots of the "diffuse distribution type" was investigated to verify that the present invention still has high-precision local target area identification capability when faced with complex sections with blurred boundaries and diffuse lesions due to extensive tissue damage.
[0238] 1. Experimental Data and Methods: Experimental data: Stereo-seq data from the internal research project on alcoholic liver injury models were used. Specifically, liver slice data from individuals with alcoholic liver injury were used as the target group (disease group), with sample number B04005C5. Similarly, liver Stereo-seq data from independent healthy individuals in the same internal research project were extracted as the healthy control group to construct the abnormal baseline model. This healthy control group data and the disease group data (sample number B04005C5) originated from different individuals.
[0239] Comparative method: The existing "Hotspot+Seurat+OPTICS" integrated algorithm is also used for comparison.
[0240] 2. Test Results and Analysis: (1) Verification of the internal process of the algorithm of this invention: combined with Figure 15 (B04005C5 chip analysis flowchart) shows that when faced with liver injury tissue with large-area diffuse inflammation, simple gene set scores or basic IID distribution are insufficient to clearly delineate the lesion boundaries.
[0241] Figure 15 The image in section A (basic score results) shows the initial inflammatory baseline activity scores of each cell calculated based on the set of inflammation-associated genes. Since this sample is a diffuse lesion model, it can be seen that the initial scores exhibit a large-area, broad distribution of high expression (reddish tinge) across almost the entire tissue section, with extremely blurred lesion boundaries. Clinically, it is impossible to directly locate the target area based on this image.
[0242] Figure 15 Figure B (IID distribution results) shows the spatial inflammation score coefficient (IID) distribution results obtained after spatial-expression coupling and spatial smoothing transformation. After spatial topological smoothing, locally isolated random technical noise is initially suppressed, and the signal shows certain regional contiguous features. However, due to the high overall inflammation baseline, the visual features of large-area diffusion are still obvious.
[0243] Figure 15 C (GEV Anomaly Detection Results): This section demonstrates the results of abnormal cell screening after performing a significance test on the healthy baseline using a Generalized Extreme Value Distribution (GEV) model. This is the core advantage of this invention in dealing with diffuse samples: the model automatically adapts to and raises the background inflammatory baseline, successfully and extremely sensitively separating the core cell populations with statistically extreme abnormalities (extreme value characteristics) in the degree of inflammation from a large-scale "pan-inflammatory" background, thereby significantly filtering out irrelevant background diffuse signals.
[0244] Figure 15Mid-D (Hotspot Delineation Results): This section demonstrates the precise identification and delineation of inflammatory hotspot regions after spatial community clustering based on the aforementioned extreme abnormal cells and calculation of Local Enrichment Score (LDES). (Comparison) Figure 15 As can be seen from A, the diffuse visual noise has been completely eliminated, and the algorithm has successfully and accurately located several of the most pathologically critical and densely inflammatory target areas in the extensively damaged tissue.
[0245] The results of this experiment (see [link]) Figure 15 (C) This paper fully demonstrates the necessity of introducing the "GEV-based anomaly detection" step. From a theoretical perspective, inflammatory hotspots are complex features composed of multiple gene variables. The quantitative indicators derived from them (such as IID) typically do not follow a symmetrical normal distribution in biological samples. The traditional normal distribution assumption is insufficient to cover complex heterogeneous backgrounds, easily leading to serious statistical biases. Since this approach focuses on the significant increase of inflammatory signals under extreme pathological conditions, these high-score points mathematically belong to the category of maxima. Therefore, this approach creatively introduces a generalized extreme value distribution (GEV) model for modeling, which can capture the long-tail features of the score distribution more accurately than the normal distribution. Figure 15 C and Figure 15 The comparison between A and B shows that the abnormal baseline established by the GEV model can accurately isolate statistically significant extreme abnormal signals from a large-scale "diffuse inflammation" background, effectively eliminating the interference of normal physiological fluctuations, thereby achieving high-gain extraction of core lesions. Figure 15 The visualized workflow intuitively demonstrates the necessity of introducing GEV anomaly detection. (Comparison) Figure 15 A (basic score) and Figure 15 As shown in the middle image (results after GEV testing), in complex samples like alcoholic liver injury with extensive diffuse inflammation, the initial score almost covers the entire image, failing to delineate the core region. However, the GEV model, through adaptive modeling and adjustment of the background baseline, successfully isolated and located the truly pathologically extreme abnormal cell locations from the diffuse signal. This intuitive visual difference fully demonstrates the crucial role of the GEV model in improving the accuracy of identifying complex pathological environments.
[0246] It is worth noting that this invention creatively uses the Generalized Extreme Value (GEV) distribution instead of the traditional normal distribution when constructing the abnormal baseline model. This is because the inflammatory hotspot index is a complex scalar resulting from the deep coupling of multiple gene variables, and its essence is to identify cell sites that exhibit extreme bias under specific pathological conditions. From a statistical perspective, the identification of inflammatory hotspots falls into the category of 'maximums'. Inflammatory signals in biological tissue sections often exhibit significant long-tail characteristics under healthy physiological conditions and do not conform to the assumption of a symmetrical normal distribution. If a normal distribution is blindly used for modeling, the identification results will suffer from severe false positive bias because the extreme probabilities at the tail cannot be accurately captured. By using the GEV model to model this part of the maximum value data, the upper limit of score fluctuation under healthy conditions can be quantified extremely accurately, thus providing the most solid mathematical benchmark for subsequent screening of truly pathologically significant abnormal cells.
[0247] (2) Horizontal comparison results with the comparative example: combined with Figure 16 , Figure 17 and Figure 18 The results clearly show the HE pathology images (with local magnification indicators), the identification results of the present invention, and the identification results of the comparative examples.
[0248] Figure 16 (Gold Standard Reference Image) shows the original H&E (hematoxylin-eosin) pathological staining image of this diffuse sample, which also includes magnified indications of local areas. This image objectively reflects the harsh section condition of "blurred boundaries and diffuse lesions" caused by extensive tissue damage, serving as the "gold standard" reference for evaluating the target area localization ability of various algorithms in complex diffuse backgrounds.
[0249] Figure 17 (Identification Results of the Invention) This paper presents the hotspot delineation results output by the inflammation hotspot identification method provided in the embodiments of the present invention. Comparison Figure 16 As can be seen, even in complex tissue environments with extremely high overall inflammatory baselines and diffuse lesions, the method of this invention can still resist interference from globally diffuse signals and accurately locate and purify several local core areas of severe inflammation. This fully verifies the extremely high recognition accuracy and spatial noise robustness of this invention in complex microenvironments.
[0250] Figure 18 (Comparative Recognition Results) shows the recognition results output after parallel processing using a comparative algorithm (i.e., the existing Hotspot+Seurat+OPTICS integrated algorithm). Compared to... Figure 16 and Figure 17In contrast, this comparative model, lacking adaptive background density calibration and extreme anomaly screening mechanisms, resulted in a large area of red coverage across the entire image (i.e., a severe "starry sky" false positive phenomenon). This result failed to effectively distinguish between core severe lesions and minor background fluctuations, completely losing its clinical significance in guiding local pathological diagnosis.
[0251] Experiments showed that, against a diffusely distributed tissue background, the comparative model (Hotspot+Seurat+OPTICS) lacked an adaptive background density calibration mechanism, resulting in the entire image being covered by red-marked abnormal points, exhibiting a severe "starry sky" false positive phenomenon, and completely losing its significance in guiding local pathological diagnosis.
[0252] On the contrary, the solution of the present invention can still resist the interference of global diffuse signals and locate the local area where severe inflammation has actually occurred with extreme precision, demonstrating extremely high local area recognition accuracy and spatial noise resistance robustness.
[0253] refer to Figure 19 This application embodiment also provides an inflammatory hotspot identification device, comprising: an acquisition module 10, used to acquire spatial transcriptome data of a tissue to be tested; wherein the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue and the expression level of inflammation-related genes; a calculation module 20, used to calculate the spatial inflammation score coefficient of each cell to be tested based on the spatial coordinates and the expression level; a screening module 30, used to construct an abnormal baseline model based on healthy control group data, and to screen out abnormal cells from the cells to be tested by combining the abnormal baseline model and the spatial inflammation score coefficient; a clustering module 40, used to cluster the abnormal cells to obtain at least one cell cluster; and an identification module 50, used to determine the inflammatory hotspots in the tissue to be tested based on the spatial enrichment characteristics of the cell clusters.
[0254] It is understood that the device in this embodiment corresponds to the inflammatory hotspot identification method in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.
[0255] This application also provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the inflammation hotspot identification method described in any of the foregoing embodiments.
[0256] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0257] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.
[0258] This application also provides a computer storage medium storing a computer program, which, when executed on a processor, implements the inflammation hotspot identification method according to any one of the foregoing embodiments.
[0259] The computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0260] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0261] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0262] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0263] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for identifying inflammatory hotspots, characterized in that, include: Acquire spatial transcriptome data of the tissue to be tested; wherein, the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue to be tested and the expression levels of inflammation-related genes; Based on the spatial coordinates and the expression level, the spatial inflammation score coefficient of each of the tested cells is calculated; An abnormal baseline model was constructed based on data from a healthy control group, and abnormal cells were screened from the test cells by combining the abnormal baseline model with the spatial inflammation score coefficient. The abnormal cells are clustered to obtain at least one cell cluster; Based on the spatial enrichment characteristics of the cell clusters, inflammatory hotspots in the tissue under test are identified.
2. The method for identifying inflammatory hotspots as described in claim 1, characterized in that, The abnormal baseline model is constructed based on data from a healthy control group. Abnormal cells are then screened from the test cells using the abnormal baseline model and the spatial inflammation score coefficient. These abnormal cells include: The spatial inflammation score coefficient of cells in the healthy control group was obtained, and extreme value sample data were extracted according to cell type; The generalized extreme value distribution model was fitted using the extreme value sample data; Substitute the spatial inflammation score coefficient of the cell under test into the generalized extreme value distribution model to calculate the weighted marginal p-value of the cell under test. Cells whose weighted marginal p-value is less than a preset significance threshold are identified as abnormal cells; and / or, The extraction of extreme value sample data by cell type includes: For each cell type in the healthy control group, the spatial inflammation score coefficients of each cell contained therein were ranked. Extract the spatial inflammation score coefficient of the cells ranking in the top preset percentage, and use this as the extreme value sample data corresponding to that cell type; and / or, The calculation of the spatial inflammation score coefficient for each cell under test based on the spatial coordinates and the expression level includes: For each of the cells to be tested, a local neighborhood network is determined based on its spatial coordinates; The spatial-expression coupling potential is calculated by combining the expression levels of inflammation-related genes in the test cells and the local neighborhood network. A spatial smoothing transformation is performed on the space-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the tested cells; and / or, The step of determining inflammatory hotspots in the tissue under test based on the spatial enrichment characteristics of the cell clusters includes: For each of the cell clusters, calculate its local background density and the density of outliers within the cluster. The local enrichment score is calculated based on the local background density, the density of outliers within the cluster, and a preset smoothing constant. Cell clusters with local enrichment scores greater than a preset enrichment threshold are identified as inflammatory hotspots; and / or, The process of clustering the abnormal cells to obtain at least one cell cluster includes: The abnormal cells are clustered using a space-attribute weighted community clustering algorithm, and the expression for the resulting clustering set is as follows: ; Wherein, P represents the set of points of the abnormal cells; C k The cluster represents the k-th cell cluster obtained by clustering; K represents the total number of cell clusters containing abnormal cells that are greater than or equal to the preset minimum number of points; N represents the set of noise points that have not formed cell clusters.
3. The method for identifying inflammatory hotspots as described in claim 2, characterized in that, The step of substituting the spatial inflammation score coefficient of the test cells into the generalized extreme value distribution model to calculate the weighted marginal p-value of the test cells includes: Determine the proportional weight of each cell type within the spatial neighborhood of the cell to be tested; For each cell type, the spatial inflammation score coefficient of the cell under test is substituted into the cumulative distribution function of its corresponding generalized extreme value distribution model to calculate the conditional p-value; Based on the proportional weights corresponding to each cell type, all the condition p-values are weighted and summed to obtain the weighted marginal p-value of the cell to be tested. Preferably, the weighted marginal p-value of the cell to be tested is calculated using the following expression: ; Where, p j The weighted marginal p-value represents the cell j being tested; G represents the set of cell clusters or types; g represents a specific cell cluster or type within set G; π jg The probability weight representing the cell group or type g to which the test cell j belongs; IID j G represents the spatial inflammation score coefficient of cell j being tested; g (IID j This represents the cumulative distribution function value calculated by substituting the spatial inflammation score coefficient of the cell j under test into the generalized extreme value distribution model corresponding to the cell cluster or type g. Preferably, the expression for calculating the cumulative distribution function value is: ; Where Gg(x) represents the cumulative distribution function value of the generalized extreme value distribution model corresponding to cell cluster or type g; x represents the input spatial inflammation score coefficient; μ g ξ represents the location parameter of the generalized extreme value distribution model corresponding to cell cluster or type g; g The shape parameter representing the cell cluster or type g; σ g The scale parameter represents the cell cluster or type g; exp represents an exponential function with the natural constant as the base; and / or, The healthy control group data includes n c The record information of each cell; the record information of the i-th cell is represented as: ; in, Represents the sample name to which the i-th cell belongs; The cell name representing the i-th cell; The spatial inflammation score coefficient representing the i-th cell; Represents the group or type to which the i-th cell belongs, and 1 ≤ i ≤ n c ; Furthermore, for each cell type g in the healthy control group, the expression for extracting the extreme value sample data set corresponding to that cell type is: ; Where Ctrlg represents the initial candidate set of extreme value sample data corresponding to cell type g; and / or, The calculation of the spatial-expression coupling potential value by combining the expression levels of inflammation-related genes in the test cells and the local neighborhood network includes: The basal inflammatory activity of the test cells was calculated using local functional activity functions. Based on the spatial distance and gene expression similarity between cells in the local neighborhood network, the network decay weights are determined; The spatial-expression coupling potential value is calculated based on the inflammatory baseline activity and the network decay weights.
4. The method for identifying inflammatory hotspots as described in claim 3, characterized in that, The determination of network decay weights based on the spatial distance and gene expression similarity between cells in the local neighborhood network includes: The distance decay factor is calculated based on the spatial Euclidean distance between the cell under test and its neighboring cells in its local neighborhood network. Based on the correlation between the gene expression vectors of the cell under test and the neighboring cells, an expression similarity factor is calculated; The product of the distance decay factor and the expression similarity factor is determined as the network decay weight; and / or, The formula for calculating the basic inflammatory activity of the cells under test is as follows: ; Wherein, F(e) represents the basic inflammatory activity of the test cell e; τ represents the set of pathways related to the inflammatory response; T represents any pathway in the set of pathways; ω T The weights representing pathway T; GSVA T (e) represents the assessment value for evaluating pathway T activity; and / or, The expression for calculating the space-expression coupling potential energy value is as follows: ; Among them, SECP i F(e) represents the spatial-expression coupling potential of cell i under test; i ) represents the basic inflammatory activity of test cell i; k represents the number of surrounding cells in the local neighborhood network of test cell i; w ij The network decay weights represent the relationship between cell i and its surrounding cells j. α represents the expression similarity factor, and the following conditions are met: ; where sim(e i ,e j The expression correlation between cell i and surrounding cell j is represented by ).
5. The method for identifying inflammatory hotspots as described in claim 4, characterized in that, The network attenuation weight w ij The calculation expression is: ; Where (x,y) i The spatial coordinates (x, y) represent the coordinates of cell i to be tested. j The spatial coordinates of the surrounding cell j; d represents the square of the Euclidean distance between the two coordinates; d represents the minimum distance between all cells in the spatial transcriptome data; exp represents an exponential function with the natural constant as the base.
6. The method for identifying inflammatory hotspots as described in claim 2, characterized in that, The spatial smoothing transformation of the spatial-expression coupling potential value to obtain the spatial inflammation score coefficient for each of the tested cells includes: Calculate the vector of representations after spatial smoothing transformation: ; in, The vector represents the expression level of a specific inflammation-associated gene g after spatial smoothing transformation; SECP represents the vector obtained after normalizing the spatial-expression coupling potential values of each cell; e g This represents the initial expression vector of the gene g; Based on the expression vector, the spatial inflammation score coefficient is calculated: ; Among them, IID i The spatial inflammation score coefficient represents cell i; m represents the total number of inflammation-related genes; k represents the kth inflammation-related gene. The expression level of the k-th inflammation-associated gene in cell i after spatial smoothing transformation is represented by the expression level vector. Extract from.
7. The method for identifying inflammatory hotspots as described in claim 2, characterized in that, The formula for calculating the density of outliers within a cluster is: ; ; Where, λ sig.k Representing cell cluster C k The density of outliers within the cluster; |C k | Represents cell cluster C k The number of abnormal cells contained within; A k Representing cell cluster C k The area of the hull; Area(ConcaveHull(·)) represents a function to solve for the area of the hull of a point set in space; and / or, The method for calculating the local background density includes: determining the cell cluster C k The distance from the center of the cell to its farthest abnormal cell is used as the actual radius R. k With the actual radius R k The defined circular region is designated as the local background region; the expression for calculating the local background density is as follows: ; Where, λ bg.k Representing cell cluster C k The local background density; N abnormal,k The value R represents the total number of abnormal cells within the local background region; π represents pi; where the actual radius R... k The calculation expression is: ; Among them, R k Represents cell cluster C k The actual radius; Representative cell cluster C k The maximum value of the calculated results for all abnormal cells; (x,y) represents cell cluster C. k The spatial coordinates of any abnormal cell in the array, where x is the abscissa and y is the ordinate; Representing cell cluster C k The center x-coordinate; Representing cell cluster C k The center ordinate; The expression for calculating the central abscissa is as follows: ; and / or, the expression for calculating the center ordinate is: ; in, Representing cell cluster C k The calculated x-coordinate of the center is output. Representing cell cluster C k The calculated center ordinate is output. Representing cell cluster C k The total number of abnormal cells contained within; Represents cell cluster C k Sum the x or y coordinates of all abnormal cells; and / or, The expression for calculating the local enrichment fraction is: ; Among them, LDES k λ represents the local enrichment fraction of cell cluster k; sig.k The density of outliers within cell cluster k; λ bg.k γ represents the local background density of cell cluster k; γ represents the preset smoothing constant, and γ is a constant greater than 0.
8. A device for identifying inflammatory hotspots, characterized in that, include: The acquisition module is used to acquire spatial transcriptome data of the tissue to be tested; wherein, the spatial transcriptome data includes the spatial coordinates of each cell to be tested in the tissue to be tested and the expression levels of inflammation-related genes; The calculation module is used to calculate the spatial inflammation score coefficient of each of the test cells based on the spatial coordinates and the expression level; The screening module is used to construct an abnormal baseline model based on data from a healthy control group, and to screen out abnormal cells from the cells to be tested by combining the abnormal baseline model and the spatial inflammation score coefficient. A clustering module is used to cluster the abnormal cells to obtain at least one cell cluster; The identification module is used to determine the inflammatory hotspots in the tissue to be tested based on the spatial enrichment characteristics of the cell clusters.
9. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the inflammatory hotspot identification method according to any one of claims 1-7.
10. A computer storage medium, characterized in that, It stores a computer program that, when executed on a processor, implements the method for identifying inflammatory hotspots according to any one of claims 1-7.