Screening method of lung adenocarcinoma immune cell subpopulation network markers

By using the KLoop algorithm and the Graph Attention Network (GAT) model, the challenges of screening network markers and analyzing spatial relationships of immune cell subsets in lung adenocarcinoma were solved. This enabled fine segmentation of immune cell subsets and quantification of their functional states, improving the accuracy and reliability of tumor microenvironment analysis.

CN121545585APending Publication Date: 2026-02-17SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511713476.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In the study of immune cells in lung adenocarcinoma, existing technologies, such as RNA-seq data analysis, are costly, complex, and fail to reflect the complex gene regulatory network of lung adenocarcinoma. They are also unable to effectively screen for immune cell subset network markers and analyze cell spatial relationships.

Method used

We constructed a network biomarker screening process for immune cell subpopulations using the KLoop algorithm, and combined it with a graph attention network (GAT) cell classification model to perform spatial relationship analysis of immune cell subpopulations, including data preprocessing, network biomarker screening, and spatial map acquisition, and verified the effectiveness of the biomarkers.

Benefits of technology

It enables fine segmentation and functional state quantification of immune cell subsets, breaking through the resolution limitations of traditional methods, providing more accurate tumor microenvironment analysis, and providing guidance for personalized treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545585A_ABST
    Figure CN121545585A_ABST
Patent Text Reader

Abstract

The invention provides a lung adenocarcinoma immune cell subset network marker screening method. According to the lung adenocarcinoma immune cell subset network marker screening method, a lung adenocarcinoma immune cell subset network marker screening model with specificity and accuracy is constructed, and a network marker capable of accurately representing immune cell subset characteristics is screened out; meanwhile, by establishing a high-interpretability cell classification model based on the network markers and adopting the relevance between the network markers and cell classification results, the cell classification performance is improved, and effective and reliable verification of the network markers is achieved. Finally, a lung adenocarcinoma immune cell subset spatial relationship analysis algorithm based on the network marker regulatory relationship is constructed, the limitation of a high-cost spatial imaging technology is broken through, the correlation between the regulatory relationship between the network markers and immune cell spatial distribution is disclosed, and clues are provided for understanding a cell interaction mechanism in a tumor microenvironment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of bioinformatics and artificial intelligence, in particular to a lung adenocarcinoma immune cell subpopulation network marker screening method. BACKGROUND

[0002] Lung adenocarcinoma is one of the highest incidence and mortality cancers, due to its high invasiveness and poor prognosis, which makes early diagnosis and treatment of lung adenocarcinoma particularly important. There is currently a certain amount of technical accumulation in the study of lung adenocarcinoma immune cell characteristics.

[0003] In the prior art, RNA sequencing (RNA-seq) is one of high-throughput sequencing technologies, which is widely used in the study of lung adenocarcinoma cell characteristics. Bulk RNA sequencing (bulk RNA-seq), single-cell RNA sequencing (single-cell RNA-seq, scRNA-seq) and spatial transcriptomics RNA sequencing (spatial transcriptomics RNA-seq, stRNA-seq) technologies have become important tools for analyzing lung adenocarcinoma tumor microenvironment, providing multi-dimensional gene expression information for mining lung adenocarcinoma immune cell characteristics.

[0004] Among the three sequencing data, bulk RNA-seq is cost-effective, mature in technology, and suitable for large-scale sample analysis, but cannot reveal single-cell heterogeneity and spatial distribution; scRNA-seq can analyze the gene expression characteristics of single cells, helping to identify cell differentiation and rare cell types, but has the problems of high cost and complex data analysis; although stRNA-seq can obtain gene expression and cell spatial location information at the same time, spatial relationship analysis relies on high-cost technology (such as CosMx imaging, IMC), and the resolution of the obtained data has not reached the single-cell level, lacking systematic analysis of the correlation between gene regulation relationship and cell spatial distribution. Due to the limitations of the separate use of the three RNA-seq data, it is difficult to reflect the complex gene regulation network of lung adenocarcinoma. Although the above-mentioned prior art can better characterize cell spatial characteristics, the research methods are complex and the cost is high. Therefore, how to construct a more scientific and effective lung adenocarcinoma immune cell subpopulation network marker screening and cell spatial relationship analysis method is still an urgent direction to be improved in the field. SUMMARY

[0005] The purpose of the present application is to provide a lung adenocarcinoma immune cell subpopulation network marker screening method to solve at least one technical problem in the prior art.

[0006] To solve the above technical problems, the present application provides a lung adenocarcinoma immune cell subpopulation network marker screening method, comprising the following steps: S1. Construct a screening process for network biomarkers of immune cell subsets in lung adenocarcinoma based on the KLoop algorithm; The network biomarker screening process includes raw data preprocessing, cell-network biomarker matrix construction, KLoop clustering, and network biomarker screening. S2. Construct a cell classification model based on graph attention network (GAT); The cell classification model consists of three parts: a representation layer, a aggregation layer, and a linear classification layer, and is used to verify the effectiveness of the network markers. S3, Spatial Relationship Analysis of Immune Cell Subpopulations; The spatial relationship analysis of immune cell subsets includes spatial data preprocessing, spatial map acquisition and determination of network marker regulatory relationships, and spatial relationship analysis of immune cell subsets.

[0007] Furthermore, S1 specifically refers to: S11. Raw data preprocessing: Load and preprocess the annotated lung adenocarcinoma immune cell gene expression data and candidate network biomarkers; S12. Constructing a cell-network marker 01 matrix: Divide the lung adenocarcinoma immune cell subset into multiple initial cell clusters, extract the matrix elements of each initial cell cluster, and characterize the candidate network markers expressed by the lung adenocarcinoma immune cell subset through the matrix elements. S13, KLoop Clustering: An iterative clustering algorithm is used to segment each lung adenocarcinoma immune cell subpopulation until the initial cell clusters after segmentation meet the preset conditions, and the final cell clusters are extracted. S14. Network marker screening: Calculate the expression ratio of each candidate network marker, and screen the candidate network markers that simultaneously meet the requirements of high accuracy and high specificity as the final network markers.

[0008] Furthermore, S13 specifically includes: S131, Given a set containing A dataset of data points Each data point have Each feature dimension, i.e. ; S132, Random selection The center of each initial cell cluster ,in Indicates the first The center position vectors of the initial cell clusters; Calculate each of the data points With the center of each of the initial cell clusters distance and the data points Assigned to the nearest center In the initial cell cluster where it is located; wherein, the distance It is represented as: ; S133, the data points After allocation, based on the data points in each of the initial cell clusters Update the center of the initial cell cluster The location; the center of the initial cell cluster. Update Center The update center Characterized as:

[0009] in, Indicates the first The data points contained in the initial cell clusters gather, This represents the data point corresponding to the initial cell cluster. Quantity; S134. Apply the k-means method to each initial cell cluster to divide it into two smaller clusters. Repeat this process until the resulting cell clusters satisfy the following conditions: the number of cells in the cluster is ≥ 1% of the total number of cells, and there are valid network markers in the cluster.

[0010] Furthermore, S14 specifically includes: S141. Calculate the expression percentage of the candidate network markers. ;

[0011] in, This represents the ratio of cells co-expressing the candidate network markers to the total number of cells in the cell cluster; Represented in the initial cell cluster The CCP expressed the aforementioned candidate network markers The number of cells; Represented in the initial cell cluster The total number of cells in the body; S142. Screening of network identifiers; Based on the expressed proportion The network markers are screened, wherein the screened network markers must simultaneously meet the requirements of high accuracy and high specificity; The high accuracy rate is indicated in the context of the initial cell cluster. Chinese expression proportion No less than 80%; The high specificity corresponds to the initial cell cluster. The network markers mentioned above are the Three Sigma Law. ;

[0012] in, The network markers represent different initial cell clusters. The average percentage of expression, The network markers represent different initial cell clusters. The standard deviation of the expression proportion.

[0013] Furthermore, S2 specifically refers to: S21. The representation layer of the cell classification model specifically involves: constructing a cell graph structure, treating cells as graph nodes; S22. The aggregation layer of the cell classification model is specifically: a graph attention network is used to calculate attention weights and update node features; S23. The linear classification layer of the cell classification model is specifically: extracting feature representations to classify and predict cell types.

[0014] Furthermore, S21 specifically includes: S211. Construct a cell gene expression matrix using all the network markers described; S212. Calculate the cosine similarity between each cell, and connect the cells with a similarity greater than 0.75 to obtain the cell-cell adjacency matrix. S213, Introducing Hyperparameters To calculate the mixed adjacency matrix To balance the weights between the co-expression information of network markers and gene expression in the cells, and to construct a cell graph structure; the hybrid adjacency matrix for:

[0015] Where I is the identity matrix, It is an adjacency matrix. and In comparison, the introduction of hyperparameters It is a diagonal element.

[0016] Furthermore, S22 specifically includes: S221. Perform a linear transformation on each node in the aggregation layer; S222. An attention head is used to capture the structural information in the cell graph and consider the state of the neighboring nodes of each node to calculate the attention score. After activation function and normalization processing, the final attention weight is obtained. ; ; S223, Based on the attention weights The updated node representation is obtained by weighted summation of the features of the neighboring nodes. ;

[0017] in, Represents the node In the embedding representation at the k-th layer, ReLU is a non-linear activation function. Represents the weight matrix. Represents a node and The weight of the edge between them. Indicates the first Layer bias in a multilayer neural network.

[0018] Furthermore, S23 specifically includes: S231. Based on the linear classification layer, extract the feature representation to classify and predict the cell type. For each cell... Predict the probability that it belongs to each type of immune cell. ; ; S232. Use cross-entropy loss to predict the difference between class distribution and label, and calculate the cross-entropy loss function. Simultaneously, the cross-entropy loss function is adjusted by using inverse class frequencies plus logarithmic transformation. Category weights in ;

[0019]

[0020] in, The cells are The true category distribution; It is the class probability predicted by the model; The weight of cell category c is represented by totalNum; totalNum represents the total number of cells in the dataset. This represents the number of cells in cell category c; S233, The objective function of the cell classification model is: ; .

[0021] Furthermore, S3 specifically refers to: S31. Spatial data preprocessing: Quality control, standardization, and dimensionality reduction of spatial RNA-seq data; S32. Spatial map acquisition and regulatory relationship determination: Construct a shared nearest neighbor graph, divide the cell subpopulations, acquire the spatial map, and determine the regulatory relationship between the network markers; S33. Spatial Relationship Analysis: Calculate the distance between the cell clusters where a regulatory relationship exists / does not exist, verify the significance of the differences using a T-test, and infer the spatial relationship.

[0022] Furthermore, S32 specifically includes: S321. Determine the appropriate number of influencing components based on the ranking of the percentage of variance explained by each principal component; S322, Based on Euclidean distance in PCA space, using The nearest neighbor algorithm obtains the nearest neighbors around each cell. Calculate cell similarity among neighboring nodes. To construct a shared nearest neighbor graph; simultaneously, to iteratively calculate the modularity among different cell populations, and then to divide the cells into modules to obtain cell subpopulations. ;

[0023]

[0024] in, This is the number of nearest neighbor nodes used in the K-Nearest Neighbors (KNN) algorithm; the default value is 20. and Represents cells and Shared nearest neighbor node Rank among all neighboring nodes This represents the sum of all edge weights in a shared nearest neighbor graph (SNN graph). and They represent the cells respectively. and The sum of the weights of all connected edges; We used manual annotation methods to annotate immune cell subsets to obtain a spatial atlas of immune cells in lung adenocarcinoma. S323. Based on the lung adenocarcinoma gene regulatory network, given two network markers, determine whether there is a pair of genes in the gene regulatory network between the two network markers. If they exist, it indicates that there is a regulatory relationship between the two network markers; otherwise, there is no regulatory relationship.

[0025] On the other hand, the present invention also provides a lung adenocarcinoma immune cell subset network biomarker screening system for performing the aforementioned lung adenocarcinoma immune cell subset network biomarker screening method, the system comprising: A module for screening network biomarkers of immune cell subsets in lung adenocarcinoma: used to perform raw data preprocessing, cell-network biomarker matrix construction, KLoop clustering, and network biomarker screening; Cell classification module of graph attention network: used to verify the effectiveness of the network markers; The immune cell subset spatial relationship analysis module is used to perform spatial data preprocessing, spatial map acquisition and network marker regulatory relationship determination, and immune cell subset spatial relationship analysis.

[0026] Furthermore, the cell classification module of the graph attention network is configured to construct a cell graph structure, treating cells as graph nodes; use the graph attention network to calculate attention weights and update node features; and extract feature representations to classify and predict cell types.

[0027] Furthermore, the immune cell subpopulation spatial relationship analysis module is configured to perform quality control, standardization, and dimensionality reduction on spatial RNA-seq data; construct a shared nearest neighbor graph, divide the cell subpopulations, obtain the spatial map, determine the regulatory relationships between the network markers; and calculate the distance between cell clusters with / without regulatory relationships, verify the significance of the differences through a T-test, and infer the spatial relationships.

[0028] By adopting the above technical solution, the present invention has the following beneficial effects: The lung adenocarcinoma immune cell subpopulation network biomarker screening scheme of the present invention uses a density clustering algorithm to further subdivide large immune cell subpopulations into biologically significant local microenvironment subclusters based on the actual spatial distribution of cells, rather than simply treating all cells of the same type as homogeneous; this allows us to focus on smaller, potentially more homogeneous local regions. Simultaneously, by calculating the average expression level or activity score of network markers in each local microenvironment sub-cluster, the functional state of immune cells within that local region was directly quantified. As representatives of gene regulatory networks, the activity scores of network markers can more accurately reflect the complex functional state of cells, rather than the expression of a single gene. Furthermore, by combining spatial visualization technology, the quantified local functional status score is directly mapped onto the spatial atlas of tumor tissue, enabling researchers and clinicians to clearly identify which regions within the tumor exhibit different functional states of the same immune cell subset. This approach overcomes the limitations of existing technologies in capturing subtle functional differences of the same immune cell subset in different spatial microenvironments, providing more refined and instructive information for accurately identifying tumor local immune escape mechanisms, discovering new therapeutic targets, and developing personalized treatment strategies for specific tumor microenvironment regions. Attached Figure Description

[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0030] Figure 1 This is a flowchart of the lung adenocarcinoma immune cell subset network marker screening method of the present invention; Figure 2 for Figure 1 Flowchart of S1; Figure 3 for Figure 2 Diagram of the 01 matrix construction of cell-network markers in S12; Figure 4 This is a schematic diagram of the cell diagram structure in S21; Figure 5 This is a schematic diagram of the attention network framework in S22. Figure 6 for Figure 1 Graph of the algorithm for spatial relationship analysis of immune cell subsets in S3 lung adenocarcinoma; Figure 7 This is a flowchart of the process for determining the relationship between spatial map acquisition and network marker regulation in S32. Detailed Implementation

[0031] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Those skilled in the art should understand that the following specific embodiments or implementation methods are a series of optimized configurations listed to further explain the specific content of the invention. These configuration methods can be combined or used in conjunction with each other, unless the invention explicitly states that some or a specific embodiment or implementation method cannot be associated with or used in conjunction with other embodiments or implementation methods. Furthermore, the following specific embodiments or implementation methods are merely optimized configurations and are not intended to limit the scope of protection of the invention.

[0033] Currently, there are three core issues in the study of immune cell subsets in lung adenocarcinoma.

[0034] One approach is the screening of network biomarkers. Traditional methods for screening network biomarkers often use a single lung adenocarcinoma gene as a biomarker, which is insufficient to reflect the complex role of gene regulatory networks in the development and progression of lung adenocarcinoma.

[0035] Secondly, network biomarker computation and validation. The effectiveness of network biomarkers needs to be validated through specific application scenarios. As a fundamental step in scRNA-seq analysis, cell type identification directly depends on the biomarker's ability to characterize cell features. Therefore, the cell type identification process based on network biomarkers can be regarded as an important validation method for the effectiveness of network biomarkers.

[0036] In existing cell classification models, while machine learning or deep learning methods can improve classification accuracy by automatically learning data features, the black-box nature of the feature learning process makes it difficult to correlate the effects of specific biomarkers. Even with excellent classification results, the effectiveness of specific network biomarkers cannot be clearly proven, making them unreliable means of validating network biomarkers. In contrast, cell classification models constructed with the assistance of biomarkers can directly link classification results to the characterizing ability of biomarkers: if the model can accurately identify cell types, it proves that the biomarkers can effectively distinguish cell characteristics. Thirdly, spatial relationship analysis. Studying the spatial relationships of cells—that is, the interactions and positional relationships between cells and their microenvironment and other cells—is crucial for understanding tumor progression and immune responses.

[0037] In recent years, with the development of spatial RNA-seq, many researchers have begun to combine single-cell RNA-seq and spatial RNA-seq data to analyze the spatial relationships of lung adenocarcinoma cells. While the aforementioned techniques can effectively characterize cellular spatial features, they are complex and costly.

[0038] Based on this, this application provides a screening scheme for immune cell subset network biomarkers in lung adenocarcinoma. Please refer to [link / reference]. Figure 1 As shown, the method for screening network biomarkers of immune cell subsets in lung adenocarcinoma includes the following steps: S1. Construct a screening process for network biomarkers of immune cell subsets in lung adenocarcinoma based on the KLoop algorithm; The network biomarker screening process includes raw data preprocessing, cell-network biomarker matrix construction, KLoop clustering, and network biomarker screening. S2. Construct a cell classification model based on graph attention network (GAT); The cell classification model consists of three parts: a representation layer, a aggregation layer, and a linear classification layer, and is used to validate the effectiveness of network markers. S3, Spatial Relationship Analysis of Immune Cell Subpopulations; Spatial relationship analysis of immune cell subsets includes spatial data preprocessing, spatial map acquisition and determination of network marker regulatory relationships, and spatial relationship analysis of immune cell subsets.

[0039] The following instruction manual will provide further explanation for S1 to S3.

[0040] Please see Figure 2 As shown, S1 is specifically: S11. Raw data preprocessing: Load and preprocess the annotated lung adenocarcinoma immune cell gene expression data and candidate network biomarkers; S12. Constructing a cell-network biomarker 01 matrix: The lung adenocarcinoma immune cell subset is divided into multiple initial cell clusters, and the matrix elements of each initial cell cluster are extracted. The candidate network biomarkers expressed by the lung adenocarcinoma immune cell subset are characterized by the matrix elements. S13, KLoop Clustering: An iterative clustering algorithm is used to segment each lung adenocarcinoma immune cell subpopulation until the initial cell clusters after segmentation meet the preset conditions, and the final cell clusters are extracted. S14. Network biomarker screening: Calculate the expression percentage of each candidate network biomarker, and screen candidate network biomarkers that simultaneously meet the requirements of high accuracy and high specificity as the final network biomarkers.

[0041] S11 specifically involves loading annotated lung adenocarcinoma immune cell gene expression data and candidate network markers, and grouping the quality-controlled and standardized lung adenocarcinoma immune cell gene expression data by cell type.

[0042] Please see Figure 3 As shown, S12 specifically refers to the cell-network marker 01 matrix, which divides the immune cell subset into several cell clusters. The rows of the cell-network marker 01 matrix correspond to different cells in the immune cell subset, and the columns of the cell-network marker 01 matrix represent the network markers corresponding to the immune cell subset.

[0043] If cells in an immune cell subset co-express a certain network marker, the element at the corresponding position is assigned a value of 1; otherwise, it is assigned a value of 0. This is used to find the network marker that accurately represents the cell cluster. Co-expression means that the expression intensity of all genes in the network in the cell is greater than 0.

[0044] S13 specifically refers to: S131, Given a set containing A dataset of data points Each data point have Each feature dimension, i.e. ; S132, Random selection The center of each initial cell cluster ,in Indicates the first The center position vector of each initial cell cluster; that is ; Calculate each data point With the center of each initial cell cluster distance and data points Assigned to the nearest center In the initial cell cluster; where, the distance It is represented as: ; S133, Data Point After allocation, based on the data points in each initial cell cluster Update the center of the initial cell cluster Location; center of the initial cell cluster Update Center Update Center Characterized as:

[0045] in, Indicates the first Data points contained in an initial cell cluster gather, This represents the data point in the corresponding initial cell cluster. Quantity; S134. Apply the k-means method to each initial cell cluster to divide it into two smaller clusters. Repeat this process until the resulting cell clusters meet the following conditions: the number of cells in the cluster is ≥ 1% of the total number of cells, and there are effective network markers in the cluster. In S134, before applying the k-means method to each initial cell cluster to divide it into two smaller clusters, the steps of allocating data points in S132 and updating the cluster center in S133 need to be repeated continuously until the update change of the cluster center is less than the preset threshold or the maximum number of iterations is reached; then, the k-means method is applied to each cell cluster to divide it into two smaller clusters, and the clusters satisfy the above conditions.

[0046] Furthermore, S14 specifically refers to: S141. Calculate the expression percentage of candidate network biomarkers. ;

[0047] in, This represents the ratio of cells co-expressing candidate network markers to the total number of cells in a cell cluster; Represents the initial cell cluster CCP expresses candidate online logo The number of cells; Represents the initial cell cluster The total number of cells in the body; S142. Screening of network identifiers; Based on expression percentage Screening of network identifiers is carried out, in which the screened network identifiers must simultaneously meet the requirements of high accuracy and high specificity; High accuracy indicates accuracy in the corresponding initial cell clusters. Chinese expression proportion No less than 80%; High specificity indicates the corresponding initial cell cluster Network Icons Three Sigma Law ;

[0048] in, Indicating network markers in different initial cell clusters The average percentage of expression, Indicating network markers in different initial cell clusters The standard deviation of the expression proportion.

[0049] In this application, S2 is used to verify the direct association between biomarkers and cell types. In S2, cells are used as nodes of a graph neural network to construct a cell classification model based on a graph attention network (GAT). The graph attention network is used to learn cell representations and supervised learning is performed on lung adenocarcinoma immune cell data with known cell types to achieve prediction of cell types in new datasets, thereby verifying the effectiveness of the network biomarkers.

[0050] Specifically, S2 is: S21. The representation layer of the cell classification model specifically involves constructing a cell graph structure, treating cells as graph nodes. S22. The aggregation layer of the cell classification model specifically employs a graph attention network to calculate attention weights and update node features. S23. The linear classification layer of the cell classification model is specifically: extracting feature representations to classify and predict cell types.

[0051] Please see Figure 4 As shown, S21 treats cells as graph nodes and determines whether there are connecting edges between cells based on their similarity; specifically, S21 is as follows: S211. Construct a cell gene expression matrix using all network markers; S212. Calculate the cosine similarity between each cell and connect edges between cells with a similarity greater than 0.75 to obtain the cell-cell adjacency matrix. S213, Introducing Hyperparameters To calculate the mixed adjacency matrix To balance the weights between co-expression information of network markers and gene expression in cells, and to construct a cell graph structure; a hybrid adjacency matrix was used. for:

[0052] Where I is the identity matrix, It is an adjacency matrix. and In comparison, introducing hyperparameters The diagonal elements are used, and the elements that are 1 in the original cell-network marker 01 matrix are transformed into... Elements that were previously 0 remain unchanged.

[0053] Please see Figure 5 As shown, S22 employs a graph attention network. For each node in the aggregation layer, it calculates the attention weights between itself and its neighboring nodes by considering the states of its neighboring nodes, and then updates the features of the cell node.

[0054] In this application, S22 specifically refers to: S221. Perform a linear transformation on each node in the aggregation layer; S222. An attention head is used to capture structural information in the cell graph and consider the states of each node's neighboring nodes to calculate the attention score. After activation function and normalization processing, the final attention weight is obtained. ; ; in, For neighboring nodes With the central node The original correlation score between them; As the central node All neighboring nodes (including) The sum of the original association scores (inclusive); LeakyReLU is the activation function; k represents the index variable used to traverse nodes. All neighboring nodes, referring to the neighbor set in general. any one of the members, Represents shared, learnable parameters. , Let represent the original feature vectors of nodes i and j. W represents the collective characteristic vectors of all other neighboring cells in the microenvironment surrounding the central cell; W represents the weight matrix.

[0055] S223, Based on attention weights The updated node representation is obtained by weighted summation of the features of neighboring nodes. ;

[0056] in, Represents a node In the embedding representation at the k-th layer, ReLU is a non-linear activation function. Represents the weight matrix. Represents a node and The weight of the edge between them. Indicates the first Layer bias in a multilayer neural network.

[0057] Furthermore, S23 specifically refers to: S231. Based on the feature representation extracted from the linear classification layer, cell type classification prediction is performed. For each cell... Predict the probability that it belongs to each type of immune cell. ; ; S232. Use cross-entropy loss to predict the difference between class distribution and label, and calculate the cross-entropy loss function. Simultaneously, the cross-entropy loss function is adjusted by using inverse class frequencies plus logarithmic transformation. Category weights in ;

[0058]

[0059] in, It is a cell The true category distribution; It is the class probability predicted by the model; The weight of cell category c is represented by totalNum; totalNum represents the total number of cells in the dataset. This represents the number of cells in cell category c; S233, The objective function of the cell classification model is: ; .

[0060] Please see Figure 6 As shown, S3 specifically refers to: S31. Spatial data preprocessing: Quality control, standardization, and dimensionality reduction of spatial RNA-seq data; S32. Spatial map acquisition and regulatory relationship determination: Construct a shared nearest neighbor graph, divide cell subpopulations, acquire spatial maps, and determine the regulatory relationships between network markers; S33. Spatial Relationship Analysis: Calculate the distance between cell clusters with / without regulatory relationships, verify the significance of the differences using a T-test, and infer the spatial relationships.

[0061] S31 specifically involves using raw spatial RNA-seq data from lung adenocarcinoma as input, and filtering cells to obtain cell data that meets quality standards based on indicator 1: the number of genes detected in a single cell is greater than 50 and less than 700; and indicator 2: the proportion of mitochondrial genome read is less than 20%.

[0062] Furthermore, the raw spatial RNA-seq data is standardized to obtain the correct relative gene expression abundance among cells, thereby avoiding sampling effects; the standardized values ​​of the raw spatial RNA-seq data are... ;

[0063] in, Represents raw spatial RNA-seq data; This represents the sum of gene counts in a single cell; This is the scaling factor, which is a constant value of 10000.

[0064] Furthermore, gene expression was scaled using the normalized values ​​of spatial RNA-seq data to prevent genes with high expression values ​​from masking genes with low but significant expression values; gene expression scaling. for:

[0065] in, This represents the standardized values ​​of spatial RNA-seq data; This represents the mean value of the gene across all cells; This represents the standard deviation of the gene across all cells.

[0066] Finally, principal component analysis was used to reduce the dimensionality of gene expression data in order to extract the top 30 influencing components that are closely related to cell type.

[0067] Please see Figure 7 As shown, S32 is specifically used for: S321. Determine the appropriate number of influencing components based on the ranking of the percentage of variance explained by each principal component; S322, Based on Euclidean distance in PCA space, using The nearest neighbor algorithm finds the nearest neighbors around each cell. Calculate cell similarity among neighboring nodes. To construct a shared nearest neighbor graph, the modularity between different cell populations is iteratively calculated, and then modules are partitioned to obtain cell subpopulations. ;

[0068]

[0069] in, This is the number of nearest neighbor nodes used in the K-Nearest Neighbors (KNN) algorithm; the default value is 20. and Represents cells and Shared nearest neighbor node Rank among all neighboring nodes This represents the sum of all edge weights in a shared nearest neighbor graph (SNN graph). and They represent the cells respectively. and The sum of the weights of all connected edges; We used manual annotation methods to annotate immune cell subsets to obtain a spatial atlas of immune cells in lung adenocarcinoma. S323. Based on the lung adenocarcinoma gene regulatory network, given two network markers, determine whether there is a pair of genes in the gene regulatory network between the two network markers. If they exist, it indicates that there is a regulatory relationship between the two network markers; otherwise, there is no regulatory relationship.

[0070] Furthermore, S4 specifically involves: dividing the immune cell subsets into cell clusters based on network markers; and using Euclidean distance to calculate the distance between cell clusters corresponding to network markers that do not have a regulatory relationship (group0).

[0071] Furthermore, the distance between cell clusters corresponding to network markers with regulatory relationships was calculated (group1); a T-test was used to statistically verify whether the mean of group0 being greater than the mean of group1 was statistically significant.

[0072] Finally, if the mean of group0 is significantly greater than that of group1, it proves that there is a correlation between the regulatory relationship between network markers and the spatial distance between immune cell subsets; otherwise, there is no correlation between the regulatory relationship between network markers and the spatial distance between immune cell subsets, thus enabling the analysis of the spatial relationship between immune cell subsets.

[0073] The lung adenocarcinoma immune cell subset network biomarker screening method of this invention employs the KLoop algorithm to ensure that the expression rate of the screened biomarkers in the corresponding cell clusters is higher than 80%, and conforms to the three Sigma law, thus possessing both specificity and accuracy. Simultaneously, by establishing a cell classification model based on graph attention networks, at the cell type level, it significantly outperforms traditional cell classification models such as MYGAT, SingleR, ScPred, and scmap in four evaluation metrics: accuracy (0.974), precision (0.973), recall (0.975), and F1 score (0.974). Compared with the best performance of traditional models, it improves by approximately 3.83%, 3.84%, 3.83%, and 4.3% in accuracy, precision, recall, and F1 score, respectively. Furthermore, spatial relationship analysis of immune cell subsets: the distance between cell clusters with / without regulatory relationships was statistically verified by T test (P<0.05), which improved the reliability of spatial relationship inference results to 99.5%. This breakthrough overcomes the limitation of insufficient resolution of traditional spatial RNA-seq data and can more accurately quantify the spatial distance association between cell clusters, providing a more efficient and reliable quantitative analysis method for the study of cell interaction mechanisms in the tumor microenvironment.

[0074] Furthermore, the solution provided by this invention can also be implemented through a modular system approach to execute the lung adenocarcinoma immune cell subset network biomarker screening method provided in the above embodiments.

[0075] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for screening network biomarkers of immune cell subsets in lung adenocarcinoma, characterized in that, Includes the following steps: S1. Construct a screening process for network biomarkers of immune cell subsets in lung adenocarcinoma based on the KLoop algorithm; The network biomarker screening process includes raw data preprocessing, cell-network biomarker matrix construction, KLoop clustering, and network biomarker screening. S2. Construct a cell classification model based on graph attention network; The cell classification model consists of three parts: a representation layer, a aggregation layer, and a linear classification layer, and is used to verify the effectiveness of the network markers. S3, Spatial Relationship Analysis of Immune Cell Subpopulations; The spatial relationship analysis of immune cell subsets includes spatial data preprocessing, spatial map acquisition and determination of network marker regulatory relationships, and spatial relationship analysis of immune cell subsets.

2. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 1, characterized in that, Specifically, S1 is: S11. Raw data preprocessing: Load and preprocess the annotated lung adenocarcinoma immune cell gene expression data and candidate network biomarkers; S12. Constructing a cell-network marker 01 matrix: Divide the lung adenocarcinoma immune cell subset into multiple initial cell clusters, extract the matrix elements of each initial cell cluster, and characterize the candidate network markers expressed by the lung adenocarcinoma immune cell subset through the matrix elements. S13, KLoop Clustering: An iterative clustering algorithm is used to segment each lung adenocarcinoma immune cell subpopulation until the initial cell clusters after segmentation meet the preset conditions, and the final cell clusters are extracted. S14. Network marker screening: Calculate the expression ratio of each candidate network marker, and screen the candidate network markers that simultaneously meet the requirements of high accuracy and high specificity as the final network markers.

3. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 2, characterized in that, Specifically, S13 is: S131, Given a set containing A dataset of data points Each data point have Each feature dimension, i.e. ; S132, Random Selection The center of each initial cell cluster ,in Indicates the first The center position vectors of the initial cell clusters; Calculate each of the data points With the center of each of the initial cell clusters distance and the data points Assigned to the nearest center In the initial cell cluster where it is located; wherein, the distance It is represented as: ; S133, the data points After allocation, based on the data points in each of the initial cell clusters Update the center of the initial cell cluster The location; the center of the initial cell cluster. Update Center The update center Characterized as: in, Indicates the first The data points contained in the initial cell clusters gather, This represents the data point corresponding to the initial cell cluster. Quantity; S134. Apply the k-means method to each initial cell cluster to divide it into two smaller clusters. Repeat this process until the resulting cell clusters satisfy the following conditions: the number of cells in the cluster is ≥ 1% of the total number of cells, and there are valid network markers in the cluster.

4. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 2, characterized in that, Specifically, S14 is: S141. Calculate the expression percentage of the candidate network markers. ; in, This represents the ratio of cells co-expressing the candidate network markers to the total number of cells in the cell cluster; Represented in the initial cell cluster The CCP expressed the aforementioned candidate network markers The number of cells; Represented in the initial cell cluster The total number of cells in the body; S142. Screening of network identifiers; Based on the expressed proportion The network markers are screened, wherein the screened network markers must simultaneously meet the requirements of high accuracy and high specificity; The high accuracy rate is indicated in the context of the initial cell cluster. Chinese expression proportion No less than 80% The high specificity corresponds to the initial cell cluster. The aforementioned network markers, Three Sigma Law ; in, The network markers represent different initial cell clusters. The average percentage of expression, The network markers represent different initial cell clusters. The standard deviation of the expression proportion.

5. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 1, characterized in that, Specifically, S2 is: S21. The representation layer of the cell classification model specifically involves: constructing a cell graph structure, treating cells as graph nodes; S22. The aggregation layer of the cell classification model is specifically: a graph attention network is used to calculate attention weights and update node features; S23. The linear classification layer of the cell classification model is specifically: extracting feature representations to classify and predict cell types.

6. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 5, characterized in that, Specifically, S21 is: S211. Construct a cell gene expression matrix using all the network markers described; S212. Calculate the cosine similarity between each cell, and connect the cells with a similarity greater than 0.75 to obtain the cell-cell adjacency matrix. S213, Introducing Hyperparameters To calculate the mixed adjacency matrix To balance the weights between the co-expression information of network markers and gene expression in the cells, and to construct a cell graph structure; the hybrid adjacency matrix for: Where I is the identity matrix, It is an adjacency matrix. and In comparison, the introduction of hyperparameters It is a diagonal element.

7. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 5, characterized in that, Specifically, S22 is: S221. Perform a linear transformation on each node in the aggregation layer; S222. An attention head is used to capture the structural information in the cell graph and consider the state of the neighboring nodes of each node to calculate the attention score. After activation function and normalization processing, the final attention weight is obtained. ; ; S223, Based on the attention weights The updated node representation is obtained by weighted summation of the features of the neighboring nodes. ; in, Represents the node In the embedding representation at the k-th layer, ReLU is a non-linear activation function. Represents the weight matrix. Represents a node and The weight of the edge between them. Indicates the first Layer bias in a multilayer neural network.

8. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 5, characterized in that, Specifically, S23 is: S231. Based on the linear classification layer, extract the feature representation to classify and predict the cell type. For each cell... Predict the probability that it belongs to each type of immune cell. ; ; S232. Use cross-entropy loss to predict the difference between class distribution and label, and calculate the cross-entropy loss function. Simultaneously, the cross-entropy loss function is adjusted by using inverse class frequencies plus logarithmic transformation. Category weights in ; in, The cells are The true category distribution; It is the class probability predicted by the model; The weight of cell category c is represented by totalNum; totalNum represents the total number of cells in the dataset. This represents the number of cells in cell category c; S233, The objective function of the cell classification model is: ; 。 9. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 1, characterized in that, Specifically, S3 is: S31. Spatial data preprocessing: Quality control, standardization, and dimensionality reduction of spatial RNA-seq data; S32. Spatial map acquisition and regulatory relationship determination: Construct a shared nearest neighbor graph, divide the cell subpopulations, acquire the spatial map, and determine the regulatory relationship between the network markers; S33. Spatial Relationship Analysis: Calculate the distance between the cell clusters where a regulatory relationship exists / does not exist, verify the significance of the differences using a T-test, and infer the spatial relationship.

10. The method for screening lung adenocarcinoma immune cell subset network markers according to claim 9, characterized in that, Specifically, S32 is as follows: S321. Determine the appropriate number of influencing components based on the ranking of the percentage of variance explained by each principal component; S322, Based on Euclidean distance in PCA space, using The nearest neighbor algorithm obtains the nearest neighbors around each cell. Calculate cell similarity among neighboring nodes. To construct a shared nearest neighbor graph; simultaneously, to iteratively calculate the modularity among different cell populations, and then to divide the cells into modules to obtain cell subpopulations. ; in, This represents the number of nearest neighbors found using the K-nearest neighbor algorithm. and Represents cells and Shared nearest neighbor node Rank among all neighboring nodes This represents the sum of the edge weights in the shared nearest neighbor graph. and They represent the cells respectively. and The sum of the weights of all connected edges; We used manual annotation methods to annotate immune cell subsets to obtain a spatial atlas of immune cells in lung adenocarcinoma. S323. Based on the lung adenocarcinoma gene regulatory network, given two network markers, determine whether there is a pair of genes in the gene regulatory network between the two network markers. If they exist, it indicates that there is a regulatory relationship between the two network markers; otherwise, there is no regulatory relationship.