Method, device and equipment for screening cell population sensitive drugs and storage medium

By integrating single-cell and spatial transcriptome data, and using clustering algorithms and cell marker genes for repeated refinement, feature genes are extracted and feature association analysis is performed. This solves the problems of unreasonable data structure division and insufficient feature association in existing technologies, and achieves more accurate drug screening.

CN121789757APending Publication Date: 2026-04-03SHENZHEN TRADITIONAL CHINESE MEDICINE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing integrated analysis of single-cell and spatial transcriptome data suffers from unreasonable data structure division and lack of characteristic correlations between omics, resulting in overly data-driven cell clustering and regional division, and one-sided information on characteristic gene responses.

Method used

By acquiring single-cell transcriptome and spatial transcriptome data, performing preliminary segmentation and noise reduction, and then using a preset clustering algorithm for grouping and division, combined with cell marker genes for repeated clustering correction, cell and regional characteristic genes are extracted, feature association analysis is performed, and sensitive drugs that can target specific tissue regions are screened.

Benefits of technology

This improves the accuracy of data segmentation and the reliability of characteristic genes, ensuring the precision of drug screening. It can better reflect the characteristics of cell clusters and tissue regions, thus enhancing the reliability of drug screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789757A_ABST
    Figure CN121789757A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and equipment for screening cell population sensitive drugs and a storage medium, and the method comprises the steps: obtaining single cell transcriptome data and space transcriptome data, and dividing the data to obtain cell population groups and tissue region groups; respectively extracting cell characteristic genes and regional characteristic genes from the cell population group and the tissue region group, and carrying out characteristic correlation analysis to obtain common characteristic genes; and based on the common characteristic gene, performing biological calculation in combination with pre-acquired cell line drug intervention and expression data, and screening sensitive drugs capable of targeting a cell population in a specific tissue area corresponding to the common characteristic gene. According to the method, the single cell transcriptome data and the space transcriptome data can be integrated, feature association is carried out, the common feature genes with consistent changes between omics data under the same condition are extracted, then sensitive drugs are screened from cell populations in a specific tissue area on the basis of the common feature genes, and the screening result is high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics, and in particular to a method, apparatus, device, and storage medium for screening cell population-sensitive drugs. Background Technology

[0002] The spatial distribution of cell populations within a tissue region is particularly important for their function. For example, cell populations located at the periphery of tumor tissue and those located within the tumor tissue exhibit significantly different biological behaviors and cellular functions. Targeted drug intervention in specific tissue regions holds promise for controlling disease progression, and integrating single-cell transcriptomics and spatial transcriptomics can capture the distribution of cell populations in different tissue regions. Based on this, establishing drug screening methods has significant clinical application value.

[0003] Single-cell transcriptome data provides single-cell precision unattainable by traditional techniques, while spatial transcriptome data provides spatial localization of cells within tissue regions. However, single-cell transcriptome data loses the spatial location information of each individual cell, and existing spatial transcriptome data is still insufficient to achieve spatial localization with the precision of a single cell within a tissue region. Therefore, integrating single-cell and spatial transcriptome data can enable the characterization of the spatial distribution of cell populations within a tissue region.

[0004] Currently, integrated single-cell and spatial transcriptome analyses are broadly categorized into three types: 1) Deconvolution: This involves inferring the cell type composition captured in each well of the spatial transcriptome data. For example, SPOTlight uses a non-negative matrix method for decomposition and inference. 2) Mapping: This specifies which wells in the spatial transcriptome data belong to which cell type. For example, Seurat uses tag transfer to anchor cell type attribution based on the similarity of characteristic gene expression. 3) Spatial receptor-ligand interactions: This involves combining single-cell transcriptome data with real physical distances to infer receptor-ligand pairs that interact between cells.

[0005] However, existing integrated analysis of single-cell and spatial transcriptome data has certain flaws, mainly including: (1) the clustering of cell populations in single-cell transcriptome data and the division of tissue regions in spatial transcriptome data do not take into account the information of the correspondence between cells and tissue structures, resulting in cell clustering and region division being too data-driven; (2) the extraction of characteristic genes from each omics data only considers single-omics data, the information of characteristic gene responses is relatively one-sided, and there is a lack of intrinsic correlation between different omics. Summary of the Invention

[0006] In view of this, this application provides a method, apparatus, device and storage medium for screening cell population-sensitive drugs to solve the problems of unreasonable data structure division and lack of feature correlation between omics in existing single-cell and spatial transcription data integration methods.

[0007] To address the aforementioned technical problems, this application provides a method for screening cell population-sensitive drugs, comprising: acquiring single-cell transcriptome data and spatial transcriptome data and performing preliminary segmentation to obtain preliminary cell population groups and preliminary tissue region groups; using cell marker genes, repeatedly clustering and correcting the preliminary cell population groups and preliminary tissue region groups to obtain final cell population groups and final tissue region groups; extracting cell characteristic genes and region characteristic genes from the final cell population groups and final tissue region groups, and performing feature association analysis to obtain common characteristic genes that exhibit the same expression pattern among cell populations and their localized tissue regions; and based on the common characteristic genes, performing biological computation in conjunction with pre-acquired cell line drug intervention and expression data to screen for sensitive drugs that can target cell populations within specific tissue regions corresponding to the common characteristic genes.

[0008] As a further improvement of this application, the acquisition of single-cell transcriptome data and spatial transcriptome data and their preliminary division to obtain preliminary cell population grouping and preliminary tissue region grouping include: acquiring single-cell transcriptome data and spatial transcriptome data and performing noise reduction; using a preset clustering algorithm to cluster the single-cell transcriptome data to obtain preliminary cell population grouping, and using a preset clustering algorithm to divide the spatial transcriptome data to obtain preliminary tissue region grouping.

[0009] As a further improvement of this application, cell marker genes are used to repeatedly cluster and correct the preliminary cell population grouping and preliminary tissue region grouping to obtain the final cell population grouping and final tissue region grouping. This includes: extracting preliminary cell characteristic genes from each preliminary cell population grouping based on differences in cell gene expression, and extracting preliminary region characteristic genes from each preliminary tissue region grouping based on differences in gene expression between different tissue regions; classifying cell populations by cell type based on cell marker genes of different cell types; confirming the cell type corresponding to each preliminary cell population based on the overlap between preliminary cell characteristic genes and each type of marker gene, and confirming the region type corresponding to each preliminary tissue region based on the overlap between preliminary region characteristic genes and each type of marker gene; obtaining the tissue region location of all preliminary cell populations on spatial transcriptome data; obtaining the number of region types distributed within the same cell population and the number of cell populations distributed within the same tissue region based on the tissue region location; further dividing the same cell population based on the number of region types, and further dividing the same tissue region based on the number of cell populations, until each tissue region corresponds one-to-one with each cell population, thus obtaining the final cell population grouping and final tissue region grouping.

[0010] As a further improvement of this application, cell-specific genes and region-specific genes are extracted from the final cell population grouping and the final tissue region grouping, respectively, and feature association analysis is performed to obtain common feature genes that exhibit the same expression pattern among cell populations and their corresponding tissue regions. This includes: extracting final cell-specific genes and final region-specific genes from each final cell population and each final tissue region, and taking the union of the two to obtain a feature gene set; selecting target feature genes one by one from the gene set, and calculating the gene expression value of the target feature gene under two different preset conditions, where the two different preset conditions include two different cell populations or two different tissues. The region and gene expression values ​​include the first gene expression value of the target characteristic gene in a specific cell population under a preset condition and the second gene expression value in a tissue region where the specific cell population is located, as well as the third gene expression value of the target characteristic gene in another specific cell population under another preset condition and the fourth gene expression value in a tissue region where the target characteristic gene is located in another specific cell population. Based on the gene expression values, the expression difference of the target characteristic gene between two different preset conditions is confirmed. When the expression difference confirms that the expression pattern of the target characteristic gene is consistent in the single-cell transcriptome and the spatial transcriptome, the target characteristic gene is confirmed as a common characteristic gene between the two different preset conditions.

[0011] As a further improvement of this application, based on common characteristic genes, and combined with pre-acquired cell line drug intervention and expression data, biological computation is performed to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to common characteristic genes. This includes: identifying all target common characteristic genes corresponding to cell populations in specific tissue regions; acquiring cell line drug intervention data and gene expression data; performing correlation analysis on the drug intervention data of all candidate drugs and the gene expression data of the target common characteristic genes in all cell lines to identify candidate drugs sensitive to the expression of the target common characteristic genes, thus obtaining a first drug candidate set; using the activity value of the target common characteristic genes in all cell lines to identify target cell lines with high activity; analyzing the drug intervention data corresponding to the target cell lines to identify candidate drugs sensitive to the expression of the target cell lines, thus obtaining a second drug candidate set; and taking the intersection of the first and second drug candidate sets to obtain sensitive drugs that can target cell populations in specific tissue regions.

[0012] As a further improvement of this application, the method of identifying target cell lines with high activity by utilizing the activity values ​​of target common characteristic genes in all cell lines includes: calculating the average expression value of the target common characteristic gene in each cell type using the gene expression value of each target common characteristic gene in each cell line; calculating the activity value of each target common characteristic gene in all cell lines using the average expression value of each target common characteristic gene and the gene expression value in each cell type; constructing an activity value matrix of target common characteristic genes and cell lines based on the activity values; clustering the activity value matrix to identify the cluster with the highest activity value, and identifying all cell lines in the cluster as target cell lines.

[0013] As a further improvement of this application, the drug intervention data corresponding to the target cell lines are analyzed to identify candidate drugs that are sensitive to the expression of the target cell lines, thereby obtaining a second drug candidate set. This includes: obtaining bioavailability data from the drug intervention data corresponding to all target cell lines; arranging all bioavailability data in descending order from high to low; confirming whether the arrangement of all bioavailability data is continuous based on the fluctuation range between all adjacent bioavailability data; if continuous, using the median of all bioavailability data as the final bioavailability data; if discontinuous, performing statistical distribution on all bioavailability data and using the result of the statistical distribution as the final bioavailability data; and confirming whether the candidate drugs are sensitive to the expression of the target cell lines based on the final bioavailability data, thereby obtaining a second drug candidate set.

[0014] To address the aforementioned technical problems, another technical solution adopted in this application is: providing an apparatus for screening cell population-sensitive drugs, comprising: an acquisition module for acquiring single-cell transcriptome data and spatial transcriptome data and performing preliminary segmentation to obtain preliminary cell population groups and preliminary tissue region groups; a grouping module for repeatedly clustering and correcting the preliminary cell population groups and preliminary tissue region groups using cell marker genes to obtain final cell population groups and final tissue region groups; a feature association module for extracting cell characteristic genes and region characteristic genes from the final cell population groups and final tissue region groups, and performing feature association analysis to obtain common characteristic genes that exhibit the same expression pattern among cell populations and their corresponding tissue regions; and a drug screening module for performing biological computation based on the common characteristic genes and combined with pre-acquired cell line drug intervention and expression data to screen for sensitive drugs that can target cell populations within specific tissue regions corresponding to the common characteristic genes.

[0015] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a computer device, the computer device including a processor and a memory coupled to the processor, the memory storing program instructions, and when the program instructions are executed by the processor, causing the processor to perform the steps of the method for screening cell population sensitive drugs as described above.

[0016] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a storage medium storing program instructions capable of implementing the method for screening cell population sensitive drugs as described above.

[0017] The beneficial effects of this application are as follows: The method for screening cell population-sensitive drugs in this application first obtains the global structure of two omics data. It adopts both the overall structure and characteristics of the single-cell transcriptome data and, with the help of the structure and characteristics of the two omics data, interactively optimizes the internal structure of the single-cell transcriptome data and extracts features, maximizing the reduction of noise information. This ensures that the data segmentation can reflect the characteristics of cell populations and tissue regions. Furthermore, in integrating single-cell transcriptome data and spatial transcriptome data, the two omics are integrated to perform feature association and extract common feature genes that change consistently between omics data under the same conditions. This makes the integration of single-cell transcriptome data and spatial transcriptome data better, and the data segmentation can reflect the characteristics of cell populations and tissue regions. Then, based on these common feature genes, sensitive drugs that can target cell populations in specific tissue regions corresponding to the common feature genes are screened, making the screened sensitive drugs more reliable. Attached Figure Description

[0018] Figure 1 This is a schematic flowchart of a method for screening cell population-sensitive drugs according to an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram illustrating the epithelial cell population segmentation method for screening cell population-sensitive drugs according to an embodiment of the present invention;

[0020] Figure 3 This is a schematic diagram of the tissue region distribution of epithelial cells in spatial transcriptome data, as shown in the original literature;

[0021] Figure 4 This is a schematic diagram of the spatial localization of different epithelial cell populations in tissue regions identified by the method for screening cell population-sensitive drugs according to an embodiment of the present invention. Figure 5 This is a schematic diagram of the functional modules of the device for screening cell population-sensitive drugs according to an embodiment of the present invention;

[0022] Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0023] Figure 7 This is a schematic diagram of the structure of the storage medium according to an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0025] The terms "first," "second," and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0026] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0027] Figure 1 This is a schematic flowchart of a method for screening cell population-sensitive drugs according to an embodiment of the present invention. It should be noted that if substantially the same results are obtained, the method of the present invention is not necessarily identical. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, the method for screening cell population-sensitive drugs includes the following steps:

[0028] Step S101: Obtain single-cell transcriptome data and spatial transcriptome data and perform preliminary segmentation to obtain preliminary cell population grouping and preliminary tissue region grouping.

[0029] Specifically, after obtaining single-cell transcriptome data and spatial transcriptome data, the single-cell transcriptome data and spatial transcriptome data are processed, and the processed data of the two omics are divided to obtain preliminary cell population grouping and preliminary tissue region grouping.

[0030] Furthermore, step S101 specifically includes:

[0031] 1. Acquire single-cell transcriptome data and spatial transcriptome data and perform noise reduction.

[0032] Specifically, after obtaining single-cell transcriptome and spatial transcriptome data, using R and Python environments, conventional software (such as Seurat) and custom programming, noise data caused by library construction was removed, including double or multi-cell, free contaminating RNA, and mitochondrial contamination. For mitochondrial contamination, the proportion of mitochondrial genes removed was considered to be higher than 20%.

[0033] 2. Use a preset clustering algorithm to group single-cell transcriptome data to obtain preliminary cell population groups, and use the preset clustering algorithm to divide spatial transcriptome data to obtain preliminary tissue region groups.

[0034] Specifically, for cell populations, this embodiment uses Seurat software or unsupervised clustering algorithms (such as NMF) to group cells and obtain the general structure of cell composition in single-cell transcriptome data; for tissue regions, this embodiment uses Seurat software or unsupervised clustering algorithms (such as NMF) to divide tissue regions and obtain the general structure of tissue regions in spatial transcriptome data.

[0035] Step S102: Using cell marker genes, perform repeated clustering corrections on the preliminary cell population grouping and preliminary tissue region grouping to obtain the final cell population grouping and final tissue region grouping.

[0036] Specifically, after obtaining preliminary cell population groups and preliminary tissue region groups, the cell populations are classified based on cell marker genes of different cell types. Then, the preliminary cell population groups and preliminary tissue region groups are repeatedly clustered and corrected, thereby interactively optimizing the internal structure of single-omics data and extracting features, minimizing noise information, and ensuring that the data division can reflect the characteristics of cell groups and tissue regions.

[0037] Furthermore, step S102 specifically includes:

[0038] 1. Preliminary cell characteristic genes were extracted from each preliminary cell population based on differences in cell gene expression, and preliminary regional characteristic genes were extracted from each preliminary tissue region based on differences in gene expression between different tissue regions.

[0039] 2. Based on cell marker genes of different cell types, classify cell populations by cell type.

[0040] 3. Based on the overlap between preliminary cell characteristic genes and each type of marker gene, determine the cell type corresponding to each preliminary cell population, and based on the overlap between preliminary regional characteristic genes and each type of marker gene, determine the region type corresponding to each preliminary tissue region.

[0041] Specifically, in this embodiment, preliminary cell characteristic genes are compared with marker genes for each cell type to obtain the overlap between the preliminary cell characteristic genes and each type of marker gene. When the overlap is greater than or equal to a preset overlap, the preliminary cell population corresponding to the preliminary cell characteristic gene is classified as the corresponding cell type. If there is no overlap, the preliminary cell population is classified as an unknown cell type. Similarly, preliminary region characteristic genes are processed and classified in the same way. Thus, the preliminary cell population is divided into M types, and the preliminary tissue regions are divided into N types. The preset overlap is preferably set to 2.

[0042] 4. Obtain the tissue region localization of all preliminary cell populations on spatial transcriptome data.

[0043] 5. Based on tissue region localization, obtain the number of regional types of the same cell population distribution, and obtain the number of cell populations distributed within the same tissue region.

[0044] 6. Based on the number of region types, the same cell population is further subdivided, and based on the number of cell populations, the same tissue region is further subdivided until each tissue region corresponds one-to-one with each cell population, thus obtaining the final cell population grouping and the final tissue region grouping.

[0045] Specifically, in this embodiment, the number of cell populations divided by single-cell transcriptome data is set as S, and the number of regions divided by spatial transcriptome data is set as T. Therefore, for each category M, it includes at least one cell population C. ik Where 1 ≤ i ≤ M; 1 ≤ k ≤ S. For each N category, it includes at least one organizational region C. jg Where 1≤j≤N; 1≤g≤T. Then, the tissue regions of all cell populations on the spatial transcriptome data were obtained using the Seurat software package or SPOTlight software. For C ikIf it is distributed across n types of regions, where 2 ≤ n ≤ N, then C is reclassified according to the n types of regions. i k is divided into n groups; for C jg If there are m types of cells distributed in this region, where 2≤m≤M, then C is reclassified according to m types of cells. jg The tissue is divided into m regions. This reclassification process is repeated until n=1 and m=1, meaning that each cell population uniquely corresponds to a tissue region of that type, thus avoiding ambiguous spatial localization of cell populations. Through this interactive clustering, the final cell population groups and final tissue region groups are obtained.

[0046] Step S103: Extract cell characteristic genes and region characteristic genes from the final cell population grouping and the final tissue region grouping respectively, and perform characteristic association analysis to obtain common characteristic genes that show the same expression pattern among cell populations and their localized tissue regions.

[0047] Specifically, the final cell population grouping and the final tissue region grouping are obtained. Cell characteristic genes and region characteristic genes are extracted from the final cell population grouping and the final tissue region grouping, respectively. Then, feature association analysis is performed to obtain common characteristic genes that show the same expression pattern among cell populations and their localized tissue regions.

[0048] In this process, for the cell population categories and tissue regions obtained in the above steps, the Seurat package of the R statistical platform is used to calculate the differentially expressed genes of each cell population and each tissue region compared with other populations and regions. The differentially expressed genes are filtered using logFoldChange>1 and the corrected P value is less than 0.01. The genes obtained by filtering are the cell characteristic genes and the region characteristic genes.

[0049] Furthermore, step 103 specifically includes:

[0050] 1. Extract the final cell characteristic genes and the final region characteristic genes from each final cell population and each final tissue region respectively, and take the union of the two to obtain the characteristic gene set.

[0051] 2. Select target feature genes one by one from the gene set, and calculate the gene expression value of the target feature gene under two different preset conditions. The two different preset conditions include two different cell populations or two different tissue regions. The gene expression value includes the first gene expression value of the target feature gene in a specific cell population and the second gene expression value in the tissue region where the specific cell population is located under one preset condition, and the third gene expression value of the target feature gene in another specific cell population and the fourth gene expression value in the tissue region where the other specific cell population is located under another preset condition.

[0052] 3. Based on gene expression values, confirm the expression differences of the target characteristic gene between two different preset conditions. When the expression difference confirms that the expression pattern of the target characteristic gene is consistent in the single-cell transcriptome and the spatial transcriptome, the target characteristic gene is confirmed as a common characteristic gene between the two different preset conditions.

[0053] Specifically, in this embodiment, it is assumed that for single-cell transcriptome data, there are c populations of cells, and for spatial transcriptome data, there are r regions. Preset conditions i and j corresponding to two different omics are randomly selected from each omics, including C. i and C j , (1≤i≤c; 1≤j≤c); R i and R j , (1≤i≤r; 1≤j≤r), where C i R i Corresponding conditions i, C j R j The corresponding condition j represents a specific cell type and its corresponding tissue region location. Based on cell characteristic genes and region characteristic genes, target characteristic genes S are selected one by one from their union to obtain their gene expression values ​​under each condition: S ci (first gene expression value), S cj (Second gene expression value), S ri (Third gene expression value), S rj (Fourth gene expression value) is determined using a t-test to assess the expression difference between conditions i and j. If the expression pattern is consistent across all omics, it is identified as a common characteristic gene between conditions i and j. All common characteristic genes are obtained by examining each characteristic gene in the union of cellular and regional characteristic genes.

[0054] Step S104: Based on common characteristic genes, combined with pre-acquired cell line drug intervention and expression data, perform biological computation to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to common characteristic genes.

[0055] It should be noted that traditional drug screening lacks information on the precise cell populations targeted and their distribution within specific tissue regions. The entire screening process is somewhat unpredictable, and the specific cell populations and tissue regions targeted by the screened drugs remain unknown, while also consuming significant human and financial resources. Therefore, to improve the accuracy of drug screening, this embodiment uses common characteristic genes, combined with pre-acquired cell line drug intervention and expression data, to perform biological computation to screen for sensitive drugs that can target cell populations within specific tissue regions corresponding to the common characteristic genes.

[0056] Furthermore, step 104 specifically includes:

[0057] 1. Identify all common target characteristic genes corresponding to cell populations within a specific tissue region.

[0058] 2. Obtain drug intervention data and gene expression data for cell lines.

[0059] Specifically, cell line drug intervention and expression data can be obtained from the publicly available CCLE (Cancer Cell Line Encyclopedia) or GDSC (Genomics of Drug Sensitivity in Cancer) databases. These databases include drug intervention data for over a thousand drugs in each cell line, as well as gene expression profile data for each cell line. In the drug intervention data, the AUC (Area Under Curve) value represents the bioavailability of each drug.

[0060] 3. Correlation analysis was performed on the drug intervention data of all candidate drugs and the gene expression data of the target common characteristic gene in all cell lines to identify candidate drugs that are sensitive to the expression of the target common characteristic gene, thus obtaining the first drug candidate set.

[0061] Specifically, this embodiment extracts any target common characteristic gene from a cell population within a given tissue region, calculates the Pearson correlation coefficient between the expression value of the target common characteristic gene in all cell lines and the AUC value corresponding to each drug in the same cell line, and calculates the relevant significance P-value. Under the threshold condition of P-value ≤ 0.05, if the Pearson correlation coefficient is greater than or equal to 0.3, the drug is considered to be sensitive to the expression of the gene; if the Pearson correlation coefficient is less than 0.3, the drug is considered to be resistant to the expression of the gene.

[0062] 4. Identify high-activity target cell lines by utilizing the activity values ​​of common characteristic genes in all cell lines.

[0063] Specifically, this embodiment involves two aspects: firstly, screening for drugs that are sensitive to characteristic genes based on their gene expression; and secondly, obtaining cell lines with high activity based on characteristic genes, and then screening for sensitive drugs based on the drug responses of these cell lines. Therefore, this embodiment also needs to utilize the activity values ​​of the target common characteristic gene in all cell lines to confirm the target cell lines with high activity.

[0064] Furthermore, the step of identifying high-activity target cell lines by utilizing the activity values ​​of the target common characteristic genes in all cell lines specifically includes:

[0065] 4.1 Calculate the average expression value of the common characteristic gene of the target in each cell line using the gene expression value of each common characteristic gene of the target in each cell line.

[0066] 4.2 Calculate the activity value of each target common characteristic gene in all cell lines using the average expression value of each target common characteristic gene and the gene expression value in each cell type.

[0067] 4.3 Construct an activity value matrix of the target common characteristic genes and cell lines based on the activity values.

[0068] 4.4 Cluster the activity value matrix to identify the cluster with the highest activity value, and select all cell lines in the cluster as target cell lines.

[0069] Specifically, extract all common characteristic genes of a cell population within a given tissue region, with a set number of M genes. Assume there are N cell types to which the cell line belongs, and for each cell type j, 1 ≤ j ≤ N, including n. j For a cell line k, for a characteristic gene i, where 1 ≤ i ≤ M, its gene expression value in cell line k is G. ik where 1≤k≤n j Its average expression value A in this cell type ij for:

[0070]

[0071] Its activity value R in all cell lines ij for:

[0072]

[0073] Among them, G ij This represents the gene expression value for each cell type.

[0074] Based on the above calculation method, the activity value matrix of all common characteristic genes in all cell lines was obtained. Then, unsupervised clustering was performed using pheatmap in the statistical computing platform R to select clusters composed of cell lines with high activity values, which have the same expression characteristics as the cell population of a given tissue region.

[0075] 5. Analyze the drug intervention data corresponding to the target cell line, identify candidate drugs that are sensitive to the expression of the target cell line, and obtain a second set of drug candidates.

[0076] Specifically, after obtaining the target cell line, based on the drug intervention data corresponding to the target cell line, candidate drugs that are sensitive to the expression of the target cell line are identified, thus obtaining a second set of drug candidates.

[0077] Further, the steps of analyzing drug intervention data corresponding to the target cell lines to identify candidate drugs that are sensitive to the expression of the target cell lines and obtaining a second drug candidate set specifically include:

[0078] 5.1 Obtain bioavailability data from the drug intervention data corresponding to all target cell lines.

[0079] 5.2 Sort all bioavailability data in descending order from high to low.

[0080] 5.3. Based on the fluctuation range between all adjacent bioavailability data, determine whether the arrangement of all bioavailability data is continuous. If it is continuous, the median of all bioavailability data is used as the final bioavailability data. If it is not continuous, perform statistical distribution on all bioavailability data and use the result of the statistical distribution as the final bioavailability data.

[0081] 5.4. Based on the final bioavailability data, confirm whether the candidate drugs are sensitive to expression in the target cell line to obtain the second set of drug candidates.

[0082] Specifically, the AUC values ​​(bioavailability data) of the drug intervention data obtained above for the target cell lines are extracted. The AUC values ​​of each drug in the corresponding cell lines are sorted from highest to lowest. If the ranking values ​​are continuous and uninterrupted, the cell lines are divided according to the median AUC value: cell lines with high AUC values ​​are sensitive to the drug, and those with low AUC values ​​are resistant. If the AUC ranking values ​​are discontinuous, a mixture model is used to separate the AUC values ​​based on their statistical distribution (such as normal, Gaussian, negative binomial, etc.) to identify drugs that are sensitive or resistant to specific cell lines.

[0083] In this embodiment, any two adjacent AUC values ​​below the preset fluctuation range are considered as continuous and uninterrupted, while any two adjacent AUC values ​​above the preset fluctuation range are considered as continuous and interrupted.

[0084] 6. Take the intersection of the first drug candidate set and the second drug candidate set to obtain a sensitive drug that can target cell populations in a specific tissue region.

[0085] Specifically, this embodiment uses drug screening based on characteristic genes of cell populations within a specific tissue region. It approaches drug screening from two angles: firstly, it screens for drugs that exhibit sensitivity to characteristic genes and their expression; secondly, it obtains cell lines with high activity based on these characteristic genes and then screens for sensitive drugs based on the drug responses of these cell lines. These two methods effectively screen for sensitive drugs that can act on both the cell population and the characteristic genes of the cell group, maximizing the fidelity of obtaining candidate sensitive drugs.

[0086] Furthermore, to verify the feasibility of this embodiment, this embodiment uses published single-cell transcriptome data and spatial transcriptome data (Reference: Nature Communications, 2022, 13:6823) as input data for verification. Specifically, please refer to [the relevant documentation / reference]. Figures 2-4 This embodiment is applied to single-cell transcriptome data and spatial transcriptome data that have been reported in the literature. Figure 2 The epithelial cell populations defined in this embodiment are shown, including cell populations 1, 2, 3, 4, 5, and 6; Figure 3 The original literature presented a tissue region distribution map of epithelial cells based on spatial transcriptome data. Figure 4 The spatial localization of different epithelial cell populations identified in this embodiment is shown in the tissue regions, with the regional distribution of their localization circled by dashed circles. It can be seen that this embodiment performs detailed cell population segmentation based on single-cell transcriptome data and spatial transcriptome data reported in the literature. However, the original literature does not provide as detailed a cell population segmentation as this embodiment, which captures more differences between cells. Furthermore, the cell distribution regions in the tissue regions described in the original literature are relatively general. In contrast, the different cell populations identified in this embodiment form more regional distributions in the tissue regions, better reflecting the tissue structure and fully reflecting the differences between tissue regions. In addition, sensitive drugs for cell population 2 were screened during the validation process, and the screening results are shown in Table 1 below.

[0087] Table 1

[0088] Sensitive drugs Drug description Vecuronium Acetylcholine receptor antagonist Olanzapine Dopamine receptor antagonist Argatroban Thrombin inhibitor Doxepin Histamine receptor antagonist SIB-1893 Glutamate receptor antagonist

[0089] Based on the fact that the single-cell transcriptome and spatial transcriptome data were from breast cancer, a literature review revealed that the screened drug Argatroban has been reported as a sensitive drug for breast cancer cells. This fully demonstrates the reliability of the drug screened in this embodiment. In summary, this embodiment is feasible in cell population segmentation, tissue region localization, and screening of cell populations sensitive to drugs within specific tissue regions. Furthermore, the results are superior to those reported in existing literature, and the screened sensitive drugs are consistent with clinical practice. The method for screening cell population-sensitive drugs in this embodiment first obtains the global structure of two omics data. It adopts both the overall structure and features of the single-cell transcriptome data and, with the help of the structures and features of the two omics data, interactively optimizes the internal structure of the single-cell transcriptome data and extracts features to minimize noise information. This ensures that the data segmentation reflects the characteristics of cell populations and tissue regions. Furthermore, in integrating single-cell transcriptome data and spatial transcriptome data, the two omics are integrated to perform feature association and extract common feature genes that show consistent changes between the omics data under the same conditions. This results in better integration of single-cell transcriptome data and spatial transcriptome data, and the data segmentation reflects the characteristics of cell populations and tissue regions. Then, based on these common feature genes, sensitive drugs that can target cell populations in specific tissue regions corresponding to the common feature genes are screened, making the screened sensitive drugs more reliable.

[0090] Figure 5 This is a schematic diagram of the functional modules of the device for screening cell population-sensitive drugs according to an embodiment of the present invention. Figure 5 As shown, the device 20 for screening cell population-sensitive drugs includes an acquisition module 21, a grouping module 22, a feature association module 23, and a drug screening module 24.

[0091] The acquisition module 21 is used to acquire single-cell transcriptome data and spatial transcriptome data and perform preliminary division to obtain preliminary cell population grouping and preliminary tissue region grouping.

[0092] Grouping module 22 is used to repeatedly cluster and correct the preliminary cell population grouping and preliminary tissue region grouping using cell marker genes to obtain the final cell population grouping and final tissue region grouping.

[0093] The feature association module 23 is used to extract cell feature genes and region feature genes from the final cell population grouping and the final tissue region grouping, respectively, and perform feature association analysis to obtain common feature genes that show the same expression pattern among cell populations and their localized tissue regions.

[0094] The drug screening module 24, based on common characteristic genes, combines pre-acquired cell line drug intervention and expression data to perform biological computation to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to common characteristic genes.

[0095] Optionally, the acquisition module 21 performs the operation of acquiring single-cell transcriptome data and spatial transcriptome data and performing preliminary segmentation to obtain preliminary cell population grouping and preliminary tissue region grouping. Specifically, this includes: acquiring single-cell transcriptome data and spatial transcriptome data and performing noise reduction; using a preset clustering algorithm to cluster the single-cell transcriptome data to obtain preliminary cell population grouping, and using a preset clustering algorithm to segment the spatial transcriptome data to obtain preliminary tissue region grouping.

[0096] Optionally, the grouping module 22 performs repeated clustering corrections on the preliminary cell population grouping and preliminary tissue region grouping using cell marker genes to obtain the final cell population grouping and final tissue region grouping. Specifically, this includes: extracting preliminary cell characteristic genes from each preliminary cell population grouping based on differences in cell gene expression, and extracting preliminary region characteristic genes from each preliminary tissue region grouping based on differences in gene expression between different tissue regions; classifying cell populations by cell type based on cell marker genes of different cell types; confirming the cell type corresponding to each preliminary cell population based on the overlap between preliminary cell characteristic genes and each type of marker gene, and confirming the region type corresponding to each preliminary tissue region based on the overlap between preliminary region characteristic genes and each type of marker gene; obtaining the tissue region location of all preliminary cell populations on spatial transcriptome data; obtaining the number of region types distributed within the same cell population and the number of cell populations distributed within the same tissue region based on the tissue region location; further dividing the same cell population based on the number of region types, and further dividing the same tissue region based on the number of cell populations, until each tissue region corresponds one-to-one with each cell population, thus obtaining the final cell population grouping and final tissue region grouping.

[0097] Optionally, the feature association module 23 performs the following operations: extracting cell-specific genes and region-specific genes from the final cell population grouping and the final tissue region grouping, respectively, and performing feature association analysis to obtain common feature genes that exhibit the same expression pattern among cell populations and their corresponding tissue regions. Specifically, this includes: extracting final cell-specific genes and final region-specific genes from each final cell population and each final tissue region, and taking the union of the two to obtain a feature gene set; selecting target feature genes one by one from the gene set, and calculating the gene expression value of the target feature gene under two different preset conditions, where the two different preset conditions include two different cell populations or two... In different tissue regions, gene expression values ​​include the first gene expression value of the target characteristic gene in a specific cell population under a preset condition and the second gene expression value in the tissue region where the specific cell population is located, as well as the third gene expression value of the target characteristic gene in another specific cell population under another preset condition and the fourth gene expression value in the tissue region where the other specific cell population is located. Based on the gene expression values, the expression difference of the target characteristic gene between the two different preset conditions is confirmed. When the expression difference confirms that the expression pattern of the target characteristic gene is consistent in the single-cell transcriptome and the spatial transcriptome, the target characteristic gene is confirmed as a common characteristic gene between the two different preset conditions.

[0098] Optionally, the drug screening module 24 performs biological computation based on common characteristic genes and pre-acquired cell line drug intervention and expression data to screen sensitive drugs that can target cell populations in specific tissue regions corresponding to common characteristic genes. Specifically, this includes: identifying all target common characteristic genes corresponding to cell populations in specific tissue regions; acquiring cell line drug intervention data and gene expression data; performing correlation analysis on the drug intervention data of all candidate drugs and the gene expression data of the target common characteristic genes in all cell lines to identify candidate drugs sensitive to the expression of the target common characteristic genes, thus obtaining a first drug candidate set; using the activity value of the target common characteristic genes in all cell lines to identify target cell lines with high activity; analyzing the drug intervention data corresponding to the target cell lines to identify candidate drugs sensitive to the expression of the target cell lines, thus obtaining a second drug candidate set; and taking the intersection of the first drug candidate set and the second drug candidate set to obtain sensitive drugs that can target cell populations in specific tissue regions.

[0099] Optionally, the drug screening module 24 performs the operation of identifying target cell lines with high activity using the activity values ​​of the target common characteristic gene in all cell lines. Specifically, this includes: calculating the average expression value of the target common characteristic gene in each cell type using the gene expression value of each target common characteristic gene in each cell line; calculating the activity value of each target common characteristic gene in all cell lines using the average expression value of each target common characteristic gene and the gene expression value in each cell type; constructing an activity value matrix of the target common characteristic gene and cell lines based on the activity values; clustering the activity value matrix to identify the cluster with the highest activity value, and selecting all cell lines in the cluster as target cell lines.

[0100] Optionally, the drug screening module 24 performs the following operations: analyzing the drug intervention data corresponding to the target cell line, identifying candidate drugs that are sensitive to the expression of the target cell line, and obtaining a second drug candidate set. Specifically, this includes: obtaining bioavailability data from the drug intervention data corresponding to all target cell lines; sorting all bioavailability data in descending order from high to low; determining whether the arrangement of all bioavailability data is continuous based on the fluctuation range between all adjacent bioavailability data; if continuous, using the median of all bioavailability data as the final bioavailability data; if discontinuous, performing statistical distribution on all bioavailability data and using the result of the statistical distribution as the final bioavailability data; and determining whether the candidate drugs are sensitive to the expression of the target cell line based on the final bioavailability data, thereby obtaining a second drug candidate set.

[0101] For further details regarding the implementation of the technical solutions for each module in the apparatus for screening cell population-sensitive drugs in the above embodiments, please refer to the description in the method for screening cell population-sensitive drugs in the above embodiments, which will not be repeated here.

[0102] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0103] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 6 As shown, the computer device 30 includes a processor 31 and a memory 32 coupled to the processor 31. The memory 32 stores program instructions. When the program instructions are executed by the processor 31, the processor 31 performs the method steps for screening cell population sensitive drugs as described in any of the above embodiments.

[0104] The processor 31 can also be referred to as a Central Processing Unit (CPU). The processor 31 may be an integrated circuit chip with signal processing capabilities. The processor 31 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.

[0105] See Figure 7 , Figure 7 This is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention. The storage medium of this embodiment stores program instructions 41 capable of implementing the above-described method for screening cell population-sensitive drugs. These program instructions 41 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed computer devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0107] Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for screening cell population-sensitive drugs, characterized in that, It includes: Single-cell transcriptome data and spatial transcriptome data were acquired and preliminarily divided to obtain preliminary cell population grouping and preliminary tissue region grouping. Using cell marker genes, the preliminary cell population grouping and the preliminary tissue region grouping are repeatedly clustered and corrected to obtain the final cell population grouping and the final tissue region grouping. Cellular characteristic genes and regional characteristic genes are extracted from the final cell population grouping and the final tissue region grouping, respectively, and feature association analysis is performed to obtain common characteristic genes that show the same expression pattern among cell populations and their localized tissue regions. Based on the common characteristic genes, combined with pre-acquired cell line drug intervention and expression data, biological computation is performed to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to the common characteristic genes.

2. The method for screening cell population-sensitive drugs according to claim 1, characterized in that, The acquisition of single-cell transcriptome data and spatial transcriptome data, followed by preliminary segmentation to obtain preliminary cell population grouping and preliminary tissue region grouping, includes: Acquire the single-cell transcriptome data and the spatial transcriptome data and denoise them; The single-cell transcriptome data is grouped using a preset clustering algorithm to obtain the preliminary cell population grouping, and the spatial transcriptome data is divided using the preset clustering algorithm to obtain the preliminary tissue region grouping.

3. The method for screening cell population-sensitive drugs according to claim 1, characterized in that, The process involves repeatedly clustering and refining the preliminary cell population grouping and the preliminary tissue region grouping using cell marker genes to obtain the final cell population grouping and the final tissue region grouping, including: Preliminary cell characteristic genes were extracted from each preliminary cell population based on differences in cell gene expression, and preliminary regional characteristic genes were extracted from each preliminary tissue region based on differences in gene expression between different tissue regions. Cell populations are classified into cell types based on cell marker genes of different cell types; The cell type corresponding to each preliminary cell population is determined based on the number of overlaps between the preliminary cell characteristic genes and each type of marker gene, and the region type corresponding to each preliminary tissue region is determined based on the number of overlaps between the preliminary region characteristic genes and each type of marker gene. Obtain the tissue region localization of all preliminary cell populations on spatial transcriptome data; Based on the tissue region location, the number of regional types of the same cell population distribution and the number of cell populations distributed within the same tissue region are obtained; The same cell population is further subdivided based on the number of the region types, and the same tissue region is further subdivided based on the number of the cell populations, until each tissue region corresponds one-to-one with each cell population, thus obtaining the final cell population grouping and the final tissue region grouping.

4. The method for screening cell population-sensitive drugs according to claim 3, characterized in that, The cell-specific genes and region-specific genes are extracted from the final cell population grouping and the final tissue region grouping, respectively, and feature association analysis is performed to obtain common feature genes that exhibit the same expression pattern among cell populations and their corresponding tissue regions, including: The final cell characteristic genes and the final region characteristic genes were extracted from each final cell population and each final tissue region, and the union of the two was taken to obtain the characteristic gene set. Target feature genes are selected one by one from the gene set, and the gene expression values ​​of the target feature genes are calculated under two different preset conditions. The two different preset conditions include two different cell populations or two different tissue regions. The gene expression values ​​include the first gene expression value of the target feature gene in a specific cell population and the second gene expression value in the tissue region where the specific cell population is located under one preset condition, and the third gene expression value of the target feature gene in another specific cell population and the fourth gene expression value in the tissue region where the other specific cell population is located under another preset condition. Based on the gene expression value, the expression difference of the target characteristic gene between the two different preset conditions is confirmed. When the expression difference confirms that the expression pattern of the target characteristic gene is consistent in the single-cell transcriptome and the spatial transcriptome, the target characteristic gene is confirmed as a common characteristic gene between the two different preset conditions.

5. The method for screening cell population-sensitive drugs according to claim 1, characterized in that, The process of using pre-acquired cell line drug intervention and expression data for biological computation based on the common characteristic genes to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to the common characteristic genes includes: Identify all target common characteristic genes corresponding to cell populations within a specific tissue region; Obtain drug intervention data and gene expression data for cell lines; Correlation analysis was performed on the drug intervention data of all candidate drugs and the gene expression data of the target common characteristic gene in all cell lines to identify candidate drugs that are sensitive to the expression of the target common characteristic gene, thus obtaining the first drug candidate set; The target cell lines with high activity were identified by using the activity values ​​of the target common characteristic genes in all cell lines. Analyze the drug intervention data corresponding to the target cell line to identify candidate drugs that are sensitive to the expression of the target cell line, and obtain a second drug candidate set; The intersection of the first drug candidate set and the second drug candidate set is used to obtain the sensitive drug that can target the cell population in the specific tissue region.

6. The method for screening cell population-sensitive drugs according to claim 5, characterized in that, The method of identifying high-activity target cell lines by utilizing the activity values ​​of the target common characteristic genes in all cell lines includes: The average expression value of the target common characteristic gene in each cell type was calculated using the gene expression value of each target common characteristic gene in each cell line; The activity values ​​of each common characteristic gene of each target were calculated in all cell lines using the average expression value of each common characteristic gene of each target and the gene expression value of each cell type; Construct an activity value matrix of the target common characteristic gene and the cell line based on the activity value; Cluster the activity value matrix to identify the cluster with the highest activity value, and select all cell lines in the cluster as the target cell lines.

7. The method for screening cell population-sensitive drugs according to claim 5, characterized in that, The analysis of drug intervention data corresponding to the target cell line identifies candidate drugs that are sensitive to the expression of the target cell line, resulting in a second drug candidate set, including: Obtain bioavailability data from the drug intervention data for all target cell lines; Sort all bioavailability data in descending order from highest to lowest; The continuity of all bioavailability data is determined based on the fluctuation range between all adjacent bioavailability data. If they are continuous, the median of all bioavailability data is used as the final bioavailability data. If they are not continuous, a statistical distribution is performed on all bioavailability data, and the result of the statistical distribution is used as the final bioavailability data. Based on the final bioavailability data, it is confirmed whether the candidate drugs are sensitive to expression in the target cell line, thus obtaining the second drug candidate set.

8. An apparatus for screening cell population-sensitive drugs, characterized in that, It includes: The acquisition module is used to acquire single-cell transcriptome data and spatial transcriptome data and perform preliminary segmentation to obtain preliminary cell population grouping and preliminary tissue region grouping. The grouping module is used to repeatedly cluster and correct the preliminary cell population grouping and the preliminary tissue region grouping using cell marker genes to obtain the final cell population grouping and the final tissue region grouping. The feature association module is used to extract cell feature genes and region feature genes from the final cell population grouping and the final tissue region grouping, respectively, and perform feature association analysis to obtain common feature genes that show the same expression pattern among cell populations and their localized tissue regions. The drug screening module, based on the common characteristic gene, combines pre-acquired cell line drug intervention and expression data to perform biological calculations to screen for sensitive drugs that can target cell populations in specific tissue regions corresponding to the common characteristic gene.

9. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, the memory storing program instructions that, when executed by the processor, cause the processor to perform the steps of the method for screening cell population-sensitive drugs as described in any one of claims 1-7.

10. A storage medium, characterized in that, The device stores program instructions capable of implementing the method for screening cell population-sensitive drugs as described in any one of claims 1-7.