Gene classification method and device based on gene expression data and electronic equipment

CN121617478BActive Publication Date: 2026-08-07BEIJING DINGCHENG PEPTIDE SOURCE BIOINFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DINGCHENG PEPTIDE SOURCE BIOINFORMATION TECHNOLOGY CO LTD
Filing Date
2025-12-02
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]因此不同样本基因的表达水平会有波动,已知技术通过设置固定阈值的方法来判定标记基因的类型,结果并不准确

Benefits of technology

处理器,用于调用和执行所述计算机指令以实现本申请实施例各个方面所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617478B_ABST
    Figure CN121617478B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a gene classification method and device based on gene expression data and electronic equipment, and belongs to the technical field of bioinformatics and the field of data processing. The method comprises the following steps: determining the expression value of a gene in different cell clusters to obtain an expression value data set corresponding to the gene; determining a target expression value threshold of the gene according to the expression value distribution in the expression value data set; and classifying the gene in the cell cluster based on the target expression value threshold. The embodiment of the application can utilize the distribution rule of the gene expression value, improve the objectivity and accuracy of biological classification, effectively cope with the fluctuation caused by the batch effect on the gene expression level, and ensure the consistency of gene classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of bioinformatics and data processing, and more specifically, to a gene classification method, apparatus, and electronic device based on gene expression data. Background Technology

[0002] Cellular RNA sequencing technology can detect gene expression profiles at single-cell resolution. During cell annotation, cell type determination is required based on clustering results and the expression status of marker genes in each cell cluster.

[0003] Current single-cell sequencing technologies are subject to batch effects, where gene expression levels fluctuate across different experimental batches. Specific reasons for batch effects may include: differences in PBMC isolation procedures from blood, variations in experimental reagents and consumables, differences in equipment and experimental environment, and differences in experimental procedures and personnel.

[0004] Therefore, gene expression levels fluctuate across different samples. Known techniques for determining marker gene type by setting fixed thresholds are inaccurate. Furthermore, relying solely on "expert experience" to determine whether marker gene expression is negative or positive is subjective, making comparisons and reproducibility difficult. Moreover, methods based on human experience are ill-suited for large-scale automated analysis, limiting their application in large research projects. Summary of the Invention

[0005] This application provides a gene classification method, apparatus, and electronic device based on gene expression data, for objective, accurate, and adaptive classification of genes.

[0006] According to a first aspect of the embodiments of this application, a gene classification method based on gene expression data is provided, the method comprising: Determine the expression values ​​of the gene in different cell clusters to obtain the expression value dataset corresponding to the gene; Based on the expression value distribution in the expression value dataset, determine the target expression value threshold for the gene; The genes in the cell cluster are classified based on the target expression value threshold.

[0007] According to a second aspect of the embodiments of this application, a gene classification method applicable to different experimental batches is provided. The method includes: in different experimental batches, gene classification is performed using the method provided in the first aspect of the embodiments of this application, so as to effectively address the problem of gene expression value fluctuations caused by batch effects.

[0008] According to a third aspect of the embodiments of this application, a cell annotation method is provided, the method comprising: For each gene in the cell cluster, gene classification is performed using the method described in the first aspect of the embodiments of this application; Based on the results of the gene classification, cell type annotation is performed on the cells in the cell cluster.

[0009] According to a fourth aspect of the embodiments of this application, a gene classification device based on gene expression data is provided, the device comprising: The dataset determination module is used to determine the expression values ​​of genes in different cell clusters and obtain the expression value dataset corresponding to the genes. The expression value threshold determination module determines the target expression value threshold of the gene based on the expression value distribution in the expression value dataset. A classification module is used to classify the genes in the cell cluster based on the target expression value threshold.

[0010] According to a fifth aspect of the embodiments of this application, an electronic device is provided, comprising: Memory, used to store one or more computer instructions; A processor is used to invoke and execute the computer instructions to implement the methods described in various aspects of the embodiments of this application.

[0011] According to a sixth aspect of the embodiments of this application, a computer-readable storage medium is provided, storing one or more computer instructions that, when executed, implement the methods described in various aspects of the embodiments of this application.

[0012] According to a seventh aspect of the embodiments of this application, a computer program product is provided that, when executed, implements the methods described in various aspects of the embodiments of this application.

[0013] By employing the embodiments of this application, the target expression value threshold for genes determined based on the expression value distribution in the expression value dataset can more accurately reflect the distribution pattern of expression values, thereby improving the objectivity and accuracy of biological classification. Furthermore, compared to existing technologies that classify based on fixed thresholds, this method not only has higher accuracy but also effectively addresses the fluctuations in gene expression levels caused by batch effects, ensuring the consistency of gene classification. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating a gene classification method based on gene expression data according to an embodiment of this application; Figure 2 This is a flowchart illustrating a method for determining a target expression value threshold based on the expression value distribution in an expression value dataset, according to an embodiment of this application. Figure 3 This is a schematic diagram of a gene sorting device according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0016] It should be understood that "multiple" as mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish the same or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., do not necessarily mean different. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or devices.

[0017] First, the terms or nouns involved in the embodiments of this application will be introduced.

[0018] PBMC (Peripheral Blood Mononuclear Cell): Peripheral blood mononuclear cells, including lymphocytes and monocytes, are an important source of samples for immunological research.

[0019] Single-cell sequencing is a technique that performs sequencing analysis of the genome, transcriptome, etc., at the level of a single cell, and can reveal cellular heterogeneity.

[0020] Gene expression level refers to the number of mRNA molecules produced by the transcription of a specific gene in a cell, reflecting the degree of gene activity.

[0021] Batch effect: Technical deviations caused by non-biological factors such as experimental time, reagent batches, and operators.

[0022] Cell clustering: The process of grouping cells based on the similarity of their gene expression profiles.

[0023] Cell annotation: Identifying the biological function of cell clusters based on known marker genes.

[0024] Expression rate: In single-cell data, the proportion of cells in a cell population that express a specific gene.

[0025] Adaptive Threshold: A classification boundary value automatically calculated based on the characteristics of data distribution.

[0026] The above describes the terminology that may be involved in the embodiments of this application. Next, the embodiments of this application and related technologies will be illustrated with examples.

[0027] This application provides a gene classification method based on gene expression data, referring to... Figure 1 The flowchart shown is a gene classification method based on gene expression data. The method includes the following processing steps.

[0028] 100: Determine the expression value of the gene in different cell clusters to obtain the corresponding gene expression value dataset.

[0029] Optionally, in one implementation of this embodiment, the expression value is used to quantify gene expression, including: gene expression level, gene expression rate (i.e., the proportion of cells in which the gene is detected to be expressed) or similar parameters. The similar parameters can be parameters derived from gene expression level or gene expression rate to quantify gene expression.

[0030] In one specific application of this embodiment, the expression value dataset of each gene in the cell cluster can be determined separately.

[0031] 102: Determine the target expression value threshold for the gene based on the expression value distribution in the expression value dataset.

[0032] 104: Classify genes in cell clusters based on target expression value thresholds.

[0033] For example, genes whose expression values ​​do not exceed the target expression value threshold are classified as Category 1 (e.g., negative), and genes whose expression values ​​exceed the target expression value threshold are classified as Category 2 (e.g., positive).

[0034] The method provided in this embodiment, which determines the target expression value threshold for genes based on the expression value distribution in the expression value dataset, can more accurately reflect the distribution pattern of expression values, thereby improving the objectivity and accuracy of biological classification. Furthermore, compared to existing technologies that classify based on fixed thresholds, it not only has higher accuracy but also effectively addresses the fluctuations in gene expression levels caused by batch effects, ensuring the consistency of gene classification.

[0035] Optionally, in one implementation of this embodiment, before processing 104, the validity of the target expression value threshold is determined based on whether the target expression value threshold falls within the threshold range corresponding to the gene; if the target expression value threshold is not within the threshold range, the target expression value threshold is set to the threshold setting value corresponding to the gene.

[0036] By adopting this approach, the target expression threshold can be refined by setting a threshold range corresponding to the gene and comparing the target expression threshold with this threshold range, thus ensuring its biological interpretability.

[0037] To reasonably determine the threshold range and threshold setting value corresponding to genes, different versions of the threshold range and threshold setting value can be designed based on the same dataset. For the cell clusters obtained after cell clustering, the gene classification method provided in this application, along with different versions of the threshold range and threshold setting value, are used for gene classification. Then, cell type annotation is performed based on the gene classification results. Finally, the cell type annotation results are compared with the known cell type labels in the dataset to determine the optimal threshold range and threshold setting value for the annotation results.

[0038] Optionally, in the embodiments of this application, the gene classification method can have multiple execution logics in determining the target expression value threshold of a gene based on the expression value distribution in the expression value dataset (examples below illustrate methods such as specifying quantiles and class intervals). To achieve better results, the method can also perform the following processing: based on the validation dataset, gene classification is performed using different execution logics, and cell type annotation is performed based on the gene classification results. Then, the cell type annotation results are compared with the cell type labels of the validation dataset to determine the execution logic that makes the cell type annotation results optimal (with the highest matching degree with the cell annotation labels of the validation dataset).

[0039] It should be noted that for different technical details under the same execution logic (e.g., the selection of quantiles in the specified quantile method, the number of sub-intervals in the class interval method), multiple different versions can be designed, and then compared and verified based on the validation dataset to determine the optimal version.

[0040] In addition, in other embodiments of this application, gene classification can be performed using different execution logics, threshold ranges, and threshold settings based on the validation dataset, and cell type annotation can be performed based on the gene classification results; the cell type annotation results are compared with the cell type labels of the validation dataset to determine the optimal combination of execution logic, threshold range, and threshold settings that results in the optimal cell type annotation results.

[0041] It should be noted that the terms "optimal", "highest", and "best" mentioned in the embodiments of this application do not refer to absolute maximum or minimum values, but rather to maximum or minimum values ​​determined by comparison in a finite number of experiments or a finite amount of data.

[0042] Figure 2 This is a flowchart illustrating a method for determining a target expression value threshold based on the expression value distribution in an expression value dataset, according to an embodiment of this application. (Refer to...) Figure 2 The method includes the following processing steps.

[0043] 200: Divide the expression value dataset into multiple consecutive sub-intervals.

[0044] For example, the dataset can be divided into 20 sub-intervals. The specific number of sub-intervals is not limited in this embodiment. "Continuous" means that the expression value in a later sub-interval is greater than the expression value in an earlier sub-interval, and adjacent sub-intervals are connected end-to-end. For example, the expression value dataset can be divided into multiple equally spaced sub-intervals.

[0045] 202: Based on the expression value distribution characteristics and interval distribution characteristics of each sub-interval, a candidate expression value threshold is selected from the candidate expression value thresholds that makes the difference between the first-class interval and the second-class interval reach a set condition as the target expression value threshold. The candidate expression value threshold is the boundary value that divides the expression value dataset into the first-class interval and the second-class interval, where the sub-intervals within the first-class interval are continuous, and the sub-intervals within the second-class interval are also continuous. "Reaching a set condition" is, for example, maximizing the difference between the first-class interval and the second-class interval.

[0046] The method provided in this embodiment selects a target expression value threshold from candidate expression value thresholds that results in a difference between the first and second class intervals meeting a predetermined condition. This allows for more accurate gene classification based on the target expression value threshold. Furthermore, the method provided in this embodiment specifically determines the target expression value threshold corresponding to each gene based on the expression value dataset. Therefore, the target expression value threshold better reflects the objective laws governing gene expression values. Compared to existing technologies that classify based on fixed thresholds, this method not only has higher accuracy but also effectively addresses batch effects that cause fluctuations in gene expression levels, ensuring consistency.

[0047] Optionally, in one implementation of this embodiment, the expression value distribution feature of a sub-interval is the proportion of the number of expression values ​​in the sub-interval to the total number of expression values ​​in the dataset, and the interval distribution feature of the sub-interval is the boundary point or midpoint of the sub-interval. This way, both the expression value distribution feature of the sub-interval and the interval distribution feature can be used to represent the distribution of expression values ​​in the dataset, providing a feature foundation for accurate classification based on the target expression value threshold. It should be noted that regardless of whether a boundary point or a midpoint is selected, all sub-intervals are consistent; for example, all may select the left boundary point or all may select the midpoint.

[0048] Optionally, in one implementation of this embodiment, the following process can be used in process 202 to determine the degree of difference between class intervals.

[0049] First, determine the feature values ​​of the sub-intervals based on their interval distribution characteristics and expression value distribution characteristics. For example, calculate the product of the interval distribution characteristics and expression value distribution characteristics of the sub-intervals as the feature values ​​of the sub-intervals.

[0050] Then, based on the feature values ​​of the sub-intervals within the first-class interval, the feature values ​​of the first-class interval are determined. For example, the feature values ​​of the first-class interval are determined according to the feature values ​​of the sub-intervals within the first-class interval and the weights of the first-class interval.

[0051] Next, based on the feature values ​​of the sub-intervals within the second-type interval, the feature values ​​of the second-type interval are determined. For example, the feature values ​​of the second-type interval are determined according to the feature values ​​of the sub-intervals within the second-type interval and the weights of the second-type interval.

[0052] Then, based on the feature values ​​of the first type of interval and the feature values ​​of the second type of interval, the degree of difference between the first type of interval and the second type of interval is determined. For example, the degree of difference is determined by comprehensively considering the feature values ​​and weights of the first type of interval, as well as the feature values ​​and weights of the second type of interval.

[0053] This implementation provides a framework for determining the features of sub-intervals, the features of class intervals, and the degree of difference between class intervals. This framework can accurately determine the degree of difference between class intervals, thus providing a logical basis for accurately determining the threshold of the target expression value.

[0054] For example, the weight of a class interval, which refers to either a first-class interval or a second-class interval, can be determined as follows: The sum of the expression value distribution characteristics of each sub-interval within the class interval is calculated to obtain the weight of the class interval, which is related to the sum. For example, the sum can be directly used as the weight of the class interval, or the weight of the class interval can be positively correlated with the sum.

[0055] For example, the feature value of a class interval can be determined by calculating the ratio of the sum of the feature values ​​of each subinterval in the class interval to the weight of the class interval, and using this ratio as the feature value of the class interval.

[0056] For example, the degree of difference between two class intervals can be determined based on the following formula.

[0057]

[0058] in, To represent the degree of difference, the difference value, This represents the weight of the first type of interval. This indicates the weight of the second type of interval. Represents the characteristic values ​​of the first type of interval. This represents the characteristic value of the second type of interval.

[0059] This application also provides a method for determining a target expression value threshold for a gene based on the expression value distribution in an expression value dataset. The method first arranges the expression values ​​in the dataset in order, for example, from smallest to largest. Then, it selects the expression value at a predetermined quantile (e.g., median, upper quartile) as the target expression value threshold. The specific value of the predetermined quantile can be obtained experimentally based on the approach provided in this application. This application does not impose specific limitations on this.

[0060] The method provided in this application embodiment can adaptively determine the expression value threshold associated with the expression value distribution, compared with the fixed threshold used in the prior art, and effectively cope with the fluctuations in gene expression levels caused by batch effects.

[0061] Optionally, in one implementation of this application embodiment, before selecting the expression value at a set quantile as the target expression value threshold, end data in the expression value dataset can be deleted to reduce extreme value interference. For example, all minimum and maximum value data in the expression value dataset can be deleted, and then the expression value at the set quantile can be selected as the target expression threshold.

[0062] This application also provides a gene classification method applicable to different experimental batches. The method includes classifying genes using the method provided in the preceding embodiments of this application in each experimental batch. This effectively addresses the fluctuations in gene expression levels caused by batch effects, ensuring the consistency of gene classification results across different batches and the consistency of subsequent cell annotation results.

[0063] This application also provides a cell annotation method, which includes: classifying each gene in a cell cluster using the method provided in the previous embodiments of this application; and annotating the cells in the cell cluster with cell types based on the gene classification results.

[0064] For example, suppose a cell cluster contains x cells. Among the y genes corresponding to this cell cluster, the classification (negative or positive) of y1 genes will determine the cell type annotation of the x cells in the cell cluster. Using the gene classification method provided in this application, the classification of y1 genes can be determined, and then cell type annotation can be performed.

[0065] Figure 3 This is a schematic diagram of a gene classification device based on gene expression data according to an embodiment of this application, with reference to... Figure 3 The gene classification device includes the following modules.

[0066] The dataset determination module is used to determine the expression values ​​of genes in different cell clusters, thereby obtaining the corresponding expression value datasets for the genes.

[0067] The expression value threshold determination module determines the target expression value threshold for genes based on the expression value distribution in the expression value dataset.

[0068] The classification module is used to classify genes in cell clusters based on a target expression value threshold.

[0069] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. (Refer to...) Figure 4 The electronic device includes a memory 50 and a processor 10.

[0070] The memory is used to store one or more computer instructions.

[0071] A processor for calling and executing computer instructions to implement at least one of the gene classification methods mentioned in the foregoing embodiments of this application.

[0072] Optionally, in one implementation of this embodiment, such as Figure 4 As shown, the electronic device includes a processor 10, at least one communication bus 20, a user interface 30, at least one external communication interface 40, and a memory 50. The communication bus 20 is configured to enable communication between these components. The user interface 30 may include a display screen, and the external communication interface 40 may include standard wired and wireless interfaces.

[0073] This application also provides a computer-readable storage medium storing one or more computer instructions, which, when executed, implement at least one of the methods provided in this application.

[0074] This application also provides a computer program product, which, when run, implements at least one of the methods provided in this application.

[0075] The descriptions of the computer program products, computer-readable storage media, and electronic devices described above are similar to those of the method embodiments described above, and have similar beneficial effects. For any technical details not disclosed in the apparatus, computer program products, computer-readable storage media, and electronic devices provided in the embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0076] The embodiments of this application have been described above. The specific implementation of the embodiments of this application in practical applications will be illustrated below.

[0077] This solution provides an adaptive threshold calculation method based on data distribution characteristics for objectively and reproducibly classifying gene expression status in single-cell transcriptome data. This method can replace fixed thresholds or subjective experience-based judgments. By analyzing the distribution characteristics of gene expression levels in a cell population, it automatically calculates an optimal discrimination threshold, thereby binarizing the gene expression status into "positive" (expressed) or "negative" (not expressed). Each gene undergoes the following steps to calculate its adaptive threshold. The specific process is as follows: Step 1: Gene Expression Data Preprocessing First, obtain the expression level data of the target gene in a cell cluster. Expression level can be measured in various ways, such as normalized expression level, expression rate (i.e., the proportion of cells in which the gene is detected as expressed), or other indicators that quantify gene expression. The purpose of this step is to prepare a numerical vector for each gene to be classified, containing the gene's expression level across all target cells. This vector will serve as the basis for subsequent threshold calculations.

[0078] Step 2: Adaptive threshold calculation based on data distribution Adaptive threshold calculation based on the data's inherent distribution can effectively address batch effects on gene expression level fluctuations. Compared to fixed or empirically judged thresholds, this calculation method is subjective, ensuring reproducibility and supporting the needs of large-scale automated analysis. The method provided here is suitable for finding the optimal threshold to classify data into two categories. Its mathematical principles and calculation process are as follows: ① Data Distribution Statistics: Treat the gene expression level vector obtained in step 1 as a dataset. First, determine a suitable numerical range (usually from the minimum to the maximum value of the dataset), and discretize this range into several (denoted as N) consecutive intervals (or "bins"). Next, count the number of data points falling into each interval, thus forming a histogram distribution of the gene expression level. Let the data value corresponding to the i-th interval be... The frequency of data points within this interval is If the total number of data points is M, then the probability of each data value appearing is... .

[0079] For example, suppose the expression range of a certain gene is (0, 10), and it is discretized into 5 intervals: (0, 2), [2, 4), [4, 6), [6, 8), [8, 10). Then the expression values ​​of these 5 intervals are... We can take their respective midpoints, that is =1、 =3、 =5、 = 7、 =9.

[0080] ② Traverse candidate thresholds and calculate the variance between class intervals: Assume a candidate threshold The entire dataset was split into two classes: Class C0 (containing all expression levels less than or equal to...). The data points can be considered as a potential "negative" group) and the C1 class (including all expression levels greater than 1). These data points can be considered as potential "positive" groups.

[0081] Calculate class weights: Weights of class C0 For all less than or equal to The sum of the probabilities of the data points, i.e. Weight of class C1 for .

[0082] Calculate the mean of the class interval (a type of eigenvalue): the mean of class C0. for The mean of class C1 for .

[0083] Calculating the variance between class intervals: The variance between two classes (between-class variance) is defined as follows: This variance measures the variance of the dataset in... After segmentation, the degree of separation between the two classes is considered. The larger the variance, the better the two classes are distinguished.

[0084] ③ Determine the optimal threshold: Iterate through all possible candidate thresholds (Typically, for each boundary value of an interval), calculate the corresponding variance. The choice makes the inter-class variance The largest candidate threshold This serves as the final adaptive threshold (i.e., the target expression value threshold). Mathematically, this threshold ensures that the difference between the two classes is maximized when the data is divided into two categories.

[0085] Step 3: Refinement of Adaptive Threshold examine Does it fall within a pre-defined, reasonable range (e.g., expression rate between 5% and 90%)? If the threshold is too low (e.g., close to 0), it may lead to misclassifying background noise as positive expression; if the threshold is too high, it may miss true low-level expression. If the threshold exceeds a reasonable range, it is adjusted to a preset default value (e.g., 0.45 or other experience-based values) to ensure that the threshold is biologically interpretable.

[0086] In other possible applications, the threshold range can be adjusted based on experiments or experience, or different correction strategies can be adopted (e.g., correction based on statistical distribution rather than a fixed range, specifying the interval of quantiles based on data distribution, etc.), as long as the threshold is verified and corrected based on biological rationale.

[0087] Step 4: Evaluation of Adaptive Threshold Calculation Results To ensure the effectiveness and superiority of the adaptive threshold calculation method, a systematic evaluation can be conducted. The evaluation primarily includes visual verification and performance quantification comparison based on downstream tasks.

[0088] ①Visual review Generate an expression level distribution map (such as a histogram or density map) of each gene in the cell population, with the horizontal axis representing the range of values ​​in each interval and the vertical axis representing the frequency. Clearly label the adaptive threshold calculated using the aforementioned method on the map. This visualization method helps researchers intuitively verify the reasonableness of the threshold, especially when the data exhibits a multimodal distribution, skewed distribution, or other complex distribution patterns. It can quickly identify potential biases in the algorithm and provide an intuitive basis for subsequent algorithm optimization.

[0089] ②Performance evaluation based on downstream annotation tasks Select a validation dataset with clearly defined cell type labels (ideally at least 30 samples), and run different versions of the adaptive threshold calculation method and the refinement method in parallel, keeping other analysis parameters consistent. Select the gene set that performs best in the annotation results (e.g., the one with the highest match to the cell type labels) as the final version.

[0090] Next, in order to more clearly illustrate the solutions of the relevant embodiments of this application, the application of the embodiments of this application will be illustrated with examples in conjunction with downstream analysis.

[0091] 1. Data input: After clustering, calculate the expression rate of marker genes for each cluster of cells.

[0092] 2. Adaptive threshold calculation: For example, use the method mentioned above to calculate adaptive thresholds for all selected marker genes. For example, the threshold for CD3D is 0.31, the threshold for CD4 is 0.35, and the threshold for CD8 is 0.55.

[0093] 3. Threshold Refinement: Check if the adaptive threshold obtained in step 2 is within a preset reasonable range (e.g., expression rate of 5% to 90%). If the threshold is outside the range, correct it to the preset default value (e.g., 0.45); if it is within the range, use it directly. In this example, the thresholds obtained are all within a reasonable range, so they are used.

[0094] 4. Adaptive threshold determination: The expression status of the marker gene is binarized and classified using a refined adaptive threshold.

[0095] 5. Cell type annotation: Based on the determination results of multiple marker genes in step 4, and combined with the known cell type determination rules (e.g., CD3D positive, CD4 positive, CD8 negative indicates CD4+ T cells), complete the cell type annotation for this cell cluster.

[0096] 6. Downstream analysis: Based on the annotation results, subsequent biological analyses such as differential expression analysis and functional enrichment analysis are performed.

[0097] Those skilled in the art will understand that the various embodiments of this application can be applied to fields such as clinical immune monitoring, disease mechanism research, drug development, and basic immunology research. Specifically, clinical immune monitoring involves assessing a patient's immune status through single-cell sequencing analysis of PBMC samples, providing a basis for immunotherapy; disease mechanism research aims to deepen understanding of cellular heterogeneity in the development of diseases such as tumor immunity and autoimmune diseases; drug development assesses the impact of drugs on specific immune cell subsets during drug screening, improving drug development efficiency; and basic immunology research explores the mechanisms of immune cell differentiation, development, and functional regulation.

[0098] In the above embodiments of this application, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The steps illustrated in the related flowcharts can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown here. In other words, the order of steps described in the foregoing embodiments is merely an example. Reasonable adjustments to the order of steps based on the content of the embodiments of this application are also within the protection scope of the embodiments of this application.

[0099] The sequence numbers or order of description of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0100] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0101] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0103] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0104] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the scene data of the current frame in the 3D virtual scene involved in the embodiments of this application, the client's device information, and the scene interaction information are all obtained with full authorization.

[0105] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A gene classification method based on gene expression data, characterized in that, The method includes: Determine the expression values ​​of the gene in different cell clusters to obtain the expression value dataset corresponding to the gene; Based on the expression value distribution in the expression value dataset, the target expression value threshold for the gene is determined, including: The expression value dataset is divided into multiple consecutive sub-intervals. Based on the expression value distribution characteristics and interval distribution characteristics of each sub-interval, a candidate expression value threshold is selected from the candidate expression value thresholds that makes the difference between the first type interval and the second type interval reach a set condition as the target expression value threshold. The candidate expression value threshold is the boundary value that divides the expression value dataset into the first type interval and the second type interval, where the sub-intervals in the first type interval are continuous, and the sub-intervals in the second type interval are also continuous. The degree of difference is calculated using the following formula: , To represent the degree of difference, This represents the weight of the first type of interval. This represents the weight of the second type of interval. Represents the feature values ​​of the first type of interval. The characteristic value of the second type of interval; The validity of the target expression value threshold is determined based on whether the target expression value threshold falls within the threshold range corresponding to the gene. If the target expression value threshold is not within the threshold range, the target expression value threshold is set to the threshold setting value corresponding to the gene. The genes in the cell cluster are classified based on the target expression value threshold.

2. The method according to claim 1, characterized in that, The expression value is used to quantify gene expression, including gene expression level or gene expression rate.

3. The method according to claim 1, characterized in that, The expression value distribution feature of the sub-interval is the proportion of the number of expression values ​​in the sub-interval to the total number of expression value datasets, and the interval distribution feature of the sub-interval is the boundary point or midpoint of the sub-interval.

4. The method according to claim 1, characterized in that, The step of selecting candidate expression value thresholds from the candidate expression value thresholds based on the expression value distribution characteristics and interval distribution characteristics of each of the sub-intervals, such that the difference between the first type interval and the second type interval reaches a set condition, as the target expression value threshold, includes: The feature values ​​of the sub-intervals are determined based on the interval distribution characteristics and expression value distribution characteristics of the sub-intervals; Based on the feature values ​​of the sub-intervals within the first type of interval, determine the feature values ​​of the first type of interval; Based on the feature values ​​of the sub-intervals within the second type of interval, determine the feature values ​​of the second type of interval; Based on the feature values ​​of the first type of interval and the feature values ​​of the second type of interval, the degree of difference between the first type of interval and the second type of interval is determined.

5. The method according to claim 4, characterized in that, Determining the feature values ​​of the sub-intervals based on their interval distribution characteristics and expression value distribution characteristics includes: The product of the interval distribution characteristics and the expression value distribution characteristics of the sub-interval is calculated as the feature value of the sub-interval; Determining the feature values ​​of the first type of interval based on the feature values ​​of the sub-intervals within the first type of interval includes: The feature values ​​of the first type of interval are determined based on the feature values ​​of the sub-intervals within the first type of interval and the weights of the first type of interval. Determining the feature values ​​of the second type of interval based on the feature values ​​of the sub-intervals within the second type of interval includes: The feature values ​​of the second type of interval are determined based on the feature values ​​of the sub-intervals within the second type of interval and the weights of the second type of interval.

6. The method according to claim 5, characterized in that, The method further includes: The weights of the class intervals are determined using the following method, where the class interval refers to either the first class interval or the second class interval: The sum of the expression value distribution features of each sub-interval in the class interval is calculated to obtain the weight of the class interval, and the weight of the class interval is related to the sum. The feature values ​​of the class intervals are determined in the following manner: The sum of the feature values ​​of each sub-interval in the class interval is calculated as the ratio of the weight of the class interval, and this ratio is used as the feature value of the class interval.

7. The method according to claim 1, characterized in that, The method includes various execution logics for determining the target expression value threshold of the gene based on the expression value distribution in the expression value dataset; The method includes multiple threshold ranges and threshold setting values; The method further includes: Based on the validation dataset, gene classification was performed using different execution logics, threshold ranges, and threshold settings, and cell type annotation was performed based on the gene classification results. The results of the cell type annotation are compared with the cell type labels of the validation dataset to determine the execution logic, threshold range, and threshold setting that optimize the results of the cell type annotation.

8. A gene classification method applicable to different experimental batches, characterized in that, The method includes: classifying genes using the method described in any one of claims 1-7 in different experimental batches.

9. A cell annotation method, characterized in that, The method includes: For each gene in the cell cluster, gene classification is performed using the method described in any one of claims 1-7; Based on the results of the gene classification, cell type annotation is performed on the cells in the cell cluster.

10. A gene classification device based on gene expression data, characterized in that, The device includes: The dataset determination module is used to determine the expression values ​​of genes in different cell clusters and obtain the expression value dataset corresponding to the genes. The expression value threshold determination module determines the target expression value threshold for the gene based on the expression value distribution in the expression value dataset, including: The expression value dataset is divided into multiple consecutive sub-intervals. Based on the expression value distribution characteristics and interval distribution characteristics of each sub-interval, a candidate expression value threshold is selected from the candidate expression value thresholds that makes the difference between the first type interval and the second type interval reach a set condition as the target expression value threshold. The candidate expression value threshold is the boundary value that divides the expression value dataset into the first type interval and the second type interval, where the sub-intervals in the first type interval are continuous, and the sub-intervals in the second type interval are also continuous. The validity of the target expression value threshold is determined by whether it falls within the threshold range corresponding to the gene. If the target expression value threshold is not within the threshold range, then the target expression value threshold is set to the threshold setting value corresponding to the gene. The degree of difference is calculated using the following formula: , To represent the degree of difference, This represents the weight of the first type of interval. This represents the weight of the second type of interval. Represents the feature values ​​of the first type of interval. The characteristic value of the second type of interval; A classification module is used to classify the genes in the cell cluster based on the target expression value threshold.

11. An electronic device, characterized in that, The electronic device includes: Memory, used to store one or more computer instructions; A processor for invoking and executing the computer instructions to implement the method as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The system stores one or more computer instructions that, when executed, implement the method as described in any one of claims 1-9.

13. A computer program product, characterized in that, The computer program product is executed to implement the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Cell type determination method and device, storage medium and electronic device

    CN112289379A