Rapid and stable single cell type labeling method
By reducing the dimensionality of the sequencing data of individual cells under different hyperparameters, and analyzing the cell rationality index and cell mixing degree in the data point chart, the problem of inaccurate dimensionality reduction results caused by unreasonable hyperparameter selection in the existing technology is solved, and more accurate single-cell type annotation is achieved.
Patent Information
- Application Number
- CN202510503735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the existing single-cell type annotation method, due to the unreasonable selection of hyperparameters, the dimensionality reduction results of cell data are not accurate enough.
By identifying the cell categories of different cells and reducing the dimensionality of the sequencing data of individual cells under different set hyperparameters for different values, the data point map before and after dimensionality reduction is obtained. According to the difference in the number distribution of data points in the neighborhood of each cell corresponding to each cell in the data point map, the rationality index for each cell is determined. Combining the differences between boundary and non-boundary regions, the degree of cell mixing in each cell type is determined, thereby determining the accuracy of hyperparameters.
The setting hyperparameters for the optimal value are determined adaptively, which effectively improves the accuracy of the dimensionality reduction results during single-cell type annotation process.
Smart Images

Figure CN120015135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a fast and stable single cell type annotation method. Background Art
[0002] The single-cell sequencing technology scRNA-seq generates data with very high dimensionality, including the expression levels of thousands of genes in single cells. By labeling single cell types, it can help researchers discover clusters of specific cell types.
[0003] In the prior art, in the process of single cell type annotation, the umap-learn tool is often used to perform UMAP dimensionality reduction on single cell data. However, there are many hyperparameters in the process of UMAP dimensionality reduction, such as the number of neighbors, minimum distance, etc. The rationality of setting these hyperparameters will seriously affect the accuracy of the dimensionality reduction results of the cell data. At present, the hyperparameters in the process of UMAP dimensionality reduction are mainly used to achieve dimensionality reduction of single cell data by calling default values. However, since different dimensionality reduction parameters should be selected for different cell data, when the hyperparameters determined by calling the default values are unreasonable, the dimensionality reduction results of the cell data will be inaccurate. Summary of the invention
[0004] The purpose of the present invention is to provide a fast and stable single cell type labeling method to solve the problem that the dimensionality reduction results of cell data are not accurate enough due to unreasonable selection of hyperparameters used in the existing single cell type labeling process.
[0005] In order to solve the above technical problems, in a first aspect, the present invention provides a fast and stable single cell type labeling method, comprising the following steps: Identify cell types of different cells based on sequencing data of different cells; Perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtain data point graphs before and after dimensionality reduction, where each data point in the data point graph corresponds to a cell; According to the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction, the rationality index of each cell is determined. The rationality index reflects the rationality of the setting of the current value of the hyperparameter for each cell. Determine the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, and determine the degree of cell mixing in the cell area of each cell type based on the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area, and the distribution of the data points of the cell type in the boundary area and the data points corresponding to other cell types; Determining the accuracy of the set hyperparameters at different values based on the overall distribution level of cell mixing of cell regions of all cell types at different values of the set hyperparameters; The set hyperparameter with the value corresponding to the maximum accuracy is taken as the set hyperparameter with the optimal value, and the data point graph after dimensionality reduction corresponding to the set hyperparameter with the optimal value is taken as the optimal dimensionality reduction result.
[0006] In combination with the first aspect above, in some possible implementations, determining the rationality index of each cell includes: Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before dimensionality reduction, and obtain a first cell type density; Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph after dimensionality reduction, and obtain the second cell type density; determining a difference between the density of the second cell type and the density of the first cell type to obtain a first difference; The product of the first difference and the second cell type density is determined to obtain a first product, and the first product is used as a rationality indicator for each cell.
[0007] In combination with the first aspect above, in some possible implementations, determining the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction includes: Determine each boundary data point of the cell region formed by the data points corresponding to each cell type in the data point map after dimensionality reduction; Construct the k-neighborhood of each boundary data point, and then determine the neighborhood radius of each boundary data point; The minimum neighborhood radius among all neighborhood radiuses is used as the boundary distance; The area within the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction and within the boundary distance to the boundary line of the cell area is taken as the boundary area, and the other areas within the cell area except the boundary area are taken as the non-boundary area.
[0008] In combination with the first aspect above, in some possible implementations, further determining the neighborhood radius of each boundary data point includes: Determine the distance between each data point in the k-neighborhood of each boundary data point and the corresponding boundary data point as the candidate distance for each boundary data point; The maximum distance among all candidate distances of each boundary data point is determined, and the maximum distance is used as the neighborhood radius of each boundary data point.
[0009] In combination with the first aspect above, in some possible implementations, determining the degree of cell mixing in the cell region of each cell type includes: In the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the boundary clarity of the boundary area corresponding to each cell type is determined according to the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type; In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the degree of cell mixing of each other cell type in the boundary area corresponding to each cell type is determined according to the difference in the number of data points corresponding to the cell type and the data points corresponding to each other cell type, and the area proportion of the data points corresponding to each other cell type; The degree of cell mixing in the cell region of each cell type is determined according to the degree of cell mixing of all other cell types in the border region corresponding to each cell type and the boundary clarity of the border region corresponding to each cell type.
[0010] In combination with the first aspect above, in some possible implementations, determining the boundary clarity of the boundary area corresponding to each cell type includes: In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the average value of the rationality index of the cells corresponding to all the data points corresponding to the cell type to obtain the first rationality index mean value; In the non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the average value of the rationality index of the cells corresponding to all the data points corresponding to the cell type to obtain the second rationality index mean; Determining the boundary clarity of the boundary area corresponding to each cell type according to the difference between the mean value of the first rationality index and the mean value of the second rationality index; The difference between the mean of the first rationality index and the mean of the second rationality index is determined to obtain a second difference, and the second difference is subjected to negative correlation normalization processing, so as to obtain the boundary clarity of the boundary area corresponding to each cell type.
[0011] In combination with the first aspect above, in some possible implementations, determining the cell mixing degree of each other cell type in the boundary region corresponding to each cell type includes: In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the difference between the number of data points corresponding to the cell type and the number of data points corresponding to each other cell type to obtain a third difference, and perform negative correlation processing on the third difference to obtain a difference processing result; Determine the product of the difference processing result and the area proportion of the data points corresponding to each other cell type in the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction, and obtain a second product, and use the second product as the cell mixing degree of each other cell type in the boundary area corresponding to each cell type.
[0012] In combination with the first aspect above, in some possible implementations, determining the degree of cell mixing in the cell region of each cell type includes: Determine the maximum cell mixing degree among the cell mixing degrees of all other cell types in the boundary region corresponding to each cell type under different values of the set hyperparameters; The ratio of the maximum cell mixing degree to the boundary definition of the boundary area corresponding to each cell type was determined, and the normalized value of the ratio was used as the cell mixing degree of the cell area of each cell type.
[0013] In conjunction with the first aspect, in some possible implementations, determining the accuracy of setting hyperparameters with different values includes: Determine the average of the cell mixing degree of the cell regions of all cell types in the data point graph after dimensionality reduction with different values of the set hyperparameters, and obtain the average cell mixing degree; The average cell mixing degree was negatively correlated to obtain the accuracy of the set hyperparameters with different values.
[0014] In combination with the first aspect above, in some possible implementations, the umap algorithm is used to perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and the hyperparameter is set to the number of neighbors.
[0015] In order to solve the above technical problems, in a second aspect, the present invention also provides a fast and stable single cell type labeling device, the device comprising: A cell type identification module, used to identify the cell types of different cells based on the sequencing data of different cells; A data point graph acquisition module is used to perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtain data point graphs before and after dimensionality reduction, where each data point in the data point graph corresponds to a cell; A rationality acquisition module is used to determine the rationality index of each cell according to the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction. The rationality index reflects the rationality of the setting of the current value of the hyperparameter for each cell; A mixing degree acquisition module is used to determine the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction, and determine the cell mixing degree of the cell area of each cell type according to the difference between the rationality indexes of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area, and the distribution of the data points of the cell type in the boundary area and the data points corresponding to other cell types; An accuracy acquisition module, for determining the accuracy of different values of the set hyperparameters according to the overall distribution level of the degree of cell mixing of the cell regions of all cell types under different values of the set hyperparameters; The screening module is used to take the setting hyperparameter with the value corresponding to the maximum accuracy as the setting hyperparameter with the optimal value, and take the reduced-dimensional data point graph corresponding to the setting hyperparameter with the optimal value as the optimal dimensionality reduction result.
[0016] To solve the above technical problems, in a third aspect, the present invention also provides a fast and stable single cell type annotation system, comprising a memory and a processor. The memory is used to store an executable computer program, and the processor is used to call and run the executable computer program from the memory, so that the system executes the method in the above first aspect or any possible implementation of the first aspect.
[0017] To solve the above technical problems, in a fourth aspect, the present invention further provides a computer program product, which includes: a computer program code, when the computer program code runs on a computer, enables the computer to execute the method in the above first aspect or any possible implementation of the first aspect.
[0018] To solve the above technical problems, in a fifth aspect, the present invention further provides a computer-readable storage medium, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the above first aspect or any possible implementation of the first aspect.
[0019] The present invention has the following beneficial effects: the present invention identifies the cell types of different cells based on the sequencing data of different cells, and performs dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtains a data point diagram before and after dimensionality reduction, wherein each data point in the data point diagram corresponds to a cell; the data point diagram before and after dimensionality reduction processing is analyzed, and the rationality index of each cell is determined according to the difference between the number distribution of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point diagram before and after dimensionality reduction; the degree of cell mixing of the cell region of each cell type under different set hyperparameters is determined, and then the difference between the rationality indexes of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area of the cell region constituted by the data points corresponding to each cell type in the data point diagram after dimensionality reduction, as well as the distribution of the data points of the cell type to other cell types in the boundary area, is combined to determine the degree of cell mixing of the cell region of each cell type, and finally determine the accuracy of the set hyperparameters with different values; the set hyperparameter corresponding to the value with the maximum accuracy is used as the set hyperparameter with the optimal value, and the data point diagram after dimensionality reduction corresponding to the set hyperparameter with the optimal value is used as the optimal dimensionality reduction result. The present invention performs dimensionality reduction processing on the sequencing data of single cells under different set hyperparameters, and analyzes the obtained data point graphs before and after dimensionality reduction, so as to adaptively determine the optimal set hyperparameters, thereby effectively improving the accuracy of the dimensionality reduction results in the single cell type labeling process. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A flowchart of the steps of a fast and stable single cell type labeling method according to an embodiment of the present invention; Figure 2 A structural diagram of sequencing data of a cell according to an embodiment of the present invention; Figure 3 This is a schematic diagram of data point distribution determined after dimensionality reduction processing is performed on sequencing data of different cells with different values of the number of neighbors using the umap algorithm in an embodiment of the present invention; Figure 4 A schematic diagram of a boundary area and a non-boundary area of a cell area according to an embodiment of the present invention; Figure 5Schematic diagram of the structure of a fast and stable single cell type labeling device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to clearly illustrate the technical features of the present invention, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0023] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.
[0024] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0025] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0026] It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0027] Although operations or steps are described in a specific order in the drawings in the embodiments of the present invention, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or that all the operations or steps shown must be performed to obtain the desired results. In the embodiments of the present invention, these operations or steps may be performed in series; these operations or steps may also be performed in parallel; or some of these operations or steps may be performed.
[0028] At the same time, it is understood that the data involved in the technical solution of the present invention (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the relevant laws, regulations and relevant provisions. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meanings as those generally understood by technicians in the technical field of the present invention, and all parameters or indicators in the formulas involved in the present invention are normalized values that eliminate the influence of dimensions.
[0029] In order to solve the problem that the dimensionality reduction results of cell data are not accurate due to unreasonable selection of hyperparameters used in the existing single cell type labeling process, an embodiment of the present invention provides a fast and stable single cell type labeling method, which effectively improves the accuracy of the dimensionality reduction results of cell data by adaptively determining the optimal values of the set hyperparameters.
[0030] A fast and stable single cell type labeling method provided by an embodiment of the present invention will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 FIG. 4 is a schematic diagram showing a basic process of a fast and stable single cell type labeling method provided by an embodiment of the present invention. Figure 1 As shown, the method specifically comprises the following steps: Step S100: Identify the cell types of different cells based on the sequencing data of different cells.
[0032] Specifically, a pre-trained cell type recognition model is obtained, sequencing data of different cells are input into the cell type recognition model, the cell types of different cells are recognized using the cell type recognition model, and the cell types of different cells are output, thereby realizing cell category recognition of different cells.
[0033] In this embodiment, training data is obtained from an open source database cell marker, and the training data refers to sequencing data of different cells. The training data is input into the transformer network constituting the cell type recognition model to match the gene features of the cells. The matching result is used as the feature of the single cell and input into the linear classifier. The linear classifier is a fully connected layer, and the activation function is softmax. The linear classifier will output a vector, and the dimension of the vector is the number of all cell categories. The value of each dimension in the vector is the possibility of the labeled cell category of the single cell. The value of each dimension in the vector is normalized, and the largest value in these dimensions is selected as the cell category of the current single cell. A cell category label is set for each cell corresponding to the training data, and the cell category obtained by the cell type recognition model is compared with the cell category label of each cell corresponding to the training data. The cell type recognition model is continuously iterated and optimized to improve the accuracy of recognition, and finally complete the training of the cell type recognition model.
[0034] Users upload sequencing data of different cells through the front-end interface. The sequencing data includes gene expression matrix, gene name, and cell barcode. Figure 2 The structure diagram of sequencing data is shown. Based on the sequencing data of different cells, the cell type recognition model composed of the transformer network is used to complete the type recognition of single cells, and the output vector and the cell category label of the corresponding single cell are recorded at the same time, thereby realizing the cell category recognition of different cells.
[0035] It should be understood that the above is only a specific method for identifying the cell types of different cells given in this embodiment. Other methods or models can also be used to achieve cell type identification of different cells, which will not be elaborated here.
[0036] Step S200: Perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtain data point graphs before and after dimensionality reduction, where each data point in the data point graph corresponds to a cell.
[0037] Specifically, the above-mentioned cell type recognition model composed of a transformer network is used to match the genetic features of each cell to complete the type recognition of single cells. However, the sequencing data of single cells often have high-dimensional characteristics. There are thousands of gene expression information in a cell. It is impossible to directly observe and understand the relationship between cells in a high-dimensional space. Therefore, it is necessary to use the UMAP algorithm to reduce the dimensionality of the sequencing data of single cells, so as to facilitate more direct observation of how different types of single cell populations are distributed and the distance between them. It helps to grasp the heterogeneity and similarity of cell populations as a whole, distinguish the clustering morphology of different cell subpopulations in low-dimensional space, and complete the single cell type labeling.
[0038] In the process of using the UMAP algorithm to reduce the dimension of single-cell sequencing data, the number of neighbors n_neighbors is an important hyperparameter in the UMAP algorithm, which has a significant impact on the dimension reduction results, and its default value is 15. If the number of neighbors is set too small, cells that originally belonged to the same cell subpopulation will be over-dispersed in the low-dimensional space after dimensionality reduction, and cannot form tight and clear clusters, resulting in the inability to intuitively distinguish different types of cell populations in the data after dimensionality reduction, making the entire cell distribution chaotic; on the contrary, if the number of neighbors is set too large, different cell subpopulations will be merged, resulting in cells that should have been distinguished from each other overlapping after dimensionality reduction by the UMAP algorithm, and the purpose of showing the differences in cell populations cannot be achieved.
[0039] Therefore, in order to determine the optimal number of neighbors, the output dimension of the UMAP algorithm n_components=2 is set, and the UMAP algorithm is used to reduce the dimensionality of the sequencing data of different cells under different values of the number of neighbors, and the data point graphs before and after dimensionality reduction are obtained, and each data point in the data point graph corresponds to a cell. Figure 3 The figure shows the distribution of data points obtained after dimensionality reduction processing of sequencing data of different cells using the UMAP algorithm with different values of the number of neighbors.
[0040] Among them, when using the UMAP algorithm to perform dimensionality reduction processing on sequencing data of different cells under different values of the number of neighbors, the interval of the number of neighbors and the traversal step can be set to obtain the corresponding neighbor number traversal sequence, so as to use the UMAP algorithm to perform dimensionality reduction processing on sequencing data of different cells under each number of neighbors in the neighbor number traversal sequence, and obtain the data point graph before and after each dimensionality reduction corresponding to each number of neighbors in the neighbor number traversal sequence.
[0041] Step S300: Determine the rationality index of each cell based on the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction. The rationality index reflects the rationality of the setting of the current value of the set hyperparameter for each cell.
[0042] Specifically, for each cell, that is, each cell barcode in the sequencing data, the cell barcode is the unique identifier of the cell, set the k value, and obtain the k-neighborhood corresponding to the data point corresponding to the cell in the data point map before dimensionality reduction. It should be noted that this neighborhood reflects the genetic similarity between cells. For example, if the neighborhood of the data point corresponding to cell A contains data points corresponding to cells B, C, etc., it indicates that these cells have a high degree of similarity in gene expression levels and belong to a common expression of a specific gene combination. At the same time, obtain the k-neighborhood corresponding to the data point corresponding to the cell in the data point map after dimensionality reduction.
[0043] Under a certain number of neighbors, if the cells can still be closely clustered together to form relatively independent clusters after dimensionality reduction, it means that the number of neighbors is appropriately selected. Therefore, the cells in the neighborhood after dimensionality reduction show consistency of cell types, which means that the currently selected number of neighbors will not cause a mixture of different cell types in the umap dimensionality reduction result, and different cells can be well distinguished, so the current number of neighbors is appropriately selected.
[0044] Furthermore, the rationality index of each cell is determined based on the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction, and the implementation steps include: Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before dimensionality reduction, and obtain a first cell type density; Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph after dimensionality reduction, and obtain the second cell type density; determining a difference between the density of the second cell type and the density of the first cell type to obtain a first difference; The product of the first difference and the second cell type density is determined to obtain a first product, and the first product is used as a rationality indicator for each cell.
[0045] In the above steps, in this embodiment, for any cell in the data point graph after dimensionality reduction , the second cell type density is determined by the following formula: ; Where: Represents the cell after dimensionality reduction density of the second cell type; Represents the cell after dimensionality reduction The corresponding data point is in the neighborhood of the cell is the number of data points of the same cell class; Represents the cell after dimensionality reduction The number of data points in the neighborhood of the corresponding data point; Indicates the sign after dimensionality reduction.
[0046] In the above formula, the cell after dimensionality reduction Density of the second cell type The larger the value of The stronger the consistency of cell types in the neighborhood of the corresponding data point, the greater the probability that the number of neighbors belongs to the optimal number of neighbors.
[0047] In the same way, the cells before dimensionality reduction can be determined The density of the first cell type is , Indicates the sign before dimensionality reduction.
[0048] Considering that cells identified as the same general type also have internal heterogeneity, there may be different cell subpopulations, and the differences between these cell subpopulations sometimes overlap with other types of cells. For example, in the tumor microenvironment, different T cell subpopulations in immune cells, such as regulatory T cells and cytotoxic T cells, have different functions and gene expression characteristics, and the characteristics of some of these subpopulations may be similar to other types of cells such as tumor-associated fibroblasts in some aspects, resulting in these cells of different identification types appearing in each other's neighborhoods when constructing neighborhoods. This reflects the complex heterogeneity of the cell population itself, which is a normal phenomenon caused by differences in the intrinsic characteristics of cells.
[0049] Therefore, considering that the above differences will be reflected in cell genes, if there are obvious category differences in the neighborhood of the cell before dimensionality reduction, and the consistency in the neighborhood of the cell after dimensionality reduction is also poor, it is considered a normal phenomenon; if there is a strong inconsistency in the cell neighborhood before and after dimensionality reduction of the sequencing data, but the consistency is still significantly improved after dimensionality reduction, it can also be explained that after dimensionality reduction under the current number of neighbors, there is a better cell cluster performance in the dimensionality reduction space, thereby making up for the defects in the above-mentioned type density.
[0050] In this embodiment, for any cell after dimensionality reduction, , and its rationality index is determined by the following formula: ; Where: Represents cells rationality indicators; Represents the cell after dimensionality reduction density of the second cell type; Represents the cell before dimensionality reduction The density of the first cell type.
[0051] In the above formula, Represents cells before and after dimensionality reduction The degree of improvement in the consistency of cell types in the neighborhood is calculated by comparing the degree of improvement with the cell type consistency after dimensionality reduction. Density of the second cell type The product of , thus obtaining the cell Rationality index The rationality index of a cell reflects the rationality of the current number of neighbors for the cell, preventing the same type of cells from diverging after dimensionality reduction due to a small number of neighbors, thereby affecting the accuracy of dimensionality reduction for cell judgment.
[0052] Step S400: Determine the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point diagram after dimensionality reduction, and determine the degree of cell mixing in the cell area of each cell type based on the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area, as well as the distribution of the data points of the cell type and the corresponding data points of other cell types in the boundary area.
[0053] Specifically, in the process of finding the optimal number of neighbors, if the optimal number of neighbors for dimensionality reduction by the UMAP algorithm is exceeded, the number of neighbors will be seriously too large, so that although the cells of the same type show obvious aggregation, the rationality of the cells calculated above is too large, but the cells of different types will also show obvious aggregation, resulting in inaccurate representation of different types of cells, which leads to inaccurate structure of cells obtained after dimensionality reduction. For example, if the number of neighbors is too large, the originally different cell subpopulations will be synthesized into a larger group in the space after dimensionality reduction, which is manifested as different types of cells mixed together in the space after dimensionality reduction. Therefore, it is also necessary to combine the type information of these cells to further judge the rationality of the number of neighbors. It is necessary to consider the density characteristics of each cell to determine the rationality, that is, while considering the characteristics of cells of the same type, it is also necessary to consider the ability to distinguish between different types of cells to prevent the number of neighbors from being too large.
[0054] When different types of cells are mixed, it is mainly manifested in the boundary part of the cell area composed of the data points corresponding to these different types of cells. Therefore, it is necessary to first obtain the boundary area and non-boundary area of the cell area composed of the data points corresponding to each cell type after dimensionality reduction.
[0055] Furthermore, the above steps of determining the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction include: Determine each boundary data point of the cell region formed by the data points corresponding to each cell type in the data point map after dimensionality reduction; Construct the k-neighborhood of each boundary data point, and then determine the neighborhood radius of each boundary data point; The minimum neighborhood radius among all neighborhood radiuses is used as the boundary distance; The area within the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction and within the boundary distance to the boundary line of the cell area is taken as the boundary area, and the other areas within the cell area except the boundary area are taken as the non-boundary area.
[0056] According to the above steps, the data points corresponding to all cells of the same type in the data point map after dimensionality reduction are obtained, and the convex hull corresponding to the data points corresponding to all cells is obtained by using the Graham algorithm to obtain the position of the cell area formed by the data points corresponding to all cells of this type. The above operation is performed on the data points corresponding to all cells of different types to obtain the position and area information of the cell area formed by the data points corresponding to different types of cells.
[0057] Select any cell type a as the analysis object, obtain the data points whose corresponding data points constitute the boundary of the cell region as the boundary data points of the cell type a, obtain the k-neighborhood after dimensionality reduction constructed by each boundary data point, and determine the neighborhood radius at this time, take the minimum value of the neighborhood radius of each boundary data point, that is, the minimum neighborhood radius as the boundary distance, take the area within the boundary distance in the cell region constituted by the corresponding data points of the cells of cell type a as the boundary area of the cell region constituted by the corresponding data points of the cells of cell type a, and record it as At the same time, the area in the cell area except the boundary area is taken as the boundary area of the cell area composed of the corresponding data points of the cells of cell type a, and recorded as The process of determining the k-neighborhood belongs to the prior art and will not be described here. Figure 4 It shows that the cell corresponding data points of a certain cell type constitute the boundary area and non-border area of the cell area, where 1 represents the cell area, 2 represents the boundary area, and 3 represents the non-border area.
[0058] Among them, in this embodiment, the step of determining the neighborhood radius includes: determining the distance between each data point in the k-neighborhood of each boundary data point and the corresponding boundary data point as the candidate distance for each boundary data point; determining the maximum distance among all the candidate distances for each boundary data point, and using the maximum distance as the neighborhood radius of each boundary data point.
[0059] For any cell type a, the corresponding data points constitute the boundary area of the cell area , if the boundary area The cells corresponding to all the data points in the image are still mainly cell type a, and for the boundary areas with a certain degree of mixing, the fewer the mixed boundary areas, the smaller the degree of boundary mixing, which means that the current number of neighbors is reasonably selected; on the contrary, if the number of neighbors is too large, there will be a large number of data points corresponding to cells of other cell types in the boundary area, and the more obvious the mixing between cells of different cell types, the deeper the degree of boundary mixing.
[0060] Furthermore, the above-mentioned determination of the degree of cell mixing in the cell region of each cell type comprises the following steps: In the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the boundary clarity of the boundary area corresponding to each cell type is determined according to the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type; In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the degree of cell mixing of each other cell type in the boundary area corresponding to each cell type is determined according to the difference in the number of data points corresponding to the cell type and the data points corresponding to each other cell type, and the area proportion of the data points corresponding to each other cell type; The degree of cell mixing in the cell region of each cell type is determined according to the degree of cell mixing of all other cell types in the border region corresponding to each cell type and the boundary clarity of the border region corresponding to each cell type.
[0061] For the above steps, in the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point diagram after dimensionality reduction, if the difference between the rationality indicators of the cells corresponding to the data points of the cell type is small, it means that the cell distribution of the cell type in the boundary area and the non-boundary area is relatively similar, and the possibility of mixing multiple other types of cells is small. At this time, the boundary clarity of the boundary area corresponding to the cell type is higher.
[0062] Furthermore, in the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the boundary clarity of the boundary area corresponding to each cell type is determined according to the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type, that is: in the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the average value of the rationality indicators of the cells corresponding to all the data points of the cell type is determined to obtain the first rationality indicator mean; in the non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the average value of the rationality indicators of the cells corresponding to all the data points of the cell type is determined to obtain the second rationality indicator mean; according to the difference between the first rationality indicator mean and the second rationality indicator mean, the boundary clarity of the boundary area corresponding to each cell type is determined.
[0063] For the above steps, the data points corresponding to any cell type a constitute the boundary area of the cell area and non-border areas , determine the boundary area The mean of the rationality index of cells corresponding to all data points of cell type a is obtained to obtain the first rationality index mean, which is recorded as , and determine the non-boundary area The mean of the rationality index of cells corresponding to all data points of cell type a is obtained to obtain the mean of the second rationality index, which is recorded as According to the mean of the first rationality index The mean of the second rationality index The difference between the two determines the boundary area corresponding to cell type a The smaller the difference, the larger the value of boundary clarity.
[0064] In this embodiment, according to the first rationality indicator mean The mean of the second rationality index The difference between them is used to determine the boundary area corresponding to cell type a by the following formula Boundary clarity: ; Where: Indicates the boundary area corresponding to cell type a The clarity of the boundaries; Represents the normalization function.
[0065] In the above formula, if the first rationality index mean The mean of the second rationality index The smaller the difference between them, the smaller the boundary area corresponding to cell type a. The cells corresponding to the inner data points and the non-border area The cells corresponding to the inner data points have a high similarity, so the possibility of mixing multiple types of cells is smaller, and the clarity of the boundary area is higher, and the corresponding The larger the value of .
[0066] In the above manner, the boundary clarity of the border region corresponding to each cell type can be determined.
[0067] After determining the boundary clarity of the boundary area corresponding to each cell type, for any cell type a, if the boundary area of the cell area formed by the data points corresponding to the cells of cell type a The more data points corresponding to cells of cell type a are present, the more significant the boundary is. If the area of data points corresponding to cells of other cell types is large, and the number of data points corresponding to cells of other cell types is close to the number of data points corresponding to cells of cell type a, it means that the boundary area The more mixed the data points corresponding to cells of other cell types are, the less significant the boundary will be.
[0068] Furthermore, the above-mentioned determination of the degree of cell mixing of each other cell type in the boundary area corresponding to each cell type includes the following steps: In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the difference between the number of data points corresponding to the cell type and the number of data points corresponding to each other cell type to obtain a third difference, and perform negative correlation processing on the third difference to obtain a difference processing result; Determine the product of the difference processing result and the area proportion of the data points corresponding to each other cell type in the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction, and obtain a second product, and use the second product as the cell mixing degree of each other cell type in the boundary area corresponding to each cell type.
[0069] In view of the above steps, in this embodiment, for any cell type a, the boundary area of the cell area formed by the data points corresponding to the cells is , determine the boundary area The cell region composed of all corresponding data points of any other cell type b in Since the cell area The process of determining is exactly the same as the process of determining the area constituted by the data points corresponding to the cells of cell type a, and will not be repeated here.
[0070] The boundary area of the cell area formed by the data points corresponding to the cells of cell type a In the cell area, the corresponding data points of all cells of cell type b constitute The area proportion, as well as the difference between the number of data points corresponding to cell type a and the number of data points corresponding to cell type b, the boundary area corresponding to cell type a is determined by the following formula The degree of cell mixing within cell type b: ; Where: Indicates the boundary area corresponding to cell type a the degree of cell mixing within cell type b; Represents the boundary area of the cell area composed of data points corresponding to cells of cell type a In the cell type b, the area of the cell region constituted by the corresponding data points of the cells; represents the area of the boundary region corresponding to cell type a; Respectively represent the boundary area of the cell area composed of the data points corresponding to the cells of cell type a In the figure, the percentage of data points corresponding to cell type a and the percentage of data points corresponding to cell type b; Represents the normalization function.
[0071] In the above formula, Indicates the boundary area corresponding to cell type a The cell area consisting of the data points corresponding to cell type b The area ratio reflects the boundary area corresponding to cell type a. The greater the value, the more serious the mixing degree. At the same time, the boundary area of the cell area formed by the data points corresponding to the cells of cell type a The percentage of data points corresponding to cell type a The proportion of data points corresponding to cell type b The smaller the difference, the closer the boundary area corresponding to cell type a is. The cells of inner cell type b are more mixed.
[0072] According to the above method, the boundary area corresponding to cell type a can be determined Correspondingly, the degree of cell mixing of each other cell type in the boundary area corresponding to each cell type can be determined.
[0073] Based on the cell mixing degree of all other cell types in the boundary region corresponding to each cell type and the boundary clarity of the boundary region corresponding to each cell type, the cell mixing degree of the cell region of each cell type can be determined. Further, determining the cell mixing degree of the cell region of each cell type includes: determining the maximum cell mixing degree among the cell mixing degrees of all other cell types in the boundary region corresponding to each cell type under different set hyperparameters; determining the ratio of the maximum cell mixing degree to the boundary clarity of the boundary region corresponding to each cell type, and using the normalized value of the ratio as the cell mixing degree of the cell region of each cell type.
[0074] With respect to the above steps, in this embodiment, for any cell type a corresponding to the boundary region The cell mixing degree of other cell types in the region is selected, and the largest cell mixing degree is selected as the boundary area corresponding to cell type a. The degree of cell mixing. At the same time, the boundary area corresponding to cell type a The higher the boundary clarity, the smaller the degree of cell mixing should be. Therefore, the boundary area corresponding to cell type a is used The boundary clarity of the cell type a corresponds to the boundary area The degree of cell mixing is corrected to finally obtain the boundary area corresponding to cell type a The degree of cell mixing: ; Where: Indicates the boundary area corresponding to cell type a The degree of cell mixing; Indicates the boundary area corresponding to cell type a The maximum cell mixing degree among the cell mixing degrees of other cell types in Indicates the boundary area corresponding to cell type a The degree of cell mixing with any other cell type b in the Indicates taking the maximum value function, that is, taking the boundary area corresponding to cell type a The maximum value among all other cell types in the cell population; Represents the normalization function.
[0075] In the above manner, the cell mixing degree of the cell region of each cell type can be determined, and the cell mixing degree reflects the accuracy of the current neighbor number selection. The higher the cell mixing degree, the lower the accuracy of the previous neighbor number selection.
[0076] Step S500: determining the accuracy of the set hyperparameters with different values according to the overall distribution level of the cell mixing degree of the cell regions of all cell types under the set hyperparameters with different values.
[0077] Specifically, at each value of the number of neighbors in the neighbor number traversal sequence, the cell mixing degree of the cell regions of all cell types is comprehensively considered to determine the accuracy of the set hyperparameters of the current value. When the cell regions of different cell types have a small degree of cell mixing, it means that the accuracy of the current value of the number of neighbors is higher.
[0078] Furthermore, the overall distribution level of the cell mixing degree of the cell regions of all cell types under the above-mentioned different values of the set hyperparameters is determined to determine the accuracy of the set hyperparameters with different values, and the implementation steps include: Determine the maximum cell mixing degree among the cell mixing degrees of all other cell types in the boundary region corresponding to each cell type under different values of the set hyperparameters; The ratio of the maximum cell mixing degree to the boundary definition of the boundary area corresponding to each cell type was determined, and the normalized value of the ratio was used as the cell mixing degree of the cell area of each cell type.
[0079] For the above steps, in this embodiment, according to the overall distribution level of the cell mixing degree of the cell regions of all cell types under different values of the set hyperparameters, the accuracy of the set hyperparameters with different values is determined by the following formula: ; Where: Indicates the accuracy of the set hyperparameters for each value, that is, the accuracy of the number of neighbors for each value; Represents the average cell mixing degree, that is, the cell mixing degree of the cell region of all cell types under each value of the set hyperparameter The average value of .
[0080] In the above manner, the accuracy of the number of neighbors for each value, that is, the accuracy of the set hyperparameters for each value, can be determined.
[0081] Step S600: The setting hyperparameter corresponding to the value of maximum accuracy is used as the setting hyperparameter with the optimal value, and the dimensionality-reduced data point graph corresponding to the setting hyperparameter with the optimal value is used as the optimal dimensionality reduction result.
[0082] Specifically, for each value of the number of neighbors in the neighbor number traversal sequence, the number of neighbors with the greatest accuracy is selected as the optimal value of the number of neighbors, and the data point graph after dimensionality reduction at this time is used as the best dimensionality reduction result. Based on the best dimensionality reduction result, single cell type labeling is completed, such as using different colors or symbols to label different types of cells in the best dimensionality reduction result, analyzing the potential differences between different cells, so as to perform more accurate clustering analysis to discover new cell subpopulations, etc.
[0083] Based on the same inventive concept, the embodiment of the present invention also provides a fast and stable single cell type labeling device, such as Figure 5 As shown, the device comprises: A cell type identification module, used to identify the cell types of different cells based on the sequencing data of different cells; A data point graph acquisition module is used to perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtain data point graphs before and after dimensionality reduction, where each data point in the data point graph corresponds to a cell; A rationality acquisition module is used to determine the rationality index of each cell according to the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction. The rationality index reflects the rationality of the setting of the current value of the hyperparameter for each cell; A mixing degree acquisition module is used to determine the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction, and determine the cell mixing degree of the cell area of each cell type according to the difference between the rationality indexes of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area, and the distribution of the data points of the cell type in the boundary area and the data points corresponding to other cell types; An accuracy acquisition module, for determining the accuracy of different values of the set hyperparameters according to the overall distribution level of the degree of cell mixing of the cell regions of all cell types under different values of the set hyperparameters; The screening module is used to take the setting hyperparameter with the value corresponding to the maximum accuracy as the setting hyperparameter with the optimal value, and take the reduced-dimensional data point graph corresponding to the setting hyperparameter with the optimal value as the optimal dimensionality reduction result.
[0084] It should be noted that the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above.
[0085] Based on the same inventive concept, an embodiment of the present invention also provides a fast and stable single-cell type labeling system, which includes: a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the system can execute any one of the fast and stable single-cell type labeling methods introduced above.
[0086] The embodiment of the present invention can divide the system into functional modules according to the above method example. For example, each functional module can be corresponded to, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0087] Based on the same inventive concept, an embodiment of the present invention further provides a computer program product, which includes: a computer program code, which, when executed on a computer, enables the computer to execute any one of the aforementioned fast and stable single cell type labeling methods.
[0088] Based on the same inventive concept, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program code. When the computer program code runs on a computer, the computer executes any one of the fast and stable single-cell type labeling methods introduced above.
[0089] It should be noted that the above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A fast and stable single cell type annotation method, characterized in that: The following steps are involved: Identify cell types of different cells based on sequencing data of different cells; Perform dimensionality reduction processing on the sequencing data of a single cell under different set hyperparameters, and obtain data point graphs before and after dimensionality reduction, where each data point in the data point graph corresponds to a cell; According to the difference between the number distribution of data points of the same cell category as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before and after dimensionality reduction, the rationality index of each cell is determined. The rationality index reflects the rationality of the setting of the current value of the hyperparameter for each cell. Determine the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, and determine the degree of cell mixing in the cell area of each cell type based on the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type in the boundary area and the non-boundary area, and the distribution of the data points of the cell type in the boundary area and the data points corresponding to other cell types; Determining the accuracy of the set hyperparameters at different values based on the overall distribution level of cell mixing of cell regions of all cell types at different values of the set hyperparameters; The set hyperparameter with the value corresponding to the maximum accuracy is taken as the set hyperparameter with the optimal value, and the data point graph after dimensionality reduction corresponding to the set hyperparameter with the optimal value is taken as the optimal dimensionality reduction result.
2. A fast and stable single cell type labeling method according to claim 1, characterized in that: Determine the plausibility indicators for each cell, including: Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph before dimensionality reduction, and obtain a first cell type density; Determine the proportion of data points of the same cell type as the cell in the neighborhood of the data point corresponding to each cell in the data point graph after dimensionality reduction, and obtain the second cell type density; determining a difference between the density of the second cell type and the density of the first cell type to obtain a first difference; The product of the first difference and the second cell type density is determined to obtain a first product, and the first product is used as a rationality indicator for each cell.
3. A fast and stable single cell type labeling method according to claim 1, characterized in that: Determine the boundary area and non-boundary area of the cell area composed of the data points corresponding to each cell type in the data point map after dimensionality reduction, including: Determine each boundary data point of the cell region formed by the data points corresponding to each cell type in the data point map after dimensionality reduction; Construct the k-neighborhood of each boundary data point, and then determine the neighborhood radius of each boundary data point; The minimum neighborhood radius among all neighborhood radiuses is used as the boundary distance; The area within the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction and within the boundary distance to the boundary line of the cell area is taken as the boundary area, and the other areas within the cell area except the boundary area are taken as the non-boundary area.
4. A fast and stable single cell type labeling method according to claim 3, characterized in that: Then determine the neighborhood radius of each boundary data point, including: Determine the distance between each data point in the k-neighborhood of each boundary data point and the corresponding boundary data point as the candidate distance for each boundary data point; The maximum distance among all candidate distances of each boundary data point is determined, and the maximum distance is used as the neighborhood radius of each boundary data point.
5. A fast and stable single cell type labeling method according to claim 1, characterized in that: Determine the degree of cellular intermixing of the cell regions for each cell type, including: In the boundary area and non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the boundary clarity of the boundary area corresponding to each cell type is determined according to the difference between the rationality indicators of the cells corresponding to the data points corresponding to the cell type; In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, the degree of cell mixing of each other cell type in the boundary area corresponding to each cell type is determined according to the difference in the number of data points corresponding to the cell type and the data points corresponding to each other cell type, and the area proportion of the data points corresponding to each other cell type; The degree of cell mixing in the cell region of each cell type is determined according to the degree of cell mixing of all other cell types in the border region corresponding to each cell type and the boundary clarity of the border region corresponding to each cell type.
6. A fast and stable single cell type labeling method according to claim 5, characterized in that: Determine the boundary definition of the border regions corresponding to each cell type, including: In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the average value of the rationality index of the cells corresponding to all the data points corresponding to the cell type to obtain the first rationality index mean value; In the non-boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the average value of the rationality index of the cells corresponding to all the data points corresponding to the cell type to obtain the second rationality index mean; Determining the boundary clarity of the boundary area corresponding to each cell type according to the difference between the mean value of the first rationality index and the mean value of the second rationality index; The difference between the mean of the first rationality index and the mean of the second rationality index is determined to obtain a second difference, and the second difference is subjected to negative correlation normalization processing, so as to obtain the boundary clarity of the boundary area corresponding to each cell type.
7. A fast and stable single cell type labeling method according to claim 5, characterized in that: Determine the degree of intermixing of cells of each cell type within the boundary region corresponding to each cell type, including: In the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point graph after dimensionality reduction, determine the difference between the number of data points corresponding to the cell type and the number of data points corresponding to each other cell type to obtain a third difference, and perform negative correlation processing on the third difference to obtain a difference processing result; Determine the product of the difference processing result and the area proportion of the data points corresponding to each other cell type in the boundary area of the cell area constituted by the data points corresponding to each cell type in the data point map after dimensionality reduction, and obtain a second product, and use the second product as the cell mixing degree of each other cell type in the boundary area corresponding to each cell type.
8. A fast and stable single cell type labeling method according to claim 5, characterized in that: Determine the degree of cellular intermixing of the cell regions for each cell type, including: Determine the maximum cell mixing degree among the cell mixing degrees of all other cell types in the boundary region corresponding to each cell type under different values of the set hyperparameters; The ratio of the maximum cell mixing degree to the boundary definition of the boundary area corresponding to each cell type was determined, and the normalized value of the ratio was used as the cell mixing degree of the cell area of each cell type.
9. A fast and stable single cell type labeling method according to claim 1, characterized in that: Determine the accuracy of setting hyperparameters for different values, including: Determine the average of the cell mixing degree of the cell regions of all cell types in the data point graph after dimensionality reduction with different values of the set hyperparameters, and obtain the average cell mixing degree; The average cell mixing degree was negatively correlated to obtain the accuracy of the set hyperparameters with different values.
10. A fast and stable single cell type labeling method according to claim 1, characterized in that: The umap algorithm was used to perform dimensionality reduction on the sequencing data of a single cell under different values of the set hyperparameters, and the hyperparameter was set to the number of neighbors.
Citation Information
Patent Citations
Differential analysis method and system based on single cell samples of mixed experimental group and control group
CN114864003A
Recognition method and device suitable for cells in cell cluster and electronic equipment
CN114897872A
Cross-species single cell annotation method
CN118298926A
Database visualization method based on single cell
CN119673278A
Systems and methods for identifying cell-associated barcodes in mutli-genomic feature data from single-cell partitions
US20220076780A1