Clustering processing method and apparatus, electronic device, and computer-readable storage medium

By using multiple clustering algorithms and outlier detection, combined with the Leiden algorithm and deep learning feature representation algorithm, the inaccuracy of clustering results caused by a single clustering algorithm is solved, and higher-precision clustering processing is achieved.

CN114298125BActive Publication Date: 2026-05-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-10-22
Publication Date
2026-05-29

Smart Images

  • Figure CN114298125B_ABST
    Figure CN114298125B_ABST
Patent Text Reader

Abstract

The application discloses a clustering processing method and device, electronic equipment and a computer readable storage medium, and belongs to the technical field of data processing. The method comprises the following steps: obtaining a plurality of first clusters and a plurality of second clusters, the plurality of first clusters are obtained by performing clustering processing on a plurality of object data based on a first clustering algorithm, at least one object data is included in each first cluster, the plurality of second clusters are obtained by performing clustering processing on the plurality of object data based on a second clustering algorithm, and at least one object data is included in each second cluster; performing outlier detection processing on the plurality of object data to obtain an outlier detection result of each object data; and determining a plurality of third clusters based on the plurality of first clusters, the plurality of second clusters and the outlier detection result of each object data, and at least one object data is included in each third cluster. The application can improve the accuracy of a clustering result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a clustering processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With the rise of big data, data processing technologies have become increasingly important, and clustering is one of the key technologies in data processing. Clustering can group individual object data in a dataset into multiple clusters, with each cluster containing at least one object data.

[0003] In related technologies, a single clustering algorithm is typically used to cluster the various object data in a dataset. However, due to the wide variety of clustering algorithms available—including, but not limited to, the Leiden algorithm, the Louvain algorithm, and the current state-of-the-art (SOTA) algorithm based on deep learning feature representation—using only one clustering algorithm to cluster multiple object data results in low accuracy. Summary of the Invention

[0004] This application provides a clustering processing method, apparatus, electronic device, and computer-readable storage medium, which can be used to solve the problem of low accuracy of clustering results in related technologies. The technical solution includes the following contents.

[0005] On the one hand, embodiments of this application provide a clustering processing method, the method comprising:

[0006] Multiple first clusters and multiple second clusters are obtained. The multiple first clusters are obtained by clustering multiple object data based on a first clustering algorithm. Each first cluster includes at least one object data. The multiple second clusters are obtained by clustering the multiple object data based on a second clustering algorithm. Each second cluster includes at least one object data.

[0007] Outlier detection processing is performed on the multiple object data to obtain the outlier detection results for each object data.

[0008] Based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the object data, multiple third clusters are determined, and each third cluster includes at least one object data.

[0009] On the other hand, embodiments of this application provide a clustering processing apparatus, the apparatus comprising:

[0010] The acquisition module is used to acquire multiple first clusters and multiple second clusters. The multiple first clusters are obtained by clustering multiple object data based on a first clustering algorithm, and each first cluster includes at least one object data. The multiple second clusters are obtained by clustering the multiple object data based on a second clustering algorithm, and each second cluster includes at least one object data.

[0011] The detection module is used to perform outlier detection processing on the multiple object data to obtain the outlier detection results for each object data.

[0012] The determination module is used to determine multiple third clusters based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the object data, wherein each third cluster includes at least one object data.

[0013] In one possible implementation, the detection module is used to perform outlier detection processing on each object data in any one of the plurality of second clusters, and obtain outlier detection results for each object data in the any one of the second clusters.

[0014] The determining module is used to determine multiple third clusters based on the outlier detection results of the multiple first clusters and the object data in each of the multiple second clusters.

[0015] In one possible implementation, the determining module is configured to, for any one of the plurality of second clusters, in response to the presence of first object data in each object data of the plurality of second clusters, determine the first cluster corresponding to the plurality of second clusters, wherein the outlier detection result of the first object data is non-outlier object data; determine that the first cluster corresponding to the plurality of second clusters is a third cluster; and in response to the first object data not belonging to the third cluster, add the first object data to the third cluster.

[0016] In one possible implementation, the determining module is configured to determine a cross-tabulation based on the plurality of first clusters and the plurality of second clusters, wherein a row of data in the cross-tabulation represents the object data in a first cluster and a column of data in the cross-tabulation represents the object data in a second cluster; and determine the first cluster corresponding to any one of the second clusters from the plurality of first clusters based on the cross-tabulation.

[0017] In one possible implementation, a first cluster corresponds to a first cluster identifier, a second cluster corresponds to a second cluster identifier, and an object data includes an object identifier.

[0018] The determining module is used to determine the object identifier corresponding to each cluster identifier set based on the first cluster identifier of each first cluster, the object identifier of each object data contained in each first cluster, the second cluster identifier of each second cluster, and the object identifier of each object data contained in each second cluster, wherein a cluster identifier set includes one first cluster identifier and one second cluster identifier; and to determine the cross-tabulation based on the number of object identifiers corresponding to each cluster identifier set.

[0019] In one possible implementation, the cross-tab includes multiple non-zero data, wherein the non-zero data represents the number of identical object data contained in the first cluster corresponding to the row containing the non-zero data and the second cluster corresponding to the column containing the non-zero data.

[0020] The determining module is used to determine the largest non-zero data from the non-zero data contained in the column corresponding to any second cluster in the cross table; and to determine the first cluster corresponding to the row containing the largest non-zero data as the first cluster corresponding to any second cluster.

[0021] In one possible implementation, the determining module is configured to, for any one of the plurality of second clusters, in response to the presence of second object data in each object data of the any one of the second clusters, determine that the first cluster to which a second object data belongs is a third cluster, and the outlier detection result of the second object data is outlier object data.

[0022] In one possible implementation, the detection module is used to perform dimensionality reduction processing on the object data to obtain dimensionality-reduced object data; and to perform outlier detection processing on the dimensionality-reduced object data to obtain outlier detection results for the object data.

[0023] In one possible implementation, a dimensionality-reduced object data corresponds to at least two dimensionality-reduced data;

[0024] The detection module is used to, for any dimensionality-reduced object data, merge at least two types of dimensionality-reduced data corresponding to the any dimensionality-reduced object data to obtain merged data corresponding to the any dimensionality-reduced object data; and perform outlier detection processing on the merged data corresponding to each dimensionality-reduced object data to obtain outlier detection results for each object data.

[0025] In one possible implementation, the detection module is used to perform outlier detection processing on each object data in any one of the plurality of first clusters to obtain outlier detection results for each object data in the first cluster.

[0026] The determining module is used to determine multiple third clusters based on the outlier detection results of the multiple second clusters and the object data in each of the multiple first clusters.

[0027] In one possible implementation, the determining module is configured to: determine at least two clustering results based on the outlier detection results of the plurality of first clusters, the plurality of second clusters, and the object data, wherein one clustering result includes a plurality of fourth clusters and each fourth cluster includes at least one object data; determine an evaluation index for each clustering result, wherein the evaluation index is used to characterize the accuracy of the clustering result; and determine the plurality of third clusters based on the clustering result corresponding to the largest evaluation index among the evaluation indices of the clustering results.

[0028] In one possible implementation, the object data is a gene expression matrix of a cell;

[0029] The acquisition module is used to cluster the gene expression matrices of multiple cells using the Leiden algorithm to obtain multiple first clusters; and to cluster the gene expression matrices of the multiple cells using a current-level algorithm based on deep learning feature expression to obtain multiple second clusters.

[0030] In one possible implementation, the acquisition module is further configured to acquire multiple fifth clusters obtained by clustering the multiple object data based on the third clustering algorithm, wherein each fifth cluster includes at least one object data.

[0031] The determining module is further configured to determine multiple sixth clusters based on the outlier detection results of the multiple third clusters, the multiple fifth clusters, and the object data, wherein each sixth cluster includes at least one object data.

[0032] On the other hand, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores at least one piece of program code, which is loaded and executed by the processor to enable the electronic device to implement any of the clustering processing methods described above.

[0033] On the other hand, a computer-readable storage medium is also provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the at least one piece of program code being loaded and executed by a processor to enable a computer to implement any of the clustering processing methods described above.

[0034] On the other hand, a computer program or computer program product is also provided, wherein the computer program or computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement any of the above-mentioned clustering processing methods.

[0035] The technical solution provided in this application has at least the following beneficial effects:

[0036] The technical solution provided in this application embodiment is to determine multiple third clusters based on multiple first clusters, multiple second clusters, and outlier detection results of each object data. The first and second clusters are obtained by clustering multiple object data using different clustering algorithms, thereby realizing the use of two clustering algorithms to cluster multiple object data, which can improve the accuracy of clustering results. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a schematic diagram of the implementation environment of a clustering processing method provided in an embodiment of this application;

[0039] Figure 2 This is a flowchart of a clustering processing method provided in an embodiment of this application;

[0040] Figure 3 This is a clustering diagram of the Leiden algorithm provided in an embodiment of this application;

[0041] Figure 4 This is a clustering diagram of a Deep Embedding For Single-cell Clustering (DESC) algorithm provided in an embodiment of this application;

[0042] Figure 5 This is a flowchart of a cell data clustering processing method provided in an embodiment of this application;

[0043] Figure 6 This is a schematic diagram of the clustering results of single-cell data corresponding to the Leiden algorithm provided in an embodiment of this application;

[0044] Figure 7 This is a schematic diagram of the clustering results of single-cell data corresponding to the DESC algorithm provided in an embodiment of this application;

[0045] Figure 8 This is a schematic diagram of the clustering result of single-cell data after clustering processing, provided in an embodiment of this application;

[0046] Figure 9 This is a schematic diagram of the actual clustering result of single-cell data provided in an embodiment of this application;

[0047] Figure 10 This is a schematic diagram of the structure of a clustering processing device provided in an embodiment of this application;

[0048] Figure 11 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0049] Figure 12 This is a schematic diagram of the structure of a server provided in an embodiment of this application;

[0050] Figure 13 This is a clustering result diagram of single-cell data corresponding to the Leiden algorithm provided in an embodiment of this application;

[0051] Figure 14 This is a clustering result diagram of single-cell data corresponding to the DESC algorithm provided in an embodiment of this application;

[0052] Figure 15 This is a clustering result diagram of single-cell data after clustering processing, provided in an embodiment of this application.

[0053] Figure 16 This is a diagram showing the actual clustering result of single-cell data provided in an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0055] Figure 1 This is a schematic diagram illustrating the implementation environment of a clustering processing method provided in an embodiment of this application, such as... Figure 1 The implementation environment shown includes an electronic device 11, and the clustering processing method in this embodiment can be executed by the electronic device 11. Exemplarily, the electronic device 11 may include at least one of a terminal device or a server.

[0056] The terminal device can be at least one of a smartphone, game console, desktop computer, tablet computer, e-book reader, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, and laptop computer.

[0057] The server can be a single server, a server cluster consisting of multiple servers, or any of the following: a cloud computing platform or a virtualization center. This application embodiment does not limit this. The server can communicate with terminal devices via a wired or wireless network. The server can have functions such as data processing, data storage, and data transmission and reception, which are not limited in this application embodiment.

[0058] The clustering method provided in this application is based on artificial intelligence (AI) technology. AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions.

[0059] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0060] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, and intelligent transportation. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0061] Based on the above implementation environment, this application provides a clustering processing method to... Figure 2 The flowchart shown in this embodiment of the application illustrates a clustering processing method. This method can be implemented by... Figure 1 The electronic device 11 in the middle performs the operation. For example... Figure 2 As shown, the method includes steps 201-203.

[0062] Step 201: Obtain multiple first clusters and multiple second clusters. The multiple first clusters are obtained by clustering multiple object data based on the first clustering algorithm. Each first cluster includes at least one object data. The multiple second clusters are obtained by clustering multiple object data based on the second clustering algorithm. Each second cluster includes at least one object data.

[0063] In this embodiment, the first clustering algorithm and the second clustering algorithm are different clustering algorithms. For example, the first clustering algorithm and the second clustering algorithm are any two algorithms selected from Leiden algorithm, Louvain algorithm, SOTA algorithm based on deep learning feature representation, K-means clustering algorithm, DESC algorithm, etc. The object data includes, but is not limited to, cell data, point cloud data, video data, etc.

[0064] In this application embodiment, the number of first clusters obtained by clustering multiple object data based on the first clustering algorithm is not limited, nor is the number of second clusters obtained by clustering multiple object data based on the second clustering algorithm. The number of second clusters may be the same as or different from the number of first clusters. For example, the number of first clusters may be 51, and the number of second clusters may be 51 or 33.

[0065] In one possible implementation, the object data is the gene expression matrix of cells; obtaining multiple first clusters and multiple second clusters includes: using the Leiden algorithm to cluster the gene expression matrices of multiple cells to obtain multiple first clusters; and using a current-level algorithm based on deep learning feature representation to cluster the gene expression matrices of multiple cells to obtain multiple second clusters.

[0066] In this embodiment, the object is a cell, and the cell data is the cell's gene expression matrix. Each row of the gene expression matrix represents the expression of a gene under different environmental conditions or at different time points. Each column represents the expression of all genes under different conditions or samples (such as cells, tissues, experimental conditions, treatment factors, etc.). The data in each cell represents the expression level of a specific gene in a specific cell.

[0067] When clustering gene expression matrices of multiple cells, the Leiden algorithm is used to cluster the cell data. The Leiden algorithm obtains the final clustering result (i.e., multiple first clusters) through multiple stages of clustering. Each stage of clustering includes moving cell processing and refining processing. Moving cell processing is the initial clustering process for multiple cell data, resulting in multiple original clusters for that stage, each containing at least one cell data point. Refining processing is the process of further clustering the original clusters for that stage, resulting in multiple updated clusters for that stage, each containing at least one cell data point.

[0068] Please see Figure 3 , Figure 3 This is a clustering diagram of the Leiden algorithm provided in an embodiment of this application, wherein, Figure 3 Only two stages of clustering processing are shown, namely the first stage and the second stage. First, a gene expression matrix is ​​used to represent each cell to be clustered. Then, based on the gene expression matrices of each pair of cells, the weights between the data points of each pair of cells are determined, resulting in... Figure 3 The value 'a' shown here includes multiple cell data points and the weights between every two cell data points.

[0069] For the first stage, cell-shifting processing is performed on cell 'a' to perform preliminary clustering of multiple cell data, resulting in multiple original clusters corresponding to the first stage, such as... Figure 3 As shown in b, cell data of the same color in b represent cell data from the same original cluster, while cell data of different colors represent cell data from different original clusters. Next, the multiple original clusters (i.e., b) corresponding to the first stage are refined to re-cluster the cell data within the same original cluster, resulting in multiple updated clusters, such as... Figure 3 As shown in c, one of the original clusters in c contains cell data of different colors, indicating that during the refinement process, the cell data in that original cluster was clustered into at least two clusters. Then, a second stage of clustering processing is performed based on the clustering results shown in c.

[0070] Regarding the second phase, such as Figure 3 As shown in d, d includes five clusters, which are the updated clusters corresponding to the first stage, i.e., the clustering result shown in c. The lines between clusters represent the weights between them. Moving cell processing is applied to d to perform preliminary clustering of multiple cell data, obtaining multiple original clusters corresponding to the second stage, such as... Figure 3 As shown in the diagram e, clusters of the same color in e represent the same original cluster, while clusters of different colors represent different original clusters. Next, the multiple original clusters (i.e., e) corresponding to the second stage are refined to re-cluster the cell data within the same original cluster, resulting in multiple updated clusters, such as... Figure 3 As shown in f, one of the original clusters in f contains at least two clusters of the same color, indicating that the original cluster was not clustered with other original clusters during the refining process. Then, the next stage (i.e., the third stage) of clustering is performed based on the clustering results shown in f, or the clustering results shown in f are used as the final clustering results, resulting in multiple first clusters.

[0071] As can be seen from the above description of the Leiden algorithm, in each stage of clustering, the Leiden algorithm first performs a cell-moving process to initially cluster all cell data, and then performs a refinement process to cluster all cell data again. Therefore, the Leiden algorithm can achieve global clustering of cell data and is a global clustering algorithm.

[0072] When clustering multiple cell datasets, the DESC algorithm is used. DESC is a state-of-the-art (SOTA) algorithm based on deep learning feature representation. The DESC algorithm inputs the gene expression matrix of each cell into an autoencoder, which encodes the gene expression matrix of each cell to obtain the gene expression features of each cell. Based on these gene expression features, the clustering result of each cell dataset is determined. The clustering result of a cell dataset is the cluster to which that cell belongs. Based on the gene expression features of each cell, the probability of each cell belonging to its cluster and the batch number of each cell dataset can also be determined. Specifically, the DESC algorithm determines the cluster to which the cell belongs by determining the probability of it belonging to each cluster; this probability is denoted as the maximum probability, and the batch number of the cell dataset is the sampling batch number.

[0073] like Figure 4 As shown, Figure 4 This is a schematic diagram of a clustering algorithm for the DESC algorithm provided in an embodiment of this application. First, the gene expression matrix of each cell is input to an autoencoder, which encodes the gene expression matrix of each cell to obtain the gene expression features of each cell. Then, clustering is performed based on the gene expression features of each cell to obtain the clustering results. When implementing the DESC algorithm based on a model, the model parameters are optimized based on the clustering results and a loss function to update the model. Using the updated model and the gene expression features of each cell, clustering is performed on the cell data to obtain the clustering results. The loss function is not limited.

[0074] Optionally, the loss function is Figure 4The expression Loss = KL(P||Q) is shown, where Loss is the loss function, KL represents the KL divergence (Kullback-Leibler Divergence), also called relative entropy, P is the true probability distribution (i.e., the probability that cell data belongs to each cluster), and Q is the pseudo-distribution of P.

[0075] It should be noted that the output information includes, but is not limited to, clustering results. For example... Figure 4 As shown, the output information consists of three parts. The first part is the clusters (i.e., the clustering results). Clusters numbered 0-5 represent 6 clusters (i.e., the second cluster), and each cluster contains at least one cell data point. In other words, the clusters are the clustering results of all cell data points. The second part is the probability of each cell data point. By determining the probability of a cell data point belonging to each cluster, the maximum probability of that cell data point is determined, and thus the probability of a cell data point within a cluster is determined as this maximum probability. The third part is the batch of cell data. Each cluster contains at least one batch of cell data points.

[0076] As can be seen from the above description of the DESC algorithm, the DESC algorithm determines the clustering result of the gene expression matrix of a cell (i.e., cell data) based on the gene expression matrix of a cell. In this way, the clustering result of each cell data is determined, which has a strong clustering ability for individual samples and is a local clustering processing algorithm.

[0077] Step 202: Perform outlier detection processing on multiple object data to obtain the outlier detection results for each object data.

[0078] In this embodiment, an outlier detection (OD) algorithm is used to perform outlier detection processing on multiple object data sets to obtain outlier detection results for each object data set. The outlier detection result for any object data set is either an outlier object data set or a non-outlier object data set. This embodiment does not limit the outlier detection algorithm.

[0079] For example, outlier detection algorithms are based on Copula-based outlier detection (COPOD) algorithms. Copula is a statistical probability function used to model multidimensional cumulative distributions and also to model dependencies between multiple random variables (RVs).

[0080] In this embodiment, PyOD (a library for detecting outliers in data) of Python (a computer programming language) provides various outlier detection algorithms, including the COPOD algorithm. The COPOD algorithm has three advantages. The first advantage is that it does not require distance calculation between samples, resulting in fast execution. The second advantage is that it does not require parameter tuning and can be directly called. The third advantage is that its outlier detection performance is significantly better than other outlier detection algorithms.

[0081] Because the COPOD algorithm has the above three advantages, the embodiments of this application adopt the COPOD algorithm, which can conveniently and quickly perform outlier detection processing on multiple object data, and the outlier detection results of each object data are highly accurate.

[0082] In this embodiment of the application, when performing outlier detection processing on multiple object data, outlier detection processing can be performed on multiple original object data. For example, if the original object data is the gene expression matrix of cells, then outlier detection processing can be performed on the gene expression matrices of multiple cells.

[0083] Because the original object data has a high dimensionality, outlier detection is slow. Therefore, the original object data can be processed to improve the speed of outlier detection. Optionally, outlier detection processing can be performed on multiple object datasets to obtain outlier detection results for each object dataset. This includes: performing dimensionality reduction processing on each object dataset to obtain dimensionality-reduced object datasets; and performing outlier detection processing on each dimensionality-reduced object dataset to obtain outlier detection results for each object dataset.

[0084] For any given object data, dimensionality reduction is performed to decrease its size and volume, resulting in dimensionality-reduced object data. This process is repeated to obtain various dimensionality-reduced object data sets. Then, an outlier detection algorithm is used to detect outliers in each of these sets, yielding the outlier detection results. The outlier detection result for any given dimensionality-reduced object data set is the outlier detection result for the corresponding object data set.

[0085] It is understood that there are multiple ways to perform dimensionality reduction on any object data, and the embodiments of this application do not limit the dimensionality reduction method of object data.

[0086] Optionally, Principal Component Analysis (PCA) can be used to reduce the dimensionality of any object data to obtain dimensionality-reduced object data. For example, PCA can be used to reduce the dimensionality of the gene expression matrix of cells (which typically includes 20,000 to 30,000 data points) to 50 data points.

[0087] Optionally, dimensionality reduction can be performed on any object data using neural network technology to obtain dimensionality-reduced object data. For example, dimensionality reduction can be performed on the gene expression matrix of cells using a feature extraction network, reducing the amount of data from the original 20,000 to 30,000 data points to 32 data points.

[0088] Since there are multiple ways to perform dimensionality reduction on any object data, at least two dimensionality reduction methods can be used to reduce the dimensionality of any object data, resulting in dimensionality-reduced object data. In this case, one dimensionality-reduced object data corresponds to at least two dimensionality-reduced data types. Outlier detection processing is performed on each dimensionality-reduced object data to obtain outlier detection results for each object data type. This includes: for any dimensionality-reduced object data, merging the at least two dimensionality-reduced data types corresponding to that object data to obtain merged data corresponding to that object data; and performing outlier detection processing on the merged data corresponding to each dimensionality-reduced object data to obtain outlier detection results for each object data type.

[0089] Since at least two dimensionality reduction methods are used to reduce the dimensionality of any object data, the dimensionality reduction object data corresponding to the object data includes the dimensionality reduction object data corresponding to each dimensionality reduction method. The dimensionality reduction object data corresponding to any one dimensionality reduction method is a type of dimensionality reduction data.

[0090] For any given dimensionality reduction object data, at least two corresponding dimensionality reduction data are sequentially concatenated, or operations such as addition, subtraction, multiplication, and division are performed on the at least two corresponding dimensionality reduction data to merge the corresponding dimensionality reduction object data, resulting in merged data for that dimensionality reduction object data. In this way, the merged data for each dimensionality reduction object data is obtained. Then, based on an outlier detection algorithm, outlier detection processing is performed on the merged data corresponding to each dimensionality reduction object data to obtain the outlier detection results for each object data.

[0091] First, merge at least two types of dimensionality-reduced data corresponding to each dimensionality-reduced object data. Then, perform outlier detection processing on the merged data corresponding to each dimensionality-reduced object data. This reduces the amount of computation, increases the speed of outlier detection, and improves the accuracy of outlier detection.

[0092] It should be noted that when performing outlier detection processing on multiple object data based on outlier detection algorithms, in addition to the aforementioned outlier detection processing on multiple original object data, multiple dimensionality-reduced object data, or the merged data corresponding to multiple dimensionality-reduced object data, other information of each object data can also be combined for outlier detection processing. For example, based on the outlier detection algorithm, the probability of each object data belonging to the first cluster, and the probability of each object data belonging to the second cluster, outlier detection processing can be performed on multiple original object data to improve the accuracy of outlier detection. In application, the settings can be flexibly configured according to the application scenario and human experience, and are not limited in the embodiments of this application.

[0093] Step 203: Based on the outlier detection results of multiple first clusters, multiple second clusters, and each object data, determine multiple third clusters, each of which includes at least one object data.

[0094] In this embodiment, multiple third clusters are determined based on the outlier detection results of multiple first clusters, multiple second clusters, and each object data. Since the first cluster is obtained based on the first clustering algorithm and the second cluster is obtained based on the second clustering algorithm, this embodiment can determine multiple third clusters based on the first clustering algorithm and the second clustering algorithm, thereby improving the accuracy of the clustering results.

[0095] Taking the Leiden algorithm as the first clustering algorithm and the DESC algorithm as the second clustering algorithm as examples, the following explanation is provided. The Leiden algorithm is a global clustering algorithm capable of globally clustering cell data, but it can lead to inaccurate clustering results for local cell data. The DESC algorithm, on the other hand, is a local clustering algorithm with strong clustering capabilities for individual samples, but it does not consider the relationships between cell data, resulting in poor accuracy in the clustering results. The clustering method described in this application can determine multiple third clusters based on the Leiden and DESC algorithms, combining the advantages of the Leiden algorithm for global clustering of cell data with the strong clustering capability of the DESC algorithm for individual samples. This achieves complementarity between the Leiden and DESC algorithms, improving the accuracy of the clustering results.

[0096] In this embodiment, multiple third clusters can be determined based on outlier detection results from multiple first clusters, multiple second clusters, and individual object data, denoted as implementation method A1. Alternatively, multiple third clusters can also be determined based on outlier detection results from multiple first clusters, multiple second clusters, and individual object data, denoted as implementation method A2. Implementation methods A1 and A2 will be described in detail below.

[0097] Implementation method A1 involves performing outlier detection processing on multiple object data sets to obtain outlier detection results for each object data set. This includes: for any one of the multiple second clusters, performing outlier detection processing on each object data set within that second cluster to obtain outlier detection results for each object data set within that second cluster; and determining multiple third clusters based on the outlier detection results of the multiple first clusters, the multiple second clusters, and each object data set. This includes: determining multiple third clusters based on the outlier detection results of each object data set within the multiple first clusters and the multiple second clusters.

[0098] In this embodiment, for any second cluster, outlier detection is performed on the object data within that second cluster using an outlier detection algorithm, resulting in outlier detection results for each object data item in that second cluster. This method determines the outlier detection results for each object data item in each second cluster. Then, based on the outlier detection results of multiple first clusters and each object data item in each second cluster, multiple third clusters are determined.

[0099] Optionally, based on the outlier detection results of multiple first clusters and object data in each of the second clusters, multiple third clusters are determined, including: for any second cluster among the multiple second clusters, in response to the existence of first object data in each object data of any second cluster, determining the first cluster corresponding to any second cluster from the multiple first clusters, wherein the outlier detection result of the first object data is non-outlier object data; determining the first cluster corresponding to any second cluster as a third cluster; and in response to the first object data not belonging to a third cluster, adding the first object data to a third cluster.

[0100] In this embodiment of the application, for any object data in any second cluster, when the outlier detection result of the object data is non-outlier object data, the object data is first object data. At this time, the first cluster corresponding to the second cluster is determined from multiple first clusters as the main cluster corresponding to the second cluster, and the main cluster corresponding to the second cluster is determined as a third cluster.

[0101] It should be noted that for any second cluster, the first cluster to which any first object data in that second cluster belongs may be the same as or different from the primary cluster corresponding to that second cluster. When the first cluster to which any first object data in that second cluster belongs is the same as the primary cluster corresponding to that second cluster, it indicates that the first object data belongs to the primary cluster corresponding to that second cluster, that is, the first object data belongs to a third cluster. When the first cluster to which any first object data in that second cluster belongs is different from the primary cluster corresponding to that second cluster, adding the first object data to the primary cluster corresponding to that second cluster is equivalent to adding the first object data to a third cluster.

[0102] For example, multiple first clusters are designated as clusters 0 to 11, and multiple second clusters are designated as clusters 0 to 9. Object data S15 and S20 both correspond to cluster 2, and object data S15 corresponds to cluster 7, while object data S20 corresponds to cluster 5. The outlier detection results for both object data S15 and object data S20 are non-outlier object data. In this case, both object data S15 and object data S20 are first object data. From clusters 0 to 11, the primary cluster corresponding to cluster 2 (i.e., the first cluster) is determined to be cluster 5. Therefore, cluster 5 is considered a third cluster (denoted as cluster 5). For object data S15, since it corresponds to cluster 7, it is equivalent to cluster 7 being different from cluster 5 (i.e., first cluster 5). Therefore, object data S15 is added to cluster 5. For object data S20, since object data S20 corresponds to the first cluster 5, it is equivalent to the first cluster 5 to which object data S20 belongs being the same as the third cluster 5. Therefore, object data S20 already belongs to the third cluster 5.

[0103] As can be seen from the above, for any second cluster, the method of the embodiments of this application can make each first object data in the second cluster belong to the main cluster corresponding to the second cluster, thereby obtaining a third cluster. That is to say, each first object data in a second cluster belongs to the same third cluster.

[0104] Optionally, determining the first cluster corresponding to any second cluster from multiple first clusters includes: determining a cross-tabulation based on multiple first clusters and multiple second clusters, where a row of data in the cross-tabulation represents the data of each object in a first cluster, and a column of data in the cross-tabulation represents the data of each object in a second cluster; and determining the first cluster corresponding to any second cluster from multiple first clusters based on the cross-tabulation.

[0105] A crosstab is a categorized summary table that includes several rows and several columns. In this embodiment, a crosstab is determined based on multiple first clusters and multiple second clusters, where each row of the crosstab corresponds to a specific first cluster, and each column corresponds to a specific second cluster.

[0106] Optionally, a first cluster corresponds to a first cluster identifier, a second cluster corresponds to a second cluster identifier, and an object data includes an object identifier; determining a cross-tabulation based on multiple first clusters and multiple second clusters includes: determining the object identifier corresponding to each cluster identifier set based on the first cluster identifier of each first cluster, the object identifier of each object data contained in each first cluster, the second cluster identifier of each second cluster, and the object identifier of each object data contained in each second cluster, wherein a cluster identifier set includes a first cluster identifier and a second cluster identifier; and determining the cross-tabulation based on the number of object identifiers corresponding to each cluster identifier set.

[0107] Both the first cluster identifier and the second cluster identifier include, but are not limited to, at least one of numbers, characters, symbols, etc. The object identifier of the object data includes, but is not limited to, at least one of numbers, characters, symbols, etc. The first cluster identifier, the second cluster identifier, and the object identifier of the object data can be the same, different, or any two of the three can be the same, without any limitation.

[0108] For example, the first cluster identifier includes the numbers 0 to 11, the second cluster identifier includes the numbers 0 to 9, and the object identifier of the object data consists of the character S and numbers, including S1 to SN, where N is a positive integer greater than 1.

[0109] In this embodiment of the application, the object identifier of an object data is unique and used to identify the object data. As can be seen from step 201, multiple object data are clustered into multiple first clusters, and these multiple object data are clustered into multiple second clusters. Since each first cluster corresponds to a first cluster identifier and each second cluster corresponds to a second cluster identifier, the correspondence between each first cluster identifier and the object identifier of each object data, as well as the correspondence between each second cluster identifier and the object identifier of each object data, can be determined.

[0110] As shown in Tables 1 and 2 below, Table 1 is a correspondence table between a first cluster identifier and an object identifier of object data provided in an embodiment of this application, and Table 2 is a correspondence table between a second cluster identifier and an object identifier of object data provided in an embodiment of this application.

[0111] Table 1

[0112]

[0113] Table 2

[0114]

[0115] As shown in Tables 1 and 2, the object identifiers of the N object data are S1, S2, S3, S4...SN, where N is a positive integer. When clustering the N object data into multiple first clusters, object identifier S1 corresponds to cluster identifier 1, object identifier S2 corresponds to cluster identifier 0, object identifier S3 corresponds to cluster identifier 11, object identifier S4 corresponds to cluster identifier 0, and object identifier SN corresponds to cluster identifier 3. Similarly, when clustering the N object data into multiple second clusters, object identifier S1 corresponds to cluster identifier 9, object identifier S2 corresponds to cluster identifier 0, object identifier S3 corresponds to cluster identifier 1, object identifier S4 corresponds to cluster identifier 7, and object identifier SN corresponds to cluster identifier 5.

[0116] In this embodiment, based on the correspondence between each first cluster identifier and the object identifier of each object data, and the correspondence between each second cluster identifier and the object identifier of each object data, the object identifiers corresponding to each set of cluster identifiers for each object data are determined. The object identifiers of each object data corresponding to a set of cluster identifiers are the object identifiers of the same object data corresponding to the first and second cluster identifiers in that set; that is, the same object identifiers corresponding to the first and second cluster identifiers in that set. Then, based on the number of object identifiers corresponding to each object data for each set of cluster identifiers, a cross-tabulation is determined.

[0117] For example, based on Tables 1 and 2, it can be determined that there are 8389 object identifiers corresponding to the cluster identifier set {1, 1}, 4034 object identifiers corresponding to the cluster identifier set {1, 2}, and 1 object identifier corresponding to the cluster identifier set {1, 3}, etc., where 'a' in the cluster identifier set {a, b} is the first cluster identifier and 'b' is the second cluster identifier. The cross-tabulation table shown in Table 3 can be determined by the number of object identifiers corresponding to each cluster identifier set. Table 3 is a schematic diagram of a cross-tabulation table provided in an embodiment of this application.

[0118] Table 3

[0119]

[0120] In the cross-tabulation table shown in Table 3, each row corresponds to a first cluster, with cluster identifiers ranging from 0 to 11, totaling 12 first clusters. Each column corresponds to a second cluster, with cluster identifiers ranging from 0 to 9, totaling ten second clusters. Table 3 shows that for the object data with second cluster identifier 2 obtained after clustering multiple object data using the second clustering algorithm, when clustering multiple object data using the first clustering algorithm, the object data with second cluster identifier 2 is scattered across first cluster identifiers 1, 4, and 5.

[0121] In this embodiment, a first cluster corresponding to each second cluster is determined from multiple first clusters based on a cross-tabulation. The cross-tabulation includes multiple non-zero data points, where each non-zero data point represents the number of identical object data points contained in the first cluster corresponding to the row containing the non-zero data point and the second cluster corresponding to the column containing the non-zero data point. Determining the first cluster corresponding to any second cluster from multiple first clusters based on the cross-tabulation includes: determining the largest non-zero data point from among the non-zero data points contained in the column corresponding to any second cluster in the cross-tabulation; and determining the first cluster corresponding to the row containing the largest non-zero data point as the first cluster corresponding to any second cluster.

[0122] Any data in the cross-tabulation is either a target character or a non-zero data. In the embodiments of this application and the following embodiments, the target character is 0, an empty character (i.e., empty), or a special character (such as the symbol &, the string Nnnn, etc.). Any non-zero data in the cross-tabulation represents the number of identical object data contained in the first cluster corresponding to the row where the data is located and the second cluster corresponding to the column where the data is located.

[0123] For any second cluster, which corresponds to a column of data in the crosstab, the largest non-zero data is determined from the non-zero data contained in the column corresponding to the second cluster in the crosstab, and the first cluster corresponding to the row containing the largest non-zero data is determined as the first cluster corresponding to the second cluster.

[0124] In this embodiment of the application, the second cluster can be used as the Assistant Cluster Result (ACR), and the first cluster can be used as the Main Cluster Result (MCR). The main cluster result and the assistant cluster result satisfy the formula shown below.

[0125]

[0126] Where i is the column number, i.e., a second cluster. Assuming the second cluster identifiers are 0 to M (M is a positive integer), the number of second clusters is M+1. The value of i can be any number from 0 to M, such as 0 to 9 in Table 3. CM(i) is the first cluster corresponding to the row containing the largest non-zero data in column i, i.e., the mapping cluster (ClusterMapping, CM) corresponding to column i, which is the first cluster. INDEX is a function that returns the values ​​in the table. ACR(i) are the non-zero data in column i, and MAX(ACR(i)) is the largest non-zero data in column i.

[0127] For example, as shown in Table 4 below, Table 4 is the fourth column of Table 3 (i.e. the column identified as 2 in the second cluster).

[0128]

[0129] Since the largest non-zero data in the column corresponding to the second cluster identifier 2 is 4034, and the row containing 4034 corresponds to the first cluster identifier 1, CM(2) = 1, that is, the row containing the largest non-zero data in the second column corresponds to the first cluster identifier 1 of the first cluster. In other words, the second cluster identifier 2 (corresponding to a second cluster) corresponds to the first cluster identifier 1 (corresponding to a first cluster).

[0130] Based on the above principles, for Table 3, we can determine that: CM(0)=0, CM(1)=1, CM(2)=1, CM(3)=7, CM(4)=5, CM(5)=2, CM(6)=11, CM(7)=3, CM(8)=9, CM(9)=10.

[0131] After determining the first cluster corresponding to any second cluster, the first cluster corresponding to any second cluster is designated as a third cluster, and each first object data in any second cluster belongs to this third cluster. For example, object data S15 corresponds to first cluster identifier 7 and second cluster identifier 2, and object data S15 is first object data. Since second cluster identifier 2 corresponds to first cluster identifier 1, the first cluster corresponding to first cluster identifier 1 is determined as a third cluster, and object data S15 is added to this third cluster, making object data S15 belong to this third cluster. In this way, the first cluster corresponding to each second cluster is determined as a third cluster, and each first object data in each second cluster belongs to the third cluster corresponding to each second cluster.

[0132] In this embodiment of the application, based on the outlier detection results of the object data in the multiple first clusters and each of the second clusters, multiple third clusters are determined, including: for any second cluster among the multiple second clusters, in response to the existence of second object data in each object data in the multiple second clusters, the first cluster to which a second object data belongs is determined as a third cluster, and the outlier detection result of the second object data is the outlier object data.

[0133] In this embodiment, for any object data in any second cluster, if the outlier detection result of the object data is outlier object data, then the object data is considered second object data. At this time, the first cluster to which the second object data belongs is determined as a third cluster, thereby establishing the correspondence between the second object data and the third cluster. In this way, the correspondence between each second object data in the second cluster and the third cluster is determined, thus establishing the correspondence between each second object data in each second cluster and the third cluster.

[0134] For example, multiple first clusters are designated as clusters 0 to 11, and multiple second clusters are designated as clusters 0 to 9. Object data S19 and S200 both correspond to second cluster 2, and object data S19 corresponds to first cluster 7, while object data S200 corresponds to first cluster 5. The outlier detection results for both object data S19 and object data S200 are outlier object data. In this case, object data S19 and object data S200 are both second object data. First cluster 7 is determined to be a third cluster (denoted as third cluster 7) to establish the correspondence between object data S19 and third cluster 7. Similarly, first cluster 5 is determined to be a third cluster (denoted as third cluster 5) to establish the correspondence between object data S200 and third cluster 5. In other words, first cluster 7 and first cluster 5 are two third clusters.

[0135] Implementation method A2 involves performing outlier detection processing on multiple object data to obtain outlier detection results for each object data, including: for any one of the multiple first clusters, performing outlier detection processing on each object data in that first cluster to obtain outlier detection results for each object data in that first cluster; and determining multiple third clusters based on the outlier detection results of the multiple first clusters, multiple second clusters, and each object data, including: determining multiple third clusters based on the outlier detection results of each object data in the multiple second clusters and each first cluster.

[0136] In this embodiment, for any first cluster, outlier detection is performed on the object data within that first cluster using an outlier detection algorithm, resulting in outlier detection results for each object data within that first cluster. This method determines the outlier detection results for each object data within each first cluster. Then, based on the outlier detection results of multiple second clusters and the object data within each first cluster, multiple third clusters are determined.

[0137] Optionally, based on the outlier detection results of multiple second clusters and each object data in each first cluster, multiple third clusters are determined, including: for any first cluster among the multiple first clusters, in response to the existence of third object data in each object data in any first cluster, determining the second cluster corresponding to any first cluster from the multiple second clusters, wherein the outlier detection result of the third object data is non-outlier object data; determining the second cluster corresponding to any first cluster as a third cluster; in response to the third object data not belonging to a third cluster, adding the third object data to a third cluster.

[0138] Optionally, based on the outlier detection results of multiple second clusters and each object data in each first cluster, multiple third clusters are determined, including: for any one of the multiple first clusters, in response to the existence of fourth object data in each object data in any one first cluster, determining the second cluster to which a fourth object data belongs as a third cluster, and the outlier detection result of the fourth object data as outlier object data.

[0139] For an explanation of implementation method A2, please refer to the explanation of implementation method A1 above. The principles of implementation method A1 and implementation method A2 are similar, and will not be repeated here.

[0140] In one possible implementation, multiple third clusters are determined based on outlier detection results of multiple first clusters, multiple second clusters, and each object data. This includes: determining at least two clustering results based on outlier detection results of multiple first clusters, multiple second clusters, and each object data, where each clustering result includes multiple fourth clusters, and each fourth cluster includes at least one object data; determining evaluation metrics for each clustering result, where the evaluation metrics characterize the accuracy of the clustering results; and determining multiple third clusters based on the clustering result corresponding to the largest evaluation metric among the evaluation metrics of each clustering result.

[0141] In this embodiment, based on outlier detection results from multiple first clusters, multiple second clusters, and each object's data, at least two clustering results can be determined using at least two methods. For each clustering result, an evaluation index is calculated. The calculation method for the evaluation index is not limited here; for example, the evaluation index is the Adjusted Rand Index (ARI) or Normalized Mutual Information (NMI). The ARI value is [-1, 1], with a higher value indicating higher accuracy of the clustering result. The NMI value is [0, 1], with a higher value indicating higher accuracy of the clustering result.

[0142] After calculating the evaluation index for each clustering result, the maximum evaluation index is determined from all the evaluation indexes, and the clustering result corresponding to the maximum evaluation index is selected. Each fourth cluster in the clustering result corresponding to the maximum evaluation index is a multiple third cluster.

[0143] In one possible implementation, after determining multiple third clusters based on outlier detection results of multiple first clusters, multiple second clusters, and each object data, the method further includes: obtaining multiple fifth clusters obtained by clustering multiple object data using a third clustering algorithm, wherein each fifth cluster includes at least one object data; and determining multiple sixth clusters based on outlier detection results of multiple third clusters, multiple fifth clusters, and each object data, wherein each sixth cluster includes at least one object data.

[0144] The third clustering algorithm is not limited in this embodiment. For example, the third clustering algorithm includes, but is not limited to, the Leiden algorithm, the Louvain algorithm, the SOTA algorithm based on deep learning feature representation, the K-means clustering algorithm, and the DESC algorithm. The first, second, and third clustering algorithms are three different clustering algorithms.

[0145] This application does not limit the number of fifth clusters obtained based on the third clustering algorithm. The number of fifth clusters may be the same as or different from the number of first clusters, the number of fifth clusters may be the same as or different from the number of second clusters, the number of fifth clusters may be the same as or different from the number of third clusters, or the number of fifth clusters may be the same as or different from the number of fourth clusters.

[0146] In this embodiment, multiple sixth clusters are determined based on the outlier detection results of multiple third clusters, multiple fifth clusters, and each object data. For details, please refer to the relevant descriptions of steps 202 and 203. The implementation principles of the two are similar and will not be repeated here.

[0147] Understandably, after clustering multiple object data using three or more clustering algorithms, this embodiment first determines multiple third clusters based on the clustering results of two clustering algorithms (i.e., multiple first clusters and multiple second clusters), following steps 201-203. Then, based on the multiple third clusters and the clustering results of a third clustering algorithm (i.e., multiple fifth clusters), multiple sixth clusters are determined. Subsequently, based on the multiple sixth clusters and the clustering results of a fourth clustering algorithm, multiple seventh clusters are determined, and so on, to achieve the goal of using three or more clustering algorithms to determine the clustering results of multiple object data, thereby improving the accuracy of the clustering results.

[0148] The above method determines multiple third clusters based on the outlier detection results of multiple first clusters, multiple second clusters, and each object data. The first and second clusters are obtained by clustering multiple object data using different clustering algorithms, thus realizing the use of two clustering algorithms to cluster multiple object data and improving the accuracy of clustering results.

[0149] The clustering processing method of this application embodiment has been described in detail above from the perspective of method and steps. The following is a detailed explanation in conjunction with a specific scenario. The scenario of this application embodiment involves clustering multiple cell data based on two different clustering algorithms, where each cell data point is a single cell. The number of cells in this scenario is 24,679, and the source of these 246.79 million cell data points is not limited here.

[0150] like Figure 5 As shown, Figure 5 This is a flowchart illustrating a cell data clustering method provided in an embodiment of this application, where the cell data is the cell gene expression matrix. First, the cell gene expression matrix is ​​preprocessed. On one hand, preprocessing the cell gene expression matrix reduces its dimensionality, resulting in dimensionality-reduced object data. On the other hand, the preprocessed cell gene expression matrix is ​​then clustered. In this embodiment, the Leiden algorithm is used to cluster multiple preprocessed cell gene expression matrices, resulting in multiple first clusters. A deep embedding single-cell clustering algorithm is then used to cluster multiple preprocessed cell gene expression matrices, resulting in multiple second clusters. Next, a cross-tabulation is obtained based on the multiple first clusters and multiple second clusters. Each column of the cross-tabulation represents the cell data in a second cluster, and each row of the cross-tabulation represents the cell data in a first cluster.

[0151] It should be noted that when using the Leiden algorithm to cluster the gene expression matrices of multiple preprocessed cells, and / or when using the deep embedding single-cell clustering algorithm to cluster the gene expression matrices of multiple preprocessed cells, the gene expression matrices of the preprocessed cells may be subjected to dimensionality reduction. In this case, the gene expression matrices of the preprocessed cells after dimensionality reduction are the dimensionality-reduced object data.

[0152] In this embodiment, for any column in the cross-tabulation, which corresponds to at least one dimensionality-reduced object data, outlier detection processing is performed on each dimensionality-reduced object data corresponding to that column to obtain outlier detection results for each dimensionality-reduced object data corresponding to that column. In this way, outlier detection results are obtained for each column of the cross-tabulation corresponding to each dimensionality-reduced object data. Then, based on the outlier detection results for each column of the cross-tabulation corresponding to each dimensionality-reduced object data, multiple third clusters are determined.

[0153] related Figure 5 The explanation can be found in the relevant descriptions of steps 201 to 203. The implementation principles of the two are similar, and will not be repeated here.

[0154] This embodiment uses the Leiden algorithm to cluster 24,679 single-cell data points, obtaining the clustering results corresponding to the Leiden algorithm (i.e., the multiple first clusters mentioned above). Please refer to... Figure 6 or Figure 13 , Figure 6 This is a schematic diagram illustrating the clustering results of single-cell data corresponding to the Leiden algorithm provided in an embodiment of this application. Figure 13 This is a clustering result diagram of single-cell data corresponding to a Leiden algorithm provided in an embodiment of this application. Figure 6 It can be seen that using the Leiden algorithm to cluster 24,679 single-cell data points yields 16 first clusters. The first cluster identifiers of these 16 first clusters range from 0 to 15, and each first cluster contains at least one single-cell data point.

[0155] This application embodiment also uses the DESC algorithm to cluster 24,679 single-cell data points, obtaining the clustering results corresponding to the DESC algorithm (i.e., the multiple second clusters mentioned above). Please refer to [link to relevant documentation]. Figure 7 or Figure 14 , Figure 7 This is a schematic diagram of the clustering results of single-cell data corresponding to the DESC algorithm provided in an embodiment of this application. Figure 14 This is a clustering result diagram of single-cell data corresponding to a DESC algorithm provided in an embodiment of this application. Figure 7 It can be seen that the DESC algorithm was used to cluster 24,679 single-cell data to obtain 12 second clusters. The second clusters of these 12 second clusters are identified as 0 to 11, and each second cluster includes at least one single-cell data.

[0156] Subsequently, embodiments of this application are based on 16 first clusters (i.e. Figure 6 The clustering results shown) and 12 second clusters (i.e. Figure 7 The clustering results shown are used to obtain a cross-tabulation. The outlier detection results for each column of the cross-tabulation corresponding to the single-cell data are determined. Based on the outlier detection results for each column of the cross-tabulation corresponding to each single-cell data, multiple third clusters are determined. Please refer to [link to relevant documentation]. Figure 8 or Figure 15 , Figure 8 This is a schematic diagram illustrating the clustering result of single-cell data after clustering processing, provided in an embodiment of this application. Figure 15 This is a clustering result diagram of single-cell data after clustering processing, provided in an embodiment of this application. The clustering result of the single-cell data after clustering processing is multiple third clusters. Figure 8As can be seen, by using the clustering processing method provided in this application embodiment to cluster the clustering results of single-cell data corresponding to the Leiden algorithm and the clustering results of single-cell data corresponding to the DESC algorithm, 16 third clusters are obtained. The third cluster identifiers of these 16 third clusters are 0 to 15, and each third cluster includes at least one single-cell data.

[0157] Please see Figure 9 or Figure 16 , Figure 9 This is a schematic diagram illustrating the actual clustering results of single-cell data provided in an embodiment of this application. Figure 16 This is a diagram illustrating the actual clustering results of single-cell data provided in an embodiment of this application. Figure 9 It is evident that the actual clustering results of the 24,679 single-cell data points are 8 clusters. These 8 clusters are: the cluster corresponding to FCGR3A+ monocyte data, the cluster corresponding to CD14+ monocyte data (a type of osteoclast precursor cell), the cluster corresponding to dendritic cell data, the cluster corresponding to megakaryocyte data, the cluster corresponding to CD4 T cell data (a type of lymphocyte), the cluster corresponding to B cell data (a type of lymphocyte), the cluster corresponding to CD8 T cell data (a type of lymphocyte), and the cluster corresponding to natural killer cell (NK cell) data.

[0158] contrast Figure 6 and Figure 9 It can be seen that the 16 first clusters obtained by the Leiden algorithm after clustering single-cell data differ significantly from the actual clustering results of the single-cell data, such as CD4 T cell data. Figure 6 The data is evenly divided into four first clusters, labeled 0, 1, 8, and 9 respectively. Similarly, in comparison... Figure 7 and Figure 9 It can be seen that the 12 second clusters obtained by the DESC algorithm after clustering the single-cell data differ somewhat from the actual clustering results of the single-cell data, such as the CD4 T cell data. Figure 7 The data is evenly divided into four second clusters, labeled 0, 6, 5, and 9 respectively. In contrast... Figure 8 and Figure 9 It can be seen that, using the clustering processing method provided in the embodiments of this application, for... Figure 6 and Figure 7 The clustering results shown are obtained after processing. Figure 8 The 16 third clusters in the data are relatively close to the actual clustering results of single-cell data, such as CD4 T cell data. Figure 8Most of them are divided into a third cluster, which is identified as 0. A portion is unequally divided into three third clusters, identified as 1, 8 and 9 respectively.

[0159] This application embodiment also uses ARI and NMI to evaluate the accuracy of the clustering results of single-cell data corresponding to the Leiden algorithm (hereinafter referred to as the clustering results of the Leiden algorithm), the accuracy of the clustering results of single-cell data corresponding to the DESC algorithm (hereinafter referred to as the clustering results of the DESC algorithm), and the accuracy of the clustering results of single-cell data obtained after using the clustering processing method provided in this application embodiment to cluster the clustering results of the Leiden algorithm and the clustering results of the DESC algorithm (hereinafter referred to as the clustering results of this application embodiment), as shown in Table 5 below.

[0160]

[0161] As can be clearly seen from Table 5, the clustering results of this embodiment show improvements in both ARI and NMI values ​​compared to the Leiden algorithm, and also show significant improvements in both ARI and NMI values ​​compared to the DESC algorithm. Since higher ARI and NMI values ​​generally indicate higher clustering accuracy, the clustering results of this embodiment are significantly superior to those of the Leiden and DESC algorithms. Therefore, the clustering processing method provided in this embodiment can improve the accuracy of clustering results.

[0162] Figure 10 The diagram shown is a structural schematic of a clustering processing device provided in an embodiment of this application. Figure 10 As shown, the device includes:

[0163] The acquisition module 1001 is used to acquire multiple first clusters and multiple second clusters. The multiple first clusters are obtained by clustering multiple object data based on a first clustering algorithm. Each first cluster includes at least one object data. The multiple second clusters are obtained by clustering multiple object data based on a second clustering algorithm. Each second cluster includes at least one object data.

[0164] The detection module 1002 is used to perform outlier detection processing on multiple object data to obtain the outlier detection results for each object data.

[0165] The determination module 1003 is used to determine multiple third clusters based on the outlier detection results of multiple first clusters, multiple second clusters and each object data, wherein each third cluster includes at least one object data.

[0166] In one possible implementation, the detection module 1002 is used to perform outlier detection processing on the object data of any one of the multiple second clusters, and obtain the outlier detection results of the object data of any one of the second clusters.

[0167] The determination module 1003 is used to determine multiple third clusters based on the outlier detection results of the data of each object in multiple first clusters and each second cluster.

[0168] In one possible implementation, the determining module 1003 is configured to, for any one of a plurality of second clusters, in response to the presence of first object data in each object data of any one of the multiple second clusters, determine the first cluster corresponding to any one of the multiple second clusters, wherein the outlier detection result of the first object data is non-outlier object data; determine that the first cluster corresponding to any one of the second clusters is a third cluster; and in response to the first object data not belonging to a third cluster, add the first object data to a third cluster.

[0169] In one possible implementation, the determining module 1003 is used to determine a cross-tabulation based on multiple first clusters and multiple second clusters, wherein a row of data in the cross-tabulation represents the data of each object in a first cluster, and a column of data in the cross-tabulation represents the data of each object in a second cluster; and to determine the first cluster corresponding to any second cluster from the multiple first clusters based on the cross-tabulation.

[0170] In one possible implementation, a first cluster corresponds to a first cluster identifier, a second cluster corresponds to a second cluster identifier, and an object data includes an object identifier.

[0171] The determination module 1003 is used to determine the object identifiers corresponding to each cluster identifier set based on the first cluster identifiers of each first cluster, the object identifiers of each object data contained in each first cluster, the second cluster identifiers of each second cluster, and the object identifiers of each object data contained in each second cluster. A cluster identifier set includes one first cluster identifier and one second cluster identifier. The cross-tabulation is determined based on the number of object identifiers corresponding to each cluster identifier set.

[0172] In one possible implementation, the crosstab includes multiple non-zero data, where the non-zero data represents the number of identical object data contained in the first cluster corresponding to the row containing the non-zero data and the second cluster corresponding to the column containing the non-zero data.

[0173] The determination module 1003 is used to determine the largest non-zero data from the non-zero data contained in the column corresponding to any second cluster in the cross-tabulation; and to determine the first cluster corresponding to the row containing the largest non-zero data as the first cluster corresponding to any second cluster.

[0174] In one possible implementation, the determining module 1003 is used to determine, for any one of a plurality of second clusters, in response to the existence of second object data in each object data of any second cluster, to determine that the first cluster to which a second object data belongs is a third cluster, and the outlier detection result of the second object data is outlier object data.

[0175] In one possible implementation, the detection module 1002 is used to perform dimensionality reduction processing on each object data to obtain each dimensionality-reduced object data; and to perform outlier detection processing on each dimensionality-reduced object data to obtain the outlier detection results of each object data.

[0176] In one possible implementation, a dimensionality-reduced object data corresponds to at least two dimensionality-reduced data;

[0177] The detection module 1002 is used to merge at least two types of dimensionality-reduced data corresponding to any dimensionality-reduced object data to obtain merged data corresponding to any dimensionality-reduced object data; and to perform outlier detection processing on the merged data corresponding to each dimensionality-reduced object data to obtain outlier detection results for each object data.

[0178] In one possible implementation, the detection module 1002 is used to perform outlier detection processing on each object data in any one of the multiple first clusters to obtain the outlier detection results of each object data in any one of the first clusters.

[0179] The determination module 1003 is used to determine multiple third clusters based on the outlier detection results of multiple second clusters and the object data in each first cluster.

[0180] In one possible implementation, the determining module 1003 is used to determine at least two clustering results based on outlier detection results of multiple first clusters, multiple second clusters, and each object data. One clustering result includes multiple fourth clusters, and each fourth cluster includes at least one object data. The module also determines an evaluation index for each clustering result, which is used to characterize the accuracy of the clustering result. Based on the clustering result corresponding to the largest evaluation index among the evaluation indices of each clustering result, the module determines multiple third clusters.

[0181] In one possible implementation, the object data is the gene expression matrix of the cell;

[0182] The acquisition module 1001 is used to cluster the gene expression matrices of multiple cells using the Leiden algorithm to obtain multiple first clusters; and to cluster the gene expression matrices of multiple cells using a current-level algorithm based on deep learning feature expression to obtain multiple second clusters.

[0183] In one possible implementation, the acquisition module 1001 is further configured to acquire multiple fifth clusters obtained by clustering multiple object data based on the third clustering algorithm, wherein a fifth cluster includes at least one object data.

[0184] The determination module 1003 is also used to determine multiple sixth clusters based on the outlier detection results of multiple third clusters, multiple fifth clusters and each object data, wherein a sixth cluster includes at least one object data.

[0185] The aforementioned device determines multiple third clusters based on the outlier detection results of multiple first clusters, multiple second clusters, and various object data. The first and second clusters are obtained by clustering multiple object data using different clustering algorithms. This achieves the use of two clustering algorithms to determine the clustering results of multiple object data, thereby improving the accuracy of the clustering results and the accuracy of subsequent data processing.

[0186] It should be understood that the above Figure 10 The provided device, in implementing its functions, is only illustrated by the division of the above-described functional modules. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.

[0187] Figure 11 A structural block diagram of a terminal device 1100 provided in an exemplary embodiment of this application is shown. The terminal device 1100 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, laptop computer, or desktop computer. The terminal device 1100 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0188] Typically, terminal device 1100 includes a processor 1101 and a memory 1102.

[0189] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0190] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 are used to store at least one instruction, which is executed by the processor 1101 to implement the clustering processing method provided in the method embodiments of this application.

[0191] In some embodiments, the terminal device 1100 may also optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1103 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.

[0192] Peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1101 and memory 1102. In some embodiments, processor 1101, memory 1102 and peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1101, memory 1102 and peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0193] The radio frequency (RF) circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0194] Display screen 1105 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1105 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1101 for processing. In this case, display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, disposed on the front panel of terminal device 1100; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 1100 or in a folded design; in still other embodiments, display screen 1105 may be a flexible display screen, disposed on a curved or folded surface of terminal device 1100. Furthermore, display screen 1105 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0195] The camera assembly 1106 is used to acquire images or videos. Optionally, the camera assembly 1106 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0196] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1101 for processing, or input to the radio frequency circuit 1104 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 1100. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.

[0197] The positioning component 1108 is used to locate the current geographical location of the terminal device 1100 in order to enable navigation or LBS (Location Based Service). The positioning component 1108 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.

[0198] Power supply 1109 is used to supply power to the various components in terminal device 1100. Power supply 1109 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1109 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0199] In some embodiments, the terminal device 1100 further includes one or more sensors 1110. The one or more sensors 1110 include, but are not limited to: an accelerometer 1111, a gyroscope 1112, a pressure sensor 1113, a fingerprint sensor 1114, an optical sensor 1115, and a proximity sensor 1116.

[0200] Accelerometer 1111 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 1100. For example, accelerometer 1111 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1101 can control display screen 1105 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1111. Accelerometer 1111 can also be used for games or for acquiring user motion data.

[0201] The gyroscope sensor 1112 can detect the orientation and rotation angle of the terminal device 1100. The gyroscope sensor 1112 can work in conjunction with the accelerometer sensor 1111 to collect the user's 3D movements on the terminal device 1100. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0202] The pressure sensor 1113 can be disposed on the side bezel of the terminal device 1100 and / or on the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side bezel of the terminal device 1100, it can detect the user's grip signal on the terminal device 1100, and the processor 1101 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0203] The fingerprint sensor 1114 is used to collect the user's fingerprint. The processor 1101 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 1114, or the fingerprint sensor 1114 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 1101 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1114 can be located on the front, back, or side of the terminal device 1100. When the terminal device 1100 has a physical button or manufacturer logo, the fingerprint sensor 1114 can be integrated with the physical button or manufacturer logo.

[0204] An optical sensor 1115 is used to collect ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 based on the ambient light intensity collected by the optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 based on the ambient light intensity collected by the optical sensor 1115.

[0205] The proximity sensor 1116, also known as a distance sensor, is typically located on the front panel of the terminal device 1100. The proximity sensor 1116 is used to detect the distance between the user and the front of the terminal device 1100. In one embodiment, when the proximity sensor 1116 detects that the distance between the user and the front of the terminal device 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from a screen-on state to a screen-off state; when the proximity sensor 1116 detects that the distance between the user and the front of the terminal device 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from a screen-off state to a screen-on state.

[0206] Those skilled in the art will understand that Figure 11 The structure shown does not constitute a limitation on the terminal device 1100, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0207] Figure 12 This is a schematic diagram of the server structure provided in the embodiments of this application. The server 1200 can vary considerably due to different configurations or performance. It may include one or more processors 1201 and one or more memories 1202, wherein the one or more memories 1202 store at least one piece of program code. This at least one piece of program code is loaded and executed by the one or more processors 1201 to implement the clustering processing methods provided in the above-described method embodiments. For example, the processor 1201 is a CPU. Of course, the server 1200 may also have wired or wireless network interfaces, keyboards, and input / output interfaces for input and output. The server 1200 may also include other components for implementing device functions, which will not be elaborated here.

[0208] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one piece of program code that is loaded and executed by a processor to enable an electronic device to implement any of the clustering processing methods described above.

[0209] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0210] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer instruction that is loaded and executed by a processor to enable the computer to implement any of the clustering processing methods described above.

[0211] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0212] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0213] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A clustering processing method, characterized in that, The method includes: Multiple first clusters and multiple second clusters are obtained. The multiple first clusters are obtained by clustering multiple object data based on a first clustering algorithm. Each first cluster includes at least one object data. The multiple second clusters are obtained by clustering the multiple object data based on a second clustering algorithm. Each second cluster includes at least one object data. The object data includes a gene expression matrix of a cell. The gene expression matrix includes data with multiple rows and columns. The rows of the gene expression matrix represent the expression of a gene under different environmental conditions or at different time points. The columns of the gene expression matrix represent the expression of a gene under different conditions or samples. The data in any row and any column represents the expression level of a gene in a cell. Outlier detection processing is performed on the multiple object data to obtain the outlier detection results for each object data. Based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the object data, multiple third clusters are determined, and each third cluster includes at least one object data.

2. The method according to claim 1, characterized in that, The outlier detection process performed on the multiple object data to obtain outlier detection results for each object data includes: For any one of the plurality of second clusters, outlier detection processing is performed on each object data in the first second cluster to obtain the outlier detection results for each object data in the first second cluster. Based on the outlier detection results of the plurality of first clusters, the plurality of second clusters, and the object data, a plurality of third clusters are determined, including: Based on the outlier detection results of the object data in the multiple first clusters and each of the second clusters, multiple third clusters are determined.

3. The method according to claim 2, characterized in that, Based on the outlier detection results of the multiple first clusters and the object data in each of the second clusters, multiple third clusters are determined, including: For any one of the plurality of second clusters, in response to the presence of first object data in each object data of the plurality of second clusters, the first cluster corresponding to the plurality of second clusters is determined from the plurality of first clusters, and the outlier detection result of the first object data is non-outlier object data; The first cluster corresponding to any one of the second clusters is determined to be a third cluster; In response to the first object data not belonging to the third cluster, the first object data is added to the third cluster.

4. The method according to claim 3, characterized in that, Determining the first cluster corresponding to any one of the second clusters from the plurality of first clusters includes: A cross-tabulation is determined based on the plurality of first clusters and the plurality of second clusters, wherein a row of data in the cross-tabulation represents the data of each object in a first cluster, and a column of data in the cross-tabulation represents the data of each object in a second cluster; The first cluster corresponding to any one of the second clusters is determined from the plurality of first clusters based on the cross-tabulation.

5. The method according to claim 4, characterized in that, The first cluster corresponds to a first cluster identifier, the second cluster corresponds to a second cluster identifier, and object data includes an object identifier; The determination of the cross-tabulation based on the plurality of first clusters and the plurality of second clusters includes: Based on the first cluster identifier of each first cluster, the object identifier of each object data contained in each first cluster, the second cluster identifier of each second cluster, and the object identifier of each object data contained in each second cluster, the object identifier corresponding to each cluster identifier set is determined. A cluster identifier set includes a first cluster identifier and a second cluster identifier. The cross-tabulation is determined based on the number of object identifiers corresponding to each cluster identifier set.

6. The method according to claim 4, characterized in that, The cross table includes multiple non-zero data, and the non-zero data represents the number of identical object data contained in the first cluster corresponding to the row where the non-zero data is located and the second cluster corresponding to the column where the non-zero data is located. The step of determining the first cluster corresponding to any one of the second clusters from the plurality of first clusters based on the cross-tabulation includes: From the non-zero data contained in the column corresponding to any second cluster in the cross table, determine the largest non-zero data; The first cluster corresponding to the row containing the largest non-zero data is determined as the first cluster corresponding to any second cluster.

7. The method according to claim 2, characterized in that, Based on the outlier detection results of the multiple first clusters and the object data in each of the second clusters, multiple third clusters are determined, including: For any one of the plurality of second clusters, in response to the existence of second object data in each object data in any one of the second clusters, the first cluster to which a second object data belongs is determined to be a third cluster, and the outlier detection result of the second object data is outlier object data.

8. The method according to any one of claims 1 to 7, characterized in that, The outlier detection process performed on the multiple object data to obtain outlier detection results for each object data includes: The dimensionality reduction process is performed on the data of each object to obtain the dimensionality-reduced object data; Outlier detection processing is performed on the data of each dimensionality reduction object to obtain the outlier detection results of each object data.

9. The method according to claim 8, characterized in that, One dimensionality reduction object data corresponds to at least two dimensionality reduction data; The outlier detection process performed on the data of each dimensionality-reduced object to obtain the outlier detection results for each object data includes: For any dimensionality reduction object data, at least two dimensionality reduction data corresponding to the dimensionality reduction object data are merged to obtain the merged data corresponding to the dimensionality reduction object data. Outlier detection processing is performed on the merged data corresponding to each of the dimensionality-reduced object data to obtain the outlier detection results for each object data.

10. The method according to any one of claims 1 to 7, characterized in that, The outlier detection process performed on the multiple object data to obtain outlier detection results for each object data includes: For any one of the plurality of first clusters, outlier detection processing is performed on each object data in the first cluster to obtain the outlier detection results of each object data in the first cluster. Based on the outlier detection results of the plurality of first clusters, the plurality of second clusters, and the object data, a plurality of third clusters are determined, including: Based on the outlier detection results of the multiple second clusters and the object data in each of the first clusters, multiple third clusters are determined.

11. The method according to any one of claims 1 to 7, characterized in that, Based on the outlier detection results of the plurality of first clusters, the plurality of second clusters, and the object data, a plurality of third clusters are determined, including: Based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the object data, at least two clustering results are determined. One clustering result includes multiple fourth clusters, and each fourth cluster includes at least one object data. Determine the evaluation index for each clustering result, wherein the evaluation index is used to characterize the accuracy of the clustering result; Based on the clustering result corresponding to the largest evaluation index among the evaluation indicators of each clustering result, the plurality of third clusters are determined.

12. The method according to any one of claims 1 to 7, characterized in that, The object data is the gene expression matrix of the cell; The acquisition of multiple first clusters and multiple second clusters includes: The gene expression matrix of multiple cells was clustered using the Leiden algorithm to obtain multiple first clusters; The gene expression matrices of the multiple cells are clustered using a current-level algorithm based on deep learning feature representation to obtain multiple second clusters.

13. The method according to any one of claims 1 to 7, characterized in that, After determining multiple third clusters based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the data of each object, the method further includes: Obtain multiple fifth clusters obtained by clustering the multiple object data based on the third clustering algorithm, wherein each fifth cluster includes at least one object data; Based on the outlier detection results of the multiple third clusters, the multiple fifth clusters, and the object data, multiple sixth clusters are determined, and each sixth cluster includes at least one object data.

14. A clustering processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire multiple first clusters and multiple second clusters. The multiple first clusters are obtained by clustering multiple object data based on a first clustering algorithm. Each first cluster includes at least one object data. The multiple second clusters are obtained by clustering the multiple object data based on a second clustering algorithm. Each second cluster includes at least one object data. The object data includes a gene expression matrix of a cell. The gene expression matrix includes data with multiple rows and columns. The rows of the gene expression matrix represent the expression of a gene under different environmental conditions or at different time points. The columns of the gene expression matrix represent the expression status of the gene under different conditions or samples. The data in any row and any column represents the expression level of a gene in a cell. The detection module is used to perform outlier detection processing on the multiple object data to obtain the outlier detection results for each object data. The determination module is used to determine multiple third clusters based on the outlier detection results of the multiple first clusters, the multiple second clusters, and the object data, wherein each third cluster includes at least one object data.

15. The apparatus according to claim 14, characterized in that, The detection module is used to perform outlier detection processing on each object data in any one of the plurality of second clusters, and obtain the outlier detection result of each object data in the any one of the second clusters. The determining module is used to determine multiple third clusters based on the outlier detection results of the multiple first clusters and the object data in each of the multiple second clusters.

16. The apparatus according to claim 15, characterized in that, The determining module is configured to, for any one of the plurality of second clusters, in response to the presence of first object data in each object data of the plurality of second clusters, determine the first cluster corresponding to the plurality of second clusters, wherein the outlier detection result of the first object data is non-outlier object data; The first cluster corresponding to any second cluster is determined to be a third cluster; in response to the first object data not belonging to the third cluster, the first object data is added to the third cluster.

17. The apparatus according to claim 16, characterized in that, The determining module is configured to determine a cross-tabulation based on the plurality of first clusters and the plurality of second clusters, wherein a row of data in the cross-tabulation represents the data of each object in a first cluster, and a column of data in the cross-tabulation represents the data of each object in a second cluster; and determine the first cluster corresponding to any second cluster from the plurality of first clusters based on the cross-tabulation.

18. The apparatus according to claim 17, characterized in that, The first cluster corresponds to a first cluster identifier, the second cluster corresponds to a second cluster identifier, and object data includes an object identifier; The determining module is used to determine the object identifier corresponding to each cluster identifier set based on the first cluster identifier of each first cluster, the object identifier of each object data contained in each first cluster, the second cluster identifier of each second cluster, and the object identifier of each object data contained in each second cluster, wherein a cluster identifier set includes one first cluster identifier and one second cluster identifier; and to determine the cross-tabulation based on the number of object identifiers corresponding to each cluster identifier set.

19. The apparatus according to claim 17, characterized in that, The cross table includes multiple non-zero data, and the non-zero data represents the number of identical object data contained in the first cluster corresponding to the row where the non-zero data is located and the second cluster corresponding to the column where the non-zero data is located. The determining module is used to determine the largest non-zero data from the non-zero data contained in the column corresponding to any second cluster in the cross table; The first cluster corresponding to the row containing the largest non-zero data is determined as the first cluster corresponding to any second cluster.

20. The apparatus according to claim 15, characterized in that, The determining module is used to determine, for any one of the plurality of second clusters, in response to the existence of second object data in each object data of the any one of the second clusters, to determine that the first cluster to which a second object data belongs is a third cluster, and the outlier detection result of the second object data is outlier object data.

21. The apparatus according to any one of claims 14 to 20, characterized in that, The detection module is used to perform dimensionality reduction processing on the object data to obtain dimensionality-reduced object data; and to perform outlier detection processing on the dimensionality-reduced object data to obtain outlier detection results for the object data.

22. The apparatus according to claim 21, characterized in that, One dimensionality reduction object data corresponds to at least two dimensionality reduction data; The detection module is used to, for any dimensionality-reduced object data, merge at least two types of dimensionality-reduced data corresponding to the any dimensionality-reduced object data to obtain merged data corresponding to the any dimensionality-reduced object data; and perform outlier detection processing on the merged data corresponding to each dimensionality-reduced object data to obtain outlier detection results for each object data.

23. The apparatus according to any one of claims 14 to 20, characterized in that, The detection module is used to perform outlier detection processing on each object data in any one of the plurality of first clusters, and obtain the outlier detection result of each object data in the first cluster. The determining module is used to determine multiple third clusters based on the outlier detection results of the multiple second clusters and the object data in each of the multiple first clusters.

24. The apparatus according to any one of claims 14 to 20, characterized in that, The determining module is configured to determine at least two clustering results based on the outlier detection results of the plurality of first clusters, the plurality of second clusters, and the object data, wherein one clustering result includes a plurality of fourth clusters and each fourth cluster includes at least one object data; determine an evaluation index for each clustering result, wherein the evaluation index is used to characterize the accuracy of the clustering result; and determine the plurality of third clusters based on the clustering result corresponding to the largest evaluation index among the evaluation indices of the clustering results.

25. The apparatus according to any one of claims 14 to 20, characterized in that, The object data is the gene expression matrix of the cell; The acquisition module is used to cluster the gene expression matrices of multiple cells using the Leiden algorithm to obtain multiple first clusters; and to cluster the gene expression matrices of the multiple cells using a current-level algorithm based on deep learning feature expression to obtain multiple second clusters.

26. The apparatus according to any one of claims 14 to 20, characterized in that, The acquisition module is further configured to acquire multiple fifth clusters obtained by clustering the multiple object data based on the third clustering algorithm, wherein each fifth cluster includes at least one object data. The determining module is further configured to determine multiple sixth clusters based on the outlier detection results of the multiple third clusters, the multiple fifth clusters, and the object data, wherein each sixth cluster includes at least one object data.

27. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one piece of program code, which is loaded and executed by the processor to enable the electronic device to implement the clustering processing method as described in any one of claims 1 to 13.

28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to enable the computer to implement the clustering processing method as described in any one of claims 1 to 13.

29. A computer program product, characterized in that, The computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement the clustering processing method as described in any one of claims 1 to 13.