Feature extraction method and apparatus for cell grouping, and cell grouping method and apparatus

By perturbing and coding fusion of the cell feature matrix and the adjacency matrix, the target embedding matrix is ​​generated, which solves the problem that traditional methods cannot extract deep cell characteristics and achieves more accurate cell populations.

WO2025123187A1PCT designated stage expired Publication Date: 2025-06-19SHENZHEN HUADA GENE INST +1
3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/137958
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Traditional methods cannot effectively extract deep-seated characteristics of cell gene expression, resulting in inaccurate cell population results.

Method used

By perturbing the cell feature matrix and adjacency matrix of the sample cells, different feature and adjacency matrix views are generated, and these views are encoded and fused using the feature encoder to be trained to generate the target embedding matrix, and the trained feature encoder is obtained through decoding and iterative optimization.

Benefits of technology

Accurate extraction of the characteristics required for cell population is achieved, and the accuracy and depth of cell population is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023137958_19062025_PF_FP_ABST
    Figure CN2023137958_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a feature extraction method and apparatus for cell grouping, and a cell grouping method and apparatus. The feature extraction method for cell grouping comprises: performing disturbance on a cell feature matrix of a sample cell, so as to obtain a first feature matrix and a second feature matrix; performing disturbance on an adjacency matrix of the sample cell, so as to obtain a first adjacency matrix and a second adjacency matrix; by means of a feature encoder to be trained, performing encoding on a first image to obtain a first embedding matrix, and performing encoding on a second image to obtain a second embedding matrix; fusing the first embedding matrix with the second embedding matrix, so as to obtain a target embedding matrix; performing decoding on the target embedding matrix, so as to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; and on the basis of at least one of a first difference and a second difference, iteratively optimizing parameters of said feature encoder until an iteration stop condition is met, so as to obtain a trained feature encoder, wherein the feature encoder is used for extracting features required for cell grouping.
Need to check novelty before this filing date? Find Prior Art

Description

Cell clustering feature extraction method, cell clustering method and device Technical Field

[0001] The present application relates to the fields of computer technology and artificial intelligence technology, and in particular to a cell clustering feature extraction method, a cell clustering method, and a cell clustering device. Background Art

[0002] With the development of artificial intelligence technology, it has become possible to automatically cluster cells. The purpose of cell clustering is to divide cells into different clusters based on the similarity and dissimilarity of their characteristics, ensuring that cells within each cluster are as similar as possible and cells in different clusters are as different as possible.

[0003] Traditional methods typically use principal component analysis (PCA) to represent cell features, reducing the dimensionality of cell features while minimizing information loss. However, this method fails to extract deeper features of gene expression, resulting in inaccurate cell clustering based on this information.

[0004] Summary of the Invention

[0005] Based on this, it is necessary to provide a cell clustering feature extraction method and device, a cell clustering method and device, a computer device and a computer-readable storage medium that can accurately extract the features required for cell clustering in order to address the above technical problems.

[0006] In a first aspect, the present application provides a method for extracting cell clustering features. The method comprises:

[0007] Perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix;

[0008] perturbing the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix;

[0009] Encoding the first graph using a feature encoder to be trained to obtain a first embedding matrix, and encoding the second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix;

[0010] Fusing the first embedding matrix and the second embedding matrix to obtain a target embedding matrix;

[0011] Decoding the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix;

[0012] Based on at least one of the first difference and the second difference, iteratively optimize the parameters of the feature encoder to be trained until the iteration stop condition is met, thereby obtaining a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0013] In a second aspect, the present application also provides a cell clustering method. The method comprises:

[0014] Inputting a graph consisting of a cell feature matrix and an adjacency matrix of cells to be clustered into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained based on a first graph and a second graph; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells;

[0015] Clustering is performed on the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix to obtain cell clustering results of the cells to be clustered.

[0016] In a third aspect, the present application further provides a cell clustering feature extraction device. The device comprises:

[0017] A first perturbation module is used to perturb the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix;

[0018] a second perturbation module, configured to perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix;

[0019] an encoding module, configured to encode a first graph using a feature encoder to be trained to obtain a first embedding matrix, and encode a second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix;

[0020] a fusion module, configured to fuse the first embedding matrix and the second embedding matrix to obtain a target embedding matrix;

[0021] a decoding module, configured to decode the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix;

[0022] A parameter optimization module is configured to iteratively optimize the parameters of the feature encoder to be trained based on at least one of the first difference and the second difference until an iteration stop condition is met, thereby obtaining a trained feature encoder; the feature encoder is configured to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; and the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0023] In a fourth aspect, the present application further provides a cell clustering device. The device comprises:

[0024] An embedding matrix determination module is configured to input a graph consisting of a cell feature matrix and an adjacency matrix of cells to be clustered into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained based on a first graph and a second graph; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells;

[0025] A clustering module is used to cluster the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix to obtain cell clustering results of the cells to be clustered.

[0026] In a fifth aspect, the present application further provides a computer device. The computer device includes a memory and one or more processors, wherein the memory stores a computer program, and when the computer program is executed by the processor, the one or more processors execute the steps of the cell clustering feature extraction method or cell clustering method described in each embodiment of the present application.

[0027] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes one or more processors to perform the steps of the cell clustering feature extraction method or cell clustering method described in each embodiment of the present application.

[0028] The above-mentioned cell clustering feature extraction method and device, cell clustering method and device, computer equipment and computer-readable storage medium perturb the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix, and perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix, which can obtain richer features under different views of the graph, and then use the feature encoder to be trained to encode the first and second images of different views with rich information respectively, and fuse the encoding results of the first and second images to obtain a target embedding matrix, so that the target embedding matrix can have richer and deeper features, and then decode the target embedding matrix to obtain a reconstructed cell feature matrix and a reconstructed adjacency matrix to determine the reconstruction loss, so that the features required for cell clustering can be accurately extracted through the feature encoder that can extract deep features. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0030] FIG1 is a schematic diagram of a process for extracting cell clustering features according to an embodiment;

[0031] FIG2 is a schematic diagram of the overall process of a method for extracting cell clustering features in one embodiment;

[0032] FIG3 is a schematic diagram of a process for cell clustering in one embodiment;

[0033] FIG4 is a structural block diagram of a cell clustering feature extraction device according to an embodiment;

[0034] FIG5 is a structural block diagram of a cell clustering feature extraction device in another embodiment;

[0035] FIG6 is a block diagram of a cell clustering device according to an embodiment;

[0036] FIG7 is a block diagram of a cell clustering device according to another embodiment;

[0037] FIG8 is a diagram showing the internal structure of a computer device according to one embodiment;

[0038] FIG9 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0040] In an exemplary embodiment, as shown in FIG1 , a method for extracting cell clustering features is provided. This embodiment uses the method applied to a computer device as an example. The computer device may be a terminal or a server. The terminal may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, and the like. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, and the like. The server may be implemented as an independent server or a server cluster consisting of multiple servers. In this embodiment, the method comprises the following steps:

[0041] Step 102 : perturb the cell characteristic matrix of the sample cells to obtain a first characteristic matrix and a second characteristic matrix.

[0042] The cell feature matrix is ​​a matrix used to characterize the features of each sample cell.

[0043] In an exemplary embodiment, the characteristics of the sample cells represented by the cell characteristic matrix may include gene characteristics and adjacency relationship characteristics of the sample cells.

[0044] In an exemplary embodiment, the cell signature matrix can be derived from a gene expression matrix and an adjacency matrix of sample cells. The gene expression matrix of sample cells is a matrix used to characterize the genes of each sample cell. The adjacency matrix of sample cells is a matrix used to characterize the adjacency relationships between sample cells.

[0045] In an exemplary embodiment, the adjacency matrix of the sample cells can be determined based on the spatial position matrix of the sample cells. The adjacency matrix of the sample cells is used to characterize the distance similarity between the sample cells. The computer device can determine the distance similarity between the sample cells based on the spatial position matrix of the sample cells, and obtain the adjacency matrix of the sample cells based on the distance similarity. The spatial position matrix of the sample cells is a matrix used to characterize the spatial position of each sample cell.

[0046] In one exemplary embodiment, rows in a sample cell gene expression matrix represent sample cells, and columns represent genes. Rows in a sample cell spatial position matrix represent sample cells, and columns represent coordinate values ​​of the sample cells. In other embodiments, the reverse can be true: columns in a sample cell gene expression matrix represent sample cells, and rows represent genes. Columns in a sample cell spatial position matrix represent sample cells, and rows represent coordinate values ​​of the sample cells.

[0047] In an exemplary embodiment, each row in the adjacency matrix of the sample cells represents the distance similarity between a sample cell node and each sample cell node. For example, the i-th row represents the distance similarity between sample cell node i and each sample cell node, and the i-th row and the j-th column represent the distance similarity between sample cell node i and sample cell node j.

[0048] In an exemplary embodiment, the computer device may perform different perturbations on the cell characteristic matrix of the sample cells to obtain a first characteristic matrix and a second characteristic matrix.

[0049] Step 104 : perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix.

[0050] In an exemplary embodiment, the computer device may perform different perturbations on the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix.

[0051] Step 106: Encode the first graph using a feature encoder to be trained to obtain a first embedding matrix, and encode the second graph to obtain a second embedding matrix; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix.

[0052] In an exemplary embodiment, the first Among them, X1 represents the first characteristic matrix, A m Represents the first adjacency matrix. The second graph Among them, X2 represents the second characteristic matrix, A d Represents the second adjacency matrix.

[0053] In an exemplary embodiment, the computer device may input the first image and the second image into corresponding feature encoders to be trained for encoding, thereby obtaining a first embedding matrix and a second embedding matrix, respectively. Parameters of the feature encoders to be trained corresponding to the first image and the second image are shared.

[0054] Step 108: Fusing the first embedding matrix and the second embedding matrix to obtain a target embedding matrix.

[0055] In an exemplary embodiment, the computer device may linearly add the first embedding matrix and the second embedding matrix to obtain a target embedding matrix.

[0056] Step 110: Decode the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix.

[0057] In an exemplary embodiment, the computer device may input the target embedding matrix into a decoder for decoding to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix.

[0058] In an exemplary embodiment, the computer device may input the target embedding matrix into a decoder for decoding to obtain a reconstructed cell feature matrix, and then determine a reconstructed adjacency matrix based on the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix. In an exemplary embodiment, the computer device may perform an inner product between the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix to obtain a reconstructed adjacency matrix.

[0059] In an exemplary embodiment, the decoder may be a decoder based on a graph convolutional neural network.

[0060] Step 112, iteratively optimize the parameters of the feature encoder to be trained based on at least one of the first difference and the second difference until the iteration stop condition is met, thereby obtaining a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0061] In an exemplary embodiment, the computer device may determine a reconstruction loss value based on at least one of the first difference and the second difference, and iteratively optimize the parameters of the feature encoder to be trained based on the reconstruction loss value until an iteration stop condition is met, thereby obtaining a trained feature encoder.

[0062] In an exemplary embodiment, the computer device may determine a feature matrix reconstruction loss value based on the first difference, determine an adjacency matrix reconstruction loss value based on the second difference, and determine a reconstruction loss value based on the sum of the feature matrix reconstruction loss value and the adjacency matrix reconstruction loss value.

[0063] In an exemplary embodiment, the computer device can determine the feature matrix reconstruction loss value based on the 2 norm of the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells. In an exemplary embodiment, the feature matrix reconstruction loss value L REC-F The calculation formula can be:

[0064] Where N is the number of sample cells and X is the cell feature matrix of the sample cells. is the reconstructed cell feature matrix.

[0065] In an exemplary embodiment, the computer device may determine the adjacency matrix reconstruction loss value based on the 2-norm of the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cell. In an exemplary embodiment, the adjacency matrix reconstruction loss value L REC-A The calculation formula can be:

[0066] Where N is the number of sample cells and A is the adjacency matrix of sample cells. is the reconstructed adjacency matrix.

[0067] In an exemplary embodiment, the reconstruction loss value L REC It can be calculated using the following formula: REC =L REC-F +L REC-A

[0068] In an exemplary embodiment, the iteration stopping condition may include that the reconstruction loss value is less than or equal to a first preset loss threshold. In other embodiments, the iteration stopping condition may include that the number of iterations is greater than or equal to a preset number threshold.

[0069] The above-mentioned cell clustering feature extraction method perturbs the cell feature matrix of the sample cells to obtain the first feature matrix and the second feature matrix, and perturbs the adjacency matrix of the sample cells to obtain the first adjacency matrix and the second adjacency matrix, which can obtain richer features under different views of the graph, and then uses the twin feature encoders to be trained to encode the first and second images of different views with rich information respectively, and fuses the encoding results of the first and second images to obtain the target embedding matrix, so that the target embedding matrix can have richer and deeper features, and then decodes the target embedding matrix to obtain the reconstructed cell feature matrix and the reconstructed adjacency matrix to determine the reconstruction loss, and can accurately train the feature encoder, so that the features required for cell clustering can be accurately extracted through the feature encoder that can extract deep features.

[0070] In an exemplary embodiment, perturbing the cell characteristic matrix of the sample cells to obtain a first characteristic matrix and a second characteristic matrix includes: acquiring a first noise matrix and a second noise matrix; and perturbing the cell characteristic matrix of the sample cells according to the first noise matrix and the second noise matrix, respectively, to obtain a first characteristic matrix and a second characteristic matrix.

[0071] In an exemplary embodiment, the computer device may obtain the first noise matrix and the second noise matrix by randomly sampling from a Gaussian distribution.

[0072] In an exemplary embodiment, the computer device may multiply the first noise matrix with the cell characteristic matrix of the sample cells to obtain a first characteristic matrix, and multiply the second noise matrix with the cell characteristic matrix of the sample cells to obtain a second characteristic matrix.

[0073] In the above embodiment, the cell feature matrix of the sample cells is perturbed according to the first noise matrix and the second noise matrix respectively to obtain the first feature matrix and the second feature matrix, thereby obtaining a richer feature matrix.

[0074] In an exemplary embodiment, perturbing the adjacency matrix of the sample cells to obtain the first adjacency matrix and the second adjacency matrix includes: determining the distance similarity between each sample cell node based on the adjacency matrix of the sample cells; deleting the edges connecting the sample cell nodes whose distance similarity meets a preset condition from the adjacency matrix of the sample cells to obtain the first adjacency matrix; determining the importance of each sample cell node based on the adjacency matrix of the sample cells, and adjusting the distance similarity between each sample cell node based on the importance to obtain the second adjacency matrix.

[0075] Among them, one sample cell node corresponds to one sample cell.

[0076] In an exemplary embodiment, the edges connecting sample cell nodes under the preset conditions may be edges connecting sample cells whose distance similarity between the sample cells is less than or equal to a preset similarity threshold, or edges connecting sample cells corresponding to a preset number of distance similarities selected in ascending order, or edges connecting sample cells corresponding to a preset percentage of distance similarities selected in ascending order. For example, edges connecting sample cells corresponding to the top 10% of distance similarities selected in ascending order may be deleted.

[0077] In an exemplary embodiment, the computer device may determine the importance of each sample cell node according to the adjacency matrix of the sample cells using a personalized PageRank (PPR) algorithm, and adjust the distance similarity between each sample cell node according to the importance to obtain a second adjacency matrix.

[0078] In the above embodiment, edges connecting sample cell nodes whose distance similarity meets preset conditions are deleted from the sample cell adjacency matrix to obtain a first adjacency matrix. The importance of each sample cell node is determined based on the sample cell adjacency matrix, and the distance similarity between each sample cell node is adjusted based on the importance to obtain a second adjacency matrix. This can produce a richer adjacency matrix. By introducing the personalized PageRank (PPR) algorithm to generate the second adjacency matrix, the remote information capture capability of the shallow network structure model is improved, which further enhances the clustering capability.

[0079] In an exemplary embodiment, before perturbing the cell feature matrix of the sample cells to obtain the first feature matrix and the second feature matrix, the method further includes: obtaining a spatial position matrix and a gene expression matrix of the sample cells; the spatial position matrix is ​​used to characterize the spatial position of each sample cell; the gene expression matrix is ​​used to characterize the genes of each sample cell; based on the spatial position matrix, determining the adjacency matrix of the sample cells; and combining the adjacency matrix and the gene expression matrix to obtain the cell feature matrix of the sample cells.

[0080] In an exemplary embodiment, the computer device may combine the adjacency matrix and the gene expression matrix to obtain a cell feature matrix of the sample cells.

[0081] For example, suppose the gene expression matrix is ​​an N*M matrix and the adjacency matrix is ​​an N*N matrix. Where N represents the number of sample cells and M represents the number of genes. The splicing combination method is: gene expression matrix | adjacency matrix, that is, the adjacency matrix is ​​placed on the right side of the gene expression matrix for splicing, or the adjacency matrix can be placed on the left side of the gene expression matrix for splicing. In other embodiments, assuming that the gene expression matrix is ​​an M*N matrix and the adjacency matrix is ​​an N*N matrix, the adjacency matrix can be placed below or above the gene expression matrix for splicing.

[0082] In the above embodiment, the adjacency matrix of the sample cells is determined based on the spatial position matrix, and the adjacency matrix and the gene expression matrix are combined to obtain the cell feature matrix of the sample cells, thereby obtaining various feature information of the sample cells.

[0083] In an exemplary embodiment, encoding the first image by a feature encoder to be trained to obtain a first embedding matrix, and encoding the second image to obtain a second embedding matrix includes: taking the first image and the second image as images to be encoded, respectively, inputting the feature matrix and the adjacency matrix in the image to be encoded into the feature encoder to be trained corresponding to the image to be encoded, so as to encode the image to be encoded layer by layer; sharing parameters between the feature encoders to be trained corresponding to the first image and the second image; in the process of encoding the image to be encoded layer by layer, taking the first encoding layer of the feature encoder to be trained as the current layer, and taking the feature matrix input to the first encoding layer as the embedding matrix input to the current layer, and according to the embedding matrix input to the current layer and the adjacency matrix The product of the normalized matrices determines the fusion matrix of the current layer; the fusion matrix of the current layer and the embedding matrix input to the current layer are weightedly fused according to the weights in the current layer to obtain the embedding matrix output by the current layer; the embedding matrix output by the current layer is used as the input of the next layer, and the next layer is used as the new current layer; the step of determining the fusion matrix of the current layer according to the product of the normalized matrix of the embedding matrix input to the current layer and the adjacency matrix and subsequent steps are returned to execute, and the embedding matrix output by the last layer in the feature encoder to be trained is used as the embedding matrix corresponding to the image to be encoded; wherein, when the image to be encoded is the first image, the corresponding embedding matrix is ​​the first embedding matrix; when the image to be encoded is the second image, the corresponding embedding matrix is ​​the second embedding matrix.

[0084] In an exemplary embodiment, the computer device may weight the product of the embedding matrix of the current layer and the normalized matrix of the adjacency matrix by the weight matrix of the current layer to obtain a first result, weight the embedding matrix of the current layer by the weight matrix of the current layer to obtain a second result, and then add the nonlinear activation function value of the first result and the nonlinear activation function value of the second result to obtain the embedding matrix output by the current layer.

[0085] In an exemplary embodiment, the formula for the first layer of the feature encoder is as follows:

[0086] in, A m and A d are the first adjacency matrix and the second adjacency matrix respectively. m and D d The first adjacency matrix A m and the second adjacency matrix A d The degree matrix of . I is the identity matrix. and The first adjacency matrix A m and the second adjacency matrix A d The normalized matrix of . and is the weight matrix of the feature encoder layer l. b (l) is the bias vector of the feature encoder layer l. σ is the nonlinear activation function. For example, the nonlinear activation function can be ReLU or Tanh. is the encoding result of the lth layer of the feature encoder corresponding to the first image. is the encoding result of the l-1th layer of the feature encoder corresponding to the first figure. is the encoding result of the lth layer of the feature encoder corresponding to the second figure. is the encoding result of the feature encoder layer l-1 corresponding to the second figure. When the number of encoder layers l = 1, is the first input image, The second input image.

[0087] In the above embodiment, since there is no need to mine negative samples, only the positive samples need to be encoded through the encoder of the twin network structure. Therefore, compared with the method of using contrastive learning, the space utilization is reduced, and it is achieved that the space utilization is reduced while learning richer feature representations.

[0088] In an exemplary embodiment, decoding the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix includes: inputting the adjacency matrix of the sample cells and the target embedding matrix into the decoder to decode the target embedding matrix layer by layer; in the process of decoding the target embedding matrix layer by layer, using the first decoding layer of the decoder as the current layer, and using the target embedding matrix input to the first decoding layer as the input of the current layer; weighting the normalized product of the input of the current layer and the adjacency matrix of the sample cells to obtain a decoding result output by the current layer, using the decoding result output by the current layer as the input of the next layer, and using the next layer as the new current layer, returning to execute the step of weighting the normalized product of the input of the current layer and the adjacency matrix of the sample cells to obtain a decoding result output by the current layer and subsequent steps, using the decoding result output by the last layer in the decoder as the reconstructed cell feature matrix; determining the reconstructed adjacency matrix based on the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix.

[0089] In an exemplary embodiment, during the decoding process, the computer device can weight the normalized product of the input of the current layer and the adjacency matrix of the sample cell with the weight matrix of the current layer, and then take the nonlinear activation function value of the weighted result to obtain the decoding result output by the current layer.

[0090] In an exemplary embodiment, the calculation formula for the decoding result of the k-th layer decoder is:

[0091] Among them, H (k) is the decoding result of the kth layer of the decoder. (k-1) is the decoding result of the k-1th layer of the decoder. σ is the nonlinear activation function. A is the adjacency matrix of the sample cell. I is the identity matrix, and D is the degree matrix of A. is the normalized value of A. W (k) is the parameter matrix of the k-th layer of the decoder.

[0092] In an exemplary embodiment, the computer device may perform an inner product on the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix to obtain a reconstructed adjacency matrix.

[0093] In the above embodiment, the decoder performs layer-by-layer decoding to obtain a reconstructed cell feature matrix and a reconstructed adjacency matrix, so that the reconstruction loss can be accurately calculated. According to the reconstruction loss value, the low-dimensional feature embedding can take into account both gene expression information and spatial position information, thereby improving the accuracy of the trained feature encoder.

[0094] In an exemplary embodiment, iteratively optimizing the parameters of the feature encoder to be trained based on at least one of the first difference and the second difference includes: determining a reconstruction loss value based on at least one of the first difference and the second difference; determining a target loss value based on at least one of a clustering guidance loss value and a de-redundancy loss value, and the reconstruction loss value; iteratively optimizing the parameters of the feature encoder to be trained based on the target loss value; wherein the clustering guidance loss value is determined based on the difference between the soft allocation matrix and the target distribution matrix; the soft allocation matrix is ​​determined based on the clustering result obtained by clustering based on the target embedding matrix; the target distribution matrix is ​​obtained by normalizing the soft allocation matrix; and the de-redundancy loss value is determined based on the difference between the first embedding matrix and the second embedding matrix.

[0095] In an exemplary embodiment, the computer device may determine the target loss value according to the sum of at least one of the clustering guidance loss value and the de-redundancy loss value and the reconstruction loss value.

[0096] In an exemplary embodiment, the target loss value L may be calculated as follows: L = L REC +L C +L RR

[0097] Among them, L REC is the reconstruction loss value. L C is the clustering guidance loss value. L RR is the de-redundancy loss value.

[0098] In an exemplary embodiment, the iteration stopping condition may include that the target loss value is less than or equal to a second preset loss threshold. In other embodiments, the iteration stopping condition may include that the number of iterations is greater than or equal to a preset number threshold.

[0099] In the above embodiment, the target loss value is determined based on at least one of the clustering guidance loss value and the de-redundancy loss value, and the reconstruction loss value. According to the clustering guidance loss value, the feature embedding related to the clustering task can be effectively learned. According to the de-redundancy loss value, the redundant information in the embedding can be eliminated, and a distinguishable embedding is generated for each sample cell node. According to the reconstruction loss value, the low-dimensional feature embedding can take into account both gene expression information and spatial position information, thereby further improving the accuracy of the trained feature encoder.

[0100] In an exemplary embodiment, the steps for determining the clustering guidance loss value include: clustering according to the target embedding matrix to obtain reference cluster centers; determining a soft assignment matrix between the vectors corresponding to each sample cell in the target embedding matrix and the reference cluster centers; the soft assignment matrix is ​​used to represent the probability of the vector being assigned to each reference cluster center; normalizing the soft assignment matrix to generate a target distribution matrix; and determining the clustering guidance loss value based on the difference between the distributions of the soft assignment matrix and the target distribution matrix.

[0101] Among them, each row in the target embedding matrix is ​​a vector corresponding to a sample cell.

[0102] In an exemplary embodiment, the computer device may cluster the vectors corresponding to each sample cell in the target embedding matrix to obtain reference cluster centers.

[0103] In an exemplary embodiment, the computer device may use Student's t distribution to calculate a soft assignment matrix between the vectors corresponding to each sample cell in the target embedding matrix and the reference cluster center.

[0104] In an exemplary embodiment, the element p in the i-th row and j-th column of the target distribution matrix P is ij It can be calculated by the following formula:

[0105] Among them, q ij is the element in row i and column j in the soft assignment matrix Q. j' represents the index of a column in the soft assignment matrix.

[0106] In an exemplary embodiment, the computer device may calculate the clustering guidance loss value using KL divergence based on the soft assignment matrix and the target distribution matrix. The formula is as follows:

[0107] In the above embodiment, a soft assignment distribution and a target distribution are generated based on the clustering results of the target embedding matrix, and then the two distributions are aligned by using a clustering guidance loss value to guide network learning, thereby improving the accuracy of the feature encoder and effectively learning feature embeddings related to the clustering task.

[0108] In an exemplary embodiment, the de-redundancy loss value includes a first de-redundancy loss value; the step of determining the de-redundancy loss value includes: determining a node similarity matrix based on the similarity between the first embedding matrix and the second embedding matrix; and determining the first de-redundancy loss value based on the difference between the node similarity matrix and the identity matrix.

[0109] In an exemplary embodiment, the similarity can be cosine similarity. N It can be calculated by the following formula:

[0110] Where H1 is the first embedding matrix. H2 is the second embedding matrix. T represents the transpose of H2, and || || represents the modular operation of the matrix.

[0111] In an exemplary embodiment, the computer device may calculate the MSE of the node similarity matrix and the identity matrix to determine the first de-redundancy loss value. RR-N It can be calculated by the following formula:

[0112] Where N is the number of sample cells. N is the node similarity matrix. I is the identity matrix.

[0113] In the above embodiment, a node similarity matrix is ​​determined based on the similarity between the first embedding matrix and the second embedding matrix, and a first de-redundancy loss value is determined based on the difference between the node similarity matrix and the identity matrix. This decorrelation operation can enable the network to reduce redundant information between sample cell nodes in the latent space, thereby making the learned embedding more discriminative, enhancing the robustness of the feature encoder, and solving the problem of representation collapse existing in traditional methods.

[0114] In an exemplary embodiment, the de-redundancy loss value also includes a second de-redundancy loss value; the step of determining the de-redundancy loss value also includes: mapping the first embedding matrix and the second embedding matrix into cluster-level embeddings respectively to obtain a first cluster-level embedding matrix and a second cluster-level embedding matrix; determining a cluster-level similarity matrix based on the similarity between the first cluster-level embedding matrix and the second cluster-level embedding matrix; and determining the second de-redundancy loss value based on the difference between the cluster-level similarity matrix and the unit matrix.

[0115] Among them, cluster-level embedding is obtained by clustering the vectors of each sample cell node in the embedding matrix and averaging the vectors in the same cluster.

[0116] In an exemplary embodiment, a computer device may utilize a read function The first embedding matrix and the second embedding matrix are mapped to the cluster-level embedding to obtain a first cluster-level embedding matrix and a second cluster-level embedding matrix.

[0117] In an exemplary embodiment, the similarity can be cosine similarity. Cluster-level similarity matrix S F It can be calculated by the following formula:

[0118] Where Z1 is the first cluster-level embedding matrix, and Z2 is the second cluster-level embedding matrix.

[0119] In an exemplary embodiment, the computer device may calculate the MSE of the cluster-level similarity matrix and the identity matrix to determine the second de-redundancy loss value. RR-F It can be calculated by the following formula:

[0120] Where D is the dimension of the target embedding matrix H. F is the cluster-level similarity matrix. I is the identity matrix.

[0121] In an exemplary embodiment, the computer device may determine the de-redundancy loss value according to the sum of the first de-redundancy loss value and the second de-redundancy loss value. RR The calculation formula is: RR =L RR-N +L RR-F

[0122] In the above embodiment, mapping node embeddings to cluster-level embeddings can reduce redundant information at the feature level, thereby making the learned embeddings more discriminative.

[0123] In an exemplary embodiment, the method further includes: inputting a graph consisting of a cell feature matrix and an adjacency matrix of the cells to be clustered into a trained feature encoder to obtain an extracted embedding matrix; clustering the embedding vectors corresponding to each cell to be clustered in the extracted embedding matrix to obtain a cell clustering result of the cells to be clustered.

[0124] The extracted embedding matrix refers to the embedding matrix output by the feature encoder after encoding the input image. The cell clustering result of the cells to be clustered refers to the result of dividing the cells to be clustered into multiple cell clusters.

[0125] In an exemplary embodiment, the cells to be grouped may be cells other than sample cells. In other embodiments, the cells to be grouped may be sample cells themselves.

[0126] In an exemplary embodiment, each row of the extracted embedding matrix represents an embedding vector corresponding to a cell to be clustered. The computer device may cluster the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix, and group cells to be clustered corresponding to vectors belonging to the same cluster in the clustering results into the same cell group.

[0127] In the above embodiment, the embedding matrix of the cells to be clustered is extracted by a feature encoder that has been trained to extract deep features, and then clustering is performed based on the extracted embedding matrix that has learned deep features, which can improve the accuracy of cell clustering.

[0128] In an exemplary embodiment, the method further includes: clustering the embedding vectors corresponding to each sample cell in the target embedding matrix to obtain a cell clustering result of the sample cells.

[0129] The cell clustering result of the sample cells refers to the result of dividing the sample cells into multiple cell clusters.

[0130] In the above embodiment, when a group of sample cells needs to be clustered, a feature encoder that can extract deep-level features can be trained based on the cell feature matrix and adjacency matrix of the sample cells, so that the obtained target embedding matrix can reflect the deep-level features, thereby directly clustering according to the target embedding matrix to obtain accurate cell clustering results of the sample cells.

[0131] As shown in Figure 2, it is a schematic diagram of the overall process of the cell cluster feature extraction method in each embodiment of the present application. First, the cell feature matrix of the sample cells is disturbed by the first noise matrix N1 and the second noise matrix N2 respectively to obtain the first feature matrix X1 and the second feature matrix X2, and the adjacency matrix A is disturbed differently to obtain the first adjacency matrix A. m and the second adjacency matrix A d . The first picture and the second picture Input the parameter-shared feature encoder to be trained to obtain the first embedding matrix H1 and the second embedding matrix H2, and linearly add the first embedding matrix H1 and the second embedding matrix H2 to obtain the target embedding matrix H. Then, the target embedding matrix H is input to the decoder to obtain the reconstructed cell feature matrix and the reconstructed adjacency matrix To determine the reconstruction loss value L REC Determine the de-redundancy loss value L according to the similarity between the first embedding matrix H1 and the second embedding matrix H2RR The clustering guidance loss value L is determined based on the result of clustering the target embedding matrix H. C According to the reconstruction loss value L REC De-redundancy loss value L RR and clustering guidance loss value L C The parameters of the feature encoder are optimized to finally obtain a trained feature encoder, which can be used to extract the features required for cell clustering.

[0132] In an exemplary embodiment, as shown in FIG3 , a cell clustering method is provided, and this embodiment is illustrated by applying the method to a computer device. The computer device may be a terminal or a server. The terminal may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, and the like. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, and the like. The server may be implemented as an independent server or a server cluster consisting of multiple servers. In this embodiment, the method comprises the following steps:

[0133] In step 302, a graph consisting of a cell feature matrix and an adjacency matrix of cells to be clustered is input into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained based on the first graph and the second graph; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of the sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells.

[0134] Step 304 : Clustering the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix to obtain cell clustering results of the cells to be clustered.

[0135] The above-mentioned cell clustering method extracts the embedding matrix of the cells to be clustered through a feature encoder that has been trained to extract deep-level features, and then performs clustering based on the extracted embedding matrix that has learned deep-level features, which can improve the accuracy of cell clustering.

[0136] In an exemplary embodiment, the training steps of the feature encoder include: perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix; perturbing the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix; encoding the first graph to be trained to obtain a first embedding matrix, and encoding the second graph to obtain a second embedding matrix; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; fusing the first embedding matrix and the second embedding matrix to obtain a target embedding matrix; decoding the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; iteratively optimizing the parameters of the feature encoder to be trained according to at least one of the first difference and the second difference until the iteration stop condition is met to obtain a trained feature encoder; the feature encoder is used to extract the features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0137] In the above embodiment, the cell feature matrix of the sample cells is perturbed to obtain a first feature matrix and a second feature matrix, and the adjacency matrix of the sample cells is perturbed to obtain a first adjacency matrix and a second adjacency matrix, so that richer features can be obtained under different views of the graph. Then, the first graph and the second graph of different views with rich information are encoded respectively by the twin feature encoders to be trained, and the encoding results of the first graph and the second graph are fused to obtain a target embedding matrix, so that the target embedding matrix can have richer and deeper features. The target embedding matrix is ​​then decoded to obtain a reconstructed cell feature matrix and a reconstructed adjacency matrix, and then the reconstruction loss is determined. The feature encoder can be accurately trained, so that the features required for cell clustering can be accurately extracted through the feature encoder that can extract deep features.

[0138] The feature encoder model was evaluated on data from four mouse embryonic developmental stages: E9.5 E1S1, E9.5 E2S2, E9.5 E2S3, and E9.5 E2S4, as well as human breast cancer (10x Visium) data. Widely used evaluation metrics were used: the Adjusted Rand Index (ARI), which ranges from -1 to 1, with values ​​closer to 1 indicating a positive performance; the Normalized Mutual Information (NMI), which ranges from 0 to 1, with values ​​closer to 1 indicating a positive performance; and the Fowlkes-Mallows Index (FMI), which ranges from 0 to 1, with values ​​closer to 1 indicating a positive performance. These three metrics measure the degree of fit between two distributions; larger values ​​indicate a closer fit between the clustering results and the ground truth.

[0139] Table 1. Test results of the model on the mouse embryo E9.5 E1S1 stage data

[0140] As can be seen from Table 1, the algorithm proposed in this application is far superior to the existing algorithm model in most indicators.

[0141] Table 2. Test results of the model on the mouse embryo E9.5 E2S2 stage data

[0142] As can be seen from Table 2, the algorithm proposed in this application is far superior to the existing algorithm model in all indicators.

[0143] Table 3. Test results of the model on the mouse embryo E9.5 E2S3 stage data

[0144] As can be seen from Table 3, the application proposed by this patent is far superior to the existing algorithm model in all indicators.

[0145] Table 4. Test results of the model on the mouse embryo E9.5 E2S4 stage data

[0146] It can be seen from Table 4 that the algorithm proposed in this application is far superior to the existing algorithm model in all indicators.

[0147] Table 5. Test results of the model on human breast cancer (10x Visium) data

[0148] It can be seen from Table 5 that the algorithm proposed in this application is far superior to the existing algorithm model in all indicators.

[0149] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0150] Based on the same inventive concept, the present application also provides a cell clustering feature extraction device for implementing the cell clustering feature extraction method mentioned above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more cell clustering feature extraction device embodiments provided below can be found in the above-mentioned limitations of the cell clustering feature extraction method and will not be repeated here.

[0151] In an exemplary embodiment, as shown in FIG4 , a cell cluster feature extraction device 400 is provided, comprising: a first perturbation module 402 , a second perturbation module 404 , an encoding module 406 , a fusion module 408 , a decoding module 410 , and a parameter optimization module 412 , wherein:

[0152] The first perturbation module 402 is configured to perturb the cell characteristic matrix of the sample cells to obtain a first characteristic matrix and a second characteristic matrix.

[0153] The second perturbation module 404 is configured to perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix.

[0154] The encoding module 406 is configured to encode the first graph using a feature encoder to be trained to obtain a first embedding matrix, and encode the second graph to obtain a second embedding matrix; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix.

[0155] A fusion module 408 is configured to fuse the first embedding matrix and the second embedding matrix to obtain a target embedding matrix;

[0156] The decoding module 410 is configured to decode the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix.

[0157] The parameter optimization module 412 is used to iteratively optimize the parameters of the feature encoder to be trained based on at least one of the first difference and the second difference until the iteration stopping condition is met, thereby obtaining a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0158] In an exemplary embodiment, the first perturbation module 402 is further configured to obtain a first noise matrix and a second noise matrix; and to perturb the cell feature matrix of the sample cells according to the first noise matrix and the second noise matrix, respectively, to obtain a first feature matrix and a second feature matrix.

[0159] In an exemplary embodiment, the second perturbation module 404 is further configured to determine the distance similarity between each sample cell node based on the adjacency matrix of the sample cells; delete the edges connecting the sample cell nodes whose distance similarity meets a preset condition from the adjacency matrix of the sample cells to obtain a first adjacency matrix; determine the importance of each sample cell node based on the adjacency matrix of the sample cells, and adjust the distance similarity between each sample cell node based on the importance to obtain a second adjacency matrix.

[0160] In an exemplary embodiment, the first perturbation module 402 is further used to obtain a spatial position matrix and a gene expression matrix of the sample cells; the spatial position matrix is ​​used to characterize the spatial position of each sample cell; the gene expression matrix is ​​used to characterize the genes of each sample cell; based on the spatial position matrix, an adjacency matrix of the sample cells is determined; and the adjacency matrix and the gene expression matrix are combined to obtain a cell feature matrix of the sample cells.

[0161] In an exemplary embodiment, the encoding module 406 is further configured to use the first image and the second image as images to be encoded, respectively, and input the feature matrix and the adjacency matrix of the image to be encoded into the feature encoder to be trained corresponding to the image to be encoded, so as to encode the image to be encoded layer by layer; parameters are shared between the feature encoders to be trained corresponding to the first image and the second image; in the process of encoding the image to be encoded layer by layer, the first encoding layer of the feature encoder to be trained is used as the current layer, and the feature matrix input to the first encoding layer is used as the embedding matrix input to the current layer, and the fusion matrix of the current layer is determined according to the product of the embedding matrix input to the current layer and the normalized matrix of the adjacency matrix; The fusion matrix of the current layer and the embedding matrix input to the current layer are weightedly fused according to the weights in the current layer to obtain the embedding matrix output by the current layer, the embedding matrix output by the current layer is used as the input of the next layer, and the next layer is used as the new current layer, and the step of determining the fusion matrix of the current layer according to the product of the normalized matrix of the embedding matrix input to the current layer and the adjacency matrix and subsequent steps are returned to execute, and the embedding matrix output by the last layer in the feature encoder to be trained is used as the embedding matrix corresponding to the image to be encoded; wherein, when the image to be encoded is the first image, the corresponding embedding matrix is ​​the first embedding matrix; when the image to be encoded is the second image, the corresponding embedding matrix is ​​the second embedding matrix.

[0162] In an exemplary embodiment, it is also used to input the adjacency matrix of the sample cells and the target embedding matrix into the decoder to decode the target embedding matrix layer by layer; in the process of decoding the target embedding matrix layer by layer, the first decoding layer of the decoder is used as the current layer, and the target embedding matrix input to the first decoding layer is used as the input of the current layer; the normalized product of the input of the current layer and the adjacency matrix of the sample cells is weighted to obtain the decoding result output by the current layer, the decoding result output by the current layer is used as the input of the next layer, and the next layer is used as the new current layer, and the step of weighting the normalized product of the input of the current layer and the adjacency matrix of the sample cells to obtain the decoding result output by the current layer and the subsequent steps are returned, and the decoding result output by the last layer in the decoder is used as the reconstructed cell feature matrix; the reconstructed adjacency matrix is ​​determined according to the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix.

[0163] In an exemplary embodiment, the parameter optimization module 412 is further used to determine a reconstruction loss value based on at least one of the first difference and the second difference; determine a target loss value based on at least one of the clustering guidance loss value and the de-redundancy loss value, and the reconstruction loss value; iteratively optimize the parameters of the feature encoder to be trained based on the target loss value; wherein, the clustering guidance loss value is determined based on the difference between the soft allocation matrix and the target distribution matrix; the soft allocation matrix is ​​determined based on the clustering result obtained by clustering based on the target embedding matrix; the target distribution matrix is ​​obtained by normalizing the soft allocation matrix; and the de-redundancy loss value is determined based on the difference between the first embedding matrix and the second embedding matrix.

[0164] In an exemplary embodiment, the parameter optimization module 412 is further used to perform clustering based on the target embedding matrix to obtain reference cluster centers; determine a soft assignment matrix between the vectors corresponding to each sample cell in the target embedding matrix and the reference cluster centers; the soft assignment matrix is ​​used to characterize the probability that the vectors are assigned to each reference cluster center; normalize the soft assignment matrix to generate a target distribution matrix; and determine a clustering guidance loss value based on the difference between the distributions of the soft assignment matrix and the target distribution matrix.

[0165] In an exemplary embodiment, the parameter optimization module 412 is further configured to determine a node similarity matrix based on the similarity between the first embedding matrix and the second embedding matrix; and determine a first de-redundancy loss value based on the difference between the node similarity matrix and the identity matrix.

[0166] In an exemplary embodiment, the parameter optimization module 412 is further used to map the first embedding matrix and the second embedding matrix into cluster-level embeddings, respectively, to obtain a first cluster-level embedding matrix and a second cluster-level embedding matrix; determine a cluster-level similarity matrix based on the similarity between the first cluster-level embedding matrix and the second cluster-level embedding matrix; and determine a second de-redundancy loss value based on the difference between the cluster-level similarity matrix and the unit matrix.

[0167] In an exemplary embodiment, as shown in FIG5 , the cell clustering feature extraction device 400 further includes a clustering module 414 .

[0168] The encoding module 406 is further configured to input the graph consisting of the cell feature matrix and the adjacency matrix of the cells to be clustered into the trained feature encoder to obtain an extracted embedding matrix.

[0169] The clustering module 414 is further configured to cluster the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix to obtain cell clustering results of the cells to be clustered.

[0170] In an exemplary embodiment, the clustering module 414 is further configured to cluster the embedding vectors corresponding to the sample cells in the target embedding matrix to obtain cell clustering results of the sample cells.

[0171] In an exemplary embodiment, as shown in FIG6 , a cell clustering device 600 is provided, comprising: an embedding matrix determination module 602 and a clustering module 604 , wherein:

[0172] The embedding matrix determination module 602 is used to input the graph consisting of the cell feature matrix and the adjacency matrix of the cells to be clustered into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained based on the first graph and the second graph; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of the sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells.

[0173] The clustering module 604 is used to cluster the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix to obtain cell clustering results of the cells to be clustered.

[0174] In an exemplary embodiment, the cell clustering device 600 further includes: a model training module 606 .

[0175] The model training module 606 is also used to perturb the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix; perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix; encode the first graph through the feature encoder to be trained to obtain a first embedding matrix, and encode the second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix; fuse the first embedding matrix and the second embedding matrix to obtain a target embedding matrix; decode the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; iteratively optimize the parameters of the feature encoder to be trained according to at least one of the first difference and the second difference until the iteration stop condition is met to obtain a trained feature encoder; the feature encoder is used to extract the features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

[0176] The cell clustering feature extraction device and each module in the cell clustering device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0177] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be shown in Figure 7. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a cell clustering feature extraction method or a cell clustering method is implemented.

[0178] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be shown in Figure 8. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a cell clustering feature extraction method or a cell clustering method is implemented. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0179] Those skilled in the art will understand that the structure shown in Figure 7 or Figure 8 is merely a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0180] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0181] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0182] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0183] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0184] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0185] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for extracting cell population characteristics, characterized in that, The method includes: perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix; perturbing the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix; encoding the first graph through a feature encoder to be trained to obtain a first embedding matrix, and encoding the second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix; fusing the first embedding matrix and the second embedding matrix to obtain a target embedding matrix; decoding the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; iteratively optimizing the parameters of the feature encoder to be trained according to at least one of a first difference and a second difference until an iteration stop condition is met, to obtain a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

2. The method according to claim 1, characterized in that, The perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix includes: obtaining a first noise matrix and a second noise matrix; perturbing the cell feature matrix of the sample cells according to the first noise matrix and the second noise matrix respectively to obtain a first feature matrix and a second feature matrix.

3. The method according to claim 1 or 2, characterized in that, The perturbing the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix includes: determining the distance similarity between each sample cell node according to the adjacency matrix of the sample cells; deleting the edges connected between the sample cell nodes with distance similarity meeting a preset condition from the adjacency matrix of the sample cells to obtain a first adjacency matrix; determining the importance of each sample cell node according to the adjacency matrix of the sample cells, and adjusting the distance similarity between each sample cell node according to the importance to obtain a second adjacency matrix.

4. The method according to any one of claims 1 to 3, characterized in that, Before the perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix, the method further includes: obtaining a spatial position matrix and a gene expression matrix of the sample cells; the spatial position matrix is used to characterize the spatial position of each sample cell; the gene expression matrix is used to characterize the gene of each sample cell; determining the adjacency matrix of the sample cells according to the spatial position matrix; combining the adjacency matrix and the gene expression matrix to obtain the cell feature matrix of the sample cells.

5. The method according to any one of claims 1 to 4, characterized in that, The encoding the first graph through a feature encoder to be trained to obtain a first embedding matrix, and encoding the second graph to obtain a second embedding matrix includes: Take the first graph and the second graph as the graphs to be encoded respectively, and input the feature matrix and the adjacency matrix in the graph to be encoded into the feature encoder to be trained corresponding to the graph to be encoded, so as to perform layer-by-layer encoding on the graph to be encoded; the parameters between the feature encoders to be trained corresponding to the first graph and the second graph are shared; During the process of performing layer-by-layer encoding on the graph to be encoded, take the first encoding layer of the feature encoder to be trained as the current layer, and take the feature matrix input to the first encoding layer as the embedding matrix input to the current layer, and determine the fusion matrix of the current layer according to the product of the embedding matrix input to the current layer and the normalized matrix of the adjacency matrix; Perform weighted fusion on the fusion matrix of the current layer and the embedding matrix input to the current layer according to the weights in the current layer to obtain the embedding matrix output by the current layer, take the embedding matrix output by the current layer as the input of the next layer, and take the next layer as the new current layer, return to execute the step of determining the fusion matrix of the current layer according to the product of the embedding matrix input to the current layer and the normalized matrix of the adjacency matrix and subsequent steps, and take the embedding matrix output by the last layer in the feature encoder to be trained as the embedding matrix corresponding to the graph to be encoded; Among them, when the graph to be encoded is the first graph, the corresponding embedding matrix is the first embedding matrix; when the graph to be encoded is the second graph, the corresponding embedding matrix is the second embedding matrix.

6. The method according to any one of claims 1 to 5, characterized in that, The decoding the target embedding matrix to obtain at least one of the reconstructed cell feature matrix and the reconstructed adjacency matrix includes: Input the adjacency matrix of the sample cells and the target embedding matrix into the decoder to perform layer-by-layer decoding on the target embedding matrix; During the process of performing layer-by-layer decoding on the target embedding matrix, take the first decoding layer of the decoder as the current layer, and take the target embedding matrix input to the first decoding layer as the input of the current layer; Perform weighting on the product of the input of the current layer and the normalization of the adjacency matrix of the sample cells to obtain the decoding result output by the current layer, take the decoding result output by the current layer as the input of the next layer, and take the next layer as the new current layer, return to execute the step of performing weighting on the product of the input of the current layer and the normalization of the adjacency matrix of the sample cells to obtain the decoding result output by the current layer and subsequent steps, and take the decoding result output by the last layer in the decoder as the reconstructed cell feature matrix; Determine the reconstructed adjacency matrix according to the reconstructed cell feature matrix and the transpose of the reconstructed cell feature matrix matrix.

7. The method according to any one of claims 1 to 6, characterized in that, The iteratively optimizing the parameters of the feature encoder to be trained according to at least one of the first difference and the second difference includes: Determine the reconstruction loss value according to at least one of the first difference and the second difference; Determine the target loss value according to at least one of the clustering guidance loss value and the redundancy removal loss value, and the reconstruction loss value; Iteratively optimize the parameters of the feature encoder to be trained according to the target loss value; Among them, the clustering guidance loss value is determined according to the difference between the soft assignment matrix and the target distribution matrix; the soft assignment matrix is determined according to the clustering result obtained by clustering based on the target embedding matrix; the target distribution matrix is obtained by normalizing the soft assignment matrix; The redundancy removal loss value is determined according to the difference between the first embedding matrix and the second embedding matrix.

8. The method according to claim 7, wherein, The steps for determining the clustering guidance loss value include: Clustering according to the target embedding matrix to obtain a reference clustering center; Determining a soft assignment matrix between the vectors corresponding to each sample cell in the target embedding matrix and the reference clustering center; the soft assignment matrix is used to represent the probabilities that the vectors are respectively assigned to each of the reference clustering centers; Normalizing the soft assignment matrix to generate a target distribution matrix; Determining the clustering guidance loss value according to the difference between the distributions of the soft assignment matrix and the target distribution matrix.

9. The method according to claim 7 or 8, wherein, The redundancy removal loss value includes a first redundancy removal loss value; the steps for determining the redundancy removal loss value include: Determining a node similarity matrix according to the similarity between the first embedding matrix and the second embedding matrix; Determining the first redundancy removal loss value according to the difference between the node similarity matrix and the identity matrix.

10. The method according to claim 9, wherein, The redundancy removal loss value further includes a second redundancy removal loss value; the steps for determining the redundancy removal loss value further include: Mapping the first embedding matrix and the second embedding matrix into cluster-level embeddings respectively to obtain a first cluster-level embedding matrix and a second cluster-level embedding matrix; Determining a cluster-level similarity matrix according to the similarity between the first cluster-level embedding matrix and the second cluster-level embedding matrix; Determining the second redundancy removal loss value according to the difference between the cluster-level similarity matrix and the identity matrix.

11. The method according to any one of claims 1 to 10, wherein, The method further includes: Inputting the graph composed of the cell feature matrix and the adjacency matrix of the cells to be grouped into the trained feature encoder to obtain an extracted embedding matrix; Clustering the embedding vectors corresponding to each of the cells to be grouped in the extracted embedding matrix to obtain the cell grouping result of the cells to be grouped.

12. The method according to any one of claims 1 to 10, wherein, The method further includes: Clustering the embedding vectors corresponding to each of the sample cells in the target embedding matrix to obtain the cell grouping result of the sample cells.

13. A method for cell population, wherein, The method includes: Inputting the graph composed of the cell feature matrix and the adjacency matrix of the cells to be grouped into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained according to a first graph and a second graph; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of the sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells; Clustering the embedding vectors corresponding to each of the cells to be grouped in the extracted embedding matrix to obtain the cell grouping result of the cells to be grouped.

14. The method according to claim 13, wherein, The training steps of the feature encoder include: Perturbing the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix; Perturbing the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix; Encoding the first graph through the feature encoder to be trained to obtain a first embedding matrix, and encoding the second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix; Fusing the first embedding matrix and the second embedding matrix to obtain a target embedding matrix; Decoding the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; Iteratively optimizing the parameters of the feature encoder to be trained according to at least one of a first difference and a second difference until the iteration stop condition is met, to obtain a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

15. A cell population feature extraction device, characterized in that The device includes: A first perturbation module, configured to perturb the cell feature matrix of the sample cells to obtain a first feature matrix and a second feature matrix; A second perturbation module, configured to perturb the adjacency matrix of the sample cells to obtain a first adjacency matrix and a second adjacency matrix; An encoding module, configured to encode the first graph through the feature encoder to be trained to obtain a first embedding matrix, and encode the second graph to obtain a second embedding matrix; the first graph includes the first feature matrix and the first adjacency matrix; the second graph includes the second feature matrix and the second adjacency matrix; A fusion module, configured to fuse the first embedding matrix and the second embedding matrix to obtain a target embedding matrix; A decoding module, configured to decode the target embedding matrix to obtain at least one of a reconstructed cell feature matrix and a reconstructed adjacency matrix; A parameter optimization module, configured to iteratively optimize the parameters of the feature encoder to be trained according to at least one of a first difference and a second difference until the iteration stop condition is met, to obtain a trained feature encoder; the feature encoder is used to extract features required for cell clustering; the first difference is the difference between the reconstructed cell feature matrix and the cell feature matrix of the sample cells; the second difference is the difference between the reconstructed adjacency matrix and the adjacency matrix of the sample cells.

16. A cell population device, characterized in that The device includes: An embedding matrix determination module, configured to input a graph composed of a cell feature matrix and an adjacency matrix of cells to be clustered into a pre-trained feature encoder to obtain an extracted embedding matrix; the pre-trained feature encoder is pre-trained according to a first graph and a second graph; the first graph includes a first feature matrix and a first adjacency matrix; the second graph includes a second feature matrix and a second adjacency matrix; the first feature matrix and the second feature matrix are obtained by perturbing the cell feature matrix of sample cells; the first adjacency matrix and the second adjacency matrix are obtained by perturbing the adjacency matrix of the sample cells. A clustering module, configured to cluster the embedding vectors corresponding to the cells to be clustered in the extracted embedding matrix respectively to obtain a cell clustering result of the cells to be clustered.

17. A computer device, comprising a memory and one or more processors, the memory storing computer-readable instructions, characterized in that When the one or more processors execute the computer-readable instructions, the steps of the method according to any one of claims 1 to 14 are implemented.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are adapted to be loaded and executed by one or more processors to perform the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Training method of neural network model

    CN114003960A

  • High-precision single cell clustering method and system based on marker gene and ensemble learning

    CN115512772A

  • Training method and apparatus for neutral network for image recognition

    US20170061246A1