Training of cell feature extraction model, cell feature extraction method and device

By constructing a cell feature extraction model and using a neural network model to perform data augmentation and feature extraction on the reference cell map, the problem of accuracy in cellular material group data feature extraction was solved, and efficient feature extraction was achieved in noisy environments.

CN116978004BActive Publication Date: 2025-12-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211231869.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-12-05
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

How to effectively extract the material composition data features of cells for analysis and processing is an urgent problem to be solved in the existing technology.

Method used

By constructing a cell feature extraction model, using a neural network model to perform data augmentation on the reference cell image, generating a first cell image and a second cell image, and performing feature extraction, the cell feature extraction model is trained to extract features of the target cell.

Benefits of technology

It enhances the logic and organization of the data, improves the accuracy of feature extraction and noise resistance, and ensures accurate extraction of cell features in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116978004B_ABST
    Figure CN116978004B_ABST
Patent Text Reader

Abstract

The application discloses a kind of training of cell feature extraction model, cell feature extraction method and device, belong to biological technology field.Method includes: obtaining reference cell graph, the node of reference cell graph is characterized the material group data of sample cell, the edge of reference cell graph is characterized the correlation of the sample cell corresponding to both ends node;Data enhancement is carried out to reference cell graph to obtain first cell graph and second cell graph;The first feature of each sample cell, second feature is obtained by neural network model to first cell graph, second cell graph is carried out feature extraction;Based on the first feature and second feature of each sample cell, neural network model is trained to obtain cell feature extraction model.Because the accuracy of the first feature and second feature of sample cell is higher, and eliminate noise to a certain extent, therefore, the cell feature extraction model obtained based on the first feature and second feature of sample cell can extract accurate cell feature, and have certain anti-noise performance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of biotechnology, and in particular to a training method of a cell feature extraction model, a cell feature extraction method and device. BACKGROUND

[0002] In the field of biotechnology, a cell is a basic building unit of a living being. There are proteome, genome, metabolome and other substance groups in a cell. Any substance group data can be used to describe a cell. By analyzing the substance group data, the physiological state and function of the cell can be known, which is helpful for elucidating disease mechanism and preventing and treating diseases.

[0003] Generally, when analyzing the substance group data of a cell, cell features are extracted from the substance group data of the cell for analysis and processing based on the cell features. Therefore, how to extract cell features becomes a technical problem to be solved. SUMMARY

[0004] The present application provides a training method of a cell feature extraction model, a cell feature extraction method and device, which can be used to solve the problems in the related art. The technical solution includes the following contents.

[0005] In one aspect, a training method of a cell feature extraction model is provided, which includes:

[0006] obtaining a reference cell graph of a sample data set, the sample data set including substance group data of a plurality of sample cells, the reference cell graph including a plurality of nodes and a plurality of edges, any node representing the substance group data of any sample cell, and any edge representing that the sample cells corresponding to the two end nodes of the edge are related cells;

[0007] performing data enhancement processing on the reference cell graph to obtain a first cell graph and a second cell graph, the first cell graph and the second cell graph being different cell graphs;

[0008] extracting features of the sample cells from the first cell graph by a neural network model to obtain first features of the sample cells;

[0009] extracting features of the sample cells from the second cell graph by the neural network model to obtain second features of the sample cells;

[0010] training the neural network model based on the first features and the second features of the sample cells to obtain a cell feature extraction model, the cell feature extraction model being used to extract cell features of a target cell.

[0011] In another aspect, a cell feature extraction method is provided, which includes:

[0012] obtain a target cell graph of a target dataset, the target dataset comprising material group data of a plurality of target cells, the target cell graph comprising a plurality of nodes and a plurality of edges, any node representing material group data of any target cell, and any edge representing that the target cells corresponding to the two nodes of the edge are related;

[0013] extract features of the target cell graph by a cell feature extraction model to obtain cell features of each target cell, the cell feature extraction model being trained according to the training method of the cell feature extraction model.

[0014] In another aspect, a training device of a cell feature extraction model is provided, and the device comprises:

[0015] an obtaining module configured to obtain a reference cell graph of a sample dataset, the sample dataset comprising material group data of a plurality of sample cells, the reference cell graph comprising a plurality of nodes and a plurality of edges, any node representing material group data of any sample cell, and any edge representing that the sample cells corresponding to the two nodes of the edge are related cells;

[0016] a data enhancement module configured to perform data enhancement processing on the reference cell graph to obtain a first cell graph and a second cell graph, the first cell graph and the second cell graph being different cell graphs;

[0017] a feature extraction module configured to extract features of the first cell graph by a neural network model to obtain first features of each sample cell, and extract features of the second cell graph by the neural network model to obtain second features of the each sample cell;

[0018] a training module configured to train the neural network model based on the first features and the second features of the each sample cell to obtain a cell feature extraction model, the cell feature extraction model being used to extract cell features of target cells.

[0019] In a possible implementation, the obtaining module is configured to determine the nodes of an initial cell graph based on the material group data of each sample cell, calculate a correlation coefficient between any two sample cells based on the material group data of the any two sample cells, and add an edge between the nodes corresponding to the any two sample cells in the initial cell graph if the correlation coefficient between the any two sample cells is greater than a correlation threshold, to obtain a reference cell graph.

[0020] In a possible implementation, the obtaining module is configured to obtain a third cell graph, the third cell graph being the initial cell graph or obtained by adjusting the initial cell graph at least once; perform feature extraction on the third cell graph by using the neural network model to obtain third features of the sample cells; determine a correlation probability between each two sample cells based on the third features of the sample cells; and adjust the third cell graph based on the correlation probability between each two sample cells to obtain the reference cell graph.

[0021] In a possible implementation, the obtaining module is configured to, for any first correlation probability greater than a first reference probability in the correlation probability between each two sample cells, adjust the reference cell graph by adjusting an edge between nodes of the two sample cells corresponding to the any first correlation probability in the third cell graph.

[0022] In a possible implementation, the obtaining module is configured to, for any second correlation probability less than a second reference probability in the correlation probability between each two sample cells, adjust the reference cell graph by adjusting an edge between nodes of the two sample cells corresponding to the any second correlation probability in the third cell graph.

[0023] In a possible implementation, the data enhancement module is configured to perform mask processing on material group data of the sample cells represented by each node in the reference cell graph to obtain the first cell graph, or perform adjustment processing on edges in the reference cell graph to obtain the first cell graph, or perform mask processing on material group data of the sample cells represented by each node in the reference cell graph and perform adjustment processing on edges in the reference cell graph to obtain the first cell graph.

[0024] In a possible implementation, the training module is configured to determine a plurality of positive sample pairs, any positive sample pair including a first feature of any sample cell and a second feature of the any sample cell; for the any positive sample pair, determine a feature similarity between the first feature of the any sample cell and the second feature of the any sample cell, and take the feature similarity between the first feature of the any sample cell and the second feature of the any sample cell as a similarity of the any positive sample pair; and train the neural network model based on the similarity of each positive sample pair to obtain the cell feature extraction model.

[0025] In a possible implementation, the training module is configured to determine a plurality of negative sample pairs, each negative sample pair including a first feature of one sample cell and a second feature of another sample cell; determine, for each negative sample pair, a feature similarity between the first feature of the one sample cell and the second feature of the another sample cell, and take the feature similarity between the first feature of the one sample cell and the second feature of the another sample cell as a similarity of the negative sample pair; and train the neural network model based on the similarity of each negative sample pair to obtain the cell feature extraction model.

[0026] In a possible implementation, the apparatus further includes:

[0027] The feature extraction module is further configured to perform feature extraction on the reference cell map by using the neural network model to obtain reference features of each sample cell.

[0028] The clustering module is configured to perform clustering processing on the plurality of sample cells based on the reference features of each sample cell to obtain a plurality of sample clustering clusters, and each sample clustering cluster includes at least one sample cell.

[0029] The training module is configured to train the neural network model based on the plurality of sample clustering clusters, the first features of each sample cell, and the second features of each sample cell to obtain the cell feature extraction model.

[0030] In a possible implementation, the training module is configured to, for each sample clustering cluster, extract a cell prototype in the sample clustering cluster from each sample cell included in the sample clustering cluster; determine a loss of the sample clustering cluster based on the cell prototype in the sample clustering cluster and other cells in the sample clustering cluster, the other cells in the sample clustering cluster being sample cells in the sample clustering cluster other than the cell prototype in the sample clustering cluster; and train the neural network model based on the loss of each sample clustering cluster, the first features of each sample cell, and the second features of each sample cell to obtain the cell feature extraction model.

[0031] In a possible implementation, the training module is configured to: for any one sample cell, determine neighboring cells of the any one sample cell based on the reference cell graph, wherein the neighboring cells of the any one sample cell are sample cells in the reference cell graph that have edges with the any one sample cell; determine an updated sample cluster to which the any one sample cell belongs, by using a sample cluster to which the any one sample cell belongs and sample clusters to which the neighboring cells of the any one sample cell belong; determine a distance index of the any one sample cell based on the sample cluster to which the any one sample cell belongs and the updated sample cluster to which the any one sample cell belongs, wherein the distance index of the any one sample cell is used to represent a distance between the any one sample cell and a cluster center of the sample cluster to which the any one sample cell belongs; and select a sample cell with a minimum distance index from sample cells included in the any one sample cluster based on the distance indexes of the sample cells included in the any one sample cluster, and take the sample cell with the minimum distance index as a cell prototype in the any one sample cluster.

[0032] In another aspect, a device for extracting cell features is provided, and the device includes:

[0033] An obtaining module is configured to obtain a target cell graph of a target data set, wherein the target data set includes material group data of a plurality of target cells, and the target cell graph includes a plurality of nodes and a plurality of edges, wherein any one node represents material group data of any one target cell, and any one edge represents a correlation between target cells corresponding to two end nodes of the any one edge.

[0034] A feature extraction module is configured to perform feature extraction on the target cell graph by using a cell feature extraction model to obtain cell features of the target cells, wherein the cell feature extraction model is obtained by training according to the method for training a cell feature extraction model.

[0035] In a possible implementation, the obtaining module is configured to: determine nodes of an original cell graph based on material group data of the target cells; for any two target cells, calculate a correlation coefficient between the any two target cells based on the material group data of the any two target cells; and if the correlation coefficient between the any two target cells is greater than a correlation threshold, add an edge between nodes corresponding to the any two target cells in the original cell graph to obtain the target cell graph.

[0036] In a possible implementation, the acquisition module is configured to acquire a fourth cell map, the fourth cell map being the original cell map or a cell map obtained by adjusting the original cell map at least once; perform feature extraction on the fourth cell map by using the cell feature extraction model to obtain fourth features of the target cells; determine a correlation probability between each two target cells based on the fourth features of the target cells; and adjust the fourth cell map based on the correlation probability between each two target cells to obtain the target cell map.

[0037] In a possible implementation, the apparatus further includes:

[0038] a clustering module configured to perform clustering processing on the target cells based on the cell features of the target cells to obtain target clustering clusters, and each target clustering cluster including at least one target cell;

[0039] a prototype extraction module configured to extract a cell prototype in each target clustering cluster from the target cells included in the target clustering cluster.

[0040] In a possible implementation, the prototype extraction module is configured to, for each target cell, determine neighboring cells of the target cell based on the target cell map, the neighboring cells of the target cell being target cells in the target cell map that have edges with the target cell; determine an updated target clustering cluster to which the target cell belongs, by using a target clustering cluster to which the target cell belongs and target clustering clusters to which the neighboring cells of the target cell belong; determine a distance index of the target cell based on a target clustering cluster to which the target cell belongs and the updated target clustering cluster to which the target cell belongs, the distance index of the target cell being used to represent a distance between the target cell and a clustering center of the target clustering cluster to which the target cell belongs; and select a target cell with a minimum distance index from the target cells included in the target clustering cluster based on the distance index of each target cell included in the target clustering cluster, and take the target cell with the minimum distance index as the cell prototype in the target clustering cluster.

[0041] In another aspect, an electronic device is provided, which includes a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to implement the training method of the cell feature extraction model or the extraction method of the cell features described above.

[0042] In another aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one computer program, the at least one computer program being loaded and executed by a processor to enable an electronic device to implement the training method of the cell feature extraction model or the extraction method of the cell feature.

[0043] In another aspect, a computer program or computer program product is provided, the computer program or computer program product storing at least one computer program, the at least one computer program being loaded and executed by a processor to enable an electronic device to implement the training method of the cell feature extraction model or the extraction method of the cell feature.

[0044] The technical solutions provided in the present application at least bring the following beneficial effects:

[0045] The technical solutions provided in the present application are to describe the material group data of a plurality of sample cells and the information of whether each two sample cells are related by referring to a reference cell map, which enhances the logicality and orderliness of the data and lays a foundation for subsequent data processing. The first cell map and the second cell map are obtained by performing data enhancement processing on the reference cell map, so that the first cell map and the second cell map have noise. When the first cell map and the second cell map are extracted by the neural network model, the model can extract relatively accurate features based on the logicality and orderliness of the data in the presence of noise. Therefore, the accuracy of the first feature and the second feature of the sample cells is high, and when the neural network model is trained based on the first feature and the second feature of each sample cell to obtain the cell feature extraction model, the cell feature extraction model can extract accurate cell features, and the anti-noise performance of the cell feature extraction model is enhanced. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0047] Figure 1 is an implementation environment schematic diagram of a cell feature extraction model training method or a cell feature extraction method provided by an embodiment of the present application;

[0048] Figure 2 is a flowchart of a cell feature extraction model training method provided by an embodiment of the present application;

[0049] Figure 3is a flowchart of a cell feature extraction method provided by an embodiment of the present application;

[0050] Figure 4 is a training schematic diagram of a cell feature extraction model provided by an embodiment of the present application;

[0051] Figure 5 is an extraction schematic diagram of a cell prototype provided by an embodiment of the present application;

[0052] Figure 6 is a structural schematic diagram of a training device of a cell feature extraction model provided by an embodiment of the present application;

[0053] Figure 7 is a structural schematic diagram of a cell feature extraction device provided by an embodiment of the present application;

[0054] Figure 8 is a structural schematic diagram of a terminal device provided by an embodiment of the present application;

[0055] Figure 9 is a structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0057] Figure 1 is an implementation environment schematic diagram of a cell feature extraction model training method or a cell feature extraction method provided by an embodiment of the present application, as shown in the figure, the implementation environment includes a terminal device 101 and a server 102. Wherein, the cell feature extraction model training method or the cell feature extraction method in the embodiments of the present application can be executed by the terminal device 101, or can be executed by the server 102, or can be executed by the terminal device 101 and the server 102 together. Figure 1

[0058] The terminal device 101 can be a smart phone, a game console, a desktop computer, a tablet computer, a laptop computer, a smart television, a smart vehicle device, a smart voice interaction device, a smart home appliance, etc. The server 102 can be a server, or a server cluster composed of multiple servers, or any one of a cloud computing platform and a virtualization center, which is not limited in the embodiments of the present application. The server 102 can be connected with the terminal device 101 through a wired network or a wireless network. The server 102 can have functions of data processing, data storage and data transceiving, which are not limited in the embodiments of the present application. The number of the terminal device 101 and the server 102 is not limited, and can be one or more. ​

[0059] The training method of the cell feature extraction model or the cell feature extraction method provided by the embodiments of the present application can be implemented based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology and application system of using digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0060] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, etc.

[0061] In the field of biotechnology, cells are the basic building blocks of life. In general, the omics data of the cell can be analyzed to know the physiological state and function of the cell. When analyzing the omics data of the cell, the omics data of the cell needs to be extracted to obtain cell features, so as to analyze and process based on the cell features. Based on this, how to train a cell feature extraction model to extract features from the omics data of the cell has become a technical problem to be solved.

[0062] The embodiments of the present application provide a training method of a cell feature extraction model, which can be applied to the above-mentioned implementation environment, and can train a cell feature extraction model to extract features from the omics data of the cell. Figure 2 As shown in the flowchart of the training method of the cell feature extraction model provided by the embodiments of the present application, for the convenience of description, the terminal device 101 or the server 102 executing the training method of the cell feature extraction model in the embodiments of the present application is called an electronic device, and the method can be executed by the electronic device. As shown in the flowchart, the method includes the following steps. Figure 2

[0063] ​In step 201, a reference cell graph of a sample data set is obtained, the sample data set including material group data of a plurality of sample cells, and the reference cell graph including a plurality of nodes and a plurality of edges, any node representing material group data of any sample cell, and any edge representing that sample cells corresponding to two nodes at two ends of the edge are related cells.

[0064] In the embodiment, the sample data set includes material group data of a plurality of sample cells, and the material group data of any sample cell is at least one of proteome data of the sample cell, genome data of the sample cell, and metabolome data of the sample cell. The material group data includes data of a plurality of substances, for example, the proteome data includes data of a plurality of proteins, the genome data includes data of a plurality of genes, and the metabolome data includes data of a plurality of metabolites. Optionally, the sample data set is an expression matrix of cell material groups, the expression matrix being a two-dimensional matrix, each row of the expression matrix representing material group data of a sample cell, and each column of the expression matrix representing data of a substance.

[0065] The initial cell graph of the sample data set can be constructed based on the sample data set, and the initial cell graph is used as the reference cell graph (see the following implementation manner A1), or the initial cell graph is adjusted to obtain the reference cell graph (see the following implementation manner A2). The reference cell graph is a graph structure including a plurality of nodes and a plurality of edges, and one node represents material group data of one sample cell. Therefore, the number of nodes in the reference cell graph is equal to the number of sample cells, and edges can exist or not exist between any two nodes. When edges exist between any two nodes, it is indicated that sample cells corresponding to the two nodes are related. Therefore, one edge represents that sample cells corresponding to two nodes at two ends of the edge are related.

[0066] In the implementation manner A1, the reference cell graph is the initial cell graph. In this case, step 201 includes: determining each node of the initial cell graph based on material group data of each sample cell; for any two sample cells, calculating a correlation coefficient between the two sample cells based on material group data of the two sample cells; and if the correlation coefficient between the two sample cells is greater than a correlation threshold, adding an edge between nodes corresponding to the two sample cells in the initial cell graph to obtain the reference cell graph.

[0067] In the embodiment, the material group data of each sample cell can be used as each node in the initial cell graph, or the material group data of the sample cell is encoded to obtain an encoded representation of the sample cell, and the encoded representation of each sample cell is used as each node in the initial cell graph. The initial cell graph is a graph structure including a plurality of nodes, and one node represents material group data of one sample cell. Therefore, the number of nodes in the initial cell graph is equal to the number of sample cells.

[0068] This application does not limit the method of encoding the material composition data of sample cells. For example, principal component analysis (PCA) can be performed on the material composition data of sample cells to extract principal component data from the material composition data of sample cells and obtain the encoded characterization of sample cells.

[0069] For any two sample cells, the correlation coefficient between them can be calculated using the material composition data or the encoded characterization of these two sample cells, according to a correlation formula. The larger the correlation coefficient, the stronger the correlation between the two sample cells. Therefore, the correlation coefficient between two sample cells characterizes the degree of correlation between them. This application does not limit the correlation formula or the numerical value of the correlation coefficient. For example, the correlation formula may be the Pearson correlation formula, the Spearman correlation formula, etc., and the correlation coefficient may be a value greater than or equal to 0 and less than or equal to 1.

[0070] The correlation coefficient between any two sample cells can be thresholded. Based on the thresholding result, it can be determined whether to add an edge between the nodes corresponding to any two sample cells in the initial cell graph. The initial cell graph includes multiple edges; an edge may or may not exist between any two nodes. When an edge exists between two nodes, it indicates that the sample cells corresponding to these two nodes are correlated. Therefore, an edge represents the correlation between the sample cells corresponding to the two nodes at both ends of the edge.

[0071] Optionally, for any two sample cells, if the correlation coefficient between the two sample cells is greater than the correlation threshold, then the thresholding result between the two sample cells is determined to be a first value (e.g., the first value is 1), and an edge can be added between the nodes corresponding to the two sample cells in the initial cell graph; if the correlation coefficient between the two sample cells is not greater than the correlation threshold, then the thresholding result between the two sample cells is determined to be a second value (e.g., the second value is 0), and in this case, no edge needs to be added between the nodes corresponding to the two sample cells in the initial cell graph.

[0072] This application does not limit the value of the correlation threshold; for example, the correlation threshold is 0.7. It should be noted that a larger correlation threshold results in fewer edges in the initial cell graph, making it appear sparser; conversely, a smaller correlation threshold results in more edges in the initial cell graph, making it appear denser. In this application, the correlation threshold is a hyperparameter, and its value remains unchanged during the training of the cell feature extraction model.

[0073] In implementation manner A2, the reference cell graph is obtained by adjusting the initial cell graph. In this case, step 201 comprises: obtaining a third cell graph, the third cell graph being the initial cell graph or obtained by adjusting the initial cell graph at least once; performing feature extraction on the third cell graph by using the neural network model to obtain third features of the sample cells; determining the correlation probability between each two sample cells based on the third features of the sample cells; and adjusting the third cell graph based on the correlation probability between each two sample cells to obtain the reference cell graph.

[0074] In the embodiments of the present application, the initial cell graph can be taken as the third cell graph, or the third cell graph can be obtained by adjusting the initial cell graph at least once according to implementation manner A2. The third cell graph comprises a plurality of nodes and a plurality of edges, one node representing the substance group data of one sample cell, the number of nodes in the third cell graph being equal to the number of sample cells, and edges possibly existing between any two nodes or not existing between any two nodes. When edges exist between two nodes, it indicates that the sample cells corresponding to the two nodes are correlated, that is, one edge represents that the sample cells corresponding to the two nodes at the ends of the edge are correlated.

[0075] The third cell graph can be input into the neural network model to perform feature extraction on each node in the third cell graph by using the neural network model to obtain node features of each node in the third cell graph. Since any node in the third cell graph represents the substance group data of one sample cell, performing feature extraction on the nodes in the third cell graph is equivalent to performing feature extraction on the substance group data of the sample cells, and the node features of the nodes in the third cell graph are equivalent to the third features of the sample cells. The manner of performing feature extraction on the third cell graph can be seen from the description of step 203 below, and the implementation principles are the same, which will not be described here.

[0076] The embodiments of the present application do not limit the model structure, model size, etc. of the neural network model. For example, the neural network model is an initial network model, in which case the model structure, size, etc. of the neural network model are the same as those of the initial network model. The embodiments of the present application do not limit the model structure, size, etc. of the initial network model. For example, the initial network model is a Graph Convolution Network (GCN) model, and the activation function in the GCN model can be a Rectified Linear Unit (ReLU) function, a Sigmoid function, etc. Alternatively, the neural network model can also be a model obtained by training the initial network model at least once according to steps 201 to 205.

[0077] Corresponding to any two sample cells, the third feature of the two sample cells can be used to calculate the feature similarity between the third features of the two sample cells according to the calculation formula of the feature similarity. The greater the feature similarity between the third features of the two sample cells, the more relevant the two sample cells are; the smaller the feature similarity between the third features of the two sample cells, the less relevant the two sample cells are. Therefore, the feature similarity between the third features of the two sample cells is equivalent to the correlation probability between the two sample cells.

[0078] The embodiments of the present application do not limit the calculation formula of the feature similarity. For example, the calculation formula of the feature similarity is the calculation formula of the cosine similarity, the calculation formula of the Mahalanobis distance, the calculation formula of the Pearson correlation coefficient, etc.

[0079] After calculating the correlation probability between each two sample cells, for any two sample cells, the edge between the nodes corresponding to the two sample cells in the third cell graph can be adjusted by adding, deleting, or remaining unchanged based on the correlation probability between the two sample cells. In this way, the reference cell graph is obtained by adjusting the third cell graph.

[0080] In one possible implementation, if the correlation probability between the two sample cells is greater than the probability threshold, the edge between the nodes corresponding to the two sample cells in the third cell graph exists, and the corresponding adjustment processing is addition or remaining unchanged. If the correlation probability between the two sample cells is not greater than the probability threshold, the edge between the nodes corresponding to the two sample cells in the third cell graph does not exist, and the corresponding adjustment processing is deletion or remaining unchanged.

[0081] In another possible implementation, the third cell graph is adjusted based on the correlation probability between each two sample cells to obtain the reference cell graph, including at least one of the following three implementation manners, which are denoted as implementation manner B1 to implementation manner B3.

[0082] Implementation manner B1, for any first correlation probability greater than the first reference probability in the correlation probability between each two sample cells, the edge between the nodes of the two sample cells corresponding to the first correlation probability in the third cell graph exists, and the reference cell graph is obtained.

[0083] For any two sample cells, if the correlation probability between the two sample cells is greater than the first reference probability, the correlation probability between the two sample cells is the first correlation probability, and an edge between the nodes of the two sample cells in the third cell graph can be adjusted. Embodiments of the present application do not limit the value of the first reference probability. Illustratively, the first reference probability is a set value, or the correlation probabilities between each two sample cells are sorted, and the first quantity of sorted correlation probabilities are taken as the first reference probability. For example, the correlation probabilities between each two sample cells are sorted in descending order, and the N+1th sorted correlation probability is taken as the first reference probability, in which case the first N sorted correlation probabilities are all first correlation probabilities. Therefore, edges exist between the nodes of the two sample cells corresponding to the N first correlation probabilities in the third cell graph.

[0084] In implementation manner B2, for any second correlation probability less than the second reference probability among the correlation probabilities between each two sample cells, an edge between the nodes of the two sample cells corresponding to any second correlation probability in the third cell graph is adjusted to be non-existent, to obtain the reference cell graph.

[0085] For any two sample cells, if the correlation probability between the two sample cells is less than the second reference probability, the correlation probability between the two sample cells is the second correlation probability, and an edge between the nodes of the two sample cells in the third cell graph can be adjusted. Embodiments of the present application do not limit the value of the second reference probability. Illustratively, the value of the second reference probability is less than or equal to the first reference probability. The second reference probability is a set value, or the correlation probabilities between each two sample cells are sorted, and the second quantity of sorted correlation probabilities are taken as the second reference probability. For example, the correlation probabilities between each two sample cells are sorted in ascending order, and the N+1th sorted correlation probability is taken as the second reference probability, in which case the first N sorted correlation probabilities are all second correlation probabilities. Therefore, no edges exist between the nodes of the two sample cells corresponding to the N second correlation probabilities in the third cell graph.

[0086] In implementation manner B3, for any first correlation probability greater than the first reference probability among the correlation probabilities between each two sample cells, an edge between the nodes of the two sample cells corresponding to any first correlation probability in the third cell graph is adjusted to exist, and for any second correlation probability less than the second reference probability among the correlation probabilities between each two sample cells, an edge between the nodes of the two sample cells corresponding to any second correlation probability in the third cell graph is adjusted to be non-existent, to obtain the reference cell graph.

[0087] The description of implementation manner B3 can be found in implementation manner B1 and implementation manner B2, which have similar implementation principles and will not be described here again.

[0088] Optionally, the initial cell graph is taken as a reference cell graph, the neural network model is trained for the first time by using the reference cell graph, then the initial cell graph is adjusted, the adjusted initial cell graph is taken as a reference cell graph, and the neural network model is trained for the second time by using the reference cell graph, and so on. That is, after the neural network model is trained each time by using the reference cell graph, the reference cell graph is adjusted, the adjusted reference cell graph is taken as a reference cell graph for the next time of training, and the neural network model is trained for the next time by using the reference cell graph. By continuously adjusting the reference cell graph, the accuracy of the reference cell graph can be improved, and the noise reduction of the reference cell graph can be realized.

[0089] In step 202, data enhancement processing is performed on the reference cell graph to obtain a first cell graph and a second cell graph.

[0090] In the embodiment of the application, the reference cell graph can be subjected to one time of data enhancement processing to obtain a first cell graph. The reference cell graph is subjected to one time of data enhancement processing again to obtain a second cell graph, or the first cell graph is subjected to one time of data enhancement processing again to obtain a second cell graph. The first cell graph and the second cell graph are different cell graphs, and the determination manner of the first cell graph and the second cell graph is similar, and the determination manner of the first cell graph is taken as an example for description.

[0091] Data enhancement processing is a means for expanding data by introducing noise, cropping data, rotating data (such as rotating an image), and the like. In the embodiment of the application, the data enhancement processing on the reference cell graph can be modifying the reference cell graph to obtain a new cell graph, so that the method obtains a new cell graph based on the reference cell graph on the basis of obtaining the reference cell graph, thereby realizing data expansion. Exemplarily, since the reference cell graph includes nodes and edges, the reference cell graph can be modified to obtain a new cell graph by deleting or adding nodes, deleting or adding edges, and the like, and the new cell graph can be the first cell graph and the second cell graph described above.

[0092] In a possible implementation, the "data enhancement processing is performed on the reference cell graph to obtain a first cell graph" in step 202 includes: performing mask processing on substance group data of sample cells represented by each node in the reference cell graph to obtain the first cell graph; or performing adjustment processing on edges in the reference cell graph to obtain the first cell graph; or performing mask processing on substance group data of sample cells represented by each node in the reference cell graph and performing adjustment processing on edges in the reference cell graph to obtain the first cell graph.

[0093] Optionally, since the nodes in the reference cell graph can be the omics data of the sample cells or the encoded representation of the sample cells, the omics data of the sample cells corresponding to each node in the reference cell graph or the encoded representation of the sample cells can be randomly masked to remove part / all of the data corresponding to part / all of the nodes in the reference cell graph. In this way, the omics data of the sample cells represented by each node in the reference cell graph can be masked, and the cell graph obtained after the masking can be taken as the first cell graph.

[0094] For any two nodes in the reference cell graph, if there is an edge between the two nodes, the edge between the two nodes can be kept or deleted. If there is no edge between the two nodes, an edge can be added between the two nodes. In this way, the edges in the reference cell graph can be kept, deleted, or added, and the cell graph obtained after the adjustment can be taken as the first cell graph.

[0095] In application, on the one hand, the omics data of the sample cells represented by each node in the reference cell graph can be masked, and on the other hand, the edges in the reference cell graph can be adjusted, and the cell graph obtained after the masking and the adjustment can be taken as the first cell graph. The order of the masking and the adjustment is not limited in the embodiments of the present application. That is, the omics data of the sample cells represented by each node in the reference cell graph can be masked first, and then the edges in the reference cell graph can be adjusted, or the edges in the reference cell graph can be adjusted first, and then the omics data of the sample cells represented by each node in the reference cell graph can be masked.

[0096] In step 203, the first cell graph is subjected to feature extraction by a neural network model to obtain the first features of the sample cells.

[0097] In the embodiments of the present application, the sample data set is processed into the first cell graph, and all sample cells and the correlation between each two sample cells can be systematically described in the first cell graph. The first features of all sample cells corresponding to the batch can be obtained by inputting all sample cells into the neural network model at one time and using the neural network model to extract the features of the first cell graph, which avoids batch effect caused by inputting all sample cells into the neural network model in batches.

[0098] The first cell graph can be input into the neural network model, and the neural network model can perform feature extraction on the first cell graph to obtain the first features of the sample cells. Optionally, the neural network model can update the first cell graph multiple times during feature extraction. At each update, for any node in the current first cell graph, the neural network model determines the neighboring nodes of the node in the current first cell graph, and updates the node using the neighboring nodes and the node. In this way, the nodes in the current first cell graph can be updated to obtain an updated first cell graph. The nodes in the last updated first cell graph correspond to the first features of the sample cells.

[0099] Since the neural network model updates the nodes of the sample cells by integrating the neighboring nodes of the nodes during the determination of the first features of the sample cells, the neural network model can automatically fill the omics data of the sample cells represented by the nodes based on the omics data of the sample cells represented by the neighboring nodes to some extent, thereby alleviating the problem of data missing in the omics data of the sample cells represented by the nodes in the first cell graph due to the masking of the omics data of the sample cells represented by the nodes in the reference cell graph, and effectively improving the data quality.

[0100] In one possible implementation, the neural network model includes multiple network layers, and the multiple network layers are used to update the first cell graph multiple times. In this case, one network layer corresponds to one update. For example, the neural network model includes three network layers, and each network layer is used to update the current first cell graph to obtain an updated first cell graph. For example, the current first cell graph is updated using an activation function such as a rectified linear unit (ReLU) function.

[0101] Optionally, the first adjacency matrix can be used to represent the edges in the first cell graph, and the first feature matrix can be used to represent the nodes in the first cell graph, where the first feature matrix includes the omics data of the sample cells corresponding to the nodes in the first cell graph or the encoded representation of the sample cells corresponding to the nodes in the first cell graph.

[0102] The first adjacency matrix is used to represent whether the sample cells corresponding to each pair of nodes in the first cell graph are related. Since an edge represents that the sample cells corresponding to the two nodes at the ends of the edge are related, the first adjacency matrix can describe each edge in the first cell graph. The first adjacency matrix includes multiple rows and multiple columns, and each row and each column represents a sample cell. The first adjacency matrix includes multiple data, and each data can be first data or second data. When the data is first data (for example, 1), it represents that the sample cell corresponding to the row where the first data is located and the sample cell corresponding to the column where the first data is located are related. When the data is second data (for example, 0), it represents that the sample cell corresponding to the row where the second data is located and the sample cell corresponding to the column where the second data is located are not related. The data at the intersection of the row where a sample cell is located and the column where the sample cell is located is first data.

[0103] Exemplarily, the first adjacency matrix and the first feature matrix can be input into the neural network model, and the neural network model can update the first feature matrix multiple times based on the first adjacency matrix and the first feature matrix. Each time of updating outputs an updated first feature matrix, and the last updated first feature matrix includes the first features of the sample cells.

[0104] In step 204, the second cell graph is subjected to feature extraction by using the neural network model, to obtain second features of the sample cells. The implementation of step 204 can refer to the description of step 203, and the implementation principles are similar, which will not be described here.

[0105] In step 205, the neural network model is trained based on the first features and the second features of the sample cells, to obtain a cell feature extraction model, which is used to extract cell features of a target cell.

[0106] In the embodiments of the present application, the loss of the neural network model can be determined based on the first features of the sample cells and the second features of the sample cells. The neural network model is trained based on the loss of the neural network model, to obtain a trained neural network model. If the trained neural network model meets a training end condition, the trained neural network model is taken as the cell feature extraction model. If the trained neural network model does not meet the training end condition, the trained neural network model is taken as the neural network model for the next training, and the neural network model is subjected to the next training in the manner of steps 201 to 205, until the cell feature extraction model is obtained.

[0107] The embodiments of the present application do not limit the training end condition. For example, the training end condition is that the number of training reaches a set number, for example, the number of training is 500. Alternatively, the training end condition is that the gradient descent of the loss of the neural network model is within a set range.

[0108] In the embodiments of the present application, step 205 includes the following three implementation manners, denoted as implementation manner C1 to implementation manner C3.

[0109] In implementation manner C1, a plurality of positive sample pairs are determined, any positive sample pair including the first feature of any sample cell and the second feature of any sample cell; for any positive sample pair, the feature similarity between the first feature of any sample cell and the second feature of any sample cell is determined, and the feature similarity between the first feature of any sample cell and the second feature of any sample cell is taken as the similarity of any positive sample pair; and the neural network model is trained based on the similarity of each positive sample pair to obtain the cell feature extraction model.

[0110] In the embodiments of the present application, the first feature of a sample cell and the second feature of the sample cell are taken as a positive sample pair. Therefore, the positive sample pair is the same sample cell based on the reference cell graph (i.e., the first cell graph and the second cell graph) after two data enhancement processes.

[0111] For any positive sample pair, the feature similarity between the first feature of the sample cell in the positive sample pair and the second feature of the sample cell can be calculated, which can be cosine similarity, distance, relative entropy, etc. The feature similarity is taken as the similarity of the positive sample pair, and the loss of the positive sample pair is obtained based on the similarity of the positive sample pair, for example, any one of power calculation, square root calculation, exponential calculation, etc. is performed on the similarity of the positive sample pair to obtain the loss of the positive sample pair. The sum or average value, etc. of the loss of each positive sample pair is taken as the loss of the neural network model, so as to train the cell feature extraction model by using the loss of the neural network model.

[0112] In the process of training the cell feature extraction model based on the loss of each positive sample pair, the model is continuously adjusted, so that the first feature of the sample cell and the second feature of the sample cell output by the model become more and more similar. That is, in the training process, the similarity between the first feature of the sample cell and the second feature of the sample cell in the positive sample pair is continuously increased, so that the model can output similar features when different data enhancement processes are performed on the material group data of the sample cell, the robustness of the model to different data enhancement processes is improved, and thus the features output by the model can overcome the problem of data noise.

[0113] Optionally, the first feature of each sample cell and the second feature of each sample cell can be projected into a feature space by using the same projection operator to obtain the projected first feature of each sample cell and the projected second feature of each sample cell. The structure of the projection operator is not limited in the embodiments of the present application. For example, the projection operator is a fully connected neural network with at least one layer (e.g., two layers), and the activation function in the projection operator can be a rectified linear unit (ReLU) function, a Sigmoid function, etc. The similarity of the positive sample pair is obtained by calculating the feature similarity between the projected first feature of the sample cell and the projected second feature of the sample cell, and the cell feature extraction model is trained based on the similarity of each positive sample pair.

[0114] Suppose the first feature of the sample cell is u and the second feature of the sample cell is v, the cosine similarity between u and v is cos(u, v). After projecting u and v into a feature space by using the same projection operator g(), the projected first feature of the sample cell g(u) and the projected second feature of the sample cell g(v) can be obtained. At this time, the cosine similarity between g(u) and g(v) is cos(g(u), g(v)). Since the expression ability of cos(g(u), g(v)) is stronger than that of cos(u, v), projecting the first feature of the sample cell and the second feature of the sample cell by using the same projection operator and calculating the similarity between the projected first feature of the sample cell and the projected second feature of the sample cell can better measure the similarity between the first feature of the sample cell and the second feature of the sample cell, and enhance the representation ability of the cell feature extraction model.

[0115] In implementation manner C2, a plurality of negative sample pairs are determined, and each negative sample pair includes the first feature of one sample cell and the second feature of another sample cell. For each negative sample pair, the feature similarity between the first feature of one sample cell and the second feature of another sample cell is determined, and the feature similarity between the first feature of one sample cell and the second feature of another sample cell is taken as the similarity of each negative sample pair. The neural network model is trained based on the similarity of each negative sample pair to obtain the cell feature extraction model.

[0116] In the embodiments of the present application, the first feature of one sample cell and the second feature of another sample cell are taken as a negative sample pair. Therefore, the negative sample pair is different representations of different sample cells based on two data enhanced reference cell maps (i.e., the first cell map and the second cell map).

[0117] For any one negative sample pair, the feature similarity between the first feature of one sample cell in the negative sample pair and the second feature of another sample cell can be calculated, which can be cosine similarity, distance, relative entropy, etc. The feature similarity is taken as the similarity of the negative sample pair, and based on the similarity of the negative sample pair, the loss of the negative sample pair is obtained, for example, any one calculation such as logarithmic calculation, inverse calculation, etc. of the similarity of the negative sample pair, to obtain the loss of the negative sample pair. The sum or average value, etc. of the losses of each negative sample pair is taken as the loss of the neural network model, so as to train the cell feature extraction model by using the loss of the neural network model.

[0118] In the process of training the cell feature extraction model based on the loss of each negative sample pair, the model is continuously adjusted so that the first feature of one sample cell and the second feature of another sample cell output by the model are more and more dissimilar. That is, in the training process, the similarity between the first feature of one sample cell and the second feature of another sample cell in the negative sample pair is continuously reduced, so that the model can output dissimilar features when different data enhancement processing is performed on the substance group data of the two sample cells, the robustness of the model to different data enhancement processing is improved, so that the features output by the model can overcome the problem of data noise.

[0119] Optionally, in addition to being able to take the first feature of one sample cell and the second feature of another sample cell as a negative sample pair, the first feature of one sample cell and the first feature of another sample cell can also be taken as a negative sample pair, the feature similarity between the first feature of one sample cell and the first feature of another sample cell is calculated to obtain the similarity of the negative sample pair, and the second feature of one sample cell and the second feature of another sample cell are taken as a negative sample pair, the feature similarity between the second feature of one sample cell and the second feature of another sample cell is calculated to obtain the similarity of the negative sample pair. The neural network model is trained by the similarity of the above negative sample pair, so that the model can output dissimilar features when the same data enhancement processing is performed on the substance group data of the two sample cells, and the accuracy of the model in feature extraction is improved.

[0120] Optionally, the feature similarity between the first feature of one sample cell in the negative sample pair and the second feature of another sample cell after projection can be calculated to obtain the similarity of the negative sample pair, so as to train the cell feature extraction model based on the similarity of each negative sample pair and enhance the representation ability of the cell feature extraction model.

[0121] In an implementation C3, a plurality of positive sample pairs and a plurality of negative sample pairs are determined, any one positive sample pair including a first feature of any one sample cell and a second feature of any one sample cell, and any one negative sample pair including the first feature of one sample cell and the second feature of another sample cell; for any one positive sample pair, a feature similarity between the first feature of any one sample cell and the second feature of any one sample cell is determined, and the feature similarity between the first feature of any one sample cell and the second feature of any one sample cell is taken as a similarity of any one positive sample pair; for any one negative sample pair, a feature similarity between the first feature of one sample cell and the second feature of another sample cell is determined, and the feature similarity between the first feature of one sample cell and the second feature of another sample cell is taken as a similarity of any one negative sample pair; based on the similarities of the respective positive sample pairs and the similarities of the respective negative sample pairs, the neural network model is trained to obtain the cell feature extraction model.

[0122] In the embodiments of the present application, the determination manner of the positive sample pairs and the similarities of the positive sample pairs can be seen from the description of the implementation C1, and the determination manner of the negative sample pairs and the similarities of the negative sample pairs can be seen from the description of the implementation C2. Based on the similarities of the respective positive sample pairs and the similarities of the respective negative sample pairs, a contrastive loss of the positive and negative sample pairs can be determined, and the contrastive loss of the positive and negative sample pairs is taken as a loss of the neural network model to train the neural network model to obtain the cell feature extraction model.

[0123] Optionally, for any one positive sample pair, a negative sample pair related to the positive sample pair can be determined from the respective negative sample pairs, and the negative sample pair related to the positive sample pair includes the first feature of the sample cell included in the positive sample pair or the second feature of the sample cell included in the positive sample pair. Based on the similarity of the positive sample pair and the similarity of the negative sample pair related to the positive sample pair, a loss of the positive sample pair is determined. For example, the loss of the positive sample pair is determined according to the following formula (1) and formula (2).

[0124] L i = l(u i ,v i ) + l(v i ,u i ) Formula (1)

[0125]

[0126] wherein the i th positive sample pair includes the first feature u i of the i th sample cell and the second feature v i of the i th sample cell. L i represents the loss of the i th positive sample pair. The loss of the i th positive sample pair includes u i and vi The loss between (i.e., l(u)) i ,v i )) and v i with u i The loss between (i.e., l(v)) i ,u i Formula (2) above shows l(u) i ,v i The method of determining l(v) i ,u i The method of determining ) and l(u) i ,v i The method for determining ) is similar and will not be repeated here.

[0127] In formula (2), log represents the logarithmic sign, e represents the exponential sign, and τ is a hyperparameter. cos(u i ,v i ) characterization u i With v i The cosine similarity between them (i.e., the similarity between the i-th positive sample pair). i With v k A negative sample pair can be formed, which is a negative sample pair associated with the i-th positive sample pair, cos(u i ,v k ) characterization u i With v k The cosine similarity between them (i.e., the similarity between negative sample pairs related to the i-th positive sample pair). i with u k Alternatively, a negative sample pair can be formed, which is a negative sample pair associated with the i-th positive sample pair, cos(u i ,u k ) characterization u i with u k The cosine similarity between them (i.e., the similarity between negative sample pairs associated with the i-th positive sample pair).

[0128] After calculating the loss of each positive sample pair, the contrast loss of the positive and negative sample pairs can be determined based on the loss of each positive sample pair. The contrast loss of the positive and negative sample pairs is used as the loss of the neural network model, so as to train the cell feature extraction model using the loss of the neural network model. For example, the contrast loss of the positive and negative sample pairs is determined based on the loss of each positive sample pair according to the formula (3) shown below.

[0129]

[0130] Where L represents the contrast loss of positive and negative sample pairs, and N represents the number of sample cells, which also represents the number of positive sample pairs. icharacterize the loss of the i-th positive sample pair.

[0131] In a possible implementation, after step 201, steps 206 to 207 are further included.

[0132] In step 206, the reference features of the sample cells are extracted from the reference cell image by using the neural network model. The implementation of step 206 can refer to the description of step 203, and the implementation principles are similar, which will not be repeated here.

[0133] In step 207, the sample cells are clustered based on the reference features of the sample cells, to obtain sample clustering clusters, and each sample clustering cluster includes at least one sample cell.

[0134] The embodiments of the present application do not limit the clustering method. For example, for any two sample cells, the feature similarity between the reference features of the two sample cells is calculated, which can be cosine similarity, distance, relative entropy, etc. If the feature similarity between the reference features of the two sample cells is greater than the similarity threshold, it means that the two sample cells are probably similar cells, and the two sample cells can be divided into the same sample clustering cluster. If the feature similarity between the reference features of the two sample cells is not greater than the similarity threshold, it means that the two sample cells are probably dissimilar cells, and the two sample cells can be divided into different sample clustering clusters. In this way, the sample cells can be clustered to obtain sample clustering clusters.

[0135] Optionally, the community discovery algorithm, such as Louvain algorithm or K-means clustering algorithm (K-means), can be used to cluster the sample cells based on the reference features of the sample cells to obtain sample clustering clusters.

[0136] Taking K-means as an example of the community discovery algorithm, in the embodiments of the present application, the initial clustering centers of the sample clustering clusters can be obtained, and the distances between the reference features of the sample cells and the initial clustering centers of the sample clustering clusters can be calculated. For any sample cell, the minimum distance is determined from the distances between the reference features of the sample cell and the initial clustering centers of the sample clustering clusters, and the sample cell is clustered in the sample clustering cluster corresponding to the minimum distance. In this way, the sample cells can be clustered in the sample clustering clusters.

[0137] If the clustering end condition is met, for example, the number of clustering reaches a set number, or the distance between the feature average value calculated based on the reference features of each sample cell in any sample clustering cluster and the initial clustering center of the sample clustering cluster is within a set range, the clustering ends, and each sample clustering cluster is obtained.

[0138] If the clustering end condition is not met, the feature average value is calculated based on the reference features of each sample cell in any sample clustering cluster, the feature average value is taken as the initial clustering center of the sample clustering cluster, each sample cell is clustered in each sample clustering cluster in the clustering manner of the K-means, until the clustering end condition is met, the clustering ends, and each sample clustering cluster is obtained.

[0139] In the case where each sample clustering cluster is obtained, step 205 includes step 2051.

[0140] In step 2051, the neural network model is trained based on each sample clustering cluster, the first features of each sample cell, and the second features of each sample cell, and a cell feature extraction model is obtained.

[0141] In the embodiments of the present application, on the one hand, at least one loss in the loss of the positive sample pair, the loss of the negative sample pair, and the contrast loss of the positive and negative sample pair can be determined based on the first features of each sample cell and the second features of each sample cell, and on the other hand, the loss of each sample clustering cluster can be determined based on each sample clustering cluster. The loss of the neural network model is determined based on the at least one loss and the loss of each sample clustering cluster, so as to train the neural network model by using the loss of the neural network model, and obtain the cell feature extraction model.

[0142] In a possible implementation, step 2051 includes steps D1 to D3.

[0143] In step D1, for any sample clustering cluster, a cell prototype in the sample clustering cluster is extracted from each sample cell included in the sample clustering cluster.

[0144] In the embodiments of the present application, the cell prototype in the sample clustering cluster is one or more sample cells included in the sample clustering cluster, and the distance between the cell prototype in the sample clustering cluster and the clustering center of the sample clustering cluster is smaller than the distance between other cells in the sample clustering cluster and the clustering center of the sample clustering cluster.

[0145] In a possible implementation, for any sample cluster, a feature mean value is calculated based on the reference features of the sample cells in the sample cluster, and the feature mean value is taken as the cluster center of the sample cluster. The distance between the reference feature of each sample cell in the sample cluster and the cluster center of the sample cluster is calculated, and the sample cell with a distance less than a reference distance is taken as the cell prototype in the sample cluster.

[0146] In another possible implementation, the step D1 includes: determining, for any sample cell, the neighboring cells of the sample cell based on the reference cell graph, the neighboring cells of the sample cell being the sample cells in the reference cell graph having edges with the sample cell; determining, by using the sample cluster to which the sample cell belongs and the sample clusters to which the neighboring cells of the sample cell belong, the updated sample cluster to which the sample cell belongs; determining, based on the sample cluster to which the sample cell belongs and the updated sample cluster to which the sample cell belongs, the distance index of the sample cell, the distance index of the sample cell being used to represent the distance between the sample cell and the cluster center of the sample cluster to which the sample cell belongs; and selecting, based on the distance indexes of the sample cells included in any sample cluster, the sample cell with the minimum distance index from the sample cells included in any sample cluster, and taking the sample cell with the minimum distance index as the cell prototype in the sample cluster.

[0147] For any sample cell, the node corresponding to the sample cell in the reference cell graph is determined, and the sample cell corresponding to the node at the other end of the edge with one end being the node of the sample cell is taken as the neighboring cell of the sample cell. For example, in the reference cell graph, the node A is connected with the nodes B, C and D, and the node B is connected with the nodes A and C, the neighboring cells of the sample cell corresponding to the node A are the sample cells corresponding to the nodes B, C and D, and the neighboring cells of the sample cell corresponding to the node B are the sample cells corresponding to the nodes A and C.

[0148] Optionally, the sample cluster to which the sample cell belongs and the sample cluster to which the neighboring cell of the sample cell belongs are the same, and the updated sample cluster to which the sample cell belongs is the sample cluster to which the sample cell belongs. Alternatively, the number of the neighboring cells of the sample cell is at least one, and the sample cluster to which the sample cell belongs and the sample cluster to which any neighboring cell of the sample cell belong can be the same or different, and the sum of the number of the sample cell and the neighboring cells included in each sample cluster is calculated, and the sample cluster with the maximum sum of data is determined as the updated sample cluster to which the sample cell belongs.

[0149] Optionally, a first matrix is used to represent the sample cluster to which the sample cell belongs. The first matrix includes a plurality of data, each data corresponding to a sample cluster, that is, the number of data in the first matrix is equal to the number of sample clusters. Since the sample cell belongs to one of the sample clusters, the data corresponding to the sample cluster to which the sample cell belongs in the first matrix is the third data (for example, 1), and the data corresponding to the sample cluster to which the sample cell does not belong in the first matrix is the fourth data (for example, 0).

[0150] Based on the same principle, a second matrix is used to represent the sample cluster to which the neighboring cell of the sample cell belongs. The first matrix and the second matrix can be weighted and summed to obtain a target matrix, and the updated sample cluster to which the sample cell belongs is determined based on the target matrix.

[0151] Then, according to the cross-entropy function, the cross-entropy of the sample cell is calculated using the sample cluster to which the sample cell belongs and the updated sample cluster to which the sample cell belongs. In the embodiment of the present application, the smaller the cross-entropy of the sample cell, the closer the sample cluster to which the sample cell belongs and the updated sample cluster to which the sample cell belongs, and the more likely the sample cell is the cluster center of the sample cluster to which the sample cell belongs, that is, the farther the sample cell is from the cluster boundary of the sample cluster to which the sample cell belongs. Therefore, the cross-entropy of the sample cell can be used as a distance index of the sample cell to measure the distance between the sample cell and the cluster center of the sample cluster to which the sample cell belongs.

[0152] In the above manner, the distance index of each sample cell can be determined. Since each sample cluster includes at least one sample cell, the sample cell with the smallest distance index can be selected from each sample cell included in any sample cluster, and the sample cell with the smallest distance index is used as the cell prototype in the sample cluster.

[0153] In application, the distance index of each sample cell can be sorted in ascending order, and the sample cells belonging to different sample clusters are selected in order according to the sorting order, and the selected sample cell of any sample cluster is used as the cell prototype in the sample cluster. For example, there are K (K is a positive integer) sample clusters, and the sample cells belonging to the K sample clusters are selected in order according to the sorting order, and the kth (k is a positive integer taking a value of 1 to K) sample cell is used as the cell prototype in the kth sample cluster.

[0154] Optionally, for any sample cluster, a sample cell with the largest number of neighboring cells can be determined from the sample cells included in the sample cluster as a cell prototype in the sample cluster after the neighboring cells of each sample cell are determined.

[0155] Step D2, determining a loss of any sample cluster based on the cell prototype in the sample cluster and other cells in the sample cluster, the other cells in the sample cluster being sample cells in the sample cluster other than the cell prototype in the sample cluster.

[0156] In the embodiments of the present application, the feature distance between the reference feature of the cell prototype in any sample cluster and the reference feature of any other cell in the sample cluster can be calculated, and the loss of the sample cluster can be determined based on the feature distance between the reference feature of the cell prototype in the sample cluster and the reference feature of each other cell in the sample cluster.

[0157] Optionally, the prototype contrast loss is determined according to the formula (4) as shown below, and the prototype contrast loss is obtained based on the average of the losses of each sample cluster.

[0158]

[0159] wherein, The prototype contrast loss is represented by L. N represents the number of sample clusters. K represents the number of sample cells in the i th sample cluster. h i The reference feature of the cell prototype in the i th sample cluster is represented by h. k The reference feature of the k th sample cell in the i th sample cluster is represented by c k, wherein the k th sample cell can be the cell prototype or other cells. τ is a hyperparameter, ∑ is a summation symbol, log is a logarithm symbol, and exp is a symbol of an exponential function with a natural constant e as a base. Wherein, The loss of the i th sample cluster is represented by L i.

[0160] Step D3, training the neural network model based on the loss of each sample cluster, the first feature of each sample cell and the second feature of each sample cell to obtain a cell feature extraction model.

[0161] In the embodiments of the present application, at least one of the loss of the positive sample pair, the loss of the negative sample pair, and the contrast loss of the positive and negative sample pair can be determined based on the first feature of each sample cell and the second feature of each sample cell. The loss of each sample clustering cluster is weighted and calculated to obtain the loss of the neural network model. Alternatively, the prototype contrast loss is determined based on the loss of each sample clustering cluster, and the loss of each sample clustering cluster and the prototype contrast loss are weighted and calculated to obtain the loss of the neural network model. The neural network model is trained using the loss of the neural network model to obtain the cell feature extraction model.

[0162] Since the cell prototype in any sample clustering cluster is close to the clustering center of the sample clustering cluster, the clustering accuracy of the cell prototype in the sample clustering cluster is higher than that of other cells in the sample clustering cluster, and the reference feature of the cell prototype in the sample clustering cluster is more robust. Therefore, when the loss of the sample clustering cluster is calculated using the reference feature of the cell prototype in the sample clustering cluster and the reference feature of other cells in the sample clustering cluster, and the cell feature extraction model is trained using the loss of the sample clustering cluster, the reference feature of other cells in the sample clustering cluster can be close to the reference feature of the cell prototype in the sample clustering cluster, that is, the reference feature of other cells in the sample clustering cluster is enhanced using the reference feature of the cell prototype in the sample clustering cluster, thereby removing the noise in the reference feature of other cells in the sample clustering cluster, so that the reference feature of other cells in the sample clustering cluster has high robustness and the noise problem in the sample data set is alleviated.

[0163] In the embodiments of the present application, the sample data set does not need to be labeled, so the embodiments of the present application are based on unsupervised training of the neural network model based on the unlabeled data set, and a converged cell feature extraction model is obtained, and the cell feature extraction model has good anti-noise performance. During training, an optimizer (for example, an Adam optimizer) can be used for parameter update, and the learning rate is set to 0.001 for training. Through self-supervised training of the neural network model, automatic training is realized in a data-driven manner, avoiding problems caused by human analysis errors.

[0164] It can be understood that the loss of the neural network model in the embodiments of the present application can also be determined based on the target contrast loss, for example, the loss of the neural network model is determined based on the first feature and the second feature of each sample cell, the similarity of each positive sample pair, the similarity of each negative sample pair, at least one of each sample clustering cluster, and the target contrast loss. The determination method of the target contrast loss is introduced below.

[0165] Exemplarily, a reference feature average value is determined based on the reference features of the respective sample cells. The reference feature average value and the reference feature of any sample cell form a reference sample pair, which is a positive sample pair. In addition, a first feature average value is determined based on the first features of the respective sample cells. The first feature average value and the first feature of any sample cell form a first sample pair, which is a negative sample pair. Similarly, a second feature average value is determined based on the second features of the respective sample cells. The second feature average value and the second feature of any sample cell form a second sample pair, which is also a negative sample pair.

[0166] Based on at least one of the respective reference sample pairs, the respective first sample pairs and the respective second sample pairs, a target contrast loss is determined. For example, a similarity between the reference feature average value and the reference feature of the sample cell in any reference sample pair is calculated, obtaining a similarity of the reference sample pair. In the same way, the similarities of the respective first sample pairs and the respective second sample pairs can be determined. Based on at least one of the similarities of the respective reference sample pairs, the similarities of the respective first sample pairs and the similarities of the respective second sample pairs, the target contrast loss is determined.

[0167] By training the neural network model through the target contrast loss, the similarity of the reference sample pairs (i.e. the positive sample pairs) can be narrowed, and the similarities of the first sample pairs and the second sample pairs (i.e. the negative sample pairs) can be widened, thereby improving the accuracy of the features output by the model.

[0168] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the sample data set and the like involved in the present application are obtained under full authorization.

[0169] The above method describes the substance group data of multiple sample cells and the information whether each two sample cells are related by referring to the reference cell graph, enhances the logicality and orderliness of the data, and lays a foundation for subsequent data processing. The first cell graph and the second cell graph are obtained by performing data enhancement processing on the reference cell graph, so that the first cell graph and the second cell graph have noise. When the neural network model extracts features from the first cell graph and the second cell graph, the model can extract relatively accurate features based on the logicality and orderliness of the data in the presence of noise. Therefore, the accuracy of the first features and the second features of the sample cells is high, and when the neural network model is trained based on the first features and the second features of each sample cell to obtain a cell feature extraction model, the cell feature extraction model can extract accurate cell features, thereby enhancing the noise resistance of the cell feature extraction model.

[0170] The embodiment of the present application also provides a cell feature extraction method, which can be applied to the above implementation environment, and can extract features from the substance group data of cells, thereby realizing cell coding. Figure 3 As shown in the flowchart of the cell feature extraction method provided by the embodiment of the present application, for the convenience of description, the terminal device 101 or the server 102 executing the cell feature extraction method in the embodiment of the present application is referred to as an electronic device, and the method can be executed by the electronic device. As shown in the flowchart, the method includes the following steps. Figure 3

[0171] Step 301: Obtain a target cell graph of a target data set.

[0172] The target data set includes substance group data of multiple target cells, and the target cell graph includes multiple nodes and multiple edges. Any node represents the substance group data of any target cell, and any edge represents the correlation between the target cells corresponding to the two end nodes.

[0173] In the embodiment of the present application, the content of the target data set can be described in step 201 about the sample data set, and the content of the target cell graph can be described in step 201 about the sample cell graph. The implementation principles are similar, and will not be described here. The implementation mode of step 301 can include implementation mode E1 and implementation mode E2.

[0174] In implementation mode E1, step 301 includes: determining the nodes of an original cell graph based on the substance group data of each target cell; for any two target cells, calculating the correlation coefficient between the two target cells based on the substance group data of the two target cells; if the correlation coefficient between the two target cells is greater than a correlation threshold, adding an edge between the nodes corresponding to the two target cells in the original cell graph to obtain the target cell graph. ​

[0175] In this embodiment of the application, the content of the original cell diagram can be found in the description of the initial cell diagram above. Therefore, the content of implementation method E1 can be found in the description of implementation method A1. The implementation principles of the two are similar and will not be repeated here.

[0176] In implementation E2, step 301 includes: obtaining a fourth cell map, which is the original cell map or obtained by adjusting the original cell map at least once; extracting features from the fourth cell map using a cell feature extraction model to obtain the fourth features of each target cell; determining the correlation probability between each pair of target cells based on the fourth features of each target cell; and adjusting the fourth cell map based on the correlation probability between each pair of target cells to obtain the target cell map.

[0177] In this embodiment of the application, the content of the fourth cell diagram can be found in the description of the third cell diagram above. Therefore, the content of implementation method E2 can be found in the description of implementation method A2. The implementation principles of the two are similar, and will not be repeated here.

[0178] Step 302: Extract features from the target cell image using a cell feature extraction model to obtain the cell features of each target cell.

[0179] In this embodiment of the application, the cell feature extraction model is based on and Figure 2 The relevant cell feature extraction model was trained using a specific training method. The feature extraction process for the target cell image is described in step 203; the underlying principles are similar and will not be repeated here.

[0180] In one possible implementation, steps 303 to 304 are included after step 302.

[0181] Step 303: Based on the cell characteristics of each target cell, cluster the multiple target cells to obtain each target cluster, and each target cluster includes at least one target cell.

[0182] In this embodiment of the application, the content of clustering multiple target cells can be found in the above description of clustering multiple sample cells, that is, the content of step 303 can be found in the description of step 207. The implementation principles of the two are similar, and will not be repeated here.

[0183] Step 304: For any target cluster, extract the cell prototype from each target cell included in the target cluster.

[0184] Optionally, the step 304 comprises: determining, for any target cell, neighboring cells of the any target cell based on the target cell graph, the neighboring cells of the any target cell being target cells in the target cell graph having edges with the any target cell; determining an updated target cluster to which the any target cell belongs, using the target cluster to which the any target cell belongs and target clusters to which the neighboring cells of the any target cell belong; determining a distance index of the any target cell based on the target cluster to which the any target cell belongs and the updated target cluster to which the any target cell belongs, the distance index of the any target cell being used to represent a distance between the any target cell and a cluster center of the target cluster to which the any target cell belongs; and selecting, from target cells included in the any target cluster, a target cell having a minimum distance index among distance indexes of the target cells included in the any target cluster, and taking the target cell having the minimum distance index as a cell prototype in the any target cluster.

[0185] In the embodiments of the present application, the content of the step 304 can be seen from the description of the step D1, and the implementation principles are similar, which will not be described here again.

[0186] In the embodiments of the present application, extracting the cell prototypes in the target clusters is equivalent to obtaining a plurality of cell prototypes, and each cell prototype belongs to a different cell type. Subsequently, the cell type of any cell can be determined based on the cell features of the cell prototypes, the mass data is deconvoluted, and the like, thereby facilitating the correlation analysis of diseases, and playing an important role in improving the quality of mass data of cells and mining clinical information.

[0187] For example, when the cell type of any cell is determined based on the cell features of the cell prototypes, the mass data of the cell can be feature-extracted by using a cell feature extraction model to obtain the cell features of the cell. The feature similarity between the cell features of the cell and the cell features of each cell prototype is calculated, and the cell type of the cell prototype with the maximum feature similarity is determined as the cell type of the cell.

[0188] When the mass data is deconvoluted based on the cell features of the cell prototypes, the mass data to be detected can be obtained, the mass data to be detected is data obtained by detecting a clinical sample tissue, and represents the expression value of the sum of all cell types in the clinical sample tissue. It can be understood as the sum of the results of each cell type multiplied by its corresponding proportion. Since each cell prototype corresponds to a cell type, the proportion of each cell prototype can be obtained by decomposing the mass data to be detected based on the cell features of the cell prototypes, thereby obtaining different cell types included in the mass data to be detected, and achieving the purpose of deconvoluting the mass data.

[0189] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the target data set involved in the present application is obtained under full authorization.

[0190] The cell feature extraction model in the above method can extract accurate cell features and has strong anti-noise performance. Therefore, the cell features of each target cell obtained by the cell feature extraction model for feature extraction on the target cell graph are highly accurate, which is beneficial to improve the accuracy of subsequent data processing.

[0191] The above describes the training method of the cell feature extraction model and the extraction method of the cell features from the perspective of method steps. Next, the training process of the cell feature extraction model in the present application will be introduced. Please refer to Figure 4 , Figure 4 is a training schematic diagram of a cell feature extraction model provided by the present application.

[0192] In the present application, the sample data set includes proteomic data of multiple sample cells, and the proteomic data of any sample cell includes multiple proteins. Therefore, the sample data set can be represented as a two-dimensional matrix, each row of the two-dimensional matrix corresponds to a sample cell, and each column of the two-dimensional matrix corresponds to a protein. Based on the two-dimensional matrix, an initial cell graph can be constructed, and the initial cell graph is taken as a reference cell graph. The construction method of the initial cell graph can be found in the description of the above implementation manner A1, and will not be repeated here.

[0193] The reference cell graph is subjected to data enhancement processing to obtain a first cell graph and a second cell graph. The data enhancement processing method can be found in the description of step 202, and will not be repeated here.

[0194] The first cell graph can be input into a neural network model, and the neural network model is used to extract features from the first cell graph to obtain first features of each sample cell. A projection operator is used to project the first features of each sample cell into a certain feature space to obtain the projected first features of each sample cell. The method of extracting features from the first cell graph can be found in the description of step 203, and the related content of the projection operator has been introduced above, and will not be repeated here.

[0195] Based on the same processing as the first cell map, the second cell map can be input into another neural network model, and the second cell map is feature extracted by the neural network model to obtain the second features of each sample cell. The second features of each sample cell are projected into the feature space by another projection operator to obtain the second features of each sample cell after projection.

[0196] The first features and the second features of each sample cell after projection are used to determine the contrast loss of the positive and negative sample pairs. The weights of the above two neural network models are shared, which is equivalent to the same neural network model, and the weights of the above two projection operators are shared, which is equivalent to the same projection operator.

[0197] In addition, the reference cell map can be input into the neural network model, and the reference cell map is feature extracted by the neural network model to obtain the reference features of each sample cell. Based on the reference features of each sample cell, the clustering processing is performed on each sample cell to obtain each sample clustering cluster, and the cell prototype in each sample clustering cluster is extracted from the sample cells included in each sample clustering cluster. The cell prototype in each sample clustering cluster is used to determine the prototype contrast loss. The content related to the clustering processing, the extraction method of the cell prototype, and the determination method of the prototype contrast loss are described above, and will not be described here.

[0198] In the embodiments of the present application, the neural network model can be trained based on the contrast loss of the positive and negative sample pairs and the prototype contrast loss to obtain the trained neural network model. If the trained neural network model meets the training end condition, the trained neural network model is used as the cell feature extraction model; if the trained neural network model does not meet the training end condition, the neural network model needs to be trained again based on the reference cell map until the cell feature extraction model is obtained.

[0199] When the neural network model is trained again based on the reference cell map, on the one hand, the trained neural network model can be used as the neural network model for the next training, and on the other hand, the reference cell map can be adjusted to obtain the adjusted reference cell map, which is used as the reference cell map for the next training to train the neural network model for the next time.

[0200] In adjusting the reference cell graph, a correlation probability between each two sample cells can be calculated based on the reference features of the two sample cells, and the reference cell graph is adjusted based on the correlation probability between each two sample cells, for example, an edge is added between the nodes of the two sample cells in the reference cell graph or an edge is deleted in the reference cell graph, to obtain an adjusted reference cell graph. Wherein, the manner of adjusting the reference cell graph can be seen from the description of the implementation manner A2 above, and the third cell graph in the implementation manner A2 is equivalent to the reference cell graph herein, and the implementation principles are the same, which will not be repeated here.

[0201] After the cell feature extraction model is trained, the cell feature extraction model can be used to extract cell features, and subsequent analysis and processing can be performed based on the cell features. For example, please refer to Figure 5 , Figure 5 is a schematic diagram of extracting a cell prototype provided by an embodiment of the present application.

[0202] In the embodiment of the present application, a target data set can be obtained. The target data set includes proteomic data of a plurality of target cells, and the proteomic data of any target cell includes a plurality of proteins. Therefore, the target data set can be represented as a two-dimensional matrix, each row of the two-dimensional matrix corresponds to a target cell, and each column of the two-dimensional matrix corresponds to a protein. Based on the two-dimensional matrix, a target cell graph can be constructed. The target cell graph can be an original cell graph constructed based on the two-dimensional matrix, or can be obtained by adjusting the original cell graph. The construction method of the original cell graph is similar to that of the initial cell graph, and the adjustment method of the original cell graph is similar to that of the reference cell graph, which will not be repeated here.

[0203] The target cell graph can be input into the cell feature extraction model, and the cell feature extraction model can be used to extract features of the target cell graph to obtain cell features of each target cell. Based on the cell features of each target cell, clustering processing is performed on each target cell to obtain each target cluster, and a cell prototype in each target cluster is extracted from the target cells included in each target cluster. The content of the clustering processing and the extraction method of the cell prototype are described above, which will not be repeated here.

[0204] In Figure 4In the training process shown, the protein group data of multiple sample cells and the information of whether each two sample cells are related are described by referring to the reference cell graph, so that the logicality and orderliness of the data are enhanced, and a foundation is laid for subsequent data processing. The first cell graph and the second cell graph are obtained by performing data enhancement processing on the reference cell graph, so that the first cell graph and the second cell graph have noise. When the neural network model is used to extract features from the first cell graph and the second cell graph, the model can extract relatively accurate features based on the logicality and orderliness of the data in the presence of noise. Therefore, the accuracy of the first features and the second features of the sample cells is high, which leads to high accuracy of the contrast loss of the positive and negative sample pairs. The cell feature extraction model is trained by using the contrast loss of the positive and negative sample pairs, so that the model can extract accurate cell features and enhance the anti-noise performance of the cell feature extraction model.

[0205] In addition, since the cell prototypes in the sample clustering cluster are close to the clustering center of the sample clustering cluster, the reference features of the cell prototypes in the sample clustering cluster are more robust. By using the cell prototypes in the sample clustering cluster to determine the prototype contrast loss and training the cell feature extraction model using the prototype contrast loss, the reference features of other cells in the sample clustering cluster can be enhanced using the reference features of the cell prototypes in the sample clustering cluster, thereby removing the noise in the reference features of the other cells in the sample clustering cluster, so that the reference features of the other cells in the sample clustering cluster have high robustness and the noise problem in the sample data set is alleviated.

[0206] Figure 6 The structure of the training device of the cell feature extraction model provided by the embodiment of the application is shown in the structure diagram. Figure 6 The device comprises:

[0207] The acquisition module 601 is configured to acquire a reference cell graph of a sample data set. The sample data set comprises protein group data of multiple sample cells. The reference cell graph comprises multiple nodes and multiple edges. Any node represents the protein group data of any sample cell. Any edge represents that the sample cells corresponding to the two end nodes of the edge are related cells.

[0208] The data enhancement module 602 is configured to perform data enhancement processing on the reference cell graph to obtain a first cell graph and a second cell graph. The first cell graph and the second cell graph are different cell graphs.

[0209] The feature extraction module 603 is configured to extract features from the first cell graph by using a neural network model to obtain first features of each sample cell; and extract features from the second cell graph by using the neural network model to obtain second features of each sample cell.

[0210] The training module 604 is configured to train the neural network model based on the first features and the second features of the sample cells, to obtain a cell feature extraction model, and the cell feature extraction model is configured to extract the cell features of the target cell.

[0211] In a possible implementation, the obtaining module 601 is configured to determine the nodes of the initial cell graph based on the omics data of the sample cells; for any two sample cells, calculate a correlation coefficient between the any two sample cells based on the omics data of the any two sample cells; if the correlation coefficient between the any two sample cells is greater than a correlation threshold, add an edge between the nodes corresponding to the any two sample cells in the initial cell graph, to obtain a reference cell graph.

[0212] In a possible implementation, the obtaining module 601 is configured to obtain a third cell graph, the third cell graph being the initial cell graph or obtained by adjusting the initial cell graph at least once; perform feature extraction on the third cell graph by using the neural network model, to obtain third features of the sample cells; determine a correlation probability between any two sample cells based on the third features of the sample cells; and adjust the third cell graph based on the correlation probability between any two sample cells, to obtain the reference cell graph.

[0213] In a possible implementation, the obtaining module 601 is configured to, for any first correlation probability greater than a first reference probability in the correlation probability between any two sample cells, adjust the nodes of the any two sample cells corresponding to the any first correlation probability in the third cell graph, to obtain the reference cell graph.

[0214] In a possible implementation, the obtaining module 601 is configured to, for any second correlation probability less than a second reference probability in the correlation probability between any two sample cells, adjust the nodes of the any two sample cells corresponding to the any second correlation probability in the third cell graph, to obtain the reference cell graph.

[0215] In a possible implementation, the data enhancement module 602 is configured to perform mask processing on the omics data of the sample cells represented by the nodes in the reference cell graph, to obtain the first cell graph; or perform adjustment processing on the edges in the reference cell graph, to obtain the first cell graph; or perform mask processing on the omics data of the sample cells represented by the nodes in the reference cell graph, and perform adjustment processing on the edges in the reference cell graph, to obtain the first cell graph.

[0216] In a possible implementation, the training module 604 is configured to determine a plurality of positive sample pairs, each of the positive sample pairs including the first feature of any sample cell and the second feature of any sample cell; determine, for each of the positive sample pairs, a feature similarity between the first feature of any sample cell and the second feature of any sample cell, and take the feature similarity between the first feature of any sample cell and the second feature of any sample cell as a similarity of the positive sample pair; and train the neural network model based on the similarity of each of the positive sample pairs to obtain the cell feature extraction model.

[0217] In a possible implementation, the training module 604 is configured to determine a plurality of negative sample pairs, each of the negative sample pairs including the first feature of one sample cell and the second feature of another sample cell; determine, for each of the negative sample pairs, a feature similarity between the first feature of one sample cell and the second feature of another sample cell, and take the feature similarity between the first feature of one sample cell and the second feature of another sample cell as a similarity of the negative sample pair; and train the neural network model based on the similarity of each of the negative sample pairs to obtain the cell feature extraction model.

[0218] In a possible implementation, the apparatus further includes:

[0219] The feature extraction module 603 is further configured to perform feature extraction on the reference cell image by using the neural network model to obtain reference features of each sample cell.

[0220] The clustering module is configured to perform clustering processing on the plurality of sample cells based on the reference features of each sample cell to obtain a plurality of sample clustering clusters, and each of the sample clustering clusters includes at least one sample cell.

[0221] The training module 604 is configured to train the neural network model based on each of the sample clustering clusters, the first features of each sample cell, and the second features of each sample cell to obtain the cell feature extraction model.

[0222] In a possible implementation, the training module 604 is configured to, for each of the sample clustering clusters, extract a cell prototype in the sample clustering cluster from each sample cell included in the sample clustering cluster; determine a loss of the sample clustering cluster based on the cell prototype in the sample clustering cluster and other cells in the sample clustering cluster, the other cells in the sample clustering cluster being sample cells in the sample clustering cluster other than the cell prototype in the sample clustering cluster; and train the neural network model based on the loss of each of the sample clustering clusters, the first features of each sample cell, and the second features of each sample cell to obtain the cell feature extraction model.

[0223] In a possible implementation, the training module 604 is configured to, for any sample cell, determine neighboring cells of the any sample cell based on the reference cell graph, the neighboring cells of the any sample cell being sample cells in the reference cell graph that have edges with the any sample cell; determine an updated sample cluster to which the any sample cell belongs, by using a sample cluster to which the any sample cell belongs and sample clusters to which the neighboring cells of the any sample cell belong; determine a distance index of the any sample cell based on the sample cluster to which the any sample cell belongs and the updated sample cluster to which the any sample cell belongs, the distance index of the any sample cell being used to represent a distance between the any sample cell and a cluster center of the sample cluster to which the any sample cell belongs; and select, from the sample cells included in any sample cluster, a sample cell with a minimum distance index, and take the sample cell with the minimum distance index as a cell prototype in the any sample cluster.

[0224] The device described above describes the material group data of the plurality of sample cells and the information about whether each two sample cells are related by using the reference cell graph, and enhances the logicality and orderliness of the data, and lays a foundation for subsequent data processing. The first cell graph and the second cell graph are obtained by performing data enhancement processing on the reference cell graph, so that the first cell graph and the second cell graph have noise. When the neural network model is used to perform feature extraction on the first cell graph and the second cell graph, the model can extract relatively accurate features based on the logicality and orderliness of the data in the case that the data has noise. Therefore, the first features and the second features of the sample cells are relatively accurate, and when the neural network model is trained based on the first features and the second features of the sample cells to obtain the cell feature extraction model, the cell feature extraction model can extract accurate cell features, and the anti-noise performance of the cell feature extraction model is enhanced.

[0225] It should be understood that the above Figure 6 When the device provided in the present application is implemented, the above-mentioned division of functional modules is used as an example, and in actual applications, the above-mentioned functions can be completed by different functional modules, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be described here.

[0226] Figure 7 As shown in FIG. 6, which is a structural schematic diagram of a cell feature extraction device provided in an embodiment of the present application, as shown in FIG. 7, the device includes: Figure 7 As shown in FIG. 6, which is a structural schematic diagram of a cell feature extraction device provided in an embodiment of the present application, as shown in FIG. 7, the device includes:

[0227] The acquisition module 701 is configured to acquire a target cell graph of a target data set, the target data set comprising metabolomic data of a plurality of target cells, and the target cell graph comprising a plurality of nodes and a plurality of edges, any node representing metabolomic data of any target cell, and any edge representing correlation between target cells corresponding to a pair of nodes at two ends of the edge.

[0228] The feature extraction module 702 is configured to perform feature extraction on the target cell graph by using a cell feature extraction model to obtain cell features of the target cells, the cell feature extraction model being trained according to the training method of the cell feature extraction model.

[0229] In a possible implementation, the acquisition module 701 is configured to determine the nodes of the original cell graph based on the metabolomic data of the target cells; for any two target cells, calculate a correlation coefficient between the two target cells based on the metabolomic data of the two target cells; if the correlation coefficient between the two target cells is greater than a correlation threshold, add an edge between the nodes corresponding to the two target cells in the original cell graph to obtain the target cell graph.

[0230] In a possible implementation, the acquisition module 701 is configured to acquire a fourth cell graph, the fourth cell graph being the original cell graph or obtained by adjusting the original cell graph at least once; perform feature extraction on the fourth cell graph by using the cell feature extraction model to obtain fourth features of the target cells; determine a correlation probability between any two target cells based on the fourth features of the target cells; and adjust the fourth cell graph based on the correlation probability between any two target cells to obtain the target cell graph.

[0231] In a possible implementation, the apparatus further comprises:

[0232] The clustering module is configured to perform clustering processing on the target cells based on the cell features of the target cells to obtain target clustering clusters, any target clustering cluster comprising at least one target cell.

[0233] The prototype extraction module is configured to, for any target clustering cluster, extract a cell prototype in the target clustering cluster from the target cells included in the target clustering cluster.

[0234] In a possible implementation, the prototype extraction module is configured to: determine, for each target cell, neighboring cells of the target cell based on the target cell graph, the neighboring cells of the target cell being target cells in the target cell graph having edges with the target cell; determine an updated target cluster to which the target cell belongs, based on a target cluster to which the target cell belongs and target clusters to which the neighboring cells of the target cell belong; determine a distance index of the target cell, based on the target cluster to which the target cell belongs and the updated target cluster to which the target cell belongs, the distance index of the target cell being used to represent a distance between the target cell and a cluster center of the target cluster to which the target cell belongs; and select, from target cells included in each target cluster, a target cell having a minimum distance index, as a cell prototype in the target cluster.

[0235] The cell feature extraction model in the device can extract accurate cell features and has strong anti-noise performance. Therefore, the cell features of each target cell obtained by performing feature extraction on the target cell graph by the cell feature extraction model are accurate, which is beneficial to improving the accuracy of subsequent data processing.

[0236] It should be understood that the above Figure 7 When the device provided by the present application implements its functions, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to the needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the device and method embodiments provided by the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.

[0237] Figure 8 A structure block diagram of a terminal device 800 provided by an example embodiment of the present application is shown. The terminal device 800 includes a processor 801 and a memory 802.

[0238] The processor 801 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 801 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 801 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 801 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 801 can further include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0239] The memory 802 can include one or more computer-readable storage media that can be non-transitory. The memory 802 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one computer program for being executed by the processor 801 to implement the training method of the cell feature extraction model or the extraction method of the cell feature provided by the method embodiments in the present application.

[0240] In some embodiments, the terminal device 800 can also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 803 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera component 806, an audio circuit 807, and a power supply 808.

[0241] The peripheral interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802 and the peripheral interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802 and the peripheral interface 803 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.

[0242] The radio frequency circuit 804 is used to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 can also include NFC (Near Field Communication) related circuitry, and the present application is not limited in this regard.

[0243] The display screen 805 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 is further configured to capture touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. In this case, the display screen 805 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 805 can be one, disposed on the front panel of the terminal device 800; in other embodiments, the display screen 805 can be at least two, respectively disposed on different surfaces of the terminal device 800 or in a folding design; in other embodiments, the display screen 805 can be a flexible display screen, disposed on a curved surface or a folding surface of the terminal device 800. Even, the display screen 805 can also be disposed in an irregular shape, i.e., a special-shaped screen. The display screen 805 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.

[0244] The camera assembly 806 is configured to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 806 can further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0245] The audio circuit 807 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 801 for processing or to the radio frequency circuit 804 to realize voice communication. The microphone can be multiple for the purpose of stereo sound collection or noise reduction, and arranged at different parts of the terminal device 800. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert an electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker can be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert an electrical signal into a sound wave audible to humans, but also convert an electrical signal into a sound wave inaudible to humans for ranging purposes. In some embodiments, the audio circuit 807 can also include a headphone jack.

[0246] The power supply 808 is used to supply power to various components in the terminal device 800. The power supply 808 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0247] In some embodiments, the terminal device 800 further includes one or more sensors 809. The one or more sensors 809 include, but are not limited to, an acceleration sensor 811, a gyroscope sensor 812, a pressure sensor 813, an optical sensor 814, and a proximity sensor 815.

[0248] The acceleration sensor 811 can detect the acceleration in three coordinate axes of the coordinate system established by the terminal device 800. For example, the acceleration sensor 811 can be used to detect the components of gravitational acceleration in three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 811. The acceleration sensor 811 can also be used for game or user motion data collection.

[0249] The gyroscope sensor 812 can detect the body direction and rotation angle of the terminal device 800, and the gyroscope sensor 812 can collect 3D actions of the user on the terminal device 800 in cooperation with the acceleration sensor 811. The processor 801 can realize the following functions according to the data collected by the gyroscope sensor 812: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.

[0250] The pressure sensor 813 can be disposed on the side bezel of the terminal device 800 and / or on the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side bezel of the terminal device 800, it can detect the user's grip signal on the terminal device 800, and the processor 801 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0251] An optical sensor 814 is used to collect ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 based on the ambient light intensity collected by the optical sensor 814. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera assembly 806 based on the ambient light intensity collected by the optical sensor 814.

[0252] The proximity sensor 815, also known as a distance sensor, is typically located on the front panel of the terminal device 800. The proximity sensor 815 is used to detect the distance between the user and the front of the terminal device 800. In one embodiment, when the proximity sensor 815 detects that the distance between the user and the front of the terminal device 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from a screen-on state to a screen-off state; when the proximity sensor 815 detects that the distance between the user and the front of the terminal device 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from a screen-off state to a screen-on state.

[0253] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the terminal device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0254] Figure 9A structural diagram of a server provided in the embodiments of the present application is shown in FIG. 9. The server 900 can have great differences due to different configurations or performances, and can include one or more processors 901 and one or more memories 902. The one or more memories 902 store at least one computer program, which is loaded and executed by the one or more processors 901 to implement the training method of the cell feature extraction model or the extraction method of the cell feature provided in any of the above method embodiments. The processor 901 is exemplarily a CPU. Of course, the server 900 can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for implementing device functions, which are not described herein.

[0255] In exemplary embodiments, a computer readable storage medium is also provided, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above training method of the cell feature extraction model or the extraction method of the cell feature.

[0256] Optionally, the above computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0257] In exemplary embodiments, a computer program or computer program product is also provided, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to enable an electronic device to implement any of the above training method of the cell feature extraction model or the extraction method of the cell feature.

[0258] It should be understood that “multiple” mentioned herein refers to two or more. “And / or” describes the association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The character “ / ” generally represents that the associated objects before and after it are in an “or” relationship.

[0259] The above sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0260] The above only describes exemplary embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for a cell feature extraction model, characterized in that, The method includes: Obtain a reference cell diagram of a sample dataset, wherein the sample dataset includes material composition data of multiple sample cells, and the reference cell diagram includes multiple nodes and multiple edges, wherein any node represents the material composition data of any sample cell, and any edge represents that the sample cells corresponding to the nodes at both ends of any edge are related cells. The reference cell map is subjected to data augmentation processing to obtain a first cell map and a second cell map, wherein the first cell map and the second cell map are different cell maps; The first feature of each sample cell is obtained by extracting features from the first cell map using a neural network model; The second cell map is feature extracted using the neural network model to obtain the second features of each sample cell; The neural network model is used to extract features from the reference cell map to obtain reference features for each sample cell. Based on the reference features of each sample cell, the multiple sample cells are clustered to obtain each sample cluster, and each sample cluster includes at least one sample cell. Based on the clusters of each sample and the first and second features of each sample cell, the neural network model is trained to obtain a cell feature extraction model, which is used to extract cell features of the target cell.

2. The method according to claim 1, characterized in that, The reference cell map for obtaining the sample dataset includes: The nodes of the initial cell map are determined based on the material composition data of each sample cell. For any two sample cells, calculate the correlation coefficient between the two sample cells based on the material composition data of the two sample cells; If the correlation coefficient between any two sample cells is greater than the correlation threshold, then an edge is added between the nodes corresponding to the two sample cells in the initial cell graph to obtain a reference cell graph.

3. The method according to claim 1, characterized in that, The reference cell map for obtaining the sample dataset includes: Obtain a third cell map, which is the initial cell map or obtained by adjusting the initial cell map at least once; The neural network model is used to extract features from the third cell map to obtain the third features of each sample cell. Based on the third feature of each sample cell, the correlation probability between each pair of sample cells is determined. Based on the correlation probability between each pair of sample cells, the third cell map is adjusted to obtain the reference cell map.

4. The method according to claim 3, characterized in that, The adjustment of the third cell map based on the correlation probability between every two sample cells to obtain the reference cell map includes: For any first correlation probability greater than the first reference probability among the correlation probabilities between any two sample cells, adjust the nodes of the two sample cells corresponding to any first correlation probability in the third cell graph to have an edge, and obtain the reference cell graph.

5. The method according to claim 3, characterized in that, The adjustment of the third cell map based on the correlation probability between every two sample cells to obtain the reference cell map includes: For any second correlation probability that is less than the second reference probability among the correlation probabilities between any two sample cells, adjust the nodes of the two sample cells corresponding to any second correlation probability in the third cell graph so that there is no edge between them, and obtain the reference cell graph.

6. The method according to claim 1, characterized in that, The reference cell map is subjected to data augmentation processing to obtain a first cell map, including: The material composition data of the sample cells represented by each node in the reference cell diagram are masked to obtain the first cell diagram; Alternatively, the edges in the reference cell diagram can be adjusted to obtain the first cell diagram. Alternatively, the material composition data of the sample cells represented by each node in the reference cell diagram can be masked, and the edges in the reference cell diagram can be adjusted to obtain the first cell diagram.

7. The method according to claim 1, characterized in that, The process of training the neural network model based on the first and second features of each sample cell to obtain a cell feature extraction model includes: Multiple positive sample pairs are identified, wherein each positive sample pair includes a first feature of a sample cell and a second feature of the sample cell; For any positive sample pair, determine the feature similarity between the first feature and the second feature of any sample cell, and use the feature similarity between the first feature and the second feature of any sample cell as the similarity of the positive sample pair. The neural network model is trained based on the similarity of each positive sample pair to obtain a cell feature extraction model.

8. The method according to claim 1, characterized in that, The process of training the neural network model based on the first and second features of each sample cell to obtain a cell feature extraction model includes: Identify multiple negative sample pairs, where each negative sample pair includes a first feature of one sample cell and a second feature of another sample cell; For any negative sample pair, determine the feature similarity between the first feature of the one sample cell and the second feature of the other sample cell, and use the feature similarity between the first feature of the one sample cell and the second feature of the other sample cell as the similarity of the negative sample pair. The neural network model is trained based on the similarity of each negative sample pair to obtain a cell feature extraction model.

9. The method according to claim 1, characterized in that, The process of training the neural network model based on the clusters of each sample, the first feature of each sample cell, and the second feature of each sample cell to obtain a cell feature extraction model includes: For any sample cluster, extract the cell prototype from each sample cell included in the sample cluster; Based on the cell prototype in any sample cluster and the other cells in any sample cluster, determine the loss of any sample cluster, wherein the other cells in any sample cluster are sample cells in any sample cluster other than the cell prototype in any sample cluster. The neural network model is trained based on the loss of each sample cluster, the first feature of each sample cell, and the second feature of each sample cell to obtain a cell feature extraction model.

10. The method according to claim 9, characterized in that, The step of extracting cell prototypes from each sample cell included in any sample cluster includes: For any sample cell, the neighboring cells of the sample cell are determined based on the reference cell diagram. The neighboring cells of the sample cell are the sample cells in the reference cell diagram that have an edge with the sample cell. Using the sample cluster to which any sample cell belongs and the sample cluster to which the neighboring cells of any sample cell belong, the updated sample cluster to which the sample cell belongs is determined; Based on the sample cluster to which any sample cell belongs and the updated sample cluster to which any sample cell belongs, a distance index for any sample cell is determined. The distance index for any sample cell is used to characterize the distance between any sample cell and the cluster center of the sample cluster to which the sample cell belongs. Based on the distance index of each sample cell included in any sample cluster, the sample cell with the smallest distance index is selected from each sample cell included in any sample cluster, and the sample cell with the smallest distance index is used as the cell prototype in any sample cluster.

11. A method for extracting cell features, characterized in that, The method includes: Obtain the target cell graph of the target dataset, wherein the target dataset includes material composition data of multiple target cells, and the target cell graph includes multiple nodes and multiple edges, wherein any node represents the material composition data of any target cell, and any edge represents the target cell related to the nodes at both ends of the edge; The target cell image is subjected to feature extraction by a cell feature extraction model to obtain the cell features of each target cell. The cell feature extraction model is trained according to the training method of the cell feature extraction model according to any one of claims 1 to 10.

12. The method according to claim 11, characterized in that, The acquisition of the target cell map of the target dataset includes: The nodes of the original cell map are determined based on the material composition data of each target cell; For any two target cells, the correlation coefficient between the two target cells is calculated based on the material composition data of the two target cells. If the correlation coefficient between any two target cells is greater than the correlation threshold, then an edge is added between the nodes corresponding to any two target cells in the original cell graph to obtain the target cell graph.

13. The method according to claim 11, characterized in that, The acquisition of the target cell map of the target dataset includes: A fourth cell image is obtained, which is either the original cell image or obtained by adjusting the original cell image at least once. The fourth cell map is extracted using the cell feature extraction model to obtain the fourth feature of each target cell; Based on the fourth feature of each target cell, the correlation probability between any two target cells is determined. Based on the correlation probability between each pair of target cells, the fourth cell map is adjusted to obtain the target cell map.

14. The method according to claim 11, characterized in that, The method further includes: Based on the cellular characteristics of each target cell, the multiple target cells are clustered to obtain each target cluster, and each target cluster includes at least one target cell. For any target cluster, extract the cell prototype from each target cell included in the target cluster.

15. The method according to claim 14, characterized in that, The step of extracting cell prototypes from each target cell included in any target cluster includes: For any target cell, the neighboring cells of the target cell are determined based on the target cell graph. The neighboring cells of the target cell are the target cells in the target cell graph that have an edge with the target cell. Using the target cluster to which any target cell belongs and the target cluster to which the neighboring cells of any target cell belong, determine the updated target cluster to which any target cell belongs; Based on the target cluster to which any target cell belongs and the updated target cluster to which any target cell belongs, a distance index for any target cell is determined. The distance index for any target cell is used to characterize the distance between any target cell and the cluster center of the target cluster to which the target cell belongs. Based on the distance index of each target cell included in any target cluster, select the target cell with the smallest distance index from each target cell included in any target cluster, and use the target cell with the smallest distance index as the cell prototype in any target cluster.

16. A training device for a cell feature extraction model, characterized in that, The device includes: The acquisition module is used to acquire a reference cell diagram of a sample dataset. The sample dataset includes material composition data of multiple sample cells. The reference cell diagram includes multiple nodes and multiple edges. Any node represents the material composition data of any sample cell, and any edge represents that the sample cells corresponding to the nodes at both ends of any edge are related cells. The data augmentation module is used to perform data augmentation processing on the reference cell map to obtain a first cell map and a second cell map, wherein the first cell map and the second cell map are different cell maps; The feature extraction module is used to extract features from the first cell map using a neural network model to obtain the first features of each sample cell; and to extract features from the second cell map using the neural network model to obtain the second features of each sample cell. The feature extraction module is also used to extract features from the reference cell map through the neural network model to obtain reference features for each sample cell; The clustering module is used to perform clustering processing on the multiple sample cells based on the reference features of each sample cell to obtain each sample cluster, wherein each sample cluster includes at least one sample cell. The training module is used to train the neural network model based on the clusters of each sample and the first and second features of each sample cell to obtain a cell feature extraction model, which is used to extract cell features of the target cell.

17. The apparatus according to claim 16, characterized in that, The acquisition module is used to determine each node of the initial cell map based on the material composition data of each sample cell; for any two sample cells, the correlation coefficient between the two sample cells is calculated based on the material composition data of the two sample cells; if the correlation coefficient between the two sample cells is greater than the correlation threshold, an edge is added between the nodes corresponding to the two sample cells in the initial cell map to obtain a reference cell map.

18. The apparatus according to claim 16, characterized in that, The acquisition module is used to acquire a third cell map, which is an initial cell map or obtained by adjusting the initial cell map at least once; to extract features from the third cell map using the neural network model to obtain the third features of each sample cell; to determine the correlation probability between each pair of sample cells based on the third features of each sample cell; and to adjust the third cell map based on the correlation probability between each pair of sample cells to obtain the reference cell map.

19. The apparatus according to claim 18, characterized in that, The acquisition module is used to adjust the nodes of the two sample cells corresponding to any first correlation probability that is greater than the first reference probability in the correlation probability between any two sample cells to obtain the reference cell graph.

20. The apparatus according to claim 18, characterized in that, The acquisition module is used to adjust the nodes of the two sample cells corresponding to any second correlation probability that is less than the second reference probability in the correlation probability between each pair of sample cells to ensure that there is no edge between them, so as to obtain the reference cell graph.

21. The apparatus according to claim 16, characterized in that, The data enhancement module is used to mask the material composition data of the sample cells represented by each node in the reference cell diagram to obtain the first cell diagram; or, to adjust the edges in the reference cell diagram to obtain the first cell diagram; or, to mask the material composition data of the sample cells represented by each node in the reference cell diagram and to adjust the edges in the reference cell diagram to obtain the first cell diagram.

22. The apparatus according to claim 16, characterized in that, The training module is used to determine multiple positive sample pairs, each positive sample pair including a first feature and a second feature of a sample cell; for each positive sample pair, the feature similarity between the first feature and the second feature of the sample cell is determined, and the feature similarity between the first feature and the second feature of the sample cell is used as the similarity of the positive sample pair; the neural network model is trained based on the similarity of each positive sample pair to obtain a cell feature extraction model.

23. The apparatus according to claim 16, characterized in that, The training module is used to determine multiple negative sample pairs, where each negative sample pair includes a first feature of one sample cell and a second feature of another sample cell. For each negative sample pair, the feature similarity between the first feature of one sample cell and the second feature of another sample cell is determined, and the feature similarity between the first feature of one sample cell and the second feature of another sample cell is used as the similarity of the negative sample pair. The neural network model is trained based on the similarity of each negative sample pair to obtain a cell feature extraction model.

24. The apparatus according to claim 16, characterized in that, The training module is used to extract cell prototypes from each sample cell included in any sample cluster for any sample cluster. Based on the cell prototype in any sample cluster and the other cells in any sample cluster, the loss of any sample cluster is determined. The other cells in any sample cluster are sample cells in any sample cluster other than the cell prototype in the sample cluster. Based on the loss of each sample cluster, the first feature of each sample cell, and the second feature of each sample cell, the neural network model is trained to obtain a cell feature extraction model.

25. The apparatus according to claim 24, characterized in that, The training module is used to determine the neighboring cells of any sample cell based on the reference cell diagram, wherein the neighboring cells of any sample cell are sample cells in the reference cell diagram that have an edge with the sample cell. Using the sample cluster to which any given sample cell belongs and the sample clusters to which its neighboring cells belong, an updated sample cluster to which the given sample cell belongs is determined. Based on the sample cluster to which the given sample cell belongs and the updated sample cluster to which the given sample cell belongs, a distance index for the given sample cell is determined. The distance index of the given sample cell is used to characterize the distance between the given sample cell and the cluster center of the sample cluster to which the given sample cell belongs. Based on the distance indices of each sample cell included in the given sample cluster, the sample cell with the smallest distance index is selected from the sample cells included in the given sample cluster, and the sample cell with the smallest distance index is used as the cell prototype in the given sample cluster.

26. A device for extracting cell characteristics, characterized in that, The device includes: The acquisition module is used to acquire the target cell graph of the target dataset. The target dataset includes material group data of multiple target cells. The target cell graph includes multiple nodes and multiple edges. Any node represents the material group data of any target cell, and any edge represents the target cell related to the nodes at both ends of the edge. The feature extraction module is used to extract features from the target cell map using a cell feature extraction model to obtain the cell features of each target cell. The cell feature extraction model is trained according to the training method of the cell feature extraction model according to any one of claims 1 to 10.

27. The apparatus according to claim 26, characterized in that, The acquisition module is used to determine each node of the original cell graph based on the material composition data of each target cell; for any two target cells, the correlation coefficient between the two target cells is calculated based on the material composition data of the two target cells; if the correlation coefficient between the two target cells is greater than the correlation threshold, an edge is added between the nodes corresponding to the two target cells in the original cell graph to obtain the target cell graph.

28. The apparatus according to claim 26, characterized in that, The acquisition module is used to acquire a fourth cell image, which is the original cell image or obtained by adjusting the original cell image at least once; to extract features from the fourth cell image using the cell feature extraction model to obtain the fourth features of each target cell; and to determine the correlation probability between each pair of target cells based on the fourth features of each target cell. Based on the correlation probability between each pair of target cells, the fourth cell map is adjusted to obtain the target cell map.

29. The apparatus according to claim 26, characterized in that, The device further includes: The clustering module is used to perform clustering processing on the multiple target cells based on the cell characteristics of each target cell to obtain various target clusters, wherein each target cluster includes at least one target cell; The prototype extraction module is used to extract cell prototypes from each target cell included in the target cluster for any given target cluster.

30. The apparatus according to claim 29, characterized in that, The prototype extraction module is used to determine the neighboring cells of any target cell based on the target cell graph, wherein the neighboring cells of any target cell are target cells in the target cell graph that have an edge with the target cell. Using the target cluster to which any target cell belongs and the target cluster to which its neighboring cells belong, an updated target cluster to which the target cell belongs is determined. Based on the target cluster to which the target cell belongs and the updated target cluster to which the target cell belongs, a distance index for the target cell is determined. The distance index of the target cell is used to characterize the distance between the target cell and the cluster center of the target cluster to which the target cell belongs. Based on the distance indices of each target cell included in the target cluster, a target cell with the smallest distance index is selected from the target cells included in the target cluster, and the target cell with the smallest distance index is used as the cell prototype in the target cluster.

31. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to implement the training method of the cell feature extraction model as described in any one of claims 1 to 10 or to implement the cell feature extraction method as described in any one of claims 11 to 15.

32. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable an electronic device to implement the training method of the cell feature extraction model as described in any one of claims 1 to 10 or the cell feature extraction method as described in any one of claims 11 to 15.

33. A computer program product, characterized in that, The computer program product stores at least one computer program, which is loaded and executed by a processor to enable the electronic device to implement the training method of the cell feature extraction model as described in any one of claims 1 to 10 or the cell feature extraction method as described in any one of claims 11 to 15.

Citation Information

Patent Citations

  • Omics data processing method and device, electronic equipment and storage medium

    CN114429786A