Cancer tissue classification method, device, electronic device and storage medium

By constructing the gene feature matrix and gene adjacency matrix and using graph convolutional neural network to process neighbor relationships in different orders, the limitations and low accuracy of cancer tissue classification in the existing technology are solved, and a more efficient and comprehensive classification effect is achieved.

CN114154557BActive Publication Date: 2025-06-10CENTRAL UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111313250.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-06-10
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

The prior art has limitations and one-sidedness when categorizing cancer tissues, and the classification accuracy is low, especially ignoring the characteristics of the nodes themselves and the node relationship between neighbor distances in different orders.

Method used

By obtaining the gene data of the tissue to be examined, a gene feature matrix and gene adjacency matrix are constructed, and inputting them into the graph convolution neural network for processing. The multi-layer graph convolution network layer aggregates the neighbor relationships of different orders, and finally inputs the classifier for classification.

Benefits of technology

It improves the accuracy and comprehensiveness of cancer tissue classification, can handle gene network characteristics more effectively, and enhances the performance and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154557B_ABST
    Figure CN114154557B_ABST
Patent Text Reader

Abstract

The present invention provides a cancer tissue classification method, apparatus, electronic device and storage medium. The method includes: obtaining gene data corresponding to a tissue set to be examined; the tissue to be examined includes a plurality of tissue samples to be examined; determining a gene feature matrix and a gene adjacency matrix according to the gene data; inputting the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain a plurality of graph convolutional network layers; aggregating the plurality of graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result; inputting the aggregation result into a classifier for classification to obtain a diagnosis result. This solution has a high accuracy rate for cancer tissue classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gene recognition, and particularly relates to a cancer tissue classification method, device, electronic device and storage medium. Background Art

[0002] The classification research of cancer diseases is a complex problem. With the rapid development of high-throughput sequencing, the roles of gene expression profiles and gene network technologies have become increasingly prominent, and it also provides strong support for the diagnosis and decision-making of cancer patients. On the one hand, gene expression profiles can classify diseases for samples. On the other hand, the role of gene networks lies in the description of the relationships between genes. If a certain gene mutates, it will amplify the impact through the gene network. In past studies, researchers often only targeted some genes, while gene networks can involve the regulatory relationships between multiple genes and have an amplification effect. Therefore, considering multi-order neighbor gene networks and combining and analyzing the information of different-order neighbor gene networks can help us classify diseases.

[0003] From the perspective of network methods, the common related effects between nodes in a network are network centrality indicators. Since the network centrality method only evaluates the importance between nodes based on the positions of the nodes in the network, this method ignores the characteristics of the nodes themselves and cannot consider the relationships between nodes at different-order neighbor distances. Currently, many machine learning algorithms are also used in disease detection work. Classic machine learning algorithms such as logistic regression, support vector machine (SVM) classification algorithms, random forests, and feedforward neural networks directly classify and predict samples based on the gene expression profiles of the samples, and these methods cannot further process network features.

[0004] When using the existing single gene expression profile and network centrality method for cancer tissue classification, there are limitations and one-sidedness, and the classification accuracy is low. Summary of the Invention

[0005] The purpose of the embodiments of this specification is to provide a cancer tissue classification method, device, electronic device and storage medium.

[0006] To solve the above technical problems, the embodiments of the present application are implemented as follows:

[0007] In a first aspect, the present application provides a cancer tissue classification method, which includes:

[0008] Obtain gene data corresponding to a set of tissues to be examined; the tissues to be examined include several tissue samples to be examined;

[0009] Determine a gene feature matrix and a gene adjacency matrix according to the gene data;

[0010] Input the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers;

[0011] Aggregate the multiple graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result;

[0012] Input the aggregation result into a classifier for classification to obtain a diagnostic result.

[0013] In one embodiment, the gene data includes gene expression profile data and corresponding gene relationship network data;

[0014] Determine the gene feature matrix and the gene adjacency matrix according to the gene data, including:

[0015] Determine the gene feature matrix according to the gene expression profile data;

[0016] Construct the gene adjacency matrix according to the gene relationship network data.

[0017] In one embodiment, each tissue sample to be detected includes several features, and the features of all tissue samples to be detected constitute a feature matrix;

[0018] Determine the gene feature matrix according to the gene expression profile data, including:

[0019] Normalize the feature matrix to obtain a sparse matrix;

[0020] Store the sparse matrix to obtain the gene feature matrix.

[0021] In one embodiment, the gene relationship network data includes network nodes and network edges, the network nodes are gene features, and the network edges represent the relationships between the network nodes;

[0022] Construct the gene adjacency matrix according to the gene relationship network data, including:

[0023] Determine the gene adjacency matrix according to the network nodes and network edges.

[0024] In one embodiment, input the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers, including:

[0025] Transform the gene adjacency matrix A through the Laplacian matrix L = D - A, and after normalization, it is: where I n is the identity matrix;

[0026] Use the renormalization method to transform into where is The corresponding degree matrix;

[0027] According to the above conversion, the node features of the (l + 1)-th hidden layer of the graph convolutional neural network are obtained as:

[0028]

[0029] where σ is the activation function, and b (l) is the bias value of the l-th layer;

[0030] Let By stacking multiple layers of graph convolutional neural networks, the relationships of multi-order neighbors are obtained:

[0031]

[0032] …

[0033]

[0034] …

[0035]

[0036] where H (1) , H (2) ,..., H (l) , H (out) are the node features of each hidden layer, b (in) ,..., b (l) , b (out) are the bias values of each layer, Y is the output data, and f(·) is the softmax(·) function.

[0037] In one embodiment, multiple graph convolutional network layers are aggregated through an enhanced graph convolutional neural network, and splicing aggregation or attention-weighted splicing aggregation is adopted.

[0038] In one embodiment, according to the gene data, a gene feature matrix and a gene adjacency matrix are determined, including:

[0039] Preprocess the gene data to obtain the preprocessed gene data;

[0040] According to the preprocessed gene data, determine the gene feature matrix and the gene adjacency matrix.

[0041] In a second aspect, the present application provides a cancer tissue classification device, and the device includes:

[0042] An acquisition module, configured to acquire gene data corresponding to a set of tissues to be examined; the tissues to be examined include a plurality of tissue samples to be examined;

[0043] A determination module, configured to determine a gene feature matrix and a gene adjacency matrix according to the gene data;

[0044] A processing module, configured to input a gene feature matrix and a gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers;

[0045] An aggregation module, configured to aggregate multiple graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result;

[0046] A classification module, configured to input the aggregation result into a classifier for classification to obtain a diagnosis result.

[0047] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cancer tissue classification method as in the first aspect.

[0048] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the cancer tissue classification method as in the first aspect.

[0049] As can be seen from the technical solutions provided by the embodiments of the present specification above, this solution can fuse gene expression profile data and gene relationship network data through a graph convolutional neural network, solving the limitations and one-sidedness of existing methods for cancer tissue classification and the defect of low classification accuracy. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a schematic flowchart of the cancer tissue classification method provided by the present application;

[0052] Figure 2 It is a schematic structural diagram of the cancer tissue classification device provided by the present application;

[0053] Figure 3 It is a schematic structural diagram of the electronic device provided by the present application. Detailed Embodiments

[0054] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0055] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0056] Without departing from the scope or spirit of this application, various improvements and changes can be made to the specific implementation manners of this application's specification, which are obvious to those skilled in the art. Other implementation manners obtained from this application's specification are obvious to those skilled in the art. This application's specification and embodiments are merely exemplary.

[0057] Regarding the use of "comprising", "including", "having", "containing", etc. in this article, they are all open-ended terms, that is, they are meant to include but not be limited to.

[0058] In this application, "parts" are counted by mass parts unless otherwise specified.

[0059] The following further details the present invention in conjunction with the accompanying drawings and embodiments.

[0060] Referring to Figure 1 , which shows a schematic flow diagram applicable to the cancer tissue classification method provided in the embodiments of this application.

[0061] As Figure 1 shown, the cancer tissue classification method may include:

[0062] S110. Obtain gene data corresponding to the tissue set to be examined; the tissue to be examined includes a number of tissue samples to be examined.

[0063] Specifically, the tissue samples to be examined are cancer cells, and in this embodiment, breast cancer cells are used as an example for illustration. The tissue set to be examined is a collection of a number of tissue samples to be examined.

[0064] Gene data includes gene expression profile data and corresponding gene relationship network data. The gene expression profile data includes gene characteristic attributes and label information of corresponding samples. The gene relationship network data includes information such as gene network abbreviation numbers, names, gene attributes, and corresponding network edges.

[0065] Gene expression profile data can be obtained from gene databases, such as triple-class breast cancer Breast-A and quadruple-class breast cancer Breast-B, etc. The data can be obtained from the Broad Institute of Harvard University.

[0066] Gene relationship network data can be obtained from GIANT2.0 (Genome-wide Analysis of gene Networks in Tissues). This server hosts large-scale human tissue-specific gene expression networks, including 1,448,412 network edges corresponding to the abbreviation numbers, names, and gene attributes of each gene network.

[0067] Exemplarily, the number of samples obtained is as follows: there are 98 tissue samples of breast cancer type A (Breast-A), with 98×1,214 gene characteristics and a 1,214×1,214 gene network; the number of samples of breast cancer type B (Breast-B) is 49, with 49×1,214 gene characteristics and a 1,214×1,214 gene network.

[0068] S120. Determine a gene feature matrix and a gene adjacency matrix according to the gene data.

[0069] In one embodiment, S120 determining a gene feature matrix and a gene adjacency matrix according to the gene data may include:

[0070] Determine a gene feature matrix according to the gene expression profile data;

[0071] Construct a gene adjacency matrix according to the gene relationship network data.

[0072] Optionally, the gene expression profile data includes labels of corresponding tissue samples to be detected; each tissue sample to be detected includes several characteristics, and the characteristics of all tissue samples to be detected constitute a feature matrix.

[0073] Determining a gene feature matrix according to the gene expression profile data may include:

[0074] Normalize the feature matrix to obtain a sparse matrix;

[0075] Store the sparse matrix to obtain a gene feature matrix.

[0076] Specifically, assume that there are n tissue samples to be tested, each of which has m features, forming an n×m feature matrix. Normalize the n×m feature matrix, specifically: sum each row of the incoming feature matrix, take the inverse and perform dot multiplication. Then store the normalized feature matrix. Since some of the features have a value of 0, the feature matrix is ​​still a sparse matrix after normalization. Use matrix sparseness to store it, and get a feature matrix. The gene feature matrix is ​​recorded as X. Using matrix sparseness to store sparse matrices can improve computational efficiency.

[0077] It is understandable that when processing the characteristics of the tissue samples to be tested, the sample labels can be encoded. That is, since the data set composed of the labels of all tissue samples to be tested is multi-classified, the labels can be encoded using the one-hot encoding method, that is, using an N-bit state register to encode N states, each state has an independent register bit, and only one bit is valid at any time.

[0078] Optionally, the gene relationship network data includes network nodes and network edges, the network nodes are gene features, and the network edges represent the relationship between the network nodes.

[0079] According to the gene relationship network data, the gene adjacency matrix is ​​constructed, including:

[0080] According to the network nodes and network edges, the gene adjacency matrix is ​​determined.

[0081] Specifically, gene relationship network data consists of network nodes and network edges, where network nodes represent genes and key factors that regulate gene expression, and network edges are the relationships between network nodes. Network nodes are gene features, not samples.

[0082] According to the network nodes and network edges, we can get an undirected graph G = (V, E), where V = (v 1 ,v 2 ,…,v m ) is the set of network nodes, where the number of network nodes is the number of gene features m, and E is the relationship between network nodes, that is, the set of network edges.

[0083] The relationship between network node i and network node j is represented by a ij Indicates that if there is a correlation between network node i and network node j, then a ij =0, if there is no correlation, then a ij =1, which measures the connection strength between any two network nodes.

[0084] The relationships between all network nodes and network edges are stored as a matrix, which is the gene adjacency matrix A∈R corresponding to the undirected graphm×m 。

[0085] It can be understood that the gene adjacency matrix can be normalized by the degree matrix, D = diag(d 1 , d 2 , …, d m ), where d i = ∑ j a ij 。

[0086] To process the data for validity and consistency, the obtained original gene data is comprehensively sorted out (i.e., preprocessed).

[0087] In one embodiment, according to the gene data, a gene feature matrix and a gene adjacency matrix are determined, including:

[0088] Preprocess the gene data to obtain the preprocessed gene data;

[0089] According to the preprocessed gene data, determine the gene feature matrix and the gene adjacency matrix.

[0090] Specifically, to preprocess the gene data, first is data cleaning, mainly including duplicate value processing (deleting duplicate data), missing value processing (before modeling, reviewing and deleting attributes with excessive missing values and filling in the mean values of outliers with fewer missing values using a model), and outlier processing; second is data integration, performing operations such as data normalization.

[0091] It can be understood that when determining the gene feature matrix and the gene adjacency matrix according to the gene data, generally, the gene feature matrix and the gene adjacency matrix are determined based on the preprocessed gene data.

[0092] S130. Input the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers.

[0093] Specifically, the difference between the graph convolutional neural network in this application and the traditional graph convolutional network (GCN) is that the gene relationship network data describes the relationships between genes and cannot directly apply the traditional sample relationship model to the gene relationship network data. That is, each person has a fixed number of corresponding gene expression situations, and genes can be understood as the features corresponding to each sample. The graph convolutional neural network constructed in this application is about features, and the corresponding model formula should be changed.

[0094] For the gene feature matrix X and the gene adjacency matrix A input into the graph convolutional neural network, the graph convolutional neural network requires frequency domain conversion, and the gene adjacency matrix A needs to be transformed by means of the Laplacian matrix L = D - A. After normalization, it is: Among them, I n is the identity matrix.

[0095] To avoid the problems of gradient explosion and gradient vanishing caused by repeated application of this operator, a renormalization method is used to transform into where A self-loop is added to the original connection, so that it not only contains the information of neighbors but also adds its own information. is the corresponding degree matrix; after the above processing, the computational complexity of the graph convolutional neural network calculation can be greatly reduced.

[0096] According to the above input and changes, the node features of the (l + 1)-th hidden layer of the graph convolutional neural network can be obtained as:

[0097]

[0098] where σ is the activation function, and b (l) is the bias value of the l-th layer.

[0099] To simplify the notation, let Based on the above notations, larger-scale neighborhood information can be obtained by stacking multiple layers of GCN, aggregating the relationships of more-order neighbors. The stacked graph convolutional network layers further express node information by aggregating neighbor topological relationship information (one layer is for first-order neighbors, two layers are for second-order neighbors, and so on to obtain more-order neighbor information). The specific construction is as follows:

[0100]

[0101] …

[0102]

[0103] …

[0104]

[0105] where H (1) , H (2) ,..., H (l) , H (out) are the node features of each hidden layer. The dimension of the node features of each hidden layer is the same as the gene feature dimension, which is 1214. b (in) ,..., b (l) , b (out) are the bias values of each layer, Y is the output data, and f(·) is the softmax(·) function.

[0106] S140. Aggregate multiple graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result.

[0107] For gene networks, neighbors of different orders may have different angular effects. Some genes that deviate from the core of the gene network may require the effect of higher-order graph convolution, while for genes closer to the core, fewer layers may be sufficient. Therefore, this application will establish a multi-layer graph convolution model and fuse multiple graph convolutional network layers into one layer through methods such as concatenation aggregation and attention mechanism weighted concatenation aggregation, and use a feedforward neural network as the final output layer for feature dimensionality reduction.

[0108] Normally, a graph convolutional neural network only takes the result of the final layer in the output stage and ignores the intermediate feature representations in the front of the network. Inspired by the densenet model, this application combines the node features of all hidden layers of the graph convolutional layer instead of only using the last layer, which can improve the network performance and generalization ability. Therefore, the enhanced graph convolutional network (EGCN) of this application is used to combine all hidden layers learned by the GCN to enhance the network effect. This application will denote the output of each graph convolutional network layer as H (1) ,H (2) ,...,H (l) ,H (out) as {H (r)}, r = {1, 2,..., q}, where q is the total number of layers, and aggregate and enhance the representations of all graph convolutional network layers through the enhanced graph convolutional neural network.

[0109] This application shows two ways to perform aggregation: concatenation aggregation and attention weighted concatenation aggregation.

[0110] Among them, concatenation aggregation is:

[0111] F cat = cat(H (1) ,H (2) ,...,H (q)

[0112] Concatenate all graph convolutional network layers, where cat(·) is the concatenation function, directly concatenating the node features of each hidden layer, which is to average-aggregate the neighbors of each order of gene nodes and does not perform adaptive range neighbor aggregation for each gene node. The dimension of F cat is the sum of the dimensions of each layer.

[0113] Among them, attention weighted concatenation aggregation is to use the attention mechanism to calculate the corresponding attention scores for each graph convolutional node, and calculate the output of gene nodes based on the attention scores and the node features of each hidden layer, so that the appropriate neighbor range is selected when gene nodes perform neighbor aggregation.

[0114] Exemplarily, the output of the t-th layer graph convolutional node e is Input into the LSTM model, and each node in the model will obtain its corresponding backward expression After dimensionality reduction using the Linear layer, the attention scores of the corresponding nodes are obtained through softmax The final output of node e is

[0115] The attention weighted splicing aggregation is:

[0116] where a is any node.

[0117] It can be understood that the aggregation result after aggregating the enhanced graph convolutional neural network can also be input into the fully connected layer for dimensionality reduction:

[0118] F out = σ(F cat W F + b F )

[0119] where σ is the activation function Relu(·), W F is the learned weight parameter, b F is the bias parameter to be learned, and the fully connected layer further reduces the dimensionality and integrates the aggregated information.

[0120] S150. Input the aggregation result into the classifier for classification to obtain the diagnosis result.

[0121] Specifically, the aggregation result is classified using an SVM classifier.

[0122] The design uses an SVM multi-classifier. Since the samples are multi-class, it is completed using a one-vs-rest SVM classifier. Denote the number of classes as l, and the i-th training sample is (x i , y i ), y i ∈ 1,..., l, there are a total of m samples, and l SVM models are constructed. The k-th model takes the k-th class as the positive sample and the others as negative samples. The optimization problem of the k-th model is as follows:

[0123]

[0124] s.t. (w k ) T φ(x i ) + b k ≥ 1 - ξ i k , if yi = k

[0125] (w k ) T φ(x k ) + b k ≤ -1 + ξ i k ,if y i ≠ k

[0126] ξ i k ,≥ 0, i = 1,..., m

[0127] where C, ξ i are penalty parameters, and φ represents the non - linear mapping from the input space to the feature space.

[0128] After solving, the k - th decision function is obtained as follows:

[0129] (w k ) T φ(x t ) + b k

[0130] Classify x into the class with the maximum decision function value:

[0131]

[0132] Use the SVM classifier to replace the commonly used softmax classifier in the graph convolutional neural network model, that is, use the fully - connected network connected after the enhanced graph convolutional network model as the input of the multi - class SVM classifier, optimize the parameters of the SVM classifier through multiple trainings, and finally obtain an SVM classifier that can efficiently classify breast cancer types, so as to accurately classify the types of breast cancer in tissues, and finally output the classification results.

[0133] It can be understood that according to the EGCN structure, after the final aggregation, if a certain part plays a more important role, then the weight between it and the (Output) layer should be greater than that of the other part. Since the splicing of network layers does not affect the connection of weights, the importance of each layer of the neural network can be measured by the connection weights. This application uses the relative importance score of weights to express the ratio of the importance of a specific layer to the entire network, thereby analyzing the importance of a certain layer in the EGCN and better analyzing and obtaining the optimal neighbor order.

[0134] Relative Importance Score of Weights RIS:

[0135]

[0136] Among them, 1 = 1, 2, 3.. represents the 1st, 2nd, 3rd... layer corresponding to the graph convolutional layer in the splicing layer, and w cat→out refers to the weights of the overall splicing layer and the output layer, ||x|| 1 is the 1-norm. By comparing the sum of the absolute values of the weights of a certain layer with the sum of the absolute values of all weights in the splicing layer, the importance ratio of a certain layer of the neural network is obtained, so as to understand the relative importance of graph convolution of a certain layer.

[0137] In this embodiment, the importance of each layer of graph convolutional layer in the enhanced graph convolutional neural network is evaluated through the relative importance score, which helps to further optimize the model effect.

[0138] In the embodiment of the present application, the graph convolutional neural network can fuse gene expression profile data and gene relationship network data, solve the limitations and one-sidedness of existing methods for cancer tissue classification, and the defect of low classification accuracy. In addition, by combining the memory feature with the topological relationship of genes through the enhanced graph convolutional neural network, and aggregating the neighbor node relationships of different orders of gene nodes at the same time, the transmission and amplification relationship of the gene regulatory network is better considered.

[0139] Refer to Figure 2 , which shows a schematic structural diagram of a cancer tissue classification device described according to an embodiment of the present application.

[0140] As Figure 2 shown, the cancer tissue classification device may include:

[0141] An acquisition module 210, configured to acquire gene data corresponding to a set of tissues to be detected; the tissues to be detected include a plurality of tissue samples to be detected;

[0142] A determination module 220, configured to determine a gene feature matrix and a gene adjacency matrix according to the gene data;

[0143] A processing module 230, configured to input the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain a plurality of graph convolutional network layers;

[0144] An aggregation module 240, configured to aggregate the plurality of graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result;

[0145] A classification module 250, configured to input the aggregation result into a classifier for classification to obtain a diagnosis result.

[0146] Optionally, the gene data includes gene expression profile data and corresponding gene relationship network data;

[0147] The determination module 220 is further configured to:

[0148] Determine a gene feature matrix according to the gene expression profile data;

[0149] Construct a gene adjacency matrix based on gene relationship network data.

[0150] Optionally, each tissue sample to be tested includes several features, and the features of all tissue samples to be tested form a feature matrix.

[0151] The determination module 220 is further configured to:

[0152] Normalize the feature matrix to obtain a sparse matrix.

[0153] Store the sparse matrix to obtain a gene feature matrix.

[0154] Optionally, the gene relationship network data includes network nodes and network edges. The network nodes are gene features, and the network edges represent the relationships between the network nodes.

[0155] The determination module 220 is further configured to:

[0156] Determine a gene adjacency matrix according to the network nodes and network edges.

[0157] Optionally, the processing module 230 is further configured to:

[0158] Transform the gene adjacency matrix A through the Laplacian matrix L = D - A, and after normalization, it is: where I n is the identity matrix;

[0159] Use the renormalization method to transform into where is the corresponding degree matrix;

[0160] According to the above transformation, the feature of the hidden layer node of the (l + 1)-th layer of the graph convolutional neural network is:

[0161]

[0162] where σ is the activation function, and b (l) is the bias value of the l-th layer;

[0163] Let By stacking multiple layers of graph convolutional neural networks, the relationships of multi-order neighbors are obtained:

[0164]

[0165] …

[0166]

[0167] …

[0168]

[0169] Among them, H (1) , H (2) , …, H (l) , H (out) are the node features of each hidden layer, b (in) , ..., b (l) , b (out) are the bias values of each layer, Y is the output data, and f(·) is the softmax(·) function.

[0170] Optionally, multiple graph convolutional network layers are aggregated through an enhanced graph convolutional neural network, and splicing aggregation or attention-weighted splicing aggregation is adopted.

[0171] Optionally, the determining module 220 is further configured to:

[0172] Preprocess gene data to obtain preprocessed gene data;

[0173] Determine a gene feature matrix and a gene adjacency matrix according to the preprocessed gene data.

[0174] The cancer tissue classification device provided in this embodiment can execute the embodiments of the above method, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0175] Figure 3 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. As Figure 3 shown, a schematic structural diagram of an electronic device 300 suitable for implementing the embodiments of the present application is shown.

[0176] As Figure 3 shown, the electronic device 300 includes a central processing unit (CPU) 301, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage section 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0177] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 306 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required so that a computer program read therefrom is installed into the storage section 308 as required.

[0178] In particular, according to an embodiment of the present disclosure, the processes described above with reference to Figure 1 can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the above-described cancer tissue classification method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311.

[0179] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code includes one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions denoted by the blocks may occur in a different order than that denoted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0180] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0181] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0182] As another aspect, the present application also provides a storage medium. This storage medium can be the storage medium included in the aforementioned device in the above embodiments; it can also exist independently and be a storage medium not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to execute the cancer tissue classification method described in the present application.

[0183] The storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0184] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0185] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

Claims

1. A method for classifying cancer tissues, characterized in that, the method includes: obtaining gene data corresponding to a set of tissues to be examined; the tissues to be examined include a number of tissue samples to be examined; determining a gene feature matrix and a gene adjacency matrix according to the gene data; inputting the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers, including: Transform the gene adjacency matrix A through the Laplacian matrix L = D - A, and after standardization, it is as follows: where D is the degree matrix corresponding to A, and I n is the identity matrix; Convert to using the renormalization method into where is the corresponding degree matrix according to the above conversion, the node features of the (l + 1)-th hidden layer of the graph convolutional neural network are obtained as: Among them, σ is the activation function, and b (l) is the bias value of the l-th layer; Let By stacking multiple layers of the graph convolutional neural network, the relationships of multi-order neighbors are obtained: Among them, X is the gene feature matrix, W is the weight matrix, and H (1) , H (2) ,..., H (l) , H (out) are the node features of each hidden layer, b (in) ,..., b (l) , b (out) are the bias values of each layer, Y is the output data, and f(·) is the softmax(·) function; aggregating the multiple graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result; inputting the aggregation result into a classifier for classification to obtain a diagnosis result.

2. The method according to claim 1, characterized in that, the gene data includes gene expression profile data and corresponding gene relationship network data; the determining of the gene feature matrix and the gene adjacency matrix according to the gene data includes: determining the gene feature matrix according to the gene expression profile data; constructing the gene adjacency matrix according to the gene relationship network data.

3. The method according to claim 2, characterized in that, each tissue sample to be examined includes a number of features, and the features of all the tissue samples to be examined constitute a feature matrix; the determining of the gene feature matrix according to the gene expression profile data includes: normalizing the feature matrix to obtain a sparse matrix; storing the sparse matrix to obtain the gene feature matrix.

4. The method according to claim 2, characterized in that, the gene relationship network data includes network nodes and network edges, the network nodes are gene features, and the network edges represent the relationships between the network nodes; the constructing of the gene adjacency matrix according to the gene relationship network data includes: determining the gene adjacency matrix according to the network nodes and the network edges.

5. The method according to any one of claims 1-4, characterized in that, the aggregating of the multiple graph convolutional network layers through an enhanced graph convolutional neural network adopts splicing aggregation or attention-weighted splicing aggregation.

6. The method according to any one of claims 1-4, characterized in that, the determining of the gene feature matrix and the gene adjacency matrix according to the gene data includes: preprocessing the gene data to obtain preprocessed gene data; determining the gene feature matrix and the gene adjacency matrix according to the preprocessed gene data.

7. A cancer tissue classification device, characterized in that, the device includes: an obtaining module, configured to obtain gene data corresponding to a set of tissues to be examined; the tissues to be examined include a number of tissue samples to be examined; a determining module, configured to determine a gene feature matrix and a gene adjacency matrix according to the gene data; a processing module, configured to input the gene feature matrix and the gene adjacency matrix into a graph convolutional neural network to obtain multiple graph convolutional network layers, including: The gene adjacency matrix A is transformed by the Laplacian matrix L = D - A and after standardization, it is as follows: where D is the degree matrix corresponding to A, and I n is the identity matrix; Use the renormalization method to transform into where is the corresponding degree matrix; according to the above conversion, the node features of the (l + 1)-th hidden layer of the graph convolutional neural network are obtained as: Among them, σ is the activation function, and b (l) is the bias value of the l-th layer; Let By stacking multiple layers of the graph convolutional neural network, the relationships of multi-order neighbors are obtained: Among them, X is the gene feature matrix, W is the weight matrix, and H (1) , H (2) ,..., H (l) , H (out) are the node features of each hidden layer, b (in) ,..., b (l) , b (out) are the bias values of each layer, Y is the output data, and f(·) is the softmax(·) function; an aggregating module, configured to aggregate the multiple graph convolutional network layers through an enhanced graph convolutional neural network to obtain an aggregation result; A classification module, configured to input the aggregation result into a classifier for classification to obtain a diagnosis result.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the cancer tissue classification method according to any one of claims 1-6 is implemented.

9. A readable storage medium, having stored thereon a computer program, wherein, when the program is executed by a processor, the cancer tissue classification method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Supervised multi-perspective human synergistic lethal gene prediction method based on graph convolutional network

    CN110473592A

  • Community-enhanced graph convolutional neural network method

    CN111476261A