Tumor gene typing prediction method and system based on multi-modal data graph structure
By constructing a three-level pathological structure atlas and an improved graph attention mechanism, the problems of low data utilization efficiency and multi-level heterogeneity representation in existing tumor subtyping methods are solved, achieving high-precision tumor genotyping prediction and continuous model optimization.
Patent Information
- Application Number
- CN202511133750.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-14
Smart Images

Figure CN120636519B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor typing, and in particular to a method and system for predicting tumor gene typing based on a multimodal data graph structure. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Tumor classification is crucial for precision medicine and the development of personalized treatment strategies. Existing tumor classification methods primarily rely on single data modalities, such as gene expression profiles, gene mutation information, or pathological images. These methods have limitations in data utilization efficiency, classification accuracy, and characterization of the tumor microenvironment.
[0004] With the development of high-throughput omics and digital pathology technologies, existing technologies have enabled the simultaneous acquisition of multi-omics data and pathological image information from the same sample. However, existing technologies still lack an efficient and unified analytical framework for the effective integration of multimodal data, modeling of complex tumor structures, and deep mining of structure-expression association information. Furthermore, traditional graph-based modeling methods often focus only on single-level or single-scale structures, making it difficult to simultaneously represent the multi-level spatial heterogeneity of cells, the immune microenvironment, and the entire tissue. Summary of the Invention
[0005] To address the above problems, the present invention proposes a tumor genotyping prediction method and system based on multimodal data graph structure. By constructing a three-level pathological structure map, integrating multi-omics features, and adopting an improved graph attention mechanism for graph representation learning, it can more comprehensively mine the structural expression information of tumors and improve the accuracy and generalization ability of genotyping.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for predicting tumor genotyping based on a multimodal data graph structure, comprising the following steps:
[0008] Acquiring multimodal data of multiple tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features;
[0009] The pathology image is input into the Hyper-RBAT-Net network, which automatically segments and classifies the cell instances, TLS regions, and tissue blocks in the pathology image, outputs the corresponding structured mask image, and constructs a three-level pathology structure atlas based on the mask image.
[0010] The three-level pathological structure map is jointly embedded with the sample multi-omics features to construct a fusion graph structure. The graph representation of the fusion graph structure is learned using graph attention and main neighborhood aggregation mechanisms to obtain the sample embedding representation.
[0011] Use sample embedding representation to predict tumor typing, define the loss function, train the tumor typing prediction model, and use the trained tumor typing prediction model to predict tumor genotyping.
[0012] As an optional implementation, the Hyper-RBAT-Net network includes a residual enhancement structure that fuses channel attention and spatial attention.
[0013] As an optional implementation, the three-level pathological structure atlas includes a cell layer, a TLS layer, and a tissue layer, wherein the cell layer nodes are single cells, and the edge weights are calculated by measuring the spatial adjacency relationship using the Gaussian kernel; the TLS layer nodes are TLS clusters, and the edge weights are calculated based on morphological and density similarity; the tissue layer nodes are tissue blocks, and the edge weights are calculated based on gene expression heterogeneity.
[0014] As an optional implementation, the graph attention and primary neighborhood aggregation mechanism simultaneously considers multiple primary neighborhood statistics when aggregating nodes and introduces independent attention weights for each statistic.
[0015] As an optional implementation, the loss function is composed of a cross entropy loss, a distillation regularization term, a gradient smoothing regularization term, and a synthetic prior alignment term, and the weights of each component are adjusted by hyperparameters.
[0016] As an optional implementation method, it also includes saving the parameters and structure of the trained tumor typing prediction model to support rapid typing prediction of subsequent new samples. At the same time, when receiving new annotated data, the fine-tuning module can be called for incremental training to continuously optimize model performance.
[0017] In a second aspect, the present invention provides a tumor genotyping prediction system based on a multimodal data graph structure, comprising:
[0018] A data acquisition module is configured to: acquire multimodal data of a plurality of tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features;
[0019] The three-level pathology structure atlas construction module is configured to: input the pathology image into the Hyper-RBAT-Net network, automatically segment and classify the cell instances, TLS regions, and tissue blocks in the pathology image, output the corresponding structured mask image, and construct a three-level pathology structure atlas based on the mask image;
[0020] The sample embedding representation extraction module is configured to jointly embed the three-level pathological structure map with the sample multi-omics features to construct a fusion graph structure, and use graph attention and primary neighborhood aggregation mechanisms to learn the graph representation of the fusion graph structure to obtain the sample embedding representation;
[0021] The model training and prediction module is configured to: use sample embedding representation to predict tumor typing, define a loss function, train the tumor typing prediction model, and use the trained tumor typing prediction model to predict tumor genotyping.
[0022] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0023] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.
[0024] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which implements the method described in the first aspect when executed by a processor.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] This paper proposes a tumor genotyping prediction method based on a multimodal data graph structure. First, by constructing a three-level pathological structure graph (P3-Graph) consisting of cell, TLS, and tissue layers, this method systematically incorporates key structural information from the tumor microenvironment, particularly the spatial distribution characteristics of tertiary lymphoid structures (TLS), effectively improving the modeling capabilities of complex tumor tissue structures. Second, the proposed Hyper-RBAT-Net network, with its multi-scale attention residual properties, efficiently identifies multiple tissue components in pathological images and automatically generates structured mask images, providing high-quality input for subsequent graph construction. Furthermore, the proposed graph attention and principal neighbor aggregation mechanism (HA-PNA) fuses the graph structure with multi-omics data, enhancing the model's ability to jointly perceive expression heterogeneity and spatial topology, thereby obtaining biologically interpretable embedding representations. Compared to traditional methods, this method not only achieves higher classification accuracy but also exhibits excellent generalization and scalability, making it suitable for rapid prediction of new samples and online model fine-tuning. It also possesses continuous learning capabilities, meeting the clinical demand for efficient, accurate, and interpretable tumor typing technologies.
[0027] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0029] Figure 1 Schematic diagram of the structure of the Hyper-RBAT-Net network of the present invention;
[0030] Figure 2 The tumor pathology image segmentation mask and the three-level pathology structure diagram of the present invention;
[0031] Figure 3 This is a data flow chart of the tumor gene typing prediction method based on the multimodal data graph structure of the present invention. DETAILED DESCRIPTION
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0033] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0034] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but includes other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0036] Example 1
[0037] like Figure 3 As shown, this embodiment provides a tumor genotyping prediction method based on a multimodal data graph structure, comprising the following steps:
[0038] S1 acquires multimodal data of multiple tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features;
[0039] S2 inputs the pathological image into the Hyper-RBAT-Net network, automatically segments and classifies the cell instances, TLS regions, and tissue blocks in the pathological image, outputs the corresponding structured mask image, and constructs a three-level pathological structure atlas based on the mask image;
[0040] S3 jointly embeds the three-level pathological structure map with the sample multi-omics features to construct a fusion graph structure. It uses graph attention and main neighborhood aggregation mechanisms to learn the graph representation of the fusion graph structure and obtain the sample embedding representation.
[0041] S4 uses sample embedding representation to predict tumor typing, defines a loss function, trains the tumor typing prediction model, and uses the trained tumor typing prediction model to predict tumor genotyping.
[0042] The solution of the present invention is explained in detail below.
[0043] Step S1 first acquires multimodal raw data from multiple tumor samples to be tested. This data includes HE-stained pathological slide images, the sample's gene expression matrix, and optional mutational load and methylome information, combined with clinically annotated genotyping labels as supervisory signals. Compared to traditional methods, this method introduces a data architecture that deeply integrates multi-omics data with spatial pathology images, breaking the limitations of a single modality at the sample level and providing a complete multidimensional information foundation for subsequent deep atlas modeling.
[0044] Step S2 constructs a pathology image recognition network Hyper-RBAT-Net based on the improved residual attention module to achieve fine segmentation and structured annotation of cell instances, TLS regions, and tissue blocks in pathology images. The innovation lies in the proposed residual enhancement structure that integrates channel attention and spatial attention. Its feature map expression formula is:
[0045] ;
[0046] Among them, the input feature tensor of RBAT-Net feature map is recorded as , represents the basic features extracted from the original pathological image. After channel attention and spatial attention enhancement, the output feature tensor obtained by network fusion is expressed as . is a dynamic gating coefficient, which is a learnable parameter with a value in the interval [0, 1] and is used to adjust the fusion weight between channel attention and spatial attention. and They are channel attention and spatial attention enhancement features, specifically defined as:
[0047] ;
[0048] ;
[0049] The feature representation after channel attention enhancement is , whose calculation depends on the input features Perform global average pooling (denoted as ), and then through two linear transformations, namely two weight matrices and , the middle passes through the ReLU activation function, and the final output is normalized by the Sigmoid function and Do element-wise product (Hadamard product, denoted as ), thereby achieving channel weight adjustment. Spatial attention enhancement feature The calculation of Average pooling is performed on the channel dimension ( ) and max pooling ( ), concatenate the two (using the symbol After that, a size of The convolution kernel Then, after Sigmoid activation, the original feature Multiply element by element to enhance the response of spatially significant regions. This module innovatively adopts dynamic factors Adaptively adjust local and global attention contributions to effectively improve the segmentation accuracy of cell instances and tiny TLS regions, and enhance the network's generalization ability in cross-sample data.
[0050] The mask image obtained by segmentation is used to construct a three-level pathological structure map (P3-Graph), which includes the cell layer, TLS layer, and tissue layer. Each node in the cell layer corresponds to a single cell, and the edge weight is calculated using the Gaussian kernel metric spatial adjacency relationship. The edge weight calculation formula is:
[0051] ;
[0052] where x i and x j Represents cells i and cells j The center of mass position of σ cell is the bandwidth parameter of the Gaussian kernel, which is used to control the degree of spatial proximity attenuation.
[0053] The nodes in the TLS layer are TLS clusters. The edge weight is based on the similarity of shape and density. The edge weight calculation formula is:
[0054] ;
[0055] Among them, hi and h j Represents TLS clusters i and j The shape and density eigenvector of D. ij Indicates a TLS blob i and j The spatial distance between them. TLS is the bandwidth parameter of the Gaussian kernel, which is used to adjust the decay rate of feature similarity and spatial distance.
[0056] The tissue layer nodes are tissue blocks, and the edge weights are based on gene expression heterogeneity. The edge weight calculation formula is:
[0057] ;
[0058] Among them, g i and g j Represents organizational blocks i and j The gene expression feature vector of D ij Indicates the organization block i and j The spatial distance between them. θ ij Indicates the organization block i and j The angle between the main axis directions. σ tissue is the bandwidth parameter of the Gaussian kernel, which is used to adjust the decay rate of gene expression heterogeneity and spatial distance. λ is the adjustment factor, which is used to control the weight of directional information in edge weight calculation.
[0059] Step S3: The tertiary structure map is jointly embedded with the sample multi-omics features to construct a fusion graph structure. The innovative graph attention and primary neighborhood aggregation mechanism (HA-PNA) is used for graph representation learning. This mechanism considers multiple primary neighborhood statistics (mean, maximum, standard deviation, harmonic mean) when aggregating nodes and introduces independent attention weights for each statistic. The calculation formula is:
[0060] ;
[0061] ;
[0062] Among them, each node The current representation vector of , and its updated representation is The update process is completed by a multi-layer perceptron (MLP) to nonlinearly transform the neighbor node information. The set of neighbor nodes is denoted as , a variety of statistical functions are used to aggregate the representation of its neighbors, forming a set of statistical functions , including mean, maximum (max), standard deviation (std) and harmonic mean (harmonic-mean). Each statistical function Neighbor node features After application, the corresponding statistical features are calculated and then multiplied by the attention weight calculated by the attention mechanism The attention weight is based on the node With node The splicing of features (denoted as ) and the parameter vector The inner product is calculated and normalized using the softmax function to measure the relative importance of neighboring nodes under a certain statistic. Finally, the aggregation results of all statistical methods are concatenated and fed into the MLP to obtain the updated representation of the node.
[0063] Step S4: Input the node embedding representation obtained from the fusion graph representation learning into the typing prediction module, and output the sample genotyping label through a multi-layer perceptron and softmax classifier. The present invention innovatively models the typing label as a product of conditional joint probabilities:
[0064] ;
[0065] Among them, the global representation of the entire graph structure is , which is usually obtained through graph-level pooling. Based on this embedding vector, multiple softmax classifiers are used to classify different genotype labels. Modeling, each classifier has a corresponding weight matrix and bias The predicted probability of each label is , the joint prediction probability of all labels is the product of these outputs, that is, , Represents the total number of label categories. This model structure can achieve conditional independence modeling between multiple labels and adapt to the common label co-occurrence and non-mutually exclusive characteristics in genotyping tasks.
[0066] The classification model is trained using the cross-entropy loss function, and its performance is evaluated using multiple indicators such as accuracy, F1 value, and AUROC. To address the issue of continuous model optimization, this paper proposes a training strategy that combines a memory distillation regularization term, with the total loss defined as:
[0067] ;
[0068] Among them, the total loss It consists of four parts:
[0069] 1. Cross Entropy Loss : Measuring the difference between the current model prediction result and the true label is the most core supervisory signal.
[0070] 2. Historical consistency regularization term : The KL divergence constraint is used to keep the output of the current model consistent with the output of the previous stage model to prevent catastrophic forgetting. The coefficient β controls the strength of “remembering the past”.
[0071] 3. Gradient smoothing regularization term : Penalizes drastic gradient changes of the loss parameters to avoid "oscillatory" forgetting during model updates. The coefficient γ determines the degree of smoothing.
[0072] 4. Synthetic Prior Alignment : Using the maximum mean difference (MMD) to align the distribution of the current model with the "synthetic prior distribution" generated from Gaussian noise, it is equivalent to "pre-storing" possible future data features into the current representation space in advance. The coefficient δ controls the strength of "pre-loading" future information.
[0073] The trained model parameters, structure, and graph generation rules are stored for rapid classification prediction of new samples. The system features a fine-tuning module that automatically triggers incremental training upon receipt of newly annotated data, enabling seamless knowledge updates. This design supports end-to-end deployment and is compatible with existing pathology and multi-omics data platforms, demonstrating strong engineering feasibility and clinical application potential.
[0074] The genotyping prediction method of the present invention is applied to the typing prediction of lung cancer and breast cancer respectively.
[0075] Multimodal fusion classification prediction based on lung cancer samples:
[0076] This example uses lung cancer patients as the subjects, and the data sources include a tertiary hospital and the TCGA database, which contains a total of 1,200 cases. First, according to step S1, multimodal data of each patient are collected: the pathological section image is stained with HE, with a resolution of 0.25 μm / pixel; the gene expression matrix contains approximately 18,000 genes; mutation load and methylation group data are optional omics features; all samples are classified and labeled by three senior pathologists. Image data preprocessing includes color standardization and removal of background artifacts, and gene data is batch effect corrected (ComBat) and standardized. In step S2, the constructed pathological image recognition module is based on the improved Hyper-RBAT-Net network (such as Figure 1 As shown in Figure 2), the network introduces dual-path attention fusion and dynamic gating mechanism on the conventional residual structure, through the formula:
[0077] ;
[0078] ;
[0079] in, represents the input image feature map, is the global average pooling result, and are weight and bias terms respectively, Represents the Sigmoid function, the final output is Acts as a dynamic fusion coefficient to regulate attention pathways. and Represent the output of spatial attention and channel attention respectively, and the two are based on Weighted fusion is used to improve the adaptability and accuracy of pathological feature expression.
[0080] The model outputs a structured mask image (such as Figure 2 As shown in Figure 2, the instance segmentation of cells, TLS, and tissue blocks is achieved. A three-level pathological structure map P3-Graph is generated based on the mask: the cell layer uses the center of mass as the node, and the edge weight is based on the spatial Gaussian kernel:
[0081] ;
[0082] in Represents cells and The center of mass position, is the Gaussian kernel parameter, which is used to control the degree of spatial proximity attenuation. The tissue layer further integrates expression heterogeneity based on spatial adjacency. The tissue layer uses functional tissue blocks as nodes, and the edge weight is calculated by combining expression heterogeneity and spatial adjacency:
[0083] ;
[0084] in Represents the spatial distance between tissue blocks, is the distance control parameter, are the expression feature vectors of tissue blocks, Represents cosine similarity, which can measure the consistency between expression patterns.
[0085] P3-Graph and multi-omics features are jointly embedded into a fusion graph structure and input into the HA-PNA graph representation learning module, which is implemented through multi-statistic aggregation and multi-head attention:
[0086] ;
[0087] in, is a node The neighbor set of is the feature representation of neighbor nodes, Represents various statistical aggregation functions (including mean, maximum, standard deviation, and harmonic mean), is the attention weight corresponding to the aggregation type, This mechanism improves the model's ability to perceive differences in structural relationships and expression patterns by weighted fusion of statistical features from different neighborhoods and further processing them using a multi-layer perceptron (MLP).
[0088] The loss function is defined as
[0089] ,
[0090] in, is the cross entropy loss; is the historical consistency distillation term; is the gradient smoothing regularization term; is the synthetic prior alignment term. The hyperparameters β, γ, and δ are determined by grid search on the validation set.
[0091] Figure 1 Schematic diagram of the Hyper-RBAT-Net network structure, including a four-stage encoder, a dual-path attention residual module, and a dynamic gating unit; Figure 2 Schematic diagram of the tumor pathology image segmentation mask and three-level pathology structure atlas, showing cells, TLS, and tissue layers.
[0092] Multimodal fusion typing prediction based on breast cancer samples:
[0093] This example is based on breast cancer patient sample data, totaling 900 cases, including public databases and clinically collected samples. The data pattern in step S1 is consistent with that in Example 1. First, 700 cases are selected as initial training data, and the initial model training is completed according to steps S2 to S4. Hyper-RBAT-Net segmentation is constructed in the same way as P3-Graph, but in order to adapt to the characteristics of breast cancer tissue, the spatial Gaussian edge weight σ parameter is adaptively tuned. When 200 new samples are added, the model adopts an incremental learning strategy without the need for overall retraining, and is jointly optimized through cross entropy and distillation loss:
[0094] ;
[0095] in is the cross entropy loss; is the historical consistency distillation term; is the gradient smoothing regularization term; is the synthetic prior alignment term. The hyperparameters β, γ, and δ are determined by grid search on the validation set. The incremental training process introduces adaptive edge scaling in the HA-PNA module:
[0096]
[0097] in, For the The node feature matrix of the layer, is the normalized adjacency matrix, is the linear transformation weight, is a learnable scaling coefficient matrix. This mechanism allows the model to dynamically adjust the influence of adjacent edges within the graph structure, thereby enhancing its ability to model heterogeneous neighborhood relationships. The final model achieved an AUROC of 0.92 on incremental data and 0.93 on the validation set, demonstrating excellent generalization and continuous learning capabilities.
[0098] Example 2
[0099] This embodiment provides a tumor genotyping prediction system based on a multimodal data graph structure, including:
[0100] A data acquisition module is configured to: acquire multimodal data of a plurality of tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features;
[0101] The three-level pathology structure atlas construction module is configured to: input the pathology image into the Hyper-RBAT-Net network, automatically segment and classify the cell instances, TLS regions, and tissue blocks in the pathology image, output the corresponding structured mask image, and construct a three-level pathology structure atlas based on the mask image;
[0102] The sample embedding representation extraction module is configured to jointly embed the three-level pathological structure map with the sample multi-omics features to construct a fusion graph structure, and use graph attention and primary neighborhood aggregation mechanisms to learn the graph representation of the fusion graph structure to obtain the sample embedding representation;
[0103] The model training and prediction module is configured to: use sample embedding representation to predict tumor typing, define a loss function, train the tumor typing prediction model, and use the trained tumor typing prediction model to predict tumor genotyping.
[0104] It should be noted that the above modules correspond to the steps in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules can be executed in a computer system as part of the system.
[0105] In further embodiments, there is also provided:
[0106] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed by the processor, wherein when the computer instructions are executed by the processor, the method in embodiment 1 is performed. For the sake of brevity, no further details are given here.
[0107] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0108] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method in embodiment 1 is completed.
[0109] The method in Example 1 can be directly executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, it will not be described in detail here.
[0110] A computer program product includes a computer program, wherein the computer program implements the method in embodiment 1 when executed by a processor.
[0111] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions contained in program modules, which are executed in a device on a real or virtual processor of a target to perform the process / method described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided between program modules as needed. The machine-executable instructions for the program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0112] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0113] In the context of the present invention, computer program code or related data can be carried by any appropriate carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, and the like.
[0114] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0115] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A tumor genotyping prediction method based on a multimodal data graph structure, characterized in that: The following steps are involved: Acquiring multimodal data of multiple tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features; The pathology image is input into the Hyper-RBAT-Net network, which automatically segments and classifies the cell instances, TLS regions, and tissue blocks in the pathology image, outputs the corresponding structured mask image, and constructs a three-level pathology structure atlas based on the mask image. The three-level pathological structure map is jointly embedded with the sample multi-omics features to construct a fusion graph structure. The graph representation of the fusion graph structure is learned using graph attention and main neighborhood aggregation mechanisms to obtain the sample embedding representation. Use sample embedding representation to predict tumor typing, define the loss function, train the tumor typing prediction model, and use the trained tumor typing prediction model to predict tumor genotyping; The Hyper-RBAT-Net network includes a residual enhancement structure that integrates channel attention and spatial attention. Its feature mapping expression formula is: ; Where, is the input feature tensor of the RBAT-Net feature map, which represents the basic features extracted from the original pathological image. After channel attention and spatial attention enhancement, the output feature tensor obtained by network fusion is expressed as , is a dynamic gating coefficient used to adjust the fusion weight between channel attention and spatial attention. and They are channel attention and spatial attention enhancement features, specifically defined as: ; ; Where, is the feature representation after channel attention enhancement, represents the Hadamard product, and is the weight matrix, ReLU is the activation function, represents the global average pooling operation, To enhance the features of spatial attention, Represents a splicing operation, represents the average pooling operation, represents the maximum pooling operation, Indicates the size The convolution kernel of The three-level pathological structure atlas includes the cell layer, TLS layer, and tissue layer. The cell layer nodes are single cells, and the edge weights are calculated by Gaussian kernel metric spatial adjacency. The TLS layer nodes are TLS clusters, and the edge weights are calculated based on morphological and density similarity. The tissue layer nodes are tissue blocks, and the edge weights are calculated based on gene expression heterogeneity. The loss function consists of cross entropy loss, distillation regularization term, gradient smoothing regularization term and synthetic prior alignment term, and the weights of each component are adjusted through hyperparameters.
2. The tumor genotyping prediction method based on multimodal data graph structure according to claim 1, characterized in that: The graph attention and primary neighborhood aggregation mechanisms simultaneously consider multiple primary neighborhood statistics when aggregating nodes and introduce independent attention weights for each statistic.
3. The tumor genotyping prediction method based on multimodal data graph structure according to claim 1, characterized in that: It also includes saving the parameters and structure of the trained tumor classification prediction model, supporting the rapid classification prediction of subsequent new samples. At the same time, when receiving new annotated data, the fine-tuning module can be called for incremental training to continuously optimize model performance.
4. A tumor genotyping prediction system based on a multimodal data graph structure, characterized in that: The method for predicting tumor genotyping based on a multimodal data graph structure according to any one of claims 1 to 3 comprises: A data acquisition module is configured to: acquire multimodal data of a plurality of tumor samples to be detected, wherein the multimodal data includes sample pathological images and sample multi-omics features; The three-level pathology structure atlas construction module is configured to: input the pathology image into the Hyper-RBAT-Net network, automatically segment and classify the cell instances, TLS regions, and tissue blocks in the pathology image, output the corresponding structured mask image, and construct a three-level pathology structure atlas based on the mask image; The sample embedding representation extraction module is configured to jointly embed the three-level pathological structure map with the sample multi-omics features to construct a fusion graph structure, and use graph attention and primary neighborhood aggregation mechanisms to learn the graph representation of the fusion graph structure to obtain the sample embedding representation; The model training and prediction module is configured to: use sample embedding representation to predict tumor typing, define a loss function, train the tumor typing prediction model, and use the trained tumor typing prediction model to predict tumor genotyping.
5. An electronic device, characterized in that: The method comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 3 is completed.
6. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, complete the method according to any one of claims 1 to 3.
7. A computer program product, characterized in that The invention comprises a computer program, which is used to implement the method according to any one of claims 1 to 3 when executed by a processor.
Citation Information
Patent Citations
Brain tumor image segmentation method, system and device and storage medium
CN114581662A
Breast tumor BI-RADS grading method based on multi-modal graph neural network
CN118430790A