A method and device for spatial single-cell gene multi-omics quantitative prediction

Through the AIHE multi-omics deep learning model, H&E stained sections were converted into WSI format image blocks, single-cell information was extracted using pre-trained models, and multi-omics data recognition was combined with EfficientNet-B3 backbone network, which solved the problem of high cost and long time for single-cell whole transcript sequencing in the existing technology, and achieved efficient single-cell multi-omics quantitative prediction and real-time multi-omics analysis.

CN119229956BActive Publication Date: 2025-08-26NANNING YUANDONG LIFE MEDICAL RESEARCH & DEVELOPMENT CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411083617.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-08-26
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

The existing spatial multiomics technology has high cost and long time to sequence whole transcripts at the single-cell level, and is unable to process fresh frozen sections, limited throughput, the potential of H&E stained images is not fully utilized, there are inconsistencies and human errors in sample preparation and staining processes, limiting the applicability of tissue volume and biomarker detection.

Method used

Using AIHE multi-omics deep learning model, single-cell information is extracted by converting H&E stained sections into WSI format image blocks, and single-cell information is extracted using the pre-trained HoverNet model, and multi-omics data recognition training is carried out in combination with the EfficientNet-B3 backbone network to identify single-cell quantitative data.

Benefits of technology

It improves the recognition level of H&E stained images, reduces research costs, and realizes multiomic quantitative prediction at the single-cell level, which can provide real-time multiomic prediction during the surgery, supporting real-time clinical guidance and intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229956B_ABST
    Figure CN119229956B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for spatial single-cell gene multi-omics quantitative prediction, relating to the field of single-cell identification technology. The method constructs an AIHE multi-omics deep learning model for H&E-stained sections. First, multi-omics data is converted into binary variable data, and the H&E-stained sections are processed into WSI format image blocks. Regional labels are set for cells in the WSI format image blocks. Then, the WSI format image blocks containing the regional labels of single cells are used as input of the AIHE multi-omics deep learning model, and the converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition. After training, the AIHE multi-omics deep learning model identifies single-cell quantitative data corresponding to the H&E-stained sections to be identified, thereby improving the recognition level of the H&E-stained images and reducing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the field of single-cell identification technology, and specifically to a method and device for spatial single-cell gene multi-omics quantitative prediction. Background Art

[0002] Spatial multi-omics technologies are rapidly advancing, with resolution increasing from bulk and point-level to single-cell and even subcellular dimensions. Cutting-edge imaging technologies such as MSI, VisiumHD, and Xenium have already achieved quantitative measurements at the micron level. However, they still face many unresolved technical challenges. Among them, spatial transcriptomics has been hailed by Nature as the leading sequencing technology of the 2020s, with potential to become third-generation sequencing. However, with the continuous improvement in spatial resolution, a major challenge has emerged: a significant increase in sequencing costs and time. Furthermore, the requirements for tissue samples have become more stringent. For example, the state-of-the-art technologies, VisiumHD and Xenium, are only compatible with formalin-fixed paraffin-embedded (FFPE) sections and cannot process fresh-frozen (FF) sections obtained during surgery. Furthermore, their 5-10 day sequencing cycle often limits their use to post-treatment analysis, hindering real-time guidance and intervention during treatment. Another major challenge is throughput limitations. While VisiumHD technology can sequence the entire transcriptome, its accuracy still falls short of true single-cell levels. On the other hand, Xenium technology, which uses optical imaging to measure transcripts at the nanoscale, is limited by its reliance on multiple rounds of fluorescent probe hybridization. This limitation limits its ability to sequence only approximately 500 genes, rather than the entire transcriptome. Therefore, there remains a gap in technologies capable of performing whole-transcriptome sequencing with spatial precision at the single-cell level.

[0003] In routine pathology practice, hematoxylin and eosin (H&E) staining is one of the most commonly used staining methods, leveraging its ability to interact with tissue components to visualize tissue structure under the microscope. However, this method faces several challenges. First, inconsistencies and human error during tissue preparation and staining are common problems. Second, H&E staining requires significant processing time, expensive imaging systems, and substantial procedural fees. Furthermore, time constraints for sampling are another significant obstacle. Time and cost considerations can limit the volume of tissue available for staining, and once tissue is irreversibly stained, the value of the biopsy sample may be compromised, making it unsuitable for subsequent biomarker testing. Advances in whole-slide imaging (WSI) technology now allow high-magnification scans containing tens of billions of pixels to be acquired in a matter of hours. However, the full potential of H&E-stained images has not yet been fully realized. Currently, H&E-stained images are primarily used to observe basic nuclear morphology and distribution. If other molecular features need to be assessed, more complex and costly techniques must be employed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a method and device for spatial single-cell gene multi-omics quantitative prediction.

[0005] The present invention solves the above-mentioned technical problem with the following technical solution: A method for spatial single-cell gene multi-omics quantitative prediction comprises the following steps:

[0006] Obtaining H&E stained sections and their corresponding multi-omics data from a basic database, converting the multi-omics data into binary variable data, and processing the H&E stained sections into WSI format image blocks;

[0007] The pre-trained HoverNet model is used to extract single-cell information from the WSI format image blocks, and the nuclear region corresponding to the single-cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block;

[0008] An AIHE multi-omics deep learning model is constructed using EfficientNet-B3 as the backbone network. WSI format image blocks containing regional labels of single cells are used as inputs of the AIHE multi-omics deep learning model. The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model.

[0009] The H&E-stained sections to be identified were input into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E-stained sections to be identified.

[0010] On the basis of the above technical solution, the present invention can also be improved as follows.

[0011] Furthermore, the H&E stained sections and their corresponding multi-omics data are obtained from the basic database, the multi-omics data are converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, specifically:

[0012] H&E-stained sections and their corresponding multi-omics data were obtained from two basic databases, the TCGA database and the GTEx database;

[0013] performing batch effect removal processing on the multi-omics data, performing median truncation processing on the multi-omics data after the batch effect removal processing, and classifying the multi-omics data after the median truncation processing using a multi-label binary classification method to obtain binary variable data;

[0014] The H&E stained sections are converted into WSI format images, and the WSI format images are optimized, and the optimized WSI format images are cut into blocks according to a set size and set pixels to obtain WSI format image blocks.

[0015] Furthermore, the pre-trained HoverNet model is used to extract single cell information from the WSI format image blocks, and the nuclear region corresponding to the single cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block, specifically:

[0016] Using the pre-trained HoverNet model, single-cell information is extracted from WSI format image blocks, including cell nucleus position and morphology information;

[0017] The nuclear region corresponding to the cell nucleus position and morphology information is expanded to the adjacent region by a morphological expansion operator to obtain the region label of each single cell in the WSI format image block. The morphological expansion operator is:

[0018]

[0019] The morphological expansion operator is represented by expanding each point x in the WSI format image block A to (B) x , recorded as As a region label set, the region label set includes a region label of each single cell in a WSI format image block.

[0020] Furthermore, the WSI format image blocks containing the regional labels of single cells are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data are used as the training target of the AIHE multi-omics deep learning model, and the AIHE multi-omics deep learning model is trained for multi-omics data recognition, specifically:

[0021] Assume that the WSI format image block contains m cells, randomly mask mn cells in the WSI format image block as background noise in the feature dimension, and the remaining n cells are trained using the corresponding region labels;

[0022] The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model, and the model parameters of the EfficientNet-B3 backbone network are updated using the AdamW optimizer with a pre-set learning rate, including:

[0023] During each round of training, the WSI format image block is copied into two images, namely img1 and img0. In img1, all cells are masked by the regional label of each single cell, so that img1 contains background noise unrelated to the cell, while img0 retains n cells and masks other cells, so that img0 contains background noise and information of n cells.

[0024] Input the img1 image and the img0 image into the EfficientNet-B3 backbone network, extract the corresponding image features, record them as feature1 and feature0, and calculate the feature difference between feature1 and feature0 to obtain the difference between the two features;

[0025] The difference is input into a fully connected layer connected to the output of the EfficientNet-B3 backbone network to output various types of single-cell quantitative data.

[0026] Another technical solution of the present invention to solve the above technical problem is as follows: a device for spatial single-cell gene multi-omics quantitative prediction, comprising:

[0027] Obtaining H&E stained sections and their corresponding multi-omics data from a basic database, converting the multi-omics data into binary variable data, and processing the H&E stained sections into WSI format image blocks;

[0028] The pre-trained HoverNet model is used to extract single-cell information from the WSI format image blocks, and the nuclear region corresponding to the single-cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block;

[0029] An AIHE multi-omics deep learning model is constructed using EfficientNet-B3 as the backbone network. WSI format image blocks containing regional labels of single cells are used as inputs of the AIHE multi-omics deep learning model. The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model.

[0030] The H&E-stained sections to be identified were input into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E-stained sections to be identified.

[0031] Another technical solution of the present invention to solve the above-mentioned technical problem is as follows: a device for spatial single-cell gene multi-omics quantitative prediction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for spatial single-cell gene multi-omics quantitative prediction as described above is implemented.

[0032] The beneficial effects of the present invention are: constructing an AIHE multi-omics deep learning model for H&E stained sections, first converting multi-omics data into binary variable data, and processing the H&E stained sections into WSI format image blocks, and setting regional labels for cells in the WSI format image blocks, then using the WSI format image blocks containing single-cell regional labels as input of the AIHE multi-omics deep learning model, and using the converted binary variable data as the training target of the AIHE multi-omics deep learning model, performing multi-omics data recognition training on the AIHE multi-omics deep learning model, and after training, the AIHE multi-omics deep learning model identifies the single-cell quantitative data corresponding to the H&E stained sections to be identified, thereby improving the recognition degree of H&E stained images and reducing research costs compared to existing immunohistochemistry (IHC) and immunofluorescence (IF). BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of the process of a method for spatial single-cell multi-omics quantitative prediction provided by an embodiment of the present invention;

[0034] Figure 2 A schematic diagram illustrating the process of processing H&E-stained sections into WSI format images according to an embodiment of the present invention;

[0035] Figure 3 A schematic diagram of training an AIHE multi-omics deep learning model provided in an embodiment of the present invention;

[0036] Figure 4 A schematic diagram of the prediction training of single cell quantitative data provided by an embodiment of the present invention;

[0037] Figure 5 Schematic diagram of single cell quantitative data prediction using the trained AIHE multi-omics deep learning model provided in an embodiment of the present invention;

[0038] Figure 6 A block diagram of a device for spatial single-cell multi-omics quantitative prediction provided by an embodiment of the present invention;

[0039] Figure 7 Schematic diagram showing the survival-DCA curve showing the prognostic performance of AIHE on recurrence-free survival (RFS) in the TCGA-COAD cohort;

[0040] Figure 8 Schematic diagram showing the survival-DCA curve showing the prognostic performance of AIHE on overall survival (OS) in the TCGA-COAD cohort;

[0041] Figure 9 Waterfall plot of the true mutation spectrum and AIHE-predicted mutation spectrum for TCGA-COAD cohort samples;

[0042] Figure 10 Schematic diagram of the predictive performance of AIHE on TCGA-COAD cohort samples using mainstream commercial cancer gene mutation detection panels (MSK-IMPACT (MSK) and FoundationOne CDx (F1CDX));

[0043] Figure 11 Schematic diagram of the composite ROC curve showing the prediction performance of AIHE for cancer-related gene mutations in COAD cohort samples;

[0044] Figure 12 Schematic diagram showing the predictive performance of AIHE for TMB in TCGA-COAD cohort samples for the composite ROC-calibration-DCA curve;

[0045] Figure 13 Schematic diagram showing the composite ROC-calibration-DCA curve for the prediction performance of AIHE for MSI in TCGA-COAD cohort samples;

[0046] Figure 14 Schematic diagram of the point-level spatial transcriptome data generated based on AIHE for TCGA-AZ-4313 samples;

[0047] Figure 15 Schematic diagram of the point-level spatial transcriptome data generated based on AIHE for TCGA-AY-6196 samples;

[0048] Figure 16 Schematic diagram of learning spatial single-cell transcriptomes from H&E staining of pathological sections for AIHE of TCGA-AZ-4313 and TCGA-AY-6196 samples;

[0049] Figure 17 Schematic diagram of learning spatial single-cell transcriptomes from H&E staining of pathological sections for AIHE of TCGA-3L-AA1B sample;

[0050] Figure 18 Schematic diagram of AIHE for spatial single-cell proteomics prediction;

[0051] Figure 19 Schematic diagram of performing spatial single-cell multi-omics predictions for AIHE. DETAILED DESCRIPTION

[0052] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0053] like Figure 1 As shown, an embodiment of the present invention provides a method for spatial single-cell gene multi-omics quantitative prediction, comprising the following steps:

[0054] Obtaining H&E stained sections and their corresponding multi-omics data from a basic database, converting the multi-omics data into binary variable data, and processing the H&E stained sections into WSI format image blocks;

[0055] The pre-trained HoverNet model is used to extract single-cell information from the WSI format image blocks, and the nuclear region corresponding to the single-cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block;

[0056] An AIHE multi-omics deep learning model is constructed using EfficientNet-B3 as the backbone network. WSI format image blocks containing regional labels of single cells are used as inputs of the AIHE multi-omics deep learning model. The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model.

[0057] The H&E-stained sections to be identified were input into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E-stained sections to be identified.

[0058] It should be understood that the AIHE multi-omics deep learning model principle is:

[0059] A. Figure 2 The figure shows a common workflow for obtaining WSI data in current clinical practice. Both FFPE and FF sections can provide high-resolution WSI scans.

[0060] B. The contrast region-guided training process, where each image patch is considered to be a subset of the task-relevant (cell-related) and task-irrelevant background noise. To eliminate the influence of background noise at a higher feature level, a model backbone is used to extract features from image patches containing and not containing cells, and these features are subtracted. The difference between these features is ultimately used for multi-omics prediction.

[0061] C. During the training phase, AIHE uses tissue-level multi-omics data as training targets, which are considered as the average of a randomly sampled small cell subset.

[0062] D. In the inference stage, AIHE makes predictions for each individual cell while preserving the spatial information of the cell, thereby obtaining in situ multi-omics prediction data from histopathological images.

[0063] In the above embodiment, an AIHE multi-omics deep learning model for H&E stained sections is constructed. First, the multi-omics data is converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, and regional labels are set for the cells in the WSI format image blocks. Then, the WSI format image blocks containing the regional labels of single cells are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition. After training, the AIHE multi-omics deep learning model identifies the single-cell quantitative data corresponding to the H&E stained sections to be identified, thereby improving the recognition degree of H&E stained images and reducing research costs compared to existing immunohistochemistry (IHC) and immunofluorescence (IF).

[0064] Preferably, the H&E stained sections and their corresponding multi-omics data are obtained from a basic database, the multi-omics data are converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, specifically:

[0065] H&E-stained sections and their corresponding multi-omics data were obtained from two basic databases, the TCGA database and the GTEx database;

[0066] performing batch effect removal processing on the multi-omics data, performing median truncation processing on the multi-omics data after the batch effect removal processing, and classifying the multi-omics data after the median truncation processing using a multi-label binary classification method to obtain binary variable data;

[0067] The H&E stained sections are converted into WSI format images, and the WSI format images are optimized, and the optimized WSI format images are cut into blocks according to a set size and set pixels to obtain WSI format image blocks.

[0068] Specifically, this study was conducted in accordance with the Declaration of Helsinki. H&E-stained sections and their corresponding multi-omics data were obtained based on the TCGA database and the GTEx database, of which TCGA mainly includes tumor samples and GTEx mainly includes normal samples.

[0069] To reduce the error rate, the continuous variable multi-omics data was truncated at the median, thus converting the prediction problem into a multi-label binary classification problem (high expression level vs. low expression level). In addition, genes with too sparse expression in the transcriptome (expression level of 0 in more than half of the samples) and genes with mutation rate less than 3% were deleted and uniformly regarded as binary variables of whether they had mutated.

[0070] like Figure 2 As shown, for WSI format images, high-magnification scanned images with billions of pixels were obtained. In addition, all H&E stained slides and their corresponding multi-omics data were collected from the TCGA and GTEx multi-organ databases and converted into matrix format. As of December 2022, data collection has been completed, including all 29,886 WSIs in the TCGA database and all 25,713 WSIs in the GTEx database. It is worth noting that in order to enable the model to predict intraoperative pathological images, FFPE slides and FF slides in the TCGA database were included in the development. The conversion process is as follows: Resected Tissue is HE stained and sent to Tissue Slicer to obtain Paraffin (FFPE slide) or Fresh Frozen (FF slide), which is then sent to Microscopy Slide Scanner to obtain WSI format images.

[0071] The image was uniformly sliced ​​into 512x512 pixel blocks at a size of 256x256μm, resulting in a total of 110,463,388 image blocks. To accurately localize single cells, the pre-trained HoverNet was used to extract the position and morphology of all cell nuclei. Morphological expansion techniques were then used to expand the nuclear region to its adjacent regions, which served as the region label for each single cell.

[0072] Preferably, the pre-trained HoverNet model is used to extract single cell information from the WSI format image blocks, and the nuclear region corresponding to the single cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block, specifically:

[0073] Using the pre-trained HoverNet model, single-cell information is extracted from WSI format image blocks, including cell nucleus position and morphology information;

[0074] The nuclear region corresponding to the cell nucleus position and morphology information is expanded to the adjacent region by a morphological expansion operator to obtain the region label of each single cell in the WSI format image block. The morphological expansion operator is:

[0075]

[0076] The morphological expansion operator is represented by expanding each point x in the WSI format image block A to (B) x , recorded as As a region label set, the region label set includes a region label of each single cell in a WSI format image block.

[0077] Preferably, the WSI format image blocks containing the regional labels of single cells are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data are used as the training target of the AIHE multi-omics deep learning model, and the AIHE multi-omics deep learning model is trained for multi-omics data recognition, specifically as follows:

[0078] Assume that the WSI format image block contains m cells, randomly mask mn cells in the WSI format image block as background noise in the feature dimension, and the remaining n cells are trained using the corresponding region labels;

[0079] The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model, and the model parameters of the EfficientNet-B3 backbone network are updated using the AdamW optimizer with a pre-set learning rate, including:

[0080] During each round of training, the WSI format image block is copied into two images, namely img1 and img0. In img1, all cells are masked by the regional label of each single cell, so that img1 contains background noise unrelated to the cell, while img0 retains n cells and masks other cells, so that img0 contains background noise and information of n cells.

[0081] Input the img1 image and the img0 image into the EfficientNet-B3 backbone network, extract the corresponding image features, record them as feature1 and feature0, and calculate the feature difference between feature1 and feature0 to obtain the difference between the two features;

[0082] The difference is input into a fully connected layer connected to the output of the EfficientNet-B3 backbone network to output various types of single-cell quantitative data. Specifically, the fully connected layer is a binary classifier.

[0083] In the above example, to mitigate the impact of noise, the information in the H&E image is divided into two parts: the single-cell components relevant to the prediction task and the irrelevant background noise. A contrast region guidance strategy is used to guide the model to focus on elements relevant to the task.

[0084] Specifically, during training, mn cells are randomly masked in image patches containing m cells, thus treating them as background noise in the feature dimension. The remaining n cells are trained directly using WSI-level labels, where n values ​​are randomly sampled during training. The model architecture uses EfficientNet-B3 as the backbone network. Given the noisy nature of the training data labels, the AdamW optimizer uses a learning rate of 1e-5.

[0085] like Figure 3 As shown in the figure, during the training phase, mn cells in an image patch containing m cells are randomly masked, thereby excluding them as background noise in the feature dimension. The remaining n cells are directly trained using WSI-level labels, allowing the model to learn the average expression of these n cells, where n values ​​are randomly sampled during training. This training process allows the model to learn the statistical differences in cell distribution between different samples under some noisy labels, while also being able to make predictions for a small number of cells or individual cells.

[0086] In terms of model architecture, EfficientNet-B3 was used as the backbone network. Unlike previous studies, this backbone network was incorporated into the training process to develop a molecular feature extractor for H&E-stained cell images. Due to the noisy nature of the training data labels, the model was trained on all image patches in just one epoch using the AdamW optimizer with a learning rate of 1e-5 to avoid memory constraints and overfitting issues associated with multiple epochs of training.

[0087] AdamW is an optimizer that uses the gradient descent principle and has been packaged into the open-source deep learning framework PyTorch. The required parameters are a learning rate of 1e-5 and the number of training cycles. The specific optimization process is encapsulated in the software package, which means that the model parameters are continuously updated using data during training.

[0088] like Figure 3-5 As shown in the figure, the model training process is:

[0089] During each round of training, a HE patch and its corresponding cell region are sampled and subjected to morphological expansion (i.e., miorphology expansion). This image is then copied into two copies, img0 and img1. img1 masks out all cells using the aforementioned cell region (i.e., mask-out), while img0 retains n cells and masks out the rest. At this point, img1 contains only background noise unrelated to cells (i.e., HE Patch without Cells—Background Noise), while img0 contains background noise and information about n cells (i.e., HE Patch with n Cells Left). Both are then passed through the EfficientNet backbone network to extract corresponding image features, denoted as feature0 (Cell Features) and feature1 (Background Features). Feature1 is then subtracted from feature0 to obtain the difference between the two features. This subtraction removes the background noise information shared by both features, resulting in the difference representing the unmasked cell information, which is the average information of the n cells mentioned above. This difference feature will be fed into the subsequent fully connected layer to predict multi-omics information, including transcriptome, mRNA, miRNA, proteome, methylation, single nucleotide polymorphism, copy number variation (Transcriptome, Proteomics, Methylomics, Single-Nucleotide Polymorphism, Copy Number Variatior) and other information.

[0090] In addition to directly predicting high-dimensional multi-omics data, models pre-trained on these data have become the foundation for models used in H&E-stained images. Due to their rich molecular-level annotation information, they achieve higher performance in downstream tasks compared to current models based on unsupervised pre-training. When applied to downstream prediction tasks, the classification layer used for multi-omics prediction is removed and replaced with a new classification layer tailored for the specific task. Furthermore, the model is fine-tuned for one cycle on the target data using a low learning rate of 1e-6.

[0091] like Figure 4-Figure 5As shown in the figure, the hematoxylin and eosin staining artificial intelligence model (i.e., AIHE multi-omics deep learning model) is a multi-omics deep learning model for H&E staining images. This model extends the multi-omics data analysis that was previously only feasible at the patient or tissue level to the single-cell level, enabling the prediction of multi-omics quantitative data, including transcriptome, mRNA, miRNA, proteome, methylation, single nucleotide polymorphism, copy number variation (Transcriptome, Proteomics, Methylomics, Single-Nucleotide Polymorphism, Copy Number Variatior) and other information. SNPs and copy number variations (CNVs) are performed on each cell using only H&E staining images. A comprehensive benchmark test was conducted on the spatial transcriptome data of clinically collected tumor tissue samples, demonstrating the excellent predictive ability of AIHE. At the same time, since H&E staining images are the most commonly used data type in clinical pathology practice, the model can be seamlessly integrated with the pathology workflow, and multi-omics prediction results can be obtained within tens of minutes after acquiring WSI data. When combined with FF sections, it can provide immediate multi-omics predictions during surgery, enabling real-time guidance and intervention during the surgical process. Furthermore, AIHE has demonstrated its ability to identify significant changes in gene expression within pathologist-annotated spatial domains and predict the enrichment of specific genes in complex biological tissues. This breakthrough unlocks the biological information implicit in H&E staining images, making a significant contribution to researchers' ability to reveal the multi-level molecular characteristics and underlying molecular mechanisms of disease.

[0092] Next, the reliability and feasibility of this model are demonstrated through the application of the AIHE multi-omics deep learning model.

[0093] 1) AIHE performs well in clinical prognosis:

[0094] Based on the TCGA and GTEx databases, multi-omics data and corresponding clinical information of the colon cancer (COAD) clinical cohort were obtained. Figure 7-Figure 8As shown in the figure, risk_score represents the high- and low-score groups corresponding to the risk score, and thresholdprobability represents the threshold probability (or risk probability). The AIHE model's prognostic assessment for both tumors and non-tumor patients was analyzed using receiver operating characteristic (ROC) curves, calibration curves, and direct current analysis (DCA) curves. The results showed that the AIHE model had an area under the curve (AUC) of approximately 1 for both tumors and non-tumor patients (AUC = 0.991). The model's calibration curve approximated the diagonal line. Within the threshold range of 0.1-0.8, the net clinical benefit of intervention based on the model's predicted probability exceeded the net clinical benefit of no treatment or total treatment, confirming the excellent diagnostic performance of the AIHE model for both COAD and non-COAD samples. Furthermore, the AIHE model demonstrated relatively excellent predictive performance for COAD lymph node metastasis and distant metastasis, with AUC values ​​of 0.857 and 0.801, respectively. The model's calibration curve approximated the diagonal line. Further survival analysis demonstrated that the AIHE model had a significant prognostic impact on recurrence-free survival (RFS) and overall survival (OS) in the TCGA cohort samples from this study, with patients exhibiting a high risk score exhibiting a poorer prognosis. In summary, these analyses preliminarily validated the prognostic efficacy of the AIHE model in a clinical cohort.

[0095] 2) AIHE can accurately identify gene mutations, making it a candidate alternative to panel sequencing:

[0096] Next, we evaluated the predictive performance of AIHE on the mutation spectrum. First, we compared the mutation frequencies of genes in the TCGA-COAD cohort samples with those predicted by AIHE, and found that they were highly similar, such as Figure 9 AIHE has shown excellent predictive performance in mainstream commercial cancer gene mutation detection panels, including MSK-IMPACT (MSK) and FoundationOneCDx (F1CDX). Figure 10 In addition, the AUC values ​​of AIHE for predicting MYH4, GCN1L1, and F5 gene mutations were all higher than 0.8, and the overall AUC value for predicting all gene mutations exceeded 0.7, as shown in Figure 2. Figure 11 This indicates that AIHE has the potential to accurately identify gene mutations and preliminarily confirms its accuracy in predicting gene mutations.

[0097] AIHE also provides transcriptomic predictions of WSI, which is very useful in various clinical settings. To demonstrate the clinical application of this transcriptomic prediction, TMB and MSI in the TCGA-COAD cohort were directly examined using WSI as a diagnostic use case. The composite ROC-calibration-DCA triplet curve showed that AIHE's TMB and MSI predictions were very accurate, with AUC values ​​of 0.953 and 0.864, respectively. Figure 12-13In summary, this analysis demonstrates that AIHE is comparable to mainstream commercial cancer gene mutation detection panels in predicting mutations, TMB, and MSI in tumor samples. It can accurately identify gene mutations and has the potential to become a viable alternative to panel sequencing.

[0098] 3) Based on AIHE, spot-level spatial transcriptome data can be generated:

[0099] Figure 14 -A is the H&E staining, spatial annotation clusters, spatial region annotations and UMAP annotation maps of the COAD cohort, as shown in Figure 14 As shown in Figure 1, the spatial transcriptome of the TCGA-COAD cohort (TCGA-AZ-4313) was predicted using AIHE and annotated into multiple spatial regions based on the expression of different marker genes. In addition, the expression of marker genes was spatialized and densitized and compared with H&E staining images identified by pathologists to verify the accuracy of AIHE prediction.

[0100] like Figure 14 As shown, Figure 14 -B is the expression of tumor area markers (EPCAM and MYC) in the COAD cohort; Figure 14 -C is the expression of cancer stem cell region markers (POU5F1 and CD24) in the COAD cohort; Figure 14 -D is the expression of angiogenic area markers (CD34 and VEGFA) in the COAD cohort; Figure 14 -E is the expression of immune region markers (CD4 and CD8A) in the COAD cohort; this includes tumor regions (EPCAM, MYC) (e.g. Figure 14 -B), tumor stem regions (POU5F1, CD24) (as shown Figure 14 -C), angiogenesis area (CD34, VEGFA) (as shown Figure 14 -D) and immune regions (CD4, CD8) (as shown Figure 14 These results demonstrate the high performance of AIHE in predicting spatial transcriptome expression.

[0101] like Figure 15 As shown, Figure 15 -A is the H&E staining, spatial annotation clusters, spatial region annotations, and UMAP annotation maps of the COAD cohort; Figure 15 -B is the expression of TLS regional markers (CR2 and CD38) in the COAD cohort; Figure 15-C is the spatial true annotation map and AIHE predicted annotation map of the BRCA cohort, which shows the correlation between the true expression values ​​of BRCA tumor-related markers (PTBP1, GNAS and CRABP1) and the AIHE predicted expression values.

[0102] In addition, AIHE identified tertiary lymphoid structure (TLS) regions in the TCGA-COAD cohort (TCGA-AY-6196), such as Figure 15 -A, and the spatial regions of expression of their marker genes (CR2, CD38) overlapped with the H&E staining images, as shown in Figure 15 -B. The predictive performance of AIHE was further validated in the TCGA-Breast Cancer (TCGA-BRCA) cohort, and it was found that the spatial true annotation map and AIHE predicted annotation map of this cohort were significant, and the true expression values ​​of tumor-related markers (PTBP1, GNAS, and CRABP1) were significantly positively correlated with the AIHE predicted expression values ​​(p < 0.001). Figure 15 -C as shown.

[0103] Taken together, these results demonstrate that AIHE can generate spot-level spatial transcriptome data, accurately predict different functional regions within tumors, and help experts identify tumor lesions faster and more accurately.

[0104] 4) AIHE has the potential to learn spatial single-cell transcriptomes from H&E staining of pathological sections:

[0105] like Figure 16 As shown, Figure 16 -A is the spatial single-cell atlas of the COAD cohort (TCGA-AZ-4313), mapping spatial single-cell types; Figure 16 -B is the single-cell atlas of the COAD cohort (TCGA-AZ-4313), mapping spatial single cell types; Figure 16 -C is a spatial single-cell map of the expression of related genes (CD4, CD8A, CD34, VEGFA, ACTA2 and MYC); Figure 16 -D is the spatial single-cell atlas of the COAD cohort (TCGA-AY-6196), mapping spatial single-cell types; Figure 16 -E is the single-cell atlas of the COAD cohort (TCGA-AY-6196), mapping spatial single cell types; Figure 16 -F is a spatial single-cell map of the expression of related genes (CD38, MS4A1, CD4 and CD8A).

[0106] We further explored the ability of AIHE to learn spatial single-cell transcriptomes from H&E-stained pathological sections. AIHE predicted single-cell transcriptome expression in the TCGA-COAD cohort using H&E-stained images. Dimensionality reduction clustering was then used to create a single-cell atlas, onto which the expression of region-specific marker genes was mapped. The results showed that in the TCGA-COAD cohort, the single-cell transcriptomes predicted by AIHE clustered into cell clusters consistent with the functional regions of the spatial transcriptome. In addition, the annotated cell ecology of different samples showed significant differences. For example, CD4 T cells, CD8 T cells, colon cancer cells, endothelial cells, and fibroblasts were annotated in the TCGA-AZ-4313 sample, as shown in Figure 2. Figure 16 AC shown; B cell samples were also identified in TCGA-AY-6196, such as Figure 16 As shown in DF, this indicates that different H&E staining images can reflect different cell ecosystems.

[0107] like Figure 17 As shown, Figure 17 -A is the H&E staining, spatial annotation clusters, spatial region annotations, and UMAP annotation images of the COAD cohort (TCGA-3L-AA1B). Figure 17 -B is the expression of angiogenesis markers (CD34 and VEGFA) in the COAD cohort (TCGA-3L-AA1B); Figure 17 -C is the expression of fibrosis markers (ACTA2 and FAP) in the COAD cohort (TCGA-3L-AA1B); Figure 17 -D is the spatial single-cell atlas of the COAD cohort (TCGA-3L-AA1B), mapping spatial single-cell types; Figure 17 -E is the single-cell atlas of the COAD cohort (TCGA-3L-AA1B), mapping spatial single cell types; Figure 17 -F is a spatial single-cell map of the expression of related genes (CD34, VEGFA, ACTA2 and FAP).

[0108] In addition, the spatial regions of immune cells, tumors, and cancer stem cells were annotated in the TCGA-3L-AA1B sample, e.g. Figure 17 AC. AIHE prediction further subdivides these regions into different cell clusters, as shown in Figure 17 As shown in Figure DF, the accuracy of regional annotation is limited in areas with complex cellular ecology. However, further cell-level annotation by AIHE can overcome this limitation and enhance the understanding of the biological information inherent in H&E-stained images. In summary, these findings preliminarily confirm the potential of AIHE to learn spatial single-cell transcriptomes from H&E-stained pathological sections.

[0109] 5) AIHE can be used for spatial single-cell multi-omics prediction:

[0110] like Figure 18 As shown, Figure 18 -A is the H&E staining-spatial density map-IHC staining map showing the expression of PR, ER and HER2 protein levels in the TCGA breast cancer cohort. Figure 18 -B is H&E staining-spatial density map-IHC staining map showing the expression of proSPC and CD31 protein levels in the TCGA-lung adenocarcinoma cohort.

[0111] Based on H&E staining images, we used AIHE to predict multi-omics data, including spatial single-cell proteome, miRNA and methylation. In the TCGA-breast cancer cohort (25) ( Figure 18 A) and lung adenocarcinoma cohort (26) ( Figure 18 In B), AIHE-predicted single-cell proteome expression closely matches actual H&E expression.

[0112] like Figure 19 As shown, Figure 19 -A is a spatial single-cell atlas mapping spatial single-cell types and subpopulations, and a single-cell atlas mapping spatial single-cell types and subpopulations of the COAD cohort (TCGA-3L-AA1B); Figure 19 -B is the expression of spatial single-cell mapping proteins (PDL1, JAK2, SMAD1 and TIGAR); Figure 19 -C is a spatial single-cell map of miRNA expression (hsa-mir-16-1 and hsa-mir-421); Figure 19 -D is a spatial single-cell map of the expression of methylated genes (HOXA5 and CDKN2B); Figure 19 -E: spatial single-cell mapping of SNP (CTNNB1 and KRAS) expression; Figure 19 -F is a spatial single-cell atlas mapping CNV (JAK2 and CES1P1) expression.

[0113] Furthermore, in the TCGA-COAD cohort samples (TCGA-3L-AA1B), AIHE accurately matched the same cell types and subpopulations in both spatial single-cell maps and single-cell maps ( Figure 19 A). The markers that are highly expressed in these subpopulations have been validated in multiple omics studies. Notably, proteins such as PDL1, JAK2, SMAD1, and TIGAR are commonly expressed in colon cancer subpopulations ( Figure 19B). In addition, these subpopulations consistently showed high expression of miRNAs (hsa-mir-16-1 and hsa-mir-421), acetylated genes (HOXA5 and CDKN2B), SNPs (CTNNB1 and KRAS), and CNVs (JAK2 and CES1P1) ( Figure 19 CF). These results collectively indicate that AIHE also has the potential to predict other spatial single-cell omics.

[0114] With the rapid development of precision oncology, combining clinical features, imaging results, pathological examination and molecular testing is crucial for selecting appropriate clinical tumor treatment methods. However, the information in histological images is often complex and ambiguous, which poses a challenge for pathologists to perform robust, reproducible and efficient analysis. Artificial intelligence has been widely used in the analysis of histological images. AI can directly predict molecular changes in routine histopathology sections and effectively extract relevant clinical information. This ability has been successfully verified in many studies. These methods extend the use of H&E-stained tissue sections beyond traditional tumor diagnosis and subtype identification, and directly predict the source of molecular changes. The present invention successfully developed and verified AIHE, which can accurately predict multi-omics quantitative data of each cell in pathological images with high spatial resolution by combining spatial information with histological multi-view neighborhood features.

[0115] H&E-stained images are crucial for identifying different types of tissue and their morphological changes, particularly in cancer diagnosis. Therefore, examining H&E-stained images typically involves observing cells, and identifying cellular ecological signatures within these slides could significantly improve the causal interpretability and accuracy of AI models. Numerous studies have shown that tissue sections often contain a wide range of information, reflecting changes from the cellular level to the tissue, system, and even the entire body, representing the common effects and trends of cellular ecological changes. However, traditional AI analysis of H&E-stained pathology slides often lacks biological interpretability. The advantage of AIHE is that it extends multi-omics data analysis, previously feasible only at the patient or tissue level, to the single-cell level. This allows for the prediction of multi-omics quantitative data for each cell using only H&E-stained images. By leveraging the potential of the spatial single-cell transcriptome of H&E-stained pathology slides, AIHE eliminates the need for additional sequencing processes, significantly reducing treatment costs for patients.

[0116] Furthermore, AIHE has demonstrated the ability to identify significant changes in gene expression within spatial domains annotated by pathologists and predict the enrichment of specific genes in complex biological tissues. The present invention uses AIHE to generate spatial transcriptome data for COAD samples and annotate them into multiple spatial regions based on different marker genes. AIHE's predictions are highly consistent with the annotations of clinical pathologists. Previous scanning flow cytometry experiments have shown that distinguishing between two types of T cells or B cells based solely on morphology is very challenging. AIHE can spatially resolve genes specifically expressed by T cells or B cells, thereby distinguishing these cells. In addition to immune cells, AIHE can also identify angiogenesis areas, fibroblast areas, tumor areas, and cancer stem cell areas, providing virtual multiplex staining for WSI. This capability can make AIHE a major tool for medical diagnosis and prediction. Through real-time prediction of intraoperative frozen sections, AIHE enables surgeons to understand the nature of the lesion during surgery, thereby helping to promptly determine the scope of surgery and make appropriate treatment decisions, helping to improve the accuracy and safety of the surgical procedure.

[0117] In summary, it has been demonstrated that AIHE can not only predict multi-omics quantitative data for each cell from H&E staining images, but can also be used to predict patient prognosis, gene mutations, TMB, and MSI status, thereby assisting physicians in assessing disease progression and assisting immunotherapy. In addition, AIHE can also generate point-level spatial transcriptome data from H&E staining images, exploring the potential of spatial single-cell multi-omics, thereby unlocking the biological information embedded in H&E staining images and enabling more patients to receive precision medicine. Despite these advances, this study still has certain limitations: all data generation is based on predictive simulations, and the resulting bioinformatics analysis results can only be presented descriptively. Therefore, all research results should be experimentally verified to improve the accuracy and scientific validity of the conclusions.

[0118] like Figure 6 As shown, an embodiment of the present invention further provides a device for spatial single-cell gene multi-omics quantitative prediction, comprising:

[0119] An image preprocessing module is used to obtain H&E stained sections and their corresponding multi-omics data from a basic database, convert the multi-omics data into binary variable data, and process the H&E stained sections into WSI format image blocks;

[0120] A label processing module is used to extract single cell information from the WSI format image blocks using the pre-trained HoverNet model, and expand the nuclear region corresponding to the single cell information to the adjacent region to obtain the regional label of each single cell in the WSI format image block;

[0121] A model training module is used to build an AIHE multi-omics deep learning model using EfficientNet-B3 as the backbone network, using WSI format image blocks containing single-cell regional labels as input to the AIHE multi-omics deep learning model, and using the converted binary variable data as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model.

[0122] The identification module is used to input the H&E stained slice to be identified into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E stained slice to be identified.

[0123] Preferably, in the image preprocessing module, H&E stained sections and their corresponding multi-omics data are obtained from the basic database, the multi-omics data are converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, specifically:

[0124] H&E-stained sections and their corresponding multi-omics data were obtained from two basic databases, the TCGA database and the GTEx database;

[0125] performing batch effect removal processing on the multi-omics data, performing median truncation processing on the multi-omics data after the batch effect removal processing, and classifying the multi-omics data after the median truncation processing using a multi-label binary classification method to obtain binary variable data;

[0126] The H&E stained sections are converted into WSI format images, and the WSI format images are optimized, and the optimized WSI format images are cut into blocks according to a set size and set pixels to obtain WSI format image blocks.

[0127] Preferably, in the label processing module, a pre-trained HoverNet model is used to extract single cell information from the WSI format image blocks, and the nuclear region corresponding to the single cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block, specifically:

[0128] Using the pre-trained HoverNet model, single-cell information is extracted from WSI format image blocks, including cell nucleus position and morphology information;

[0129] The nuclear region corresponding to the cell nucleus position and morphology information is expanded to the adjacent region by a morphological expansion operator to obtain the region label of each single cell in the WSI format image block. The morphological expansion operator is:

[0130]

[0131] The morphological expansion operator is represented by expanding each point x in the WSI format image block A to (B) x , recorded as As a region label set, the region label set includes a region label of each single cell in a WSI format image block.

[0132] Preferably, in the model training module, the WSI format image blocks containing single cell region labels are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data are used as the training target of the AIHE multi-omics deep learning model to perform multi-omics data recognition training on the AIHE multi-omics deep learning model, specifically:

[0133] Assume that the WSI format image block contains m cells, randomly mask mn cells in the WSI format image block as background noise in the feature dimension, and the remaining n cells are trained using the corresponding region labels;

[0134] The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model, and the model parameters of the EfficientNet-B3 backbone network are updated using the AdamW optimizer with a pre-set learning rate, including:

[0135] During each round of training, the WSI format image block is copied into two images, namely img1 and img0. In img1, all cells are masked by the regional label of each single cell, so that img1 contains background noise unrelated to the cell, while img0 retains n cells and masks other cells, so that img0 contains background noise and information of n cells.

[0136] Input the img1 image and the img0 image into the EfficientNet-B3 backbone network, extract the corresponding image features, record them as feature1 and feature0, and calculate the feature difference between feature1 and feature0 to obtain the difference between the two features;

[0137] The difference is input into a fully connected layer connected to the output of the EfficientNet-B3 backbone network to output various types of single-cell quantitative data.

[0138] An embodiment of the present invention also provides a device for spatial single-cell gene multi-omics quantitative prediction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for spatial single-cell gene multi-omics quantitative prediction as described above is implemented.

[0139] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0140] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for spatial single-cell gene multi-omics quantitative prediction, characterized in that: The steps include: Obtaining H&E stained sections and their corresponding multi-omics data from a basic database, converting the multi-omics data into binary variable data, and processing the H&E stained sections into WSI format image blocks; The pre-trained HoverNet model is used to extract single-cell information from the WSI format image blocks, and the nuclear region corresponding to the single-cell information is expanded to the adjacent region to obtain the regional label of each single cell in the WSI format image block; An AIHE multi-omics deep learning model is constructed using EfficientNet-B3 as the backbone network. WSI format image blocks containing regional labels of single cells are used as inputs of the AIHE multi-omics deep learning model. The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model. Input the H&E-stained slice to be identified into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E-stained slice to be identified; The H&E stained sections and their corresponding multi-omics data are obtained from the basic database, the multi-omics data are converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, specifically: H&E-stained sections and their corresponding multi-omics data were obtained from two basic databases, the TCGA database and the GTEx database; performing batch effect removal processing on the multi-omics data, performing median truncation processing on the multi-omics data after the batch effect removal processing, and classifying the multi-omics data after the median truncation processing using a multi-label binary classification method to obtain binary variable data; Converting the H&E stained sections into WSI format images, optimizing the WSI format images, and dicing the optimized WSI format images according to a set size and set pixels to obtain WSI format image blocks; The pre-trained HoverNet model is used to extract single cell information from the WSI format image blocks, and the nuclear area corresponding to the single cell information is expanded to the adjacent area to obtain the regional label of each single cell in the WSI format image block, specifically: Using the pre-trained HoverNet model, single-cell information is extracted from WSI format image blocks, including cell nucleus position and morphology information; The nuclear region corresponding to the cell nucleus position and morphology information is expanded to the adjacent region by a morphological expansion operator to obtain the region label of each single cell in the WSI format image block. The morphological expansion operator is: The morphological expansion operator is represented by expanding each point x in the WSI format image block A to (B) x , recorded as As a region label set, the region label set includes a region label of each single cell in a WSI format image block; The WSI format image blocks containing the regional labels of single cells are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data are used as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition, specifically as follows: Assume that the WSI format image block contains m cells, randomly mask mn cells in the WSI format image block as background noise in the feature dimension, and the remaining n cells are trained using the corresponding region labels; The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model, and the model parameters of the EfficientNet-B3 backbone network are updated using the AdamW optimizer with a pre-set learning rate, including: During each round of training, the WSI format image block is copied into two images, namely img1 and img0. In img1, all cells are masked by the regional label of each single cell, so that img1 contains background noise unrelated to the cell, while img0 retains n cells and masks other cells, so that img0 contains background noise and information of n cells. Input the img1 image and the img0 image into the EfficientNet-B3 backbone network, extract the corresponding image features, record them as feature1 and feature0, and calculate the feature difference between feature1 and feature0 to obtain the difference between the two features; The difference is input into a fully connected layer connected to the output of the EfficientNet-B3 backbone network to output various single-cell quantitative data; The fully connected layer is a binary classifier.

2. A device for spatial single-cell gene multi-omics quantitative prediction, characterized by: include: An image preprocessing module is used to obtain H&E stained sections and their corresponding multi-omics data from a basic database, convert the multi-omics data into binary variable data, and process the H&E stained sections into WSI format image blocks; A label processing module is used to extract single cell information from the WSI format image blocks using the pre-trained HoverNet model, and expand the nuclear region corresponding to the single cell information to the adjacent region to obtain the regional label of each single cell in the WSI format image block; A model training module is used to build an AIHE multi-omics deep learning model using EfficientNet-B3 as the backbone network, using WSI format image blocks containing single-cell regional labels as input to the AIHE multi-omics deep learning model, and using the converted binary variable data as the training target of the AIHE multi-omics deep learning model. The AIHE multi-omics deep learning model is trained for multi-omics data recognition to obtain a trained AIHE multi-omics deep learning model. The recognition module is used to input the H&E stained slice to be identified into the trained AIHE multi-omics deep learning model to obtain single-cell quantitative data corresponding to the H&E stained slice to be identified; In the image preprocessing module, H&E stained sections and their corresponding multi-omics data are obtained from the basic database, the multi-omics data are converted into binary variable data, and the H&E stained sections are processed into WSI format image blocks, specifically: H&E-stained sections and their corresponding multi-omics data were obtained from two basic databases, the TCGA database and the GTEx database; performing batch effect removal processing on the multi-omics data, performing median truncation processing on the multi-omics data after the batch effect removal processing, and classifying the multi-omics data after the median truncation processing using a multi-label binary classification method to obtain binary variable data; Converting the H&E stained sections into WSI format images, optimizing the WSI format images, and dicing the optimized WSI format images according to a set size and set pixels to obtain WSI format image blocks; In the label processing module, the pre-trained HoverNet model is used to extract single cell information from the WSI format image blocks, and the nuclear area corresponding to the single cell information is expanded to the adjacent area to obtain the regional label of each single cell in the WSI format image block, specifically: Using the pre-trained HoverNet model, single-cell information is extracted from WSI format image blocks, including cell nucleus position and morphology information; The nuclear region corresponding to the cell nucleus position and morphology information is expanded to the adjacent region by a morphological expansion operator to obtain the region label of each single cell in the WSI format image block. The morphological expansion operator is: The morphological expansion operator is represented by expanding each point x in the WSI format image block A to (B) x , recorded as As a region label set, the region label set includes a region label of each single cell in a WSI format image block; In the model training module, the WSI format image blocks containing single cell region labels are used as the input of the AIHE multi-omics deep learning model, and the converted binary variable data are used as the training target of the AIHE multi-omics deep learning model to perform multi-omics data recognition training on the AIHE multi-omics deep learning model, specifically: Assume that the WSI format image block contains m cells, randomly mask mn cells in the WSI format image block as background noise in the feature dimension, and the remaining n cells are trained using the corresponding region labels; The converted binary variable data is used as the training target of the AIHE multi-omics deep learning model, and the model parameters of the EfficientNet-B3 backbone network are updated using the AdamW optimizer with a pre-set learning rate, including: During each round of training, the WSI format image block is copied into two images, namely img1 and img0. In img1, all cells are masked by the regional label of each single cell, so that img1 contains background noise unrelated to the cell, while img0 retains n cells and masks other cells, so that img0 contains background noise and information of n cells. Input the img1 image and the img0 image into the EfficientNet-B3 backbone network, extract the corresponding image features, record them as feature1 and feature0, and calculate the feature difference between feature1 and feature0 to obtain the difference between the two features; The difference is input into a fully connected layer connected to the output of the EfficientNet-B3 backbone network to output various types of single-cell quantitative data.

3. A device for spatial single-cell gene multi-omics quantitative prediction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for spatial single-cell gene multi-omics quantitative prediction as described in claim 1 is implemented.

Citation Information

Patent Citations

  • Medical image feature recognition prediction model

    CN112990214A

  • Method for revealing cellular molecular mechanism of occurrence and development of glaucoma

    CN118443950A