A tumor tissue chip high-throughput detection method, device, equipment and medium

CN122761978APending Publication Date: 2026-09-15贵州图辑科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610626731.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0006]本申请提供了一种肿瘤组织芯片高通量检测方法、装置、设备及介质,解决了现有TMA技术自动化不足,复用性和拓展性受限的技术问题

Benefits of technology

本申请中,提供了一种肿瘤组织芯片高通量检测方法、装置、设备及介质,法通过预训练图像—基因表达模型,由H&E全切片直接预测空间转录组并生成虚拟IHC图,自动筛选候选取芯区域、生成坐标集并映射至石蜡块,结合编码规则实现取芯、染色、扫描、解码及定量分析的全流程自动化,无需人工目测选区;基于通用H&E图像预训练模型,仅需更换目标生物标志物面板即可生成新的虚拟IHC图和取芯规则;通过构建含不同患者、多种标志物的混合TMA块及标准化编码体系,实现同一TMA平台在多基因面板、多预测模型及多队列间的灵活复用与扩展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122761978A_ABST
    Figure CN122761978A_ABST
Patent Text Reader

Abstract

The application discloses a tumor tissue chip high-throughput detection method, device, equipment and medium. The method directly predicts spatial transcriptome from H&E full section through a pre-trained image-gene expression model, generates a virtual IHC graph, automatically screens a candidate core region, generates a coordinate set and maps to a paraffin block, and realizes full-process automation of core taking, staining, scanning, decoding and quantitative analysis in combination with a coding rule, without manual visual selection. Based on a general H&E image pre-training model, a new virtual IHC graph and core taking rule can be generated only by replacing a target biomarker panel. Through construction of a mixed TMA block containing different patients and multiple markers and a standardized coding system, flexible reuse and expansion of the same TMA platform among multiple gene panels, multiple prediction models and multiple queues are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of digital pathology, spatial omics and molecular pathology detection technology, and in particular to a high-throughput detection method, device, equipment and medium for tumor tissue microarrays. Background Technology

[0002] Spatial transcriptomics (ST) technology can measure the spatial expression profiles of hundreds to thousands of genes on tissue sections, providing important information for tumor microenvironment analysis and precision medicine. However, existing commercial ST platforms suffer from high costs, complex processes, and limited throughput, making it difficult to routinely perform on large-scale clinical samples.

[0003] In recent years, predicting spatial transcriptome data from conventional hematoxylin and eosin (H&E) stained sections using deep learning models has become a research hotspot. Methods such as DiffGene, which train image-gene joint models on paired H&E-spatial transcriptome data (e.g., Xenium, Visium), can simultaneously model pathological image features and multi-gene expression in a unified latent space. This allows for accurate regression of sequenced genes and extrapolation completion across different gene panels, thus obtaining a "virtual spatial transcriptome" or "virtual IHC" map. These methods significantly reduce experimental costs and make it possible to obtain approximate ST information from conventional H&E whole sections.

[0004] Despite the progress made in virtual transcriptomics technology, the following shortcomings still exist: 1. Insufficient automation in TMA technology: Although the existing tumor tissue microarray (TMA) technology can arrange multiple tissue microarrays on a single slide for parallel staining, the microarray area still mainly relies on manual visual selection by pathologists. It lacks automatic docking and coding management with virtual spatial omics results, making it difficult to form an end-to-end automated workflow. 2. Limited reusability and scalability: Existing TMA designs are usually only designed for a single cohort or a single task, making it difficult to flexibly reuse and scale to new gene panels or new prediction models.

[0005] Therefore, there is an urgent need for a high-throughput method that tightly integrates virtual spatial transcriptome prediction, automatic region selection, TMA design and encoding, and IHC batch validation to achieve systematic validation of multiple patients and multiple biomarkers, and significantly reduce experimental costs and manpower input. Summary of the Invention

[0006] This application provides a high-throughput detection method, device, equipment, and medium for tumor tissue microarrays, which solves the technical problems of insufficient automation, limited reusability, and limited scalability of existing TMA technology.

[0007] In view of this, the first aspect of this application provides a high-throughput detection method for tumor tissue microarrays, the method comprising: Step S101: Pre-train an image-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; Step S102: Use the H&E whole slice to be tested as input to the image-gene expression prediction model to generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. Step S103: Based on the preset target biomarker panel, generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map; Step S104: In the virtual IHC expression intensity map, candidate core regions for each target biomarker are selected according to the preset core retrieval rules, and a first coordinate set is generated based on the spatial coordinates of the candidate core regions. The first coordinate set is then used to establish a first correspondence between the patient number, biomarker number, and candidate priority. Step S105: After mapping the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and coding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as coding rules in the database. Step S106: After core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, construct several mixed TMA blocks containing different patients under the same marker number, and perform immunohistochemical staining on each mixed TMA block. Step S107: After scanning each mixed TMA block after staining to obtain the corresponding real IHC image, decode the real IHC image according to the encoding rules to obtain the correspondence between the core well position in each mixed TMA block and the patient number and biomarker number. Step S108: After performing image analysis on each core well location to obtain the actual IHC quantitative results, the actual IHC quantitative results are evaluated and compared with the virtual IHC expression intensity map of the corresponding patient number to generate a verification report containing multiple patients and multiple target biomarkers.

[0008] Optionally, step S101 specifically includes: In the training set of H&E whole slices and spatial transcriptome data with back-to-back relationships, the H&E whole slices are divided into several image patches, and the spatial transcriptome sequencing points located in the same image patch are aggregated to obtain pseudobulk gene expression vectors. Based on several image patches and their corresponding pseudo-bulk gene expression vectors, an open gene embedding encoder and a diffusion-based image feature extraction network are pre-trained to obtain an image-gene expression prediction model. The open gene embedding encoder has permutation invariance to gene sets of different panels.

[0009] Optionally, step S103 further includes: The virtual IHC expression intensity map was sequentially normalized, smoothed, thresholded, and then pseudo-colorized. The virtual IHC expression intensity maps of each gene are clustered according to the virtual expression vectors of multiple genes to obtain the virtual molecular region division.

[0010] Optionally, the preset core extraction rules in step S104 include at least one of the following: high expression threshold rule, expression gradient boundary rule, tumor-stromal junction rule, spatial clustering anomaly rule, or virtual molecular region boundary alignment rule.

[0011] Optionally, the immunohistochemical staining of each mixed TMA block in step S106 specifically involves performing immunohistochemical staining on each mixed TMA block within the same batch according to the target biomarker type.

[0012] Optionally, the image analysis of each core hole position in step S108 to obtain the true IHC quantitative result is specifically as follows: Image analysis was performed on each well location, and at least one quantitative indicator among IHC staining intensity, positive staining area ratio, average optical density, or H-score was calculated to obtain the true IHC quantitative results.

[0013] Optionally, the comparison and evaluation of the actual IHC quantitative results with the virtual IHC expression intensity map corresponding to the patient number in step S108 specifically includes: The real IHC quantitative results are paired with virtual IHC expression intensity maps corresponding to patient numbers and biomarker numbers. At least one evaluation index among correlation coefficient, mean square error, or classification consistency is calculated to obtain the validation results of the target biomarker in multiple patients.

[0014] A second aspect of this application provides a high-throughput detection device for tumor tissue microarrays, the device comprising: The pre-training unit is used to pre-train a picture-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; The first processing unit is used to take the H&E whole slice to be tested as input to the image-gene expression prediction model and generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. The second processing unit is used to generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map based on a preset target biomarker panel. The candidate core retrieval unit is used to screen candidate core retrieval regions for each target biomarker in the virtual IHC expression intensity map according to preset core retrieval rules, and generate a first coordinate set based on the spatial coordinates of the candidate core retrieval regions, and construct a first correspondence between the first coordinate set and the patient number, biomarker number and candidate priority. The encoding unit is used to map the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and encoding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as the encoding rule in the database. The hybrid TMA block construction unit is used to construct several hybrid TMA blocks containing different patients under the same marker number after core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, and to perform immunohistochemical staining on each hybrid TMA block. The decoding unit is used to scan each stained mixed TMA block to obtain the corresponding real IHC image, and then decode the real IHC image according to the encoding rules to obtain the correspondence between the core hole position in each mixed TMA block and the patient number and biomarker number. The verification unit is used to perform image analysis on each core well location, obtain the real IHC quantitative results, evaluate and compare the real IHC quantitative results with the virtual IHC expression intensity map of the corresponding patient number, and generate a verification report containing multiple patients and multiple target biomarkers.

[0015] A third aspect of this application provides a high-throughput detection device for tumor tissue microarrays, the device comprising a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the steps of the high-throughput detection method for tumor tissue microarrays as described in the first aspect above, according to the instructions in the program code.

[0016] A fourth aspect of this application provides a computer-readable storage medium for storing program code for executing the high-throughput tumor tissue microarray detection method described in the first aspect above.

[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: This application provides a high-throughput detection method, device, equipment, and medium for tumor tissue microarrays. The method uses a pre-trained image-gene expression model to directly predict the spatial transcriptome from H&E whole slides and generate a virtual IHC map. It automatically selects candidate core regions, generates coordinate sets, and maps them to paraffin blocks. Combined with coding rules, it automates the entire process of core extraction, staining, scanning, decoding, and quantitative analysis without the need for manual visual selection. Based on a general H&E image pre-trained model, only the target biomarker panel needs to be changed to generate a new virtual IHC map and core extraction rules. By constructing a hybrid TMA block containing different patients and multiple biomarkers and a standardized coding system, the same TMA platform can be flexibly reused and expanded across multiple gene panels, multiple prediction models, and multiple cohorts. Attached Figure Description

[0018] Figure 1 This is a flowchart of the high-throughput detection method for tumor tissue microarrays in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the high-throughput tumor tissue chip detection device in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the high-throughput tumor tissue chip detection device in the embodiments of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0020] For easier understanding, please refer to Figure 1 , Figure 1 This is a flowchart of the high-throughput detection method for tumor tissue microarrays in the embodiments of this application, as follows: Figure 1 As shown, specifically: Step S101: Pre-train an image-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; Step S101 specifically includes: In the training set of H&E whole slices and spatial transcriptome data with back-to-back relationships, the H&E whole slices are divided into several image patches, and the spatial transcriptome sequencing points located in the same image patch are aggregated to obtain pseudobulk gene expression vectors. Based on several image patches and their corresponding pseudo-bulk gene expression vectors, an open gene embedding encoder and a diffusion-based image feature extraction network are pre-trained to obtain an image-gene expression prediction model. The open gene embedding encoder has permutation invariance to gene sets of different panels.

[0021] It should be noted that paired H&E-stained whole-slice images (WSI) and spatial transcriptome (ST) data were collected. H&E (hematoxylin-eosin staining) is the gold standard staining method for pathological diagnosis; hematoxylin stains the cell nucleus blue, while eosin stains the cytoplasm and matrix pink.

[0022] High-resolution WSI images (typically 100,000 x 100,000 pixels or more) are divided into small image blocks (e.g., 224 x 224 pixels) to facilitate deep learning processing.

[0023] The gene expression values ​​of ST sequencing spots falling within the same image patch are averaged to obtain a pseudo-bulk gene expression vector. ST technologies (such as 10x Genomics Visium) deploy thousands of capture regions with a diameter of 55 μm on the tissue, each region containing transcriptome information of 1-10 cells.

[0024] Genes are processed using a set input method, and permutation invariance is achieved through an attention mechanism—that is, changing the order of input genes does not affect the output. It supports any number and combination of gene panels without retraining the network structure.

[0025] Leveraging the powerful generative capabilities and fine-grained feature capture capabilities of the diffusion model, pathological morphological features related to gene expression are extracted from H&E images.

[0026] Step S102: Use the H&E whole slice to be tested as input to the image-gene expression prediction model to generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. It should be noted that the H&E WSI to be tested is input into the pre-trained model, and the model predicts multi-gene expression values ​​for each image patch, which are then recombined according to spatial location into a spatial transcriptomic map corresponding to the original tissue morphology. This map overlays molecular-level gene expression information while maintaining the morphological resolution of H&E.

[0027] Step S103: Based on the preset target biomarker panel, generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map; Step S103 also includes: The virtual IHC expression intensity map was sequentially normalized, smoothed, thresholded, and then pseudo-colorized. The virtual IHC expression intensity maps of each gene are clustered according to the virtual expression vectors of multiple genes to obtain the virtual molecular region division.

[0028] It should be noted that, based on the preset target biomarker panel (such as PD-L1, HER2, Ki-67, etc.), the predicted expression values ​​of the corresponding genes are extracted from the spatial transcriptome map to generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map.

[0029] The virtual IHC expression intensity map is processed sequentially by normalization, smoothing, thresholding, and pseudo-coloring to map gene expression values ​​to a preset range. Gaussian filtering is applied to eliminate prediction noise and maintain spatial continuity. Expression intensity cutoff is set to distinguish positive / negative regions. Expression intensity is mapped to color (e.g., blue - low expression, red - high expression) to generate a visual image similar to a real IHC.

[0030] Spatial clustering of multi-gene expression vectors (such as K-means and Leiden algorithms) can identify microenvironment regions with similar molecular characteristics (such as immune hot zones and metabolically active zones) to assist in subsequent core extraction strategies.

[0031] Step S104: In the virtual IHC expression intensity map, candidate core regions for each target biomarker are selected according to the preset core retrieval rules, and a first coordinate set is generated based on the spatial coordinates of the candidate core regions. The first coordinate set is then used to establish a first correspondence between the patient number, biomarker number, and candidate priority. The preset core extraction rules in step S104 include at least one of the following: high expression threshold rule, expression gradient boundary rule, tumor-stromal junction rule, spatial clustering anomaly rule, or virtual molecular region boundary alignment rule.

[0032] It should be noted that multi-dimensional core extraction rules are applied to the virtual IHC map to automatically select representative areas: By selecting the top 10% of expression values ​​using a high expression threshold rule, we can ensure the capture of strong positive signals. By expressing gradient boundary rules, core samples are taken at locations of drastic changes in expression intensity to study the transitional region of tumor heterogeneity; Core samples were taken at the tumor-stroma interface using the tumor-stromal boundary rules to study the interactions of the immune microenvironment. By using spatial clustering anomaly rules to select "abnormal islands" that are significantly different from the molecular characteristics of the surrounding area, drug-resistant clones can be discovered. By using virtual molecular region boundary alignment rules to extract cores at cluster boundaries, we can verify the differences in protein expression in different molecular subregions.

[0033] Generate the first coordinate set (pixel coordinates on the original H&E slice) and establish a 3D correspondence: Patient ID, Biomarker ID, and candidate priority (ranked based on rule matching degree).

[0034] Step S105: After mapping the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and coding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as coding rules in the database. It should be noted that since spatial transcriptome data usually come from adjacent slices (3-5 μm thick) of H&E sections, they need to be processed through: Tissue contour matching is used to extract the outer contour of the tissue for shape registration. Image registration algorithms (such as SIFT feature matching and deep learning registration networks) map the first coordinate set to the adjacent unstained slices (paraffin block sections used for TMA coring) to generate the second coordinate set.

[0035] The perforation layout is optimized based on the second coordinate set to avoid excessive concentration of perforations in the same patient (batch effect), balance the spatial distribution of each marker in the TMA block, and reserve control wells (positive / negative control tissues).

[0036] By constructing a second correspondence between TMA block row and column positions (such as A01, B03) and patient ID + marker ID + core retrieval serial number, and storing this second correspondence as an encoding rule in a database (such as SQL or a blockchain evidence storage system), the entire process can be ensured to be traceable.

[0037] Step S106: After core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, construct several mixed TMA blocks containing different patients under the same marker number, and perform immunohistochemical staining on each mixed TMA block. In step S106, the immunohistochemical staining of each mixed TMA block is specifically performed as follows: each mixed TMA block is subjected to immunohistochemical staining within the same batch according to the type of target biomarker.

[0038] It should be noted that a tissue microarray instrument (such as Beecher Instruments MTA-1) is used to extract cores (0.6-2.0 mm in diameter) from the donor paraffin block according to the second coordinate set, and then arranges them on the recipient block.

[0039] The same TMA block contains: different patients (multi-center cohort) and the same biomarker number (e.g., all PD-L1 samples are grouped into one block for easy staining in the same batch).

[0040] IHC staining is performed in batches according to the type of marker to eliminate batch effects such as antibody batch, incubation time, and color development conditions, thus ensuring the comparability of results.

[0041] Step S107: After scanning each mixed TMA block after staining to obtain the corresponding real IHC image, decode the real IHC image according to the encoding rules to obtain the correspondence between the core well position in each mixed TMA block and the patient number and biomarker number. It should be noted that by using digital pathology scanners (such as Leica Aperio and Hamamatsu NanoZoomer) to acquire high-resolution TMA images, and by recognizing the TMA grid layout and combining it with the coding rules in the database, each core image is automatically matched with the patient ID and marker ID, without the need for manual label verification.

[0042] Step S108: After performing image analysis on each core well location to obtain the actual IHC quantitative results, the actual IHC quantitative results are evaluated and compared with the virtual IHC expression intensity map of the corresponding patient number to generate a verification report containing multiple patients and multiple target biomarkers.

[0043] In step S108, image analysis is performed on the position of each core hole to obtain the actual IHC quantitative results, specifically as follows: Image analysis was performed on each well location, and at least one quantitative indicator among IHC staining intensity, positive staining area ratio, average optical density, or H-score was calculated to obtain the true IHC quantitative results.

[0044] Step S108 involves evaluating and comparing the actual IHC quantitative results with the virtual IHC expression intensity map corresponding to the patient number. The real IHC quantitative results are paired with virtual IHC expression intensity maps corresponding to patient numbers and biomarker numbers. At least one evaluation index among correlation coefficient, mean square error, or classification consistency is calculated to obtain the validation results of the target biomarker in multiple patients.

[0045] The technical solution of this application fundamentally solves the limitations of TMA technology in terms of manual dependence and reusability through a closed-loop design of "virtual prediction → intelligent core sampling → coding tracking → quantitative verification", and provides a scalable technical framework for high-throughput pathological biomarker verification.

[0046] Based on the high-throughput tumor tissue microarray detection method provided in the embodiments of this application, this application provides two application examples for explanation: Application Example 1: Data preparation and preprocessing: Human colorectal cancer tissue sections with paired H&E whole slices and Xenium spatial transcriptome data were selected. The H&E whole slices were divided into several non-overlapping image patches of 224×224 pixels at 20× magnification. Xenium probe points within each image patch were aggregated to generate pseudobulk gene expression vectors containing more than 500 genes.

[0047] Model architecture and training: An open gene embedding encoder and a conditional diffusion U-Net were jointly pre-trained. The open gene embedding encoder is permutation-invariant to gene sets from different panels, supporting arbitrary numbers and combinations of gene inputs. The conditional diffusion U-Net performs bidirectional modeling between the image space and the gene expression space, predicting multi-gene expression vectors given image patches and generating consistent image features given gene vectors, thereby strengthening the alignment between images and genes. The image-gene expression prediction model was pre-trained based on several image patches and their corresponding pseudo-bulk gene expression vectors.

[0048] The new diagnostic H&E whole slice is input into the trained image-gene expression prediction model to obtain the multi-gene prediction values ​​of each image patch on the whole slice, which are then recombined into a spatial transcriptome map containing the multi-gene expression predictions for each spatial location.

[0049] Based on a pre-defined panel of target biomarkers (such as CD24, SLAMF1, etc.), predicted expression values ​​of corresponding genes are extracted from the spatial transcriptome map to generate virtual IHC expression intensity maps for each gene. The virtual IHC expression intensity maps are then subjected to normalization, smoothing, thresholding, and pseudo-coloring. The virtual IHC expression intensity maps of each gene are clustered according to multi-gene virtual expression vectors to obtain virtual molecular region divisions.

[0050] In the virtual IHC expression intensity map, candidate core regions for each target biomarker are selected according to preset core extraction rules. These preset core extraction rules include at least one of the following: high expression threshold rule (selecting regions with top expression values), expression gradient boundary rule (core extraction at locations of drastic intensity changes), tumor-stromal boundary rule (prioritizing core extraction at the tumor invasion front), and spatial clustering anomaly rule (identifying anomalous regions with significantly different molecular characteristics). Threshold segmentation and connected component analysis are performed on the virtual IHC heatmap, and the coordinates of several candidate core center points are obtained through minimum spacing constraints, generating a first coordinate set. A first correspondence is established between the first coordinate set and the patient ID, biomarker ID, and candidate priority.

[0051] A low-magnification scan was performed on an unstained slide adjacent to the diagnostic H&E slide. Rigid or non-rigid image registration algorithms were used to map the first coordinate set to the unstained slide coordinate system, generating a second coordinate set. Based on the second coordinate set, the TMA layout and encoding rules for each target biomarker were determined: a 10×10 TMA grid layout was designed, with row numbers encoded as patient IDs, column numbers encoded as core retrieval numbers, and the entire set of core wells corresponding to a single biomarker number. A second correspondence was constructed between the row and column positions of each core well in the TMA block and the patient ID, biomarker number, and core retrieval number, and this second correspondence was stored in the database as an encoding rule. A core retrieval list was generated based on the encoding rule and imported into a semi-automatic core retrieval device.

[0052] Core samples were taken from FFPE (paraffin-embedded) blocks from different patients at specified coordinates. Core posts were then inserted into corresponding receiving paraffin blocks to construct CD24-specific hybrid TMA blocks (containing several core wells from different patients under the same biomarker number). The above process was repeated for other biomarkers to obtain multiple biomarker-specific hybrid TMA blocks. Immunohistochemical staining was performed on each hybrid TMA block within the same batch according to the target biomarker type to eliminate batch effects.

[0053] Scanning the stained mixed TMA blocks yields corresponding real IHC images. The analysis software automatically generates a TMA grid based on row and column coding, core spacing, and coding rules stored in the database, and automatically assigns a label of "Patient ID - Marker Number - Core Retrieval Number" to each core, thus achieving automatic decoding of the real IHC images.

[0054] Image analysis was performed on each well location, and quantitative indicators such as the proportion of DAB staining area, average optical density, and H-score were extracted through colorimetric thresholding to obtain the true IHC quantitative results. The true IHC quantitative results were paired with virtual IHC expression intensity maps of corresponding patient and biomarker numbers in the database. Evaluation indicators such as correlation coefficient and mean squared error were calculated to obtain the overall validation results of the biomarker in multiple patients, generating a validation report containing multiple patients and multiple target biomarkers. If significant deviations are found in certain patients or batches, tissue quality, staining batches, or model prediction shifts can be traced back and used for subsequent model retraining.

[0055] Application Example 2: Repeat the above process on other solid tumor cohorts such as breast cancer and lung cancer. Since the open gene embedding encoder in step S101 has permutation invariance, there is no need to retrain the network structure. It is only necessary to include the paired H&E-spatial transcriptome data of the new cancer type into the training set for fine-tuning, or directly apply the pre-trained model to the inference of the new cancer type, which can support virtual spatial transcriptome prediction for different tumor types.

[0056] Core samples of different diameters are selected based on the characteristics of the biomarkers and the degree of tumor heterogeneity. 0.6 mm diameter: Suitable for biomarkers with high homogeneity or small biopsy samples. 1.0 mm diameter: Suitable for tumors with high heterogeneity, ensuring representativeness. For tumors with higher heterogeneity, multiple core retrieval centers can be designed on the same patient (such as covering different spatial areas like the tumor core area, the invasion front area, and the peritumoral stroma area), and multiple points of information can be recorded by adding the core retrieval sequence dimension.

[0057] High-throughput comparative analysis across cohorts and tumor types is achieved through a unified encoding rule (row number = patient ID, column number = core sample number, block number = biomarker number) and automatic decoding software. The encoding rule is stored in a database, supporting flexible reuse and expansion when adding new cancer types or biomarker panels, without requiring the reconstruction of the TMA infrastructure.

[0058] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of the high-throughput tumor tissue chip detection device in the embodiments of this application, as shown below. Figure 2 As shown, specifically: Pre-training unit 201 is used to pre-train an image-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; The first processing unit 202 is used to take the H&E whole slice to be tested as input to the image-gene expression prediction model and generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. The second processing unit 203 is used to generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map based on a preset target biomarker panel. The candidate core retrieval unit 204 is used to screen candidate core retrieval areas of each target biomarker in the virtual IHC expression intensity map according to the preset core retrieval rules, and generate a first coordinate set based on the spatial coordinates of the candidate core retrieval areas, and construct a first correspondence between the first coordinate set and the patient number, biomarker number and candidate priority. The encoding unit 205 is used to map the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and encoding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as the encoding rule in the database. The hybrid TMA block construction unit 206 is used to construct several hybrid TMA blocks containing different patients under the same marker number after core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, and to perform immunohistochemical staining on each hybrid TMA block. The decoding unit 207 is used to scan each mixed TMA block after staining to obtain the corresponding real IHC image, and then decode the real IHC image according to the encoding rules to obtain the correspondence between the core hole position and the patient number and biomarker number in each mixed TMA block. The verification unit 208 is used to perform image analysis on each core well location, obtain the real IHC quantitative results, evaluate and compare the real IHC quantitative results with the virtual IHC expression intensity map of the corresponding patient number, and generate a verification report containing multiple patients and multiple target biomarkers.

[0059] Another embodiment of the present invention provides a high-throughput detection device for tumor tissue microarrays, such as... Figure 3 As shown, device 10 includes: One or more processors 110 and memory 120, Figure 3 The following description uses a processor 110 as an example. The processor 110 and the memory 120 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0060] Processor 110 is used to perform various control logics of device 10, and can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Furthermore, processor 110 can also be any conventional processor, microprocessor, or state machine. Processor 110 can also be implemented as a combination of computing devices, such as a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors combined with DSP and / or any other such configuration.

[0061] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the method for constructing the multilingual phoneme representation model in this embodiment of the invention. The processor 110 executes various functional applications and data processing of the device 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, thereby implementing the method for constructing the multilingual phoneme representation model in the above-described method embodiment.

[0062] The memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created according to the use of the device 10. Furthermore, the memory 120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 120 may optionally include memory remotely located relative to the processor 110, and these remote memories may be connected to the device 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0063] One or more units are stored in memory 120, and when executed by one or more processors 110, perform the following steps: In view of this, the first aspect of this application provides a high-throughput detection method for tumor tissue microarrays, the method comprising: Step S101: Pre-train an image-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; Step S102: Use the H&E whole slice to be tested as input to the image-gene expression prediction model to generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. Step S103: Based on the preset target biomarker panel, generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map; Step S104: In the virtual IHC expression intensity map, candidate core regions for each target biomarker are selected according to the preset core retrieval rules, and a first coordinate set is generated based on the spatial coordinates of the candidate core regions. The first coordinate set is then used to establish a first correspondence between the patient number, biomarker number, and candidate priority. Step S105: After mapping the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and coding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as coding rules in the database. Step S106: After core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, construct several mixed TMA blocks containing different patients under the same marker number, and perform immunohistochemical staining on each mixed TMA block. Step S107: After scanning each mixed TMA block after staining to obtain the corresponding real IHC image, decode the real IHC image according to the encoding rules to obtain the correspondence between the core well position in each mixed TMA block and the patient number and biomarker number. Step S108: After performing image analysis on each core well location to obtain the actual IHC quantitative results, the actual IHC quantitative results are evaluated and compared with the virtual IHC expression intensity map of the corresponding patient number to generate a verification report containing multiple patients and multiple target biomarkers.

[0064] This invention provides a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by one or more processors, they implement any of the embodiments of the high-throughput detection method for tumor tissue microarray described above.

[0065] As examples, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory components or memories disclosed in the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.

[0066] This application provides a high-throughput detection method, device, equipment, and medium for tumor tissue microarrays. The method uses a pre-trained image-gene expression model to directly predict the spatial transcriptome from H&E whole slides and generate a virtual IHC map. It automatically selects candidate core regions, generates coordinate sets, and maps them to paraffin blocks. Combined with coding rules, it automates the entire process of core extraction, staining, scanning, decoding, and quantitative analysis without the need for manual visual selection. Based on a general H&E image pre-trained model, only the target biomarker panel needs to be changed to generate a new virtual IHC map and core extraction rules. By constructing a hybrid TMA block containing different patients and multiple biomarkers and a standardized coding system, the same TMA platform can be flexibly reused and expanded across multiple gene panels, multiple prediction models, and multiple cohorts.

[0067] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0068] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0069] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0070] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0071] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0072] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0073] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0074] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

[0075] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.

Claims

1. A high-throughput detection method for tumor tissue microarrays, characterized in that, include: Step S101: Pre-train an image-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; Step S102: Use the H&E whole slice to be tested as input to the image-gene expression prediction model to generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. Step S103: Based on the preset target biomarker panel, generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map; Step S104: In the virtual IHC expression intensity map, candidate core regions for each target biomarker are selected according to the preset core retrieval rules, and a first coordinate set is generated based on the spatial coordinates of the candidate core regions. The first coordinate set is then used to establish a first correspondence between the patient number, biomarker number, and candidate priority. Step S105: After mapping the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and coding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as coding rules in the database. Step S106: After core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, construct several mixed TMA blocks containing different patients under the same marker number, and perform immunohistochemical staining on each mixed TMA block. Step S107: After scanning each mixed TMA block after staining to obtain the corresponding real IHC image, decode the real IHC image according to the encoding rules to obtain the correspondence between the core well position in each mixed TMA block and the patient number and biomarker number. Step S108: After performing image analysis on each core well location to obtain the actual IHC quantitative results, the actual IHC quantitative results are evaluated and compared with the virtual IHC expression intensity map of the corresponding patient number to generate a verification report containing multiple patients and multiple target biomarkers.

2. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, Step S101 specifically includes: In the training set of H&E whole slices and spatial transcriptome data with back-to-back relationships, the H&E whole slices are divided into several image patches, and the spatial transcriptome sequencing points located in the same image patch are aggregated to obtain pseudobulk gene expression vectors. Based on several image patches and their corresponding pseudo-bulk gene expression vectors, an open gene embedding encoder and a diffusion-based image feature extraction network are pre-trained to obtain an image-gene expression prediction model. The open gene embedding encoder has permutation invariance to gene sets of different panels.

3. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, Step S103 further includes: The virtual IHC expression intensity map was sequentially normalized, smoothed, thresholded, and then pseudo-colorized. The virtual IHC expression intensity maps of each gene are clustered according to the virtual expression vectors of multiple genes to obtain the virtual molecular region division.

4. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, The preset core extraction rules in step S104 include at least one of the following: high expression threshold rule, expression gradient boundary rule, tumor-stromal junction rule, spatial clustering anomaly rule, or virtual molecular region boundary alignment rule.

5. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, In step S106, the immunohistochemical staining of each mixed TMA block is specifically performed as follows: each mixed TMA block is subjected to immunohistochemical staining in the same batch according to the type of target biomarker.

6. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, In step S108, image analysis is performed on the position of each core hole to obtain the actual IHC quantitative result, specifically as follows: Image analysis was performed on each well location, and at least one quantitative indicator among IHC staining intensity, positive staining area ratio, average optical density, or H-score was calculated to obtain the true IHC quantitative results.

7. The high-throughput detection method for tumor tissue microarrays according to claim 1, characterized in that, In step S108, the evaluation and comparison of the actual IHC quantitative results with the virtual IHC expression intensity map corresponding to the patient number specifically involves: The real IHC quantitative results are paired with virtual IHC expression intensity maps corresponding to patient numbers and biomarker numbers. At least one evaluation index among correlation coefficient, mean square error, or classification consistency is calculated to obtain the validation results of the target biomarker in multiple patients.

8. A high-throughput detection device for tumor tissue microarrays, characterized in that, include: The pre-training unit is used to pre-train a picture-gene expression prediction model based on paired H&E whole slices and spatial transcriptome data; The first processing unit is used to take the H&E whole slice to be tested as input to the image-gene expression prediction model and generate a spatial transcriptome map containing multi-gene expression predictions for each spatial location. The second processing unit is used to generate a virtual IHC expression intensity map of each gene in the spatial transcriptome map based on a preset target biomarker panel. The candidate core retrieval unit is used to screen candidate core retrieval areas for each target biomarker in the virtual IHC expression intensity map according to preset core retrieval rules, and generate a first coordinate set based on the spatial coordinates of the candidate core retrieval areas, and construct a first correspondence between the first coordinate set and the patient number, biomarker number and candidate priority. The encoding unit is used to map the first coordinate set to the second coordinate set on the adjacent unstained slices through tissue contour matching or image registration algorithm, determine the TMA layout and encoding rules of each target biomarker according to the second coordinate set, construct the second correspondence between the row and column positions of each core well in the TMA block and the patient number, biomarker number and core retrieval sequence number, and store the second correspondence as the encoding rule in the database. The hybrid TMA block construction unit is used to construct several hybrid TMA blocks containing different patients under the same marker number after core extraction from paraffin-embedded blocks of multiple patients according to the second coordinate set, and to perform immunohistochemical staining on each hybrid TMA block. The decoding unit is used to scan each stained mixed TMA block to obtain the corresponding real IHC image, and then decode the real IHC image according to the encoding rules to obtain the correspondence between the core hole position in each mixed TMA block and the patient number and biomarker number. The verification unit is used to perform image analysis on each core well location, obtain the real IHC quantitative results, evaluate and compare the real IHC quantitative results with the virtual IHC expression intensity map of the corresponding patient number, and generate a verification report containing multiple patients and multiple target biomarkers.

9. A high-throughput detection device for tumor tissue microarrays, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the high-throughput detection method for tumor tissue microarrays according to any one of claims 1-7, based on the instructions in the program code.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code for executing the high-throughput detection method for tumor tissue microarrays according to any one of claims 1-7.