Space protein expression map reconstruction method and device and storage medium

By combining multi-channel images and strip cutting data from two tissue slices with a deep learning network, the high cost and complex process of existing technologies are solved, and efficient and high-precision spatial protein map reconstruction is achieved.

CN121983146APending Publication Date: 2026-05-05WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
Filing Date
2026-01-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for achieving high-resolution map reconstruction in spatial proteomics require multiple tissue slices and complex experimental procedures, and rely on cross-omics association assumptions, resulting in high experimental costs, low efficiency, and potential impact on reconstruction accuracy.

Method used

By combining multi-channel high-resolution images of two tissue slices and tissue strip cutting data with a deep learning network, the protein expression levels in the tissue region can be predicted through training the deep learning network, thus achieving high-precision reconstruction.

Benefits of technology

It achieves high-precision protein map reconstruction, reduces the number of tissue sections, simplifies the experimental procedure, avoids cross-omics hypothesis dependence, and improves reconstruction efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983146A_ABST
    Figure CN121983146A_ABST
Patent Text Reader

Abstract

The invention discloses a reconstruction method and a reconstruction device of a space protein expression map and a storage medium. Comprising the following steps: receiving a multi-channel high-resolution tissue image of a tissue region obtained based on tissue slices; receiving a polymerization true value of the expression quantity of each target protein on each first-direction tissue band and each second-direction tissue band, wherein the polymerization true value is obtained after the first-direction tissue band and the second-direction tissue band are cut on the first tissue slice and the second tissue slice; a deep learning network is trained based on the aggregation true values of the protein expression quantities on the tissue strips in the first direction and the second direction and the multi-channel high-resolution tissue image, and the trained deep learning network predicts the expression quantity of each target protein on each image unit based on the multi-channel high-resolution tissue image; and reconstructing a two-dimensional protein expression map. According to the method, the two-dimensional space protein map can be reconstructed with high precision only by using as few as two tissue slices and the aggregation value of the protein expression quantity on the tissue strip which is relatively easy to obtain without the protein expression quantity of each unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of space omics analysis technology. More specifically, this application relates to a method, apparatus, and storage medium for reconstructing a space protein expression map. Background Technology

[0002] Spatial-resolved proteomics enables high-resolution measurements of protein expression and distribution in situ within tissues, which is crucial for a deeper understanding of the tissue microenvironment, cellular heterogeneity, and intercellular interactions. However, achieving deep (thousands of proteins) and high-resolution (micrometer-scale) protein mapping of whole-tissue sections has been a significant challenge in this field. The core bottleneck lies in the throughput limitations of mass spectrometry and the non-amplifiable nature of protein molecules, making point-by-point scanning of the entire tissue impractical in terms of time and cost.

[0003] To overcome this bottleneck, a representative advanced technology in this field utilizes microfluidic chips to generate parallel microchannels at different angles on two adjacent tissue slices, and performs in-situ lysis and proteomics analysis on the tissue within the channels to obtain two sets of orthogonal protein projection data. Microfluidic chips are complex to operate, often have low processing yields (e.g., less than 30%), submicron channels are prone to blockage and deformation, and chip and debugging costs are high. Furthermore, this technology requires a "cross-omics transfer learning" strategy, meaning that a third adjacent slice is needed for other spatial omics analyses, such as hematoxylin and eosin (H&E) imaging or spatial transcriptomics as a reference map; or, pathologists need to delineate categories such as tumor areas and stroma areas on histological images such as H&E imaging, and the model learns known spatial patterns of different categories to guide the reconstruction of the protein map. Therefore, this technology requires a large number of tissue sections (three consecutive images), has a complex experimental procedure (two sections require microfluidic chip operation, and the other section requires a separate spatial omics experiment as a reference), and even when using histological images from H&E imaging, clinicians still need to segment and label different regions on the images. This not only increases the complexity of the experiment and creates additional manpower, but more importantly, its reliance on the strong assumption of cross-omics association (i.e., assuming that the spatial pattern of the reference omics or histological image is highly correlated with the spatial pattern of the target proteome) may have a significant impact on the reconstruction accuracy, especially when there is a spatial inconsistency between the protein and its transcript, which is not uncommon in organisms.

[0004] Another representative cutting-edge technology employs a sparse sampling strategy, using laser microdissection to cut a series of parallel tissue bands at different angles from multiple (e.g., eight) consecutive tissue sections for proteomics analysis. The reconstruction algorithm uses a deep learning model to reconstruct a two-dimensional protein map through pure mathematical inversion using these projection data from multiple angles and sections. To obtain a sufficient number of projection angles to ensure reconstruction stability, this technique requires up to eight consecutive sections. This not only increases the experimental workload but also results in a final map that is an "average projection" of these eight sections (up to 80 µm thick), potentially masking fine biological structures and cellular heterogeneity at a single level, thus affecting the accuracy of protein map reconstruction.

[0005] Therefore, although the above-mentioned technologies have greatly promoted the development of spatial proteomics technology and improved the reconstruction of protein maps to a higher resolution, no protein map reconstruction technology with higher resolution has yet been found in the field that can simplify the experimental process, require fewer tissue sections, and does not require manual delineation or rely on cross-omics association assumptions. Summary of the Invention

[0006] This application addresses the aforementioned deficiencies in the prior art. There is a need for a method, apparatus, and storage medium for reconstructing spatial protein expression maps, which can achieve high-precision reconstruction of spatial protein expression maps using fewer tissue sections, with a simpler experimental procedure, without requiring clinicians to draw any form of histological image, and without relying on cross-omics association hypotheses.

[0007] According to a first aspect of this application, a method for reconstructing a spatial protein expression map is provided, comprising: receiving a multi-channel high-resolution tissue image containing a tissue region obtained based on at least one tissue slice; receiving aggregated ground truth values ​​of the expression levels of each target protein on each first-direction tissue band obtained by cutting the tissue region of a first tissue slice into tissue bands in a first direction, and aggregated ground truth values ​​of the expression levels of each target protein on each second-direction tissue band obtained by cutting the tissue region of a second tissue slice into tissue bands in a second direction; training a deep learning network based on the aggregated ground truth values ​​of the expression levels of each target protein on each first-direction tissue band and each second-direction tissue band, and the multi-channel high-resolution tissue image; and using the trained deep learning network based on the multi-channel high-resolution tissue image to predict the expression levels of each target protein on each image unit of the tissue region and use this prediction for reconstructing a two-dimensional protein expression map of the tissue region.

[0008] According to a second aspect of this application, a spatial protein expression map reconstruction apparatus is provided, comprising an interface and at least one processor. The interface is configured to receive a multi-channel high-resolution tissue image containing a tissue region obtained based on at least one tissue slice; aggregated true values ​​of the expression levels of each target protein on each first-direction tissue band obtained by cutting the tissue region of the first tissue slice into tissue bands; and aggregated true values ​​of the expression levels of each target protein on each second-direction tissue band obtained by cutting the tissue region of the second tissue slice into tissue bands. The at least one processor is configured to execute the steps of the spatial protein expression map reconstruction method according to various embodiments of this application.

[0009] According to a third aspect of this application, a non-transitory computer-readable storage medium is provided, having stored thereon computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the steps of the method for reconstructing spatial protein expression maps according to various embodiments of this application are performed.

[0010] The spatial protein expression map reconstruction method, reconstruction device, and storage medium provided in the various embodiments of this application can complete the high-precision reconstruction of protein spatial maps using at least two tissue slices. Specifically, one tissue slice can be used to acquire a multi-channel high-resolution tissue image of the tissue region, eliminating the need for time-consuming and costly point-by-point scanning of the tissue region. Instead, a tissue slice is used to cut tissue bands from different directions. Based on the tissue band cutting, the aggregated values ​​of protein expression levels on each tissue band are obtained relatively easily. Then, using the acquired aggregated values ​​of protein expression levels in the tissue bands and the multi-channel high-resolution tissue image, a deep learning network is trained. After training, the protein expression levels of each unit in the tissue region can be accurately predicted based solely on the multi-channel high-resolution tissue image, thereby achieving high-precision reconstruction of two-dimensional protein maps. Compared to existing technologies, the reconstruction method in this application requires as few as two tissue slices, thus avoiding obscuring the fine biological structures and cellular heterogeneity at a single level. Furthermore, it eliminates the need for complex and low-yield microfluidic chip operations, eliminates the need for clinicians to perform any region segmentation / tissue type identification and delineation on histological images, and does not rely on cross-omics association hypotheses. With the simplest possible experimental procedure and the lowest possible number of tissue slices required, it achieves high-precision reconstruction of spatial protein expression maps.

[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application.

[0012] It should be understood that the foregoing general description and the following detailed description are merely illustrative and explanatory, and are not intended to limit the scope of the claimed invention. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A schematic flowchart illustrating a method for reconstructing a spatial protein expression map according to an embodiment of this application is shown.

[0015] Figure 2 A schematic block diagram showing a portion of the components of a deep learning network according to an embodiment of this application is provided.

[0016] Figure 3 A schematic diagram illustrating the training steps of a deep learning network according to an embodiment of this application is shown.

[0017] Figure 4 A schematic diagram illustrating a method for obtaining training input data for a deep learning network according to an embodiment of this application is shown.

[0018] Figure 5 This diagram illustrates a method for reconstructing spatial protein expression maps and an exemplary workflow according to embodiments of this application.

[0019] Figure 6 A partial block diagram of a spatial protein expression map reconstruction apparatus according to an embodiment of this application is shown. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the described embodiments of this application without creative effort are within the scope of protection of this application.

[0021] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. Words such as "comprising" or "including" mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, without excluding other elements or objects.

[0022] The terms "first," "second," and similar words used in this application do not indicate any order, quantity, or importance, but are merely used for distinction. Words such as "including" or "comprising" mean that the element preceding the word encompasses the elements listed after it, and do not exclude the possibility of encompassing other elements as well. The execution order of the steps in the method described in conjunction with the accompanying drawings in this application is not intended to be limiting. As long as the logical relationship between the steps is not affected, several steps can be integrated into a single step, a single step can be decomposed into multiple steps, and the execution order of the steps can be changed according to specific needs.

[0023] It should also be understood that the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Furthermore, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.

[0024] To keep the following description of the embodiments of this application clear and concise, detailed descriptions of known functions and known components are omitted.

[0025] According to embodiments of this application, a method for reconstructing a spatial protein expression map is provided. Figure 1 A schematic flowchart illustrating a method for reconstructing a spatial protein expression map according to an embodiment of this application is shown.

[0026] like Figure 1 As shown, in step S101, a multi-channel high-resolution tissue image containing a tissue region, obtained based on at least one tissue slice, can be received. Here, "high resolution" means that the physical size corresponding to each pixel in the tissue image is less than or equal to a certain value. For example, if a spatial reconstruction resolution of 12 micrometers is desired, a high-resolution image can be considered as a multi-channel microscopic image with a pixel size of no more than 0.5 micrometers, etc. However, it is not limited to this. If the imaging device allows, the pixel size can be even smaller, which will not be listed here.

[0027] In some embodiments, the multi-channel high-resolution tissue image can be acquired, for example, by performing various immunofluorescence stainings on the tissue section. More specific methods for acquiring multi-channel high-resolution tissue images will be described in detail later, and the acquisition process will be combined with... Figure 4Please provide a detailed explanation.

[0028] Next, in step 102, for example, the aggregated true values ​​of the expression levels of each target protein on each first-direction tissue band obtained by cutting the first tissue slice into tissue regions in a first direction, and the aggregated true values ​​of the expression levels of each target protein on each second-direction tissue band obtained by cutting the second tissue slice into tissue regions in a second direction, may be received. In some embodiments, the first tissue slice and the second tissue slice may be two adjacent tissue slices.

[0029] In some embodiments, the first direction may be orthogonal to the second direction to ensure the integrity of the tissue section structure during strip cutting, minimize tissue deformation, and improve the accuracy of subsequent staining and microscopic observation. In other embodiments, the first and second directions are not limited to being orthogonal; they may be cut at angles other than 90° (such as 45° or 30°). This allows the resulting strips to be more suitable for exposing specific structures in specific research scenarios (such as tracing nerve fiber paths, analyzing anisotropic tissues such as muscle or fibrocartilage), and so on. Regardless of whether the first and second directions are orthogonal, it must be ensured that each spatial sub-region (i.e., the basic unit for reconstructing the protein expression map) obtained after cutting the tissue region is acquired at least once in both the first and second directions. The geometry of the defined spatial sub-regions may be, for example, quadrilaterals, hexagons, or other polygons; this application does not limit this, and the specific shape can be determined based on the tissue section. If other methods are used to cut each cell individually, it is necessary to ensure that no two identical cells appear in a single row or column. Specific cutting methods will not be listed here. The specific implementation methods for strip cutting of organizational regions will be discussed later. Figure 4 An example is provided.

[0030] Then, in step 103, the deep learning network is trained based on the aggregated true values ​​of the expression levels of each target protein in each first-direction tissue band and each second-direction tissue band, and the received multi-channel high-resolution tissue image. In some embodiments, the deep learning network can be constructed based on a composite system structure such as a multi-layer convolutional neural network + attention module + multi-layer perceptron, etc. This application does not limit this, and exemplary implementations will be provided in the following sections. Figure 5 The details are described in the text.

[0031] In some embodiments, multi-channel high-resolution tissue images can be acquired based on the same tissue slice. In this way, high-resolution anatomical images (such as cell nuclear staining images) on the same slice are used as structural prior information and trained together with the aggregated values ​​of protein expression levels on tissue bands. This allows the trained deep learning network to have a stronger ability to accurately locate protein expression to specific cells or subcellular structures.

[0032] Finally, in step 104, based on the previously received multi-channel high-resolution tissue images, a trained deep learning network can be used to predict the expression levels of each target protein in each image unit of the tissue region and to reconstruct the protein expression map in the two-dimensional space of the tissue region.

[0033] According to the spatial protein expression map reconstruction method described in the embodiments of this application, high-precision reconstruction of the protein spatial map can be achieved using at least two tissue slices. One tissue slice can be used to obtain a multi-channel high-resolution tissue image of the tissue region, eliminating the need for time-consuming and costly point-by-point scanning of the tissue region. Instead, tissue slices are cut from different directions using a single tissue slice. Based on these cuts, aggregated values ​​of protein expression levels on each tissue slice are obtained relatively easily. Then, based on the multi-channel high-resolution tissue image, these aggregated values ​​of protein expression levels are used as supervisory information for training a deep learning model. After training, the protein expression levels of each unit in the tissue region can be accurately predicted using only the multi-channel high-resolution tissue image, thereby achieving high-precision reconstruction of the two-dimensional protein map. Compared to existing technologies, the reconstruction method in this application requires as few as two tissue slices, thus avoiding obscuring the fine biological structures and cellular heterogeneity at a single level. Furthermore, it does not require the use of complex and low-yield microfluidic chips, nor does it rely on cross-omics association assumptions. Instead, it directly utilizes endogenous, multi-dimensional biological information (such as multi-channel immunofluorescence images) from the same slice as high-resolution structural prior information to guide the training and prediction of deep learning models, thereby achieving more efficient, accurate, and biologically intuitive high-resolution protein map reconstruction.

[0034] Unlike existing technologies that typically require pathologists / clinicians to delineate categories such as tumor areas and stromal areas on images so that the model can learn known spatial patterns of different categories to guide protein atlas reconstruction, in the embodiments of this application, no identification or labeling related to cell phenotype or tissue structure morphology is required on the multi-channel high-resolution tissue images before training the deep learning network using the multi-channel high-resolution tissue images, and before predicting protein expression atlases using the multi-channel high-resolution tissue images. Whether this identification or labeling related to cell phenotype or tissue structure morphology is performed manually or automatically by a computer, it is not a step required by the embodiments of this application. Therefore, this not only avoids cumbersome procedures and additional requirements for personnel, fundamentally eliminating reliance on manual image annotation, but also avoids the negative impact of inaccurate annotation on model training and prediction.

[0035] In some embodiments, multi-channel high-resolution tissue images are acquired as follows: First, reference protein markers can be designed based on the tissue spatial structure, cell type heterogeneity, and complementarity between protein markers in the tissue region; then, the tissue sections are stained with immunofluorescent antibodies using each reference protein marker; finally, multi-channel high-resolution tissue images are acquired based on the immunofluorescently stained tissue sections.

[0036] Taking liver tissue as an example, during the process of filing this application, the inventors discovered through numerous experiments that, under normal circumstances, using only as few as three immunofluorescence staining images obtained with a limited number of complementary expression patterns, it is possible to accurately infer the reference protein markers for the overall protein expression distribution. Combined with one nuclear marker for immunofluorescence antibody staining, this captures crucial spatial structural information in liver tissue sections. Specifically, using liver tissue as an example, one particular staining method involves selecting three complementary protein markers for immunofluorescence staining of tissue sections, combined with one nuclear marker, forming a total of four marker channels to obtain multi-channel, high-resolution tissue images. These three protein markers can be selected based on prior anatomical knowledge, specifically for different spatial regions of the liver, including but not limited to markers for the central venous region, portal venous region, and hepatocyte-specific markers. In other embodiments of this application, when the tissue region is a different anatomical location or contains different tissue types, fewer or more complementary protein markers can be selected to stain the tissue sections with immunofluorescence antibodies to obtain the required multi-channel high-resolution tissue images. Experimental results show that, even without prior knowledge about the tissue region, accurate prediction of the spatial expression of at least some proteins can be achieved when selecting the staining method. In the embodiments of this application, a non-targeted "complementary marker staining image," which has never been used in the prior art, is employed to invert or reconstruct the entire protein spatial map. A deep learning model is trained using multi-channel images with limited complementary information as input. The generalization ability of the trained deep learning model in terms of tissue spatial structure, cell type, etc., is utilized to infer the full protein expression distribution more accurately, thereby greatly reducing experimental costs and improving the efficiency of protein map reconstruction. It is understood that, depending on the application requirements, this can be extended to dozens or even hundreds of multi-channel markers. Increasing the number of channels can, to some extent, help further improve the accuracy of protein map reconstruction. Furthermore, if at least some prior knowledge about the correlation between tissue regions and tissue spatial structure and cell type heterogeneity can be combined, and complementary protein markers can be selected more precisely for immunofluorescence antibody staining, the accuracy of protein spatial expression prediction and the types of target proteins that can be accurately predicted can be further improved.

[0037] Figure 2 A schematic block diagram illustrating a portion of the components of a deep learning network according to an embodiment of this application is shown. Figure 2 As shown, the deep learning network 20 may sequentially include a feature extraction module 201 and a protein-specific prediction module 202, wherein the feature extraction module 201 is shared by all target proteins, while the protein-specific prediction module 202 contains an independent prediction head corresponding to each target protein.

[0038] exist Figure 2 Based on the structure of the deep learning network 20 shown, and having obtained the aggregated true values ​​of the expression levels of each target protein in each first-direction tissue band and each second-direction tissue band, as well as multi-channel high-resolution tissue images, the deep learning network can be trained based on the aggregated true values ​​of the expression levels of each target protein in each first-direction tissue band and each second-direction tissue band, and the multi-channel high-resolution tissue images. Specifically, for example, it can be trained as follows: Figure 3 Perform each step as shown.

[0039] like Figure 3 As shown, in step 301, the tissue region in the multi-channel high-resolution tissue image can first be divided into multiple image units according to the cutting method of the first direction tissue strip and the second direction tissue strip.

[0040] Then, in step 302, the deep learning network can be trained multiple times until the training objective is reached or a specified number of rounds are reached. In each round of training, a specified number of first-direction and second-direction tissue bands are randomly selected. The image units corresponding to the selected tissue bands and the aggregated ground truth values ​​of the expression levels of each target protein on the tissue bands are used as training samples, and then... Figure 3 The deep learning network is trained using sub-steps S3021-S3024.

[0041] In sub-step S3021: the feature extraction module is used to generate the feature encoding vector corresponding to each image unit in the current first-direction tissue strip / second-direction tissue strip.

[0042] In sub-step S3022: For each target protein, based on the feature encoding vector corresponding to each image unit, the protein-specific prediction module generates a predicted value of the expression level of the target protein in each image unit in the current first-direction tissue band / second-direction tissue band by each independent prediction head.

[0043] In sub-step S3023: For each target protein, a first loss function is calculated based on the sum of the predicted expression levels of the target protein in each image unit of the current first-direction tissue band / second-direction tissue band, and the aggregated true value of the expression level of the target protein in the current first-direction tissue band / second-direction tissue band.

[0044] In sub-step S3024: Based on the first loss function of each target protein in each first direction tissue band and each second direction tissue band, calculate the second loss function for training the deep learning network, and determine whether the training objective has been achieved based on the second loss function.

[0045] In the training process of the aforementioned deep learning network, the number of first-direction and second-direction tissue strips selected in each training round can be the same or different. This application does not impose any restrictions on this, but preferably, the number of both can be the same, so that the composition of training samples in each round can be more balanced.

[0046] In some embodiments, the reconstruction method according to this application further includes generating an effective tissue region mask based on the multi-channel high-resolution tissue image. Dividing the tissue region in the multi-channel high-resolution tissue image into multiple image units according to the first and second direction tissue strips further includes: dividing the tissue region in the multi-channel high-resolution tissue image into multiple image units according to the first and second direction tissue strips; applying the effective tissue region mask to the multiple image units such that the multiple image units only include image units within the effective tissue region. This avoids the model erroneously learning regions that do not contain effective tissue, interfering with model training speed and convergence, and reducing the prediction accuracy of the trained deep learning network.

[0047] Furthermore, the step of predicting the expression levels of each target protein in each image unit of the tissue region based on the multi-channel high-resolution tissue image and using a trained deep learning network for reconstructing a protein expression map in the two-dimensional space of the tissue region can further include: inputting each image unit in the effective tissue region into the trained deep learning network in a preset order to generate a proteome abundance matrix composed of the expression levels of each target protein in each image unit of the effective tissue region; and reconstructing a protein expression map in the two-dimensional space of the effective tissue region based on the protein abundance matrix. In some embodiments, the protein abundance matrix can be directly used as the protein expression map in the two-dimensional space of the tissue region. In other embodiments, tools such as Python can be used to convert the protein abundance matrix into ANDATA format (a data format commonly used in single-cell spatial transcriptomics), etc. This application does not limit this, as long as it can meet the requirements of possible downstream steps such as bioinformatics analysis.

[0048] In some embodiments, with the support of a laser microdissection system (LCM), the first-direction tissue strip cutting and the second-direction tissue strip cutting can reach the single-cell scale. Thus, when the first-direction tissue strip cutting and the second-direction tissue strip cutting reach the single-cell scale, the reconstruction method of this application embodiment can reconstruct the protein expression map of the tissue region in two-dimensional space at the single-cell scale. LCM technology outperforms microfluidic chips in all aspects of spatial resolution, cutting precision, operational flexibility, and sample compatibility: Under optimized parameters (such as laser power, focused spot, and slice thickness), LCM cutting can achieve a minimum band width of 5-10 μm, meeting the high-resolution sampling requirements of single-cell omics analysis; it does not require pre-designed microchannel structures and photolithography, and can be directly operated on conventional paraffin-embedded or frozen sections, making it suitable for various tissue types. In particular, with the assistance of a high-resolution microscope, it is possible to determine the angle between rows and columns and the shape of the cut spatial units based on the distribution of heterogeneous tissues (such as tumor margins and neural nuclei), which is more conducive to the accurate reconstruction of spatial maps; LCM is a pure photothermal cutting method, avoiding cell damage or protein denaturation caused by fluid shear force or channel wall friction in microfluidics; LCM-cut strips do not require additional transfer steps in microfluidics and can be directly collected into standard centrifuge tubes or PCR tubes, reducing the risk of sample loss and seamlessly connecting to downstream mass spectrometry analysis processes, which also makes the measurement of protein expression levels more accurate.

[0049] In some embodiments, when a new target protein needs to be added, only an independent prediction head corresponding to the new target protein needs to be added to the protein-specific prediction module. Then, the deep learning network with the added independent prediction head is further trained using the aggregated true values ​​of the expression levels of the new target protein in each first-direction tissue band and each second-direction tissue band, and the multi-channel high-resolution tissue image. Next, based on the multi-channel high-resolution tissue image, the supplementally trained deep learning network can predict the expression level of the new target protein in each image unit of the tissue region and use it to reconstruct the protein expression map in the two-dimensional space of the tissue region. This greatly increases the scalability of the reconstruction method according to the embodiments of this application.

[0050] Figure 4 This diagram illustrates a method for acquiring training input data for a deep learning network according to an embodiment of this application. As mentioned above, the training input data for a deep learning network includes at least two main categories of data: multi-channel high-resolution tissue images and aggregated ground truth values ​​of the expression levels of the target protein in various first / second direction tissue bands obtained after striping the tissue region. Figure 4 As shown, the acquisition of the two types of data can be performed in the following steps.

[0051] Step 1, Tissue Section Acquisition: Animal care and tissue preparation are required. All animal experiments in the embodiments of this application followed the guidelines of the Westlake University Animal Experiment Management and Use Committee and were approved. Wild-type mice aged 2-4 months were selected and deeply anesthetized with 1% sodium pentobarbital, then perfused with 1×PBS and 4% paraformaldehyde (PFA). Liver tissue was collected and fixed in 4% PFA solution at 4°C for 6 hours. After dehydration with a gradient of 75%, 95%, and 100% ethanol for 30 minutes each step, the tissue was embedded in paraffin at 60°C. Two tissue sections with a thickness of 10 μm were cut from the embedded tissue using a rotary microtome, preferably two continuous sections, and attached to poly-L-lysine-treated slides. One of the sections, as shown... Figure 4 The first tissue section can be stained with hematoxylin and eosin (H&E), and after obtaining multi-channel high-resolution tissue images of the tissue area using methods such as stereomicroscopy, subsequent processing steps can be performed. The other section, such as... Figure 4 The second tissue section can be processed directly without staining.

[0052] Step 2, Tissue linear swelling + immunofluorescence imaging: The first and second tissue sections, after dewaxing and rehydration, were incubated with anchoring solution (0.1 mg / mL NSA, 100 mM MES, pH 6.0) at room temperature in the dark for 1 hour, followed by washing three times with 100 mM MOPS buffer. Monomer solution was then added and incubated at 4°C for 12 hours, followed by polymerization in a nitrogen-filled vacuum oven at 37°C for 2 hours. The resulting tissue-hydrogel complex was autoclaved at 95°C for 8 hours to denature proteins, and then washed three times with 1×PBS to allow for higher-precision immunofluorescence imaging of the swollen tissue sections. After swollen samples, blocking solution was applied at 37°C for 1 hour. Then, 200 μL of primary antibody (multiple complementary / dissimilar primary antibodies can be stained simultaneously as needed) was added, and incubation was performed overnight at 4°C, followed by three 20-min washes with PBS; then, the corresponding secondary antibody was added, and incubation was performed at room temperature for 2 hours, followed by three 10-min washes with PBS. Finally, DAPI staining was performed for 10 min, followed by three 5-min washes with PBS. The first tissue section was placed under a Nikon microscope (20× or 40×) and imaged at a resolution of at least 0.5 micrometers using different secondary antibodies to obtain the multi-channel immunofluorescence images required for deep learning network training. After the required tissue images were acquired, a z-axis scan was set, and both tissue sections were expanded to their maximum magnification (approximately 4-5 times linear size) in ddH2O. The two tissue sections were then stained with Coomassie Brilliant Blue and fixed onto an LCM steel frame membrane using dry glue. Coomassie Brilliant Blue staining is used for visualization of the expanded samples, effectively confirming the cutting angle and facilitating recovery and sample preparation.

[0053] Step 3, LCM-based tissue strip cutting: Based on the Coomassie brilliant blue staining image, tissue strip cutting is performed on the two tissue sections. Preferably, the first tissue section is sampled and cut along the x-axis (first direction, also referred to as row in this text), and the second tissue section is sampled and cut along the y-axis (second direction, also referred to as column in this text), which is orthogonal to the x-axis. The resulting tissue strips are collected using LCM, such as... Figure 4 As shown, the tissue band width is as narrow as 12 μm, meeting the high-resolution sampling requirements for single-cell omics analysis of most tissue types. Next, each tissue band was used as a sample and processed using a filter-assisted in-gel digestion method: trypsin was added to a final concentration of 5 ng / μL, and the sample was incubated in a humidified chamber at 37°C for 12 hours. Subsequently, peptides were eluted sequentially with 2% ACN / 0.1% TFA, 70% ACN / 0.1% TFA, and 100% ACN buffer. The collected peptides were concentrated under vacuum and analyzed by LC-MS / MS (liquid chromatography-tandem mass spectrometry) to obtain high-resolution data on the width of the single-cell tissue bands and the mass-to-charge ratio / intensity values ​​of each spatially paired mouse liver protein. These data served as the aggregated true values ​​of the expression levels (i.e., abundance) of each target protein on the first / second direction tissue bands, and a protein row / column aggregated expression matrix (also known as a row / column aggregated expression profile) was constructed. Each element in the matrix represents the aggregated expression level of a specific target protein in a particular row / column.

[0054] Step 4: Obtain tissue region mask based on high-resolution tissue image: Generate a binary matrix with the same dimension as the row and column aggregated expression spectrum output by LCM tissue segmentation, used to identify the effective region where the tissue sample is located. A value of 1 indicates an organized region, and 0 indicates an unorganized region. In this embodiment, for example, an algorithm in OpenCV can be used to identify the tissue region mask in the high-resolution tissue image.

[0055] Through the above workflow, two main categories of paired data types are provided for protein reconstruction map reconstruction of the deep learning network in this application embodiment, including: (a) multi-channel high-resolution histological images (i.e., multi-channel immunofluorescence maps) containing tissue regions, (b) high-quality protein row / column aggregated expression matrices corresponding to tissue slices, and (c) tissue region masks corresponding to slice images.

[0056] Figure 5 This diagram illustrates an exemplary workflow of a method for reconstructing spatial protein expression maps according to embodiments of this application. Figure 5The diagram employs a left-middle-right overall flow, combined with a top-bottom partitioning to distinguish between the training and inference modes. The left area contains the data input for the deep learning network's training / inference modes, the central area represents the core model architecture, and the right area shows the final output. The top flow represents the training process, and the bottom flow represents the inference process. Note that... Figure 5 The specific implementation methods of each part are exemplary and not restrictive.

[0057] like Figure 5 As shown, in the deep learning network, the feature extraction module is implemented as a shared multi-scale convolutional neural network (Multi-Scale CNN) feature extractor, while the protein-specific prediction module consists of multiple specific gated protein heads, the same number as the target protein.

[0058] Among them, the shared multi-scale convolutional neural network feature extractor receives the size N corresponding to a single pixel. M pixels (e.g., 25) 25 or 20 The network takes high-resolution image patches (e.g., 25) as input. Internally, the network consists of three cascaded convolutional stages (Convolutional Neural Network Block 1, Convolutional Neural Network Block 2, and Convolutional Neural Network Block 3), each containing a convolutional layer, batch normalization (BatchNorm), and SiLU activation function, with progressive downsampling as depth increases.

[0059] like Figure 5 As shown, the shared multi-scale feature extractor does not only output the features of the last layer, but branches are drawn from the ends of the three convolutional neural network blocks respectively. After 1x1 convolutional projection, adaptive average pooling and flattening, the final output is a three-dimensional feature vector, including: (1) shallow features F shallow : It preserves the high-frequency texture and edge information of the image; (2) Mid-level features F mid : Capture local cell structure information; (3) Deep features F deep These three feature vectors encode highly abstract organizational semantics and contextual information. Together, they constitute the multi-scale feature pyramid of the image patch.

[0060] Figure 5 Each specific gating prediction head corresponds to a target protein to be analyzed. This independent configuration of prediction heads addresses the issue that different target proteins have varying requirements for image features. For example, structural proteins may rely more on image texture information, while functional proteins may depend more on tissue semantics, and so on. Taking n target proteins as an example, the internal composition and processing flow of the specific gating prediction head are as follows: Step 1, Gated Routing: The deep features are received as input through a sub-network consisting of a fully connected layer (MLP) and a Softmax activation function, and three normalized weight coefficients g1, g2, and g3 are output, satisfying g1+g2+g3=1.

[0061] Step 2, Adaptive Fusion: Using the above weights, the shallow, middle, and deep features of the input are weighted and summed according to the following formula (1) to generate a personalized feature vector specific to the target protein. F fused : F fused = g1 F shallow + g2 F mid + g3 F deep (1) Step 3, Attention Enhancement and Prediction: The fused features undergo channel recalibration through an attention module, and finally pass through a multilayer perceptron (e.g., ...). Figure 5 The fully connected layer shown in the figure predicts the non-negative scalar protein expression value of the pixel.

[0062] Figure 5 The training process of a deep learning network is also illustrated. For example, the deep learning network can be trained according to steps 501-506 below.

[0063] In step 501, data sampling and batch processing are performed. For a training batch, multiple samples will be randomly selected, each sample representing a row or column of a specific target protein (e.g., the i-th row / column of protein P1), such as... Figure 5 As shown in "1. Sample Sampling" in the figure. This dataset determines the valid pixel coordinates of all pixels belonging to the tissue region in the i-th row / column based on the tissue region mask (e.g., ...). Figure 5 (See "2. Extract and input image patch" in the middle).

[0064] In step 502, image patch extraction is performed. For each valid pixel coordinate determined in step 501, for example, an image patch such as 43x43 pixels is extracted from the corresponding location in the high-resolution tissue image.

[0065] In step 503, shared multi-scale feature extraction is performed. All image patches extracted from the i-th row / column are batch-input into the shared multi-scale feature extractor to generate a dictionary containing shallow, medium, and deep features for each image patch.

[0066] In step 504, protein-specific expression level prediction is performed. The extracted multi-scale features are input into the target protein P (P1 to P2). n In the gated prediction head, after gating fusion and attention mechanisms, the predicted value of each pixel in the row / column is output.

[0067] In step 505, predicted value aggregation and loss function calculation are performed. The predicted values ​​v1-v of all valid pixels in the i-th row / column are aggregated. m The summation is performed to obtain a row / column "predicted value aggregation". Then, the model minimizes the composite loss function according to the following equation (2). L total To update the model parameters during training: L total = L MAE + λ1 L Corr + λ2 L TV (2) in, L MAE The absolute error loss is calculated by aggregating the predicted values. sum With aggregated ground truth sum The L1 distance between them, this loss term ensures the accuracy of the total protein expression; L Corr For Pearson correlation coefficient loss, when the number of pixels sampled in a training session is greater than 1, the Pearson correlation between the predicted vector and the distribution of the true pixel values ​​corresponding to that row / column (if there is sparse prior knowledge) or the historical iteration trend is calculated. In the case of only aggregating true value labels, this loss term mainly strengthens the model's learning of spatial distribution pattern through the "orthogonal pressure" mechanism (i.e., cross-validation of row prediction and column prediction at the same pixel), preventing the model from deceiving the summation task by simply "averaging" the values. L TV The total variational regularization loss is used to calculate the sum of the absolute values ​​of the differences between adjacent pixels in the prediction vector. This loss term is used to constrain the spatial smoothness of the prediction results, prevent the generation of isolated noise points, and conform to the prior knowledge that protein expression in biological tissues has local continuity.

[0068] In step 506, the model parameters are updated. The calculated composite loss is backpropagated. The error signal simultaneously updates the weights of the protein P-specific gated prediction head and the shared multi-scale feature extractor. This embodiment employs the Adam optimizer and sets different learning rates for the shared multi-scale feature extractor and the specific gated prediction head to achieve more stable training. This process iterates over all rows and columns of samples for all proteins until the deep learning network model converges.

[0069] Specifically, during the training process according to embodiments of this application, a pixel (r, c) simultaneously belongs to the r-th row and the c-th column, and its prediction result participates in two orthogonal aggregation calculations sequentially. The deep learning network model cleverly handles this process during training through the following mechanism, achieving self-consistency and convergence of the results without requiring any additional correlation processing steps: The first is randomized training: in each training epoch, the order of all samples (i.e., all rows and columns of all target proteins) is completely shuffled. The model does not systematically "learn all rows first and then learn columns," but rather receives supervision information from different directions in a random order.

[0070] Secondly, robust convergence is achieved through "orthogonal pressure": when the model adjusts its prediction for pixel (r,c) by fitting the sum of the r-th row, this adjustment is quickly tested when learning the c-th column. This continuous "pull" and "calibration" from two orthogonal directions can be called "orthogonal pressure." This pressure forces the model to learn a globally consistent and direction-independent mapping relationship between "biological state fingerprint" and "expression value." Ultimately, the model's weights converge to a stable state that optimally satisfies all row and column constraints simultaneously, ensuring that the prediction for any pixel is robust and self-consistent.

[0071] Thirdly, it has multi-scale perception capabilities: by integrating shallow, medium and deep features, the model can identify both tiny subcellular structures (contributed by shallow features) and macroscopic tissue partitions (contributed by deep features), significantly improving the detail clarity of the reconstructed map.

[0072] Fourth is intelligent feature selection: the gating mechanism allows the model to "adapt to the protein". For example, for widely distributed cytoskeletal proteins, the model will automatically assign higher weights to shallow texture features; while for enzymes expressed in specific functional regions, the model will rely more on deep semantic features.

[0073] Fifth is the dual precision of form and value: introducing L Corr Pearson correlation coefficient loss and L TVThe total variational regularization loss solves the chessboard effect or fuzziness problem that is easily caused by traditional summation constraints alone, making the reconstruction results not only numerically accurate, but also more realistic and natural in morphology.

[0074] like Figure 5 As shown, once the deep learning network model has been trained, the following steps can be followed to generate a high-resolution protein map covering the entire tissue region.

[0075] Step 1: Switch the shared multi-scale feature extractor and the protein-specific gating prediction head to evaluation mode.

[0076] Step 2: Global scan. For each pixel coordinate (r, c) in the entire effective tissue region defined by the tissue region mask, extract the corresponding 43x43 image patch from the high-resolution image and generate a global feature matrix. Input all these image patches in batches into the trained shared multi-scale feature extractor to generate a feature vector matrix covering the entire tissue region.

[0077] Step 3: Parallel prediction. For each protein P that needs to be predicted, its corresponding specific gating prediction head is applied to the entire feature vector matrix, so that the expression value of all pixels can be calculated at once, thereby generating a complete high-resolution protein abundance matrix of the target protein.

[0078] Step 4: Masking. The generated expression matrix is ​​masked using a tissue region mask to remove predicted values ​​from the background region, resulting in the final, accurate two-dimensional high-resolution protein expression map.

[0079] According to embodiments of this application, a device for reconstructing spatial protein expression maps is also provided. Figure 6 A partial block diagram of a spatial protein expression map reconstruction apparatus according to an embodiment of this application is shown.

[0080] like Figure 6 As shown, the reconstruction apparatus 600 includes at least an interface 601 and at least one processor 602. The interface 601 may be configured, for example, to receive a multi-channel high-resolution tissue image containing a tissue region obtained based on at least one tissue slice, aggregated true values ​​of the expression levels of each target protein on each first-direction tissue band obtained by cutting the tissue region of the first tissue slice into tissue bands in a first direction, and aggregated true values ​​of the expression levels of each target protein on each second-direction tissue band obtained by cutting the tissue region of the second tissue slice into tissue bands in a second direction.

[0081] In other embodiments, the reconstruction apparatus 600 may also include a storage area (not shown), which may be used, for example, to store pre-trained and trained deep learning networks, as well as other related images and data.

[0082] In some embodiments, interface 601 may include, for example, a network cable connector, a cable connector, a serial connector, a USB connector, a parallel connector, a high-speed data transmission adapter such as fiber optic, USB 3.0, or Xunlei, a wireless network adapter such as a WiFi adapter, or a telecommunications (3G, 4G / LTE, etc.) adapter. In some embodiments, interface 601 may, for example, directly receive required data such as multi-channel high-resolution tissue images containing tissue regions from a device (not shown) such as a high-resolution microscope or a high-definition camera, and may also acquire pre-trained deep learning networks, images of tissue regions, aggregated ground truth data of protein expression levels on tissue bands in the first and second directions from other storage media.

[0083] In some embodiments, at least one processor 602 may be configured to perform the steps of the spatial protein expression map reconstruction method according to various embodiments of this application, the specific implementation of which has been described above. Figures 1-5 The details have been explained in detail, so I will not repeat them here.

[0084] In some embodiments, at least one processor 602 may be a processing device including more than one general-purpose processing device, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be more than one special-purpose processing device, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a system-on-a-chip (SoC), etc.

[0085] According to embodiments of this application, a non-transitory computer-readable storage medium is also provided, on which computer-executable instructions are stored, wherein when the computer-executable instructions are executed by a processor, the various steps of the spatial protein expression map reconstruction method according to various embodiments of this application are performed.

[0086] In addition, this storage medium can also be used to store multi-channel high-resolution tissue images received via interface high-resolution microscopes and other devices, as well as aggregated true value data of protein expression levels on tissue bands in the first and second directions, etc., which will not be listed here.

[0087] In some embodiments, the aforementioned non-transitory computer-readable storage medium may be, for example, read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash drives or other forms of flash memory, cache, registers, static memory, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape cassette or other magnetic storage devices, or any other possible non-transitory medium used to store information or instructions that can be accessed by a computer device.

[0088] According to the spatial protein expression map reconstruction method, reconstruction apparatus, and storage medium of various embodiments of this application, aggregated supervision signals (e.g., the row sum and column sum of protein abundance) are used as the training target of the deep learning network model, instead of using pixel-level real value data. The training data is easier to obtain and the values ​​are more accurate. By using a convolutional neural network feature extraction module shared among multiple target molecules (such as proteins) and multiple independent prediction head networks that correspond one-to-one with specific target proteins, specific target features can be extracted for various types of proteins. The predicted protein abundance expression values ​​after the model is trained are more accurate. Through the careful design of the loss function and training process adapted to the various protein expression prediction tasks of this application, it is ensured that the nonlinear relationship between the underlying, complex spatial organization features and protein expression can be automatically learned from high-resolution microscopic images. Compared with traditional mathematical interpolation methods, the abundance matrix predicted by the trained deep learning network model can generate a more refined, accurate, morphologically more natural, and biologically realistic high-resolution expression map. Furthermore, the reconstruction method according to the embodiments of this application cleverly avoids the dependence on expensive and hard-to-obtain pixel-level ground truth maps. By using more readily available aggregated data (row and column sums) as a supervision signal, it greatly reduces the data requirements and experimental costs and technical barriers of spatial proteomics research, including labor and equipment. In addition, the "shared feature extractor + independent prediction head" architecture adopted in this application makes the model more computationally efficient when processing multiple proteins. Moreover, when a new protein needs to be modeled, there is no need to consume new tissue slices or retrain the entire large system from scratch. Only the shared module needs to be fixed, and a lightweight new prediction head needs to be added and trained, which combines high efficiency and good scalability.

[0089] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, which will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered illustrative only, and the true scope and spirit are indicated by the full scope of the claims and their equivalents.

[0090] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments may be used by those skilled in the art upon reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of the application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being able to be combined with each other in various combinations or arrangements. The scope of this application should be determined by reference to the claims and the full scope of their equivalents.

[0091] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.

Claims

1. A method for reconstructing a spatial protein expression map, characterized in that, include: Receive a multi-channel, high-resolution tissue image containing tissue regions, obtained based on at least one tissue slice; Receive the aggregated true value of the expression level of each target protein on each first direction tissue band obtained after cutting the tissue region into a first direction tissue band in the first tissue section, and the aggregated true value of the expression level of each target protein on each second direction tissue band obtained after cutting the tissue region into a second direction tissue band in the second tissue section. The deep learning network is trained based on the aggregated true values ​​of the expression levels of each target protein in each first-direction tissue band and each second-direction tissue band, as well as the multi-channel high-resolution tissue image. Based on the multi-channel high-resolution tissue images, a trained deep learning network is used to predict the expression levels of each target protein in each image unit of the tissue region and to reconstruct the protein expression map in the two-dimensional space of the tissue region.

2. The reconstruction method according to claim 1, characterized in that, Before training the deep learning network using the multi-channel high-resolution tissue images, and before predicting protein expression profiles using the multi-channel high-resolution tissue images, no identification or labeling related to cell phenotype or tissue structure morphology is performed on the multi-channel high-resolution tissue images, wherein the identification or labeling related to cell phenotype or tissue structure morphology is performed manually or automatically by a computer.

3. The reconstruction method according to claim 1 or 2, characterized in that, The deep learning network sequentially includes a feature extraction module and a protein-specific prediction module. The feature extraction module is shared by all target proteins, while the protein-specific prediction module contains an independent prediction head corresponding to each target protein. The training of the deep learning network based on the aggregated true values ​​of the expression levels of each target protein in each first-direction tissue band and each second-direction tissue band, and the multi-channel high-resolution tissue image specifically includes: According to the cutting method of the first direction tissue strip and the second direction tissue strip, the tissue region in the multi-channel high-resolution tissue image is divided into multiple image units; The deep learning network is trained in multiple rounds until the training objective or a specified number of rounds is reached. In each round of training, a specified number of tissue bands in the first and second directions are randomly selected. The image units corresponding to the selected tissue bands and the aggregated ground truth values ​​of the expression levels of each target protein on the tissue bands are used as training samples, and the following steps are performed: Step 1: Use the feature extraction module to generate the feature encoding vector corresponding to each image unit in the current first-direction tissue strip / second-direction tissue strip; Step 2: For each target protein, based on the feature encoding vector corresponding to each image unit, the protein-specific prediction module generates a predicted value of the expression level of the target protein in each image unit in the current first-direction tissue band / second-direction tissue band using each independent prediction head. Step 3: For each target protein, calculate the first loss function based on the sum of the predicted expression levels of the target protein in each image unit of the current first-direction tissue band / second-direction tissue band, and the aggregated true value of the expression level of the target protein in the current first-direction tissue band / second-direction tissue band. Step 4: Based on the first loss function of each target protein in each first direction tissue band and each second direction tissue band, calculate the second loss function for training the deep learning network, and determine whether the training objective has been achieved based on the second loss function.

4. The reconstruction method according to claim 3, characterized in that, The reconstruction method further includes generating an effective tissue region mask based on the multi-channel high-resolution tissue image. The step of dividing the tissue region in the multi-channel high-resolution tissue image into multiple image units according to the cutting method of the first direction tissue strip and the second direction tissue strip further includes: dividing the tissue region in the multi-channel high-resolution tissue image into multiple image units according to the cutting method of the first direction tissue strip and the second direction tissue strip; applying the effective tissue region mask to the multiple image units, so that the multiple image units only include image units in the effective tissue region.

5. The reconstruction method according to claim 4, characterized in that, The step of using a trained deep learning network to predict the expression levels of each target protein in each image unit of the tissue region based on the multi-channel high-resolution tissue image and then using this prediction for the reconstruction of a two-dimensional protein expression map of the tissue region further includes: Each image unit in the effective tissue region is input into a trained deep learning network in a preset order to generate a proteome abundance matrix composed of the expression levels of each target protein in each image unit of the effective tissue region. Based on the protein abundance matrix, a two-dimensional protein expression map of the effective tissue region is reconstructed.

6. The reconstruction method according to claim 1 or 2, characterized in that, The first-direction tissue strip cutting and the second-direction tissue strip cutting can achieve single-cell scale, and, When tissue strip cutting in the first direction and the second direction reaches the single-cell scale, it is possible to reconstruct the protein expression map of the two-dimensional space of the tissue region at the single-cell scale.

7. The reconstruction method according to claim 3, characterized in that, The reconstruction method further includes, in the case of adding a target protein, Add an independent prediction head corresponding to the newly added target protein to the protein-specific prediction module; The deep learning network with the newly added independent prediction head was further trained using the aggregated true values ​​of the expression levels of the newly added target protein in each first-direction tissue band and each second-direction tissue band, and the multi-channel high-resolution tissue image. Based on the multi-channel high-resolution tissue images, a pre-trained deep learning network is used to predict the expression levels of newly added target proteins in each image unit of the tissue region and to reconstruct the protein expression map of the tissue region in two-dimensional space.

8. The reconstruction method according to claim 1 or 2, characterized in that, The multi-channel high-resolution tissue images were acquired as follows: Reference protein markers are designed based on the tissue spatial structure of the tissue region, cell type heterogeneity, and complementarity between protein markers. The tissue sections were stained with immunofluorescent antibodies using various reference protein markers. Multi-channel high-resolution tissue images were obtained from tissue sections stained with immunofluorescence antibodies.

9. A device for reconstructing a spatial protein expression map, characterized in that, include: An interface configured to receive: a multi-channel high-resolution tissue image containing a tissue region obtained based on at least one tissue slice; aggregated true values ​​of the expression levels of each target protein on each first-direction tissue band obtained by cutting the tissue region of the first tissue slice into tissue bands in a first direction; and aggregated true values ​​of the expression levels of each target protein on each second-direction tissue band obtained by cutting the tissue region of the second tissue slice into tissue bands in a second direction. At least one processor is configured to perform the steps of the method for reconstructing a spatial protein expression map according to any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the steps of the method for reconstructing a spatial protein expression map according to any one of claims 1 to 8 are performed.