A method for constructing a healthy wuzhishan pig multi-omics spatial atlas based on a molecular biological information analysis processing system
By constructing a set of molecular anchor points specific to the Wuzhishan pig breed and performing anchor point correction and cross-slice spatial registration, the problems of insufficient accuracy and consistency in the construction of multi-omics spatial maps of healthy Wuzhishan pigs in existing technologies have been solved, and more accurate and stable multi-omics spatial map construction has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANIMAL HUSBANDRY & VETERINARY RES INST OF HAINAN ACAD OF AGRI SCI
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies, when constructing multi-omics spatial maps of healthy Wuzhishan pigs, struggle to improve the accuracy of cross-transcriptome, proteome, and metabolome spatial association results and the consistency of cross-individual results based on continuous slice spatial integration, especially with insufficient consideration of the genetic background of specific pig breeds.
By constructing a breed-specific molecular anchor set for Wuzhishan pigs, anchor point correction was performed on spatial transcriptomic, proteomic, and metabolomic data based on a molecular bioinformatics analysis and processing system. Cross-slice spatial registration was performed in conjunction with histological relay sections. An anchor point relay triads were constructed within spatial microdomain units, and conserved spatial microdomains were extracted among different individuals to construct a multi-omics spatial map of healthy Wuzhishan pigs.
It improves the accuracy of spatial correlation results between different molecular layers, weakens the interference of individual differences on map construction results, reduces the proportion of metabolic features falsely retained, improves the authenticity and credibility of three-layer joint analysis results, enhances the boundary accuracy and functional purity of spatial micro-domain division, and constructs a map with cross-tissue functional connection and three-dimensional structural expression capabilities.
Smart Images

Figure CN122337356A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics, spatial omics data processing, and molecular map construction, and particularly to a method for constructing a multi-omics spatial map of healthy Wuzhishan pigs based on a molecular bioinformatics analysis and processing system. Background Technology
[0002] With the development of spatial omics technology, the demand for obtaining multi-level molecular information from the same tissue sample and performing joint analysis is constantly increasing. In practical applications, due to limitations in throughput, detection type, or sample carrying capacity of a single slice, sequential slices from the same tissue block are often used to perform different omics detections to obtain multi-level data such as transcriptomics, proteomics, and metabolomics. Therefore, how to establish stable spatial correspondences between different sequential slices and realize joint analysis of multi-omics data on this basis has become a focus of related technical fields.
[0003] Existing technologies such as SpatialEx / SpatialEx+ propose using histological images as universal anchor points to integrate spatial molecular data from consecutive tissue sections. This type of technology establishes mapping relationships between different sections using histological morphological information, thereby achieving alignment and integration of spatial data across sections, and is applicable to non-overlapping or heterogeneous tissue sections. Therefore, the main technical approach of these existing technologies revolves around "consecutive sections—histological anchoring—cross-section spatial integration" to address the problem of spatial correspondence and joint analysis when multiple omics data originate from different sections.
[0004] However, the key technical focus of the aforementioned existing technologies lies primarily in the spatial integration between different consecutive slices, with an emphasis on slice-level spatial mapping and the multi-omics fusion framework itself. For local pig breeds like healthy Wuzhishan pigs with specific genetic backgrounds, relying solely on histological morphological information to establish cross-slice correspondences when constructing multi-omics spatial maps still struggles to adequately ensure the accuracy of spatial association results between different molecular levels and the consistency of map results across different individuals. In other words, while existing technologies can achieve multi-omics spatial integration at the consecutive slice level, there is still room for improvement in constructing multi-omics spatial maps with better cross-individual stability for specific pig breeds.
[0005] Therefore, the main technical problem that the existing technology still needs to solve is: when constructing multi-omics spatial maps for healthy Wuzhishan pigs, how to improve the accuracy of spatial association results across transcriptomics, proteomics and metabolomics and the consistency of results across individuals based on continuous slice spatial integration. Summary of the Invention
[0006] To overcome the aforementioned technical deficiencies, the present invention aims to provide a method for constructing a multi-omics spatial map of healthy Wuzhishan pigs based on a molecular bioinformatics analysis and processing system. This application constructs a breed-specific molecular anchor set for Wuzhishan pigs, performs anchor point correction on spatial transcriptome, spatial proteome, and spatial metabolome data based on this set, performs cross-slice spatial registration using histological relay sections, constructs anchor point relay triplets within spatial microdomain units, and extracts conserved spatial microdomains among different individuals, thereby constructing a multi-omics spatial map of healthy Wuzhishan pigs.
[0007] This invention discloses a method for constructing a multi-omics spatial map of healthy Wuzhishan pigs based on a molecular bioinformatics analysis and processing system, comprising the following steps:
[0008] S1. Obtain whole-genome variation data from at least 3 healthy Wuzhishan pigs, and construct a Wuzhishan pig breed-specific molecular anchor set by combining pig pan-genome reference and Wuzhishan pig full-length transcript data. The Wuzhishan pig breed-specific molecular anchor set shall include at least breed-conserved single nucleotide variant sites, breed-conserved structural variant sites, variant-supported splicing sites, and variant-supported coding peptide sites.
[0009] S2. Obtain tissue block samples of at least one target tissue from each healthy Wuzhishan pig, and prepare consecutive adjacent slices from the same tissue block for each tissue block sample. Consecutive adjacent slices include at least histological relay slices, spatial transcriptome slices, spatial proteome slices and spatial metabolome slices.
[0010] S3. Based on the Wuzhishan pig breed-specific molecular anchor set, the molecular bioinformatics analysis and processing system performs anchor point correction on the raw data generated by spatial transcriptome slices, spatial proteome slices and spatial metabolome slices respectively to obtain the corrected transcription feature matrix, corrected protein feature matrix and corrected metabolic feature matrix.
[0011] S4. Using histological relay slices as displacement transfer layers, cross-slice spatial registration was performed on spatial transcriptome slices, spatial proteome slices, and spatial metabolome slices, and the registration results were uniformly projected to spatial micro-domain units.
[0012] S5. For each spatial micro-domain unit, only when the corrective transcription feature carrying the same molecular anchor point, the corrective protein feature encoded by the corrective transcription feature and covering the same molecular anchor point, and the corrective metabolic feature directly related to the enzyme-catalyzed reaction corresponding to the corrective protein feature are detected simultaneously, the corrective transcription feature, the corrective protein feature and the corrective metabolic feature are identified as the anchor point relay triplet.
[0013] S6. Spatial registration of homologous tissue regions was performed on the anchor point relay triplet among different healthy Wuzhishan pigs in the same target tissue, and conservative spatial micro-domains were extracted based on repetition, spatial continuity and consistency of reaction order.
[0014] S7. Construct a multi-omics spatial map of healthy Wuzhishan pigs based on the conservative spatial microdomains of each target organization.
[0015] Preferably, constructing a Wuzhishan pig breed-specific molecular anchor set includes: mapping the whole genome variation data of each healthy Wuzhishan pig to a pig linear reference genome and a pig pan-genome reference, respectively, and screening out structural variation breakpoints among exon variation sites, splice boundary variation sites, insertion / deletion variation sites, and breed-conserved structural variation sites that are repeated in no less than 70% of healthy Wuzhishan pigs and are simultaneously supported by Wuzhishan pig full-length transcript data, as the Wuzhishan pig breed-specific molecular anchor set.
[0016] Preferably, the Wuzhishan pig breed-specific molecular anchor set also meets the following conditions: at least a portion of the breed-conserved single nucleotide variant sites, variant-supported splicing connection sites, and structural variant breakpoints are located in the coding region, splicing boundary region, or promoter region, and correspond to transcript structural change information that is repeated among different healthy Wuzhishan pigs.
[0017] Preferably, the transcription feature correction in anchor point correction includes: for transcription features whose alignment and localization results given by the pig linear reference genome and the pig pan-genome reference are inconsistent, a graph reference path containing the Wuzhishan pig breed-specific molecular anchor set is preferentially used for relocalization, and the anchor transcript model is reconstructed based on the Wuzhishan pig full-length transcript data to obtain the corrected transcription feature matrix.
[0018] Preferably, the protein feature correction in anchor point correction includes: constructing an anchor point coding sequence library based on the anchor point transcript model, and rematching the peptide features generated by spatial proteome slicing to the anchor point coding sequences in the anchor point coding sequence library, retaining only the peptide features that cross the amino acid change region corresponding to the variety's conserved single nucleotide variation site, cross the splice connection site supported by the variation, or cover the adjacent region of the structural variation breakpoint in the variety's conserved structural variation site, as the protein features in the corrected protein feature matrix.
[0019] Preferably, the metabolic feature correction in anchor point correction includes: retaining a metabolic feature as a metabolic feature in the correction metabolic feature matrix only when a certain metabolic feature corresponds to the direct substrate or direct product of the enzyme-catalyzed reaction characterized by the correction protein feature, and the metabolic feature, the correction protein feature, and the correction transcription feature are located in the same spatial micro-domain unit.
[0020] Preferably, cross-slice spatial registration includes: using the histological relay slice as the relay reference, selecting at least three common anatomical landmarks from the microvascular boundary, glandular duct boundary, connective tissue boundary and muscle bundle boundary, constructing a segmented displacement field from the histological relay slice to the spatial transcriptome slice, spatial proteome slice and spatial metabolome slice, and projecting the data in the spatial transcriptome slice, spatial proteome slice and spatial metabolome slice to the spatial microdomain unit based on the segmented displacement field.
[0021] Preferably, the boundary of a spatial micro-domain unit is established only when the following conditions are met simultaneously: there is a morphological boundary in the histological relay section, and at least two of the corrected transcription feature matrix, the corrected protein feature matrix, and the corrected metabolic feature matrix have omics boundaries at corresponding positions.
[0022] Preferably, the anchor relay triad is established only under the following conditions: the corrected transcriptional feature, the corrected protein feature, and the corrected metabolic feature correspond to the same reaction entry, or to two adjacent reaction entries, and the corrected metabolic feature is the substrate or product of the enzyme-catalyzed reaction characterized by the corrected protein feature.
[0023] Preferably, when the same anchor point relay triplets appear simultaneously in the current spatial micro-domain unit and its first ring of adjacent spatial micro-domain units, and maintain the same upstream to downstream enrichment direction in the current spatial micro-domain unit and its first ring of adjacent spatial micro-domain units, while there is no opposite enrichment direction in the reverse adjacent spatial micro-domain units relative to the upstream to downstream enrichment direction, the anchor point relay triplets are determined as stable anchor point relay units.
[0024] Preferably, the conservative spatial microdomain is extracted in the following way: spatial registration is performed on the spatial microdomain units of different healthy Wuzhishan pigs in the same target tissue, and only the region in which no less than 80% of the healthy Wuzhishan pigs contain the same stable anchor point relay unit, and the centroid offset between the corresponding spatial microdomain units does not exceed one-third of the length of the short side of the smallest circumscribed rectangle of the smaller spatial microdomain unit, is retained as the conservative spatial microdomain.
[0025] Preferably, multiple stable anchor relay units that share the same anchor transcript model or the same anchor coding sequence within each conserved spatial microdomain are aggregated to form a pathway core module, and the pathway core module is used as the main annotation unit of the conserved spatial microdomain.
[0026] Preferably, the pathway core modules in at least two different target tissues of the same healthy Wuzhishan pig are compared. When the conserved spatial microdomains in different target tissues contain the same pathway core modules and the upstream to downstream arrangement order of each stable anchor relay unit is consistent, a cross-tissue homologous functional axis is established between different target tissues.
[0027] Preferably, for adjacent consecutive slices of the same tissue block sample, when the spatially registered spatial micro-domain units in adjacent consecutive slices contain the same pathway core module, and their projected overlap area accounts for more than 60% of the area of the smaller spatial micro-domain unit, the corresponding spatial micro-domain units in adjacent consecutive slices are merged into three-dimensional conservative micro-pillars.
[0028] Preferably, constructing a multi-omics spatial map of healthy Wuzhishan pigs includes: using three-dimensional conserved micropillars as intra-tissue nodes and cross-tissue homologous functional axes as cross-tissue connection edges to construct a multi-omics spatial map of healthy Wuzhishan pigs that simultaneously includes intra-tissue three-dimensional spatial relationships and cross-tissue functional relationships. In each intra-tissue node, the source information of the corresponding Wuzhishan pig breed-specific molecular anchor set and the upstream-to-downstream arrangement information of the stable anchor relay units constituting the core module of the pathway are associated and recorded. When a new healthy Wuzhishan pig sample is imported, consistency verification is first performed based on the whole-genome variation data of the new healthy Wuzhishan pig sample and the Wuzhishan pig breed-specific molecular anchor set. Then, anchor point correction, anchor point relay triplet identification, stable anchor relay unit determination, and pathway core module identification are performed on the new healthy Wuzhishan pig sample. Only when the pathway core module formed by the new sample appears repeatedly in at least two other new healthy Wuzhishan pigs is the conserved spatial microdomain corresponding to the pathway core module and the three-dimensional conserved micropillars formed therefrom incorporated into the multi-omics spatial map of healthy Wuzhishan pigs.
[0029] Compared with existing technologies, the above technical solution has the following advantages:
[0030] 1. In existing technologies, spatial integration schemes for continuous sections mainly rely on histological images to establish spatial correspondences between sections. While this can achieve spatial alignment between different sections, it lacks a unified constraint on the true correspondences between different molecular layers, tailored to the specific genetic background of the breed. This can easily lead to inaccurate associations between the transcriptional, proteomic, and metabolic layers. This invention constructs a set of Wuzhishan pig breed-specific molecular anchor points and performs anchor point correction on spatial transcriptomic, proteomic, and metabolomic data based on this set. Simultaneously, it constructs an anchor point relay triad within spatial micro-domain units, thereby effectively improving the accuracy of spatial association results between different molecular layers.
[0031] 2. Existing spatial omics integration methods typically focus more on the integration analysis of single samples or single slices, neglecting the consistency of results between different individuals. This makes the map results susceptible to individual differences, local noise, or random signals. This invention extracts conserved spatial microdomains among different healthy Wuzhishan pigs and screens stable anchor relay units based on repetition, spatial continuity, and consistency of reaction order. This effectively reduces the interference of random individual differences on the map construction results, thereby improving the repeatability and stability of multi-omics spatial maps among different healthy Wuzhishan pig individuals.
[0032] 3. In existing technologies, metabolomics data are prone to introducing numerous false-positive retention signals during spatial integration due to the complex sources of metabolites and significant diffusion effects, thus affecting the reliability of the overall analysis results. This invention limits the corrected metabolic features to substrates or products directly related to the enzymatic reactions characterized by the corrected protein features, and requires them to correspond to the same spatial micro-domain unit as the corrected transcriptional features and the corrected protein features. Based on this, spatial metabolomics data is screened, thereby significantly reducing the proportion of falsely retained metabolic features and improving the authenticity and reliability of the three-layer combined analysis results.
[0033] 4. In existing technologies, spatial region division often relies solely on histological morphological boundaries or a single omics boundary, which can easily lead to overly wide boundaries, boundary drift, or functional mixing. This invention, in the process of establishing spatial micro-domain units, simultaneously introduces morphological boundaries from histological relay sections and at least two types of omics boundaries as dual constraints. This ensures that the resulting spatial micro-domain units possess both structural rationality and molecular functional consistency, thereby improving the boundary accuracy and functional purity of the spatial micro-domain division results.
[0034] 5. Existing spatial omics schemes typically focus more on local spatial relationships within a single tissue, making it difficult to simultaneously reflect functional correspondences between different tissues and the three-dimensional structural continuity between consecutive slices. This invention establishes cross-tissue homologous functional axes by comparing core pathway modules in different target tissues, and forms three-dimensional conserved micropillars by merging corresponding spatial micro-domain units in adjacent consecutive slices. Thus, the final constructed multi-omics spatial map of healthy Wuzhishan pigs simultaneously possesses the ability to express cross-tissue functional connections and three-dimensional spatial structures within tissues.
[0035] 6. In existing technologies, map updates often involve directly incorporating new samples, which can easily introduce random signals, local anomalies, or individual-specific noise into the map, affecting the stability of existing maps. This invention, when incorporating new samples, first performs consistency verification with the Wuzhishan pig breed-specific molecular anchor set. Then, it requires that the core pathway modules formed by the new samples appear repeatedly in at least two other newly added healthy Wuzhishan pigs before incorporating the corresponding conserved spatial microdomains and three-dimensional conserved micropillars into the existing map. This improves the quality control capability and long-term stability during incremental map updates.
[0036] 7. Since the core of this invention is not limited to a specific tissue, but revolves around the technical main lines of Wuzhishan pig breed-specific molecular anchor point set, cross-layer anchor point correction, cross-slice spatial registration, stable anchor point relay unit screening, and conservative spatial micro-domain extraction, this invention is not only applicable to target tissues such as liver, spleen, and jejunum, but can also be extended to other healthy Wuzhishan pig tissues such as lung tissue and longissimus dorsi muscle, and has a good tissue applicability and application scalability. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the overall process of a method for constructing a multi-omics spatial map of healthy Wuzhishan pigs based on a molecular bioinformatics analysis and processing system according to the present invention.
[0038] Figure 2 This is a schematic diagram illustrating the relationship between consecutive adjacent sections and histological relay sections in this invention;
[0039] Figure 3 This is a schematic diagram of the process for constructing a specific molecular anchor set for the Wuzhishan pig breed and correcting cross-layer anchor points in this invention.
[0040] Figure 4 This is a schematic diagram of the spatial micro-domain unit, the anchor point relay triplet, and the formation of the conservative spatial micro-domain in this invention;
[0041] Figure 5 This is a schematic diagram of the three-dimensional conservative micropillars and cross-tissue homologous functional axes in this invention;
[0042] Figure 6 This is a curve comparing the accuracy of three-layer spatial association between Embodiment 1 and Comparative Example 1 of the present invention;
[0043] Figure 7 This is a bar chart comparing the key performance indicators of Embodiment 1 and Comparative Example 1 of the present invention. Detailed Implementation
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment focuses on the joint analysis and processing of spatial transcriptome data, spatial proteome data, and spatial metabolome data generated by healthy Wuzhishan pigs under continuous adjacent slice conditions. The key is to construct a Wuzhishan pig breed-specific molecular anchor set to complete cross-layer anchor point correction, cross-slice spatial registration, anchor point relay triplet identification, stable anchor point relay unit screening, conserved spatial microdomain extraction, pathway core module aggregation, cross-tissue homology functional axis establishment, three-dimensional conserved micropillar formation, and map incremental update, thereby constructing a multi-omics spatial map of healthy Wuzhishan pigs.
[0045] Example 1:
[0046] This embodiment selected six healthy Wuzhishan pigs as the basic sample, numbered Pig No. 1, Pig No. 2, Pig No. 3, Pig No. 4, Pig No. 5, and Pig No. 6, aged 5.5 to 6.5 months and weighing 18.4 kg to 21.7 kg. All samples underwent clinical examination, routine blood tests, blood biochemical tests, and nucleic acid testing for major porcine pathogens before collection to confirm that none of the healthy Wuzhishan pigs showed significant abnormalities. Liver, spleen, and jejunum were collected from each healthy Wuzhishan pig as target tissues, and tissue blocks with intact tissue structure were excised from each target tissue. Peripheral blood samples were also collected for whole-genome variation data acquisition. To ensure comparability between different omics data, tissue block samples of each target tissue were selected from the same adjacent anatomical location, and the time from excision to liquid nitrogen flash freezing was controlled within 30 minutes. Hereinafter, without ambiguity, "tissue block sample of the target tissue" will be referred to simply as "tissue block sample".
[0047] First, a breed-specific molecular anchor set for Wuzhishan pigs was constructed. High-throughput whole-genome resequencing was performed on genomic samples from six healthy Wuzhishan pigs, with an average sequencing depth of 31.6-fold. Total ribonucleic acid was extracted from liver, spleen, and jejunum tissues, and full-length transcriptome sequencing was performed. Subsequently, the obtained whole-genome variation data were mapped to a pig linear reference genome and a pig pan-genome reference, respectively, to screen for structural breakpoints among exon variant sites, splice boundary variant sites, insertion / deletion variant sites, and breed-conserved structural variant sites. Cross-validation was performed using Wuzhishan pig full-length transcriptome data. The screening criteria were: candidate anchors must be repeated in at least five of the six healthy Wuzhishan pigs, have structural support in the Wuzhishan pig full-length transcriptome data, and at least some candidate anchors must be located in coding regions, splice boundary regions, or promoter regions, corresponding to repeatable transcriptome structural changes among different healthy Wuzhishan pigs. Candidate sites meeting the above criteria were included in the Wuzhishan pig breed-specific molecular anchor set.
[0048] In this embodiment, the proportion of candidate anchor points is determined using the following formula:
[0049]
[0050] in, This indicates the proportion of candidate anchor points. This indicates the number of healthy Wuzhishan pigs detected at this candidate anchor point. This represents the total number of healthy Wuzhishan pigs. In this embodiment, For example, if 5 healthy Wuzhishan pigs were detected at a certain liver candidate site, the proportion of that candidate anchor site would be calculated as follows:
[0051]
[0052] Since 83.33% is higher than the 70% screening threshold set in this embodiment, this candidate site can proceed to the subsequent verification process. If, in conjunction with the full-length transcript data of Wuzhishan pigs, it is confirmed that it corresponds to changes in transcript structure, then this candidate site will ultimately be included in the Wuzhishan pig breed-specific molecular anchor set. After screening, Table 1 shows the statistical results of the Wuzhishan pig breed-specific molecular anchor set in this embodiment.
[0053] Table 1 Statistical Results of Wuzhishan Pig Breed-Specific Molecular Anchor Set
[0054] Target organization Number of conserved single nucleotide variant sites in a variety Number of splice junction sites supported by the variant Number of insertion / deletion variant sites Number of structural variation breakpoints Number of coding peptide sites supported by the mutation liver 4286 362 286 131 1068 spleen 4015 341 251 126 1037 jejunum 4542 389 309 143 1124 After deduplication 10184 876 691 318 2746
[0055] As shown in Table 1, insertion / deletion variants were identified and counted separately as an important component of the Wuzhishan pig breed-specific molecular anchor set in this embodiment. 286, 251, and 309 insertion / deletion variants were obtained from the liver, spleen, and jejunum, respectively, totaling 691 after deduplication. Further cross-validation with the full-length transcript data of Wuzhishan pigs revealed that some of these insertion / deletion variants corresponded to changes in the length of the first exon, elongation of the 5' untranslated region, or changes in the position of splice junction sites. This indicates that insertion / deletion variants are not merely simple variation information but can also have a real impact on transcript structure.
[0056] To further illustrate the correspondence between promoter region anchors and transcript structural changes, this embodiment selects a representative promoter region example from a liver tissue sample. In pig liver sample No. 1, a 7-base insertion / deletion variant was detected approximately 182 bases upstream of the main transcription start site in the promoter region of a lipid metabolism-related gene. This insertion / deletion variant was repeated in 5 out of 6 healthy Wuzhishan pigs, representing a occurrence rate of 83.33%. Further analysis using Wuzhishan pig full-length transcript data revealed that in samples carrying this insertion / deletion variant, the 5' end of the main transcript of the lipid metabolism-related gene was extended upstream by 41 nucleotides relative to the general reference model, forming a stable first-terminal transcript structural change, while the connection between the first and downstream exons remained uninterrupted. This demonstrates that breed-specific variations in promoter regions can also serve as components of the Wuzhishan pig breed-specific molecular anchor set and can be validated in Wuzhishan pig full-length transcript data through transcript structural changes. Table 2 provides illustrative statistics for this instance of the starter subregion.
[0057] Table 2. Representative examples of promoter region variations supporting transcript structural alterations.
[0058] Target organization Mutation location type Variant forms The number of occurrences in 6 healthy Wuzhishan pigs Occurrence rate Observed transcript structural changes liver Startup sub-region 7 base insertion / deletion variant sites 5 83.33% The 5' end of the main transcript is extended by 41 nucleotides, forming a stable head-end structural alteration.
[0059] After constructing the Wuzhishan pig breed-specific molecular anchor set, consecutive adjacent sections from the same tissue block sample were prepared for each target tissue block sample. A total of 5 consecutive adjacent sections were prepared for each tissue block sample: the first was a histological relay section with a thickness of 5 μm; the second was a spatial transcriptome section with a thickness of 10 μm; the third was a spatial proteome section with a thickness of 8 μm; the fourth was a spatial metabolome section with a thickness of 12 μm; and the fifth was a backup tissue verification section with a thickness of 5 μm. The spacing between adjacent consecutive sections was controlled within 15 μm. Histological images of the histological relay sections were obtained using hematoxylin-eosin staining; raw transcriptional features of the spatial transcriptome sections were obtained using a capture probe array; raw peptide features of the spatial proteome sections were obtained using antibody labeling combined with imaging mass spectrometry; and raw metabolic features of the spatial metabolome sections were obtained using matrix-assisted laser desorption / ionization imaging mass spectrometry. Figure 2 The relative relationships between the above consecutive adjacent sections and the relay role of histological relay sections are shown.
[0060] Anchor point correction was then performed. For the raw transcriptional features generated from spatial transcriptome slices, they were first mapped to both the pig linear reference genome and the pig pan-genome reference genome. Transcriptional features with inconsistent alignment results were screened out. Then, a graph reference path containing the Wuzhishan pig breed-specific molecular anchor set was preferentially used for relocation. Anchor point transcript models were reconstructed using Wuzhishan pig full-length transcript data. Only transcriptional features that could be stably mapped to the anchor point transcript model and matched the Wuzhishan pig breed-specific molecular anchor set were retained, forming a corrected transcriptional feature matrix. For the raw peptide features generated from spatial proteome slices, an anchor point coding sequence library was formed based on the anchor point transcript model. The raw peptide features were then re-matched to the anchor point coding sequences in the library. Only peptide features covering amino acid changes corresponding to conserved single nucleotide variant sites, splicing sites supported by variants, or adjacent regions of structural variant breakpoints were retained, forming a corrected protein feature matrix. For the raw metabolic features generated from spatial metabolomics slices, only those metabolic features meeting the following three conditions are retained: First, the metabolic feature corresponds to the direct substrate or direct product of the enzymatic reaction characterized by the corrected protein feature; second, the metabolic feature, corrected protein feature, and corrected transcription feature correspond to the same spatial microdomain unit after cross-slice spatial registration; third, the metabolic feature is repeated in at least 5 of the 6 healthy Wuzhishan pigs. This process yields the corrected metabolic feature matrix. Figure 3 The cross-layer anchor point correction logic is shown.
[0061] This embodiment provides a specific peptide example. In a representative region of a pig liver tissue block sample, the anchor coding sequence generated after translation from a certain anchor transcript model contains an amino acid substitution segment caused by a breed-conserved single nucleotide variant site. After re-matching the original peptide features, a peptide of 12 amino acid residues in length is obtained. This peptide completely covers the amino acid substitution segment and appears in the spatial proteome slice along with the corrected transcriptional and corrected metabolic features in the same region. Therefore, this peptide feature is preserved and included in the corrected protein feature matrix. This specific example illustrates that the correspondence between the anchor coding sequence and the peptide feature is not an abstract description but can be directly supported by actual detection data.
[0062] After obtaining the corrected transcriptional feature matrix, corrected protein feature matrix, and corrected metabolic feature matrix, cross-slice spatial registration is performed. Specifically, using histological relay slices as relay benchmarks, at least three of the four common anatomical landmarks—microvascular boundaries, glandular duct boundaries, connective tissue boundaries, and muscle bundle boundaries—are selected as the basis for cross-slice spatial registration in tissue block samples from different target tissues. Specifically, in liver tissue block samples, microvascular boundaries, bile duct boundaries, and connective tissue boundaries are preferred, with the bile duct boundaries belonging to the glandular duct boundary category; in spleen tissue block samples, microvascular boundaries, connective tissue boundaries, and splenic trabecular muscle bundle boundaries are preferred, with the splenic trabecular muscle bundle boundaries belonging to the muscle bundle boundary category; and in jejunal tissue block samples, microvascular boundaries, intestinal glandular duct boundaries, and muscle layer muscle bundle boundaries are preferred, with the intestinal glandular duct boundaries belonging to the glandular duct boundary category and the muscle layer muscle bundle boundaries belonging to the muscle bundle boundary category.
[0063] Subsequently, segmented displacement fields were established using selected common anatomical landmarks, pointing histological relay slices to spatial transcriptome, spatial proteome, and spatial metabolome slices, respectively. Data from different slices were then uniformly projected onto the same coordinate frame. Spatial micro-domains were not directly divided according to a fixed grid, but rather established under the dual constraints of morphological and omics boundaries. That is, spatial micro-domains were defined only when the histological relay slice had a morphological boundary at the corresponding position, and at least two of the corrected transcriptional feature matrix, corrected proteome feature matrix, and corrected metabolic feature matrix had omics boundaries at the corresponding positions. Figure 4 The left side illustrates how this type of spatial micro-domain unit is formed.
[0064] After forming spatial microdomains, anchor relay triads were identified within each microdomain. Anchor relay triads were required to simultaneously satisfy the following relationships: the corrected transcriptional feature, corrected protein feature, and corrected metabolic feature corresponded to the same reaction entry, or to two adjacent reaction entries, and the corrected metabolic feature was the substrate or product of the enzymatic reaction characterized by the corrected protein feature. Through this screening, representative anchor relay triads were obtained in the liver, spleen, and jejunum. Table 3 lists some representative examples.
[0065] Table 3 shows some representative examples of anchor point relay triplets.
[0066] Target organization Anchor transcript model corresponding gene regions Anchor coding sequence corresponds to protein function Corrected metabolic characteristic type Reaction item relationship Location of space liver Fatty acid oxidation-related transcript model Mitochondrial acyl-CoA oxidation-related protein Acyl-CoA downstream metabolites Same reaction entry Perihepatic cord area spleen Immune response-related transcript model Antigen processing-related proteins Peptide cleavage downstream metabolites Adjacent reaction entries adjacent areas of splenic trabeculae jejunum Transport-related membrane protein transcript model Transporter protein complex subunit Small molecule substrate-related metabolites Same reaction entry Adjacent areas of intestinal glands
[0067] After obtaining the anchor point relay triplets, stable anchor point relay units are further screened. Specifically, each anchor point relay triplet must appear simultaneously in the current spatial micro-domain unit and its first ring of adjacent spatial micro-domain units, maintain the same upstream-to-downstream enrichment direction in both the current and first rings of adjacent spatial micro-domain units, and have no opposite enrichment direction in the reverse-adjacent spatial micro-domain units relative to this upstream-to-downstream enrichment direction. Only when all of the above conditions are met is the corresponding anchor point relay triplet identified as a stable anchor point relay unit.
[0068] Subsequently, conservative spatial microdomains were extracted among different healthy Wuzhishan pigs. First, spatial microdomain units from six healthy Wuzhishan pigs within the same target tissue were spatially registered to homologous tissue regions. Then, the repetition of stable anchor relay units in each candidate region was examined. A candidate region was considered a conservative spatial microdomain only if it contained the same stable anchor relay unit in at least five healthy Wuzhishan pigs, and the centroid offset between corresponding spatial microdomain units did not exceed one-third of the shorter side length of the smallest circumscribed rectangle of the corresponding spatial microdomain unit with the smaller area. This step involves the calculation of spatial offset, and this embodiment uses the following standard mathematical expression:
[0069]
[0070] in, This represents the centroid offset distance between two corresponding spatial micro-domain units. Represents the centroid coordinates of the first spatial micro-domain unit. This represents the centroid coordinates of the second spatial micro-domain unit.
[0071] For example, in a candidate region of the liver, the centroid coordinates of the first spatial micro-domain unit are... The centroid coordinates of the second spatial micro-domain unit are The centroid offset distance is calculated as follows:
[0072]
[0073]
[0074]
[0075]
[0076] If the length of the short side of the minimum circumscribed rectangle of the smaller of the two corresponding spatial micro-domain units is... The allowable offset threshold is:
[0077]
[0078] because Therefore, this candidate region satisfies the centroid offset constraint and can continue to participate in the identification of conservative spatial micro-domains.
[0079] After the conserved spatial microdomains are determined, multiple stable anchor relay units sharing the same anchor transcript model or the same anchor coding sequence within each conserved spatial microdomain are aggregated to form pathway core modules, which are then used as the main annotation units for that conserved spatial microdomain. Subsequently, pathway core modules in at least two different target tissues of the same healthy Wuzhishan pig are compared. When the conserved spatial microdomains in different target tissues contain the same pathway core module, and the stable anchor relay units constituting these pathway core modules have a consistent upstream-to-downstream arrangement order, cross-tissue homologous functional axes are established between the corresponding target tissues.
[0080] Furthermore, for adjacent consecutive slices of the same tissue block sample, if the spatial micro-domain units after cross-slice spatial registration contain the same pathway core module, and the projected overlap area accounts for more than 60% of the area of the smaller spatial micro-domain unit, then the corresponding spatial micro-domain units in the adjacent consecutive slices are merged into a three-dimensional conservative micro-pillar. The projected overlap area ratio is calculated using the following formula:
[0081]
[0082] in, Indicates the proportion of the projected overlapping area. This represents the projected overlap area between two spatial micro-domains that have undergone cross-slice spatial registration. This represents the area of the smaller of the two spatial micro-domain units.
[0083] For example, in two adjacent consecutive slices of a jejunum sample, the projected overlap area of two spatial micro-domains after cross-slice spatial registration is: The area of the smaller one is The ratio of the projected overlapping area is calculated as follows:
[0084]
[0085] Since 65.0% is higher than the merging threshold of 60%, these two spatial micro-domain units can be merged into a three-dimensional conservative micro-pillar. Figure 5 The construction relationship of three-dimensional conservative micropillars and cross-tissue homologous functional axes is shown.
[0086] After forming three-dimensional conserved micropillars and cross-tissue homologous functional axes, a multi-omics spatial map of healthy Wuzhishan pigs is constructed using the three-dimensional conserved micropillars as intra-tissue nodes and the cross-tissue homologous functional axes as cross-tissue connection edges. Information on the source of Wuzhishan pig breed-specific molecular anchor sets and the upstream-to-downstream arrangement of stable anchor relay units constituting the core pathway modules are recorded in each intra-tissue node. To verify the incremental update capability of the map, this embodiment further introduces three new healthy Wuzhishan pig samples, and the same processing procedure is repeated for the liver, spleen, and jejunum. Before the new samples are incorporated into the map, consistency verification with the Wuzhishan pig breed-specific molecular anchor set is performed, followed by anchor point correction, anchor relay triplet identification, stable anchor relay unit determination, and core pathway module identification. Only when the core pathway module formed by the new sample appears repeatedly in at least two other new healthy Wuzhishan pigs is the conserved spatial microdomain corresponding to that core pathway module and the resulting three-dimensional conserved micropillars incorporated into the existing map.
[0087] Table 4 presents the overall results of map construction and incremental updates in this embodiment.
[0088] Table 4: Map Construction and Incremental Update Results
[0089] project liver spleen jejunum total Number of initial conservative space micro-fields 18 15 21 54 Number of core modules in the initial path 24 20 27 71 Number of cross-organizational homologous functional axes 6 5 7 18 Initial number of three-dimensional conservative micropillars 31 26 35 92 Number of new sample candidate pathway core modules 9 7 11 27 Number of core pathway modules included after repeated occurrence screening 4 3 5 12 The number of newly added three-dimensional conservative micropillars after inclusion 5 4 6 15
[0090] Single-link example:
[0091] To further clarify the complete processing chain from raw data to final map nodes, this embodiment uses a representative region from a pig liver tissue block sample as an example for continuous illustration. The coordinates of this representative region in the histological relay section range from 1.20 mm to 1.96 mm horizontally and from 2.05 mm to 2.84 mm vertically. A total of 1287 original transcriptional features were detected in this region in the original spatial transcriptome sections. Among them, 214 original transcriptional features showed inconsistent localization results after comparison with the pig linear reference genome and the pig pan-genome reference. After relocalizing these 214 original transcriptional features using a graph reference path containing a set of molecular anchors specific to the Wuzhishan pig breed, 163 transcriptional features were stably mapped to the anchor transcript model. These were then merged with other original transcriptional features within the same region, resulting in a final corrected number of 934 transcriptional features for this region.
[0092] In the same region, 462 original peptide features were detected in the spatial proteome sections. After re-matching with anchor coding sequences in the anchor coding sequence library, 121 peptide features that met the conditions of amino acid change segment coverage, splice connection site coverage, or structural variation breakpoint adjacent segment coverage were retained, forming the corresponding set of corrected protein features for this region. In the corresponding region of the spatial metabolome sections, a total of 389 original metabolic features were detected. Among them, 46 metabolic features that have a direct substrate or direct product relationship with the enzyme-catalyzed reaction characterized by the corrected protein features and correspond to the same spatial micro-domain unit after cross-slice spatial registration with the corrected protein features and corrected transcription features were identified, ultimately forming the corresponding set of corrected metabolic features for this region.
[0093] After cross-layer anchor point correction, cross-slice spatial registration and spatial micro-domain unit division were performed on this region, resulting in three spatial micro-domain units. The largest spatial micro-domain unit, with an area of 0.84 square millimeters, is located at the boundary between the hepatic cord and the bile duct. For this spatial micro-domain unit, after biochemical reaction item matching, seven candidate anchor point relay triads were identified. Four of these triads met the criteria of the same reaction item or two adjacent reaction items. Further screening for neighborhood direction consistency retained three stable anchor point relay units. Among these three stable anchor point relay units, two shared the same anchor point transcript model and had the same upstream-to-downstream arrangement order, thus they were aggregated into the same pathway core module. This pathway core module was repeated in five out of six healthy Wuzhishan pig livers and met the centroid shift threshold, therefore this region was identified as a conserved spatial micro-domain. Furthermore, in consecutive adjacent slices adjacent to this region, the projected overlap area of spatial micro-domain units with the same pathway core module was 68.4%, therefore they were merged into three-dimensional conserved micropillars. Finally, a cross-tissue homologous functional axis is established between this three-dimensional conserved micropillar and a conserved spatial microdomain in the spleen containing the same pathway core module. This axis is recorded as an intra-tissue node and cross-tissue functional connection unit in the multi-omics spatial map of healthy Wuzhishan pigs. Through this single-link example, the complete processing process of this embodiment, from raw data, anchor point correction, spatial registration, anchor point relay triad identification, stable anchor point relay unit screening, conserved spatial microdomain extraction, pathway core module aggregation, to the formation of the three-dimensional conserved micropillar and the cross-tissue homologous functional axis, can be clearly seen.
[0094] This embodiment explains the basis for setting the key thresholds. Regarding the 70% repetition rate used in screening the "Wuzhishan pig breed-specific molecular anchor set," this embodiment tested five thresholds—50%, 60%, 70%, 80%, and 90%—in preliminary experiments. The results showed that when the threshold was below 70%, although the number of candidate anchors increased significantly, the mismatch rate in subsequent anchor transcript model reconstruction increased rapidly. When the threshold was above 80%, although the reliability of candidate anchors further improved, the number of effective anchors that could be used for shared analysis across different target tissues decreased significantly. Considering both the number and stability of anchors, this embodiment ultimately selected 70% as the basic screening threshold.
[0095] Regarding the condition of "at least 5 healthy Wuzhishan pigs appearing repeatedly", this embodiment conducted a pre-comparison in a sample of 6 healthy Wuzhishan pigs. It was found that when the number of repeated individuals was less than 5, the recurrence rate of the conservative spatial micro-domain fluctuated greatly, while when it reached 5 or more, the recurrence rate of the conservative spatial micro-domain increased significantly and the fluctuation decreased. Therefore, at least 5 individuals were selected as the recurrence threshold for conservative spatial micro-domain extraction.
[0096] Regarding the condition that "the centroid offset does not exceed one-third of the short side length of the smallest circumscribed rectangle of the smaller area," this embodiment compares three thresholds: one-quarter, one-third, and one-half. The results show that using one-quarter further tightens the spatial offset, but significantly reduces the number of candidate regions that can form conservative spatial micro-domains; using one-half increases the number of candidate regions, but leads to region boundary expansion and functional module mixing. After comprehensive comparison, the one-third threshold achieves a better balance between reproducibility, boundary fit, and functional purity.
[0097] Regarding the condition that "the projected overlapping area occupies more than 60% of the area of the smaller spatial micro-domain unit," this embodiment compared three thresholds: 50%, 60%, and 70%. The results show that at the 50% threshold, the number of three-dimensional conservative micropillars increases, but the reproducibility rate of three-dimensional results in repeated experiments decreases; at the 70% threshold, the reproducibility rate improves, but the number of three-dimensional conservative micropillars is significantly insufficient. Considering both the three-dimensional result reproducibility rate and the number of retained micropillars, this embodiment ultimately adopts 60% as the merging threshold.
[0098] The condition that "the core pathway module formed by the newly added sample must be repeated in at least two other newly added healthy Wuzhishan pigs" is primarily intended to prevent the erroneous inclusion of functional modules in a single newly added sample due to accidental sampling disturbances, detection noise, or individual differences. Preliminary experiments show that if repetition in only one newly added healthy Wuzhishan pig is required, the probability of isolated micropillars and unstable cross-tissue connections appearing in the incremental map is relatively high; when repetition in at least two other newly added healthy Wuzhishan pigs is required, the stability of the incremental nodes is significantly improved. Therefore, this embodiment uses this threshold for incremental map update control.
[0099] To make the above parameter settings more intuitive, Table 5 presents the comparison results of the preliminary experiments under different parameter thresholds.
[0100] Table 5 Comparison Results of Preliminary Experiments for Key Parameters
[0101] Parameter Items Alternative thresholds Three-layer spatial correlation accuracy Conservative spatial micro-domain reproducibility 3D result reproduction rate Remark Anchor point appearance ratio threshold 50% 84.1% 69.4% 63.7% High number of candidates, high mismatch rate Anchor point appearance ratio threshold 60% 88.3% 76.2% 70.1% Stability has improved Anchor point appearance ratio threshold 70% 92.6% 86.9% 81.5% Best overall effect Anchor point appearance ratio threshold 80% 93.4% 87.8% 82.0% The number of anchor points has been significantly reduced. Anchor point appearance ratio threshold 90% 93.8% 88.1% 82.3% Insufficient anchor point coverage Spatial offset threshold One-quarter of the length of the shorter side 93.0% 82.1% 80.8% Too few areas are preserved Spatial offset threshold One-third of the length of the shorter side 92.6% 86.9% 81.5% Best overall effect Spatial offset threshold Half the length of the shorter side 89.7% 84.0% 76.2% Regional mixing increases Overlap area threshold 50% 92.4% 86.5% 73.4% Merging the width Overlap area threshold 60% 92.6% 86.9% 81.5% Best overall effect Overlap area threshold 70% 92.8% 87.2% 83.1% The number of three-dimensional retaining columns is too small.
[0102] Example 2:
[0103] To further verify the applicability of the technical solution of this invention under different tissue combinations and parameter conditions, this embodiment selected eight healthy Wuzhishan pigs as sample sources, aged 6.0 to 7.2 months and weighing 20.1 to 24.6 kg. All healthy Wuzhishan pigs were confirmed healthy after clinical examination, blood biochemistry tests, and porcine pathogen nucleic acid testing. Unlike Example 1, this embodiment selected liver, lung tissue, and the longissimus dorsi muscle as target tissues, and processed tissue block samples from each target tissue. The average depth of whole-genome resequencing was 33.8 times, and the full-length transcript sequencing coverage was higher than in Example 1. The construction principle of the Wuzhishan pig breed-specific molecular anchor set was the same as in Example 1, but a 75% repeatability threshold was used in the anchor occurrence ratio screening, and in the repeatability screening of new samples, repeatability was required in at least three additional healthy Wuzhishan pigs before inclusion in the atlas.
[0104] For section setup, the thickness of histological relay sections was 5 micrometers, spatial transcriptome sections were 9 micrometers, spatial proteome sections were 8 micrometers, and spatial metabolome sections were 10 micrometers. The spacing between adjacent consecutive sections was controlled within 12 micrometers. For cross-section spatial registration, microvascular boundaries, bile duct boundaries, and connective tissue boundaries were used in liver tissue samples; microvascular boundaries, bronchiolar boundaries, and connective tissue boundaries were used in lung tissue samples; and microvascular boundaries, connective tissue boundaries, and muscle bundle boundaries were used in longissimus dorsi muscle tissue samples. The remaining processing steps, including anchor point correction, spatial microdomain unit establishment, anchor point relay triad identification, stable anchor point relay unit screening, conserved spatial microdomain extraction, pathway core module aggregation, cross-tissue homologous functional axis establishment, three-dimensional conserved micropillar formation, and atlas incremental update, were the same as in Example 1.
[0105] The results of Example 2 show that, even with different combinations of target tissues and slightly different threshold settings, it is still possible to stably form conservative spatial microdomains, pathway core modules, cross-tissue homologous functional axes, and three-dimensional conservative micropillars. Table 6 provides a statistical overview of the results of Example 2.
[0106] Table 6, Summary of Results Statistics for Example 2
[0107] project liver lung tissue Longissimus dorsi muscles total Number of initial conservative space micro-fields 20 17 19 56 Number of core modules in the initial path 26 22 24 72 Number of cross-organizational homologous functional axes 7 6 6 19 Initial number of three-dimensional conservative micropillars 34 29 31 94 Number of new sample candidate pathway core modules 11 8 10 29 Number of core pathway modules included after repeated occurrence screening 5 4 4 13 The number of newly added three-dimensional conservative micropillars after inclusion 6 5 5 16
[0108] Example 2 illustrates that the technical solution of the present invention is not only applicable to the target tissue group of liver, spleen and jejunum. Under the premise of keeping the core steps such as Wuzhishan pig breed-specific molecular anchor point set, cross-layer anchor point correction, cross-slice spatial registration, stable anchor point relay unit screening and conservative spatial micro-domain extraction unchanged, it can still be extended to other target tissue combinations.
[0109] Comparative Example 1:
[0110] Comparative Example 1 employs the continuous slice spatial integration approach described in the background technique, which uses histological images as universal anchors. It only achieves cross-slice spatial registration and multi-omics integration between different consecutive adjacent slices based on histological relay slices. It does not construct a Wuzhishan pig breed-specific molecular anchor set, reconstruct anchor transcript models, establish anchor coding sequence libraries, or impose retention conditions on metabolic features that require "direct substrates or direct products that are co-located with transcriptional and protein features in the same spatial microdomain unit." Furthermore, it does not introduce screening logic for stable anchor relay units, conserved spatial microdomains, pathway core modules, cross-tissue homology functional axes, or three-dimensional conserved micropillars. Comparative Example 1 and this embodiment use the same source of 6 healthy Wuzhishan pigs, the same target tissues, the same slice order, and the same detection platform.
[0111] The evaluation metrics for this implementation method and Comparative Example 1 include: three-layer spatial association accuracy, proportion of erroneously retained metabolic features, cross-individual conserved spatial microdomain reproducibility rate, spatial microdomain boundary consistency coefficient, and three-dimensional result reproducibility rate. The three-layer spatial association accuracy is calculated using the following formula:
[0112]
[0113] in, Indicates the accuracy of the three-layer spatial association. This indicates the number of three-level related entries that, after manual review and cross-verification with the database, are indeed the same reaction entry or adjacent reaction entries. This indicates the total number of three-level related entries that are ultimately retained.
[0114] For example, in the liver sample statistics of this embodiment, the total number of three-level association entries retained was 500, of which 463 entries were determined to be correctly associated after review. Therefore, the accuracy rate of the three-level spatial association was calculated as follows:
[0115]
[0116] In the statistical analysis of the liver samples in Comparative Example 1, a total of 500 three-level association entries were ultimately retained, of which 367 were correctly associated. Therefore, the accuracy rate of the three-level spatial association is:
[0117]
[0118] Table 7 summarizes the key effects of this embodiment and Comparative Example 1.
[0119] Table 7 compares the effects of this embodiment with Comparative Example 1.
[0120] Comparison items This implementation method Comparative Example 1 Three-layer spatial correlation accuracy 92.6% 73.4% Proportion of metabolic features mistakenly retained 8.3% 27.8% Cross-individual conservative spatial micro-domain recurrence rate 86.9% 58.7% Spatial micro-domain boundary coincidence coefficient 0.84 0.63 3D result reproduction rate 81.5% 55.2%
[0121] Comparative Example 2:
[0122] Comparative Example 2 retains the steps of constructing the Wuzhishan pig breed-specific molecular anchor set, reconstructing the anchor transcript model, establishing the anchor coding sequence library, and cross-slice spatial registration, based on Example 1. However, in the metabolic feature retention rules, it no longer requires that the metabolic feature must be the direct substrate or direct product of the enzymatic reaction characterized by the corrector protein feature. Instead, it only requires that the metabolic feature, the corrector protein feature, and the corrector transcription feature correspond to the same spatial micro-domain unit after cross-slice spatial registration. In other words, Comparative Example 2 removes the key constraint of "direct substrate or direct product" to verify the contribution of this technical feature to the overall effect.
[0123] The results showed that the number of metabolic features entering the final three-layer association analysis in Comparative Example 2 increased significantly, but the proportion of falsely retained metabolic features increased sharply, and the accuracy of the three-layer spatial association decreased. This indicates that the condition of "direct substrate or direct product" has a significant effect on controlling the retention of false positives in the spatial metabolic layer.
[0124] Comparative Example 3:
[0125] Comparative Example 3 retains the steps of constructing the Wuzhishan pig breed-specific molecular anchor set, cross-layer anchor correction, and cross-slice spatial registration based on Example 1, and retains the restriction of "direct substrate or direct product." However, in the screening of stable anchor relay units, the constraint that "the current spatial microdomain unit and its first-ring adjacent spatial microdomain units maintain the same upstream-to-downstream enrichment direction, and there is no opposite enrichment direction in the reverse-adjacent spatial microdomain units" is removed. Instead, the existence of an anchor relay triplet in the current spatial microdomain unit is used as the criterion for determining a stable anchor relay unit. This comparative example is used to verify the contribution of the "directional consistency and reverse exclusion" screening logic to the stability of cross-individual conservative spatial microdomains.
[0126] The results show that Comparative Example 3 can obtain more candidate stable anchor relay units in the early stage, but after repeated comparisons among different healthy Wuzhishan pigs, the reproduction rate of its conservative spatial micro-domain and the reproduction rate of the three-dimensional results both decreased significantly. This indicates that screening based solely on local colocalization will introduce more functional chains with unstable spatial propagation directions, thereby affecting the stable formation of conservative spatial micro-domains and three-dimensional conservative micro-pillars.
[0127] Table 8 presents the comparison results between Example 1, Comparative Example 1, Comparative Example 2, and Comparative Example 3.
[0128] Table 8 compares the effects of Example 1 with multiple comparative examples.
[0129] Comparison items Example 1 Comparative Example 1 Comparative Example 2 Comparative Example 3 Three-layer spatial correlation accuracy 92.6% 73.4% 81.2% 85.1% Proportion of metabolic features mistakenly retained 8.3% 27.8% 19.6% 11.7% Cross-individual conservative spatial micro-domain recurrence rate 86.9% 58.7% 72.4% 69.3% Spatial micro-domain boundary coincidence coefficient 0.84 0.63 0.79 0.82 3D result reproduction rate 81.5% 55.2% 67.8% 63.9%
[0130] Figure 6 Furthermore, a comparison curve of the accuracy of the three-layer spatial association between this embodiment and Comparative Example 1 on six healthy Wuzhishan pig individuals is provided; from Figure 6 It can be seen that the accuracy of the three-layer spatial association of this embodiment on six healthy Wuzhishan pig individuals is consistently higher than that of Comparative Example 1, indicating that this embodiment does not only achieve improvement on individual samples, but also has good repeatability on different individuals. Figure 7 A bar chart comparing the key performance indicators of this embodiment and Comparative Example 1 is provided, which can more intuitively reflect that this embodiment is superior to Comparative Example 1 in terms of accuracy of three-layer spatial association, cross-individual conservative spatial micro-domain reproducibility, spatial micro-domain boundary matching coefficient, and three-dimensional result reproducibility. At the same time, it is significantly lower than Comparative Example 1 in terms of the proportion of erroneously retained metabolic features.
[0131] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A method for constructing a multi-omics spatial map of healthy Wuzhishan pigs based on a molecular bioinformatics analysis and processing system, characterized in that, Includes the following steps: S1. Obtain whole genome variation data from at least 3 healthy Wuzhishan pigs, and construct a Wuzhishan pig breed-specific molecular anchor set by combining pig pangenome reference and Wuzhishan pig full-length transcript data. The Wuzhishan pig breed-specific molecular anchor set includes at least breed-conserved single nucleotide variant sites, breed-conserved structural variant sites, variant-supported splicing sites, and variant-supported coding peptide sites. S2. Obtain tissue block samples of at least one target tissue from each healthy Wuzhishan pig, and prepare consecutive adjacent slices from the same tissue block for each tissue block sample. The consecutive adjacent slices include at least histological relay slices, spatial transcriptome slices, spatial proteome slices and spatial metabolome slices. S3. The molecular bioinformatics analysis and processing system performs anchor point correction on the raw data generated by the spatial transcriptome slices, spatial proteome slices and spatial metabolome slices based on the Wuzhishan pig breed-specific molecular anchor point set, to obtain the corrected transcription feature matrix, the corrected protein feature matrix and the corrected metabolic feature matrix. S4. Using the histological relay slice as a displacement transfer layer, perform cross-slice spatial registration on the spatial transcriptome slice, the spatial proteome slice, and the spatial metabolome slice, and project the registration results uniformly to the spatial micro-domain unit. S5. For each of the aforementioned spatial micro-domain units, only when the corrective transcription feature carrying the same molecular anchor point, the corrective protein feature encoded by the corrective transcription feature and covering the same molecular anchor point, and the corrective metabolic feature directly related to the enzyme-catalyzed reaction corresponding to the corrective protein feature are simultaneously detected, the corrective transcription feature, the corrective protein feature, and the corrective metabolic feature are identified as the anchor point relay triad. S6. Among different healthy Wuzhishan pigs in the same target tissue, the anchor point relay triplet is spatially registered with the homologous tissue region, and conservative spatial micro-domains are extracted based on repetition, spatial continuity and reaction sequence consistency. S7. Construct a multi-omics spatial map of healthy Wuzhishan pigs based on the conserved spatial microdomains of each target organization.
2. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 1, characterized in that, The construction of the Wuzhishan pig breed-specific molecular anchor set includes: mapping the whole genome variation data of each healthy Wuzhishan pig to the pig linear reference genome and the pig pan-genome reference, respectively; screening out exon variation sites, splice boundary variation sites, insertion / deletion variation sites, and structural variation breakpoints among the breed-conserved structural variation sites that are repeated in no less than 70% of healthy Wuzhishan pigs and are simultaneously supported by the full-length transcript data of Wuzhishan pigs, as the Wuzhishan pig breed-specific molecular anchor set.
3. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 2, characterized in that, The Wuzhishan pig breed-specific molecular anchor set also satisfies the following conditions: at least a portion of the breed-conserved single nucleotide variant sites, the variant-supported splicing sites, and the structural variant breakpoints are located in the coding region, splicing boundary region, or promoter region, and correspond to transcript structural change information that is repeatedly observed among different healthy Wuzhishan pigs.
4. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 1, characterized in that, The transcription feature correction in the anchor point correction includes: for transcription features whose alignment and localization results are inconsistent with those given by the pig linear reference genome and the pig pan-genome reference, a graph reference path containing the Wuzhishan pig breed-specific molecular anchor set is preferentially used for relocalization, and the anchor transcript model is reconstructed based on the Wuzhishan pig full-length transcript data to obtain the corrected transcription feature matrix.
5. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 4, characterized in that, The protein feature correction in the anchor point correction includes: constructing an anchor point coding sequence library based on the anchor point transcript model, and re-matching the peptide features generated by the spatial proteome slices to the anchor point coding sequences in the anchor point coding sequence library, retaining only the peptide features that cross the amino acid change segment corresponding to the variety's conserved single nucleotide variant site, cross the splice connection site supported by the variant, or cover the adjacent segment of the structural variant breakpoint in the variety's conserved structural variant site, as the protein features in the corrected protein feature matrix.
6. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 5, characterized in that, The metabolic feature correction in the anchor point correction includes: retaining a metabolic feature as a metabolic feature in the corrected metabolic feature matrix only when a certain metabolic feature corresponds to the direct substrate or direct product of the enzyme-catalyzed reaction characterized by the corrected protein feature, and the metabolic feature, the corrected protein feature, and the corrected transcription feature are located in the same spatial micro-domain unit.
7. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 1, characterized in that, The cross-slice spatial registration includes: using the histological relay slice as a relay reference, selecting at least three common anatomical landmarks from the microvascular boundary, glandular duct boundary, connective tissue boundary, and muscle bundle boundary, constructing a segmented displacement field from the histological relay slice to the spatial transcriptome slice, the spatial proteome slice, and the spatial metabolome slice, and projecting the data in the spatial transcriptome slice, the spatial proteome slice, and the spatial metabolome slice to the spatial microdomain unit based on the segmented displacement field.
8. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 7, characterized in that, The boundary of the spatial micro-domain unit is established only when the following conditions are met simultaneously: there is a morphological boundary in the histological relay section, and at least two of the corrected transcription feature matrix, the corrected protein feature matrix, and the corrected metabolic feature matrix have omics boundaries at corresponding positions.
9. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 1, characterized in that, The anchor relay triad is established only under the following conditions: the corrected transcriptional feature, the corrected protein feature, and the corrected metabolic feature correspond to the same reaction entry, or to two adjacent reaction entries, and the corrected metabolic feature is the substrate or product of the enzymatic reaction characterized by the corrected protein feature.
10. The method for constructing a multi-omics spatial map of healthy Wuzhishan pigs according to claim 9, characterized in that, When the same anchor point relay triplet appears simultaneously in the current spatial micro-domain unit and its first ring of adjacent spatial micro-domain units, and maintains the same upstream to downstream enrichment direction in the current spatial micro-domain unit and its first ring of adjacent spatial micro-domain units, while there is no opposite enrichment direction in the reverse adjacent spatial micro-domain units relative to the upstream to downstream enrichment direction, the anchor point relay triplet is determined as a stable anchor point relay unit.