A method and system for matching spatial transcriptome and spatial metabolome spatial information
By selecting identification points in the spatial transcriptome and performing HE staining, combined with registration and coordinate transformation, the problem of spatial distribution differences between the spatial transcriptome and the spatial metabolome was solved, point-to-point matching and data association were achieved, and the accuracy of the analysis was improved.
Patent Information
- Application Number
- CN202211282284.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Due to the different detection platforms and experimental operations of the spatial transcriptome and the spatial metabolome, their spatial distributions vary greatly, making it difficult to align spatial positions and match point-to-point data, and thus difficult to perform association analysis.
By selecting identification points within the detection range of spatial transcriptome samples, obtaining barcode information and performing HE staining identification, combining MSIReader software for registration, calculating a unified coordinate system, and performing coordinate transformation through the final rotation angle and scaling ratio, spatial information matching between the spatial transcriptome and the spatial metabolome is achieved.
Point-to-point matching between spatial transcriptome and spatial metabolome was achieved, which improved the accuracy and reliability of data association and provided more reliable data support for subsequent association analysis.
Smart Images

Figure CN115798587B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image data processing, and relates to a method and system for matching spatial information of a spatial transcriptome and a spatial metabolome. Background Art
[0002] Spatial transcriptome sequencing is the result of the in-depth development of single-cell transcriptome technology. It can simultaneously obtain the spatial location information and gene expression data of cells without the need to prepare cell suspensions. It further promotes the study of the true gene expression of cells in situ in tissues and provides an important research tool for many fields such as tissue cell function, microenvironment interaction, developmental lineage tracing, and disease pathology.
[0003] Spatial transcriptomics, an emerging cutting-edge technology developed closely on the heels of high-throughput single-cell transcriptomics, preserves the spatial location of tissues while simultaneously analyzing the transcriptome of tissue sections. Spatial transcriptomics is a technique used to spatially analyze RNA-seq data, thereby analyzing all mRNAs within a single tissue section. Spatial barcoded reverse transcription primers are systematically attached to the chip surface, enabling the encoding and acquisition of spatial location information during sample processing and subsequent sequencing. When a tissue cryosection is attached to the spatial transcriptomics chip, the barcoded primers bind and capture adjacent mRNAs from the tissue. The captured mRNAs undergo reverse transcription, and the resulting cDNA contains the spatial barcode. Analysis of the sequence of the spatial barcodes in the sequencing results allows the sequence of each mRNA transcript to be mapped back to its starting position within the tissue section.
[0004] Spatial metabolomics, a type of mass spectrometry imaging technology, can directly obtain information on the structure, content, and spatial distribution of a large number of known and unknown endogenous metabolites and exogenous drugs from biological tissues. It does not require chemical or radioactive labeling or complex sample pretreatment, and offers the advantages of high specificity, high throughput, and preservation of spatial information.
[0005] Spatial metabolomics combines mass spectrometry imaging with metabolomics, leveraging the ability of mass spectrometry imaging to accurately identify and locate the differential distribution of multiple metabolites within tissues and even cells. Combined with metabolomics, it enables in-depth metabolomics analysis of target microregions and the ability to determine metabolite species and concentrations, enabling the detection of the spatial distribution of metabolites in biological samples. Spatial metabolomics allows for direct, label-free scanning and analysis of biological tissue sections, locating the spatial location of the scanned points using the scanned information.
[0006] The transcriptome is an important method for analyzing gene expression within an organism, while the metabolome is the foundation and direct embodiment of an organism's phenotype. Metabolites are the final result of gene transcription under internal and external regulation and are the material basis of an organism's phenotype. In the era of systems biology research, biological developmental processes are complex and varied, and gene regulatory networks are intricate. Using a single omics approach to systems biology is often one-sided. Correlation analysis using spatial transcriptomics and metabolomics can analyze internal changes in organisms at both the cause and effect levels, identify key pathways associated with metabolite changes, construct core regulatory networks, and comprehensively explore biological growth, development, and stress mechanisms, providing a holistic understanding of biological problems.
[0007] However, due to the different spatial information identifiers of the spatial transcriptome and the spatial metabolome, different experimental operation steps of the slice samples, and their non-identical slice samples and different spatial distributions, it is difficult to align the spatial positions of the data of the two different systems and match the point-to-point data, and it is even more difficult to conduct correlation analysis between the spatial transcriptome and the spatial metabolome.
[0008] Specifically, due to the differences in the detection platforms and experimental operations of the spatial transcriptome and the spatial metabolome, the spatial distributions of the two sets are different. The spatial transcriptome uses a chip as the detection platform, with a chip detection area of 6.5mm×6.5mm, and a spot as the smallest detection unit. The diameter of a single spot is 55μm, and the distance between the spot centers is 100μm. The upper and lower layers are staggered. The spatial metabolome uses mass spectrometry as the detection platform. Its detection area is not restricted and can reach a maximum of 10cm×10cm. The pixel is used as the smallest detection unit. The common size of a single pixel is 100μm×100μm, and the pixels are distributed in a square combination. The different spatial distributions of the spatial transcriptome and the spatial metabolome make it difficult to associate spots with pixels, and even more difficult to associate spots with the underlying data within the pixels. Summary of the Invention
[0009] In order to address the deficiencies in the prior art, the present invention aims to provide a method and system for spatial information matching between spatial transcriptome and spatial metabolome.
[0010] The present invention proposes a spatial information matching method for a spatial transcriptome and a spatial metabolome. The spatial information matching method unifies the spatial information identifiers of the spatial transcriptome and the spatial metabolome, and converts the spatial coordinate system of the spatial metabolome into the spatial coordinates of the spatial transcriptome, thereby achieving spatial information matching between the spatial transcriptome and the spatial metabolome. The method comprises the following steps:
[0011] Step (1) selecting 6 or more spots evenly distributed within the sample detection range of the spatial transcriptome (selecting and distributing them around and in the middle of the entire sample detection range so that the selected spots are extended to the entire sample detection range) as identification spots, obtaining the barcode information corresponding to the identification spots and marking the selected spots on the corresponding HE staining map;
[0012] HE staining is one of the most basic and widely used technical methods in histology and pathology teaching and scientific research. The pathological characteristics presented by HE staining are associated with the spatial transcriptome.
[0013] There are a total of 4 spatial transcriptome chips on the expression slide of the spatial transcriptome; the detection area size of each spatial transcriptome chip is 6.5mm×6.5mm, with a total of 4992 spots, the diameter of a single spot is 55μm, and the distance between each two adjacent spots and the center of the spot is 100μm; the spot is the smallest unit detected in the spatial transcriptome.
[0014] The barcode is a spatial barcode in the Visium spatial transcriptome technology. Each spot has its own unique barcode code, and the spatial position information of the corresponding spot can be obtained through the spatial barcode.
[0015] The sample after HE staining can still be used for spatial transcriptome detection without affecting the spatial transcriptome data. The HE staining image can match the spatial position of the spatial transcriptome data, and spot identification can be performed on the HE staining image.
[0016] Step (2), using the HE staining image of the marked points in step 1 to align with the spatial metabolomics image, and obtaining the spatial coordinate position information of the corresponding pixel points of the marked points on the spatial metabolomics after alignment;
[0017] The pixel point is the smallest unit detected in the spatial metabolome; the resolution of the smallest unit detected in the spatial metabolome varies according to the mass spectrometry scanning rate and the moving platform rate, and the higher the rate, the greater the scanning resolution.
[0018] The resolution of the smallest unit detected in the spatial metabolome represents the distance between pixels; different resolutions result in different distances between pixels; and the conversion of spatial information in the spatial metabolome is affected by the resolution.
[0019] The registration operation refers to rotating, scaling, and flipping the HE staining image based on the spatial metabolomics imaging map through the MSIReader software, matching the spatial metabolomics imaging map and the HE staining map, and obtaining the spatial coordinate information of the identification site on the HE staining map on the spatial metabolomics imaging map through the software.
[0020] Step (3), converting the barcode spatial information of the spatial transcriptome and the spatial metabolome spatial information into a unified spatial information identifier;
[0021] The conversion into a unified spatial information identifier refers to performing operations on the two sets of spatial numbers x and y in the spatial information corresponding to the barcode barcode in the spatial transcriptome and the two sets of spatial numbers x and y in the spatial metabolome spatial information, respectively, to obtain a unified spatial coordinate system using a unified standard (i.e., converting the coordinates of the spatial transcriptome and the spatial metabolome into a micron coordinate system).
[0022] The conversion formula of x space number in the spatial transcriptome is as follows:
[0023] UTX=(TransX / 2+0.5)*transresolution;
[0024] The conversion formula of y space number in spatial transcriptome is as follows:
[0025] UTY=TransY*sqrt(0.75)*transresolution;
[0026] Among them, UTX and UTY are the spatial coordinates of the unified standard of spatial transcriptome data; TransX and TransY are the two sets of spatial numbers of x and y in the spatial transcriptome; transresolution is the numerical value of the spatial transcriptome resolution; sqrt is the square root operation.
[0027] The conversion formula of the x space number in the spatial metabolome is as follows:
[0028] UMX=MetaX*metaresolution;
[0029] The conversion formula of y space number in the spatial metabolome is as follows:
[0030] UMY=MetaY*metaresolution;
[0031] Among them, UMX and UMY are the spatial coordinates after the spatial metabolome data are standardized; MetaX and MetaY are the two sets of spatial numbers of x and y in the spatial metabolome, respectively; metaresolution is the numerical value of the spatial metabolome resolution.
[0032] Step (4), calculating the distances between different identification points in the spatial transcriptome; calculating the distances between different identification points in the spatial metabolome; and calculating the scaling ratio of the distances between the identification points corresponding to the spatial transcriptome and the spatial metabolome;
[0033] In order to evaluate the accuracy of the corresponding marker points in the spatial transcriptome and the spatial metabolome and to correct the position differences of the marker points caused by the error of the experimental instrument, the scaling ratio is calculated;
[0034] The calculation formula of the scaling ratio is as follows:
[0035] Ratio=sqrt((UTXa-UTXb) 2 +(UTYa-UTYb) 2 ) / sqrt((UMXa-UMXb) 2 +(UMYa-UMYb) 2 )
[0036] Among them, UTXa, UTYa, UTXb, and UTYb are the X and Y coordinate values of landmark points a and b in the spatial transcriptome after unified spatial coordinates; UMXa, UMYa, UMXb, and UMYb are the X and Y coordinate values of landmark points a and b in the spatial metabolome after unified spatial coordinates; Ratio is the scaling ratio of the distance between landmark points a and b in the spatial transcriptome to the distance between landmark points a and b in the spatial metabolome.
[0037] After obtaining the scaling ratios between each pair of landmark points, the scaling ratio data needs to be screened. The screening operation includes: eliminating data with scaling ratios greater than 1.1 and less than 0.91, while ensuring that the number of remaining scaling ratio data is ≥4, and averaging the remaining scaling ratio data to obtain the final scaling ratio for spatial metabolome conversion; if the number of remaining scaling ratio data is <4, return to step (1) and repeat.
[0038] Step (5), judging whether the spatial transcriptome and the spatial metabolome samples are flipped during the detection process by the rotation angle of the lines connecting each pair of identification points in the spatial transcriptome and the spatial metabolome, and calculating the final rotation angle between the spatial transcriptome and the spatial metabolome detection samples;
[0039] The marker points are the marker points after the scaling ratio screening; if the deviation of all the rotation angles between the marker points is within 5°, it means that the sample has not been flipped; if the deviation of all the rotation angles between the marker points is outside 5°, the conversion formula of the y space number in the spatial metabolome in step (3) is modified to:
[0040] UMY=-MetaY*resolution,
[0041] Then, the rotation angles between each pair of landmark points in the spatial transcriptome and spatial metabolome detection samples were calculated. After obtaining all rotation angles with a deviation within 5°, the average of all rotation angles was calculated to obtain the final rotation angle.
[0042] The calculation formula of the rotation angle is as follows:
[0043] Angle=arctan((UMYa-UMYb) / (UMXa-UMXb))-arctan((UTYa-UTYb) / (UTXa-UTXb));
[0044] Among them, UTXa, UTYa, UTXb, and UTYb are the X and Y coordinate values of the landmark points a and b in the spatial transcriptome after unified spatial coordinates; UMXa, UMYa, UMXb, and UMYb are the X and Y coordinate values of the landmark points a and b in the spatial metabolome after unified spatial coordinates; arctan is the calculation of the inverse tangent function; Angle is the rotation angle between the line connecting the landmark points a and b in the spatial transcriptome and the line connecting the landmark points a and b in the spatial metabolome.
[0045] Step (6) uses the conversion space identification method to transform the spatial metabolome coordinates through the three parameters of the final rotation angle, the final scaling ratio, and the coordinates of the center identification point of the spatial transcriptome, so that the transformed spatial metabolome coordinates match the spatial transcriptome coordinates.
[0046] The spatial transcriptome center identification point refers to the center point position of the spatial transcriptome detection chip;
[0047] The conversion space identification method refers to converting the spatial position of each pixel in the spatial metabolome into the corresponding coordinates in the spatial transcriptome through spatial coordinate transformation. The final scaling ratio and final rotation angle obtained in step (4) and step (5) also satisfy the following calculation formula:
[0048] Ratio'=sqrt((UMX'-UTXo) 2 +(UTY'-UTYo) 2 ) / sqrt((UMX-UMXo) 2 +(UMY-UMYo) 2 );
[0049] Angle'=arctan((UMY-UMYo) / (UMX-UMXo))-arctan((UMY'-UTYo) / (UMX'-UTXo));
[0050] Among them, Ratio' is the final scaling ratio; Angle' is the final rotation angle; UTXo and UTYo are the spatial information of the center point position of the spatial transcriptome detection chip after the unified coordinate system; UMXo and UMYo are the spatial information of the center point position of the spatial transcriptome detection chip in the unified coordinate system of the corresponding position of the spatial metabolome; UMX and UMY are the X and Y coordinate values of the spatial metabolome after the unified spatial coordinates; UMX' and UMY' are the X and Y coordinate values after the spatial metabolome matches the spatial transcriptome.
[0051] The obtained final rotation angle, final scaling ratio, and coordinate parameters of the center marker point of the spatial transcriptome can be used to obtain the X and Y coordinate values UMX' and UMY' after the spatial metabolome matches the spatial transcriptome.
[0052] Since the spatial transcriptome and metabolome analysis systems only recognize the corresponding spatial information patterns, in order to make their spatial information mutually recognizable, the unified and related coordinate system is restored and converted into the spatial information system of the spatial transcriptome and metabolome;
[0053] The calculation formula for converting the spatial information of the unified coordinate system into the spatial information system of the spatial transcriptome is as follows:
[0054] UMX'=(TransX' / 2+0.5)*transresolution;
[0055] UMY'=TransY'*sqrt(0.75)*transresolution;
[0056] Among them, UMX' and UMY' are the spatial information after the spatial metabolome is matched with the spatial transcriptome; transresolution is the value of the spatial transcriptome resolution; sqrt is the square root operation; TransX' and TransY' are the two sets of spatial numbers of x and y in the spatial transcriptome after the spatial metabolome is matched with the spatial transcriptome;
[0057] The calculation formula for converting the spatial information of the unified coordinate system into the spatial metabolomics spatial information system is as follows:
[0058] UMX'=MetaX'*metaresolution;
[0059] UMY'=MetaY'*metaresolution;
[0060] Among them, UMX' and UMY' are the spatial information after the spatial metabolome is matched with the spatial transcriptome; metaresolution is the value of the spatial metabolome resolution; MetaX' and MetaY' are the two sets of spatial numbers x and y in the spatial metabolome spatial information after the spatial metabolome is matched with the spatial transcriptome spatial information.
[0061] The present invention also proposes a system for implementing the above-mentioned spatial information matching method between spatial transcriptome and spatial metabolome, and an application of the above-mentioned spatial information matching method in spatial information matching between spatial transcriptome and spatial metabolome.
[0062] The spatial information matching system of the spatial transcriptome and the spatial metabolome includes a spatial information acquisition module, a spatial coordinate unification module, a spatial parameter calculation module, and a spatial information conversion module;
[0063] The spatial information acquisition module aligns the HE staining image containing the spatial transcriptome information with the marked points with the spatial metabolome imaging image to obtain the spatial coordinate position information of the marked points on the spatial metabolome;
[0064] The spatial coordinate unification module converts the spatial information of the spatial transcriptome and the spatial metabolome into a unified spatial information identifier;
[0065] The spatial parameter calculation module calculates the final scaling ratio and final rotation angle between the spatial transcriptome and the spatial metabolome;
[0066] The spatial information conversion module uses the three parameters of the final rotation angle, the final scaling ratio, and the coordinates of the center mark point of the spatial transcriptome to convert the spatial metabolome coordinates using a conversion space mark method, so that the converted spatial metabolome coordinates match the spatial transcriptome coordinates.
[0067] The spatial information matching method of the present invention is combined with a subsequent data association method to obtain a method for point-to-point matching of spatial transcriptomes and spatial metabolomes.
[0068] The data association method fits the spatial metabolome data and integrates the fitted data into the spatial transcriptome spatial distribution pattern, ultimately achieving data association between the spatial metabolome and the spatial transcriptome at the point-to-point level, including the following steps:
[0069] Step I: fitting the spatial metabolome data to improve the pixel resolution of the spatial metabolome through fitting;
[0070] The fitting refers to amplifying the pixels in the spatial metabolome as needed, and amplifying according to different fitting methods depending on whether the original pixels are located in the sample area and / or the non-sample area.
[0071] If there are only two pixels, both located in the sample area, the calculation method for the middle pixel after the two pixels are amplified is as follows:
[0072] The underlying data of the amplified pixel between two pixels is calculated according to the following formula:
[0073] MZ1ab=(MZ1a+MZ1b) / 2;
[0074] MZ2ab=(MZ2a+MZ2b) / 2;
[0075] …
[0076] MZnab=(MZna+MZnb) / 2;
[0077] Among them, pixels a and b are adjacent pixels; MZ1a, MZ2a, …, MZna are the expression intensities of the CRP of MZ1, MZ2, …, MZn corresponding to pixel a; MZ1b, MZ2b, …, MZnb are the expression intensities of the CRP of MZ1, MZ2, …, MZn corresponding to pixel b; MZ1ab, MZ2ab, …, MZnab are the expression intensities of the CRP of MZ1, MZ2, …, MZn corresponding to the amplified pixels between pixels a and b; the amplified pixels between pixels a and b are abbreviated as {ab}.
[0078] If there are only two pixels, and one pixel is in the sample area and the other is in the non-sample area, the calculation method for the middle pixel after the two pixels are amplified is as follows:
[0079] The underlying data of the amplified pixel between two pixels is calculated according to the following formula:
[0080] MZ1ab=0;
[0081] MZ2ab=0;
[0082] …
[0083] MZnab = 0;
[0084] Among them, pixels a and b are adjacent pixels; MZ1ab, MZ2ab, …, MZnab are the expression intensities of the MZ1, MZ2, …, MZn mass-to-nuclear ratios of the pixels amplified between pixels a and b, respectively; the pixels amplified between pixels a and b are abbreviated as 0, indicating pixels in the non-sample area.
[0085] If there are only two pixels, and both pixels are non-sample area pixels, the calculation method for the middle pixel after the two pixels are amplified is as follows:
[0086] The underlying data of the amplified pixel between two pixels is calculated according to the following formula:
[0087] MZ1ab=0;
[0088] MZ2ab=0;
[0089] …
[0090] MZnab = 0;
[0091] Among them, pixels a and b are adjacent pixels; MZ1ab, MZ2ab, …, MZnab are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to the amplified pixels between pixels a and b, respectively; the amplified pixels between pixels a and b are abbreviated as 0, indicating that the amplified pixels are pixels in the non-sample area.
[0092] Step II: Associating the pixels of the spatial metabolome fitting with the spots of the spatial transcriptome;
[0093] The association refers to corresponding the pixel points after spatial metabolome fitting with the spots of the spatial transcriptome; the pixel points are attributed by calculating the distance between the center position of the pixel point after spatial metabolome fitting and the center position of the spot of the spatial transcriptome; if the distance between the center position of the pixel point after spatial metabolome fitting and the center position of a spot of the spatial transcriptome is less than the diameter of the spot of the spatial transcriptome, then the pixel point after spatial metabolome fitting corresponds to the spot in the spatial transcriptome.
[0094] Step III, sum the underlying pixel data corresponding to each spot;
[0095] The sum operation is to integrate the underlying data of multiple pixels belonging to the same spot into the underlying data of one spot, which is calculated by the following formula:
[0096] MZ1spotA=(MZ1a+MZ1b+…+MZ1x) / j;
[0097] MZ2spotA=(MZ2a+MZ2b+…+MZ2x) / j;
[0098] …
[0099] MZnspotA=(MZna+MZnb+…+MZnx) / j;
[0100] Among them, MZ1a, MZ2a, …, MZna are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point a associated with spot A; MZ1b, MZ2b, …, MZnb are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point b associated with spot A; MZ1x, MZ2x, …, MZnx are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point x associated with spot A; j is the number of pixels associated with spot A; MZ1spotA, MZ2spotA, …, MZnspotA are the underlying data after the summation of the spatial metabolomes corresponding to spot A.
[0101] The underlying data refers to the expression intensity data corresponding to mz at each pixel in a spatial metabolome analysis. mz is the ratio of the mass number to the charge number of a charged particle. It is substance-specific data and is usually referred to as mz.
[0102] Step IV: Obtain the spatial metabolome data corresponding to each spot of the spatial transcriptome.
[0103] Spatial metabolome data are processed using a combination of x and y spatial numbers, and the corresponding expression intensity is assigned to each mz in the new spatial coordinate position. The expression intensity is the relative expression value of the spatial metabolite at each pixel corresponding to mz. The value indicates the relative content, with a larger value indicating a higher relative content at the corresponding pixel.
[0104] Spatial transcriptome data is mapped one-to-one with the spatial metabolite coordinates using barcodes. It also includes the gene count value for each transcribed gene at the corresponding coordinate. Gene counts represent the number of transcripts detected for the corresponding gene in the spatial transcriptome. Transcript expression is determined by the count value. Larger count values correspond to higher transcriptome expression.
[0105] The present invention also proposes a system for implementing the above-mentioned method for associating spatial transcriptome and spatial metabolome data, and the application of the above-mentioned data association method in associating spatial transcriptome and spatial metabolome data.
[0106] The spatial transcriptome and spatial metabolome data association system includes a spatial metabolism data fitting module and a spatial transcription and spatial metabolism data association module;
[0107] The spatial metabolic data fitting module fits the spatial metabolome data and associates it with the spatial transcriptome spots;
[0108] The spatial transcription and spatial metabolic data association module performs sum operation on the underlying pixel data corresponding to each spot to obtain the spatial metabolic group data corresponding to each spot of the spatial transcription group.
[0109] The above-mentioned point-to-point matching method of spatial transcriptome and spatial metabolome achieves point-to-point matching of spatial transcriptome and spatial metabolome by matching the spatial information of spatial transcriptome and spatial metabolome and associating the data.
[0110] The steps include:
[0111] Step 1: align the HE staining image containing the spatial transcriptome information with the marked points with the spatial metabolome imaging map to obtain the spatial coordinate position information of the marked points on the spatial metabolome;
[0112] The identification points are all located within the sample detection range of the spatial transcriptome, and the number of the identification points is ≥6; the identification points contain barcode information, which represents the spatial information of the identification points.
[0113] After the HE staining image and the spatial metabolomics imaging map are registered, the HE staining image and the spatial metabolomics imaging map are matched to obtain the spatial coordinate information of the identification point on the spatial metabolomics imaging map.
[0114] Step 2: Convert the spatial information of the spatial transcriptome and the spatial metabolome into a unified spatial information identifier;
[0115] The spatial coordinate systems of the spatial transcriptome and the spatial metabolome are unified by respectively performing operations on the two sets of spatial numbers x and y in the spatial information corresponding to the barcode in the spatial transcriptome and the two sets of spatial numbers x and y in the spatial information of the spatial metabolome and converting them into a micron coordinate system.
[0116] Step 3: Calculate the final scaling ratio and final rotation angle between the spatial transcriptome and the spatial metabolome;
[0117] The scaling ratios between each pair of landmark points were calculated, and the obtained scaling ratio data were screened to eliminate the data that did not meet the requirements. The remaining scaling ratio data were averaged to obtain the final scaling ratio for spatial metabolome conversion.
[0118] The rotation angle between the line connecting any two identification points in the spatial transcriptome and the line connecting two corresponding identification points in the spatial metabolome was calculated; the rotation angles of all the lines connecting any two points were calculated according to the principle of permutation and combination, and the average of all the obtained rotation angles was calculated to obtain the final rotation angle.
[0119] Step 4: Using the final rotation angle, final scaling ratio, and spatial transcriptome center identification point coordinates, the spatial metabolome coordinates are transformed using the transformation spatial identification method so that the transformed spatial metabolome coordinates match the spatial transcriptome coordinates.
[0120] Step 5: Fit the spatial metabolome data and associate them with the spatial transcriptome spots;
[0121] After amplifying the pixels in the spatial metabolome, the pixels are assigned according to the distance between the center position of the pixel after fitting the spatial metabolome and the center position of the spot in the spatial transcriptome.
[0122] Step 6: Add the underlying pixel data corresponding to each spot to obtain the spatial metabolome data corresponding to each spot in the spatial transcriptome.
[0123] The underlying data of multiple pixels belonging to the same spot are summed to obtain the spatial metabolome data corresponding to each spot of the spatial transcriptome.
[0124] The present invention also proposes a system for implementing the above-mentioned point-to-point matching method of spatial transcriptome and spatial metabolome, and the application of the above-mentioned point-to-point matching method in the association analysis of spatial transcriptome and spatial metabolome.
[0125] The spatial transcriptome and spatial metabolome point-to-point matching system includes a spatial information acquisition module, a spatial coordinate unification module, a spatial parameter calculation module, a spatial information conversion module, a spatial metabolic data fitting module, and a spatial transcription and spatial metabolic data association module;
[0126] The spatial information acquisition module aligns the HE staining image containing the spatial transcriptome information with the marked points with the spatial metabolome imaging image to obtain the spatial coordinate position information of the marked points on the spatial metabolome;
[0127] The spatial coordinate unification module converts the spatial information of the spatial transcriptome and the spatial metabolome into a unified spatial information identifier;
[0128] The spatial parameter calculation module calculates the final scaling ratio and final rotation angle between the spatial transcriptome and the spatial metabolome;
[0129] The spatial information conversion module uses the final rotation angle, the final scaling ratio, and the coordinates of the center mark point of the spatial transcriptome to convert the spatial metabolome coordinates using a conversion space mark method, so that the converted spatial metabolome coordinates match the spatial transcriptome coordinates;
[0130] The spatial metabolic data fitting module fits the spatial metabolome data and associates it with the spatial transcriptome spots;
[0131] The spatial transcription and spatial metabolic data association module performs sum operation on the underlying pixel data corresponding to each spot to obtain the spatial metabolic group data corresponding to each spot of the spatial transcription group.
[0132] The beneficial effects of the present invention include: the present invention provides a method and system for spatial information matching between spatial transcriptome and spatial metabolome, which, by combining with subsequent data association methods, realizes point-to-point matching between spatial transcriptome and spatial metabolome, reduces the manual selection of regional matching between spatial transcriptome and spatial metabolome, and greatly increases the amount of associated data after point-to-point matching, providing more reliable data support for subsequent association analysis between spatial transcriptome and spatial metabolome. BRIEF DESCRIPTION OF THE DRAWINGS
[0133] Figure 1 Schematic diagram showing the process of the point-to-point matching method of the spatial transcriptome and the spatial metabolome of the present invention.
[0134] Figure 2 Schematic diagram of the spatial transcriptome chip detection method.
[0135] Figure 3 Schematic diagram of the spatial metabolome mass spectrometry detection method.
[0136] Figure 4 Schematic diagram showing the process of spatial information matching between spatial transcriptome and spatial metabolome.
[0137] Figure 5 Schematic diagram showing the direction and coordinate identification method of spatial metabolomics mass spectrometry detection.
[0138] Figure 6 Schematic diagram showing spot distribution and HE staining in the spatial transcriptome in Example 1 of the matching method.
[0139] Figure 7 Schematic diagram showing spot identification of spatial transcriptome selection in matching method Example 1.
[0140] Figure 8 This figure shows the spatial metabolomics image of the unmarked points in Example 1 of the matching method.
[0141] Figure 9 A schematic diagram showing the marking of identification points after matching the spatial transcriptome HE staining map with the spatial metabolome imaging map in Example 1 of the matching method.
[0142] Figure 10A mass spectrometry image showing the conversion of the spatial metabolome information into the spatial transcriptome information in Example 1 of the matching method.
[0143] Figure 11 Schematic diagram showing the flow of the method for associating spatial transcriptome and spatial metabolome data.
[0144] Figure 12 Schematic diagram showing the spatial distribution of spatial transcriptome detection.
[0145] Figure 13 Schematic diagram showing the spatial distribution of spatial metabolome detection and data storage methods.
[0146] Figure 14 A diagram showing the different superposition effects of the spatial distribution of the spatial transcriptome and the spatial metabolome.
[0147] Figure 15 This figure shows the schematic diagram of the calculation after two fittings when all 2*2 pixels of the spatial metabolome are sample areas.
[0148] Figure 16 This figure shows the schematic diagram of the calculation after two fittings when one pixel in the 2*2 pixel spatial metabolome is a non-sample area.
[0149] Figure 17 This figure shows the schematic diagram of the calculation after two fittings when two pixels in the 2*2 pixel spatial metabolome are non-sample areas.
[0150] Figure 18 This figure shows the schematic diagram of the calculation after two fittings when three pixels in the 2*2 pixel area of the spatial metabolome are non-sample areas.
[0151] Figure 19 This figure shows the schematic diagram of the calculation after two fittings when all 2*2 pixels of the spatial metabolome are non-sample areas.
[0152] Figure 20 Schematic diagram showing the association between spatial metabolome pixels and spatial transcriptome spots.
[0153] Figure 21 Schematic diagram showing the results of the second-order fitting sample area and non-sample area.
[0154] Figure 22 Schematic diagram showing the association between pixels and spots in the sample area and non-sample area.
[0155] Figure 23 Schematic diagram showing the visualization of the spatial metabolome data after conversion into spatial transcriptome distribution in Example 2.
[0156] Figure 24Schematic diagram showing the changes in the imaging image after spatial metabolomics data fitting in Example 2.
[0157] Figure 25 A schematic diagram showing a system for implementing the above-mentioned method for point-to-point matching of spatial transcriptome and spatial metabolome according to the present invention. DETAILED DESCRIPTION
[0158] The present invention is further described in detail with reference to the following specific examples and accompanying drawings. The processes, conditions, experimental methods, etc. for implementing the present invention, except for those specifically mentioned below, are common knowledge and common common sense in the art and are not particularly limited by the present invention.
[0159] The present invention proposes a spatial information matching method for spatial transcriptome and spatial metabolome, which can be subsequently combined with a data association method to achieve a point-to-point matching function between spatial transcriptome and spatial metabolome.
[0160] The spatial information matching method is used to transform the spatial coordinate position information of the spatial transcriptome and the spatial metabolome, so that the spatial coordinate positions of the two different omics systems can be matched. The data association method is used to process the underlying data in the spatial transcriptome and the spatial metabolome, so that the data of the two different omics systems can be associated.
[0161] Specifically, the spatial information matching method of spatial transcriptome and spatial metabolome is used to solve the problem of spatial position mismatch caused by different spatial information identifiers of spatial transcriptome and spatial metabolome, different experimental operation steps of different slice samples and non-identical slice samples.
[0162] The spatial transcriptome and spatial metabolome data association method is used to solve the problem that the spatial distribution of data with different spatial distributions cannot be associated in spatial position due to the different staggered distribution of the upper and lower layers of the spatial transcriptome and the square combination distribution of the spatial metabolome.
[0163] The spatial information matching method between the spatial transcriptome and the spatial metabolome and the data association method are combined and run to finally achieve point-to-point matching between the spatial transcriptome and the spatial metabolome.
[0164] Specifically, the spatial information matching method refers to unifying the spatial information identifiers of the spatial transcriptome and the spatial metabolome, and converting the spatial coordinate system of the spatial metabolome into the spatial coordinates of the spatial transcriptome to achieve spatial information matching between the spatial transcriptome and the spatial metabolome.
[0165] like Figure 4 As shown, the following steps are included:
[0166] (1) Select six or more spots within the sample detection range of the spatial transcriptome as identification points, obtain the barcode spatial information corresponding to the identification spots, and identify the spots on the corresponding HE (hematoxylin-eosin) staining image. Six is the minimum number of identification points. If there are fewer than six identification points, the matching effect is poor, and the identification points are scattered throughout the chip detection area.
[0167] (2) Use the HE staining image with the marked points to align with the spatial metabolome image to obtain the spatial coordinate position information of the corresponding pixel points of the marked points on the spatial metabolome. The spatial metabolome detection resolution is usually 100μm, 40μm, and 20μm, that is, the distance between pixels is 100μm, 40μm, and 20μm.
[0168] (3) Convert the barcode spatial information of the spatial transcriptome and the spatial metabolome spatial information into a unified spatial information identifier.
[0169] (4) Calculate the distances between different identification points of the spatial transcriptome; calculate the distances between different identification points of the spatial metabolome; and calculate the scaling ratio of the distances between the identification points corresponding to the spatial transcriptome and the spatial metabolome. Since the spatial coordinate systems of the spatial transcriptome and the spatial metabolome have been unified, the theoretical scaling ratio is 1. However, since the samples of the spatial metabolome and the spatial transcriptome are adjacent slice samples, there are differences in the actual sample detection. In order to prevent the scaling ratio from being different from 1 due to errors in point selection caused by different samples or instrument errors, it is necessary to perform scaling ratio correction in subsequent operations.
[0170] (5) Eliminate scaling ratios greater than 1.1 and less than 0.91. After elimination, the number of remaining scaling ratios must be greater than or equal to 4 (if it does not meet the requirements, restart from step (1)). Calculate the mean of the remaining scaling ratios to obtain the final scaling ratio, which is used as the final scaling ratio of the spatial transcriptome and the spatial metabolome. The final scaling ratio is used to correct data caused by errors in site selection and experimental operation.
[0171] (6) Use the spatial information of the identification points that meet the scaling ratio conditions to determine whether the spatial transcriptome and spatial metabolome samples are flipped during detection and calculate the final rotation angle between the two detection samples. The final rotation angle is the mean of all rotation angles obtained by calculating any two marker points, that is, the rotation angles of two points in the spatial transcriptome and the corresponding two points in the spatial metabolome are calculated. Since there are at least 4 or more identification points for the calculation of the rotation angle, more than 6 combinations of rotation angles will be obtained in the end. After calculating the rotation angles of all combinations of identification points of the spatial transcriptome and the spatial metabolome, it is also necessary to determine whether the rotation angle deviation is ≤5°. If so, perform the subsequent mean calculation step. Otherwise, return to step (3) and start again. The rotation angle deviation refers to the absolute difference between the rotation angle of each combination and the final rotation angle.
[0172] (7) The spatial metabolome coordinates are transformed using the transformation spatial identification method through the three parameters of the final rotation angle, the final scaling ratio, and the coordinates of the center mark point of the spatial transcriptome, so that the transformed spatial metabolome coordinates are finally matched with the spatial transcriptome coordinates.
[0173] The data association method refers to fitting the spatial metabolome data, integrating the fitted data into the spatial distribution pattern of the spatial transcriptome, and finally realizing the data association between the spatial metabolome and the spatial transcriptome at the point-to-point level.
[0174] like Figure 11 As shown, the following steps are included:
[0175] (1) The spatial metabolome data were fitted twice. After fitting, the pixel resolution of the spatial metabolome was improved from the original 100 μm × 100 μm to 25 μm × 25 μm.
[0176] (2) Correlate the pixels of the spatial metabolome fitting with the spots of the spatial transcriptome;
[0177] (3) Perform sum operation on the underlying pixel data corresponding to each spot;
[0178] (4) Obtain the spatial metabolome data corresponding to each spot in the spatial transcriptome.
[0179] Example 1
[0180] Matching method
[0181] (1) Select 6 spots as identification points within the sample detection range of the spatial transcriptome, obtain the barcode information corresponding to the identification spot, and mark the spot on the corresponding HE staining map (see Figure 6 and Figure 7 ). Figure 6Schematic diagram of spot distribution and HE staining in the spatial transcriptome; Figure 7 This is the distribution map of the six spot points selected in the spatial transcriptome.
[0182] Table 1 Barcode and spatial information of spatial transcriptome markers
[0183] Barcode Marking point TransX TransY GGTAGTGCTCGCACCA a 72 34 GATGGCGCACACATTA b 29 19 CTGGACGCAGTCCGGC c 13 51 CACAAGAAAGATATTA d 112 14 AATAGAATCTGTTTCA e 86 48 CTGCAAGCACGTTCCG f 111 61
[0184] In Table 1, Barcode is the spatial barcode in spatial transcriptome technology. Each spot has its own unique barcode code, and the spatial location information of the corresponding spot can be obtained through the spatial barcode; TransX and TransY are the two sets of spatial numbers x and y in the spatial transcriptome;
[0185] (2) Use the HE staining image with the marked points to align with the spatial metabolome image to obtain the spatial coordinate position information of the corresponding pixel points of the marked points on the spatial metabolome after alignment (see Figure 8 and Figure 9 ). Figure 8 This is the spatial metabolomics image of the unlabeled points; Figure 9 Schematic diagram including identification points after matching the spatial transcriptome HE staining map and the spatial metabolome imaging map.
[0186] Table 2 Spatial information table of spatial metabolomics identification points
[0187]
[0188]
[0189] In Table 2, MetaX and MetaY are two sets of spatial numbers of x and y in the spatial metabolome.
[0190] (3) Convert the barcode spatial information of the spatial transcriptome and the spatial metabolome spatial information into a unified spatial information identifier.
[0191] Table 3 Unified space information table
[0192] Marking point UTX UTY UMX UMY a 3650 2944.5 4100 5200 b 1500 1645.5 5200 7500 c 700 4416.7 2200 8000 d 5650 1212.4 6200 3400 e 4350 4156.9 3000 4300 f 5600 5282.8 1900 2900
[0193] In Table 3, UTX and UTY are the spatial coordinates of the unified standard of spatial transcriptome data; UMX and UMY are the spatial coordinates of the unified standard of spatial metabolome data.
[0194] (4) Calculate the distances between different identification points of the spatial transcriptome; calculate the distances between different identification points of the spatial metabolome; and calculate the scaling ratio of the distances between the identification points corresponding to the spatial transcriptome and the spatial metabolome.
[0195] Table 4 Calculation results of the scaling ratio of two identification points
[0196] Two-by-two marking points Idle distance Space generation distance Zoom ratio a—b 2513.0 2549.5 1.01 a—c 3297.0 3383.8 1.03 a—d 2645.8 2765.9 1.05 a—e 1400.0 1421.3 1.02 a—f 3044.7 3182.8 1.05 b—c 2884.4 3041.4 1.05 b—d 4172.5 4220.2 1.01 b—e 3798.7 3883.3 1.02 b—f 5480.9 5661.3 1.03 c—d 5896.6 6095.9 1.03 c—e 3659.2 3785.5 1.03 c—f 4975.9 5108.8 1.03 d—e 3218.7 3324.2 1.03 d—f 4070.6 4329.0 1.06 e—f 1682.3 1780.5 1.06
[0197] (5) Scaling ratios greater than 1.1 and less than 0.91 were eliminated. After elimination, the number of remaining scaling ratios must be greater than or equal to 4. The mean of the remaining scaling ratios was calculated to obtain the final scaling ratio, which was used as the final scaling ratio of the spatial transcriptome and the spatial metabolome. The final scaling ratio was used for point selection, which was used to eliminate points that introduced excessive errors and to correct data caused by experimental operation errors. The final scaling ratio was 1.03.
[0198] (6) The spatial information of the identified points that meet the scaling conditions is used to determine whether the spatial transcriptome and spatial metabolome samples are flipped during the detection process and to calculate the final rotation angle between the two detected samples. The final rotation angle is 82.9°. The absolute difference between the rotation angle of each combination and the final rotation angle is called the rotation angle deviation. The final rotation angle is 82.9°, and the rotation angle of each combination is ≤5° from the final rotation angle, which meets the rotation angle deviation threshold.
[0199] Table 5 Calculation results of rotation angles of two marking points
[0200] Two-by-two marking points Rotation angle a—b 84.4 a—c 82.4 a—d 81.5 a—e 80.7 a—f 83.6 b—c 83.4 b—d 82.2 b—e 83.1 b—f 84.1 c—d 81.9 c—e 81.9 c—f 83.3 d—e 81.9 d—f 82.7 e—f 86.1 Final rotation angle 82.9
[0201] (7) The spatial metabolome coordinates are transformed using the transformation space identification method through the three parameters of the final rotation angle, the final scaling ratio, and the coordinates of the center mark of the spatial transcriptome, so that the transformed spatial metabolome coordinates are matched with the spatial transcriptome coordinates. Figure 10 As shown, Figure 10 This is the mass spectrometry imaging diagram after the spatial metabolome information is converted into the spatial transcriptome information.
[0202] Table 6 Conversion of some spatial metabolome coordinates into spatial transcriptome information
[0203]
[0204]
[0205] In Table 6, MetaX and MetaY are the two sets of spatial numbers of x and y in the spatial metabolome; UMX and UMY are the spatial coordinates after the spatial metabolome data are standardized; UMX' and UMY' are the spatial information after the spatial metabolome is matched with the spatial transcriptome; MetaX' and MetaY' are the two sets of spatial numbers of x and y in the spatial metabolome spatial information after the spatial metabolome is matched with the spatial transcriptome spatial information; TransX' and TransY' are the two sets of spatial numbers of x and y in the spatial transcriptome spatial information after the spatial metabolome is matched with the spatial transcriptome spatial information.
[0206] Example 2
[0207] Data association methods
[0208] In the data association method, the following example illustrates how fitting can improve the pixel resolution of the spatial metabolome:
[0209] like Figure 15 As shown in the figure, if the pixels are arranged in a 2*2 pattern and are all sample area pixels, the calculation method for the middle pixel after the four pixels are amplified is as follows:
[0210] The underlying data of the amplified pixel in the middle of the 2*2 pixel point is calculated according to the following formula:
[0211] MZ1abcd=(MZ1a+MZ1b+MZ1c+MZ1d) / 4;
[0212] MZ2abcd=(MZ2a+MZ2b+MZ2c+MZ2d) / 4;
[0213] …
[0214] MZnabcd=(MZna+MZnb+MZnc+MZnd) / 4;
[0215] Among them, pixel points a, b, c, and d are arranged in 2*2 order; MZ1a, MZ2a, …, MZna are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point a; MZ1b, MZ2b, …, MZnb are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point b; MZ1c, MZ2c, …, MZnc are the expression intensities of the mass-to-nuclear ratios of MZ1, MZ2, …, MZn corresponding to pixel point c. MZ1d, MZ2d, …, MZnd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel d; MZ1abcd, MZ2abcd, …, MZnabcd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to the intermediate pixels after amplification between pixels a, b, c, d; the amplified pixels between pixels a, b, c, d are abbreviated as {abcd}.
[0216] like Figure 16 As shown in the figure, if the pixels are arranged in a 2*2 pattern and one of the pixels is located in the non-sample area, the calculation method for the middle pixel after the four pixels are amplified is as follows:
[0217] The underlying data of the amplified pixel in the middle of the 2*2 pixel point is calculated according to the following formula, where pixel a is a pixel in the non-sample area:
[0218] MZ1abcd=(MZ1b+MZ1c+MZ1d) / 3;
[0219] MZ2abcd=(MZ2b+MZ2c+MZ2d) / 3;
[0220] …
[0221] MZnabcd=(MZnb+MZnc+MZnd) / 3;
[0222] Among them, pixels a, b, c, and d are arranged in 2*2 order; MZ1b, MZ2b, …, MZnb are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel b; MZ1c, MZ2c, …, MZnc are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel c; MZ1d, MZ2d, …, MZnd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel d; MZ1abcd, MZ2abcd, …, MZnabcd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to the amplified intermediate pixels between pixels a, b, c, and d; the amplified pixels between pixels a, b, c, and d are abbreviated as {bcd}.
[0223] like Figure 17 As shown in the figure, if the pixels are arranged in a 2*2 pattern and two pixels are located in the non-sample area, the calculation method for the middle pixel after the four pixels are amplified is as follows:
[0224] The underlying data of the amplified pixel in the middle of the 2*2 pixel point is calculated according to the following formula, where pixels a and b are pixels in the non-sample area:
[0225] MZ1abcd=(MZ1c+MZ1d) / 2;
[0226] MZ2abcd=(MZ2c+MZ2d) / 2;
[0227] …
[0228] MZnabcd=(MZnc+MZnd) / 2;
[0229] Among them, pixels a, b, c, and d are arranged in 2*2 order; MZ1c, MZ2c, …, MZnc are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel c; MZ1d, MZ2d, …, MZnd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to pixel d; MZ1abcd, MZ2abcd, …, MZnabcd are the expression intensities of the CPR of MZ1, MZ2, …, MZn corresponding to the amplified intermediate pixels between pixels a, b, c, and d; the amplified pixels between pixels a, b, c, and d are abbreviated as {cd}.
[0230] like Figure 18 As shown in the figure, if the pixels are arranged in a 2*2 pattern, and three of the pixels are located in the non-sample area, the calculation method for the middle pixel after the four pixels are amplified is as follows:
[0231] The underlying data of the amplified pixel in the middle of the 2*2 pixel point is calculated according to the following formula, where pixels a, b, and c are pixels in the non-sample area:
[0232] MZ1abcd=0;
[0233] MZ2abcd=0;
[0234] …
[0235] MZnabcd=0;
[0236] Pixels a, b, c, and d are arranged in a 2*2 pattern; MZ1abcd, MZ2abcd, …, MZnabcd are the expression intensities of the CNRs of the intermediate pixels between the four pixels a, b, c, and d after amplification, corresponding to MZ1, MZ2, …, MZn; the amplified pixels between the pixels a, b, c, and d are abbreviated as 0, indicating that the amplified pixels are non-sample area pixels.
[0237] like Figure 19 As shown in the figure, if the pixels are arranged in a 2*2 pattern and are all in the non-sample area, the calculation method for the middle pixel after the four pixels are amplified is as follows:
[0238] The underlying data of the amplified pixel in the middle of the 2*2 pixel point is calculated according to the following formula:
[0239] MZ1abcd=0;
[0240] MZ2abcd=0;
[0241] …
[0242] MZnabcd=0;
[0243] Pixels a, b, c, and d are arranged in a 2*2 pattern; MZ1abcd, MZ2abcd, …, MZnabcd are the expression intensities of the CNRs of the intermediate pixels between the four pixels a, b, c, and d after amplification, corresponding to MZ1, MZ2, …, MZn; the amplified pixels between the pixels a, b, c, and d are abbreviated as 0, indicating that the amplified pixels are non-sample area pixels.
[0244] Example 3
[0245] Correlating spatial metabolome data with spatial transcriptome data in mouse kidney
[0246] (1) The spatial metabolome data of mouse kidney were fitted twice. After fitting, the pixel resolution of the spatial metabolome was improved from the original 100 μm × 100 μm to 25 μm × 25 μm; Figure 24 )
[0247] (2) Correlate the pixels of the fitted spatial metabolome with the spots of the spatial transcriptome;
[0248] Table 7 Correlation table between some spatial metabolome coordinates and spatial transcriptome barcodes
[0249]
[0250]
[0251] The barcode in Table 7 is the spatial barcode in spatial transcriptomics technology. Each spot has its own unique barcode code, and the spatial location information of the corresponding spot can be obtained through the spatial barcode; the spatial metabolome coordinates are the combination of two sets of spatial numbers x and y in the spatial metabolome.
[0252] (3) Perform sum operation on the underlying pixel data corresponding to each spot ( Figure 23 ); the underlying data refers to the expression intensity data corresponding to mz at each pixel in the spatial metabolome analysis. mz is the ratio of the mass number to the charge number of a charged particle and is substance-specific data. It is usually referred to as mz.
[0253] (4) Obtain the spatial metabolome data corresponding to each spot in the spatial transcriptome.
[0254] Table 8 Partial spatial metabolomics data
[0255] mz 1-5 2-4 3-5 5-5 7-5 71.01222 1400.171 2714.815 1905.272 1492.15 1368.167 72.71514 259.7251 515.8836 798.2261 715.2259 472.7763 72.99158 109.9754 1056.229 986.7844 1435.146 535.1709 73.02797 817.4667 0 631.1629 895.4172 454.4463 74.0232 27.17561 434.8098 135.878 63.1555 252.8567 75.00722 403.1064 787.5696 806.4386 759.8611 831.8674 75.45135 662.9783 756.3464 722.308 576.1749 635.5394 78.95747 9850.157 8887.678 10582.36 10793.19 10528.72 79.95577 1097.258 1293.686 1267.888 1538.529 1882.602 83.04869 523.8511 0 300.4212 590.7473 322.4308
[0256] In Table 8, mz represents the mass-to-charge ratio data for the spatial metabolome; 1-5, 2-4, 3-5, 5-5, and 7-5 represent the expression intensities of the mass-to-charge ratios for the spatial number combinations of x and y in the spatial metabolome. The expression intensity is the relative expression value of the spatial metabolite corresponding to mz at each pixel. The numerical value indicates the relative content, with larger values indicating higher relative content at the corresponding pixel.
[0257] Table 9 Partial spatial transcriptome data
[0258]
[0259]
[0260] In Table 9, the Gene Name column is the gene name in the spatial transcriptome; the values in the GCCACCCATTCCACTT, TACTCACAACGTAGTA, GTTCGGTGTGGATTTA, GGAGACATTCACGGGC, and TAGAACGCCAGTAACG columns are the count values of the corresponding genes encoded by the Barcode of the spatial transcriptome.
[0261] The count value is the number of transcripts detected for the corresponding Gene Name in the spatial transcriptome. Transcript expression quantification uses the count value to determine the level of transcriptome expression. A larger count value corresponds to a higher transcriptome expression.
[0262] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the appended claims.
Claims
1. A method for matching spatial information of spatial transcriptome and spatial metabolome, characterized in that: The spatial information matching method unifies the spatial information identifiers of the spatial transcriptome and the spatial metabolome, and converts the spatial coordinate system of the spatial metabolome into the spatial coordinates of the spatial transcriptome, thereby achieving spatial information matching between the spatial transcriptome and the spatial metabolome, comprising the following steps: Step (1) selecting 6 or more evenly distributed spots within the sample detection range of the spatial transcriptome as identification spots, obtaining the barcode information corresponding to the identification spots, and marking the selected spots on the corresponding HE staining map; Step (2), using the HE staining image of the marked points in step 1 to align with the spatial metabolomics image, and obtaining the spatial coordinate position information of the corresponding pixel points of the marked points on the spatial metabolomics after alignment; Step (3), converting the barcode spatial information of the spatial transcriptome and the spatial metabolome spatial information into a unified spatial information identifier; Step (4), calculating the distances between different identification points in the spatial transcriptome; calculating the distances between different identification points in the spatial metabolome; and calculating the scaling ratio of the distances between the identification points corresponding to the spatial transcriptome and the spatial metabolome; In step (4), the scaling ratio is calculated in order to evaluate the accuracy of the corresponding marker points in the spatial transcriptome and the spatial metabolome and to correct the position differences of the marker points caused by the error of the experimental instrument; The calculation formula of the scaling ratio is as follows: Ratio=sqrt((UTXa-UTXb) 2 +(UTYa-UTYb) 2 ) / sqrt((UMXa-UMXb) 2 +(UMYa-UMYb) 2 ) Among them, UTXa, UTYa, UTXb, and UTYb are the X and Y coordinate values of the landmark points a and b in the spatial transcriptome after the spatial coordinates are unified; UMXa, UMYa, UMXb, and UMYb are the X and Y coordinate values of the landmark points a and b in the spatial metabolome after the spatial coordinates are unified; Ratio is the scaling ratio of the distance between the landmark points a and b in the spatial transcriptome to the distance between the landmark points a and b in the spatial metabolome; After obtaining the scaling ratios between each pair of landmarks, the scaling ratio data needs to be screened. The screening operation includes: eliminating data with scaling ratios greater than 1.1 and less than 0.91, while ensuring that the number of remaining scaling ratio data is ≥4, and averaging the remaining scaling ratio data to obtain the final scaling ratio for spatial metabolome conversion; if the number of remaining scaling ratio data is <4, return to step (1) and repeat; Step (5), judging whether the spatial transcriptome and the spatial metabolome samples are flipped during the detection process by the rotation angle of the lines connecting each pair of identification points in the spatial transcriptome and the spatial metabolome, and calculating the final rotation angle between the spatial transcriptome and the spatial metabolome detection samples; Step (6), using the final rotation angle, the final scaling ratio, and the coordinates of the center mark point of the spatial transcriptome to transform the spatial metabolome coordinates using the conversion space mark method, so that the transformed spatial metabolome coordinates match the spatial transcriptome coordinates; In step (6), the spatial transcriptome center identification point refers to the center point position of the spatial transcriptome detection chip; The conversion space identification method refers to converting the spatial position of each pixel in the spatial metabolome into the corresponding coordinates in the spatial transcriptome through spatial coordinate transformation. The final scaling ratio and final rotation angle obtained in step (4) and step (5) also satisfy the following calculation formula: Ratio'=sqrt((UMX'-UTXo) 2 +(UTY'-UTYo) 2 ) / sqrt((UMX-UMXo) 2 +(UMY-UMYo) 2 ); Angle'=arctan((UMY-UMYo) / (UMX-UMXo))-arctan((UMY'-UTYo) / (UMX'-UTXo)); Among them, Ratio' is the final scaling ratio; Angle' is the final rotation angle; UTXo and UTYo are the spatial information of the center point position of the spatial transcriptome detection chip after the unified coordinate system; UMXo and UMYo are the spatial information of the center point position of the spatial transcriptome detection chip in the unified coordinate system of the corresponding position of the spatial metabolome; UMX and UMY are the X and Y coordinate values of the spatial metabolome after the unified spatial coordinates; UMX' and UMY' are the X and Y coordinate values after the spatial metabolome matches the spatial transcriptome.
2. The spatial information matching method according to claim 1, wherein: There are a total of 4 spatial transcriptome chips on the expression slide of the spatial transcriptome; the detection area size of each spatial transcriptome chip is 6.5mm×6.5mm, with a total of 4992 spots, the diameter of a single spot is 55μm, and the distance between each two adjacent spots and the center of the spot is 100μm; the spot is the smallest unit detected in the spatial transcriptome.
3. The spatial information matching method according to claim 1, wherein: In step (1), the barcode is a spatial barcode in the Visium spatial transcriptome technology. Each spot has its own unique barcode code, and the spatial location information of the corresponding spot can be obtained through the spatial barcode; and / or, In step (1), the sample after HE staining can still be used for spatial transcriptome detection without affecting the spatial transcriptome data. The HE staining image can match the spatial position of the spatial transcriptome data, and spot identification can be performed on the HE staining image.
4. The spatial information matching method according to claim 1, wherein: In step (2), the pixel point is the smallest unit detected in the spatial metabolome; the resolution of the smallest unit detected in the spatial metabolome varies according to the mass spectrometry scanning rate and the moving platform speed, and the higher the speed, the greater the scanning resolution; and / or, In step (2), the registration operation refers to rotating, scaling, and flipping the HE staining image based on the spatial metabolomics imaging map through the MSIReader software, matching the spatial metabolomics imaging map and the HE staining map, and obtaining the spatial coordinate information of the identification site on the HE staining map on the spatial metabolomics imaging map through the software.
5. The spatial information matching method according to claim 4, wherein: The resolution of the smallest unit detected in the spatial metabolome represents the distance between pixels; different resolutions result in different distances between pixels; and the conversion of spatial information in the spatial metabolome is affected by the resolution.
6. The spatial information matching method according to claim 1, wherein: In step (3), the conversion into a unified spatial information identifier refers to performing operations on the two sets of spatial numbers x and y in the spatial information corresponding to the barcode barcode in the spatial transcriptome and the two sets of spatial numbers x and y in the spatial information of the spatial metabolome, respectively, and converting the coordinates of the spatial transcriptome and the spatial metabolome into a micron coordinate system to obtain a unified spatial coordinate system.
7. The spatial information matching method according to claim 6, wherein: The conversion formula of x space number in the spatial transcriptome is as follows: UTX=(TransX / 2+0.5)*transresolution; The conversion formula of y space number in spatial transcriptome is as follows: UTY=TransY*sqrt(0.75)*transresolution; Among them, UTX and UTY are the spatial coordinates of the unified standard of spatial transcriptome data; TransX and TransY are the two sets of spatial numbers of x and y in the spatial transcriptome; transresolution is the numerical value of the spatial transcriptome resolution; sqrt is the square root operation.
8. The spatial information matching method according to claim 6, wherein: The conversion formula of the x space number in the spatial metabolome is as follows: UMX=MetaX*metaresolution; The conversion formula of y space number in the spatial metabolome is as follows: UMY=MetaY*metaresolution; Among them, UMX and UMY are the spatial coordinates after the spatial metabolome data are standardized; MetaX and MetaY are the two sets of spatial numbers of x and y in the spatial metabolome, respectively; metaresolution is the numerical value of the spatial metabolome resolution.
9. The spatial information matching method according to claim 1, wherein: In step (5), the marker points are the marker points after the scaling ratio screening; if the deviation of all the rotation angles of the lines between the marker points is within 5°, it means that the sample has not been flipped; if the deviation of all the rotation angles of the lines between the marker points is outside 5°, the conversion formula of the y space number in the spatial metabolome in step (3) is modified to: UMY=-MetaY*resolution, Then, the rotation angles between each pair of landmark points in the spatial transcriptome and spatial metabolome detection samples were calculated. After obtaining all rotation angles with a deviation within 5°, the average of all rotation angles was calculated to obtain the final rotation angle. The calculation formula of the rotation angle is as follows: Angle=arctan((UMYa-UMYb) / (UMXa-UMXb))-arctan((UTYa-UTYb) / (UTXa-UTXb)); Among them, UTXa, UTYa, UTXb, and UTYb are the X and Y coordinate values of the landmark points a and b in the spatial transcriptome after unified spatial coordinates; UMXa, UMYa, UMXb, and UMYb are the X and Y coordinate values of the landmark points a and b in the spatial metabolome after unified spatial coordinates; arctan is the calculation of the inverse tangent function; Angle is the rotation angle between the line connecting the landmark points a and b in the spatial transcriptome and the line connecting the landmark points a and b in the spatial metabolome.
10. The spatial information matching method according to claim 9, wherein: The obtained final rotation angle, final scaling ratio, and coordinate parameters of the center marker point of the spatial transcriptome can be used to obtain the X and Y coordinate values UMX' and UMY' after the spatial metabolome matches the spatial transcriptome.
11. The spatial information matching method according to claim 10, wherein: Since the spatial transcriptome and metabolome analysis systems only recognize the corresponding spatial information patterns, in order to make their spatial information mutually recognizable, the unified and related coordinate system is restored and converted into the spatial information system of the spatial transcriptome and metabolome; The calculation formula for converting the spatial information of the unified coordinate system into the spatial information system of the spatial transcriptome is as follows: UMX'=(TransX' / 2+0.5)*transresolution; UMY'=TransY'*sqrt(0.75)*transresolution; Among them, UMX' and UMY' are the X and Y coordinate values after the spatial metabolome is matched with the spatial transcriptome; transresolution is the value of the spatial transcriptome resolution; sqrt is the square root operation; TransX' and TransY' are the two sets of spatial numbers of x and y in the spatial transcriptome after the spatial metabolome is matched with the spatial transcriptome; The calculation formula for converting the spatial information of the unified coordinate system into the spatial metabolomics spatial information system is as follows: UMX'=MetaX'*metaresolution; UMY'=MetaY'*metaresolution; Among them, metaresolution is the numerical value of the spatial metabolome resolution; MetaX' and MetaY' are two sets of spatial numbers of x and y in the spatial metabolome spatial information after the spatial metabolome matches the spatial transcriptome spatial information.
12. An application of the spatial information matching method according to any one of claims 1 to 11 in spatial information matching of spatial transcriptome and spatial metabolome.
13. A system for implementing the spatial information matching method according to any one of claims 1 to 11, characterized in that: The system includes a spatial information acquisition module, a spatial coordinate unification module, a spatial parameter calculation module, and a spatial information conversion module; The spatial information acquisition module aligns the HE staining image containing the spatial transcriptome information with the marked points with the spatial metabolome imaging image to obtain the spatial coordinate position information of the marked points on the spatial metabolome; The spatial coordinate unification module converts the spatial information of the spatial transcriptome and the spatial metabolome into a unified spatial information identifier; The spatial parameter calculation module calculates the final scaling ratio and final rotation angle between the spatial transcriptome and the spatial metabolome; The spatial information conversion module uses the three parameters of the final rotation angle, the final scaling ratio, and the coordinates of the center mark point of the spatial transcriptome to convert the spatial metabolome coordinates using a conversion space mark method, so that the converted spatial metabolome coordinates match the spatial transcriptome coordinates.
14. A hardware system or computer-readable storage medium for implementing the method according to any one of claims 1 to 11, characterized in that: The hardware system includes: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 11 is implemented; when the computer program is executed by the processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Method and system for association analysis of transcriptome and metabolome data
CN109979527A
Spatial transcriptome data processing method and device and computer readable storage medium
CN114267414A