Systems, methods, apparatuses, and media for spatial omics-based marker screening
The spatial omics biomarker screening system has solved the problem of low-abundance biomarkers being masked in fecal biomarker research, enabling more accurate screening of disease biomarkers and advancing the development of personalized medicine research.
Patent Information
- Application Number
- CN202411892203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing fecal biomarker studies have neglected the spatial characteristics of fecal heterogeneity, resulting in the masking of low-abundance biomarkers. Furthermore, different sampling locations lead to poor repeatability, making it difficult to effectively screen biomarkers for diseases such as colorectal cancer.
A biomarker screening system based on spatial omics was adopted. Through spatial mapping and analysis modules, molecular type and location data were obtained, molecular-level spatial mapping and differential analysis were performed, and a custom target analysis region was defined to screen differentially expressed molecules as candidate biomarkers.
It can detect more low-abundance biomarkers in samples that are masked by homogenization, providing a rich database of potential biomarkers for personalized medicine research and improving the accuracy and repeatability of screening.
Smart Images

Figure CN119943162B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biology, and in particular to a marker screening system, method, device and medium based on spatial omics. BACKGROUND
[0002] With the development of precision medicine, people will collect the corresponding metabolites of the human body according to different diseases for detection, and then different markers in the sample are used to investigate the incidence of different diseases. For example, human feces is a kind of biological sample that is easy to obtain and low in cost, which contains a variety of host proteins, microorganisms, food and other biological molecules. When it passes along the gastrointestinal tract, it will incorporate various substances secreted by the intestinal environment, so that various intestinal-related disease biomarkers can be mined from human feces to detect the corresponding incidence, including colorectal cancer (CRC), inflammatory bowel disease, diabetes and obesity, etc.
[0003] In the existing fecal marker research, small aliquot of feces (for example, 200 mg or less) is usually collected, which is directly stirred and mixed for sample preparation and mass spectrometry analysis. However, this ignores the research value of fecal heterogeneity. Human feces does not represent a single colon environment, but represents the entire gastrointestinal environment. From the physiological formation and discharge process of feces, it can be known that feces gradually forms along the entire digestive tract, and many small intestine-related substances are distributed inside the feces, and the outer layer of the feces is mainly related to the colon microenvironment.
[0004] Therefore, the practice of sampling and homogenizing different spatial positions in the feces to obtain data representing the entire feces has limitations. For example, mixing the feces will cause many CRC-related low-abundance proteins distributed in the outer layer of the feces to be covered up, and due to the different sampling positions of the feces, the CRC fecal markers screened out differ greatly and have poor reproducibility. SUMMARY
[0005] The main purpose of the embodiments of the present application is to provide a marker screening system, method, device and medium based on spatial omics, which can detect more low-abundance markers that have been covered up due to homogenization in the sample through a spatial mapping module and a spatial analysis module, provide a rich potential biomarker database for corresponding research diseases, and promote the development of personalized medicine research.
[0006] To achieve the above purpose, a first aspect of the embodiments of the present application provides a marker screening system based on spatial omics, comprising:
[0007] a spatial mapping module and a spatial analysis module, the spatial analysis module being in communication connection with the spatial mapping module;
[0008] The spatial mapping module is configured to obtain a target molecule type, a spatial omics dataset of a target sample, and a spatial position dataset, and perform spatial mapping of the target sample at a molecular level according to the target molecule type, the spatial omics dataset, and the spatial position dataset to obtain molecular distribution data of the target sample.
[0009] The spatial mapping module is further configured to correspondingly draw a molecular distribution map of the target sample according to the molecular distribution data.
[0010] The spatial analysis module is configured to obtain pathological information of a target disease and molecular distribution maps of a plurality of target samples, and self-define target analysis regions corresponding to the target samples in each of the molecular distribution maps according to the pathological information.
[0011] The spatial analysis module is further configured to perform differential analysis and screening on molecules in each of the target analysis regions according to a statistical test method, to obtain differential molecules, and determine the differential molecules as candidate markers of the target disease.
[0012] Further, in some embodiments, performing differential analysis and screening on molecules in each of the pathological feature distribution regions according to the statistical test method to obtain differential molecules includes:
[0013] Summing the abundance values of the molecules of the same molecule type in each of the pathological feature distribution regions to obtain a molecule total abundance set of each of the pathological feature distribution regions;
[0014] Normalizing the molecule total abundance set of each of the pathological feature distribution regions to obtain a normalized molecule total abundance set of each of the pathological feature distribution regions;
[0015] Performing differential analysis and screening on the normalized molecule total abundance set of each of the pathological feature distribution regions to obtain differential molecules between each of the pathological feature distribution regions.
[0016] Further, in some embodiments, determining the target analysis regions corresponding to the target samples in each of the molecular distribution maps according to the pathological information includes:
[0017] Determining a pathological feature of the target disease according to the pathological information;
[0018] Determining a pathological feature distribution region set of each of the target samples in each of the molecular distribution maps according to the pathological feature, the pathological feature distribution region set including a plurality of pathological feature distribution regions of different region ranges;
[0019] Self-defining a pathological feature distribution region in the pathological feature distribution region set of each of the target samples as a target analysis region corresponding to the target sample to determine the target analysis regions corresponding to the target samples.
[0020] Further, in some embodiments, according to the target molecule type, the spatial omics dataset and the spatial position dataset, the spatial mapping of the target sample at the molecular level is performed to obtain the molecular distribution data of the target sample, including:
[0021] According to the target molecule type, the target spatial omics data of the target molecule in the spatial omics dataset is determined;
[0022] The target spatial omics dataset is normalized to obtain a normalized target spatial omics dataset;
[0023] According to the target molecule type, the target spatial position data of the target molecule in the spatial position dataset is determined;
[0024] The normalized target spatial omics data and the target spatial position data are spatially correlated and mapped to obtain the molecular distribution data.
[0025] Further, in some embodiments, the marker screening system further comprises a sampling module, which is in communication connection with the spatial mapping module, wherein:
[0026] The sampling module is used to obtain the target sample, extract molecules from the target sample to obtain a sample molecule solution;
[0027] The sampling module is further used to perform mass spectrometry detection on the sample molecule solution to obtain the spatial omics data and the spatial position dataset of the target sample.
[0028] Further, in some embodiments, when the sample molecule solution is a protein molecule solution, the molecules are extracted from the target sample to obtain sample molecules, including:
[0029] The target sample is frozen and sliced to obtain a sample slice;
[0030] The sample slice is divided into a plurality of sample microparticles with spatial position information by using a micro support;
[0031] According to the pressure cycle technology, the sample microparticles are subjected to protein lysis treatment to obtain a protein lysis solution;
[0032] According to the Automated SP3 technology, the protein lysis solution is subjected to high-throughput and automated enzymatic digestion purification treatment to obtain a protein peptide solution;
[0033] The protein peptide solution is transferred to Evotips to obtain a protein molecule solution.
[0034] Further, in some embodiments, when the sample molecule solution is a lipid molecule solution, the molecules are extracted from the target sample to obtain sample molecules, including:
[0035] Freezing and slicing the target sample to obtain a sample slice;
[0036] Segmenting the sample slice into a plurality of sample microparticles retaining spatial position information by using a micro support;
[0037] Performing lipid dissolution treatment on the sample microparticles by using a lipid extraction solution, and taking supernatant, to obtain a lipid molecule solution after concentration and redissolution.
[0038] To achieve the above object, a second aspect of the embodiment of the present application proposes a marker screening method based on spatial omics, comprising:
[0039] Obtaining spatial omics data sets and spatial position data sets of a target molecule type and a plurality of target samples, and performing spatial mapping at a molecular level on each target sample according to the target molecule type, each spatial omics data set and each spatial position data set, to obtain molecular distribution data of each target sample;
[0040] According to each molecular distribution data, drawing a molecular distribution graph of each target sample;
[0041] Obtaining pathological information of a target research disease, and according to the pathological information, defining a target analysis region corresponding to each target sample in each molecular distribution graph;
[0042] According to a statistical test method, performing differential analysis and screening on molecules in each target analysis region, to obtain differential molecules, and determining the differential molecules as markers of the target research disease.
[0043] To achieve the above object, a third aspect of the embodiment of the present application proposes an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the marker screening method of the second aspect of the embodiment when executing the computer program.
[0044] To achieve the above object, a fourth aspect of the embodiment of the present application proposes a storage medium, which is a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the marker screening method of the second aspect of the embodiment.
[0045] In the embodiments of the present application, the following beneficial effects are achieved: the present application comprises a spatial mapping module and a spatial analysis module, the spatial analysis module is in communication connection with the spatial mapping module, wherein the spatial mapping module is configured to acquire a target molecule type, a spatial omics data set and a spatial position data set of a target sample, and perform spatial mapping of the target sample at a molecular level according to the target molecule type, the spatial omics data set and the spatial position data set, so as to obtain molecular distribution data of the target sample; the spatial mapping module is further configured to correspondingly draw a molecular distribution map of the target sample according to the molecular distribution data; the spatial analysis module is configured to acquire pathological information of a target research disease and molecular distribution maps of a plurality of target samples, and according to the pathological information, define target analysis regions corresponding to the target samples in each molecular distribution map; the spatial analysis module is further configured to perform differential analysis and screening on molecules in each target analysis region according to a statistical test method, so as to obtain differential molecules, and determine the differential molecules as candidate markers of the target research disease, thereby being capable of detecting more low-abundance markers that are hidden due to homogenization in the past, providing a rich potential biomarker database for the corresponding research disease, and promoting the development of personalized medical research. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 is an architecture diagram of a marker screening system based on spatial omics provided by some embodiments of the present application;
[0047] Figure 2 is an operation flowchart of the marker screening system provided by some embodiments of the present application;
[0048] Figure 3 is a flowchart of the spatial mapping module performing spatial mapping of the target sample at a molecular level provided by some embodiments of the present application;
[0049] Figure 4 is a fecal spatial distribution map of HP, TF and TG drawn by the spatial mapping module provided by some embodiments of the present application;
[0050] Figure 5 is a flowchart of the spatial analysis module determining a target analysis region provided by some embodiments of the present application;
[0051] Figure 6 is a flowchart of the spatial mapping module performing differential analysis on molecules according to a statistical test method provided by some embodiments of the present application;
[0052] Figure 7 is a result schematic diagram of the spatial analysis module calculating differential molecules of the outermost layer and the whole region of the feces provided by some embodiments of the present application;
[0053] Figure 8is a result diagram of the correlation and co-localization characteristics of the differential protein molecules provided by some embodiments of the present application;
[0054] Figure 9 is a flowchart of protein molecule extraction on a target sample provided by some embodiments of the present application;
[0055] Figure 10 is a flowchart of lipid molecule extraction on a target sample provided by some embodiments of the present application;
[0056] Figure 11 is a flowchart of the marker screening method based on spatial omics provided by some embodiments of the present application;
[0057] Figure 12 is a hardware structure schematic diagram of an electronic device provided by some embodiments of the present application. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0059] In the description of the present application, it should be understood that the orientation description, such as the orientation or position relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element indicated must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0060] It should also be noted that in the description of the present application, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, more than, etc. are understood as not including the number, above, below, etc. are understood as including the number. If it is described as first, second, it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or the sequence of indicated technical features.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0062] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "exemplary embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the exemplary description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0063] With the development of precision medicine, people will collect the corresponding metabolites of the human body according to different diseases for detection, and then through different markers in the sample to exclude the incidence of different diseases, for example, human feces is a kind of easily accessible and low-cost biological sample, which contains a variety of host proteins, microorganisms, food and other biological molecules, and when it passes along the gastrointestinal tract, it will mix with various substances secreted by the intestinal environment, so that various intestinal related disease biomarkers can be mined through human feces to detect the corresponding incidence, including colorectal cancer (Colorectal Cancer, CRC), inflammatory bowel disease, diabetes and obesity, etc.
[0064] However, in the existing fecal marker research, small aliquot of feces (for example, 200 mg or less) is usually collected, which is directly stirred and mixed for sample preparation and mass spectrometry analysis, however, this ignores the research value of fecal heterogeneity, and has certain limitations, for example, mixing the feces will cause many CRC related low abundance proteins distributed in the outer layer of the feces to be covered, and due to the different sampling positions of the feces, the CRC fecal markers screened out have large differences and poor repeatability.
[0065] Based on this, the application comprises a spatial mapping module and a spatial analysis module, the spatial analysis module is in communication connection with the spatial mapping module, wherein the spatial mapping module is used to acquire a target molecule type, a spatial omics data set and a spatial position data set of a target sample, and perform spatial mapping of the target sample at a molecular level according to the target molecule type, the spatial omics data set and the spatial position data set, so as to obtain molecular distribution data of the target sample; the spatial mapping module is further used to correspondingly draw a molecular distribution map of the target sample according to the molecular distribution data; the spatial analysis module is used to acquire pathological information of a target research disease and molecular distribution maps of a plurality of target samples, and determine a target analysis region corresponding to each target sample in each molecular distribution map according to the pathological information; the spatial analysis module is further used to perform differential analysis and screening on molecules in each target analysis region according to a statistical test method, so as to obtain differential molecules, and determine the differential molecules as candidate markers of the target research disease, thereby being capable of detecting more low-abundance markers that are hidden due to homogenization in the past, providing a rich potential biomarker database for the corresponding research disease, and promoting the development of personalized medical research.
[0066] The embodiment of the application provides a marker screening system and method based on spatial omics, equipment and medium thereof, and a storage medium of the system.
[0067] In a first aspect, a marker screening system based on spatial omics is described.
[0068] Referring to Figure 1 and Figure 2 , it is shown that Figure 1 is an architecture diagram of the marker screening system based on spatial omics provided by some embodiments of the application, Figure 2 is an operation flow diagram of the marker screening system provided by some embodiments of the application, the marker screening system 100 comprises a spatial mapping module 101, a spatial analysis module 102 and a sampling module 103, the spatial analysis module 102 is in communication connection with the spatial mapping module 101, and the sampling module 103 is in communication connection with the spatial mapping module 101.
[0069] The spatial mapping module 101 is used to acquire a target molecule type, a spatial omics data set and a spatial position data set of a target sample, and perform spatial mapping of the target sample at a molecular level according to the target molecule type, the spatial omics data set and the spatial position data set, so as to obtain molecular distribution data of the target sample.
[0070] The target molecule type is a molecule type that is the focus of research on the target disease in the sample, such as a recognized protein related to colorectal cancer (CRC), and the lipid molecule type includes a haptoglobin molecule type (HP), a transferrin type (TF), and a triglyceride lipid molecule type (TG).
[0071] In addition, the spatial omics dataset includes spatial omics data of all molecules in the target sample, and the spatial position dataset includes spatial position data of all molecules in the target sample.
[0072] In an optional embodiment, referring to Figure 3 as shown, Figure 3 is a flowchart of the spatial mapping module provided by another embodiment of the present application for performing spatial mapping of a target sample at the molecular level. The spatial mapping can include, but is not limited to, steps S301 to S304.
[0073] Step S301: According to the target molecule type, determine the target spatial omics data of the target molecule in the spatial omics dataset.
[0074] In an optional embodiment, when the haptoglobin molecule type is input into the spatial mapping module 101, the target spatial omics data of the haptoglobin molecule can be determined in the spatial omics dataset.
[0075] In an optional embodiment, when the transferrin molecule type is input into the spatial mapping module 101, the target spatial omics data of the transferrin molecule can be determined in the spatial omics dataset.
[0076] In an optional embodiment, when the triglyceride lipid molecule type is input into the spatial mapping module 101, the target spatial omics data of the triglyceride lipid molecule type can be determined in the spatial omics dataset.
[0077] Step S302: Normalize the target spatial omics dataset to obtain a normalized target spatial omics dataset.
[0078] In an optional embodiment, the target spatial omics data of the haptoglobin molecule is normalized to obtain the normalized target spatial omics data.
[0079] In an optional embodiment, the target spatial omics data of the transferrin molecule is normalized to obtain the normalized target spatial omics data.
[0080] In an optional embodiment, the target spatial omics data of the triglyceride lipid molecule is normalized to obtain the normalized target spatial omics data.
[0081] Step S303: According to the target molecule type, the target spatial position data of the target molecule is determined in the spatial position data set.
[0082] In an optional embodiment, the target spatial position data of the haptoglobin molecule is determined in the spatial position data set when the haptoglobin molecule type is input in the spatial mapping module 101.
[0083] In an optional embodiment, the target spatial position data of the transferrin molecule is determined in the spatial position data set when the transferrin molecule type is input in the spatial mapping module 101.
[0084] In an optional embodiment, the target spatial position data of the triglyceride lipid molecule is determined in the spatial position data set when the triglyceride lipid molecule type is input in the spatial mapping module 101.
[0085] Step S304: The normalized target spatial omics data is spatially correlated and mapped with the target spatial position data to obtain molecule distribution data.
[0086] In an optional embodiment, the normalized target spatial omics data of the haptoglobin molecule type is spatially correlated and mapped with the corresponding target spatial position data to obtain molecule distribution data of the haptoglobin molecule type.
[0087] In an optional embodiment, the normalized target spatial omics data of the transferrin molecule type is respectively spatially correlated and mapped with the corresponding target spatial position data to obtain molecule distribution data of the transferrin molecule type.
[0088] In an optional embodiment, the normalized target spatial omics data of the triglyceride lipid molecule type is respectively spatially correlated and mapped with the corresponding target spatial position data to obtain molecule distribution data of the triglyceride lipid molecule type.
[0089] Further, the spatial mapping module 101 is further used to draw a molecule distribution map of the target sample according to the molecule distribution data.
[0090] In an optional embodiment, the molecule distribution data of the haptoglobin molecule, the transferrin molecule, and the triglyceride lipid molecule is used to draw a molecule distribution map of the target sample, wherein the molecule distribution map includes the overall distribution of the haptoglobin molecule, the transferrin molecule, and the triglyceride lipid molecule.
[0091] Referring to Figure 4 as shown, Figure 4This involves using the spatial mapping module provided in other embodiments of this application to draw a spatial distribution map of HP, TF, and TG feces, from... Figure 4 It can be seen that stool samples collected from people with colorectal cancer (i.e. Figure 4 CRC samples) and fecal samples collected from healthy individuals (i.e. Figure 4 Compared with CT samples, it can be seen that HP molecules and TF molecules associated with colorectal cancer are mainly distributed in the outer ring of feces, proving that the heterogeneity of feces has spatial characteristics and significance.
[0092] Furthermore, the spatial analysis module 102 is used to acquire pathological information of the target disease and molecular distribution maps of multiple target samples, and to customize the target analysis area corresponding to each target sample in each molecular distribution map based on the pathological information.
[0093] Among them, multiple target samples can be four fecal samples collected from healthy people and three fecal samples collected from people with colorectal cancer.
[0094] In one alternative embodiment, refer to Figure 5 As shown, Figure 5 This is a flowchart of a spatial analysis module for determining a target analysis region provided in some embodiments of this application. The method for determining the target analysis region may include, but is not limited to, steps S501 to S503.
[0095] Step S501: Based on the pathological information, determine the pathological characteristics of the target disease.
[0096] In one alternative embodiment, when the target disease to be studied is colorectal cancer, the pathological features of colorectal cancer are determined based on pathological information such as the location, characteristics, and distribution of related protein molecules. These pathological features can be characterized by the fact that the outer layer of human feces is very close to the colonic microenvironment, which has a better diagnostic effect.
[0097] Step S502: Based on the pathological characteristics, determine the set of pathological characteristic distribution regions for each target sample in each molecular distribution map.
[0098] The set of pathological feature distribution areas includes multiple pathological feature distribution areas with different regional ranges.
[0099] In one alternative embodiment, based on the pathological characteristics of the outer layer of human feces being very close to the colonic microenvironment, the pathological characteristic distribution areas of different regions of each target sample are determined in each molecular distribution map.
[0100] Step S503: In each set of pathological feature distribution regions of each target sample, a pathological feature distribution region is selected as a target analysis region of the corresponding target sample, to determine the target analysis region corresponding to each target sample.
[0101] In an optional embodiment, in the set of pathological feature distribution regions corresponding to one fecal sample, a pathological feature distribution region covering the outermost layer is selected as the target analysis region of the fecal sample, and in the set of pathological feature distribution regions corresponding to another fecal sample, a pathological feature distribution region covering the outermost layer is selected as the target analysis region of the fecal sample.
[0102] In another optional embodiment, in the set of pathological feature distribution regions corresponding to one fecal sample, a pathological feature distribution region covering the entire region is selected as the target analysis region of the fecal sample, and in the set of pathological feature distribution regions corresponding to another fecal sample, a pathological feature distribution region covering the entire region is selected as the target analysis region of the fecal sample.
[0103] Further, the spatial analysis module 102 is also used to perform differential analysis and screening on the molecules in each target analysis region according to a statistical test method, to obtain differential molecules, and determine the differential molecules as candidate markers for the target research disease.
[0104] In an optional embodiment, referring to FIG. 6, Figure 6 Figure 6 FIG. 6 is a flowchart of a differential analysis method of the spatial mapping module according to some embodiments of the present application, which can include but is not limited to steps S601 to S603.
[0105] Step S601: The abundance values of the same molecule type in each pathological feature distribution region are summed up to obtain a set of total abundance of molecules in each pathological feature distribution region.
[0106] The set of total abundance of molecules includes the total abundance values of different molecule types.
[0107] In an optional embodiment, the abundance values of the same molecule type in each pathological feature distribution region covering the outermost layer are summed up to obtain a set of total abundance of molecules of the pathological feature distribution region covering the outermost layer,
[0108] In another optional embodiment, the abundance values of the same molecule type in the pathological feature distribution region covering the entire region are summed up to obtain a set of total abundance of molecules of the pathological feature distribution region covering the entire region.
[0109] Step S602: Normalizing the molecular sum abundance set of each pathological feature distribution region to obtain a normalized molecular sum abundance set of each pathological feature distribution region.
[0110] In an alternative embodiment, the molecular sum abundance set of the pathological feature distribution region covering the outermost layer is normalized to obtain a plurality of normalized molecular sum abundance sets of the pathological feature distribution region covering the outermost layer.
[0111] In an alternative embodiment, the molecular sum abundance set of the pathological feature distribution region covering the whole region is normalized to obtain a plurality of normalized molecular sum abundance sets of the pathological feature distribution region covering the whole region.
[0112] Step S603: Performing differential analysis and screening on each normalized molecular sum abundance set to obtain a set of molecular difference values between each pathological feature distribution region.
[0113] It should be noted that the statistical test method can be a fold change method, a t-test method, a wilcoxn test method, or a limma differential analysis method, which is not limited in the present application.
[0114] In an alternative embodiment, the limma differential analysis method is used to perform differential analysis and screening on each normalized molecular sum abundance set of the pathological feature distribution region covering the outermost layer to obtain a differential molecule between each pathological feature distribution region covering the outermost layer.
[0115] In another alternative embodiment, the limma differential analysis method is used to perform differential analysis and screening on the normalized molecular sum abundance set of the pathological feature distribution region covering the whole region to obtain a differential molecule between each pathological feature distribution region covering the whole region.
[0116] Referring to Figure 7 , Figure 7 is a schematic diagram of the result of calculating the differential molecules of the outermost layer and the whole region of the feces by the spatial analysis module according to some embodiments of the present application, from which Figure 7 It can be seen from the limma differential analysis result of part A in Figure 7As shown in Part B, the differentially expressed protein molecules screened from the outermost layer of feces also include those screened from the mixed feces from the entire region. Notably, in the results with adjust-p values < 0.05, the outermost fecal sample also screened out four potential fecal biomarkers that had not been previously reported in the literature: IGHG1 protein, ANXA1 protein, FGA protein, and FABP5 protein, which are respectively associated with cancer migration, immune regulation, blood coagulation, and fatty acid transport.
[0117] Furthermore, refer to Figure 8 As shown, Figure 8 This is a schematic diagram showing the correlation and co-localization characteristics of differentially expressed protein molecules provided in some embodiments of this application. Figure 8 Part A shows that the outermost fecal markers have a strong correlation (Spearman correlation analysis, adjust-p < 0.05, r > 0.7). From Figure 8 The differential protein distribution map in section B shows that three previously unreported potential CRC fecal biomarkers detected in the outermost fecal sample co-localize with HP, the most widely used clinical adjunctive diagnostic biomarker, and are mainly distributed in the outer layer of feces. Therefore, the fecal mixing procedure would mask these unevenly distributed low-abundance protein molecules, while the spatial analysis module of this application can detect more low-abundance biomarkers in the sample that were previously masked by homogenization.
[0118] Furthermore, the sampling module 103 is used to acquire the target sample, perform molecular extraction on the target sample, and obtain a sample molecular solution.
[0119] In one possible implementation, when the sample molecular solution is a protein molecular solution, refer to Figure 9 As shown, Figure 9 This is a flowchart of protein molecule extraction from a target sample provided in some embodiments of this application. The protein molecule extraction method may include, but is not limited to, steps S901 to S905.
[0120] Step S901: Freeze and slice the target sample to obtain sample slices.
[0121] In one possible embodiment, target samples were obtained from 10 subjects, and the target samples were formed feces (7 control subjects and 3 CRC subjects). The collected formed feces were frozen at -80°C and then sliced at approximately the middle of the feces to obtain frozen fecal cross-sections with a thickness of approximately 2-3 mm.
[0122] Step S902: Use a microscaffold to divide the sample slice into several sample microparticles that retain spatial location information.
[0123] Specifically, by using a microstent containing 30x30 micropores, each with a cross-sectional size of 1.2mmx1.2mm, the fecal cross-section is cut into several fecal particles that retain positional information, obtaining sample particles.
[0124] Step S903: According to the pressure cycling technique, the sample section is subjected to protein lysis treatment, obtaining a protein lysis solution.
[0125] In one possible embodiment, 40ul of protein lysis solution and a portion of the fecal particles described above are added to a single PCT-Microtube microtube, which is tightly capped with a PCT-MicroPestle plug, and then placed in a PCT instrument. Then, according to the set reaction parameters, the sample is subjected to a cyclic pressure from 1 atmosphere to about 3000 times atmospheric pressure to lyse and extract proteins. Finally, after the cycle is completed, the fecal protein lysis solution is aspirated to obtain the protein lysis solution of the fecal sample.
[0126] Step S904: According to the Automated SP3 technique, the protein lysis solution is subjected to high-throughput, automated enzymatic purification treatment to obtain a protein peptide solution.
[0127] In one possible embodiment, according to the automated process of the SP3 automated liquid handling system (Bravo), 2ul of the protein lysis solution is diluted 20 times and heated for 15 minutes before being added to a 96-well plate. Then, according to the measured CRC fecal protein concentration range, the Bravo instrument automatically adds 5ul of mixed magnetic beads and 45ul of 100% ACN to each well, and incubates for 18 minutes with shaking. After the protein and magnetic beads are fully combined, the deep well plate is placed on a magnetic stand for 5 minutes to adsorb and precipitate, and the supernatant is removed. Then, the plate is washed twice with 200ul / well of 80% ethanol and once with 171.5ul / well of 100% acetonitrile. After removing the washing solution, the magnetic beads are resuspended in 35ul / well of enzyme buffer and 10ul / well of trypsin solution, and the plate is sealed with a PCR film and placed in a shaker for overnight enzymatic digestion. The enzymatic digestion conditions are 37°C constant temperature and 220rpm enzymatic digestion and purification treatment, obtaining a protein peptide solution.
[0128] Among them, the mixed magnetic beads described above are 50ul Sera-Mag Speed magnetic beads A and 50ul Fisher Scientific magnetic beads B mixed and resuspended in 1ml ddH2O to achieve a final working concentration of 5ug / ul.
[0129] The trypsin solution described above has the following composition: 20ug of trypsin (Promega) dissolved in 4ml of enzyme buffer.
[0130] The enzyme buffer component is 0.65ml pure water, 0.25ml Tris, and 0.1ml ACN.
[0131] Step S905: Transfer the protein peptide segment solution to the Evotips to obtain a protein molecule solution.
[0132] In one possible embodiment, according to the automatic process of the above-mentioned Bravo instrument, 6ul of 5% trifluoroacetic acid is automatically added to the protein peptide segment solution after overnight enzymolysis, the enzymolysis is quickly acidified and stopped, then the magnetic beads are fixed on the magnetic stand, the supernatant is transferred to the activated Evotips, the protein peptide segment solution is recovered and loaded, and a protein molecule solution is obtained.
[0133] The activation and loading of the peptide segment of the Evotips are as follows: Step A: 20ul of buffer B is added to each Evotip, and centrifuged at 700rcf for 1 minute; Step B: Evotips are activated by soaking in isopropanol until all the tips are off-white; Step C: 20ul of buffer A is added to each Evotip, and centrifuged at 700rcf for 1 minute again; Step D: After the enzyme-stopped peptide segment solution is transferred to the Evotips, it is centrifuged at 800rcf for 1 minute; Step E: 50ul of buffer A is added to each Evotip, and centrifuged at 800rcf for 1 minute; Step F: 100ul of buffer A is added to each Evotip and instant centrifugation is performed to move the liquid down to the membrane, keeping the Evotips wet.
[0134] It should be noted that the buffer A component is 99.9% ddH2O, 0.1% FA, and the buffer B component is 80% ACN, 20% water, and 0.1% FA.
[0135] The buffer A component is 99.9% ddH2O, 0.1% FA, and the buffer B component is 80% ACN, 20% water, and 0.1% FA.
[0136] In one possible implementation, when the sample molecule solution is a lipid molecule solution, referring to Figure 10 As shown in Figure 10 is a flowchart provided by some embodiments of the present application for lipid molecule extraction of a target sample. The lipid molecule extraction method can include but is not limited to steps S1001 to S1003.
[0137] Step S1001: Freeze and slice the target sample to obtain a sample slice.
[0138] In one possible embodiment, 10 target samples of subjects are obtained, and the target samples are shaped feces (7 control samples and 3 CRC samples), the collected shaped feces are cut into about 2-3 mm thick frozen feces cross-sections after being frozen at -80°C.
[0139] Step S1002: The sample slice is divided into a plurality of sample microparticles that retain spatial position information by using a micro support.
[0140] Specifically, the feces cross-section is cut into a plurality of feces particles that retain spatial position information by using a micro support containing 30x30 micropores, each micropore having a cross-sectional size of 1.2mmx1.2mm, to obtain the sample microparticles.
[0141] Step S1003: The sample slice is subjected to lipid dissolution treatment by using a lipid extraction solution, and the supernatant is taken to obtain a lipid molecule solution after concentration and resuspension.
[0142] In one possible embodiment, chloroform-methanol mixed solution (chloroform:methanol = 3:1) is added to the feces particles, and after oscillation for 2 min, centrifugation is performed at 4°C and 16000g for 10 min to take the supernatant, which is concentrated and stored at -80°C to obtain the lipid molecule solution.
[0143] It should be noted that the lipid molecule solution needs to be resuspended with 50ul chloroform-methanol mixed solution (chloroform:methanol = 5:13) before mass spectrometry detection.
[0144] Further, the sampling module 103 is further configured to perform mass spectrometry detection on the sample molecule solution to obtain spatial omics data and spatial position data set of the target sample.
[0145] In one possible embodiment, when the sample molecule solution is a protein molecule solution, Evosepone liquid chromatograph and Orbitrap Fusion Lumos Tribrid mass spectrometer coupled with FAIMS source are used to perform mass spectrometry detection on the protein molecule solution in Evotips to obtain spatial omics data and spatial position data set corresponding to the protein molecules, wherein, in the mass spectrometry detection process, a positive mode is adopted, and the spray voltage is 2200V. The analysis column is Acclaim PepMap100 C18 column (75mmx250mm, 2mm, Thermo Scientific), Evosep one SPD60 preprogrammed gradient is used, the flow rate is 1ul / min, the corresponding mobile phase A is 0.1% formic acid, and the corresponding mobile phase B is 0.1% formic acid, 80% acetonitrile and 20% pure water.
[0146] In an alternative embodiment, when the sample molecule solution is a lipid molecule solution, a Thermo Scientific Dionex UltiMate 3000 rapid separation liquid chromatograph coupled with an Orbitrap Exploris 480 mass spectrometer is used to analyze the reconstituted lipid molecule solution for mass spectrometry detection to obtain the spatial omics data and spatial position data set corresponding to the lipid molecule solution. During the mass spectrometry detection, the spray voltage is 3500V, the analysis column is C18, the flow rate is 0.2ml / min, the corresponding mobile phase A is 60% acetonitrile, 10mM ammonium acetate, 0.1% acetic acid and 40% pure water, and the corresponding mobile phase B is 90% isopropanol, 10% acetonitrile, 10mM ammonium acetate and 0.1% acetic acid.
[0147] In a second aspect, the embodiments of the present application also provide a spatial omics-based marker screening method. Referring to Figure 11 , Figure 11 A flowchart of the spatial omics-based marker screening method provided by some embodiments of the present application is shown. The spatial omics-based marker screening method can include, but is not limited to, steps S1101 to S1104.
[0148] Step S1101: Obtain the target molecule type, the spatial omics data set and the spatial position data set of a plurality of target samples, and perform spatial mapping at the molecular level for each target sample according to the target molecule type, each spatial omics data set and each spatial position data set to obtain the molecular distribution data of each target sample.
[0149] Step S1102: Draw a molecular distribution map of each target sample according to each molecular distribution data.
[0150] Step S1103: Obtain the pathological information of the target research disease, and according to the pathological information, determine the target analysis region corresponding to each target sample in each molecular distribution map;
[0151] Step S1104: Perform differential analysis and screening on the molecules in each target analysis region according to a statistical test method to obtain differential molecules, and determine the differential molecules as markers of the target research disease.
[0152] The marker screening method and the marker screening system 100 described above are based on the same inventive concept. The above process describes the marker screening method of the embodiments of the present application. The target molecule type, the spatial omics data set and the spatial position data set of the target sample are obtained, and the spatial mapping of the target sample at the molecular level is performed according to the target molecule type, the spatial omics data set and the spatial position data set, so as to obtain the molecular distribution data of the target sample. Then, the molecular distribution map of the target sample is drawn according to the molecular distribution data. Next, the pathological information of the target research disease and the molecular distribution maps of the target samples are obtained, and the target analysis area corresponding to each target sample is determined in each molecular distribution map according to the pathological information. Finally, the molecules in each target analysis area are subjected to differential analysis and screening according to a statistical test method, so as to obtain the differential molecules, and the differential molecules are determined as the markers of the target research disease. Thus, more low-abundance markers that are hidden due to homogenization in the past can be detected in the sample, a rich potential biomarker database for the corresponding research disease is provided, and the development of personalized medical research is promoted.
[0153] The embodiments of the present application also provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. The processor implements the marker screening method described above when executing the computer program. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a vehicle-mounted computer, etc.
[0154] Please refer to Figure 12 , Figure 12 is a hardware structure schematic diagram of an electronic device provided by some embodiments of the present application. The electronic device includes:
[0155] The processor 1201 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc. The processor 1201 is used to execute related programs to implement the QSPI serial port switching method and / or the cache data reading method provided by the embodiments of the present application.
[0156] The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1202 and are called and executed by the processor 1201 to implement the QSPI serial port switching method and / or the cache data reading method provided by the embodiments of the present application.
[0157] The input / output interface 1203 is configured to realize information input and output.
[0158] The communication interface 1204 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).
[0159] The bus 1205 is configured to transmit information between various components (for example, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204) of the device.
[0160] The processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are connected to each other through the bus 1205 to realize the communication connection between the device.
[0161] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the marker screening method provided by the embodiments of the present application.
[0162] The memory is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0163] The embodiments described in the specification are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0164] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0165] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0166] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0167] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and above-described drawings of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0168] It should be understood that, in the application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases of only A, only B, and A and B existing at the same time, wherein A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, "a and b", "a and c", "b and c", or "a and b and c", wherein a, b, and c can be single or multiple.
[0169] In several embodiments provided in the application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are only illustrative, for example, the division of the above units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0170] The units described above as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0171] In addition, the functional units in each embodiment of the application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0172] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer accessible storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program storage media.
[0173] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not limited to the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. A biomarker screening system based on spatial omics, characterized in that, include: A spatial mapping module and a spatial analysis module, wherein the spatial analysis module is communicatively connected to the spatial mapping module; The spatial mapping module is used to acquire the target molecule type, the spatial omics dataset and the spatial location dataset of the target sample, and perform molecular-level spatial mapping on the target sample based on the target molecule type, the spatial omics dataset and the spatial location dataset to obtain the molecular distribution data of the target sample. The spatial mapping module is also used to draw a molecular distribution map of the target sample based on the molecular distribution data. The spatial analysis module is used to acquire pathological information of the target disease under study and molecular distribution maps of multiple target samples, and to customize the target analysis region corresponding to each target sample in each molecular distribution map based on the pathological information. The spatial analysis module is also used to perform differential analysis and screening of molecules in each of the target analysis regions according to statistical testing methods, obtain differentially expressed molecules, and identify the differentially expressed molecules as candidate biomarkers for the target disease under study. Specifically, based on the pathological information, the target analysis region corresponding to each target sample is custom-defined in each of the molecular distribution maps, including: Based on the pathological information, the pathological characteristics of the target disease under study are determined; Based on the pathological characteristics, the pathological characteristic distribution region set of each target sample is determined in each of the molecular distribution maps, and the pathological characteristic distribution region set includes multiple pathological characteristic distribution regions of different ranges. In each of the target samples, a pathological feature distribution region is selected as the target analysis region for that target sample, thereby determining the target analysis region for each target sample.
2. The marker screening system according to claim 1, characterized in that, The differential analysis and screening of molecules in each of the pathological feature distribution regions are performed according to statistical testing methods to obtain differentially expressed molecules, including: The abundance values of molecules of the same molecular type in each of the pathological feature distribution regions are summed to obtain the total abundance set of molecules in each of the pathological feature distribution regions. The molecular sum abundance sets of each of the pathological feature distribution regions are normalized to obtain the normalized molecular sum abundance sets. Differential analysis and screening were performed on the normalized molecular sum abundance sets to obtain the differential molecules among the distribution regions of the pathological features.
3. The marker screening system according to claim 1, characterized in that, The step of performing molecular-level spatial mapping on the target sample based on the target molecule type, the spatial omics dataset, and the spatial location dataset to obtain the molecular distribution data of the target sample includes: Based on the target molecule type, determine the target spatial omics data of the target molecule in the spatial omics dataset; The target spatial omics data are normalized to obtain normalized target spatial omics data; Based on the target molecule type, determine the target spatial location data of the target molecule in the spatial location dataset; The molecular distribution data is obtained by spatially associating the normalized target spatial omics data with the target spatial location data.
4. The marker screening system according to claim 1, characterized in that, The system further includes a sampling module, which is communicatively connected to the spatial mapping module, wherein: The sampling module is used to acquire the target sample, perform molecular extraction on the target sample, and obtain a sample molecular solution; The sampling module is also used to perform mass spectrometry detection on the sample molecular solution to obtain the spatial omics data and spatial location dataset of the target sample.
5. The marker screening system according to claim 4, characterized in that, When the sample molecular solution is a protein molecular solution, the step of extracting molecules from the target sample to obtain sample molecules includes: The target sample was frozen and sliced to obtain sample slices; The sample slice is divided into several sample microparticles that retain spatial location information using a microscaffold. The sample particles were subjected to protein lysis treatment using pressure cycling technology to obtain a protein lysis buffer; The protein lysis buffer was subjected to high-throughput, automated enzymatic hydrolysis and purification using Automated SP3 technology to obtain a protein peptide solution. The protein peptide solution was transferred to Evotips to obtain the protein molecule solution.
6. The marker screening system according to claim 4, characterized in that, When the sample molecular solution is a lipid molecular solution, the step of extracting molecules from the target sample to obtain sample molecules includes: The target sample was frozen and sliced to obtain sample slices; The sample slice is divided into several sample microparticles that retain spatial location information using a microscaffold. The sample particles were dissolved using a lipid extraction solution, and the supernatant was collected, concentrated, and reconstituted to obtain the lipid molecular solution.
7. A biomarker screening method based on spatial omics, characterized in that, include: Acquire the target molecule type, spatial omics datasets and spatial location datasets of multiple target samples, and perform molecular-level spatial mapping on each target sample according to the target molecule type, each of the spatial omics datasets and each of the spatial location datasets to obtain the molecular distribution data of each target sample; Molecular distribution maps of each target sample are drawn based on the molecular distribution data. Obtain pathological information of the target disease under study, and based on the pathological information, define the target analysis region corresponding to each target sample in each molecular distribution map in a custom manner; According to statistical testing methods, differential analysis and screening are performed on molecules in each of the target analysis regions to obtain differentially expressed molecules, and the differentially expressed molecules are identified as biomarkers for the target disease under study. Specifically, based on the pathological information, the target analysis region corresponding to each target sample is custom-defined in each of the molecular distribution maps, including: Based on the pathological information, the pathological characteristics of the target disease under study are determined; Based on the pathological characteristics, the pathological characteristic distribution region set of each target sample is determined in each of the molecular distribution maps, and the pathological characteristic distribution region set includes multiple pathological characteristic distribution regions of different ranges. In each of the target samples, a pathological feature distribution region is selected as the target analysis region for that target sample, thereby determining the target analysis region for each target sample.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the marker screening method as described in claim 7.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the marker screening method as described in claim 7.
Citation Information
Patent Citations
Molecular marker detection product based on PDX / PDTX tumor living tissue biological sample and database and preparation method thereof
CN116593262A
Rapid tumor tissue identification method based on fingerprint spectrogram of lipids on tissue surface
WO2020259187A1