Marker screening system, method and equipment based on spatial omics and medium thereof
Through a spatial omics-based marker screening system, spatial omics analysis of fecal samples was solved, and the problem of low abundance marker masking caused by ignoring fecal heterogeneity in the prior art was solved, and more comprehensive and accurate marker screening was achieved, providing a rich biomarker database for personalized medicine.
Patent Information
- Application Number
- CN202411892203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In the existing fecal marker research, samples obtained by direct stirring and mixing ignore the heterogeneity of feces, resulting in many low-abundance CRC-related proteins distributed in the outer layer of feces being masked, and different sampling positions lead to large differences in markers and poor repeatability.
The spatial oscillary screening system is adopted to perform spatial oscillary analysis of the samples through the spatial mapping module and the spatial analysis module, obtain the target molecular type, spatial oscillary data set and spatial location data set, perform spatial mapping at the molecular level, draw the molecular distribution map, and determine the target analysis area based on pathological information, and screen differential molecules through statistical testing methods.
It can detect more low-abundance markers masked by homogenization in the samples, providing a rich database of potential biomarkers for the corresponding disease research, and promoting the development of personalized medical research.
Smart Images

Figure CN119943162A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the biological field, and in particular to a marker screening system, method, device and medium based on spatial omics. Background Art
[0002] With the development of precision medicine, people will collect corresponding metabolites from the human body for testing according to different diseases, and then use different markers in the samples to check the incidence of different diseases. For example, human feces is an easily accessible and low-cost biological sample that contains a variety of host proteins, microorganisms, food and other biological molecules. When it passes down the gastrointestinal tract, it will be mixed with various substances in the intestinal environment and secretions. Therefore, human feces can be used to mine biomarkers of various intestinal-related diseases with high throughput to detect the corresponding incidence, including colorectal cancer (CRC), inflammatory bowel disease, diabetes and obesity.
[0003] In existing fecal marker studies, small aliquots of feces (e.g., 200 mg or less) are usually collected and directly mixed for sample preparation and mass spectrometry analysis, but this ignores the research value of fecal heterogeneity. Human feces do not represent a single colon environment, but the entire gastrointestinal environment. From the physiological formation and excretion process of feces, we can know that feces are gradually formed along the entire digestive tract. There are many small intestine-related substances distributed inside the feces, and the outer layer of feces is mainly related to the colon microenvironment.
[0004] Therefore, there are limitations to sampling from different spatial locations in the stool and homogenizing to obtain data representing the entire stool. For example, mixing the stool will cause many CRC-related low-abundance proteins distributed in the outer layer of the stool to be masked, and due to the different sampling locations of the stool, the CRC stool markers screened out vary greatly and have poor reproducibility. Summary of the invention
[0005] The main purpose of the embodiments of the present application is to propose a marker screening method, device and medium, system, device and medium based on spatial omics, and perform spatial omics analysis on samples through a spatial mapping module and a spatial analysis module, so as to detect more low-abundance markers in the samples that were previously masked by homogenization, provide a rich potential biomarker database for the corresponding diseases being studied, and promote the development of personalized medical research.
[0006] To achieve the above objectives, the first aspect of the embodiments of the present application proposes a marker screening system based on spatial omics, comprising: A spatial mapping module and a spatial analysis module, wherein the spatial analysis module is in communication connection with the spatial mapping module; The spatial mapping module is used to obtain the target molecule type, the spatial omics data set and the spatial position data set of the target sample, and perform molecular-level spatial mapping of the target sample according to the target molecule type, the spatial omics data set and the spatial position data set to obtain the molecular distribution data of the target sample; The spatial mapping module is also used to draw a molecular distribution map of the target sample according to the molecular distribution data; The spatial analysis module is used to obtain the pathological information of the target disease and the molecular distribution maps of multiple target samples, and to define the target analysis area corresponding to each target sample in each molecular distribution map according to the pathological information; The spatial analysis module is also used to perform differential analysis and screening on the molecules in each target analysis area according to statistical test methods, obtain differential molecules, and determine the differential molecules as candidate markers for the target disease.
[0007] Furthermore, in some embodiments, the molecules in each pathological characteristic distribution area are differentially analyzed and screened according to a statistical test method to obtain differential molecules, including: The abundance values of molecules of the same molecular type in each pathological characteristic distribution area are summed up respectively to obtain the total abundance set of molecules in each pathological characteristic distribution area; Normalizing the molecular sum abundance sets of each pathological characteristic distribution area to obtain each normalized molecular sum abundance set; The normalized total abundance sets of molecules were differentially analyzed and screened to obtain differential molecules between the distribution areas of various pathological characteristics.
[0008] Furthermore, in some embodiments, determining the target analysis region corresponding to each target sample in each molecular distribution map according to the pathological information includes: Based on the pathological information, determine the pathological characteristics of the target disease; According to the pathological characteristics, a pathological characteristic distribution area set of each target sample is determined in each molecular distribution map, and the pathological characteristic distribution area set includes pathological characteristic distribution areas of multiple different regional ranges; A pathological characteristic distribution region is customarily selected from the pathological characteristic distribution region set of each target sample as the target analysis region corresponding to the target sample, so as to determine the target analysis region corresponding to each target sample.
[0009] Furthermore, in some embodiments, according to the target molecule type, the spatial omics data set and the spatial position data set, the target sample is spatially mapped at the molecular level to obtain the molecular distribution data of the target sample, including: According to the target molecule type, the target spatial omics data of the target molecule is determined in the spatial omics data set; Normalizing the target space omics data set to obtain a normalized target space omics data set; According to the target molecule type, determining the target spatial position data of the target molecule in the spatial position data set; The normalized target spatial omics data and target spatial position data are spatially correlated and mapped to obtain molecular distribution data.
[0010] Furthermore, in some embodiments, the marker screening system further comprises a sampling module, which is communicatively connected to the spatial mapping module, wherein: The sampling module is used to obtain the target sample, perform molecular extraction on the target sample, and obtain a sample molecular solution; The sampling module is also used to perform mass spectrometry detection on the sample molecular solution to obtain the spatial omics data and spatial position data set of the target sample.
[0011] Furthermore, in some embodiments, when the sample molecule solution is a protein molecule solution, performing molecular extraction on the target sample to obtain the sample molecules includes: Freezing and slicing the target sample to obtain sample slices; Using a micro-support to divide the sample slice into a plurality of sample particles with certain spatial position information; According to the pressure cycle technology, the sample particles are subjected to protein lysis treatment to obtain a protein lysis solution; According to the Automated SP3 technology, the protein lysate is subjected to high-throughput, automated enzymatic purification to obtain a protein peptide solution; Transfer the protein peptide solution to Evotips to obtain a protein molecule solution.
[0012] Furthermore, in some embodiments, when the sample molecule solution is a lipid molecule solution, performing molecular extraction on the target sample to obtain the sample molecules includes: Freezing and slicing the target sample to obtain sample slices; Using a micro-support to divide the sample slice into a number of sample particles with preserved spatial position information; The sample particles are subjected to lipid dissolution treatment using a lipid extract and the supernatant is taken and concentrated and redissolved to obtain a lipid molecule solution.
[0013] To achieve the above purpose, the second aspect of the embodiment of the present application proposes a marker screening method based on spatial omics, comprising: Obtaining target molecule types, spatial omics data sets and spatial position data sets of multiple target samples, and performing molecular-level spatial mapping on each target sample according to the target molecule types, each spatial omics data set and each spatial position data set to obtain molecular distribution data of each target sample; Draw a molecular distribution map of each target sample according to each molecular distribution data; Obtain pathological information of the target disease, and according to the pathological information, define the target analysis area corresponding to each target sample in each molecular distribution map in a customized manner; According to the statistical test method, the molecules in each target analysis area are differentially analyzed and screened to obtain differential molecules, and the differential molecules are determined as markers for the target research disease.
[0014] To achieve the above-mentioned purpose, the third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the marker screening method of the above-mentioned second aspect embodiment is implemented.
[0015] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the marker screening method of the above-mentioned second aspect embodiment is implemented.
[0016] The embodiments of the present application have the following beneficial effects: the present application uses a spatial mapping module and a spatial analysis module, and the spatial analysis module is communicatively connected with the spatial mapping module, wherein the spatial mapping module is used to obtain the target molecule type, the spatial omics data set and the spatial position data set of the target sample, and according to the target molecule type, the spatial omics data set and the spatial position data set, perform spatial mapping of the target sample at the molecular level to obtain the molecular distribution data of the target sample; the spatial mapping module is also used to draw a molecular distribution map of the target sample according to the molecular distribution data; the spatial analysis module is used to obtain the pathological information of the target research disease and the molecular distribution maps of multiple target samples, and according to the pathological information, customize the target analysis area corresponding to each target sample in each molecular distribution map; the spatial analysis module is also used to perform differential analysis and screening on the molecules in each target analysis area according to a statistical test method to obtain differential molecules, and determine the differential molecules as candidate markers for the target research disease, thereby being able to detect more low-abundance markers that were previously masked by homogenization in the sample, providing a rich potential biomarker database for the corresponding research disease, and promoting the development of personalized medical research. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1is an architectural diagram of a marker screening system based on spatial omics provided in some embodiments of the present application; Figure 2 is an operation flow chart of a marker screening system provided in some embodiments of the present application; Figure 3 is a flow chart of a spatial mapping module provided in other embodiments of the present application performing molecular-level spatial mapping on a target sample; Figure 4 The spatial mapping module provided in some other embodiments of the present application is used to draw the fecal spatial distribution map of HP, TF and TG; Figure 5 is a flow chart of determining a target analysis area by a spatial analysis module provided in some embodiments of the present application; Figure 6 is a flowchart of a spatial mapping module performing differential analysis on molecules according to a statistical test method, provided in some embodiments of the present application; Figure 7 This is a schematic diagram of the results of calculating the difference molecules between the outermost layer and the entire area of feces by the spatial analysis module provided in some embodiments of the present application; Figure 8 It is a result schematic diagram of the correlation and co-localization characteristics of differential protein molecules provided in some embodiments of the present application; Fig. 9 is a flow chart of protein molecule extraction of a target sample provided by some embodiments of the present application; Fig.10 is a flow chart of extracting lipid molecules from a target sample provided by some embodiments of the present application; Fig.11 A flow chart of a marker screening method based on spatial omics provided in some embodiments of the present application; Fig.12 It is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] In the description of the present application, it should be understood that descriptions involving orientation, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0020] It should also be noted that in the description of this application, "several" means more than one, "more" means more than two, "greater than", "less than", "exceed", etc. are understood to exclude the number itself, and "above", "below", "within", etc. are understood to include the number itself. If there is a description of "first" or "second", it is only used to distinguish the technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0022] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0023] With the development of precision medicine, people will collect corresponding metabolites from the human body for testing according to different diseases, and then use different markers in the samples to check the incidence of different diseases. For example, human feces is an easily accessible and low-cost biological sample that contains a variety of host proteins, microorganisms, food and other biological molecules. When it passes down the gastrointestinal tract, it will be mixed with various substances in the intestinal environment and secretions. Therefore, human feces can be used to mine biomarkers of various intestinal-related diseases with high throughput to detect the corresponding incidence, including colorectal cancer (CRC), inflammatory bowel disease, diabetes and obesity.
[0024] In existing studies on fecal markers, small aliquots of feces (e.g., 200 mg or less) are usually collected and directly stirred for sample preparation and mass spectrometry analysis. However, this ignores the research value of fecal heterogeneity and has certain limitations. For example, mixing feces will cause many CRC-related low-abundance proteins distributed in the outer layer of feces to be masked, and due to different sampling locations of feces, the screened CRC fecal markers vary greatly and have poor repeatability.
[0025] Based on this, the present application uses a spatial mapping module and a spatial analysis module, and the spatial analysis module is communicatively connected with the spatial mapping module, wherein the spatial mapping module is used to obtain the target molecule type, the spatial omics data set and the spatial position data set of the target sample, and according to the target molecule type, the spatial omics data set and the spatial position data set, perform spatial mapping of the target sample at the molecular level to obtain the molecular distribution data of the target sample; the spatial mapping module is also used to draw a molecular distribution map of the target sample according to the molecular distribution data; the spatial analysis module is used to obtain the pathological information of the target research disease and the molecular distribution maps of multiple target samples, and according to the pathological information, customize the target analysis area corresponding to each target sample in each molecular distribution map; the spatial analysis module is also used to perform differential analysis and screening on the molecules in each target analysis area according to the statistical test method, obtain differential molecules, and determine the differential molecules as candidate markers for the target research disease, so as to detect more low-abundance markers that were previously masked by homogenization in the sample, provide a rich potential biomarker database for the corresponding research disease, and promote the development of personalized medical research.
[0026] The embodiments of the present application provide a marker screening system, method, device and medium thereof, system and storage medium based on spatial omics, which are specifically described through the following embodiments.
[0027] In the first aspect, a marker screening system based on spatial omics in an embodiment of the present application is first described.
[0028] Reference Figure 1 and Figure 2 As shown, Figure 1 is an architectural diagram of a marker screening system based on spatial omics provided in some embodiments of the present application, Figure 2 It is an operation flow chart of a marker screening system provided in some embodiments of the present application. The marker screening system 100 includes a spatial mapping module 101, a spatial analysis module 102, and a sampling module 103. The spatial analysis module 102 is communicatively connected to the spatial mapping module 101, and the sampling module 103 is communicatively connected to the spatial mapping module 101.
[0029] Among them, the spatial mapping module 101 is used to obtain the target molecule type, the spatial omics dataset and the spatial position dataset of the target sample, and perform molecular-level spatial mapping of the target sample according to the target molecule type, the spatial omics dataset and the spatial position dataset to obtain the molecular distribution data of the target sample.
[0030] Among them, the target molecule type is the molecule type that is focused on in the sample for the target research disease, such as the recognized related proteins of colorectal cancer (CRC), and the lipid molecule types include haptoglobin molecule type (haptoglobin, HP), transferrin type (Transferrin, TF), and triglyceride lipid molecule type (Triglyceride, TG).
[0031] Furthermore, the spatial omics data set includes the spatial omics data of all molecules in the target sample, and the spatial position data set includes the spatial position data of all molecules in the target sample.
[0032] In an optional embodiment, referring to Figure 3 As shown, Figure 3 This is a flow chart of a spatial mapping module provided in other embodiments of the present application performing molecular-level spatial mapping on a target sample. The spatial mapping may include but is not limited to steps S301 to S304.
[0033] Step S301: Determine target spatial omics data of the target molecule in the spatial omics data set according to the target molecule type.
[0034] In an optional embodiment, the type of haptoglobin molecule is input into the spatial mapping module 101, and the target spatial omics data of the haptoglobin molecule can be determined in the spatial omics data set.
[0035] In an optional embodiment, the type of transferrin molecule is input into the spatial mapping module 101, and the target spatial omics data of the transferrin molecule can be determined in the spatial omics data set.
[0036] In an optional embodiment, the triglyceride lipid molecule type is input into the spatial mapping module 101, and target spatial omics data of the triglyceride lipid molecule type can be determined in the spatial omics data set.
[0037] Step S302: normalizing the target space omics data set to obtain a normalized target space omics data set.
[0038] In an optional embodiment, the target space omics data of the haptoglobin molecule is normalized to obtain normalized target space omics data.
[0039] In an optional embodiment, the target space omics data of the transferrin molecule is normalized to obtain the normalized target space omics data.
[0040] In an optional embodiment, the target space omics data of triglyceride lipid molecules are normalized to obtain normalized target space omics data.
[0041] Step S303: Determine the target spatial position data of the target molecule in the spatial position data set according to the target molecule type.
[0042] In an optional embodiment, the type of haptoglobin molecule is input into the spatial mapping module 101, and the target spatial position data of the haptoglobin molecule can be determined in the spatial position data set.
[0043] In an optional embodiment, the type of transferrin molecule is input into the spatial mapping module 101, and the target spatial position data of the transferrin molecule can be determined in the spatial position data set.
[0044] In an optional embodiment, the triglyceride lipid molecule type is input into the spatial mapping module 101, and the target spatial position data of the thyroglobulin molecule can be determined in the spatial position data set.
[0045] Step S304: spatial correlation mapping is performed on the normalized target spatial omics data and the target spatial position data to obtain molecular distribution data.
[0046] In an optional embodiment, the normalized target spatial omics data of the haptoglobin molecule type and the corresponding target spatial position data are spatially correlated and mapped to obtain the molecular distribution data of the haptoglobin molecule type.
[0047] In an optional embodiment, the normalized target spatial omics data of the transferrin molecule type and the corresponding target spatial position data are spatially correlated and mapped to obtain the molecular distribution data of the transferrin molecule type.
[0048] In an optional embodiment, the normalized target spatial omics data of the triglyceride lipid molecule type and the corresponding target spatial position data are spatially correlated and mapped to obtain molecular distribution data of the triglyceride lipid molecule type.
[0049] Furthermore, the spatial mapping module 101 is also used to draw a molecular distribution map of the target sample according to the molecular distribution data.
[0050] In an optional embodiment, the molecular distribution data of haptoglobin molecules, transferrin molecules, and triglyceride lipid molecules correspond to the molecular distribution map of the target sample, wherein the molecular distribution map includes the overall distribution of haptoglobin molecules, transferrin molecules, and triglyceride lipid molecules.
[0051] Reference Figure 4 As shown, Figure 4The spatial distribution map of HP, TF and TG in feces is drawn by using the spatial mapping module provided in other embodiments of the present application. Figure 4 It can be seen that stool samples collected from people with colorectal cancer (i.e. Figure 4 CRC samples from healthy individuals (i.e. Figure 4 Compared with the CT samples in the feces, it can be seen that the HP molecules and TF molecules related to colorectal cancer are mainly distributed in the outer circle of feces, proving that the heterogeneity of feces has spatial characteristics and significance.
[0052] Furthermore, the spatial analysis module 102 is used to obtain pathological information of the target disease and molecular distribution maps of multiple target samples, and to user-defined target analysis regions corresponding to each target sample in each molecular distribution map according to the pathological information.
[0053] The multiple target samples may be 4 stool samples collected from healthy people and 3 stool samples collected from people suffering from colorectal cancer.
[0054] In an optional embodiment, referring to Figure 5 As shown, Figure 5 This is a flowchart of a spatial analysis module determining a target analysis area provided by some embodiments of the present application. The method for determining a target analysis area may include but is not limited to steps S501 to S503.
[0055] Step S501: Determine the pathological characteristics of the target disease according to the pathological information.
[0056] In an optional embodiment, when the target research disease is colorectal cancer, the pathological characteristics of colorectal cancer are determined based on the pathological information of colorectal cancer, such as the site of onset, characteristics of onset, and distribution of related protein molecules. The pathological characteristics can be characterized as the outer layer of human feces being very close to the colon microenvironment, which has a better diagnostic effect.
[0057] Step S502: According to the pathological characteristics, the pathological characteristic distribution area set of each target sample is determined in each molecular distribution map.
[0058] The pathological characteristic distribution area set includes a plurality of pathological characteristic distribution areas in different area ranges.
[0059] In an optional embodiment, based on the pathological feature characterized in that the outer layer of human feces is very close to the colon microenvironment, the pathological feature distribution areas of different regional ranges of each target sample are determined in each molecular distribution map.
[0060] Step S503: Customize and select a pathological characteristic distribution region from the pathological characteristic distribution region set of each target sample as the target analysis region corresponding to the target sample, so as to determine the target analysis region corresponding to each target sample.
[0061] In an optional embodiment, in a set of pathological characteristic distribution areas corresponding to a stool sample, the pathological characteristic distribution area covering the outermost layer is custom-selected as the target analysis area of the stool sample, and then, in a set of pathological characteristic distribution areas corresponding to another stool sample, the pathological characteristic distribution area covering the outermost layer is custom-selected as the target analysis area of the stool sample.
[0062] In another optional embodiment, in a set of pathological characteristic distribution areas corresponding to a stool sample, a pathological characteristic distribution area covering the entire area is custom-selected as the target analysis area of the stool sample, and then, in a set of pathological characteristic distribution areas corresponding to another stool sample, a pathological characteristic distribution area covering the entire area is custom-selected as the target analysis area of the stool sample.
[0063] Furthermore, the spatial analysis module 102 is also used to perform differential analysis and screening on the molecules in each target analysis area according to a statistical test method, obtain differential molecules, and determine the differential molecules as candidate markers for the target disease.
[0064] In an optional embodiment, referring to Figure 6 As shown, Figure 6 Some embodiments of the present application provide a flowchart of a spatial mapping module performing differential analysis on molecules according to a statistical test method. The differential analysis method may include but is not limited to steps S601 to S603.
[0065] Step S601: summing up the abundance values of molecules of the same molecular type in each pathological characteristic distribution area to obtain a total abundance set of molecules in each pathological characteristic distribution area.
[0066] The molecular total abundance set includes total abundance values of molecules of different molecular types.
[0067] In an optional embodiment, the abundance values of molecules of the same molecular type in each of the outermost pathological characteristic distribution regions are summed to obtain a plurality of molecular sum abundance sets covering the outermost pathological characteristic distribution regions. In another optional embodiment, the abundance values of molecules of the same molecular type in the pathological characteristic distribution area covering the entire region are summed to obtain a total abundance set of molecules in the pathological characteristic distribution area covering the entire region.
[0068] Step S602: normalizing the molecular sum abundance sets of each pathological characteristic distribution region to obtain each normalized molecular sum abundance set.
[0069] In an optional embodiment, the molecular sum abundance set covering the outermost pathological characteristic distribution region is normalized to obtain a plurality of normalized molecular sum abundance sets covering the outermost pathological characteristic distribution region.
[0070] In an optional embodiment, the molecular sum abundance set of the pathological characteristic distribution area covering the entire region is normalized to obtain a plurality of normalized molecular sum abundance sets of the pathological characteristic distribution area covering the entire region.
[0071] Step S603: performing difference analysis and screening on each normalized molecular sum abundance set to obtain a set of molecular difference values between each pathological characteristic distribution area.
[0072] It should be noted that the statistical test method may be a fold change method, a t-test method, a Wilcoxon test method, or a limma difference analysis method, and this application does not make any specific limitation.
[0073] In an optional embodiment, the limma difference analysis method is used to perform difference analysis and screening on the normalized total abundance set of target molecules covering the outermost pathological characteristic distribution area to obtain the difference molecules between the pathological characteristic distribution areas covering the outermost layer.
[0074] In another optional embodiment, the normalized total abundance set of molecules in the pathological characteristic distribution areas covering the entire region is differentially analyzed and screened by the limma differential analysis method to obtain differential molecules between the pathological characteristic distribution areas covering the entire region.
[0075] Reference Figure 7 As shown, Figure 7 is a schematic diagram of the results of calculating the difference molecules between the outermost layer and the entire area of feces by the spatial analysis module provided in some embodiments of the present application, from Figure 7 The results of the limma difference analysis in Part A show that the number of markers screened in the outermost layer of feces is about 2.5 times higher than that of the mixed fecal samples in the entire region. Figure 7As can be seen in Part B, the differential protein molecules screened out from the outermost layer of feces also include the differential protein molecules screened out from the mixed feces in the entire region. It is worth noting that in the results with an adjust-p value of <0.05, the outermost layer of feces samples also screened out four potential fecal markers that have not been reported in related literature before: IGHG1 protein molecule, ANXA1 protein molecule, FGA protein molecule, FABP5 protein molecule, which are related to cancer migration, immune regulation, blood coagulation and fatty acid transport, respectively.
[0076] Further, refer to Figure 8 As shown, Figure 8 is a result schematic diagram of the correlation and co-localization characteristics of differential protein molecules provided in some embodiments of the present application, Figure 8 Part A shows that the outermost fecal markers have a strong correlation (Spearman correlation analysis, adjusted-p<0.05, r>0.7). Figure 8 From the differential protein molecule distribution diagram in part B, it can be seen that the three unreported potential CRC stool markers detected in the outermost layer of stool samples co-localize with HP, the most widely used clinical auxiliary group diagnostic marker, and are mainly distributed in the outer layer of stool. Therefore, the experimental operation of mixing stool will mask these unevenly distributed low-abundance protein molecules, and the spatial analysis module of this application can detect more low-abundance markers in the sample that were previously masked by homogenization.
[0077] Furthermore, the sampling module 103 is used to obtain a target sample, perform molecular extraction on the target sample, and obtain a sample molecule solution.
[0078] In a possible embodiment, when the sample molecule solution is a protein molecule solution, referring to Fig. 9 As shown, Fig. 9 This is a flow chart of protein molecule extraction from a target sample provided by some embodiments of the present application. The protein molecule extraction method may include but is not limited to steps S901 to S905.
[0079] Step S901: freezing and slicing the target sample to obtain sample slices.
[0080] In one possible embodiment, target samples are obtained from 10 subjects, and the target samples are formed stool (7 control subjects and 3 CRC subjects). The collected formed stool is frozen at -80°C and then sliced at about the middle of the stool to obtain frozen stool cross-sections with a thickness of about 2-3 mm.
[0081] Step S902: using a micro-stent to divide the sample slice into a plurality of sample particles with preserved spatial position information.
[0082] Specifically, by using a micro-stent containing 30×30 micro-holes, each micro-hole having a cross-sectional size designed to be 1.2 mm×1.2 mm, the feces cross-section is cut into a number of feces particles with retained position information to obtain sample microparticles.
[0083] Step S903: Perform protein lysis on the sample slices according to the pressure cycle technology to obtain a protein lysis solution.
[0084] In a possible embodiment, 40ul of protein lysis solution and a portion of the fecal microparticles are added to a single PCT-Microtube, which is tightly capped with a PCT-MicroPestle stopper and placed in a PCT instrument. Then, according to the set reaction parameters, a cyclic pressure from 1 atmosphere to about 3000 times the atmospheric pressure is applied to the sample to lyse and extract the protein. Finally, after the cycle is completed, the fecal protein lysis solution is aspirated to obtain a protein lysis solution of the fecal sample.
[0085] Step S904: According to the Automated SP3 technology, the protein lysate is subjected to high-throughput, automated enzymatic purification treatment to obtain a protein peptide solution.
[0086] In a possible embodiment, according to the automated process of the SP3 automated liquid handling system (Bravo), 2ul of protein lysate was diluted 20 times and heated for 15min before being added to a 96-well plate. Then, according to the measured CRC fecal protein concentration range, the Bravo instrument automatically added 5ul of mixed magnetic beads and 45ul of 100% ACN to each well, and incubated for 18min with shaking; after the protein and the magnetic beads were fully combined, the deep-well plate was placed on a magnetic rack for adsorption precipitation for 5min, and the supernatant was removed, followed by washing twice with 200μl / well of 80% ethanol and once with 171.5μl / well of 100% acetonitrile. After removing the washing solution, the magnetic beads were resuspended in 35ul / well of enzyme buffer and 10ul / well of trypsin solution, and the plate was sealed with a PCR film and placed in a shaker for enzymatic hydrolysis overnight. The enzymatic hydrolysis conditions were constant temperature at 37°C and 220rpm for enzymatic hydrolysis and purification to obtain a protein peptide solution.
[0087] The mixed magnetic beads are 50ul Sera-Mag Speed magnetic beads A and 50ul Fisher Scientific magnetic beads B mixed and resuspended in 1ml ddH2O to achieve a final working concentration of 5ug / ul.
[0088] The composition of the above trypsin solution is 20 μg of trypsin (Promega) dissolved in 4 ml of enzyme buffer.
[0089] The enzyme buffer composition is 0.65 ml of pure water, 0.25 ml of Tris, and 0.1 ml of ACN.
[0090] Step S905: Transfer the protein peptide solution to Evotips to obtain a protein molecule solution.
[0091] In a possible embodiment, according to the automated process of the Bravo instrument, 6ul of 5% trifluoroacetic acid is automatically added to the protein peptide solution after overnight enzymatic hydrolysis, and the enzymatic hydrolysis is quickly acidified and stopped. Subsequently, the magnetic beads are fixed on the magnetic rack, the supernatant is transferred to the activated Evotips, and the protein peptide solution is recovered and loaded to obtain a protein molecule solution.
[0092] Among them, the activation and peptide loading steps of Evotips are as follows: Step A: add 20μl buffer B to each Evotips and centrifuge at 700 rcf for 1 minute; Step B: Evotips are activated by soaking in isopropanol until all tips are grayish white; Step C: add 20μl buffer A to each Evotips and centrifuge again at 700 rcf for 1 minute; Step D: transfer the stopped enzymatic peptide solution to Evotips and centrifuge at 800 rcf for 1 minute; Step E: add 50ul buffer A to each Evotips and centrifuge at 800 rcf for 1 minute; Step F: add 100μl buffer A to each Evotips and centrifuge briefly to move the liquid down to the membrane and keep the Evotips moist.
[0093] It should be noted that the components of the buffer A are 99.9% ddH2O and 0.1% FA, and the components of the buffer B are 80% ACN, 20% water, and 0.1% FA.
[0094] The components of the buffer A are 99.9% ddH2O and 0.1% FA, and the components of the buffer B are 80% ACN, 20% water and 0.1% FA.
[0095] In one possible embodiment, when the sample molecule solution is a lipid molecule solution, referring to Fig.10 As shown, Fig.10 This is a flow chart of extracting lipid molecules from a target sample provided by some embodiments of the present application. The lipid molecule extraction method may include but is not limited to steps S1001 to S1003.
[0096] Step S1001: freezing and slicing the target sample to obtain sample slices.
[0097] In one possible embodiment, target samples are obtained from 10 subjects, and the target samples are formed stool (7 control subjects and 3 CRC subjects). The collected formed stool is frozen at -80°C and then sliced at about the middle of the stool to obtain frozen stool cross-sections with a thickness of about 2-3 mm.
[0098] Step S1002: using a micro-stent to divide the sample slice into a plurality of sample particles with preserved spatial position information.
[0099] Specifically, by using a micro-stent containing 30×30 micro-holes, each micro-hole having a cross-sectional size designed to be 1.2 mm×1.2 mm, the feces cross-section is cut into a number of feces particles with retained position information to obtain sample microparticles.
[0100] Step S1003: The sample slices are subjected to lipid dissolution treatment using a lipid extract and the supernatant is taken and concentrated and redissolved to obtain a lipid molecule solution.
[0101] In a possible embodiment, a chloroform-methanol mixture (chloroform:methanol=3:1) is added to the fecal pellets, shaken for 2 minutes, and centrifuged at 4°C, 16,000g, for 10 minutes to obtain the supernatant, which is concentrated and frozen at -80°C to obtain a lipid molecule solution.
[0102] It should be noted that the lipid molecule solution needs to be re-dissolved with 50ul of chloroform-methanol mixture (chloroform: methanol = 5:13) before mass spectrometry detection.
[0103] Furthermore, the sampling module 103 is also used to perform mass spectrometry detection on the sample molecule solution to obtain spatial omics data and spatial position data set of the target sample.
[0104] In an optional embodiment, when the sample molecule solution is a protein molecule solution, an Evosepone liquid chromatograph and an Orbitrap Fusion Lumos Tribrid mass spectrometer coupled to a FAIMS source are used to perform mass spectrometry detection on the protein molecule solution in Evotips to obtain spatial omics data and spatial position data sets corresponding to the protein molecules, wherein a positive electrode mode is used during mass spectrometry detection, and the spray voltage is 2200 V. The analytical column is an Acclaim PepMap100 C18 column (75 mm × 250 mm, 2 mm, 100 Å, Thermo Scientific), and a pre-programmed gradient of Evosepone SPD60 is used, with a flow rate of 1 ul / min, and the corresponding mobile phase A is 0.1% formic acid, and the corresponding mobile phase B is 0.1% formic acid, 80% acetonitrile and 20% pure water. In an optional embodiment, when the sample molecule solution is a lipid molecule solution, a ThermoScientific Dionex UltiMate 3000 rapid separation liquid chromatograph coupled to an Orbitrap Exploris 480 mass spectrometer is used to analyze the reconstituted lipid molecule solution for mass spectrometry detection to obtain spatial omics data and spatial position data sets corresponding to the lipid molecule solution, wherein during the mass spectrometry detection process, the spray voltage is 3500V, the analytical column is C18, the flow rate is 0.2ml / min, the corresponding mobile phase A is 60% acetonitrile, 10mM ammonium acetate, 0.1% acetic acid and 40% pure water, and the corresponding mobile phase B is 90% isopropanol, 10% acetonitrile, 10mM ammonium acetate and 0.1% acetic acid.
[0105] In the second aspect, the present application embodiment also provides a marker screening method based on spatial omics, referring to Fig.11 , Fig.11 A flow chart of a marker screening method based on spatial omics provided in some embodiments of the present application. The marker screening method based on spatial omics may include but is not limited to steps S1101 to S1104.
[0106] Step S1101: Obtain target molecule types, spatial omics datasets and spatial position datasets of multiple target samples, and perform molecular-level spatial mapping on each target sample according to the target molecule types, each spatial omics dataset and each spatial position dataset to obtain molecular distribution data of each target sample.
[0107] Step S1102: Draw a molecular distribution map of each target sample according to each molecular distribution data.
[0108] Step S1103: Obtaining pathological information of the target disease, and customizing the target analysis region corresponding to each target sample in each molecular distribution map according to the pathological information; Step S1104: According to the statistical test method, the molecules in each target analysis area are differentially analyzed and screened to obtain differential molecules, and the differential molecules are determined as markers of the target research disease.
[0109] The above-mentioned marker screening method and the above-mentioned marker screening system 100 are based on the same inventive concept. The above-mentioned process describes the marker screening method of an embodiment of the present application, by obtaining the target molecule type, the spatial omics data set and the spatial position data set of the target sample, and according to the target molecule type, the spatial omics data set and the spatial position data set, the target sample is spatially mapped at the molecular level to obtain the molecular distribution data of the target sample, and then, the molecular distribution map of the target sample is drawn according to the molecular distribution data; then, the pathological information of the target research disease and the molecular distribution maps of multiple target samples are obtained, and according to the pathological information, the target analysis area corresponding to each target sample is customized in each molecular distribution map; finally, according to the statistical test method, the molecules in each target analysis area are differentially analyzed and screened to obtain differential molecules, and the differential molecules are determined as markers of the target research disease, so that more low-abundance markers that have been masked by homogenization in the past can be detected in the sample, providing a rich potential biomarker database for the corresponding research disease, and promoting the development of personalized medical research.
[0110] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above marker screening method when executing the computer program. The electronic device can be any smart terminal including a mobile phone, a tablet computer, a car computer, etc.
[0111] See also Fig.12 , Fig.12 : is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application, and the electronic device includes: The processor 1201 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the QSPI serial port transfer method and / or cache data reading method provided in the embodiments of the present application; The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1202 can store an operating system and other applications. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1202, and the processor 1201 calls and executes the QSPI serial port transfer method and / or cache data reading method provided in the embodiments of this application; Input / output interface 1203, used to implement information input and output; The communication interface 1204 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.); A bus 1205 that transmits information between various components of the device (e.g., the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204); The processor 1201 , the memory 1202 , the input / output interface 1203 and the communication interface 1204 are connected to each other in communication within the device via the bus 1205 .
[0112] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it is the marker screening method provided in the embodiment of the present application.
[0113] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0114] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0115] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0116] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0118] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0119] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0120] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0121] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0122] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0123] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0124] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A marker screening system based on spatial omics, characterized in that: include: A spatial mapping module and a spatial analysis module, wherein the spatial analysis module is communicatively connected with the spatial mapping module; The spatial mapping module is used to obtain the target molecule type, the spatial omics data set and the spatial position data set of the target sample, and perform molecular-level spatial mapping on the target sample according to the target molecule type, the spatial omics data set and the spatial position data set to obtain the molecular distribution data of the target sample; The spatial mapping module is also used to draw a molecular distribution map of the target sample according to the molecular distribution data; The spatial analysis module is used to obtain pathological information of the target disease and molecular distribution maps of the plurality of target samples, and to user-definedly determine the target analysis region corresponding to each target sample in each molecular distribution map according to the pathological information; The spatial analysis module is also used to perform differential analysis and screening on the molecules in each target analysis area according to a statistical test method to obtain differential molecules, and determine the differential molecules as candidate markers for the target disease.
2. The marker screening system according to claim 1, characterized in that According to the statistical test method, the molecules in each of the pathological characteristic distribution areas are differentially analyzed and screened to obtain differential molecules, including: The abundance values of molecules of the same molecular type in each of the pathological characteristic distribution areas are summed up respectively to obtain a total abundance set of molecules in each of the pathological characteristic distribution areas; Normalizing the molecular sum abundance sets of each of the pathological characteristic distribution areas to obtain each normalized molecular sum abundance set; The normalized total abundance sets of molecules are differentially analyzed and screened to obtain the differential molecules between the distribution areas of the pathological characteristics.
3. The marker screening system according to claim 1, characterized in that The step of customizing, according to the pathological information, the target analysis region corresponding to each of the target samples in each of the molecular distribution maps comprises: Determining the pathological characteristics of the target research disease according to the pathological information; According to the pathological characteristics, determining a pathological characteristic distribution area set of each target sample in each of the molecular distribution maps, wherein the pathological characteristic distribution area set includes pathological characteristic distribution areas of multiple different area ranges; In the pathological characteristic distribution area sets of each target sample, one pathological characteristic distribution area is customized to be selected as the target analysis area corresponding to the target sample, so as to determine the target analysis area corresponding to each target sample.
4. The marker screening system according to claim 1, characterized in that The step of performing molecular-level spatial mapping on the target sample according to the target molecule type, the spatial omics data set, and the spatial position data set to obtain molecular distribution data of the target sample includes: According to the target molecule type, determining target spatial omics data of the target molecule in the spatial omics data set; Normalizing the target space omics data to obtain normalized target space omics data; Determining target spatial position data of the target molecule in the spatial position data set according to the target molecule type; The normalized target spatial omics data and the target spatial position data are spatially correlated and mapped to obtain the molecular distribution data.
5. The marker screening system according to claim 1, characterized in that The system further comprises a sampling module, which is communicatively connected with the spatial mapping module, wherein: The sampling module is used to obtain the target sample, perform molecular extraction on the target sample, and obtain a sample molecule solution; The sampling module is also used to perform mass spectrometry detection on the sample molecule solution to obtain the spatial omics data and spatial position data set of the target sample.
6. The marker screening system according to claim 5, characterized in that When the sample molecule solution is a protein molecule solution, the step of performing molecular extraction on the target sample to obtain sample molecules includes: Freezing and slicing the target sample to obtain sample slices; Using a micro-support to divide the sample slice into a plurality of sample particles with retained spatial position information; According to the pressure cycle technology, the sample particles are subjected to protein lysis treatment to obtain a protein lysis solution; According to the Automated SP3 technology, the protein lysate is subjected to high-throughput, automated enzymatic hydrolysis and purification treatment to obtain a protein peptide solution; The protein peptide solution is transferred to Evotips to obtain the protein molecule solution.
7. The marker screening system according to claim 5, characterized in that When the sample molecule solution is a lipid molecule solution, the step of performing molecular extraction on the target sample to obtain sample molecules includes: Freezing and slicing the target sample to obtain sample slices; Using a micro-support to divide the sample slice into a plurality of sample particles with retained spatial position information; The lipid extract is used to dissolve the sample particles and the supernatant is taken, concentrated and redissolved to obtain the lipid molecule solution.
8. A marker screening method based on spatial omics, characterized in that: include: Acquire a target molecule type, a spatial omics data set and a spatial position data set of multiple target samples, and perform molecular-level spatial mapping on each of the target samples according to the target molecule type, each of the spatial omics data sets and each of the spatial position data sets to obtain molecular distribution data of each of the target samples; Draw a molecular distribution map of each target sample according to each molecular distribution data; Acquire pathological information of the target disease to be studied, and according to the pathological information, define the target analysis region corresponding to each of the target samples in a customized manner in each of the molecular distribution maps; According to the statistical test method, the molecules in each of the target analysis areas are differentially analyzed and screened to obtain differential molecules, and the differential molecules are determined as markers for the target research disease.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the marker screening method according to claim 8 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the marker screening method according to claim 8 is implemented.
Citation Information
Patent Citations
Screening method of liver cirrhosis anion markers and application
CN110850074A
Disease type discrimination method based on pathology and living tissue clinical diagnosis big data
CN112863665A
Marker for identifying and diagnosing thyroid follicular tumor and application thereof
CN115112745A
Molecular marker detection product based on PDX / PDTX tumor living tissue biological sample and database and preparation method thereof
CN116593262A
Rapid tumor tissue identification method based on fingerprint spectrogram of lipids on tissue surface
WO2020259187A1