Method of determining endometrial receptivity and application thereof
The method measures endometrial implantation-related gene expression using biomarkers and machine learning to accurately determine the implantation window, addressing inaccuracies in existing IVF-ET methods and enhancing IVF success.
Patent Information
- Application Number
- JP2025076579
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-22
- Filing Date
- 2025-05-02
- Publication Date
- 2025-08-07
AI Technical Summary
Existing methods for determining endometrial implantation in IVF-ET are inaccurate, leading to increased risks of implantation failure due to unclear implantation window periods, particularly for women with repeated implantation failure or secondary infertility disorders.
A method involving the measurement of endometrial implantation-related gene expression levels in samples such as endometrial tissue, uterine fluid, or vaginal secretions, using biomarkers identified through RNA-seq sequencing and machine learning algorithms to determine the implantation window period.
Significantly reduces the error rate in determining endometrial implantation status, improving the success rate of IVF-ET by providing accurate and non-invasive assessment of the implantation window.
Smart Images

Figure 2025116004000127 
Figure 2025116004000128 
Figure 2025116004000129
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of biomedicine, and in particular to a method for determining endometrial implantation and its applications. [Background technology]
[0002] During human reproduction, a fertilized egg is located in the mother's uterus, attaches, implants, and ultimately develops into a mature fetus. The implantation process significantly influences the success of pregnancy, and successful clinical pregnancy requires good endometrial implantation (ER) and high-quality embryos, as well as the simultaneous development of the endometrium and embryo. ER refers to the endometrium's ability to receive an embryo. Embryo implantation is only possible within a short, specific period, known as the "implantation window." In adult females, this corresponds to days 20–24 of the menstrual cycle or days 6–8 after ovulation.
[0003] In the field of in vitro fertilization and embryo transfer (IVF-ET), women with repeated implantation failure or other secondary infertility disorders have an inaccurate time period for the implantation window period. Calculating the implantation window period according to the menstrual cycle or ovulation date significantly increases the risk of implantation failure. Although the specific mechanism of ER is unknown, it is clear that insufficient ER is one of the important reasons for embryo implantation failure in IVF-ET.
[0004] Therefore, there is an urgent need to find stable, non-invasive and accurate markers and methods for assessing endometrial implantation, which can help medical professionals clearly determine the ER status and help patients accurately find the implantation window period, which is of great importance in promoting the success rate of IVF-ET. Summary of the Invention
[0005] Contents of the invention The purpose of the present invention is to provide a stable, non-invasive and accurate marker and method for assessing endometrial implantation, thereby helping medical professionals clearly determine the ER status and help patients accurately find the implantation window period, which is of great importance to promoting the success rate of IVF-ET.
[0006] A first aspect of the present invention is to provide a method for determining endometrial implantation, comprising the steps of: (a) providing a sample; (b) measuring the expression level of endometrial implantation-related genes in the sample; (c) comparing the expression level of the endometrial implantation-related gene obtained in step (b) with a predetermined value, thereby determining endometrial implantation.
[0007] In another preferred embodiment, the sample is selected from the group consisting of endometrial tissue, uterine fluid, uterine washings, vaginal exfoliated cells, vaginal secretions, endometrial biopsy products, serum, plasma, or combinations thereof.
[0008] In another preferred embodiment, the expression level of the endometrial implantation-associated gene obtained in step (b) is higher than a predetermined value, indicating the presence of endometrial implantation.
[0009] In another preferred embodiment, the expression level of an endometrial implantation-associated gene comprises the expression level of a cDNA of the endometrial implantation-associated gene.
[0010] In another preferred embodiment, the samples are derived from periods LH+n, LH+n+2, LH+n+4, where n is 3-7, preferably n is 4-6.
[0011] In another preferred embodiment, the sample is derived from the following period: day n after ovulation, where n is 3-7, preferably n is 4-6.
[0012] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A.
[0013] [Table 1-1]
[0014] [Table 1-2]
[0015] [Table 1-3]
[0016] [Table 1-4]
[0017] [Table 1-5]
[0018] [Table 1-6]
[0019] [Table 1-7]
[0020] [Table 1-8]
[0021] [Table 1-9]
[0022] [Table 1-10]
[0023] [Table 1-11]
[0024] [Table 1-12]
[0025] [Table 1-13]
[0026] [Table 1-14]
[0027] [Table 1-15]
[0028] [Table 1-16]
[0029] [Table 1-17]
[0030] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 40 genes selected from Table A.
[0031] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 147 genes selected from Table A.
[0032] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 259 genes selected from Table A.
[0033] In another preferred embodiment, the endometrial implantation-associated genes further comprise 5 to 200 genes.
[0034] In another preferred embodiment, the endometrial implantation-associated genes further comprise one or more genes selected from Table B.
[0035] [Table 2-1]
[0036] [Table 2-2]
[0037] [Table 2-3]
[0038] [Table 2-4]
[0039] [Table 2-5]
[0040] [Table 2-6]
[0041] [Table 2-7]
[0042] [Table 2-8]
[0043] [Table 2-9]
[0044] [Table 2-10]
[0045] [Table 2-11]
[0046] [Table 2-12]
[0047] In another preferred embodiment, the endometrial implantation-associated genes further comprise additional genes, bringing the total number of genes to 10,000.
[0048] In a second aspect of the invention, a set of biomarkers is provided, which set comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A.
[0049] In another preferred embodiment, the set of biomarkers comprises at least 40 genes selected from Table A.
[0050] In another preferred embodiment, the set of biomarkers comprises at least 147 genes selected from Table A.
[0051] In another preferred embodiment, the set of biomarkers comprises at least 259 genes selected from Table A.
[0052] In another preferred embodiment, the set of biomarkers further comprises between 5 and 200 genes.
[0053] In another preferred embodiment, the set of biomarkers further comprises additional genes, so that the total number of genes reaches 10,000.
[0054] In another preferred embodiment, the set of biomarkers is used to determine endometrial implantation or for the manufacture of a kit or reagent, wherein the kit or reagent is used to assess or diagnose (including early and / or auxiliary diagnosis) the endometrial implantation status of a test subject.
[0055] In another preferred embodiment, the biomarker or set of biomarkers is obtained from endometrial tissue, uterine fluid, uterine washings, vaginal exfoliated cells, vaginal secretions, endometrial biopsy products, serum, or plasma samples.
[0056] In another preferred embodiment, one or more biomarkers selected from Table A are increased compared to a predetermined value, indicating endometrial implantation in the test subject.
[0057] In another preferred embodiment, each biomarker is identified by a method selected from the group consisting of RT-qPCR, RT-qPCR chip, next generation sequencing, expression profile chip, methylation chip, third generation sequencing, or a combination thereof.
[0058] In another preferred embodiment, the set is used to assess the endometrial implantation status of a test subject.
[0059] A third aspect of the present invention provides a combination of reagents for determining endometrial implantation status, wherein the combination of reagents comprises a reagent for detecting each biomarker in the set of the second aspect of the present invention.
[0060] In another preferred embodiment, the reagents comprise materials for detecting each biomarker in the set of the second aspect of the invention using a method selected from the group consisting of RT-qPCR, RT-qPCR chip, next generation sequencing, expression profile chip, methylation chip, third generation sequencing, or a combination thereof.
[0061] A fourth aspect of the present invention provides a kit, the kit comprising the set of the second aspect of the present invention and / or the reagent combination of the third aspect of the present invention.
[0062] A fifth aspect of the present invention provides the use of a set of biomarkers for the manufacture of a kit for use in assessing the endometrial implantation status of a test subject, wherein the set of biomarkers comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A.
[0063] In another preferred embodiment, the evaluation and diagnosis comprises the following steps: (1) providing a sample obtained from a subject to be detected and detecting the level of each biomarker in the set in the sample; and (2) A step of comparing the level detected in (1) with a predetermined value.
[0064] In another preferred embodiment, the sample is selected from the group consisting of endometrial tissue, uterine fluid, uterine washings, vaginal exfoliated cells, vaginal secretions, endometrial biopsy products, serum, plasma, or combinations thereof.
[0065] In another preferred embodiment, one or more biomarkers selected from Table A are increased compared to a predetermined value, indicating endometrial implantation in the subject under test.
[0066] In another preferred embodiment, the method further comprises the step of treating the sample prior to step (1).
[0067] A sixth aspect of the present invention provides a method for determining endometrial implantation in a subject to be detected, comprising the steps of: (1) providing a sample from a subject to be detected and detecting the level of each biomarker in a set in the sample, wherein the set comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A; (2) A step of comparing the level detected in (1) with a predetermined value.
[0068] A seventh aspect of the present invention provides a system for assessing the endometrial implantation status of a subject to be detected, comprising: (a) an endometrial implantation status signature input module, the input module being used to input an endometrial implantation status signature of a subject, wherein the endometrial implantation status signature comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A; (b) a processing module for determining endometrial implantation status, wherein the processing module inputs characteristics of the endometrial implantation status according to a preset criterion, thereby performing an evaluation process to obtain a score of the endometrial implantation status; further, the module compares the score of the endometrial implantation status with a predetermined value, thus obtaining an auxiliary diagnosis result, wherein if the score of the endometrial implantation status is higher than the predetermined value, it is suggested that the subject has endometrial implantation; (c) an auxiliary diagnostic result output module, the output module being used to output the auxiliary diagnostic result;
[0069] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 40 genes selected from Table A.
[0070] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 147 genes selected from Table A.
[0071] In another preferred embodiment, the endometrial implantation-associated genes comprise at least 259 genes selected from Table A.
[0072] In another preferred embodiment, the endometrial implantation-associated genes further comprise 5 to 200 genes. In another preferred embodiment, the endometrial implantation-associated genes further comprise additional genes so that the total number of genes reaches 10,000.
[0073] In another preferred embodiment, the subject is a human.
[0074] In another preferred embodiment, the score comprises: (a) a score for a single feature; and (b) a sum of the scores for multiple features.
[0075] In another preferred embodiment, the feature input module is selected from the group consisting of a sample collector, a sample storage tube, a cell lysis and nucleic acid sample extraction kit, an RNA nucleic acid reverse transcription amplification kit, an NGS library construction kit, a library quantification kit, a sequencing kit, or a combination thereof.
[0076] In another preferred embodiment, the processing module for determining endometrial implantation status includes a processor and a memory, and the memory stores scoring data of ER status based on characteristics of endometrial implantation status.
[0077] In another preferred embodiment, the output module includes a reporting system.
[0078] It should be understood that each of the above-mentioned technical features of the present invention or each technical feature specifically described below (e.g., embodiments) may be combined with each other, thereby constituting novel or preferred technical solutions that would not be described one by one in this specification due to space constraints. [Brief explanation of the drawings]
[0079] [Figure 1] FIG. 1 is a distribution diagram showing the cDNA amplification products of the samples in Example 4 of the present invention. [Figure 2] FIG. 2 shows an overview of the process of the supervised learning method used in the seventh embodiment of the present invention. [Figure 3] FIG. 3 shows a flow diagram of data processing in the seventh embodiment of the present invention. [Figure 4] FIG. 4 is a schematic diagram showing the results after detection of patients in Example 7 of the present invention. [Figure 5] FIG. 5 is a diagram showing a target detection mode in Example 7 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] Detailed Description of the Embodiments Through extensive and detailed research, the present inventors have unexpectedly discovered biomarkers for determining endometrial implantation and their combinations. Specifically, the present invention discloses a set of biomarkers, which can be used to evaluate or diagnose the endometrial implantation status of a test subject (including early diagnosis and / or auxiliary diagnosis), which can significantly reduce the error rate. Therefore, the present invention has important application value. Based on the above, the present inventors have completed the present invention.
[0081] term Terms used herein have meanings commonly understood by those of ordinary skill in the art. However, to better understand the present invention, some definitions and related terms are set forth below.
[0082] According to the present invention, the term "ER" refers to the ability of the endometrium to accept an embryo, the implantation of which is only tolerated by the endometrium for a specific short period of time.
[0083] According to the present invention, the term "set of biomarkers" refers to one biomarker or a combination of two or more biomarkers.
[0084] According to the present invention, the levels of biomarkers are identified by RT-qPCR, RT-qPCR chip, next generation sequencing, expression profile chip, methylation chip, third generation sequencing, or other methods.
[0085] According to the present invention, the term "biomarker," also known as "biological marker," refers to a measurable indicator of an individual's biological state. Such a biomarker can be any substance in an individual that is relevant to the individual's specific biological state (e.g., disease), such as a nucleic acid marker (e.g., DNA), a protein marker, a cytokine marker, a chemokine marker, a carbohydrate marker, an antigen marker, an antibody marker, a species marker (species / genus marker), a functional marker (KO / OG marker), etc. After measurement and evaluation, biomarkers are typically used to examine normal biological processes, pathogenic processes, or pharmacological responses to therapeutic interventions, making them useful in many scientific fields.
[0086] According to the present invention, the term "individual" refers to an animal, in particular a mammal, such as a primate, preferably a human.
[0087] According to the present invention, the term "plasma" refers to the liquid component of whole blood. Depending on the separation method used, plasma may be completely free of cellular components, contain varying amounts of platelets and / or small amounts of other cellular components.
[0088] According to the present invention, terms such as "a / an," "one," and "such" not only refer to a singular individual / unit, but also include general categories that can identify specific embodiments.
[0089] It should be noted that the terms provided herein are merely explained to those skilled in the art to better understand the present invention, but are not to be construed as limiting the present invention.
[0090] Detection Method In the present invention, a substance for detecting each biomarker in the set of the present invention by a method selected from the group consisting of RT-qPCR, RT-qPCR chip, next-generation sequencing, expression profile chip, methylation chip, third-generation sequencing, or a combination thereof.
[0091] kit In the present invention, the kit comprises the set according to the second aspect of the invention and / or the reagent combination according to the third aspect of the invention.
[0092] Predetermined value In the present invention, the predetermined value refers to the ER period (i.e., the scoring value obtained by scoring the endometrium using artificial intelligence or decision tree C4.5 algorithm, hidden Markov model (HMM), neural network backpropagation (BP), support vector machine (SVM), and various cluster analysis algorithms (simple clustering, hierarchical clustering, K-means clustering, self-organizing feature map, fuzzy clustering, Bayesian classifier, k-nearest neighbor, neural network method, decision tree method, voting classification method, principal component analysis (PCA), etc.
[0093] Evaluation method In another preferred embodiment, the method of the present invention can calculate a weighted comprehensive score by the formula S = W1S1 + W2S2 + WSi + WnSn, where W1, W2... Wn are weights, and S1, S2... Sn are the scores for each marker.
[0094] Preferably, the weights may be based on the analytical values in Table 9. For example, for the assessment of endometrial implantation, any weight (e.g., W1) may be the analytical value of the corresponding marker in Table 9.
[0095] In a preferred embodiment, S of the detected population 対象 =W1S1+W2S2+WiS3+·····WnSn
[0096] If the detected population of interest is greater than a predetermined value, it indicates that the object has an ER state.
[0097] The experimental results of the present invention show that the markers of the present invention can significantly reduce the error rate and significantly improve the accuracy of determining or diagnosing ER status.
[0098] A method for constructing an analytical model to determine endometrial implantation In the present invention, the method for constructing an analytical model for determining endometrial implantation includes the following steps:
[0099] High-sensitivity RNA reverse transcription and amplification from cDNA was performed, and then the cDNA was subjected to library construction for next-generation sequencing. After sequencing, the expression profile information of the sample was constructed from the sequencing output data. Bioinformatics analysis and classification identified the ER status and accurately determined the window period for embryo implantation into the endometrium, achieving personalized and accurate decisions.
[0100] In a preferred embodiment, the present invention provides a method for constructing an analytical model for determining endometrial implantation, comprising the following steps: (1) Collect samples from healthy women with different menstrual cycles, extract RNA, perform RNA reverse transcription, and amplify cDNA; (2) constructing a cDNA library for high-throughput sequencing; (3) By using reinforcement learning, unsupervised learning, or supervised learning methods, the expression levels of various genes in multiple samples are compared with different markers to obtain differentially expressed genes. Build an analytical model.
[0101] This invention relies on a highly sensitive RNA reverse transcription and cDNA amplification process and is based on RNA-seq sequencing, which allows for the acquisition of multiple expression profile information from endometrium, uterine fluid, or other reproductive endocrinology-related body fluids or exfoliation samples from patients. These samples are then subjected to ultra-high-dimensional classification and typing using bioinformatics, statistics, and machine learning methods based on different sampling periods, sampling methods, and expression profile features. ER status is determined according to different types.
[0102] The task of supervised learning is to learn from a model so that it can map from any given input to a predicted outcome, thus achieving high-dimensional predictive analytics.
[0103] The "multiple samples" in "training" in step (3) refer to samples from the same source but different individuals, such as 102 samples of endometrial tissue in the implantation period, 205 samples of endometrial tissue in the pre-implantation period, and 300 samples of endometrial tissue in the post-implantation period.
[0104] Preferably, the different menstrual cycles in step (1) are three cycles.
[0105] Preferably, the middle period of the three periods is LH+7, or day 5 after ovulation.
[0106] The test sample used in the present invention is obtained from an endometrial biopsy of a female subject in the natural menstrual cycle, performed on day 7 after the occurrence of the luteinizing hormone (LH) peak (LH+7), or from an endometrial biopsy of a female subject in a hormone replacement therapy (HRT) cycle, performed on day 5 after ovulation (P+5). This test is not suitable for use in patients with endometrial lesions (including endometrial adhesions, endometrial polyps, bronchial tuberculosis, etc.), hydrosalpinx, subjects who have not undergone proximal tubal ligation, submucosal uterine fibroids, uterine fibroids that protrude into the uterine cavity or adenomyoma, or patients with endometriosis (phases III to IV).
[0107] Preferably, each of the three periods further includes a period of 1 to 3 days, preferably a period of 2 days, before and after the intermediate period.
[0108] Preferably, the intermediate period is the implantation period, 1 to 3 days before the intermediate period is the pre-implantation period, and 1 to 3 days after the intermediate period is the post-implantation period.
[0109] In the present invention, the grouping criteria for the "training set," also known as the "training data" or "training set," used in the model are a variety of healthy Chinese women in natural cycles with no past medical history or primary infertility, and women with a mean age of 19-25 kg / m 2 In addition to grouping and testing a large number of samples and accumulating clinical outcomes of cases, we also mastered the expression profiles or exfoliation of endometrial tissue, uterine fluid, or other reproductive endocrinology-related fluids during the "implantation window" (defined as the "implantation window" by tracking clinical outcomes during this period, when embryo implantation occurs and the embryo can effectively implant and develop). Meanwhile, we also sampled cases two days before and after the "implantation window" to obtain corresponding RNA-seq data, and defined the labels as the "implantation window" and "post-implantation period," respectively.
[0110] In the present invention, samples from the same individual at different menstrual cycles are obtained and sequenced, and the gene expression profile characteristics from the samples at different periods are compared, which can better indicate differentially expressed genes and thus reduce the occurrence of false positives.
[0111] Preferably, the sample in step (1) comprises any one or combination of at least two of fundal endometrial tissue, uterine fluid, or vaginal scraping.
[0112] In the present invention, the sample may be an endometrial biopsy product, or may be a patient's uterine fluid obtained by non-invasive means, such as uterine washings, vaginal exfoliated cells, and even vaginal secretions. The sample origins are wide-ranging, sample collection is simple and rapid, and the degree of female compliance is promoted, along with improved accuracy of gene expression profile characteristics, by verifying different origins of samples from the same individual. Uterine washings, vaginal exfoliated cells, and vaginal secretions have small sample sizes, and conventional detection methods require a large number of samples, so endometrial biopsy products must be selected. However, the present invention can meet testing requirements with a small amount of sample, thereby expanding the range of sample types and reducing pain and discomfort for the subject.
[0113] Preferably the sample volume of fundal endometrial tissue is greater than 5 mg, preferably 5-10 mg, for example 5 mg, 6 mg, 7 mg, 8 mg, 9 mg or 10 mg.
[0114] Preferably, the sample volume of uterine fluid is greater than 10 μL, preferably 10-15 μL, for example 10 μL, 11 μL, 12 μL, 13 μL, 14 μL or 15 μL.
[0115] Preferably the sample volume of the vaginal scraping is greater than 5 mg, preferably 5-10 mg, for example 5 mg, 6 mg, 7 mg, 8 mg, 9 mg or 10 mg.
[0116] In the present invention, the sample volume of the sample is smaller, the accuracy rate is higher, and the implantation state of the sample can be accurately predicted without a large number of samples.
[0117] Preferably, the cDNA in the library in step (2) has a concentration of 5 ng / μL or more.
[0118] Preferably, the sequencing in step (2) comprises RNA-Seq sequencing and / or qPCR sequencing.
[0119] In the present invention, the RNA-Seq sequencing method is superior to ChIP sequencing, detecting 2 to 8 times more differentially expressed genes than ChIP sequencing. In terms of the detection accuracy of low-abundance genes, the qPCR validation rate of RNA-Seq is 5 times higher than that of ChIP. In terms of the accuracy of differential expression fold change, the qPCR correlation of RNA-Seq is 14% higher than that of ChIP. The RNA-Seq or qPCR sequencing methods used in the present invention are common technical procedures for those skilled in the art.
[0120] Preferably, the sequencing in step (2) has a read length greater than 45 nt.
[0121] Preferably, the sequencing in step (2) has a number of reads of 2.5 million or more.
[0122] In the present invention, sequencing with a read length of 45 bases or more and a read number of 2.5 million or more meets the sequencing requirements.
[0123] In the present invention, the sequencing read length and read number are specifically selected, which reduces experimental duration and cost while ensuring accuracy.
[0124] Preferably, the present invention further comprises a data pre-processing step before step (3).
[0125] Preferably, the data pre-processing step involves normalization by gene length and sequencing depth.
[0126] Preferably, the normalization method comprises any one or combination of at least two of RPKM, TPM or FPKM, preferably FPKM.
[0127] In this invention, fragments per kilobase million (FPKM) is used for normalization for gene length and sequencing depth to eliminate the influence of sequencing depth after obtaining different labels of RNA-Seq data. Although RPKM and trans reads per million (TPM) are similar normalization methods, FPKM is preferably used in this invention.
[0128] FPKM is more flexible and easier to commercialize, and is suitable for paired-end or single-read sequencing libraries. RPKM is only suitable for single-read libraries. TPM values can reflect the ratio of reads of specific genes compared, so the values can be directly used for comparison between samples, but it is more tedious in the process, slower in operation, and has lower efficiency in batch analysis.
[0129] In the supervised learning of the present invention, a model is constructed using training data labeled with the pre-implantation period, implantation window period, and post-implantation period, and the model obtained by training can be used to predict the implantation status of unknown data (referring to a new sample). For example, the uterine fluid expression profile from one sample is newly input, and the machine learning model is used to determine the implantation status of the sample.
[0130] Preferably, the marker is an expression profile characteristic of endometrial tissue, uterine fluid or vaginal shedding under different implantation conditions.
[0131] In the present invention, expression profiles of samples from different sources, including any one or combination of at least two of endometrial tissue, uterine fluid or vaginal exfoliation under different implantation conditions, are obtained for the same individual, thus significantly improving the reliability of predicting outcomes.
[0132] Preferably, the method for analyzing differentially expressed genes is to find all genes with FPKM>0 in each sample, and screen the intersection of differentially expressed genes between pre-implantation and implantation, pre-implantation and post-implantation periods to meet p-value<0.05, and fold change F>2 or fold change<0.5.
[0133] In the present invention, sample screening always eliminates the influence of genes expressed at high or low levels in different markers on the analysis model, thus ensuring the elimination of overfitting conditions while achieving good fitting effects in the subsequent analysis model.
[0134] Preferably, the supervised learning method in step (3) comprises any one or a combination of at least two of Naive Bayes, Decision Tree, Logistic Regression, KNN or Support Vector Machine (SVM), preferably SVM.
[0135] SVM does not have many constraints on the original data and does not require prior information. There is a huge amount of data obtained by RNA-Seq sequencing, and different genes show different expression. SVM can maintain ultra-multidimensional (ultra-high dimensional) analysis, which can result in more accurate models.
[0136] Preferably, the SVM comprises the following script: library(e1071) svm.model<-svm(data.class~.,data4,kernel='linear') summary(svm.model) table(data$class,predict(svm.model,data4,type="data.class")) mydata=read.table(file.choose(),header=T,row.names=1) mydata2=log2(mydata+1) predict(svm.model,mydata2) table(shdata3$shdata.class,predict(svm.model,shdata3,type="shdata.class")) ## table(testdata$data.class,predict(svm.model,testdata,type="data.class")) shdata=read.table(file.choose(),header=T,row.names=1) shdata2=log2(shdata[,-length(shdata[1,])]+1) shdata3=data.frame(shdata2,shdata$class) predict(svm.model,shdata3) table(shdata3$shdata.class,predict(svm.model,shdata3,type="shdata.class"))
[0137] Preferably, the method for constructing an analytical model for determining endometrial implantation specifically comprises the following steps: (1) Samples were collected from healthy women during the pre-implantation, implantation, and post-implantation periods, and RNA extraction, RNA reverse transcription, and cDNA amplification were performed, respectively; (2) construct a cDNA library, enrich and purify the library, and then perform high-throughput sequencing to a concentration of 5 ng / μL or more; with a read length of more than 45 nt and a read count of 2.5 million or more; (3) Using the expression profile characteristics of endometrial tissue, uterine fluid, and vaginal exfoliation obtained in step (2) as markers under different implantation conditions, the marker data were used for model training using supervised learning methods to compare the expression of different genes with different markers. After obtaining RNA-Seq sequence data for different markers, normalization by gene length and sequencing depth was performed using the FPKM method, thereby comparing the expression of different genes in multiple samples with different markers, excluding the influence of genes that are always highly or lowly expressed in different markers on the model; differentially expressed genes were analyzed to obtain the model. (4) Using an SVM learning model, an analytical model is constructed to predict the implantation status of unknown samples using training data labeled with the pre-implantation period, implantation period, and post-implantation period, and the sampling time is repeatedly adjusted within multiple periods to improve the accuracy of the determination, after which an implantation test is performed.
[0138] In the present invention, RNAseq is used to obtain sequencing run output data, and expression levels are normalized using FPKM. Differences in gene expression levels between different time periods in the same individual are compared, and genes with significant differences are marked as differentially expressed genes. Genes with the most significant expression differences are identified from the same sample during the preimplantation, implantation, and postimplantation periods. The "differentially expressed genes" are then optimized across multiple individuals and different time periods. A feature library for the three time periods is constructed using machine learning. After normalizing expression levels, new samples to be tested are automatically classified according to the machine's feature assessment. Our model-building method utilizes a large number of training sets, fully training the machine with differentially expressed genes with significant specificity consistent with SVM, and improving model accuracy by tracking clinical outcomes over time.
[0139] In the present invention, in order to obtain individual implantation periods from different individuals, the sampling time can be repeatedly adjusted within multiple periods. For example, if the first detection is on LH+5, the test for the next period is postponed by two days and obtained on LH+7; if the first detection is on LH+9, the test for the next period is conducted two days earlier and obtained on LH+7.
[0140] The present invention has the following main advantages: (a) The biomarkers of the present invention can be used to accurately determine ER status, greatly reducing the error rate and therefore having great application value. (b) The present invention is based on an RNA-seq sequencing method that relies on a highly sensitive RNA reverse transcription and cDNA amplification process to obtain multiple expression profile information from endometrium, uterine fluid, or other reproductive endocrinology-related body fluids or exfoliation samples from patients. These samples are subjected to ultra-high-dimensional classification and typing using bioinformatics, statistics, and machine learning methods based on different sampling periods, sampling manners, and expression profile features. ER status is determined according to different types. (c) The present invention provides a method for building a model used to determine endometrial implantation using gene expression profile features that integrate and optimize RNA extraction, reverse transcription, cDNA purification, and library construction, and performs corresponding data processing, analysis, and model building, thereby significantly improving the accuracy of determining endometrial implantation. (d) The method of the present invention is characterized by the simplicity and convenience afforded by the rapid process and short duration, along with the low toxicity afforded by non-invasive uterine fluid biopsy.
[0141] The present invention will be further described in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention, but are not used to limit or restrict the scope of the present invention. Any experimental methods not specified in the detailed conditions of the following examples should generally follow conventional conditions or conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are calculated by weight.
[0142] Unless otherwise specified, the reagents and materials used in the examples of the present invention are commercially available products.
[0143] material Qiagen RNeasy Micro kit produced by Qiagen under serial number 74004; MALBAC Platinum Microscale RNA Amplification Kit manufactured by Yikon Genomic with serial number KT110700724; DNA Clean & Concentrator-5, manufactured by Zymo and with serial number OfD4014; A gene sequencing library kit (Illumina transposase method) manufactured by Yikon Genomics with serial number KT100801924; Agilent 2100 Bioanalyzer; Highly sensitive DNA chips; The materials used in the following examples are not limited to the above list, and other similar materials can be substituted, and the equipment not under the specified conditions is subjected to the conventional conditions or conditions recommended by the manufacturer. Those skilled in the art should acquire and use the relevant knowledge of conventional materials and equipment. [Example]
[0144] Example 1 Sample pretreatment and RNA extraction Sampling was performed by a gynecologist or a specialist qualified to perform specimen biopsies. Samples included endometrial tissue from the uterine fundus (>5 mg), uterine fluid (>10 μl) or other reproductive endocrinologically related fluids (>10 μl), and vaginal scrapings (>5 mg). After biopsy, the tissue, scraping, or fluid was completely immersed in RNA preservation solution (approximately 20 μL RNA) as soon as possible. Samples were stored in a refrigerator at -20°C or -80°C before transport.
[0145] The RNA of endometrial tissue was extracted using Qiagen RNeasy Micro Kit manufactured by Qiagen, and the specific method was as follows. 1. Experimental preparation: Clean your hands, pipettors, and benchtop with RNAse scavenger and nucleic acid remover (Note: All extraction steps should be performed in an area free of RNAse contamination). 2. 70% ethanol solution (350 μL for each sample) and 80% ethanol solution (500 μL for each sample) were prepared with deionized water and absolute ethanol. 3.4 volumes of absolute ethanol were added to Buffer RPE (provided in the kit) and mixed homogeneously for further use. 4. 10 μL of β-mercaptoethanol was taken and added to 990 μL of buffer RLT in a fume cupboard (use 350 μL of mixed solution for each sample) and mixed evenly for further use (Note: Prepare immediately before use for each extraction). 5. 50 μL of DNase-free water was added to the lyophilized DNase I powder (provided in the kit), mixed uniformly by inverting five times, and left at room temperature for 2 minutes. The resulting liquid (marked as DNase I solution) was subpackaged and stored in a -20°C refrigerator for further use, with no more than three freeze-thaw cycles. 6. 10 μL of DNase I solution (obtained from step 5) was taken into an EP tube without RNA enzyme, then 70 μL of buffer RDD was added, mixed evenly with a pipette, and then placed on ice for further use (marked as DNase I mixed solution). The solution was prepared immediately before use for each extraction. 7. The samples were removed from the -80°C refrigerator and placed on ice to thaw, then the samples, along with the storage fluid, were transferred to 1.5 mL RNase-free EP tubes. 8. Add 350 μL of β-mercaptoethanol-containing buffer RLT (obtained in step 4) to the EP tube, vortex for 30 seconds, and centrifuge in a flash. 9.350 μL of 70% ethanol was added successively to the EP tube, vortexed for 30 seconds, and centrifuged in a flash. 10. 650 μL of the supernatant was carefully absorbed (do not absorb any undissolved tissue fragments) and transferred to an adsorption column (provided in the kit). 11. Centrifuge at 14,000 xg for 30 seconds and discard the waste liquid from the collection cannula. 12.350 μL of buffer solution RW1 (included in the kit) was added to the adsorption column, and the column was centrifuged at 14,000×g for 30 seconds, and the waste liquid in the recovery cannula was discarded. 13.80 μL of the DNase I mixed solution was carefully added to the center of the adsorption column and incubated at room temperature for 20 minutes. 14.350 μL of buffer solution RW1 (included in the kit) was added to the adsorption column, and the column was centrifuged at 14,000×g for 30 seconds, and the waste liquid in the recovery cannula was discarded. 15. 500 μL of absolute ethanol-containing buffer RPE (obtained in step 3) was added to the adsorption column, centrifuged at 14,000 × g for 30 seconds, and the waste liquid in the collection cannula was discarded. 16. Add 500 μL of 80% ethanol (obtained in step 2) to the adsorption column, centrifuge at 14,000 × g for 30 seconds, and discard the waste liquid from the recovery cannula. 17. The empty adsorption column was inserted into the collection cannula for 14,000 x g centrifugation for 2 minutes, after which the collection cannula was discarded. 18. The adsorption column was inserted into a new 1.5 mL RNase-free EP tube, the tube was opened, and the tube was allowed to dry in air for 1 minute. 19.21 μL of RNase-free water was carefully added to the center of the adsorption column, incubated at room temperature for 1 minute, and centrifuged at 17,000 × g for 2 minutes to recover the liquid, i.e., RNA. 20.1 μL of RNA was collected and quantified using the Qubit RNA HS kit. The quantified RNA was stored in a refrigerator at -80°C for further use.
[0146] Example 2 RNA reverse transcription and cDNA amplification 1. Experimental preparation: Both hands, the pipetter, and the super clean bench tabletop were cleaned with an RNA enzyme scavenger and a nucleic acid remover. RNase- and nucleic acid-free tips, 1.5 mL EP tubes, and 0.2 mL PCR tubes were prepared in the super clean bench and exposed to ultraviolet light for 30 minutes. (Steps 2 to 11 are performed in an RNase- and nucleic acid-free super clean bench.) 2. The RNA sample and lysis buffer (provided in the kit) were placed on ice, thawed, vortexed, centrifuged in a flash, and then placed on ice for further use. 3.2 μL of RNA was collected and placed in a 1.5 mL EP tube for each sample. The RNA samples were diluted to approximately 5 ng / μL with RNase-free water according to the measured RNA concentration of the sample. 4. 1.5 μL lysis buffer and 3.5 μL diluted RNA sample were added sequentially to a 0.2 mL PCR tube. 5. For the reverse transcription negative control (RT-NC), replace the RNA sample from step 4 with 3.5 μL RNase-free water. 6.13.3μL * The RT buffer (reaction section + pipetting loss) was collected and placed into a new 0.2 mL PCR tube (the volume of each PCR tube was 50 μL or less). 7. After pre-heating the PCR amplifier, place the PCR tubes from steps 4-6 into the PCR amplifier and incubate at 72°C for 3 minutes. 8. Immediately place the PCR tube on ice, incubate for at least 2 minutes, and then flash centrifuge. 9. The RT enzyme mix (provided in the kit) was removed from the -20°C refrigerator, flash centrifuged, avoiding shaking, and placed on ice for further use. 10.1.7μL * (Reaction section + pipetting loss) After incubation at 72°C, the RT enzyme mixture was added to the RT buffer (step 6), the mixture was mixed evenly by slight pipetting, and then placed on ice for use. The mixture was marked "RT mix." Absorb 11.15 µL of RT mixture (step 10) and add it to the PCR tube containing the RNA sample or RT-NC (steps 4-5), then pipette slightly and mix evenly, centrifuge in a flash, and place on ice. 12. Samples were incubated on a preheated PCR amplifier, conditions shown in Table 1.
[0147] [Table 3]
[0148] 13. The PCR mix (included in the kit) was removed from the refrigerator at -20°C. The PCR mix was placed on ice, thawed, inverted, mixed evenly, centrifuged in a flash, placed on ice, and used further. 14.30 μL of PCR mixture was added to each reverse transcription reaction product, mixed slightly with a pipette, mixed evenly, centrifuged in a flash, and then placed on ice for further use. 15. A reverse transcription negative control (RT-NC) was prepared by replacing the reverse transcription reaction product from step 3 with 20 μL of RNase-free water. 16. Samples were incubated on a preheated PCR amplifier, the conditions of which are shown in Table 2.
[0149] [Table 4]
[0150] Note: The number of amplification cycles can be increased or decreased appropriately depending on the sample; advice on adaptive adjustments is provided in Table 3.
[0151] [Table 5]
[0152] Example 3 Purification of amplification products 1. Experimental preparation: Ensure that the experimental area is separate for the reverse transcription and amplification steps. Clean both hands, pipettors, and the experimental table with a nucleic acid remover. 2.4 volumes of absolute ethanol was added to the buffer RPE, inverted and mixed evenly for further use. 3. The PCR product was placed on ice for 2 minutes and then centrifuged in a flash for further use. 4. 50 μL of DNA binding buffer and 50 μL of PCR product were added consecutively to a 1.5 mL EP tube. Then, the mixture was vortexed evenly and centrifuged in a flash. 5. An adsorption column was inserted into the collection cannula, and 300 μL of the mixed solution from step 4 was transferred to the adsorption column. 6. Centrifugation was carried out at 14,000 x g for 30 seconds. 7. 200 μL of ethanol-containing washing buffer was added to the adsorption column and centrifuged at 14,000 × g for 30 seconds. 8. Repeat step 7 once and discard the waste liquid from the collection cannula. 9. The adsorption column was inserted into the collection cannula and centrifuged at 14,000 x g for 2 minutes, after which the collection cannula was discarded. 10. The adsorption column was inserted into a new 1.5 mL EP tube, opened, and allowed to dry in air for 1 minute. 11.30 μL of elution buffer (provided by DNA Clean & Concentrator-5) was carefully added to the center of the adsorption column, incubated at room temperature for 1 minute, and centrifuged at 17,000 × g for 2 minutes to collect the liquid, i.e., the RNA amplification product. 12.1 μL of RNA was collected and quantified using the Qubit DNA HS Kit. The negative control RT-NC, used during reverse transcription, should have a post-amplification concentration of less than 2 ng / μL, and the negative control PCR-NC, used during PCR, should have a post-amplification concentration of less than 0.4 ng / μL. The cDNA amplification product concentration of the sample should exceed 40 ng / μL. Subsequent sequencing steps were not performed on the RT-NC and PCR-NC samples.
[0153] Example 4 Quality control of amplification products 1 μl of the purified cDNA amplification product was taken and diluted appropriately for detection. The instruction manual is shown in the high-sensitivity DNA chip instruction manual, and the results are shown in Figure 1.
[0154] In general, the cDNA amplification products of the samples were distributed within 400–10,000 bp, with the main peak located around 2,000 bp, as shown in Figure 1. The cDNA used in this application complied with the quality requirements.
[0155] Example 5 Transposition and Library Construction 1. DNA Fragmentation (1) The fragmentation buffer was removed from -20°C, thawed at room temperature, shaken evenly, centrifuged in a flash, and then waited. The reaction system was shown in Table 4 according to the number of samples (N).
[0156] [Table 6]
[0157] (2) 9.5 μL of the above mixed fragment solution was taken and subpackaged into a 0.2 mL PCR tube, followed by flash centrifugation. 0.5 μL (approximately 10 ng) of MALBAC amplified product was taken and added to each PCR tube containing the 9.5 μL mixed fragment solution from the previous step. The mixed solution was then vortexed to mix evenly and then flash centrifugation was performed. (3) Place the prepared reaction system into the PCR amplifier, select "On, 105°C" for "heated lid", screw down the heated lid, and follow the reaction procedure shown in Table 5.
[0158] [Table 7]
[0159] 2. Library Enrichment (1) The amplification buffer was removed from -20°C, thawed at room temperature, shaken evenly, centrifuged with a flash, and then waited. The reaction system was shown in Table 6 according to the number of samples (N).
[0160] [Table 8]
[0161] (2) The mixture was vortexed to mix well and centrifuged in a flash. (3) 12 μL of each mixed amplification solution prepared above was taken and added to the fragmentation product in "Step 1.1." 3 μL of tag primer was added to each of the above reactions, vortexed, mixed well, and centrifuged in a flash. (4) The serial number of the tag primer corresponding to each sample was recorded (Note: there are 24 tag primers, one tag primer was added to each reaction, and samples in the same running batch should have different, non-overlapping tag primers). (5) Place the prepared reaction system into the PCR amplifier, select "On, 105°C" for "heated lid", screw down the lid, and follow the reaction procedure shown in Table 7.
[0162] [Table 9]
[0163] 3. Library Purification (1) Remove the magnetic beads from 4°C and place them at room temperature to equilibrate. Vortex the magnetic beads for 20 seconds to thoroughly blend them into a homogenous solution. (2) 20 μL of each constructed library was placed in a new 1.5 mL centrifuge tube, and 0.6× resuspended magnetic beads were added (for example, if the initial library volume was 20 μL, 12 μL of magnetic beads was added). The tube was vortexed thoroughly, left at room temperature for 5 minutes, and then centrifuged in a flash. (3) Centrifugation in a flash: Place the centrifuge tube on a magnetic frame to separate the magnetic beads from the supernatant for about 5 minutes, and when the solution is clear, place the centrifuge tube on the magnetic frame, carefully open the tube cap to prevent spillage, and then carefully transfer the supernatant to a new 1.5 mL centrifuge tube (not to absorb the magnetic beads), and then discard the magnetic beads (do not discard the supernatant). (4) Magnetic beads were resuspended in the supernatant in a volume 0.15 times the volume of the initial library (for example, if the volume of the initial library was 20 μL, 3 μL of magnetic beads were added), vortexed, mixed uniformly, centrifuged in a flash, and left at room temperature for 5 minutes. (5) The centrifuge tube was placed on a magnetic frame to separate the magnetic beads from the supernatant. After about 5 minutes, the liquid became clear, and the supernatant was carefully absorbed and discarded (careful not to discard the magnetic beads, and pipetting was performed). (6) The centrifuge tube was fixed on a magnetic frame, and approximately 200 μL of freshly prepared 80% ethanol was added to the centrifuge tube (Note: The ethanol was added carefully along with the tube wall to prevent the magnetic beads from scattering, so that the magnetic beads were immersed in the ethanol). After leaving it at room temperature for 30 seconds, the supernatant was carefully removed. (7) The previous step was repeated once. (8) The centrifuge tube was fixed on a magnetic frame and left at room temperature for about 10 minutes to allow the ethanol to completely evaporate. (9) The centrifuge tube was removed from the magnetic frame, 17.5 μL of eluent was added, and the mixture was vortexed to completely resuspend the magnetic beads. The mixture was then flash centrifuged and left at room temperature for 5 minutes. The centrifuge tube was then placed on the magnetic frame to separate the magnetic beads from the liquid. After about 5 minutes, the solution became clear, and 15 μL of the supernatant was carefully absorbed into a new centrifuge tube (taking care not to absorb the magnetic beads) and stored at -20°C.
[0164] 4. Library Quality Control The purified libraries could be individually quantified using the Qubit dsDNA HS Assay Kit, and their concentrations were generally above 5 ng / ul (real-time quantitative PCR can be performed to confirm qualification for high-quality sequencing results).
[0165] Example 6 Sequencing Run See the experimental procedure description in the Illumina test kit. Sequencing strategy: Both single-end and double-end methods were available, with read lengths over 45 nt and 2.5 million reads guaranteed.
[0166] Example 7 Data Processing Steps Using expression profiles from samples from volunteers with defined clinical outcomes as training data, an analytical model was constructed using machine learning. Then, location data was input and preliminary determination was performed using the analytical model as follows. 1. Use a supervised learning method (as shown in Figure 2), in which the "training data" used has labels that are expression profile features of endometrial tissue, uterine fluid or other reproductive endocrinology-related fluids or exfoliation under various implantation conditions. 2. The grouping criteria for determining the findings as "training data" or "training set" were healthy Chinese women with different natural cycles, without a history of or primary infertility, and with a body mass index of 19-25 kg / m 2After obtaining the RNA-seq data of the three labels, FPKM was used for normalization by gene length and sequencing depth to eliminate the influence of sequencing depth. During the normalization process, the expression levels of different genes from samples with different labels were compared, thereby analyzing the differential expression of genes. Specifically, the following was done: all genes with FPKM>0 in each sample were found, and the intersection of differentially expressed genes between "preimplantation period" vs. "implantation period", "preimplantation period" vs. "postimplantation period", and "implantation period" vs. "postimplantation period" was screened from all genes belonging to the training set samples with FPKM>0 (any gene with p-value<0.05 and fold change>2 or fold change<0.5 met the selection criteria). Thus, 12,734 differentially expressed genes were obtained, and were named ENSG00000000003, ENSG00000 104881, ENSG00000128928, ENSG00000151116, ENSG00000171222, ENSG00000198961, ENSG0000 0261732, ENSG00000000419, ENSG00000104883, ENSG00000128944, ENSG00000151117, ENSG0000 ENSG00000198964, ENSG00000261760, etc. The above gene codes definitively correspond only to genes in the NCBI database (https: / / www.ncbi.nlm.nih.gov / ). Although a large number of differently expressed genes are available, no detailed description is provided herein, limiting the length of the specification. 3. In the supervised learning of the present invention, a model was constructed using training data labeled with pre-implantation cycles, implantation window cycles, and post-implantation cycles. The model obtained by training can be used to predict the implantation status of unknown data (referring to new samples). For example, a uterine fluid expression profile from a case was newly input, and the machine learning model was used to determine the implantation status of the sample (as shown in Figure 3). 4. Using SVM, samples corresponding to the labels of 3 were used as input variables for building the training set. The script is as follows: library(e1071) svm.model<-svm(data.class~.,data4,kernel='linear') summary(svm.model) table(data$class,predict(svm.model,data4,type="data.class")) mydata=read.table(file.choose(),header=T,row.names=1) mydata2=log2(mydata+1) predict(svm.model,mydata2) table(shdata3$shdata.class,predict(svm.model,shdata3,type="shdata.class")) ## table(testdata$data.class,predict(svm.model,testdata,type="data.class")) shdata=read.table(file.choose(),header=T,row.names=1) shdata2=log2(shdata[,-length(shdata[1,])]+1) shdata3=data.frame(shdata2,shdata$class) predict(svm.model,shdata3) table(shdata3$shdata.class,predict(svm.model,shdata3,type="shdata.class")) 5. According to the results of SVM, the implantation status of the unknown sample was defined, the endometrial implantation of the sample was determined, and the embryo transfer (as shown in Figure 4) was directed according to the implantation status. To obtain a good pregnancy outcome, the sampling time could be adjusted multiple times in several cycles, and the implantation test was performed (as shown in Figure 5).
[0167] Example 8 Clinical validation The analytical model constructed in this invention was used to predict the risk of infertility in different infertile individuals aged 23 to 39 years. Based on the SVM results, the implantation status of unknown samples was defined to determine the endometrial implantation status of the samples. Embryo transfer was instructed based on the implantation status, and pregnancy outcomes were recorded. The results are shown in Tables 8-1, 8-2, and 8-3.
[0168] [Table 10-1]
[0169] [Table 10-2]
[0170] [Table 10-3]
[0171] [Table 11-1]
[0172] [Table 11-2]
[0173] [Table 11-3]
[0174] [Table 12-1]
[0175] [Table 12-2]
[0176] [Table 12-3]
[0177] The sample period, sample type, medical history, and infertility type represent the clinical information of the sample; the RNA concentration, cDNA concentration, sequencing data size, unique mapping ratio, and exon ratio represent the sequencing quality control information; the support vector classification (SVC) represents the implantation status determined by machine learning; and the implantation method refers to the adjustment of the implantation time based on the results of the machine learning analysis. Tables 8-1, 8-2, and 8-3 show that clinical validation showed that 19 infertile women achieved successful pregnancy by using the model of the present invention to predict the embryo implantation time, demonstrating that the model of the present invention has very high accuracy and is therefore useful for promoting medical advances.
[0178] The results of the determination of the markers of the present invention are shown in Table 9.
[0179] [Table 13-1]
[0180] [Table 13-2]
[0181] [Table 13-3]
[0182]
Table 13-4
[0183]
Table 13-5
[0184]
Table 13-6
[0185]
Table 13-7
[0186]
Table 13-8
[0187]
Table 13-9
[0188]
Table 13-10
[0189]
Table 13-11
[0190]
Table 13-12
[0191]
Table 13-13
[0192]
Table 13-14
[0193]
Table 13-15
[0194]
Table 13-16
[0195]
Table 13-17
[0196]
Table 13-18
[0197]
Table 13-19
[0198]
Table 13-20
[0199]
Table 13-21
[0200]
Table 13-22
[0201]
Table 13-23
[0202]
Table 13-24
[0203]
Table 13-25
[0204]
Table 13-26
[0205]
Table 13-27
[0206]
Table 13-28
[0207]
Table 13-29
[0208]
Table 13-30
[0209]
Table 13-31
[0210]
Table 13-32
[0211]
Table 13-33
[0212]
Table 13-34
[0213]
Table 13-35
[0214]
Table 13-36
[0215]
Table 13-37
[0216]
Table 13-38
[0217]
Table 13-39
[0218]
Table 13-40
[0219]
Table 13-41
[0220]
Table 13-42
[0221]
Table 13-43
[0222]
Table 13-44
[0223]
Table 13-45
[0224]
Table 13-46
[0225]
Table 13-47
[0226]
Table 13-48
[0227]
Table 13-49
[0228]
Table 13-50
[0229]
Table 13-51
[0230]
Table 13-52
[0231]
Table 13-53
[0232]
Table 13-54
[0233] [Table 13-55]
[0234] [Table 13-56]
[0235] [Table 13-57]
[0236] [Table 13-58]
[0237] [Table 13-59]
[0238] [Table 13-60]
[0239] [Table 13-61]
[0240] [Table 13-62]
[0241] As shown in Table 9, in the present invention, only a single marker is used, but as long as these markers exceed a predetermined value, they can be used as useful auxiliary determination or diagnostic information of endometrial implantation, and therefore, particularly for early and / or auxiliary diagnosis.
[0242] Furthermore, by making a comprehensive judgment using the multiple markers shown in Table 9, the error rate (only 17.5%) can be further significantly reduced, improving the accuracy of the judgment.
[0243] When multiple markers (including an additional 5–200 genes) are sampled, the error rate can be as low as 16.5%.
[0244] When the number of markers used is 10,000, the error rate is 6.7%.
[0245] For error rate, clinical outcome is the gold standard, and the date of successful embryo transfer is the implantation period. Accuracy rate = number of cases in which the gestational age was determined by method / total number of successful embryo transfers; error rate = number of cases in which the gestational age was not determined by method / number of successful embryo transfers.
[0246] It can therefore be seen from the above that the markers of the present invention have very high predictive value, particularly when used in combination, which can further reduce the error rate in the determination of endometrial implantation status, thereby improving the accuracy of the determination.
[0247] Furthermore, the markers in Table 9 were further screened according to Table 11 to obtain 40 markers, 147 markers and 259 markers with very low error rates and high accuracy effects, where blank refers to an analytical value of <4.
[0248] [Table 14]
[0249] Furthermore, it can be seen from Table 10 that the genes in Table 9 of the present invention are core genes. Based on this, the newly added genes improve the accuracy but have low weight.
[0250] [Table 15]
[0251] In summary, the present invention relies on a highly sensitive RNA reverse transcription and cDNA amplification process and is based on RNA-seq sequencing, thereby obtaining multiple expression profile information from endometrium, uterine fluid, or other reproductive endocrinology-related body fluids or exfoliation samples from patients. The sampling period and sampling method of these samples are specifically selected to obtain different expression profile features, which are then subjected to ultra-high-dimensional classification and typing using bioinformatics, statistics, and machine learning methods. Depending on the type, the ER status is determined. In the present invention, the expression profile features are used in a model to determine the ER status before embryo implantation, thus accurately determining the endometrial implantation window period.
[0252] The applicant states that the present invention describes detailed methods through the above examples, but is not limited to the above detailed methods. That is, it does not mean that the present invention must depend on the above detailed methods for implementation. Those skilled in the art should know that improvements to the present invention corresponding to each raw material of the product of the present invention, the addition of auxiliary components, the selection of specific means, etc. are included in the protection scope and disclosure scope of the present invention.
[0253] All documents mentioned herein are incorporated by reference in this application to the same extent as if each document were incorporated by reference separately. It should be understood that those skilled in the art, after reading the above teachings of the present invention, may make any changes or modifications to the present invention, and that equivalent forms thereof are also within the scope defined by the claims of this application.
Claims
1. 1. A method for determining endometrial implantation, comprising the steps of: (a) providing a sample; (b) measuring the expression level of an endometrial implantation-associated gene in the sample; (c) comparing the expression level of the endometrial implantation-related gene obtained in step (b) with a predetermined value, thereby determining endometrial implantation. The above method, comprising:
2. The method of claim 1, wherein the expression level of the endometrial implantation-related gene obtained in step (b) is higher than a predetermined value, which indicates the presence of endometrial implantation.
3. 2. The method of claim 1, wherein the samples are from the following periods: LH+n, LH+n+2, LH+n+4, where n is 3-7, preferably n is 4-6.
4. 2. The method of claim 1, wherein the endometrial implantation-associated genes comprise at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A: 【Table 1-1】 【Table 1-2】 【Table 1-3】 【Table 1-4】 【Table 1-5】 【Table 1-6】 【Table 1-7】 【Table 1-8】 【Table 1-9】 【Table 1-10】 【Table 1-11】 【Table 1-12】 【Table 1-13】 【Table 1-14】 【Table 1-15】 【Table 1-16】 【Table 1-17】
5. A set of biomarkers that comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A.
6. A combination of reagents used to determine endometrial implantation status, comprising a reagent for detecting each biomarker in the set of claim 5.
7. A kit comprising the set according to claim 5 and / or the combination of reagents according to claim 6.
8. Use of a set of biomarkers for the manufacture of a kit, said kit being used to assess the endometrial implantation status of a subject to be detected, said set of biomarkers comprising at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A.
9. 1. A method for assessing endometrial implantation status of a test subject, comprising: (1) providing a sample from a subject and detecting the level of each biomarker in a sample set, the level comprising at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A; and (2) comparing the level detected in (1) with a predetermined value; The above method, comprising:
10. 1. A system for assessing endometrial implantation status in a subject, comprising: (a) an endometrial implantation status feature input module, the input module being used to input the endometrial implantation status feature of a specimen to be detected, wherein the endometrial implantation status feature of the specimen to be detected comprises at least 70%, preferably at least 80%, more preferably at least 90%, more preferably at least 95% of the genes selected from Table A; (b) an endometrial implantation state determination processing module, wherein the processing module performs an evaluation process on the input characteristics of the endometrial implantation state according to a preset criterion, thereby obtaining a score of the endometrial implantation state; and further compares the score of the endometrial implantation state with a predetermined value, thereby obtaining an auxiliary diagnosis result, wherein if the score of the endometrial implantation state is higher than the predetermined value, it is suggested that the subject has endometrial implantation; and (c) an auxiliary diagnostic result output module, the output module being used to output the auxiliary diagnostic result; The system described above.