Method, model and marker for non-invasive determination of endometrial receptivity

By non-invasively obtaining the adhesion trace tissue of embryo pre-transplantation or transplant tube walls, RNA extraction and sequencing are performed, and a receptive prediction model is established in combination with machine learning algorithms, which solves the error and invasive problems of endometrial receptivity in the existing technology, and achieves efficient and accurate non-invasive prediction.

CN114517232BActive Publication Date: 2025-06-17YIKON GENOMICS (SUZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210253498.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-06-17
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The methods used in the prior art to determine endometrial receptivity are incorrect and are an invasive biopsy method, resulting in low maternal injury and compliance and prolonging the pregnancy time.

Method used

The adhesion trace tissue of embryo pre-transplant or transplant tube walls is obtained by non-invasively, cell extraction and RNA extraction are performed, cDNA library is constructed and sequenced through next-generation sequencing technology, and a receptive prediction model is established in combination with machine learning algorithms.

Benefits of technology

Non-invasive samples were obtained, which reduced maternal damage and pain, improved the accuracy and efficiency of endometrial receptivity prediction, and shortened pregnancy time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114517232B_ABST
    Figure CN114517232B_ABST
Patent Text Reader

Abstract

Methods, models and markers for non-invasive determination of endometrial receptivity. The present invention provides a method for predicting endometrial receptivity using samples obtained by non-invasive sampling, a method for establishing a prediction model using the samples, the established prediction model, markers associated with endometrial receptivity in the samples, and corresponding kits.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to the fields of biology and clinical assisted reproduction, and more particularly to a non-invasive method for determining endometrial receptivity, and markers, prediction models, and corresponding kits and systems used in this method. Background Art

[0002] During the human reproductive process, the fertilized egg localizes, adheres, and implants in the maternal uterus and finally develops into a mature fetus. The implantation process has an important impact on successful pregnancy. Successful clinical pregnancy requires not only high-quality embryos but also good endometrial receptivity (ER) and synchronous development of the endometrium and the embryo. Endometrial receptivity is a physiological phenomenon, referring to the ability of the endometrium to accept the embryo. Only during a short specific period can the endometrium allow embryo implantation, and this period is called the "implantation window period". For adult women, it corresponds to the 20th - 24th day of the menstrual cycle or 6 - 8 days after ovulation.

[0003] The specific mechanism of endometrial receptivity is not yet fully understood at present, but it is clear that the endometrium being in the non-receptive phase is one of the important reasons for embryo implantation failure in in vitro fertilization-embryo transfer (IVF-ET). The success rate of embryo transfer is very important, directly related to whether all the previous efforts and costs can succeed or end in failure. However, the traditional method using the menstrual cycle window period is not always feasible. Even for a considerable number of subjects, such as women with repeated implantation failures or women suffering from other secondary infertility, the time node of the implantation window period is very inaccurate. If the implantation window period is still calculated according to the menstrual cycle or the ovulation day, there will be a great risk of transplantation failure.

[0004] Clinically, there are already various methods for judging endometrial receptivity, such as transvaginal ultrasound, serum estrogen and progesterone levels, endometrial biopsy, etc. to judge the time of the endometrial implantation window, but there are certain errors in all of them. With the progress of molecular biology, especially the completion of the Human Genome Project, the clinical application of using genes as molecular markers to assist in judging various physiological and pathological states has become more and more mature. However, in the prediction of endometrial receptivity, all existing methods have a significant drawback, that is, the samples they use are all obtained from biopsy sampling of endometrial tissue. This invasive method brings unnecessary damage to the mother and also has potential unstable factors, thus reducing the compliance of the subjects. Moreover, the damage causes the biopsy cycle to be unable to perform embryo transfer, prolonging the time for the subjects to achieve pregnancy and the ART cycle. In addition, the additional operation itself is also a way that consumes resources and increases the pain of the subjects.

[0005] Researchers have been pursuing simpler and less invasive examination methods, and the prediction of endometrial receptivity is no exception. Clinically, there is an urgent need for an efficient and non-invasive method for diagnosing endometrial receptivity. Brief Introduction of the Invention

[0006] The present invention provides a method for non-invasively sampling and judging endometrial receptivity.

[0007] In the first aspect, the present invention provides a method for obtaining a living sample to be analyzed in a non-invasive manner. In a specific embodiment, the non-invasive manner is that a trace amount of tissue adheres to the wall of the transplantation tube during embryo pre-transplantation. In another specific embodiment, the non-invasive manner is that a trace amount of tissue adheres to the wall of the transplantation tube during embryo transplantation.

[0008] In one embodiment, the number of cells in the sample obtained by the non-invasive method is less than 1000. In a more specific embodiment, the number of cells in the sample obtained by the non-invasive method is less than 900, 800, 700, 600, 500, 400, 300 or 200. In a preferred embodiment, the number of cells in the sample obtained by the non-invasive method is between 400 and 300.

[0009] In a specific embodiment, the method of the present invention further includes transferring the cells and tissues on the wall of the tube to a sample collection tube by the method of intercepting and flushing. In a specific embodiment, in a preferred embodiment, a tissue sample preservation solution is pre-added to the sample collection tube.

[0010] In one embodiment, the method of the present invention further includes extracting total RNA from a sample. In a specific embodiment, the extraction process uses the Yikon MALBAC Platinum MicroRNA Amplification Kit (KT110700724) and is carried out according to the instructions provided by the manufacturer. In another specific embodiment, the extraction process includes the following steps: a) Centrifuge the sample at high speed (1600g for 2 minutes for the first time; 3000g for 3 minutes for the second time; 13000rpm for 5 minutes for the third time) to precipitate all cells. b) Remove the supernatant, add 200 μL × 2 tissue sample preservation solution to resuspend the cells, and centrifuge at 3000g for 5 minutes; 200 μL × 1 time, leaving about 50 μL. c) Add 300 μL of cell lysate and shake vigorously for 30 seconds to 1 minute; centrifuge instantaneously and discard the supernatant. d) Add 350 μL of freshly prepared 70% ethanol and mix well by shaking; centrifuge instantaneously and discard the supernatant. e) Add 650 μL of the above mixed liquid to the adsorption column and centrifuge at 13000 × g for 1 minute, discarding the centrifugation waste liquid. f) Add 500 μL of washing buffer to the adsorption column membrane and centrifuge at 13000 × g for 1 minute, discarding the centrifugation waste liquid. g) Carefully drip 100 μL of DNA digesting enzyme onto the adsorption column membrane and incubate at 25 degrees for 15 minutes. h) Add 500 μL of washing buffer and centrifuge at 13000 × g for 1 minute, discarding the centrifugation waste liquid; add 500 μL of washing buffer to the adsorption column again and centrifuge at 13000 × g for 1 minute, discarding the centrifugation waste liquid. i) Add 700 μL of 80% ethanol to the tube and centrifuge at 13000 × g for 2 minutes, discarding the centrifugation waste liquid; centrifuge at 13000 × g for 1 minute again and discard the collection tube. j) Insert the adsorption column into a 1.5 mL centrifuge tube, open the cap and air dry for 1 minute; add 20 μL of DEPC water, incubate at 25 degrees for 1 minute, centrifuge at 13000 × g for 2 minutes, and repeat once, collecting a total of 40 μL of eluate, which is the total RNA. In a more specific embodiment, almost all of the RNA in the trace sample obtained by non-invasive sampling is extracted, thereby providing a sufficient amount of purified RNA nucleic acid sample for subsequent detection.

[0011] In one embodiment, the method of the present invention further includes reverse transcribing the mRNA in the RNA nucleic acid sample to generate a cDNA library. In a specific embodiment, the reverse transcription uses the Yikon MALBAC Platinum MicroRNA Amplification Kit (KT110700724) and is carried out according to the instructions provided by the manufacturer.

[0012] In one embodiment, the method of the present invention further includes further processing the cDNA library to construct a sequencing library. In a specific embodiment, the further processing uses the Yikon Fragmentation Kit (KT100804248) and the Gene Sequencing Library Kit (KT100804048) and is carried out according to the instructions provided by the manufacturer.

[0013] In one embodiment, the method of the present invention further includes performing sequencing on the resulting library. In a specific embodiment, the sequencing method is next-generation sequencing (NGS). In another specific embodiment, the sequencing method is Sanger sequencing or third-generation sequencing (e.g., nanopore sequencing). In a preferred embodiment of the present invention, the Illumina Nextseq sequencing system and the supporting Next-seq High-Output Kit are used.

[0014] In a second aspect, the present invention provides a method for establishing a receptivity prediction model. In one embodiment, the present invention provides a method for analyzing sequencing results using a machine learning algorithm and establishing a receptivity prediction model.

[0015] In a specific embodiment, the machine learning algorithm analysis includes the following steps: I) Classify the samples into receptive and non-receptive categories according to the gold standard (actual clinical outcome), preferably split them into two sets, a training set and a test set, and use the sample data of the training set for the next step; II) Denoise and debias all the data of the training set samples to generate fragment count alignment files for each gene; III) Screen and collect differentially expressed genes and use them as features to be identified; IV) Convert the data to TPM as the normalized expression data of each gene to eliminate the influence of the difference in the amount of sample data off the machine; V) Input the expression data of the features to be identified (genes) into the machine learning algorithm, and perform supervised learning by the machine learning algorithm to obtain the threshold, sensitivity, and importance of each feature for judging receptivity, thereby obtaining a receptivity prediction model. Preferably, verify the accuracy of the established model in the test set.

[0016] In one embodiment, the model includes one or more sub-models, and different sub-models contain all or different parts of the aforementioned differentially expressed genes.

[0017] In a specific embodiment, step II) specifically includes:

[0018] a) Perform the following processing on the sequencing data off the machine for each sample

[0019] i. Process the data off the machine with trimmomatic (version 0.33) to remove adapter sequences and filter low-quality reads to generate clean sequenced data

[0020] ii. Align the clean sequencing data generated in the previous step to the human reference genome (version GRCh38) using HISAT2 (version 2.0.5) to generate an alignment file.

[0021] iii. Sort the alignment file generated in the previous step and mark the repetitive sequences to generate the final alignment file.

[0022] iv. Process the alignment file generated in the previous step using htseq_count. This step also requires a gene annotation file, the Homo_sapiens.GRCh38.84.gtf file, which is downloaded from Ensembl and contains information such as the name and location coordinates of each gene. Htseq_count can count how many sequences are in the region of each gene based on this information (for paired-end sequencing, reads belonging to the same fragment are only counted once), and generate a genecount file.

[0023] b) Combine the genecount files of all samples in a batch to generate a total genecount file, where each row represents a sample, each column represents a gene, and the data is the number of fragments of the gene in the sample.

[0024] In a specific embodiment, the step III) specifically includes the following selection methods:

[0025] i. Exclude genes related to rRNA and mitochondria with greater subsequent interference.

[0026] ii. For the training set, perform gene differential expression analysis to screen out the differentially expressed genes in the two types of samples, obtaining 3193 genes.

[0027] i ii. Further perform variable screening on the 3193 genes screened in the previous step, obtaining a total of 250 genes as candidate features for the next model training. The detailed list is shown in Table 2.

[0028] In a more specific embodiment, the variable screening is to select genes with FDR < 0.001 from the differentially expressed genes, obtaining 1270 genes. Use the cpm function of edgeR to output the logTMM value of each gene in each sample, calculate the pearson correlation between these genes pairwise, and then screen the genes according to the following steps.

[0029] 1) Sort the genes in ascending order of FDR.

[0030] 2) Mark the top gene in the list as selected.

[0031] 3) Mark genes with a correlation greater than or equal to 0.6 with any of the selected genes as excluded;

[0032] 4) Repeat steps 2) and 3) until the end of the list;

[0033] 5) Construct a model with the genes obtained by screening, evaluate the accuracy using the 10-fold cross-validation method, and sequentially delete genes that have no effect on the accuracy of the model. Finally, 250 genes are obtained.

[0034] In a more specific embodiment, step ii. above uses edgeR (version 3.34.0) for differential gene expression analysis, and the screening conditions are logFoldChange > 0, FDR < 0.05, including the following steps:

[0035] a) edgeR requires two input files. The first is the genecount file, which contains the number of reads for each gene in each sample, and the second is the sample information file, which contains the category to which each sample belongs;

[0036] b) After edgeR reads in the input files, it first normalizes the number of reads for each gene in each sample. When normalizing, genes with stable expression in all samples are selected;

[0037] c) edgeR then searches for genes with differential expression in different categories based on the category information of the samples. The significance of the difference is represented by the p-value after FDR correction. The smaller the p-value, the more significant the difference.

[0038] In a specific embodiment, step IV) specifically includes converting the genecount data to TPM to exclude the influence of the sample throughput data volume. The calculation formula is as follows:

[0039]

[0040] where Ni is the reads count on the i-th gene and Li is the length of the i-th gene.

[0041] In a specific embodiment, the machine learning algorithm in step V) is the random forest algorithm.

[0042] In one embodiment, the method of the present invention further includes evaluating and / or validating indicators such as sensitivity, specificity, accuracy, and ROC of the established receptivity prediction model in the split independent test set.

[0043] In a specific embodiment, the method of the present invention further includes validating the established receptivity prediction model in at least two test sets.

[0044] In a third aspect, the present invention also provides a biomarker set for predicting endometrial receptivity to guide embryo transfer.

[0045] In one embodiment, the biomarker set of the present invention includes some or all of the differentially expressed genes identified in step III) above, for example, any number between more than 5 and less than or equal to 250. In a more specific embodiment, in order to accurately predict endometrial receptivity to guide embryo transfer, the biomarker set of the present invention includes at least 10 gene members. In an even more specific embodiment, the biomarker set of the present invention includes at least 11 gene members, at least 12 gene members, at least 13 gene members, at least 14 gene members, at least 15 gene members, at least 16 gene members, at least 17 gene members, at least 18 gene members, at least 19 gene members, at least 20 gene members, at least 25 gene members, at least 30 gene members, at least 35 gene members, at least 40 gene members, at least 45 gene members, at least 50 gene members, at least 55 gene members, at least 60 gene members, at least 65 gene members, at least 70 gene members, at least 80 gene members, at least 90 gene members, at least 100 gene members, at least 110 gene members, at least 120 gene members, at least 130 gene members, at least 140 gene members, at least 150 gene members, at least 160 gene members, at least 170 gene members, at least 180 gene members, at least 190 gene members, at least 200 gene members, at least 210 gene members, at least 220 gene members, at least 230 gene members, at least 240 gene members, or at least 250 gene members. In one embodiment, the biomarker set of the present invention includes 5 genes, 6 genes, 7 genes, 8 genes, 9 genes, 10 genes, 15 genes, 20 genes, 25 genes, 30 genes, 40 genes, 50 genes, 60 genes, 70 genes, 80 genes, 90 genes, 100 genes, 110 genes, 120 genes, 130 genes, 140 genes, 150 genes, 160 genes, 170 genes, 180 genes, 190 genes, 200 genes, 210 genes, 220 genes, 230 genes, 240 genes, or 250 genes with the highest weights in the receptivity prediction model established in step V) above.

[0046] In one embodiment, the biomarker set of the present invention comprises at least the following genes: ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6.

[0047] In a fourth aspect, the present invention provides uses and methods based on the biomarker set.

[0048] In one embodiment, the present invention provides the use of the above-mentioned biomarker set or its specific detection tool for preparing an assay composition, assay reagent, and / or assay device for detecting and determining the endometrial receptivity status of a subject, wherein the assay composition, assay reagent, and / or assay device can detect the expression level of each member of the above-mentioned biomarker set in a sample from the subject; and calculate the expression level through the receptivity prediction model of the present invention, and the obtained result is used to judge and diagnose the endometrial receptivity status.

[0049] In one embodiment, the present invention provides an endometrial receptivity prediction method, which includes obtaining the expression levels of all gene members of the biomarker set by using a substance or method for specifically measuring the level or amount of the biomarker. In a specific embodiment, the substance for measuring the level or amount of the biomarker is a composition, reagent, or device. In a specific embodiment, the substance for measuring the level or amount of the biomarker is a specific primer and / or probe molecule. In a more specific embodiment, the composition or reagent for measuring the level or amount of the biomarker further comprises one or more of a tag sequence, an adapter sequence, a fluorescent molecule, a pyrophosphate molecule, and a labeling molecule. In a specific embodiment, the method for measuring the level or amount of the biomarker is a PCR method. In a more specific embodiment, the PCR method is a commonly used method in the art such as real-time fluorescence quantitative PCR, isothermal amplification PCR, etc. In a specific embodiment, the method for measuring the level or amount of the biomarker is a sequencing method. In a more specific embodiment, the sequencing method is one or more of next-generation sequencing (NGS), Sanger sequencing, and third-generation sequencing (e.g., nanopore sequencing).

[0050] In one embodiment, the reagent for measuring the level or amount of the biomarker does not contain a reagent for specifically detecting genes other than the members of the biomarker set of the present invention.

[0051] In one embodiment, the method for determining the level or amount of a biomarker does not determine the level or amount of genes other than the members of the biomarker set of the present invention.

[0052] In one embodiment, the present invention provides a method for predicting endometrial receptivity, wherein the sample used in the method is obtained in a non-invasive manner, that is, the acquisition of the sample does not cause damage to the endometrium of the subject. In a preferred embodiment, the present invention provides an integrated non-invasive method for predicting endometrial receptivity. The term "integrated" means that this method does not need to be carried out separately, but is carried out simultaneously during embryo pre-transfer or embryo transfer operations. The sampling process is integrated into the above operations, so that the sampling cycle is the transplantation cycle. For example, when the transplantation fails in a cycle and the test results indicate that it is due to the endometrium, the next cycle of transplantation can be guided with reference to the test results without waiting, greatly shortening the pregnancy time of the pregnant woman.

[0053] In another aspect, the present invention provides a kit for the above method, which kit contains reagents for determining the level or amount of each biomarker in the biomarker combination of the present invention in a sample, such as specific reagents and general reagents. Brief Description of the Drawings

[0054] Figure 1 The curve graph of shows the ROC curve obtained by the feature10 sub-model in the independent test set.

[0055] Figure 2 The curve graph of shows the ROC curve obtained by the feature30 sub-model in the independent test set.

[0056] Figure 3 The curve graph of shows the ROC curve obtained by the feature50 sub-model in the independent test set.

[0057] Figure 4 The curve graph of shows the ROC curve obtained by the feature70 sub-model in the independent test set.

[0058] Figure 5 The curve graph of shows the ROC curve obtained by the feature90 sub-model in the independent test set.

[0059] Figure 6 The curve graph of shows the ROC curve obtained by the feature110 sub-model in the independent test set.

[0060] Figure 7 The curve graph of shows the ROC curve obtained by the feature130 sub-model in the independent test set.

[0061] Figure 8 The curve graph shows the ROC curve obtained by the feature150 sub - model in the independent test set.

[0062] Figure 9 The curve graph shows the ROC curve obtained by the feature170 sub - model in the independent test set.

[0063] Figure 10 The curve graph shows the ROC curve obtained by the feature190 sub - model in the independent test set.

[0064] Figure 11 The curve graph shows the ROC curve obtained by the feature210 sub - model in the independent test set.

[0065] Figure 12 The curve graph shows the ROC curve obtained by the feature230 sub - model in the independent test set.

[0066] Figure 13 The curve graph shows the ROC curve obtained by the feature250 sub - model in the independent test set. DETAILED DESCRIPTION OF THE INVENTION

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For the purposes of the present invention, the following terms are defined below.

[0068] All numerical names, such as pH, temperature, time, concentration, and molecular weight, including ranges, are approximate values that vary in increments of 0.1 (+) or (-). It should be understood that although not always explicitly stated, all numerical names are preceded by the term "about". The term "about" when used in conjunction with a numerical value means to cover numerical values within a range having a lower limit that is 5% less than the specified numerical value and an upper limit that is 5% greater than the specified numerical value.

[0069] The term "and / or" when used to connect two or more alternatives should be understood to mean any one of the alternatives or any two or more of the alternatives.

[0070] As used herein, the term "comprising" or "including" means including the recited elements, integers, or steps, but not excluding any other elements, integers, or steps. In this document, when the term "comprising" or "including" is used, unless otherwise specified, the case consisting of the recited elements, integers, or steps is also covered. For example, when referring to an antibody variable region "comprising" a specific sequence, it is also intended to cover an antibody variable region consisting of that specific sequence.

[0071] It should also be understood that, although not always explicitly stated, the reagents described herein are merely exemplary and their equivalents are known in the art.

[0072] As used herein, the term "endometrial receptivity" refers to the ability of the endometrium to accept an embryo. Only during a short, specific period does the endometrium have the greatest ability to accept an embryo, allowing embryo implantation. This period is called the receptive period and may also be referred to as the implantation window, receptive state, etc. Relatively, when in a period that mostly cannot accept embryo implantation, it is called the non-receptive period, pre-receptive period, non-receptive state, etc. Unless otherwise specified, these terms can be used interchangeably herein.

[0073] Currently, if the receptive state of the endometrium can be confirmed, in most cases, the embryo can be successfully implanted by transplanting embryos at different stages. If in the receptive state, then transplanting embryos at the D5 (5th day after fertilization) stage has a high probability of successful implantation; while in the pre-receptive state, transplanting embryos at the D3 (3rd day after fertilization) stage has a high probability of successful implantation.

[0074] The terms "nucleic acid" and "polynucleotide" are used interchangeably and refer to polymeric forms of nucleotides of any length, deoxyribonucleotides or ribonucleotides or their analogs. Polynucleotides can have any three-dimensional structure and can perform any function. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, ESTs or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides can contain modified nucleotides such as methylated nucleotides and nucleotide analogs. If present, the modification of the nucleotide structure can be carried out before or after the assembly of the polymer. The nucleotide sequence can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with labeled components. The term also refers to double-stranded and single-stranded molecules. Unless otherwise specified or required, any polynucleotide embodiment of the present invention includes the double-stranded form and each of the two complementary single-stranded forms known or predicted to constitute the double-stranded form.

[0075] The terms "marker" or "biomarker" are used interchangeably herein.

[0076] As used herein, the term "marker" refers to a class of substances that are derived from in vivo or in vitro samples and whose presence levels in the samples can be relatively easily detected using existing or conventional experimental methods and tools in the art. Such presence levels (e.g., high or low, or presence or absence) are associated with a specific physiological or pathological state. Therefore, the specific physiological or pathological state can be inferred or assisted in being inferred by obtaining the presence levels of the marker. A marker can be any substance present in the body, such as, but not limited to, nucleic acids, polysaccharides, proteins, inorganic or organic small molecules, or their polymers or hybrids (e.g., glycoproteins, phosphorylated proteins, methylated nucleic acid sequences, etc.), or their encoding genes (e.g., for proteins). In some embodiments of the present invention, the marker is a specific gene, which is reflected by changes in the expression level. Preferably, the expression level is reflected as the content of mRNA. In some embodiments of the present invention, the specific physiological state to be inferred refers to endometrial receptivity, i.e., after embryo transfer, embryo implantation will occur with a high probability. In some further embodiments, the sample is a non-invasive sample. In some preferred embodiments, the expression level of the gene is obtained by sequencing.

[0077] "Gene" refers to a polynucleotide containing at least one open reading frame (ORF) that can encode a specific polypeptide or protein after transcription and translation. The polynucleotide sequence can be used to identify larger fragments or full-length coding sequences related to the gene. Methods for isolating larger fragment sequences are known to those skilled in the art.

[0078] "Gene expression" or alternatively "gene product" refers to the nucleic acid or amino acid (e.g., peptide or polypeptide) produced when a gene is transcribed and translated.

[0079] When used in the context of polynucleotide manipulation, "probe" refers to an oligonucleotide provided as a reagent for detecting a target that may be present in a sample of interest by hybridizing with the target. Generally, the probe contains a label or a method by which a label can be attached before or after the hybridization reaction. Suitable labels include, but are not limited to, radioisotopes, fluorescent dyes, chemiluminescent compounds, dyes, and proteins including enzymes.

[0080] A "primer" is a short polynucleotide, usually having a free 3'-OH group, which binds to a target or "template" that may be present in a sample of interest by hybridizing to the target, thereby facilitating the polymerization of a polynucleotide complementary to the target. "Polymerase chain reaction" ("PCR") is a reaction that uses a "pair of primers" or "set of primers" consisting of an "upstream" and a "downstream" primer, a catalyst for polymerization such as a DNA polymerase and generally a thermostable polymerase, to replicate copies of the target polynucleotide. The PCR method is well known in the art and is taught, for example, in PCR: A Practical Approach, M. MacPherson et al., IRL Press at Oxford University Press (1991). The entire process of generating replicated copies of a polynucleotide, such as PCR or gene cloning, is collectively referred to herein as "replication". Primers can also be used as probes in hybridization reactions, such as DNA or RNA blot analysis (Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition (1989)).

[0081] As used herein, "expression" refers to the process of transcription of DNA into mRNA and / or the subsequent translation of the transcribed mRNA into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in eukaryotic cells.

[0082] When "differential expression" is applied to a gene, it refers to the production of a difference in the mRNA transcribed and / or translated from that gene or the protein product encoded by that gene. A differentially expressed gene can be overexpressed or underexpressed compared to the expression level in control cells. However, as used herein, overexpression is an increase in gene expression and is generally at least 1.25-fold or, alternatively, at least 1.5-fold or, alternatively, at least 2-fold, or alternatively, at least 3-fold or alternatively, at least 4-fold the expression detected in control counterpart cells or tissues. As used herein, underexpression is a decrease in gene expression and is generally less than at least 1.25-fold or alternatively, at least 1.5-fold or alternatively, at least 2-fold or alternatively, at least 3-fold or alternatively, at least 4-fold the expression detected in control counterpart cells or tissues. The term "differentially expressed" also refers to those in which expression is detected in cancer cells or cancer tissues and the expression is present, but not detectable in control cells.

[0083] The term "cDNA" refers to complementary DNA, which is prepared from mRNA molecules present in a cell or organism using an enzyme such as reverse transcriptase. A "cDNA" library is a collection of all the mRNA molecules present in a cell or organism, where all of the mRNA is converted to cDNA molecules using reverse transcriptase and then inserted into a "vector" (another DNA molecule that can continue to replicate after addition of foreign DNA). Exemplary vectors for libraries include bacteriophages, viruses that infect bacteria such as lambda phage. The library can then be probed for a specific cDNA of interest (and the resulting mRNA).

[0084] Measurement of gene expression

[0085] Gene expression can be detected by any suitable method, which includes, for example, detecting the amount of mRNA transcribed from a gene or the amount of cDNA reverse transcribed from the mRNA transcribed from a gene or the amount of polypeptide or protein encoded by a gene. Based on a sample or a modified high-throughput assay, these methods can be performed on the sample.

[0086] For example, using known techniques, any one of gene copy number, transcription, or translation can be determined. For example, amplification methods such as PCR can be used. The general procedures of PCR are taught in MacPherson et al., PCR: A Practical Approach, (IRL Press at Oxford University Press (1991)). However, the PCR conditions for each application reaction are determined empirically. Many parameters affect the success of the reaction. In particular, the annealing temperature and time, extension time, Mg 2 + and / or ATP concentration, pH, and relative concentration of primers, template, and deoxyribonucleotides. After amplification, the resulting DNA fragments can be detected by agarose gel electrophoresis followed by staining with ethidium bromide and ultraviolet illumination.

[0087] In one embodiment, hybridized nucleic acids are detected by detecting one or more labels attached to the sample nucleic acid. Labels can be incorporated by any of a variety of methods well known to those skilled in the art. However, in one aspect, the label is incorporated simultaneously during the amplification step of preparing the sample nucleic acid. Thus, for example, polymerase chain reaction (PCR) using labeled primers or labeled nucleotides can provide labeled amplification products. In an alternative embodiment, transcription amplification as described above is performed using labeled nucleotides (e.g., fluorescein-labeled UTP and / or CTP) to incorporate the label into the transcribed nucleic acid.

[0088] Alternatively, the label can be added directly to the original nucleic acid sample (e.g., mRNA, polyA, mRNA, cDNA, etc.) or directly to the amplification product after amplification is completed. Methods for ligating the label to the nucleic acid are well known to those skilled in the art and include, for example, nick translation or end labeling (e.g., using labeled RNA) through the action of a kinase on the nucleic acid, followed by binding (ligating) a nucleic acid linker bound to the sample nucleic acid to the label (e.g., fluorophore).

[0089] Detection of the label is well known to those skilled in the art. Thus, for example, a radioactive label can be detected using a film or a scintillation counter, and a fluorescent label can be detected using a light detector that detects emitted light. Enzyme labels are typically detected by providing an enzyme and a substrate and detecting the reaction product resulting from the action of the enzyme on the substrate, and calorimetric labels are detected by simple visual inspection of the color label.

[0090] Thus, a "reagent for determining the level or amount of a biomarker" refers to a reagent that can be used to quantify or measure the level or amount of a biomarker in a sample. Based on the sequences of the biomarkers provided by the present invention, such reagents can be easily designed or obtained by conventional methods well known in the art. For example, such reagents include, but are not limited to, PCR primers that can be used to quantify or measure the level or amount of a biomarker by, for example, real-time PCR; probes that can be used to quantify or measure the level or amount of a biomarker by, for example, quantitative Southern blotting; microarrays (e.g., gene chips) that can be used to quantify or measure the level or amount of a biomarker, etc. Additionally, as is known in the art, second-generation sequencing methods or third-generation sequencing methods can also be used to quantify or measure the level or amount of a biomarker. Thus, such reagents can also be commercially available reagents for performing second-generation sequencing methods or third-generation sequencing methods.

[0091] The term human reference genome refers to a file that records the base composition at each position on human chromosomes, generated by the Human Genome Project, and GRCh38 is the 38th version. Homo_sapiens.GRCh38.84.gtf is a gene annotation file that matches the GRCh38 human reference genome and records information such as genes and their positions on the reference genome.

[0092] The term Reads, also known as read segments, refers to a very large number of sequences included in the NGS output data, and each sequence is called a read.

[0093] In paired-end sequencing, a fragment is sequenced once at each of the left and right ends, generating two reads, and these two reads belong to the same DNA fragment

[0094] The term Trimmomatic refers to a tool for processing NGS raw data, http: / / www.usadellab.org / cms / ?page=trimmomatic 。

[0095] The term HISAT2 refers to a tool for aligning reads to the human reference genome, see http: / / daehwankimlab.github.io / hisat2 / 。

[0096] The term htseq_count refers to a tool for calculating the number of reads on genes, https: / / htseq.readthedocs.io / en / master / index.html 。

[0097] A commonly used tool for differential gene expression analysis by edgeR, https: / / www.bioconductor.org / packages / release / bioc / vignettes / edgeR / inst / doc / edgeRUsersGuide.pdf 。

[0098] The term LogFoldChange refers to the measure of the difference in the expression of a gene between two groups of samples in differential expression analysis. The calculation method is to divide the mean (or median) of the expression in group 1 by the mean (or median) of the expression in group 2, and then perform a log2 transformation.

[0099] Unless otherwise specified, FDR in this article refers to the value after correcting the p-value with FDR in edgeR, and the correction method is the Benjamini-Hochberg method.

[0100] The term TPM or its full name: Transcripts Per Million reads, is an indicator for relative quantification when the difference in the amount of raw sequencing data is too large due to a large difference in the sequencing sample size. TPM first addresses the issue of the difference in the number of reads caused by gene length, and then addresses the issue of the difference in the number of reads caused by sequencing depth. When emphasizing accuracy or quantifying the expression level of target genes, TPM is the most effective, and its application makes the relative expression levels between different sequencing samples comparable. The TPM calculation formula is as follows (Ni is the reads count on the i-th gene, and Li is the length of the i-th gene):

[0101]

[0102] Machine learning

[0103] The term "prediction model" refers to a specific mathematical model obtained by applying a prediction method to a data set. In the embodiments detailed herein, such a data set is obtained from measurements of gene expression levels in non-invasive samples of subjects, where the classification (receptive or pre-receptive) of each sample is known. This model can be used to classify samples of unknown receptivity status as receptive or pre-receptive. Optionally, such classification is a probability prediction (i.e., a ratio or percentage that is to be interpreted as a probability). The exact details of how these gene-specific measurements are combined to produce the classification and probability prediction depend on the specific mechanism used to construct the prediction model.

[0104] The term "machine learning algorithm", as commonly understood in the art, refers to a specific strategy that enables a computer to achieve a certain artificial intelligence goal. Common algorithmic strategies include inductive learning, analogical learning, analytical learning, etc. According to the learning form, it is further divided into supervised learning and unsupervised learning. The former refers to the process of using a set of samples with known results to adjust the parameters of the algorithm classifier so that it gradually approaches and reaches the required performance. Commonly used supervised learning algorithms in this field include Naive Bayes, decision tree, k-nearest neighbor, neural network, support vector machine, random forest, logistic regression, least squares method, adboost algorithm, hidden Markov model, etc. Unless otherwise specified, machine learning algorithms in this article all use supervised learning algorithms. Those skilled in the art can know that since the present invention seeks to discover internal laws, there is no definite requirement for the number of training samples required by the machine learning algorithm, that is, samples with known results, and the conventional and obtainable sample sizes can meet the needs of training the algorithm.

[0105] Unless otherwise specified, the term "random forest" refers to the random forest model, which is an integrated supervised learning algorithm. By building multiple decision tree models and then fusing them together, a more accurate and stable model can be obtained, and this model can also obtain good results when using default parameters. It is one of the commonly used models in this field currently. An implementation example of the random forest model in the R language can be seen at: https: / / cran.r-project.org / web / packages / randomForest / randomForest.pdf In some preferred embodiments of the present invention, the machine learning algorithm is the random forest model.

[0106] In some specific embodiments of the present invention, the commonly used hyperparameters of the random forest model are as follows:

[0107] a) n_estimators, the number of decision trees constructed. Generally, the larger the number, the better, but the calculation time will also increase accordingly. When the number of decision trees reaches a certain value, the effect will not get better anymore.

[0108] b) max_features: The number of features randomly selected when constructing each decision tree. The smaller this value, the lower the similarity between decision trees.

[0109] c) max_samples: The number or proportion of samples used when constructing each decision tree. The lower this value, the lower the similarity between decision trees.

[0110] d) class_weight: The proportion of samples in each class. By default, the proportion of samples in all classes is the same. For imbalanced samples, "balanced" can be used.

[0111] e) criterion: The metric used to evaluate the quality of each split when constructing a decision tree.

[0112] f) max_depth: The maximum depth of each decision tree.

[0113] g) min_samples_split: The minimum number of samples required for a decision tree to split.

[0114] h) min_samples_leaf: The minimum number of samples in each node of a decision tree.

[0115] Unless otherwise specified, in this article, the training of the model refers to the process of finding a set of optimal hyperparameters for the construction of the final model. In some specific embodiments of the present invention, the methods used for model training are as follows:

[0116] i. Random method: Randomly combine hyperparameters, search for a period of time, and then select the optimal parameters. Use accuracy to evaluate the model during hyperparameter tuning;

[0117] ii. Train the model with default parameters;

[0118] iii. Select the optimal method among the two methods. The evaluation metric is the average accuracy of 10-fold cross-validation (10-fold cv).

[0119] The term Scikit-learn refers to machine learning software based on the Python language. See https: / / scikit- learn.org / stable / .

[0120] The term Caret refers to software used for hyperparameter tuning of machine learning models. See https: / / topepo.github.io / caret / index.html .

[0121] The full name of the ROC curve is the Receiver Operating Characteristic curve, which is a comprehensive index that reflects the continuous variables of sensitivity and specificity and shows the relationship between the two through graphing. The AUC (area under the ROC curve) is the area under the ROC curve. The larger the AUC, the better, indicating that the classification accuracy of the test is higher.

[0122] In one aspect, the present invention provides a method for obtaining a living sample to be analyzed in a non-invasive manner. In a specific embodiment, the non-invasive manner is to adhere trace tissue to the wall of the transplantation tube during embryo pre-transplantation. In another specific embodiment, the non-invasive manner is to adhere trace tissue to the wall of the transplantation tube during embryo transplantation.

[0123] In one embodiment, the number of cells in the sample obtained by the non-invasive method is less than 1000. In a more specific embodiment, the number of cells in the sample obtained by the non-invasive method is less than 900, 800, 700, 600, 500, 400, 300 or 200. In a preferred embodiment, the number of cells in the sample obtained by the non-invasive method is between 400 and 300.

[0124] In a specific embodiment, the method of the present invention further includes transferring the cells and tissues on the tube wall to the sample collection tube by the method of intercepting and flushing. In a specific embodiment, in a preferred embodiment, a tissue sample preservation solution is pre-added to the sample collection tube.

[0125] In one embodiment, the method of the present invention further includes extracting total RNA from a sample. In a specific embodiment, the extraction process uses the Yikon MALBAC Platinum Micro RNA Amplification Kit (KT110700724) and is carried out according to the instructions provided by the manufacturer. In another specific embodiment, the extraction process includes the following steps: a) Centrifuge the sample at high speed (1600g for the first time for 2 minutes; 3000g for the second time for 3 minutes; 13000rpm for the third time for 5 minutes) to precipitate all cells. b) Remove the supernatant, add 200μL×2 tissue sample preservation solution to resuspend the cells, and centrifuge at 3000g for 5 minutes; 200μL×1 time, leaving about 50μL. c) Add 300μL of cell lysate and shake vigorously for 30 seconds to 1 minute; centrifuge instantaneously and discard the supernatant. d) Add 350μL of freshly prepared 70% ethanol and mix well by shaking; centrifuge instantaneously and discard the supernatant. e) Add 650μL of the above mixed liquid to the adsorption column and centrifuge at 13000×g for 1 minute, discarding the centrifugation waste liquid. f) Add 500μL of washing buffer to the adsorption column membrane and centrifuge at 13000×g for 1 minute, discarding the centrifugation waste liquid. g) Carefully drip 100μL of DNA digestion enzyme onto the adsorption column membrane and incubate at 25 degrees for 15 minutes. h) Add 500μL of washing buffer and centrifuge at 13000×g for 1 minute, discarding the centrifugation waste liquid; add 500μL of washing buffer to the adsorption column again and centrifuge at 13000×g for 1 minute, discarding the centrifugation waste liquid. i) Add 700μL of 80% ethanol to the tube and centrifuge at 13000×g for 2 minutes, discarding the centrifugation waste liquid; centrifuge at 13000×g for 1 minute again and discard the collection tube. j) Insert the adsorption column into a 1.5mL centrifuge tube, open the lid and air dry for 1 minute; add 20μL of DEPC water, incubate at 25 degrees for 1 minute, centrifuge at 13000×g for 2 minutes, and repeat once. A total of 40μL of eluate is collected, which is the total RNA. In a more specific embodiment, almost all of the RNA in the trace sample obtained by non-invasive sampling is extracted, thereby providing a sufficient amount of purified RNA nucleic acid sample for subsequent detection.

[0126] In one embodiment, the method of the present invention further includes reverse transcribing messenger RNA in the RNA nucleic acid sample to generate a cDNA library. In a specific embodiment, the reverse transcription uses the Yikon MALBAC Platinum Micro RNA Amplification Kit (KT110700724) and is carried out according to the instructions provided by the manufacturer. In a specific embodiment, the concentration of the cDNA library obtained from the non-invasive sample of the present invention is about 10-20ng / μL, about 20-30ng / μL, about 30-40ng / μL, about 40-50ng / μL, about 50-60ng / μL, about 60-70ng / μL, about 70-80ng / μL, about 80-90ng / μL.

[0127] In one embodiment, the method of the present invention further includes further processing the cDNA library to construct a sequencing library. In a specific embodiment, the further processing uses the Yikon Fragmentation Kit (KT100804248) and the Gene Sequencing Library Kit (KT100804048) and is carried out according to the instructions provided by the manufacturer.

[0128] In another alternative embodiment, the cell lysis method, and / or RNA extraction method, and / or reverse transcription method, and / or library amplification method, and / or fragmentation method, and / or sequencing library construction method used in the method of the present invention refer to the previous application of the applicant's Chinese patent application with the application number 201810019044.7, the entire content of which is incorporated herein by reference.

[0129] In one embodiment, the method of the present invention further includes performing sequencing on the obtained library. In a specific embodiment, the sequencing method is next-generation sequencing (NGS). In another specific embodiment, the sequencing method is Sanger sequencing, or third-generation sequencing (e.g., nanopore sequencing). In a preferred embodiment of the present invention, the Illumina Nextseq sequencing system and the supporting Next-seq High-Output Kit are used.

[0130] In another aspect, the present invention provides a method for establishing a receptivity prediction model. In one embodiment, the present invention provides a method for analyzing the sequencing results by machine learning algorithms and establishing a receptivity prediction model.

[0131] In a specific embodiment, the machine learning algorithm analysis includes the following steps: I) Classify the samples into receptive and non-receptive categories according to the gold standard (actual clinical outcome), and preferably split them into two sets, a training set and a test set, and take the sample data of the training set for the next step; II) Denoise and debias all the data of the training set samples to generate a fragment count alignment file for each gene; III) Screen and collect the genes with differential expression and use them as the features to be identified; IV) Convert the data to TPM as the normalized expression level data for each gene, thereby eliminating the influence of the difference in the amount of sample data off the machine; V) Input the expression level data of the features to be identified (genes) into the machine learning algorithm, and the machine learning algorithm performs supervised learning to obtain the threshold, sensitivity, and importance of each feature for judging receptivity, thereby obtaining a receptivity prediction model. Preferably, the established model is verified for accuracy in the test set.

[0132] In one embodiment, the model includes one or more sub-models, and different sub-models contain all or different parts of the aforementioned differentially expressed genes.

[0133] In a specific embodiment, step (II) specifically includes:

[0134] c) The following processing is performed on the sequencing data of each sample respectively

[0135] i. The sequencing data is processed using trimmomatic (version 0.33) to remove adapter sequences and filter low-quality reads, generating cleaned and clean sequencing data.

[0136] ii. The clean sequencing data generated in the previous step is aligned to the human reference genome (version GRCh38) using HISAT2 (version 2.0.5) to generate an alignment file.

[0137] iii. The alignment file generated in the previous step is sorted, and duplicate sequences are marked to generate the final alignment file.

[0138] iv. The alignment file generated in the previous step is processed using htseq_count. This step also requires the use of a gene annotation file, the Homo_sapiens.GRCh38.84.gtf file, which is downloaded from Ensembl and contains information such as the name and position coordinates of each gene. Htseq_count can count how many sequences are in the region where each gene is located based on this information (for paired-end sequencing, reads belonging to the same fragment are only counted once), generating a genecount file.

[0139] d) The genecount files of all samples in a batch are merged to generate a total genecount file, where each row is a sample and each column is a gene, and the data therein is the number of fragments belonging to the gene in the sample. In a specific embodiment, step (III) specifically includes the following selection methods:

[0140] i. Exclude genes related to rRNA and mitochondria that cause greater interference later.

[0141] ii. For the training set, perform gene differential expression analysis to screen out the differentially expressed genes in the two types of samples, obtaining 3193 genes.

[0142] iii. The 3193 genes screened in the previous step are further screened by variables, and a total of 250 genes are obtained as candidate features for the next model training. The detailed list is shown in Table 2.

[0143] In a more specific embodiment, the variables are screened to select genes with FDR < 0.001 from the differentially expressed genes, resulting in 1270 genes. The logTMM values of each gene in each sample are output using the cpm function of edgeR, and the pearson correlation between these genes is calculated pairwise. Then, the genes are screened according to the following steps

[0144] 1) Sort the genes in ascending order of FDR;

[0145] 2) Mark the top gene in the list as selected;

[0146] 3) Mark the genes with a correlation greater than or equal to 0.6 with any of the selected genes as excluded;

[0147] 4) Repeat steps 2) and 3) until the end of the list;

[0148] 5) Construct a model with the screened genes and evaluate the accuracy using the 10-fold cross-validation method. Successively delete the genes that have no effect on the accuracy of the model, and finally obtain 250 genes.

[0149] In an even more specific embodiment, step ii. above uses edgeR (version 3.34.0) for gene differential expression analysis, and the screening conditions are logFoldChange > 0, FDR < 0.05, including the following steps:

[0150] a) edgeR requires two input files. The first is the genecount file, which contains the number of reads of each gene in each sample, and the second is the sample information file, which contains which category each sample belongs to;

[0151] b) After edgeR reads in the input files, it first normalizes the number of reads of each gene in each sample. When normalizing, genes with stable expression in all samples are selected;

[0152] c) edgeR then looks for genes with different expressions in different categories according to the category information of the samples. The significance of the difference is represented by the p-value after FDR correction. The smaller the p-value, the more significant the difference.

[0153] In a specific embodiment, step IV) specifically includes converting the genecount data to TPM to exclude the influence of the sample throughput data volume. The calculation formula is as follows:

[0154]

[0155] Where Ni is the read count on the i-th gene, and Li is the length of the i-th gene.

[0156] In a specific embodiment, the machine learning algorithm in step V) is a random forest algorithm.

[0157] In one embodiment, the method of the present invention further includes evaluating and / or validating indicators such as sensitivity, specificity, accuracy, and ROC of the established receptivity prediction model in the split independent test set.

[0158] In a specific embodiment, the method of the present invention further includes validating the established receptivity prediction model in at least two test sets.

[0159] In another aspect, the present invention also provides a biomarker set, which is used to predict endometrial receptivity to guide embryo transfer.

[0160] In one embodiment, the biomarker set of the present invention comprises some or all of the differentially expressed genes identified in step III) above, for example, any number between more than 5 and less than or equal to 250. In a more specific embodiment, in order to accurately predict endometrial receptivity to guide embryo transfer, the biomarker set of the present invention includes at least 10 gene members. In an even more specific embodiment, the biomarker set of the present invention includes at least 11 gene members, at least 12 gene members, at least 13 gene members, at least 14 gene members, at least 15 gene members, at least 16 gene members, at least 17 gene members, at least 18 gene members, at least 19 gene members, at least 20 gene members, at least 25 gene members, at least 30 gene members, at least 35 gene members, at least 40 gene members, at least 45 gene members, at least 50 gene members, at least 55 gene members, at least 60 gene members, at least 65 gene members, at least 70 gene members, at least 80 gene members, at least 90 gene members, at least 100 gene members, at least 110 gene members, at least 120 gene members, at least 130 gene members, at least 140 gene members, at least 150 gene members, at least 160 gene members, at least 170 gene members, at least 180 gene members, at least 190 gene members, at least 200 gene members, at least 210 gene members, at least 220 gene members, at least 230 gene members, at least 240 gene members, or at least 250 gene members. In one embodiment, the biomarker set of the present invention comprises 5 genes, 6 genes, 7 genes, 8 genes, 9 genes, 10 genes, 15 genes, 20 genes, 25 genes, 30 genes, 40 genes, 50 genes, 60 genes, 70 genes, 80 genes, 90 genes, 100 genes, 110 genes, 120 genes, 130 genes, 140 genes, 150 genes, 160 genes, 170 genes, 180 genes, 190 genes, 200 genes, 210 genes, 220 genes, 230 genes, 240 genes, or 250 genes with the highest weights in the receptivity prediction model established according to step V) above.

[0161] In one embodiment, the biomarker set of the present invention comprises at least the following genes: ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6. In a preferred embodiment, the biomarker set of the present invention comprises the following genes: ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6, and further sequentially comprises 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, or 240 gene markers numbered 11 - 250 in Table 2.

[0162] In a particular aspect, the endometrial receptivity prediction model of the present invention comprises one or more sub - models, wherein different sub - models are respectively related to the different biomarker sets described above.

[0163] In yet another aspect, the present invention provides uses and methods based on the biomarker set.

[0164] In one embodiment, the present invention provides the use of the above - mentioned biomarker set or its specific detection tool for preparing an assay composition, assay reagent, and / or assay device for detecting and determining the endometrial receptivity status of a subject, wherein the assay composition, assay reagent, and / or assay device can detect the expression level of each member of the above - mentioned biomarker set in a sample from the subject; and calculate the expression level through the receptivity prediction model of the present invention, and the obtained result is used to judge and diagnose the endometrial receptivity status.

[0165] In one embodiment, the present invention provides a method for predicting endometrial receptivity, the method comprising obtaining the expression levels of all gene members of the biomarker set by using a substance or method for specifically measuring the level or amount of the biomarker. In a specific embodiment, the substance for measuring the level or amount of the biomarker is a composition, reagent or device. In a specific embodiment, the substance for measuring the level or amount of the biomarker is a specific primer and / or probe molecule. In a more specific embodiment, the composition or reagent for measuring the level or amount of the biomarker further comprises one or more of a tag sequence, an adapter sequence, a fluorescent molecule, a pyrophosphate molecule, and a labeling molecule. In a specific embodiment, the method for measuring the level or amount of the biomarker is a PCR method. In a more specific embodiment, the PCR method is a method commonly used in the art such as real-time fluorescence quantitative PCR, isothermal amplification PCR, etc. In a specific embodiment, the method for measuring the level or amount of the biomarker is a sequencing method. In a more specific embodiment, the sequencing method is one or more of next-generation sequencing (NGS), Sanger sequencing, and third-generation sequencing (e.g., nanopore sequencing).

[0166] In one embodiment, the reagent for measuring the level or amount of the biomarker does not contain a reagent for specifically detecting genes other than the members of the biomarker set of the present invention.

[0167] In one embodiment, the method for measuring the level or amount of the biomarker does not measure the level or amount of genes other than the members of the biomarker set of the present invention.

[0168] In one embodiment, the present invention provides a method for predicting endometrial receptivity, wherein the sample used in the method is obtained in a non-invasive manner, that is, the acquisition of the sample does not cause damage to the endometrium of the pregnant woman. In a preferred embodiment, the present invention provides an integrated non-invasive method for predicting endometrial receptivity, and the "integrated" means that this method does not need to be carried out separately, but is carried out simultaneously during embryo pre-transfer or embryo transfer operations, and the sampling process is integrated into the above operations, so that the sampling cycle is the transplantation cycle. For example, when the cycle transplantation fails and the test result indicates an endometrial cause, the next cycle of transplantation can be guided with reference to the test result without waiting, which greatly shortens the pregnancy time of the pregnant woman.

[0169] In yet another aspect, the present invention provides a kit for the above-mentioned method, the kit comprising reagents for determining the level or amount of each biomarker in the biomarker combination of the present invention in a sample, such as specific reagents and general reagents.

[0170] The following examples illustrate, by way of example and not limitation, the embodiments of the present invention and their technical effects. Other embodiments, methods, and products are within the capabilities of those of ordinary skill in the fields of molecular diagnosis and molecular biology, and thus need not be described in detail herein. Other embodiments falling within the scope of the art are considered to be part of the present invention. Example

[0171] Example 1 : Establish a method for non-invasive sampling of the endometrium in subjects waiting for embryo transfer

[0172] The inventors accidentally discovered that, as one of the important guiding steps for embryo transfer, embryo pre-transfer, the trace tissue cells adhered during its operation are sufficient for gene differential expression analysis related to endometrial receptivity.

[0173] To determine the stability and reliability of the above discovery, the inventors collected trace tissue cells adhered to the operating tubes of embryo pre-transfer from 243 subjects clinically.

[0174] (After the (pre-)transfer operation, the transplantation tube wall was intercepted and rinsed, so that the extremely trace endometrial tissue or cells adhered to it were transferred into a sample collection tube. The tissue sample preservation solution (ComWin Biotech, CW0592M) in the sample collection tube can protect the tissue cell sample to ensure the subsequent extraction of total RNA.

[0175] Clinical sampling scenario 1 (a total of 12 cases): Visible tissue was brought out by the transplantation tube. The inner and outer walls of the transplantation catheter were rinsed with PBS at 4 °C in a culture dish. Using a nuclease-free pipette tip, small pieces of endometrial tissue were picked and transferred into an ERT tissue preservation tube (1.5 mL) (the tube contains 0.3 ml of preservation solution).

[0176] Clinical sampling scenario 2 (a total of 231 cases): No visible tissue. Through the inner tube rinsing method combined with the interception method, a very small amount of cells were recovered. Inner tube rinsing method: Replace the brand-new 1 mL syringe, inject 100 - 150 μL of PBS to rinse the transplantation catheter, and directly inject the rinsing solution into an ERT tissue preservation tube (1.5 mL) (the tube contains 0.5 ml of preservation solution). The above steps can be repeated once according to the actual situation. Interception method: Intercept about 5 mm of the outer tube and 8 - 10 mm of the inner tube of the transplantation catheter and directly place them into an ERT tissue preservation tube (1.5 mL) (the tube contains 0.3 ml of preservation solution).

[0177] Subsequently, using the trace cell RNA extraction method, total RNA was extracted as completely as possible. And a cDNA library was obtained by reverse transcription and amplification using the MALBAC Platinum MicroRNA Amplification Kit (KT110700724) independently developed by Yikon. The extraction method was carried out according to the following steps:

[0178] a) Centrifuge the sample at high speed (the first time at 1600g for 2 minutes; the second time at 3000g for 3 minutes; the third time at 13000rpm for 5 minutes) to precipitate all cells.

[0179] b) Remove the supernatant, add 200μL×2 tissue sample preservation solution to resuspend the cells, and centrifuge at 3000g for 5 minutes; 200μL×1 time, leaving about 50μL.

[0180] c) Add 300μL of cell lysate, shake vigorously for 30 seconds to 1 minute; centrifuge instantaneously and discard the supernatant. d) Add 350μL of freshly prepared 70% ethanol, shake well; centrifuge instantaneously and discard the supernatant.

[0181] e) Add 650μL of the above mixed liquid to the adsorption column, centrifuge at 13000×g for 1 minute, and discard the centrifugation waste liquid.

[0182] f) Add 500μL of washing buffer to the adsorption column membrane, centrifuge at 13000×g for 1 minute, and discard the centrifugation waste liquid.

[0183] g) Carefully drop 100μL of DNA digest on the adsorption column membrane and incubate at 25°C for 15 minutes.

[0184] h) Add 500μL of washing buffer, centrifuge at 13000×g for 1 minute, and discard the centrifugation waste liquid; add 500μL of washing buffer to the adsorption column again, centrifuge at 13000×g for 1 minute, and discard the centrifugation waste liquid.

[0185] i) Add 700μL of 80% ethanol to the tube, centrifuge at 13000×g for 2 minutes, and discard the centrifugation waste liquid; centrifuge at 13000×g for 1 minute again and discard the collection tube.

[0186] j) Insert the adsorption column into a 1.5mL centrifuge tube, open the lid and air dry for 1 minute; add 20μL of DEPC water, incubate at 25°C for 1 minute, centrifuge at 13000×g for 2 minutes, and repeat once, collecting a total of 40μL of eluate, which is the total RNA. The extraction can also be carried out according to the kit instructions. Reverse transcription is carried out according to the kit (KT110700724) instructions, or according to the reverse transcription operation method described in the instructions of the Chinese invention patent application with the application number 201810019044.7.

[0187] The cDNA library was obtained as shown in Table 1 below:

[0188] Table 1

[0189] cDNA preparation Sequencing library preparation Final prediction result Number of successful cases 243 228 224 Total number of cases 243 243 243 Success rate 100.0% 93.8% 92.2%

[0190] The results showed that reliable cDNA libraries with quality could be successfully obtained from all 243 trace cells. And in the subsequent methods, 228 cases were successfully subjected to library construction and sequencing, and 224 cases finally obtained accurate receptivity prediction results according to the technical solution of the present invention (see the subsequent examples).

[0191] The inventor further measured the nucleic acid concentration of the obtained cDNA libraries, and deduced the number of cells in the non-invasive samples. Two groups of samples, namely the conventional method group and the non-invasive sample group, were set for comparison:

[0192] The average concentration of the cDNA libraries obtained from 243 non-invasive samples was 45.9±30.4 ng / μL;

[0193] The conventional method group was samples obtained by invasive methods such as endometrial cell biopsy and endometrial exfoliated cell sampling clinically. There were 7 cases in total. 10 ng was taken from their total RNA for cDNA library construction. The average concentration of the obtained cDNA libraries was 114.5±12.4 ng / μL.

[0194] Based on the cell amount corresponding to 10 ng of total RNA in the conventional method group, which was about 2,300 - 2,700 cells / sample, combined with the efficiency curve of cDNA amplification, it was calculated that the cell amount in the non-invasive sample group was about 300 - 1,000 cells / sample.

[0195] Therefore, through this example, the inventor established a new method for obtaining endometrial receptivity prediction samples by non-invasive sampling method, and verified its stability and reliability.

[0196] For subjects undergoing embryo pre-transplantation operations, samples can be obtained by the non-invasive sampling method of the present invention, and endometrial receptivity can be predicted to guide subsequent embryo transplantation, greatly improving the success rate.

[0197] For subjects who have already undergone embryo transplantation operations, samples can also be obtained by the non-invasive sampling method of the present invention, and endometrial receptivity can be predicted. If this transplantation fails, the prediction results can be referred to guide subsequent embryo transplantation, obtaining more information with as few operations as possible, and reducing the pain of the subjects.

[0198] Due to the relatively scarce cell content in the samples and different cell types from conventional samples, new challenges are brought to the establishment of the predictive model. The subsequent method steps will be established and improved through more following examples.

[0199] Example 2: Amplification and sequencing of samples obtained by non-invasive sampling

[0200] For the cDNA library obtained in Example 1, a sequencing library was constructed using the Yikang Fragmentation Kit (KT100804248) and the Gene Sequencing Library Kit (KT100804048). The specific operation steps were carried out according to the instructions, or according to the reverse transcription operation method described in the specification of the Chinese invention patent application with the application number 201810019044.7.

[0201] For the obtained sequencing library, sequencing was performed using the Illumina Nextseq sequencing system and the supporting Next-seq High-Output Kit kit, and the off-machine data was obtained. Those skilled in the art can also use any common sequencing scheme in the art to sequence the obtained library.

[0202] Example 3 : Screen and analyze differentially expressed genes from sequencing data as biomarkers of endometrial receptivity Biological set, and establish a prediction model using the random forest algorithm.

[0203] The samples described in Example 1 were classified according to the final actual clinical outcome into the receptive phase and the pre-receptive phase (i.e., non-receptive phase), and the specific classification method was as follows:

[0204] If the embryo transplanted at D5 of the subject from which the sample was derived implanted successfully subsequently, it indicated that the test sample was in a receptive state at the time of transplantation; if the embryo did not implant successfully, the test sample was in a pre-receptive state.

[0205] If the embryo transplanted at D3 of the subject from which the sample was derived implanted successfully subsequently, it indicated that the test sample was in a pre-receptive state at the time of transplantation; if the embryo did not implant successfully, the test sample was in a receptive state.

[0206] The samples were numbered with D5 + serial number or D3 + serial number, and their receptive states were marked according to the gold standard of clinical outcome. And the 243 samples were divided into two groups: a training set including 211 samples for supervising and training the machine learning algorithm to establish a prediction model; and a test set including 32 samples for subsequently independently testing the established model.

[0207] 3.1 Identification of differentially expressed genes

[0208] The off-machine sequencing data obtained in Example 2 was used to analyze the relative expression level of each gene, and then compare the gene expression levels between the two types of samples, so as to identify the genes with differential expression.

[0209] Specifically, the off-machine data of each sample was processed as follows:

[0210] i. The off-machine data was processed using trimmomatic (version 0.33) to remove adapter sequences, filter low-quality reads, and generate clean sequenced data.

[0211] ii. The clean sequenced data generated in the previous step was aligned to the human reference genome (version GRCh38) using HISAT2 (version 2.0.5) to generate an alignment file.

[0212] iii. The alignment file generated in the previous step was sorted, and duplicate sequences were marked to generate the final alignment file.

[0213] iv. The alignment file generated in the previous step was processed using htseq_count. This step also requires a gene annotation file, namely the Homo_sapiens.GRCh38.84.gtf file, which was downloaded from Ensembl and contains information such as the name and location coordinates of each gene. Htseq_count can count how many sequences are in the region where each gene is located based on this information (for paired-end sequencing, reads belonging to the same fragment are only counted once), generating a genecount file.

[0214] v. The genecount files of all samples in a batch were merged to generate a total genecount file, where each row represents a sample, each column represents a gene, and the data is the number of fragments belonging to that gene in that sample.

[0215] Thus, the relative expression level of each gene in each specific sample was obtained. Furthermore, differentially expressed genes with application significance were identified through the following specific steps:

[0216] i. Exclude genes related to rRNA and mitochondria that have greater interference in the follow-up.

[0217] ii. For the training set, edgeR (version 3.34.0) was used for gene differential expression analysis:

[0218] a) EdgeR requires two input files. The first is the genecount file, which contains the number of reads of each gene in each sample, and the second is the sample information file, which contains which category each sample belongs to (i.e., receptive or pre-receptive);

[0219] b) After edgeR reads the input files, it first normalizes the number of reads of each gene in each sample, and selects genes with stable expression in all samples during normalization;

[0220] c) Then, edgeR searches for genes with differential expression in different categories based on the category information of the samples. The significance of the difference is represented by the p-value after FDR correction. The smaller the p-value, the more significant the difference. The screening conditions are set as logFoldChange > 0 and FDR < 0.05. Genes with differences between the two types of samples are screened out, resulting in 3,193 genes.

[0221] iii. Further variable screening is performed on the 3,193 genes screened in the previous step. The variable screening is to select genes with FDR < 0.001 from the differentially expressed genes, resulting in 1,270 genes. The logTMM value of each gene in each sample is output using the cpm function of edgeR. The pearson correlation between these genes is calculated pairwise, and then the genes are screened according to the following steps:

[0222] 1) Sort the genes in ascending order of FDR;

[0223] 2) Mark the top gene in the list as selected;

[0224] 3) Mark genes with a correlation greater than or equal to 0.6 with any of the selected genes as excluded;

[0225] 4) Repeat steps 2) and 3) until the end of the list;

[0226] 5) Build a model with the screened genes and evaluate the accuracy using the 10-fold cross-validation method. Successively delete genes that have no impact on the accuracy of the model. Finally, 250 genes are obtained.

[0227] 3.2 Establish a prediction model using differentially expressed genes

[0228] Using the expression data of the 250 differentially expressed genes obtained in the previous step as features, a prediction model for judging the endometrial receptivity status from gene expression data is further established using machine learning algorithms.

[0229] 1. Data processing:

[0230] Convert the genecount data to TPM to exclude the influence of the sample throughput data volume. The calculation formula is as follows:

[0231]

[0232] where Ni is the reads count on the i-th gene and Li is the length of the i-th gene.

[0233] 2. Model Selection: The random forest model is an ensemble learning method that builds multiple decision tree models and then combines them to obtain a more accurate and stable model. Moreover, this model can also achieve good results when using default parameters and is one of the commonly used models currently.

[0234] (1). The commonly used hyperparameters of the random forest model are as follows:

[0235] a) n_estimators, the number of decision trees to be constructed. Generally, the larger the number, the better, but the calculation time will also increase accordingly. When the number of decision trees reaches a certain value, the effect will not improve anymore.

[0236] b) max_features, the number of features randomly selected when constructing each decision tree. The smaller this value, the lower the similarity between decision trees.

[0237] c) max_sample, the number or proportion of samples used when constructing each decision tree. The lower this value, the lower the similarity between decision trees.

[0238] d) class_weight, the proportion of samples in each class. By default, the proportion of samples in all classes is the same. For unbalanced samples, "balanced" can be adopted.

[0239] e) criterion, the metric used to evaluate the quality of each split when constructing a decision tree.

[0240] f) max_depth, the maximum depth of each decision tree.

[0241] g) min_samples_split, the minimum number of samples required for a decision tree to split.

[0242] h) min_samples_leaf, the minimum number of samples in each node of a decision tree.

[0243] (2). Model Training: The process of model training is to find a set of optimal hyperparameters for the construction of the final model. The methods used in this training are as follows

[0244] i. Random method: Randomly combine the hyperparameters, search for a period of time, and then select the optimal parameters. When tuning the parameters, evaluate the model using accuracy respectively.

[0245] ii. Train the model with default parameters.

[0246] iii. Select the optimal method among the two methods. The evaluation metric is the average accuracy of 10-fold cross-validation (10-fold cv).

[0247] The following software was used when training the model: the reported caret (6.0_88), randomForest (4.6_14), and scikit-learn (0.24.1).

[0248] Through the random forest algorithm, machine learning was carried out on the split training set, and the importance ranking of individual genes is shown in Table 2 below:

[0249] Table 2 (the values in the table are TPM)

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258] As shown in the above table, by implementing the random forest algorithm, the TPM values of the above differentially expressed genes were used to randomly generate multiple decision trees and obtain prediction results, and the prediction results were compared with the actual sample results to train the algorithm. After a sufficient machine learning process, the present invention successfully established a prediction model for endometrial receptivity for non-invasive samples. The model was trained with the samples in the training set and was able to obtain correct prediction results for all samples. Moreover, through evaluation in an independent test set, its high accuracy, that is, its universality for other samples, was confirmed, which has clinical application significance.

[0259] As shown in the results of Table 2, the model gave a ranking from high to low according to the importance of the identified individual genes in the prediction. The higher the gene is ranked, the higher the accuracy of the judgment when it is used alone as a marker. For example, the gene ELK4 (ENSG00000158711) ranked first can achieve an accuracy of 60% when judged according to its expression level in the obtained model.

[0260] Obviously, the complex physiological regulation processes in the body necessarily involve the participation of multiple genes or multiple pathways. Therefore, the 60% prediction accuracy obtained using a single biomarker is clearly insufficient. By constructing sub-models by sequentially adding gene features in order, the inventors found that a sub-model including at least the top 10 gene members (i.e., ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6) can achieve a relatively high accuracy and has application value.

[0261] In other words, in the samples obtained by the non-invasive sampling method discovered and adopted in the present invention, the differential expression of the above 250 genes reflects the receptive state of the samples. Among them, the expression levels of 10 genes, namely ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6, are most correlated with the receptive state and can be used as molecular biomarkers for the receptive state in non-invasive samples. If more biomarkers from the above 250 genes are incorporated for comprehensive consideration based on these 10 genes, more accurate prediction results may be obtained.

[0262] To verify this conclusion identified in the training set, according to the previous experimental design, the model of the present invention was used for a separately divided test set to test the accuracy of the model of the present invention in samples different from the training set (i.e., to further verify the universality of the biomarker set and the prediction model).

[0263] Example 4 : Validate the endometrial receptivity biomarker set and prediction model in the test set 。

[0264] An independent test set was used to evaluate the sensitivity, specificity, accuracy, ROC and other indicators of the model trained in Example 3.

[0265] Feature 10 refers to the sub-model composed of the differentially expressed genes ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6 ranked 1-10 in Table 2, and its prediction accuracy was evaluated. The model of Feature 30 is composed of the genes numbered 1-30 in Table 2, and so on. In this example, 13 sub-models of Feature 10, Feature 30, Feature 50, Feature 70, Feature 90, Feature 110, Feature 130, Feature 150, Feature 170, Feature 190, Feature 210, Feature 230, and Feature 250 were calculated (Feature 250 includes all markers).

[0266] The results are as follows and Figure 1 - Figure 13 shown below.

[0267] feature 10

[0268] Model output result

[0269]

[0270] Model accuracy

[0271]

[0272] feature 30

[0273] Model output result

[0274]

[0275]

[0276] Model accuracy

[0277]

[0278] feature 50

[0279] Model output result

[0280]

[0281]

[0282] Model accuracy

[0283]

[0284] feature 70

[0285] Model output result

[0286]

[0287]

[0288] Model accuracy

[0289]

[0290] feature 90

[0291] Model output result

[0292]

[0293]

[0294] Model accuracy

[0295]

[0296] feature 110

[0297] Model output result

[0298]

[0299]

[0300] Model accuracy

[0301]

[0302] feature 130

[0303] Model output result

[0304]

[0305]

[0306] Model accuracy

[0307]

[0308] feature 150

[0309] Model output result

[0310]

[0311]

[0312] Model accuracy

[0313]

[0314] feature 170

[0315] Model output result

[0316]

[0317]

[0318] Model accuracy

[0319]

[0320] feature 190

[0321] Model output result

[0322]

[0323] Model accuracy

[0324]

[0325] feature 210

[0326] Model output result

[0327]

[0328]

[0329] Model accuracy

[0330]

[0331] feature 230

[0332] Model output result

[0333]

[0334]

[0335] Model accuracy

[0336]

[0337] feature 250

[0338] Model output result

[0339]

[0340]

[0341] Model accuracy

[0342]

[0343] As shown in the above results, the sub-model composed of the top 10 markers already has a prediction sensitivity as high as 89% and a prediction specificity of 85%, and the prediction accuracy is 87.5% (i.e., accurately predicting the receptivity status of 28 out of 32 samples), which has sufficient clinical significance.

[0344] Based on these 10 markers, continuing to incorporate more markers ranked later into the model can also improve the detection accuracy to a certain extent.

[0345] Thus, it can be seen that the differentially expressed gene markers identified using the training set in Example 3 also showed an association with endometrial receptivity in the test set, and a stable and reliable prediction model was established by evaluating the weights of each gene using a machine learning algorithm.

[0346] In summary, the present invention for the first time discovers that the gene expression information carried by trace non-invasive samples is sufficient to predict endometrial receptivity, and proposes a method for predicting endometrial receptivity using non-invasive samples. The present invention also for the first time establishes a set of markers and a prediction model for predicting endometrial receptivity using non-invasive samples. Under the condition of hardly increasing the pain of the subjects, it greatly improves the success rate of subsequent embryo transfer of the subjects and shortens the time to pregnancy, having significant clinical significance.

Claims

1. A biomarker combination for predicting the endometrial receptivity of a subject, characterized in that Composed of 10 - 250 genes selected from Table 2, and the combination contains at least the genes numbered 1 - 10 in Table 2, namely ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB and TDRD6 , and optionally further contains one or more of the genes numbered 11 - 250 in Table 2.

2. Use of a reagent for specifically detecting the expression level or amount of the biomarker combination according to claim 1 in the preparation of a kit for predicting the endometrial receptivity of a subject.

3. A method for establishing a non-invasive prediction model for endometrial receptivity, comprising the following steps: i. By intercepting and flushing, transfer the trace tissue cells adhered to the transplant tube wall during pre-embryo transplantation or embryo transplantation into a sample collection tube to obtain a sample from the subject; ii. Extract RNA from the sample and reverse transcribe to obtain a cDNA library, and sequence the cDNA libraries of all samples; iii. Use the sequencing data in step ii. for a machine learning algorithm, where step iii. further includes the following specific steps: I) Mark the samples as receptive or non-receptive categories according to the outcome of embryo transplantation of the subjects from whom the samples are derived, and split them into two sets: a training set and a test set, and take the sample data of the training set for the next step; II) Denoise and de-bias all the data of the training set samples to generate a fragment count alignment file for each gene; III) Screen and collect genes with differential expression and use them as features to be identified; IV) Convert the raw sequencing data into TPM, which is the normalized expression level data of each gene, so as to eliminate the influence of the difference in the amount of sample sequencing data; V) Input the expression level data of the features to be identified into the machine learning algorithm, and perform supervised learning by the machine learning algorithm to obtain the threshold, sensitivity, and importance of each feature for judging receptivity, so as to obtain a receptivity prediction model, wherein, III) specifically includes: (i). Exclude genes related to rRNA and mitochondria with greater subsequent interference, (ii). For the training set, perform gene differential expression analysis and screen out the genes with differences between the two types of samples, (iii). Further perform variable screening on the genes screened in the previous step, where the variable screening is to select genes with FDR < 0.001 from the differentially expressed genes, use the cpm function of edgeR to output the logTMM value of each gene in each sample, calculate the pearson correlation between these genes pairwise, and then screen genes according to the following steps (1) Sort the genes in ascending order of FDR; (2) Mark the top gene in the list as selected; (3) Mark the genes with a correlation greater than or equal to 0.6 with any of the selected genes as excluded; (4) Repeat (2) and (3) until the end of the list; (5) Construct a model with the screened genes, evaluate the accuracy using the 10-fold cross-validation method, and sequentially delete the genes that have no effect on the accuracy of the model. Finally, determine that the genes used for the prediction model are selected from 10 - 250 genes in Table 2, and at least include the genes numbered 1 - 10 in Table 2, namely ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB, and TDRD6, and optionally also include one or more genes numbered 11 - 250 in Table 2.

4. The method for establishing a non-invasive prediction model for endometrial receptivity according to claim 3, wherein the machine learning algorithm is a random forest.

5. A combined product of polynucleotides for predicting the endometrial receptivity of a subject, characterized in that The combined product consists of polynucleotides of 10 - 250 genes in Table 2, and at least contains the polynucleotides of the genes numbered 1 - 10 in Table 2, namely ELK4, PRKACB, PHB2, GM2A, PPA1, NCEH1, CAPG, DDIT4, LDHB and TDRD6 and optionally further contains polynucleotides of one or more genes numbered 11 - 250 in Table 2. The polynucleotides are differentially expressed in non-invasive samples of different endometrial receptivity states, and the polynucleotides are endometrial receptivity gene markers.

Citation Information

Patent Citations

  • Single-cell RNA (ribonucleic acid) reverse transcription and library construction method

    CN108103055A

  • Method for judging endometrial receptivity and application thereof

    CN110042156A

  • Novel embryo transfer tube for obtaining endometrium

    CN113288377A