Intelligent transcriptome analysis system based on maternal-fetal interface biomarkers
The intelligent transcriptome analysis system solves the problem that traditional methods cannot fully reflect the biological information of the maternal-fetal interface, improves sample diversity and data accuracy, and reveals gene expression patterns and interactions in the maternal-fetal interface.
Patent Information
- Application Number
- CN202510508185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional biomarker analysis methods are insufficient to fully reflect the complex biological information at the maternal-fetal interface, and are limited by sample size and experimental techniques, making it difficult to effectively reveal gene expression patterns at the maternal-fetal interface.
Design an intelligent transcriptome analysis system, including a maternal-fetal RNA cryopreservation module, a maternal-fetal genome sequencing and computation module, a transcriptome differential expression determination module, and a transcriptome differential expression network construction module. By collecting and classifying blood samples from early, mid, and late pregnancy, RNA extraction, cryopreservation, genome sequencing, and differential expression analysis are performed to construct a gene expression network.
To ensure sample diversity and representativeness, improve the reliability and breadth of research results, maintain the integrity of RNA, improve the accuracy and reliability of gene expression data, reveal the complex physiological and molecular interactions at the maternal-fetal interface, and provide genomic data support.
Smart Images

Figure CN120452521B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biomedical detection, and particularly relates to an intelligent transcriptome analysis system based on biomarkers of the maternal-fetal interface. BACKGROUND
[0002] The maternal-fetal interface includes the interaction between the maternal endometrium and the placenta. This site not only plays a decisive role in fetal development, but also involves the regulation of the maternal immune system. Therefore, in-depth understanding of the biological mechanisms of the maternal-fetal interface is of great significance for improving the early prevention of pregnancy diseases. In the study of the maternal-fetal interface, the detection and analysis of biomarkers is a key link. With the development of high-throughput sequencing technology (such as RNA-seq), transcriptome analysis has become an effective means to reveal gene expression patterns and discover potential biomarkers. However, traditional biomarker analysis methods, such as enzyme-linked immunosorbent assay (ELISA) and fluorescent quantitative PCR, can usually only detect a few known markers, and are limited by sample size and experimental techniques, making it difficult to fully reflect the complex biological information of the maternal-fetal interface. SUMMARY
[0003] Therefore, it is necessary to provide an intelligent transcriptome analysis system based on biomarkers of the maternal-fetal interface to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, an intelligent transcriptome analysis system based on biomarkers of the maternal-fetal interface comprises the following modules:
[0005] A maternal-fetal RNA low-temperature storage module is used to obtain a maternal-fetal interface biological blood sample set, which includes not less than 500 blood samples for each stage of early pregnancy, mid-pregnancy and late pregnancy, and divides the maternal-fetal interface biological blood sample set into three sample groups for RNA extraction and low-temperature storage to generate corresponding maternal-fetal blood RNA low-temperature sample groups;
[0006] A maternal-fetal genome sequencing calculation module is used to construct maternal-fetal blood RNA sample libraries through corresponding maternal-fetal blood RNA low-temperature sample groups, which include maternal-fetal RNA gene sample groups and maternal-fetal RNA transcription sample groups, and perform single-end sequencing calculation on the corresponding maternal-fetal RNA gene sample groups in the maternal-fetal blood RNA sample libraries to obtain RNA sample base quality scores;
[0007] A transcriptome differential expression determination module is used to perform reference gene mapping screening on the maternal-fetal RNA gene sample groups based on the RNA sample base quality scores to generate a maternal-fetal RNA reference genome; and determine differential expression genes based on the maternal-fetal RNA reference genome to generate a maternal-fetal RNA differential expression genome;
[0008] The Transcriptional Differential Expression Network Construction Module is used to perform target gene clustering and enrichment analysis on the differentially expressed maternal RNA genome to obtain the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set. It searches for expression pattern associations between target genes through a preset database, and performs gene expression network analysis on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern associations between target genes to generate the maternal-fetal target gene expression network.
[0009] Furthermore, the maternal-fetal RNA cryopreservation module includes the following functions:
[0010] Obtain a biological blood sample set from the maternal-fetal interface, including no less than 500 blood samples for each stage of early pregnancy, mid-pregnancy, and late pregnancy;
[0011] The maternal-fetal interface biological blood sample set was divided into three corresponding sample groups according to the experimental requirements. One sample was placed in liquid nitrogen for rapid freezing, one sample was placed in 5 ml of 10% formaldehyde fixative and sealed with paraffin for later use, and one sample was placed in 1.5 ml of RNAlaUer preservation solution, centrifuged and sealed to generate three corresponding maternal-fetal biological blood sample groups.
[0012] By passing the corresponding maternal-fetal biological blood sample group into the corresponding kit for RNA extraction, a maternal-fetal blood RNA sample group is generated.
[0013] The RNA concentration of the maternal-fetal blood RNA sample group was measured by utilizing the specific binding between fluorescent dye and RNA molecules. Based on the RNA concentration, the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group was calculated to obtain the integrity RIN number of the maternal-fetal blood RNA sample.
[0014] Based on the integrity RIN number corresponding to the maternal-fetal blood RNA sample, the integrity quality control of the corresponding RNA samples in the maternal-fetal blood RNA sample group was carried out to screen out RNA samples with an integrity RIN number greater than 7 that meet the integrity quality requirements. The screened RNA samples that meet the integrity quality requirements were aliquoted into RNase-free cryovials and stored in an ultra-low temperature freezer at -80°C to generate the corresponding maternal-fetal blood RNA low-temperature sample group.
[0015] Furthermore, the calculation of integrity indicators for corresponding RNA samples within the maternal-fetal blood RNA sample group based on RNA concentration includes:
[0016] RNA electrophoresis analysis was performed on corresponding RNA samples within the maternal-fetal blood RNA sample group based on RNA concentration to generate corresponding maternal-fetal blood RNA electrophoresis patterns.
[0017] The main peak and secondary peak values of RNA electrophoresis for each RNA sample were obtained by using maternal-fetal blood RNA electrophoresis patterns.
[0018] Based on the main peak and secondary peak of RNA electrophoresis corresponding to the RNA sample, the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group is calculated. The height ratio between the main peak and secondary peak of RNA electrophoresis is calculated to obtain the corresponding RIN number, so as to obtain the integrity RIN number of the maternal-fetal blood RNA sample.
[0019] Furthermore, the maternal-fetal genome sequencing computation module includes the following functions:
[0020] By dividing the corresponding maternal-fetal blood RNA low-temperature sample group into the corresponding maternal-fetal RNA gene sample group and the maternal-fetal RNA to be transcribed sample group;
[0021] The maternal-fetal RNA transcriptome was generated by reverse transcription of the maternal-fetal RNA sample group using an RNA amplification kit.
[0022] The cDNA fragments corresponding to the maternal fetal RNA transcriptome were randomly fragmented, repaired at the ends, A-tailed and ligated with adapters. The processed maternal fetal RNA transcriptome and maternal fetal RNA gene genome were purified and library constructed to generate a maternal fetal blood RNA sample library.
[0023] Single-end sequencing calculations were performed on the corresponding maternal-fetal RNA genome samples in the maternal-fetal blood RNA sample library to obtain the corresponding RNA sample base quality fraction.
[0024] Furthermore, the single-end sequencing calculation of the corresponding maternal-fetal RNA genome sample group within the maternal-fetal blood RNA sample library includes:
[0025] The QubiU-dsDNA-HS kit was used to quantitatively screen the corresponding maternal fetal RNA gene sample groups in the maternal fetal blood RNA sample library to generate corresponding representative samples for maternal fetal RNA gene quantification.
[0026] By setting the corresponding sequencing read length to 140bp, and performing single-end sequencing on a representative sample of maternal fetal RNA gene quantification based on the sequencing read length, the maternal fetal RNA gene expression sequencing coding region was generated.
[0027] Base quality assessment calculations were performed on the coding regions of maternal-fetal RNA gene expression sequencing to obtain the corresponding RNA sample base quality scores.
[0028] Furthermore, the calculation of base quality assessment for the coding region of maternal-fetal RNA gene expression sequencing includes:
[0029] Statistical analysis of base distribution in the coding region of maternal-fetal RNA gene expression sequencing was performed to statistically analyze the number and proportion of A, U, C and G bases within the coding region window, thereby obtaining the base distribution of the maternal-fetal RNA gene coding region.
[0030] Based on the base distribution in the coding region of maternal RNA genes, base mismatch assessment was performed on each base in the coding region of maternal RNA gene expression sequencing to obtain the probability of base mismatch in maternal RNA genes.
[0031] The hydrogen bond strength and spatial structure free energy between each base were obtained by sequencing the expression region of maternal RNA gene. The base signal intensity was evaluated based on the hydrogen bond strength and spatial structure free energy between each base to obtain the base signal intensity of maternal RNA gene.
[0032] Based on the probability of base mismatch at maternal-fetal RNA gene base positions and the base signal intensity of maternal-fetal RNA gene, the base quality of the corresponding bases in the coding region of maternal-fetal RNA gene expression sequencing is calculated to obtain the base quality score of the corresponding RNA sample.
[0033] Furthermore, the transcriptome differential expression determination module includes the following functions:
[0034] Reference gene mapping screening of maternal-fetal RNA genome samples was performed based on the base mass fraction of RNA samples to generate a maternal-fetal RNA reference genome.
[0035] Sequence alignment calculations were performed between the corresponding RNA transcription samples in the maternal fetal RNA transcription sample group based on the reference genes corresponding to each reference gene in the maternal fetal RNA reference genome to obtain the sequence alignment similarity E value between each RNA transcription sample and the corresponding reference gene.
[0036] Based on the sequence similarity E value between each RNA transcription sample and the corresponding reference gene, the maternal-fetal RNA transcription sample group was screened for reference similarity samples. The top 10 RNA transcription sample sequences with a sequence similarity E value less than or equal to 1e-3 were selected by comparison and screening, and maternal-fetal RNA transcription sample sequences similar to the reference genome were generated.
[0037] GO annotation was extracted from maternal fetal RNA transcription sample sequences that were similar to the reference genome to extract GO entries related to the target gene, thus obtaining maternal fetal RNA transcription GO annotation sample sequences.
[0038] Differentially expressed genes were identified between two maternal RNA transcription samples corresponding to the GO annotation sequence of maternal RNA transcription. The fold change in gene expression level and the P-value between the two maternal RNA transcription samples were analyzed to determine the significance. Transcription genes with a fold change in gene expression level greater than 4 and a P-value less than 0.05 were identified as differentially expressed genes to generate a differentially expressed maternal RNA genome.
[0039] Furthermore, the reference gene mapping screening specifically involves comparing the base quality fraction of RNA samples based on a preset base quality threshold. If the base quality fraction of an RNA sample is greater than or equal to the preset base quality threshold, then the corresponding RNA sample within its maternal-fetal RNA gene sample group is mapped to the corresponding reference genome; otherwise, no mapping is performed, thereby generating a maternal-fetal RNA reference genome.
[0040] Furthermore, the transcriptional differential expression network building module includes the following functions:
[0041] Cluster analysis of target genes in the differentially expressed maternal-fetal RNA genome was performed to generate a set of target genes for differentially expressed maternal-fetal RNA.
[0042] By analyzing the set of target genes for differential expression in maternal-fetal RNA, the main metabolic pathways and signal transduction pathways involved by differentially expressed genes in each gene category were obtained.
[0043] Gene enrichment analysis was performed on the gene categories corresponding to the differentially expressed genes in the maternal-fetal RNA differential expression target gene set based on the main metabolic pathways and signal transduction pathways involved in each gene category, so as to obtain the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set.
[0044] The expression pattern associations between target genes are found using a pre-defined database.
[0045] Gene expression network analysis was performed on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern association between target genes, in order to generate the maternal-fetal target gene expression network.
[0046] Furthermore, the target gene clustering analysis of the differentially expressed maternal-fetal RNA genome includes:
[0047] Quantitative statistical analysis was performed on each gene sample within the differentially expressed maternal-fetal RNA genome to obtain the sample expression level and gene expression level of each differentially expressed maternal-fetal RNA gene sample.
[0048] Based on the expression levels of each differentially expressed maternal RNA gene sample and the gene expression levels, target gene clustering analysis is performed on each gene sample within the differentially expressed maternal RNA genome. Hierarchical clustering is performed by calculating the Euclidean distance between each gene sample from both the expression level and gene expression level to generate a set of differentially expressed maternal RNA target genes.
[0049] The beneficial effects of this invention are:
[0050] The intelligent transcriptome analysis system based on maternal-fetal interface biomarkers proposed in this invention comprises a maternal-fetal RNA cryogenic storage module, a maternal-fetal genome sequencing and computation module, a transcriptome differential expression determination module, and a transcriptional differential expression network construction module. Compared with existing technologies, the advantages of this application lie in ensuring sample diversity and representativeness by collecting at least 500 blood samples from different gestational stages (early, mid, and late pregnancy), thereby improving the reliability and breadth of research results. This process involves rigorous sample collection and classification to ensure that the number of samples meets the requirements for statistical analysis. Simultaneously, RNA extraction and cryogenic storage effectively preserve the integrity of the RNA, preventing degradation during storage. Cryogenic storage extends the lifespan of the samples, allowing for analysis at different time points while maintaining their original biological characteristics. RNA samples preserved in this way provide stable materials for subsequent transcriptome research, thus providing effective genomic data support for analyzing the maternal-fetal interface, maternal-fetal interaction, and fetal health. Secondly, after RNA extraction and cryogenic storage, the next step is to construct a high-quality maternal-fetal blood RNA sample library. This library provides the foundation for subsequent gene expression analysis and transcriptome sequencing. Crucially, during library construction, the proper division and classification of maternal-fetal RNA genomic and transcriptomic sample groups is essential. Single-end sequencing of the maternal-fetal RNA genomic sample group yields the base quality score of each gene. These scores are key indicators for evaluating sequencing data quality, reflecting errors and their impact during sequencing. This process allows for the selection of high-quality RNA samples, ensuring subsequent data analysis is conducted at low noise levels, thus improving the accuracy and reliability of gene expression data. Next, after obtaining the base quality scores of the maternal-fetal RNA samples, reference gene mapping is performed. Gene mapping compares these RNA samples with existing reference genomes, identifying the genes present in each sample and their expression patterns. Based on these reference genomes, further screening for differentially expressed genes helps researchers identify specific gene expression patterns at the maternal-fetal interface, revealing the complex physiological and molecular interactions between the mother and fetus.Finally, after obtaining the differentially expressed genomes of maternal-fetal RNA, target gene clustering and enrichment analysis are necessary steps to further understand gene function and biological significance. Clustering analysis of differentially expressed genes can group genes with similar expression patterns together, thereby revealing the biological processes or pathways they participate in. This clustering analysis can not only help researchers understand the changes in gene expression between the mother and fetus at different stages of pregnancy, but also reveal key regulatory factors under different biological conditions. Enrichment analysis helps researchers discover the distribution of target genes in different biological pathways, thereby further exploring the potential roles of these genes in fetal development, immune regulation, or metabolic regulation. In addition, based on the association of expression patterns among these target genes, maternal-fetal target gene expression networks can be constructed. Gene expression network analysis can reveal the interrelationships and regulatory effects between genes, providing a visual tool for studying the molecular mechanisms of maternal-fetal interaction. By analyzing gene expression networks, we can delve into the interactions and regulatory mechanisms between key genes, thus comprehensively reflecting the complex biological information at the maternal-fetal interface. Attached Figure Description
[0051] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0052] Figure 1 This is a schematic diagram of the modules of the intelligent transcriptome analysis system based on maternal-fetal interface biomarkers of the present invention;
[0053] Figure 2 for Figure 1 A functional flowchart of the maternal-fetal RNA cryopreservation module;
[0054] Figure 3 for Figure 1 A functional flowchart of the maternal-fetal genome sequencing computation module. Detailed Implementation
[0055] The technical system of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0056] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor systems and / or microcontroller systems.
[0057] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0058] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides an intelligent transcriptome analysis system based on maternal-fetal interface biomarkers, the system comprising the following modules:
[0059] The maternal-fetal RNA cryopreservation module is used to obtain a maternal-fetal interface biological blood sample set, including no less than 500 blood samples for each stage of early pregnancy, mid pregnancy and late pregnancy. The maternal-fetal interface biological blood sample set is divided into three sample groups for simultaneous RNA extraction and cryopreservation to generate corresponding maternal-fetal blood RNA cryopreservation sample groups.
[0060] The maternal-fetal genome sequencing calculation module is used to construct a maternal-fetal blood RNA sample library from the corresponding maternal-fetal blood RNA low-temperature sample group, which includes the maternal-fetal RNA gene sample group and the maternal-fetal RNA transcript sample group, and to perform single-end sequencing calculation on the corresponding maternal-fetal RNA gene sample group in the maternal-fetal blood RNA sample library to obtain the corresponding RNA sample base quality fraction.
[0061] The transcriptome differential expression determination module is used to perform reference gene mapping screening on maternal fetal RNA gene sample groups based on the base quality fraction of RNA samples to generate a maternal fetal RNA reference genome; and to determine differentially expressed genes on maternal fetal RNA transcriptomes based on the maternal fetal RNA reference genome to generate a maternal fetal RNA differentially expressed genome.
[0062] The Transcriptional Differential Expression Network Construction Module is used to perform target gene clustering and enrichment analysis on the differentially expressed maternal RNA genome to obtain the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set. It searches for expression pattern associations between target genes through a preset database, and performs gene expression network analysis on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern associations between target genes to generate the maternal-fetal target gene expression network.
[0063] In the embodiments of this invention, please refer to Figure 1 The diagram shown is a schematic representation of the modules of the intelligent transcriptome analysis system based on maternal-fetal interface biomarkers of the present invention. In this example, the intelligent transcriptome analysis system based on maternal-fetal interface biomarkers includes the following modules:
[0064] S1: Maternal-fetal RNA cryopreservation module, used to obtain a maternal-fetal interface biological blood sample set, including no less than 500 blood samples for each stage of early pregnancy, mid pregnancy and late pregnancy, and divide the maternal-fetal interface biological blood sample set into three sample groups for simultaneous RNA extraction and cryopreservation to generate corresponding maternal-fetal blood RNA cryopreservation sample groups.
[0065] In this embodiment of the invention, the collection of maternal-fetal interface biological blood samples was carried out at the sample collection center of a professional obstetrics and gynecology hospital with the informed consent of the pregnant women and the approval of the hospital's ethics committee. Samples were collected for three stages: early pregnancy (1-12 weeks), mid pregnancy (13-27 weeks), and late pregnancy (28 weeks and beyond). 5ml disposable vacuum blood collection tubes were used, and professional nurses drew blood samples from the pregnant women's elbow veins. For example, during the early pregnancy sample collection, two blood collection sessions were arranged each day, one in the morning and one in the afternoon, with 5 nurses assigned to each session. Over two weeks, 530 early pregnancy blood samples were successfully collected. Similarly, blood samples were collected from 510 cases in the second trimester and 520 cases in the third trimester to obtain a maternal-fetal interface biological blood sample set. In the laboratory, each blood sample was divided into three portions. Using a sterile pipette, approximately 2 ml of blood sample was precisely pipetted and quickly placed in a liquid nitrogen tank (maintained at -196°C) for rapid freezing. Another approximately 2 ml blood sample was added to a sample bottle containing 5 ml of 10% formaldehyde fixative, gently shaken to mix, and then sealed with paraffin. The last approximately 1 ml blood sample was added to a centrifuge tube containing 1.5 ml of RNAlater preservation solution, centrifuged (set to 3000 rpm, centrifuged for 5 minutes), and then sealed. Each sample was clearly labeled, including the pregnant woman's information, gestational age, and sample type. The portion of the sample containing RNAlater preservation solution was selected and processed using Qiagen's RNeasy Mini... RNA extraction was performed using the kit. Following the kit instructions, the sample was mixed with lysis buffer to lyse the cells. RNA separation, washing, and elution were then performed sequentially. The extracted RNA sample was immediately transferred to an RNase-free cryovial and stored in an ultra-low temperature freezer at -80°C to generate a low-temperature sample set of maternal-fetal blood RNA.
[0066] S2: Maternal-fetal genome sequencing calculation module, used to construct a maternal-fetal blood RNA sample library from the corresponding maternal-fetal blood RNA low-temperature sample group, including the maternal-fetal RNA gene sample group and the maternal-fetal RNA transcript sample group, and to perform single-end sequencing calculation on the corresponding maternal-fetal RNA gene sample group in the maternal-fetal blood RNA sample library to obtain the corresponding RNA sample base quality fraction.
[0067] In this embodiment of the invention, cryopreserved tubes from the maternal-fetal blood RNA cryopreservation sample group were removed from a -80°C ultra-low temperature freezer in a molecular biology laboratory. 70% of the samples were assigned to the maternal-fetal RNA gene sample group, and 30% to the maternal-fetal RNA transcription sample group. Sterile pipettes were used for precise manipulation and labeling. For the maternal-fetal RNA to be transcribed sample group, the Thermo Fisher Scientific SuperScript IV First-Strand Synthesis System RNA amplification kit was used for reverse transcription. Taking a specific sample as an example, 10 μl of RNA sample was added to a reaction tube containing reverse transcription reaction buffer, dNTPs, random primers, and reverse transcriptase, adjusting the total volume to 20 μl. After mixing, the sample was placed in a PCR instrument and reverse transcribed according to the kit's recommended program (25°C for 10 minutes, 50°C for 60 minutes, and 85°C for 5 minutes) to generate cDNA, i.e., the maternal-fetal RNA transcription sample group. For the maternal-fetal RNA gene sample group, Covaris was used... The cDNA was randomly fragmented using an S220 focused ultrasound disruptor, with a power of 140W, a working time of 30 seconds, and 5 cycles. Then, the NEB Next Ultra II End-Repair / dA-Tailing Module kit was used for end repair (37°C for 30 minutes), A-tailing (70°C for 30 minutes), and adapter ligation (20°C for 15 minutes). The maternal-fetal RNA genome and the processed maternal-fetal RNA transcriptome were then purified using the Qiagen MinElute PCR Purification Kit to remove impurities and primer dimers. Library construction was then performed using the Illumina TruSeq DNA PCR-Free Library Preparation Kit, involving end repair, A-tailing, adapter ligation, fragment selection, and PCR amplification to generate a maternal-fetal blood RNA sample library. Finally, single-end sequencing of the maternal-fetal RNA genome was performed using the Illumina HiSeq 2500 sequencing platform in a high-throughput sequencing laboratory. The sample was diluted to 5pM, the sequencing read length was set to 140bp, and it was loaded onto the sequencing chip using a dedicated sample loading device. The sequencing platform used sequencing-by-synthesis technology, and the base types were determined by detecting the fluorescence signal generated by the binding of fluorescently labeled dNTPs to the template strand. After sequencing, the data was processed using the Illumina BaseSpace software that came with the sequencing platform to calculate the signal intensity and quality value of each base position, and finally obtain the base quality fraction of the RNA sample.
[0068] S3: Transcriptome differential expression determination module, used to perform reference gene mapping screening on maternal fetal RNA gene sample groups based on the base quality fraction of RNA samples to generate maternal fetal RNA reference genome; and to determine differentially expressed genes on maternal fetal RNA transcriptomes based on maternal fetal RNA reference genome to generate maternal fetal RNA differential expression genome.
[0069] In this embodiment of the invention, the maternal-fetal RNA gene sample group is screened for reference gene mapping based on the base quality score of the RNA samples using BWA (Burrows-Wheeler Aligner) software. A preset base quality threshold of 70% is used. The base quality score of each maternal-fetal RNA gene sample is compared with the threshold. For samples that reach or exceed 70%, sequence alignment mapping is performed using BWA software with the human reference genome (GRCh38 version) as a reference. For example, if there are 100 samples in the maternal-fetal RNA gene sample group, and 65 samples meet the base quality score requirement, the BWA software maps these 65 samples to the corresponding positions in the reference genome, generating maternal-fetal RNA gene samples. Using the maternal-fetal RNA reference genome, the edgeR package in R was used on a data analysis workstation to identify differentially expressed genes in maternal-fetal RNA transcriptome samples based on the maternal-fetal RNA reference genome. The results were then imported into the R environment to construct an experimental design matrix using the edgeR package, setting grouping information, and employing a precise test algorithm to perform significance analysis on the two groups, calculating the fold change in gene expression levels and the p-value. The screening criteria were set as a fold change in gene expression levels greater than 4 and a p-value less than 0.05. Transcriptional genes meeting these criteria were selected as differentially expressed genes, ultimately generating a maternal-fetal RNA differentially expressed genome.
[0070] S4: Transcriptional Differential Expression Network Construction Module, used to perform target gene clustering and enrichment analysis on maternal-fetal RNA differential expression genomes to obtain the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set; to find the expression pattern association between target genes through a preset database, and to perform gene expression network analysis on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern association between target genes, so as to generate the maternal-fetal target gene expression network.
[0071] In this embodiment of the invention, the clusterProfiler package in R language is used to perform target gene clustering analysis on the differentially expressed maternal-fetal RNA genome. By using the pamk function and employing a partitioning clustering method, genes with similar expression patterns are grouped into one class by calculating the expression similarity between genes (such as Euclidean distance). The number of clusters is set to 5, for example, 100 differentially expressed genes are divided into 5 categories to generate a set of target genes for differentially expressed maternal-fetal RNA. Next, using the clusterProfiler package, gene enrichment analysis was performed on the gene categories corresponding to the differentially expressed genes in the maternal-fetal RNA differential expression target gene set, based on the main metabolic pathways and signal transduction pathways involved in each gene category. The enrichKEGG function was used, referencing the KEGG database, to analyze the enrichment degree of each gene category in specific metabolic and signal transduction pathways. For example, for a gene category involved in the "PI3K-Akt signaling pathway," an enrichment score was calculated to assess its significance in that pathway. The gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set was generated and stored in the form of visual charts (such as bar charts and bubble charts). This data was then accessed on the data analysis workstation via STRING (Search Tool for the Retrieval of Interacting Genes). The Genes / Proteins database was used to search for expression pattern associations between target genes. A list of all target genes was extracted, converted to a STRING database-recognizable format, and entered into the database search interface. The database integrates protein-protein interaction and gene co-expression data, showing, for example, that gene A and gene B exhibit positively correlated expression patterns in multiple cell types. The expression pattern association information between target genes was organized into a network table. Finally, Cytoscape software was used to perform gene expression network analysis on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern associations between target genes. The data was imported into Cytoscape software, with genes as nodes, gene expression associations as edges, and gene pathways as node attributes. Parameters such as node color, shape, and edge thickness and color were set to display gene relationships and pathway distribution. For example, gene nodes participating in the same gene pathway were set to the same color, and the edge thickness was adjusted according to the strength of gene expression associations, ultimately generating a maternal-fetal target gene expression network.
[0072] Furthermore, the maternal-fetal RNA cryopreservation module includes the following functions:
[0073] Obtain a biological blood sample set from the maternal-fetal interface, including no less than 500 blood samples for each stage of early pregnancy, mid-pregnancy, and late pregnancy;
[0074] The maternal-fetal interface biological blood sample set was divided into three corresponding sample groups according to the experimental requirements. One sample was placed in liquid nitrogen for rapid freezing, one sample was placed in 5 ml of 10% formaldehyde fixative and sealed with paraffin for later use, and one sample was placed in 1.5 ml of RNAlaUer preservation solution, centrifuged and sealed to generate three corresponding maternal-fetal biological blood sample groups.
[0075] By passing the corresponding maternal-fetal biological blood sample group into the corresponding kit for RNA extraction, a maternal-fetal blood RNA sample group is generated.
[0076] The RNA concentration of the maternal-fetal blood RNA sample group was measured by utilizing the specific binding between fluorescent dye and RNA molecules. Based on the RNA concentration, the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group was calculated to obtain the integrity RIN number of the maternal-fetal blood RNA sample.
[0077] Based on the integrity RIN number corresponding to the maternal-fetal blood RNA sample, the integrity quality control of the corresponding RNA samples in the maternal-fetal blood RNA sample group was carried out to screen out RNA samples with an integrity RIN number greater than 7 that meet the integrity quality requirements. The screened RNA samples that meet the integrity quality requirements were aliquoted into RNase-free cryovials and stored in an ultra-low temperature freezer at -80°C to generate the corresponding maternal-fetal blood RNA low-temperature sample group.
[0078] As an embodiment of the present invention, reference is made to... Figure 2 As shown, Figure 1 A functional flowchart of the maternal-fetal RNA cryopreservation module is shown. In this embodiment, the maternal-fetal RNA cryopreservation module includes the following functions:
[0079] S11: Obtain a biological blood sample set from the maternal-fetal interface, including no less than 500 blood samples for each stage of early pregnancy, mid-pregnancy and late pregnancy;
[0080] In this embodiment of the invention, the collection of maternal-fetal interface biological blood samples is carried out in a professional obstetrics and gynecology medical institution after communicating with and obtaining approval from the hospital's ethics committee and with the informed consent of the pregnant woman. Samples are collected in three stages: early pregnancy (1-12 weeks), mid pregnancy (13-27 weeks), and late pregnancy (28 weeks and beyond). At least 500 blood samples are planned to be collected in each stage. In the blood collection room, professionally trained nurses use disposable vacuum blood collection tubes to draw 5ml of venous blood from the pregnant woman's elbow vein. For example, in the early pregnancy sample collection, blood is collected twice a day, in the morning and afternoon, within a week, with 5 nurses operating simultaneously each time. After two weeks, 520 early pregnancy blood samples were successfully collected. The collected blood samples are immediately labeled with the pregnant woman's basic information (name initials, age, gestational week) and the time of blood collection. They are then placed in a special sample transport box equipped with ice packs to maintain a low temperature environment and quickly transported to the laboratory for further processing, thereby obtaining a maternal-fetal interface biological blood sample set.
[0081] S12: Divide the maternal-fetal interface biological blood sample set into three corresponding sample groups according to the experimental requirements. One sample is placed in liquid nitrogen for rapid freezing, one sample is placed in 5ml of 10% formaldehyde fixative and sealed with paraffin for later use, and one sample is placed in 1.5ml of RNAlaUer preservation solution, centrifuged and sealed to generate three corresponding maternal-fetal biological blood sample groups.
[0082] In this embodiment of the invention, by preparing the corresponding experimental equipment and reagents in the sample processing area of the laboratory, for the obtained maternal-fetal interface biological blood sample set, each blood sample is divided into three corresponding portions according to the experimental requirements. Using a sterile pipette, approximately 2 ml of blood sample is accurately drawn and quickly placed into a pre-cooled liquid nitrogen tank for rapid freezing. The temperature of the liquid nitrogen tank is maintained at -196°C to ensure that the sample is frozen in a very short time to preserve the original state of the sample to the greatest extent. Another approximately 2 ml of blood sample is drawn and slowly added to a special sample bottle containing 5 ml of 10% formaldehyde fixative. The bottle is gently shaken to mix the blood and fixative thoroughly. Then, the sample bottle is sealed with paraffin and placed on a sample storage rack for later use. Finally, take about 1 ml of blood sample and add it to a centrifuge tube containing 1.5 ml of RNAlaUer preservation solution. Place the centrifuge tube in a centrifuge, set the speed to 3000 rpm, and centrifuge for 5 minutes. After centrifugation, seal the centrifuge tube with sealing film to generate three corresponding maternal-fetal biological blood sample groups. Place them in the corresponding sample storage area and label them for later retrieval.
[0083] S13: By passing the corresponding maternal-fetal biological blood sample group into the corresponding kit for RNA extraction, a maternal-fetal blood RNA sample group is generated.
[0084] In this embodiment of the invention, RNA extraction is performed in a clean bench in a molecular biology laboratory using the corresponding RNeasy Mini KiU kit. This kit contains various reagents and consumables required for RNA extraction. For each maternal-fetal biological blood sample group, taking the sample centrifuged and sealed in 1.5 ml of RNAlaUer preservation solution as an example, the centrifuge tube is opened, and the blood sample and RNAlaUer preservation solution are transferred to the dedicated lysis buffer provided by the kit. The cells are thoroughly lysed using a pipette. Then, following the steps in the kit instructions, different reagents are added sequentially to perform RNA separation, washing, and elution operations. For example, ethanol is added to promote RNA precipitation. The RNA precipitate is separated by centrifugation, and then the RNA precipitate is washed with 70% ethanol to remove impurities. Finally, RNA is eluted from the adsorption column using RNase-free water. This process is performed on each maternal-fetal biological blood sample group to generate a maternal-fetal blood RNA sample group. The extracted RNA samples are then temporarily placed on an ice box in preparation for the next step of concentration and integrity testing.
[0085] S14: The RNA concentration of the maternal-fetal blood RNA sample group was measured by utilizing the specific binding between fluorescent dye and RNA molecules, and the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group was calculated based on the RNA concentration to obtain the integrity RIN number of the maternal-fetal blood RNA sample.
[0086] In this embodiment of the invention, the RNA concentration of the maternal-fetal blood RNA sample group was measured using a NanoDrop 2000 spectrophotometer in the operating area of a laboratory spectrophotometer. This instrument calculates the RNA concentration based on Beer-Lambert's law by detecting the absorbance of RNA molecules at a specific wavelength. 1-2 μl of the extracted RNA sample was placed on the spectrophotometer's detection platform, and the instrument automatically measured and displayed the RNA concentration value. For example, the measurement result for a certain RNA sample showed a concentration of 500 ng / μl. Simultaneously, an AgilenU 2100 bioanalyzer was used to calculate the integrity index of the RNA samples. This instrument utilizes microfluidic chip technology to electrophoretically separate the RNA samples on the chip. By detecting the distribution of RNA bands, the integrity RIN number (RNA Integrity Number) of the maternal-fetal blood RNA sample was calculated. For example, after analysis, the integrity RIN number of this RNA sample was 8.5. This concentration measurement and integrity index calculation were performed on each RNA sample within the maternal-fetal blood RNA sample group, and the results were recorded in the experimental data record table.
[0087] S15: Based on the integrity RIN number corresponding to the maternal-fetal blood RNA sample, the integrity quality control of the corresponding RNA samples in the maternal-fetal blood RNA sample group is performed to screen out RNA samples with integrity RIN numbers greater than 7 that meet the integrity quality requirements. The screened RNA samples that meet the integrity quality requirements are aliquoted into RNase-free cryovials and stored in an ultra-low temperature freezer at -80℃ to generate the corresponding maternal-fetal blood RNA low-temperature sample group.
[0088] In this embodiment of the invention, the integrity quality of RNA samples within the maternal-fetal blood RNA sample group is controlled based on the recorded integrity RIN number of the maternal-fetal blood RNA samples in the sample storage area. The integrity RIN number of each RNA sample is checked individually, and RNA samples with an integrity RIN number greater than 7 are selected. These samples are deemed to have qualified integrity quality. Using a sterile pipette, 100 μl of the selected qualified RNA samples are precisely aspirated and dispensed into RNase-free cryovials. Each cryovial is clearly labeled with information including sample number, RNA concentration, and integrity RIN number. The cryovials containing the qualified RNA samples are then placed in a cryovial box and quickly transferred to a -80°C ultra-low temperature freezer for storage. This generates a corresponding maternal-fetal blood RNA low-temperature sample group, providing high-quality RNA samples for subsequent intelligent transcriptome analysis based on maternal-fetal interface biomarkers.
[0089] Furthermore, the calculation of integrity indicators for corresponding RNA samples within the maternal-fetal blood RNA sample group based on RNA concentration includes:
[0090] RNA electrophoresis analysis was performed on corresponding RNA samples within the maternal-fetal blood RNA sample group based on RNA concentration to generate corresponding maternal-fetal blood RNA electrophoresis patterns.
[0091] In this embodiment of the invention, RNA electrophoresis analysis of corresponding RNA samples within a maternal-fetal blood RNA sample group is performed using a vertical electrophoresis system in the molecular biology experimental area of the laboratory. An RNA sample is selected from the extracted maternal-fetal blood RNA sample group, which is temporarily stored in an icebox. First, a 1% agarose gel is prepared. 1g of agarose powder is accurately weighed and added to 100ml of 1×UAE (Uris-acetic acid-EDUA) buffer. The mixture is heated and stirred in a microwave oven until the agarose is completely dissolved. After cooling to approximately 50°C, an appropriate amount of nucleic acid dye (such as GoldView) is added, and the mixture is gently shaken. The mixture is poured into a gel mold, a comb is inserted, and after the gel solidifies, the comb is carefully removed. The gel is placed in an electrophoresis tank, and 1×UAE buffer is added until it covers the gel. Using a pipette, 10μl of RNA sample is thoroughly mixed with 2μl of 6× loading buffer and slowly added to the sample wells of the gel. Simultaneously, RNA molecular weight standards (such as UhermoFisher) are added to adjacent sample wells. Using ScienUific's 1kb Plus DNA Ladder as a reference, the power was turned on and the voltage was set to 120V for 30 minutes of electrophoresis. During electrophoresis, RNA molecules migrate towards the positive electrode under the influence of the electric field, and RNA fragments of different sizes are separated on the gel due to their different migration rates. After electrophoresis, the gel was placed in a gel imaging system (such as Bio-Rad's GelDoc XR+SysUem) for imaging, ultimately generating the corresponding maternal-fetal blood RNA electrophoresis pattern.
[0092] Preferably, the main peak and secondary peak of RNA electrophoresis for each RNA sample are obtained by using maternal-fetal blood RNA electrophoresis patterns;
[0093] In this embodiment of the invention, the maternal-fetal blood RNA electrophoresis pattern is analyzed using professional image analysis software (such as ImageJ) in the laboratory data analysis area to obtain the main peak and secondary peak of RNA electrophoresis for each RNA sample. The ImageJ software is opened, the previously generated maternal-fetal blood RNA electrophoresis pattern is imported, and the "gel analysis" function under the "analysis" menu of the software is used to first calibrate the lane where the RNA molecular weight standard is located, and set the known molecular weight band position so that the software can accurately identify the size of RNA fragments corresponding to different positions. Then, the lane where the RNA sample is located is selected for analysis. The software generates a curve of gray value versus migration distance by detecting the gray value distribution of the electrophoretic bands. On the curve, obvious peaks represent concentrated areas of different RNA fragments. By observing the curves, the highest peak was identified as the main peak of RNA electrophoresis, representing the most abundant RNA fragment. For example, the main peak of RNA electrophoresis for this sample appeared at a migration distance of 2 cm, and the corresponding RNA fragment size was determined to be 28S rRNA after calibration. The secondary peaks were relatively low but still significant peaks. For example, the secondary peak that appeared at a migration distance of 3 cm corresponded to 18S rRNA. The migration distances corresponding to the main and secondary peaks of RNA electrophoresis, as well as the RNA fragment size obtained according to calibration, were recorded in the experimental data recording table.
[0094] Preferably, the integrity index of the corresponding RNA samples in the maternal-fetal blood RNA sample group is calculated based on the main peak and secondary peak of RNA electrophoresis corresponding to the RNA sample. The height ratio between the main peak and secondary peak of RNA electrophoresis is calculated to obtain the corresponding RIN number, so as to obtain the integrity RIN number corresponding to the maternal-fetal blood RNA sample.
[0095] In this embodiment of the invention, the integrity index of corresponding RNA samples within the maternal-fetal blood RNA sample group is calculated by using the RNA electrophoresis main peak and secondary peak data of RNA samples recorded in a previous file in the data analysis area. According to industry-standard calculation methods, the corresponding RIN number is obtained by calculating the height ratio between the RNA electrophoresis main peak and secondary peak. The corresponding file is opened in Excel, and a new column is added to record the RIN number. Assuming the grayscale value of the RNA electrophoresis main peak is Pa and the grayscale value of the RNA electrophoresis secondary peak is Pb, the height ratio is calculated in Excel using the formula "=Pa / Pb". For example, if the grayscale value of the RNA electrophoresis main peak of a sample is 800 and the grayscale value of the RNA electrophoresis secondary peak is 400, the calculated height ratio is 2. This height ratio is converted into the RIN number using a pre-established standard curve. Assuming that according to the standard curve, the RIN number corresponding to a height ratio of 2 is 8.0, the calculated integrity RIN number corresponding to the maternal-fetal blood RNA sample is filled into the newly added column, finally obtaining the integrity RIN number for all samples.
[0096] Furthermore, the maternal-fetal genome sequencing computation module includes the following functions:
[0097] By dividing the corresponding maternal-fetal blood RNA low-temperature sample group into the corresponding maternal-fetal RNA gene sample group and the maternal-fetal RNA to be transcribed sample group;
[0098] The maternal-fetal RNA transcriptome was generated by reverse transcription of the maternal-fetal RNA sample group using an RNA amplification kit.
[0099] The cDNA fragments corresponding to the maternal fetal RNA transcriptome were randomly fragmented, repaired at the ends, A-tailed and ligated with adapters. The processed maternal fetal RNA transcriptome and maternal fetal RNA gene genome were purified and library constructed to generate a maternal fetal blood RNA sample library.
[0100] Single-end sequencing calculations were performed on the corresponding maternal-fetal RNA genome samples in the maternal-fetal blood RNA sample library to obtain the corresponding RNA sample base quality fraction.
[0101] As an embodiment of the present invention, reference is made to... Figure 3 As shown, Figure 1 A functional flowchart of the maternal-fetal genome sequencing computation module is shown in this embodiment. The maternal-fetal genome sequencing computation module includes the following functions:
[0102] S21: By dividing the corresponding maternal fetal blood RNA low-temperature sample group into the corresponding maternal fetal RNA gene sample group and the maternal fetal RNA to be transcribed sample group.
[0103] In this embodiment of the invention, cryovials are retrieved from the maternal-fetal blood RNA cryopreservation sample group stored at -80°C in the sample processing area of the laboratory. According to a pre-set ratio, for example, 70% of the samples are assigned to the maternal-fetal RNA gene sample group and 30% to the maternal-fetal RNA transcription sample group. Using a sterile pipette, the appropriate volume of RNA sample is precisely aspirated from each cryovial. For example, if a cryovial contains 100 μl of maternal-fetal blood RNA cryopreservation sample, 70 μl is transferred to a new cryovial labeled "Maternal-fetal RNA Gene Sample Group_Sample Number," and the remaining 30 μl is transferred to a cryovial labeled "Maternal-fetal RNA Transcription Sample Group_Sample Number." This operation is performed on all samples in the maternal-fetal blood RNA cryopreservation sample group to ensure accurate sample division. After division, the cryovials of the maternal-fetal RNA gene sample group and the maternal-fetal RNA transcription sample group are placed on different sample racks and clearly labeled for subsequent retrieval.
[0104] S22: The maternal-fetal RNA transcriptome was generated by reverse transcription of the maternal-fetal RNA transcriptome using an RNA amplification kit.
[0105] In this embodiment of the invention, Uhermo Fisher Scientific's SuperScripU IV FirsU-SUrand Synuhesis SysUem was selected in the molecular biology experimental region. The RNA amplification kit is used to reverse transcribe maternal RNA samples. Select a cryovial from the maternal RNA sample group (e.g., "Maternal RNA Sample Group_001") and pipette 10 μl of RNA sample into a reaction tube containing reverse transcription buffer, dNUPs, random primers, and reverse transcriptase. Adjust the total volume to 20 μl, gently mix, and place the tube in a PCR instrument. Perform the reverse transcription reaction according to the kit's recommended procedure: first, incubate at 25°C for 10 minutes to anneal the primers and RNA template; then, incubate at 50°C for 60 minutes for cDNA synthesis; finally, heat at 85°C for 5 minutes to inactivate the reverse transcriptase. Repeat this reverse transcription process for each sample in the maternal RNA sample group to generate the corresponding cDNA. Collect the generated cDNA products in a new cryovial labeled "Maternal RNA Transcription Sample Group_Sample Number" and store at -20°C for further processing.
[0106] S23: Randomly break down, repair the ends of, add A tails and ligates the cDNA fragments corresponding to the maternal fetal RNA transcriptome, and purify and construct a library of maternal fetal blood RNA samples.
[0107] In this embodiment of the invention, the cDNA fragments corresponding to the maternal-fetal RNA transcriptome are subjected to a series of treatments in the sample processing area. This involves randomly fragmenting the cDNA using a Covaris S220 focused ultrasound disruptor with ultrasound parameters set to 140W power, 30 seconds working time, and 5 cycles, resulting in fragments with an average length of approximately 300-500 bp. Then, the NEB NexU UlUra II End-Repair / dA-Uailing Module kit is used for end repair, A-tailing, and adapter ligation. Following the kit instructions, the appropriate reagents are added sequentially, and incubation is performed at 37°C for 30 minutes for end repair, at 70°C for 30 minutes for A-tailing, and at 20°C for 15 minutes for adapter ligation. For the maternal-fetal RNA gene sample, a corresponding purification kit (such as Qiagen MinEluUe PCR) is also used. The maternal fetal RNA transcriptome and maternal fetal RNA gene genome were purified using Illumina UruSeq DNAPCR-Free Library PreparaUiU to remove impurities and primer dimers. The processed maternal fetal RNA transcriptome and maternal fetal RNA gene genome were then mixed and library construction was performed using Illumina UruSeq DNAPCR-Free Library PreparaUiU. After steps such as end repair, A-tailing, adapter ligation, fragment screening, and PCR amplification, a maternal fetal blood RNA sample library was finally generated. The library products were collected in dedicated library tubes and stored at -20°C.
[0108] S24: Perform single-end sequencing calculations on the corresponding maternal-fetal RNA gene sample genome within the maternal-fetal blood RNA sample library to obtain the corresponding RNA sample base quality fraction.
[0109] In an embodiment of the present invention, in a high-throughput sequencing laboratory, single-end sequencing calculation is performed on the maternal-fetal RNA gene sample group corresponding in the maternal-fetal blood RNA sample library. A library tube containing the maternal-fetal RNA gene sample group is taken out from the maternal-fetal blood RNA sample library, and the library product is diluted to a suitable concentration, such as 10 pM. A dedicated sample loading device is used to load the diluted library product onto a sequencing chip. The sequencing platform adopts the sequencing by synthesis (SBS) technology. During the sequencing process, fluorescently labeled dNUPs will bind to the template strand according to the base complementary pairing principle. The sequencer determines the base type at each position by detecting the fluorescence signal. After the sequencing is completed, the sequencing data is processed using the data analysis software (such as Illumina BaseSpace) supporting the sequencing platform. The software calculates the signal intensity and quality value at each base position to obtain the corresponding RNA sample base quality score. For example, after analysis of a certain sample, the proportion of its RNA sample base quality score reaching Q30 (indicating a base recognition error rate of 0.1%) is 90%, and finally the corresponding RNA sample base quality score is obtained.
[0110] Further, the single-end sequencing calculation of the maternal-fetal RNA gene sample group corresponding in the maternal-fetal blood RNA sample library includes:
[0111] Using the QubiU-dsDNA-HS kit to perform QubiU quantitative screening on the maternal-fetal RNA gene sample group corresponding in the maternal-fetal blood RNA sample library to generate a corresponding maternal-fetal RNA gene quantitative representative sample;
[0112] In this embodiment of the invention, in the sample quantification analysis area of the laboratory, a library tube containing a maternal fetal RNA gene sample group is taken from a maternal fetal blood RNA sample library stored at -20°C. The sample is quantitatively screened using the Uhermo Fisher Scientific QubiU dsDNAHS (High Sensitive Detection) kit. A QubiU fluorometer and its matching detection tubes are prepared. Following the kit instructions, the QubiU working solution is first prepared by mixing the QubiU dsDNAHS reagent and QubiU buffer at a ratio of 1:200. 1 μl of sample is taken from the maternal fetal RNA gene sample library tube and added to a solution containing 199 μl of... In the QubiU working solution detection tube, gently vortex to mix, and then place the detection tube into the QubiU fluorometer. The instrument utilizes the principle that the fluorescence intensity is proportional to the DNA content after the fluorescent dye specifically binds to double-stranded DNA (in this case, cDNA and other double-stranded nucleic acids in the maternal-fetal RNA gene sample group). It accurately measures the nucleic acid concentration in the sample. For example, if the concentration of the maternal-fetal RNA gene sample group in a certain library tube is measured to be 50 ng / μl, and samples with concentrations within this range are selected as representative samples for maternal-fetal RNA gene quantification based on a pre-set concentration range (assuming it is 30-80 ng / μl), the corresponding representative samples for maternal-fetal RNA gene quantification are finally generated.
[0113] Preferably, by setting the corresponding sequencing read length to 140bp, and performing single-end sequencing on a representative sample of maternal-fetal RNA gene quantification based on the sequencing read length, the maternal-fetal RNA gene expression sequencing coding region is generated.
[0114] In this embodiment of the invention, single-end sequencing of representative maternal fetal RNA gene quantification samples was performed using the Illumina HiSeq 2500 sequencing platform in a high-throughput sequencing laboratory. The corresponding cryopreservation tubes were removed from a -20°C freezer, and the samples were diluted to a suitable concentration for sequencing, such as 5 pM. The sequencing read length was set to 140 bp in the sequencing platform's software. This means the sequencer will read 140 bases of sequence information sequentially from one end of the RNA fragment. A dedicated sample loading device, such as Illumina's HiSeq X Five ClusUer, was used. The SUaUion sequencing platform loads diluted maternal-fetal RNA gene quantitative representative samples into designated channels of a sequencing chip. After the sequencing platform is started, sequencing-by-synthesis (SBS) technology is used. Under the action of RNA polymerase, fluorescently labeled dNUPs are added one by one to the newly synthesized RNA strand according to the base complementary pairing principle. Each time a dNUP is added, the sequencer determines the type of base at that position by detecting the fluorescence signal and records it. After a series of reactions and detection processes, each RNA fragment in the maternal-fetal RNA gene quantitative representative sample is sequenced, and finally the maternal-fetal RNA gene expression sequencing coding region is generated.
[0115] Preferably, the base quality assessment calculation is performed on the coding region of maternal-fetal RNA gene expression sequencing to obtain the corresponding RNA sample base quality score.
[0116] In this embodiment of the invention, the base quality of maternal fetal RNA gene expression sequencing coding region data is evaluated using FasUQC software on a data analysis workstation. The FasUQC software is opened, and the file "sample number_maternal fetal RNA gene expression sequencing coding region.fasUq" is imported into the software. FasUQC software evaluates base quality by analyzing the signal intensity and quality value of each base position in the sequencing data. For each base, its quality value is calculated using the formula Q = -10 × log0. 10 (P) is calculated, where P is the probability of misidentifying the base. For example, if the probability of misidentifying a base is P = 0.001, then its quality value Q = -10 × log 10 (0.001) = 30, which means reaching the Q30 quality level. The software performs this calculation on all bases in the coding region data of maternal-fetal RNA gene expression sequencing and counts the proportion of bases at different quality levels. For example, after analysis, the proportion of bases that reach Q30 (base recognition error rate of 0.1%) in this sample is 85%, and finally the corresponding RNA sample base quality score is obtained.
[0117] Furthermore, the calculation of base quality assessment for the coding region of maternal-fetal RNA gene expression sequencing includes:
[0118] Statistical analysis of base distribution in the coding region of maternal-fetal RNA gene expression sequencing was performed to statistically analyze the number and proportion of A, U, C and G bases within the coding region window, thereby obtaining the base distribution of the maternal-fetal RNA gene coding region.
[0119] In this embodiment of the invention, base distribution statistical analysis of maternal fetal RNA gene expression sequencing coding region data is performed on a data analysis workstation using the PyUhon programming language in conjunction with the BiopyUhon library. The corresponding file is read from the sequencing platform's storage device and converted into a BiopyUhon-processable sequence object. The coding region window size is set to 100 bp, and the window is slid by 10 bp increments starting from the sequence start position. For the sequence within each window, BiopyUhon's built-in functions are used to statistically analyze A, U, C, and G. The number of bases, for example, within a certain window, is counted to have 25 A bases, 20 U bases, 30 C bases, and 25 G bases. The proportion of each base is calculated as follows: the proportion of A bases is 25 ÷ 100 = 25%, the proportion of U bases is 20 ÷ 100 = 20%, the proportion of C bases is 30 ÷ 100 = 30%, and the proportion of G bases is 25 ÷ 100 = 25%. The base count and proportion data within these windows are compiled into a table and stored in an Excel file named after the sample number, thus obtaining the base distribution of the maternal-fetal RNA gene coding region.
[0120] Preferably, base mismatch assessment is performed on each base in the coding region of the maternal RNA gene expression sequencing based on the base distribution of the maternal RNA gene, so as to obtain the probability of base mismatch in the maternal RNA gene.
[0121] In this embodiment of the invention, based on the obtained base distribution data of the coding region of the maternal-fetal RNA gene, a self-compiled Python script is used to evaluate the base position mismatch between each base in the coding region of the maternal-fetal RNA gene expression sequencing. Data is read from the corresponding file, and the actual base distribution within the window is compared with the known normal base distribution rules (such as A and U, C and G pairing in the human genome, and the overall ratio is relatively stable). For example, if the number of A and C pairings is large in a certain window, exceeding the normal range, the base position mismatch probability of that window is obtained by calculating the ratio of the actual number of mismatched base pairs to the total number of bases in the window. Assuming that in a 100bp window, the number of mismatched base pairs should normally be less than 5 pairs, but 10 pairs are actually detected, then the base position mismatch probability of that window is 10 ÷ 100 = 10%. This calculation is performed for all windows, and the base position mismatch probabilities of each window are summarized and stored in a text file named after the sample number, finally obtaining the base position mismatch probability of the maternal-fetal RNA gene.
[0122] Preferably, the hydrogen bond strength and spatial structure free energy between each base are obtained by sequencing the coding region of the maternal RNA gene expression, and the base signal intensity is evaluated based on the hydrogen bond strength and spatial structure free energy between each base to obtain the maternal RNA gene base signal intensity.
[0123] In this embodiment of the invention, using a data analysis workstation and professional molecular biology simulation software such as NUPACK, the hydrogen bond strength and spatial structure free energy between each base are obtained from the coding region of maternal-fetal RNA gene expression sequencing. Based on this, the base signal intensity is evaluated. The corresponding file is converted into a NUPACK-recognizable sequence format and imported into the software. NUPACK software, based on the physicochemical properties of nucleic acid molecules, calculates the hydrogen bond strength between each base by simulating the interactions between bases. For example, A and U form two hydrogen bonds, and their hydrogen bond strength is relatively weak. Assuming the simulation calculation shows that the hydrogen bond strength of the A / U pair is 5 kcal... l / mol; three hydrogen bonds are formed between C and G, and the hydrogen bond strength is relatively strong. The hydrogen bond strength of the CG pair is 7 kcal / mol. At the same time, the software calculates the spatial structure free energy of the entire coding region sequence. This value reflects the stability of the sequence in maintaining a specific spatial structure. Based on the hydrogen bond strength and spatial structure free energy, an evaluation model is established. For example, the hydrogen bond strength and spatial structure free energy are weighted and summed (assuming the weight of hydrogen bond strength is 0.6 and the weight of spatial structure free energy is 0.4) to obtain the signal intensity value of each base. These maternal-fetal RNA gene base signal intensity values are recorded in a text file named after the sample number to finally obtain the maternal-fetal RNA gene base signal intensity.
[0124] Preferably, the base quality of the corresponding bases in the coding region of the maternal RNA gene expression sequencing is calculated based on the probability of base mismatch at the maternal fetal RNA gene base position and the base signal intensity of the maternal fetal RNA gene, so as to obtain the base quality score of the corresponding RNA sample.
[0125] In this embodiment of the invention, a self-developed PyUhon program is used to perform base quality assessment calculations on the corresponding bases in the coding region of the maternal RNA gene expression sequencing based on the base mismatch probability and base signal intensity of the maternal RNA gene. The mismatch probability threshold is set to 5%, and the signal intensity threshold is set to 6 kcal / mol (these thresholds are determined based on a large amount of experimental data and experience). For each base in the coding region, if the base mismatch probability in its window is less than 5% and the base signal intensity is greater than 6 kcal / mol, the base is considered to be of good quality and is recorded as a high-quality base; otherwise, it is recorded as a low-quality base. The number of high-quality bases is counted, and their proportion of the total number of bases is calculated to obtain the corresponding RNA sample base quality score. For example, if a sample has a total of 10,000 bases in its coding region, and 8,000 of them are high-quality bases, then the base quality fraction of the RNA sample is 8,000 ÷ 10,000 = 80%. This RNA sample base quality fraction result is recorded in a new text file named after the sample number, providing more comprehensive and accurate base quality data for subsequent intelligent transcriptome analysis based on maternal-fetal interface biomarkers.
[0126] Furthermore, the transcriptome differential expression determination module includes the following functions:
[0127] Reference gene mapping screening of maternal-fetal RNA genome samples was performed based on the base mass fraction of RNA samples to generate a maternal-fetal RNA reference genome.
[0128] In this embodiment of the invention, the bioinformatics analysis software BWA (Burrows-Wheeler Aligner) is used on a data analysis workstation to perform reference gene mapping screening on maternal-fetal RNA gene sample groups based on the base quality fraction of RNA samples. This is done to obtain the base quality fraction of RNA samples from previous studies. A preset base quality threshold of 70% is used. The base quality fraction of each RNA sample within each maternal-fetal RNA gene sample group is compared with this threshold. Using BWA software and a human reference genome (such as the GRCh38 version) as a reference, RNA samples with a base quality fraction greater than or equal to 70% are mapped to the corresponding positions in the reference genome using the software's sequence alignment algorithm. For example, if a maternal-fetal RNA gene sample group contains 100 RNA samples, and 60 samples have a base quality fraction of 70% or higher, BWA software aligns and maps these 60 samples to the reference genome. The successfully mapped samples are then compiled into a new dataset, ultimately generating a maternal-fetal RNA reference genome.
[0129] Preferably, sequence alignment calculations are performed between each RNA transcription sample in the maternal-fetal RNA transcription sample group based on each reference gene in the maternal-fetal RNA reference genome to obtain the sequence alignment similarity E value between each RNA transcription sample and the corresponding reference gene.
[0130] In this embodiment of the invention, the sequence alignment software Bowtie2 is used to perform sequence alignment calculations between each reference gene in the maternal fetal RNA reference genome and each RNA transcription sample in the maternal fetal RNA transcription sample group to obtain the reference gene sequence and each RNA transcription sample. The Bowtie2 software uses techniques such as Burrows-Wheeler transformation to accurately align each RNA transcription sample with the reference gene in the reference genome. During the alignment process, the software calculates the sequence alignment similarity E value between each RNA transcription sample and the corresponding reference gene. The E value represents the expected number of times a similar or better alignment result is obtained under random conditions. For example, after an RNA transcription sample is aligned with a reference gene, the sequence alignment similarity E value calculated by the Bowtie2 software is 1e-5. The sequence alignment similarity E values of all RNA transcription samples and their corresponding reference genes are recorded in an Excel file named after the sample number.
[0131] Preferably, the maternal-fetal RNA transcriptome is screened for reference similarity samples based on the sequence similarity E value between each RNA transcriptome sample and the corresponding reference gene, so as to screen out the top 10 RNA transcriptome sequences with a sequence similarity E value less than or equal to 1e-3, and generate maternal-fetal RNA transcriptome sequences that are similar to the reference genome.
[0132] In this embodiment of the invention, a self-developed Python script is used to screen maternal-fetal RNA transcription sample groups based on the sequence alignment similarity E value between each RNA transcription sample and the corresponding reference gene. The script iterates through the sequence alignment similarity E values of all RNA transcription samples, sets the screening condition to a sequence alignment similarity E value less than or equal to 1e-3, and selects the top 10 RNA transcription sample sequences that meet this condition. For example, in 1000 RNA transcription sample sequences, 30 sequences have an E value less than or equal to 1e-3. The script sorts the sequences by E value from smallest to largest, selects the top 10, organizes these 10 RNA transcription sample sequences into a new file, and finally generates maternal-fetal RNA transcription sample sequences that are similar to the reference genome.
[0133] Preferably, GO annotation extraction is performed on maternal fetal RNA transcription sample sequences that are similar to the reference genome to extract GO entries related to the target gene, thereby obtaining maternal fetal RNA transcription GO annotation sample sequences;
[0134] In this embodiment of the invention, the DAVID (Database for Annotation, Visualization and Integrated Discovery) online tool is used on a data analysis workstation to extract GO annotations from maternal-fetal RNA transcription sample sequences that are similar to the reference genome. The previously selected corresponding sequence is copied into the input box of the DAVID tool, the species is selected as human, and the task is submitted. The DAVID tool analyzes the input RNA transcription sample sequence based on its built-in gene annotation database and extracts GO (Gene Ontology) entries related to the target gene. GO entries include annotation information in three aspects: biological process, molecular function, and cellular composition. For example, for a certain RNA transcription sample sequence, the DAVID tool analyzes and finds that it is related to GO entries such as "positive regulation of cell proliferation" (biological process) and "DNA binding" (molecular function). The extracted GO annotation information is organized into a table, and finally, the maternal-fetal RNA transcription GO annotation sample sequence is generated.
[0135] Preferably, differentially expressed genes are identified between two maternal RNA transcription samples corresponding to the GO annotation sample sequence of maternal RNA transcription. The fold change in gene expression level and the P-value between the two maternal RNA transcription samples are analyzed for significance. Transcription genes with a fold change in gene expression level greater than 4 and a P-value less than 0.05 between the maternal RNA transcription samples are identified as differentially expressed genes to generate a maternal RNA differential expression genome.
[0136] In this embodiment of the invention, the edgeR package in R language is used to identify differentially expressed genes between two corresponding maternal RNA transcription samples within the GO annotation sequence of maternal RNA transcription. The data is imported into the R environment, and an experimental design matrix is constructed using functions in the edgeR package to set grouping information. For example, the maternal RNA transcription GO annotation sequence is divided into two groups, each containing several RNA transcription samples. The precision test algorithm of the edgeR package is used to perform significance analysis on the two groups of samples, and the fold change in gene expression levels and the p-value between the two maternal RNA transcription samples are calculated. The screening criteria are set as a fold change in gene expression levels greater than 4 and a p-value less than 0.05. Transcription genes that meet the criteria are selected as differentially expressed genes. For example, after analysis, among 1000 genes, 50 genes have a fold change in expression levels greater than 4 and a p-value less than 0.05. These differentially expressed genes are then organized into a new dataset, and finally, a maternal RNA differentially expressed genome is generated.
[0137] Furthermore, the reference gene mapping screening specifically involves comparing the base quality fraction of RNA samples based on a preset base quality threshold. If the base quality fraction of an RNA sample is greater than or equal to the preset base quality threshold, then the corresponding RNA sample within its maternal-fetal RNA gene sample group is mapped to the corresponding reference genome; otherwise, no mapping is performed, thereby generating a maternal-fetal RNA reference genome.
[0138] Furthermore, the transcriptional differential expression network building module includes the following functions:
[0139] Cluster analysis of target genes in the differentially expressed maternal-fetal RNA genome was performed to generate a set of target genes for differentially expressed maternal-fetal RNA.
[0140] In this embodiment of the invention, the clusterProfiler package in R language is used to perform target gene clustering analysis on the differentially expressed maternal-fetal RNA genome on a data analysis workstation. The differentially expressed gene data is imported into the R environment, and the pamk function in the clusterProfiler package is used to cluster the genes using a partitioning clustering method. During the clustering process, the expression similarity between genes is calculated, and genes with similar expression patterns are grouped into one class. For example, the number of clusters is set to 5. The function divides 100 differentially expressed genes into 5 different categories according to the difference in gene expression levels. The clustering results are organized into a data frame format, with the genes in each category as a subset, and finally a set of differentially expressed maternal-fetal RNA target genes is generated.
[0141] Preferably, the main metabolic pathways and signal transduction pathways involved by differentially expressed genes in each gene category are obtained through maternal-fetal RNA differential expression target gene set analysis;
[0142] In this embodiment of the invention, the main metabolic pathways and signal transduction pathways involved by differentially expressed genes in each gene category are obtained through maternal-fetal RNA differential expression target gene set analysis using the Metascape online tool. The corresponding gene category data are extracted, and the gene list for each gene category is copied to the input box of the Metascape tool. The species is selected as human, and the analysis task is submitted. The Metascape tool performs functional enrichment analysis on the input genes based on its integrated multiple authoritative databases, such as KEGG (Kyoto Encyclopedia of Genes and Genomes) and Reactome. For example, for 20 genes in a certain gene category, Metascape analysis shows that these genes are mainly involved in the "PI3K-Akt signaling pathway" (signal transduction pathway) and the "glycolysis / gluconeogenesis metabolic pathway," and the information on the main metabolic pathways and signal transduction pathways corresponding to each gene category is organized into a tabular form.
[0143] Preferably, gene enrichment analysis is performed on the gene categories corresponding to the differentially expressed genes in the maternal-fetal RNA differential expression target gene set based on the main metabolic pathways and signal transduction pathways involved in each gene category, so as to obtain the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set.
[0144] In this embodiment of the invention, gene enrichment analysis is performed on the corresponding gene categories within the maternal-fetal RNA differential expression target gene set based on the main metabolic pathways and signal transduction pathways involved by differentially expressed genes in each gene category using the clusterProfiler package in the R language. The data is imported into the R environment, and for each gene category, the enrichKEGG function in the clusterProfiler package is used with the KEGG database as a reference to analyze the enrichment degree of the gene category in specific metabolic and signal transduction pathways. For example, for a gene category involved in the "PI3K-Akt signaling pathway", the function evaluates the significance of the gene category in the signaling pathway by calculating the enrichment score (such as enrichment fold, p-value, etc.). The gene pathway distribution results corresponding to each gene category are organized into visual charts (such as bar charts and bubble charts) to show the enrichment of each gene category in different pathways, and intuitively present the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set.
[0145] Preferably, the expression pattern associations between target genes are found through a preset database;
[0146] In this embodiment of the invention, by accessing a preset gene database, such as the STRING (Search Tool for the Retrieval of Interacting Genes / Proteins) database, the expression pattern associations between target genes are searched. A list of all target genes is extracted from the corresponding file. After converting the gene names into a format recognizable by the STRING database, the gene list is entered into the database search interface for retrieval. The STRING database integrates protein-protein interaction data and gene co-expression data from multiple data sources. For example, for gene A and gene B, the database shows that they exhibit a positively correlated expression pattern in multiple cell types, that is, when the expression level of gene A increases, the expression level of gene B also tends to increase. The expression pattern association information between target genes is organized into a network table to record information such as the association strength and association type between each gene and other genes.
[0147] Preferably, gene expression network analysis is performed on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern association between target genes, so as to generate a maternal-fetal target gene expression network.
[0148] In this embodiment of the invention, gene expression network analysis is performed on the gene pathway distribution corresponding to each gene category in the maternal-fetal target gene set based on the expression pattern association between target genes using Cytoscape software. This obtains gene pathway distribution data and reads the corresponding gene expression association information. This data is then imported into Cytoscape software, with genes as nodes, gene expression associations as edges, and gene pathways as node attributes. By setting different node colors, shapes, and edge thicknesses and colors, the interrelationships between genes and their distribution in different pathways are visually displayed. For example, gene nodes participating in the same gene pathway are set to the same color, and the edge thickness is adjusted according to the gene expression association strength. After layout adjustment and parameter optimization, the maternal-fetal target gene expression network is finally generated.
[0149] Furthermore, the target gene clustering analysis of the differentially expressed maternal-fetal RNA genome includes:
[0150] Quantitative statistical analysis was performed on each gene sample within the differentially expressed maternal-fetal RNA genome to obtain the sample expression level and gene expression level of each differentially expressed maternal-fetal RNA gene sample.
[0151] In this embodiment of the invention, quantitative statistical analysis of various gene samples within the differentially expressed maternal-fetal RNA genome is performed using the R language and its rich bioinformatics analysis packages. The gene sample data are imported into the R environment, and the DESeq2 package is used. This package can perform standardization and differential expression analysis on RNA sequencing data. First, a data object containing sample information and a gene expression count matrix is constructed. The sample information includes sample source, experimental conditions, etc., and the gene expression count matrix records the expression level of each gene in different samples. The DESeq2 package uses a built-in algorithm to further refine the original gene expression counts. Standardization is performed to eliminate technical differences between samples. For example, for a certain maternal-fetal RNA differentially expressed gene sample, its expression count in sample A is 100 and its expression count in sample B is 120 in the original data. After standardization by the DESeq2 package, the corrected sample expression level is obtained, which is 95 in sample A and 115 in sample B. At the same time, the package can also calculate the overall gene expression level of each gene in all samples. For example, the average expression level of gene X in all samples is 105. The sample expression levels and gene expression levels corresponding to each maternal-fetal RNA differentially expressed gene sample are organized into a data frame format.
[0152] Preferably, target gene clustering analysis is performed on each gene sample within the differentially expressed maternal RNA genome based on the sample expression level and gene expression level of each differentially expressed maternal RNA gene sample. This is done by calculating the Euclidean distance between each gene sample from both the sample expression level and gene expression level to perform hierarchical clustering and generate a set of differentially expressed maternal RNA target genes.
[0153] In this embodiment of the invention, by continuing to use the R language, based on the sample expression levels and gene expression levels corresponding to each differentially expressed gene sample of maternal-fetal RNA, the pvclust package is used to perform target gene clustering analysis among the gene samples within the differentially expressed genome of maternal-fetal RNA. The data is imported into the R environment. The pvclust package provides powerful hierarchical clustering capabilities. For each gene sample, at the sample expression level, the sum of squares of the differences in expression levels between it and other gene samples in different samples is calculated, and the square root is taken to obtain the Euclidean distance at the sample expression level; at the gene expression level, the sum of squares of the differences in the overall gene expression levels between this gene sample and other gene samples is calculated, and the square root is taken to obtain the Euclidean distance at the gene expression level. The Euclidean distances at both levels are combined as an indicator to measure the similarity between gene samples. For example, gene sample A and gene sample B have an Euclidean distance of 5 at the sample expression level and an Euclidean distance of 3 at the gene expression level. Using specific weights (assuming a sample expression level weight of 0.6 and a gene expression level weight of 0.4), the combined distance is calculated as 5 × 0.6 + 3 × 0.4 = 4.2. Based on these combined distances, the pvclust package uses a hierarchical clustering algorithm to progressively merge gene samples with similar distances, ultimately generating a dendritic clustering graph. According to the structure of the dendritic graph and pre-defined clustering criteria (such as a cluster height threshold), the gene samples are divided into different categories, ultimately generating a set of differentially expressed target genes for maternal-fetal RNA. These include upregulated genes such as CXCL8, CXCL2, CXCL3, BMP7, CXCR2, TNF, IL6, GH1, and GH2, and downregulated genes such as CXCL14, CXCL13, CCR4, CCR9, EGF, and VEGFD.
[0154] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0155] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. An intelligent transcriptome analysis system based on biomarkers of the maternal-fetal interface, characterized by, The method comprises the following modules: A maternal-fetal RNA low-temperature storage module is used to obtain a maternal-fetal interface biological blood sample set, wherein the sample set comprises not less than 500 blood samples in each of an early pregnancy stage, a middle pregnancy stage and a late pregnancy stage, and the maternal-fetal interface biological blood sample set is divided into three sample groups for RNA extraction and low-temperature storage to generate corresponding maternal-fetal blood RNA low-temperature sample groups; A maternal-fetal genome sequencing calculation module is used to construct a maternal-fetal blood RNA sample library from the corresponding maternal-fetal blood RNA low-temperature sample groups, wherein the sample library comprises a maternal-fetal RNA gene sample group and a maternal-fetal RNA transcription sample group, and single-end sequencing calculation is performed on the maternal-fetal RNA gene sample group in the maternal-fetal blood RNA sample library to obtain corresponding RNA sample base quality scores; wherein the maternal-fetal genome sequencing calculation module comprises the following functions: The maternal-fetal RNA gene sample group and a maternal-fetal RNA transcription sample group are divided from the corresponding maternal-fetal blood RNA low-temperature sample groups; RNA reverse transcription is performed on the maternal-fetal RNA transcription sample group by using an RNA amplification kit to generate the maternal-fetal RNA transcription sample group; The cDNA fragments of the maternal-fetal RNA transcription sample group are randomly broken, end-repaired, A-tailed and ligated with adaptors, and the processed maternal-fetal RNA transcription sample group and the maternal-fetal RNA gene sample group are purified and library constructed to generate the maternal-fetal blood RNA sample library; Single-end sequencing calculation is performed on the maternal-fetal RNA gene sample group in the maternal-fetal blood RNA sample library to obtain corresponding RNA sample base quality scores; wherein the calculation comprises: QubiU quantification screening is performed on the maternal-fetal RNA gene sample group in the maternal-fetal blood RNA sample library by using a QubiU-dsDNA-HS kit to generate corresponding maternal-fetal RNA gene quantification representative samples; Single-end sequencing is performed on the maternal-fetal RNA gene quantification representative samples based on a sequencing read length of 140 bp to generate a maternal-fetal RNA gene expression sequencing coding region; Base quality evaluation calculation is performed on the maternal-fetal RNA gene expression sequencing coding region to obtain corresponding RNA sample base quality scores; wherein the calculation comprises: Base distribution statistical analysis is performed on the maternal-fetal RNA gene expression sequencing coding region to statistically analyze the number and proportion of A, U, C and G bases in the coding region window to obtain maternal-fetal RNA gene coding region base distribution; Base position mismatch evaluation is performed on each base in the maternal-fetal RNA gene expression sequencing coding region based on the maternal-fetal RNA gene coding region base distribution to obtain maternal-fetal RNA gene base position mismatch probability; The hydrogen bond binding strength and spatial structure free energy between each base are obtained from the maternal-fetal RNA gene expression sequencing coding region, and base signal strength evaluation is performed according to the hydrogen bond binding strength and spatial structure free energy between each base to obtain maternal-fetal RNA gene base signal strength; According to the base position mismatch probability of the maternal-fetal RNA gene and the base signal intensity, base quality evaluation and calculation are performed on the corresponding bases in the coding region of the maternal-fetal RNA gene expression sequencing to obtain the corresponding RNA sample base quality score; The transcriptome differential expression determination module is configured to perform reference gene mapping screening on the maternal-fetal RNA gene sample group based on the RNA sample base quality score to generate a maternal-fetal RNA reference genome, and perform differential expression gene determination on the maternal-fetal RNA transcript sample group based on the maternal-fetal RNA reference genome to generate a maternal-fetal RNA differential expression genome. The transcriptome differential expression determination module includes the following functions: Perform reference gene mapping screening on the maternal-fetal RNA gene sample group based on the RNA sample base quality score to generate a maternal-fetal RNA reference genome. Perform sequence alignment calculation between each RNA transcript sample in the maternal-fetal RNA transcript sample group and the corresponding reference gene in the maternal-fetal RNA reference genome to obtain a sequence alignment similarity E value between each RNA transcript sample and the corresponding reference gene. Perform reference similar sample screening on the maternal-fetal RNA transcript sample group based on the sequence alignment similarity E value between each RNA transcript sample and the corresponding reference gene to screen out the top 10 RNA transcript sample sequences with a corresponding sequence alignment similarity E value less than or equal to 1e-3, and generate maternal-fetal RNA transcript sample sequences similar to the reference genome. Perform GO annotation extraction on the maternal-fetal RNA transcript sample sequences similar to the reference genome to extract corresponding GO entries related to the target gene and obtain maternal-fetal RNA transcript GO annotation sample sequences. Perform differential expression gene determination between the corresponding two maternal-fetal RNA transcript samples in the maternal-fetal RNA transcript GO annotation sample sequences to significantly analyze the gene expression quantity difference fold and P value between the two maternal-fetal RNA transcript samples, and identify the transcription gene corresponding to the gene expression quantity difference fold greater than 4 and the P value less than 0.05 between the two maternal-fetal RNA transcript samples as a differential expression gene to generate a maternal-fetal RNA differential expression genome. The transcript differential expression network construction module is configured to perform target gene clustering and enrichment analysis on the maternal-fetal RNA differential expression genome to obtain the gene pathway distribution corresponding to each gene classification in the maternal-fetal target gene set, find the expression pattern association between the target genes through a preset database, and perform gene expression network analysis on the gene pathway distribution corresponding to each gene classification in the maternal-fetal target gene set based on the expression pattern association between the target genes to generate a maternal-fetal target gene expression network.
2. The smart transcriptome analysis system based on the biomarkers of the maternal-fetal interface according to claim 1, characterized in that, The maternal-fetal RNA low-temperature storage module includes the following functions: Obtain a maternal-fetal interface biological blood sample set, which includes no less than 500 blood samples for each stage of early pregnancy, mid-pregnancy, and late pregnancy. The maternal-fetal interface biological blood sample set is divided into corresponding three sample groups according to experimental requirements, one of which is quickly frozen in liquid nitrogen, one is placed in 5ml 10% formaldehyde fixing solution and paraffin stored for standby, and one is centrifuged and sealed in a 1.5ml RNAlaUer preservation solution to generate three corresponding maternal-fetal biological blood sample groups; The corresponding maternal-fetal biological blood sample group is put into the corresponding kit for RNA extraction to generate a maternal-fetal blood RNA sample group; The RNA concentration of the maternal-fetal blood RNA sample group is measured by specific binding between fluorescent dyes and RNA molecules, and the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group is calculated based on the RNA concentration to obtain the integrity RIN number of the corresponding RNA sample of the maternal-fetal blood RNA sample; Based on the integrity RIN number of the maternal-fetal blood RNA sample, the integrity quality control of the corresponding RNA sample in the maternal-fetal blood RNA sample group is carried out to screen out the RNA sample with integrity RIN number greater than 7 and integrity quality qualified, and the screened integrity quality qualified RNA sample is subpackaged into an RNase-free cryotube and moved to a-80℃ ultra-low temperature refrigerator for storage, to generate a corresponding maternal-fetal blood RNA low temperature sample group.
3. The smart transcriptome analysis system based on the biomarkers of the maternal-fetal interface according to claim 2, characterized in that, The integrity index calculation of the corresponding RNA sample in the maternal-fetal blood RNA sample group based on the RNA concentration includes: RNA electrophoresis analysis of the corresponding RNA sample in the maternal-fetal blood RNA sample group based on the RNA concentration is carried out to generate a corresponding maternal-fetal blood RNA electropherogram; The RNA electrophoresis main peak value and the RNA electrophoresis secondary peak value of each RNA sample are obtained from the maternal-fetal blood RNA electropherogram; According to the RNA electrophoresis main peak value and the RNA electrophoresis secondary peak value of the RNA sample, the integrity index of the corresponding RNA sample in the maternal-fetal blood RNA sample group is calculated to calculate the height ratio between the RNA electrophoresis main peak value and the RNA electrophoresis secondary peak value to obtain the corresponding RIN number, to obtain the integrity RIN number of the corresponding RNA sample of the maternal-fetal blood RNA sample.
4. The smart transcriptome analysis system based on the biomarkers of the maternal-fetal interface according to claim 1, characterized in that, The reference gene mapping screening specifically compares and judges the base quality score of the RNA sample according to the preset base quality threshold, if the base quality score of the RNA sample is greater than or equal to the preset base quality threshold, then the corresponding RNA sample in the maternal-fetal RNA gene sample group is mapped into the corresponding reference genome, otherwise it is not mapped, to screen and generate a maternal-fetal RNA reference genome.
5. The smart transcriptome analysis system based on the biomarkers of the maternal-fetal interface according to claim 1, wherein, The transcriptional differential expression network construction module includes the following functions: Target gene clustering analysis is performed on the maternal-fetal RNA differential expression genome to generate a maternal-fetal RNA differential expression target gene set; The main metabolic pathways and signal transduction pathways in which the differential expression genes participate are obtained through analysis of the maternal-fetal RNA differential expression target gene set; The gene enrichment analysis is performed on the gene classification corresponding to the main metabolic pathways and signal transduction pathways in which the differentially expressed genes participate, based on the gene classification corresponding to the target gene set of the maternal-fetal RNA differential expression, to obtain the gene pathway distribution corresponding to each gene classification in the target gene set of the maternal-fetal RNA differential expression; The expression pattern correlation between the target genes is found by searching a preset database; The gene expression network analysis is performed on the gene pathway distribution corresponding to each gene classification in the target gene set of the maternal-fetal RNA differential expression, based on the expression pattern correlation between the target genes, to generate the maternal-fetal target gene expression network.
6. The smart transcriptome analysis system based on the biomarkers of the maternal-fetal interface according to claim 5, characterized in that, The target gene clustering analysis on the maternal-fetal RNA differential expression gene set includes: The quantitative statistical analysis is performed on each gene sample in the maternal-fetal RNA differential expression gene set, to obtain the sample expression quantity and the gene expression quantity corresponding to each maternal-fetal RNA differential expression gene sample; The target gene clustering analysis is performed between each gene sample in the maternal-fetal RNA differential expression gene set, based on the sample expression quantity and the gene expression quantity corresponding to each maternal-fetal RNA differential expression gene sample, to hierarchically cluster by calculating the Euclidean distance between each gene sample from two aspects of the sample expression quantity and the gene expression quantity, to generate the target gene set of the maternal-fetal RNA differential expression.
Citation Information
Patent Citations
Application of key gene in in-vitro regulation of oocyte GVBD and interference fragment
CN116640734A
Application of TNFSF4 (Tumor Necrosis Factor Factor Factor 4) as premature delivery noninvasive detection marker, screening method and targeted detection method
CN116769897A