Database construction method, device, equipment and medium based on Nephila clavata spiders
By constructing a database of Bangluo bride spider, the problem of insufficient spider-omics data is solved, and a comprehensive analysis of rapid acquisition of genomics and other omics information is achieved, and a variety of data visualization tools are provided.
Patent Information
- Application Number
- CN202310897382.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-07-20
AI Technical Summary
The existing technology lacks methods to quickly obtain spider genome and other omics data, resulting in insufficient omics-related information for spiders.
A database based on the Bangluo bride spider was constructed. By assembling high-quality chromosomal genomes, annotating transcriptomes and protein sequences, determining gene expression values and co-expression relationships, cell cluster classification and spatial structure analysis, and a comprehensive database was constructed in combination with the database management system.
It has achieved rapid acquisition of comprehensive information on spider genome and other omics, provided a variety of data visualization tools, and supported multi-level biological data exploration and analysis.
Smart Images

Figure CN116894027B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of databases, and particularly to a method, device, equipment and medium for constructing a database based on Nephila clavata spiders. Background Art
[0002] Spiders are abundant arthropod predators, including more than 50,000 extant species. All spiders produce silk, a natural high-performance protein fiber that is crucial for the survival and reproduction of spiders. Spider silk fibers have special properties, including high tensile strength and toughness, low density, and biocompatibility. Many researchers have tried to mimic the production and spinning process of natural spider silk proteins to fabricate artificial materials with spider silk properties. However, current research on spider silk formation is mostly based on physics and materials research, which only provides partial spider silk properties.
[0003] In addition, current data information about spiders mainly focuses on the identification of spider species, lacking omics-related data of spiders, further lacking spider transcriptome data and other omics data.
[0004] In summary, how to conveniently and quickly obtain comprehensive information on the genomes and other omics of spiders is an urgent problem to be solved currently. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for constructing a database based on Nephila clavata spiders, which can conveniently and quickly obtain comprehensive information on the genomes and other omics of spiders, and the specific solutions are as follows:
[0006] In the first aspect, the present application discloses a method for constructing a database based on Nephila clavata spiders, including:
[0007] Assembling the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition;
[0008] Annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of each species closely related to the Nephila clavata spiders to obtain an initial annotation file, and deleting the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file;
[0009] Filtering the sequenced general transcriptomes to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences;
[0010] Classifying cell clusters of the sequenced single-cell transcriptomes to obtain classified cell clusters, and performing functional annotation on the classified cell clusters to obtain an annotated transcriptome;
[0011] Preprocess the sequenced spatial transcriptome, and cluster the sequenced spatial transcriptome based on the preprocessing results to determine the cell distribution within the spatial structure;
[0012] According to the database management system, construct a database based on Nephila clavata spiders based on the target annotation file, the expression values, the co-expression relationships, the annotated transcriptome, and the cell distribution within the spatial structure.
[0013] Optionally, assembling the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition includes:
[0014] Filter and delete all low-quality long read sequences in the sequenced gene sequences corresponding to Nephila clavata spiders that do not meet the third preset high-quality condition;
[0015] Perform de novo assembly on the filtered high-quality long read sequences to obtain primary contigs;
[0016] Correct the primary contigs to obtain corrected contigs;
[0017] Use the short read sequences in the sequenced gene sequences corresponding to Nephila clavata spiders to polish the corrected contigs to obtain polished contigs;
[0018] Use the paired-end Hi-C short reads in the sequenced gene sequences to trim the polished contigs to obtain trimmed contigs, and use the trimmed contigs to construct pseudochromosomes, and then determine the gene position information based on the pseudochromosomes;
[0019] Map the gene position information to the polished contigs to obtain the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition.
[0020] Optionally, annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of Nephila clavata spiders and the protein sequences of various species closely related to Nephila clavata spiders to obtain an initial annotation file includes:
[0021] Perform de novo annotation on the chromosomal genome based on the general transcriptomes corresponding to several tissues of Nephila clavata spiders, and at the same time perform homology annotation on the chromosomal genome based on the protein sequences of various species closely related to Nephila clavata spiders to obtain an initial annotation file.
[0022] Optionally, the screening criteria include that the proportion of repetitive sequences corresponding to a single gene sequence is less than a first target proportion, the length of the protein sequence corresponding to a single gene sequence is greater than a target length, the FPKM value of a single gene sequence in more than a target number of the tissues is greater than a target value, the single gene sequence is from the various species, and the identity between the full-length transcriptome sequencing data and the single gene sequence is higher than a second target proportion.
[0023] Optionally, determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences includes:
[0024] Aligning the high-quality transcriptome sequencing sequences with the chromosomal genome to obtain an alignment result representing the corresponding relationship between each position in the chromosomal genome and each gene sequence in the high-quality transcriptome sequencing sequences;
[0025] Sorting the alignment result according to the positional relationship of the chromosomal genome and storing it as a target result in a target format;
[0026] Calculating the expression value and co-expression relationship of each gene sequence in the high-quality transcriptome sequencing sequences based on the Pearson correlation coefficient and the target result.
[0027] Optionally, classifying the sequenced single-cell transcriptome to obtain classified cell clusters includes:
[0028] Removing low-quality cells in the sequenced single-cell transcriptome that do not meet the preset cell conditions, and obtaining a first gene expression matrix for each cell with the column coordinates being cells and the row coordinates being genes corresponding to the other cells except the low-quality cells;
[0029] Clustering the other cells based on the first gene expression matrix to obtain classified cell clusters.
[0030] Optionally, preprocessing the sequenced spatial transcriptome and clustering the sequenced spatial transcriptome based on the preprocessing result to determine the cell distribution within the spatial structure includes:
[0031] Removing low-quality reads in the sequenced spatial transcriptome, and obtaining a second gene expression matrix for each cell with the column coordinates being cells and the row coordinates being genes corresponding to the other reads except the low-quality reads;
[0032] Clustering the cells corresponding to the other reads based on the second gene expression matrix to determine the cell distribution within the spatial structure.
[0033] In a second aspect, the present application discloses a database construction device based on Nephila clavata spiders, including:
[0034] An assembly module for assembling the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality conditions;
[0035] A first annotation module for annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of various species closely related to the Nephila clavata spiders to obtain an initial annotation file;
[0036] A deletion module for deleting the annotations corresponding to the gene sequences that do not meet the screening criteria in the initial annotation file to obtain a target annotation file;
[0037] A filtering module for filtering the sequenced general transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality conditions, and determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences;
[0038] A classification module for classifying cell clusters of the sequenced single-cell transcriptome to obtain classified cell clusters;
[0039] A second annotation module for functionally annotating the classified cell clusters to obtain an annotated transcriptome;
[0040] A clustering module for preprocessing the sequenced spatial transcriptome and clustering the sequenced spatial transcriptome based on the preprocessing results to determine the cell distribution within the spatial structure;
[0041] A database construction module for constructing a database based on Nephila clavata spiders according to a database management system and based on the target annotation file, the expression value and the co-expression relationship, the annotated transcriptome, and the cell distribution within the spatial structure.
[0042] In a third aspect, the present application discloses an electronic device, including:
[0043] A memory for storing a computer program;
[0044] A processor for executing the computer program to implement the aforementioned method for constructing a database based on Nephila clavata spiders.
[0045] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned method for constructing a database based on Nephila clavata spiders is implemented.
[0046] It can be seen that the present application assembles the chromosomal genome of Nephila clavata that meets the first preset high-quality condition; based on the general transcriptomes corresponding to several tissues of the Nephila clavata and the protein sequences of each species closely related to the Nephila clavata, the chromosomal genome is annotated to obtain an initial annotation file, and the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file are deleted to obtain a target annotation file; the sequenced general transcriptome is filtered to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and the expression value and co-expression relationship of each gene sequence are determined based on the high-quality transcriptome sequencing sequences; the sequenced single-cell transcriptome is classified into cell clusters after classification, and the classified cell clusters are functionally annotated to obtain an annotated transcriptome; the sequenced spatial transcriptome is preprocessed, and based on the preprocessing results, the sequenced spatial transcriptome is clustered to determine the cell distribution within the spatial structure; according to the database management system, and based on the target annotation file, the expression value and the co-expression relationship, the annotated transcriptome, and the cell distribution within the spatial structure, a database based on Nephila clavata is constructed. Thus, the present application constructs a comprehensive spider database to facilitate the rapid acquisition of comprehensive information on the genome and other omics of spiders. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0048] Figure 1 It is a flowchart of a method for constructing a database based on Nephila clavata disclosed in the present application;
[0049] Figure 2 It is a flowchart of a specific method for constructing a database based on Nephila clavata disclosed in the present application;
[0050] Figure 3 It is a schematic structural diagram of a device for constructing a database based on Nephila clavata disclosed in the present application;
[0051] Figure 4 It is a structural diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] Spiders are abundant arthropod predators, including more than 50,000 extant species. All spiders produce silk, a natural high-performance protein fiber that is crucial for the survival and reproduction of spiders. Silk fibers have special properties, including high tensile strength and toughness, low density, and biocompatibility. Many researchers have attempted to mimic the production and spinning process of natural silk proteins to fabricate artificial materials with silk properties. However, current research on silk formation is mostly based on physics and materials research, which only provides partial silk properties.
[0054] In addition, current data information about spiders mainly focuses on the identification of spider species, lacking omics-related data of spiders, and further lacking spider transcriptome data and other omics data.
[0055] For this reason, the embodiments of the present application propose a database construction scheme based on Nephila clavata spiders, which can conveniently and quickly obtain comprehensive information on the genomes and other omics of spiders.
[0056] The embodiments of the present application disclose a method for constructing a database based on Nephila clavata spiders. Refer to Figure 1 as shown, the method includes:
[0057] Step S11: Assemble the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition.
[0058] In this embodiment, the chromosomal genome that meets the first preset high-quality condition also refers to the chromosomal genome at the high-quality chromosomal level.
[0059] Step S12: Based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of each species closely related to the Nephila clavata spiders, annotate the chromosomal genome to obtain an initial annotation file, and delete the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file.
[0060] In this embodiment, the initial annotation file is a file including the chromosomal genome and each annotation obtained after annotating the chromosomal genome.
[0061] Step S13: Filter the sequenced ordinary transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality conditions, and determine the expression values and co-expression relationships of each gene sequence based on the high-quality transcriptome sequencing sequences.
[0062] In this embodiment, determining the expression values and co-expression relationships of each gene sequence based on the high-quality transcriptome sequencing sequences includes: aligning the high-quality transcriptome sequencing sequences with the chromosomal genome to obtain an alignment result representing the correspondence between each position in the chromosomal genome and each gene sequence in the high-quality transcriptome sequencing sequences; sorting the alignment result according to the positional relationship of the chromosomal genome and storing it as a target result in a target format; calculating the expression values and co-expression relationships of each gene sequence in the high-quality transcriptome sequencing sequences based on the Pearson correlation coefficient and the target result.
[0063] It should be noted that the fastp V0.19.5 software is used to remove low-quality gene sequences in the ordinary transcriptome to filter the sequenced ordinary transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality conditions. The HISAT2 V2.1.0 software based on default parameters is used to align the high-quality transcriptome sequencing sequences with the chromosomal genome to obtain an alignment result representing the correspondence between each position in the chromosomal genome and each gene sequence in the high-quality transcriptome sequencing sequences. The Samtools V1.9 software is used to sort the alignment result according to the positional relationship of the chromosomal genome and store it as a target result in a target format. Finally, the ballroot V2.16.0 is used to calculate the expression values and co-expression relationships of each gene sequence in the high-quality transcriptome sequencing sequences based on the Pearson correlation coefficient and the target result. It should be noted that the second preset high-quality conditions are determined according to the set parameters of the fastp software.
[0064] It should be noted that fastp is a fast, flexible, and versatile data preprocessing tool for quality control and trimming of sequencing data; HISAT2 for transcriptome analysis is a fast and sensitive sequence alignment software; the target format is the bam format; the Samtools software is a tool software for operating sam and bam files, which can perform binary viewing, format conversion, sorting, and merging of alignment files, and complete the statistical summary of alignment results by combining information such as flag and tag in the sam format; ballroot is software for calculating expression values and co-expression relationships.
[0065] Step S14: Classify the sequenced single-cell transcriptome to obtain classified cell clusters, and perform functional annotation on the classified cell clusters to obtain an annotated transcriptome.
[0066] In this embodiment, the cell cluster classification of the single-cell transcriptome after sequencing to obtain the classified cell clusters includes: removing the low-quality cells in the single-cell transcriptome after sequencing that do not meet the preset cell conditions, and obtaining the first gene expression matrix of each cell with the column coordinates as cells and the row coordinates as genes for the other cells outside the low-quality cells; clustering the other cells based on the first gene expression matrix to obtain the classified cell clusters.
[0067] It should be noted that the cellranger v4.0.0 software is used to remove the low-quality cells in the single-cell transcriptome after sequencing that do not meet the preset cell conditions, and obtain the first gene expression matrix of each cell with the column coordinates as cells and the row coordinates as genes for the other cells outside the low-quality cells. In this process, the preset cell conditions include that the cell number and gene UMI count are within the median range ± retaining 2 times the median absolute deviation (MAD, Median Absolute Deviation). In addition, in this process, Doublet Finder v2.0.3 is also used to remove doublets (≥2 cells in one oil droplet), that is, the preset cell conditions include no doublet cells; the Seurat v4.0.6 software package is used to cluster the other cells based on the first gene expression matrix to obtain the classified cell clusters. It should be noted that the Monocle v2.22.0 software package can also be used to predict the developmental trajectories of all cell clusters corresponding to the classified cell clusters.
[0068] It should be noted that cellranger is a single-cell data analysis tool; the Doublet Finder software package is used to detect doublets in single-cell RNA sequencing data (single-cell transcriptome after sequencing) to remove doublets; Seurat is a toolbox for quality control, analysis, and exploration of single-cell RNA sequencing data; trajectory analysis is mainly through pseudotime analysis represented by the Monocle software; UMI (Unique Molecular Identifier): a tagged gene, which is a marker used to identify and track a single molecule in gene detection. UMI is a 12-nt nucleotide sequence, and each mRNA is randomly linked to a UMI, so different UMIs can be counted, and finally the number of mRNAs can be counted; the median absolute deviation (MAD) is a standard used to describe the variability of univariate samples in quantitative data; the number of other cells (high-quality cells) for clustering is 9349.
[0069] In this embodiment, functional annotation of the classified cell clusters is performed to obtain the annotated transcriptome, that is, KEGG and GO functional analyses are performed using various weighted maker genes after clustering.
[0070] Step S15: Preprocess the sequenced spatial transcriptome, and cluster the sequenced spatial transcriptome based on the preprocessing result to determine the cell distribution within the spatial structure.
[0071] In this embodiment, preprocessing the sequenced spatial transcriptome and clustering the sequenced spatial transcriptome based on the preprocessing result to determine the cell distribution within the spatial structure includes: removing low-quality reads corresponding to the sequenced spatial transcriptome, and obtaining a second gene expression matrix for each cell with column coordinates as cells and row coordinates as genes corresponding to the other reads except the low-quality reads; clustering the cells corresponding to the other reads based on the second gene expression matrix to determine the cell distribution within the spatial structure.
[0072] It should be noted that the spaceranger v1.3.1 mkref and count programs are used to remove low-quality reads corresponding to the sequenced spatial transcriptome, and obtain a second gene expression matrix for each cell with column coordinates as cells and row coordinates as genes corresponding to the other reads except the low-quality reads; Seurat v4.0.6 is used; the cells corresponding to the other reads are clustered based on the second gene expression matrix to determine the cell distribution within the spatial structure.
[0073] It should be noted that the spaceranger software and count programs can convert the original fastq data of the sequenced spatial transcriptome into an expression matrix after quality control, alignment, quantification and other steps; this clustering process is spot clustering.
[0074] Step S16: Construct a database based on Nephila clavata spiders according to the database management system, and based on the target annotation file, the expression values and the co-expression relationships, the annotated transcriptome and the cell distribution within the spatial structure.
[0075] In this embodiment, the database management system is MySQL; spiderDB is used to combine the MySQL database management system with a dynamic web interface, which is written in Python, HTML (Hyper Text Markup Language), CSS (Cascading Style Sheets), Javascript and jQuery; jQuery is a fast and concise JavaScript framework. The entire project is open access and can be used by anyone, and is configured on an Ubuntu (V18.04) Linux machine with an Apache2 server; spiderDB is a web-based tool.
[0076] In this embodiment, the web pages corresponding to the constructed database can be divided into: gene information, eFP, subcellular localization, gene co-expression, gene family, single-cell transcriptome after sequencing, JBrowser, chromosome localization, collinearity, orthologous genes, and spatial single-cell, etc.
[0077] In summary, the constructed database can provide comprehensive information on the genome of Nephila clavata and other omics, and combine multiple data visualization tools into the same interface, allowing users to explore biological data at multiple levels and comprehensively analyze genes.
[0078] It should be noted that the Nephila clavata spider can be replaced by other spiders.
[0079] It can be seen that this application assembles the chromosomal genome of Nephila clavata spiders that meet the first preset high-quality conditions; based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of various species closely related to the Nephila clavata spiders, the chromosomal genome is annotated to obtain an initial annotation file, and the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file are deleted to obtain a target annotation file; the general transcriptomes after sequencing are filtered to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality conditions, and the expression values and co-expression relationships of each gene sequence are determined based on the high-quality transcriptome sequencing sequences; the single-cell transcriptomes after sequencing are classified into cell clusters, and the classified cell clusters are functionally annotated to obtain an annotated transcriptome; the spatial transcriptomes after sequencing are preprocessed, and based on the preprocessing results, the spatial transcriptomes after sequencing are clustered to determine the cell distribution within the spatial structure; according to the database management system, and based on the target annotation file, the expression values and co-expression relationships, the annotated transcriptome, and the cell distribution within the spatial structure, a database based on Nephila clavata spiders is constructed. Thus, it can be seen that this application constructs a comprehensive spider database to facilitate the rapid acquisition of comprehensive information on the genomes and other omics of spiders.
[0080] This embodiment of the application discloses a specific method for constructing a database based on Nephila clavata spiders. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. See Figure 2 as shown, specifically including:
[0081] Step S21: Filter and delete all low-quality long read sequences in the sequenced gene sequences corresponding to the Nephila clavata spiders that do not meet the third preset high-quality condition; perform initial assembly on the filtered high-quality long read sequences to obtain primary contigs; correct the primary contigs to obtain corrected contigs; use the short read sequences in the sequenced gene sequences corresponding to the Nephila clavata spiders to polish the corrected contigs to obtain polished contigs.
[0082] In this embodiment, Canu v2.2 with default parameters is used to perform initial assembly on the filtered high-quality long read sequences to obtain primary contigs; SMARTdenovo is used to correct the primary contigs, and then based on this correction, Racon v1.5.0 is used for three rounds of correction to obtain corrected contigs; Pilon v1.24 is used to polish the corrected contigs based on the short read sequences to obtain polished contigs. It should be noted that the fastp V0.19.5 software can be used to filter and delete all low-quality long read sequences in the sequenced gene sequences corresponding to the Nephila clavata spiders that do not meet the third preset high-quality condition; the third preset high-quality condition is the condition determined according to the parameters of the fastp V0.19.5 software.
[0083] It should be noted that Canu is a genome assembly software; SMARTdenovo is a de novo assembler; Racon is a software for constructing consensus sequences, and its feature is high speed; Pilon is a genome error correction software.
[0084] Step S22: Use the paired-end Hi-C short reads in the sequenced gene sequences to trim the polished contigs to obtain trimmed contigs, and use the trimmed contigs to construct pseudochromosomes, and then determine gene position information based on the pseudochromosomes; map the gene position information to the polished contigs to obtain the chromosomal genome of the Nephila clavata spiders that meets the first preset high-quality condition.
[0085] In this embodiment, Hic pro v2.10 is used to trim the polished contigs based on the paired-end Hi-C short reads to obtain trimmed contigs; 3D-DNA v180419 is used and the trimmed contigs are used to construct pseudochromosomes, and then the Juicebox v1.9 tool is used to manually correct the pseudochromosomes to determine gene position information, and based on Juicer v1.6.2, the gene position information is mapped to the polished contigs to obtain the chromosomal genome of the Nephila clavata spiders that meets the first preset high-quality condition.
[0086] It should be noted that HiC-Pro is an efficient Hi-C data analysis software; 3D-DNA is a simple and convenient Hi-C processing software that can lift contigs to the chromosomal level; the Juicebox v1.9 tool is a graphical interface tool that can conveniently view and display Hi-C maps; Juicer is a tool for assisting genome assembly.
[0087] Step S23: Based on the general transcriptomes corresponding to several tissues of the Nephila clavata spider and the protein sequences of each species closely related to the Nephila clavata spider, annotate the chromosomal genome to obtain an initial annotation file, and delete the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file.
[0088] In this embodiment, based on the general transcriptomes corresponding to several tissues of the Nephila clavata spider and the protein sequences of each species closely related to the Nephila clavata spider, annotating the chromosomal genome to obtain an initial annotation file includes: de novo annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spider, and at the same time performing homology annotation on the chromosomal genome based on the protein sequences of each species closely related to the Nephila clavata spider to obtain an initial annotation file.
[0089] It should be noted that each species closely related to the Nephila clavata spider includes four species: A. ventricosus (Araneus ventricosus, the orb-weaving spider), A. bruennichi (Argiope bruennichi, the wasp spider), T. antipodiana (Trichonephila antipodiana, the antipodal nephila spider), and T. clavipes (Trichonephila clavipes, the golden silk orb-weaver).
[0090] It should be noted that de novo annotation and homology annotation are performed using the AUGUSTUS V3.2.3 software; AUGUSTUS is a software for predicting gene structures in eukaryotic genomes.
[0091] In this embodiment, the screening criteria include that the proportion of the repetitive sequence corresponding to a single gene sequence is less than the first target proportion, the length of the protein sequence corresponding to a single gene sequence is greater than the target length, the FPKM value of a single gene sequence in more than the target number of tissues is greater than the target value, a single gene sequence comes from each of the species, and the identity between the full-length transcriptome sequencing data and a single gene sequence is higher than the second target proportion.
[0092] It should be noted that the first target proportion is 50%; the target length is 50 aa; the target quantity is 10; the target value is 0.1; the second target proportion is 95%; FPKM is Fragments Per Kilobase of exon model per Million mapped fragments, that is, the fragments of transcripts per thousand bases per million mapped reads; fragments refer to each nucleic acid fragment used for sequencing.
[0093] Step S24: Filter the sequenced ordinary transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality conditions, and determine the expression values and co-expression relationships of each gene sequence based on the high-quality transcriptome sequencing sequences.
[0094] Among them, for the more specific processing process of step S24, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0095] Step S25: Classify the sequenced single-cell transcriptome to obtain classified cell clusters, and perform functional annotation on the classified cell clusters to obtain an annotated transcriptome.
[0096] Among them, for the more specific processing process of step S25, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0097] Step S26: Preprocess the sequenced spatial transcriptome, and perform clustering on the sequenced spatial transcriptome based on the preprocessing results to determine the cell distribution within the spatial structure.
[0098] Among them, for the more specific processing process of step S26, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0099] Step S27: According to the database management system, construct a database based on Nephila clavata spiders based on the target annotation file, the expression values and co-expression relationships, the annotated transcriptome, and the cell distribution within the spatial structure.
[0100] In this embodiment, the web page (module) corresponding to the constructed database can be divided into: gene information, eFP, subcellular localization, gene co-expression, gene family, sequenced single-cell transcriptome, JBrowser, chromosome localization, collinearity, orthologous genes, and spatially sequenced single-cell transcriptome, etc.
[0101] In this embodiment, the gene information module displays basic information on multiple aspects of the selected gene, such as gene ID (Identity document), protein domain function description, gene distribution, and gene location. In addition, there are also some functional annotations from multiple databases, such as KO (KEGG Orthology), GO (Glycollic oxidase), KOG, Pfam, and kegg enzymes, as well as Pub-Med links related to the gene. In addition, the structure diagram shows that exons, untranslated regions (UTRs), and introns present the gene structure of the selected gene.
[0102] In this embodiment, the eFP module presents the expression values of the selected gene in 17 different tissues in different colors.
[0103] In this embodiment, in the subcellular localization module, the database shows the predicted subcellular localization of proteins using ng-LOC, which is an n-gram-based Bayesian classifier that can predict the subcellular localization of proteins in prokaryotes and eukaryotes. The predicted localization of the gene in the cell picture is shown, and the color gradient represents the confidence score of finding the selected gene in a given compartment.
[0104] In this embodiment, gene co-expression relationships are predicted based on transcriptome data, and a co-expression network is generated. The gene co-expression module provides a view of co-expressed genes by showing co-expressed genes in a network with interactive analysis tools.
[0105] In this embodiment, to study the evolution of *Nephila clavata*, the database contains a comprehensive gene comparison and evolution dataset. The gene family module provides a user-friendly graphical view showing the gene structure and Pfam domain pattern diagram connected to a bootstrap similarity dendrogram.
[0106] In this embodiment, after sequencing, the single-cell transcriptome module shows the distribution of each gene in each cell cluster.
[0107] In this embodiment, the JBrowser module shows the gene distribution in the species sequence, including exons, introns, cds, and reference sequences.
[0108] In this embodiment, the chromosome localization module provides a friendly way to show the positions of the selected gene and its family on the chromosome based on a chromosome viewer.
[0109] In this embodiment, the synteny module shows the synteny of chromosomes between *Nephila clavata* and other spider species, as well as the positions of the selected genes on the chromosomes.
[0110] In this embodiment, the orthologous gene module shows which orthogonal clusters contain the selected genes.
[0111] In this embodiment, the single-cell transcriptome module after spatial sequencing shows the distribution and expression of each gene in different spaces.
[0112] In this embodiment, multiple data visualization tools are combined into the same interface, allowing users to explore biological data at multiple levels and comprehensively analyze genes. Additionally, this database can be applied to the study of the major ampullate glands of spiders, providing important data for elucidating the silk-spinning mechanism of spiders. This database can also be applied to multi-omics studies such as genomics, transcriptomics, proteomics, single-cell transcriptomics after sequencing, and spatial transcriptomics after sequencing, laying a foundation for other spider molecular research.
[0113] It can be seen that this application filters and deletes all low-quality long read sequences in the sequenced gene sequences corresponding to the Nephila clavata spiders that do not meet the third preset high-quality condition; performs initial assembly on the filtered high-quality long read sequences to obtain primary contigs; corrects the primary contigs to obtain corrected contigs; polishes the corrected contigs using the short read sequences in the sequenced gene sequences corresponding to the Nephila clavata spiders to obtain polished contigs; trims the polished contigs using the paired-end Hi-C short reads in the sequenced gene sequences to obtain trimmed contigs, and constructs pseudochromosomes using the trimmed contigs, and then determines gene position information based on the pseudochromosomes; maps the gene position information to the polished contigs to obtain the chromosomal genome of the Nephila clavata spiders that meets the first preset high-quality condition. Annotates the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of various species closely related to the Nephila clavata spiders to obtain an initial annotation file, and deletes the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file; filters the sequenced general transcriptomes to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and determines the expression values and co-expression relationships of each gene sequence based on the high-quality transcriptome sequencing sequences; classifies cell clusters for the single-cell transcriptomes after sequencing to obtain classified cell clusters, and performs functional annotation on the classified cell clusters to obtain an annotated transcriptome; preprocesses the spatial transcriptomes after sequencing, and clusters the spatial transcriptomes after sequencing based on the preprocessing results to determine the cell distribution within the spatial structure; constructs a database based on the Nephila clavata spiders according to the database management system, based on the target annotation file, the expression values and the co-expression relationships, the annotated transcriptome, and the cell distribution within the spatial structure. Thus, it can be seen that this application constructs a comprehensive spider database to facilitate the rapid acquisition of comprehensive information on the genomes and other omics of spiders.
[0114] Correspondingly, the embodiment of the present application also discloses a database construction device based on the Nephila clavata spider. Refer to Figure 3 as shown, the device includes:
[0115] An assembly module 11 for assembling the chromosomal genome of the Nephila clavata spider that meets the first preset high-quality condition;
[0116] A first annotation module 12 for annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spider and the protein sequences of each species closely related to the Nephila clavata spider to obtain an initial annotation file;
[0117] A deletion module 13 for deleting the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file;
[0118] A filtering module 14 for filtering the sequenced general transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences;
[0119] A classification module 15 for classifying cell clusters of the sequenced single-cell transcriptome to obtain classified cell clusters;
[0120] A second annotation module 16 for functionally annotating the classified cell clusters to obtain an annotated transcriptome;
[0121] A clustering module 17 for preprocessing the sequenced spatial transcriptome, and clustering the sequenced spatial transcriptome based on the preprocessing result to determine the cell distribution within the spatial structure;
[0122] A database construction module 18 for constructing a database based on the Nephila clavata spider according to a database management system and based on the target annotation file, the expression value and the co-expression relationship, the annotated transcriptome, and the cell distribution within the spatial structure.
[0123] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0124] It can be seen that the present application constructs a comprehensive spider database to facilitate the rapid acquisition of comprehensive information on the genomes and other omics of spiders.
[0125] Furthermore, the embodiment of the present application also provides an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of the present application.
[0126] Figure 4 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the database construction method based on Nephila spiders disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0127] In this embodiment, the power supply 26 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed on it here; the input / output interface 24 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are made here.
[0128] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include a computer program 221, and the storage method may be short-term storage or permanent storage. Among them, in addition to the computer program that can be used to complete the database construction method based on Nephila spiders executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 221 may further include a computer program that can be used to complete other specific tasks.
[0129] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the database construction method based on Nephila spiders disclosed above.
[0130] For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0131] The various embodiments in this application are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference may be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference may be made to the description in the method part for related parts.
[0132] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0133] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.
[0134] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0135] The above has introduced in detail a method, apparatus, device, and storage medium for constructing a database based on Nephila clavata spiders provided by this application. Specific examples are used in this document to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for constructing a database based on the spider Nephila clavata, characterized in that, Comprising: Assembling the chromosomal genome of Nephila clavata that meets the first preset high-quality condition; Annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata and the protein sequences of each species closely related to the Nephila clavata to obtain an initial annotation file, and deleting the annotations corresponding to each gene sequence that does not meet the screening criteria in the initial annotation file to obtain a target annotation file; Filtering the sequenced general transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences; Classifying cell clusters of the sequenced single-cell transcriptome to obtain a classified transcriptome containing several classified cell clusters, and performing functional annotation on the classified cell clusters to obtain an annotated transcriptome; Preprocessing the sequenced spatial transcriptome, and clustering the sequenced spatial transcriptome based on the preprocessing results to determine the cell distribution within the spatial structure; Constructing a database based on Nephila clavata according to a database management system, and based on the target annotation file, the expression value and the co-expression relationship, the annotated transcriptome, and the cell distribution within the spatial structure.
2. The database construction method based on Nephila clavata spiders according to claim 1, characterized in that The assembling the chromosomal genome of Nephila clavata that meets the first preset high-quality condition includes: Filtering and deleting all low-quality long read sequences in the sequenced gene sequences corresponding to the Nephila clavata that do not meet the third preset high-quality condition; Performing de novo assembly on the filtered high-quality long read sequences to obtain primary contigs; Correcting the primary contigs to obtain corrected contigs; Polishing the corrected contigs using the short read sequences in the sequenced gene sequences corresponding to the Nephila clavata to obtain polished contigs; Trimming the polished contigs using paired-end Hi-C short reads to obtain trimmed contigs, and constructing pseudochromosomes using the trimmed contigs, and then determining gene position information based on the pseudochromosomes; Mapping the gene position information to the polished contigs to obtain the chromosomal genome of Nephila clavata that meets the first preset high-quality condition.
3. The database construction method based on Nephila clavata spiders according to claim 1, wherein The annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata and the protein sequences of each species closely related to the Nephila clavata to obtain an initial annotation file includes: Performing de novo annotation on the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata, and simultaneously performing homology annotation on the chromosomal genome based on the protein sequences of each species closely related to the Nephila clavata to obtain an initial annotation file.
4. The method for constructing a database based on Nephila clavata spiders according to any one of claims 1 to 3, characterized in that The screening criteria include that the proportion of repetitive sequences corresponding to a single gene sequence is less than the first target proportion, the length of the protein sequence corresponding to a single gene sequence is greater than the target length, the FPKM value of a single gene sequence in more than the target number of the tissues is greater than the target value, the single gene sequence comes from the respective species, and the identity between the full-length transcriptome sequencing data and the single gene sequence is higher than the second target proportion.
5. The database construction method based on Nephila clavata spiders according to claim 1, characterized in that Determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences includes: Aligning the high-quality transcriptome sequencing sequences with the chromosomal genome to obtain an alignment result representing the corresponding relationship between each position in the chromosomal genome and each gene sequence in the high-quality transcriptome sequencing sequences; Sorting the alignment result according to the positional relationship of the chromosomal genome and storing it as a target result in a target format; Calculating the expression value and co-expression relationship of each gene sequence in the high-quality transcriptome sequencing sequences based on the Pearson correlation coefficient and the target result.
6. The database construction method based on Nephila clavata spiders according to claim 1, characterized in that, Classifying the sequenced single-cell transcriptome into classified cell clusters includes: Removing low-quality cells in the sequenced single-cell transcriptome that do not meet the preset cell conditions, and obtaining a first gene expression matrix of each cell with the column coordinates being cells and the row coordinates being genes for the other cells outside the low-quality cells; Clustering the other cells based on the first gene expression matrix to obtain classified cell clusters.
7. The database construction method based on Nephila clavata spiders according to claim 1, wherein Preprocessing the sequenced spatial transcriptome and clustering the sequenced spatial transcriptome based on the preprocessing result to determine the cell distribution within the spatial structure includes: Removing low-quality reads in the sequenced spatial transcriptome, and obtaining a second gene expression matrix of each cell with the column coordinates being cells and the row coordinates being genes for the other reads outside the low-quality reads; Clustering the cells corresponding to the other reads based on the second gene expression matrix to determine the cell distribution within the spatial structure.
8. A database construction device based on the Nephila clavata spider, characterized in that, Including: An assembly module for assembling the chromosomal genome of Nephila clavata spiders that meets the first preset high-quality condition; A first annotation module for annotating the chromosomal genome based on the general transcriptomes corresponding to several tissues of the Nephila clavata spiders and the protein sequences of the respective species closely related to the Nephila clavata spiders to obtain an initial annotation file; A deletion module for deleting the annotations corresponding to each gene sequence in the initial annotation file that do not meet the screening criteria to obtain a target annotation file; A filtering module for filtering the sequenced general transcriptome to obtain high-quality transcriptome sequencing sequences that meet the second preset high-quality condition, and determining the expression value and co-expression relationship of each gene sequence based on the high-quality transcriptome sequencing sequences; A classification module for classifying the sequenced single-cell transcriptome into classified cell clusters; A second annotation module for functionally annotating the classified cell clusters to obtain an annotated transcriptome; A clustering module, which is used to preprocess the sequenced spatial transcriptome and cluster the sequenced spatial transcriptome based on the preprocessing results to determine the cell distribution within the spatial structure; A database construction module, which is used to construct a database based on Nephila clavata spiders according to a database management system and based on the target annotation file, the expression values and the co-expression relationships, the annotated transcriptome, and the cell distribution within the spatial structure.
9. An electronic device, characterized in that, Comprising: A memory, which is used to store a computer program; A processor, which is used to execute the computer program to implement the method for constructing a database based on Nephila clavata spiders according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the method for constructing a database based on Nephila clavata spiders according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A
Single cell transcriptome computation and analysis method and system incorporating deep learning model
WO2022188785A1