Method for identification and evaluation of the annelid fibrinolysin family
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINGGANGSHAN UNIVERSITY
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-07
AI Technical Summary
这种策略效率不高,每次只能鉴定一个蛋白序列
Smart Images

Figure CN122531484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular pharmacognosy, specifically to methods for identifying and evaluating the annelid fibrinolytic enzyme family. Background Technology
[0002] Annelids are a group of higher worms with homologous segmentation of the body, including representative species such as leeches, earthworms, worm-like insects, and sipunculids, and they have significant economic and ecological value. Researchers both domestically and internationally have discovered numerous bioactive proteins with anticoagulant, thrombolytic, anti-inflammatory, and antioxidant functions in leeches, earthworms, and other annelids. These bioactive proteins can serve as important pharmacognosy resources for maintaining cardiovascular health. For example, hirudin, found in leeches, is the most potent anticoagulant bioactive substance discovered to date; while lumbrokinase, found in earthworms, is a fibrinolytic enzyme with strong thrombolytic activity.
[0003] Annelid fibrinolytic enzymes (hereinafter referred to as cyclolysins) are encoded by a multi-gene family, making them a relatively complex protein family. Currently, the identification of cyclolysins generally employs mass spectrometry sequencing of proteins extracted from tissues. This strategy is inefficient, identifying only one protein sequence at a time. Therefore, how to efficiently identify and assess the importance of cyclolysin family members is a pressing technical problem that needs to be solved in this field. Summary of the Invention
[0004] Based on this, the present invention provides a method for identifying and evaluating the annelid fibrinolytic enzyme family, thereby solving at least one problem in the prior art.
[0005] In a first aspect, the present invention provides a method for identifying a family of fibrinolytic enzymes in annelids, comprising the following steps: Obtain reference genome data, coding sequence data (CDS), and transcriptome reads of annelids, and assemble them into transcriptome sequences (unigenes). Based on the sequence of lumbrokinase, a potential homologous sequence (SeqA) is searched in the reference genome data, and based on the potential homologous sequence (SeqA), the transcript sequence with the highest similarity is searched in the transcriptome sequence (SeqB). Based on the potential homologous sequence (SeqA) and the transcript sequence with the highest similarity (SeqB), the codon sequence of the complete protein-coding region is obtained (SeqC). Based on the codon sequence (SeqC), highly homologous sequences are searched in the genome sequence to determine the complete annelid fibrinolytic enzyme encoding gene sequence (SeqD). Redundant sequences were removed from the annelid fibrinolytic enzyme encoding gene sequence (SeqD) to obtain the annelid fibrinolytic enzyme sequence (SeqE).
[0006] In some optional embodiments, the annelid's reference genome data, coding sequence data (CDS), and transcriptome reads are sourced from the GenBank database and assembled into transcriptome sequences (unigene) using Trinity software.
[0007] In some optional embodiments, the potential homologous sequence refers to a homologous sequence with a sequence similarity of >=30% to lumbrokinase. When searching for potential homologous sequences (SeqA), lumbrokinase can be used as the bait sequence, and the exonerate software can be used for the search.
[0008] In some optional embodiments, the transcript sequence with the highest similarity (SeqB) refers to a transcript sequence with a similarity of >= 95% to the potential homologous sequence (SeqA). When searching for the highest similarity transcript sequence (SeqB), each SeqA can be used as a bait sequence, and the search can be performed using the blastn software in the BLAST package.
[0009] In some optional embodiments, the codon sequence (SeqC) includes sequences such as the coding region sequence, start codon, and stop codon. For example, the potential homologous sequence (SeqA) and the transcript sequence with the highest similarity (SeqB) can be imported into MEGA software for sequence alignment to determine information such as the coding region sequence, start codon, and stop codon, thereby obtaining the complete codon sequence (SeqC) of the protein coding region.
[0010] In some optional embodiments, the highly homologous sequence refers to a homologous sequence with a similarity of >= 95% to the codon sequence (SeqC). When searching for highly homologous sequences, the codon sequence (SeqC) can be used as bait sequences, and the search can be performed using the exonerate software. Based on the highly homologous sequence, combined with the GT-AG rule and data such as start codon, stop codon, exons, and introns, the complete annelid fibrinolytic enzyme encoding gene sequence (SeqD) is obtained.
[0011] In some optional embodiments, when removing redundant sequences from the annelid fibrinolytic enzyme encoding gene sequence (SeqD), MEGA software can be used for sequence alignment, combining sequence similarity and coordinates in a reference genome to remove redundant sequences.
[0012] Secondly, the present invention provides a method for evaluating gene expression of the fibrinolytic enzyme family of annelids, comprising the following steps: Sequence name information (ID) is obtained based on the annelid fibrinolytic enzyme sequence (SeqE) and coding sequence data (CDS). Remove all matching sequences corresponding to the sequence name information (ID) from the coded sequence data (CDS), and merge the remaining coded sequence data (CDS) with the annelid fibrinolytic enzyme sequence (SeqE) to obtain the reference sequence dataset (SeqF). The TPM (Transcripts per million) value was obtained based on the reference sequence dataset (SeqF) and transcriptome read data. The TPM values of the annelid plasminogen lysate sequences (SeqE) corresponding to each transcriptome read were combined, and the mean TPM value of the coding region of each annelid plasminogen lysate gene was calculated as its relative expression level data.
[0013] In some optional embodiments, the blastp software in the BLAST package is used to compare the encoded sequence data (CDS) with the annelid fibrinolytic enzyme sequence (SeqE) as the bait sequence to obtain the sequence name information (ID) of the matching sequence.
[0014] In some alternative embodiments, the grep program in the seqkit software is used to remove all matching sequences corresponding to the sequence name information (ID) from the encoded sequence data (CDS).
[0015] In some optional embodiments, 3-5 transcriptome reads are randomly selected, and each transcriptome read is aligned to the reference sequence dataset (SeqF) using Salmon software to obtain the TPM value.
[0016] Thirdly, the present invention provides a method for evaluating the protein expression of an annelid fibrinolytic enzyme family, comprising the following steps: Total protein was extracted from live annelids; The protein was enzymatically hydrolyzed using trypsin and LysC enzyme, and the resulting peptides were desalted. The desalted peptides were separated by liquid chromatography, and each separated sample was mixed with an internal standard (iRT) at a preset volume ratio for mass spectrometry detection. The reference sequence dataset (SeqF) was translated into protein sequences and merged with the internal standard (iRT) sequence as the target for a database search. The relative expression levels of each type of fibrinolytic enzyme in annelids were then calculated.
[0017] In some optional embodiments, the extraction of total protein from live annelids includes: homogenizing the live annelids, filtering the supernatant, extracting the total protein with a saturated phenol-Tris-HCl (7.8) solution; precipitating the protein with a 0.1 M ammonium acetate-methanol solution, and washing the protein successively with methanol and acetone.
[0018] In some optional embodiments, the enzymatic hydrolysis of the protein using trypsin and LysC enzyme comprises: mixing the protein with chloroacetic acid, trypsin, and LysC enzyme, and hydrolyzing at 37°C and 1500 rpm for 1-2 hours with shaking.
[0019] In some optional embodiments, the desalination employs SOLA. TM The procedure was performed using an SPE 96-well plate.
[0020] In some optional embodiments, the volume ratio of each separated sample to internal standard (iRT) is 1:20.
[0021] In some optional embodiments, the database search is performed using DIA-NN software.
[0022] Because of the adoption of the above technical solutions, the embodiments of the present invention have at least the following beneficial effects: (1) By making full use of the genomic and transcriptomic data on public databases, and through repeated comparison and searching, the vast majority of cyclolysin gene family members can be identified (dozens of cyclolysin family members can be obtained from each annelid), with high sensitivity; (2) By combining software comparison with manual verification, the accuracy of the identified sequence can be guaranteed to the maximum extent, while avoiding false positives or sequence redundancy, resulting in high reliability. (3) By using transcriptome data from public databases and proteome data from our own sequencing, we evaluated the expression levels of cyclolysin family members from the perspectives of genes and proteins, respectively, resulting in more objective, reliable and accurate results. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the process for identifying and evaluating the cyclolysin family in an embodiment of the present invention.
[0024] Figure 2 This is a sequence alignment result of three representative cyclolysins from lumbrokinase and broad-bodied golden thread leeches in an embodiment of the present invention.
[0025] Figure 3 This is a diagram showing the sequence alignment results of lumbrokinase and four representative cyclosporine sequences from Eisenia fetida in this embodiment of the invention.
[0026] Figure 4 This is a diagram showing the sequence alignment results of lumbrokinase and three representative cyclosporines from *Ulva uniflora* in an embodiment of the present invention.
[0027] Figure 5 This is a diagram showing the sequence alignment results of five representative cyclosporine compounds from lumbrokinase and Sipunculus nudus in an embodiment of the present invention. Detailed Implementation
[0028] The following will provide a clear and complete description of the concept and technical effects of the present invention, so as to fully explain the purpose, solution and effects of the present invention.
[0029] like Figure 1 As shown, the following examples illustrate the identification and evaluation of the annelid fibrinolytic enzyme (cyclolysin) family according to steps 1-3: 1. Cyclolysin sequence identification (1) Acquisition of genome and transcriptome data: The reference genome data (genome) of the target annelid species and all coding region sequence data (CDS) from structural annotation were retrieved and downloaded from GenBank; at the same time, all transcriptome reads files of this species were downloaded and assembled using Trinity software to obtain transcriptome sequences (unigene). (2) Homologous sequence search of cyclolysin: Using lumbrokinase (GenBank No. AAN28692.1) as bait sequence, the exonerate software was used to search for all potential homologous sequences (similarity >= 30%) in the genome, denoted as SeqA; using each of the above SeqA as bait sequence, the blastn software in the BLAST package was used to search for the unigene sequence with the highest similarity (similarity >= 95%) in the transcriptome sequence, denoted as SeqB; (3) Cyclolysin gene sequence identification: The SeqA and SeqB sequences were imported into MEGA software for sequence alignment to determine the coding region sequence, start codon, stop codon, and other information, and the complete codon sequence of the protein coding region was obtained, denoted as SeqC; the SeqC sequence was used as the bait sequence, and the exonerate software was used to search for highly homologous sequences (similarity >= 95%) in the genome. Combined with the GT-AG rule and data such as start codon, stop codon, exon, and intron, the complete cyclolysin coding gene was obtained, denoted as SeqD; (4) Removal of redundant cyclolysin sequences: All the SeqD sequences were merged and sequence alignment was performed using MEGA software. Combining sequence similarity and coordinates in the reference genome, redundant sequences were removed, and finally a usable cyclolysin SeqE dataset was obtained.
[0030] 2. Evaluation of cyclolysin gene expression (1) Gene dataset sequence retrieval: Using the blastp software in the BLAST package, the SeqE sequence is used as bait to compare the CDS sequence and obtain the sequence name information ID of the matching sequence; (2) Construction of reference gene dataset: Using the ID information, remove all matching sequences using the grep program in the seqkit software; merge the SeqE sequence with the remaining CDS sequence to obtain a new reference sequence dataset, denoted as SeqF; (3) Transcript reads sequence alignment: Randomly select 3-5 transcriptome data, use the SeqF dataset as a template, and use Salmon software to align the reads of each transcriptome to the SeqF dataset to obtain the TPM (Transcripts per million) value; (4) Cyclolysin gene expression analysis: The TPM values of the SeqE (cyclolysin) sequence corresponding to each transcriptome data were merged in an Excel table, and the average TPM value of the coding region of each cyclolysin gene was calculated as its relative expression data.
[0031] 3. Evaluation of Cyclolysin Protein Expression (1) Protein extraction and purification: live annelids were homogenized, the supernatant was filtered, and total protein was extracted with phenol-Tris-HCl (7.8) saturated solution; protein was precipitated with 0.1 M ammonium acetate-methanol solution, and the protein was washed with methanol and acetone successively. (2) Enzymatic hydrolysis and desalting: Chloroacetic acid, trypsin and LysC enzyme were added to the protein and the protein was hydrolyzed at 37°C and 1500 rpm for 1-2 hours. The peptides after hydrolysis were desalted using SOLA™ SPE 96-well plates. (3) Enzymatic hydrolysis peptide mass spectrometry detection: The enzymatically hydrolyzed peptides are separated by liquid chromatography and sequentially sent to the mass spectrometer for detection; before mass spectrometry injection, each sample is mixed with the internal standard (iRT): sample = 1:20 by volume ratio as internal standard; (4) Cyclolysin protein expression analysis: The SeqF dataset was translated into protein sequences and merged with the internal standard (iRT) sequence as the search target. The DIA-NN software was used for search analysis to calculate the relative expression level of each cyclolysin.
[0032] Example 1: Identification and evaluation of the cyclolysin family of Hirudo medicinalis Broad-bodied golden leech ( Whitmania pigra Leeches, also known as blood leeches, are listed in the Pharmacopoeia of the People's Republic of China and are an important traditional Chinese medicine for antithrombosis. The inventors downloaded genome assembly data (No. GCA_041430665.1) and raw transcriptome data (No. CRA030167) from GenBank. After sequence assembly, alignment, and redundancy removal, four complete cyclolysin sequences (Wpig01-Wpig04) were identified. Transcriptome analysis revealed that only one cyclolysin-encoding gene (Wpig04) was expressed. Live broad-bodied golden leeches were collected from Xiajiang County, Ji'an City, Jiangxi Province, and randomly divided into three groups for proteome sequencing, obtaining data from three groups: Leech 1-Leech 3. Proteome analysis showed that three cyclolysins (Wpig02, Wpig03, and Wpig04) were expressed to varying degrees (Table 1). Therefore, it can be inferred that Wpig04 has greater value (for example, it can be used as a preferred industry-academia-research development sequence), followed by Wpig02 and Wpig03, while Wpig01 has little value.
[0033] Table 1. Relative expression levels of the hibiscus lysin gene and its protein in *Hirudinia bromos*. Example 2: Identification and evaluation of the cyclosporine family in Eisenia fetida. The child loves the earthworm ( Eisenia fetida Earthworm (E. spp.) is a type of earthworm, and its extract, known as earthworm protein, is listed in the "Catalogue of New Resource Foods of the People's Republic of China". The inventors downloaded genome assembly data (No. GCA_003999395.1) and raw transcriptome data (No. PRJNA663354) from GenBank, and raw transcriptome data (No. PRJCA014735) from the National Center for Biotechnology Information. Through sequence assembly and alignment, 17 complete cyclolysin sequences (Efet01-Efet17) were identified. Transcriptome analysis revealed that all 17 cyclolysin coding genes were expressible. Live specimens of *Eisenia fetida* were collected from Zhanghuang Town, Yutai County, Jining City, Shandong Province, and randomly divided into three groups for proteome sequencing, yielding data from three groups: Earthworm 1-Efetida 3. Proteome analysis showed that only four cyclolysin proteins (Efet02, Efet06, Efet08, and Efet13) were expressed (Table 1). Combined with sequence alignment results ( Figure 2 Based on the gene and protein expression information, it is inferred that these four cyclolysins (Efet02, Efet06, Efet08, and Efet13) have relatively high research and development value.
[0034] Table 2. Relative expression levels of the cyclosporin gene and its protein in Eisenia fetida. Example 3: Identification and evaluation of the cyclolysin family in *Urtica monocyclicis* Uni-ringed thorn beetle ( Urechis unicinctus Sea cucumber, also known as sea worm, is a marine annelid rich in nutritional and medicinal value. The inventors downloaded genome assembly data (No. GCA_034190875.2) and raw transcriptome data (No. PRJNA917787 and PRJNA485379) from GenBank. Through sequence assembly, alignment, and redundancy removal, 19 complete cyclolysin sequences (Uuni01-Uuni19) were identified. Transcriptome analysis revealed that all 19 cyclolysin-coding genes were expressible. A batch of *Urocystis monocyclicis* was purchased online from a Tmall store (Ganhai Story Flagship Store, address: Haitou Town, Ganyu District, Lianyungang City, Jiangsu Province), randomly divided into 3 groups, and proteomic sequencing was performed, obtaining data from 3 groups: sea cucumber 1-sea cucumber 3. Proteomic analysis showed that only 7 cyclolysin proteins were expressed (Table 2). Combined with sequence alignment results (… Figure 3 Based on the gene and protein expression information, it is inferred that three cyclolysins (Uuni01, Uuni03, and Uuni11) have relatively high research and development value.
[0035] Table 3. Relative expression levels of the cyclolysin gene and its protein in *Ulva monocyclicis*. Example 4: Identification and evaluation of the cyclosporine family in Sipuncula lataniae Sipuntia (a type of cricket) Sipunculus nudus Also known as sandworm. The inventors downloaded genome assembly data (No. GCA_026874595.2) and raw transcriptome data (No. PRJNA592829, PRJNA777348, and PRJNA916931) from GenBank. After sequence assembly, alignment, and redundancy removal, 29 complete cyclolysin sequences (Snud01-Snud29) were identified. Transcriptome analysis revealed that 10 genes were expressed. A batch of square-patterned Sipunculus worms was purchased online from Tmall (Ganhai Story Flagship Store, address: Haitou Town, Ganyu District, Lianyungang City, Jiangsu Province), randomly divided into 3 groups, and proteome sequencing was performed, obtaining 3 groups of data: Sandworm 1-Sandworm 3. Proteome analysis detected the expression of 11 proteins (Table 4). Based on the above information, it is inferred that the five cyclosporins (Snud01, Snud11, Snud15, Snud16, and Snud17) have relatively high research and development value.
[0036] Table 4. Relative expression levels of the cyclosporine gene and its protein in Sipuncula lataniae. The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. Any embodiment that achieves the technical effects of the present invention by the same or equivalent means should fall within the protection scope of the present invention. Within the protection scope of the present invention, various modifications and variations can be made to the technical solutions and / or implementation methods.
Claims
1. A method for identifying a family of fibrinolytic enzymes from annelids, characterized in that, Includes the following steps: Obtain reference genome data, coding sequence data, and transcriptome read data of annelids, and assemble them into transcript sequences; Based on the sequence of lumbrokinase, potential homologous sequences are searched in the reference genome data, and based on the potential homologous sequences, the transcript sequence with the highest similarity is searched in the transcriptome sequence; The codon sequence of the complete protein-coding region is obtained based on the potential homologous sequence and the transcript sequence with the highest similarity. Based on the codon sequence, highly homologous sequences are searched in the genome sequence to determine the complete annelid fibrinolytic enzyme encoding gene sequence. Redundant sequences were removed from the annelid fibrinolytic enzyme encoding gene sequence to obtain the annelid fibrinolytic enzyme sequence.
2. The method according to claim 1, characterized in that, The reference genome data, coding sequence data, and transcriptome read data of the annelids were obtained from the GenBank database and assembled into transcript sequences using Trinity software.
3. The method according to claim 1, characterized in that, The potential homologous sequence refers to a homologous sequence with a sequence similarity of >=30% to lumbrokinase.
4. The method according to claim 1, characterized in that, The transcript sequence with the highest similarity refers to the transcript sequence with a similarity of >= 95% to the potential homologous sequence.
5. The method according to claim 1, characterized in that, The codon sequence includes a coding region sequence, a start codon sequence, and a stop codon sequence.
6. The method according to claim 1, characterized in that, The highly homologous sequence refers to a homologous sequence with a similarity of >=95% to the codon sequence.
7. The method according to claim 1, characterized in that, When removing redundant sequences from the gene sequence encoding fibrinolytic enzyme in the annelids, MEGA software was used for sequence alignment. The redundant sequences were removed by combining sequence similarity with coordinates in the reference genome.
8. A method for evaluating gene expression in annelid fibrinolytic enzymes, characterized in that, Includes the following steps: According to claim 1, the sequence name information is obtained from the annelid fibrinolytic enzyme sequence and coding sequence data; Remove all matching sequences corresponding to the sequence name information from the encoded sequence data, and merge the remaining encoded sequence data with the annelid fibrinolytic enzyme sequence to obtain a reference sequence dataset; TPM values were obtained based on the reference sequence dataset and transcriptome read data. The TPM values of the annelid plasminase sequences corresponding to each transcriptome read were combined, and the mean TPM value of the coding region of each annelid plasminase gene was calculated as its relative expression level data.
9. A method for evaluating the protein expression of an annelid fibrinolytic enzyme family, characterized in that, Includes the following steps: Total protein was extracted from live annelids; The protein was enzymatically hydrolyzed using trypsin and LysC enzyme, and the resulting peptides were desalted. The desalted peptides were separated by liquid chromatography, and each separated sample and internal standard were mixed at a preset volume ratio and then analyzed by mass spectrometry. The reference sequence dataset described in claim 8 is translated into protein sequences and merged with the internal standard sequence as the target for a database search. The relative expression levels of fibrinolytic enzymes of each annelid are then calculated.
10. The method according to claim 9, characterized in that, The volume ratio of each separated sample to internal standard is 1:20.