A method, device and system for comparing multiple gene sequences

By introducing automatic editing and correction steps in the gene multi-sequence alignment method, the problem that gene sequence data in the prior art does not meet the software input requirements is solved, efficient and accurate gene multi-sequence alignment is achieved, and the workload of researchers is significantly reduced.

CN114420207BActive Publication Date: 2025-05-23NORTHWEST INST OF PLATEAU BIOLOGY CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210096587.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-05-23
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Existing gene sequence alignment methods and software require standardized gene sequence data when used. However, the data obtained by actual sequencing often does not meet the input requirements of these software, resulting in researchers requiring manual editing and correction. The increased workload is not conducive to efficient gene multi-sequence alignment.

Method used

A gene multi-sequence alignment method is provided, including automatic editing and correction steps, correcting the sequence to be analyzed by setting template information and multi-sequence global alignment, forming a data set, and reducing sequencing error sites by identifying and processing antisense strand sequences, cutting and discarding sequencing vector sequences, filling sequence end defects, and deleting non-accurate region sequences.

Benefits of technology

Through automated correction and alignment processes, the workload of researchers is significantly reduced, the efficiency of gene multi-sequence alignment is improved, and the accuracy of alignment results is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114420207B_ABST
    Figure CN114420207B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of bioinformatics, and specifically relates to a gene multiple sequence alignment method, device and system. The gene multiple sequence alignment method of the present invention comprises the following steps: step 1, setting template information, wherein the template information comprises the template sequence, the position of the conserved motif and the length of the conserved motif of the gene to be analyzed; step 2, using the template information of step 1 as a reference, correcting the sequence to be analyzed and globally aligning the multiple sequences to form a data set. The present invention also provides a device and a system for applying the above-mentioned gene multiple sequence alignment method. The technical solution of the present invention can greatly reduce the workload of gene multiple sequence alignment and significantly improve the efficiency of multiple sequence alignment work. Therefore, it has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a gene multiple sequence alignment method, device and system. Background Art

[0002] Multiple gene sequence alignment is a fundamental component and important foundation of bioinformatics. The basic idea of ​​sequence alignment is to treat nucleic acid sequences and protein primary structure sequences as strings composed of basic characters, based on the universal biological principle that sequence determines structure, and structure determines function. This approach detects similarities between sequences and reveals functional, structural, and evolutionary information within biological sequences.

[0003] For example, multiple sequence comparison of the Cytb gene in schizothorax is an important system for multiple sequence alignment research. The cytochrome b gene (Cytb) is a coding gene in the mitochondrial genome of eukaryotes. Due to its single copy, moderate evolutionary rate, and abundant polymorphic sites, it is a target gene commonly used to study biological evolution and diversity. Schizothorax is a species of fish in the subfamily Schizothoracinae of the family Cyprinidae, endemic to the Qinghai-Tibet Plateau and surrounding areas. It is an important fishery resource and ecological keystone species in my country's Qinghai-Tibet Plateau. In terms of origin and evolution, schizothorax species are closely related and morphologically similar. The sequence structure of the Cytb gene contains both species-specific polymorphic sites and conserved motifs shared by schizothorax. Therefore, the Cytb gene is a commonly used target gene for distinguishing different schizothorax species and studying their phylogeny and taxonomic classification. Sequencing and sequence alignment analysis of the Cytb gene are the most important and fundamental research methods. First, PCR technology was used to amplify the Cytb gene of schizothorax fish using universal primers (L14724 and H15915) in vitro to obtain a gene sequence containing the full-length open reading frame (ORF) of the Cytb gene. Then, pyrosequencing technology was used to determine the deoxyribonucleotide composition order of the amplified product to obtain sequence information composed of different bases. Then, gene sequence software was used to compare and analyze the Cytb gene sequences of different samples to obtain genetic variation information for species identification, classification, and systematic evolution research.

[0004] Existing gene sequence alignment methods and software are commercially available, including online or PC-installed software such as CLUSTALX, MEGA, MUSCLE, and MAFFT. MEGA, which integrates multiple alignment methods such as CLUSTALX and MUSCLE, is the most commonly used sequence alignment analysis software, offering strong visualization capabilities.

[0005] Although these software programs are relatively mature, they require that the genetic sequence data they input must be in a standardized format. However, in actual work, the genetic sequences obtained by sequencing often do not meet the input requirements of these software programs. For example, the genetic sequences obtained by sequencing may have the following problems: they are positive chain sequences or antisense chain sequences, they may contain "junk" sequences after primers, adapters, and stop codons, they may include redundant sequences that are not within the comparison range, they may have defects at the 5' and 3' ends, they may contain sites with sequencing errors, etc.

[0006] Therefore, before using software to compare existing sequencing data, researchers must edit and correct the sequences based on their own research experience with related species. This significantly increases the researcher's workload and is not conducive to the efficient implementation of gene multi-sequence alignment. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention provides a gene multiple sequence alignment method, device and system, the purpose of which is to provide a gene multiple sequence alignment method including the steps of automatically editing and correcting gene sequences, thereby reducing the workload of researchers in related research work and improving the efficiency of gene multiple sequence alignment work.

[0008] A method for comparing multiple gene sequences comprises the following steps:

[0009] Step 1: setting template information, wherein the template information includes the template sequence and conserved motif information of the gene to be analyzed;

[0010] In step 2, the sequences to be analyzed are corrected and globally aligned based on the template information in step 1 to form a data set.

[0011] Preferably, in step 1, the template information is obtained through the degenerate sequence of the positive strand of the gene to be analyzed.

[0012] Preferably, in step 2, the specific steps of performing correction and global alignment of multiple sequences are as follows:

[0013] Step 2.1, using the template information in step 1 as a reference, identifying the antisense strand sequence in the sequence to be analyzed, and processing the antisense strand sequence into a sense strand sequence;

[0014] Step 2.2, using the template information from step 1 as a reference, identifying the sequencing vector sequence after the primer, adapter, and stop codon in the sequence to be analyzed after processing in step 2.1, and shearing and discarding the sequencing vector sequence;

[0015] Step 2.3, performing a global multi-sequence alignment on the sequences to be analyzed after processing in step 2.2 to form an aligned data set;

[0016] Step 2.4: Based on the template information from step 1, proofread the dataset obtained from step 2.3, fill in the missing sequences at the 5' and 3' ends, and delete the non-aligned regions.

[0017] In step 2.5, a multi-sequence global comparison is performed on the data set processed in step 2.4 to identify sequence samples containing sequencing error sites, delete sequence samples containing sequencing error sites, or adjust sequence samples containing sequencing error sites.

[0018] Preferably, in step 2.5, the method for identifying sequence samples containing sequencing error sites is: statistically analyzing the frequency of occurrence of the genotype or deletion type at each site of the sequence samples in the data set processed by step 2.4; if a certain genotype or deletion type at a certain site appears only once, then the site with the genotype or deletion type is a sequencing error site.

[0019] Preferably, in step 2, MATFF, CLUSTALX, MEGA or MUSCLE is used to perform a global multiple sequence alignment.

[0020] Preferably, the template sequence of the gene to be analyzed is an interdigitated sequence, and the interdigitated sequence is as described in SEQ ID NO.1;

[0021] There are three conserved motifs, and the sequences are AAAATTGCTAA, ATTGCCCG, and GTAATTAC;

[0022] The positions of the conserved motifs are from position 34 to position 44, from position 292 to position 299, and from position 433 to position 440 from 5'.

[0023] Preferably, the file format of the sequence to be analyzed and the data set is FASTA format.

[0024] The present invention also provides a computer device for gene multiple sequence alignment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned gene multiple sequence alignment method when executing the program.

[0025] The present invention also provides a system for gene multiple sequence alignment, comprising:

[0026] the aforementioned computer equipment;

[0027] The server is used to store and transmit the raw data of the sequence to be analyzed.

[0028] The present invention also provides a computer-readable storage medium storing a computer program for implementing the above-mentioned gene multiple sequence alignment method.

[0029] The present application provides a method for automatically correcting sequencing data and performing global multi-sequence alignment based on template information of the gene to be analyzed. In particular, in a preferred embodiment, the present invention provides template information for the Cytb gene of the schizothorax fish, which can accurately perform automatic correction and global multi-sequence alignment on the sequencing data of the schizothorax fish. The method of the present invention automatically implements the correction process that requires manual completion in the prior art through a program, greatly reducing the workload of researchers in multi-sequence alignment research and significantly improving the efficiency of related work. Therefore, the present invention has a good application prospect.

[0030] Obviously, based on the above contents of the present invention, according to common technical knowledge and customary means in this field, without departing from the above basic technical ideas of the present invention, other various forms of modifications, replacements or changes can be made.

[0031] The following further describes the above content of the present invention in detail through specific embodiments in the form of examples. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. All technologies implemented based on the above content of the present invention fall within the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a schematic diagram of the process of Example 1 of the present invention. DETAILED DESCRIPTION

[0033] It should be noted that the algorithms for data collection, transmission, storage and processing steps not specifically described in the embodiments, as well as the hardware structures, circuit connections, etc. not specifically described can all be implemented through the disclosed content of the prior art.

[0034] Example 1

[0035] This embodiment provides a system for multiple gene sequence alignment, comprising a server and a computer device. The server is used to store and transmit raw data of sequences to be analyzed, and the computer device is used for multiple gene sequence alignment and comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a method for multiple gene sequence alignment is implemented.

[0036] The method for performing multiple sequence alignment of Cytb gene of Schizothorax fish using the above system is as follows: Figure 1 The specific steps are as follows:

[0037] Step 1: Check the format of the Cytb gene sequence generated by the sequencing instrument one by one. All sample gene sequence files must be in the same folder (path). If the program detects that the sequence format is standard FASTA, the sequence information is extracted and merged into a text file generated during this step. The file name is defined as AllSeq.fas (repeatedly referred to as AllSeq.fas file). If the program detects that the sequence format is not standard FASTA, the sequence format is first converted to standard FASTA format and then merged into the AllSeq.fas file.

[0038] Step 2: The program automatically reads the template sequence and extracts conserved motifs with pre-set positions and lengths from the template sequence. The template sequence is a degenerate positive-strand sequence obtained from the cytb gene sequences of multiple schizothorax species. The program presets template sequences obtained based on previous research and analysis, and provides position and length information for identifying sequence consistency motifs (Table 1). The position and length parameters of the template sequence and motif can also be manually modified and set when necessary. The position and length parameters of the motif directly affect the normal operation of downstream analysis and the accuracy of the analysis results.

[0039] In this embodiment, the template sequence of the gene to be analyzed is an intercalated sequence, and the intercalated sequence is as described in SEQ ID NO. 1. The sequence of SEQ ID NO. 1 is as follows:

[0040]

[0041] The motif information is shown in the following table.

[0042] Table 1 Conserved motifs

[0043]

[0044] Step 3: The program automatically edits the sample sequence. The template sequence provided is used as a standard reference to detect the sample sequences in the AllSeq.fas file one by one. If the sample sequence is a positive chain sequence, no processing is performed; if the sample sequence is an antisense complementary sequence, it is automatically aligned and reversely complementary to convert it into a positive chain sequence. Next, the program automatically identifies the "junk" sequences after the primers, adapters, and stop codons of the sample sequences in the AllSeq.fas file one by one, and cuts and discards these sequences. The data processed in step 3 automatically generates a FASTA format file AllSeqEdited.fas. The first sequence in the AllSeqEdited.fas file is the template sequence, followed by the analysis sample sequence. The order of the analysis sample sequences is the same as the sample order in the AllSeq.fas file.

[0045] Step 4: The program automatically calls MATFF software to perform a global alignment of multiple sequences on AllSeqEdited.fas. If the MATFF software has been set in the environment variable, no path configuration is required. Otherwise, the absolute path of the MATFF software needs to be set in the configuration file in advance. Before executing the MATFF alignment, the program detects the number of system threads (virtual cores) and obtains the number of threads. The thread number parameter is an integer of 2 / 3 of the maximum number of threads. For example, if the maximum number of threads detected is 56, the thread number parameter is 37. The program then automatically executes MATFF software to perform a global alignment of multiple sequences on the AllSeqEdited.fas file and generates the aligned FASTA format file AllSeqEditedAligned.fas.

[0046] In addition to the MATFF software selected in this embodiment, other software capable of performing multiple sequence alignment can also be used.

[0047] Step 5: The program automatically corrects the sequence, fills in the missing sequences at the 5' and 3' ends, and deletes redundant non-aligned sequences. This step generates the AllSeqProofed.fas in FASTA format.

[0048] Step 6: The program automatically calls MATFF software to perform a global multi-sequence alignment on AllSeqProofed.fas. After the alignment is complete, the output is converted into an ordered site matrix. The frequency of the base type or deletion type at each site in each sample is examined site by site across all samples in the AllSeqProofed.fas file. If the base type or deletion type at a site in a sample appears only once across all samples, the information for that site is defined as a suspected sequencing error site and recorded in the log file. The base type with the highest frequency at that site is also recorded as the recommended correction base type for that site in the sample and also in the log file. This step ultimately generates the FASTA-formatted file AllSeqProofedAligned.fas and the log file.

[0049] The log file lists suspected sequencing error sites and suggested adjustments. This information is for reference only. You can adjust the corresponding sites in the AllSeqProofedAligned.fas file based on your specific needs. The AllSeqProofedAligned.fas file is in FASTA format and can be visualized in software such as MEGA, used for direct genetic analysis in software such as DNAsp, or converted to other formats for phylogenetic tree construction.

[0050] Using the Cytb gene sequencing data of 100 schizothorax fish as test samples, this method was used to correct and analyze the original sequenced sequences, and the results of manual correction and analysis were used as the standard for comparison. The processing and analysis process of the Cytb gene sequences of 100 schizothorax fish by this patented method took 41 seconds, while the manual processing and analysis process took 5.3 hours, significantly reducing the time consumption of this method. To verify the accuracy of this method, the manually corrected sequence of each test sample was compared with the corrected sequence obtained by this method. The resulting sequences were completely consistent, thus verifying that the accuracy of the results of this method is 100%.

[0051] As can be seen from the above examples, the present invention provides a method for multiple gene sequence alignment with an automatic correction function. Compared with the prior art, the method of the present invention omits the step of manually correcting the sequences to be analyzed, greatly reducing the workload of researchers in related fields and significantly improving the efficiency of multiple sequence alignment. Therefore, the present invention has a promising application prospect. SEQUENCE LISTING <110> Northwest Institute of Plateau Biology, Chinese Academy of Sciences <120> A gene multiple sequence alignment method, device and system <130> GY417-2021P0114522CC <160> 2 <170> PatentIn version 3.5 <210> 1 <211> 1140 <212> DNA <213> Schizothorax <400> 1 atggcaagcc tacgaaaaac tcacccccta attaaaattg ctaacagtgc actagttgac 60 ctgccagcac catccaacat ctcagcatga tgaaactttg gctctcttct aggrctatgc 120 ctagccactc aaatcctaac cggcctattc ctagccatac actatacctc agaygtttca 180 accgcattct catcagtagt ccatatttgc cgggacgtaa attacggctg actaatccgc 240 aacgtrcacg ccaacggagc atcrttcttc tttatctgya tttatataca tattgcccga 300 ggcctatact acggatctta cctctacaaa gaaacctgaa atattggtgt rgtccttctr 360 cttcttgtta tratgacggc cttcgtaggv tatgtcctgc catgrggtca aatrtctttt 420 tgaggtgcya cagtaattac maatctcctr tccgctgtrc cataygtagg tgaygttcta 480 gtccaatgga tttgaggcgg attctcagta gataaygcaa cactaacrcg attcttcgca 540 tttcactttc tatttccatt tgtaattgct gctataacca tcttrcacct cctrttttta 600 catgaaactg grtcaaayaa cccgattggs ctcaactcag acgcagataa aatccccttc 660 cacccatact ttacatataa agayttrctt ggcttcgtaa ttatactttt tttacttatg 720 cttttagcac tattttctcc gaayctgctr ggagacccag aaaacttcac ccccgccaac 780 cccctagtca caccaccaca cattaarcca gartgatatt tcctgttcgc ctatgccaty 840 ctmcgrtcya tcccaaacaa rcttggtggt gtacttgcwc tactattttc tattctrgth 900 ttaatagttg tgcctctsct tcacacctcc aagctacgag gactaacatt ccgcccaatc 960 acccaattct tattctgaac tctrgtggca gacatratta tcytracatg aattggcggc 1020 ataccagtag aacacccatt tattattatt ggacaagtcg catccgccct dtactttgca 1080 ctgtttctcr tttttatacc actagcaggg tgrgtagaaa ataaagcact ggaattagcc 1140 <210> 2 <211> 11 <212> DNA <213> Schizothorax <400> 2 aaaattgcta a 11

Claims

1. A method for multiple gene sequence alignment, It is characterized in that The steps include: Step 1, setting template information, the template information includes the template sequence of the gene to be analyzed and the information of the conserved motif; the template information is obtained by the degenerate sequence of the sense strand of the gene to be analyzed Step 2, using the template information of step 1 as a reference, calibrating the sequence to be analyzed and globally aligning multiple sequences to form a data set; the specific steps of performing calibration and globally aligning multiple sequences are as follows: Step 2.1, using the template information of step 1 as a reference, identifying the antisense strand sequence in the sequence to be analyzed, and processing the antisense strand sequence into a sense strand sequence; Step 2.2, using the template information of step 1 as a reference, identifying the sequencing vector sequence after the primer, linker and stop codon in the sequence to be analyzed after being processed in step 2.1, and cutting and discarding the sequencing vector sequence; Step 2.3, performing a global comparison of multiple sequences of the sequences to be analyzed after being processed in step 2.2 to form a compared data set; Step 2.4, using the template information of step 1 as a reference, proofread the data set obtained in step 2.3, fill in the missing sequences at the 5'- and 3'-ends, and delete the sequences in the non-aligned regions; Step 2.5, perform a global multi-sequence comparison on the data set processed in step 2.4, identify sequence samples containing sequencing error sites, delete sequence samples containing sequencing error sites, or adjust sequence samples containing sequencing error sites.

2. The gene multiple sequence alignment method according to claim 1, Features: In step 2.5, the method for identifying sequence samples containing sequencing error sites is: statistically analyzing the frequency of occurrence of the genotype or deletion type at each site of the sequence samples in the data set processed in step 2.4; if a certain genotype or deletion type at a site only appears once, then the site with the genotype or deletion type is a sequencing error site.

3. The gene multiple sequence alignment method according to claim 1, Features: In step 2, MATFF, CLUSTALX, MEGA or MUSCLE is used for global alignment of multiple sequences.

4. The gene multiple sequence alignment method according to claim 1, Features: The gene to be analyzed is the cytb gene of the schizothorax species; The template sequence of the gene to be analyzed is an intervening sequence, and the intervening sequence is as described in SEQ ID NO.1; There are three conserved motifs, and the sequences are AAAATTGCTAA, ATTGCCCG, and GTAATTAC; The positions of the conserved motifs are from position 34 to position 44, from position 292 to position 299, and from position 433 to position 440 from 5'.

5. The gene multiple sequence alignment method according to claim 1, Features: The file format of the sequence to be analyzed and the data set is FASTA format.

6. A computer device for gene multiple sequence alignment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, Features: When the processor executes the program, the gene multiple sequence alignment method according to any one of claims 1 to 5 is implemented.

7. A system for multiple gene sequence alignment, It is characterized in that include: The computer device of claim 6; The server is used to store and transmit the raw data of the sequence to be analyzed.

8. A computer-readable storage medium, Features: A computer program for implementing the gene multiple sequence alignment method according to any one of claims 1 to 5 is stored thereon.

Citation Information

Patent Citations

  • SSR typing method and system based on sequencing data

    CN110570901A

  • Method and system for analyzing and monitoring viral genome variation

    CN113593639A