Method, device, equipment and medium for evaluating genomic sequence and annotation information

By supplementing existing evaluation tools with additional evaluation criteria and evaluation methods for genome sequences and annotation information, the problem of inconsistent quality of annotation information in genome assembly data has been solved, resulting in more accurate evaluation results and higher data science value.

CN116564422BActive Publication Date: 2026-05-29BEIJING INSTITUTE OF GENOMICS CHINESE ACADEMY OF SCIENCES (CHINA NATIONAL CENTER FOR BIOINFORMATION)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INSTITUTE OF GENOMICS CHINESE ACADEMY OF SCIENCES (CHINA NATIONAL CENTER FOR BIOINFORMATION)
Filing Date
2023-05-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

The lack of high-quality genome annotation information in existing genome assembly data limits the value of data science and fails to effectively support omics analysis and research.

Method used

This paper provides a method for evaluating genome sequences and annotation information. By acquiring input files and performing an initial evaluation using an evaluation tool, and then supplementing the evaluation tool with additional evaluation criteria, the paper generates more accurate target evaluation results.

Benefits of technology

This improved the accuracy and readability of the evaluation results, enhanced the readability of error messages, and ensured that the quality of genome sequences and annotation information met the needs of scientific research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116564422B_ABST
    Figure CN116564422B_ABST
Patent Text Reader

Abstract

The application provides a method, device, equipment and medium for evaluating genome sequence and annotation information, the method comprising: obtaining an input file of a to-be-evaluated genome sequence and to-be-evaluated annotation information; performing information extraction on the input file in response to an information extraction command, obtaining different types of to-be-evaluated files from the input file, and inputting the to-be-evaluated files and a preset test file into an evaluation tool respectively to obtain a first evaluation result and a test result output by the evaluation tool; making an additional evaluation standard re-evaluate the to-be-evaluated files to obtain a second evaluation result; and generating a target evaluation result of the quality of the to-be-evaluated genome sequence and the to-be-evaluated annotation information according to the first evaluation result and the second evaluation result. The application uses an existing evaluation tool to evaluate the to-be-evaluated genome sequence and the to-be-evaluated annotation information, and at the same time, the shortcomings of the evaluation tool are made up, so that more accurate and perfect results can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene biotechnology, and more specifically, to a method, apparatus, device, and medium for evaluating genome sequence and annotation information. Background Technology

[0002] Genome sequences and genome annotations are fundamental resources for genomics-related research. Genome annotations bridge the gap between genome sequences with unknown functions and biological research on the species, and the quality of genome annotations determines the value of the genome.

[0003] With the increasing volume of genomic data each year, high-quality genomic data is an essential condition for achieving data sharing and promoting scientific research. It is a crucial prerequisite and foundation for the acquisition, dissemination, and reuse of genomic data. As the accumulated genomic data grows, rigorously controlling the quality of genome sequences and annotation information becomes a key focus of scientific research. Currently, a large amount of genome assembly data lacks genome annotation information, and the existing genome annotation information is of inconsistent quality, significantly limiting the scientific value of genome assembly data. High-quality genome annotation information can assist in omics-related analyses and research, while enhancing the value of data utilization, providing important data resources for omics analysis, comparative genomics analysis, molecular breeding, and virus tracing. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a method, apparatus, electronic device and storage medium for evaluating genome sequences and annotation information, so as to overcome the problems in the prior art.

[0005] In a first aspect, embodiments of this application provide a method for evaluating genome sequences and annotation information, the method comprising:

[0006] The input file for obtaining the genome sequence to be evaluated and the annotation information to be evaluated includes a genome sequence file, an annotation information file, and a completion information file.

[0007] The system executes an operation in response to an information extraction command, extracts information from the input file, obtains different types of files to be evaluated from the input file, and inputs the files to be evaluated and preset test files into the evaluation tool to obtain the first evaluation result and test result output by the evaluation tool.

[0008] The document to be evaluated is re-evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items to be added obtained from the analysis of the test results.

[0009] Based on the first evaluation result and the second evaluation result, a target evaluation result is generated for the quality of the genome sequence to be evaluated and the annotation information to be evaluated.

[0010] In some technical solutions of this application, the aforementioned document to be evaluated includes a nucleotide sequence file to be evaluated, and the method obtains the nucleotide sequence file to be evaluated in the following manner:

[0011] The genome sequence file in the preset format is decompressed and information is removed to obtain the initial nucleotide sequence file;

[0012] Select the file with a preset end marker from the initial nucleotide sequence file as the intermediate nucleotide sequence file;

[0013] The intermediate nucleotide sequence file is supplemented with information contained in the information file to obtain a supplemented nucleotide sequence file to be evaluated.

[0014] In some technical solutions of this application, the above method re-evaluates the nucleotide sequence file to be evaluated in the following manner:

[0015] By analyzing the test results, the additional evaluation criteria corresponding to the nucleotide sequence file to be evaluated is the sequence identification criteria defined in the definition line.

[0016] The nucleotide sequence file to be evaluated is re-evaluated using the sequence identification criteria defined in the line, resulting in a second evaluation result for the nucleotide sequence file to be evaluated.

[0017] In some technical solutions of this application, the aforementioned document to be evaluated includes a document containing annotation information to be evaluated, and the method obtains the document containing annotation information to be evaluated in the following manner:

[0018] Select the target comment information file with a preset file extension from the comment information file;

[0019] The target annotation information file is decompressed to obtain the annotation information file to be evaluated.

[0020] In some technical solutions of this application, the above method also includes:

[0021] Based on the file format of the annotation information file to be evaluated, the annotation information file to be evaluated is divided into different types; wherein, the different types include a first type and a second type;

[0022] The method re-evaluates the annotation information file to be evaluated in the following manner:

[0023] By analyzing the test results, the additional evaluation criteria for the first type of annotation information file to be evaluated are the information quantity standard and the first information content standard.

[0024] The annotation information file to be evaluated of the first type is re-evaluated using the information quantity standard and the first information content standard to obtain the second evaluation result of the annotation information file to be evaluated of the first type.

[0025] By analyzing the test results, the additional evaluation criteria for the second type of annotation information file to be evaluated are the second information content standard and the information format standard.

[0026] The second type of annotation information file to be evaluated is re-evaluated using the second information content standard and information format standard to obtain the second evaluation result of the second type of annotation information file to be evaluated.

[0027] In some technical solutions of this application, the aforementioned files to be evaluated include a template file to be evaluated, a sample metadata file to be evaluated, and a genome assembly information file to be evaluated. The method obtains the template file to be evaluated, the sample metadata file to be evaluated, and the genome assembly information file to be evaluated in the following manner:

[0028] The response information filling operation generates a corresponding filling information file based on the filling information corresponding to the information filling operation.

[0029] Extract the information from the information file and classify the information to obtain template information, sample metadata and genome assembly information;

[0030] Based on the template information, sample metadata information, and genome assembly information, and the preset format requirements, the template file to be evaluated, the sample metadata information file to be evaluated, and the genome assembly information file to be evaluated are generated.

[0031] In some technical solutions of this application, the file to be evaluated by the above method includes a nucleotide sequence file to be evaluated and an annotation information file to be evaluated. The method re-evaluates the nucleotide sequence file to be evaluated and the annotation information file to be evaluated in the following manner:

[0032] Analysis of the test results revealed that the additional evaluation criteria for the nucleotide sequence file and annotation information file to be evaluated are consistency criteria, genome contamination criteria, and genome size criteria.

[0033] The nucleotide sequence file and annotation information file to be evaluated are re-evaluated using the consistency criteria, genomic contamination criteria, and genomic size criteria to obtain a second evaluation result of the re-evaluation of the nucleotide sequence file and annotation information file to be evaluated.

[0034] Secondly, embodiments of this application provide an evaluation apparatus for genome sequence and annotation information, the apparatus comprising:

[0035] The acquisition module is used to acquire input files for the genome sequence to be evaluated and the annotation information to be evaluated; the input files include a genome sequence file, an annotation information file, and a completion information file.

[0036] The first evaluation module is used to execute operations in response to information extraction commands, extract information from the input file, obtain different types of files to be evaluated from the input file, and input the files to be evaluated and preset test files into the evaluation tool respectively to obtain the first evaluation result and test result output by the evaluation tool;

[0037] The second evaluation module is used to re-evaluate the document to be evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items to be added obtained from the analysis of the test results.

[0038] The generation module is used to generate target evaluation results for the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result.

[0039] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for evaluating genome sequence and annotation information.

[0040] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method for evaluating genome sequence and annotation information.

[0041] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0042] This application's method includes obtaining an input file containing the genome sequence to be evaluated and annotation information to be evaluated; the input file includes a genome sequence file, an annotation information file, and a completion information file; responding to an information extraction command to perform an operation, extracting information from the input file, obtaining different types of files to be evaluated from the input file, and inputting the files to be evaluated and preset test files into an evaluation tool respectively, obtaining a first evaluation result and a test result output by the evaluation tool; applying additional evaluation criteria to re-evaluate the files to be evaluated, obtaining a second evaluation result; the additional evaluation criteria are generated based on the additional items obtained from the analysis of the test results; and generating a target evaluation result for the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result. This application uses existing evaluation tools to evaluate the genome sequence to be evaluated and the annotation information to be evaluated, while also addressing the shortcomings of existing evaluation tools. The method of this application can obtain more accurate and complete target evaluation results, increasing the readability of the evaluation results and enhancing the readability of error messages.

[0043] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart illustrating an evaluation method for genome sequence and annotation information provided in an embodiment of this application is shown.

[0046] Figure 2 This illustration shows a schematic diagram of obtaining a nucleotide sequence file to be evaluated, provided by an embodiment of this application.

[0047] Figure 3 This illustration shows a schematic diagram of an evaluation device for genome sequence and annotation information provided in an embodiment of this application;

[0048] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0050] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0051] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0052] Genome sequences and genome annotations are fundamental resources for genomics-related research. Genome annotations bridge the gap between genome sequences with unknown functions and biological research on the species, and the quality of genome annotations determines the value of the genome.

[0053] With the increasing volume of genomic data each year, high-quality genomic data is an essential condition for achieving data sharing and promoting scientific research. It is a crucial prerequisite and foundation for the acquisition, dissemination, and reuse of genomic data. As the accumulated genomic data grows, rigorously controlling the quality of genome sequences and annotation information becomes a key focus of scientific research. Currently, a large amount of genome assembly data lacks genome annotation information, and the existing genome annotation information is of inconsistent quality, significantly limiting the scientific value of genome assembly data. High-quality genome annotation information can assist in omics-related analyses and research, while enhancing the value of data utilization, providing important data resources for omics analysis, comparative genomics analysis, molecular breeding, and virus tracing.

[0054] Based on this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for evaluating genome sequences and annotation information, which are described below through embodiments.

[0055] Figure 1 The diagram illustrates a flowchart of a method for evaluating genome sequence and annotation information provided in an embodiment of this application, wherein the method includes steps S101-S104; specifically:

[0056] S101. Obtain the input file for the genome sequence to be evaluated and the annotation information to be evaluated; the input file includes a genome sequence file, an annotation information file, and a completion information file;

[0057] S102. Execute the response information extraction command to extract information from the input file, obtain different types of files to be evaluated from the input file, and input the files to be evaluated and the preset test files into the evaluation tool respectively to obtain the first evaluation result and test result output by the evaluation tool.

[0058] S103. The document to be evaluated is re-evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items to be added obtained from the analysis of the test results.

[0059] S104. Generate the target evaluation result of the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result.

[0060] This application uses existing evaluation tools to evaluate the genome sequence and annotation information to be evaluated, while also addressing the shortcomings of existing evaluation tools. The method described in this application can obtain more accurate and comprehensive target evaluation results, increase the readability of the evaluation results, and enhance the readability of error messages.

[0061] The following provides a detailed description of some embodiments of this application. Unless otherwise specified, the following embodiments and features can be combined with each other. It should be noted that the methods in the embodiments of this application need to be performed under a Linux system.

[0062] S101. Obtain the input file for the genome sequence to be evaluated and the annotation information to be evaluated; the input file includes the genome sequence file, the annotation information file, and the information to be filled in.

[0063] This application, in order to evaluate the quality of the genome sequence and annotation information to be evaluated, obtains input files for the genome sequence and annotation information to be evaluated. The input files include a genome sequence file, an annotation information file, and a user-entered information file. The user-entered information file is obtained by responding to an information entry operation and generating the information file based on the entry information corresponding to the operation. The information file contains the user-entered information.

[0064] S102. Execute the response information extraction command to extract information from the input file, obtain different types of files to be evaluated from the input file, and input the files to be evaluated and the preset test files into the evaluation tool respectively to obtain the first evaluation result and test result output by the evaluation tool.

[0065] After obtaining the input file, this embodiment of the application needs to extract the file to be evaluated from the input file. There are five types of files to be evaluated in this application: nucleotide sequence file to be evaluated, annotation information file to be evaluated, template file to be evaluated, sample meta-information file to be evaluated, and genome assembly information file to be evaluated. That is, this embodiment of the application obtains five types of files to be evaluated—nucleotide sequence file to be evaluated, annotation information file to be evaluated, template file to be evaluated, sample meta-information file to be evaluated, and genome assembly information file to be evaluated—by processing three types of files: genome sequence file, annotation information file, and completion information file.

[0066] For the nucleotide sequence file to be evaluated, such as Figure 2 As shown, the embodiments of this application are obtained in the following ways:

[0067] S201. The genome sequence file in the preset format is decompressed and information is removed to obtain the initial nucleotide sequence file;

[0068] S202. Select a file with a preset end marker from the initial nucleotide sequence file as an intermediate nucleotide sequence file;

[0069] S203. Use the information contained in the information file to supplement the intermediate nucleotide sequence file to obtain the supplemented nucleotide sequence file to be evaluated.

[0070] Supplementing the intermediate nucleotide sequence file with information contained in the information file includes: adding genetic coding information contained in the information file to the definition line of the intermediate nucleotide sequence file.

[0071] In practice, for a FASTA sequence file (the nucleotide sequence file to be evaluated), the following code is executed:

[0072] 1. Decompress the .gz or .bz2-ending fasta file and extract it to the file sample.fsa. If the decompression command fails, output the error message "Error: Decompression of fasta file is invalid." to the err file.

[0073] 2. Read the FASTA file, excluding header and comment lines.

[0074] 3. For fasta files ending with '.fsa', '.fa', or '.fasta', rename them to sample.fsa.

[0075] 4. Check if the file is empty. If it is empty, output the error message "Error: fasta file is an empty file." to the err file.

[0076] 5. Read the information contained in the information file, namely the species, genetic code, topology, whether it is a plastid, whether it is a mitochondrial / chloroplast / apocrine plastid, whether it is completeness, and which chromosome it belongs to, and supplement the content of the definition line in the FASTA file.

[0077] First, define a line in the FASTA file to add species information and the corresponding genetic code for that species. If the sequence is identified as belonging to a specific chromosome, define a line in the FASTA file to add the chromosome location information and mark it with `[location=chromosome]`. If the sequence is from a mitochondria, the genetic code is changed to the chloroplast genetic code for that species, and the line is defined with `[location=mitochondrion]`. If the sequence is from a chloroplast / apocrine body, the genetic code becomes 11, and the line is defined with either `[location=chloroplast]` or `[location=apicoplast]`. If the sequence is also completeness, define a line in the FASTA file with `[completeness=complete]`. If the sequence is also circular in topology, define a line in the FASTA file with `[topology=circular]`. Simultaneously, check if the sequence has a gap region; if not, define a line in the FASTA file with `[topology=circular] gap at end, not circularized`.

[0078] The annotation information file to be evaluated is obtained in the following manner in this embodiment of the application:

[0079] From the annotation information file, select the target annotation information file with a preset file extension; decompress the target annotation information file to obtain the annotation information file to be evaluated.

[0080] In practice, for the comment information file, the specific code executes the following:

[0081] 1. Determine the file extension of the comment information file. Only files with the following extensions are accepted:

[0082] Files with extensions like .gff, .gff3, .gff.gz, .gf.bz2, .tbl, .tbl, .tbl.gz, or .tbl.bz2 are acceptable, but files ending in .tar.gz or .tar.bz2 are not. Otherwise, output the error message "Error: Allowed compressed format is .gz or .bz2, rather than packing compression in .tar.gz or .tar.bz2, etc. Please modify the compressed format. Error: The accepted annotation file is .gff or .tbl format. Your annotation file is unknown." to the err file.

[0083] 2. Decompress the comment information file ending in .gz or .bz2, and extract it to the file sample.gff or sample.tbl. If the decompression command fails, output the error message "Error: Decompression of gff / tbl file is invalid." to the err file.

[0084] The template file, sample metadata file, and genome assembly information file to be evaluated are obtained in the following ways in the embodiments of this application:

[0085] Extract the information from the information file, and classify the information to obtain template information, sample meta-information, and genome assembly information; generate the template file to be evaluated, the sample meta-information file to be evaluated, and the genome assembly information file to be evaluated based on the template information, sample meta-information, and genome assembly information and the preset format requirements.

[0086] In specific implementation, for generating the sample.sbt file (the template file to be evaluated), the code executes the following: extracting the contact person's name, institution, city, country, street, email address, and postal code filled in on the page; the data author's name, institution, city, country, street, and postal code; and the title of the published article, the author's name, journal name, year of publication, issue, volume, and page numbers, etc., and compiling them into the sample.sbt file according to the required format. For generating the sample.src file (the sample metadata file to be evaluated), the code executes the following: extracting the data metadata filled in on the page, including sample collection date, species name, country, isolate, cell type, variety, bacterial strain type, genotype, host, sex, subtype tissue type, collector, etc., and compiling them into the sample.src file in multi-column format. For generating the sample.asm file (the genome assembly information file to be evaluated), the code executes the following: extracting the assembly method and version number, assembly name, genome coverage, and sequencing technology information filled in on the page, and compiling them into the sample.asm file according to the specified format.

[0087] After obtaining the document to be evaluated, this embodiment of the application uses an evaluation tool to perform an initial evaluation of the document to be evaluated, and obtains the first evaluation result output by the evaluation tool.

[0088] The evaluation tool used here is the table2asn software developed by NCBI. It can be used to integrate FASTA sequence files and annotation information files into ASN.1 and perform quality control on genomic data. table2asn is a command-line program that converts genomic and gene sequence annotation information into ASN.1 text format files recognized by GenBank within NCBI, along with other indicative files to assist in modification. During table2asn execution, data quality control is performed. The table2asn software's quality control primarily examines the rationality, standardization, and compliance of the sequence data and gene sequence annotation information. Data that fails quality control will generate multiple different error report files. Depending on the parameters used, various optional output files can be generated, including four categories: validation files (.Val suffix), validation file summaries (.Stats suffix), difference report files (.dr suffix), and alert files (.log suffix).

[0089] In practice, not all sequencing results produce complete, continuous data. In the FASTA format, a gap is represented by a continuous sequence of "N" bases distributed between sequence regions. The interpretation of a gap, particularly the number of consecutive Ns considered as a gap, is controlled by the following parameters: -gap-type, -gaps-min, -gaps-unknown, and -l. For genome-level gap information, by default, Ns of 10 or more are considered a gap, while Ns of fewer than 10 are considered regular genomic site mutations. It also accepts the option to specify the number of consecutive Ns representing gaps of unknown length, as the Nlen and min_gap_len parameters. Therefore, the information "-gaps-unknown min_gap_len-gaps-minNlen" is added to the table2asn command line. Furthermore, for sequencing fragments on the same chromosome, if their biological order and relative orientation are known, they can be assembled into a gap sequence. It consists of islands of known sequences, separated by uncertain sequences and gaps of estimated or uncertain length. Therefore, the chain evidence type combining sequences from two gaps will also be considered. The parameter gap_linkage indicates the type of link evidence used across gaps, so the "-l gap_linkage-gap-typescaffold" information will be added to the command line.

[0090] Next, the table2asn command is executed using the previously prepared file to be evaluated:

[0091] table2asn -i sample.fsa -t sample.sbt -f sample.tbl / sample.gff -logfilesample.log -w sample.asm-src-file sample.src -M nZT-no-locus-tags-needed-gaps-unknown min_gap_len.

[0092] The error messages from the table2asn results will be displayed in the verification file, summary file, difference report file, and prompt file. This embodiment extracts the content from these files and performs a secondary interpretation of the error messages to enhance their readability. Additional quality control measures are implemented separately for any content that does not meet our quality control requirements.

[0093] S103. The document to be evaluated is re-evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items to be added obtained from the analysis of the test results.

[0094] When testing the evaluation tool with test data in this application embodiment, some defects were found in the evaluation tool, which are reflected in the test results output by the evaluation tool. Specifically, table2asn does not validate the genome size. The genome size of any species has a standardized and reasonable range, and only genome data within the specified range can pass validation. table2asn's validation of the sequence file definition lines is too broad and the validation content is not precise. For the annotation information file gff3, table2asn does not strictly enforce the requirement that the CDS length must be a multiple of 3, except for "Transl_except" or partial CDS. table2asn does not strictly validate illegal frame values ​​and strand values ​​in the annotation information file gff3. table2asn does not perform any validation or processing on contaminating sequences such as vectors, adapters, primers, and indexes present in the sequence. Therefore, the current table2asn software cannot meet the needs of high-quality evaluation of genome sequences and genome annotation information.

[0095] Through the above testing process, this application did not directly take the first evaluation result of the evaluation tool on the document to be evaluated as the final target evaluation result, but instead re-evaluated the defects of the evaluation tool and obtained a second evaluation result.

[0096] The re-evaluation process is based on the test results. By analyzing the test results, this embodiment of the application identifies the deficiencies in the evaluation tool and generates additional evaluation criteria for each deficiency. These additional evaluation criteria are then used to conduct an additional evaluation of the document to be evaluated, resulting in a second evaluation result. Since the deficiencies in the evaluation tool are multifaceted, the additional evaluation criteria generated in this embodiment of the application are also targeted.

[0097] The re-evaluation of the nucleotide sequence file to be evaluated includes:

[0098] By analyzing the test results, the additional evaluation standard corresponding to the nucleotide sequence file to be evaluated is the sequence identification standard of the defined line; the sequence identification standard of the defined line is used to re-evaluate the nucleotide sequence file to be evaluated to obtain the second evaluation result of the nucleotide sequence file to be evaluated.

[0099] In practical implementation, for the genome sequence file, the following additional evaluations were performed in this application embodiment:

[0100] (1) The sequence ID in the genome sequence file definition line will be determined in this embodiment of the application to determine whether the sequence ID contains any characters other than numbers, letters, _, -, :, *, and #. If it does, an error message will be generated: 'ERROR: Allowed characters in genome sequence ID include letters, digits, hypophens(-), underscores(_), periods(.), colons(:), asters(*), and number signs(#).'.

[0101] (2) For the genome sequence file definition line, this embodiment of the application will determine whether there is a space between the greater than sign and the sequence ID, such as '>seqID'. If there is, an error message will be generated: 'ERROR: Your submitted genome sequence file is not a valid fasta format. Fasta format file starts with '>' and sequence ID, then followed by sequence. There should be no space symbol between '>' and sequence ID.'.

[0102] (3) For the sequence ID in the definition line of the genome sequence file, this embodiment of the application will determine whether the sequence ID does not start with a letter, such as 123Contig21. If it does, an error message will be generated: 'ERROR: Genome sequence ID must start from letters.'.

[0103] The specific code executes the following:

[0104] 1. Check if the seqID in the FASTA file contains characters other than numbers, letters, underscores, hyphens, underscores, periods, colons, asterisks, and number signs. If so, send the error message "Error: Allowed characters in genome sequence ID include letters, digits, hypophens (-), underscores (_), periods (.), colons (:), asterisks (*), and number signs (#)" to the err file.

[0105] 2. Check if the seqID in the FASTA file does not start with a letter, such as 123Contig21. If it does, send an error message "Error: In Genome sequence file, the sequence ID must start from letters (e.g., >Contig21, the sequence ID is Contig21)" to the err file.

[0106] 3. Check if the definition line in the FASTA file contains content in the form of ">space seqID". If a space exists, extract the characters after the greater than sign and the space, and call this seqID. If no space exists, return the error message "Error: Your submitted genome sequence file is not a valid FASTA format. Fastaformat file starts with '>' and sequence ID, then followed by sequence. For example, '>chr1\nATCGNCGAT\nATCGNCGAT.\n One more space between '>' and 'seqid'" to the err file.

[0107] The re-evaluation of the annotation information files to be evaluated includes:

[0108] Based on the file format of the annotation information file to be evaluated, the annotation information file to be evaluated is divided into different types; wherein, the different types include a first type and a second type;

[0109] By analyzing the test results, the additional evaluation criteria for the first type of annotation information file to be evaluated are the information quantity standard and the first information content standard.

[0110] The annotation information file to be evaluated of the first type is re-evaluated using the information quantity standard and the first information content standard to obtain the second evaluation result of the annotation information file to be evaluated of the first type.

[0111] By analyzing the test results, the additional evaluation criteria for the second type of annotation information file to be evaluated are the second information content standard and the information format standard.

[0112] The second type of annotation information file to be evaluated is re-evaluated using the second information content standard and information format standard to obtain the second evaluation result of the second type of annotation information file to be evaluated.

[0113] The first type of annotation information file to be evaluated here is the GFF3 file. A GFF3 file is a nine-column, tab-separated plain text file. The first column is the sequence ID, either seqid or Sequence ID. The seqid must come from the sequence ID in the genome sequence file (FASTA file). The second column is the source of the annotation information. Many gene annotation software programs, such as EVM or AUGUSTUS, will fill this column with the annotation software name, or a dot "." can be used. The third column is the annotation feature type. If this region is a gene, it is "gene"; if it is an exon, it is "exon"; if it is a transcript, it is "transcript" or "mRNA", etc., mainly indicating the characteristics of this region. The fourth column is the start position of the feature on the sequence. Note that the gene / transcript / exon / UTR / CDS coordinates in the annotation information file should be within the sequence size range. The fifth column is the end position of the feature on the sequence. Note that the gene / transcript / exon / UTR / CDS coordinates in the annotation information file should be within the sequence size range. The sixth column is the score for that feature. For example, sequence similarity scores or reliability scores can be replaced with ".". The seventh column shows the strand information for the feature. "+" indicates a positive strand, "-" indicates a negative strand, and "." indicates uncertainty or irrelevant strand. The strand information for different features (gene / RNA / transcript / exon / CDS / UTR, etc.) within a gene should be consistent (except for trans-splicing). The eighth column shows the protein-coding shift phase information. This is generally used for CDS, with values ​​ranging from '0', '1', '2', or '.'. It indicates the shift phase of the reading frame during encoding. Except for partial genes, the phase value (codon_start) of the first CDS should be 0. The ninth column is the attribute column. The format is tag=value. Multiple tags are separated by semicolons. It may contain an ID: each independent feature requires a unique ID. The ID attribute is required for features with progeny (e.g., gene and mRNA) or features spanning multiple lines, but is optional for other features. Each feature's ID must be unique within the GFF3 file. In the case of discontinuous features (i.e., a single feature existing at multiple genomic locations), the same ID may appear on multiple lines. All lines sharing an ID must collectively represent a single feature. Name: The name of the feature. Unlike IDs, Names are not required to be unique within the file. Parents: Indicates the parent class of the feature. Parent information is used to display the hierarchy, such as "Gene>mRNA>CDS,exon,five_prime_UTR,etc".In addition to ID, Name, and Parents mentioned above, there may also be other types such as Product, Target, Note, and Ontology_term, which need to be filled in according to the specific situation.

[0114] For the sequence annotation information file gff3, this application embodiment further evaluated:

[0115] (1) The annotation information file needs to have 9 columns. If the condition is not met, an error message will be generated: 'ERROR: Gffannotation file contains and must contain nine columns. Column numbers in those features are not nine.'

[0116] (2) When the third column of the annotation information file contains features other than genes, such as mRNA, CDS, exon, etc., the parent feature ID of the feature must be specified in the ninth column, i.e., the parent information. The feature corresponding to the parent feature ID must exist. If the condition is not met, the following error messages will be generated: "ERROR: In annotation file, a RNA / transcript / exon / CDS / UTR feature must contain 'Parent' information in the ninth column. Those RNA / transcript / exon / CDS / UTR need additional 'Parent' information." and "ERROR: In annotation file, a parent feature from RNA / transcript / exon / CDS / UTR feature must exist. Those 'Parent' features can't be found in gff file."

[0117] (3) Multiple labels in the ninth column of the annotation information file need to be separated by semicolons, not other symbols. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the delimiter should be in the ninth column."

[0118] (4) In column 9 of the annotation information file, if a parent feature exists, the IDs of the parent feature and the feature in this row should not be the same. For example, ID = AS1; Parent = AS1. If the condition is not met, an error message will be generated: "ERROR: In an annotation file, a RNA / transcript / exon / CDS / UTR feature ID cannot be equal to their parent ID. Those RNA / transcript / exon / CDS / UTR feature IDs are equal to their parent IDs."

[0119] (5) In column 9 of the annotation information file, the ID attribute is required for features that have children, such as genes and mRNAs. If the condition is not met, an error message will be generated: "ERROR: The ID attribute is required for features that have children (e.g., genes and mRNAs)."

[0120] (6) In the 8th column of the annotation information file, the valid frame values ​​are 0, 1, and 2. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the 8th column is frame value of CDS. The valid frame value is 0, 1, 2."

[0121] (7) In column 7 of the annotation information file, valid strand values ​​are + or -. Except for trans-splicing, the strand information within a gene should be consistent, such as gene / RNA / transcript / exon / CDS / UTR. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the strand information should be + or -. Those lines feature strand's information is illegal."

[0122] (8) If it is a complete CDS, the frame value of the first CDS interval must be 0. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the codon_start or frame value for firstCDS should be 0 except a partial gene."

[0123] (9) Except for Transl_except and incomplete CDS, the sum of the lengths of all CDS in a transcript must be a multiple of 3. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the CDS length of a transcript should be a multiple of 3. Those CDS features should be fixed for no-triple of CDS length."

[0124] (10) If it is a complete CDS, the stop codon must be a stop codon in the codon table corresponding to that species. In this embodiment, the appropriate codon table will be selected based on the sequence species information and component information (e.g., whether it is chloroplast or mitochondria) to determine whether the stop codon of the CDS is consistent with the codon table. If the condition is not met, an error message will be generated: "ERROR: The following CDS sequence features contain illegal stopcodon."

[0125] The specific code executes the following:

[0126] 1. The annotation information file needs to have 9 columns. If this condition is not met, the following error message will be generated: "ERROR: Gffannotation file contains and must contain nine columns. Column numbers in those features are not nine."

[0127] 2. When column 3 of the annotation information file contains features other than genes, such as mRNA, CDS, exon, etc., column 9 must specify the parent feature ID, i.e., the parent information. Furthermore, the feature corresponding to the parent feature ID must exist. If this condition is not met, the following error messages will be generated: "ERROR: In annotation file, a RNA / transcript / exon / CDS / UTR feature must contain 'Parent' information in the ninth column. Those RNA / transcript / exon / CDS / UTR need additional 'Parent' information." and "ERROR: In annotation file, a parent feature from RNA / transcript / exon / CDS / UTR feature must exist. Those 'Parent' features can't be found in gff file."

[0128] 3. Multiple labels in column 9 of the annotation file must be separated by semicolons, not other symbols. If this condition is not met, an error message will be generated: "ERROR: In annotation file, the delimiter should be in the first column."

[0129] 4. In column 9 of the comment information file, if a parent feature exists, the IDs of the parent feature and the feature in this row should not be the same. For example, ID = AS1; Parent = AS1. An error message will be generated if this condition is not met.

[0130] "ERROR:In annotation file,a RNA / transcript / exon / CDS / UTR feature IDcan not be equal to their parent ID.Those RNA / transcript / exon / CDS / UTR featureIDs are equal to their parent IDs.".

[0131] 5. In column 9 of the annotation information file, the ID attribute is required for features that have children, such as genes and mRNAs. If this condition is not met, an error message will be generated: "ERROR: The ID attribute is required for features that have children (e.g., genes and mRNAs)."

[0132] 6. In the 8th column of the annotation information file, the valid frame values ​​are 0, 1, and 2. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the 8th column is frame value of CDS. The valid frame value is 0, 1, 2."

[0133] 7. In column 7 of the annotation information file, valid strand values ​​are + or -. Except for trans-splicing, the strand information within a gene should be consistent, such as gene / RNA / transcript / exon / CDS / UTR. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the strand information should be + or -. Those lines feature strand's information is illegal."

[0134] 8. If it is a complete CDS, the frame value of the first CDS interval must be 0. If the condition is not met, an error message will be generated: "ERROR: In annotation file, the codon_start or frame value for first CDS should be 0 except a partial gene."

[0135] 9. Except for Transl_except and incomplete CDS cases, the sum of the lengths of all CDS in a transcript must be a multiple of 3. If this condition is not met, the following error message will be generated: "ERROR: In annotation file, the CDS length of a transcript should be a multiple of 3. Those CDS features should be fixed for no-triple of CDS length."

[0136] 10. If it is a complete CDS, the stop codon must be a stop codon within the codon table corresponding to that species. In this embodiment, the appropriate codon table is selected based on the sequence species information and component information (e.g., whether it is a chloroplast or mitochondria), and it is determined whether the stop codon of the CDS matches the codon table. If the condition is not met, an error message is generated: "ERROR: The following CDS sequence features contain illegal stopcodon."

[0137] The second type of annotation information file to be evaluated here is the sequence annotation information file (tbl). The sequence annotation information file (tbl) has 5 columns, each separated by a tab. This file must contain: structural annotation information for coding genes, structural annotation information for non-coding genes, and functional annotation information for genes. Each feature is described using 5 columns, divided into two parts.

[0138] Part 1: This section contains the structural information of the feature within the sequence. It has three columns: the start site, the end site, and the feature name. If the feature is on the sense strand, the start site is smaller than the end site; if it's on the negative strand, the start site is larger than the end site. If the feature is a CDS or exon of an incomplete gene, there will be multiple rows of data, but the feature name will only be displayed in the third column of the first row.

[0139] Part 2 contains feature functional annotation information. Columns 4 and 5 are used, preceded by three tabs. Column 4 corresponds to the feature's qualifier, and column 5 is the qualifier's value. The qualifier is a descriptive label for the feature. If there are multiple qualifiers and their values, they are represented across multiple rows.

[0140] Before all annotations for each sequence, there is a line describing the sequence ID of the genome sequence, beginning with ">Feature", for example: >Feature scaffold_1. All annotations following this line belong to the sequence scaffold_1, and "Feature" must not be omitted. Feature and scaffold_1 are separated by a space.

[0141] For the sequence annotation information file tbl, the following aspects were additionally evaluated in this application embodiment:

[0142] 1. Check if the ">Feature" field in the .tbl file exists. If it does not exist, report the error "Error: seqid name is wrong, '>Feature SeqId table_name' is standard form. These sequence identifiers (SeqId) must be the same as that used on the sequence. The table_name is optional. Please check and fix them." to the err file.

[0143] The Qualifier key information in the 2.tbl file has a selectable range. If a feature outside this range is found, an error message will be sent to the err file: "Error: The feature line contains 3 columns. The qualifier line contains 5 columns. The first three columns of the qualifier line are blank tab-separated columns. Invalid feature information in tbl file. Please check and fix them." Currently, the Qualifier key's value range is:

[0144] 'gene','Gene','CDS','cds','mRNA','exon','five_prime_UTR','t hree_prime_UTR','rRNA','tRNA','ncRNA','tmRNA','transcript','mob ile_genetic_element','origin_of_replication','promoter','repeat_region','intron','protein',"five_prime_utr","three_prime_utr","tig","supercontig","chromosome","landmark","start_codon","stop_codon".

[0145] 3. In the .tbl file, the separator between the Qualifier key and the qualifier value should be a tab, not a space. If a space is used as the separator, the error message "Error: The feature line contains 3 columns. The qualifier line contains 5 columns. The first three columns of the qualifier line are blank tab-separated columns. Missing column or missing tab in .tblfile. Please check and fix them" will be sent to the err file.

[0146] If the second part of the 4.tbl file does not start with three tabs but with multiple spaces, the error message "Error: The feature line contains 3 columns. The qualifier line contains 5 columns. The first three columns of the qualifier line are blank tab-separated columns. Missing tab in tbl file. Please check and fix them" will be sent to the err file.

[0147] In the .tbl file, the codon value should be 1 / 2 / 3. If it is outside this range, the error message "Error: In annotation file (.tbl format), the valid codon value is 1,2,3. Those lines codon values ​​are invalid" will be sent to the err file.

[0148] The nucleotide sequence files and annotation information files to be evaluated were re-evaluated, including:

[0149] Analysis of the test results revealed that the additional evaluation criteria for the nucleotide sequence file and annotation information file to be evaluated are consistency criteria, genome contamination criteria, and genome size criteria.

[0150] The nucleotide sequence file and annotation information file to be evaluated are re-evaluated using the consistency criteria, genomic contamination criteria, and genomic size criteria to obtain a second evaluation result of the re-evaluation of the nucleotide sequence file and annotation information file to be evaluated.

[0151] In practice, the internal consistency between the genome sequence and annotation is checked. It is required that the sequence ID in the genome annotation matches the genome sequence ID, and a feature coordinate must be within the corresponding genome sequence range. If these conditions are not met, an error message is generated: "ERROR: In annotation file, sequence must be from genome sequence. Those sequences in annotation file are extra that cannot find match in genome file:", "ERROR: A sequence in assignment should be from genome sequence. Those sequences cannot find match in genome sequence:", or "ERROR: Those sequences in genome file are extra that cannot find match in tbl file:".

[0152] Next, the genome sequence will be checked for contamination. Using data collected from the UniVec database, fragments potentially originating from vector sources (vector contamination), cloned cDNA, or adapter and primer sequences commonly used in genome assembly will be identified, thus detecting sequence contamination. UniVec is a dedicated, non-redundant vector database, with data obtained from the NCBI FTP directory (ftp: / / ftp.ncbi.nlm.nih.gov / pub / UniVec). BLAST analysis with preset parameters will be performed on the vector, adapter, primer, and index sequences obtained from the UniVec database to optimally detect sequence contamination. The database will query the genome sequence to see if any sequence fragments match those in the UniVec database, determining if the genome sequence is contaminated with vectors, adapters, primers, or indexes. If contaminating sequences such as vectors, adapters, primers, or indexes are found in the sequence file, they need to be corrected, generating the error message: "We ran the sequences through the contamination screen. The screen found the following sequences might be from contamination that need to be excluded or trimmed. Please make sure they are corrected and adjust the sequence file, then resubmit it."

[0153] Finally, the genome size of the sequence file is checked and compared with the expected genome size range for that species. If the sequence is a whole-genome assembly, the genome size is within a reasonable range, and only sequences within this range are accepted. Specifically, first, confidence intervals for different species' genomes are established based on the expected genome size files provided on the NCBI website. The sequence file's length (excluding gaps and 10 Ns) is then used to determine if it is a genome sequence file for the specified species, based on the expected genome size confidence intervals for different species and the length of the sequence file excluding gaps (i.e., gaps and 10 Ns are ignored). Generally, acceptable genome size ranges are as follows: Archaea (100,000bp-15,000,000bp), Bacteria (100,000bp-15,000,000bp), Eukaryotes (100,000bp to unlimited, practically 100Gbases), Viruses (100bp-15,000,000bp), Metagenomics or Unclassified Organisms (at least 200bp, i.e., the minimum length of a whole-genome shotgun sequence in the International Nucleotide Sequence Database Collaboration).

[0154] S104. Generate the target evaluation result of the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result.

[0155] After obtaining the first evaluation result and the second evaluation result through the above method, the embodiments of this application combine the first evaluation result and the second evaluation result to generate the target evaluation result of the quality of the genome sequence to be evaluated and the annotation information to be evaluated.

[0156] In this embodiment, as an optional implementation, the genome sequence and genome annotation information of *Atropa belladonna* are used as input data, and the present method is used for quality control evaluation. First, the table2asn software is used to perform data quality control, and then the aforementioned additional evaluation criteria are used to detect the legality of the genome sequence content, the legality of the genome annotation content, the genome sequence size, and the contamination status of the genome sequence for further quality control analysis. The server used is CentOS Linux release 7.4.1708.

[0157] The results showed that the table2asn software produced a total of four error messages: “NO_ANNOTATION:501 bioseqs have no features.”, “FATAL:MISSING_GENES:140418 features have no genes.”, “Warning:valid[SEQ_FEAT.NotSpliceConsensusDonor]”, and “Error:valid[SEQ_FEAT.ShortIntron]”. To improve readability, we have revised the above text to: “Error: Sequence ID is not valid, or Sequence ID in genome sequence file does not exist in genome annotation file. NO_ANNOTATION: 501 bioseqs have no features.”, “FATAL: MISSING_GENES: 140418 features have no genes.”, “Warning: Splice junctions typically have GT as the first two bases of the intron (splice donor) and AG as the last two bases of the intron (splice acceptor). This intron does not conform to that pattern.”, “Error: Introns should belonger than 10nt.”

[0158] Using the aforementioned additional evaluation criteria, three more error messages were obtained: "ERROR: A sequencein assignment should be from genome sequence. Those sequences cannot find match in genome sequence:", "ERROR: In annotation file, the codon_start or framevalue for first CDS should be 0 except a partial gene. Those first CDSs' codon_start values ​​are invalid:", and "ERROR: Those sequences in genome file are extra that cannot find match in gff file:".

[0159] In summary, our method enables comprehensive quality control of genome sequences and annotation files, establishes a robust quality control analysis workflow, and enhances the readability of quality control results.

[0160] Figure 3 This invention provides a schematic diagram of the structure of an evaluation device for genome sequence and annotation information, comprising:

[0161] The acquisition module is used to acquire input files for the genome sequence to be evaluated and the annotation information to be evaluated; the input files include a genome sequence file, an annotation information file, and a completion information file.

[0162] The first evaluation module is used to execute operations in response to information extraction commands, extract information from the input file, obtain different types of files to be evaluated from the input file, and input the files to be evaluated and preset test files into the evaluation tool respectively to obtain the first evaluation result and test result output by the evaluation tool;

[0163] The second evaluation module is used to re-evaluate the document to be evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items to be added obtained from the analysis of the test results.

[0164] The generation module is used to generate target evaluation results for the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result.

[0165] The file to be evaluated includes a nucleotide sequence file to be evaluated, and the device obtains the nucleotide sequence file to be evaluated in the following manner:

[0166] The genome sequence file in the preset format is decompressed and information is removed to obtain the initial nucleotide sequence file;

[0167] Select the file with a preset end marker from the initial nucleotide sequence file as the intermediate nucleotide sequence file;

[0168] The intermediate nucleotide sequence file is supplemented with information contained in the information file to obtain a supplemented nucleotide sequence file to be evaluated.

[0169] The device re-evaluates the nucleotide sequence file to be evaluated in the following manner:

[0170] By analyzing the test results, the additional evaluation criteria corresponding to the nucleotide sequence file to be evaluated is the sequence identification criteria defined in the definition line.

[0171] The nucleotide sequence file to be evaluated is re-evaluated using the sequence identification criteria defined in the line, resulting in a second evaluation result for the nucleotide sequence file to be evaluated.

[0172] The document to be evaluated includes a file of annotation information to be evaluated, and the device obtains the file of annotation information to be evaluated in the following manner:

[0173] Select the target comment information file with a preset file extension from the comment information file;

[0174] The target annotation information file is decompressed to obtain the annotation information file to be evaluated.

[0175] The device further includes:

[0176] The differentiation module is used to classify the annotation information file to be evaluated into different types according to the file format of the annotation information file to be evaluated; wherein, the different types include a first type and a second type;

[0177] The device re-evaluates the annotation information file to be evaluated in the following manner:

[0178] By analyzing the test results, the additional evaluation criteria for the first type of annotation information file to be evaluated are the information quantity standard and the first information content standard.

[0179] The annotation information file to be evaluated of the first type is re-evaluated using the information quantity standard and the first information content standard to obtain the second evaluation result of the annotation information file to be evaluated of the first type.

[0180] By analyzing the test results, the additional evaluation criteria for the second type of annotation information file to be evaluated are the second information content standard and the information format standard.

[0181] The second type of annotation information file to be evaluated is re-evaluated using the second information content standard and information format standard to obtain the second evaluation result of the second type of annotation information file to be evaluated.

[0182] The file to be evaluated includes a template file to be evaluated, a sample metadata file to be evaluated, and a genome assembly information file to be evaluated. The device obtains the template file to be evaluated, the sample metadata file to be evaluated, and the genome assembly information file to be evaluated in the following manner:

[0183] The response information filling operation generates a corresponding filling information file based on the filling information corresponding to the information filling operation.

[0184] Extract the information from the information file and classify the information to obtain template information, sample metadata and genome assembly information;

[0185] Based on the template information, sample metadata information, and genome assembly information, and the preset format requirements, the template file to be evaluated, the sample metadata information file to be evaluated, and the genome assembly information file to be evaluated are generated.

[0186] The device re-evaluates the nucleotide sequence file to be evaluated and the annotation information file to be evaluated in the following manner:

[0187] Analysis of the test results revealed that the additional evaluation criteria for the nucleotide sequence file and annotation information file to be evaluated are consistency criteria, genome contamination criteria, and genome size criteria.

[0188] The nucleotide sequence file and annotation information file to be evaluated are re-evaluated using the consistency criteria, genomic contamination criteria, and genomic size criteria to obtain a second evaluation result of the re-evaluation of the nucleotide sequence file and annotation information file to be evaluated.

[0189] like Figure 4As shown, this application provides an electronic device for executing the evaluation method of genome sequence and annotation information in this application. The device includes a memory, a processor, a bus, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the evaluation method of genome sequence and annotation information.

[0190] Specifically, the aforementioned memory and processor can be general-purpose memory and processor, without any specific limitations. When the processor runs the computer program stored in the memory, it can execute the aforementioned evaluation methods for genome sequence and annotation information.

[0191] Corresponding to the evaluation method of genome sequence and annotation information in this application, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when run by a processor, executes the steps of the above-described evaluation method of genome sequence and annotation information.

[0192] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the aforementioned evaluation methods for genome sequence and annotation information.

[0193] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0194] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0195] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0196] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0197] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0198] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for evaluating genome sequence and annotation information, characterized in that, The method includes: The input file for obtaining the genome sequence to be evaluated and the annotation information to be evaluated includes a genome sequence file, an annotation information file, and a completion information file. The system executes an operation in response to an information extraction command, extracts information from the input file, obtains different types of files to be evaluated from the input file, and inputs the files to be evaluated and a preset test file into the evaluation tool to obtain the first evaluation result and the test result output by the evaluation tool; wherein, the evaluation tool is table2asn; the files to be evaluated include nucleotide sequence files and annotation information files; The file to be evaluated is re-evaluated using additional evaluation criteria to obtain a second evaluation result; the additional evaluation criteria are generated based on the additional items obtained from the analysis of the test results; wherein, the re-evaluation includes sequence identification of nucleotide sequence files, format of GFF3 files, format of tbl files, consistency evaluation, contamination evaluation, and genome size evaluation; Based on the first evaluation result and the second evaluation result, a target evaluation result is generated for the quality of the genome sequence to be evaluated and the annotation information to be evaluated.

2. The method according to claim 1, characterized in that, The file to be evaluated includes a nucleotide sequence file to be evaluated, and the method obtains the nucleotide sequence file to be evaluated in the following manner: The genome sequence file in the preset format is decompressed and information is removed to obtain the initial nucleotide sequence file; Select the file with a preset end marker from the initial nucleotide sequence file as the intermediate nucleotide sequence file; The intermediate nucleotide sequence file is supplemented with information contained in the information file to obtain a supplemented nucleotide sequence file to be evaluated.

3. The method according to claim 2, characterized in that, The method re-evaluates the nucleotide sequence file to be evaluated in the following manner: By analyzing the test results, the additional evaluation criteria corresponding to the nucleotide sequence file to be evaluated is the sequence identification criteria defined in the definition line. The nucleotide sequence file to be evaluated is re-evaluated using the sequence identification criteria defined in the line, resulting in a second evaluation result for the nucleotide sequence file to be evaluated.

4. The method according to claim 1, characterized in that, The file to be evaluated includes a file of annotation information to be evaluated, and the method obtains the file of annotation information to be evaluated in the following manner: Select the target comment information file with a preset file extension from the comment information file; The target annotation information file is decompressed to obtain the annotation information file to be evaluated.

5. The method according to claim 4, characterized in that, The method further includes: Based on the file format of the annotation information file to be evaluated, the annotation information file to be evaluated is divided into different types; wherein, the different types include the first type GFF3 format file and the second type tbl format file; The method re-evaluates the annotation information file to be evaluated in the following manner: By analyzing the test results, the additional evaluation criteria for the first type of annotation information file to be evaluated are the information quantity standard and the first information content standard. The first information content standard includes: checking whether the file has 9 columns, checking whether non-gene features contain parent information, checking whether the parent feature exists, checking whether the attribute column separator is a semicolon, checking whether the feature ID is different from the parent feature ID, checking whether features with child features have an ID attribute, checking whether the frame value of the 8th column is 0, 1 or 2, checking whether the chain value of the 7th column is + or -, checking whether the first frame value of the complete CDS is 0, checking whether the total length of the CDS is a multiple of 3, and checking whether the stop codon conforms to the species codon table. The annotation information file to be evaluated of the first type is re-evaluated using the information quantity standard and the first information content standard to obtain the second evaluation result of the annotation information file to be evaluated of the first type. By analyzing the test results, the additional evaluation criteria for the second type of annotation information file to be evaluated are the second information content standard and the information format standard. The second information content standard includes: the information format standard includes checking whether the file starts with >Feature, checking whether the Qualifier key is within a preset range, checking whether there is a tab between the Qualifier key and the value, checking whether there are three tab characters before the function annotation information part, and checking whether the password value is 1, 2 or 3. The second type of annotation information file to be evaluated is re-evaluated using the second information content standard and information format standard to obtain the second evaluation result of the second type of annotation information file to be evaluated.

6. The method according to claim 1, characterized in that, The files to be evaluated include a template file to be evaluated, a sample metadata file to be evaluated, and a genome assembly information file to be evaluated. The method obtains the template file to be evaluated, the sample metadata file to be evaluated, and the genome assembly information file to be evaluated in the following manner: The response information filling operation generates a corresponding filling information file based on the filling information corresponding to the information filling operation. Extract the information from the information file and classify the information to obtain template information, sample metadata and genome assembly information; Based on the template information, sample metadata information, and genome assembly information, and the preset format requirements, the template file to be evaluated, the sample metadata information file to be evaluated, and the genome assembly information file to be evaluated are generated.

7. The method according to claim 1, characterized in that, The file to be evaluated includes a nucleotide sequence file to be evaluated and an annotation information file to be evaluated. The method re-evaluates the nucleotide sequence file to be evaluated and the annotation information file to be evaluated in the following manner: Analysis of the test results revealed that the additional evaluation criteria for the nucleotide sequence file and annotation information file to be evaluated are consistency criteria, genome contamination criteria, and genome size criteria. The nucleotide sequence file and annotation information file to be evaluated are re-evaluated using the consistency criteria, genomic contamination criteria, and genomic size criteria to obtain a second evaluation result of the re-evaluation of the nucleotide sequence file and annotation information file to be evaluated.

8. An evaluation device for genome sequence and annotation information, characterized in that, The device includes: The acquisition module is used to acquire input files for the genome sequence to be evaluated and the annotation information to be evaluated; the input files include a genome sequence file, an annotation information file, and a completion information file. The first evaluation module is used to execute operations in response to information extraction commands, extract information from the input file, obtain different types of files to be evaluated from the input file, and input the files to be evaluated and preset test files into the evaluation tool respectively to obtain the first evaluation result and test result output by the evaluation tool; wherein, the evaluation tool is table2asn; the files to be evaluated include nucleotide sequence files and annotation information files; The second evaluation module is used to re-evaluate the file to be evaluated using additional evaluation criteria to obtain a second evaluation result. The additional evaluation criteria are generated based on the additional items obtained from the analysis of the test results. The re-evaluation includes sequence identification of nucleotide sequence files, format of GFF3 files, format of tbl files, consistency evaluation, contamination evaluation, and genome size evaluation. The generation module is used to generate target evaluation results for the quality of the genome sequence to be evaluated and the annotation information to be evaluated based on the first evaluation result and the second evaluation result.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and the processor communicates with the memory via the bus when the electronic device is in operation, and the machine-readable instructions, when executed by the processor, perform the steps of the evaluation method for genomic sequence and annotation information as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the evaluation method for genome sequence and annotation information as described in any one of claims 1 to 7.