Biological information analysis method for rapid typing of influenza A virus subtypes

By constructing a reference database using bioinformatics analysis methods and combining sequence alignment and judgment rules, rapid and accurate subtyping of influenza A virus under conditions of low sequencing depth and partial sequence deletion was achieved. This solves the problems of inefficiency and high cost of existing technologies and is applicable to the subtyping analysis of a variety of viruses.

CN121237222APending Publication Date: 2025-12-30NANJING NOIN MEDICAL TESTING LABORATORY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511807296.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies are insufficient for rapidly and accurately identifying influenza A virus subtypes under conditions of low sequencing depth and partial sequence deletion. Traditional methods are prone to missed detections and misjudgments, and are costly, making it difficult to meet the needs of rapid response and grassroots testing.

Method used

Using bioinformatics analysis methods, rapid genotyping is achieved by constructing a reference database, preprocessing sequencing data, performing sequence alignment and determination rules, and combining single sequence tendency determination with sample proportion statistics.

Benefits of technology

Achieving high-precision typing under low coverage and low sequencing depth conditions reduces costs and avoids misclassification problems of traditional methods, making it suitable for clinical pathogen detection and epidemiological surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237222A_ABST
    Figure CN121237222A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of bioinformatics, and provides a biological information analysis method for rapid typing of influenza A virus subtypes, which comprises database construction, sequencing data preprocessing, sequence comparison, sequence level subtype determination, sample level dominant subtype determination and typing result output. The method does not need to depend on full-length or high-depth sequencing data, subtype judgment can be completed only through a short sequence, and the sequencing cost is remarkably reduced; relying on a differential threshold mechanism, a good judgment effect is still kept under the conditions of low coverage and low sequencing depth; according to the method, a double-layer typing mechanism combining single-sequence tendency judgment and sample proportion statistics is adopted, typing stability and robustness are remarkably improved, calculation is completely dependent on comparison information, the method is not influenced by PCR primer mutation, and the problem of wrong typing in a traditional method is effectively avoided. The method is suitable for clinical pathogen detection, subtype determination and epidemiological monitoring, and can be expanded to typing analysis of other high-variability viruses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, and in particular relates to a bioinformatics analysis method for rapid subtyping of influenza A virus subtypes. Background Technology

[0002] Influenza A virus is a highly variable negative-sense RNA virus. Subtypes of its hemagglutinin (HA) and neuraminidase (NA) genes combine to form various viral subtypes, such as H1N1, H3N2, H5N1, and H7N9. Accurate subtyping is crucial for the prevention and treatment of influenza A: in clinical practice, different subtypes exhibit significant differences in pathogenicity, transmissibility, and sensitivity to antiviral drugs such as oseltamivir. Precise subtyping provides clinicians with clear diagnostic evidence, guides targeted medication, and reduces the incidence of severe cases. In epidemiological surveillance, subtyping can track viral mutation trends, transmission chains, and epidemiological characteristics, providing core data support for public health early warning and prevention and control strategies. Therefore, developing efficient and accurate influenza A virus subtyping technology is of great significance for clinical diagnosis, drug selection, and epidemiological surveillance.

[0003] Currently, methods for subtyping influenza A virus subtypes mainly include PCR primer-based detection methods and whole-genome alignment analysis methods (relying on next-generation sequencing (NGS) technology). The former relies on the matching degree between primers and target sequences, making it prone to missed detections and misclassifications in highly variable strains, and also difficult to handle unknown novel subtypes. The latter has stringent requirements for sample quality and sequencing depth, and is unsuitable for samples with low viral loads, nucleic acid degradation, or partial sequence deletions. Furthermore, it has long detection cycles, consumes significant computational resources, and is costly, making it difficult to meet the needs of rapid response and grassroots testing. Therefore, there is an urgent need to develop an analytical method that can rapidly and accurately identify influenza A virus subtypes even under conditions of low sequencing depth and partial sequence deletions. To this end, this invention proposes a bioinformatics analysis method for rapid subtyping of influenza A virus subtypes. Summary of the Invention

[0004] The purpose of this invention is to provide a bioinformatics analysis method for rapid subtyping of influenza A virus subtypes, aiming to solve the problems mentioned in the background art.

[0005] The objective of this invention is achieved through the following technical solution: A bioinformatics analysis method for rapid subtyping of influenza A virus subtypes includes the following steps: Step 1: Database construction; Establish a reference database that includes multiple subtypes of influenza A virus; Step 2: Sequencing data preprocessing; The raw sequencing data underwent quality control, adapter removal, and length screening, and human sequences were filtered to obtain high-quality sequences. Step 3: Sequence alignment; The preprocessed high-quality sequences are compared with the subtype sequences in the reference database to generate a comparison result file, and the comparison score of each sequence in different subtypes is extracted. Step 4: Sequence-level subtype determination; Each sequence is assigned a subtype according to the preset determination rules: Condition 1: If a sequence scores full marks when aligned to two or more different subtypes, then the sequence cannot be distinguished as a subtype and is marked as an "untyped sequence". Condition 2: If a sequence scores full marks only in one subtype, while failing to score full marks in other subtypes, then the sequence is determined to belong to that subtype; Condition 3: If no subtype has a perfect score, and the highest score of a certain subtype is higher than that of other subtypes by at least a set threshold, then the sequence is determined to be inclined to that subtype; if no subtype has a perfect score and the difference in the highest scores of each subtype is less than the set threshold, then it is marked as an "untyped sequence". Step 5: Determine the dominant subtype at the sample level; The subtype distribution of all separable sequences in the statistical sample is analyzed, the sequence proportion of each subtype is calculated, and the dominant subtype of the sample is determined. Step 6: Output the fracturing results; The output includes a typing report containing basic sample information, dominant subtype, minor subtype, the proportion of each subtype sequence, and the total number of typifiable sequences.

[0006] Furthermore, in step 1, the reference database includes characteristic fragment sequences of the hemagglutinin and neuraminidase genes of influenza A virus subtypes.

[0007] Furthermore, in step 2, Fastp software is used for quality control, connector removal, and length screening, and BWA software is used to filter human-derived sequences.

[0008] Furthermore, in step 3, BWA software is used for sequence alignment. The alignment result is in SAM file format, and the maximum alignment score is the sequencing sequence length.

[0009] Furthermore, in step 4, the initial value of the threshold is set to 5 points, which is determined based on 50bp-75bp sequencing data.

[0010] Furthermore, in step 5, if the proportion of a certain subtype sequence to all separable sequences exceeds 50%, then the subtype is determined to be the dominant subtype of the sample.

[0011] Furthermore, in step 6, the classification report supports exporting in Excel or PDF format.

[0012] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the bioinformatics analysis method for rapid subtyping of influenza A virus subtypes as described above.

[0013] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the bioinformatics analysis method for rapid subtyping of influenza A virus subtypes as described above.

[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. High-precision and rapid genotyping under low conditions: This invention does not rely on full-length or high-depth sequencing data. It can complete the subtype determination using only short sequences (such as 50bp), which greatly reduces the sequencing cost. Moreover, relying on the differential threshold mechanism of ≥5 points, it can still maintain good determination results under low coverage and low sequencing depth conditions.

[0015] 2. Technological breakthroughs: This invention adopts a two-level typing mechanism (sequence level + sample level) that combines single sequence tendency determination with sample proportion statistics, which significantly improves typing stability and robustness. At the same time, it relies entirely on alignment information for calculation, is independent of primer and probe design, and is not affected by PCR primer mutations, effectively avoiding the misclassification problem of traditional methods.

[0016] 3. Wide range of application scenarios: This invention is applicable to clinical pathogen detection, subtype determination and epidemiological monitoring, which can meet the needs of practical applications and has strong scalability. It can be extended to the typing analysis of other highly variable viruses such as coronavirus and respiratory syncytial virus. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0018] In order to provide a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention will now be described in detail below, but it should not be construed as limiting the scope of implementation of the present invention.

[0019] This invention provides a bioinformatics analysis method for rapid subtyping of influenza A virus subtypes. The flowchart of this method is shown below. Figure 1 This includes the following steps: Step 1: Database construction; Establish a reference database of characteristic fragment sequences of HA and NA genes from multiple influenza A virus subtypes (such as H1N1, H3N2, H5N1, etc.), screen and annotate characteristic regions of HA and NA genes, and dynamically update the database according to the influenza epidemic subtypes for subsequent comparative analysis.

[0020] Step 2: Sequencing data preprocessing; Fastp software was used to perform quality control, adapter removal, and length screening on the raw sequencing data; then BWA software was used to filter human sequences, retaining high-quality sequences for analysis.

[0021] Step 3: Sequence alignment; The BWA software was used to align the preprocessed high-quality sequences (e.g., 50bp) with the sequences of each subtype in the constructed reference database. The alignment results were in SAM file format, and the alignment score of each sequence in different subtypes was output. The perfect alignment score (i.e., the full score) was equal to the length of the sequencing sequence.

[0022] Step 4: Sequence-level subtype determination; Each sequence is assigned a subtype according to the preset determination rules: Condition 1: If a sequence scores full marks (e.g., 50 points) when aligned to two or more different subtypes, then the sequence cannot be distinguished as a subtype and is marked as an "untyped sequence". Condition 2: If a sequence scores full marks only in one subtype, while failing to score full marks in other subtypes, then the sequence is determined to belong to that subtype; Condition 3: If no subtype alignment score is full, and the highest alignment score of a certain subtype is higher than that of other subtypes by at least a set threshold (this threshold needs to be optimized and adjusted based on historical sample data, and the initial value is determined based on 50bp-75bp sequencing data testing), then the sequence is determined to be inclined to that subtype; if no subtype alignment score is full and the difference in the highest scores of each subtype is less than the set threshold, then it is marked as "untyped sequence".

[0023] Step 5: Determine the dominant subtype at the sample level; The subtype distribution of all separable sequences (excluding undetermined sequences) in the statistical sample is analyzed. The proportion of each subtype sequence to the total number of separable sequences is calculated. If the proportion of a certain subtype sequence exceeds 50%, the dominant subtype of the sample is determined to be that subtype.

[0024] Step 6: Output the fracturing results; Output standardized typing reports, including key information such as basic sample information, dominant subtype, minor subtype, sequence proportion of each subtype, and total number of typifiable sequences. Supports export in Excel or PDF format, making it convenient for clinicians to view and use as a basis for influenza virus surveillance data.

[0025] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0026] Example 1: Typing analysis of influenza A virus H1N1 subtype and influenza A virus H3N2 subtype; This embodiment uses the bioinformatics analysis method of the present invention to achieve rapid and accurate typing of H1N1 and H3N2 subtypes for clinically suspected influenza A infection samples. The specific operation process is as follows: (1) Sample collection and processing; Clinical influenza samples were collected, and RNA was extracted from the samples using a commercially available standard viral RNA extraction kit (strictly following the kit instructions). RNA was then reverse transcribed to obtain cDNA (complementary DNA). A high-throughput sequencing library was constructed using a standard library construction kit, with a library size of 50 bp, and the library concentration was adjusted to suit the sequencing platform requirements.

[0027] (2) Sequencing and data processing; Sequencing: The constructed library is sent to the BGI sequencing platform for high-throughput sequencing to obtain the raw sequencing data for each sample.

[0028] Sequencing data preprocessing: Fastp software (version Fastp 0.23.2; parameters set to -u 10 (allowing a maximum of 10% of base quality values ​​below Q20 in a single read), -e 30 (using the Q30 standard, i.e., the average error rate must be less than 0.001), -y -Y 60 (removing reads with a low complexity base ratio below 60%, retaining those above 60%), and -l 35 (requiring a sequencing length of at least 35bp) was used to perform quality control, adapter removal, and length screening on the raw sequencing data. Then, BWA software (version BWA-MEM2 2.2.1; parameters -M -T 25) was used to filter human sequences (referencing the hg38 database), retaining high-quality sequences for analysis.

[0029] Sequence alignment: The preprocessed 50bp high-quality sequence was aligned with the subtype sequences in the constructed reference database using BWA software (version BWA-MEM2 2.2.1, parameters -M -T 30 -a). The alignment results were in SAM file format. The alignment scores (out of 50) of each sequence in different subtypes were extracted (using Python 3.6 programming language to extract the third column (aligned reference genome) and the fourteenth column (alignment score) of each sequence in the SAM file).

[0030] In this embodiment, the reference database includes the following subtypes of HA and NA feature fragment sequences: H1N1: (GenBank ID: GCA_000865725.1); H3N2: (GenBank ID: GCA_039834415.1); Other subtypes: The database can be dynamically updated based on the actual types of influenza outbreaks.

[0031] (3) Sequence hierarchical subtype determination; Each sequence is assigned a subtype according to the preset determination rules: Condition 1: If a sequence scores 50 points when aligned to two or more different subtypes, then the sequence cannot be distinguished as a subtype and is marked as an "untyped sequence". Condition 2: If a sequence scores 50 points in only one subtype, while failing to reach full marks in other subtypes, then the sequence is determined to belong to that subtype; Condition 3: If the subtype alignment score is 50, and the highest alignment score of a certain subtype is at least 5 points higher than that of other subtypes, then the sequence is determined to be inclined to that subtype; if the subtype alignment score is 50 and the difference in the highest scores of each subtype is less than 5 points, then it is marked as an "untyped sequence".

[0032] (4) Determination of dominant subtype at the sample level; The subtype distribution of all separable sequences (excluding untyped sequences) in the statistical sample is analyzed, and the proportion of each subtype to the total number of separable sequences is calculated. In this embodiment, the H3N2 subtype accounts for 99.95% of the samples, and the H1N1 subtype accounts for 0.05%. The H3N2 subtype accounts for more than 50% and is therefore determined to be the dominant subtype of the sample.

[0033] (5) Output of typing results; The system outputs a standardized typing report, including basic sample information, the dominant subtype (H3N2), the minor subtype (H1N1), the percentage of each subtype (H3N2: 99.95%, H1N1: 0.05%), and the total number of typifiable sequences (12882). The report can be exported in Excel or PDF format. See Table 1 for the typing data exported in Excel format, and Table 2 for the detailed RNA virus detection results exported in PDF format.

[0034] Table 1

[0035] Table 2

[0036] Based on the typing results, it can be basically determined that the sample is infected with the H3N2 subtype of influenza A virus.

[0037] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the present invention.

Claims

1. A bioinformatics analysis method for rapid subtyping of influenza A virus subtypes, characterized by, The method comprises the following steps: Step 1: database construction; A reference database comprising multiple influenza A virus subtypes is established; Step 2: preprocessing of sequencing data; Quality control, adapter removal and length screening are performed on the raw sequencing data, and human sequences are filtered to obtain high-quality sequences; Step 3: sequence alignment; The preprocessed high-quality sequences are aligned with the sequences of each subtype in the reference database to generate an alignment result file, and the alignment scores of each sequence in different subtypes are extracted; Step 4: sequence-level subtype determination; The subtype of each sequence is determined according to the preset determination rule: Condition 1: if the scores of a sequence aligned to two or more different subtypes are all full marks, the sequence cannot be distinguished by subtype, and is marked as "untyped sequence"; Condition 2: if the alignment score of a sequence is full mark in a certain subtype, and the scores of other subtypes do not reach full mark, the sequence is determined to be the subtype; Condition 3: if there is no full mark for subtype alignment, and the highest alignment score of a certain subtype is at least a set threshold higher than that of other subtypes, the sequence is determined to be inclined to the subtype; if there is no full mark for subtype alignment and the difference between the highest scores of each subtype is less than the set threshold, it is marked as "untyped sequence"; Step 5: determination of the dominant subtype at the sample level; The subtype distribution of all typeable sequences in the sample is counted, the proportion of sequences of each subtype is calculated, and the dominant subtype of the sample is determined; Step 6: output of typing results; A typing report is outputted, which includes sample basic information, dominant subtype, secondary subtype, sequence proportion of each subtype, and total number of typeable sequences.

2. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 1, the reference database comprises characteristic fragment sequences of hemagglutinin and neuraminidase genes of influenza A virus subtypes.

3. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 2, Fastp software is used for quality control, adapter removal and length screening, and BWA software is used to filter human sequences.

4. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 3, BWA software is used for sequence alignment, and the alignment result format is SAM file. The full score of alignment score is the length of sequencing sequence.

5. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 4, the initial value of the threshold is 5 points, which is determined based on 50-75 bp sequencing data test.

6. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 5, if the proportion of sequences of a certain subtype in all typeable sequences is more than 50%, the subtype is determined as the dominant subtype of the sample.

7. The bioinformatic analysis method for rapid subtyping of influenza A virus subtypes according to claim 1, characterized in that, In step 6, the typing report supports Excel or PDF format export.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the influenza A virus subtype rapid typing bioinformatics analysis method according to any one of claims 1-7.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the influenza A virus subtype rapid typing bioinformatics analysis method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Influenza A virus fast typing and analyzing process

    CN105989247A

  • Blood metagenome sequencing data analysis method and device and application thereof

    CN110349630A

  • Metagenome-based human adenovirus molecular typing and tracing method and system

    CN112687344A

  • Sequence typing method based on BLAST alignment

    CN119694392A