Method and device for automatically identifying and visualizing variation information of high-throughput sequencing result

By generating hyperlinks that directly jump to gene visualization tools, the problem of cumbersome steps and susceptibility to human error in high-throughput sequencing result analysis is solved, achieving efficient gene sequence visualization.

CN120853676APending Publication Date: 2025-10-28JINAN JINYU MEDICINE JIANYAN CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511185152.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In current technologies, the analysis of high-throughput sequencing results relies on manually importing gene sequence files, which is cumbersome and prone to human error, making it difficult to meet the clinical need for rapid viewing of gene sequences.

Method used

By generating hyperlinks to high-throughput sequencing results, users can directly jump to a gene visualization tool to display the gene sequences corresponding to the mutation locations in the sequencing samples, thus achieving automatic identification and visualization.

Benefits of technology

It shortens the operational path from high-throughput sequencing results to visualizing gene sequences, improves analysis efficiency, and reduces the impact of human error.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853676A_ABST
    Figure CN120853676A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic identification and visualization method and device for variation information of a high-throughput sequencing result. The method comprises the following steps: acquiring the high-throughput sequencing result by computer equipment; the computer equipment analyzes the high-throughput sequencing result, and extracts a field value of at least one target variation information field corresponding to each sequencing sample according to the variation type corresponding to each sequencing sample; the computer equipment generates a hyperlink corresponding to the first sequencing sample according to the field value of the at least one target variation information field corresponding to the first sequencing sample, the hyperlink is used for skipping to a gene visualization tool, and a gene sequence file corresponding to the first sequencing sample is displayed in the gene visualization tool. The gene sequence corresponds to the variation position of the first sequencing sample. According to the method, the corresponding gene sequence of each sequencing sample at the variation position can be directly skipped and displayed in the gene visualization tool according to the hyperlink, so that the efficiency of analyzing the high-throughput sequencing result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method and apparatus for automatic identification and visualization of variation information in high-throughput sequencing results. Background Technology

[0002] Currently, the massive sequencing data generated by high-throughput sequencing technology provides crucial evidence for elucidating the molecular mechanisms of diseases and guiding clinical decision-making. However, effective analysis of high-throughput sequencing results currently relies on the accurate interpretation of gene sequences, a process that typically requires the use of gene visualization tools to view relevant gene sequences. However, when medical professionals use gene visualization tools to view relevant gene sequences, they usually need to manually import gene sequence files and manually search for relevant gene sequences within those files. This entire process is cumbersome, time-consuming, and susceptible to human error, making it difficult to meet the urgent clinical need for rapid gene sequence viewing. Summary of the Invention

[0003] This application discloses an automatic identification and visualization method and apparatus for variation information in high-throughput sequencing results. By generating hyperlinks corresponding to each sequencing sample in the high-throughput sequencing results, the gene sequence corresponding to the variation position of each sequencing sample can be directly jumped to and displayed in the gene visualization tool, thereby improving the efficiency of analyzing high-throughput sequencing results.

[0004] The first aspect of this application discloses a method for automatic identification and visualization of variant information in high-throughput sequencing results, applied to a computer device, the method comprising:

[0005] Computer equipment acquires high-throughput sequencing results, the high-throughput sequencing results including variation information corresponding to one or more sequencing samples, the variation information including field values ​​corresponding to multiple variation information fields;

[0006] The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target mutation information field corresponding to each sequencing sample based on the mutation type corresponding to each sequencing sample; the at least one target mutation information field includes a field used to indicate the mutation location;

[0007] The computer device generates a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample; the first sequencing sample is any of the sequencing samples, and the hyperlink is used to jump to a gene visualization tool, and the gene visualization tool displays the gene sequence in the gene sequence file corresponding to the first sequencing sample, which corresponds to the variant position of the first sequencing sample.

[0008] In some possible embodiments, the high-throughput sequencing results include multiple mutation result sets, each mutation result set including mutation information corresponding to one or more sequencing samples with the same mutation type, and multiple mutation information fields corresponding to each sequencing sample within the same mutation result set are the same;

[0009] The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type corresponding to each sequencing sample, including:

[0010] The computer device determines the mutation type corresponding to each set of mutation results;

[0011] The computer device extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set; the target variant result set is any of the variant result sets, and the extraction rules are used to define the target variant information fields to be extracted.

[0012] In some possible embodiments, the computer device determines the mutation type corresponding to each of the mutation result sets, including:

[0013] The computer device iterates through the names corresponding to the multiple mutation result sets according to the keywords corresponding to the multiple mutation types.

[0014] If the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type, the computer device determines that the mutation type corresponding to the first mutation result set is the first mutation type; the first mutation result set can be any of the mutation result sets.

[0015] In some possible embodiments, the plurality of variant types include single nucleotide variants (SNVs), copy number variants (CNVs), and fusion genes (Fusion).

[0016] The extraction rules corresponding to the SNV define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating the position of the variant base, a field indicating the start position, and a field indicating the end position.

[0017] The extraction rules corresponding to the CNV define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating start position, and a field indicating end position.

[0018] The extraction rules corresponding to the Fusion define the target mutation information fields to be extracted, including one or more of the following: the field of the first breakpoint, the field of the second breakpoint, the field of the chromosome number corresponding to the first breakpoint, the field of the start position of the first breakpoint on the corresponding chromosome, and the field of the end position of the first breakpoint on the corresponding chromosome; the first breakpoint is used to characterize the break position of one chromosome in the Fusion, and the second breakpoint is used to characterize the break position of another chromosome in the Fusion.

[0019] In some possible embodiments, the computer device extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set, including:

[0020] The computer device extracts one or more initial mutation information fields from the target mutation result set according to the extraction rules matched by the mutation type corresponding to the target mutation result set;

[0021] The computer device performs a correction process on the one or more initial mutation information fields to obtain at least one target mutation information field;

[0022] The computer device extracts the field values ​​corresponding to each sequencing sample in the target variant result set and the at least one target variant information field.

[0023] In some possible embodiments, the target variant information field may also include the file storage path of the binary sequence alignment format (BAM) file corresponding to the sequencing sample;

[0024] The computer device generates a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample, including:

[0025] The computer device fills the hyperlink template with the file storage path of the BAM file corresponding to the first sequencing sample, the field value of the field used to indicate the mutation location corresponding to the first sequencing sample, and the file storage path of the preset reference genome, to generate the hyperlink corresponding to the first sequencing sample.

[0026] In some possible embodiments, the high-throughput sequencing results are stored in a table, with the field values ​​corresponding to each sequencing sample and the plurality of variant information fields stored in the cells under the header of each variant information field.

[0027] After generating the hyperlink corresponding to the first sequencing sample, the method further includes:

[0028] The computer device associates the hyperlink with a first cell of the first sequencing sample, the first cell storing the field value corresponding to the field used to indicate the location of the mutation in the first sequencing sample; or,

[0029] The computer device fills the hyperlink into a second cell of the first sequencing sample, the second cell being a cell below the header of the field used to indicate the hyperlink.

[0030] A second aspect of this application discloses an automatic identification and visualization device for high-throughput sequencing result variation information, the device comprising:

[0031] The result acquisition module is used to acquire high-throughput sequencing results, which include variation information corresponding to one or more sequencing samples, and the variation information includes field values ​​corresponding to multiple variation information fields.

[0032] The result parsing module is used to parse the high-throughput sequencing results and extract the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type of each sequencing sample; the at least one target variant information field includes a field for variant location.

[0033] The link generation module is used to generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample; the first sequencing sample is any of the sequencing samples, and the hyperlink is used to jump to a gene visualization tool, and display the gene sequence in the gene sequence file corresponding to the first sequencing sample that corresponds to the variant position of the first sequencing sample in the gene visualization tool.

[0034] A third aspect of this application discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform an automatic identification and visualization method for high-throughput sequencing result variation information as described in any of the above embodiments.

[0035] A fourth aspect of this application discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor enables the processor to implement the method for automatic identification and visualization of high-throughput sequencing result variation information as described in any of the above embodiments.

[0036] This application provides an automatic identification and visualization method and apparatus for variant information in high-throughput sequencing results. A computer device acquires high-throughput sequencing results, which include variant information corresponding to one or more sequencing samples. Each variant information includes field values ​​corresponding to multiple variant information fields. The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type. The at least one target variant information field includes a field indicating the variant location. The computer device generates a hyperlink corresponding to the first sequencing sample based on the field values ​​of the at least one target variant information field corresponding to the first sequencing sample. The first sequencing sample can be any sequencing sample. The hyperlink is used to jump to a gene visualization tool, and the gene visualization tool displays the gene sequence in the gene sequence file corresponding to the first sequencing sample, which corresponds to the variant location of the first sequencing sample.

[0037] In this way, the computer device can determine the mutation type of each sequencing sample based on the mutation information corresponding to each sequencing sample. Then, based on the mutation type of each sequencing sample, it can extract the field value of at least one target mutation information field corresponding to each sequencing sample. This determines the information used to accurately display the gene sequence corresponding to the mutation position of each sequencing sample in the gene sequence file. The computer device can then generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target mutation information field corresponding to the first sequencing sample. Users can click this hyperlink to jump to a gene visualization tool and directly view the gene sequence corresponding to the mutation position of the first sequencing sample. This shortens the operation path from high-throughput sequencing results to visualized gene sequences and improves the efficiency of analyzing high-throughput sequencing results. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a schematic diagram illustrating the display of a gene sequence in a gene visualization tool, as provided in an embodiment of this application.

[0040] Figure 2 A flowchart illustrating an automatic identification and visualization method for high-throughput sequencing result variation information provided in this application embodiment;

[0041] Figure 3A flowchart for extracting field values ​​of at least one target variant information field corresponding to each sequencing sample, provided in an embodiment of this application;

[0042] Figure 4 A flowchart for determining the mutation type of the first mutation result set provided in this application embodiment;

[0043] Figure 5 A flowchart illustrating the correction process for each variation information field provided in this application embodiment;

[0044] Figure 6 A structural block diagram of an automatic identification and visualization device for high-throughput sequencing result variation information provided in this application embodiment;

[0045] Figure 7 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0047] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0048] Furthermore, "at least one" refers to one or more, while "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0049] The method for automatic identification and visualization of high-throughput sequencing result variation information provided in this application embodiment can be applied to computer devices, including but not limited to personal computers, tablet computers, or laptop computers.

[0050] For example, Figure 1 This is a schematic diagram illustrating the display of a gene sequence in a gene visualization tool, as provided in an embodiment of this application. Figure 1 As shown, the computer device can display the gene sequence corresponding to the mutation position of the sequencing sample in the gene sequence file in the gene visualization tool, based on the genome coordinate interval 110 corresponding to the sequencing sample, the reference genome 120, and the gene sequence file 130 corresponding to the sequencing sample.

[0051] In some embodiments, the gene visualization tool may be a desktop application (hereinafter referred to as the desktop version), a web version, or a snapshot of the Integrated Genomics Viewer (IGV). The IGV desktop version is standalone software installed on a local computer device. It can directly access local hardware resources for high-throughput sequencing data processing and visualization, enabling rapid loading and analysis. The IGV web version does not require local installation; it allows quick viewing of gene sequences on specific chromosomes by accessing a cloud server through a browser. IGV snapshots can refer to static image files generated from the current visualization interface (such as gene sequences of specific chromosomal regions, distribution of variant sites, orbital combination views, etc.) via the desktop or web version.

[0052] The genomic coordinate interval 110 refers to one or more consecutive positional ranges on a genomic sequence, used to pinpoint the specific location of a gene sequence on a chromosome. The genomic coordinate interval 110 can be used to locate specific gene segments through clearly defined start and end coordinates. For example... Figure 1 As shown, the genome coordinate interval 110 can be chr1: 156,845,385-156,845,424, which is used to represent the continuous region on chromosome 1 from base position 156,845,385 to base position 156,845,424.

[0053] Reference Genome 120 refers to a standard version of the genome sequence used for alignment and annotation of high-throughput sequencing data. For example, Reference Genome 120 could be hg19.fa, representing the 19th edition of the Human Genome Reference Sequence. This version contains the standard base sequence and gene annotation information of the human genome and is one of the commonly used reference standards in bioinformatics analysis. In addition, ".fa" indicates the FASTA format, a common text format used in bioinformatics for storing nucleic acid or amino acid sequences.

[0054] In related technologies, when users need to view the gene sequence corresponding to a sequencing sample for analysis, they typically need to manually open a gene visualization tool, load the reference genome and the gene sequence file corresponding to the sequencing sample into the tool, and then manually extract the variant information corresponding to the sequencing sample from the high-throughput sequencing results. This process involves inputting the variant information into the gene visualization tool, which then redirects to the gene sequence file corresponding to the sequence sample and displays the gene sequence corresponding to the variant position. This entire viewing process is not only complex and increases the user's workload, reducing the efficiency of high-throughput sequencing result analysis, but also requires users to manually input a great deal of information, such as the reference genome, genome coordinate range, and the file path of the gene sequence file. Incorrect input can easily lead to inaccurate visualization results, affecting subsequent analysis.

[0055] In this embodiment, the computer device can determine the mutation type of each sequencing sample based on the mutation information corresponding to each sequencing sample. Then, based on the mutation type of each sequencing sample, it extracts the field value of at least one target mutation information field corresponding to each sequencing sample. This determines the information used to accurately display the gene sequence corresponding to the mutation position of each sequencing sample in the gene sequence file corresponding to the sequencing sample. The computer device can then generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target mutation information field corresponding to the first sequencing sample. Users can then click this hyperlink to jump to a gene visualization tool and directly view the gene sequence corresponding to the mutation position of the first sequencing sample within the gene visualization tool. This shortens the operation path from high-throughput sequencing results to visualized gene sequences and improves the efficiency of analyzing high-throughput sequencing results.

[0056] like Figure 2 As shown in one embodiment, an automatic identification and visualization method for variant information in high-throughput sequencing results is provided. This method may include the following steps:

[0057] Step 202: The computer equipment acquires the high-throughput sequencing results.

[0058] High-throughput sequencing results may include variation information corresponding to one or more sequencing samples, and the variation information may include field values ​​corresponding to multiple variation information fields.

[0059] The variant information field can include one or more of the following: sequencing sample name, variant type, variant location, variant frequency, and sequencing depth. Variant types can include single nucleotide variants (SNVs), copy number variants (CNVs), and fusion genes. Variant location can refer to the specific coordinates of the variant in the genome, including genomic coordinate ranges or the position of the variant base.

[0060] In some embodiments, high-throughput sequencing technology can be used to sequence nucleotide sequences to obtain raw sequencing data. The computer equipment can then perform filtering, deduplication, quality control checks, mutation detection, and annotation on the raw sequencing data to obtain the corresponding high-throughput sequencing results, which are then stored as an EXCEL file. High-throughput sequencing technology can refer to the technology of simultaneously sequencing a large number of sequence fragments in a nucleotide sequence. High-throughput sequencing technologies may include sequencing-by-synthesis, semiconductor sequencing, and single-molecule real-time sequencing.

[0061] In some embodiments, high-throughput sequencing results are stored in a table, with the field values ​​corresponding to each sequencing sample and multiple variant information fields stored in cells below the header of each variant information field. A computer device can read the table file corresponding to the high-throughput sequencing results and obtain the variant information corresponding to each sequencing sample stored in each cell of the table file.

[0062] For example, Table 1 shows the partial variation information corresponding to the multiple sequencing samples included in the high-throughput sequencing results. As shown in Table 1, the computer device can read the table file corresponding to the high-throughput sequencing results and obtain the variant information "NMid_CDSChange", "Gene.refGene", "Chr:start", "AAChange", "Frequency", "Total Reads", and "Reads" for the three sequencing samples, as well as the field values ​​corresponding to each variant information. "NMid_CDSChange" describes the variant type or specific change in the gene coding region. For example, "NM_001785.3:c.208G>A" indicates that in the coding sequence numbered NM_001785.3, the 208th base in the coding region has mutated from guanine G to linear purine A. "Gene.refGene" represents the gene name in the database corresponding to the reference genome. "Chr:start" is the position of the mutated base. "AAChange" indicates the amino acid sequence change caused by the variant, for example, p.A70T indicates that the alanine A at position 70 of the protein has mutated to threonine T. "Frequency" is the variant frequency. "Total Reads" is the total number of reads. "Reads" refers to the sequencing depth, and "Reads" refers to the number of sequencing fragments with that variant position.

[0063] Table 1

[0064] serial number NMid_CDSChange Gene.refGene Chr:start AAChange Frequency TotalReads Reads 1 NM_001785.3:c.208G>A CDA chr1:20931474 p.A70T 62.9% 733X 461 2 NM_030662.4:c.1176C>T MAP2K2 chr19:4090623 p.P392P 48.9% 715X 353 3 NM_000251.3:c.2400A>G MSH2 chr2:47705600 p.L800L 47.8% 848X 405

[0065] Step 204: The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type corresponding to each sequencing sample.

[0066] At least one target mutation information field includes a field indicating the location of the mutation.

[0067] Understandably, not all of the multiple variation information fields included in a sequencing sample can be used to indicate the gene sequence corresponding to the variation location in the gene sequence file corresponding to the sequencing sample. Therefore, the computer device needs to extract the variation information fields that can be used to indicate the gene sequence corresponding to the variation location in the gene sequence file corresponding to the sequencing sample from all the variation information fields of the sequencing sample, and use them as target variation information fields, such as fields used to indicate the variation location.

[0068] In some embodiments, the computer device can obtain multiple sets of field keywords corresponding to multiple mutation types according to the mutation types corresponding to each sequencing sample, and determine at least one target mutation information field that matches the set of field keywords corresponding to each sequencing sample from multiple mutation information fields of each sequencing sample, and extract the field value of at least one target mutation information field corresponding to each sequencing sample.

[0069] For example, the set of field keywords for variant type SNV includes the variant type and variant base position in the gene coding region. Therefore, the computer device can determine “NMid_CDSChange” and “Chr:start” from the above 7 variant information fields corresponding to each sequencing sample based on the set of keywords, and extract the field value “NM_001785.3:c.208G>A” corresponding to “NMid_CDSChange” and the field value “chr1:20931474” corresponding to “Chr:start” for sequencing sample number 1.

[0070] It should be noted that the inherent characteristics of different variant types determine the differences in the recording format of their variant locations. For example, SNVs involve only a single base change, so the variant location corresponding to an SNV only has the chromosome number and the position of the variant base. CNVs, on the other hand, involve duplication or deletion of genomic segments, so the variant location corresponding to a CNV is usually a genomic region including the chromosome number, start position, and end position, used to define the complete range of the variant. Furthermore, in the actual presentation of high-throughput sequencing results, different institutions have certain differences in the field division of the same variant information. Taking the variant location "chr1:20931474" corresponding to an SNV as an example, it can be simplified to a single variant information field "variant location," or it can be divided into two variant information fields: "chromosome number" and "variant base position." Therefore, for sequencing samples with different variant types, the target variant information fields extracted by the computer may be the same or different.

[0071] In some embodiments, a computer device may modify the field values ​​of at least one target variant information field corresponding to each extracted sequencing sample to integrate the record format of at least one target variant information field corresponding to each sequencing sample.

[0072] Step 206: The computer device generates a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample.

[0073] The first sequencing sample is any sequencing sample. The hyperlink is used to jump to the gene visualization tool, and the gene visualization tool displays the gene sequence in the gene sequence file corresponding to the first sequencing sample, which corresponds to the mutation position of the first sequencing sample.

[0074] In some embodiments, the computer device can build a multi-process parallel computing architecture based on the multiprocessing library of Python software to generate hyperlinks corresponding to multiple sequencing samples simultaneously based on the field values ​​of at least one target variant information field corresponding to multiple sequencing samples.

[0075] In some embodiments, the computer device may store a hyperlink template, and after obtaining the field value of at least one target variant information field corresponding to the first sequencing sample, fill the field value of at least one target variant information field corresponding to the first sequencing sample into the hyperlink template to form a hyperlink corresponding to the first sequencing sample.

[0076] For example, the target variant information field may include the variant location and the file storage path of the Binary Alignment / Map (BAM) file. The hyperlink template can be pre-set to:

[0077] http: / / localhost:60151 / load? file=A&genome=hg19&locus=B.

[0078] Wherein, 60151 is the fixed fill port number, A is the file storage path of the BAM file corresponding to the sequencing sample to be filled, hg19 is the reference genome version for the fixed fill, and B is the field value corresponding to the variant position to be filled.

[0079] In one implementation, after obtaining the field value of at least one target variant information field corresponding to the first sequencing sample, the computer device can fill the field value of at least one target variant information field corresponding to the first sequencing sample, as well as the configuration parameters, into the hyperlink template to form a hyperlink corresponding to the first sequencing sample.

[0080] Configuration parameters can be parameters related to the underlying operating logic of the hyperlink-invoked tool. For example, configuration parameters may include a port number. When the port number in the configuration parameters is the same as the port number currently being listened to by the computer device, the computer device can receive control commands through that port to control the gene visualization tool. Optionally, configuration parameters may also include the version number of the reference genome or the file storage path corresponding to the reference genome, but are not limited to these.

[0081] For example, the target variant information field may include the variant location and the file storage path of the BAM file, and configuration parameters may include the port number and the version number of the reference genome. The hyperlink template can be pre-set to:

[0082] http: / / localhost:C / load? file=A&genome=D&locus=B.

[0083] Where A is the file storage path of the BAM file corresponding to the sequencing sample to be filled, B is the field value corresponding to the variant position to be filled, C is the port number to be filled, and D is the reference genome version to be filled.

[0084] The computer device can fill the hyperlink template with the field value "chr12:2539828" of the variant position corresponding to the first sequencing sample, the field value " / data / sample1.bam" of the BAM file's file storage path, the port number "60151", and the version number "hg19" of the reference genome, forming the hyperlink corresponding to the first sequencing sample as follows:

[0085] http: / / localhost:60151 / load? file= / data / sample1.bam&genome=hg19&locus=chr12:2539828.

[0086] When a user clicks the hyperlink corresponding to the first sequencing sample, the computer device responds to the click operation and jumps to the gene visualization tool at port number 60151. It automatically loads the reference genome and the corresponding BAM file, and jumps to the BAM file corresponding to the first sequencing sample based on "chr12:2539828". The gene sequence corresponding to chr12:2539828 is displayed, thus completing the visualization of the variant sequence at the chr12:2539828 site of the first sequencing sample and its alignment.

[0087] As another implementation, the computer device can fill the hyperlink template with the file storage path of the BAM file corresponding to the first sequencing sample, the field value of the field used to indicate the mutation location corresponding to the first sequencing sample, and the file storage path of the preset reference genome to generate the hyperlink corresponding to the first sequencing sample.

[0088] Optionally, the computer device may also fill the hyperlink template with the file storage path of the protein domain file corresponding to the first sequencing sample, as well as the file storage path of the file obtained after analyzing and processing the BAM file.

[0089] For example, the computer device can fill the hyperlink template with the port number, the file storage path of the BAM file corresponding to the first sequencing sample, the field value of the field used to indicate the mutation location corresponding to the first sequencing sample, the file storage path of the preset reference genome, and the file storage path of the protein domain file corresponding to the first sequencing sample, to generate the hyperlink corresponding to the first sequencing sample as follows:

[0090] http: / / localhost:60151 / load? file=\\192.168.22.206\jnngsnas\hg19\hg19_refG ene.sorted.txt,\\192.168.22.206\jnngsnas\hg19\protein_domains_hg19_hs37d5_GRCh37_v2.3.0.gff3,\\1 0.108.6.241\jnngsnas6\Report\Solid\2025\2025-130\202506\20250603_AE010401012_4P250331069US293224D 2_A\bam\STTC11391_S20.markdup.BQSR.bam,\\10.108.6.241\jnngsnas6\Report\Solid\2025\2025-130\20250 6\20250603_AE010401012_4P250331069US293224D2_A\bam\STTC11391_S20.mutect2.bam&locus=chr1:20931474.

[0091] Among them, hg19_refGene.sorted.txt can be the file corresponding to the reference genome, protein_domains_hg19_hs37d5_GRCh37_v2.3.0.gff3 can be the protein domain file corresponding to the first sequencing sample, STTC11391_S20.markdup.BQSR.bam can be the corrected BAM file after recalibrating by marking repetitive sequences and base quality values, and STTC11391_S20.mutect2.bam can be the BAM file after mutation detection using the Mutect2 tool.

[0092] By employing the above method, when a computer device jumps to the gene visualization tool via a hyperlink, the gene visualization tool can load all the core data required for the sequencing sample analysis at once, eliminating the need to manually import different files in stages during use. This avoids analysis interruptions and omissions caused by missing files or incorrect import order. In addition to visualizing the gene sequence, the visualization tool can also simultaneously view the reference genome background, protein domain distribution, and original sequencing correction results, enabling multi-dimensional data linkage analysis and improving the efficiency and accuracy of high-throughput sequencing result analysis.

[0093] In some embodiments, the computer device can monitor the gene visualization tool to obtain its operating status. When the gene visualization tool is running, the computer device can respond to a click on a hyperlink, jump to the gene visualization tool, and display the gene sequence corresponding to the mutation position of the first sequencing sample in the gene sequence file corresponding to the first sequencing sample. When the gene visualization tool is running, the computer device can provide a prompt to indicate that the current environment of the computer device does not support the automatic jump function.

[0094] As one implementation method, the gene visualization tool can be an IGV desktop client. The computer device can monitor a pre-set local port number of the IGV desktop client. When a user clicks the hyperlink corresponding to the first sequencing sample, if the computer device detects the existence of the local port number (i.e., the IGV desktop client is in monitoring mode), the computer device can respond to the click operation by constructing a hyperlink request and pushing the hyperlink request to the IGV desktop client to achieve automatic loading of the reference genome and BAM file, as well as jumping to variant locations. If the computer device does not detect the existence of the local port number, it determines that the IGV desktop client is closed, and the computer device can issue a prompt to encourage the user to open the IGV desktop client.

[0095] It should be noted that since the IGV web client typically does not have the ability to listen on local ports, it usually cannot respond to the hyperlink requests corresponding to the hyperlinks shown in the above hyperlink template. Therefore, by setting the hyperlinks with port numbers, computer devices can not only effectively distinguish the operating environments of the IGV web client and the IGV desktop client, thus avoiding call failures caused by IGV version incompatibility, but also improve cross-platform compatibility and avoid the problem of the IGV web version's limited response to hyperlink commands.

[0096] Optionally, the computer device can detect whether a gene visualization tool is installed on the computer device if no local port number is detected. If the gene visualization tool is installed on the computer device, the computer device can automatically start the gene visualization tool.

[0097] As another implementation, the gene visualization tool can be an IGV web-based application. The computer device can communicate with a cloud server storing high-throughput sequencing results. After the user clicks the hyperlink corresponding to the first sequencing sample, the computer device can launch a default browser and send an access command with the value of at least one target variant information field to the cloud server. Upon receiving the access command, the cloud server verifies and matches the value of the at least one target variant information field, thereby displaying the gene sequence corresponding to the variant position in the gene sequence file corresponding to the first sequencing sample through the default browser. Visualizing gene sequences through the IGV web-based application not only avoids the need to install the IGV desktop client, allowing users to view the gene sequence corresponding to the variant position of the sequencing sample using only a browser, but also reduces the complexity of configuring the computer environment, thus improving the efficiency of high-throughput sequencing result analysis.

[0098] In this embodiment, the computer device can determine the mutation type of each sequencing sample based on the mutation information corresponding to each sequencing sample. Then, based on the mutation type of each sequencing sample, it extracts the field value of at least one target mutation information field corresponding to each sequencing sample. This determines the information used to accurately display the gene sequence corresponding to the mutation position of each sequencing sample in the gene sequence file corresponding to the sequencing sample. The computer device can then generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target mutation information field corresponding to the first sequencing sample. Users can then click this hyperlink to jump to a gene visualization tool and directly view the gene sequence corresponding to the mutation position of the first sequencing sample within the gene visualization tool. This shortens the operation path from high-throughput sequencing results to visualized gene sequences and improves the efficiency of analyzing high-throughput sequencing results.

[0099] In some embodiments, high-throughput sequencing results include multiple variant result sets, each variant result set including variant information corresponding to one or more sequencing samples with the same variant type, and multiple variant information fields corresponding to each sequencing sample within the same variant result set are the same.

[0100] Figure 3 This is a flowchart illustrating the extraction of field values ​​for at least one target variant information field corresponding to each sequencing sample, as provided in an embodiment of this application. Figure 3As shown, the computer equipment parses the high-throughput sequencing results and extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type. This may include the following steps:

[0101] Step 301: The computer device determines the mutation type corresponding to each mutation result set.

[0102] In some embodiments, a computer device may classify the sequencing samples included in the high-throughput sequencing results according to the mutation type to obtain various mutation result sets. Each mutation result set includes mutation information corresponding to one or more sequencing samples with the same mutation type, and multiple mutation information fields corresponding to each sequencing sample within the same mutation result set are the same.

[0103] Understandably, the mutation information between sequencing samples of different mutation types is usually not the same. Therefore, computer equipment can classify each sequencing sample according to the mutation type to obtain multiple mutation result sets, with each mutation result set corresponding to a specific mutation type. Furthermore, in a single high-throughput sequencing process, since all sequencing samples use the same sequencing and data processing procedures, each sequencing sample within the same mutation result set should have the same multiple mutation information fields, but the field values ​​of each corresponding set of multiple mutation information fields will be different.

[0104] In some embodiments, a computer device can determine the name corresponding to each set of mutation results by matching the name corresponding to each set of mutation results with the names corresponding to multiple mutation types respectively. Figure 4 A flowchart for determining the mutation type of the first mutation result set provided in an embodiment of this application. For example... Figure 4 As shown, step 301 may include the following steps:

[0105] Step 402: The computer device iterates through the names corresponding to the multiple mutation result sets according to the keywords corresponding to the multiple mutation types.

[0106] Keywords corresponding to a mutation type may include the name of the mutation type and its key characteristics. A mutation type may have one or more keywords; no specific limit is imposed here.

[0107] Traversing the names corresponding to multiple mutation result sets can refer to a computer device sequentially matching the keywords corresponding to multiple mutation types with the names corresponding to multiple mutation result sets.

[0108] Because different institutions have varying naming conventions for different variant result sets—for example, a variant result set corresponding to a single nucleotide variant (SNV) could be named "SNV" or "single nucleotide variant," or simply "single base variant," or directly named according to the characteristics of the SNV, such as "a single nucleotide substitution, insertion, or deletion"—it is not limited to these names. Therefore, although they are all variant result sets corresponding to SNVs, these sets often have multiple different names. When the name of the variant result set does not reflect the variant type, computer equipment struggles to quickly determine the specific variant type. Thus, the computer equipment needs to traverse the names corresponding to multiple variant result sets to clarify the specific variant type.

[0109] Step 404: If the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type, the computer device determines the mutation type corresponding to the first mutation result set as the first mutation type.

[0110] In some embodiments, if the name corresponding to the first mutation result set contains the keyword corresponding to the first mutation type, or if the similarity between the name corresponding to the first mutation result set and the keyword corresponding to the first mutation type is greater than a similarity threshold, the computer device may determine that the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type.

[0111] In some embodiments, the computer device may match one or more mutation information fields under the first mutation result set with the keywords corresponding to the first mutation type, and determine the mutation type corresponding to the first mutation result set as the first mutation type if any mutation information field in each mutation information field matches the keyword corresponding to the first mutation type, so as to improve the accuracy and reliability of the determination of the mutation type corresponding to the first mutation result set.

[0112] By determining the mutation type corresponding to the first mutation result set in the above manner, it is beneficial to determine the extraction rules corresponding to the mutation type in subsequent steps, thereby improving the efficiency and accuracy of extracting the field value of at least one target mutation information field from the first mutation result set, and improving the efficiency and accuracy of analyzing high-throughput sequencing data.

[0113] Step 303: The computer device extracts the field value of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set.

[0114] The target mutation result set can be any mutation result set, and the extraction rules can be used to define the target mutation information fields that need to be extracted.

[0115] In some embodiments, when multiple mutation types include SNV, CNV, and Fusion, the mutation result set includes an SNV set, a CNV set, and a Fusion set, and the extraction rules may include the extraction rules corresponding to SNV, the extraction rules corresponding to CNV, and the extraction rules corresponding to Fusion.

[0116] The extraction rules corresponding to SNVs define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating the position of the variant base, a field indicating the start position, and a field indicating the end position.

[0117] It should be noted that since SNVs only involve changes to a single base, the field values ​​corresponding to the fields indicating the position of the mutated base, the field indicating the start position, and the field indicating the end position are the same. In order to ensure that the record format of each field under SNVs is consistent with the record format of each field under other mutation types, any one of the above three fields can be extracted as the target mutation information field, or all three can be extracted as at least one target mutation field, but it is not limited to this.

[0118] For example, the computer device can extract the field values ​​corresponding to the fields indicating the positions of the mutated bases and the fields indicating the starting positions of each sequencing sample in the target variant result set corresponding to the SNV, according to the extraction rules corresponding to the SNV.

[0119] The extraction rules for CNV define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating start position, and a field indicating end position.

[0120] The extraction rules for Fusion define the target mutation information fields to be extracted, including one or more of the following: the field of the first breakpoint, the field of the second breakpoint, the field of the chromosome number corresponding to the first breakpoint, the field of the start position of the first breakpoint on the corresponding chromosome, and the field of the end position of the first breakpoint on the corresponding chromosome. The first breakpoint is used to characterize the break position of one chromosome in Fusion, and the second breakpoint is used to characterize the break position of another chromosome in Fusion.

[0121] Optionally, in the case of gene fusion mutations occurring on more than two chromosomes, the extraction rules corresponding to Fusion may further include fields corresponding to the break positions of each chromosome and / or fields corresponding to the start and end positions of each chromosome at the break positions.

[0122] Optionally, the extraction rules corresponding to SNV, CNV, and Fusion may also include the file storage path of the BAM file, the version number of the reference genome, or the file storage path of the reference genome.

[0123] The above method enables computer equipment to extract the field values ​​of at least one target variant information field corresponding to each sequencing sample within each variant result set, based on extraction rules that are based on the inherent characteristics of different variant types. This helps to improve the accuracy of variant information extraction, avoid mis-extraction or omission of target variant information fields due to differences in the characteristics of different variant types, and also facilitates the subsequent integration and analysis of data of different variant types, thereby improving the overall efficiency and accuracy of high-throughput sequencing data processing.

[0124] In some embodiments, a computer device can extract the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set, through keyword matching, fuzzy matching, or regular expressions.

[0125] Because different institutions have different names for different variant information fields, for example, a field used to indicate chromosome number can be represented by either "chr" or "chromosom", when a computer device extracts the field value of at least one target variant information field according to the extraction rules, it can use one or more of the following methods, such as keyword matching, fuzzy matching, or regular expressions, to identify each variant information field in the target variant result set.

[0126] For example, for fields used to indicate chromosome numbers, keywords can be preset with multiple possible variations such as "chr" or "chromosome", fuzzy matching can be used to identify approximate expressions such as "Chrom" or "Chr#", and regular expressions can construct more flexible matching patterns, such as "chr\d+" to match standard formats such as "chr1" and compound formats such as "chr12:123456".

[0127] By using one or more methods such as keyword matching, fuzzy matching, or regular expressions to identify each variant information field in the target variant result set, it is possible to effectively address the diverse naming methods of the same variant information field used by different sequencing platforms, analysis software, or research institutions. This helps to improve the comprehensiveness and accuracy of extracting the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set.

[0128] In this embodiment, the computer device iterates through the names corresponding to multiple mutation result sets based on the keywords corresponding to multiple mutation types. When the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type, the computer device determines the mutation type corresponding to the first mutation result set as the first mutation type. This not only enables rapid association between mutation types and mutation result sets, reducing subjective errors and operation time caused by manual judgment, but also lays an accurate and orderly data foundation for subsequent extraction and analysis of mutation information, further ensuring the efficiency and reliability of the entire high-throughput sequencing data processing process.

[0129] In some embodiments, the computer device may perform modification processing on each mutation information field to obtain the target mutation information field. Figure 5 This is a flowchart illustrating the correction process for each variation information field provided in an embodiment of this application. For example... Figure 5 As shown, the computer device extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set. This may include the following steps:

[0130] Step 501: The computer device extracts one or more initial mutation information fields from the target mutation result set according to the extraction rules matched by the mutation type corresponding to the target mutation result set.

[0131] Step 501, which extracts one or more initial mutation information fields, is the same as step 303 and will not be repeated here.

[0132] Step 503: The computer device corrects one or more initial mutation information fields to obtain at least one target mutation information field.

[0133] Correction processing can include splitting and merging. Splitting refers to dividing an initial mutation information field into multiple mutation information fields, while merging refers to combining multiple initial mutation information fields into a single mutation information field. Users can determine whether to perform splitting or merging processing on the individual initial mutation information fields based on their actual needs.

[0134] For example, in the case where an initial mutation information field is a field for mutation location and the mutation type is SNV, the computer device can split the field for mutation location corresponding to SNV into a field for indicating chromosome number and a field for indicating the position of mutated base.

[0135] When multiple initial mutation information fields are respectively a field for indicating chromosome number, a field for indicating start position, and a field for indicating end position, and the mutation type is CNV, the computer device can merge the field for indicating chromosome number, the field for indicating start position, and the field for indicating end position into a field for mutation position.

[0136] In some embodiments, the computer device may perform a validity check on the field values ​​corresponding to each initial mutation information field before performing correction processing on one or more initial mutation information fields; the computer device may also perform a validity check on the field values ​​corresponding to the target mutation information field after obtaining at least one target mutation information field.

[0137] Validation can refer to checking and verifying the field values ​​corresponding to the target variant information field from multiple aspects, such as format, logic, and value range. Validation may include verifying whether the chromosome name conforms to the preset naming conventions, verifying whether the file storage path of the BAM file is valid, and verifying whether the version number of the reference genome exists.

[0138] By validating the field values ​​corresponding to the initial variant information field or the target variant information field, invalid or erroneous data caused by sequencing errors, data transmission errors, or abnormal format conversion can be detected and intercepted in a timely manner, thus avoiding the generation of erroneous hyperlinks that could affect the analysis of high-throughput sequencing results.

[0139] Step 505: The computer device extracts the field values ​​corresponding to at least one target variant information field from each sequencing sample in the target variant result set.

[0140] In some embodiments, the computer device may perform correction processing on each initial mutation information field and perform corresponding correction processing on the field values ​​corresponding to each initial mutation information field. For example, the computer device may split the field value "chr12:2539828" into "chr12" and "2539828", or merge the field values ​​"chr1", "156,845,385" and "156,845,424" into "chr1:156,845,385-156,845,424". The computer device then extracts the field value corresponding to at least one target mutation information field based on the at least one target mutation information field after correction processing.

[0141] In this embodiment, the computer device extracts one or more initial variant information fields from the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set, and performs correction processing on the one or more initial variant information fields to obtain at least one target variant information field. Thus, the field values ​​corresponding to each sequencing sample in the target variant result set and at least one target variant information field are extracted. The field values ​​corresponding to each sequencing sample and the target variant information field can be made more standardized and accurate through correction processing methods such as splitting and merging. This makes the format of the field values ​​corresponding to the target variant information field more in line with the actual analysis needs and provides a reliable basis for the subsequent generation of hyperlinks, thereby improving the efficiency and accuracy of high-throughput sequencing result analysis.

[0142] In some embodiments, high-throughput sequencing results are stored in a table. The field values ​​corresponding to each sequencing sample and multiple variant information fields are stored in the cells under the header of each variant information field. Therefore, the high-throughput sequencing results can be an EXCEL file, the multiple variant result sets included in the high-throughput sequencing results can be a sheet sub-table of the table, the name of the variant result set can be the name of the sheet sub-table, the field values ​​of the variant information fields of the sequencing samples can be the specific data in the cells of the table, and the variant information fields of the sequencing samples can be the header of the cells where the corresponding field values ​​are located.

[0143] Computer devices can save hyperlinks to high-throughput sequencing results using any of the following methods:

[0144] Method 1: The computer device fills the second cell of the first sequencing sample with hyperlinks. The second cell is the cell below the header of the field used to indicate the hyperlink.

[0145] The computer device can fill the second cell of each sequencing sample with hyperlinks. The second cell can be a field value cell that the computer device or the user has set in advance in the table to indicate the hyperlink, or it can be a table header corresponding to the hyperlink and a cell below the table header generated by the computer device in the table corresponding to the high-throughput sequencing results after the hyperlink is generated.

[0146] Table 2 shows the fields and values ​​of some variation information corresponding to some sequencing samples provided in the embodiments of this application. As shown in Table 2, for sequencing sample No. 1, the computer device can generate a hyperlink based on the port number corresponding to No. 1, the file storage path of the BAM file, the version number of the reference genome, and the variation location. This hyperlink is then filled into the last column of the sequencing sample No. 1. This allows users to view the gene sequence of sequencing sample No. 1 simply by clicking the hyperlink under the IGV Link header, which will take them to the gene visualization tool. The gene visualization tool will then display the gene sequence corresponding to the variation location of sequencing sample No. 1 in the gene sequence file. The hyperlinks for sequencing samples No. 2 and No. 3 are similar and will not be described further.

[0147] Table 2

[0148]

[0149] By filling the second cell of the first sequencing sample with hyperlinks, users can not only see the hyperlinks intuitively, but also visualize the gene sequence without needing professional skills, simply by looking at the hyperlink text displayed in the table, thus reducing the difficulty of use. Furthermore, filling the second cell of the first sequencing sample with hyperlinks can also form a standardized information presentation format, further improving the efficiency and accuracy of users' analysis of high-throughput sequencing results.

[0150] Alternatively, the computer device can also directly generate a hyperlink corresponding to the first sequencing sample in the second cell using the HYPERLINK function.

[0151] Method 2: The computer device associates the hyperlink with the first cell of the first sequencing sample. The first cell is used to store the field value corresponding to the field used to indicate the location of the mutation in the first sequencing sample.

[0152] Table 3

[0153] serial number port number BAM file storage path Reference genome version number Variation location 1 60151 / data / sample1.bam hg19 chr1:20931474 2 60151 / data / sample2.bam hg19 chr19:4090623 3 60151 / data / sample3.bam hg19 chr2:47705600

[0154] As shown in Table 3, the computer device can associate hyperlinks with the first cell corresponding to the mutation location of sequencing sample number 1. This allows users to jump to the gene visualization tool by clicking the field value in the cell corresponding to the mutation location, where the gene sequence corresponding to the mutation location in the gene sequence file of sequencing sample number 1 is displayed. It should be noted that the underlined values ​​in Table 3 indicate that the field value is associated with the hyperlink; the same applies to subsequent tables. The hyperlinks for sequencing samples number 2 and 3 are similar and will not be described further.

[0155] Alternatively, hyperlinks can be associated with other cells in the first sequencing sample, such as the sequencing sample name, variant type, or file storage path of the BAM file, without specific limitations.

[0156] By associating hyperlinks with the first cell of the first sequencing sample, users can avoid seeing the field values ​​corresponding to the hyperlinks in the table, maintaining the simplicity and consistency of the table interface. This helps users to more clearly understand the field values ​​corresponding to each variation information field of the sequencing sample, thereby improving the efficiency of high-throughput sequencing results analysis.

[0157] In some embodiments, the computer device may also construct hyperlinks corresponding to the function items of the gene visualization tool. The function items of the gene visualization tool may include clear, save, view zoom, and snapshot export, etc.

[0158] The "Clear" function closes the BAM files corresponding to one or more sequencing samples currently displayed in the gene visualization tool's interface. The "Save" function saves the BAM files corresponding to one or more sequencing samples currently displayed in the gene visualization tool's interface. The "View Zoom" function adjusts the display ratio of the gene sequence in the gene visualization tool's interface according to a preset base range. The "Snapshot Export" function saves the gene sequence view currently displayed in the gene visualization tool's interface as an image file to the local storage path of the computer device.

[0159] In some embodiments, when multiple BAM files corresponding to sequencing samples are already open in the gene visualization tool, the default behavior of the tool is to overlay the interface corresponding to the new BAM file on top of the existing interface, rather than overwriting the interface corresponding to the already loaded BAM file. Therefore, when a user clicks multiple hyperlinks consecutively, multiple BAM files may be displayed overlaid, causing confusion in the gene visualization tool and affecting analysis efficiency.

[0160] As one implementation method, the gene visualization tool can be an IGV desktop client. The computer device can construct hyperlinks corresponding to the clearing function items based on the clearing function items and their corresponding application programming interfaces (APIs) on the IGV desktop client, as follows:

[0161] http: / / localhost:60151 / execute? command=clear.

[0162] The computer device can respond to the user's click on the hyperlink to construct a clear request and push the clear request to the IGV desktop to control the IGV desktop to close all BAM files corresponding to each currently loaded sequencing sample, thereby automatically clearing the display interface of the IGV desktop.

[0163] In some embodiments, the computer device may fill the third cell of the table with hyperlinks corresponding to each function item, and the third cell may be any unoccupied cell; or, the computer device may associate the hyperlinks corresponding to each function item with the fourth cell of the first sequencing sample, and the fourth cell may be any header of the first sequencing sample in the table among multiple headers.

[0164] For example, as shown in Table 4, the computer device can associate the hyperlinks corresponding to the clear function items with the header "IGV Links" of the sequencing samples, so that users can control the IGV desktop client to close all the BAM files corresponding to each currently loaded sequencing sample by clicking the header "IGV Links", effectively avoiding the cumulative loading problem caused by users clicking the hyperlinks multiple times.

[0165] Table 4

[0166]

[0167] By adopting the above approach, multiple functions of gene visualization tools can be integrated into high-throughput sequencing results, eliminating the need for users to operate within the gene visualization tools themselves. This maintains the high degree of automation and convenience of gene visualization tools while lowering the barrier to entry for users, thus helping to improve the efficiency and convenience of analyzing high-throughput sequencing results.

[0168] The method for automatic identification and visualization of high-throughput sequencing result variant information provided in the above embodiments. Figure 6 This is a structural block diagram of a device for automatically identifying and visualizing high-throughput sequencing result variant information, provided in an embodiment of this application. Figure 6 As shown, in one embodiment, an automatic identification and visualization device 600 for high-throughput sequencing result variation information is provided. The automatic identification and visualization device 600 for high-throughput sequencing result variation information includes a result acquisition module 601, a result parsing module 602, and a link generation module 603.

[0169] The result acquisition module 601 is used to acquire high-throughput sequencing results. The high-throughput sequencing results include variant information corresponding to one or more sequencing samples. The variant information includes field values ​​corresponding to multiple variant information fields.

[0170] The result parsing module 602 is used to parse the high-throughput sequencing results and extract the field values ​​of at least one target variant information field corresponding to each sequencing sample according to the variant type of each sequencing sample; the at least one target variant information field includes a field of variant location.

[0171] The link generation module 603 is used to generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample; the first sequencing sample can be any sequencing sample, and the hyperlink is used to jump to the gene visualization tool and display the gene sequence in the gene sequence file corresponding to the first sequencing sample that corresponds to the variant position of the first sequencing sample in the gene visualization tool.

[0172] In some embodiments, the result acquisition module 601 is further configured to determine the mutation type corresponding to each mutation result set.

[0173] The result acquisition module 601 is also used to extract the field value of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rule matched by the variant type corresponding to the target variant result set; the target variant result set is any variant result set, and the extraction rule is used to define the target variant information field to be extracted.

[0174] In some embodiments, the result acquisition module 601 includes a traversal submodule and a mutation type submodule.

[0175] The traversal submodule is used to traverse the names corresponding to multiple mutation result sets based on the keywords corresponding to multiple mutation types.

[0176] The mutation type submodule is used to determine, when the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type, that the computer device determines the mutation type corresponding to the first mutation result set as the first mutation type; the first mutation result set can be any mutation result set.

[0177] In some embodiments, the result acquisition module 601 includes an extraction submodule and a correction submodule.

[0178] The extraction submodule is used to extract one or more initial mutation information fields from the target mutation result set according to the extraction rules matched by the mutation type corresponding to the target mutation result set.

[0179] The correction submodule is used to correct one or more initial mutation information fields to obtain at least one target mutation information field.

[0180] The extraction submodule is also used to extract the field values ​​corresponding to at least one target variant information field from each sequencing sample within the target variant result set.

[0181] In some embodiments, the link generation module 603 is further configured to fill the hyperlink template with the file storage path of the BAM file corresponding to the first sequencing sample, the field value of the field used to indicate the mutation location corresponding to the first sequencing sample, and the file storage path corresponding to the preset reference genome, to generate a hyperlink corresponding to the first sequencing sample.

[0182] In some embodiments, the automatic identification and visualization device 600 for high-throughput sequencing result variation information further includes a link filling module.

[0183] The link filling module is used to associate hyperlinks with the first cell of the first sequencing sample. The first cell is used to store the field value of the first sequencing sample corresponding to the field used to indicate the location of the mutation.

[0184] The link filling module is also used to fill hyperlinks into the second cell of the first sequencing sample, which is the cell below the header of the field used to indicate the hyperlink.

[0185] Figure 7 This is a structural block diagram of a computer device provided in an embodiment of this application. Figure 7 As shown, the computer device 700 may include a memory 702 and a processor 701. The memory 702 stores a computer program. When the computer program is executed by the processor 701, the computer device 700 enables the computer device 700 to implement the automatic identification and visualization method of high-throughput sequencing result variation information as described in the above embodiments.

[0186] Processor 701 may include one or more processing cores. Processor 701 connects to various parts of the computer device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, and by calling data stored in memory. Optionally, processor 701 may be implemented using at least one hardware form of digital signal processing, field-programmable gate array, or programmable logic array. Processor 701 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 701 and may be implemented separately through a communication chip.

[0187] The memory 702 may include random access memory (RAM) or read-only memory (ROM). The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described above, etc. The data storage area may also store data created during the use of the computer device.

[0188] This application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor enables the processor to implement the method for automatic identification and visualization of high-throughput sequencing result variation information as described in the above embodiments.

[0189] This application discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, it enables the processor to implement the automatic identification and visualization method for high-throughput sequencing result variation information as described in the above embodiments.

[0190] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, ROM, etc.

[0191] The above description is merely a specific example of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A method for automatic identification and visualization of variant information in high-throughput sequencing results, characterized in that, The method includes: Computer equipment acquires high-throughput sequencing results, the high-throughput sequencing results including variation information corresponding to one or more sequencing samples, the variation information including field values ​​corresponding to multiple variation information fields; The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target mutation information field corresponding to each sequencing sample based on the mutation type corresponding to each sequencing sample; the at least one target mutation information field includes a field used to indicate the mutation location; The computer device generates a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample; the first sequencing sample is any of the sequencing samples, and the hyperlink is used to jump to a gene visualization tool, and the gene visualization tool displays the gene sequence in the gene sequence file corresponding to the first sequencing sample, which corresponds to the variant position of the first sequencing sample.

2. The method according to claim 1, characterized in that, The high-throughput sequencing results include multiple mutation result sets. Each mutation result set includes mutation information corresponding to one or more sequencing samples with the same mutation type. Moreover, within the same mutation result set, multiple mutation information fields corresponding to each sequencing sample are the same. The computer device parses the high-throughput sequencing results and extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type corresponding to each sequencing sample, including: The computer device determines the mutation type corresponding to each set of mutation results; The computer device extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set; the target variant result set is any of the variant result sets, and the extraction rules are used to define the target variant information fields to be extracted.

3. The method according to claim 2, characterized in that, The computer device determines the mutation type corresponding to each of the mutation result sets, including: The computer device iterates through the names corresponding to the multiple mutation result sets according to the keywords corresponding to the multiple mutation types. If the name corresponding to the first mutation result set matches the keyword corresponding to the first mutation type, the computer device determines that the mutation type corresponding to the first mutation result set is the first mutation type; the first mutation result set can be any of the mutation result sets.

4. The method according to claim 2, characterized in that, The multiple variant types include single nucleotide variants (SNVs), copy number variants (CNVs), and fusion genes (Fusion). The extraction rules corresponding to the SNV define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating the position of the variant base, a field indicating the start position, and a field indicating the end position. The extraction rules corresponding to the CNV define the target variant information fields to be extracted, including one or more of the following: a field indicating chromosome number, a field indicating start position, and a field indicating end position. The extraction rules corresponding to the Fusion define the target mutation information fields to be extracted, including one or more of the following: the field of the first breakpoint, the field of the second breakpoint, the field of the chromosome number corresponding to the first breakpoint, the field of the start position of the first breakpoint on the corresponding chromosome, and the field of the end position of the first breakpoint on the corresponding chromosome; the first breakpoint is used to characterize the break position of one chromosome in the Fusion, and the second breakpoint is used to characterize the break position of another chromosome in the Fusion.

5. The method according to claim 3, characterized in that, The computer device extracts the field values ​​of at least one target variant information field corresponding to each sequencing sample in the target variant result set according to the extraction rules matched by the variant type corresponding to the target variant result set, including: The computer device extracts one or more initial mutation information fields from the target mutation result set according to the extraction rules matched by the mutation type corresponding to the target mutation result set; The computer device performs a correction process on the one or more initial mutation information fields to obtain at least one target mutation information field; The computer device extracts the field values ​​corresponding to each sequencing sample in the target variant result set and the at least one target variant information field.

6. The method according to claim 1, characterized in that, The target variant information field also includes the file storage path of the binary sequence alignment format (BAM) file corresponding to the sequencing sample; The computer device generates a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample, including: The computer device fills the hyperlink template with the file storage path of the BAM file corresponding to the first sequencing sample, the field value of the field used to indicate the mutation location corresponding to the first sequencing sample, and the file storage path of the preset reference genome, to generate the hyperlink corresponding to the first sequencing sample.

7. The method according to any one of claims 1 to 5, characterized in that, The high-throughput sequencing results are stored in a table, and the field values ​​corresponding to each sequencing sample and the plurality of mutation information fields are stored in the cells under the table header corresponding to each mutation information field. After generating the hyperlink corresponding to the first sequencing sample, the method further includes: The computer device associates the hyperlink with a first cell of the first sequencing sample, the first cell being used to store the field value of the first sequencing sample corresponding to the field used to indicate the mutation location; or, The computer device fills the hyperlink into a second cell of the first sequencing sample, the second cell being a cell below the header of the field used to indicate the hyperlink.

8. An automatic identification and visualization device for variant information in high-throughput sequencing results, characterized in that, The device includes: The result acquisition module is used to acquire high-throughput sequencing results, which include variation information corresponding to one or more sequencing samples, and the variation information includes field values ​​corresponding to multiple variation information fields. The result parsing module is used to parse the high-throughput sequencing results and extract the field values ​​of at least one target variant information field corresponding to each sequencing sample based on the variant type of each sequencing sample; the at least one target variant information field includes a field for variant location. The link generation module is used to generate a hyperlink corresponding to the first sequencing sample based on the field value of at least one target variant information field corresponding to the first sequencing sample; the first sequencing sample is any of the sequencing samples, and the hyperlink is used to jump to a gene visualization tool, and display the gene sequence in the gene sequence file corresponding to the first sequencing sample that corresponds to the variant position of the first sequencing sample in the gene visualization tool.

9. A computer device, characterized in that, The system includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, enables the processor to implement the automatic identification and visualization method for high-throughput sequencing result variation information as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the processor enables the processor to implement the method for automatic identification and visualization of high-throughput sequencing result variation information as described in any one of claims 1 to 7.