Sequencing data analysis method and device, computer equipment and program product

By automating the analysis of variant sites in sequencing data using computer equipment and utilizing databases to find related information and interpret content, the problem of complex operation and inconsistent interpretation standards in the sequencing report generation process has been solved, thus improving analysis efficiency and accuracy.

CN121306240APending Publication Date: 2026-01-09TIANJIN KINGMED CENT FOR CLINICAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511402330.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

The current sequencing report generation process is complex and time-consuming for medical professionals to interpret and analyze, and inconsistent interpretation standards affect the accuracy of the reports.

Method used

Information on variant sites in sequencing data is obtained through computer equipment. Relevant information and interpretations are found in the database. The frequency of reports, clinical evidence level, and interpretations of variant sites are automatically analyzed and the relevant information is displayed.

Benefits of technology

It has improved the efficiency and quality of sequencing data analysis, reduced the need for manual interpretation, provided clinical experience and data support, and enabled the accumulation and inheritance of experience in sequencing data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306240A_ABST
    Figure CN121306240A_ABST
Patent Text Reader

Abstract

The invention provides a sequencing data analysis method and device, computer equipment and a program product, and is applied to the computer equipment, and the method comprises the following steps: the computer equipment obtains sequencing data of a to-be-analyzed sample, the sequencing data comprises one or more variation sites and variation information corresponding to each variation site, and the variation information corresponds to the variation sites; the variation information comprises identification information corresponding to variation sites; the computer equipment searches associated information corresponding to each variation site in a database according to the identification information corresponding to each variation site, and the associated information comprises one or more of the report frequency of the variation site, the clinical evidence level of the variation site and the first interpretation content of the variation site; and the computer equipment displays the associated information and / or variation information corresponding to each variation site. According to the method and the device, automatic association interpretation of a plurality of variation sites can be realized, and the efficiency and the quality of analyzing the sequencing data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method, apparatus, computer equipment, and program product for analyzing sequencing data. Background Technology

[0002] In the current sequencing report generation process, medical professionals need to interpret and analyze the sequencing data after bioinformatics analysis in order to fill in the interpretation content in the sequencing report and complete the report writing. However, the entire process is not only complex and time-consuming, but also prone to inconsistencies in the interpretation standards of the same medical professional at different times, which can lead to inconsistent results in the analysis of the same sequencing data and affect the accuracy of the sequencing report. Summary of the Invention

[0003] This application discloses a method, apparatus, computer equipment, and program product for analyzing sequencing data, which can automatically associate and interpret content from multiple variant sites, thereby improving the efficiency and quality of sequencing data analysis.

[0004] A first aspect of this application discloses a method for analyzing sequencing data, applied to a computer device, the method comprising:

[0005] The computer device acquires sequencing data of the sample to be analyzed. The sequencing data includes one or more mutation sites and mutation information corresponding to each mutation site. The mutation information includes identification information corresponding to the mutation site.

[0006] The computer device searches for the associated information corresponding to each of the variant sites in the database based on the identification information corresponding to each of the variant sites. The associated information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site.

[0007] The computer device displays the association information and / or variation information corresponding to each of the mutation sites.

[0008] In some possible embodiments, the association information includes the number of reported variant sites, the level of clinical evidence for the variant site, and the first interpretation of the variant site; the computer device searches for the association information corresponding to each variant site in the database based on the identification information corresponding to each variant site, including:

[0009] The computer device searches the database for the number of reports and the level of clinical evidence corresponding to each of the variant sites based on the identification information corresponding to each of the variant sites.

[0010] The computer device identifies target variant sites among the various variant sites whose reported number is greater than or equal to a threshold number.

[0011] The computer device searches the database for the first interpretation content corresponding to each of the target mutation sites based on the identification information corresponding to each of the target mutation sites.

[0012] In some possible embodiments, the computer device searches the database for first interpretation content corresponding to each of the target variant sites based on the identification information corresponding to each of the target variant sites, including:

[0013] The computer device, based on the identification information corresponding to the first target mutation site, filters out historical reports containing the first target mutation site from multiple historical reports stored in the database, and uses them as associated historical reports; the first target mutation site can be any of the target mutation sites.

[0014] The computer device extracts the report content corresponding to the first target mutation site from each of the associated historical reports, and uses it as the first interpretation content corresponding to the first target mutation site.

[0015] In some possible embodiments, after the computer device extracts the report content corresponding to the first target variant site from each of the associated historical reports as the first interpretation content corresponding to the first target variant site, the method further includes:

[0016] The computer device generates a first hyperlink based on the report location information corresponding to the first target mutation site and the first associated historical report, and the file access path corresponding to the first associated historical report. The first hyperlink is used to access the report content of the first target mutation site in the first associated historical report. The first associated historical report is any of the associated historical reports, and the report location information is used to indicate the paragraph position of the report content of the first target mutation site in the first associated historical report.

[0017] In some possible embodiments, after the computer device displays the association information and / or variation information corresponding to each of the variant sites, the method further includes:

[0018] In response to the interpretation operation targeting the first variant site, the computer device analyzes the first interpretation content corresponding to the first variant site using a content analysis model to obtain the second interpretation content; the first variant site is any of the variant sites.

[0019] The computer device generates a content summary based on the second interpretation content and displays the content summary;

[0020] The computer device displays the second interpretation content in response to the triggered viewing operation.

[0021] In some possible embodiments, the sequencing data further includes the number of fragments corresponding to one or more genes in the sample to be analyzed; the method further includes:

[0022] The computer device acquires the control sequencing data of the control group sample corresponding to the sample to be analyzed;

[0023] The computer device determines the expression result of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data;

[0024] The computer device displays the expression results corresponding to each gene.

[0025] In some possible embodiments, the control sequencing data includes the control expression level and normal expression level range corresponding to each gene, and the expression results corresponding to each gene include a first expression result and a second expression result;

[0026] The computer device determines the expression results of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data, including:

[0027] The computer device calculates the first expression level of the first gene based on the number of fragments corresponding to the first gene, wherein the first gene is any one of the genes in the sample to be analyzed.

[0028] The computer device compares the first expression level of the first gene with the normal expression level range corresponding to the first gene to determine the first expression result corresponding to the first gene.

[0029] The computer device calculates expression difference parameters and verification values ​​based on the first expression level corresponding to the first gene and the corresponding control expression level, and determines the second expression result corresponding to the first gene based on the expression difference parameters and the verification values.

[0030] The computer device displays the expression results corresponding to each gene, including:

[0031] If the first expression result and the corresponding second expression result of the first gene are the same, the computer device displays the first expression result and / or the second expression result of the first gene.

[0032] A second aspect of this application discloses a sequencing data analysis apparatus applied to a computer device, the apparatus comprising:

[0033] The data acquisition module is used to acquire sequencing data of the sample to be analyzed. The sequencing data includes one or more mutation sites and mutation information corresponding to each mutation site. The mutation information includes identification information corresponding to the mutation site.

[0034] The data analysis module is used to search for the association information corresponding to each of the variant sites in the database based on the identification information corresponding to each of the variant sites. The association information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site.

[0035] The data display module is used to display the association information and / or variation information corresponding to each of the said mutation sites.

[0036] A third aspect of this application discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to implement the method described in any of the above embodiments.

[0037] A fourth aspect of this application discloses a computer program product comprising a computer program that, when executed by a processor, causes the processor to perform the method described in any of the above embodiments.

[0038] This application provides a method, apparatus, computer device, and program product for analyzing sequencing data. The computer device acquires sequencing data of a sample to be analyzed. The sequencing data includes one or more variant sites and variant information corresponding to each variant site. The variant information includes identification information corresponding to each variant site. Based on the identification information corresponding to each variant site, the computer device searches for association information corresponding to each variant site in a database. The association information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site. The computer device displays the association information and / or variant information corresponding to each variant site.

[0039] In this way, computer equipment can retrieve information such as the number of reports, clinical evidence level, and first interpretation of each variant site from the database based on the identification information corresponding to each variant site. This allows for the reasonable summarization and organization of the database content based on each variant site in the sequencing report, forming an effective reference system. Furthermore, by displaying the associated information and / or variant information for each variant site, the computer equipment not only eliminates the need for medical professionals to manually search and filter information related to each variant site, improving the efficiency of sequencing data analysis, but also provides clinical experience and data support for medical professionals' analysis of current sequencing data by displaying the first interpretation of each variant site in the database. This facilitates the effective accumulation and transfer of sequencing data analysis experience, contributing to improved quality of sequencing data analysis. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 A structural block diagram of a sequencing report generation system provided in this application embodiment;

[0042] Figure 2 A flowchart illustrating a method for analyzing sequencing data provided in this application embodiment;

[0043] Figure 3 This is a flowchart illustrating the process of finding the association information corresponding to each variant site, as provided in an embodiment of this application.

[0044] Figure 4 This is a flowchart illustrating the process of finding the first interpretation content corresponding to each target variant site, as provided in an embodiment of this application.

[0045] Figure 5 A flowchart for interpreting the first variant site provided in this application embodiment;

[0046] Figure 6 A schematic diagram illustrating the display of a content summary provided in an embodiment of this application;

[0047] Figure 7 A flowchart showing the expression results of each gene is provided as an embodiment of this application;

[0048] Figure 8 A flowchart provided for embodiments of this application to determine the expression result of the first gene based on the number of fragments of the first gene and control sequencing data;

[0049] Figure 9 A structural block diagram of a sequencing data analysis device provided in an embodiment of this application;

[0050] Figure 10 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0053] Furthermore, "at least one" refers to one or more, while "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0054] Figure 1 This is a structural block diagram of a sequencing report generation system provided in an embodiment of this application. Figure 1 As shown, the sequencing report generation system 100 may include modules such as an intelligent parsing module 110, a visualization module 120, a report generation module 130, and a process traceability module 140. This sequencing report generation system 100 can be deployed on a computer device to analyze sequencing data of samples to be analyzed and generate sequencing reports.

[0055] Computer equipment can be electronic devices such as computers, laptops, and tablets, or servers or cloud servers, etc., without specific limitations.

[0056] A sequencing report refers to a test report obtained after interpreting, analyzing, and evaluating the sequencing data of the sample to be analyzed.

[0057] Sequencing data can refer to the data obtained after preliminary analysis of the raw sequencing data. Raw sequencing data can be the signal data directly output by the sequencer.

[0058] The preliminary analysis process may include steps such as quality control, sequence alignment, mutation detection, and annotation. Quality control may include filtering low-quality bases or removing duplicate sequencing fragments. Sequence alignment refers to comparing the quality-controlled valid sequence fragments with the human reference genome to determine the specific location and degree of matching of each valid sequence fragment on the human reference genome. Mutation detection refers to determining the type of variation in the valid sequence fragment corresponding to the mutation site based on the differences between the valid sequence fragment and the reference sequence fragment. Variation types may include sequence variations (such as single nucleotide variants (SNVs), small insertions and small deletions), copy number variations (CNVs), and gene fusions. Annotation refers to supplementing the corresponding variation information for the detected variation sites, providing a basis for the subsequent interpretation of the variation sites.

[0059] The intelligent parsing module 110 can be used for data interpretation, content integration, and further analysis of sequencing data. The intelligent parsing module 110 may include a first parsing unit 111 and a second parsing unit 112.

[0060] The computer device can use the first parsing unit 111 to query multiple historical reports stored in the sequencing report generation system 100 based on each variant site contained in the sequencing data, so as to obtain the report content associated with each variant site in each historical report, thereby associating the report content corresponding to each historical report with the corresponding variant site to form the first interpretation content corresponding to the variant site.

[0061] The computer device can integrate and interpret the first interpretation content and mutation information corresponding to the mutation site based on the artificial intelligence-based natural language processing model deployed in the second parsing unit 112 to obtain the second interpretation content corresponding to the mutation site.

[0062] The visualization module 120 can be used to view the sequencing fragments corresponding to each variant site in the sequencing data. Computer devices can display the sequencing fragments corresponding to each variant site through the visualization module 120, and users can control the displayed sequencing fragments corresponding to each variant site through interaction with the visualization module 120.

[0063] The visualization module 120 may include a first interaction unit 121, a second interaction unit 122 and a third interaction unit 123.

[0064] The computer device can adjust the display filtering threshold based on sequencing depth through the first interaction unit 121 to control the sequencing fragments displayed in the visualization module 120. For example, the first interaction unit 121 is the CoverageDepth parameter box. The computer device can control the visualization module 120 to only display sequencing fragments smaller than the depth threshold according to the depth threshold set by the user in the CoverageDepth parameter box.

[0065] The computer device can control whether to display sequencing fragments that belong to the positive strand, the reverse strand, or the double strand in the sequencing data through the second interaction unit 122. For example, the second interaction unit 122 is a StrandSpecific option box. The computer device can control the visualization module 120 to display sequencing fragments that belong to the positive strand, the reverse strand, or the double strand according to the option selected by the user in the StrandSpecific option box.

[0066] The computer device can highlight the variant positions corresponding to each variant site in the sequencing data through the third interaction unit 123. For example, the third interaction unit 123 is the SpliceJunctions parameter box. The computer device can control the visualization module 120 to highlight or not highlight the variant positions corresponding to each variant site according to the user's selection to enable or disable the display of variant positions in the SpliceJunctions parameter box.

[0067] The report generation module 130 can be used to generate a sequencing report based on a preset sequencing report template, according to the interpretation content output by the intelligent parsing module 110 and / or the sequencing fragments displayed by the visualization module 120.

[0068] The report generation module 130 may include a generation unit 131 and a storage unit 132.

[0069] The computer device can use the generation unit 131 to fill the corresponding position in the sequencing report template with the interpretation content output by the intelligent parsing module 110 selected by the user, thereby generating a sequencing report.

[0070] The computer device can store the sequencing report generated by the generation unit 131, the interpretation content output by the intelligent parsing module 110, and the association between each sequencing report and each variant site in the database through the storage unit 132, and number the multiple sequencing reports stored in the database.

[0071] The process traceability module 140 can be used to record operational information throughout the entire process of creating, generating, modifying, and publishing sequencing reports.

[0072] The process traceability module 140 may include a recording unit 141 and a query unit 142.

[0073] Computer equipment can record the entire process of operation information through recording unit 141 based on blockchain evidence storage technology: In each step of the operation such as creation, generation, modification and publication of sequencing report, the operation information such as operator identity, operation timestamp, and data revision version is first hashed to generate a unique and irreversible hash value. Then, the hash value and operation information are packaged into an independent data block. At the same time, the hash value of the previous block is written into the header of the current block to build a storage structure of "block association and chain traceability".

[0074] The computer device can query the sequencing report corresponding to the report number entered by the user in the query unit 142, and display the operation information of the sequencing report in the entire process.

[0075] In some embodiments, the sequencing report generation system 100 further includes an adding module. The computer device can obtain the identification information corresponding to the variant site input by the user through the adding module, and obtain the variant information and interpretation content of the corresponding variant site in the sequencing data based on the identification information through the intelligent parsing module 110. Then, the variant information and interpretation content corresponding to the variant site are displayed in the area where the user inputs the identification information corresponding to the variant site.

[0076] In some embodiments, the sequencing report generation system 100 further includes an upload module. A computer device can obtain the storage file corresponding to the sequencing data uploaded by the user through the upload module, and then use the intelligent parsing module 110 to interpret, integrate, and further analyze the sequencing data contained in the storage file.

[0077] The computer device can also acquire files corresponding to clinical information uploaded by the user through the upload module. The clinical information may include patient information, doctor information, and sampling information of the sample to be analyzed. After acquiring the clinical information, the computer device can use the generation unit 131 to fill the corresponding positions in the sequencing report template with the interpretation content and clinical information output by the intelligent parsing module 110 selected by the user, so as to generate a sequencing report.

[0078] In this embodiment of the application, the computer device can obtain the sequencing data of the sample to be analyzed through the sequencing report generation system 100, search for the association information corresponding to each variant site in the database according to the identification information corresponding to each variant site, and display the association information and / or variant information corresponding to each variant site.

[0079] like Figure 2 As shown, in one embodiment, a method for analyzing sequencing data is provided, which can be applied to a computer device. The method may include the following steps:

[0080] Step 202: The computer device acquires the sequencing data of the sample to be analyzed.

[0081] Sequencing data includes one or more variant sites, as well as variant information corresponding to each variant site. The variant information includes the identification information corresponding to the variant site.

[0082] The mutation information corresponding to the mutation site may include the mutation location, mutation frequency, sequencing depth, and identification information.

[0083] Variant location refers to the precise location of a variant site within the corresponding gene sequence, identified by its chromosome number and specific base position. Mutation frequency refers to the proportion of sequencing fragments containing variant sites out of the total number of sequencing fragments. Sequencing depth refers to the number of times a variant site is covered by sequencing fragments in the sequencing data.

[0084] Identification information can be specific identifiers or characteristic descriptions corresponding to variant sites, used to distinguish different variant sites. Since the types of variants at different variant sites may differ, the presentation of identification information for different variant sites will also vary. For gene fusion variant sites, the gene name corresponding to the gene undergoing fusion can be used as identification information. For example, for a variant site formed by the fusion of gene MBP and gene CTDP1, the corresponding identification information can be set to "MBP::CTDP1", where the symbol "::" can be used to represent gene fusion. For sequence variation variant sites, the variant position can be directly used as identification information, such as "chr1:2491397", but this is not a limitation.

[0085] Step 204: The computer device searches the database for the associated information corresponding to each variant site based on the identification information corresponding to each variant site.

[0086] Related information may include one or more of the following: the number of times the variant site has been reported, the level of clinical evidence for the variant site, and the first interpretation of the variant site.

[0087] In some embodiments, the database may store multiple historical reports and the clinical evidence levels corresponding to multiple variant sites. The computer device can retrieve the identification information corresponding to each variant site from the multiple historical reports to obtain the number of reports for the variant site and / or the first interpretation content of the variant site; the computer device can match the identification information corresponding to each variant site with the multiple variant sites stored in the database to determine the clinical evidence level corresponding to each variant site.

[0088] The number of reports can include the first number of reports and the second number of reports. The first number of reports can be the number of historical reports that retrieved the identification information corresponding to the variant site, and the second number of reports can be the cumulative frequency of occurrence of the identification information corresponding to the variant site.

[0089] For example, the database may store two historical reports. A computer device can search for the identifier information corresponding to a variant site in multiple historical reports to obtain the first and second report counts corresponding to that variant site. The first report count is 2, indicating that the variant site is present in both historical reports; the second report count is 20, indicating that the variant site appears a total of 20 times in the two historical reports.

[0090] Optionally, the number of reports may also include the number of third reports for each historical report, where the number of third reports can be the frequency of occurrence of the identifier information of the variant site in each historical report. For example, in two historical reports, the number of third reports for a variant site is 18 and 2, respectively, indicating that the identifier information of the variant site appeared 18 times in the first historical report and 2 times in the second historical report.

[0091] Clinical evidence levels are used to assess the strength of the association between a variant and a disease. Clinical evidence levels can include pathogenicity, probable pathogenicity, unclear clinical significance, possibly benign, and benign. Pathogenicity indicates sufficient evidence that the variant causes disease; probable pathogenicity indicates strong but not entirely sufficient evidence that the variant is likely to cause disease; unclear clinical significance indicates insufficient or contradictory evidence that makes it impossible to determine whether the variant is pathogenic or benign; possibly benign indicates strong but not entirely sufficient evidence that the variant is likely harmless; and benign indicates sufficient evidence that the variant has no effect on health.

[0092] Optionally, the clinical evidence level can also be high, medium, or low. A higher clinical evidence level indicates more sufficient clinical evidence that the variant site causes a disease, while a lower clinical evidence level indicates less sufficient clinical evidence that the variant site causes a disease, or that there is sufficient evidence to show that the variant site has no effect on health.

[0093] In some embodiments, if the identification information corresponding to the variant site in the sequencing data is the same as the identification information corresponding to the variant site stored in the database, the computer device may use the clinical evidence level corresponding to the variant site stored in the database as the clinical evidence level corresponding to the variant site in the sequencing data. If the identification information corresponding to the variant sites stored in the database does not contain the identification information corresponding to the variant site in the sequencing data, the computer device may not mark the clinical evidence level corresponding to the variant site in the sequencing data, or may mark its clinical evidence level as a preset level.

[0094] The first interpretation may include the corresponding report content of the variant site in historical reports.

[0095] In some embodiments, if the computer device can retrieve the identification information corresponding to the variant site from the historical report, it can extract the report content corresponding to the paragraph containing the variant site from the historical report as the first interpretation content corresponding to the variant site.

[0096] Optionally, the first interpretation may also include information such as treatment recommendations and relevant literature corresponding to the variant sites stored in the database. Treatment recommendations may refer to medication regimens and symptomatic treatment measures for diseases related to the variant site; relevant literature may refer to published academic materials that support the level of clinical evidence for the variant site and information on treatment recommendations.

[0097] Step 206: The computer device displays the association information and / or variation information corresponding to each variant site.

[0098] In some embodiments, the computer device may display the association information and / or variation information corresponding to each variant site according to the variation type corresponding to each variant site.

[0099] When the mutation type corresponding to the second mutation site is a sequence variation, the computer device can display the mutation information of the second mutation site, and in response to an expansion operation targeting the second mutation site, the computer device can display the association information of the second mutation site. When the mutation type corresponding to the third mutation site is a gene fusion, the computer device can display the association information of the third mutation site, and in response to an expansion operation targeting the third mutation site, the computer device can display the mutation information of the third mutation site. The second and third mutation sites are either distinct mutation sites.

[0100] It should be noted that when the computer device displays the association information and / or variation information corresponding to each variation site, since the third variation site usually involves multiple genes, the number of variation information types of the third variation site is more than that of the second variation site. Furthermore, due to the limitations of the computer device's display size, fully displaying the association information and variation information of each variation site would affect the user's viewing experience. Therefore, the computer device can display only the variation information or association information according to the variation type of the variation site, and then display the association information or variation information when the user triggers an expansion operation.

[0101] Furthermore, when the mutation type corresponding to the third mutation site is gene fusion, the computer device can display the association information and partial mutation information of the third mutation site, and can display another part of the mutation information of the third mutation site in response to the expansion operation targeting the third mutation site.

[0102] Partial variant information may include the variant location and the number of supporting reads. The number of supporting reads can be the number of sequencing fragments in the sequencing data that contain the variant location at that variant site.

[0103] Since the mutation location can directly indicate the specific coordinates of the gene fusion in the gene sequence, and the number of supported reads can directly reflect the detection reliability of the third mutation site, displaying the mutation location and the number of supported reads allows users to quickly assess the effectiveness of the third mutation site. Therefore, this part of the mutation information is displayed first, rather than hidden.

[0104] As shown in Table 1, the computer device can display the variation information corresponding to the second variation sites of multiple variation types that are sequence variations. The variation information may include locus, gene, chromosome, transcript number, location, changes in complementary deoxyribonucleic acid (cDNA) level, changes in amino acid level, mutation frequency, sequencing depth, and mutation type.

[0105] Table 1

[0106]

[0107] In this context, "Locus" can refer to a specific location of a gene on a chromosome; "gene" can be the name of the gene at the second mutation site; "chromosome" can be the chromosome number where the gene at the second mutation site is located; "transcript number" can be the sequence number of the ribonucleic acid (RNA) produced by the transcription of the gene at the second mutation site from a database; "location" can be the specific region of the second mutation site in the transcript or genome, for example, "exon4" indicates that the second mutation site is located in exon 4; "cDNA level change" can refer to the sequence variation that occurs after DNA is transcribed into messenger RNA (mRNA), for example, "c.440G>C" indicates that the 440th base of cDNA is mutated from guanine (G) to cytosine (C); "amino acid level change" can refer to the protein sequence change caused by cDNA variation, for example, "p.G147A" indicates that the 147th amino acid of a protein is mutated from glycine (Gly, abbreviated as G) to alanine (Ala, abbreviated as A).

[0108] Optionally, the computer device may display a viewing interaction icon for the association information corresponding to each second variant site in the last column. When a user needs to view the association information, he / she can click the viewing interaction icon, and the computer device may display the association information corresponding to the second variant site in response to the triggered viewing operation.

[0109] As shown in Table 2, the computer device can display the associated information, partial variant information, and viewing interaction icons corresponding to multiple third variant sites with gene fusion as the variant type. The associated information of the third variant site may include the number of reports, reference grade, and first interpretation content. The first interpretation content may include fusion-related diseases and literature. The partial variant information may include the number of supported reads and location information. The underlined fields in Table 2 indicate fields associated with hyperlinks. The computer device can respond to hyperlink triggers by displaying the content associated with the hyperlink.

[0110] Table 2

[0111]

[0112] Among them, the number of reported times can be the number of first reports corresponding to the third variant site; the reference grade can be the clinical evidence level; the fusion related diseases and literature can be the relevant literature of the third variant site; and the location information can be the variant location of the third variant site.

[0113] Users can click the "View" interaction icon when they need to view another part of the variation information of the third variation site. The computer device can respond to the triggered view operation and display another part of the variation information corresponding to the third variation site.

[0114] In this embodiment, a computer device acquires sequencing data of a sample to be analyzed. The sequencing data includes one or more variant sites and variant information corresponding to each variant site. The variant information includes identification information corresponding to each variant site. Based on the identification information corresponding to each variant site, the computer device searches for the association information corresponding to each variant site in the database. The association information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site. The computer device displays the association information and / or variant information corresponding to each variant site. In this way, computer equipment can retrieve information such as the number of reports, clinical evidence level, and first interpretation of each variant site from the database based on the identification information corresponding to each variant site. This allows for the reasonable summarization and organization of the database content based on each variant site in the sequencing report, forming an effective reference system. Furthermore, by displaying the associated information and / or variant information for each variant site, the computer equipment not only eliminates the need for medical professionals to manually search and filter information related to each variant site, improving the efficiency of sequencing data analysis, but also provides clinical experience and data support for medical professionals' analysis of current sequencing data by displaying the first interpretation of each variant site in the database. This facilitates the effective accumulation and transfer of sequencing data analysis experience, contributing to improved quality of sequencing data analysis.

[0115] In some embodiments, since sequencing data may typically include multiple variant sites, but some variant sites have few related reports in the database, it is not necessary to analyze all variant sites. Figure 3 This is a flowchart illustrating the process of finding the association information corresponding to each variant site, as provided in an embodiment of this application. Figure 3 As shown, the computer device searches for the associated information of each variant site in the database based on the identification information corresponding to each variant site, which may include the following steps:

[0116] Step 301: The computer device searches the database for the number of reports and the level of clinical evidence corresponding to each variant site based on the identification information corresponding to each variant site.

[0117] In some embodiments, the database may store a dataset of multiple variant sites corresponding to the number of reports and the level of clinical evidence. Therefore, the computer device can directly search in the dataset based on the identification information corresponding to each variant site. If the identification information of the variant site matches the identification information of the variant site in the dataset, the number of reports and the level of clinical evidence corresponding to the variant site can be obtained without searching in multiple historical reports stored in the database, which helps to improve the search efficiency of each variant site.

[0118] Step 303: The computer device identifies the variant sites among the various variant sites whose reported number is greater than or equal to the number threshold as target variant sites.

[0119] Since the number of reports of the variant site is greater than or equal to the number threshold, it means that there are multiple historical reports about the variant site, that is, the variant site has sufficient data support. Therefore, the variant site is identified as the target variant site in order to obtain the first interpretation content corresponding to the variant site in the future.

[0120] In one implementation, the number of reports may include a first number of reports. A computer device may identify target variant sites among various variant sites whose first number of reports is greater than or equal to a threshold corresponding to the first number of reports.

[0121] If the number of first reports of a variant site is greater than or equal to the threshold corresponding to the number of first reports, it indicates that there are many historical reports involving that variant site in the database, so the variant site can be used as a target variant site; however, if the number of first reports of a variant site is less than the threshold corresponding to the number of first reports, it indicates that the database lacks sufficient knowledge of that variant site, so the variant site should not be used as a target variant site.

[0122] The threshold for the first report can be set according to actual needs. For example, the threshold for the first report can be 4 times. That is, if the number of historical reports involving the variant site is greater than or equal to 4, the variant site will be used as the target variant site, but it is not limited to this.

[0123] As another implementation, the number of reports may include a second number of reports. The computer device may identify target variant sites among the various variant sites whose second number of reports is greater than or equal to a threshold corresponding to the second number of reports.

[0124] If the number of second reports of a variant site is greater than or equal to the threshold corresponding to the number of second reports, it indicates that the identification information of the variant site appears frequently in multiple historical reports in the database. This variant site has rich clinical case records and research analysis data, so it can be used as a target variant site. However, if the number of second reports of a variant site is less than the threshold corresponding to the number of second reports, it indicates that there is insufficient data accumulation and research support for this variant site in the database. Therefore, it will not be included in the target variant site for the time being to avoid deviations in subsequent interpretation results due to insufficient data.

[0125] For example, if the second report count of variant site A is 100 and the second report count of variant site B is 5, and the threshold for the second report count is set to 20, the computer device can use variant site A as the target variant site.

[0126] As another implementation, the number of reports may include the number of third reports corresponding to each historical report. A computer device can identify a target variant site as one that has a third report count greater than or equal to a threshold corresponding to any historical report among all variant sites.

[0127] Since the third report count represents the frequency of occurrence of the variant site in each historical report, if the third report count corresponding to the variant site and any historical report is greater than or equal to the threshold corresponding to the third report count, it indicates that the historical report has conducted a relatively in-depth analysis and study of the variant site, and therefore the variant site can be used as a target variant site. However, if the third report count corresponding to the variant site and each historical report is less than the threshold corresponding to the third report count, it indicates that the historical report may have only mentioned the variant site, and the correlation between the historical report and the variant site is not high, so it is not necessary to use the variant site as a target variant site.

[0128] The threshold for the number of third reports can also be set according to actual needs. For example, the threshold for the number of third reports can be 10 times. That is, if the number of times the variant site is mentioned in a historical report is greater than or equal to 10 times, the variant site can be used as the target variant site, but it is not limited to this.

[0129] For example, the database may store two historical reports. The number of third reports corresponding to variant site A and these two historical reports is 1 and 20, respectively. The number of third reports corresponding to variant site B and these two historical reports is 3 and 4, respectively. Therefore, when the threshold for the number of third reports is set to 10, the computer device can use variant site A as the target variant site.

[0130] Optionally, the number of reports may include at least two of the first number of reports, the second number of reports, and the third number of reports, but is not limited thereto.

[0131] Step 305: The computer device searches the database for the first interpretation content corresponding to each target variant site based on the identification information corresponding to each target variant site.

[0132] The computer device can determine the associated historical reports related to the first target mutation site from multiple historical reports in the database based on the identification information corresponding to each target mutation site, and search for the report content corresponding to each first target mutation site in the associated historical reports as the first interpretation content corresponding to the first target mutation site.

[0133] The first target variant site can be any target variant site. The first interpretation content may include the target variant site and the report content corresponding to multiple historical reports.

[0134] Figure 4 This is a flowchart illustrating the process of finding the first interpretation content corresponding to each target variant site, as provided in an embodiment of this application. Figure 4 As shown, step 305 may include the following steps:

[0135] Step 402: The computer device selects historical reports containing the first target mutation site from multiple historical reports stored in the database based on the identification information corresponding to the first target mutation site, and uses them as associated historical reports.

[0136] As one implementation method, the computer device can select historical reports with a third report count greater than or equal to the corresponding threshold from multiple historical reports stored in the database based on the third report count corresponding to the first target mutation site, and use these as associated historical reports.

[0137] For example, the database may store five historical reports. The number of third reports corresponding to variant site C and the five historical reports are 1, 20, 4, 25 and 8 respectively. The computer device can determine that the historical reports with the number of third reports of 20 and 25 are the historical reports associated with variant site C.

[0138] As another implementation, the database may store the association between multiple variant sites and historical reports involving the variant sites. If the first target variant site is one of the multiple variant sites stored in the database, the computer device may directly determine the associated historical reports related to the first target variant site.

[0139] Step 404: The computer device extracts the report content corresponding to the first target variant site from each associated historical report, and uses it as the first interpretation content corresponding to the first target variant site.

[0140] The report content corresponding to the first target variant site may include text, data, tables, and images associated with the first target variant site.

[0141] The computer device can extract the report content of the paragraph where the first target variant site is located from each associated historical report, and integrate the report content corresponding to the first target variant site in each associated historical report into the first interpretation content corresponding to the first target variant site.

[0142] Optionally, the computer device can sort the position of the report content of each associated historical report in the first interpretation content based on the identification information of the first target variant site and the number of third reports in each associated historical report. The more third reports there are, the higher the position of the report content of the associated historical report in the first interpretation content.

[0143] By extracting the first interpretation content corresponding to each target variant site in the above manner, computer devices can directly search for related historical reports associated with the target variant site, avoiding the inefficiency caused by searching all historical reports. Furthermore, the first interpretation content extracted from the related historical reports has a higher correlation with the target variant site, which can effectively reduce the mixing of irrelevant information and improve the reference value of the first interpretation content.

[0144] In some embodiments, after obtaining the first interpretation content corresponding to each target variant site, when the computer device displays the association information corresponding to each target variant site, the computer device may display the number of reports, clinical evidence level and first interpretation content corresponding to each target variant site, and display the number of reports and clinical evidence level corresponding to each variant site that does not belong to the target variant site.

[0145] In this embodiment, the computer device obtains the number of reports and the level of clinical evidence corresponding to each variant site, and identifies the variant sites with a number of reports greater than or equal to the number threshold as target variant sites. Based on the identification information corresponding to each target variant site, the computer device searches for the first interpretation content corresponding to each target variant site in the database. This not only avoids the waste of computing power and efficiency loss caused by the computer device performing indiscriminate retrieval on all variant sites, but also significantly improves the accuracy and processing efficiency of the first interpretation content search. Furthermore, by distinguishing target variant sites by the number of reports, the target variant sites can be supported by quantitative and / or qualitative data, thereby improving the relevance and scientific validity of the first interpretation content.

[0146] In some embodiments, after obtaining the first interpretation content corresponding to each target variant site, the computer device can generate a first hyperlink based on the report location information corresponding to the first target variant site and the first associated historical report, and the file access path corresponding to the first associated historical report. The first hyperlink is used to access the report content of the first target variant site in the first associated historical report. The first associated historical report is any associated historical report, and the report location information is used to indicate the paragraph position of the report content of the first target variant site in the first associated historical report.

[0147] Optionally, the computer device can associate the first hyperlink with the report content corresponding to the first target variant site and the first associated historical report in the first interpretation content. By clicking on the report content corresponding to the first target variant site and the first associated historical report, the user can open and jump to the paragraph position of the report content of the first target variant site in the first associated historical report.

[0148] It should be noted that since the first interpretation content is extracted from the first association history report containing the identification information of the first target variant site, the report content may not be complete. When the user wants to view the complete association history report, by clicking the first hyperlink, the computer device can respond to the trigger operation of the first hyperlink, open the corresponding first association history report file, and jump to the corresponding paragraph position, avoiding the tedious steps of the user having to manually open the first association history report file and improving the user's interpretation efficiency.

[0149] In some embodiments, when a computer device displays the association information and / or variation information corresponding to each variant site, due to the large amount of information, numerous technical terms, and high data complexity, the computer device needs to intelligently interpret the variant sites to facilitate users' quick understanding of the variation status of the variant sites. Figure 5 This is a flowchart illustrating the interpretation of the first variant site, provided as an embodiment of this application. Figure 5 As shown, the method may further include the following steps:

[0150] Step 501: In response to the interpretation operation for the first variant site, the computer device analyzes the first interpretation content corresponding to the first variant site through a content analysis model to obtain the second interpretation content.

[0151] The first variant site can be any variant site. The content analysis model can be a natural language processing model based on artificial intelligence. The second interpretation content can be a detailed interpretation and summary of the first interpretation content.

[0152] Because the first interpretation usually contains a lot of theoretical information, which is highly specialized and loosely structured, it is difficult for users to efficiently filter out information directly related to the first variant site when reading the first interpretation. Therefore, the second interpretation can help users quickly remove irrelevant content from the first interpretation, quickly grasp the core features of the first variant site in the first interpretation, and obtain the degree of association between the first variant site and the disease, thereby reducing the time users spend sorting through the first interpretation and improving the efficiency of users' analysis of sequencing data.

[0153] Optionally, the computer device can analyze the first interpretation content and variation information corresponding to the first variant site through a content analysis model to obtain the second interpretation content corresponding to the first variant site. The second interpretation content can be a detailed interpretation and summary of the first interpretation content and variation information.

[0154] In some embodiments, the computer device may display a viewing interaction identifier corresponding to the first variant site. In response to a click operation on the viewing interaction identifier corresponding to the first variant site, the computer device may analyze the first interpretation content corresponding to the first variant site through a content analysis model to obtain the second interpretation content.

[0155] For example, as shown in Table 2, the computer device can also display an "AI Interpretation" interactive icon corresponding to each first variant site. When a user needs to quickly interpret the first variant site, they can click on the "AI Interpretation" interactive icon. In response to the interpretation operation for the first variant site, the computer device can analyze the first interpretation content and variant information corresponding to the first variant site through a content analysis model to obtain the second interpretation content corresponding to the first variant site.

[0156] Optionally, the content analysis model can be deployed locally, and the computer device can input the first interpretation content corresponding to the first variant site into the content analysis model so that the content analysis model can analyze the first interpretation content to obtain the second interpretation content; or, the content analysis model can be deployed on a cloud server, and the computer device can send the first interpretation content to the cloud server. The cloud server can then analyze the first interpretation content through the content analysis model to obtain the second interpretation content and send the second interpretation content to the computer device.

[0157] Step 503: The computer device generates a content summary based on the second interpretation content and displays the content summary.

[0158] The computer device can respond to the interpretation operation targeting the first variant site by analyzing the first interpretation content through a content analysis model to obtain the second interpretation content and the corresponding content summary, and display the content summary.

[0159] A summary refers to the extraction and concise summary of the core information from the second interpretation. Through this summary, users can quickly grasp the key points and essential information of the first variant site without having to read the complete second interpretation, thus gaining a rapid and efficient understanding of the first variant site.

[0160] Optionally, when displaying the content summary, the computer device may also display relevant literature corresponding to the first variant site and recommended follow-up experiments. Recommended follow-up experiments may refer to subsequent experimental procedures suggested to verify the authenticity, molecular characteristics, and therapeutic effects of the first variant site.

[0161] Figure 6 This is a schematic diagram illustrating the display of a content summary provided in an embodiment of this application. After a user clicks the "AI Interpretation" interactive icon corresponding to the fusion gene MBP::CTDP1, the computer device can respond to the interpretation operation for the fusion gene MBP::CTDP1, analyze the first interpretation content and variation information corresponding to the fusion gene MBP::CTDP1 through a content analysis model, obtain the second interpretation content and content summary, and display it as shown below. Figure 6 The image shows a summary of the content corresponding to the fusion gene MBP::CTDP1 and recommended follow-up experiments.

[0162] Step 505: In response to the triggered viewing operation, the computer device displays the second interpretation content.

[0163] The computer device may prioritize displaying a summary of the second interpretation content, and when the user needs to view the complete second interpretation content, the computer device may respond to the triggered viewing operation and display the second interpretation content.

[0164] In some embodiments, the computer device may display a first viewing identifier while displaying a content summary. The first viewing identifier can be used to display second interpreted content. The computer device may display the second interpreted content in response to a viewing operation on the first viewing identifier.

[0165] As shown in Table 2, the computer device can display a "View" interaction icon while displaying the content summary. Users can click the "View" interaction icon when they need to view the second interpretation of the fusion gene MBP::CTDP1. The computer device can respond to the click operation on the "View" interaction icon and display the second interpretation.

[0166] In this embodiment, the computer device, in response to an interpretation operation targeting the first variant site, analyzes the first interpretation content corresponding to the first variant site using a content analysis model to obtain second interpretation content. A content summary is generated and displayed based on the second interpretation content. When a user needs to view the second interpretation content, the computer device can respond to a triggered viewing operation and display the second interpretation content. This not only effectively filters and organizes the first interpretation content through the content analysis model, making the second interpretation content more accurate, organized, and targeted, but also allows users to quickly obtain key interpretation information about the first variant site without needing to browse the complete second interpretation content, improving information acquisition efficiency. Furthermore, by supporting users to further view the second interpretation content through a triggered viewing operation, it also meets the user's need for complete second interpretation content, optimizing the user experience when obtaining interpretation information during the analysis of the first variant site.

[0167] In some embodiments, the sequencing data also includes the number of fragments corresponding to one or more genes in the sample to be analyzed. Figure 7 A flowchart showing the expression results of each gene is provided for embodiments of this application. For example... Figure 7 As shown, the method may further include the following steps:

[0168] Step 701: The computer device acquires the control sequencing data of the control group sample corresponding to the sample to be analyzed.

[0169] Control group samples can be normal samples, i.e., samples from healthy individuals who do not have the target study disease or have not received specific treatment. Control sequencing data from control group samples can reflect the baseline level of gene expression under normal physiological conditions.

[0170] Optionally, there are multiple control group samples, so the computer device can acquire control sequencing data corresponding to each of the multiple control group samples.

[0171] Step 703: The computer device determines the expression result of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data.

[0172] The fragment number can refer to the number of sequencing fragments, while the fragment number of the first gene can refer to the total number of valid sequencing fragments in the sequencing data that can be accurately mapped to the sequence region of the first gene. The first gene can be any gene in the sequencing data of the sample to be analyzed.

[0173] Control sequencing data may include the number of control fragments for each gene. The genes involved in the sequencing data of the sample to be analyzed are the same as those involved in the control sequencing data. It is important to note that only by comparing the number of fragments for the same gene can the expression difference of that gene between the sample to be analyzed and the control sample be accurately reflected, avoiding invalid comparisons due to mismatched gene types.

[0174] The computer equipment can determine the first expression level of the first gene based on the number of fragments corresponding to the first gene in the sequencing data of the sample to be analyzed, and determine the control expression level of the first gene based on the number of control fragments corresponding to the first gene in the control sequencing data. Thus, the expression result of the first gene can be determined based on the first expression level and the control expression level.

[0175] Expression level can refer to gene expression level, which can be the abundance of the first gene transcribed into mRNA. Gene expression level can usually be quantified by the number of sequencing fragments detected by high-throughput sequencing technology.

[0176] Expression results can include upregulation, downregulation, and normal states. Upregulation means that, provided the difference between the expression level of the first gene and the control expression level is statistically significant, the expression level of the first gene is higher than that of the control. Downregulation means that, provided the difference between the expression level of the first gene and the control expression level is statistically significant, the expression level of the first gene is lower than that of the control. Normal means that there is no statistically significant difference between the expression level of the first gene and the control expression level.

[0177] In some embodiments, the control sequencing data includes the control expression level and normal expression level range for each gene, and the expression results for each gene include the first expression result and the second expression result.

[0178] Figure 8 This is a flowchart illustrating how the expression result of a first gene is determined based on the number of fragments of the first gene and control sequencing data, as provided in this embodiment of the application. Figure 8 As shown, step 703 may include the following steps:

[0179] Step 802: The computer device calculates the first expression level corresponding to the first gene based on the number of fragments corresponding to the first gene.

[0180] Standardized metrics for expression levels may include Reads Per Kilobase of transcript per Million mapped reads (RPKM), Fragments Per Kilobase of transcript per Million mapped reads (FPKM), and Transcripts Per Kilobase of exonmodel per Million mapped reads (TPM).

[0181] Since one sequencing fragment in single-end sequencing usually corresponds to one sequencing read, while one sequencing fragment in paired-end sequencing usually corresponds to two sequencing reads, RPKM is suitable for representing gene expression levels in single-end sequencing data, FPKM is suitable for representing gene expression levels in paired-end sequencing data, and TPM is suitable for representing gene expression levels in either single-end or paired-end sequencing data.

[0182] For example, when a computer device acquires sequencing data of a sample to be analyzed using paired-end sequencing technology, the first expression level of the first gene can be represented by FPKM, which can be calculated using the following formula:

[0183]

[0184] Where C represents the number of sequencing fragments that were aligned to the first gene, R represents the base length of the first gene (usually the total length of the exons of the first gene), and L represents the total number of sequencing fragments that were successfully aligned to the standard group sample in the sequencing data of the sample to be analyzed. The standard group sample can be the human genome reference set.

[0185] The computer device can calculate the first expression level of the first gene in the sample to be analyzed according to equation (1). Similarly, the computer device can also calculate the control expression level of the first gene in multiple control group samples according to equation (1).

[0186] Step 804: The computer device compares the first expression level of the first gene with the normal expression level range corresponding to the first gene to determine the first expression result corresponding to the first gene.

[0187] The normal expression range can be defined as the range of expression fluctuations that reflects the normal physiological function of the first gene. When the expression level of the first gene is within the normal expression range, the expression result of the first gene can be considered normal; when the expression level of the first gene is greater than the upper limit of the normal expression range, the expression result of the first gene can be considered upregulated; and when the expression level of the first gene is less than the upper limit of the normal expression range, the expression result of the first gene can be considered downregulated.

[0188] In some embodiments, the normal expression range can be calculated from the control expression levels corresponding to multiple control group samples.

[0189] Computer equipment can obtain the median of the control expression level for each of the multiple control group samples. Then, based on the control expression level and its median for each of the multiple control group samples, the standard deviation of the control expression level is determined. Based on the standard deviation and the median, the normal expression level range is determined.

[0190] For example, a computer device can determine the standard deviation corresponding to the control expression level using the following formula:

[0191]

[0192] Where SD represents the standard deviation of the control expression level, and X represents the standard deviation of the control expression level. i Let represent the control expression level of the i-th control group sample, i = 1, 2, ..., n, where n represents the total number of control group samples, and Median represents the median of the control expression levels among the n control group samples.

[0193] Computer equipment can determine the normal expression range as [Median-3SD, Median+3SD] based on the standard deviation and median of the control expression level.

[0194] Step 806: The computer device calculates the expression difference parameter and verification value based on the first expression level corresponding to the first gene and the corresponding control expression level, and determines the second expression result corresponding to the first gene based on the expression difference parameter and verification value.

[0195] The expression difference parameter can be used to measure the difference in the expression level of the first gene between the sample being analyzed and the control sample. If the expression difference parameter is greater than 0, it indicates that the first gene is upregulated in the sample being analyzed; if the expression difference parameter is less than 0, it indicates that the first gene is downregulated in the sample being analyzed; if the expression difference parameter is equal to 0, it indicates that there is no difference in the expression level of the first gene between the sample being analyzed and the control sample.

[0196] For example, a computer device can determine the expression difference parameter by the following formula:

[0197] F=log2(Foldchange)=log2(F1)-log2(F2) Formula (3);

[0198] Where F is the expression difference parameter, Foldchange is the fold change in expression, F1 is the first expression level of the first gene in the sample to be analyzed, and F2 is the control expression level of the first gene in the control group sample.

[0199] It is understandable that the fold change in expression is equal to F1 / F2. By taking a logarithmic transformation of the fold change in expression to base 2, the magnitude of upregulation or downregulation of the expression level of the first gene can be more intuitively displayed, thereby more accurately determining the expression result of the first gene.

[0200] The validation value can be the adjusted P-value (adj.p), which can be the original probability value obtained from a single hypothesis test. The validation value can be obtained through multiple test correction, which may include the Benjamini-Hochberg method. Since obtaining the validation value through multiple test correction is existing technology, it will not be elaborated upon here.

[0201] If adj.p < 0.05, it indicates that the expression difference of the first gene between the sample to be analyzed and the control group is statistically significant after multiple test correction; while if adj.p ≥ 0.05, it indicates that the expression difference of the first gene is not statistically significant after multiple test correction, and is insufficient to determine that the expression difference of the first gene is a real difference.

[0202] In some embodiments, a computer device may determine the second expression result corresponding to the first gene based on expression difference parameters and check values.

[0203] If the expression difference parameter is greater than 0 and the check value adj.p < 0.05, the expression result of the second gene corresponding to the first gene is determined to be upregulated; if the expression difference parameter is less than 0 and the check value adj.p < 0.05, the expression result of the second gene corresponding to the first gene is determined to be downregulated; if the check value adj.p ≥ 0.05, the expression result of the second gene corresponding to the first gene is determined to be normal.

[0204] It should be noted that the determination of upregulation or downregulation cannot rely solely on numerical differences. It is also necessary to confirm that the expression difference is not a random error or experimental noise. Therefore, after determining whether the difference between the first expression level of the first gene and the control expression level is upregulation or downregulation based on the expression difference parameter, it is also necessary to determine whether the expression difference is a real biological change based on the check value. Only when the direction of the difference is clear (expression difference parameter is not 0) and the difference is reliable (adj.p < 0.05) can the second expression result corresponding to the first gene be finally determined to be upregulation or downregulation.

[0205] By comparing the first expression level of the first gene with the normal expression level range, the first expression result of the first gene is determined. Then, based on the first expression level of the first gene and the corresponding control expression level, expression difference parameters and verification values ​​are calculated. Finally, based on these parameters and verification values, the second expression result of the first gene is determined. This approach allows for a quick and intuitive assessment of whether the first gene is within the physiologically normal expression range, starting from the deviation of the first gene from the baseline level of a healthy population. Furthermore, by combining expression difference parameters and verification values, it avoids misinterpreting accidental experimental errors as genuine biological differences in inter-group comparisons, ensuring the reliability and scientific validity of the expression results. This leads to more accurate determination of the expression results and improves the efficiency and accuracy of users interpreting sequencing data.

[0206] Step 705: The computer device displays the expression results corresponding to each gene.

[0207] In some embodiments, the computer device may display the first expression result and / or the second expression result corresponding to each gene.

[0208] Computer equipment can display only the first expression result or the second expression result for each gene, or it can display the first expression result and the second expression result for each gene simultaneously. This can avoid the information limitations that may be caused by a single expression result, thus providing users with a more comprehensive reference for analyzing sequencing data.

[0209] In some embodiments, the computer device may determine the target expression result corresponding to the first gene based on the first expression result and the second expression result corresponding to the first gene, and display one or more of the target expression result, the first expression result and the second expression result corresponding to the first gene.

[0210] If the first expression result and the corresponding second expression result of the first gene are the same, the target expression result of the first gene is determined to be either the first expression result or the second expression result; if the first expression result and the corresponding second expression result of the first gene are different, the computer device can determine that the expression result of the first gene is unreliable and does not display the target expression result of the first gene.

[0211] Determining the target expression result by using the first and second expression results ensures the reliability and accuracy of the target expression result, avoids misleading users due to potential biases or errors in a single expression result, and helps users quickly focus on reliable gene expression information, reducing the complexity and cost of interpreting sequencing data, thereby improving the accuracy and efficiency of analyzing the first gene expression status.

[0212] In some embodiments, the computer device can simultaneously display the first expression level, normal expression level range, expression difference parameters, check value, first expression result, second expression result, and target expression result corresponding to the first gene, so that users can obtain the basis for each expression result more intuitively and improve the efficiency of users' analysis of sequencing data.

[0213] As shown in Table 3, the computer device can simultaneously display the FPKM (first expression level), normal expression range [Median-3SD, Median+3SD], expression difference parameter log2 (Foldchange), checksum judgment result adj.p < 0.05, result 1 (i.e., first expression result), result 2 (i.e., second expression result), and expression result (i.e., target expression result) for multiple genes. When results 1 and 2 for genes PRAME and WT1 are the same, the computer device will display the target expression result for genes PRAME and WT1 as "normal". However, when results 1 and 2 for genes CRLF2, FLT3, MYC, MECOM, and TP53 are different, the corresponding target expression result will be displayed as "N / A", indicating that the target expression results for genes with different results 1 and 2 are unreliable.

[0214] Table 3

[0215]

[0216] In this embodiment, the computer device acquires the control sequencing data of the control group sample corresponding to the sample to be analyzed. Based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data, the expression results of each gene are determined and displayed. This allows the expression results of the sample to be analyzed to be obtained by comparing with the genes of the control group sample, avoiding the bias or error that may occur when analyzing the sample alone. This improves the accuracy and reliability of the gene expression results of the sample to be analyzed. Furthermore, by displaying the expression results of each gene, users can obtain the expression results of each gene more intuitively and clearly, making it easier to quickly identify genes with abnormal expression, thereby improving the overall efficiency of sequencing data analysis.

[0217] In some embodiments, the computer device may obtain one or more of the following: the number of chromosomes entered by the user, the transcript number of the target gene, and the rearrangement type corresponding to the immunoglobulin / T cell receptor (IG / TCR).

[0218] In one implementation method, a computer device can determine whether a sample to be analyzed has CNV based on the number of chromosomes entered by the user and the number of chromosomes involved in the sequencing data.

[0219] The number of chromosomes entered by the user can be the same as the number of chromosomes in a normal human body. If the number of chromosomes entered is the same as the number of chromosomes involved in the sequencing data, it is determined that the sample to be analyzed has not experienced CNV; conversely, if the number of chromosomes entered is different from the number of chromosomes involved in the sequencing data, it is determined that the sample to be analyzed has experienced CNV.

[0220] For example, if the user enters 46 chromosomes, and the sequencing data of the sample to be analyzed contains only 30 chromosomes, then it is determined that the sample to be analyzed has experienced CNV, and the number of chromosomes in the sequencing data of the sample to be analyzed has deviated significantly.

[0221] As another implementation method, the computer device can determine the number of the variant transcript with a specific region deletion in the target gene based on the number of the normal transcript without deletion in the target gene entered by the user, thereby obtaining the detection threshold and relative content ratio of each variant transcript, and displaying the detection threshold, relative content ratio and detection result of each variant transcript; the target gene is any gene in the sequencing data.

[0222] The relative abundance ratio can be the proportion of variant transcripts in the total number of transcripts, reflecting the relative abundance of variant transcripts.

[0223] The detection result can include detected and undetected. When the relative content of a variant transcript is greater than or equal to the detection threshold, the detection result is detected, and the computer equipment can determine that the variant corresponding to that variant transcript exists in the target gene; when the relative content is less than the detection threshold, the detection result is undetected, and the computer equipment can determine that the variant corresponding to that variant transcript does not exist in the target gene.

[0224] As shown in Table 4, the target gene can be IKZF1. The computer device can obtain four variant transcripts corresponding to the normal transcripts of the gene IKZF1 that have not been deleted, based on the normal transcripts entered by the user. The device can also obtain the cutoff (detection threshold) and ratio (relative content ratio) of each variant transcript. Based on the comparison results between the cutoff and ratio of each variant transcript, the device can determine that the detection result of each variant transcript is "not detected", that is, there are no variants corresponding to the above four variant transcripts in the gene IKZF1.

[0225] Table 4

[0226] IKZF1 transcript cutoff ratio Detection status Selected? Normal transcripts - - - □ Variant transcript 1 0.532 0.11 Not detected □ Variant Transcript 2 0.532 0.0 Not detected □ Variant Transcript 3 0.532 0.055 Not detected □ Variant Transcript 4 0.532 0.0 Not detected □

[0227] As another implementation method, the computer device can obtain the gene sequence and cloning frequency of the Complementarity Determining Region 3 (CDR3) corresponding to the rearrangement type from the sequencing data according to the rearrangement type.

[0228] Rearrangement types can include immunoglobulin heavy locus (IGH) rearrangement, immunoglobulin light chain κ or λ rearrangement, etc.

[0229] It should be noted that, since there is a fixed correspondence between the rearrangement type and the gene sequence region corresponding to CDR3, after the computer device obtains the rearrangement type entered by the user, it can extract the gene sequence corresponding to CDR3 from the gene sequence corresponding to the sequencing data of the sample to be analyzed based on the gene sequence region corresponding to CDR3 corresponding to the rearrangement type entered by the user, and calculate the corresponding cloning frequency.

[0230] In some embodiments, after obtaining the gene sequence and cloning frequency corresponding to CDR3, the computer device can display the gene sequence and cloning frequency corresponding to CDR3 in a table where the user enters the IG / TCR rearrangement type.

[0231] As shown in Table 5, after the user enters the rearrangement type of IG / TCR as IGH, the computer device can extract the gene sequence corresponding to CDR3 and the corresponding cloning frequency from the gene sequence corresponding to the sequencing data of the sample to be analyzed, and display the gene sequence corresponding to CDR3 and the corresponding cloning frequency of 7.86%.

[0232] Table 5

[0233] No. Rearrangement type CDR3 sequence Cloning frequency Selected? 1 IGH …… 7.86% □

[0234] Optionally, after entering IGH, users can continue to enter other rearrangement types of IG / TCR. The computer device can obtain the gene sequences corresponding to CDR3 and the corresponding cloning frequencies of other rearrangement types entered by the user.

[0235] The above methods enable computer equipment to automatically obtain relevant information from user input, avoiding the tedious manual searching and filling in of information, thus improving the accuracy and generation efficiency of sequencing reports. In addition, users can choose whether to input the number of chromosomes and / or the rearrangement type corresponding to IG / TCR based on the specific first interpretation content and clinical needs, thereby precisely controlling the content of the sequencing report and improving the customizability of the sequencing report.

[0236] In some embodiments, the computer device may generate a second hyperlink corresponding to the first mutation site based on the mutation information corresponding to each mutation site. The second hyperlink is used to jump to a gene visualization tool and display the gene sequence in the gene sequence file corresponding to the first mutation site, which corresponds to the mutation position of the first mutation site.

[0237] The gene visualization tool can be either a desktop application (desktop version) or a web-based application of the Integrative GenomicsViewer (IGV). The IGV desktop application is standalone software installed on a local computer. It directly utilizes local hardware resources for sequencing data processing and visualization, enabling rapid loading and analysis of gene sequences at specific variant sites. The IGV web application is a genome browser built on the IGV.js visualization engine. Users can access the IGV web application through a browser to quickly view gene sequences at specific variant sites.

[0238] The method of jumping to a gene visualization tool via the second hyperlink to display the gene sequence corresponding to the mutation location of the first mutation site is existing technology and will not be elaborated here.

[0239] In some embodiments, the gene visualization tool may be a web-based IGV application. The computer device can implement the following interactive methods in response to triggered interactive actions:

[0240] Method 1: When displaying gene sequences via the IGV web interface, the computer device can, in response to the triggered first interactive operation, display sequencing fragments in the sequencing data that meet the target conditions. The target conditions may include one of the following: the sequencing depth corresponding to each base in the sequencing fragment is less than the depth threshold, the sequencing depth corresponding to any base in the sequencing fragment is less than the depth threshold, or the average sequencing depth corresponding to the sequencing fragment is less than the depth threshold.

[0241] The depth threshold can be set according to actual needs. For example, the depth threshold can be set to 50, and the computer device can only display sequencing fragments with an average sequencing depth of less than 50.

[0242] It should be noted that computer devices can also respond to triggered interactive operations and only display sequencing fragments that do not meet the target conditions.

[0243] By setting the above target conditions and selecting sequencing fragments that meet or do not meet the target conditions as needed, sequencing fragments can be screened on demand, thereby filtering out irrelevant sequencing fragments to reduce redundant interference, adapting to different analysis needs, and improving the user's analysis efficiency.

[0244] Method 2: When displaying gene sequences via the IGV web interface, the computer device can, in response to a triggered second interactive operation, display only the sequencing fragments belonging to the positive strand or the sequencing fragments belonging to the reverse strand in the sequencing data.

[0245] The computer device can also display double-stranded sequencing fragments in response to a triggered second interactive operation, whether the sequencing fragments belong to the positive strand or the reverse strand.

[0246] By selectively displaying sequencing fragments belonging to the positive strand or the reverse strand, information redundancy and visual interference when dual-strand data are presented simultaneously can be effectively reduced, which helps improve the efficiency and accuracy of data interpretation when users analyze gene sequences.

[0247] Method 3: When the gene sequence is displayed via the IGV web interface, the computer device can respond to the triggered third interactive operation and display the mutation location of each mutation site in a specific manner.

[0248] Specific methods may include highlighting in red. For example, for a variant site of type SNV, the computer device can highlight the mutated base in red; for a variant site of type gene fusion, after enabling the display of splice linkages, the computer device can highlight the splice site corresponding to the variant site in red, facilitating the analysis of splice events in sequencing data.

[0249] Alternatively, computer devices may also display the location of each variant site in other specific ways, without specific limitations here.

[0250] By displaying the mutation location of each mutation site in a specific way, the mutation location of the mutation site is made more intuitive and obvious, which can help users quickly locate the mutation location, effectively reduce the time cost of manual search and identification, reduce the analysis error caused by site concealment, and help improve the analysis efficiency and accuracy of sequencing data.

[0251] In some embodiments, the computer device may acquire clinical information corresponding to the sample to be analyzed.

[0252] Clinical information may include patient information, physician information, and sampling information of the sample to be analyzed. Users can upload the files corresponding to the clinical information to the computer device. After the computer device obtains the files, it parses them to extract the clinical information corresponding to the sample to be analyzed.

[0253] In some embodiments, after acquiring the clinical information corresponding to the sample to be analyzed, the computer device may generate a sequencing report of the sequencing data for the sample to be analyzed in response to a triggered generation operation.

[0254] Since the automatic generation of sequencing reports from sequencing data of the sample to be analyzed using computer equipment is an existing technology, it will not be elaborated here.

[0255] In some embodiments, the computer device may, in response to a triggered selection operation, determine the variant sites for generating a sequencing report, and, in response to a triggered generation operation, generate a sequencing report for the sequencing data of the sample to be analyzed based on one or more variant sites selected by the user, as well as clinical information, expression results of each gene in the sequencing data, detection of multiple variant transcripts of the target gene, and gene sequences and cloning frequencies corresponding to IG / TCR rearrangement types, etc., as displayed by the computer device in the above embodiments.

[0256] As shown in Tables 4 and 5, users can select the detection status of variant transcripts and the gene sequences and cloning frequencies corresponding to IG / TCR rearrangement types for generating sequencing reports in the "Selected" column. Furthermore, in the tables shown in Tables 1 to 3, the last column can be set to "Selected," where users can select the variant sites and gene expression results for generating sequencing reports.

[0257] By generating a standardized and complete sequencing report with a single click based on the user's selected display content, the sequencing report generation process is fully automated. This avoids the problems of inconsistent formats and non-standard content in traditional manual operations, improving the efficiency and accuracy of sequencing report generation. Furthermore, generating sequencing reports based on the selected display content also helps to enhance the customizability of sequencing reports.

[0258] In some embodiments, after generating a sequencing report, the computer device can record the sequencing report and the various variant sites involved in the sequencing report into a database, so that when analyzing the sequencing data of a new sample to be analyzed next time, the above embodiments can be repeated to quickly retrieve the report content of the historical report for auxiliary analysis of the sequencing data of the new sample to be analyzed.

[0259] In some embodiments, the computer device may display operation information corresponding to each report in response to a triggered traceability operation. The operation information may include operator identification, timestamp, and data revision version.

[0260] Operational information can reflect the operational actions, responsible parties, and details of data changes in the entire process of creating and publishing the sequencing report corresponding to the sequencing data.

[0261] Computer equipment can use blockchain-based evidence storage technology to record operational information throughout the entire process of creating and publishing each sequencing report, ensuring that the operational information is tamper-proof, verifiable, and provides a reliable basis for subsequent traceability and verification.

[0262] In some embodiments, the computer device may respond to a query operation with an input report number by displaying all operation information of the sequencing report corresponding to the input report number, so as to facilitate users to obtain the complete revision history of the sequencing report and ensure the traceability and quality control of the sequencing report.

[0263] The sequencing data analysis method provided in the above embodiments Figure 9 This is a structural block diagram of a sequencing data analysis device provided in an embodiment of this application. Figure 9As shown, in one embodiment, a sequencing data analysis device 900 is provided, which can be applied to a computer device. The sequencing data analysis includes a data acquisition module 901, a data analysis module 902, and a data display module 903.

[0264] The data acquisition module 901 is used to acquire sequencing data of the sample to be analyzed. The sequencing data includes one or more variant sites and the variant information corresponding to each variant site. The variant information includes the identification information corresponding to the variant site.

[0265] The data analysis module 902 is used to search for the associated information of each variant site in the database based on the identification information corresponding to each variant site. The associated information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site.

[0266] The data display module 903 is used to display the association information and / or variation information corresponding to each variant site.

[0267] In some embodiments, the associated information includes the number of times the variant site is reported, the level of clinical evidence for the variant site, and the first interpretation of the variant site.

[0268] The data analysis module 902 is also used to search the database for the number of reports and the level of clinical evidence corresponding to each variant site based on the identification information corresponding to each variant site; to identify the variant sites among the variant sites whose number of reports is greater than or equal to the number threshold as target variant sites; and to search the database for the first interpretation content corresponding to each target variant site based on the identification information corresponding to each target variant site.

[0269] In some embodiments, the data analysis module 902 is further configured to, based on the identification information corresponding to the first target mutation site, select historical reports containing the first target mutation site from multiple historical reports stored in the database as associated historical reports; the first target mutation site can be any target mutation site; and extract the report content corresponding to the first target mutation site from each associated historical report as the first interpretation content corresponding to the first target mutation site.

[0270] In some embodiments, the sequencing data analysis apparatus 900 further includes a link generation module.

[0271] The link generation module is used to generate a first hyperlink based on the report location information corresponding to the first target variant site and the first associated historical report, as well as the file access path corresponding to the first associated historical report. The first hyperlink is used to access the report content of the first target variant site in the first associated historical report. The first associated historical report is any associated historical report, and the report location information is used to indicate the paragraph position of the report content of the first target variant site in the first associated historical report.

[0272] In some embodiments, the sequencing data analysis apparatus 900 further includes a content interpretation module.

[0273] The content interpretation module is used to respond to the interpretation operation for the first variant site by analyzing the first interpretation content corresponding to the first variant site through the content analysis model to obtain the second interpretation content; the first variant site can be any variant site.

[0274] The content interpretation module is also used to generate a content summary based on the second interpretation content and to display the content summary.

[0275] The data display module 903 is also used to display the second interpretation content in response to the triggered viewing operation.

[0276] In some embodiments, the data acquisition module 901 is further configured to acquire control sequencing data of the control group sample corresponding to the sample to be analyzed.

[0277] The data analysis module 902 is also used to determine the expression results of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data.

[0278] The data display module 903 is also used to display the expression results of each gene.

[0279] In some embodiments, the control sequencing data includes the control expression level and normal expression level range for each gene, and the expression results for each gene include the first expression result and the second expression result.

[0280] The computer equipment determines the expression results of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data, including:

[0281] The data analysis module 902 is also used to calculate the first expression level of the first gene based on the number of fragments corresponding to the first gene, where the first gene is any gene in the sample to be analyzed; compare the first expression level of the first gene with the normal expression level range corresponding to the first gene to determine the first expression result of the first gene; calculate the expression difference parameter and verification value based on the first expression level of the first gene and the corresponding control expression level, and determine the second expression result of the first gene based on the expression difference parameter and verification value.

[0282] The data display module 903 is also used to display the first expression result and / or the second expression result corresponding to the first gene when the first expression result and the corresponding second expression result are the same.

[0283] Figure 10 This is a structural block diagram of a computer device provided in an embodiment of this application. Figure 10 As shown, the computer device 1000 may include a memory 1002 and a processor 1001. The memory 1002 stores a computer program. When the computer program is executed by the processor 1001, the computer device 1000 implements the financial management method as described in the above embodiments.

[0284] Processor 1001 may include one or more processing cores. Processor 1001 connects to various parts within the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, and by calling data stored in memory. Optionally, processor 1001 may be implemented using at least one hardware form selected from digital signal processing, field-programmable gate arrays, and programmable logic arrays. Processor 1001 may integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 1001 and may be implemented separately using a communication chip.

[0285] The memory 1002 may include random access memory (RAM) or read-only memory (ROM). The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described above, etc. The data storage area may also store data created during the use of the electronic device.

[0286] This application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the sequencing data analysis method as described in the above embodiments.

[0287] This application discloses a computer program product, which includes a computer program, and when executed by a processor, causes the processor to implement the sequencing data analysis method described in the above embodiments.

[0288] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, ROM, etc.

[0289] The above description is merely a specific example of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A method for analyzing sequencing data, characterized in that, Applied to a computer device, the method includes: The computer device acquires sequencing data of the sample to be analyzed. The sequencing data includes one or more mutation sites and mutation information corresponding to each mutation site. The mutation information includes identification information corresponding to the mutation site. The computer device searches for the associated information corresponding to each of the variant sites in the database based on the identification information corresponding to each of the variant sites. The associated information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site. The computer device displays the association information and / or variation information corresponding to each of the mutation sites.

2. The method according to claim 1, characterized in that, The associated information includes the number of reported variant sites, the level of clinical evidence for the variant sites, and the first interpretation of the variant sites. The computer device searches for the associated information corresponding to each of the mutation sites in the database based on the identification information corresponding to each of the mutation sites, including: The computer device searches the database for the number of reports and the level of clinical evidence corresponding to each of the variant sites based on the identification information corresponding to each of the variant sites. The computer device identifies target variant sites among the various variant sites whose reported number is greater than or equal to a threshold number. The computer device searches the database for the first interpretation content corresponding to each of the target mutation sites based on the identification information corresponding to each of the target mutation sites.

3. The method according to claim 2, characterized in that, The computer device searches the database for the first interpretation content corresponding to each of the target variant sites based on the identification information corresponding to each of the target variant sites, including: The computer device, based on the identification information corresponding to the first target mutation site, filters out historical reports containing the first target mutation site from multiple historical reports stored in the database, and uses them as associated historical reports; the first target mutation site can be any of the target mutation sites. The computer device extracts the report content corresponding to the first target mutation site from each of the associated historical reports, and uses it as the first interpretation content corresponding to the first target mutation site.

4. The method according to claim 3, characterized in that, After the computer device extracts the report content corresponding to the first target variant site from each of the associated historical reports as the first interpretation content corresponding to the first target variant site, the method further includes: The computer device generates a first hyperlink based on the report location information corresponding to the first target mutation site and the first associated historical report, and the file access path corresponding to the first associated historical report. The first hyperlink is used to access the report content of the first target mutation site in the first associated historical report. The first associated historical report is any of the associated historical reports, and the report location information is used to indicate the paragraph position of the report content of the first target mutation site in the first associated historical report.

5. The method according to claim 1, characterized in that, After the computer device displays the association information and / or mutation information corresponding to each of the mutation sites, the method further includes: In response to the interpretation operation targeting the first variant site, the computer device analyzes the first interpretation content corresponding to the first variant site using a content analysis model to obtain the second interpretation content; the first variant site is any of the variant sites. The computer device generates a content summary based on the second interpretation content and displays the content summary; The computer device displays the second interpretation content in response to the triggered viewing operation.

6. The method according to claim 1, characterized in that, The sequencing data also includes the number of fragments corresponding to one or more genes in the sample to be analyzed; the method further includes: The computer device acquires the control sequencing data of the control group sample corresponding to the sample to be analyzed; The computer device determines the expression result of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data; The computer device displays the expression results corresponding to each gene.

7. The method according to claim 6, characterized in that, The control sequencing data includes the control expression level and normal expression level range for each gene, and the expression results for each gene include a first expression result and a second expression result. The computer device determines the expression results of each gene based on the number of fragments corresponding to each gene in the sample to be analyzed and the control sequencing data, including: The computer device calculates the first expression level of the first gene based on the number of fragments corresponding to the first gene, wherein the first gene is any one of the genes in the sample to be analyzed. The computer device compares the first expression level of the first gene with the normal expression level range corresponding to the first gene to determine the first expression result corresponding to the first gene. The computer device calculates expression difference parameters and verification values ​​based on the first expression level corresponding to the first gene and the corresponding control expression level, and determines the second expression result corresponding to the first gene based on the expression difference parameters and the verification values. The computer device displays the expression results corresponding to each gene, including: If the first expression result and the corresponding second expression result of the first gene are the same, the computer device displays the first expression result and / or the second expression result of the first gene.

8. An analysis device for sequencing data, characterized in that, Applied to computer equipment, the device includes: The data acquisition module is used to acquire sequencing data of the sample to be analyzed. The sequencing data includes one or more mutation sites and mutation information corresponding to each mutation site. The mutation information includes identification information corresponding to the mutation site. The data analysis module is used to search for the association information corresponding to each of the variant sites in the database based on the identification information corresponding to each of the variant sites. The association information includes one or more of the following: the number of reports of the variant site, the clinical evidence level of the variant site, and the first interpretation content of the variant site. The data display module is used to display the association information and / or variation information corresponding to each of the said mutation sites.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1 to 7.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 7.