A method for analyzing and displaying data in computational biology research

By correcting for the effects of gene length and sequencing depth using FPKM values, and combining the judgment of repetitive sequences and individual differences, the accuracy problems caused by repetitive sequences and individual variations in transcriptome data analysis have been solved, enabling accurate comparison and visualization of differences in gene expression levels.

CN120432018BActive Publication Date: 2026-03-24BURDOCK BIOTECHNOLOGY(DEZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In traditional transcriptome data analysis, short sequences are difficult to distinguish from their specific origin, leading to multiple alignment problems and reducing the accuracy of gene expression level difference data.

Method used

The FPKM value was used as a standardization indicator. The influence of repetitive sequences was eliminated and the results of gene expression level differences were corrected by the repetitive sequence region judgment strategy and the individual difference judgment strategy. The results of expression data difference analysis were visualized.

Benefits of technology

It improves the accuracy of data analysis results, ensures the accuracy of gene expression level comparisons, identifies and corrects errors caused by repetitive sequences and individual variations, and provides reliable data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432018B_ABST
    Figure CN120432018B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biological research data analysis, and discloses a kind of computing biology research data analysis display method, including extracting total RNA from biological tissue or cell, RNA reverse transcription is cDNA, constructs library, and utilizes high-throughput sequencing platform to sequence library, to obtain a large number of short sequences.The present application can effectively identify abnormal repeat regions by comparing and threshold judging the number of short sequences mapped to the same genomic coordinate interval, so as to determine whether the gene expression difference result is reliable, when the number of abnormal sequences exceeds the preset repeat sequence threshold, by eliminating the influence of repeat sequence operation, i.e. for the case of high FPKM value, remove the completely same and overlapping short sequences, and re-align the updated results, the expression distortion problem caused by repeat counting can be corrected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biological research data analysis technology, specifically a method for analyzing and displaying computational biology research data. Background Technology

[0002] Computational biology research data generally refers to data used in the field of computational biology to analyze, simulate, and interpret biological problems. This data often originates from experimental measurements but requires subsequent processing and interpretation using computational, statistical, and mathematical models. Examples include genomic data, transcriptomic data, proteomic data, metabolomic data, and structural biology data.

[0003] Transcriptome data, typically obtained through RNA sequencing, describes the expression levels of all genes in cells under different conditions. Analysis and visualization of transcriptome data can help understand changes in gene expression regulation, biological processes, and disease states. Traditional transcriptome data analysis involves comparing short sequences obtained under different conditions with a reference genome to analyze differences in gene expression levels. However, the genome contains numerous repetitive or similar sequences, and short sequences often have difficulty identifying their specific origin, leading to multiple alignment issues and unstable alignment results. This reduces the accuracy of gene expression level difference data. Summary of the Invention

[0004] To address the problems existing in the prior art, the present invention aims to provide a computational biology research data analysis and visualization method that can analyze research data when gene expression level differences are abnormal and further reduce the impact of repetitive sequence regions, thereby improving the accuracy of data analysis results visualization.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for analyzing and displaying computational biology research data, the method comprising the following steps:

[0006] Total RNA is extracted from biological tissues or cells, reverse transcribed into cDNA, a library is constructed, and the library is sequenced using a high-throughput sequencing platform to obtain a large number of short sequences.

[0007] Short sequence data are preprocessed. The FastQC tool is used to detect the quality fraction distribution, GC content, sequence duplication, and possible sequencing bias of the sequencing data in the raw data sample. Adapter contamination and low-quality regions are identified and removed, and the remaining short sequence data are marked as usable sequences.

[0008] The available sequence data is standardized to obtain its FPKM value. A reference genome is constructed and its FPKM value is obtained. The difference between the available sequence and the reference genome's FPKM value is calculated to obtain the gene expression difference results.

[0009] The magnitude of gene expression differences is obtained by analyzing gene expression differences. A threshold for the magnitude of the difference is set. When the magnitude of the difference is less than or equal to the threshold, the gene expression difference is marked as a data difference analysis result. When the magnitude of the difference is greater than the threshold, a repetitive sequence region judgment strategy is executed. This strategy is used to eliminate the influence of repetitive sequences when there are many repetitive sequences in the available sequences, so as to obtain an updated difference result and mark it as a data difference analysis result. The data difference analysis result is then visualized to achieve data presentation.

[0010] In some implementations, the repetitive sequence region determination strategy includes obtaining the number of short sequences located in the same region among the available sequences, marking them as abnormal sequences, setting a repetitive sequence threshold, comparing the number of abnormal sequences with the repetitive sequence threshold, and taking appropriate action based on the comparison result.

[0011] In some implementations, if the abnormal sequence is less than or equal to the repetitive sequence threshold, the obtained gene expression difference result is deemed reliable, marked as a data difference analysis result, and then visualized; if the abnormal sequence is greater than the repetitive sequence threshold, the operation to eliminate the influence of repetitive sequences is performed.

[0012] In some implementations, the operation to eliminate the influence of repetitive sequences specifically involves: determining the FPKM value of the available sequence and the FPKM value of the reference genome; if the FPKM value of the available sequence is greater than that of the reference genome, removing identical and overlapping short sequences within the same repetitive region, and re-aligning the available sequence with the reference genome; obtaining an updated difference result by calculating the difference between the FPKM values ​​of the available sequence and the reference genome, marking it as a data difference analysis result, and visualizing it; if the FPKM value of the available sequence is less than that of the reference genome, the obtained gene expression difference result should be deemed reliable, marked as a data difference analysis result, and subsequently visualized.

[0013] In some implementations, when the FPKM value of the available sequence is less than the FPKM value of the reference genome, an individual difference assessment is performed to determine whether individual variations exist in the genome of the sample.

[0014] In some implementations, the method for performing individual difference judgment is as follows: the available sequence is compared with the reference sequence in the reference genome to obtain the maximum similarity between the available sequence and the reference sequence in the reference genome, and a similarity threshold is set. The maximum similarity is compared with the similarity threshold to perform individual difference judgment. Available sequences with a maximum similarity less than the similarity threshold are marked as variant sequences. The total number of variant sequences is counted, a variant number threshold is set, and an appropriate response is made based on the comparison between the total number of variant sequences and the variant number threshold.

[0015] In some implementations, if the total number of variant sequences is less than the number of variants threshold, the obtained gene expression difference results are considered reliable, marked as data difference analysis results, and subsequently visualized. If the total number of variant sequences is greater than or equal to the number of variants threshold, it is determined that the genome of this sample has individual variation. Short sequences that were missed due to filtering are added to the existing data based on their quantity, and the available sequences are re-aligned with the reference genome. The difference between the available sequences and the reference genome is used to obtain the updated difference results, which are marked as data difference analysis results and visualized.

[0016] In some implementations, after adding short sequences that were missed due to filtering to the existing data based on their quantity, the number of repeating regions before adding the missed short sequences is obtained and marked as the number of primary repeating regions. The number of repeating regions after adding the missed short sequences is obtained and marked as the number of secondary repeating regions. The relationship between the number of primary repeating regions and the number of secondary repeating regions is determined and a corresponding response is taken.

[0017] In some implementations, when the number of secondary repeating regions exceeds the number of primary repeating regions, the newly generated repeating regions are marked as newly added repeating regions, and all short sequences existing in the newly added repeating regions are compared to obtain the similarity between all short sequences in the newly added repeating regions. A sequence similarity threshold is set, and when a short sequence in the newly added repeating region has one or more other short sequences whose similarity exceeds the sequence similarity threshold, all similar short sequences are regarded as the same short sequence.

[0018] The present invention further provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the above-described method for analyzing and displaying computational biology research data.

[0019] The technical solution provided by this invention has the following advantages compared with the prior art:

[0020] Firstly, this invention uses FPKM as a standardized indicator to correct for the influence of gene length and sequencing depth, making comparisons between different genes and different samples more accurate. By calculating the difference between the available sequence and the reference genome FPKM value, the changes in gene expression level can be intuitively reflected, thereby providing reliable data support for differential expression analysis.

[0021] Secondly, by comparing and judging the number of short sequences mapped to the same genomic coordinate interval, this invention can effectively identify abnormal repetitive regions, thereby clearly determining whether the gene expression difference results are reliable. When the number of abnormal sequences exceeds the preset repetitive sequence threshold, the influence of repetitive sequences is eliminated. Specifically, for cases with high FPKM values, identical and overlapping short sequences are removed, and the results are re-compared and updated. This can correct the expression distortion caused by the inclusion of repetitions.

[0022] Third, the present invention adopts an individual difference judgment strategy. By comparing the available sequence with the reference sequence to obtain the maximum similarity and setting a similarity threshold, it can accurately identify mismatch or omission problems caused by individual variation, so as to discover problems in a timely manner when there is a large amount of individual variation in key research areas.

[0023] Fourth, by comparing the number of repeating regions before and after the addition of missed sequences, this invention can promptly identify newly added repeating regions caused by the addition, and optimize the counting of short sequences in the newly added repeating regions. This reduces unreasonable fluctuations in FPKM caused by adding short sequences missed due to filtering based on their quantity on the existing data, ensuring that the FPKM value more accurately reflects the gene expression level. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating a computational biology research data analysis and visualization method according to the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It is understood that the term "a" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple, and the term "a" should not be understood as a limitation on the number.

[0027] This invention provides a method for analyzing and displaying computational biology research data, such as... Figure 1 As shown, the method includes the following steps:

[0028] Step 1: Extract total RNA from biological tissues or cells. Using reverse transcriptase, random primers or oligo-dT primers and a mixture of dNTPs, reverse transcribe the RNA into cDNA to construct a library. The generated cDNA can be amplified by PCR to target specific targets for quality testing to ensure the smooth progress of the library construction process. Then, use a high-throughput sequencing platform to sequence the library to obtain a large number of short sequences.

[0029] Step two involves preprocessing the short sequence data. The FastQC tool is used to detect the quality score distribution, GC content, sequence repetition, and possible sequencing bias of the sequencing data in the raw data sample. Adapter contamination and low-quality regions are identified and removed using data cleaning software (such as Trimmomatic and Cutadapt). The remaining short sequence data are marked as usable sequences. Removing low-quality data from the short sequences reduces data noise caused by sequencing errors, thereby reducing noise and unnecessary background interference, and ensuring that the retained sequences have higher reliability and representativeness.

[0030] Step 3: Establish a reference genome, align the available sequences with the reference genome, and standardize the available sequence data to obtain its FPKM value. FPKM (Fragments Per Kilobase of transcript per Million mapped reads) is a commonly used RNA sequencing data standardization method used to correct for the influence of gene length and total sequencing depth, thereby enabling more accurate comparison of expression levels between different genes or samples and obtaining the FPKM value of the reference genome. The difference between the available sequence and the FPKM value of the reference genome is calculated, and the absolute value of the difference is taken to obtain the gene expression difference results.

[0031] Step four involves obtaining the magnitude of gene expression differences and setting a threshold for this magnitude. The magnitude of the difference is then compared to this threshold. If the magnitude is less than or equal to the threshold, it indicates that the change in the FPKM value of the available sequence compared to the reference genome is within a reasonable range under the current detection conditions, and this gene expression difference result is marked as a data difference analysis result. If the magnitude is greater than the threshold, it indicates that the change in the FPKM value of the available sequence compared to the reference genome is significant, and such a large change is unreasonable under the current detection conditions. A repetitive sequence region detection strategy will be implemented to analyze whether there are many repetitive sequences in the available sequences. If many repetitive sequences are present, the influence of repetitive sequences will be eliminated to obtain an updated difference result, which will then be marked as a data difference analysis result. The data difference analysis results are then visualized using heatmaps, volcano plots, principal component analysis (PCA) plots, etc., to fully display the data and aid in understanding its internal structure.

[0032] The specific method for obtaining the magnitude of the difference is as follows: divide the gene expression difference result by the FPKM value of the reference genome. For example, if the FPKM value of the reference genome is 15 (meaning that in every million short sequences mapped to the reference genome, there are approximately 15 sequencing fragments per thousand bases of the gene transcript), and the FPKM value of the available sequence is 10, the gene expression difference result will be 5, that is, the magnitude of the difference is 5 ÷ 15 = 0.33. Set the magnitude threshold to 0.3. Since the magnitude of the difference is greater than the magnitude threshold, the repetitive sequence region judgment strategy should be implemented. The threshold for the magnitude of difference needs to be set based on a comprehensive consideration of experimental conditions, data characteristics, and research objectives. The threshold for the magnitude of difference should reflect the expected biological effect. When the magnitude of the difference in the result is greater than the threshold, it means that there is an abnormal and excessive change in the expression level of genes in the sample compared with the reference genome. After excluding experimental conditions and sample contamination, and after removing data quality issues through the above methods, the main reason for the abnormal magnitude of the difference in the result is that there are a large number of repetitive or similar sequences in the genome. Short sequences are often difficult to distinguish from which specific location they come from, which leads to multiple alignment problems. Therefore, it is necessary to implement a repetitive sequence region judgment strategy.

[0033] The repetitive sequence region identification strategy includes obtaining the number of short sequences located in the same region among the available sequences and marking them as abnormal sequences (i.e., repetitive sequences). Short sequences in the same region refer to short sequences whose mapping positions overlap when aligned to the reference genome, that is, short sequences that fall within the same genomic coordinate range. A repetitive sequence threshold is set, and the number of abnormal sequences is compared with the repetitive sequence threshold. The following actions are taken based on the comparison results: If the number of abnormal sequences is less than or equal to the repetitive sequence threshold, it means that the overlap of short sequences generated when the available sequences are aligned with the reference genome is within the expected range. In this case, the repetitive region has little interference with the data analysis results, and the obtained gene expression difference results should be considered reliable. This abnormal change in gene expression reflects the abnormal regulation under different biological conditions and can also indicate that this data result truly reflects the biological changes of the sample. It is marked as a data difference analysis result and subsequently visualized. If the number of abnormal sequences is greater than the repetitive sequence threshold, it means that the overlap of short sequences generated when the available sequences are aligned with the reference genome exceeds the expected range. In this case, the repetitive region has a more serious interference with the data analysis results, and the operation of eliminating the influence of repetitive sequences is performed.

[0034] The specific steps for eliminating the influence of repetitive sequences are as follows: The FPKM values ​​of the available sequence and the reference genome are compared. If the FPKM value of the available sequence is greater than that of the reference genome, it indicates that the retention and repeated inclusion of abnormal sequences in the expression count has led to excessive accumulation in repetitive regions, distorting the FPKM value. Identical and overlapping short sequences within the same repetitive region are removed, and the available sequence is re-aligned with the reference genome. The difference between the FPKM values ​​of the available sequence and the reference genome is used to obtain the updated difference result, which is then marked as a data difference analysis result and visualized. If the FPKM value of the available sequence is less than that of the reference genome, it indicates that there are no abnormal sequences causing excessive short sequence expression in this alignment. The obtained gene expression difference result should be deemed reliable, marked as a data difference analysis result, and subsequently visualized.

[0035] Furthermore, when the FPKM value of the available sequence is less than that of the reference genome, besides the issue of multiple alignments caused by repetitive sequences, individual variations exist between the sample's genome and the reference genome, leading to mismatches or inaccurate alignment of some short sequences. Therefore, individual variation judgment should be performed. Specifically, the available sequence is compared with the reference sequence in the reference genome to obtain the maximum similarity between the available sequence and the reference sequence in the reference genome. A similarity threshold is set, and the maximum similarity is compared with the similarity threshold to perform individual variation judgment. Available sequences with a maximum similarity less than the similarity threshold are marked as variant sequences, and the total number of variant sequences is counted. A variant number threshold is set. If the total number of variant sequences is less than the variant number threshold, it indicates that the alignment between the available sequence and the reference sequence in the reference genome is basically correct, and the same region... If no short sequences within the domain are missed due to missing information in the reference genome, the obtained gene expression difference results are considered reliable, marked as data difference analysis results, and subsequently visualized. If the total number of variant sequences is greater than or equal to the variant number threshold, it indicates that a large number of abnormal sequences were not correctly aligned when comparing the available sequences with the reference sequences in the reference genome, resulting in mismatches or missing alignments in the available sequences. This indicates that the genome of this sample has individual variations, which will cause the FPKM value of the available sequences to decrease incorrectly. In this case, the short sequences missed due to filtering are added to the existing data according to their number, and the available sequences are re-aligned with the reference genome. The difference between the FPKM values ​​of the available sequences and the reference genome is used to obtain updated difference results, which are marked as data difference analysis results and visualized. It should also be noted that when the total number of variant sequences is greater than or equal to the variant number threshold, if the study requires precise quantification of certain genes or regions, and these regions happen to have a large number of sites inconsistent with the reference genome, that is, individual variants occur in key study regions, thus affecting the FPKM value, then the analysis process needs to be re-executed under the same conditions to avoid the serious impact caused by individual variants.

[0036] Taking the above embodiments as an example, when executing the repetitive sequence region judgment strategy, the number of overlapping short sequences in the same region is set to 10, that is, the number of abnormal sequences is 10. The repetitive sequence threshold is set to 8. Since the number of abnormal sequences is greater than the repetitive sequence threshold, it indicates that when the available sequence is compared with the reference genome, the overlap of short sequences exceeds the expected range. At this time, the repetitive region seriously interferes with the data analysis results, so the operation of eliminating the influence of repetitive sequences is performed. Since the FPKM value of the available sequence is 10, while the FPKM value of the reference genome is 15, it indicates that there are no abnormal sequences in this alignment that would cause excessive short sequence expression. Individual difference judgment should still be performed: obtain the maximum similarity between the available sequence and the reference sequence in the reference genome, set the similarity threshold to 95%, and set 4 available sequences with a maximum similarity of 90% to the reference sequence in the reference genome. These 4 available sequences are then marked as variant sequences. Set the variant number threshold to 3. If the total number of variant sequences is greater than or equal to the variant number threshold, then the 4 short sequences that were missed due to filtering are added to the existing data according to their number. The available sequence is then re-aligned with the reference genome, and the difference between the FPKM values ​​of the available sequence and the reference genome is used to obtain the updated difference results. These results are then marked as data difference analysis results and visualized.

[0037] After the above steps, adding the short sequences missed due to filtering to the existing data based on their quantity may result in new repetitive regions when comparing the available sequences with the reference genome. A repetitive region represents the presence of multiple short sequences within the genomic coordinate range. In this case, the number of repetitive regions before adding the missed short sequences should be obtained and marked as the primary repetitive region number. The number of repetitive regions after adding the missed short sequences should be obtained and marked as the secondary repetitive region number. The relationship between the primary and secondary repetitive regions should be determined. When the number of secondary repetitive regions exceeds the number of primary repetitive regions, it indicates that adding the short sequences missed due to filtering to the existing data has generated new repetitive regions. These newly generated repetitive regions are marked as newly added repetitive regions. All short sequences within the newly added repetitive regions are compared to obtain the similarity between all short sequences within the newly added repetitive regions. A sequence similarity threshold is set. When a short sequence within a newly added repetitive region has one or more other short sequences with a similarity exceeding the sequence similarity threshold, all similar short sequences are considered as the same short sequence, meaning that the number of this short sequence is counted as 1 when calculating the FPMK value. This avoids the problem of duplicate counting introduced by supplementing missed sequences, so that the expression level is not artificially inflated or distorted when the same actual sequence is counted repeatedly during FPKM calculation. Furthermore, by comparing the similarity of all short sequences in the newly added repeating region and setting a similarity threshold, sequences that appear multiple times but whose similarity exceeds the threshold are only counted once. This effectively removes the bias caused by technical noise or comparison redundancy, and reduces the unreasonable fluctuation of FPKM caused by adding short sequences that were missed due to filtering based on their quantity on the existing data. This ensures that the FPKM value more accurately reflects the gene expression level.

[0038] In summary, this invention aims to design a computational biology research data analysis and visualization method. Addressing the difficulty in obtaining accurate data analysis results when gene expression level differentials are abnormal, this invention uses FPKM as a standardized indicator to correct for the influence of gene length and sequencing depth, making comparisons between different genes and samples more accurate. By calculating the difference in FPKM values ​​between the available sequence and the reference genome, it intuitively reflects changes in gene expression levels, thus providing reliable data support for differential expression analysis. When the expression level change between the available sequence and the reference genome exceeds a preset reasonable range, a repetitive sequence region detection strategy is automatically triggered. By comparing the number of short sequences mapped to the same genomic coordinate interval and applying a threshold, abnormal repetitive regions can be effectively identified, thus clearly determining the reliability of gene expression differential results. When the number of abnormal sequences exceeds a preset repetitive sequence threshold, the influence of repetitive sequences is eliminated. Specifically, for cases with high FPKM values, identical and overlapping short sequences are removed, and the results are re-aligned and updated, correcting the expression level distortion caused by repetition. Furthermore, this invention employs an individual difference judgment strategy. By comparing available sequences with reference sequences to obtain the maximum similarity and setting a similarity threshold, it can accurately identify mismatches or omissions caused by individual variations. This allows for timely detection of problems when there are significant individual variations in key research areas, preventing serious impacts on quantitative analysis due to inaccurate data in key regions. By comparing the number of repetitive regions before and after the addition of omitted sequences, this invention can promptly identify newly added repetitive regions caused by supplementation. It also optimizes the counting of short sequences within these newly added repetitive regions, reducing unreasonable fluctuations in FPKM caused by adding short sequences missed due to filtering based on their quantity to existing data. This ensures that FPKM values ​​more accurately reflect gene expression levels.

[0039] The processes described above with reference to the flowcharts in the embodiments disclosed in this invention can be implemented as computer software programs. Embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more conductor segments, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination thereof.

[0040] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0041] Those skilled in the art should understand that the above description is only a specific embodiment of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application.

Claims

1. A method for analyzing and displaying computational biology research data, characterized in that, The method includes the following steps: Total RNA is extracted from biological tissues or cells, reverse transcribed into cDNA, a library is constructed, and the library is sequenced using a high-throughput sequencing platform to obtain a large number of short sequences. Short sequence data are preprocessed. The FastQC tool is used to detect the quality fraction distribution, GC content, sequence duplication, and possible sequencing bias of the sequencing data in the raw data sample. Adapter contamination and low-quality regions are identified and removed, and the remaining short sequence data are marked as usable sequences. The available sequence data is standardized to obtain its FPKM value. A reference genome is constructed and its FPKM value is obtained. The difference between the available sequence and the reference genome FPKM value is calculated to obtain the gene expression difference results. The magnitude of gene expression differences is obtained by analyzing gene expression differences and setting a threshold for the magnitude of the difference. When the magnitude of the difference is less than or equal to the threshold, the gene expression difference is marked as a data difference analysis result. When the magnitude of the difference is greater than the threshold, a repetitive sequence region detection strategy is executed. This strategy is used to eliminate the influence of repetitive sequences when there are many repetitive sequences in the available sequences, so as to obtain an updated difference result and mark it as a data difference analysis result. The data difference analysis result is then visualized to achieve data presentation. The specific steps for eliminating the influence of repetitive sequences are as follows: First, compare the FPKM values ​​of the available sequence and the reference genome. If the FPKM value of the available sequence is greater than that of the reference genome, remove identical and overlapping short sequences within the same repetitive region. Then, re-align the available sequence with the reference genome, calculate the difference between the FPKM values ​​of the available sequence and the reference genome to obtain the updated difference results, and mark them as data difference analysis results for visualization. If the FPKM value of the available sequence is less than that of the reference genome, the obtained gene expression difference results should be considered reliable, marked as data difference analysis results, and subsequently visualized. When the FPKM value of the available sequence is less than that of the reference genome, an individual difference judgment is performed to determine whether there are individual variations in the sample's genome. The method for performing the individual difference judgment is as follows: the available sequence is compared with the reference sequence in the reference genome to obtain the maximum similarity between the available sequence and the reference sequence in the reference genome, and a similarity threshold is set. The maximum similarity is compared with the similarity threshold to perform the individual difference judgment. Available sequences with a maximum similarity less than the similarity threshold are marked as variant sequences. The total number of variant sequences is counted, a variant number threshold is set, and an appropriate response is made based on the comparison between the total number of variant sequences and the variant number threshold. If the total number of variant sequences is less than the number of variants threshold, the obtained gene expression difference results are considered reliable, marked as data difference analysis results, and then visualized. If the total number of variant sequences is greater than or equal to the variant number threshold, it is determined that the genome of this sample has individual variants. Short sequences that were missed due to filtering are added to the existing data according to their quantity, and the available sequences are re-aligned with the reference genome. The difference between the available sequences and the reference genome is used to obtain the updated difference results, which are marked as data difference analysis results and visualized.

2. The computational biology research data analysis and visualization method according to claim 1, characterized in that, The strategy for determining repetitive sequence regions includes obtaining the number of short sequences located in the same region among available sequences, marking them as abnormal sequences, setting a repetitive sequence threshold, comparing the number of abnormal sequences with the repetitive sequence threshold, and taking appropriate action based on the comparison result.

3. The computational biology research data analysis and visualization method according to claim 2, characterized in that, If the abnormal sequence is less than or equal to the repetitive sequence threshold, the obtained gene expression difference result is considered reliable, marked as a data difference analysis result, and then visualized. If the abnormal sequence is greater than the repetitive sequence threshold, the operation to eliminate the influence of repetitive sequences is performed.

4. The computational biology research data analysis and visualization method according to claim 3, characterized in that, After adding short sequences that were missed due to filtering to the existing data based on their quantity, the number of repeating regions before adding the missed short sequences is obtained and marked as the number of primary repeating regions. The number of repeating regions after adding the missed short sequences is obtained and marked as the number of secondary repeating regions. The relationship between the number of primary repeating regions and the number of secondary repeating regions is determined and a corresponding response is taken.

5. The computational biology research data analysis and visualization method according to claim 4, characterized in that, When the number of secondary repeating regions exceeds the number of primary repeating regions, the newly generated repeating regions are marked as newly added repeating regions. All short sequences existing in the newly added repeating regions are compared to obtain the similarity between all short sequences in the newly added repeating regions. A sequence similarity threshold is set. When a short sequence in the newly added repeating region has one or more other short sequences whose similarity exceeds the sequence similarity threshold, all similar short sequences are regarded as the same short sequence.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement a computational biology research data analysis and visualization method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Transcriptome analysis method and system without reference genome sequence

    CN112397149A

  • Gene detection data cleaning method and system based on artificial intelligence

    CN119418762A