Parallel computing and data visualization analysis system for whole genome prediction
By designing a parallel computing and data visual analysis system, the problems of low efficiency and insufficient visualization in the prior art are solved, and efficient genomic data analysis and intuitive visual display are achieved.
Patent Information
- Application Number
- CN202510437299.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing technology has low computational efficiency and long response time when processing massive genomic data, and lacks effective data visualization tools, which makes it difficult to meet the computing efficiency and data management needs of genomic research.
A parallel computing and data visual analysis system for whole genome prediction is designed, including image generation module, image analysis module, genome prediction module and visualization module. The genome data is converted into multiple images through parallel computing and data visualization algorithms, and detailed analysis and prediction are performed.
It significantly improves the efficiency and real-time nature of data analysis, enhances the accuracy and operability of genomic data analysis, and provides an intuitive visual presentation method to help researchers better understand genomic data.
Smart Images

Figure CN119964655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene image data analysis, and in particular to a parallel computing and data visualization analysis system for whole genome prediction. Background Art
[0002] With the continuous development of precision medicine, personalized medicine and genomic research, the application prospects of genomic data in disease prediction, personalized treatment, drug development and agricultural breeding are becoming more and more extensive. As one of the important tools, genomic prediction technology can accurately predict an individual's health status, genetic disease risk and other biological characteristics through comprehensive analysis of individual genomic information. However, as the scale of genomic data continues to expand, traditional computing methods and analysis tools face huge challenges. Especially when it is necessary to process massive genomic data, how to provide scientific analysis results through efficient computing means and accurate visualization technology has become the key to promoting genomic research and application. To this end, data visualization technology based on high-performance computing and innovation is becoming the core solution to this problem.
[0003] At present, most of the analysis methods for genome prediction rely on traditional stand-alone computing or inefficient distributed computing architecture, which has problems of low computing efficiency and long response time. When performing complex analysis of genomic data, especially whole genome prediction and high-throughput genomic data processing, the amount of calculation is huge. Traditional computing methods cannot effectively process these massive data, resulting in a slow analysis process and a large consumption of resources. Secondly, most of the existing genomic data visualization tools are limited to static graphic display, lack of interactivity and multi-dimensional display, and cannot effectively present the intrinsic association of high-dimensional genomic data. Especially when dealing with genomic diversity and complex gene-environment interactions, there is a lack of intuitive and easy-to-understand visualization methods. In addition, the bottleneck of data storage and processing also limits the application of existing methods in large-scale parallel computing environments, and it is difficult to meet the dual needs of computing efficiency and data management in genomic research. Therefore, when facing genomic data analysis and prediction, the existing methods are inefficient, lack processing power, and lack effective data visualization tools, which limits their wide application and in-depth research. Summary of the invention
[0004] In view of the deficiencies of the prior art, the present invention provides a parallel computing and data visualization analysis system for whole genome prediction, which solves the problems of the above-mentioned background technology.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a parallel computing and data visualization analysis system for whole genome prediction, comprising the following modules: an image generation module, an image analysis module, a genome prediction module, and a visualization module; the image generation module is used to obtain target whole genome data, and convert the target whole genome data into a whole genome sequencing image through data visualization algorithm and parallel computing, including a genome heat map, a genome variation map, and a gene structure map; the image analysis module comprises a heat map analysis unit, a variation map analysis unit, and a structure map analysis unit; the heat map analysis unit is used to perform spatial distribution and color intensity analysis on the genome heat map, identify the expression difference area of the gene under set conditions, extract the gene expression difference distribution map in the genome heat map, and obtain the gene expression difference index, which is used to reflect the change in the expression level of the gene under set conditions; the variation map analysis unit It is used to extract gene variation characteristics in the genome variation map, identify variation types and distribution patterns in the genome, obtain genome variation distribution maps and variation type frequency maps, which are used to reflect the variation hotspot areas of various variation types in the genome; the structure map analysis unit is used to perform distribution density analysis of gene regions on the gene structure map, identify the aggregation of genes in set regions in the genome, and obtain gene density indexes, which are used to reflect the distribution characteristics of genes in different genome regions; the genome prediction module is used to receive output data from the image analysis module, predict gene expression trends, variation hotspot evolution and functional regional distribution according to the machine learning model, and generate a genome prediction map; the visualization module is used to perform comprehensive visualization of gene expression difference distribution maps, gene expression difference indices, genome variation distribution maps, variation type frequency maps, gene density indexes and genome prediction maps.
[0006] Furthermore, the specific process of converting the target whole genome data into a whole genome sequencing image through data visualization algorithm and parallel computing is as follows: the target whole genome data is loaded through parallel computing, and the target whole genome data is cleaned, filtered and standardized, and outliers and missing data are processed; the target whole genome data includes gene expression data, gene variation data and gene structure data; based on the gene expression data, a gene expression heat map is generated through a heat map visualization algorithm; based on the gene variation data, a genome variation map is constructed through a variation map visualization algorithm; based on the gene structure data, a gene structure diagram is drawn through a gene structure visualization algorithm, and the various components of the gene and their spatial positions are marked; the generated genome heat map, genome variation map and gene structure diagram are integrated into a comprehensive whole genome sequencing image.
[0007] Furthermore, the genomic heat map is subjected to spatial distribution and color intensity analysis, and regions with differential expression of genes under set conditions are identified. The specific process of extracting the gene expression difference distribution map in the genomic heat map is as follows: the spatial position of each gene region in the genomic heat map is analyzed by parallel calculation, the expression density and distribution pattern of genes in different regions are evaluated, and the expression difference of genes in different regions is identified; the color intensity of each pixel in the heat map is quantitatively analyzed to determine the relationship between different color intensities and gene expression levels, and the difference in expression intensity is identified through color changes in the heat map; according to the change in color intensity, the expression difference threshold is set to identify the gene regions that show upregulation or downregulation under the set conditions; the gene expression difference distribution map is extracted from the identified expression difference regions to show the expression changes of genes in different regions and mark the changed gene regions.
[0008] Furthermore, the specific process of obtaining the gene expression difference index is as follows: by quantitatively analyzing the color intensity of each gene region in the genome heat map, the expression value of each gene in different regions is extracted. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level; for each gene in the identified expression difference region, statistical analysis is used to perform difference analysis, and by comparing the expression level of the gene under the set conditions, the expression change of each gene under the preset experimental conditions is calculated to obtain the gene expression difference index.
[0009] Furthermore, the specific process of extracting gene variation features in the genome variation map and identifying variation types and distribution patterns in the genome is as follows: extracting features from each variation region in the genome variation map, analyzing the variation types, including mutations, insertions, deletions, and copy number variations, and extracting their spatial distribution information; and obtaining variation hotspots by calculating the distribution of variation in different gene regions and chromosome regions.
[0010] Furthermore, the specific process of obtaining the genome variation distribution map and variation type frequency map is as follows: the identified gene variation types and position data are integrated to construct a genome variation distribution map to show the distribution of different variation types in the genome; the frequency of each variation type is statistically analyzed to obtain a variation type frequency map, which displays the frequency of each variation type in the genome through a graph, reflecting the frequency of different variation types and their distribution characteristics in the genome.
[0011] Furthermore, the distribution density of gene regions in the gene structure map is analyzed, and the specific process of identifying the clustering of genes in the set region in the genome is as follows: the spatial density of each gene region in the gene structure map is evaluated through parallel calculation, the analysis region set in the genome is delineated, and the gene distribution in the region is quantified, and the distribution density of genes in each region is calculated by the K-means clustering algorithm; the degree of gene clustering is evaluated by calculating the clustering index of gene distribution in the set analysis region, and high-density gene clustering regions are identified; the distribution pattern of gene regions is further identified through spatial autocorrelation analysis, and the clustering trend of different genes in the set region is analyzed in combination with gene function annotation information.
[0012] Furthermore, the specific process of obtaining the gene density index is as follows: according to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; the gene aggregation situation in different areas is quantified by statistical methods to obtain the gene density index of each area.
[0013] Furthermore, the specific process of predicting gene expression trends, mutation hotspot evolution and functional regional distribution based on the machine learning model is as follows: based on the gene expression difference index extracted by the heat map analysis unit, the dynamic change trend of gene expression under set conditions is predicted through the time series model; based on the mutation hotspot areas identified by the mutation map analysis unit, combined with the mutation type frequency map, the potential new mutation hotspot areas are predicted through the convolutional neural network; based on the gene density index of the structure map analysis unit, the enrichment areas of functionally related genes in the genome are predicted through the clustering algorithm.
[0014] Furthermore, the specific process of comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map is as follows: coordinates of gene expression difference distribution map, genome variation distribution map and gene structure map are normalized by parallel calculation, and genome positions of different data sources are mapped to a unified reference coordinate system; based on the statistical results of gene density index and variation type frequency map, a gene region weight matrix is constructed to dynamically adjust the overlay order of visualization layers; the predicted value of gene expression trend in genome prediction map is converted into a dynamic heat map, and the predicted heat map at different time points is adjusted and displayed through the timeline control, and the transparency is overlaid and compared with the gene expression heat map; in the genome variation distribution map, the predicted new variation hotspot area is overlaid on the known variation area with a pulse flashing mark, and a ring legend is generated through the variation type frequency map to display the type ratio difference of the predicted hotspot and the historical data in real time; the prediction result of the functional enrichment area in the gene structure map is converted into a 3D contour surface, which is superimposed on the two-dimensional gene density distribution map, the height of the surface is positively correlated with the gene density index, and the color gradient represents the prediction confidence.
[0015] The present invention has the following beneficial effects: (1) The parallel computing and data visualization analysis system for whole genome prediction, the image generation module is based on parallel computing and data visualization algorithms, which converts complex genome data into a variety of genome images (genome heat map, variation map, gene structure map), effectively reducing the time delay of traditional methods in data processing. Through parallel computing technology, the system can process a large amount of genome data at the same time, accelerate the image generation process, and thus greatly improve the efficiency and real-time performance of data analysis. In the image analysis module, the heat map analysis, variation map analysis and structure map analysis units accurately extract the expression differences, gene variation characteristics and distribution density of gene regions in the genome image. These analysis results can provide in-depth support for gene function research, variation detection and genome structure analysis, help researchers identify key variation hotspots and gene clustering areas in the genome, and greatly improve the accuracy and operability of genome data analysis.
[0016] (2) The parallel computing and data visualization analysis system for whole genome prediction. The genome prediction module uses an algorithm model to perform genome prediction, which can accurately predict gene variation, expression differences, and possible functional regions. By working in collaboration with other modules, the prediction results can be combined with image analysis results in real time to provide researchers with comprehensive genome information. This module greatly improves the prediction accuracy and can quickly identify key variants with potential biological significance in large-scale genome data, further optimizing research decisions and improving the efficiency and accuracy of genome research. The visualization module makes the interpretation of genome data more intuitive and easy to understand by integrating the visualization of different analysis results. This module not only optimizes the data display method, but also enhances researchers' understanding of complex information such as genome structure, expression differences, and variation distribution, which helps to quickly discover potential gene variations and key regions, and improves the flexibility and decision-making efficiency of research.
[0017] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the parallel computing and data visualization analysis system for whole genome prediction of the present invention. DETAILED DESCRIPTION
[0019] The present application embodiment solves the problems of low genome data processing efficiency, insufficient analysis accuracy and poor data integration through parallel computing and data visualization analysis system of whole genome prediction. By combining parallel computing and data visualization algorithm, the system can efficiently convert large-scale whole genome data into intuitive genome heat map, variation map and gene structure map, effectively improving the speed and accuracy of image generation. At the same time, the system automatically extracts key information such as gene expression difference, variation region and gene density using image analysis module, reduces manual intervention, and improves the automation and accuracy of data analysis. Through the comprehensive display of visualization module, researchers can intuitively observe the changes of genome data, which promotes the efficient conduct of genomic research.
[0020] The overall idea of the solution in the embodiments of this application is as follows: Obtain the target whole genome data, and convert the target whole genome data into whole genome sequencing images through data visualization algorithms and parallel computing, including genome heat maps, genome variation maps, and gene structure maps.
[0021] The spatial distribution and color intensity of the genome heat map were analyzed to identify the differential expression regions of genes under set conditions, extract the differential distribution map of gene expression in the genome heat map, and obtain the differential gene expression index.
[0022] Extract the gene variation features in the genome variation map, identify the variation types and distribution patterns in the genome, obtain the genome variation distribution map and variation type frequency map, which are used to reflect the variation hot spots of various variation types in the genome.
[0023] The gene structure map is analyzed for distribution density of gene regions, the aggregation of genes in set regions of the genome is identified, and the gene density index is obtained to reflect the distribution characteristics of genes in different genomic regions.
[0024] Comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map and gene density index.
[0025] See also Figure 1The embodiment of the present invention provides a technical solution: a parallel computing and data visualization analysis system for whole genome prediction, comprising the following modules: an image generation module, an image analysis module, a genome prediction module, and a visualization module; the image generation module is used to obtain target whole genome data, and convert the target whole genome data into a whole genome sequencing image through a data visualization algorithm and parallel computing, including a genome heat map, a genome variation map, and a gene structure map; the image analysis module comprises a heat map analysis unit, a variation map analysis unit, and a structure map analysis unit; the heat map analysis unit is used to perform spatial distribution and color intensity analysis on the genome heat map, identify the expression difference area of the gene under the set conditions, extract the gene expression difference distribution map in the genome heat map, and obtain the gene expression difference index, which is used to reflect the expression level change of the gene under the set conditions; the variation map analysis unit is used to analyze the gene expression difference index of the genome heat map. The gene variation characteristics in the genomic variation map are extracted, the variation types and distribution patterns in the genome are identified, and the genome variation distribution map and variation type frequency map are obtained to reflect the variation hotspot areas of various variation types in the genome; the structure map analysis unit is used to perform distribution density analysis of gene regions on the gene structure map, identify the aggregation of genes in set regions in the genome, and obtain the gene density index to reflect the distribution characteristics of genes in different genome regions; the genome prediction module is used to receive the output data of the image analysis module, predict the gene expression trend, variation hotspot evolution and functional region distribution according to the machine learning model, and generate a genome prediction map; the visualization module is used to comprehensively visualize the gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map.
[0026] In this embodiment, the image generation module: the main function of this module is to generate genome sequencing images from the target whole genome data through data visualization algorithms and parallel computing technology. These images include genome heat maps, genome variation maps and gene structure maps, which are used to display genome information. Parallel computing technology enables this module to efficiently process large-scale genome data and convert complex gene information into visual images suitable for subsequent analysis. In this way, the computational bottleneck of traditional genome data processing methods is broken through, and image results can be generated in a shorter time. Genome heat map: convert the numerical information of gene expression into a heat map format, and display the expression level of genes under different conditions based on the intensity of different colors. Genome variation map: present the variation data in the genome in the form of a map, revealing the type and distribution of gene variation. Gene structure map: display the spatial distribution and structural characteristics of genes in the genome, and help users understand the relationship between genes. Image analysis module: this module is responsible for detailed analysis of the above-mentioned generated genome images (such as heat maps, variation maps and gene structure maps). The image analysis module extracts important information from the image based on advanced image processing and machine learning algorithms, and further analyzes it, thereby providing users with valuable genome data interpretation. Heat map analysis unit: This unit uses image analysis algorithms to analyze the spatial distribution and color intensity in the genome heat map and identify the differential expression areas of genes under set conditions. By extracting the gene expression difference distribution map and calculating the gene expression difference index, this module can reflect the changes in the expression level of genes under specific conditions. Variation map analysis unit: This unit extracts the gene variation characteristics in the genome variation map, uses pattern recognition algorithms to identify the variation types and analyzes their distribution patterns in the genome. Finally, the genome variation distribution map and variation type frequency map are generated to reflect the hot spots of different variation types in the genome. Structure map analysis unit: This unit mainly analyzes the spatial distribution of gene structure maps, identifies the aggregation of genes in different genome regions, and uses the gene region distribution density analysis algorithm to calculate the gene density index to reflect the distribution characteristics of genes in different genome regions. The visualization module is used to comprehensively display all the data generated by the image analysis module (such as gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map and gene density index). Through this module, users can intuitively view various analysis results and effectively compare, screen and support decisions. This module uses advanced image visualization technology to present complex analysis results in easy-to-understand graphics or charts, making the genomic data analysis results clearer and more intuitive. Parallel computing refers to a computing method that uses multiple processors or computers to work simultaneously to solve problems. In the process of genomic data processing, due to the huge amount of data, traditional serial computing methods may not be able to efficiently complete the computing tasks.Parallel computing divides tasks into small pieces and assigns them to multiple processing units for parallel processing, thereby greatly improving computing efficiency. In this system, parallel computing is used to accelerate the processing of genomic data and ensure that large-scale genomic image results can be quickly generated, which is especially important when processing high-dimensional data such as genomic heat maps and variation maps. Data visualization algorithms refer to algorithms that convert data into easy-to-understand and easy-to-analyze graphics or images through specific computational methods. These algorithms select appropriate colors, shapes, sizes, and layouts to represent different dimensions of data based on the characteristics of the data, such as gene expression levels, variation types, and gene structures. In this system, data visualization algorithms are used to convert genomic data into image formats such as heat maps, variation maps, and structure maps, so that complex genomic information can be presented to researchers intuitively. Heat maps are a two-dimensional data visualization method that uses color depth or different colors to represent the size of numerical values. In genomics, heat maps are often used to display changes in gene expression. Darker colors indicate higher gene expression, and lighter colors indicate lower expression. This method facilitates the observation of differences in gene expression under different conditions and helps researchers identify key genes and potential biological laws. The genomic variation map refers to a map that graphically displays the variation in the genome (single nucleotide polymorphism SNP, insertion / deletion Indel, structural variation). In this system, the variation map analysis unit is responsible for identifying and extracting the variation characteristics in the genome and presenting its distribution in the genome. Through the variation map, researchers can quickly locate variation hotspots and then study the relationship between different variations and diseases or other biological characteristics. The variation type frequency map is a chart used to display the frequency distribution of different types of variation in the genome. Common types of variation include point mutations, insertion / deletion, structural variation, etc. This map can help researchers quickly understand which types of variation appear most frequently in the entire genome, thereby providing data support for further research. The gene density index is an indicator to measure the distribution of genes in a specific genomic region. Regions with high gene density mean that the genes in the region are more concentrated, while regions with low gene density mean that the genes are more dispersed. This index helps to analyze the distribution characteristics of genes in the entire genome and identify gene regions that may have special functions or biological significance. Image analysis algorithms are a class of algorithms used to extract information from images. In the present invention, image analysis algorithms are used to process the generated genome heat maps, variation maps and structure maps to extract valuable information. Common image analysis methods include edge detection, segmentation, feature extraction, etc., the goal is to identify and quantify key biological features from images. For example, through the heat map analysis algorithm, the differential areas of gene expression can be automatically identified and the gene expression difference index can be calculated.
[0027] Specifically, the specific process of converting the target whole genome data into a whole genome sequencing image through data visualization algorithms and parallel computing is as follows: the target whole genome data is loaded through parallel computing, and the target whole genome data is cleaned, filtered and standardized, and outliers and missing data are processed; the target whole genome data includes gene expression data, gene variation data and gene structure data; based on the gene expression data, a gene expression heat map is generated through a heat map visualization algorithm; based on the gene variation data, a genome variation map is constructed through a variation map visualization algorithm; based on the gene structure data, a gene structure diagram is drawn through a gene structure visualization algorithm, and the various components of the gene and their spatial positions are marked; the generated genome heat map, genome variation map and gene structure diagram are integrated into a comprehensive whole genome sequencing image.
[0028] In this implementation scheme, parallel computing: The target genome data volume is huge, including gene expression data, gene variation data and gene structure data, and the processing speed of single-machine serial computing may not meet the requirements. Therefore, parallel computing is used to distribute data processing tasks to multiple computing nodes or processors for simultaneous execution, which greatly accelerates the speed of data processing. Data cleaning and filtering: The target data often contains noise or error information, and useless data needs to be removed by cleaning. The filtering process helps to screen out valid data that needs to be analyzed. Standardization: Gene data from different sources or different measurement methods may have different ranges and units, so standardization is required to ensure data consistency and facilitate subsequent analysis. Outlier and missing data processing: There may be outliers (such as measurement errors) or missing values in gene data (such as some genes are not measured in some samples). These data need to be processed by reasonable methods (such as mean filling, interpolation or deletion) to avoid bias in the analysis results. Heat map visualization algorithm: This algorithm converts the difference in gene expression into color intensity according to the numerical size of gene expression data. Areas with higher gene expression values are represented by dark colors, and areas with lower values are represented by light colors, allowing researchers to intuitively see which genes have changed in expression under different conditions. Heatmaps allow researchers to identify differences in gene expression in specific regions or conditions, and then study the effects of these genes on organisms or diseases. Variant map visualization algorithm: This algorithm plots a map of genomic variation based on gene variation data (such as SNPs, insertions / deletions, etc.). Each variation type is indicated by a different symbol, color, or shape, helping researchers identify and understand variation hotspots and variation type distribution in the genome. The variation map shows the variation types and their frequencies at different locations in the genome, providing important information for further disease research or gene function research. Gene structure visualization algorithm: This algorithm plots a complete structural map of the gene based on the gene's structural data (such as exons, introns, regulatory regions, etc.), showing the various components of the gene and their spatial locations in the genome. The gene structure map helps reveal the region of gene function, the structural changes of the gene (such as splicing isoforms), and their association with function. Integrate the image: Once the genome heatmap, variation map, and gene structure map are generated separately, the next step is to integrate these image information into a comprehensive whole genome sequencing image. This image will contain multi-dimensional information, showing the expression level, variation characteristics and structural layout of genes, so that researchers can see the overall situation of the whole genome at a glance. Comprehensive visualization: By displaying these image information on the same interface, researchers can conduct a comprehensive analysis of various features of the genome. For example, differences in gene expression may be related to the location and structural characteristics of certain gene variants. Through such a comprehensive view, researchers can discover potential associations between different gene features, thereby deepening their understanding of genome function.
[0029] Specifically, the spatial distribution and color intensity of the genome heat map are analyzed, the regions with differential expression of genes under set conditions are identified, and the specific process of extracting the gene expression difference distribution map in the genome heat map is as follows: the spatial position of each gene region in the genome heat map is analyzed by parallel calculation, the expression density and distribution pattern of genes in different regions are evaluated, and the expression difference of genes in different regions is identified; the color intensity of each pixel in the heat map is quantitatively analyzed to determine the relationship between different color intensities and gene expression levels, and the difference in expression intensity is identified through color changes in the heat map; according to the change in color intensity, the expression difference threshold is set to identify the gene regions that show upregulation or downregulation under the set conditions; the gene expression difference distribution map is extracted from the identified expression difference regions to show the expression changes of genes in different regions and mark the changed gene regions.
[0030] In this embodiment, spatial position analysis: the spatial position of each gene region in the genome heat map is analyzed. The position of each gene in the map may be related to its expression pattern or characteristics. By analyzing these spatial positions, the expression density and distribution pattern of genes in different regions can be evaluated, and the expression differences of genes in these regions can be identified. Expression density: In the heat map, the expression density of genes represents the expression level of genes in a specific region. By evaluating the expression density of different gene regions, it is possible to identify which regions have more intensive gene expression and which regions have sparse gene expression, thereby revealing the spatial distribution characteristics of genes. Distribution pattern: Not only the expression density should be analyzed, but also the distribution pattern of genes in the heat map should be identified. Gene expression may present a specific pattern, such as high expression in some regions and low expression in other regions. By analyzing the distribution pattern, the difference in gene expression in different regions can be identified, which further helps to understand the relationship between the function of the gene and the region in which it is located. Color intensity analysis: The heat map represents the expression level of the gene by color. Generally, the color intensity of the heat map (such as from light to dark) is proportional to the expression level of the gene. By quantifying the color intensity of each pixel, the expression intensity of the gene in different regions can be calculated. Color-expression relationship: Determine the relationship between different color intensities and gene expression levels, for example, dark red may represent high expression, while light yellow may represent low expression. This analysis helps researchers convert color intensity into quantitative expression levels, thereby accurately identifying differences in gene expression intensity. Threshold setting: Based on changes in color intensity, a threshold for expression differences can be set. This threshold is used to distinguish between upregulated and downregulated regions of genes. By setting a specific range of values, regions where gene expression changes significantly under specific experimental conditions can be identified. Up- and down-regulation identification: After setting the threshold, it is possible to identify which genes show upregulation (i.e., increased expression) or downregulation (i.e., decreased expression) under specific conditions. This analysis helps to find gene expression changes associated with certain biological processes, diseases, or external stimuli. Differential expression region extraction: From the differential expression regions (upregulated or downregulated regions) identified in the heat map, data from these regions are further extracted to generate a gene expression differential distribution map. This map shows the expression changes of genes in different regions or conditions. Display of expression changes: In the differential distribution map, changes in gene expression are marked by different colors or graphics, helping researchers to intuitively see which genes have significantly changed in expression under specific regions or conditions. The gene expression differential distribution map can not only show the changes in expression, but also mark the changed gene regions, thereby providing data support for subsequent biological analysis.
[0031] Specifically, the specific process of obtaining the gene expression difference index is as follows: by quantitatively analyzing the color intensity of each gene region in the genome heat map, the expression value of each gene in different regions is extracted. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level; for each gene in the identified expression difference region, statistical analysis is used to perform difference analysis, and by comparing the expression level of the gene under the set conditions, the expression change of each gene under the preset experimental conditions is calculated to obtain the gene expression difference index.
[0032] In this embodiment, the color intensity of each gene region in the genome heat map is quantitatively analyzed. There is a mapping relationship between the color intensity in the heat map and the gene expression level. Assume that the color intensity of a gene region in the heat map is ,in represents the gene number, Indicates the location of the gene in different regions (e.g. regions at different time points or under different conditions). Calculate the expression value of the gene in different regions: For each gene , according to the color intensity in the heat map, calculate the gene in different regions (by Assume that the expression value of each region is , can be obtained by: ; Among them, the function Indicates color intensity and expression values Mapping relationship between them. Identify differentially expressed regions: For the identified differentially expressed regions (such as up-regulated or down-regulated regions), select gene regions with significant changes in gene expression values. The expression values of these regions vary greatly and are the focus of analysis. Differential analysis: For each gene, use statistical methods to perform differential analysis. Assume that there are two experimental conditions (condition A and condition B), condition A: the expression of the gene at the beginning of the experiment. Condition B: the expression of the gene at the end of the experiment. The expression values of the genes are respectively and ,in Indicates the gene number. Differential analysis calculates the change in gene expression by comparing the expression levels under experimental conditions. Calculate the gene expression difference index: Gene expression difference index It is calculated based on the change in gene expression values. The gene expression difference index is expressed by the following formula: ;in: It's genetic Expression values under condition A. It's genetic Expression values under condition B. and Represents the standard deviation of expression values under condition A and condition B, respectively. It indicates the expression change of genes under different experimental conditions. The larger the value, the more significant the difference in gene expression. It is used to standardize the fluctuation of expression values under different conditions and ensure the comparability of the difference index. It reflects the difference in gene expression under different conditions. The larger the value, the more significant the difference in gene expression under different conditions.
[0033] Specifically, the specific process of extracting gene variation features in the genome variation map and identifying variation types and distribution patterns in the genome is as follows: extract features from each variation region in the genome variation map, analyze the variation types, including mutations, insertions, deletions, and copy number variations, and extract their spatial distribution information; obtain variation hotspots by calculating the distribution of variation in different gene regions and chromosome regions.
[0034] In this embodiment, variant region feature extraction: each variant region is extracted from the genome variant map by parallel computing and map analysis. These variant regions refer to specific regions that have changed in the target genome sequence. Each variant region has its specific spatial position and variant characteristics, including mutation location, type, and gene or sequence information around the location. When extracting features, the system quantifies the size, type, frequency, associated genes and other information of these variant regions for further analysis. Variations are generally divided into several main types, and each type of variation may have different effects on the function and phenotype of the gene. Common types of variation include: mutation: usually refers to changes in a single nucleotide, such as point mutations. Mutations may be substitutions, insertions or deletions. Insertion: refers to the insertion of a new nucleotide sequence at a certain position in the genome, which may be caused by the introduction of exogenous DNA fragments or repeated sequences within the gene. Deletion: refers to the loss of a DNA fragment at a certain position in the genome, which usually affects the function of the gene. Copy number variation: refers to changes in the number of gene copies in certain regions of the genome, which may cause gene overexpression or deletion. The system uses a variation calling algorithm to identify these different types of variations by comparing the variation map of the genome with the reference genome. After identifying the variant types, the system will further analyze the distribution patterns of these variants on the genome. This process includes quantitative analysis of the spatial distribution of variant regions in different gene regions and chromosome regions. Spatial distribution analysis of variants helps identify whether variants tend to be concentrated in specific regions (exons, regulatory regions) or whether they are evenly distributed throughout the genome. The system calculates the distribution of each variant type in different gene regions (exons, introns, non-coding regions) and chromosomes. This analysis can help researchers discover high-risk variant regions and possible functional variant regions in the genome. By calculating the distribution density of each variant type, the system can identify variant "hotspot" regions. Variation hotspots refer to regions with high variation frequencies, which may be related to certain biological phenomena such as genetic diseases and cancer. Variation hotspot regions may be concentrated in certain specific genes, chromosome regions, or gene regulatory elements.
[0035] Specifically, the specific process of obtaining the genome variation distribution map and variation type frequency map is as follows: Integrate the identified gene variation types and position data to construct a genome variation distribution map to show the distribution of different variation types in the genome; Perform statistical analysis on the frequency of each variation type to obtain a variation type frequency map, and use a chart to show the frequency of each variation type in the genome, reflecting the frequency of different variation types and their distribution characteristics in the genome.
[0036] In this embodiment, the process of obtaining a genomic variation distribution map and a variation type frequency map is mainly divided into two steps: the construction of a genomic variation distribution map and the statistics and display of a variation type frequency map. These steps help analyze the spatial distribution of various types of variations in the genome and the frequency of variations. Construction of a genomic variation distribution map: The gene variation types identified in the previous steps and their corresponding positions (such as chromosome positions) are integrated. Each variation type (such as mutation, insertion, deletion, copy number variation, etc.) has a clear position coordinate in the genome (such as the position on the chromosome) and may affect a specific gene or gene region. Variation position calibration: In this process, the system calibrates each identified variation (such as mutation position, insertion or deletion position) at the corresponding position of the genome. For example, a mutation may be in an exon or regulatory region of a specific gene, while a copy number variation may affect the overall copy number of certain genes. Construction of a variation distribution map: A genomic variation distribution map is generated by mapping the position information of all variations to a specific region of the genome (such as a chromosome interval, a specific gene or a gene region). The map shows the spatial distribution of different types of variations in the genome, including the distribution of variations in specific chromosome regions and gene regions (such as exons and introns). The genomic variation distribution map helps researchers find areas where variations are concentrated or potential "hot spots". Counting variation type frequencies: After identifying and calibrating the location of the variation, the system will perform frequency statistics for each variation type. This process involves calculating the frequency of each variation type in all identified variations. For example, among all variations, mutations may account for 40%, insertions account for 30%, deletions account for 20%, and copy number variations account for 10%. This statistical result is used to reflect the frequency of different types of variations in the genome. Constructing variation type frequency graphs: Based on the variation frequencies obtained by statistics, the system constructs variation type frequency graphs. Frequency graphs are usually displayed in the form of bar graphs, pie charts, etc., showing the relative frequency of different variation types in the genome. This chart can clearly reflect which types of variation are more common and which are more rare in the genome being studied. Bar graph: Each bar represents a variation type, and the height of the bar represents the frequency of the variation type. Pie chart: Each sector represents a variation type, and the size of the sector represents the proportion of the variation type in all variations. Distribution feature analysis: The combination of genomic variation distribution map and variation type frequency map helps researchers understand the distribution characteristics of genomic variation. Through the distribution map, it can be found whether the variation has a clustering trend in certain specific areas (certain gene regions), and whether a certain variation type (mutation) is more common in the entire genome. Functional and pathological association: High-frequency variation types may be related to specific gene functions or to certain disease phenotypes. The variation distribution map can help researchers identify these potential association areas and provide clues for disease research.
[0037] Specifically, the specific process of performing distribution density analysis of gene regions on the gene structure map and identifying the clustering of genes in set regions in the genome is as follows: spatial density evaluation of each gene region in the gene structure map is performed through parallel calculation, the analysis region set in the genome is delineated, and the gene distribution in the region is quantified, and the distribution density of genes in each region is calculated through the K-means clustering algorithm; the degree of gene clustering is evaluated by calculating the clustering index of gene distribution in the set analysis region, and high-density gene clustering regions are identified; the distribution pattern of gene regions is further identified through spatial autocorrelation analysis, and the clustering trend of different genes in the set region is analyzed in combination with gene function annotation information.
[0038] In this embodiment, gene region definition and analysis region setting: the system will spatially divide and locate each gene region in the gene structure diagram. These gene regions may include exons, introns, regulatory regions, etc. When setting the analysis region, a fragment of a specific chromosome may be selected, or a region where a group of genes are located may be selected by functional correlation. The selection of the analysis region can be determined according to the research objectives (such as the clustering analysis of a specific gene family). Spatial density quantification: Once the analysis region is delineated, the system will quantify the gene distribution in the region. This quantification process mainly calculates the density of genes in the set region, that is, the number of genes in a unit space. For example, assuming that the set region is a certain section of a chromosome, the system will count the number of genes in the region according to the position coordinates of the genes, and then calculate the distribution density of the genes. K-means clustering: Based on the calibrated gene position data, the K-means clustering algorithm is used to analyze the distribution pattern of genes in the set region. In this step, the algorithm adjusts the center of the cluster through continuous iteration, and finally maximizes the aggregation of genes in each cluster, so that the region with higher gene density can be clearly identified. Explanation of clustering results: Through the results of K-means clustering, the system can identify the degree of clustering of genes in a specific area, and find out which areas have dense numbers of genes and which areas have sparse distribution of genes. These results help to further analyze the spatial distribution characteristics of genes and find potential areas of gene clustering. Clustering index: The clustering index is an indicator used to quantify the degree of clustering of genes in a set area. By calculating the gene distribution density of each area, combining the number of genes in the area and the area of the area, the system generates a clustering index. Areas with higher clustering indexes indicate that the distribution of genes in the area is more concentrated, and vice versa, it means that the distribution of genes in the area is more dispersed. Identification of high-density areas: Clustering index To quantify gene expression in a region The degree of aggregation within can be calculated by the spatial distance between genes. The specific formula is as follows: Formula parameter explanation: :area The total number of genes in indicates the number of genes involved in the analysis in this region. and :Gene and genes The spatial coordinates of and Represents genes and genes The position in the genome (can be a two-dimensional or three-dimensional coordinate, depending on the dimension of the data). :Gene and genes The spatial distance between gene pairs is calculated to evaluate their regional Relative position within. Euclidean distance formula: ; This formula calculates the gene and genes The straight-line distance between . : Normalization factor used to calculate the total number of gene pairs in a region. This factor represents the total number of gene pairs and ensures that the clustering index is based on the average number of gene pairs. Based on the clustering index, the system can further identify regions with high gene clustering. These high-density regions may correspond to regions that play an important role in gene function, or are related to certain genetic characteristics, disease susceptibility, etc. Spatial autocorrelation analysis: Spatial autocorrelation analysis is a statistical method used to detect spatial correlation between gene regions. Specifically, it can help identify the distribution pattern of genes in space and analyze whether there is a statistically significant correlation between genes in adjacent regions. For example, some genes may tend to be closer to each other in the genome, while some genes show a more dispersed distribution pattern. Local aggregation and diffusion: Through spatial autocorrelation analysis, researchers can understand whether there is a local aggregation phenomenon of genes (that is, some genes are concentrated in space) or a diffusion trend (that is, genes are evenly distributed throughout the region). This analysis is important for understanding the spatial structure and functional relationship of genes in the genome. Functional annotation information: The functional annotation information of genes usually includes the biological function of the gene, the metabolic pathway involved, and the association with the disease. The system can combine this functional information to further analyze the genes in the clustered region. For example, some functionally related genes (such as cancer-related genes, immune regulation genes, etc.) may show significant clustering patterns in the genome, and these clustered regions may have special biological significance. Analyze the clustering trend of different genes: Based on the results of clustering index and spatial autocorrelation analysis, combined with the functional annotation information of genes, researchers can deeply explore whether genes of specific functional categories have specific clustering trends in the genome. For example, some disease-related genes may show a high degree of clustering in the genome, thus providing clues for the study of disease mechanisms.
[0039] Specifically, the specific process of obtaining the gene density index is as follows: according to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; the gene aggregation situation in different areas is quantified by statistical methods to obtain the gene density index of each area.
[0040] In this embodiment, the calculation of the gene density index is based on the spatial distribution density of genes in the set gene region (analysis region). By quantifying the aggregation of genes in different regions, we can obtain the gene density index of each region. This density index not only reflects the distribution density of genes in space, but also can reveal the distribution characteristics of genes in different regions and their biological significance. The specific calculation process includes calculating the gene density value, statistical gene concentration, and combining the relevant formula to obtain the final gene density index. The gene density value represents the distribution density of genes in a certain set region, which can usually be calculated by the following formula: ;in: No. The gene density value of a region. No. The number of genes in a region. :No. The area or length of a region (depending on the representation of the gene structure diagram, it may be the length of a certain segment of the chromosome, or the area of a specific functional region). The larger the gene density value, the more concentrated the gene distribution in the region, and the more obvious the aggregation. The gene density index is a quantitative assessment of the gene aggregation in different regions. It is usually calculated by statistical methods. It combines the gene density value of each region and the spatial distribution characteristics of the genes in the region. The calculation formula is as follows: in: :No. The gene density index of a region. No. The number of genes in a region. :No. The area or length of a region. : Weight factor used to indicate the importance or relevance of each gene in the region. No. Region and The spatial similarity between genes indicates the degree of gene clustering in space. The larger the value, the more genes tend to cluster in this region.
[0041] Specifically, the specific process of predicting gene expression trends, mutation hotspot evolution and functional regional distribution based on the machine learning model is as follows: based on the gene expression difference index extracted by the heat map analysis unit, the dynamic change trend of gene expression under set conditions is predicted through the time series model; based on the mutation hotspot areas identified by the mutation map analysis unit, combined with the mutation type frequency map, the potential new mutation hotspot areas are predicted through the convolutional neural network; based on the gene density index of the structure map analysis unit, the enrichment areas of functionally related genes in the genome are predicted through the clustering algorithm.
[0042] In this implementation plan, gene expression trend prediction: time series modeling, task requirements: based on the gene expression difference index (dynamically changing data), predict the expression trend of genes under set time or conditions. Core algorithm adaptation, BayesLasso algorithm model role: feature screening implementation method: use Laplace prior (double exponential distribution) to compress insignificant SNP effects and screen out key regulatory sites related to the expression difference index. SNP refers to the variation of a single base (A, T, C, G) in the genome sequence. For example: select 500 high-impact sites from 500,000 SNPs as input features of the time series model. GBLUP algorithm model role: dynamic effect modeling implementation method: combine the genome relationship matrix (VanRaden method) with the time covariate to construct an extended mixed model: ; Parameter explanation: : phenotype vector; Fixed effects design matrix (e.g. experimental conditions, environmental factors); Fixed effect coefficients; : genotype matrix; : individual genomic breeding value (random effect); : Genome relationship matrix (calculated by VanRaden method); : residual term; : Genomic variance. Capture the genetic effects of gene expression differences over time and predict future trends. Capture the genetic effects of gene expression differences over time and predict future trends. RRBLUP algorithm model role: robust baseline prediction; implementation method: as a supplement to GBLUP, through ridge regression to penalize marker effects, provide conservative estimates of gene expression trends (to prevent overfitting). Mutation hotspot prediction: spatial pattern recognition. Task requirements: predict potential new hotspot areas based on known mutation hotspots and type frequencies. Core algorithm adaptation: BayesB / BayesC algorithm model role: probability feature generation, implementation method: calculate the probability of each SNP becoming a hotspot-related mutation through the π parameter (BayesC inference, BayesB preset). Generate a probability matrix as CNN input, for example: the probability of SNP hotspots in chromosome region A is 0.8, and region B is 0.05. GBLUP algorithm role: spatial feature enhancement. Implementation method: use the genomic relationship matrix as an additional channel input to CNN to help the model identify linkage disequilibrium patterns between variant sites. BayesLasso algorithm model role: auxiliary noise filtering. Implementation: After CNN prediction, perform secondary verification on high-probability hotspot areas to eliminate false positive signals (such as sparse regression verification of key SNPs). Functional enrichment region prediction: density clustering. Task requirements: Based on the gene density index, predict the clustered areas of functionally related genes. Core algorithm adaptation, GBLUP algorithm function: similarity matrix construction. Implementation: Use the genome relationship matrix to calculate the genetic similarity between regions as the weight input for spectral clustering. For example: If the genome similarity between regions A and B is 0.7, they are more likely to be classified into the same category when clustered. BayesLasso algorithm model function: landmark SNP screening. Implementation: Screen SNPs that are significantly associated with high-density areas as characteristic markers for clustering. For example: If a certain SNP appears frequently in 10 high-density areas, it will be given a higher clustering weight. RRBLUP algorithm model function: auxiliary density. Correction implementation: Input the gene density index as a phenotype into RRBLUP to estimate the marker effect, which is used to adjust the deviation in density calculation. By comprehensively analyzing the results of the above three prediction processes and integrating them using support vector machines, a comprehensive genome prediction map is generated. This map shows the overall picture of gene expression trends, mutation hotspot evolution, and functional regional distribution, helping researchers to fully understand the structure, function, and mutation characteristics of the genome. The generated map not only helps in the in-depth analysis of the genome, but also provides intuitive visualization results.
[0043] Specifically, the specific process of comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map is as follows: coordinates of gene expression difference distribution map, genome variation distribution map and gene structure map are normalized by parallel calculation, and genome positions of different data sources are mapped to a unified reference coordinate system; based on the statistical results of gene density index and variation type frequency map, a gene region weight matrix is constructed to dynamically adjust the overlay order of visualization layers; the predicted value of gene expression trend in genome prediction map is converted into a dynamic heat map, and the predicted heat map at different time points is adjusted and displayed through the timeline control, and the transparency is overlaid and compared with the gene expression heat map; in the genome variation distribution map, the predicted new variation hotspot area is overlaid on the known variation area with a pulse flashing mark, and a ring legend is generated through the variation type frequency map to display the type ratio difference of the predicted hotspot and the historical data in real time; the prediction result of the functional enrichment area in the gene structure map is converted into a 3D contour surface, which is superimposed on the two-dimensional gene density distribution map, the height of the surface is positively correlated with the gene density index, and the color gradient indicates the prediction confidence.
[0044] In this implementation scheme, the core goal of this visualization method is to integrate multiple genomic data so that researchers can intuitively understand the distribution and prediction trends of gene expression, gene variation and functional enrichment regions. The specific process is as follows: Coordinate normalization: Since genomic data comes from different data sources (such as gene expression data, gene variation data, and gene structure data), the genomic positions of these data may be based on different reference coordinates. Parallel computing is used to normalize the coordinates of all data to ensure that the information from different data sources can be mapped to a unified genomic reference coordinate system, which is convenient for subsequent integration and comparison. Construction of gene region weight matrix: Gene density index: represents the distribution density of genes in a certain region. Variant type frequency map: Statistical frequency of different types of variations (such as SNP, insertion / deletion) in a specific gene region. Through these statistical results, a gene region weight matrix is constructed to dynamically adjust the superposition order of different visualization layers to make the information of important regions more prominent. Dynamic heat map for predicting gene expression trends: Gene expression trend prediction values (from time series models) are used to generate dynamic heat maps. Through the timeline control, users can view the prediction of gene expression at different time points. The prediction heat map is overlaid with the gene expression heat map for transparency, allowing users to intuitively compare the difference between the predicted value and the experimental data and evaluate the accuracy of the prediction. Prediction and visualization of mutation hotspots: In the genomic variation distribution map, the new mutation hotspots predicted by the convolutional neural network are highlighted with pulse flashing markers to enhance visual guidance. The mutation type frequency map generates a ring legend to display the proportional difference between the predicted hotspots and historical data in the mutation type in real time, allowing researchers to quickly determine whether the new mutation conforms to the known mutation pattern. 3D contour surface displays functional enrichment regions: Based on the gene density index, the functional enrichment regions are predicted and converted into 3D contour surfaces for visualization. The surface height is proportional to the gene density index, indicating the enrichment of functionally related genes in the genome. The color gradient indicates the prediction confidence, and the high confidence area is darker, which is convenient for identifying important areas. The integration of different genomic data is achieved through coordinate normalization, and the gene region weight matrix ensures the optimization of information hierarchy. The use of dynamic heat maps, flashing markers, ring legends, 3D contour surfaces and other methods makes the visualization of gene expression, variation hotspots, and functional enrichment regions more intuitive, providing an efficient analysis tool for genomic research.
[0045] In summary, this application has at least the following effects: The parallel computing and data visualization analysis system for whole genome prediction effectively processes and analyzes target whole genome data, including gene expression, gene variation and gene structure data, through parallel computing and data visualization algorithms, providing an efficient data processing and computing platform for genomic research. The system can accurately analyze the spatial distribution and color intensity of genome heat maps, identify the differential expression areas of genes under different experimental conditions, and reflect the changes in gene expression levels under set conditions through the gene expression difference index, thereby providing strong support for gene function research. By analyzing the genome variation map, it can extract gene variation characteristics and identify variation types, such as mutations, insertions, deletions, etc. At the same time, the distribution of different variation types in the genome and their variation hotspots are displayed through the variation type frequency map, helping researchers understand the spatial distribution pattern of gene variation. The structural map analysis module can reveal the distribution characteristics and aggregation trends of genes in different regions through spatial density evaluation and clustering index calculation, providing a scientific basis for studying the regional characteristics of the genome and gene function annotation. Through the comprehensive visualization module, data such as the gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map and gene density index are displayed in a unified manner, using spatial distribution, color depth and morphological characteristics to enable researchers to intuitively observe the various characteristics of the genome and their interrelationships.
[0046] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0047] The present invention is described with reference to flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0048] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0049] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0050] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0051] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A parallel computing and data visualization analysis system for whole genome prediction, characterized by: It includes the following modules: image generation module, image analysis module, genome prediction module, and visualization module; The image generation module is used to obtain target whole genome data, and convert the target whole genome data into whole genome sequencing images, including genome heat maps, genome variation maps, and gene structure maps, through data visualization algorithms and parallel computing; The image analysis module includes a heat map analysis unit, a variation map analysis unit, and a structure map analysis unit; The heat map analysis unit is used to perform spatial distribution and color intensity analysis on the genome heat map, identify the differential expression regions of genes under set conditions, extract the gene expression differential distribution map in the genome heat map, and obtain the gene expression differential index to reflect the expression level changes of genes under set conditions; The variation map analysis unit is used to extract the gene variation features in the genome variation map, identify the variation type and the distribution pattern in the genome, obtain the genome variation distribution map and variation type frequency map, and reflect the variation hotspot areas of various variation types in the genome; The structure graph analysis unit is used to analyze the distribution density of gene regions on the gene structure graph, identify the aggregation of genes in a set region in the genome, and obtain a gene density index to reflect the distribution characteristics of genes in different genome regions; The genome prediction module is used to receive the output data of the image analysis module, predict the gene expression trend, mutation hotspot evolution and functional area distribution according to the machine learning model, and generate a genome prediction map; The visualization module is used to comprehensively visualize the gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map.
2. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of converting target whole genome data into whole genome sequencing images through data visualization algorithms and parallel computing is as follows: The target whole genome data is loaded through parallel computing, and the target whole genome data is cleaned, filtered and standardized to handle outliers and missing data; The target whole genome data includes gene expression data, gene variation data and gene structure data; Generate a gene expression heat map based on the gene expression data using a heat map visualization algorithm; According to the gene variation data, the genome variation map is constructed through the variation map visualization algorithm; Based on the gene structure data, the gene structure diagram is drawn through the gene structure visualization algorithm to mark the various components of the gene and their spatial positions; The generated genome heatmaps, genome variation maps, and gene structure maps were integrated into a comprehensive whole-genome sequencing image.
3. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of analyzing the spatial distribution and color intensity of the genome heat map, identifying the differential expression regions of genes under set conditions, and extracting the differential distribution map of gene expression in the genome heat map is as follows: Through parallel computing, the spatial position of each gene region in the genome heat map is analyzed to evaluate the expression density and distribution pattern of genes in different regions and identify the expression differences of genes in different regions. Quantitatively analyze the color intensity of each pixel in the heat map to determine the relationship between different color intensities and gene expression levels, and identify differences in expression intensity through color changes in the heat map; Based on the changes in color intensity, the expression difference threshold is set to identify the gene regions that show upregulation or downregulation under the set conditions; Gene expression difference distribution maps are extracted from the identified differentially expressed regions to display the expression changes of genes in different regions and annotate the changed gene regions.
4. The parallel computing and data visualization analysis system for whole genome prediction according to claim 3, characterized in that: The specific process of obtaining the gene expression difference index is as follows: By quantitatively analyzing the color intensity of each gene region in the genome heat map, the expression value of each gene in different regions is extracted. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level; Statistical analysis was used to perform differential analysis on each gene in the identified differentially expressed region. By comparing the expression levels of the genes under the set conditions, the expression change of each gene under the preset experimental conditions was calculated to obtain the gene expression differential index.
5. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of extracting gene variation features from the genome variation map and identifying variation types and distribution patterns in the genome is as follows: Extract features from each variant region in the genomic variation map, analyze the types of variation, including mutation, insertion, deletion, and copy number variation, and extract their spatial distribution information; By calculating the distribution of mutations in different gene regions and chromosome regions, the mutation hotspots are found.
6. The parallel computing and data visualization analysis system for whole genome prediction according to claim 5, characterized in that: The specific process of obtaining the genome variation distribution map and variation type frequency map is as follows: Integrate the identified gene variation types and location data to construct a genome variation distribution map to show the distribution of different variation types in the genome; The frequency of each variant type is statistically analyzed to obtain a variant type frequency graph, which displays the frequency of each variant type in the genome, reflecting the frequency of different variant types and their distribution characteristics in the genome.
7. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of analyzing the distribution density of gene regions on the gene structure map and identifying the aggregation of genes in a set region in the genome is as follows: Through parallel computing, the spatial density of each gene region in the gene structure map is evaluated, the analysis region set in the genome is delineated, and the gene distribution in the region is quantified. The distribution density of genes in each region is calculated by the K-means clustering algorithm; By calculating the clustering index of gene distribution in the set analysis area, the degree of gene clustering is evaluated and high-density gene clustering areas are identified; Through spatial autocorrelation analysis, the distribution patterns of gene regions are further identified, and combined with gene function annotation information, the aggregation trends of different genes in the set region are analyzed.
8. The parallel computing and data visualization analysis system for whole genome prediction according to claim 7, characterized in that: The specific process of obtaining the gene density index is as follows: According to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; The gene clustering in different regions was quantified by statistical methods to obtain the gene density index of each region.
9. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of predicting gene expression trends, mutation hotspot evolution, and functional regional distribution based on machine learning models is as follows: According to the gene expression difference index extracted by the heat map analysis unit, the dynamic change trend of gene expression under set conditions is predicted through the time series model; Based on the mutation hotspots identified by the mutation map analysis unit and combined with the mutation type frequency map, a convolutional neural network is used to predict potential new mutation hotspots. According to the gene density index of the structural graph analysis unit, the enriched regions of functionally related genes in the genome are predicted by clustering algorithm.
10. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map is as follows: Through parallel computing, coordinates of gene expression difference distribution map, genome variation distribution map, and gene structure map are normalized, and the genome positions of different data sources are mapped to a unified reference coordinate system; Based on the statistical results of gene density index and variation type frequency map, a gene region weight matrix is constructed to dynamically adjust the stacking order of visualization layers; The predicted gene expression trend values in the genome prediction map are converted into a dynamic heat map, and the predicted heat map at different time points is adjusted and displayed through the timeline control, and the transparency is superimposed and compared with the gene expression heat map; In the genome variation distribution map, the predicted new variation hotspots are overlaid on the known variation regions with pulse flashing marks, and a ring legend is generated through the variation type frequency map to display the type ratio difference between the predicted hotspots and historical data in real time; The prediction results of functional enrichment regions in the gene structure map are converted into 3D contour surfaces and superimposed on the two-dimensional gene density distribution map. The height of the surface is positively correlated with the gene density index, and the color gradient represents the prediction confidence.
Citation Information
Patent Citations
Genome-wide methylation high-throughput sequencing method
CN108034705A
Systems and methods for genomic analysis
CN108350494A
Genome abnormality visualization system and genomic abnormality visualization method
JP2024036206A
Methods for the graphical representation of genomic sequence data
US20160342737A1
Systems and Methods for Producing Quantitatively Calibrated Grayscale Values in Magnetic Resonance Images
US20180325461A1
Cited By
Genome data analysis method
CN120199337A