Parallel computing and data visualization analysis system for whole genome prediction
Through the parallel computing and data visualization analysis system of whole genome prediction, the problem of low efficiency in processing massive genomic data has been solved, efficient and accurate genomic data analysis and intuitive visualization have been achieved, and the efficiency and accuracy of genomic research have been improved.
Patent Information
- Application Number
- CN202510437299.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing technologies have low computational efficiency when processing massive genomic data. Traditional computing methods cannot effectively handle it, and there is a lack of intuitive multi-dimensional data visualization tools, making it difficult to meet the computational efficiency and data management needs of genomic research.
The parallel computing and data visualization analysis system for whole genome prediction is adopted, including image generation module, image analysis module and visualization module. It converts genomic data into multiple images through parallel computing and data visualization algorithms, combines machine learning models to perform genome prediction, and conducts comprehensive visualization display.
It improves the efficiency and accuracy of genomic data analysis, can quickly identify key mutation hotspots and gene clustering areas, enhances research flexibility and decision-making efficiency, and makes genomic data interpretation more intuitive and easy to understand.
Smart Images

Figure CN119964655B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of genetic image data analysis, in particular to a parallel computing and data visualization analysis system for whole genome prediction. BACKGROUND
[0002] With the continuous development of precision medicine, personalized medicine and genomics research, the application prospect of genomic data in the fields of disease prediction, individualized treatment, drug development and agricultural breeding is increasingly extensive. As one of the important tools, genomic prediction technology can accurately predict the health status, genetic disease risk and other biological characteristics of individuals through comprehensive analysis of individual genomic information. However, with the continuous expansion of the scale of genomic data, traditional computing methods and analysis tools are facing great challenges. Especially in the case of handling massive genomic data, how to provide scientific analysis results through efficient computing means and accurate visualization technology has become the key to promoting genomics research and application. Therefore, based on high-performance computing and innovative data visualization technology, it is becoming the core solution to this problem.
[0003] At present, the analysis method of genomic prediction mostly depends on traditional single-computer computing or inefficient distributed computing architecture, which has the problems of low computing efficiency and long response time. When genomic data is analyzed, especially when whole genome prediction and high-throughput genomic data processing are performed, the amount of calculation is huge, and traditional computing methods cannot effectively process these massive data, resulting in slow analysis process and consumption of a large amount of resources. Secondly, most of the existing genomic data visualization tools are limited to static graphical display, lack of interactivity and multi-dimensional display, and cannot effectively present the internal correlation of high-dimensional genomic data. Especially when dealing with genomic diversity and complex gene-environment interaction, there is a lack of intuitive and easy-to-understand visualization methods. In addition, the bottleneck of data storage and processing also limits the application of existing methods in large-scale parallel computing environment, making it difficult to meet the dual needs of computing efficiency and data management in genomics research. Therefore, the existing methods are low in efficiency, insufficient in processing capacity, and lack of effective data visualization tools when facing genomic data analysis and prediction, which limits their wide application and in-depth research. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a parallel computing and data visualization analysis system for whole genome prediction, which solves the problems in the above background art.
[0005] In order to achieve the above object, the application is implemented by the following technical solutions: a parallel computing and data visualization analysis system for whole genome prediction, comprising the following modules: an image generation module, an image analysis module, a genome prediction module, and a visualization module; the image generation module is used to obtain target whole genome data, and the target whole genome data is converted into whole genome sequencing images, including a genome heat map, a genome variation map, and a gene structure map, through a data visualization algorithm and parallel computing; the image analysis module comprises a heat map analysis unit, a variation map analysis unit, and a structure map analysis unit; the heat map analysis unit is used to analyze the spatial distribution and color intensity of the genome heat map, identify the expression difference region of the gene under the set condition, extract the gene expression difference distribution map in the genome heat map, and obtain the gene expression difference index for reflecting the expression level change of the gene under the set condition; the variation map analysis unit is used to extract the gene variation characteristics in the genome variation map, identify the variation type and the distribution mode in the genome, obtain the genome variation distribution map and the variation type frequency map for reflecting the variation hotspot region of each type of variation in the genome; the structure map analysis unit is used to analyze the distribution density of the gene region of the gene structure map, identify the aggregation condition of the gene in the set region of the genome, and obtain the gene density index for reflecting the distribution characteristics of the gene in different genome regions; the genome prediction module is used to receive the output data of the image analysis module, predict the gene expression trend, variation hotspot evolution, and functional region distribution according to a machine learning model, and generate a genome prediction map; and the visualization module is used to comprehensively visually display the gene expression difference distribution map, the gene expression difference index, the genome variation distribution map, the variation type frequency map, the gene density index, and the genome prediction map.
[0006] Further, the specific process of converting the target whole genome data into the whole genome sequencing image through the data visualization algorithm and the parallel computing is as follows: the target whole genome data is loaded through parallel computing, and the target whole genome data is cleaned, filtered, and standardized to process abnormal values and missing data; the target whole genome data comprises gene expression data, gene variation data, and gene structure data; the gene expression heat map is generated through a heat map visualization algorithm according to the gene expression data; the genome variation map is constructed through a variation map visualization algorithm according to the gene variation data; the structure map of the gene is drawn through a gene structure visualization algorithm according to the gene structure data, and each component and the spatial position of the gene are marked; and the generated genome heat map, genome variation map, and gene structure map are integrated into a comprehensive whole genome sequencing image.
[0007] Further, the specific process of extracting the gene expression difference distribution map in the genome heat map is as follows: the spatial position of each gene region in the genome heat map is analyzed through parallel computing, the expression density and distribution pattern of genes in different regions are evaluated, and the expression difference of genes in different regions is identified; the color intensity of each pixel in the heat map is quantitatively analyzed to determine the relationship between different color intensities and gene expression levels, and the expression intensity difference is identified through color change in the heat map; according to the change of color intensity, the expression difference threshold is set to identify the gene region that shows up-regulation or down-regulation under the set condition; the gene expression difference distribution map is extracted from the identified expression difference region to show the expression change of genes in different regions, and the changed gene region is labeled.
[0008] Further, the specific process of obtaining the gene expression difference index is as follows: the color intensity of each gene region in the genome heat map is quantitatively analyzed to extract the expression value of each gene in different regions. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level; for each gene in the identified expression difference region, difference analysis is performed using statistical analysis, the expression level of the gene under the set condition is compared, the expression change of each gene under the preset experimental condition is calculated, and the gene expression difference index is obtained.
[0009] Further, the specific process of extracting the gene variation characteristics in the genome variation map and identifying the variation type and the distribution pattern in the genome is as follows: the characteristics of each variation region in the genome variation map are extracted, the type of variation is analyzed, including mutation, insertion, deletion and copy number variation, and the spatial distribution information is extracted; by calculating the distribution of variations in different gene regions and chromosome regions, the hot spot region of the variation is obtained.
[0010] Further, the specific process of obtaining the genome variation distribution map and the variation type frequency map is as follows: the identified gene variation type and position data are integrated to construct the genome variation distribution map, which shows the distribution of different variation types in the genome; the frequency of each variation type is statistically analyzed to obtain the variation type frequency map, which shows the frequency of each variation type in the genome through charts, reflecting the frequency and distribution characteristics of different variation types in the genome.
[0011] Furthermore, the distribution density of gene regions in the gene structure map is analyzed to identify the specific process of clustering of genes in set regions in the genome as follows: spatial density of each gene region in the gene structure map is evaluated through parallel calculation, the analysis region set in the genome is delineated, and the gene distribution in the region is quantified, and the distribution density of genes in each region is calculated by the K-means clustering algorithm; the degree of gene clustering is evaluated by calculating the clustering index of the gene distribution in the set analysis region, and high-density gene clustering regions are identified; spatial autocorrelation analysis is used to further identify the distribution pattern of gene regions, and combined with gene function annotation information, the clustering trend of different genes in the set region is analyzed.
[0012] Furthermore, the specific process of obtaining the gene density index is as follows: according to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; the gene aggregation situation in different areas is quantified by statistical methods to obtain the gene density index of each area.
[0013] Furthermore, the specific process of predicting gene expression trends, mutation hotspot evolution and functional region distribution based on the machine learning model is as follows: based on the gene expression difference index extracted by the heat map analysis unit, the dynamic change trend of gene expression under set conditions is predicted through the time series model; based on the mutation hotspot areas identified by the mutation map analysis unit, combined with the mutation type frequency map, potential new mutation hotspot areas are predicted through the convolutional neural network; based on the gene density index of the structure map analysis unit, the enrichment areas of functionally related genes in the genome are predicted through the clustering algorithm.
[0014] Furthermore, the specific process of comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map is as follows: coordinates of gene expression difference distribution map, genome variation distribution map and gene structure map are normalized through parallel computing, and the genome positions of different data sources are mapped to a unified reference coordinate system; based on the statistical results of gene density index and variation type frequency map, a gene region weight matrix is constructed to dynamically adjust the overlay order of visualization layers; the predicted gene expression trend values in the genome prediction map are converted into a dynamic heat map, and the predicted heat map at different time points is adjusted and displayed through the timeline control, and the transparency is superimposed and compared with the gene expression heat map; in the genome variation distribution map, the predicted new variation hotspot area is overlaid on the known variation area with a pulse flashing mark, and a ring legend is generated through the variation type frequency map to display the type ratio difference between the predicted hotspot and the historical data in real time; the prediction results of the functional enrichment region in the gene structure map are converted into a 3D contour surface and superimposed on the 2D gene density distribution map. The height of the surface is positively correlated with the gene density index, and the color gradient represents the prediction confidence.
[0015] The present invention has the following beneficial effects:
[0016] (1) The parallel computing and data visualization analysis system for whole genome prediction, the image generation module, based on parallel computing and data visualization algorithms, converts complex genomic data into a variety of genomic images (genome heat maps, variation maps, gene structure maps), effectively reducing the time delay of traditional methods in data processing. Through parallel computing technology, the system can process a large amount of genomic data simultaneously, accelerating the image generation process, thereby greatly improving the efficiency and real-time performance of data analysis. In the image analysis module, the heat map analysis, variation map analysis and structure map analysis units accurately extract expression differences, gene variation characteristics and distribution density of gene regions in genomic images. These analysis results can provide in-depth support for gene function research, variation detection and genome structure analysis, helping researchers identify key variation hotspots and gene clustering areas in the genome, greatly improving the accuracy and operability of genomic data analysis.
[0017] (2) The parallel computing and data visualization analysis system for whole genome prediction. The genome prediction module uses an algorithm model to perform genome prediction, which can accurately predict gene variation, expression differences and possible functional regions. By working in conjunction with other modules, the prediction results can be combined with image analysis results in real time to provide researchers with comprehensive genome information. This module greatly improves the prediction accuracy and can quickly identify key variants with potential biological significance in large-scale genome data, further optimizing research decisions and improving the efficiency and accuracy of genome research. The visualization module makes the interpretation of genome data more intuitive and easy to understand by integrating the visualization of different analysis results. This module not only optimizes the data display method, but also enhances researchers' understanding of complex information such as genome structure, expression differences and variation distribution, helps to quickly discover potential gene variations and key regions, and improves the flexibility and decision-making efficiency of research.
[0018] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of the parallel computing and data visualization analysis system for whole genome prediction of the present invention. DETAILED DESCRIPTION
[0020] The embodiment of the application solves the problems of low processing efficiency, insufficient analysis accuracy and poor data integration of genomic data by a parallel computing and data visualization analysis system for whole genome prediction. By combining parallel computing and data visualization algorithms, the system can efficiently convert large-scale whole genome data into intuitive genomic heat maps, variation maps and gene structure maps, effectively improving the speed and accuracy of image generation. At the same time, the system uses an image analysis module to automatically extract key information such as gene expression differences, variation regions and gene density, reducing manual intervention and improving the automation and accuracy of data analysis. Through the comprehensive display of the visualization module, researchers can intuitively observe the changes in genomic data, promoting the efficient conduct of genomics research.
[0021] The overall idea of the scheme in the embodiment of the application is as follows:
[0022] The target whole genome data is obtained, and the target whole genome data is converted into whole genome sequencing images, including a genomic heat map, a genomic variation map and a gene structure map, by a data visualization algorithm and parallel computing.
[0023] The spatial distribution and color intensity of the genomic heat map are analyzed, the expression difference region of the gene under the set condition is identified, the gene expression difference distribution map in the genomic heat map is extracted, and the gene expression difference index is obtained.
[0024] The gene variation characteristics in the genomic variation map are extracted, the variation types and the distribution patterns in the genome are identified, the genomic variation distribution map and the variation type frequency map are obtained, and the genomic variation distribution map and the variation type frequency map are used to reflect the variation hotspot regions of various variation types in the genome.
[0025] The distribution density of the gene region in the gene structure map is analyzed, the aggregation of the gene in the set region in the genome is identified, the gene density index is obtained, and the gene density index is used to reflect the distribution characteristics of the gene in different genomic regions.
[0026] The gene expression difference distribution map, the gene expression difference index, the genomic variation distribution map, the variation type frequency map and the gene density index are comprehensively visualized.
[0027] Please refer to Figure 1The embodiment of the present application provides a technical scheme: a parallel computing and data visualization analysis system for whole genome prediction, comprising the following modules: an image generation module, an image analysis module, a genome prediction module, and a visualization module; the image generation module is used to obtain target whole genome data, and the target whole genome data is converted into whole genome sequencing images, including a genome heat map, a genome variation atlas, and a gene structure diagram, through a data visualization algorithm and parallel computing; the image analysis module comprises a heat map analysis unit, a variation atlas analysis unit, and a structure diagram analysis unit; the heat map analysis unit is used to analyze the spatial distribution and color intensity of the genome heat map, identify the expression difference region of genes under a set condition, extract the gene expression difference distribution diagram in the genome heat map, and obtain a gene expression difference index, which is used to reflect the expression level change of genes under the set condition; the variation atlas analysis unit is used to extract the gene variation characteristics in the genome variation atlas, identify the variation type and the distribution mode in the genome, obtain a genome variation distribution diagram and a variation type frequency diagram, and reflect the variation hotspot region of various types of variations in the genome; the structure diagram analysis unit is used to analyze the distribution density of the gene region of the gene structure diagram, identify the aggregation condition of genes in the set region of the genome, obtain a gene density index, and reflect the distribution characteristics of genes in different genome regions; the genome prediction module is used to receive the output data of the image analysis module, predict the gene expression trend, the variation hotspot evolution, and the functional region distribution according to a machine learning model, and generate a genome prediction atlas; and the visualization module is used to comprehensively visualize and display the gene expression difference distribution diagram, the gene expression difference index, the genome variation distribution diagram, the variation type frequency diagram, the gene density index, and the genome prediction atlas.
[0028] In this embodiment, the image generation module: the main function of this module is to generate genomic sequencing images from target whole genome data through data visualization algorithms and parallel computing techniques. These images include genomic heatmaps, genomic variation maps, and gene structure diagrams, which are used to display genomic information. Parallel computing techniques enable the module to efficiently process large-scale genomic data and transform complex genetic information into visual images suitable for subsequent analysis. In this way, the computational bottleneck of traditional genomic data processing methods is broken, and image results can be generated in a shorter time. Genomic heatmaps: numerical information of gene expression is converted into heatmap format, and expression levels of genes under different conditions are displayed based on color intensity. Genomic variation map: the variation data in the genome is presented in the form of a map, revealing the types and distribution of gene variations. Gene structure diagram: shows the spatial distribution and structural characteristics of genes in the genome, helping users understand the relationship between genes. Image analysis module: this module is responsible for detailed analysis of the above-mentioned generated genomic images (such as heatmaps, variation maps and gene structure diagrams). Based on advanced image processing and machine learning algorithms, the image analysis module extracts important information from the image and further analyzes it to provide valuable interpretation of genomic data for users. Heatmap analysis unit: this unit uses image analysis algorithms to analyze the spatial distribution and color intensity of the genomic heatmaps, identifying areas of gene expression difference under the given conditions. By extracting gene expression difference distribution map and calculating gene expression difference index, this module can reflect the expression level changes of genes under specific conditions. Variation map analysis unit: this unit extracts gene variation features from the genomic variation map and uses pattern recognition algorithms to identify variation types and analyze their distribution patterns in the genome. Finally, generate genomic variation distribution map and variation type frequency map to reflect hotspots of different variation types in the genome. Structure diagram analysis unit: this unit mainly analyzes the spatial distribution of gene structure diagram, identifies the aggregation of genes in different genomic regions, and calculates gene density index using gene region distribution density analysis algorithm to reflect the distribution characteristics of genes in different genomic regions. The visualization module is used to comprehensively display all the data generated by the image analysis module (such as gene expression difference distribution map, gene expression difference index, genomic variation distribution map, variation type frequency map and gene density index). Through this module, users can intuitively view various analysis results and effectively compare, filter and support decision-making. This module uses advanced image visualization technology to present complex analysis results in easy-to-understand graphs or charts, making genomic data analysis results clearer and more intuitive. Parallel computing refers to a computing method that solves problems by working simultaneously through multiple processors or computers. In the process of genomic data processing, due to the large amount of data, traditional serial computing methods may not be able to efficiently complete the computing task.Parallel computing divides tasks into small pieces, assigns them to multiple processing units for parallel processing, greatly improving computing efficiency. In this system, parallel computing is used to accelerate the processing of genomic data, ensuring that image results of large-scale genomes can be generated quickly, especially when dealing with high-dimensional data such as genomic heat maps, variation maps, etc. Data visualization algorithms refer to algorithms that convert data into graphics or images that are easy to understand and analyze through specific computing methods. These algorithms will select appropriate colors, shapes, sizes, and layouts to represent different dimensions of data based on the characteristics of the data, such as gene expression levels, variation types, and gene structures. In this system, data visualization algorithms are used to convert genomic data into heat maps, variation maps, and structure maps, etc. image format, so that complex genomic information can be intuitively presented to researchers. Heat map is a two-dimensional data visualization method that uses color depth or different colors to represent numerical values. In genomics, heat maps are often used to show changes in gene expression. The darker the color, the higher the gene expression, and the lighter the color, the lower the expression. This method facilitates the observation of gene expression differences under different conditions, helping researchers identify key genes and potential biological laws. Genomic variation map refers to a map that displays variations (single nucleotide polymorphisms SNP, insertions / deletions Indel, structural variations) in the genome in a graphical manner. In this system, the variation map analysis unit is responsible for identifying and extracting variation characteristics in the genome and presenting their distribution in the genome. Through the variation map, researchers can quickly locate variation hotspots and further study the relationship between different variations and diseases or other biological characteristics. Variation type frequency chart is a chart used to display the frequency distribution of different types of variations in the genome. Common variation types include point mutations, insertions / deletions, structural variations, etc. This chart can help researchers quickly understand which variation types appear most frequently in the entire genome, providing data support for further research. Gene density index is an indicator that measures the distribution of genes in a specific genomic region. Areas with high gene density mean that genes in that region are more concentrated, while areas with low gene density mean that genes are more dispersed. This index helps analyze the distribution characteristics of genes in the entire genome and identify gene regions that may have special functions or biological significance. Image analysis algorithms are a class of algorithms used to extract information from images. In this invention, image analysis algorithms are used to process the generated genomic heat maps, variation maps, and structure maps to extract valuable information. Common image analysis methods include edge detection, segmentation, feature extraction, etc., with the goal of identifying and quantifying key biological features from images. For example, through heat map analysis algorithms, you can automatically identify areas of gene expression difference and calculate gene expression difference index.
[0029] Specifically, the specific process of transforming the target whole genome data into the whole genome sequencing image through the data visualization algorithm and parallel computing is as follows: loading the target whole genome data through parallel computing, and cleaning, filtering and standardizing the target whole genome data, processing abnormal values and missing data; the target whole genome data includes gene expression data, gene variation data and gene structure data; according to the gene expression data, a gene expression heat map is generated through a heat map visualization algorithm; according to the gene variation data, a genome variation map is constructed through a variation map visualization algorithm; according to the gene structure data, a structure map of the gene is drawn through a gene structure visualization algorithm, indicating each component of the gene and its spatial position; and the generated genome heat map, genome variation map and gene structure map are integrated into a comprehensive whole genome sequencing image.
[0030] In this embodiment, parallel computing: the target whole genome data volume is huge, including gene expression data, gene variation data and gene structure data, and the single machine serial computing processing speed may not meet the demand. Therefore, parallel computing is adopted to distribute data processing tasks to multiple computing nodes or processors for simultaneous execution, greatly accelerating the speed of data processing. Data cleaning and filtering: the target data often contains noise or error information, which needs to be removed by cleaning. The filtering process helps to screen out effective data for analysis. Standardization: gene data from different sources or different measurement methods may have different ranges and units, so standardization processing is needed to ensure data consistency and facilitate subsequent analysis. Handling of outliers and missing data: there may be outliers (such as measurement errors) or missing values (such as some genes not measured in some samples) in gene data, which need to be handled by reasonable methods (such as mean filling, interpolation or deletion) to avoid bias in the analysis results. Heatmap visualization algorithm: this algorithm converts the differences in gene expression into color intensity according to the numerical size of gene expression data. Areas with high gene expression values are represented by dark colors, and areas with low gene expression values are represented by light colors, allowing researchers to visually view which genes have changed expression under different conditions. Through the heat map, researchers can identify the expression differences of genes in specific areas or conditions, and further study the effects of these genes on organisms or diseases. Variant map visualization algorithm: this algorithm draws a map of genomic variations based on gene variation data (such as SNPs, insertions / deletions, etc.). Each type of variation is marked by different symbols, colors or shapes, helping researchers identify and understand hotspots of genomic variations and distribution of variation types. The variant map shows the types and frequencies of variations at different locations in the genome, providing important information for further disease research or gene function research. Gene structure visualization algorithm: this algorithm draws a complete structure map of the gene based on gene structure data (such as exons, introns, regulatory regions, etc.), showing the various components of the gene and their spatial location in the genome. The gene structure map helps reveal the functional regions of the gene, structural variations (such as splice isoforms) and their association with function. Integrated image: once the genome heat map, variant map and gene structure map have been generated, the next step is to integrate these image information into a comprehensive whole genome sequencing image. This image will contain multi-dimensional information, showing the expression levels, variation characteristics and structural layout of genes, allowing researchers to easily view the overall situation of the whole genome. Comprehensive visualization display: by displaying these image information in the same interface, researchers can conduct comprehensive analysis of various features of the genome. For example, differences in gene expression may be related to the location and structural characteristics of certain gene variations, and through such a comprehensive view, researchers can discover potential associations between different gene features, thereby deepening their understanding of the function of the genome.
[0031] Specifically, the spatial distribution and color intensity of the genome heat map are analyzed to identify the expression difference regions of genes under the set conditions. The specific process of extracting the gene expression difference distribution map in the genome heat map is as follows: the spatial position of each gene region in the genome heat map is analyzed through parallel computing, the expression density and distribution pattern of genes in different regions are evaluated, and the expression difference of genes in different regions is identified; the color intensity of each pixel in the heat map is quantitatively analyzed to determine the relationship between different color intensities and gene expression levels, and the expression intensity difference is identified through the color change in the heat map; according to the change of color intensity, the expression difference threshold is set to identify the gene regions that show up-regulation or down-regulation under the set conditions; the gene expression difference distribution map is extracted from the identified expression difference regions to show the expression changes of genes in different regions, and the changed gene regions are labeled.
[0032] In this embodiment, spatial position analysis: The spatial positions of each gene region in the genomic heat map are analyzed. The position of each gene in the map can be related to its expression pattern or characteristics. By analyzing these spatial positions, the expression density and distribution pattern of genes within different regions can be assessed, and the expression differences of genes within these regions can be identified. Expression density: In the heat map, the expression density of a gene represents the level of gene expression in a particular region. By evaluating the expression density of different gene regions, it can be identified which regions have more intensive gene expression and which regions have less intensive gene expression, thereby revealing the spatial distribution characteristics of genes. Distribution pattern: Not only the density of expression is analyzed, but also the distribution pattern of genes in the heat map is identified. Gene expression can exhibit a specific pattern, such as high expression in certain regions and lower expression in other regions. By analyzing the distribution pattern, the differences in gene expression within different regions can be identified, further helping to understand the relationship between gene function and the region it is located in. Color intensity analysis: The heat map represents the expression level of genes through color. Usually, the color intensity of the heat map (such as from light to dark) is proportional to the expression level of genes. By quantifying the color intensity of each pixel, the expression intensity of genes within different regions can be calculated. Color and expression level relationship: Determine the relationship between different color intensities and gene expression levels, for example, deep red may represent high expression, while light yellow may represent low expression. This analysis helps researchers to convert color intensity into quantitative expression level, thereby accurately identifying the differences in expression intensity of genes. Threshold setting: Based on the changes in color intensity, a threshold for expression difference can be set. This threshold is used to distinguish the up-regulated and down-regulated regions of genes. By setting a specific numerical range, regions where gene expression changes significantly under certain experimental conditions can be identified. Up-regulation and down-regulation identification: After setting the threshold, it can be identified which genes exhibit up-regulation (i.e., increased expression) or down-regulation (i.e., decreased expression) under certain conditions. This analysis helps to find the changes in gene expression related to certain biological processes, diseases or external stimuli. Expression difference region extraction: From the heat map, the expression difference regions (up-regulated or down-regulated regions) identified are further extracted, and a gene expression difference distribution map is generated. This map shows the expression changes of genes under different regions or conditions. Displaying expression changes: In the difference distribution map, the changes in gene expression are marked by different colors or graphics, helping researchers to visually see which genes have undergone significant changes in expression under certain regions or conditions. The gene expression difference distribution map not only shows the changes in expression, but also marks the regions of the changed genes, thereby providing data support for subsequent biological analysis.
[0033] Specifically, the specific process of obtaining the gene expression difference index is as follows: by quantitatively analyzing the color intensity of each gene region in the genome heat map, the expression value of each gene in different regions is extracted. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level; for each gene in the identified expression difference region, statistical analysis is used for difference analysis, by comparing the expression level of the gene under the set condition, the expression change amount of each gene under the preset experimental condition is calculated, and the gene expression difference index is obtained.
[0034] In this embodiment, the color intensity of each gene region in the genome heat map is quantitatively analyzed. There is a mapping relationship between the color intensity in the heat map and the gene expression level. Assuming that the color intensity of a certain gene region in the heat map is , wherein represents the gene number, represents the position of the gene in different regions (such as regions at different time points or under different conditions). The expression value of the gene in different regions is calculated: for each gene , the expression value of the gene in different regions (represented by ) is calculated according to the color intensity in the heat map. Assuming that the expression value of each region is , it can be obtained by the following way: ; wherein, the function represents the mapping relationship between the color intensity and the expression value . Identify the expression difference region: for the identified expression difference region (such as the up-regulated or down-regulated region), select the gene region with significant change in gene expression value. The expression value of these regions is large, which is the focus of analysis. Difference analysis: for each gene, use statistical method for difference analysis. Assuming that under two experimental conditions (condition A and condition B), condition A: the expression of the gene at the beginning of the experiment. Condition B: the expression of the gene at the end of the experiment. The expression values of the gene are and , wherein represents the gene number. Difference analysis calculates the change amount of gene expression by comparing the expression level under the experimental condition. Calculate the gene expression difference index: the gene expression difference index is calculated based on the change amount of gene expression value. The gene expression difference index is represented by the following formula: ; wherein: is the expression value of gene under condition A. is the expression value of gene under condition B. and respectively represent the standard deviation of the expression value under condition A and condition B. The numerator represents the expression change of the gene under different experimental conditions, and the larger the value is, the more significant the expression difference of the gene is. The denominator part is used for normalizing the expression value fluctuation under different conditions, to ensure the comparability of the difference index. The difference index reflects the expression difference of the gene under different conditions, and the larger the value is, the more significant the expression difference of the gene under different conditions is.
[0035] Specifically, the specific process of extracting the gene variation characteristics in the genomic variation atlas, identifying the variation types and the distribution mode in the genome is as follows: feature extraction is performed on each variation region in the genomic variation atlas, the types of variations including mutations, insertions, deletions and copy number variations are analyzed, and the spatial distribution information thereof is extracted; by calculating the distribution of variations in different gene regions and chromosome regions, the hot spot regions of variations are obtained.
[0036] In this embodiment, the variant region feature extraction: extract each variant region from the genomic variant map through parallel computing and atlas analysis. These variant regions refer to specific regions that have undergone changes in the target genomic sequence. Each variant region has its specific spatial location and variant characteristics, including mutation position, type, and gene or sequence information around the position. When extracting features, the system quantifies the size, type, frequency, and associated genes of these variant regions for further analysis. Variants are usually classified into several main types, and each type of variant may have different effects on gene function and phenotype. Common types of variants include: mutations: usually refer to changes in a single nucleotide, such as point mutations. Mutations can be substitutions, insertions, or deletions. Insertions: refer to the insertion of new nucleotide sequences at a certain position in the genome, which may be caused by the introduction of exogenous DNA fragments or the repetition of internal sequences. Deletions: refer to the loss of DNA fragments at a certain position in the genome, which usually affects the function of the gene. Copy number variation: refers to changes in the number of gene copies in certain regions of the genome, which may cause overexpression or deletion of genes. The system identifies these different types of variants by comparing the variant map of the genome with the reference genome using variant calling algorithms. After identifying the type of variant, the system further analyzes the distribution pattern of the variant on the genome. This process includes quantitative analysis of the spatial distribution of variant regions in different gene regions and chromosome regions. Spatial distribution analysis of variants helps identify whether there is a trend of concentrated distribution of variants in specific regions (exons, regulatory regions), or whether they are evenly distributed throughout the genome. The system calculates the distribution of each variant type in different gene regions (exons, introns, non-coding regions) and chromosomes. This analysis can help researchers identify high-risk variant regions and possible functional variant regions in the genome. By calculating the distribution density of each variant type, the system can identify variant "hotspot" regions. Variant hotspots are regions with high frequency of variants, which may be related to certain genetic diseases, cancer, and other biological phenomena. Variant hotspot regions may be concentrated in certain specific genes, chromosome regions, or gene regulatory elements.
[0037] Specifically, the specific process of obtaining the genomic variant distribution map and the variant type frequency map is as follows: integrate the identified gene variant type and position data to construct a genomic variant distribution map, which shows the distribution of different variant types in the genome; statistically analyze the frequency of each variant type to obtain a variant type frequency map, which shows the frequency of each variant type in the genome through charts, reflecting the frequency and distribution characteristics of different variant types in the genome.
[0038] In this embodiment, the process of obtaining the genomic variant distribution map and the variant type frequency map mainly consists of two steps: construction of the genomic variant distribution map and statistics and display of the variant type frequency map. These steps help analyze the spatial distribution of various types of variants in the genome and the frequency of variants. Construction of the genomic variant distribution map: integrate the identified genomic variant types and their corresponding positions (such as chromosome positions) in the previous steps. Each variant type (such as mutation, insertion, deletion, copy number variation, etc.) has a clear position coordinate (for example, position on the chromosome) in the genome, and may affect a specific gene or gene region. Variant position calibration: in this process, the system calibrates each identified variant (such as mutation position, insertion or deletion position) to the corresponding position of the genome. For example, mutations may occur in exons or regulatory regions of a specific gene, while copy number variations may affect the overall copy number of certain genes. Construction of the variant distribution map: by mapping the position information of all variants to specific regions of the genome (such as chromosome intervals, specific genes or gene regions), a genomic variant distribution map is generated. This map shows the spatial distribution of different types of variants in the genome, including the distribution of variants in specific chromosomal regions, gene regions (such as exons, introns). The genomic variant distribution map helps researchers find clustering regions or potential "hotspot" regions of variants. Statistics of variant type frequency: after identifying and calibrating the position of the variant, the system will calculate the frequency of each type of variant. This process includes calculating the frequency of each type of variant in all identified variants. For example, in all variants, mutations may account for 40%, insertions account for 30%, deletions account for 20%, and copy number variations account for 10%. The statistical results reflect the frequency of different types of variants in the genome. Construction of the variant type frequency map: based on the statistical frequency of variants, the system constructs a variant type frequency map. The frequency map is usually displayed in the form of a column chart, pie chart, etc., showing the relative frequency of different types of variants in the genome. This chart can clearly reflect which types of variants are more common in the studied genome and which variants are rare. Column chart: each column represents a type of variant, and the height of the column represents the frequency of the variant type. Pie chart: each sector represents a type of variant, and the size of the sector represents the proportion of the variant type in all variants. Distribution feature analysis: the combination of the genomic variant distribution map and the variant type frequency map helps researchers understand the distribution characteristics of genomic variants. Through the distribution map, it can be found out whether the variants are clustered in certain specific regions (certain gene regions), and whether a certain type of variant (mutation) is more common in the whole genome. Functional and pathological correlation: high-frequency variants may be related to specific gene functions or certain disease phenotypes. The variant distribution map can help researchers identify these potential correlation regions, thereby providing clues for disease research.
[0039] Specifically, the specific process of identifying the aggregation of genes in the set region of the genome by analyzing the distribution density of the gene region of the gene structure map is as follows: the spatial density of each gene region in the gene structure map is evaluated by parallel computing, the set analysis region in the genome is demarcated, the gene distribution in the region is quantified, and the distribution density of the genes in each region is calculated by the K-means clustering algorithm; the aggregation degree of the genes is evaluated by calculating the aggregation index of the gene distribution in the set analysis region, and the high-density gene aggregation region is identified; the distribution rule of the gene region is further identified by spatial autocorrelation analysis, and the aggregation trend of different genes in the set region is analyzed in combination with the gene function annotation information.
[0040] In the present embodiment, the gene region definition and analysis region setting: the system will perform spatial division and positioning on each gene region in the gene structure map. These gene regions can include exons, introns, regulatory regions, etc. When setting the analysis region, a fragment of a specific chromosome can be selected, or a region where a group of genes are located can be selected according to functional correlation. The selection of the analysis region can be determined according to the research target (for example, the aggregation analysis of a specific gene family). Spatial density quantification: once the analysis region is demarcated, the system will quantify the gene distribution in the region. This quantification process mainly calculates the density of genes in the set region, that is, the number of genes in a unit space. For example, assuming that the set region is a certain segment of a certain chromosome, the system will count the number of genes in the region according to the gene position coordinates, and then calculate the distribution density of the genes. K-means clustering: based on the calibrated gene position data, the K-means clustering algorithm is applied to analyze the distribution pattern of genes in the set region. In this step, the algorithm adjusts the center of the cluster through continuous iteration, and finally makes the aggregation degree of the genes in each cluster reach the maximum, so that the regions with high gene density can be clearly identified. Cluster result interpretation: through the result of K-means clustering, the system can identify the aggregation degree of genes in a specific region, find out which regions have dense gene number and which regions have sparse gene distribution. These results are helpful for further analyzing the spatial distribution characteristics of genes and discovering potential gene aggregation regions. Aggregation index: the aggregation index is an index for quantifying the aggregation degree of genes in the set region. By calculating the gene distribution density of each region, combining the number of genes in the region and the area of the region, the system generates the aggregation index. The region with high aggregation index indicates that the gene distribution in the region is more concentrated, and vice versa, which indicates that the gene distribution in the region is more dispersed. High-density region identification: the aggregation index is used to quantify the aggregation degree of genes in the region . The specific formula is as follows: Formula parameter explanation: : region Total number of genes in the region, representing the number of genes involved in the analysis within this region, and : Gene and Gene Spatial coordinates of Gene and represent the location of Gene and Gene in the genome (can be two-dimensional or three-dimensional coordinates, depending on the dimensionality of the data). : Euclidean distance between Gene and Gene By calculating the spatial distance between gene pairs, their relative positions within the region can be assessed. The Euclidean distance formula: ; this formula calculates the straight-line distance between Gene and Gene . : Normalization factor used to calculate the total number of gene pairs within the region. This factor represents the total number of gene pairs, ensuring that the clustering index is based on the average number of gene pairs. Based on the clustering index, the system can further identify regions with higher degrees of gene clustering. These high-density regions may correspond to regions with important functional roles in genes, or are related to certain genetic characteristics, disease susceptibility, etc.
[0041] Specifically, the specific process of obtaining the gene density index is as follows: according to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; the gene aggregation situation in different areas is quantified by statistical methods to obtain the gene density index of each area.
[0042] In this embodiment, the calculation of the gene density index is based on the spatial distribution density of genes in the set gene region (analysis region). By quantifying the aggregation of genes in different regions, we can obtain the gene density index of each region. This density index not only reflects the distribution density of genes in space, but also can reveal the distribution characteristics of genes in different regions and their biological significance. The specific calculation process includes calculating the gene density value, statistical gene clustering, and combining relevant formulas to obtain the final gene density index. The gene density value represents the distribution density of genes in a certain set region, which can usually be calculated by the following formula: ;in: No. The gene density value of the region. No. The number of genes in a region. :No. The area or length of a region (depending on the representation of the gene structure diagram, it may be the length of a certain segment of the chromosome, or the area of a specific functional region). The larger the gene density value, the more concentrated the gene distribution in the region and the more obvious the aggregation. The gene density index is a quantitative assessment of the gene aggregation in different regions. It is usually calculated by statistical methods. It combines the gene density value of each region and the spatial distribution characteristics of the genes in the region. The calculation formula is as follows: in: :No. The gene density index of a region. No. The number of genes in a region. :No. The area or length of a region. : Weight factor, used to indicate the importance or relevance of each gene in the region. No. Region and The spatial similarity between genes indicates the degree of spatial clustering of genes. The larger the value, the more genes tend to cluster in this region.
[0043] Specifically, the specific process of predicting gene expression trends, mutation hotspot evolution and functional region distribution according to the machine learning model is as follows: based on the gene expression difference index extracted by the heat map analysis unit, the dynamic change trend of gene expression under the set condition is predicted through the time series model; based on the mutation hotspot region identified by the mutation map analysis unit, combined with the mutation type frequency diagram, the potential new mutation hotspot region is predicted through the convolutional neural network; based on the gene density index of the structure map analysis unit, the enrichment region of functionally related genes in the genome is predicted through the clustering algorithm.
[0044] In the present embodiment, the gene expression trend prediction: time series modeling, task requirement: based on the gene expression difference index (dynamic change data), predict the expression trend of genes under the set time or condition Core algorithm adaptation, BayesLasso algorithm model function: feature selection Implementation: Use Laplace prior (double exponential distribution) to compress insignificant SNP effects, and select key regulatory sites related to expression difference index. SNP refers to the variation of a single base (A, T, C, G) in the genomic sequence. For example: select 500 high-impact sites from 500,000 SNPs as input features for the time series model. GBLUP algorithm model function: dynamic effect modeling Implementation: Combine the genomic relationship matrix (VanRaden method) with the time covariate to construct an extended mixed model: ; Parameter interpretation: : Phenotype vector; Fixed effect design matrix (such as experimental conditions, environmental factors); Fixed effect coefficient; : Genotype matrix; : Individual genomic breeding value (random effect); : Genomic relationship matrix (VanRaden method calculation); : Residual term; : Genome variance. Capture genetic effects of gene expression differences over time, predict future trends. Capture genetic effects of gene expression differences over time, predict future trends. RRBLUP algorithm model role: Robust baseline prediction; Implementation: As a supplement to GBLUP, penalize marker effects through ridge regression, provide conservative estimates of gene expression trends (prevent overfitting). Variant hotspot prediction: Spatial pattern recognition. Task requirement: Based on known variant hotspots and type frequencies, predict potential new hotspot regions. Core algorithm adaptation: BayesB / BayesC algorithm model role: Probability feature generation, implementation: Calculate the probability of each SNP becoming a hotspot-related variant through the π parameter (BayesC inference, BayesB preset). Generate a probability matrix as a CNN input, for example: The SNP hotspot probability of chromosome region A is 0.8, and region B is 0.05. GBLUP algorithm role: Spatial feature enhancement. Implementation: Input the genomic relationship matrix as an additional channel of CNN to help the model identify linkage disequilibrium patterns between variant sites. BayesLasso algorithm model role: Auxiliary noise filtering. Implementation: After CNN prediction, perform secondary verification on high-probability hotspot regions to remove false positive signals (such as sparse regression verification of key SNPs). Functional enrichment region prediction: Density clustering. Task requirement: Based on gene density index, predict the clustering region of functionally related genes. Core algorithm adaptation, GBLUP algorithm role: Similarity matrix construction. Implementation: Use the genomic relationship matrix to calculate the genetic similarity between regions as the weight input of spectral clustering. For example: The genomic similarity between region A and B is 0.7, so it is more likely to be classified into the same class during clustering. BayesLasso algorithm model role: Signature SNP screening. Implementation: Screen SNPs significantly associated with high-density regions as clustering feature markers. For example: A SNP frequently appears in 10 high-density regions, so it is given a higher clustering weight. RRBLUP algorithm model role: Auxiliary density. Correction implementation: Input the gene density index as a phenotype into RRBLUP to estimate marker effects, used to adjust the bias in density calculation. By integrating the results of the above three prediction processes, support vector machines are used for integration to generate a comprehensive genomic prediction map. This map shows the overall picture of gene expression trends, variant hotspot evolution, and functional region distribution, helping researchers comprehensively understand the structure, function, and variation characteristics of the genome. The generated map not only helps in-depth analysis of the genome, but also provides intuitive visual results.
[0045] Specifically, the specific process of comprehensive visual display of the gene expression difference distribution map, the gene expression difference index, the genome variation distribution map, the variation type frequency map, the gene density index and the genome prediction map is as follows: the gene expression difference distribution map, the genome variation distribution map and the gene structure map are normalized by parallel computing, the genome positions of different data sources are mapped to a unified reference coordinate system; based on the statistical results of the gene density index and the variation type frequency map, a gene region weight matrix is constructed, which is used to dynamically adjust the superposition order of the visualization layers; the gene expression trend prediction value in the genome prediction map is converted into a dynamic heat map, the prediction heat map at different time points is adjusted and displayed through a time axis control, and is compared with the gene expression heat map through transparency superposition; in the genome variation distribution map, the predicted new variation hotspot area is marked with a pulse flicker on the known variation area, and a ring legend is generated through the variation type frequency map to display the type proportion difference between the prediction hotspot and the historical data in real time; the functional enrichment area prediction result in the gene structure map is converted into a 3D contour surface, which is superimposed on the two-dimensional gene density distribution map, the height of the surface is positively correlated with the gene density index, and the color gradient represents the prediction confidence.
[0046] In this embodiment, the core goal of this visualization method is to integrate multiple genomic data, enabling researchers to intuitively understand the distribution and predicted trends of gene expression, genetic variation, and functionally enriched regions. The specific process is as follows: Coordinate normalization: Since genomic data comes from different data sources (such as gene expression data, genetic variation data, and gene structure data), the genomic positions of these data may be based on different reference coordinates. Coordinate normalization is performed on all data using parallel computing to ensure that information from different data sources can be mapped to a unified genomic reference coordinate system, facilitating subsequent integration and comparison. Construction of gene region weight matrix: Gene density index: represents the distribution density of genes in a certain region. Variation type frequency graph: statistics the frequency of different types of variations (such as SNPs, insertions / deletions) in a specific gene region. Through these statistics, a gene region weight matrix is constructed to dynamically adjust the superposition order of different visualization layers, making the information of important regions more prominent. Dynamic heat map of predicted gene expression trend: gene expression trend prediction value (from time series model) is used to generate a dynamic heat map. Through the time axis control, users can view the gene expression prediction at different time points. The prediction heat map is superimposed with the gene expression heat map in transparency, allowing users to intuitively compare the differences between prediction values and experimental data, and evaluate the accuracy of prediction. Prediction and visualization of variation hotspot regions: In the genomic variation distribution map, combine the predicted new variation hotspot regions by convolutional neural network, and highlight them with pulse flicker markers to enhance visual guidance. Variation type frequency graph generates a ring legend to display the proportion difference between predicted hotspots and historical data in terms of variation type, allowing researchers to quickly determine whether new variations conform to known variation patterns. 3D contour surface display of functionally enriched regions: based on gene density index, predict functionally enriched regions and convert them into 3D contour surfaces for visualization. The surface height is proportional to the gene density index, indicating the enrichment degree of functionally related genes in the genome. The color gradient represents the prediction confidence, with darker colors in high confidence areas for easy identification of important regions. Through coordinate normalization, different genomic data is integrated, and the gene region weight matrix ensures optimal information hierarchy. Dynamic heat maps, flicker markers, ring legends, and 3D contour surfaces make the visualization of gene expression, variation hotspots, and functionally enriched regions more intuitive, providing an efficient analysis tool for genomic research.
[0047] In summary, the present application has at least the following effects:
[0048] The parallel computing and data visualization analysis system for whole genome prediction effectively processes and analyzes target whole genome data, including gene expression, gene variation and gene structure data, through parallel computing and data visualization algorithms, providing an efficient data processing and computing platform for genomics research. The system can accurately analyze the spatial distribution and color intensity of the genome heat map, identify the expression difference regions of genes under different experimental conditions, and reflect the expression level changes of genes under the set conditions through the gene expression difference index, thereby providing strong support for gene function research. By analyzing the genome variation map, gene variation characteristics can be extracted and variation types such as mutations, insertions and deletions can be identified, and at the same time, the distribution of different variation types in the genome and the variation hotspot regions can be displayed through the variation type frequency diagram, helping researchers understand the spatial distribution pattern of gene variation. The structure map analysis module can reveal the distribution characteristics and aggregation trend of genes in different regions through spatial density evaluation and aggregation index calculation, providing scientific basis for studying the regional characteristics of the genome and gene function annotation. Through the comprehensive visualization module, the gene expression difference distribution map, the gene expression difference index, the genome variation distribution map, the variation type frequency diagram and the gene density index are displayed uniformly, and the spatial distribution, color depth and morphological characteristics are used to enable researchers to intuitively observe the various characteristics of the genome and their mutual relationships.
[0049] Those skilled in the art will appreciate that embodiments of the application can be provided as methods, systems, or computer program products. Accordingly, the application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the application can be embodied in the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage media, etc.) having computer usable program code embodied therein.
[0050] The application is described with reference to flowcharts and / or block diagrams of the systems, devices (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce an apparatus that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0051] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0052] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0053] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those of skill in the art once they have the benefit of the present disclosure. Therefore, the appended claims are intended to encompass within their scope all possible variations and modifications of the preferred embodiments. 1
[0054] It is apparent that a person skilled in the art can make various changes and modifications to the application without departing from the spirit and scope thereof. Therefore, if these modifications and changes fall within the scope of the claims and their equivalents, it is intended to include them in the application.
Claims
1. A parallel computing and data visualization analysis system for whole genome prediction, characterized by: Includes the following modules: image generation module, image analysis module, genome prediction module, and visualization module; The image generation module is used to obtain target whole genome data and convert the target whole genome data into whole genome sequencing images, including genome heat maps, genome variation maps, and gene structure maps, through data visualization algorithms and parallel computing; The image analysis module includes a heat map analysis unit, a variation map analysis unit, and a structure map analysis unit; The heat map analysis unit is used to perform spatial distribution and color intensity analysis on the genome heat map, identify differential expression regions of genes under set conditions, extract the gene expression differential distribution map in the genome heat map, and obtain the gene expression differential index, which is used to reflect the changes in gene expression levels under set conditions; The variation map analysis unit is used to extract gene variation features from the genome variation map, identify variation types and distribution patterns in the genome, and obtain genome variation distribution maps and variation type frequency maps to reflect the variation hotspots of various variation types in the genome; The structure graph analysis unit is used to perform distribution density analysis of gene regions on the gene structure graph, identify the aggregation of genes in a set region in the genome, and obtain a gene density index to reflect the distribution characteristics of genes in different genome regions; The genome prediction module is used to receive the output data of the image analysis module, predict gene expression trends, mutation hotspot evolution and functional region distribution based on the machine learning model, and generate a genome prediction map; The visualization module is used to comprehensively visualize the gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map; The specific process of performing gene region distribution density analysis on the gene structure map and identifying the clustering of genes in a set region in the genome is as follows: Through parallel computing, the spatial density of each gene region in the gene structure map is evaluated, the analysis region set in the genome is delineated, and the gene distribution in the region is quantified. The distribution density of genes in each region is calculated using the K-means clustering algorithm; By calculating the clustering index of gene distribution in the set analysis area, the degree of gene clustering is evaluated and high-density gene clustering areas are identified; Through spatial autocorrelation analysis, we can further identify the distribution patterns of gene regions and analyze the clustering trends of different genes in the set region by combining gene function annotation information; The specific process of predicting gene expression trends, mutation hotspot evolution, and functional region distribution based on machine learning models is as follows: Based on the gene expression difference index extracted from the heat map analysis unit, the dynamic change trend of gene expression under set conditions is predicted through the time series model; Based on the mutation hotspots identified by the mutation map analysis unit and combined with the mutation type frequency map, a convolutional neural network is used to predict potential new mutation hotspots. According to the gene density index of the structural map analysis unit, the enrichment region of functional related genes in the genome is predicted by clustering algorithm; The specific process of comprehensive visualization of gene expression difference distribution map, gene expression difference index, genome variation distribution map, variation type frequency map, gene density index and genome prediction map is as follows: Through parallel computing, coordinate normalization is performed on the gene expression difference distribution map, genome variation distribution map, and gene structure map, and the genomic positions of different data sources are mapped to a unified reference coordinate system; Based on the statistical results of gene density index and variation type frequency map, a gene region weight matrix is constructed to dynamically adjust the overlay order of visualization layers; Convert the gene expression trend prediction values in the genome prediction map into a dynamic heat map. Use the timeline control to adjust the display of the prediction heat map at different time points, and perform transparency overlay comparison with the gene expression heat map. In the genomic variation distribution map, the predicted new variation hotspots are overlaid on the known variation regions with pulse flashing markers, and a ring legend is generated through the variation type frequency map to display the difference in type ratio between the predicted hotspots and historical data in real time; The prediction results of functional enrichment regions in the gene structure map are converted into a 3D contour surface and superimposed on the two-dimensional gene density distribution map. The height of the surface is positively correlated with the gene density index, and the color gradient represents the prediction confidence.
2. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of converting target whole genome data into whole genome sequencing images through data visualization algorithms and parallel computing is as follows: Load the target whole genome data through parallel computing, clean, filter and standardize the target whole genome data, and handle outliers and missing data; The target whole genome data includes gene expression data, gene variation data and gene structure data; Generate gene expression heatmaps based on gene expression data using heatmap visualization algorithms; Based on the gene variation data, a genome variation map is constructed using a variation map visualization algorithm; Based on the gene structure data, a gene structure diagram is drawn using a gene structure visualization algorithm to mark the various components of the gene and their spatial locations; The generated genome heatmaps, genome variation maps, and gene structure maps were integrated into a comprehensive whole-genome sequencing image.
3. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of performing spatial distribution and color intensity analysis on the genomic heat map, identifying differentially expressed regions of genes under set conditions, and extracting the differential gene expression distribution map from the genomic heat map is as follows: By analyzing the spatial position of each gene region in the genome heat map through parallel computing, the expression density and distribution pattern of genes in different regions are evaluated, and the expression differences of genes in different regions are identified; Quantify the color intensity of each pixel in the heat map to determine the relationship between different color intensities and gene expression levels, and identify differences in expression intensity through color changes in the heat map; Based on the changes in color intensity, the expression difference threshold is set to identify gene regions that show upregulation or downregulation under the set conditions; Gene expression difference distribution maps are extracted from the identified differentially expressed regions to display the expression changes of genes in different regions and mark the changed gene regions.
4. The parallel computing and data visualization analysis system for whole genome prediction according to claim 3, characterized in that: The specific process of obtaining the gene expression difference index is as follows: By quantifying the color intensity of each gene region in the genome heat map, the expression value of each gene in different regions is extracted. The expression value is the mapping relationship between the color intensity of the corresponding region in the heat map and the gene expression level. Statistical analysis was used to perform differential analysis on each gene in the identified differentially expressed regions. By comparing the expression levels of the genes under the set conditions, the expression change of each gene under the preset experimental conditions was calculated to obtain the gene expression difference index.
5. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of extracting gene variation features from the genome variation map and identifying variation types and distribution patterns in the genome is as follows: Extract features from each variant region in the genomic variation map, analyze the types of variations, including mutations, insertions, deletions, and copy number variations, and extract their spatial distribution information; By calculating the distribution of mutations in different gene regions and chromosome regions, the mutation hotspot areas are obtained.
6. The parallel computing and data visualization analysis system for whole genome prediction according to claim 5, characterized in that: The specific process of obtaining the genomic variation distribution map and variation type frequency map is as follows: Integrate the identified genetic variation types and location data to construct a genomic variation distribution map to show the distribution of different variation types in the genome; The frequency of each variant type is statistically analyzed to obtain a variant type frequency graph, which shows the frequency of each variant type in the genome, reflecting the frequency of different variant types and their distribution characteristics in the genome.
7. The parallel computing and data visualization analysis system for whole genome prediction according to claim 1, characterized in that: The specific process of obtaining the gene density index is as follows: According to the spatial distribution density of genes in the set analysis area, the gene density value of each area is calculated; The gene clustering in different regions was quantified by statistical methods to obtain the gene density index of each region.
Citation Information
Patent Citations
Systems and methods for genomic analysis
CN108350494A
Genome abnormality visualization system and genomic abnormality visualization method
JP2024036206A
Systems and Methods for Producing Quantitatively Calibrated Grayscale Values in Magnetic Resonance Images
US20180325461A1
Identifying genome features in health and disease
US20230307092A1
Peptide centric analyses
WO2023133536A2