A multi-environment chromosome variation feature spectrum mining method
By integrating chromosome sequence and three-dimensional conformation data, and combining environmental comparative machine learning and functional network perturbation models, the problem of incomplete genetic damage assessment in existing technologies has been solved, enabling accurate and dynamic risk assessment and early warning of chromosome variations under special environments.
Patent Information
- Application Number
- CN202511728327.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-11-24
AI Technical Summary
Existing technologies, when assessing chromosomal variations under special circumstances, neglect the three-dimensional spatial conformation and its higher-level functional structure, resulting in incomplete assessment of genetic damage, lack of dynamic and forward-looking risk prediction capabilities, and limitations on health monitoring and risk intervention for occupational populations in special environments.
By acquiring chromosome sequence variation and three-dimensional spatial conformation data through integrative genomics technology, a multimodal association database is constructed. An environment-comparative machine learning model is used to decouple the joint feature spectrum of environment-specific chromosome variation-conformation. The intensity of its perturbation on biological functional networks is quantified through a functional network perturbation model. Prospective extrapolation functions are integrated to simulate risk evolution paths.
It enables multi-dimensional genetic damage analysis from linear sequences to three-dimensional spatial structures, improving the accuracy and depth of genetic risk assessment and providing support for early and accurate health warnings and prospective intervention strategies.
Smart Images

Figure CN121545574B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental genomics, and in particular to a method for mining multi-environment chromosomal variation feature profiles. Background Technology
[0002] With the continuous expansion of my country's activities in aviation, polar regions, plateaus, and specialized industries, certain occupational groups are exposed to complex stress environments such as high-altitude radiation, extreme temperatures, and hypoxia. These environmental factors are widely considered important exogenous mutagens that induce damage to the genetic material of organisms. Long-term, low-dose exposure may lead to decreased chromosome stability, potentially affecting physiological functions and increasing long-term health risks. Currently, the fields of environmental medicine and genetics have begun to recognize the importance of genotoxicity assessment for these special environmental populations and have started using technologies such as chromosome karyotype analysis, fluorescence in situ hybridization, and even high-throughput sequencing to try to reveal chromosomal variation patterns under different environmental stresses. Constructing a systematic database and analysis system that links environmental parameters with genetic variation has become a cutting-edge research direction for accurately assessing occupational health risks and potentially guiding personalized protection strategies.
[0003] While existing research has been able to identify some chromosomal abnormalities under specific environmental exposures, its methodologies have significant limitations. Current methods mostly focus on one-dimensional linear sequence variations of chromosomes, generally neglecting the potential impact of environmental stress on the three-dimensional spatial conformation and higher-order functional structures of chromosomes, leading to incomplete assessments of genetic damage. At the data analysis level, existing techniques largely rely on simple inter-group frequency statistics and single variant annotation, lacking integrated analytical models that can effectively decouple complex environmental signals and quantify functional perturbations from a systems biology perspective. Furthermore, most studies are static cross-sectional surveys, unable to simulate the dynamic process of variation accumulation over exposure time, and lack prospective predictive capabilities, resulting in severely insufficient depth and early warning value in risk assessment. These shortcomings collectively restrict the ability to conduct efficient and accurate health monitoring and risk intervention for occupational populations in specific environments. Therefore, based on the above challenges, this invention proposes a multi-environment chromosomal variation feature spectrum mining method. Summary of the Invention
[0004] To address the aforementioned issues, the present invention aims to provide a multi-environment chromosomal variation feature spectrum mining method. This method systematically decouples the specific chromosomal variation-conformation joint feature spectrum induced by different special environmental exposures. Based on this, it quantitatively assesses the perturbation intensity of the joint feature spectrum on the core biological functional network from the perspective of systems biology, ultimately achieving accurate, dynamic, and forward-looking early warning of chromosomal stability damage and related health risks in occupational populations in special environments.
[0005] To achieve the above objectives, this invention provides a method for mining multi-environment chromosomal variation feature profiles. First, through integrative genomics technology, chromosomal sequence variations and three-dimensional spatial conformation data from populations in specific environments are simultaneously acquired and associated with environmental exposure parameters to construct a multimodal association database. Next, an environmentally comparative machine learning model is used to decouple a joint feature profile of chromosomal variations and conformations from the aforementioned multimodal data, ensuring strong environment specificity and robustness. Finally, by constructing a functional network perturbation model, the joint feature profile is mapped to a molecular interaction network, and its topological impact on key nodes and pathways is quantitatively calculated. This outputs a tiered functional perturbation assessment and health risk warning information, and further integrates forward-looking extrapolation functions to simulate risk evolution paths.
[0006] In a first aspect, the present invention provides a method for mining multi-environment chromosomal variation feature spectra, including: Peripheral blood samples were obtained from different groups of people exposed to special environments, and their long-term, multi-dimensional environmental exposure physicochemical parameters were collected simultaneously. An integrated genomic analysis was performed on the samples to obtain a dataset of covariates, including primary sequence variations and three-dimensional spatial conformation variations of chromosomes. Construct a relational database to perform spatiotemporal correlation mapping of the covariation dataset, environmental exposure physicochemical parameters, and individual phenotypic data; Based on the aforementioned relational database, a unique chromosome variation-conformation joint feature spectrum under different environmental pressures is decoupled through an environmental comparative machine learning model. Using the joint feature spectrum, the perturbation intensity of a specific environmental exposure on key nodes and pathways of the core biological functional network is quantitatively assessed through a functional network perturbation model, and hierarchical early warning information is generated accordingly.
[0007] Furthermore, the acquisition of the three-dimensional spatial conformational variations of chromosomes is achieved by analyzing the structural changes of the topological association domain (TAD) of chromosomes in the cell nucleus and the dynamic reconstruction of the internal chromatin loops through high-throughput chromosome conformational capture technology. The conformational variations include, but are not limited to, the weakening or strengthening of the TAD boundary strength and the formation of abnormal chromatin loops. This method can reveal structural imbalances in the epigenetic regulatory hierarchy caused by environmental factors that cannot be detected by traditional sequence analysis.
[0008] Furthermore, the environmental comparative machine learning model employs an adversarial generative network framework, which uses its discriminator component to learn and reinforce the differential features of chromosomal variation patterns among groups exposed to different environments, thereby decoupling a robust feature spectrum that is highly specific to the environment and has weak individual differences.
[0009] Furthermore, the construction of the functional network perturbation model is specifically as follows: The variants and conformational changes identified in the joint feature spectrum are mapped onto a pre-constructed functional network that covers multi-level biomolecular interactions. The overall functional disturbance intensity is quantified by calculating the topological impact scores and flow disturbance levels of the aforementioned mutations and conformational changes on critical nodes and vulnerable paths in the network.
[0010] Furthermore, the method is based on the established joint feature spectrum and functional network perturbation model. By inputting simulated changes in environmental parameters, it calculates and infers the potential evolutionary paths of chromosome variations and conformations and their corresponding functional perturbation inflection points, providing a theoretical basis for formulating critical intervention thresholds. This step extends the system's capabilities from current status assessment to future risk simulation, realizing early warning.
[0011] Furthermore, the integrative genomics analysis also includes assessing the stability of fragile chromosomal sites and analyzing whether they become hotspots for structural variations and conformational abnormalities under special environmental exposures. This approach can prioritize identifying known vulnerable regions in the genome and improve the targeting and efficiency of feature profile mining.
[0012] Secondly, the present invention also provides a multi-environment chromosome variation feature spectrum mining system, the system being based on the method described in the first aspect above, comprising: The collaborative data acquisition module is used to control multiple omics detection devices and connect to an environmental sensor network to obtain the collaborative variation dataset and environmental exposure physicochemical parameters; The data fusion and computing engine integrates the aforementioned environmental comparative machine learning model and the aforementioned functional network perturbation model at its core. The dynamic knowledge graph module is used to store the relational database and serve as the underlying network of the functional network perturbation model. The intelligent early warning and simulation terminal is used to output tiered early warning information and visualize the results of forward-looking intervention simulations.
[0013] Furthermore, the dynamic knowledge graph module has self-evolution capabilities. It can automatically optimize the weight allocation of nodes and paths in the functional network based on newly added sample data and their corresponding early warning feedback results, and discover new potential relationships, so that the knowledge system and judgment accuracy of the entire system can continuously evolve with the accumulation of data.
[0014] Furthermore, the system also includes a human-machine collaborative decision-making loop, which allows domain experts to review and semantically annotate the joint feature spectrum and early warning information generated by the system, and feed the annotated information as prior knowledge into the environmental comparative machine learning model to achieve continuous targeted optimization of the model. This mechanism deeply integrates the domain knowledge of human experts into the artificial intelligence analysis process, effectively improving the interpretability and professional credibility of the model.
[0015] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above, wherein the program specifically includes an operating instruction set for controlling a high-throughput chromosome conformation capture sequencer, and a parallel computing instruction set for running the functional network perturbation model.
[0016] This invention provides a method for mining multi-environment chromosomal variation feature profiles. This method uses integrative genomics technology to simultaneously acquire primary sequence variations and three-dimensional spatial conformational variations of chromosomes in populations exposed to specific environments, and constructs a database by associating them with environmental parameters. Then, an environmentally comparative machine learning model is used to decouple environment-specific chromosomal variation-conformation joint feature profiles from the above data. Finally, through a functional network perturbation model, the feature profiles are mapped to a biomolecular interaction network to quantitatively assess the perturbation intensity on core functional pathways. Furthermore, a prospective extrapolation and dynamic self-optimization mechanism are integrated to achieve a closed-loop analysis from genetic damage identification to health risk warning.
[0017] This method enables multi-dimensional collaborative analysis of genetic damage from linear sequences to spatial structures, systematically revealing the combined effects of environmental exposures on chromosome stability and higher functions. It significantly improves the accuracy and depth of discovering biomarkers for genetic risks related to specific environments, and advances risk assessment from single variant annotations to the quantification of perturbations in the overall biological functional network, thus providing a scientific basis and decision support for the development of early, precise health warnings and prospective intervention strategies.
[0018] Beneficial effects By implementing the multi-environment chromosome variation feature spectrum mining method provided by the present invention, the following technical effects are achieved: (1) The dimensions of genetic damage assessment are expanded from the traditional one-dimensional linear sequence to the three-dimensional intranuclear spatial structure. This can reveal the profound impact of environmental factors on the epigenetic regulatory hierarchy and the regulation of remote gene expression, thereby achieving a more comprehensive and essential analysis of genotoxicity and overcoming the significant biological information omissions that may result from relying solely on sequence variations.
[0019] (2) Decoupling environment-specific features through an adversarial machine learning framework. It can automatically and efficiently remove the interference of confounding factors such as individual genetic background, and extract a set of environment-related biomarkers with high purity and strong robustness, which greatly improves the indicative specificity and discovery accuracy of the feature spectrum for specific environmental exposures.
[0020] (3) Topological impact calculations are performed by mapping discrete genetic variations onto the biomolecular interaction network of the system. This achieves a leap from the simple association of "single variation - single gene" to "variation set - functional network - system perturbation", which can quantitatively assess the cumulative biological consequences of multi-point, minor variations, making risk assessment more systematic and predictive.
[0021] (4) It introduces a time dimension and simulation capabilities into the static risk assessment model. It elevates the system's function from current situation analysis to future scenario simulation, enabling it to predict the evolution path of risk accumulation over exposure time and identify critical inflection points. This provides key decision support for formulating intervention strategies that precede the occurrence of damage, thus achieving early warning. Attached Figure Description
[0022] To make the above-described method for mining multi-environment chromosome variation feature spectrum of the present invention more obvious and understandable, the accompanying drawings used in the specific embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the method described in this application; Figure 2 This is a flowchart illustrating the early warning and dynamic optimization process of an intelligent system. Detailed Implementation
[0024] Example 1: This embodiment aims to illustrate the complete process from sample acquisition to feature spectrum analysis, such as... Figure 1As shown. First, sample and data acquisition is the foundation of the research. This embodiment selects four distinct groups of people as research subjects: military pilots who perform long-term high-altitude flight missions (exposed to cosmic radiation, hypoxia, and circadian rhythm disruption), border patrol personnel stationed on the Qinghai-Tibet Plateau (exposed to high-altitude hypoxia, strong ultraviolet radiation, and cold environments), geological exploration team members conducting fieldwork in the Turpan region (exposed to extreme high temperatures and strong thermal radiation), and a control group of staff members of a large city government (with no special environmental exposure history). For all participants, not only were peripheral blood samples collected for genomic analysis, but more importantly, long-term and continuous environmental exposure physicochemical parameters were systematically collected through personal environmental sensors and work logs. These parameters included, but were not limited to, environmental gamma radiation dose rate, environmental oxygen partial pressure, environmental temperature and humidity fluctuation curves, and ultraviolet index, thereby ensuring that subsequent analyses had solid environmental exposure traceability.
[0025] Integrative genomic analysis was performed during the laboratory analysis phase. Genomic DNA was extracted from peripheral blood lymphocytes, and two types of assays were conducted in parallel. The first type was whole-genome sequencing, used to comprehensively capture primary sequence variations of chromosomes, including single nucleotide variants, small insertions and deletions, and large structural variations such as copy number variations and chromosomal translocations. The second type employed high-throughput chromosome conformation capture technology to comprehensively resolve the three-dimensional spatial conformation of chromosomes within the lymphocyte nucleus. This technology can reconstruct the spatial topological association domain (TAD) structure of chromatin by analyzing whole-genome chromatin interaction data. Special attention was paid to changes in the intensity of TAD boundaries and the abnormal formation or disappearance of chromatin loops, as these conformational variations are considered key spatial structural bases for regulating gene expression. Finally, a covariance dataset containing millions of data points was generated for each sample, including both linear sequence variation information and three-dimensional spatial structural information.
[0026] Subsequently, a structured, association-based chromosomal mutation database was constructed. This database is not a simple data warehouse, but rather employs a specialized data model capable of accurately spatiotemporally associating and mapping the covariance dataset of each sample with its corresponding multidimensional environmental exposure physicochemical parameters and basic individual phenotypic data. All raw variation data underwent rigorous bioinformatics processes for quality control, alignment, and annotation, and batch effect correction was performed to eliminate technical noise that might be introduced by different testing batches, providing a high-quality, standardized data foundation for subsequent machine learning analysis.
[0027] In the core of data analysis, an environmentally comparative machine learning model based on an adversarial generative network framework was introduced. In this framework, the generator is designed to produce "fake" variant spectra that simulate real data distributions, aiming to confuse the discriminator as much as possible. The discriminator's task is to be forced to learn how to accurately distinguish real sample data from different environmental exposure sources. Through this adversarial training, the discriminator gradually develops a highly refined set of feature filters, capable of keenly capturing deep, robust variant patterns that best represent different environmental stresses, while automatically ignoring variants caused by individual genetic background differences, random factors, or technological noise. After sufficient training, the model was applied to analyze data from four population groups, successfully decoupling four clearly identifiable "chromosomal variation-conformation joint feature spectra." For example, in the pilot group, the model identified a joint spectra characterized by specific chromosomal translocations and a general decrease in the stability of several TAD boundaries; while in the high-temperature environment group, a feature spectrum significantly correlated with abnormal chromatin loop remodeling associated with heat shock protein gene clusters was found. This process effectively isolates and extracts a set of high-confidence environment-specific biomarkers from highly complex biological noise.
[0028] Finally, the extracted feature spectra were biologically interpreted using a functional network perturbation model. All significant variations and conformational changes in the aforementioned joint feature spectra were systematically mapped onto a pre-constructed multi-level biomolecular functional network encompassing protein-protein interactions, signal transduction, and metabolic pathways. By calculating the topological impact scores of these variation events on key hub nodes and vulnerable pathways within the network, the overall perturbation intensity of each feature spectra on core biological functional modules could be quantified. This not only bridges the gap from discrete lists of genetic variations to systemic functional impacts but also provides mechanistic insights into how different specific environments ultimately affect an organism's physiological functions and health status through specific genetic and epigenetic mechanisms.
[0029] Example 2: This embodiment aims to concretize a hardware and software integrated platform capable of supporting and executing the method described in Embodiment 1, and to expand its dynamic and forward-looking capabilities. The system consists of four core modules and two major mechanisms. The collaborative data acquisition module, acting as the system's "sensory organs," communicates directly with the server controlling the high-throughput sequencer and chromosome conformation capture platform through a standardized application programming interface, automatically capturing raw sequencing data. Simultaneously, this module also connects to an environmental sensor network deployed in various specific working areas, continuously receiving and preprocessing environmental exposure physicochemical parameters, achieving automated aggregation and preliminary spatiotemporal alignment of multi-source, heterogeneous data.
[0030] The data fusion and computation engine is the "brain" of the system. This engine incorporates a pre-trained environmental contrastive machine learning model and a functional network perturbation model. When new sample data and associated environmental parameters are input through the acquisition module, the engine automatically initiates the analysis process: first, it performs data standardization and quality assessment; then, it calls the contrastive learning model to decouple the environmental feature spectrum corresponding to the sample; and finally, it uses the functional network perturbation model to calculate its network perturbation strength index. The entire computation process is based on a distributed computing framework to ensure efficient processing of massive amounts of genomic data.
[0031] The dynamic knowledge graph module serves as the system's "memory and knowledge base." It is not a static database, but rather an underlying architecture built using graph database technology to construct a vast network of biomolecular interactions. Nodes in the graph represent biological entities such as genes, proteins, and metabolites, while edges represent the interactions, regulatory relationships, or reactions between them. The initial graph integrates knowledge from multiple authoritative public databases. When a new sample processed by the system is confirmed to have a high-value feature profile or corresponding clinical phenotype, this module can automatically adjust the weights of relevant nodes and edges in the network. It can even autonomously discover and suggest adding new potential interactions based on implicit correlations between data, enabling the system's knowledge system to continuously evolve with data accumulation, becoming increasingly intelligent with use.
[0032] The intelligent early warning and simulation terminal is the system's "decision output interface." The intelligent system's early warning and dynamic optimization process is as follows: Figure 2 As shown, it receives the network perturbation intensity index from the computing engine and automatically generates a graded early warning report based on preset risk level thresholds. The report not only describes the main risk characteristics in text form but also uses information visualization technology to intuitively display the key biological pathways affected by the characteristic spectrum and their positions in the network. Furthermore, the terminal integrates a prospective intervention simulation mechanism. Users can input simulated environmental parameter changes on this interface, and the system will dynamically simulate and visualize the potential evolutionary trajectory of chromosomal variation load and functional network perturbation intensity under the hypothetical scenario based on an established dose-response relationship model, identifying possible dysfunction "inflection points." This provides a quantitative scientific basis for developing advanced, threshold-based health intervention measures.
[0033] Finally, the human-machine collaborative decision-making loop is a key mechanism to ensure the system's professionalism and credibility. This loop allows domain experts to review or correct the rich clinical semantic annotations of the joint feature spectrum and early warning information automatically generated by the system on the intelligent early warning and inference terminal. The results confirmed or corrected by the experts will be fed back as high-quality "calibration data" to the environmental comparative machine learning model for fine-tuning and targeted optimization. This closed-loop design successfully integrates the domain knowledge of human experts into the analysis process of artificial intelligence, effectively improving the model's interpretability, the accuracy of professional judgment, and the ability to handle complex and marginal cases.
[0034] Example 3: Building upon the previous embodiments, this example uses a high-altitude environment as a case study to demonstrate how to form a complete technological closed loop from basic research to practical application. It involves a longitudinal study of newly enlisted soldiers about to be deployed to high-altitude areas. Peripheral blood samples were collected from the soldiers before their arrival at their high-altitude base, and their basic physiological indicators were recorded. Sample collection and environmental data recording were repeated at specific time points after arrival at the high-altitude base. All samples underwent integrative genomic analysis to construct individualized chromosomal variation-conformation co-process datasets.
[0035] These longitudinal data points are input into an intelligent analysis system. The system utilizes its built-in high-altitude environmental characteristic spectrum comparison model, trained with historical data, to analyze data from each soldier at different time points. The system not only identifies known genetic variations associated with hypoxia adaptation, but more importantly, it can detect dynamic remodeling related to hypoxia exposure at the three-dimensional genome level, such as the specific changes in the TAD boundary strength of gene clusters related to erythropoiesis regulation over exposure time. The system calculates the functional network perturbation intensity corresponding to each time point sample and maps its individual risk evolution trajectory.
[0036] Based on this longitudinal data and the network disturbance intensity calculated by the system, accurate risk stratification and early warning can be achieved. The system found that soldiers who exhibited a rapid increase and sustained high level of network disturbance intensity early after arriving at high altitudes had a significantly higher risk of developing clinically diagnosed acute mountain sickness. Conversely, soldiers with lower network disturbance intensity or a rapid adaptive decrease showed good environmental adaptation. Therefore, the system can identify "high-risk maladaptive" individuals in the early stages of soldiers' arrival at high altitudes, based on their initial genetic damage and functional disturbance response patterns, and automatically generate high-level warning reports, recommending interventions such as enhanced medical monitoring, adjustment of training intensity, or consideration of early evacuation.
[0037] Furthermore, the system utilizes a proactive intervention simulation mechanism. For identified high-risk individuals, military doctors can input different intervention hypotheses, such as "hypothesize adding 2 hours of hypoxia pre-acclimatization training daily for this soldier" or "hypothesize transferring them to a lower-altitude camp." Based on established models, the system will simulate the potentially favorable shifts in the soldier's chromosomal variations and network perturbation trajectories under these interventions, thereby providing data-driven decision support for developing personalized, optimal protection and intervention plans.
Claims
1. A method for mining multi-environment chromosome variation feature spectra, characterized in that, include: Peripheral blood samples were obtained from different groups of people exposed to special environments, and their long-term, multi-dimensional environmental exposure physicochemical parameters were collected simultaneously. An integrated genomic analysis was performed on the samples to obtain a dataset of covariates, including primary sequence variations and three-dimensional spatial conformation variations of chromosomes. The original variant data in the covariance dataset are subjected to quality control, comparison and annotation, and batch effect correction. Construct a relational database to perform spatiotemporal correlation mapping of the covariation dataset, environmental exposure physicochemical parameters, and individual phenotypic data; Based on the aforementioned relational database, a unique chromosome variation-conformation joint feature spectrum under different environmental pressures is decoupled through an environmental comparative machine learning model. Using the joint feature spectrum, the perturbation intensity of a specific environmental exposure on key nodes and pathways of the core biological functional network is quantitatively assessed through a functional network perturbation model, and hierarchical early warning information is generated accordingly. Based on the newly added sample data and its corresponding early warning feedback results, the weight allocation of nodes and paths in the functional network is automatically optimized, and new potential relationships are discovered. The results of domain experts' review and semantic annotation of the joint feature spectrum and early warning information are fed back as prior knowledge to the environmental comparative machine learning model, thereby enabling continuous targeted optimization of the model.
2. The method according to claim 1, characterized in that: The acquisition of the three-dimensional spatial conformational variations of chromosomes is achieved by analyzing the structural changes of the topological association domain of chromosomes in the cell nucleus and the dynamic reconstruction of the internal chromatin loops. The conformational variations include, but are not limited to, the weakening or strengthening of the TAD boundary strength and the formation of abnormal chromatin loops.
3. The method according to claim 1, characterized in that: The environmental comparative machine learning model employs an adversarial generative network framework, which uses its discriminator component to learn and reinforce the differences in chromosomal variation patterns among groups exposed to different environments, thereby decoupling a robust feature spectrum that is highly specific to the environment and weakly individual-specific.
4. The method according to claim 1, characterized in that: The construction of the functional network perturbation model is as follows: The variants and conformational changes identified in the joint feature spectrum are mapped onto a pre-constructed functional network that covers multi-level biomolecular interactions. The overall functional disturbance intensity is quantified by calculating the topological impact scores and flow disturbance levels of the aforementioned mutations and conformational changes on critical nodes and vulnerable paths in the network.
5. The method according to claim 1, characterized in that: The method is based on the established joint feature spectrum and functional network perturbation model. It inputs simulated changes in environmental parameters and calculates and infers the potential evolutionary paths of chromosome variations and conformations and their corresponding functional perturbation inflection points, providing a theoretical basis for formulating critical intervention thresholds.
6. The method according to claim 1, characterized in that: The integrative genomic analysis also includes assessing the stability of fragile chromosomal sites and analyzing whether they become hotspots for structural variations and conformational abnormalities under specific environmental exposures.
7. A multi-environment chromosome variation feature spectrum mining system, characterized in that: The system is implemented based on the method of any one of claims 1-6, including: The collaborative data acquisition module is used to control multiple omics detection devices and connect to an environmental sensor network to obtain the collaborative variation dataset and environmental exposure physicochemical parameters; The data fusion and computing engine integrates the aforementioned environmental comparative machine learning model and the aforementioned functional network perturbation model at its core. The dynamic knowledge graph module is used to store the relational database and serve as the underlying network of the functional network perturbation model. The intelligent early warning and simulation terminal is used to output tiered early warning information and visualize the results of forward-looking intervention simulations.
8. The system according to claim 7, characterized in that: The system also includes a human-machine collaborative decision-making loop, which allows domain experts to review and semantically annotate the joint feature spectrum and early warning information generated by the system, and feed the annotated information back to the environmental comparative machine learning model as prior knowledge.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6, wherein the program specifically includes an operating instruction set for controlling a high-throughput chromosome conformation capture sequencer and a parallel computing instruction set for running the functional network perturbation model.
Citation Information
Patent Citations
Systems and methods for characterizing topological network perturbations
CN103843000A