A method for constructing a pathogenic microorganism genome database
Through high-throughput sequencing and genome database optimization technology, the problem of inaccurate analysis of genome evolution relationships in the construction of traditional pathogenic microorganism genome databases was solved, and an efficient and reasonable database was built to support the diagnosis and research of pathogenic microorganisms.
Patent Information
- Application Number
- CN202411638424.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-17
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-11-17
AI Technical Summary
Traditional methods for constructing pathogenic microorganism genome databases are inaccurate in analyzing genome evolutionary relationships, resulting in chaotic database construction.
The genome sequence of pathogenic microorganisms is obtained through high-throughput sequencing technology, and pathogen characteristic site analysis and physiological and biochemical characteristics are analyzed. The genome evolution model and relationship coding index construction are combined to optimize the cache architecture and form an efficient genome database.
The accuracy of the analysis of evolutionary relationships in pathogenic microorganism genomes has been improved, and a reasonable and efficient database has been constructed to support complex data analysis and query, providing strong support for the diagnosis, treatment and research of pathogenic microorganisms.
Smart Images

Figure CN119601093B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of genome database construction, and in particular to a method for constructing a pathogenic microorganism genome database. Background Art
[0002] The threat posed by pathogenic microorganisms to human health is gaining increasing attention. The construction of pathogenic microbial genome databases has become a crucial foundation for research into pathogenic mechanisms, drug target discovery, and disease prevention and control. Pathogens, including bacteria, fungi, viruses, and parasites, can cause a wide range of infectious and chronic diseases, such as tuberculosis, HIV / AIDS, malaria, and influenza. With the continuous advancement of high-throughput gene sequencing technology, scientists are able to more accurately obtain genomic information from pathogens. Through in-depth analysis of this information, they are revealing their genetic diversity, mutational patterns, transmission pathways, and interactions with their hosts. Currently, major research institutions and database platforms around the world have launched projects to collect and construct genomic information on pathogens, such as the International Human Genome Project and the Global Pathogen Genome Project. These databases provide crucial data support for global disease surveillance, diagnosis, and vaccine development. For example, by organizing and analyzing pathogen genomic information, researchers can discover drug-resistance genes, identify new pathogenicity factors, and even predict the evolutionary trends of pathogens, providing strong support for clinical treatment. Especially in the context of the increasingly serious problem of drug resistance, the application value of pathogen genome databases has become increasingly prominent. However, the traditional method of constructing pathogen genome databases is inaccurate in analyzing the evolutionary relationships of pathogen genomes, resulting in a problem of chaotic database construction. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for constructing a pathogenic microorganism genome database to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, a method for constructing a pathogenic microorganism genome database is provided, the method comprising the following steps:
[0005] Step S1: Obtaining a pathogenic microorganism genome sequence; performing pathogenic characteristic point analysis on the pathogenic microorganism genome sequence to obtain pathogenic characteristic point data;
[0006] Step S2: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence based on the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data; performing genome evolution relationship simulation based on the pathogen genome physiological and biochemical characteristic data to obtain genome evolution relationship data; performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data;
[0007] Step S3: performing a relationship nesting structure analysis on the pre-relationship linkage data to obtain pre-relationship linkage relationship nesting structure data; constructing a relationship coding index on the pre-relationship linkage data based on the pre-relationship linkage relationship nesting structure data to obtain a linkage relationship coding index;
[0008] Step S4: Dynamically match the linkage relationship coding index with search conditions to obtain the linkage relationship dynamic search conditions; optimize the cache architecture based on the linkage relationship dynamic search conditions to obtain the linkage relationship cache optimization architecture; and construct a genome database for the pathogenic microorganism genome sequence using the linkage relationship cache optimization architecture to obtain the pathogenic microorganism genome database.
[0009] The present invention first uses high-throughput sequencing technology to obtain the complete genome sequence of the pathogenic microorganism, providing basic data for subsequent analysis. Next, the genome sequence is analyzed for pathogenicity signature sites, identifying important genomic features associated with pathogenicity, such as pathogenicity genes and virulence factors. These signature site data reveal the pathogen's fundamental biological characteristics, such as pathogenicity and immune escape mechanisms, providing an important basis for subsequent physiological and biochemical characterization and genome evolutionary relationship simulation. In this phase, the pathogenic microorganism genome's physiological and biochemical characteristics are deeply analyzed based on the pathogenicity signature site data, extracting information such as metabolic pathways and drug resistance genes. This physiological and biochemical property data helps scientists understand the biological functions of the pathogen and its interactions with the host. Subsequently, using genome evolution models, evolutionary analysis of the genomic data is performed to reveal the evolutionary relationships, phylogenetic relationships, and mutational signatures between different microbial populations. Next, pre-relationship linkage identification is performed to explore the interdependencies between different genes, signature sites, and biological functions, providing a dynamic evolutionary path and key nodes for subsequent research. Nested structure analysis of pre-relationship linkage data reveals the hierarchical and complex relationships between different feature points, genes, and functions in microbial genomes. Deeply exploring these relationships creates structured nested relational data, helping researchers understand how various biological processes interact and identify key linkages. A relational encoding index is then constructed based on this nested structure. This encoding index tags and categorizes different relationships, enabling efficient data access and querying, laying the organizational framework for genome database construction. Dynamically matching search criteria within the linkage relationship encoding index allows for flexible extraction of relevant data based on actual needs, thus implementing an efficient retrieval mechanism. Based on the matching results, the linkage relationship cache architecture is optimized, and intelligent cache management improves data query speed and resource utilization. Ultimately, using this optimized cache architecture, a comprehensive pathogen genome database is constructed. This database not only stores pathogen genome data but also supports complex data analysis and querying, providing powerful support for the diagnosis, treatment, and research of pathogens. Therefore, the present invention is an improvement to a traditional method for constructing a pathogenic microorganism genome database, which solves the problem that the traditional method for constructing a pathogenic microorganism genome database is inaccurate in analyzing the evolutionary relationship of pathogenic microorganism genomes, thereby causing confusion in database construction. It improves the accuracy of the analysis of the evolutionary relationship of pathogenic microorganism genomes and improves the rationality of database construction.
[0010] Preferably, step S1 includes the following steps:
[0011] Step S11: obtaining the genome sequence of the pathogenic microorganism;
[0012] Step S12: filtering low-quality sequences of the pathogenic microorganism genome sequence to obtain a pathogenic microorganism genome filtered sequence;
[0013] Step S13: functional classification identification is performed on the pathogenic microorganism genome filtered sequence to obtain the microorganism genome functional classification sequence;
[0014] Step S14: performing pathogenic characteristic point analysis on the functional classification sequence of the microbial genome to obtain pathogenic characteristic point data.
[0015] The present invention uses modern high-throughput genome sequencing technologies (such as Illumina sequencing and PacBio sequencing) to obtain the complete genome sequence of pathogenic microorganisms. The genome sequence is the foundation of subsequent research, covering the microorganism's genetic information, gene expression, and potential pathogenic mechanisms. By sequencing the genome of pathogenic microorganisms, scientists can obtain data on their entire genome. This data provides an important basis for identifying characteristics such as pathogenicity, drug resistance, and tolerance, and is the starting point for disease prevention and treatment research. After obtaining the original genome sequence, some low-quality sequencing data, such as sequencing errors, sequence duplications, or contaminating sequences, are often present. If this low-quality data is not filtered, it will affect the accuracy and reliability of subsequent analysis. Therefore, this step uses quality control technology to filter out low-quality sequences, ensuring that the obtained genome sequence has high accuracy and completeness. The filtered genome data will provide a high-quality foundation for subsequent analysis. Functional classification and labeling are performed on the filtered genome sequence to identify functional regions and genes in the genome related to pathogenicity. Through gene prediction and functional annotation, genes and functional elements related to cellular metabolism, toxin production, drug resistance, immune escape, etc. are identified. This process is usually combined with public gene libraries (such as NCBI, KEGG, etc.) for comparison and annotation, with the aim of helping researchers understand the biological functions of pathogenic microorganisms, identify their potential pathogenic mechanisms, and provide important clues for subsequent feature site analysis and the formulation of disease prevention and control strategies. Based on the functional classification sequence of the microbial genome, further in-depth exploration of feature sites related to pathogenicity, such as specific pathogenic genes, virulence factors, drug resistance genes, and key genes that interact with the host. The analysis of these feature sites can reveal the biological characteristics of pathogenic microorganisms, such as pathogenic mechanisms, virulence mechanisms, and adaptive evolution. For example, certain feature sites are related to toxin production, immune escape, or drug resistance of microorganisms. Through the analysis of pathogenic feature sites, researchers can provide theoretical support for early diagnosis, drug target discovery, and vaccine design, and provide a scientific basis for controlling diseases caused by pathogenic microorganisms.
[0016] Preferably, step S2 includes the following steps:
[0017] Step S21: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence according to the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data;
[0018] Step S22: identifying the environmental weight impact of the pathogen genome physiological and biochemical characteristic data to obtain physiological and biochemical environmental weight impact data;
[0019] Step S23: performing genome evolution relationship simulation on the pathogenic microorganism genome sequence according to the physiological and biochemical environment weight impact data to obtain genome evolution relationship data;
[0020] Step S24: performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data.
[0021] Based on previously acquired pathogen signature data, the present invention analyzes the physiological and biochemical characteristics of pathogenic microorganism genomes. This process involves detailed analysis of the proteins, enzymes, metabolic pathways, and other features encoded in the genome, thereby identifying the physiological activities and metabolic characteristics of the microorganisms in their host environment. By analyzing characteristics such as metabolic pathways, energy production, toxin synthesis, and drug resistance, the pathogenic mechanisms and biological characteristics of the pathogens can be revealed. The obtained physiological and biochemical characteristic data provide a scientific basis for in-depth research on the interaction between pathogens and their hosts, pathogenesis, and drug intervention. By analyzing the physiological and biochemical characteristic data of pathogen genomes, the impact of different environmental factors on pathogenic microorganism characteristics can be identified. For example, external environmental factors such as temperature, pH, changes in nutrients, and host immune responses affect microorganisms' metabolic activity, virulence, or drug resistance. Through environmental weight analysis, the intensity of the impact of different environmental factors on microbial physiological functions is assessed, thereby obtaining physiological and biochemical environmental weighted impact data. This data helps understand how pathogens regulate their biological activities in different environments, reveals their adaptive mechanisms, and provides theoretical support for the development of environmental control strategies. Based on physiological and biochemical environmental weighted influence data, we simulate the evolutionary relationships of pathogenic microorganism genomes. By integrating the influence of environmental factors on physiological and biochemical traits, we simulate how microbial genomes evolve to adapt to new challenges under different environmental conditions. This process helps researchers understand the adaptive evolution and genetic variation of pathogenic microorganisms, revealing how genomes adjust to environmental changes or host immune pressure during long-term evolution. Genomic evolutionary relationship data provides important insights for further studying microbial evolutionary dynamics, mutation hotspots, and the evolution of drug resistance. By conducting in-depth pre-relationship linkage identification on genomic evolutionary relationship data and analyzing the relationships between different genes and between genes and physiological and biochemical traits, we can reveal how different functional modules, genes, and environmental factors within the microbial genome influence each other. For example, mutations in certain genes are closely associated with increased virulence, the development of drug resistance, or adaptation to environmental changes. Through pre-relationship linkage identification, researchers can construct complex causal chains and interaction networks within microbial genomes, providing theoretical support for a better understanding of the evolutionary pathways, drug resistance mechanisms, and interactions between pathogens and their hosts, and providing data support for the development of new drugs and vaccine designs.
[0022] Preferably, step S23 includes the following steps:
[0023] Step S231: performing spatiotemporal dimension effect intensity analysis on the physiological and biochemical environment weight impact data to obtain spatiotemporal dimension effect intensity data;
[0024] Step S232: performing pressure chemical reaction change analysis on the pathogenic microorganism genome sequence based on the spatiotemporal action intensity data to obtain pathogen spatiotemporal pressure chemical reaction change data;
[0025] Step S233: Calculating the probability of gene segment variation of the pathogenic microorganism genome sequence based on the pathogen spatiotemporal pressure chemical reaction variation data to obtain gene segment variation probability data;
[0026] Step S234: performing quantitative prediction of the variation trend of the gene segment variation probability data to obtain quantitative data of the variation trend of the gene segment;
[0027] Step S235: performing genome evolution relationship simulation on the genome sequence of the pathogenic microorganism according to the gene segment variation probability data and the gene segment variation trend quantitative data to obtain genome evolution relationship data.
[0028] The present invention analyzes the physiological and biochemical environmental weight impact data to deeply explore the degree of influence of environmental factors (such as temperature, humidity, pH value, etc.) in different time and space dimensions on pathogenic microorganisms. The analysis of the intensity of the effect of the spatiotemporal dimension can reveal the dynamic changes of environmental factors on the physiological activities of microorganisms, reflecting how the impact of the environment on microorganisms changes over time and space. These data help researchers understand how pathogenic microorganisms undergo adaptive changes or evolution under different environmental conditions, and thus provide a scientific basis for predicting the viability and pathogenicity of microorganisms in different scenarios. Based on the intensity data of the spatiotemporal dimension, the chemical reaction changes of pathogenic microorganisms under different environmental pressures are studied. These pressures usually include factors such as temperature changes, changes in nutrient concentrations, and fluctuations in oxygen content. These environmental pressures have an important impact on the metabolic reactions and physiological activities of microorganisms. Through the analysis of changes in pressure chemical reactions, researchers can identify changes in the biochemical reaction pathways and metabolic pathways of microorganisms under different pressure conditions, thereby revealing how microorganisms regulate their gene expression and functional performance when responding to environmental pressure. This data provides key information for further understanding the physiological adaptability and stress resistance mechanisms of microorganisms. By analyzing data on temporal and spatial pressure-induced chemical reactions of pathogenic microorganisms under different environmental pressures, the probability of mutation of each gene segment in the microbial genome can be inferred. This process involves using bioinformatics algorithms to combine environmental pressure, chemical reaction changes, and gene mutation frequency to calculate the probability of gene segment mutation under specific conditions. Gene segment variation is a key factor in microbial adaptive evolution. Mutation probability data can help researchers predict genetic variation and novel mutations that may occur in microorganisms under specific environments, which is crucial for understanding changes in microbial resistance and virulence. In this step, based on the gene segment variation probability data obtained in the previous step, mathematical models and prediction algorithms are used to quantitatively analyze gene segment variation trends. This process can help researchers predict the frequency and type of mutations that will occur in specific genes or genomic regions in the future. Through quantitative prediction, researchers can identify gene segments most likely to mutate and further assess the potential impact of these mutations on microbial biological characteristics (such as virulence and drug resistance). Quantitative data on variation trends provide valuable information on the evolutionary direction of pathogens, the spread of drug resistance, and disease prediction. In this step, the evolutionary relationships of pathogenic microorganism genomes are simulated by combining data on gene segment mutation probabilities and quantitative data on gene segment mutation trends. By simulating the evolution of microbial genomes under specific environmental pressures, researchers can depict the temporal trajectory of pathogenic microorganism genomes and reveal how their genomes evolve gradually with environmental pressure and the accumulation of genetic variation. This genomic evolutionary relationship data can provide a scientific basis for predicting microbial evolutionary trends, changes in pathogenicity, and the spread of drug resistance.These data will help identify the future evolutionary paths of pathogenic microorganisms and provide important references for adjusting prevention and control strategies, vaccine development, and antibiotic use.
[0029] Preferably, step S3 includes the following steps:
[0030] Step S31: performing relationship nesting structure analysis on the preceding relationship linkage data to obtain preceding linkage relationship nesting structure data;
[0031] Step S32: performing pathogenic microorganism population genetic structure analysis on the pathogen genome physiological and biochemical characteristic data according to the pre-linkage relationship nested structure data to obtain pathogenic microorganism population genetic structure data;
[0032] Step S33: constructing a relationship coding index for the preceding relationship linkage data according to the preceding linkage relationship nested structure data and the pathogenic microorganism population genetic structure data to obtain a linkage relationship coding index.
[0033] The present invention aims to identify and reveal the complex relationships between different genes or physiological characteristics and their hierarchical structure of mutual influence by performing nested structure analysis on pre-relationship linkage data. By layering and structuring multidimensional data, it is possible to clearly identify which relationships are intrinsically dependent and which are externally influenced, thereby revealing the multiple correlations of microbial genomes during their evolution. This analysis helps to understand how genes, environmental characteristics and physiological functions interact through complex networks, and provides a structured information foundation for subsequent relationship encoding and data modeling. Based on the nested structure data of pre-relationship linkage, population genetic structure analysis of the physiological and biochemical characteristics of the genomes of pathogenic microorganisms is performed. Through population genetic analysis, genetic differences between different groups or subpopulations can be assessed, and how they exhibit different adaptability or pathogenicity characteristics under environmental pressure. This step helps to reveal genetic structural characteristics such as gene diversity, genetic drift and gene flow in pathogenic microbial populations, thereby gaining a deeper understanding of the evolutionary patterns of microbial populations and the selective pressures within the populations, as well as how they interact with the environment, host and pathological characteristics. Combining the nested structure data of the pre-linked relationship and the population genetic structure data, a relational coding index is constructed to provide a clear and efficient indexing mechanism for the entire data system. The relational coding index simplifies data query and subsequent analysis by encoding and mapping complex relationships, integrating and classifying data from multiple dimensions. In this way, the connections between data can be accessed, understood, and interpreted more quickly, improving data availability and parsing efficiency. For the processing of large-scale pathogen genomic data, constructing an effective relational coding index can save computing resources in subsequent research and accelerate research progress in areas such as the evolution of pathogenic microorganisms and the transmission pathways of drug resistance.
[0034] Preferably, step S32 includes the following steps:
[0035] Step S321: performing biochemical attribute similarity comparison on the pathogen genome physiological and biochemical characteristic data according to the pre-linkage relationship nested structure data to obtain genome biochemical attribute similarity data;
[0036] Step S322: performing gene abundance calculation on the genome biochemical attribute similarity data to obtain gene similarity abundance data;
[0037] Step S323: performing genetic population differentiation analysis on the genomic biochemical attribute similarity data based on the gene similarity abundance data to obtain genetic population differentiation data;
[0038] Step S324: performing pathogenic microorganism population genetic structure analysis based on the genetic population differentiation data and the gene similarity abundance data to obtain pathogenic microorganism population genetic structure data.
[0039] This study analyzes nested structure data from pre-linked relationships to compare the physiological and biochemical properties of pathogen genomes for similarity, aiming to identify similarities in biochemical properties between different pathogen groups or individuals. This analysis helps understand the differences in biochemical characteristics between different pathogens at the genomic level, as well as the similarities and differences in their physiological functions. For example, some pathogens respond to survival challenges through similar metabolic pathways under the same environmental pressures. Therefore, biochemical property similarity data provides a foundation for subsequent genetic analysis and population structure studies. Based on the biochemical property similarity comparison results, gene abundance of relevant genomic regions is further calculated to obtain gene similarity abundance data. Gene abundance reflects the relative number or expression level of different genes or genomic segments in a pathogen population. This data can help researchers identify genes or gene clusters that are more active under specific biochemical conditions and reveal which genes or gene combinations play an important role in microbial adaptability and evolution. By calculating abundance, genes that are significantly enriched in different groups or subpopulations can be distinguished, further understanding the genetic basis and physiological characteristics of microorganisms. Combining gene abundance data with genomic biochemical similarity data allows for genetic population differentiation analysis. By statistically analyzing gene abundance differences between different populations (e.g., microbial populations in different environments or hosts), the degree of genetic differentiation within a population is assessed. This analysis helps reveal how populations with different genetic backgrounds exhibit biochemical differences, further illuminating how pathogenic microbial populations adapt to diverse environmental pressures through natural selection and genetic drift. Genetic population differentiation data provides insights into how microbial populations evolve and differentiate, helping to uncover underlying adaptive mechanisms. Based on these genetic population differentiation data and gene abundance data, the genetic structure of pathogenic microbial populations is analyzed. By comprehensively considering gene diversity, gene abundance distribution, and genetic differentiation within a population, researchers can delineate the overall genetic structure of the microbial population, revealing the frequency distribution of different genotypes within the population and their adaptive performance in specific environments. Analyzing the genetic structure of pathogenic microbial populations provides insights into the genetic diversity, gene flow patterns, and genetic connections between populations, providing important theoretical foundations for the spread of antibiotic resistance, evolutionary trends, and pathogenicity differences in pathogens.
[0040] Preferably, step S323 includes the following steps:
[0041] According to the gene similarity abundance data, the gene abundance similarity distribution matrix is constructed for the genome biochemical attribute similarity data to obtain the gene abundance similarity distribution matrix;
[0042] Calculate the Fst value of the gene abundance similarity distribution matrix to obtain the gene similarity Fst value;
[0043] The genetic population differentiation data were obtained by analyzing the genetic population similarity distribution matrix based on the gene similarity Fst value.
[0044] The present invention constructs a gene abundance similarity distribution matrix based on gene abundance similarity data. This matrix calculates the gene abundance similarity between different samples or populations, presenting the abundance distribution of each gene across different microbial populations or samples. This matrix not only helps reveal the enrichment of different genes in different populations but also helps identify genes that exhibit similar abundance characteristics under different environmental or host conditions. By structuring this similarity data into a matrix format, it facilitates subsequent more complex genetic analyses and differentiation calculations. Using the gene abundance similarity distribution matrix, the degree of genetic differentiation of gene abundance between different populations is assessed by calculating the Fst value. The Fst value is an indicator of the degree of differentiation of gene frequencies between populations; larger values indicate more significant genetic differences between populations, while smaller values indicate less genetic differences between populations. By calculating the Fst value, it is possible to further understand which genes have strong genetic isolation between different populations, thereby providing important quantitative data on the genetic structure of pathogenic microbial populations. This calculation step is key to population genetics analysis and helps reveal the adaptive differentiation of microbial populations during evolution. Based on the gene similarity Fst values calculated in the previous step, the gene abundance similarity distribution matrix is used to perform genetic population differentiation analysis. The core of this analysis is to use Fst values to identify genes or gene clusters within the genome that exhibit significant population differentiation. Differences in the abundance of these genes reflect genetic isolation between populations due to differences in selective pressure, environmental adaptation, or gene flow. By analyzing the genetic population differentiation of these genes, we can gain a deeper understanding of the genetic structure of microbial populations, revealing the distribution of different genetic subtypes within a population and their adaptability to specific ecological niches. This analysis provides important support for further research into the evolutionary processes, adaptive mechanisms, and genetic connections between populations of pathogenic microorganisms.
[0045] The beneficial effects of the present invention include first obtaining the complete genome sequence of the pathogenic microorganism through high-throughput sequencing technology, which provides basic data for subsequent analysis. Next, the genome sequence is analyzed for pathogenicity signature sites, identifying important genomic features associated with pathogenicity, such as pathogenicity genes and virulence factors. These signature site data reveal the pathogen's fundamental biological characteristics, such as pathogenicity and immune escape mechanisms, providing an important basis for subsequent physiological and biochemical characterization and genome evolutionary relationship simulation. At this stage, the pathogenic microorganism genome's physiological and biochemical characteristics are deeply analyzed based on the pathogenicity signature site data, extracting information such as metabolic pathways and drug resistance genes. This physiological and biochemical property data helps scientists understand the biological functions of the pathogenic microorganism and its interactions with the host. Subsequently, using genome evolution models, evolutionary analysis of the genomic data is performed to reveal the evolutionary relationships, phylogenetic relationships, and mutational signatures between different microbial populations. Next, pre-relationship linkage identification is performed to explore the interdependencies between different genes, signature sites, and biological functions, providing a dynamic evolutionary path and key nodes for subsequent research. Nested structure analysis of pre-relationship linkage data reveals the hierarchical and complex relationships between different feature points, genes, and functions in microbial genomes. Deeply exploring these relationships creates structured nested relational data, helping researchers understand how various biological processes interact and identify key linkages. A relational encoding index is then constructed based on this nested structure. This encoding index tags and categorizes different relationships, enabling efficient data access and querying, laying the organizational framework for genome database construction. Dynamically matching search criteria within the linkage relationship encoding index allows for flexible extraction of relevant data based on actual needs, thus implementing an efficient retrieval mechanism. Based on the matching results, the linkage relationship cache architecture is optimized, and intelligent cache management improves data query speed and resource utilization. Ultimately, using this optimized cache architecture, a comprehensive pathogen genome database is constructed. This database not only stores pathogen genome data but also supports complex data analysis and querying, providing powerful support for the diagnosis, treatment, and research of pathogens. Therefore, the present invention is an improvement to a traditional method for constructing a pathogenic microorganism genome database, which solves the problem that the traditional method for constructing a pathogenic microorganism genome database is inaccurate in analyzing the evolutionary relationship of pathogenic microorganism genomes, thereby causing confusion in database construction. It improves the accuracy of the analysis of the evolutionary relationship of pathogenic microorganism genomes and improves the rationality of database construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic flow chart of the steps of a method for constructing a pathogenic microorganism genome database;
[0047] Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG.
[0048] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0049] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0050] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0051] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0052] To achieve this, please refer to Figures 1 to 2 A method for constructing a pathogenic microorganism genome database, the method comprising the following steps:
[0053] Step S1: Obtaining a pathogenic microorganism genome sequence; performing pathogenic characteristic point analysis on the pathogenic microorganism genome sequence to obtain pathogenic characteristic point data;
[0054] Step S2: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence based on the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data; performing genome evolution relationship simulation based on the pathogen genome physiological and biochemical characteristic data to obtain genome evolution relationship data; performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data;
[0055] Step S3: performing a relationship nesting structure analysis on the pre-relationship linkage data to obtain pre-relationship linkage relationship nesting structure data; constructing a relationship coding index on the pre-relationship linkage data based on the pre-relationship linkage relationship nesting structure data to obtain a linkage relationship coding index;
[0056] Step S4: Dynamically match the linkage relationship coding index with search conditions to obtain the linkage relationship dynamic search conditions; optimize the cache architecture based on the linkage relationship dynamic search conditions to obtain the linkage relationship cache optimization architecture; and construct a genome database for the pathogenic microorganism genome sequence using the linkage relationship cache optimization architecture to obtain the pathogenic microorganism genome database.
[0057] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a flow chart showing the steps of a method for constructing a pathogenic microorganism genome database according to the present invention. In this example, the method for constructing a pathogenic microorganism genome database includes the following steps:
[0058] Step S1: Obtaining a pathogenic microorganism genome sequence; performing pathogenic characteristic point analysis on the pathogenic microorganism genome sequence to obtain pathogenic characteristic point data;
[0059] In an embodiment of the present invention, the genome sequence of the pathogenic microorganism is first obtained. The acquisition of the genome sequence relies on gene sequencing technology, such as high-throughput second-generation sequencing technology (NGS) or third-generation sequencing technology. Through these technologies, DNA or RNA in the pathogenic microorganism sample is extracted, and a specific library construction method is used to convert it into a format suitable for sequencing. The sequence data is then read by a sequencing instrument to obtain complete pathogenic microorganism genome information. After the genome sequence is obtained, pathogenic characteristic point analysis is performed next. This process involves systematic analysis of the characteristic regions in the pathogenic microorganism genome sequence. Through gene annotation software or genome analysis tools (for example, BLAST, Glimmer, GeneMark, etc.), key genes or domains in the pathogen genome are identified and these characteristic points are marked. Characteristic points include genes that affect the pathogenicity of the pathogen, such as virulence factors, drug resistance genes, metabolism-related genes, etc. Using genome analysis tools, combined with known databases (such as NCBI Gene, KEGG, etc.), the genes in the sequence are functionally annotated to identify potential pathogenic characteristic points. Ultimately, pathogenic characteristic point data are output for subsequent analysis.
[0060] Step S2: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence based on the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data; performing genome evolution relationship simulation based on the pathogen genome physiological and biochemical characteristic data to obtain genome evolution relationship data; performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data;
[0061] In the embodiment of the present invention, the physiological and biochemical characteristics of the pathogenic microorganism genome are first analyzed according to the pathogenic characteristic point data. This step is based on the pathogenic characteristic point that has been marked, and a multi-dimensional functional evaluation is performed on it by biochemical analysis tools, mainly analyzing the physiological and biochemical characteristics such as the metabolic pathway, enzyme system, and material metabolism capacity of the genome. Through gene annotation results, combined with metabolic network and biochemical reaction databases (such as KEGG, BioCyc, etc.), the metabolic function of the pathogenic microorganism is modeled, and its physiological response under different environmental conditions is evaluated. After completing the genome physiological and biochemical characteristics analysis, these physiological and biochemical data are utilized to simulate the genome evolution relationship. This process depends on the sequence alignment and phylogenetic analysis method of the gene. By comparing the genomes of different pathogenic species or different strains, the evolutionary tree of the genome is constructed. Software such as MEGA, RAxML, FastTree can be used to calculate the evolutionary tree of the genome by maximum likelihood method or Bayesian method. These softwares, by analyzing the similarity of the genomes between different species or strains, infer their common ancestor and generate genome evolutionary relationship data. Next, we identify pre-relationship linkages based on genomic evolutionary relationship data. This process involves in-depth analysis of evolutionary linkages between different genomes. We use the topological structure of the genomic evolutionary tree, combined with sequence homology between genes, to identify pre-relationships between genomes. For example, by looking at phenomena such as symbiotic evolution and horizontal gene transfer between genes, we can identify key linkage points in evolutionary relationships, ultimately generating pre-relationship linkage data.
[0062] Step S3: performing a relationship nesting structure analysis on the pre-relationship linkage data to obtain pre-relationship linkage relationship nesting structure data; constructing a relationship coding index on the pre-relationship linkage data based on the pre-relationship linkage relationship nesting structure data to obtain a linkage relationship coding index;
[0063] In an embodiment of the present invention, the relationship nesting structure analysis is first performed on the pre-relationship linkage data. This analysis is based on the obtained pre-relationship linkage data and uses graph theory methods to establish a network structure of complex relationships. Through topological analysis of the relationship data, the nested hierarchical relationship between each pathogen genome is identified. Graph theory algorithms, such as the shortest path algorithm, community detection algorithm, etc., are used to extract the nested hierarchical structure in the relationship, and ultimately obtain the pre-relationship linkage relationship nested structure data. Based on the pre-relationship linkage relationship nested structure data, a relationship coding index is further constructed. This step encodes the nested structure so that each relationship node can obtain a unique identifier. The encoding process can be implemented by a hash algorithm or an encoding method based on a B-tree or a Trie tree to ensure that each linkage relationship can be quickly retrieved. Utilizing these encoding information, an index structure of the linkage relationship is established to facilitate subsequent efficient database retrieval and update.
[0064] Step S4: Dynamically match the linkage relationship coding index with search conditions to obtain the linkage relationship dynamic search conditions; optimize the cache architecture based on the linkage relationship dynamic search conditions to obtain the linkage relationship cache optimization architecture; and construct a genome database for the pathogenic microorganism genome sequence using the linkage relationship cache optimization architecture to obtain the pathogenic microorganism genome database.
[0065] In an embodiment of the present invention, dynamic search condition matching is first performed on the linkage relationship encoding index. The goal of this step is to optimize the index query process by setting search conditions. Based on the database query engine, dynamic query conditions are designed that are adapted to the characteristics of the linkage relationship data. By using query optimization algorithms (e.g., query plan generation, index selection algorithms, etc.), the most appropriate indexing method is adjusted and matched according to different data requirements to improve query efficiency. Next, the cache architecture is optimized based on the dynamic search conditions of the linkage relationship. To accelerate the retrieval process, a caching strategy is adopted to preload and store the linkage relationship data. By utilizing the LRU (least recently used) cache replacement algorithm, etc., it is determined which data should be cached and dynamically adjusted based on real-time query requirements. By analyzing data access frequency, the cache content is further optimized to ensure efficient query response. Finally, the optimized linkage relationship cache architecture is used to construct a genome database for pathogenic microorganism genome sequences. At this stage, through the caching mechanism and indexing system, the pathogenic microorganism genome sequences are stored in an efficient database structure, ensuring that the database can quickly respond to various query requests. Database management can be achieved using a relational database (such as MySQL, PostgreSQL) or a NoSQL database (such as MongoDB). Combined with the aforementioned cache architecture, this allows for rapid retrieval, storage, and updating of genomic data. Ultimately, the pathogen genomic database is constructed, providing efficient data support for subsequent research.
[0066] Preferably, step S1 includes the following steps:
[0067] Step S11: obtaining the genome sequence of the pathogenic microorganism;
[0068] Step S12: filtering low-quality sequences of the pathogenic microorganism genome sequence to obtain a pathogenic microorganism genome filtered sequence;
[0069] Step S13: functional classification identification is performed on the pathogenic microorganism genome filtered sequence to obtain the microorganism genome functional classification sequence;
[0070] Step S14: performing pathogenic characteristic point analysis on the functional classification sequence of the microbial genome to obtain pathogenic characteristic point data.
[0071] In an embodiment of the present invention, DNA or RNA is first extracted from a pathogenic microorganism sample. Specifically, the genome of the pathogenic microorganism is obtained by molecular biology techniques in the laboratory, such as bacterial culture, virus extraction, etc. The extracted genome is converted into a library suitable for sequencing by a standardized method. After the library is constructed, the genome is sequenced using high-throughput sequencing technology (such as Illumina sequencing or PacBio long-read sequencing technology). High-throughput sequencing can read a large number of gene sequences at a time, and has high sequencing accuracy, and can comprehensively obtain the genome sequence of the pathogenic microorganism. After obtaining the raw sequencing data, in combination with relevant quality control standards, the raw data is preliminarily processed and filtered to ensure that the genome sequence obtained is accurate. The pathogenic microorganism genome sequence obtained in step S11 is filtered for low-quality sequences. First, the sequencing data is preliminarily evaluated by quality control software (such as FastQC) to check the quality indicators of the sequence, including sequencing error rate, GC content, sequence length distribution, etc. Low-quality sequences are usually characterized by high error rate, low quality score (Phred score is less than 20) and abnormal GC content. Based on this, low-quality sequences are removed using quality filtering tools (such as Trimmomatic). This process involves setting a quality threshold to remove sequences with quality scores below the set standard. Trimming also removes adapter sequences and low-quality regions to ensure the quality of the filtered genome sequences. Ultimately, the filtered sequences are referred to as pathogen genome filtered sequences, providing foundational data for subsequent analysis. Functional classification is then performed on the pathogen genome filtered sequences. First, the filtered sequences are aligned and analyzed using known gene annotation and functional databases (such as KEGG and Gene Ontology). Using sequence alignment tools (such as BLAST or DIAMOND), each gene fragment is aligned with known gene sequences in the functional database to identify the functional regions of the genes. This alignment reveals the functional categories of each gene, including metabolic genes, drug resistance genes, virulence factors, and so on. Based on the alignment results, the gene sequences are classified using functional classification algorithms (such as hierarchical clustering or graph clustering) to form a pathogen genome functional classification sequence. Further annotation based on gene function and location is required during this process to ensure accurate and comprehensive classification. The functional classification sequence of the microorganism genome is analyzed for pathogenicity characteristic sites. Based on the functional classification information obtained in step S13, the characteristic genes and sequence regions in the pathogenic microorganism genome are further focused on. These characteristic sites are closely related to the pathogenicity of the pathogenic microorganism. Pathogenicity characteristic site identification tools (such as feature extraction methods based on hidden Markov models or sequence alignments) are used to perform a refined analysis of the functional classification sequence.The identified characteristic sites include drug resistance genes, virulence factors, and immune escape-related genes, all of which play a crucial role in the pathogenicity of pathogens. During analysis, it is also necessary to compare with known pathogen feature databases (such as the CARD database and the VFDB database) to ensure accurate identification of the characteristic sites. Ultimately, the pathogen characteristic site data is obtained, which serves as the basis for the next step of analysis. This characteristic site data provides key biological information for subsequent biological property analysis of pathogens and database construction.
[0072] Preferably, step S2 includes the following steps:
[0073] Step S21: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence according to the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data;
[0074] Step S22: identifying the environmental weight impact of the pathogen genome physiological and biochemical characteristic data to obtain physiological and biochemical environmental weight impact data;
[0075] Step S23: performing genome evolution relationship simulation on the pathogenic microorganism genome sequence according to the physiological and biochemical environment weight impact data to obtain genome evolution relationship data;
[0076] Step S24: performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data.
[0077] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:
[0078] Step S21: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence according to the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data;
[0079] In an embodiment of the present invention, the pathogenic microorganism genome sequence is analyzed for physiological and biochemical characteristics based on pathogenic characteristic point data. First, detailed genome annotation is performed using identified pathogenic characteristic point data (such as metabolic-related genes, virulence genes, drug resistance genes, etc.). Each characteristic point represents a gene with a specific physiological or biochemical function, and the role of these genes in the physiological functions, metabolic pathways, and interactions with the host of the microorganism is crucial. Using gene annotation databases (such as KEGG, Reactome, etc.), characteristic points are compared and analyzed to identify related metabolic pathways, enzymatic reactions, and their roles in microbial physiological activities. For example, genes involved in important physiological activities such as amino acid metabolism, carbohydrate metabolism, and lipid synthesis in the pathogen genome can be identified. The physiological functions of these genes are associated with the relevant gene sets and biochemical networks in the existing literature to obtain the physiological and biochemical characteristics data of the pathogenic microorganism, including information such as metabolic pathways, key enzyme activities, energy metabolism, and environmental adaptability. The pathogen genome physiological and biochemical characteristics data ultimately formed provide a detailed biological background for subsequent analysis.
[0080] Step S22: identifying the environmental weight impact of the pathogen genome physiological and biochemical characteristic data to obtain physiological and biochemical environmental weight impact data;
[0081] In an embodiment of the present invention, the environmental weight impact identification is performed on the physiological and biochemical characteristics data of the pathogen genome. The goal of this process is to evaluate the degree of influence of environmental factors (such as temperature, pH value, oxygen concentration, salinity, etc.) on the physiological and biochemical characteristics of pathogenic microorganisms. First, through literature review or experimental data, the influence weights of known environmental factors on specific physiological characteristics (such as metabolic rate, enzyme activity, etc.) are collected. The influence of these environmental factors is quantified by laboratory data or historical environmental data. On this basis, a weighted scoring algorithm is used to calculate the influence of each environmental factor on each physiological and biochemical characteristic. For example, certain metabolic pathways are more active under high temperature or low oxygen environments, and therefore, these environmental factors are given higher weights in the analysis. In addition, based on the influence of changes in environmental factors on physiological responses, an environmental impact model is constructed, and physiological and biochemical environmental weight impact data are obtained by multiple regression analysis or weighted average method. In this way, it is possible to identify which environmental factors have a significant impact on the physiological processes of pathogenic microorganisms and form an environmental weight data set.
[0082] Step S23: performing genome evolution relationship simulation on the pathogenic microorganism genome sequence according to the physiological and biochemical environment weight impact data to obtain genome evolution relationship data;
[0083] In the embodiment of the present invention, according to the physiological and biochemical environment weight influence data, the genome sequence of pathogenic microorganisms is simulated by genome evolution relationship. First, by analyzing the function of each gene in the pathogenic microorganism genome and its interaction relationship, combined with the environmental weight influence data, the evolution process of the genome under different environmental conditions is simulated. The genome evolution relationship simulation is based on the mechanisms such as gene mutation, recombination, gene loss and acquisition, considering the selection pressure of the external environment, and inferring how the genome carries out adaptive evolution in environmental changes. By using evolutionary tree construction methods (such as maximum likelihood method, neighbor joining method or Bayesian inference method), the genomes of different pathogenic microorganisms are compared and analyzed to evaluate their evolutionary branches and evolutionary paths. In the simulation process, the environmental weight influence data is used as a key factor to adjust the fitness of the gene and predict the evolutionary trend of the genome under different environmental conditions. Finally, the genome evolution relationship data obtained comprise the evolutionary history, adaptive changes and evolutionary paths of each gene in the genome in a specific environment, for subsequent database construction provides detailed evolutionary background.
[0084] Step S24: performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data.
[0085] In an embodiment of the present invention, the genome evolution relationship data is subjected to pre-relation linkage identification. The core purpose of this step is to reveal the relationship linkage between different genes or genomic regions in the genome, particularly the gene pairs with covariance in the evolution process. First, by conducting an in-depth analysis of the genome evolution relationship data, the dependency or interaction relationship between genes is identified. These relationships are manifested as the co-evolution of genes, i.e., under specific environments or specific adaptation conditions, there is a close functional linkage between certain genes. In order to identify these pre-relation linkages, the connectivity analysis in association rule mining technology (such as Apriori algorithm) or graph theory is adopted to detect the close correlation between genes or genome fragments. Based on the function, expression pattern of genes and their adaptability to the environment, the gene pairs or genome regions that are dependent on each other in the genome evolution process are identified. By this method, it is possible to reveal which genes are "pre-relationships" in the evolution process, i.e., their changes directly affect the functions of other genes. Finally, the pre-relation linkage data obtained provide detailed information about the interaction of genes in the genome and the evolutionary mechanism for subsequent analysis.
[0086] Preferably, step S23 includes the following steps:
[0087] Step S231: performing spatiotemporal dimension effect intensity analysis on the physiological and biochemical environment weight impact data to obtain spatiotemporal dimension effect intensity data;
[0088] Step S232: performing pressure chemical reaction change analysis on the pathogenic microorganism genome sequence based on the spatiotemporal action intensity data to obtain pathogen spatiotemporal pressure chemical reaction change data;
[0089] Step S233: Calculating the probability of gene segment variation of the pathogenic microorganism genome sequence based on the pathogen spatiotemporal pressure chemical reaction variation data to obtain gene segment variation probability data;
[0090] Step S234: performing quantitative prediction of the variation trend of the gene segment variation probability data to obtain quantitative data of the variation trend of the gene segment;
[0091] Step S235: performing genome evolution relationship simulation on the genome sequence of the pathogenic microorganism according to the gene segment variation probability data and the gene segment variation trend quantitative data to obtain genome evolution relationship data.
[0092] In an embodiment of the present invention, the physiological and biochemical environmental weighted impact data is first converted into the effect intensity of the spatiotemporal dimension, and spatiotemporal analysis is performed. This process requires a spatiotemporal distribution analysis of environmental factors (such as temperature, humidity, light, oxygen concentration, etc.). Specifically, by collecting environmental data over a period of time and the physiological and biochemical response data of pathogenic microorganisms under different environmental conditions, a multidimensional data model containing time and space factors is constructed. Spatiotemporal analysis methods are used, such as Kriging interpolation to smooth the spatial data, and combined with time series analysis (such as moving average or exponential smoothing), to calculate the effect intensity of each environmental factor on physiological and biochemical characteristics at different time points and spatial positions. In this way, the effect intensity of environmental factors in different time and space dimensions can be obtained to form spatiotemporal effect intensity data. These data reflect the spatiotemporal distribution characteristics of environmental changes on microbial physiological responses, providing a detailed environmental background for subsequent genome evolution simulations. Pressure chemical reaction change analysis is performed based on spatiotemporal effect intensity data. First, the spatiotemporal effect intensity data is combined with the physiological response data of the pathogenic microorganism genome to identify changes in the chemical reaction pathways of the microbial genome under different environmental pressures. Through high-throughput experimental data or literature review, patterns of changes in pathogen metabolic or enzymatic reactions under specific environmental stresses (such as pH changes, temperature fluctuations, and changes in ion concentration) are collected. Then, kinetic models (such as the Michaelis-Menten kinetic model or the Hill equation) are used to simulate changes in the chemical reaction rates and reaction pathways of pathogens under different environmental stresses. Combined with the spatiotemporal effect intensity data, multivariate analyses (such as analysis of variance or principal component analysis) are used to assess the intensity of the impact of environmental stress on chemical reaction pathway changes, further clarifying the effect of each stress condition on the changes in pathogen chemical reactions. Ultimately, the resulting data on spatiotemporal chemical reaction changes under different environmental stresses can describe the adaptive changes in microbial genomes under different environmental stresses and reveal the dynamic adjustments in their metabolic pathways. The spatiotemporal chemical reaction change data are used to calculate the mutation probability of gene segments. First, based on the spatiotemporal chemical reaction change data obtained in the previous stage, gene segments that undergo changes in response under specific environmental stresses are identified. These gene segments are key genes or metabolic regulatory genes related to environmental adaptability. By performing mutation analysis on the sequences of these genes and combining them with stress conditions, we use gene mutation probability calculation methods to quantify the mutation risk of specific gene segments under different environmental pressures. For example, we use gene mutation models (such as neutral mutation models or selective mutation models) to calculate the probability of gene mutation under stress, taking into account the selective effect of environmental stress on gene mutation. During the calculation process, environmental stress data (such as temperature, pH, oxygen concentration, etc.) are combined with gene mutation data to evaluate how specific environmental factors affect the mutation frequency and type of gene mutations.This process uses Monte Carlo simulation to simulate gene variation under different environmental conditions, thereby obtaining mutation probability data for each gene segment under these conditions. Quantitative predictions of mutation trends are then made based on this gene segment mutation probability data. First, a gene mutation trend model is constructed using the gene segment mutation probability data obtained in step S233. This process uses statistical methods, such as linear regression analysis, logistic regression, or time series analysis, to predict the long-term trends of gene segment variation under given environmental pressures. Based on collected historical or experimental data, the impact of different environmental pressures on gene variation trends over different time periods is assessed. For example, environmental stressors such as temperature fluctuations and changes in oxygen concentration can be used to analyze the cumulative effects and long-term evolutionary trends of gene mutations. The model considers the dynamics of gene mutation, such as changes in mutation rate over time and increasing or decreasing mutation frequencies. Ultimately, the resulting quantitative data on gene segment variation trends will provide a basis for predicting the evolution of pathogen genomes during long-term environmental adaptation, revealing the variation patterns of gene segments under different environmental pressures. Combining the gene segment variation probability data with the quantitative data on gene segment variation trends allows for simulation of evolutionary relationships within pathogen genomes. First, using the probability data and quantitative data on gene segment mutation trends as input parameters, an evolutionary model (such as the Darwinian evolution algorithm or the genetic drift model) is used to simulate the evolution of the entire genome. This simulation considers the probability, frequency, and type of mutation of gene segments, as well as their long-term evolutionary trends under different environmental pressures. By constructing a genome evolution model, the adaptive changes of individual gene segments under different environmental pressures are evaluated, simulating the evolutionary path of the genome under long-term environmental influences. Specifically, genome evolution simulation tools (such as genetic algorithms and Monte Carlo simulations) are used to predict the adaptive changes and evolutionary trajectories of the genome under different conditions. During the simulation, combined with the quantitative data on gene segment mutation probabilities and trends, the dynamic process of genome evolution is quantitatively predicted, ultimately generating evolutionary relationship data for pathogen genomes, providing a detailed map of the adaptive evolution of pathogen genomes. These data will provide the necessary genome evolutionary context for subsequent database construction.
[0093] Preferably, step S3 includes the following steps:
[0094] Step S31: performing relationship nesting structure analysis on the preceding relationship linkage data to obtain preceding linkage relationship nesting structure data;
[0095] Step S32: performing pathogenic microorganism population genetic structure analysis on the pathogen genome physiological and biochemical characteristic data according to the pre-linkage relationship nested structure data to obtain pathogenic microorganism population genetic structure data;
[0096] Step S33: constructing a relationship coding index for the preceding relationship linkage data according to the preceding linkage relationship nested structure data and the pathogenic microorganism population genetic structure data to obtain a linkage relationship coding index.
[0097] In the embodiment of the present invention, the pre-relation linkage data is first subjected to structured analysis to reveal the hierarchy and relevance between different data items. For this reason, a hierarchical clustering algorithm is used to perform multi-level analysis on the pre-relation linkage data. In the specific operation, the pre-relation linkage data is first normalized so that the various data items have the same scale to eliminate the influence of dimensional differences. Then, a hierarchical clustering algorithm (such as single link, full link or average link clustering method) is used to cluster according to the similarity between the data items to construct a nested relationship structure. The clustering process calculates the distance metric (such as Euclidean distance, Manhattan distance, etc.) between each data point, and gradually merges the relationship data with high similarity based on the distance, and finally obtains a tree structure or a nested hierarchical structure. By carrying out a detailed analysis of the hierarchical structure of these relationship data, the deep dependency between each relationship data item can be revealed, and the pre-relation linkage relationship nested structure data can be obtained, thereby laying the foundation for subsequent data processing. First, the pre-relation linkage relationship nested structure data is used, combined with the pathogen genome physiological and biochemical characteristic data, to conduct an in-depth analysis of the genetic structure of the pathogenic microorganism population. Specifically, genetic differences between different populations are assessed using population structure analysis methods (such as Fst analysis or Wright's fixation index) based on the physiological and biochemical properties of pathogen genomes and pre-linked data. In this step, pathogen genomic data are grouped according to their physiological and biochemical properties, with microbial genomes with high similarity grouped together. Next, using population genetics theory, genetic variation indices (such as the Huffman index or gene diversity index) are calculated between different populations. This method reveals how the genetic structure of pathogen populations varies under different physiological environments and genetic characteristics, which genetic regions contribute most to population genetic differences, and obtains pathogen population genetic structure data. This analysis helps to deepen our understanding of pathogen genetic diversity and provides essential support for subsequent database construction and pathogen feature extraction. By combining pre-linked nested structure data with pathogen population genetic structure data, a coding index for pre-linked data is constructed. First, using pre-linked data and population genetic structure data, a multidimensional coding system is designed to reflect the complex interactions between relationship data. The specific implementation method uses Huffman coding or Lempel-Ziv coding based on tree-like nested data to assign a unique identifier to each relationship item, ensuring that each level of nested relationships can be accurately represented by a code. During the encoding process, different codes are assigned to different gene groups based on the results of genetic structure analysis. The codes are weighted based on factors such as gene diversity within the group or the expression level of specific genes. This ensures that the codes not only reflect the relationship hierarchy between data but also reflect the weight of genetic variation. Bitwise operations and tree-structured search algorithms are used during the encoding process to ensure efficiency and accuracy.Ultimately, the resulting linkage relationship encoding index will include the relationships between all relevant data items and their importance in the genetic structure, forming a compact and efficient index structure that facilitates rapid retrieval and processing of pathogenic microorganism genomic data.
[0098] Preferably, step S32 includes the following steps:
[0099] Step S321: performing biochemical attribute similarity comparison on the pathogen genome physiological and biochemical characteristic data according to the pre-linkage relationship nested structure data to obtain genome biochemical attribute similarity data;
[0100] Step S322: performing gene abundance calculation on the genome biochemical attribute similarity data to obtain gene similarity abundance data;
[0101] Step S323: performing genetic population differentiation analysis on the genomic biochemical attribute similarity data based on the gene similarity abundance data to obtain genetic population differentiation data;
[0102] Step S324: performing pathogenic microorganism population genetic structure analysis based on the genetic population differentiation data and the gene similarity abundance data to obtain pathogenic microorganism population genetic structure data.
[0103] In an embodiment of the present invention, biochemical property similarity analysis is performed by comparing various biochemical properties in the physiological and biochemical property data of pathogenic microorganism genomes based on the nested structure data of the pre-linked relationships. First, the physiological and biochemical properties of the pathogenic microorganism genomes are extracted as a number of numerical features, including but not limited to metabolite concentrations, enzyme activities, cell membrane structural components, metabolic pathways, etc. Next, the similarity between the physiological and biochemical properties of different genomes is calculated using Euclidean distance or Manhattan distance. Specifically, for each group of pathogenic genomes, the biochemical property data is first normalized to eliminate bias caused by different dimensions. Then, the distance between the physiological and biochemical property vectors of each genome is calculated, and cluster analysis methods are used to group them to identify genomes with similar biochemical properties. The results of these similarity comparisons will form genome biochemical property similarity data, which will be used for subsequent gene abundance calculation and genetic population analysis. Based on the genome biochemical property similarity data obtained in step S321, gene abundance calculation is performed. Gene abundance calculation quantifies the abundance of different genes by analyzing their copy number and expression levels in each genome. The specific steps are to first calibrate the presence of each gene using genomic data. Then, the abundance of each gene in the genome is calculated using the ratio of sequencing depth to gene copy number. A gene abundance calculation method is used, for example, by normalizing gene expression data to calculate the relative abundance of each gene. During the calculation process, the abundance data can be weighted based on the adaptive characteristics of each genome in different environments or conditions to ensure that environmental variation in the genome is accounted for in the analysis. After the calculation is completed, gene abundance data is obtained, which reflects the differences in biochemical properties between different genomes and provides data support for subsequent genetic population differentiation analysis. Based on the gene abundance data, genetic population differentiation analysis is performed on the genomic biochemical property similarity data. First, the gene abundance data is used to analyze the genetic differences between different genomic populations. The Fst index (fixation index) is used to assess the degree of genetic differentiation between genomic populations. The calculation of the Fst value relies on differences in gene frequency between populations. In specific implementation, the gene abundance data is first subjected to a statistical analysis of gene frequency, converted into gene frequency, and the degree of genetic differentiation between populations is assessed by comparing the differences in gene frequency between different populations. A large Fst value indicates significant genetic variation between populations. Next, methods such as principal component analysis (PCA) or genetic distance matrices (e.g., the neighbor-joining method) are used to further analyze the genetic structure of the populations. This analysis not only reveals genetic heterogeneity between populations but also provides the necessary quantitative evidence for the genetic structure of the pathogenic microbial populations, ultimately generating genetic population differentiation data. Combining this genetic population differentiation data with the gene similarity abundance data allows for analysis of the genetic structure of the pathogenic microbial populations.First, the genetic population differentiation data and gene similarity abundance data are combined for multivariate analysis to construct a genetic structure model of the pathogenic microorganism population. The specific method is to use the structural equation model (SEM) or cluster analysis method to integrate these two types of data to reveal the intrinsic relationship between gene abundance and population genetic differentiation. By analyzing the interaction between genetic differences and gene abundance between populations, the key genes and genomic characteristics that affect the genetic structure of the population can be identified. At the same time, combined with the geographical distribution and environmental characteristics of the population, the population structure can be further refined to obtain the genetic structure data of the pathogenic microorganism population. This genetic structure data not only reflects the genetic diversity of the pathogenic microorganism population, but also reveals the connection between different genetic groups, providing a key basis for the subsequent construction of the database and the characteristic analysis of pathogenic microorganisms.
[0104] Preferably, step S323 includes the following steps:
[0105] According to the gene similarity abundance data, the gene abundance similarity distribution matrix is constructed for the genome biochemical attribute similarity data to obtain the gene abundance similarity distribution matrix;
[0106] Calculate the Fst value of the gene abundance similarity distribution matrix to obtain the gene similarity Fst value;
[0107] The genetic population differentiation data were obtained by analyzing the genetic population similarity distribution matrix based on the gene similarity Fst value.
[0108] In an embodiment of the present invention, a gene abundance similarity distribution matrix is constructed based on gene similarity abundance data. First, the abundance data of each gene in the pathogenic microorganism genome is obtained. These abundance data represent the relative expression levels of different genes in different pathogenic microorganism genomes. Then, based on these abundance data, the similarity between genes is calculated. When calculating similarity, methods such as the Pearson correlation coefficient or cosine similarity can be used to measure the similarity between gene abundances. For example, the Pearson correlation coefficient can calculate the linear relationship between the abundance data of two genes in different genomes. The closer the value is to 1, the higher the similarity; while the cosine similarity is evaluated by calculating the angle between the two gene abundance vectors. The closer the value is to 1, the more similar the abundance distribution is. For each pair of genes, the abundance similarity between them is calculated, and finally a gene abundance similarity distribution matrix is obtained. Each element in the matrix represents the degree of similarity of the corresponding gene pair. The Fst value is calculated for the gene abundance similarity distribution matrix. Fst (Fixation Index) is a commonly used indicator for measuring gene frequency differences between populations. In this step, the data in the gene abundance similarity distribution matrix is used to calculate the Fst value to understand the differentiation of gene abundance between different groups. First, the abundance distribution of genes in different groups (such as microbial groups in different environments) is calculated. In this process, the gene abundance value is converted into relative frequency (or normalized abundance) to eliminate the influence of sequencing depth and other deviations. Then, for each gene, its frequency in different groups is calculated and the differences between groups are compared. Specifically, the gene frequency of each group is calculated, and these frequencies are used to calculate the degree of genetic differentiation between groups. The range of Fst values is usually between 0 and 1. The higher the Fst value, the greater the difference in gene frequency between groups, indicating that the genetic differentiation between groups is more significant. By calculating the Fst value of each gene, a quantitative basis can be provided for subsequent genetic population differentiation analysis. Genetic population differentiation analysis is performed on the gene abundance similarity distribution matrix based on the gene similarity Fst value. The purpose of this step is to analyze the genetic differentiation of different groups through the Fst value, and to identify the genetic structure and evolutionary relationship between groups in combination with the gene abundance similarity distribution matrix. First, the Fst value is used to assess the degree of genetic differentiation between populations. Based on the magnitude of the Fst value, genes showing significant genetic differentiation between populations can be identified. Genes with large Fst values are further analyzed to examine the differences in their abundance across populations and interpret them in the context of biological context. Second, cluster analysis or principal component analysis (PCA) is used to combine gene abundance similarity data with Fst values to delineate the genetic structure of populations. By analyzing gene abundance similarity and Fst values, cluster analysis can group genomes according to genetic distance, revealing genetic connections and differences between populations.Ultimately, through genetic population differentiation analysis, genetic population differentiation data can be obtained, which provides a basis for the subsequent construction of genetic maps of pathogenic microbial populations and further database construction.
[0109] Preferably, step S33 includes the following steps:
[0110] Step S331: performing a relationship structure coupling analysis on the pre-relationship linkage data according to the pre-relationship linkage relationship nested structure data and the pathogenic microorganism population genetic structure data to obtain pre-relationship structure coupling data;
[0111] Step S332: performing weighted correlation evaluation on the pre-relationship structure coupling data to obtain weighted correlation data of the pre-relationship structure;
[0112] Step S333: performing encoding processing on the weighted related data of the pre-relation structure to obtain pre-relation structure encoded data;
[0113] Step S334: constructing a relation coding index for the pre-relation structure coding data using a B+ tree algorithm to obtain a linkage relation coding index.
[0114] In an embodiment of the present invention, a relationship structure coupling analysis is performed on the pre-relationship linkage data based on the pre-relationship nested structure data and the pathogenic microorganism population genetic structure data. First, the pre-relationship linkage nested structure data is composed of interaction information between multiple pathogen genomes, which can be represented by gene interaction networks, co-expression networks, etc. The pathogenic microorganism population genetic structure data contains the genetic differences and population differentiation characteristics of different individuals in the population. The purpose of combining these two types of data for coupling analysis is to identify the potential relationship structure between genomes and between populations. In the specific operation, first, based on gene expression similarity or genotype similarity, a gene interaction network of the pathogenic microorganism population is constructed to identify the association relationship between genes. Then, the genome genetic structure data is coupled with these association relationships, and the coupling strength between different genes is quantitatively analyzed using a weighted average method. This step can be performed by calculating the coupling coefficient between genes (for example, by calculating the Pearson correlation coefficient or Spearman rank correlation coefficient between genes), thereby obtaining the coupling strength of each pair of genes, and ultimately forming the pre-relationship structure coupling data. Weighted correlation evaluation is performed based on the pre-relationship structure coupling data. The purpose of weighting the pre-relationship structure coupling data is to further quantify the strength of the relationships between different genes or populations. First, the coupling strength between each pair of genes is quantitatively assessed. Weighted values are calculated based on coupling coefficients (such as the Pearson correlation coefficient or the Spearman coefficient) to represent the degree of association between genes. These weighted correlation data are then statistically processed to generate an evaluation matrix for the relationships between genes or populations. Each element in this matrix represents the weighted correlation between genes or populations. This weighted evaluation allows for more accurate identification of genes or populations with strong inter-population relationships and further screening for genes that play a significant role in population evolution and genetic structure changes. This data will be used in subsequent steps to construct a more precise encoding index and database query system. The weighted pre-relationship structure correlation data is encoded. The core of this step is to convert the weighted correlation evaluation results into a data format that is easy to store and query, facilitating subsequent database operations. First, specific encoding rules are used to map the weighted correlation values between each pair of genes or populations in the pre-relationship structure weighted correlation data into encoded values. For example, integer coding can be used to assign strong correlation relationships to smaller integer values and weak correlation relationships to larger integer values, or vice versa. The specific coding method can be adjusted according to actual needs. The coding process can also be standardized according to a preset standard. For example, the weighted correlation value is normalized so that it falls within a specific numerical range. The final output of this step is the pre-relationship structure coding data, in which the correlation information of each gene pair is converted into a coding value to facilitate subsequent retrieval and data operations. The pre-relationship structure coding data is subjected to a relational coding index construction using a B+ tree algorithm.The B+ tree is a self-balancing tree data structure widely used in database indexing systems. In this step, the characteristics of the B+ tree are used to organize the pre-relationship structure encoded data into an efficient index tree structure to quickly retrieve and access relevant information. First, all the encoded data are sorted according to certain rules (such as lexicographic order) and inserted into the B+ tree. Each node contains a key value (i.e., the encoded value of the pre-relationship) and a pointer to the child node. When it is necessary to find the relationship of a specific gene, the relevant data can be quickly located through the B+ tree structure. Because the B+ tree has a low search complexity (logarithmic time complexity), it can significantly improve the efficiency of subsequent query operations. Ultimately, the constructed B+ tree index provides an efficient storage and retrieval method for the pre-relationship structure encoded data, supporting rapid query and update of large-scale databases.
[0115] Preferably, step S4 includes the following steps:
[0116] Step S41: performing support range query limitation on the linkage relationship code index to obtain support range query limitation data;
[0117] Step S42: performing dynamic search condition matching on the linkage relationship code index according to the support range query limit data to obtain the linkage relationship dynamic search condition;
[0118] Step S43: Optimizing the cache architecture based on the dynamic retrieval conditions of the linkage relationship and the cache database to obtain an optimized linkage relationship cache architecture;
[0119] Step S44: constructing a genome database for the pathogenic microorganism genome sequence using a relational database and a linkage relationship cache optimization architecture to obtain a pathogenic microorganism genome database.
[0120] In an embodiment of the present invention, the linkage relationship encoding index is first restricted to support range queries. The purpose of this step is to determine the valid query range for each relationship encoding value in the database to improve query efficiency and accuracy. First, based on the overall database structure, a supported query range is defined. This range is typically based on pre-defined business requirements or the strength of relationships between genomes. For example, certain genes may be associated with the biochemical properties of specific pathogens or have strong expression activity under certain conditions. Therefore, a threshold can be set to restrict the query relationship encoding values to fall within this range. This step can be implemented through interval queries. Based on the sorting properties of the B+ tree index, algorithms such as binary search can be used to efficiently determine the encoding values that meet the range conditions during the search process. In implementation, the encoding data is first scanned to determine the maximum and minimum relationship values for each gene or gene group. The query range is then further restricted based on the preset threshold, thereby forming data that supports range queries. This process ensures that only relationships relevant to the current requirements are considered during the query, avoiding interference from irrelevant data and optimizing subsequent data retrieval performance. Cache architecture optimization is performed based on the dynamic retrieval conditions of the linkage relationships and the cache database to optimize database access efficiency and reduce repeated calculations and data read operations. First, dynamic retrieval conditions affect the update and usage strategies of cached data. By analyzing the characteristics of each query, certain commonly used genomic information can be pre-cached in memory. For example, the system can cache frequently queried genes or relational data in a high-speed cache database to reduce access to the underlying relational database. The core of this step is to adjust the cache architecture based on query frequency, data access patterns, and dynamic search conditions, enabling the system to quickly respond to user queries. Cache optimization strategies involve analyzing query conditions, identifying hot data, and allocating more cache space to it within the system. Data popularity can be assessed using frequency counts and time windows. Common optimization strategies include the LRU (least recently used) and LFU (least frequently used) algorithms. During cache architecture optimization, the system adjusts the cache size and content based on specific needs to improve cache hit rates and thus enhance system response speed. Based on the linked relational cache optimization architecture obtained in the previous step and combined with the storage advantages of relational databases, a genomic database for pathogenic microorganism genome sequences is constructed. The core of this step is to integrate the cache optimization architecture with relational database storage to build an efficient and scalable database system. First, leveraging the table structure of a relational database, the genome sequence and pre-relational encoding data are stored in separate tables, connected through relationships such as foreign keys. The design of the genome sequence table and relational index table enables efficient association between genetic data and its corresponding relational index data. During the build process, frequently used data in the cache-optimized architecture is first loaded into the in-memory cache to ensure fast response to frequently accessed data.When the database needs to access uncommon genomic data, it directly accesses the relational database to ensure the persistence and integrity of data storage. During implementation, transaction management is used to ensure data consistency, and an indexing mechanism is used to improve query efficiency. For example, the genome sequence data table will be indexed based on the gene ID field to facilitate quick search. The related relational index table will also use efficient indexing methods such as B+ trees or hash tables for optimized storage, ensuring that each gene and its associated relational encoding data can be quickly queried. This database can support dynamic updates and expansions of genomic data in subsequent use, maintaining efficient and scalable queries. Ultimately, the genomic database will combine the advantages of a cache architecture and a relational database to form an efficient and stable pathogenic microorganism genome database, providing support for subsequent genomic research and data mining.
[0121] Preferably, step S43 includes the following steps:
[0122] Step S431: Estimating the cache access extreme value based on the supported range query limit data to obtain the range query cache access extreme value data;
[0123] Step S432: Designing a cache replacement strategy for the cache database based on the range query cache access extreme value data to obtain an access cache replacement strategy;
[0124] Step S433: Optimizing the cache architecture based on the access cache replacement strategy and the range query cache access extreme value data to obtain a linkage relationship cache optimization architecture.
[0125] In an embodiment of the present invention, cache access extremes are first estimated based on the data that supports range queries. The core goal of this step is to analyze data access patterns within the query range and predict which data will be frequently accessed, thereby optimizing cache utilization. In specific implementations, access patterns for genes or genomic regions queried within a specific range are first determined by analyzing previous query history and data access frequency. By analyzing the temporal nature of the data (e.g., using a sliding window algorithm), it is possible to estimate which data will be frequently accessed and thus become cache priority targets. Certain gene sequences appear in multiple queries, resulting in a high access frequency. Based on this access frequency, extreme value estimation is performed to predict which data within a given query range will have an expected extreme access frequency. This step can be performed using statistical methods such as frequency analysis or probability distribution estimation (e.g., normal distribution or Poisson distribution). These statistical techniques can infer the future access trends of each data block, ultimately determining the extreme value of range query cache access. A cache replacement strategy is then designed based on this extreme value of range query cache access. The goal of this step is to ensure that, despite limited cache space, highly frequently accessed genomic data can be stored using an effective cache management algorithm. First, using the cache access extreme value data estimated in step S431, identify which genomic sequences will have a high access frequency in future queries. To optimize cache usage efficiency, a cache replacement strategy needs to be designed. Common strategies include LRU (least recently used), LFU (least frequently used), and MFU (most frequently used). However, considering the dynamic nature of query data and the uneven access frequency, combined with the analysis results of the access extreme value data, a dynamic cache replacement algorithm based on data access estimation can be designed. For example, for genomic data with a high predicted access frequency, a preload strategy can be used to prioritize this data in the cache and set a longer cache storage period. For data with a low predicted access frequency, a shorter cache period can be set, and an LRU or LFU strategy can be used for replacement. This dynamic replacement strategy can adjust cache content according to actual access needs, thereby improving cache hit rates and reducing the system's cache replacement overhead. Cache architecture optimization is performed based on the access cache replacement strategy and range query cache access extreme value data. This step aims to ultimately form an efficient cache architecture by integrating the previously obtained access patterns and cache replacement strategies. In specific implementation, first, based on the data obtained in steps S431 and S432 and combined with the characteristics of the query load, the hierarchical structure of the cache architecture is designed. Assuming that in some cases, the query of genomic data is concentrated in a certain range (for example, the gene sequence of certain specific pathogens), then this data should be stored in the top layer of the cache first. The cache structure can be divided into multiple levels, with the top-level cache storing frequently accessed gene sequence data, and the low-frequency data being stored in the bottom-level cache.The underlying cache can adopt a more traditional replacement strategy, such as LRU. Through this hierarchical cache structure, cache resources can be effectively utilized to maintain system stability and data access efficiency without sacrificing performance. When optimizing the cache architecture, the cache update mechanism also needs to be considered. The extreme value estimation of data access (step S431) provides the necessary basis for adjusting the cache architecture, and the size and content of the cache can be dynamically adjusted. For example, in some cases, certain genomic data will become frequently accessed due to changes in experimental conditions, and the system needs to respond to these changes in a timely manner and adjust the cache content. This can be achieved through a regular cache evaluation mechanism, dynamically adjusting the cache storage strategy according to changes in access frequency, thereby further improving the overall performance of the system. Finally, the optimized cache architecture can not only improve the cache hit rate and reduce the number of database accesses, but also effectively alleviate database pressure and improve query response speed.
[0126] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0127] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a pathogenic microorganism genome database, characterized in that: The following steps are involved: Step S1: Obtaining a pathogenic microorganism genome sequence; performing pathogenic characteristic point analysis on the pathogenic microorganism genome sequence to obtain pathogenic characteristic point data; Step S2: Analyze the physiological and biochemical characteristics of the pathogenic microorganism genome sequence based on the pathogen characteristic point data to obtain the physiological and biochemical characteristic data of the pathogen genome; Perform genome evolution relationship simulation based on the physiological and biochemical characteristics of the pathogen genome to obtain genome evolution relationship data; Performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data, step S2 includes the following steps: Step S21: performing physiological and biochemical characteristic analysis on the pathogenic microorganism genome sequence according to the pathogen characteristic point data to obtain pathogen genome physiological and biochemical characteristic data; Step S22: identifying the environmental weight impact of the pathogen genome physiological and biochemical characteristic data to obtain physiological and biochemical environmental weight impact data; Step S23: performing genome evolution relationship simulation on the pathogenic microorganism genome sequence according to the physiological and biochemical environment weight impact data to obtain genome evolution relationship data. Step S23 includes the following steps: Step S231: performing spatiotemporal dimension effect intensity analysis on the physiological and biochemical environment weight impact data to obtain spatiotemporal dimension effect intensity data; Step S232: performing pressure chemical reaction change analysis on the pathogenic microorganism genome sequence based on the spatiotemporal action intensity data to obtain pathogen spatiotemporal pressure chemical reaction change data; Step S233: Calculating the probability of gene segment variation of the pathogenic microorganism genome sequence based on the pathogen spatiotemporal pressure chemical reaction variation data to obtain gene segment variation probability data; Step S234: performing quantitative prediction of the variation trend of the gene segment variation probability data to obtain quantitative data of the variation trend of the gene segment; Step S235: performing genome evolution relationship simulation on the pathogenic microorganism genome sequence based on the gene segment variation probability data and the gene segment variation trend quantitative data to obtain genome evolution relationship data; Step S24: performing pre-relation linkage identification on the genome evolution relationship data to obtain pre-relation linkage data; Step S3: performing a relationship nesting structure analysis on the pre-relationship linkage data to obtain pre-relationship linkage relationship nesting structure data; constructing a relationship coding index on the pre-relationship linkage data based on the pre-relationship linkage relationship nesting structure data to obtain a linkage relationship coding index. Step S3 includes the following steps: Step S31: performing relationship nesting structure analysis on the preceding relationship linkage data to obtain preceding linkage relationship nesting structure data; Step S32: performing pathogenic microorganism population genetic structure analysis on the pathogen genome physiological and biochemical characteristic data based on the pre-linkage relationship nested structure data to obtain pathogenic microorganism population genetic structure data. Step S32 includes the following steps: Step S321: performing biochemical attribute similarity comparison on the pathogen genome physiological and biochemical characteristic data according to the pre-linkage relationship nested structure data to obtain genome biochemical attribute similarity data; Step S322: performing gene abundance calculation on the genome biochemical attribute similarity data to obtain gene similarity abundance data; Step S323: performing genetic population differentiation analysis on the genomic biochemical attribute similarity data based on the gene similarity abundance data to obtain genetic population differentiation data. Step S323 includes the following steps: According to the gene similarity abundance data, the gene abundance similarity distribution matrix is constructed for the genome biochemical attribute similarity data to obtain the gene abundance similarity distribution matrix; Calculate the Fst value of the gene abundance similarity distribution matrix to obtain the gene similarity Fst value; The genetic population differentiation data were obtained by analyzing the genetic population differentiation based on the gene abundance similarity distribution matrix according to the gene similarity Fst value; Step S324: performing pathogenic microorganism population genetic structure analysis based on the genetic population differentiation data and the gene similarity abundance data to obtain pathogenic microorganism population genetic structure data; Step S33: constructing a relationship coding index for the pre-relationship linkage data according to the pre-relationship linkage relationship nested structure data and the pathogenic microorganism population genetic structure data to obtain a linkage relationship coding index. Step S33 includes the following steps: Step S331: performing a relationship structure coupling analysis on the pre-relationship linkage data according to the pre-relationship linkage relationship nested structure data and the pathogenic microorganism population genetic structure data to obtain pre-relationship structure coupling data; Step S332: performing weighted correlation evaluation on the pre-relationship structure coupling data to obtain weighted correlation data of the pre-relationship structure; Step S333: performing encoding processing on the weighted related data of the pre-relation structure to obtain pre-relation structure encoded data; Step S334: constructing a relational coding index for the pre-relational structure coding data using a B+ tree algorithm to obtain a linkage relation coding index; Step S4: Dynamically matching the linkage relationship coding index with search conditions to obtain the linkage relationship dynamic search conditions; optimizing the cache architecture based on the linkage relationship dynamic search conditions to obtain the linkage relationship cache optimization architecture; constructing a genome database for the pathogenic microorganism genome sequence using the linkage relationship cache optimization architecture to obtain the pathogenic microorganism genome database. Step S4 includes the following steps: Step S41: performing support range query limitation on the linkage relationship code index to obtain support range query limitation data; Step S42: performing dynamic search condition matching on the linkage relationship code index according to the support range query limit data to obtain the linkage relationship dynamic search condition; Step S43: Optimizing the cache architecture based on the dynamic retrieval conditions of the linkage relationship and the cache database to obtain an optimized linkage relationship cache architecture; Step S44: constructing a genome database for the pathogenic microorganism genome sequence using a relational database and a linkage relationship cache optimization architecture to obtain a pathogenic microorganism genome database.
2. The method for constructing a pathogenic microorganism genome database according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining the genome sequence of the pathogenic microorganism; Step S12: filtering low-quality sequences of the pathogenic microorganism genome sequence to obtain a pathogenic microorganism genome filtered sequence; Step S13: functional classification identification is performed on the pathogenic microorganism genome filtered sequence to obtain the microorganism genome functional classification sequence; Step S14: performing pathogenic characteristic point analysis on the functional classification sequence of the microbial genome to obtain pathogenic characteristic point data.
3. The method for constructing a pathogenic microorganism genome database according to claim 1, wherein: Step S43 includes the following steps: Step S431: Estimating the cache access extreme value based on the supported range query limit data to obtain the range query cache access extreme value data; Step S432: Designing a cache replacement strategy for the cache database based on the range query cache access extreme value data to obtain an access cache replacement strategy; Step S433: Optimizing the cache architecture based on the access cache replacement strategy and the range query cache access extreme value data to obtain a linkage relationship cache optimization architecture.
Citation Information
Patent Citations
Method, device and equipment for constructing multi-modal knowledge retrieval system fused with large model
CN118394978A
Microflora release selection method and system
CN118800342A
Integrated Desktop Software for Management of Virus Data
US20110022973A1
Deep learning network for evolutionary conservation
US20230207054A1