A preferred method and system for functional genetic variant sites
By acquiring chromatin accessibility distribution information across the entire genome, and combining regulatory maps and genomic information, the weight set was optimized, solving the problem of low efficiency in the selection of gene variant sites in existing technologies. This enabled efficient localization of functional gene variant sites and provided new targets for drug development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSTITUTE OF BASIC MEDICAL SCIENCES CHINESE ACADEMY OF MEDICAL SCIENCES
- Filing Date
- 2023-09-04
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for optimizing gene mutation sites are inefficient and cannot effectively locate functional sites, resulting in a lack of new targets for drug development.
By acquiring chromatin accessibility distribution information across the entire genome, performing transformation processing using convolutional blocks, and combining regulatory maps and genomic information, an index calculation, validation, and weight adjustment model is employed to optimize the initial weight value set and output the preferred functional gene variant site information.
It improved the optimization efficiency of functional gene variant sites to 60%, providing reliable new targets for drug development for human diseases.
Smart Images

Figure CN117174169B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene locus selection, and in particular to a method and system for selecting functional gene variant sites. Background Technology
[0002] Most gene variation sites are located in non-coding regions and often function by regulating the expression levels of nearby genes. However, transcriptional regulation exhibits significant tissue-cell specificity, which poses a challenge to identifying the "functional sites" that play a regulatory role.
[0003] Existing site functional optimization methods include expression quantitative trait loci (eQTLs) and chromatin interaction. However, these methods are based on bulk tissue rather than specific cell types, ignoring the significant cell type heterogeneity of site functions, resulting in consistently unsatisfactory optimization efficiency, which has long been below 40%.
[0004] The development of single-cell and functional omics technologies has greatly enriched the measurable levels and scales of biology, providing new insights for optimizing functional sites. On the one hand, functional omics technologies, such as single-cell scATAC-seq, can depict the distribution of chromatin accessibility across the entire genome, thereby inferring regions with regulatory functions in the genome and the regulatory heterogeneity among different cell types. On the other hand, combined functional omics analysis technologies, such as single-cell enhancer-gene association regulation, can depict the association between gene expression patterns and variable sites, thereby inferring functional sites affected by regulated gene expression.
[0005] However, it is crucial to effectively locate functional sites in the human genome, improve optimization efficiency, and provide reliable new targets for drug development for human diseases. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for selecting functional gene variant sites, which can improve the selection efficiency of functional gene variant sites.
[0007] To achieve the above objectives, the present invention provides the following solution:
[0008] A preferred method for identifying a functional gene variant site, the method comprising:
[0009] Obtain accessibility distribution information of chromatin across the entire genome; the accessibility distribution information includes: base sequence arrangement information within the accessibility region and genomic location information of each base;
[0010] The accessibility distribution information is transformed using convolutional blocks to obtain accessibility feature values;
[0011] Genomic information across the entire genome is determined based on regulatory maps; each coordinate in the regulatory map corresponds one-to-one with each genome within the entire genome; the genomic information includes: whether it is a regulatory region, histone modification values, transcription factor binding affinity differences, and transcriptional activity values;
[0012] An initial weight value set is determined based on the accessibility feature value and the genomic information; the initial weight value set is the set of initial weight values corresponding to each of the accessibility feature value and the genomic information.
[0013] The accessibility feature value, the genomic information, and the initial weight set are input into the site optimization model, and the optimization information is output. The optimization information includes: optimized functional gene variant site information and regulated susceptibility gene information. The optimized functional gene variant site information includes: the variant bases that meet the set threshold conditions and the corresponding site coordinates. The regulated susceptibility gene information is based on the regulatory map, which determines the genes regulated by the optimized functional gene variant sites. The susceptibility gene information includes: the gene's common name and the corresponding start and end positions in the genome.
[0014] The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model.
[0015] The index calculation model is used to determine the index value according to the accessibility feature value, the genome information and the initial weight value set, based on a set genome region length and a set step size.
[0016] The validation model is used to determine the predictive power of the indicator value based on the gold standard evaluation indicators; the evaluation indicators include: accuracy, specificity, and sensitivity;
[0017] The weight adjustment model is used for:
[0018] Determine whether to adjust the weights based on the evaluation index values in the verification model;
[0019] If the evaluation index value is not within the set threshold condition range, then each initial weight value in the initial weight value set is optimized and adjusted in a set order until the evaluation index value meets the set threshold condition.
[0020] If the evaluation index value is within the set threshold condition, the functional gene variant site information corresponding to the evaluation index value that reaches the threshold condition will be sent to the output model.
[0021] The output model is used to output the information on the functional gene variant sites and the regulated susceptibility genes to obtain optimal information.
[0022] Optionally, convolutional blocks are used to transform the accessibility distribution information to obtain accessibility feature values, specifically including:
[0023] The matrix information is determined based on the base sequence arrangement information within the accessibility region and the genomic location information of each base.
[0024] Based on the matrix information and the set chromatin accessibility value, determine the input matrix;
[0025] The input matrix is normalized and max-pooled using convolutional blocks to obtain accessibility feature values.
[0026] Optionally, an initial set of weight values is determined based on the accessibility feature values and the genomic information, specifically including:
[0027] The accessibility feature value and the genomic information are fused to obtain the fused modality value;
[0028] Based on statistical distribution, the initial weight value set is obtained by normalizing the fused modal values using a regular normalization method.
[0029] A preferred system for functional gene variant sites, the system being used to implement the preferred method for functional gene variant sites described above, the system comprising:
[0030] The acquisition module is used to acquire chromatin accessibility distribution information across the entire genome; the accessibility distribution information includes: base sequence arrangement information within the accessibility region and genomic location information of each base;
[0031] The transformation processing module is used to transform the accessibility distribution information using convolutional blocks to obtain accessibility feature values;
[0032] An information determination module is used to determine genome information across the entire genome based on a regulatory map; each coordinate in the regulatory map corresponds one-to-one with each genome within the entire genome; the genome information includes: whether it is a regulatory region, histone modification values, transcription factor binding affinity differences, and transcriptional activity values;
[0033] An initial weight value set determination module is used to determine an initial weight value set based on the accessibility feature value and the genomic information; the initial weight value set is a set of initial weight values corresponding to each of the accessibility feature value and the genomic information;
[0034] The output module is used to input the accessibility feature value, the genomic information, and the initial weight value set into the site optimization model, and output optimization information; the optimization information includes: optimized functional gene variant site information and regulated susceptibility gene information; the optimized functional gene variant site information includes: site variant bases that meet the set threshold conditions and the corresponding site coordinates; the regulated susceptibility gene information is based on the regulatory map, determining the genes regulated by the optimized functional gene variant sites; the susceptibility gene information includes: the gene's common name and the corresponding genomic start and end positions;
[0035] The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model.
[0036] The index calculation model is used to determine the index value according to the accessibility feature value, the genome information and the initial weight value set, based on a set genome region length and a set step size.
[0037] The validation model is used to determine the predictive power of the indicator value based on the gold standard evaluation indicators; the evaluation indicators include: accuracy, specificity, and sensitivity;
[0038] The weight adjustment model is used for:
[0039] Determine whether to adjust the weights based on the evaluation index values in the verification model;
[0040] If the evaluation index value is not within the set threshold condition range, then each initial weight value in the initial weight value set is optimized and adjusted in a set order until the evaluation index value meets the set threshold condition.
[0041] If the evaluation index value is within the set threshold condition, the functional gene variant site information corresponding to the evaluation index value that reaches the threshold condition will be sent to the output model.
[0042] The output model is used to output the information on the functional gene variant sites and the regulated susceptibility genes to obtain optimal information.
[0043] Optionally, the conversion processing module includes:
[0044] The matrix information determination submodule is used to determine the matrix information based on the base sequence arrangement information within the accessibility region and the genomic location information of each base.
[0045] The input matrix determination submodule is used to determine the input matrix based on the matrix information and the set chromatin accessibility value;
[0046] The accessibility feature value determination submodule is used to perform standardization and max pooling on the input matrix using convolutional blocks to obtain accessibility feature values.
[0047] Optionally, the initial weight value set determination module includes:
[0048] The fusion submodule is used to fuse the accessibility feature value and the genomic information to obtain the fused modality value;
[0049] The calculation submodule is used to perform normalization calculations based on statistical distribution and the fused modal values using a regular normalization method to obtain the initial weight value set.
[0050] An electronic device includes a memory and a processor, the memory storing a computer program, and the processor running the computer program to cause the electronic device to perform the preferred method for functional gene variant sites described above.
[0051] A storage medium storing a computer program that, when executed by a processor, implements the preferred method for the functional gene variant sites described above.
[0052] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0053] This invention provides a method and system for optimizing functional gene variant sites. The method involves: acquiring chromatin accessibility distribution information across the entire genome; performing convolutional block transformation based on the accessibility distribution information to obtain accessibility feature values; determining genomic information across the entire genome based on regulatory maps; determining an initial weight value set based on the accessibility feature values and genomic information; inputting the accessibility feature values, genomic information, and initial weight value set into a site optimization model; and outputting optimized functional gene variant site information and the susceptibility gene information it regulates. The optimized functional gene variant site information includes: site variant bases that meet set threshold conditions and their corresponding site coordinates, thereby improving the optimization efficiency of functional gene variant sites. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1A flowchart illustrating a preferred method for a functional gene variant site provided in an embodiment of the present invention;
[0056] Figure 2 This is a structural diagram of a preferred system for functional gene variant sites provided in an embodiment of the present invention.
[0057] Symbol explanation:
[0058] Acquisition Module-1, Transformation Processing Module-2, Information Determination Module-3, Initial Weight Value Set Determination Module-4, Output Module-5. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The purpose of this invention is to provide a method and system for selecting functional gene variant sites, which can improve the selection efficiency of functional gene variant sites.
[0061] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0062] Example 1
[0063] like Figure 1 As shown, this embodiment of the invention provides a preferred method for identifying functional gene variant sites, the method comprising:
[0064] Step 100: Obtain chromatin accessibility distribution information across the entire genome.
[0065] The accessibility distribution information includes: the base sequence arrangement information within the accessibility region and the genomic location information of each base.
[0066] Step 200: Use convolutional blocks to transform the accessibility distribution information to obtain accessibility feature values.
[0067] Specifically, convolutional blocks are used to transform the accessibility distribution information to obtain accessibility feature values, including:
[0068] The matrix information is determined based on the base sequence arrangement information within the accessibility region and the genomic location information of each base.
[0069] The input matrix is determined based on the matrix information and the set chromatin accessibility value.
[0070] The input matrix is normalized and max-pooled using convolutional blocks to obtain accessibility feature values.
[0071] Step 300: Determine genome-wide genomic information based on regulatory maps.
[0072] In this regulatory map, each coordinate corresponds one-to-one with each genome within the whole genome; the genomic information includes: whether it is a regulatory region, histone modification value, transcription factor binding affinity difference value, and transcriptional activity value.
[0073] Step 400: Determine the initial set of weight values based on accessibility feature values and genomic information.
[0074] The initial weight value set is the set of initial weight values corresponding to the accessibility feature value and the genomic information, respectively.
[0075] An initial set of weight values is determined based on the accessibility feature values and the genomic information, specifically including:
[0076] The fused modality value is obtained by fusing accessibility feature values and genomic information.
[0077] Based on statistical distribution, the initial weight value set is obtained by normalizing the fused modal values using a regular normalization method.
[0078] Step 500: Input the accessibility feature value, genomic information and initial weight value set into the site optimization model, and output the optimization information; the optimization information includes: the optimized functional gene variant site information and the regulated susceptibility gene information.
[0079] The preferred functional gene mutation site information includes: the mutated bases at the site that meet the set threshold conditions and the corresponding site coordinates; the regulated susceptibility gene information is based on the regulatory map to determine the genes regulated by the preferred functional gene mutation sites; the susceptibility gene information includes: the gene's common name and the corresponding start and end positions in the genome.
[0080] The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model.
[0081] The index calculation model is used to determine index values based on accessibility feature values, genomic information, and an initial set of weight values, according to a set genomic region length and a set step size.
[0082] The validation model is used to determine the predictive power of the indicator value based on the gold standard evaluation indicators, which include accuracy, specificity, and sensitivity. The weight adjustment model is used to determine whether to adjust the weights based on the evaluation indicator values in the validation model. If the evaluation indicator value is not within the set threshold condition range, the initial weight values in the initial weight value set are optimized and adjusted in a set order until the evaluation indicator value meets the set threshold condition. If the evaluation indicator value is within the set threshold condition range, the functional gene variant site information corresponding to the evaluation indicator value that meets the threshold condition is sent to the output model.
[0083] The output model is used to output information on functional gene variant sites and regulated susceptibility genes to obtain optimal information.
[0084] Specifically, in practical applications, the use of site-optimized models to assess the functionality of gene variant sites involves five levels of functional omics data, including:
[0085] Level 1: Genome-wide chromatin accessibility distribution depicted by single-cell scATAC-seq.
[0086] Level 2: Enhancer-Gene regulatory map obtained from single-cell scATAC+RNA cross-omics joint analysis.
[0087] Level 3: Distribution of histone modifications across the entire genome.
[0088] Level 4: Differential binding affinity of transcription factors across the entire genome.
[0089] Level 5: Transcriptional activity of gene variant sites generated by large-scale parallel sequencing experiments.
[0090] The algorithm employs convolution and max pooling to extract effective features at level one. The input includes the arrangement of four base sequences (A, T, C, G) within each chromatin accessibility region, mapping the base sequences and genomic locations to a 4×N 0 / 1 matrix. Based on this matrix, the chromatin accessibility (a value between 0 and 1, with higher values indicating higher accessibility) for each genomic location is used to generate the final input matrix. The generated matrix is then transformed using 8×8 convolutional blocks, followed by standardized max pooling to extract level one features. The level one feature values are ultimately presented as accessibility feature values corresponding to each genomic coordinate across the entire genome.
[0091] Level 2 data includes data obtained from Enhancer-Gene regulatory mapping analysis, covering the entire genome, with each genome coordinate corresponding to a regulatory region (yes or no), and the regulated gene (gene A, B, etc.).
[0092] Level 3 data covers the entire genome, with histone modification values (values between 0 and 1, where higher values indicate a higher degree of modification) corresponding to each genome coordinate.
[0093] Level 4 data covers the entire genome. If a genomic coordinate corresponds to two or more alleles, the corresponding transcription factor binding affinity difference is calculated. If a genomic coordinate corresponds to one allele, the transcription factor binding affinity difference is 0.
[0094] Level 5 data represents the transcriptional activity of gene variant sites generated by large-scale parallel sequencing experiments across the entire genome. If the locus coordinates do not correspond to any large-scale parallel sequencing results, the transcriptional activity is 0.
[0095] Subsequently, modal fusion was performed on the data from Level 1 and Levels 2 through 5. To ensure the contribution of each level, regularization was performed based on the statistical distribution range of the modal values, and the initial weight values of each level were finally obtained (Level 1: 0.35, Level 2: 0.7, Level 3: 0.9, Level 4: 0.5, Level 5: 0.8).
[0096] By inputting the feature values of level one, the data of levels two to five, and the corresponding initial weight values, using a 200Mb genome region length as the calculation window and a 100kb step size for the genome region, functional scoring can be performed on the variant sites in each region across the entire genome.
[0097] After obtaining the functional scores, to select the truly functional loci within each locus, all variable loci within the locus need to be ranked according to their functional scores. If the region has been reported in existing genome-wide association studies (GWAS), the number of functional loci to be selected is determined based on the number of independent GWAS signals reported in the studies. After ranking according to functional scores, the loci with the highest scores are selected first; for the second-highest scoring loci, the ratio to the highest score is calculated, and if it is higher than 70%, it is also selected; otherwise, it is excluded. This process ultimately generates the functional loci for each genomic region. Using level-two data, the genes regulated by these functional loci are identified as susceptibility genes for that region.
[0098] The model using initial weights will be validated against the gold standard functional locus dataset of 525 genomic regions established by OpenTarget Genetics. The validation process involves inputting data from the five levels mentioned above for the corresponding 525 genomic regions, applying the initial weights for each level, predicting the functional loci and susceptibility genes corresponding to the 525 loci, and finally comparing the results with the gold standard. The model is considered successfully trained only if its prediction accuracy, specificity, and sensitivity all reach above 90%.
[0099] If the model does not meet the above criteria, it is optimized through two steps. Step one: By sequentially removing information from one level, the importance of the five levels is ranked based on the change in model performance after removing that level's information. Step two: Starting with the most important level, the weight values of that level are modified in increments of 0.1, beginning with positive modifications (increasing weight values). If increasing the weight improves model performance, it is continued until no further increases are made. If increasing the weight decreases model performance, the modification direction is reversed, and the weight is decreased. Based on this strategy, the weight values of each level are optimized until the model meets the above criteria.
[0100] Finally, the optimized model will perform functional scoring on variant sites in various regions across the entire genome, select the functional sites within the regions, and identify the regulated susceptibility genes.
[0101] Example 2
[0102] like Figure 2 As shown, this embodiment of the invention provides a system for optimizing functional gene variant sites. This system is used to implement the method for optimizing functional gene variant sites in Embodiment 1. The system includes: an acquisition module 1, a conversion processing module 2, an information determination module 3, an initial weight value set determination module 4, and an output module 5.
[0103] Module 1 is used to acquire chromatin accessibility distribution information across the entire genome; the accessibility distribution information includes: base sequence arrangement information within the accessibility region and genomic location information of each base.
[0104] The transformation processing module 2 is used to transform the accessibility distribution information using convolutional blocks to obtain accessibility feature values.
[0105] Information determination module 3 is used to determine genome information across the entire genome based on the regulatory map; each coordinate in the regulatory map corresponds one-to-one with each genome within the entire genome; the genome information includes: whether it is a regulatory region, histone modification values, differences in transcription factor binding affinity, and transcriptional activity values.
[0106] The initial weight value set determination module 4 is used to determine the initial weight value set based on the accessibility feature value and the genomic information; the initial weight value set is the set of initial weight values corresponding to the accessibility feature value and the genomic information respectively.
[0107] Output module 5 is used to input accessibility feature values, genomic information, and initial weight value set into the site optimization model and output optimization information. The optimization information includes: optimized functional gene variant site information and regulated susceptibility gene information. The optimized functional gene variant site information includes: site variant bases that meet the set threshold conditions and the corresponding site coordinates. The regulated susceptibility gene information is based on the regulatory map to determine the genes regulated by the optimized functional gene variant sites. The susceptibility gene information includes: the gene's common name and the corresponding start and end positions in the genome.
[0108] The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model.
[0109] The index calculation model is used to determine index values based on accessibility feature values, genomic information, and an initial set of weight values, according to a set genomic region length and a set step size.
[0110] The validation model is used to determine the predictive power of the indicator values based on the gold standard evaluation metrics; the evaluation metrics include: accuracy, specificity, and sensitivity.
[0111] The weight adjustment model is used to determine whether to adjust the weights based on the evaluation index values in the validation model. If the evaluation index values are not within the set threshold conditions, the initial weight values in the initial weight value set are optimized and adjusted in a set order until the evaluation index values meet the set threshold conditions. If the evaluation index values are within the set threshold conditions, the functional gene variant site information corresponding to the evaluation index values that meet the threshold conditions is sent to the output model.
[0112] The output model is used to output information on functional gene variant sites and regulated susceptibility genes to obtain optimal information.
[0113] The transformation processing module 2 includes: a matrix information determination submodule, an input matrix determination submodule, and an accessibility feature value determination submodule.
[0114] The matrix information determination submodule is used to determine the matrix information based on the base sequence arrangement information within the accessibility region and the genomic location information of each base.
[0115] The input matrix determination submodule is used to determine the input matrix based on the matrix information and the set chromatin accessibility value.
[0116] The accessibility feature value determination submodule is used to standardize and max pool the input matrix using convolutional blocks to obtain accessibility feature values.
[0117] In one embodiment, the initial weight value set determination module 4 includes a fusion submodule and a calculation submodule.
[0118] The fusion submodule is used to perform fusion based on accessibility feature values and genomic information to obtain fused modality values.
[0119] The calculation submodule is used to perform normalization calculations based on statistical distribution and the fused modal values using a regular normalization method to obtain the initial weight value set.
[0120] Example 3
[0121] This invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor runs the computer program to enable the electronic device to perform the preferred method for functional gene variant sites in Embodiment 1.
[0122] In one embodiment, the present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the preferred method for functional gene variant sites in Embodiment 1.
[0123] This invention establishes a method and system for optimizing functional gene variant sites based on the single-cell chromatin accessibility landscape and Enhancer-Gene regulatory map. It can effectively locate functional sites in the human genome with an optimization efficiency of 60%, and can provide reliable new targets for drug development for human diseases.
[0124] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0125] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for selecting functional gene variant sites, characterized in that, The method includes: Obtain accessibility distribution information of chromatin across the entire genome; the accessibility distribution information includes: base sequence arrangement information within the accessibility region and genomic location information of each base; The accessibility distribution information is transformed using convolutional blocks to obtain accessibility feature values; Genomic information across the entire genome is determined based on regulatory maps; each coordinate in the regulatory map corresponds one-to-one with each genome within the entire genome; the genomic information includes: whether it is a regulatory region, histone modification values, transcription factor binding affinity differences, and transcriptional activity values; An initial weight value set is determined based on the accessibility feature value and the genomic information; the initial weight value set is the set of initial weight values corresponding to each of the accessibility feature value and the genomic information. The accessibility feature value, the genomic information, and the initial weight value set are input into the site selection model, and selection information is output. The selection information includes: information on the selected functional gene variant sites and information on the regulated susceptible genes. The selected functional gene variant site information includes: the variant bases that meet the set threshold conditions and the corresponding site coordinates. The regulated susceptible gene information is based on the regulatory map, which determines the genes regulated by the selected functional gene variant sites. The susceptible gene information includes: the gene's common name and the corresponding start and end positions in the genome. The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model. The index calculation model is used to determine the index value according to the accessibility feature value, the genome information and the initial weight value set, based on a set genome region length and a set step size. The validation model is used to determine the predictive power of the indicator value based on the gold standard evaluation indicators; the evaluation indicators include: accuracy, specificity, and sensitivity; The weight adjustment model is used for: Determine whether to adjust the weights based on the evaluation index values in the verification model; If the evaluation index value is not within the set threshold condition range, then each initial weight value in the initial weight value set is optimized and adjusted in a set order until the evaluation index value meets the set threshold condition. If the evaluation index value is within the set threshold condition range, the functional gene variation site information corresponding to the evaluation index value that reaches the threshold condition will be sent to the output model. The output model is used to output the information on the functional gene variant sites and the information on the regulated susceptibility genes to obtain selection information; The accessibility distribution information is transformed using convolutional blocks to obtain accessibility feature values, specifically including: The matrix information is determined based on the base sequence arrangement information within the accessibility region and the genomic location information of each base. Based on the matrix information and the set chromatin accessibility value, determine the input matrix; The input matrix is standardized and max-pooled using convolutional blocks to obtain accessibility feature values. An initial set of weight values is determined based on the accessibility feature values and the genomic information, specifically including: The accessibility feature value and the genomic information are fused to obtain the fused modality value; Based on statistical distribution, the initial weight value set is obtained by normalizing the fused modal values using a regular normalization method.
2. A selection system for functional gene variant sites, characterized in that, The system is used to implement the method for selecting functional gene variant sites as described in claim 1, and the system comprises: The acquisition module is used to acquire chromatin accessibility distribution information across the entire genome; the accessibility distribution information includes: base sequence arrangement information within the accessibility region and genomic location information of each base; The transformation processing module is used to transform the accessibility distribution information using convolutional blocks to obtain accessibility feature values; An information determination module is used to determine genome information across the entire genome based on a regulatory map; each coordinate in the regulatory map corresponds one-to-one with each genome within the entire genome; the genome information includes: whether it is a regulatory region, histone modification values, transcription factor binding affinity differences, and transcriptional activity values; An initial weight value set determination module is used to determine an initial weight value set based on the accessibility feature value and the genomic information; the initial weight value set is a set of initial weight values corresponding to each of the accessibility feature value and the genomic information; The output module is used to input the accessibility feature value, the genomic information, and the initial weight value set into the site selection model, and output selection information. The selection information includes: selected functional gene variant site information and regulated susceptibility gene information. The selected functional gene variant site information includes: the variant bases that meet the set threshold conditions and the corresponding site coordinates. The regulated susceptibility gene information is based on the regulatory map, determining the genes regulated by the selected functional gene variant sites. The susceptibility gene information includes: the gene's common name and the corresponding start and end positions in the genome. The site selection model includes an index calculation model, a validation model, a weight adjustment model, and an output model. The index calculation model is used to determine the index value according to the accessibility feature value, the genome information and the initial weight value set, based on a set genome region length and a set step size. The validation model is used to determine the predictive power of the indicator value based on the gold standard evaluation indicators; the evaluation indicators include: accuracy, specificity, and sensitivity; The weight adjustment model is used for: Determine whether to adjust the weights based on the evaluation index values in the verification model; If the evaluation index value is not within the set threshold condition range, then each initial weight value in the initial weight value set is optimized and adjusted in a set order until the evaluation index value meets the set threshold condition. If the evaluation index value is within the set threshold condition range, the functional gene variation site information corresponding to the evaluation index value that reaches the threshold condition will be sent to the output model. The output model is used to output the information on the functional gene variant sites and the information on the regulated susceptibility genes to obtain selection information; The conversion processing module includes: The matrix information determination submodule is used to determine the matrix information based on the base sequence arrangement information within the accessibility region and the genomic location information of each base. The input matrix determination submodule is used to determine the input matrix based on the matrix information and the set chromatin accessibility value; The accessibility feature value determination submodule is used to perform standardization and max pooling on the input matrix using convolutional blocks to obtain accessibility feature values; The initial weight value set determination module includes: The fusion submodule is used to fuse the accessibility feature value and the genomic information to obtain the fused modality value; The calculation submodule is used to perform normalization calculations based on statistical distribution and the fused modal values using a regular normalization method to obtain the initial weight value set.
3. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store a computer program, and the processor running the computer program to cause the electronic device to perform the method for selecting functional gene variant sites as described in claim 1.
4. A storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method for selecting functional gene variant sites as described in claim 1.
Citation Information
Patent Citations
Gene variation site screening method and system
CN111091867A
Corn phenotype prediction method and system
CN113593635A