HIV gene regulatory network construction method based on Boolean network model and application of HIV gene regulatory network construction method
The HIV gene regulation network is constructed based on the Boolean network model, and the experimental cost and difficulty in replication of transcriptional regulation mechanism research in complex systems is solved, and the effective construction and optimization of the HIV gene regulation network is realized, revealing the gene expression pattern and providing new therapeutic targets.
Patent Information
- Application Number
- CN202510037142.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-10
AI Technical Summary
At this stage, the research on transcriptional regulation mechanisms in complex systems has the problem of high experimental costs and difficulty in replication, and it is difficult to effectively build a gene regulation network.
Using a Boolean network model method, by screening HIV-related genes and proteins from Pubmed, directed network maps are constructed and optimized, and converted to Boolean model. Cluster analysis is used to calculate the stable state of the network, identify the expression patterns of genes among different clusters, and evaluate the regulatory effect of genes.
The effective construction and optimization of the HIV gene regulation network was achieved, revealing the possible expression patterns of genes in HIV infection, providing new HIV treatment targets, and optimizing antiviral treatment strategies.
Smart Images

Figure SMS_5 
Figure SMS_6 
Figure FDA0005235846930000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of HIV research, and particularly relates to a method for constructing an HIV gene regulatory network based on a Boolean network model and its application. Background Art
[0002] Acquired Immune Deficiency Syndrome (AIDS) is a chronic progressive infectious disease caused by the Human Immunodeficiency Virus (HIV). Currently, it remains a major public health problem worldwide, affecting not only physical health but also social relationships, mental health, quality of life, and the economy. It is estimated that 38.4 million people worldwide are infected with the HIV virus; since the beginning of the epidemic, approximately 84.2 million people have been infected and 40.1 million have died. As of the end of 2022, 1.223 million cases of HIV-infected individuals were reported alive in the country, including 689,000 HIV-infected individuals and 534,000 AIDS patients. HIV is mainly transmitted through blood, sexual contact, and mother-to-child transmission. Once infected, HIV targets the most important CD4+ T lymphocytes (cluster of differentiation 4+ T-lymphocyte) in the body's immune system, destroying a large number of these cells and causing the body to lose its immune function. HIV infection can lead to many serious health problems and has a wide impact on human health. Its weakening of the immune system increases the risk of contracting other diseases, such as pneumonia, tuberculosis, cardiovascular diseases, etc. In addition, HIV can also cause various symptoms and complications, including persistent fever, weight loss, diarrhea, swollen lymph nodes, and nervous system damage.
[0003] Gene regulatory networks are currently at the forefront of precision biology, which can help researchers better understand how genes and regulatory elements are regulated to control cellular gene expression, providing more promising molecular mechanisms for biological research. At present, the transcriptional regulatory mechanism in complex systems remains a research challenge, mainly because the experiments for sequencing protein-DNA interactions and their roles in regulation are costly and difficult to replicate. Therefore, using the method of prediction models instead of biological experiments is one of the effective methods. For example, the inference of gene regulatory networks (GRNs). The concept of gene regulatory networks can be traced back to the regulatory research on the lactose operon in the 1960s. After decades of development, a very large number of types of methods have been used to construct gene regulatory networks, including Boolean network models, Bayesian network models, ordinary differential equation models, etc. By comparing the genomes between Homo sapiens and yeast, it can be concluded that the complexity of life does not come from the number of genes, but the nature and dynamics of the interactions between genes - gene regulatory networks, which play a central role in understanding the mechanisms of gene expression regulation, complex diseases, and cellular heterogeneity. Researchers can use gene regulatory networks to understand biological processes from a global perspective, which can help reveal the molecular mechanisms of disease development, identify key regulatory factors and involved signaling pathways, and solve many problems related to major diseases such as cancer. In recent years, scholars at home and abroad have reconstructed gene regulatory networks through computer and biological knowledge, simulated the dynamic behavior of gene regulatory networks, and revealed the mechanisms, patterns, and the structure and function of regulatory networks. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for constructing an HIV gene regulatory network based on a Boolean network model and its application.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] A method for constructing an HIV gene regulatory network based on a Boolean network model, including Step 1: Screen genes and proteins related to HIV and the regulatory relationships between genes and genes, genes and proteins, and proteins and proteins from Pubmed, construct a directed network graph based on this, and optimize the network; subsequently, define a Boolean polynomial for each node and convert the network into a Boolean model.
[0007] In Step 1:
[0008] The keywords for screening literature from Pubmed were: (HIV OR AIDS) AND (promote OR activate OR induce OR stimulate OR response OR recruit OR enrich OR inhibit OR suppress OR degrade OR block) AND (CD4 cell);
[0009] To construct a directed network diagram, Cytoscape or Gephi software was used. Optimization involved removing isolated nodes and low-correlation nodes to determine the nodes to be included in the model. Nodes represent genes or proteins, and edges represent their regulatory relationships;
[0010] Boolean polynomials were defined for each node using the logical operations OR, AND, and NOT or the modulo 2 arithmetic operations addition and multiplication.
[0011] In the Boolean model, each gene or protein node was set as a Boolean variable, and Boolean rules were used to describe the state change of each node based on the states of its upstream and downstream genes.
[0012] The corresponding relationship of the Boolean polynomials in Step 1 was given by the following formula:
[0013]
[0014] where, x i : represents the i-th Boolean variable, x j : represents the j-th Boolean variable.
[0015] The above method for constructing an HIV gene regulatory network further includes Step 2: Based on the Boolean model, using the clustering analysis method, Maple software was used to calculate the stable state of the network and identify the expression patterns of genes among different clusters.
[0016] In Step 2: The SAS nearest centroid sorting algorithm was used for clustering analysis to determine the steady-state expression pattern of the network; one node corresponds to a Boolean functional expression, and a new Boolean polynomial was defined. When and only when g[i] = 1, x[i] is a fixed point:
[0017] g[i]: = (f[i] + x[i] + 1) mod 2
[0018] Thus, calculating the stable state of the network became calculating the product of g[i] Dividing m f into several subsets for calculation, and finally summarizing the results of each subset, all possible stable states of the network were obtained;
[0019] Among them, f[i] represents the Boolean function of the i-th node, g[i] represents the fixed-point function of the i-th node, and m f represents the product of the fixed-point functions of all nodes in the network, and n represents the total number of nodes in the network;
[0020] By calculating the ratio of the expressions (0 or 1) of each node in the cluster, genes with differential expression between clusters are identified.
[0021] The above method for constructing an HIV gene regulatory network further includes Step 3: Evaluating the regulatory effect of genes based on an improved method for calculating Boolean model polynomials.
[0022] In Step 3: Calculate the regulatory effect of each gene, evaluate the impact of a specific gene on the overall dynamics and steady state of the network, set each gene to 0 or 1, and observe the expression states of other genes; for each node, 4 sets are calculated to determine the control effect of this node, and the iteration stops when these 4 sets no longer increase; the possible control effects of node k are provided by the following 4 sets:
[0023] A(k, 0): The set of genes with a final state of 0 when the target gene k is set to 0;
[0024] A(k, 1): The set of genes with a final state of 1 when the target gene k is set to 1;
[0025] B(k, 0): The set of genes with a final state of 1 when the target gene k is set to 0;
[0026] B(k, 1): The set of genes with a final state of 0 when the target gene k is set to 1.
[0027] The above method for constructing an HIV gene regulatory network further includes Step 4: Downloading the gene expression matrix of CD4+ T cell samples infected with HIV from the public database GEO to verify the reliability of the model.
[0028] In Step 4: If a gene corresponds to several probes, select the probe with the highest variability for subsequent analysis to ensure data accuracy, and the screening conditions for the data set are: ① The sample size is greater than 10; ② It contains a GPL annotation file; ③ The GPL annotation file contains Gene Symbol; The probe with the largest variation is selected with reference to the coefficient of variation, CV = standard deviation / mean; Use Foldchange analysis for verification, and Foldchange calculates the expression quantity ratio of a gene under two different conditions.
[0029] The application of the above method for constructing an HIV gene regulatory network in studying HIV drug treatment targets, where the HIV drug treatment targets are key genes or proteins involved in HIV infection or replication.
[0030] Including the following Step 1
[0031] In view of the existing problems in current HIV research, the inventor established a method for constructing an HIV gene regulatory network based on a Boolean network model by combining gene regulatory networks. The method includes Step 1: Screening genes and proteins related to HIV and the regulatory relationships between genes, between genes and proteins, and between proteins from Pubmed, constructing a directed network diagram based on this, and optimizing the network; subsequently, defining Boolean polynomials for each node and converting the network into a Boolean model. The present invention innovatively established a method for constructing an HIV gene regulatory network, and further optimized the selection of network nodes and the expression of regulatory rules based on the Boolean model, enhancing the biological authenticity of the model. At the same time, by calculating the steady state of the network and combining cluster analysis, the possible expression patterns of different genes during HIV infection were revealed. In addition, an improved algorithm based on the polynomial representation of the Boolean model was proposed to accurately analyze the control effects of genes. Finally, the reliability of the model was verified using the GEO database, potential therapeutic targets with differential expression were identified, and a new idea for the selection of targeted gene combinations in HIV treatment was provided. In summary, the present invention provides an effective method for studying the dynamic information of the network, which will help to reveal the regulatory mechanisms of biological systems, thereby providing a theoretical basis for disease diagnosis, treatment, and drug development. In the research on new HIV drug targets, the present invention also helps to reveal the regulatory relationships during the HIV infection stage and replication process, identify key genes and proteins involved in HIV infection and replication, and contribute to the discovery of new therapeutic targets and the optimization of antiviral treatment strategies. Brief Description of the Drawings
[0032] Figure 1 It is a directed network diagram of HIV regulation. In the figure: Circles represent genes or proteins and their abbreviations, black arrows represent activation effects, and black blunt lines represent inhibitory effects.
[0033] Figure 2 It is a Boolean regulation network diagram of HIV. In the figure: Squares (nodes) represent genes or proteins and their abbreviations, thick green arrows represent input signals, thick brown arrows represent output signals; black lines represent activation effects, and red blunt-ended lines represent inhibitory effects; red diamonds represent AND connections where all relevant nodes need to be jointly activated to trigger downstream events.
[0034] Figure 3 It is a graph of the results of cluster analysis. In the figure: Three colors are used to represent three different categories identified in the cluster analysis, showing the distribution of the cluster results.
[0035] Figure 4 It is a graph of the regulatory effects of genes in the HIV regulatory network.
[0036] Figure 5It is a comparison chart of the simulation control effects of three different datasets in the GEO database. Detailed implementation method
[0037] I. Construction method
[0038] Step 1: Screen for HIV-related genes, proteins, and the regulatory relationships between genes and genes, genes and proteins, and proteins and proteins from Pubmed. Based on this, construct a directed network graph and optimize the network. Subsequently, define a Boolean polynomial for each node and convert the network into a Boolean model to improve the stability and interpretability of the model. Among them,
[0039] The directed network graph is constructed using software such as Cytoscape or Gephi. The optimization is to remove isolated nodes and low-correlation nodes, determine the nodes to be included in the model. The nodes represent genes or proteins, and the edges represent their regulatory relationships. The selection of nodes and edges is optimized to enhance the biological relevance of the model;
[0040] In the Boolean model, each gene or protein node is set as a Boolean variable (0 represents non-expression or inactivity, 1 represents expression or activity). Boolean rules are used to describe the state change of each node based on the states of its upstream and downstream genes.
[0041] The corresponding relationship of the Boolean polynomial is given by the following formula:
[0042]
[0043] where, x i : represents the i-th Boolean variable, x j : represents the j-th Boolean variable.
[0044] When defining the Boolean polynomial for each node, it should be noted that when a gene is regulated by multiple genes and the regulatory effects are different, it is necessary to judge from a biological perspective whether the gene is mainly in a promoting or inhibitory role. According to different regulatory effects, the expression of the Boolean polynomial will be different. Therefore, it is necessary to consult relevant literature to examine the dominant regulatory direction of the gene. When it is uncertain whether the defined functional formula is correct, an error correction table can be drawn for comparison.
[0045] Step 2: Based on the Boolean model, use the clustering analysis method to calculate the stable state of the network using Maple software and identify the expression patterns of genes among different clusters. Among them,
[0046] Use the SAS nearest centroid sorting algorithm for clustering analysis to determine the steady-state expression pattern of the network; one node corresponds to a Boolean functional formula, define a new Boolean polynomial, and x[i] is a fixed point if and only if g[i] = 1:
[0047] g[i] := (f[i] + x[i] + 1) mod 2
[0048] Thus, calculating the steady state of the network has been transformed into calculating the product of g[i] Due to m f The equation contains many variables, resulting in a very slow calculation speed. Therefore, m f is divided into several subsets for calculation. Finally, the results of each subset are summarized to obtain all possible steady states of the network. In the formula, f[i] represents the Boolean function of the i-th node, g[i] represents the fixed-point function of the i-th node, m f represents the product of the fixed-point functions of all nodes in the network, and n represents the total number of nodes in the network.
[0049] By calculating the ratio of the expressions (0 or 1) of each node in the cluster, genes with differential expression between clusters are identified.
[0050] Step 3: Based on the improved polynomial calculation method of the Boolean model, evaluate the regulatory effects of genes. Among them,
[0051] Calculate the regulatory effect of each gene, evaluate the impact of a specific gene on the overall dynamics and steady state of the network. Set each gene to 0 or 1 and observe the expression status of other genes. Through regulatory effect analysis, key genes that have a significant impact on network stability or steady state are identified. For each node, 4 sets are calculated to determine the control effect of that node. When none of these 4 sets increase anymore, the iteration stops. The possible control effects of node k are provided by the following 4 sets:
[0052] A(k, 0): The set of genes whose final state is 0 when the target gene k is set to 0;
[0053] A(k, 1): The set of genes whose final state is 1 when the target gene k is set to 1;
[0054] B(k, 0): The set of genes whose final state is 1 when the target gene k is set to 0;
[0055] B(k, 1): The set of genes whose final state is 0 when the target gene k is set to 1.
[0056] Step 4: Download the gene expression matrix of CD4+ T cell samples infected with HIV from the public database GEO to verify the reliability of the model. Among them,
[0057] If a gene corresponds to several probes, select the probe with the highest variability for subsequent analysis to ensure data accuracy. The probe with the largest variation is selected based on the coefficient of variation, CV = standard deviation / mean. Moreover, the screening conditions for the dataset are as follows: ① The sample size is greater than 10; ② It contains a GPL annotation file; ③ The GPL annotation file contains Gene Symbol.
[0058] Use Foldchange analysis for verification. Foldchange calculates the expression ratio of a gene under two different conditions. In this study, the expression ratio of the gene under two different conditions, the high-expression group and the low-expression group, was analyzed. The grouping criterion was the median of the gene expression level.
[0059] II. Application Example
[0060] Referring to the above construction method, the specific implementation is as follows:
[0061] 2.1 Construct a network diagram and convert it into a Boolean model
[0062] 2.1.1 Screen HIV-related genes and proteins
[0063] Conduct a literature search on Pubmed. The search terms are (HIV OR AIDS) AND (promote OR activate OR induce OR stimulate OR response OR recruit OR enriche OR inhibit OR suppresse OR degrade OR block) AND (CD4 cell).
[0064] As a result, 34,785 relevant literatures were screened from Pubmed. HIV-related gene-gene, gene-protein, and protein-protein interactions, as well as the corresponding genes and proteins, were screened out from the literatures. For the screened regulatory relationship pairs, it is necessary to distinguish the source nodes and target nodes.
[0065] The screening criteria were whether the following keywords were clearly mentioned in the literature: "promotes" (or "activates" or "induces" or "stimulates" or "responses" or "recruits" or "enriches" or "inhibits") or "suppresses" (or "degrades" or "blocks"). According to the above screening criteria, a total of 179 pairs of nodes were screened out. Table 1 summarizes some genes and proteins and their interactions, including source nodes, target nodes, cell types, interaction relationships, and corresponding references. After data cleaning, 125 pairs of nodes were included in the cytoscape software to construct a network diagram ( Figure 1 ). In the network diagram, it was found that 6 pairs of nodes had no connection with the nodes in the network, so these nodes were excluded, and the remaining 119 pairs of nodes were included in the model ( Figure 2 ). In the figure, thick green arrows represent inputs, thick brown arrows represent outputs, black lines represent activation, red lines represent inhibition, and red diamonds represent AND connections. All relevant nodes need to act together to trigger downstream events.
[0066] Table 1 HIV-related genes and proteins and their related interaction relationships
[0067]
[0068] Table 1 lists HIV-related genes and proteins and their roles in the regulatory network, indicating their activation or inhibition effects.
[0069] 2.1.2 Define functional relationships for each node and transform them into a Boolean model
[0070] Use logical operations OR, AND, and NOT, or modulo 2 arithmetic operations addition and multiplication, to complete the calculation of Boolean variables and polynomials. Table 2 lists the Boolean function representations of some nodes. For example:
[0071] ① JAK1 := IL-2 indicates that JAK1 is regulated by IL-2 and has a positive regulatory effect;
[0072] ② CXCR3 := GATA-1 + 1 indicates that CXCR3 is regulated by GATA-1 and has a negative regulatory effect, where "+1" represents taking the inverse, that is, GATA-1 inhibits the expression of CXCR3.
[0073] These Boolean function expressions represent the regulatory relationships between genes or proteins through logical operations and transform them into polynomial functions that can be used for Boolean model analysis.
[0074] Table 2 Polynomial functions of nodes in the Boolean model
[0075]
[0076] Table 2 lists the corresponding polynomial functions of the nodes in the HIV regulatory network, which are used to describe the Boolean relationships of each node.
[0077] 2.2 Calculate the steady state of the network based on the Boolean model and perform cluster analysis
[0078] 2.2.1 Calculate the steady state of the Boolean network
[0079] The steady state of the Boolean network means that under different regulatory conditions, the regulatory network will eventually reach a certain equilibrium state, reflecting the on (1) or off (0) of gene expression. These steady states can predict the expression patterns of genes under specific regulatory conditions. Based on the constructed Boolean model, the steady state of the network is calculated using Maple software. By simplifying functional expressions, performing piecewise operations, and integrating the results, the computational efficiency is improved, and finally 1,717,248 steady states are calculated.
[0080] 2.2.2 Perform cluster analysis on the steady states
[0081] After the steady state calculation is completed, the nearest centroid sorting algorithm in SAS is used to perform cluster analysis in three steps:
[0082] Step 1: Use proc standard to standardize all 91 variables to a mean of 0 and a standard deviation of 1. During the analysis, it is found that some genes such as mTOR, JAK1, JAK3, CCR5, MIG, CXCL10, etc. are always 0 or 1, and these nodes that are always 0 or 1 need to be deleted in the subsequent cluster analysis.
[0083] Step 2: Use proc fastclus for cluster analysis. All the remaining variables are divided into multiple sets through the Euclidean distance algorithm, and this set is named "Clust".
[0084] Step 3: Use the CANDISC and GPLOT programs to generate a graphical representation of the clustering results ( Figure 3 ). The analysis results show that the steady states are clustered into three different groups, consisting of three clusters with sizes of 581,471, 255,273, and 880,604 respectively.
[0085] 2.2.3 Statistical analysis based on the cluster analysis results
[0086] According to the results of cluster analysis, statistical analysis was performed on the expression ratios of each gene. The results showed that there were significant differences in the expression of genes such as PLA2G1B, IL-7, IL-4, GATA3, STAT6, ERK, IFN–α, TRAIL, TRIM5-α, TRIM22, MEK, Raf, Ras, PKR, P53, AKT, mTORC1, PI3K, PTEN, Egr-1, and Foxo3a among different clusters, suggesting that they may be HIV-related Marker genes.
[0087] 2.3 Analysis of gene regulatory effects
[0088] Based on the calculation method of Boolean polynomials, the regulatory effects of genes were analyzed. During the treatment process, targeting a small number of genes is a common strategy, so understanding the control effects of these genes will help select appropriate combinations of target genes.
[0089] 2.3.1 Calculation of regulatory effects
[0090] As Figure 4 shown, each gene was set to 0 or 1 respectively, and 4 sets were calculated to analyze the expression of other genes when the target gene was 0 or 1. This figure shows how each gene interacts with each other in the constructed HIV regulatory network, reflecting the regulatory relationships of genes. The horizontal markers in the figure represent "control genes", while the vertical markers represent "controlled genes".
[0091] 2.3.2 Perturbation analysis
[0092] By perturbing each gene, its effect on the controlled genes was observed. Each column in the table represents a gene, showing the regulatory effect of this gene on other genes after perturbation. The colors of the specific matrix entries are represented as follows:
[0093] ① Dark green box: means that node j and node i are co-expressed;
[0094] ② Green box: means that when node i is "0", node j is not a constant, and when node i is "1", it causes node j to be "1";
[0095] ③ Green square, indicating that when node i is "0", node j is "0", and when node i is "1", node j is not a constant;
[0096] ④ Dark red square, indicating that the co-expression of node j and node i is opposite;
[0097] ⑤ Red box, indicating that when node i is "0", node j is not a constant, and when node i is "1", then j is "0"; ⑥ Light red box, indicating that when node i is "0", then j is "1", and when node i is "1", then j is not a constant;
[0098] ⑦ The yellow square indicates unclear regulations.
[0099] ⑧ The white box indicates that the control gene has no effect on the corresponding column gene.
[0100] 2.3.3 Result Interpretation
[0101] It can be found from the figure that TGFB1, MTOR, FOXP3, JAK1, LCK, IL23A, CXCL9, CXCL10, IFNG, IL2, PLA2G1B, IL15, TLR2, TLR5, and BCL2L1 are positively regulated by many other genes, that is, when the corresponding regulatory gene is "ON", they will also be "ON".
[0102] 2.4 Verifying the Boolean Model Using the GEO Database
[0103] 2.4.1 Downloading and Processing the Datasets
[0104] Download the corresponding datasets: GSE6740, GSE9927, and GSE14280 from the GEO database ( https: / / www.ncbi.nlm.nih.gov / , with the keyword search being (HIV OR AIDS) AND (CD4+ T cells)), including GPL annotation files and gene expression matrices. Subsequently, convert the gene names in the original data into unified gene identifiers for subsequent analysis, and process the missing data, such as filling or deleting the missing values, to ensure the accuracy of data analysis.
[0105] 2.4.2 Foldchange Analysis
[0106] When several probes correspond to one gene, select the probe with the largest variation. For each dataset, divide the expression of the selected regulatory gene into two levels: "high" and "low", and compare the expression changes of other target genes. Figure 5Part of the genes in 3 datasets were selected to simulate the control effect, with the "control genes" (14 genes) marked horizontally and the "controlled genes" (14 genes) marked vertically. Green indicates a control simulation that meets at least one dataset, and yellow indicates a control simulation that does not meet any dataset. The numbers in each square represent the coincidence rate. For example, 2(1) means that there are 2 datasets that show significant differences from the Foldchange analysis, and 1 of these 2 datasets shows consistency with the control simulation. As can be seen from the figure, the control simulation has a good effect, and the results of GEO data analysis support the simulation results of the Boolean model. This indicates that the Boolean model can effectively reflect the regulatory relationships between genes, providing an important basis for subsequent research. Through the above steps, the effectiveness of the Boolean model in simulating gene regulation was verified, providing data support for the study of HIV-related genes.
Claims
1. A method for constructing HIV gene regulatory network based on Boolean network model, characterized in that The method includes the following steps: first, screening HIV-related genes and proteins and the regulatory relationships between genes, genes and proteins, and proteins and proteins from Pubmed, and then constructing a directed network graph based on the relationships, and then optimizing the network; The network is then converted into a Boolean model by defining Boolean polynomials for each node.
2. The method for constructing HIV gene regulatory network according to claim 1, characterized in that In step one: The keywords for screening literature from Pubmed were: (HIV OR AIDS) AND (promote OR activate OR induce OR stimulate OR response OR recruit OR enriche OR inhibit OR suppresseOR degradeOR block) AND (CD4 cell); The directed network graph was constructed using Cytoscape or Gephi software. The optimization was to remove isolated nodes and low-correlation nodes and determine the nodes to be included in the model. The nodes represented genes or proteins, and the edges represented their regulatory relationships. Boolean polynomials are defined for each node using logical operations OR, AND, and NOT or modulo-2 arithmetic operations addition and multiplication. Each gene or protein node in the Boolean model is set as a Boolean variable, and Boolean rules are used to describe the state changes of each node according to the states of its upstream and downstream genes.
3. The method for constructing HIV gene regulatory network according to claim 2, characterized in that The Boolean polynomial correspondence described in step 1 is given by the following formula: Among them, x i : represents the i-th Boolean variable, x j : represents the j-th Boolean variable.
4. The method for constructing HIV gene regulatory network according to claim 1, characterized in that It also includes step 2: based on the Boolean model, the cluster analysis method is used to calculate the stable state of the network using Maple software to identify the expression pattern of genes among different clusters.
5. The method for constructing HIV gene regulatory network according to claim 4, characterized in that In step 2: cluster analysis is performed using the SAS nearest centroid sorting algorithm to determine the steady-state expression pattern of the network; one node corresponds to a Boolean function, and a new Boolean polynomial is defined. If and only if g[i] = 1, x[i] is a fixed point: g[i]:=(f[i]+x[i]+1)mod2 As a result, the steady state of the computation network is transformed into computing the product of g[i] M f Divide the calculation into several subsets, and finally summarize the results of each subset to obtain all possible steady states of the network; Among them, f[i]: represents the Boolean function of the i-th node, g[i]: represents the fixed point function of the i-th node, m f : represents the product of the fixed point functions of all nodes in the network, n: represents the total number of nodes in the network; Genes that are differentially expressed between clusters are identified by calculating the ratio of expression (0 or 1) for each node in the cluster.
6. The method for constructing HIV gene regulatory network according to claim 1, characterized in that The method also includes step three: evaluating the regulatory effect of genes based on an improved Boolean model polynomial calculation method.
7. The method for constructing HIV gene regulatory network according to claim 6, characterized in that In step 3: the regulatory effect of each gene is calculated to evaluate the impact of specific genes on the overall dynamics and steady state of the network. Each gene is set to 0 or 1 to observe the expression status of other genes. For each node, 4 sets are calculated to determine the control effect of the node. When these 4 sets no longer increase, the iteration stops. The possible control effect of node k is provided by the following 4 sets: A(k, 0): the set of genes whose final state is 0 when the target gene k is set to 0; A(k, 1): the set of genes whose final state is 1 when the target gene k is set to 1; B(k, 0): the set of genes whose final state is 1 when the target gene k is set to 0; B(k, 1): The set of genes whose final state is 0 when the target gene k is set to 1.
8. The method for constructing HIV gene regulatory network according to claim 1, characterized in that It also includes step 4: downloading the gene expression matrix of HIV-infected CD4+T cell samples from the public database GEO to verify the reliability of the model.
9. The method for constructing HIV gene regulatory network according to claim 8, characterized in that In step 4: If a gene corresponds to several probes, the probe with the highest variability is selected for subsequent analysis to ensure data accuracy, and the screening conditions of the data set are: ① The sample size is greater than 10; ② Contains GPL annotation file; ③ The GPL annotation file contains GeneSymbol; The probe with the largest variability is selected based on the coefficient of variation, CV = standard deviation / mean; Foldchange analysis is used for verification, and Foldchange calculates the expression ratio of a gene under two different conditions.
10. Application of the method for constructing HIV gene regulatory network according to any one of claims 1 to 9 in studying HIV drug therapeutic targets, characterized in that: The HIV drug treatment targets are key genes or proteins involved in HIV infection or replication.