A method and system for identifying early obesity gene markers

By constructing a topological structure of gene nodes and regulatory edges, and combining the temporal evolution trends of gene transcription levels and regulatory edge weights, abnormal changes in obesity-related gene markers are identified. This solves the problem of insufficient identification accuracy in existing technologies and enables more precise gene marker screening and personalized intervention.

CN122493957APending Publication Date: 2026-07-31PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIV
Filing Date
2026-06-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for identifying obesity-related gene markers face challenges such as massive data volumes, high analytical complexity, insufficient screening accuracy, inability to effectively identify regulatory network relationships between genes, and lack of dynamic gene transcription monitoring, which affects the formulation of early disease prediction and intervention strategies.

Method used

By constructing the topological structure of gene nodes and regulatory edges, and combining the temporal evolution trend of gene transcription level distribution and regulatory edge weights, we can identify abnormal time slices where transcription level surge intervals and regulatory edge weights change synchronously. We can analyze the regulatory consistency and diffusion direction at the junction of gene nodes, screen out abnormal gene nodes, assess the risk level of gene markers, and output a list of abnormal gene marker risk identification.

Benefits of technology

It improves the accuracy of gene marker screening, reduces false positive results, optimizes the precision of intervention programs, and ensures the personalization and precision of disease risk assessment and intervention measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493957A_ABST
    Figure CN122493957A_ABST
Patent Text Reader

Abstract

This invention relates to the field of bioinformatics, specifically to a method and system for early identification of obesity gene markers. The method includes the following steps: constructing a gene topology, extracting transcriptional distribution and edge weight evolution, identifying abnormal time slices, analyzing diffusion direction and gradient consistency to screen abnormal nodes, extracting intensity matrices and functional features to screen linkage sets, assessing deviation magnitude to label risk levels, matching response lists to screen adjustment nodes, and outputting a gene intervention response adjustment instruction set. In this invention, by constructing a topological structure of gene nodes and regulatory edges, changes in gene markers are accurately identified, improving screening accuracy, overcoming the dependence of traditional methods on massive amounts of data, analyzing regulatory consistency at gene node boundaries, meticulously capturing complex regulatory relationships between genes, and optimizing the screening process by assessing the deviation magnitude of gene marker risk levels, reducing false positives, and ensuring the personalization and precision of disease risk assessment and intervention measures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics technology, and in particular to a method and system for early identification of obesity gene markers. Background Technology

[0002] The field of bioinformatics primarily involves using computer science and statistical methods to analyze and interpret biological data, especially data analysis related to genomics, transcriptomics, and proteomics. This field encompasses the acquisition, processing, storage, and interpretation of biological data, relying on big data technologies, algorithm development, and machine learning methods to reveal complex mechanisms and disease associations within organisms. Core aspects include gene sequence analysis, gene transcription data analysis, protein and gene function prediction, and genome-wide association studies. Bioinformatics has wide applications, including disease research, drug development, and precision medicine. Traditional methods for identifying obesity gene markers involve analyzing individual genomic information to identify obesity-related gene markers to help predict obesity risk. Traditional methods employ gene association analysis techniques, such as genome-wide association studies (GWAS), which compare genetic differences between obese and non-obese individuals to identify gene loci associated with obesity. These methods rely on large-scale data collection and complex statistical analysis to identify potential gene markers, and face challenges such as massive data volumes, high analytical complexity, and insufficient accuracy.

[0003] Current technologies for processing large-scale genomic data rely on traditional genome-wide association studies (GWAS) for gene marker screening. While these methods can identify some obesity-related gene loci, they suffer from high data analysis complexity, large sample size requirements, and insufficient screening precision. Due to the high heterogeneity of genomic data, GWAS methods cannot effectively identify regulatory network relationships between genes, leading to inaccurate identification of some gene markers and the omission of certain biologically significant markers. Furthermore, existing methods lack dynamic gene transcription monitoring, failing to reflect the changing trends of gene regulation at different time points and easily overlooking the temporal characteristics of gene transcription. This can affect the early prediction and intervention strategies for diseases such as obesity, resulting in reduced effectiveness in clinical applications. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and to propose an early obesity gene marker identification method and system.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: an early obesity gene marker identification method, comprising the following steps: S1: Construct the topology of gene nodes and regulatory edges in the genomic dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, identify abnormal time slices where transcriptional level surge intervals and regulatory edge weights change synchronously, and generate a candidate gene marker set. S2: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of the gene node, analyze the regulatory consistency of the two at the gene node boundary, and combine the gene neighborhood connectivity to screen gene nodes with abnormal connectivity and regulatory diffusion to form a gene regulation abnormal marker layer. S3: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features between gene nodes, analyze the fluctuation of regulation intensity and the stability of functional patterns, screen and match abnormal gene nodes, and obtain a set of linked gene markers. S4: Based on the position number of the linked gene marker set, analyze the transcription density trend corresponding to the gene node, assess the degree of deviation from the healthy gene transcription baseline curve, mark the marker risk level of the gene node according to the deviation range, and output a gene marker risk abnormality identification list.

[0006] As a further aspect of the present invention, the candidate gene marker set includes transcriptional level burst gene numbers, regulatory edge weight abnormal sites, gene node coordinate markers, and time-slice mutation markers; the gene regulation abnormal marker layer includes regulatory diffusion abnormal gene markers, gene neighborhood connectivity units, gene boundary consistency blocks, and abnormal regulatory distribution grids; the linkage gene marker set includes regulatory abnormal continuous genes, functionally irregular regions, regulatory mutation overlap regions, and linkage abnormal gene codes; and the gene marker risk abnormality identification list includes risk level labels, gene response deviation values, local regulatory abnormality indicators, and transcription density offset levels.

[0007] As a further aspect of the present invention, the step of obtaining the candidate gene marker set specifically includes: S111: Based on the topological structure of gene nodes and regulatory edges in the genome dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, calculate the difference between two types of data within the same gene node, and obtain the trend value of the difference between transcriptional level and regulatory edge weight. S112: Based on the trend value of the difference between the transcriptional level and the regulatory side weight, identify the deviation value between the growth trend value and the time evolution trajectory of the regulatory side weight in the transcriptional level curve, superimpose the two types of values ​​in time periods, extract the time slice intervals where the growth trend exceeds the benchmark value and the deviation value exceeds the set threshold, and generate a high-frequency mutation time slice set. S113: For the high-frequency mutation time slice set, match the corresponding gene node number and spatial coordinate information, extract the corresponding gene node position in the high-frequency mutation time slice set, and generate a candidate gene marker set.

[0008] As a further aspect of the present invention, the step of obtaining the gene regulation abnormality marker layer specifically includes: S211: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of the gene node, extract the projection trajectory of the two at the gene node boundary, identify the distribution number and aggregation degree of the boundary point in the gene node, and obtain the gene boundary consistency map. S212: Based on the gene boundary consistency map, filter boundary regions with a aggregation degree higher than a preset aggregation degree threshold, compare with the spatial boundary of the genome dataset, identify continuous boundary clusters belonging to the same gene node, and obtain the abnormal regulatory diffusion regions within the gene node. S213: Call the abnormal region of regulation diffusion within the gene node, perform integrated analysis on gene boundary consistency, regulation diffusion dispersion, regulation intensity balance and regulation edge weight delay, calculate regulation intensity balance parameters, match gene nodes according to the response blocks in the layer, and form a gene regulation abnormality marker layer.

[0009] As a further aspect of the present invention, the step of obtaining the linked gene marker set specifically includes: S311: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features of the numbered gene nodes in the layer, align the data in the gene nodes with timestamps, identify the regulation intensity fluctuation value and functional mode stability offset, and obtain the gene local regulation response feature set. S312: Based on the gene local regulatory response feature set, jointly analyze the fluctuation of regulatory intensity and the stability of functional mode within the gene node, calculate the coupling feature value of regulatory behavior, screen gene units with regulatory behavior coupling degree in the layer, and establish a spatial distribution map of coordinated regulatory behavior response. S313: Call the spatial distribution map of the coordinated response of the regulatory behavior, cluster the gene nodes in the coupled feature value layer that exceed the coordinated recognition benchmark, label the node codes and coordinates corresponding to the continuous abnormal gene nodes, and obtain the set of linked gene markers.

[0010] As a further aspect of the present invention, the steps for obtaining the gene marker risk abnormality identification list are as follows: S411: Based on the position number of the linked gene marker set, extract the transcription density curve of the gene node under the specified number, perform time uniform processing, identify the density change per unit time, and obtain the gene regulation abnormal change rate set. S412: Based on the set of abnormal gene regulation change rates, identify the density distribution curve of the healthy gene stage, compare the current density change sequence with the reference curve, identify the gene regulation deviation level, extract and mark gene nodes whose deviation level exceeds the upper limit of the warning, and obtain the set of gene nodes with deviation surge. S413: Based on the set of deviation and spurious gene nodes, bind the deviation level value of each gene node to the position number in the genome dataset, sort them according to risk level, and output a list of gene marker risk anomalies.

[0011] As a further aspect of the present invention, the method further includes step S5: S5: Call the gene marker risk anomaly identification list, identify the corresponding number of gene node in the genome dataset, retrieve the preset intervention response unit list and gene protection priority sequence, screen the gene node numbers that need to adjust the response coverage, and output the gene intervention response adjustment instruction set; The gene intervention response adjustment instruction set includes adjusting the target gene node number, response level adjustment parameters, protection priority comparison items, and linkage intervention trigger types.

[0012] As a further aspect of the present invention, the step of obtaining the gene intervention response adjustment instruction set specifically includes: S511: Call the gene marker risk anomaly identification list, extract the gene node number in the genome dataset, map the gene node risk level value to the regional coordinate boundary, identify the gene node information corresponding to the risk grading standard, and generate a gene marker risk distribution map. S512: Based on the gene marker risk distribution map, extract the preset intervention strategy configuration parameters and response effect levels, match the gene marker risk levels with the intervention response levels, identify the unit numbers with insufficient response coverage, and obtain a gene marker response risk disconnect list. S513: Based on the gene marker response risk disconnect list, extract the key gene node numbers that need to improve the response coverage according to the priority sequence number of the gene protection sequence, output the adjustment control parameters linked with the original intervention unit in sequence, and output the gene intervention response adjustment instruction set.

[0013] The early obesity gene marker identification system is used to perform the above-described early obesity gene marker identification method. The system includes: The topology construction module is based on the topological structure of gene nodes and regulatory edges in the genome dataset. It compares the transcriptional level growth trend and the regulatory edge weight deviation value within the same time period, screens the synchronous mutation intervals of the two, extracts the gene node number and spatial coordinates, summarizes the abnormal time slices and gene node numbers, and generates a candidate gene marker set. Based on the candidate gene marker set, the gene localization module identifies the consistency between the direction of regulatory diffusion and the gradient vector of regulatory edge weights, marks the gene boundary numbers, matches the genome dataset, extracts the number range of gene nodes with abnormal regulatory diffusion, and establishes a gene regulation abnormality marker layer. Based on the gene regulation abnormality marker layer, the regulation linkage module retrieves the continuous data sequence of the regulation intensity matrix and local functional features in the region, determines the connectivity of the regulation abnormality boundary and the stability of the functional pattern, marks the gene node numbers that meet the linkage threshold of both, and outputs the linkage gene marker set. The risk warning module analyzes the transcription density trend of the corresponding gene node and the degree of deviation from the healthy gene transcription baseline curve based on the gene node number of the linked gene marker set, extracts the gene node number that deviates from the trend, completes the level identification according to the risk grading standard, and generates a gene marker risk abnormality identification list. Based on the gene marker risk anomaly identification list, the intervention optimization module finds the corresponding position number of the risk level gene node in the genome dataset, retrieves the current intervention response unit configuration list, compares the gene protection priority with the current response level, filters the gene node numbers that need to be updated, and outputs the gene intervention response adjustment instruction set.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This invention constructs a topological structure of gene nodes and regulatory edges, combining the temporal evolution trends of gene transcriptional level distribution and regulatory edge weights to accurately identify the changing patterns of gene markers. This significantly improves the accuracy of candidate gene marker screening, avoiding the limitations of traditional methods that rely on massive amounts of data. By combining the regulatory diffusion direction of gene nodes with the gradient vector of regulatory edge weights, the invention analyzes the regulatory consistency at the gene node boundaries, enabling a more detailed capture of complex regulatory relationships between genes and enhancing the biological significance of gene markers. Furthermore, by assessing the deviation of gene marker risk levels, the gene marker screening process is further optimized, reducing false positive results caused by errors in traditional methods. Simultaneously, by retrieving the intervention response unit list and comparing gene protection priority sequences, the intervention plan becomes more precise, ensuring the personalization and accuracy of disease risk assessment and intervention measures. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a flowchart illustrating the process of obtaining the candidate gene marker set in this invention. Figure 3 This is a flowchart illustrating the process of obtaining the gene regulation abnormality marker layer in this invention. Figure 4 This is a flowchart illustrating the process of obtaining the linked gene marker set in this invention. Figure 5 This is a flowchart illustrating the process of obtaining the list of gene marker risk anomalies in this invention. Figure 6 This is a flowchart illustrating the process of obtaining the gene intervention response adjustment instruction set in this invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0018] Example 1: Please see Figure 1 This invention provides a technical solution: a method for early identification of obesity gene markers, comprising the following steps: S1: Construct the topology of gene nodes and regulatory edges in the genomic dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, identify abnormal time slices where transcriptional level surge intervals and regulatory edge weights change synchronously, and generate a candidate gene marker set. S2: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of gene nodes, analyze the regulatory consistency of the two at the gene node boundary, and combine the gene neighborhood connectivity to screen gene nodes with abnormal connectivity and regulatory diffusion to form a gene regulation abnormal marker layer. S3: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features between gene nodes, analyze the fluctuation of regulation intensity and the stability of functional patterns, screen and match abnormal gene nodes, and obtain a set of linked gene markers. S4: Based on the position number of the linked gene marker set, analyze the transcription density trend corresponding to the gene node, assess the degree of deviation from the healthy gene transcription baseline curve, mark the marker risk level of the gene node according to the deviation range, and output a list of gene marker risk abnormalities. S5: Call the gene marker risk anomaly identification list, identify the corresponding number of gene node in the genome dataset, retrieve the preset intervention response unit list and gene protection priority sequence, filter the gene node numbers that need to adjust the response coverage, and output the gene intervention response adjustment instruction set.

[0019] The candidate gene marker set includes transcriptional level mutation gene numbers, abnormal regulatory edge weights, gene node coordinate markers, and time-slice mutation markers. The gene regulation aberration marker layer includes abnormal regulatory diffusion gene markers, gene neighborhood connectivity units, gene boundary consistency blocks, and abnormal regulatory distribution grids. The linkage gene marker set includes continuous genes with abnormal regulation, functionally irregular regions, overlapping regulatory mutation regions, and linkage aberration gene codes. The gene marker risk aberration identification list includes risk level labels, gene response deviation values, local regulatory aberration indicators, and transcription density offset levels. The gene intervention response adjustment instruction set includes adjusting target gene node numbers, response level adjustment parameters, protection priority comparison items, and linkage intervention trigger types.

[0020] Please see Figure 2 The specific steps for obtaining the candidate gene marker set are as follows: S111: Based on the topological structure of gene nodes and regulatory edges in the genome dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, calculate the difference between two types of data within the same gene node, and obtain the trend value of the difference between transcriptional level and regulatory edge weight. Based on the topological structure of gene nodes and regulatory edges in genomic data, the human transcription factor ESR1 (ENSG00000091831) and its regulated breast cancer-related genes BCL2 and MYC were selected. The relative mRNA transcription of ESR1 at four time points (T0, T1, T2, and T3, with an interval of 12 h) was measured by qRT-PCR and normalized to the range of 0–100. Meanwhile, ChIP-seq data were used to calculate the binding affinity scores (0–1) of ESR1 with the BCL2 and MYC regulatory edges, which were used as regulatory edge weights. To assess the matching degree between the transcriptional level and the regulatory edge weights, the absolute difference between the normalized transcriptional value of ESR1 and the corresponding edge weight was calculated for each time point, and then averaged to obtain the average difference trend value of ESR1. Taking T0 as an example, the normalized transcriptional level of ESR1 was 70, and the edge weights of BCL2 and MYC were 0.70 and 0.65, respectively, with differences of 69.3 and 69.35, respectively. The average value was 69.325. Following this method, the values ​​for T1 were 74.225, T2 was 81.175, and T3 was 77.15. Table 1: Trend values ​​of ESR1 gene transcription level and regulatory side weight differences; Referring to Table 1, the normalized transcriptional level of the ESR1 gene at time T0 is 70, the regulatory edge weight with BCL2 is 0.70, and the regulatory edge weight with MYC is 0.65. The average trend value of the variance was calculated by summing the difference between 70 and 0.70 and the difference between 70 and 0.65, and then dividing by 2, resulting in 69.325. At time T1, the normalized transcriptional level is 75, the regulatory edge weight with BCL2 is 0.80, and the regulatory edge weight with MYC is 0.75. The difference between 75 and 0.80 and the difference between 75 and 0.75 was summed, and then divided by 2, resulting in 74.225. This process was repeated to obtain the average trend value of the variance of the ESR1 gene at each time point.

[0021] S112: Based on the trend value of the difference between the transcriptional level and the regulatory side weight, identify the deviation value between the growth trend value and the time evolution trajectory of the regulatory side weight in the transcriptional level curve, superimpose the two types of values ​​in time periods, extract the time slice intervals where the growth trend exceeds the baseline value and the deviation value exceeds the set threshold, and generate a high-frequency mutation time slice set. Based on the trends of ESR1 transcription levels and regulatory edge weight differences at various time points, the transcriptional growth trend value was first calculated: obtained by dividing the difference in normalized transcription levels between adjacent time points by 12h. For example, T0–T1 was (75–70) / 12≈0.417 units / h, and T1–T2 was (82–75) / 12≈0.583 units / h. Next, the regulatory edge weight deviation value was calculated: using the historical average weight benchmark of 0.68 as a reference, the absolute difference between the regulatory edge weight at each time point and the benchmark was calculated, and the maximum value was taken as the deviation value. For example, at T1, BCL2 was |0.80–0.68|=0.12, MYC was |0.75–0.68|=0.07, and the deviation value was 0.12. Subsequently, the growth trend and deviation value were combined for judgment: a growth trend > 0.5 units / h was considered exceeding the benchmark (based on the upper limit of the 95% confidence interval for long-term monitoring of healthy cells), and a deviation value > 0.10 was considered a regulatory abnormality (based on network stability analysis). The results showed that the growth trend in the T0–T1 interval did not exceed the threshold (0.417), but the deviation value reached 0.12; while the growth trend in the T1–T2 interval was 0.583, with a deviation value of 0.15, both exceeding the threshold. Therefore, it was identified as an abnormally active period, generating a high-frequency mutation time slice set (T1–T2), and locating the key gene node ESR1.

[0022] S113: For high-frequency mutation time slices, match the corresponding gene node number and spatial coordinate information, extract the corresponding gene node position within the high-frequency mutation time slice, and generate a candidate gene marker set. For high-frequency mutation time slices, which include time slice intervals (T1 to T2) and gene node ESR1, the corresponding gene node numbers and spatial coordinate information are matched. Specifically, the gene node number ENSG00000091831 of ESR1 and its spatial coordinate information in the human genome are matched, for example, the chromosome 6q25.1 region, with a specific start position of 151.694.764bp and an end position of 152.118.506bp. The corresponding gene node positions within the high-frequency mutation time slices are extracted to determine the precise genomic location of the ESR1 gene exhibiting high-frequency mutation behavior within that time slice interval. Since the entire region is identified as abnormal, the entire chromosome 6q25.1 region is extracted to generate a candidate gene marker set. This set contains the gene ESR1 number and its precise genomic location information, specifically ESR1, ENSG00000091831, chromosome 6q25.1, 151.694.764-152.118.506bp.

[0023] Please see Figure 3 The specific steps for obtaining the gene regulation abnormality marker layer are as follows: S211: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of the gene node, extract the projection trajectory of the two at the gene node boundary, identify the distribution number and aggregation degree of the boundary point in the gene node, and obtain the gene boundary consistency map. Based on a candidate gene marker set (including ESR1 and its genomic location), the regulatory diffusion direction and regulatory edge weight gradient vector of ESR1 are extracted. The regulatory diffusion direction is determined by analyzing its effect on target genes. For example, ESR1 exhibits activation regulation on both BCL2 and MYC, with the direction being ESR1 pointing towards the target gene. The regulatory edge weight gradient vector is calculated from the rate of change of weights within genomic regions. For instance, if the weight between ESR1 and BCL2 increases from 0.75 to 0.85, the gradient vector represents the direction and magnitude of this increase. Further, projection operations are performed at gene node boundaries, projecting the regulatory direction vector onto the weight gradient vector of the target gene promoter region to quantify their consistency. Projection values ​​range from -1 to 1, with values ​​close to 1 indicating high consistency. The boundary point is defined as a region with a projection value > 0.7. The degree of aggregation is assessed by statistically analyzing the number of occurrences within a 5000 bp window in the genome. For example, if 12 junctions are found within a window and the average value is 5, then the aggregation is considered significant, forming a gene junction consistency map that covers the genomic junction regions and aggregation degree information of ESR1 and its targets (such as BCL2 and MYC).

[0024] S212: Based on the gene boundary consistency map, filter the boundary regions with a aggregation degree higher than the preset aggregation degree threshold, compare with the spatial boundary of the genome dataset, identify continuous boundary clusters belonging to the same gene node, and obtain the abnormal regulatory diffusion regions within the gene node. Based on the aggregation degree information of ESR1 gene regulatory regions in the gene boundary consistency map, boundary regions with aggregation degrees higher than a preset aggregation degree threshold are screened. The average level refers to an average aggregation degree of 6 boundary points per 5000 bp within the current genome region. If the aggregation degree of a certain region between ESR1 and BCL2 reaches 12 boundary points per 5000 bp, the region is selected. By comparing the spatial boundaries of the genome dataset, the selected high-aggregation boundary regions are spatially compared with the known coding regions, regulatory regions, and promoter regions of the ESR1 gene and its target genes. By performing precise boundary comparisons, continuous and clustered regions belonging to the same gene node are identified. High-aggregation regions that are spatially continuous and clearly belong to the ESR1 regulatory module are found within the ESR1 regulatory range. For example, if a continuous region from 152.100.000bp to 152.200.000bp is detected in the downstream region of ESR1, and its aggregation degree exceeds the average level, and the regulatory effect in this region clearly originates from the ESR1 gene, it is identified as an abnormal region of regulatory diffusion, thus obtaining an abnormal region of regulatory diffusion within the gene node ESR1.

[0025] S213: Invoke the abnormal regulatory diffusion regions within gene nodes, and perform integrated analysis on gene boundary consistency, regulatory diffusion dispersion, regulatory intensity balance, and regulatory edge weight delay, using the following formula: ; Calculate the equilibrium parameters of regulatory intensity, match gene nodes according to the response blocks in the layer, and form a gene regulation abnormality marker layer; in, Represents the equilibrium parameter of regulation intensity. Representing the The regulatory intensity of individual gene nodes, This represents the mean of the regulatory strength across all gene nodes. The mean value representing the consistency of all gene node boundaries. Representing the The amount of consistent adjustment in regulation at each gene node. The weighted parameter representing the consistency between the control intensity and the boundary. The total number of gene nodes; The regulatory diffusion aberration region within the ESR1 gene node is invoked. The dispersion of regulatory diffusion is measured by calculating the standard deviation of the range of ESR1's regulatory influence on its multiple target genes. The balance of regulatory intensity is calculated using the following formula: The regulatory edge weight delay is obtained by measuring the time interval between ESR1 transcriptional changes and their target gene responses, as shown in the formula. This represents a balance parameter of regulatory intensity, quantifying the equilibrium state between the intensity of gene regulation and its boundary consistency. Representing the The regulatory strength of a gene node is the average weight of that gene across all regulatory edges. The mean value representing the regulatory strength of all gene nodes is the sum of the strengths of all genes. The arithmetic mean, Representing the The regulatory consistency adjustment of each gene node is standardized and adjusted based on the gene boundary consistency map (S211 results), reflecting the concentration and stability of its regulatory influence at the genome boundary. The mean value representing the consistent adjustment of regulation across all gene nodes is the average value of all genes. The arithmetic mean, A weighted parameter representing the balance between the deviation of regulatory intensity from the mean and the deviation of regulatory consistency from the mean is used to balance the impact of the deviation of regulatory intensity from the mean. The contribution value can be set to 1.2. Analysis of multiple sets of biological experimental data revealed that when... When the value is set to 1.2, the formula achieves the optimal balance between sensitivity and specificity in identifying anomalous regions. The total number of gene nodes refers to the total number of gene nodes involved in this analysis. The advantage of this formula lies in its ability to control gene regulation intensity. Consistency with regulation The square root operation of the product of the deviations from the mean can comprehensively assess the overall balance of gene regulation. When the intensity of gene regulation deviates significantly from the average level or its regulatory consistency is abnormal, The value will increase accordingly, accurately reflecting the degree of abnormality in its regulatory balance, further improving the comprehensiveness and accuracy of abnormal region identification; taking ESR1 as an example, assuming that the abnormal region includes There are 10 gene nodes, namely ESR1, BCL2, and MYC, and the regulatory strength of these gene nodes is (…). ) and the adjustment amount of the regulation consistency ( As shown in Table 2, the weight parameters ; Table 2: Gene node regulatory intensity and consistency adjustment amount; First, calculate the mean of the regulatory strength of all gene nodes. and the mean of the adjustment amount of the regulation consistency ; ; ; Next, the regulatory strength equilibrium parameters of the ESR1 gene node were calculated. : ; Continue calculating the equilibrium parameters of regulatory intensity at the BCL2 gene node. : ; Calculate the regulatory strength equilibrium parameters of the MYC gene node : ; The regulatory strength equilibrium parameters R values ​​for ESR1, BCL2, and MYC were 0.3844, 0.2704, and 0.0343, respectively. These calculated regulatory strength equilibrium parameter values ​​were compared with a preset abnormality threshold of 0.35. This threshold was determined based on the upper limit of the 99% confidence interval for gene regulatory equilibrium parameters under normal physiological conditions. A value greater than 0.35 indicates an abnormality in regulatory balance. Among them, the R value of ESR1 is 0.3844, which is greater than 0.35, indicating that ESR1 has an abnormality in regulatory balance. The R value of BCL2 is 0.2704, which is less than 0.35, and the R value of MYC is 0.0343, which is less than 0.35. Therefore, ESR1 is identified as a gene node with an abnormality in regulatory balance. Gene node matching is performed according to the response blocks in the layer, and the ESR1 gene is marked to the corresponding abnormal regulatory region to form a gene regulation abnormality marker layer. In this layer, the ESR1 gene and its position in the genome are marked as abnormal.

[0026] Please see Figure 4 The specific steps for obtaining the set of linked gene markers are as follows: S311: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features of numbered gene nodes in the layer, align the data in the gene nodes with timestamps, identify the regulation intensity fluctuation value and functional mode stability offset, and obtain the gene local regulation response feature set. Based on a gene regulation aberration marker layer, which includes ESR1 genes marked as having abnormal regulation, the regulatory intensity matrix and local functional features of ESR1 gene nodes within the layer are extracted. The ESR1 regulatory intensity matrix records the intensity of its regulatory effect on its target genes (such as BCL2 and MYC) at five time points (T0, T1, T2, T3, and T4, 12 hours apart), with intensity values ​​ranging from 0 to 100. Local functional features are extracted to extract the activity scores of signaling pathways involved by ESR1 (range 0 to 1). All data within the ESR1 gene nodes are time-stamp aligned to ensure that regulatory intensity and functional feature data are analyzed at the same time points. The fluctuation values ​​of regulatory intensity and the stability shift of functional patterns are identified. The regulatory intensity fluctuation values ​​are calculated by calculating all time points in the ESR1 regulatory intensity matrix. The standard deviation of the data points is obtained. For example, the ESR1 regulation intensity sequence for BCL2 has an average value of 78. Its standard deviation is calculated by summing the squares of the differences between each value in the sequence and the average value, dividing by the sequence length of 5, and then taking the square root, resulting in approximately 5.25. The functional mode stability shift is obtained by calculating the standard deviation of the ESR1 signaling pathway activity score sequence, which is [0.90, 0.92, 0.95, 0.91, 0.93] with an average value of 0.922. Its standard deviation is calculated by summing the squares of the differences between each value in the sequence and the average value, dividing by the sequence length of 5, and then taking the square root, resulting in approximately 0.0172. The gene local regulatory response feature set is obtained, which includes the ESR1 regulation intensity fluctuation value of 5.25 and the functional mode stability shift of 0.0172.

[0027] S312: Based on the gene local regulatory response feature set, a joint analysis of the volatility of regulatory intensity and the stability of functional patterns within gene nodes is performed, using the following formula: ; Calculate the coupling feature value of regulatory behavior, screen gene units with high coupling degree of regulatory behavior in the layer, and establish a spatial distribution map of coordinated response of regulatory behavior; , Indicates the coupled characteristic value of regulatory behavior. It is a regulatory stabilizing factor for the k-th gene node. The fluctuation in the regulatory intensity of the k-th gene node, Represents the regulatory strength of the k-th gene node. Represents the regulatory strength of the k-th gene node. This represents the mean of the regulatory strength across all gene nodes. Represents the total number of gene nodes; Based on the gene local regulatory response feature set, where the regulatory intensity variability of ESR1 is 5.25 and the functional mode stability shift is 0.0172, a joint analysis of the regulatory intensity variability and functional mode stability within the ESR1 gene node was performed, and the regulatory behavior coupling feature value was calculated using the following formula: In the formula, This represents the coupling characteristic value of regulatory behavior, quantifying the degree of synergistic response between gene regulatory fluctuations and functional stability. For the first The regulatory stability factor of gene nodes is a shift in functional pattern stability; a smaller value indicates higher stability, and vice versa. Representing the The variability in the regulatory intensity of gene nodes, i.e., the standard deviation of their regulatory intensity. Representing the The regulatory strength of gene nodes is the average regulatory strength over the analysis period. Representing the The regulatory strength of gene nodes, here related to The meaning is the same; in the denominator, it is used to calculate the dispersion of the overall strength. The mean value representing the regulatory strength of all gene nodes is the sum of the strengths of all genes. The arithmetic mean, The total number of gene nodes refers to the total number of gene nodes involved in this analysis. The advantage of this formula lies in its ability to manipulate the fluctuations in gene regulation intensity. With the intensity of regulation and stability factor Summing the products of these ratios and dividing by the standard deviation of the total regulatory strength can effectively identify genes with large fluctuations in regulatory strength under conditions of functional instability. The value will increase significantly, thus accurately revealing the abnormal collaborative response behavior of key genes; assuming that this analysis involves These are three gene nodes, ESR1, G2, and G3, which regulate stability factors (…). ), volatility of regulation intensity ( ) and regulatory intensity ( / As shown in Table 3; Table 3: Gene node regulatory characteristics parameters; First, calculate the mean of the regulatory strength of all gene nodes. : ; Next, calculate the numerator of the formula: ESR1 item: ; G2 item: ; G3 item: ; Sum of numerators: ; Then calculate the denominator of the formula: ; ; ; Summing the denominators and taking the square root: ; Finally, the coupling characteristic value of the regulatory behavior is calculated. : ; Regulatory behavior coupling eigenvalues The threshold value was 2535.79. Gene units in the screening layer whose regulatory behavior coupling exceeded the co-identification benchmark were selected. The co-identification benchmark threshold was set to 1000. This threshold was determined by analyzing gene regulatory networks in normal cells. The value distribution, with the upper limit of its 99% confidence interval as the boundary for judging high coupling, ESR1 The value is much greater than 1000, indicating that the fluctuation of the regulatory intensity of the ESR1 gene exhibits an abnormal behavior of high coupling with the stability of its functional pattern. A spatial distribution map of the coordinated response of regulatory behavior is established, ESR1 is marked as a highly coupled gene, and its location information in the genome is displayed.

[0028] S313: Call the spatial distribution map of the coordinated response of regulatory behavior, cluster the gene nodes in the coupled feature value layer that exceed the coordinated recognition benchmark, label the node codes and coordinates of the continuous abnormal gene nodes, and obtain the set of linked gene markers. The spatial distribution map of the coordinated response to regulatory behavior was used. The ESR1 gene was identified as a highly coupled gene with a coupling eigenvalue G of 2535.79 and a co-coordination benchmark of 1000. Gene nodes in the coupling eigenvalue layer that exceeded the co-coordination benchmark were clustered. Since the G value of ESR1 (2535.79) is higher than the co-coordination benchmark of 1000, it indicates an abnormal coordinated response in its regulatory behavior. If a gene (such as G2) also exceeds this benchmark, and the spatial distance between the two genes in the genome is less than 500 kb, then cluster analysis is performed on them. For example, if ESR1 is located in the 6q25.1 region of chromosome 6, and G2 is located in the 6q25.3 region, then... If the distance between them is 200kb, then ESR1 and G2 form a cluster. The node codes and coordinates corresponding to the consecutive abnormal gene nodes are labeled. In the cluster, the code ENSG00000091831 of the gene node ESR1 and its genomic coordinates 151.694.764-152.118.506bp are identified and recorded, as well as the consecutive and abnormal genes (such as G2), to obtain the linked gene marker set. This set includes the ESR1 gene, gene number ENSG00000091831, genomic location chromosome 6q25.1 and 151.694.764-152.118.506bp.

[0029] Please see Figure 5 The specific steps for obtaining the list of abnormal genetic marker risks are as follows: S411: Based on the position number of the linked gene marker set, extract the transcription density curve of the gene node under the specified number, perform time uniform processing, identify the density change per unit time, and obtain the set of abnormal gene regulation change rates. Based on the positional numbering of the linked gene marker set, the transcription density curve of the ESR1 gene node under the specified number was extracted. This density curve was obtained by real-time quantitative analysis of the concentration of ESR1 mRNA in the cell nucleus. The relative transcription density values ​​were obtained at four time points (T0, T1, T2, and T3, 12-hour intervals), and time-uniform processing was performed to ensure consistency in data acquisition and preprocessing across all time points. The density change per unit time was identified, and the rate of change of ESR1 gene transcription density at each time point was calculated. For example, if the ESR1 transcription density at time T0 is 100 relative units and at time T1 is 110 relative units, then the density change from T0 to T1 is calculated by... The density change rate was calculated by dividing the difference between 110 and 100 by 12 hours, yielding approximately 0.83 relative units per hour. The density change rate for the time period from T1 to T2 was calculated by dividing the difference between 130 and 110 by 12 hours, yielding approximately 1.67 relative units per hour. The density change rate for the time period from T2 to T3 was calculated by dividing the difference between 150 and 130 by 12 hours, yielding approximately 1.67 relative units per hour. A set of abnormal gene regulation change rates was obtained, which contains the unit time density change of the ESR1 gene in each time period. For example, the unit time density change of the ESR1 gene in each time period is [0.83, 1.67, 1.67] relative units per hour.

[0030] S412: Based on the abnormal change rate set of gene regulation, identify the density distribution curve of the healthy gene stage, compare the current density change sequence with the reference curve, identify the deviation level of gene regulation, extract and mark gene nodes whose deviation level exceeds the upper limit of the warning, and obtain the set of gene nodes with deviation surge. Based on the abnormal change rate set of ESR1 gene regulation, the reference density distribution curve for the healthy gene stage is obtained by statistically analyzing the transcriptional density change rate of ESR1 gene under physiological conditions in a large number of healthy cells. For example, the average density change rate of ESR1 gene in a healthy state is 0.5 relative units per hour, with a standard deviation of 0.2, and its 95% confidence interval is 0.1 to 0.9 relative units per hour. The current density change sequence is compared with the reference curve, and the statistical distance (e.g., using Euclidean distance) between the current density change rate sequence of ESR1 and the healthy reference curve is compared to identify the level of gene regulation deviation. The deviation level is divided into 5 levels: 0.1-0.9 (Level 1, normal fluctuation), 1.0-1.5 (Level 2, slight deviation), 1.6... -2.0 (Level 3, moderate deviation), 2.1-2.5 (Level 4, significant deviation), greater than 2.5 (Level 5, severe deviation). Gene nodes with deviation levels exceeding the upper warning bound are extracted and marked. The upper warning bound is set to Level 2. Gene nodes with deviation levels exceeding 2 are marked as abnormal. The change rate of ESR1 in the time period from T0 to T1 is 0.83, which is Level 1. The change rate in the time period from T1 to T2 is 1.67, which is Level 3. The change rate in the time period from T2 to T3 is 1.67, which is Level 3. Since the deviation level of ESR1 (maximum 3) exceeds the upper warning bound (Level 2), the ESR1 gene is marked as a deviation mutation gene node. The deviation mutation gene node set is obtained, which includes the ESR1 gene and its deviation level 3.

[0031] S413: Based on the set of gene nodes that deviate from the spurious gene node, bind the deviation level value of each gene node to the position number in the genome dataset, sort them according to risk level, and output a list of gene marker risk anomalies. Based on the set of deviation-proliferation gene nodes, including the ESR1 gene and its deviation level 3, the deviation level value of each gene node is bound to its location number in the genome dataset. The deviation level 3 of the ESR1 gene is associated with its genomic location information (chromosome 6q25.1, 151.694.764-152.118.506bp). The nodes are sorted by risk level. Here, only the ESR1 gene node is identified as a deviation-proliferation, so it constitutes the sorting result itself. If multiple gene nodes exist, they are sorted from high to low deviation level. A list of gene marker risk anomalies is output, which clearly lists the ESR1 gene number, genomic location, and its corresponding risk level. Specifically, it lists the ESR1 gene, gene number ENSG00000091831, genomic location chromosome 6q25.1 and 151.694.764-152.118.506bp, and risk level 3.

[0032] Please see Figure 6The specific steps for obtaining the gene intervention response adjustment instruction set are as follows: S511: Call the gene marker risk anomaly identification list, extract the gene node number in the genome dataset, map the gene node risk level value to the regional coordinate boundary, identify the gene node information corresponding to the risk grading standard, and generate a gene marker risk distribution map. The gene marker risk anomaly identification list is retrieved, which includes the ESR1 gene with a risk level of 3. The ESR1 gene node ID ENSG00000091831 in the genome dataset is extracted. The risk level value of 3 of the ESR1 gene node is mapped to the genomic region coordinate boundary (chromosome 6q25.1, 151.694.764-152.118.506bp). Specifically, the risk level is used as the color intensity of a heatmap to plot the risk distribution on the genome coordinate axis. The gene node information corresponding to the risk grading standard is identified. The risk grading standard is set according to the risk level. For example, risk level 1 corresponds to "low protection", level 2 corresponds to "intermediate protection", level 3 corresponds to "high protection", level 4 corresponds to "very high protection", and level 5 corresponds to "critical protection". The risk level of the ESR1 gene is 3, so its corresponding risk grading standard is "high protection". A gene marker risk distribution map is generated. This map visualizes the location of the ESR1 gene in the genome and its corresponding "high protection" level.

[0033] S512: Based on the gene marker risk distribution map, extract the preset intervention strategy configuration parameters and response effect levels, match the gene marker risk level with the intervention response level, identify the unit number with insufficient response coverage, and obtain the gene marker response risk disconnect list. Based on the gene marker risk distribution map, where the ESR1 gene is marked as "high protection", the pre-set intervention strategy configuration parameters and response effect levels are extracted. For example, for the ESR1 gene, the existing intervention regimen "estrogen receptor antagonist therapy" (intervention response unit number: IRU-ESR1-001) has a response level of "moderately effective", defined as: 1 (ineffective), 2 (inefficient), 3 (moderately effective), 4 (highly effective), and 5 (completely effective). The gene marker risk level and intervention response level are matched, and the ESR1 risk grading standard (high protection, corresponding to level 3) is aligned with the "estrogen receptor antagonist therapy" level. The response level of "hormone receptor antagonist therapy" (moderately effective, corresponding to level 3) is compared to identify the unit number with insufficient response coverage. In this example, the risk grading criteria match the intervention response level (3 to 3), indicating that the current intervention protocol has sufficient coverage. Therefore, this intervention unit is not identified as a unit with insufficient response coverage. If the risk level of ESR1 is "very high protection" (level 4), while the intervention response level is still "moderately effective" (level 3), then IRU-ESR1-001 will be identified as having insufficient response coverage, resulting in a gene marker response risk mismatch list. In this example, the list is empty.

[0034] S513: Based on the gene marker response risk disconnect list and the priority sequence number of the gene protection sequence, extract the key gene node numbers that need to improve the response coverage, output the adjustment control parameters linked with the original intervention unit in sequence, and output the gene intervention response adjustment instruction set. According to the gene marker response risk disconnect list, which is empty in this example, if the list contains gene nodes (for example, if the S512 result shows ESR1 protection level 4 and intervention response level 3, then ESR1 will appear in this list), the key gene node numbers that need to improve response coverage are extracted based on the level numbers in the gene protection priority sequence. The priority sequence is set as follows: level 5 (critical protection) is better than level 4 (very high protection) is better than level 3 (high protection). If ESR1 exists in the list (because risk level 4 is higher than response level 3), then ESR1 is identified as a key gene node that needs to improve response coverage. The system outputs adjustment control parameters linked to the original intervention unit in sequence, based on the point number. For example, for ESR1, to improve its response level from 3 to 4, the adjustment control parameter can be set to "increase the dose of estrogen receptor antagonist therapy by 50%" or "increase the number of adjuvant chemotherapy cycles". The adjustment parameter is executed in conjunction with the original intervention unit IRU-ESR1-001, outputting a gene intervention response adjustment instruction set. This instruction set contains a specific intervention adjustment plan for ESR1, such as the ESR1 gene, the intervention response unit number IRU-ESR1-001, and the adjustment control parameter of increasing the dose of estrogen receptor antagonist by 50%.

[0035] The early obesity gene marker identification system is used to perform the above-mentioned early obesity gene marker identification method. The system includes: The topology construction module is based on the topological structure of gene nodes and regulatory edges in the genome dataset. It compares the transcriptional level growth trend and the regulatory edge weight deviation value within the same time period, screens the synchronous mutation intervals of the two, extracts the gene node number and spatial coordinates, summarizes the abnormal time slices and gene node numbers, and generates a candidate gene marker set. The gene localization module identifies the consistency between the direction of regulatory diffusion and the gradient vector of regulatory edge weights based on a candidate gene marker set, marks the gene boundary number, matches the genome dataset, extracts the number range of gene nodes with abnormal regulatory diffusion, and establishes a gene regulation abnormality marker layer. The regulatory linkage module is based on the gene regulation abnormality marker layer. It retrieves the continuous data sequence of the regulatory intensity matrix and local functional features in the region, judges the connectivity of the regulatory abnormality boundary and the stability of the functional pattern, marks the gene node number that meets the linkage threshold of the two, and outputs the linkage gene marker set. The risk warning module analyzes the transcription density trend of the corresponding gene node and the degree of deviation from the healthy gene transcription baseline curve based on the gene node number of the linked gene marker set, extracts the gene node number that deviates from the trend, completes the level identification according to the risk grading standard, and generates a list of abnormal gene marker risks. The intervention optimization module uses the gene marker risk anomaly identification list to find the corresponding position number of the risk level gene node in the genome dataset, retrieves the current intervention response unit configuration list, compares the gene protection priority with the current response level, filters the gene node numbers that need to be updated, and outputs the gene intervention response adjustment instruction set.

[0036] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A method for identifying early obesity gene markers, characterized by, Includes the following steps: S1: Construct the topology of gene nodes and regulatory edges in the genomic dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, identify abnormal time slices where transcriptional level surge intervals and regulatory edge weights change synchronously, and generate a candidate gene marker set. S2: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of the gene node, analyze the regulatory consistency of the two at the gene node boundary, and combine the gene neighborhood connectivity to screen gene nodes with abnormal connectivity and regulatory diffusion to form a gene regulation abnormal marker layer. S3: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features between gene nodes, analyze the fluctuation of regulation intensity and the stability of functional patterns, screen and match abnormal gene nodes, and obtain a set of linked gene markers. S4: Based on the position number of the linked gene marker set, analyze the transcription density trend corresponding to the gene node, assess the degree of deviation from the healthy gene transcription baseline curve, mark the marker risk level of the gene node according to the deviation range, and output a gene marker risk abnormality identification list.

2. The method of identifying early obesity gene markers according to claim 1, wherein The candidate gene marker set includes transcriptional-level mutation gene numbers, regulatory edge weight anomaly sites, gene node coordinate markers, and time-slice mutation markers. The gene regulation anomaly marker layer includes regulatory diffusion anomaly gene markers, gene neighborhood connectivity units, gene boundary consistency blocks, and anomaly regulatory distribution grid. The linkage gene marker set includes regulatory anomaly continuous genes, functionally irregular regions, regulatory mutation overlap regions, and linkage anomaly gene codes. The gene marker risk anomaly identification list includes risk level labels, gene response deviation values, local regulatory anomaly indicators, and transcription density offset levels.

3. The method for early obesity gene marker identification according to claim 1, characterized in that, The specific steps for obtaining the candidate gene marker set are as follows: S111: Based on the topological structure of gene nodes and regulatory edges in the genome dataset, extract the temporal evolution trend of transcriptional level distribution and regulatory edge weight in gene node attributes, calculate the difference between two types of data within the same gene node, and obtain the trend value of the difference between transcriptional level and regulatory edge weight. S112: Based on the trend value of the difference between the transcriptional level and the regulatory side weight, identify the deviation value between the growth trend value and the time evolution trajectory of the regulatory side weight in the transcriptional level curve, superimpose the two types of values ​​in time periods, extract the time slice intervals where the growth trend exceeds the benchmark value and the deviation value exceeds the set threshold, and generate a high-frequency mutation time slice set. S113: For the high-frequency mutation time slice set, match the corresponding gene node number and spatial coordinate information, extract the corresponding gene node position in the high-frequency mutation time slice set, and generate a candidate gene marker set.

4. The method for early obesity gene marker identification according to claim 3, characterized in that, The specific steps for obtaining the gene regulation abnormality marker layer are as follows: S211: Based on the candidate gene marker set, extract the regulatory diffusion direction and regulatory edge weight gradient vector of the gene node, extract the projection trajectory of the two at the gene node boundary, identify the distribution number and aggregation degree of the boundary point in the gene node, and obtain the gene boundary consistency map. S212: Based on the gene boundary consistency map, filter boundary regions with a aggregation degree higher than a preset aggregation degree threshold, compare with the spatial boundary of the genome dataset, identify continuous boundary clusters belonging to the same gene node, and obtain the abnormal regulatory diffusion regions within the gene node. S213: Call the abnormal region of regulation diffusion within the gene node, perform integrated analysis on gene boundary consistency, regulation diffusion dispersion, regulation intensity balance and regulation edge weight delay, calculate regulation intensity balance parameters, match gene nodes according to the response blocks in the layer, and form a gene regulation abnormality marker layer.

5. The method for early obesity gene marker identification according to claim 4, characterized in that, The specific steps for obtaining the linked gene marker set are as follows: S311: Based on the gene regulation abnormality marker layer, extract the regulation intensity matrix and local functional features of the numbered gene nodes in the layer, align the data in the gene nodes with timestamps, identify the regulation intensity fluctuation value and functional mode stability offset, and obtain the gene local regulation response feature set. S312: Based on the gene local regulatory response feature set, jointly analyze the fluctuation of regulatory intensity and the stability of functional mode within the gene node, calculate the coupling feature value of regulatory behavior, screen gene units with regulatory behavior coupling degree in the layer, and establish a spatial distribution map of coordinated regulatory behavior response. S313: Call the spatial distribution map of the coordinated response of the regulatory behavior, cluster the gene nodes in the coupled feature value layer that exceed the coordinated recognition benchmark, label the node codes and coordinates corresponding to the continuous abnormal gene nodes, and obtain the set of linked gene markers.

6. The method for early obesity gene marker identification according to claim 5, characterized in that, The specific steps for obtaining the gene marker risk abnormality identification list are as follows: S411: Based on the position number of the linked gene marker set, extract the transcription density curve of the gene node under the specified number, perform time uniform processing, identify the density change per unit time, and obtain the gene regulation abnormal change rate set. S412: Based on the set of abnormal gene regulation change rates, identify the density distribution curve of the healthy gene stage, compare the current density change sequence with the reference curve, identify the gene regulation deviation level, extract and mark gene nodes whose deviation level exceeds the upper limit of the warning, and obtain the set of gene nodes with deviation surge. S413: Based on the set of deviation and spurious gene nodes, bind the deviation level value of each gene node to the position number in the genome dataset, sort them according to risk level, and output a list of gene marker risk anomalies.

7. The method for early obesity gene marker identification according to claim 1, characterized in that, The method also includes step S5: S5: Call the gene marker risk anomaly identification list, identify the corresponding number of gene node in the genome dataset, retrieve the preset intervention response unit list and gene protection priority sequence, screen the gene node numbers that need to adjust the response coverage, and output the gene intervention response adjustment instruction set; The gene intervention response adjustment instruction set includes adjusting the target gene node number, response level adjustment parameters, protection priority comparison items, and linkage intervention trigger types.

8. The method for early obesity gene marker identification according to claim 7, characterized in that, The specific steps for obtaining the gene intervention response adjustment instruction set are as follows: S511: Call the gene marker risk anomaly identification list, extract the gene node number in the genome dataset, map the gene node risk level value to the regional coordinate boundary, identify the gene node information corresponding to the risk grading standard, and generate a gene marker risk distribution map. S512: Based on the gene marker risk distribution map, extract the preset intervention strategy configuration parameters and response effect levels, match the gene marker risk levels with the intervention response levels, identify the unit numbers with insufficient response coverage, and obtain a gene marker response risk disconnect list. S513: Based on the gene marker response risk disconnect list, extract the key gene node numbers that need to improve the response coverage according to the priority sequence number of the gene protection sequence, output the adjustment control parameters linked with the original intervention unit in sequence, and output the gene intervention response adjustment instruction set.

9. An early obesity gene marker identification system, characterized in that, The system is used to implement the early obesity gene marker identification method according to any one of claims 1-8, the system comprising: The topology construction module is based on the topological structure of gene nodes and regulatory edges in the genome dataset. It compares the transcriptional level growth trend and the regulatory edge weight deviation value within the same time period, screens the synchronous mutation intervals of the two, extracts the gene node number and spatial coordinates, summarizes the abnormal time slices and gene node numbers, and generates a candidate gene marker set. Based on the candidate gene marker set, the gene localization module identifies the consistency between the direction of regulatory diffusion and the gradient vector of regulatory edge weights, marks the gene boundary numbers, matches the genome dataset, extracts the number range of gene nodes with abnormal regulatory diffusion, and establishes a gene regulation abnormality marker layer. Based on the gene regulation abnormality marker layer, the regulation linkage module retrieves the continuous data sequence of the regulation intensity matrix and local functional features in the region, determines the connectivity of the regulation abnormality boundary and the stability of the functional pattern, marks the gene node numbers that meet the linkage threshold of both, and outputs the linkage gene marker set. The risk warning module analyzes the transcription density trend of the corresponding gene node and the degree of deviation from the healthy gene transcription baseline curve based on the gene node number of the linked gene marker set, extracts the gene node number that deviates from the trend, completes the level identification according to the risk grading standard, and generates a gene marker risk abnormality identification list. Based on the gene marker risk anomaly identification list, the intervention optimization module finds the corresponding position number of the risk level gene node in the genome dataset, retrieves the current intervention response unit configuration list, compares the gene protection priority with the current response level, filters the gene node numbers that need to be updated, and outputs the gene intervention response adjustment instruction set.