Cancer gene key regulation module identification method, system, device and storage medium

By constructing a cancer gene regulatory network, introducing high-order topology modeling and random walk centrality indicators, and combining a dynamic redundancy removal strategy, the identification bias caused by redundancy between modules in traditional methods is solved. This enables accurate identification and evaluation of key regulatory modules in cancer networks, supporting disease mechanism analysis and treatment strategy design.

CN121034665BActive Publication Date: 2026-02-24TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511548381.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-24
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Traditional high-order centrality indicators fail to adequately consider redundancy and functional overlap between modules when identifying cancer gene regulatory modules, leading to ranking bias in key cluster identification and making it difficult to accurately identify functional core clusters in cancer networks.

Method used

By collecting patient tissue sample data, a cancer gene regulatory network was constructed. High-order topology modeling methods and high-order random walk centrality indicators were introduced, and combined with dynamic redundancy removal strategies, key regulatory modules were screened out, redundant clusters were removed, and the identification accuracy was improved.

Benefits of technology

It significantly improves the ability to identify key regulatory modules, enabling more accurate assessment of the importance of each regulatory module in cancer networks, and providing solid data support for disease mechanism analysis and network intervention strategy design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034665B_ABST
    Figure CN121034665B_ABST
Patent Text Reader

Abstract

The application provides a cancer gene key regulation module identification method, system, device and storage medium, comprising: collecting tissue sample data of a patient; performing transcriptome sequencing on the collected tissue samples respectively, and performing data processing on the obtained raw sequencing data to generate a gene node set; constructing a cancer gene regulation network based on the gene node set, and constructing a simplex network by introducing a high-order topological modeling method; performing cluster centrality evaluation on the high-order clusters of the simplex network based on a set high-order random walk centrality index, and dynamically removing redundant clusters from the modules based on a dynamic redundancy removal strategy, to output a key regulation module ranking result. The application can significantly improve the identification ability of the key regulation module, thereby effectively evaluating the importance of each regulation module in the cancer network, and providing solid data support for disease mechanism analysis and network intervention strategy design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of cancer bioinformatics technology, and in particular relates to a method, system, device and storage medium for identifying key regulatory modules of cancer genes. Background Technology

[0002] In tumor systems biology research, identifying key regulatory modules related to cancer is crucial for understanding disease mechanisms, discovering potential targets, and developing precision treatment strategies. In recent years, gene regulatory networks (GRNs) have become one of the core models in cancer bioinformatics research. However, compared to general biological networks, cancer regulatory networks often exhibit significant higher-order structural features—such as transcription factor regulatory modules, co-expressed gene sets, and protein complexes. These structures typically exist in the form of n-simplexes, highly overlapping with each other, collectively forming functional "hotspots" within the regulatory network.

[0003] Traditional high-order centrality metrics (such as HOC, HOD, and HOP) often employ static structure measurement methods, which frequently fail to adequately consider redundancy and functional overlap between modules, potentially leading to ranking biases in key cluster identification. For instance, in regulatory networks such as those for breast or lung cancer, there are numerous overlapping structures sharing core genes (such as TP53, MYC, and BRCA1). Without redundancy control, these structures can easily generate a centrality amplification effect, obscuring truly independent core clusters. Summary of the Invention

[0004] In view of this, this application aims to provide a method, system, device and storage medium for identifying key regulatory modules of cancer genes, in order to solve at least one of the above problems.

[0005] To achieve the above objectives, the technical solution of this application is implemented as follows:

[0006] In a first aspect, this application provides a method for identifying key regulatory modules of cancer genes, characterized by comprising:

[0007] Collect tissue sample data from patients, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples;

[0008] Transcriptome sequencing was performed on the collected tissue samples, and the raw sequencing data was processed to generate a set of gene nodes.

[0009] A cancer gene regulatory network is constructed based on the gene node set, and a simplex network is constructed by introducing a high-order topology modeling method.

[0010] The higher-order random walk centrality index is used to evaluate the clique centrality of the simplex network, and redundant clusters are dynamically removed from the modules based on a dynamic redundancy removal strategy to filter and output the ranking results of key control modules. The higher-order random walk centrality index determines the walk rules based on the defined transition probabilities and scores the modules according to the centrality quantification formula.

[0011] Secondly, based on the same inventive concept, this application also provides a cancer gene key regulatory module identification system, comprising:

[0012] The data acquisition module is configured to collect tissue sample data from patients, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples.

[0013] The data processing module is configured to perform transcriptome sequencing on the collected tissue samples and process the raw sequencing data to generate a set of gene nodes.

[0014] The network construction module is configured to construct a cancer gene regulatory network based on the set of gene nodes, and to construct a simplex network by introducing a high-order topology modeling method.

[0015] The output module is configured to evaluate the clique centrality of the higher-order cliques of the simplex network based on a set higher-order random walk centrality index, and to dynamically remove redundant clusters from the modules based on a dynamic redundancy removal strategy, so as to filter and output the ranking results of key control modules; wherein, the higher-order random walk centrality index determines the walk rules based on the defined transition probabilities and scores the modules according to the centrality quantification formula.

[0016] Thirdly, based on the same inventive concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect.

[0017] Fourthly, based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method as described in the first aspect.

[0018] Compared with existing technologies, the method, system, device, and storage medium for identifying key regulatory modules of cancer genes described in this application have the following advantages:

[0019] The method for identifying key regulatory modules of cancer genes described in this application can significantly improve the ability to identify key regulatory modules, thereby effectively assessing the importance of each regulatory module in the cancer network and providing solid data support for disease mechanism analysis and network intervention strategy design. Attached Figure Description

[0020] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0021] Figure 1 This is a flowchart of a method for identifying key regulatory modules of cancer genes according to an embodiment of this application;

[0022] Figure 2 This is a flowchart of the random walk strategy described in the embodiments of this application;

[0023] Figure 3 This is a flowchart of the dynamic redundancy removal strategy described in the embodiments of this application;

[0024] Figure 4 This is a schematic diagram of the structure of a cancer gene key regulatory module identification system according to an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of the hardware structure of the electronic device described in an embodiment of this application. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0027] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0028] The embodiments of this application are described in detail below with reference to the accompanying drawings.

[0029] To address the problems identified in the background section, this embodiment proposes a method for identifying key regulatory modules of cancer genes. This method introduces a structural overlap threshold by simulating the dynamic access frequency between structural clusters, identifying and eliminating repetitive regulatory modules with excessive redundancy, thereby effectively suppressing centrality distortion caused by structural redundancy. Compared to traditional methods, this method can more accurately and stably characterize the functional core regions in cancer gene regulatory networks, providing a more biologically explanatory basis for subsequent mechanism research and target screening.

[0030] Please see Figure 1 As shown, this identification method specifically includes the following steps:

[0031] Step S101: Collect tissue sample data from the patient, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples.

[0032] Specifically, in this embodiment, ethically approved tissue sample data are collected from clinical partner hospitals or public biobanks, covering three groups:

[0033] Normal tissue sample (no cancer);

[0034] Early-stage cancer tissue sample (initial lesion);

[0035] Tissue samples from advanced cancer (where the tumor has spread or invaded).

[0036] It should be noted that all samples should come from the same type of cancer and have a unified clinical stage label to ensure the scientific rigor and controllability of the modeling comparison.

[0037] Step S102: Transcriptome sequencing is performed on the collected tissue samples, and the raw sequencing data is processed to generate a set of gene nodes.

[0038] Specifically, in this embodiment, high-throughput RNA sequencing (RNA-seq) is performed on the three types of samples in step S101 to obtain raw sequencing data, and the raw sequencing data is then processed. The data processing procedure includes the following steps:

[0039] Step S201: Perform quality control (FastQC) on the raw sequencing data to evaluate indicators such as base quality, adapter contamination, and GC content;

[0040] Step S202: Clean the data obtained in step S201, remove adapter sequences and low-quality regions, and generate high-quality sequencing fragments.

[0041] Step S203: Align the fragments obtained in step S202 to the human reference genome, and calculate the expression level of each gene using featureCounts (a quantitative tool used to count the number of reads for each gene), and normalize it to TPM or FPKM format;

[0042] Step S204: Use DESeq2 to perform differential expression analysis on samples from different groups and screen for significantly differentially expressed genes with FDR < 0.05 and |log2FC| > 1;

[0043] Step S205: Using the JASPAR database, identify known transcription factors and their regulatory target genes among differentially expressed genes, and extract regulatory-related nodes;

[0044] Step S206: Finally, a set of nodes is formed. This is used to construct subsequent cancer gene regulatory networks.

[0045] Step S103: Construct a cancer gene regulatory network based on the gene node set, and construct a simplex network by introducing a high-order topology modeling method.

[0046] Specifically, in this embodiment, the node set generated based on step S102... Construct an edge set And assign weights The specific steps are as follows:

[0047] Step S301: For any gene node pair An undirected edge is established between nodes if any of the following conditions are met:

[0048] Regulation relationship edge: if node It is a transcription factor, and its counterparts exist in the database. If there is evidence of regulation, then establish the regulatory boundary;

[0049] Shared expression relationship edge: If the Pearson correlation coefficient between two nodes in the expression matrix satisfies Then, establish a shared edge;

[0050] If none of the above conditions are met, no edge is created;

[0051] Step S302, Edge Weights The biological strength and expression confidence of an edge are defined as follows:

[0052] If a regulatory relationship exists, then ,( To adjust the bias term, this embodiment uses 0.2).

[0053] otherwise, ;

[0054] Step S303: Finally, construct a weighted undirected graph based on the above edge set and weights. ,say This represents the cancer gene regulatory network of the corresponding sample.

[0055] The development and progression of cancer involves the coordinated regulation of multiple genes, transcription factors, and signaling pathways. In the tumor microenvironment, regulatory factors often do not function independently, but rather form regulatory clusters in a modular and highly coordinated manner, jointly driving carcinogenic processes such as abnormal cell proliferation, apoptosis escape, and immune escape. Therefore, traditional network modeling methods that characterize gene pair relationships on an edge-by-edge basis are insufficient to reveal the complex coordinated structures in cancer regulation.

[0056] To delve deeper into these higher-order regulatory structures (such as transcription factor co-expression complexes, co-expression modules, and protein complexes), this embodiment is based on an existing cancer gene regulatory network. We introduce a high-order topology modeling method to construct a simplex network and identify the potential functional modules in the network. The specific method is as follows:

[0057] Step S401, from the diagram The group that identifies all levels is a fully connected subgraph consisting of several regulatory gene nodes, representing a possible co-regulatory module.

[0058] Step S402: Form a simplex set from all cliques of different orders, and further construct a simplex complex structure to express the high-order topological nesting relationships in the cancer regulation network. Define the simplex structure of the cancer regulation network as follows:

[0059]

[0060] In the formula, Indicates all A set of order groups, that is, a set of individual gene regulatory nodes; Indicates all The set of order groups, by A set of fully connected nodes represents a collection of collaboratively regulated structures (such as cancer-driving modules, protein-protein interaction complexes, etc.). Represents all in the network The number of order groups.

[0061] Step S403: Define the topological inclusion relationships and correlation matrices between cliques of different orders, specifically including:

[0062] Define the hierarchical inclusion relationships between different order groups to construct an association matrix for describing cancer gene regulatory networks. Groups and The topological inclusion relationship between order groups is expressed by the following formula:

[0063]

[0064] In the formula, Represents all in the network The number of tier groups, This represents the correlation matrix.

[0065] in, The matrix elements representing the incidence matrix, indicating Group Is it included? Group In this context, it is defined as:

[0066] .

[0067] Step S404, Construction Adjacency matrix between order cliques This is used to characterize the degree of structural overlap between control modules, and its formula is as follows:

[0068]

[0069] In the formula, Represents all in the network The number of strata;

[0070] in, Adjacency matrix The matrix elements in the matrix represent the clique Hetuan The number of overlapping gene nodes between them is defined as:

[0071] .

[0072] Step S104: Based on the set higher-order random walk centrality index, evaluate the clique centrality of the higher-order cliques of the simplex network, and dynamically remove redundant clusters from the modules based on the dynamic redundancy removal strategy to filter and output the ranking results of key control modules; wherein, the higher-order random walk centrality index determines the walk rules based on the defined transition probability, and scores the modules according to the centrality quantification formula.

[0073] Specifically, in this embodiment, the importance of key regulatory modules in a cancer network should not only depend on their local connectivity structure, but also on their "dynamic access potential" within the entire network. To this end, this embodiment introduces a high-order walk centrality index based on high-order topology. The importance of the group is assessed globally, such as... Figure 2 As shown, the specific method is as follows:

[0074] Step S501: Define the transition probability

[0075] From the current regulatory module, i.e., group Jump to adjacent module group The probability of them being related is directly proportional to the number of genes they share:

[0076]

[0077] Example: If There are 3 adjacent modules , , ,in, and They share two genes. and They share two genes. and If they share one gene, then , , .

[0078] Step S502, Random walk strategy

[0079] Initialization parameters: =Step size, =Number of walks;

[0080] Randomly select a starting module ;

[0081] Walk rule: Each step is based on the transition probability. Select the next module;

[0082] Step length control: Wandering One step is one round;

[0083] Frequency statistics: Execution Each visit counts the number of times each module is accessed. .

[0084] Step S503, Centrality Quantification

[0085] Module Importance score is defined as:

[0086] This is to reflect its global reachability in higher-order control networks.

[0087] Step S504: Sorting of Control Modules

[0088] All modules are sorted by score Sort in descending order to obtain the list of candidate key modules. .

[0089] Because cancer regulation modules often contain a large number of highly overlapping sub-clusters, which may lead to analytical redundancy and functional duplication, this embodiment designs a dynamic redundancy removal strategy to screen for a representative set of non-redundant modules, such as... Figure 3 As shown, the specific steps include the following:

[0090] Step S601, Input parameters

[0091] The module list obtained from step S504 The format is Threshold : Upper limit of node overlap ratio.

[0092] Step S602, Initialization

[0093] empty set : Module used to store the filtered and retained items.

[0094] Step S603, Module Filtering Strategy

[0095] Traversal Each module in Calculation module and Maximum overlap rate of selected modules:

[0096]

[0097] like If so, then the module will be retained and updated. ;

[0098] Otherwise, it will be removed.

[0099] In addition, to improve screening efficiency, this embodiment also provides an optimization method, namely, initializing an empty set. (Create an empty collection) = (used to store reserved module nodes), executes a module filtering strategy for each module to be processed. The following judgment process will be executed:

[0100] like Keep this module directly: Update the set of covered nodes: ← ;

[0101] like If so, the maximum overlap rate is calculated for further filtering.

[0102] Step S604: Output and analyze the results.

[0103] Final output module set This refers to the key non-redundant regulatory structures in the cancer network. By combining clinical analysis and differential expression analysis results, we screened high-order key regulatory nodes that are closely related to cancer progression.

[0104] This embodiment also includes a clinical analysis based on the above content, as detailed below:

[0105] Clinical correlation analysis of the key regulatory module ranking results obtained in step S604 revealed significant differences in the importance ranking of key regulatory modules in the cancer gene regulatory networks constructed from normal tissues, early-stage cancer tissues, and late-stage cancer tissues (denoted as G1, G2, and G3, respectively). For example, in G1, most higher-order regulatory modules were concentrated in gene clusters related to cell cycle regulation and DNA repair, with high centrality scores. In G2, some modules involved in the regulation of inflammatory factors or immune signaling pathways saw an increase in ranking, suggesting that early-stage carcinogenesis may be accompanied by the initial expansion of regulatory abnormalities. Furthermore, in G3, several regulatory modules related to tumor metastasis, angiogenesis, and apoptosis escape (such as modules containing key nodes like VEGFA, MMP9, and TP53) were assigned higher ranking scores.

[0106] This difference indicates that as cancer progresses, key functional regions of the gene regulatory network undergo structural remodeling and functional shift, particularly a change in regulatory focus from homeostasis maintenance to enhanced tumor-specific regulation. Further ranking, comparison, and statistical analysis of multiple samples, combined with clinical labels and patient outcome data, are expected to aid in tumor staging, identification of key pathological mechanisms, and target screening, providing fundamental data support for personalized precision medicine.

[0107] This application combines the biological structural characteristics of cancer gene regulatory networks, utilizes the regulatory relationships of regulatory factors and co-expression networks, and is based on... By constructing high-order simplex network structures using clusters, combining topological overlap relationships between nodes, and capturing the access frequency of high-order structures by modeling inter-cluster transition probabilities, and dynamically eliminating redundant clusters based on set overlap thresholds, the ability to identify key regulatory modules can be significantly improved. This effectively assesses the importance of each regulatory module in cancer networks, providing solid data support for disease mechanism analysis and network intervention strategy design.

[0108] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0109] Based on the same inventive concept, and corresponding to the methods of any of the above embodiments, the embodiments of this application also provide a cancer gene key regulatory module identification system.

[0110] like Figure 4 As shown, the cancer gene key regulatory module identification system includes:

[0111] The data acquisition module 11 is configured to collect tissue sample data from patients, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples.

[0112] The data processing module 12 is configured to perform transcriptome sequencing on the collected tissue samples and process the raw sequencing data to generate a set of gene nodes.

[0113] Network building module 13 is configured to build a cancer gene regulatory network based on a set of gene nodes, and to build a simplex network by introducing a high-order topology modeling method.

[0114] The output module 14 is configured to evaluate the clique centrality of the high-order cliques of the simplex network based on a set high-order random walk centrality index, and to dynamically remove redundant clusters from the modules based on a dynamic redundancy removal strategy, so as to filter and output the ranking results of key control modules; wherein, the high-order random walk centrality index determines the walk rules based on the defined transition probability and scores the modules according to the centrality quantification formula.

[0115] For ease of description, the above system is described by dividing it into various modules based on their functions. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.

[0116] The system described in the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0117] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.

[0118] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0119] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0120] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0121] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0122] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0123] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0124] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0125] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0126] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.

[0127] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0128] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0129] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.

[0130] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0131] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0132] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for identifying a cancer gene key regulatory module, characterized in that, include: Collect tissue sample data from patients, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples; Transcriptome sequencing was performed on the collected tissue samples, and the raw sequencing data was processed to generate a set of gene nodes. A cancer gene regulatory network is constructed based on the gene node set, and a simplex network is constructed by introducing a high-order topology modeling method. The higher-order random walk centrality index is used to evaluate the clique centrality of the simplex network, and redundant clusters are dynamically removed from the modules based on a dynamic redundancy removal strategy to filter and output the ranking results of key control modules. The higher-order random walk centrality index determines the walk rules based on the defined transition probabilities and scores the modules according to the centrality quantification formula. The dynamic redundancy removal strategy includes: input candidate key regulatory module list and a threshold value ; Initialize empty set ; traversing each module in the list of candidate key regulatory modules calculating a maximum overlap rate of the module with the selected modules in the list of candidate key regulatory modules, wherein the maximum overlap rate is calculated according to the following formula:​​   ; in response to then the retention and update module, including: ; Final output module set .

2. The method of claim 1, wherein, The process involves performing transcriptome sequencing on the collected tissue samples and processing the raw sequencing data to form a set of gene nodes, including: High-throughput RNA sequencing was performed on the collected normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples to obtain the corresponding raw sequencing data. The raw sequencing data is subjected to quality control and data cleaning to obtain high-quality sequencing fragments; The sequencing fragments were aligned to the human reference genome, gene expression levels were calculated, and differential expression analysis was used to screen different raw sequencing data to identify significantly differentially expressed genes. The JASPAR database was used to identify known transcription factors and their regulatory target genes in differentially expressed genes, extract regulatory-related nodes, and generate a set of gene nodes.

3. The method of claim 1, wherein, The method for constructing the cancer gene regulatory network includes: An edge set is constructed from the generated set of gene nodes and weights are assigned. An undirected graph is constructed based on the edge set and the weights. The edge set includes regulatory relationship edges and co-expression relationship edges.

4. The method of claim 3, wherein, The construction of the cancer gene regulatory network based on the gene node set, including the construction of a simplex network by introducing a higher-order topology modeling method, includes: Identify cliques of all orders from an undirected graph, form a simplex set of all cliques of different orders, and construct a simplex network, wherein the simplex network includes an incidence matrix and an adjacency matrix, including: Define the hierarchical inclusion relationships between groups of different orders and construct an association matrix, the formula of which is: ; wherein denotes the number of all clusters of order k in the network, denotes the number of all clusters of order k in the network, denotes the incidence matrix; wherein is a matrix element of the association matrix, indicating whether the order cluster is contained in the order cluster , which is defined as: ; Constructing An adjacency matrix between clusters, where the adjacency matrix is formulated as: ; wherein represents the number of clusters in the network represents the number of clusters in the network represents the adjacency matrix wherein the elements are the matrix elements in the adjacency matrix, representing the number of gene nodes between the clusters , which is defined as: 。 5. The method according to claim 4, characterized in that, The evaluation of the clique centrality of the simplex network based on the set higher-order random walk centrality index includes: Initialize parameters, including walking step size and number of walks; A starting module is randomly selected, and the next module is selected according to the defined transition probability. A random walk is performed according to the walk step size and the number of walks, and the number of visits is recorded. The centrality score of each module is calculated using the centrality quantification formula, and all modules are sorted in descending order according to their centrality scores to obtain a list of candidate key regulatory modules.

6. The method according to claim 5, characterized in that, The transition probability formula is: ; In the formula, This represents the transition probability value. This represents a matrix element in the adjacency matrix.

7. A system for identifying key regulatory modules of cancer genes, characterized in that, include: The data acquisition module is configured to collect tissue sample data from patients, wherein the tissue sample data includes at least normal tissue samples, early-stage cancer tissue samples, and late-stage cancer tissue samples. The data processing module is configured to perform transcriptome sequencing on the collected tissue samples and process the raw sequencing data to generate a set of gene nodes. The network construction module is configured to construct a cancer gene regulatory network based on the set of gene nodes, and to construct a simplex network by introducing a high-order topology modeling method. The output module is configured to evaluate the clique centrality of the higher-order cliques of the simplex network based on a set higher-order random walk centrality index, and to dynamically remove redundant clusters from the modules based on a dynamic redundancy removal strategy, so as to filter and output the ranking results of key control modules; wherein, the higher-order random walk centrality index determines the walk rules based on the defined transition probabilities and scores the modules according to the centrality quantification formula. The dynamic redundancy removal strategy includes: Input a list of candidate key control modules and threshold ; Initialize an empty set ; Traverse the list of candidate key control modules Each module in Calculate its relationship with The maximum overlap rate of the selected modules is given by the formula:   ; In response to Then the module will be retained and updated, including: ; Final output module set .

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1-6.

9. A non-transitory computer-readable storage medium, characterized in that, in, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to perform the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Construction method and system for reproductive medicine database

    CN119964656A

  • Deep learning cancer risk prediction method and system based on methylation sequence

    CN120148653A