Multi-target multi-task optimization method for identifying network biomarkers of individual patients

By constructing multi-objective multi-task optimization problems, combining dynamic network marker theory and multi-objective network control theory, identifying disease signals and drug targets of cancer individuals, the problem of difficulty in efficient identification at the same time in the existing technology is solved, and higher recognition accuracy and knowledge transfer effect are achieved.

CN120030870AActive Publication Date: 2025-05-23ZHENGZHOU UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411869938.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-23
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing cancer individual biomarker recognition technologies are difficult to efficiently identify disease signals and drug targets simultaneously, and existing methods have problems with inefficiency in personalized dynamic data analysis.

Method used

Dynamic network marker theory and multi-objective network control theory are adopted to construct an unconstrained disease signal detection model and a constrained drug target recognition model, and construct it as a multi-objective multi-task optimization problem. Two types of markers are identified through multi-objective multi-task optimization method, and useful knowledge between the two models is explored and utilized to improve the identification accuracy of markers.

Benefits of technology

By modeling disease signal detection and drug target recognition as multi-task optimization problems, the algorithm's solution effect on two problems is improved, the marker recognition accuracy is improved, and the effectiveness of knowledge transfer is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030870A_ABST
    Figure CN120030870A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target multi-task optimization method for identifying network biomarkers of an individual patient, which comprises the following steps of: simultaneously identifying a drug target marker and a disease signal marker from a gene network of the individual patient, and forming a multi-target multi-task optimization model according to a dynamic network marker theory and a multi-target network control theory; and then identifying the network biomarker by adopting a multi-target multi-task evolutionary algorithm. According to the invention, a dynamic network marker theory and a multi-target network control theory are utilized to respectively construct an unconstrained disease signal detection model and a constrained drug target identification model, a multi-target and multi-task optimization problem is further constructed, and then the multi-target and multi-task optimization method is utilized to identify two types of markers. Useful knowledge between the two models can be explored and utilized, and meanwhile the recognition precision of the two types of markers is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cancer individual biomarker identification, and in particular to a multi-objective and multi-task optimization method for identifying individual patient network biomarkers. Background Art

[0002] Cancer is a leading cause of death. Identifying disease signals and developing new drugs are two important research tasks. For the task of disease signal detection, personalized dynamic marker theory can use single-sample network technology to utilize the high-dimensional dynamic data of individual patients, thus avoiding the difficulty of research and analysis brought by large-scale data. Specifically, the patient's genetic data is constructed into a personalized gene interaction network (PGIN), and then a set of genes is identified from the PGIN as markers by analyzing the interaction between genes. These markers can be further used to analyze disease signals. For drug development, an important task is to identify drug targets. Some network-based methods identify markers from PGIN and further identify targets through drug activation. Among them, structured network control criteria is a popular method because it can accurately describe how to control the transition of PGIN between normal and disease states. Summary of the invention

[0003] In order to solve the above problems, the present invention utilizes dynamic network marker theory and multi-objective network control theory to respectively construct an unconstrained disease signal detection model and a constrained drug target recognition model, and further constructs them into a multi-objective multi-task optimization problem. Then, the multi-objective multi-task optimization method of the invention is used to identify the two types of markers, which can explore and utilize the useful knowledge between the two models and improve the recognition accuracy of the two types of markers.

[0004] The technical solution adopted by the present invention is: a multi-objective and multi-task optimization method for identifying individual patient network biomarkers, comprising the following steps:

[0005] S1, the disease signal detection task and drug target identification task are constructed as a multi-objective and multi-task optimization problem, the formula is:

[0006] {x 1 ,x 2 ,…,x k}=arg min(F 1 (x 1 ),F 2 (x 2 ),…,F k (x k ))

[0007] sG 1 (x 1 ),G2 (x 2 ),…,G k (x k )

[0008] x k ={x k,1 ,x k,2 ,…,x k,NP},k∈{1,2,…,K}

[0009] Where: k is the number of optimization tasks; F k is the objective function of the kth optimization task, G k is the constraint condition of the kth optimization task, x k represents the decision variable of the kth optimization task, which contains NP individuals;

[0010] The disease signal detection task T 1 The formula is:

[0011]

[0012] Where: f 1 (x) represents the minimum number of nodes; f 2 (x) represents the minimization form of maximizing the number of prior nodes; V represents the number of nodes in PGIN; V p represents the number of prior nodes in PGIN; x 1,i is a 0 / 1 variable. If the i-th node is selected, then x i is equal to 1, otherwise it is equal to 0; l i is an indicator variable. If the i-th node is a priori node, then l i is equal to 1, otherwise it is equal to 0.

[0013] Drug target identification task T 2 The formula is:

[0014]

[0015] In the minimum dominating set model (MDS), represents the index of the neighbor node of the ith node; if the ith node and the jth node are connected by an edge, then the jth node is called the neighbor of the ith node; in the Nonlinear Control Model of Undirected Network Algorithms (NCUA), |E| represents the number of edges in PGIN, e j represents the jth edge in E; when the selected node is j When connected, e j is equal to 1, otherwise it is equal to 0;

[0016] S2, randomly initialize (NP-2) individuals as P11 ;

[0017] S3, generate two single-objective optimal solutions as P 12 ;

[0018] S4, Merge P 11 and P 12 , as T 1 The initial population P 1 , and at T 1 Evaluation P 1 ;

[0019] S5, randomly generate three populations of size NP, respectively as T 2 The three initial populations P 2 , P 21 and P 22 , and at T 2 Three populations were evaluated above; among them, P 2 For optimizing two objectives and constraints, P 21 For optimizing the first objective and constraint, P 22 Used to optimize the second objective and constraints;

[0020] S6, setting the algebraic counter Cnt equal to 1;

[0021] S7, P 1 Use the binary genetic algorithm to generate NP / 2 offspring individuals, denoted as OP 1 , and at T 1 Review OP 1 ;

[0022] S8,P 2 Produce offspring, recorded as OP 2 ;

[0023] S9,P 21 and P 22 Use the binary genetic algorithm to generate NP / 4 offspring individuals, denoted as OP 21 and OP 22 ;

[0024] S10, P 1 Use the unconstrained non-dominated selection strategy to select from the set {P 1 , OP 1} select NP individuals as the new P 1 ;

[0025] S11, evaluate OP on T2 2 , OP 21 and OP 22 ;

[0026] S12, P2 Using the epsilon method from the set {P 2 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 2 ;

[0027] S13, P 21 Using the epsilon method from the set {P 21 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 21 , where all individuals' f 2 The value is set to 0;

[0028] S14, P 22 Using the epsilon method from the set {P 22 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 22 , where all individuals' f 1 The value is set to 0;

[0029] S15, Cnt=Cnt+1;

[0030] S16: If the maximum number of evaluations is reached, the algorithm stops and outputs P 1 and P 2 ; Otherwise, jump to S7.

[0031] Furthermore, the specific steps of S3 include:

[0032] S31, generates a |V|-dimensional row vector of all zeros, denoted as A, which is the optimal solution to the first objective;

[0033] S32, generates a |V|-dimensional row vector of all zeros, denoted as B, where all dimensional values ​​corresponding to the prior nodes are set to 1, which is the optimal solution for the second objective.

[0034] Further, the specific steps of S8 are:

[0035] S81, if the remainder of Cnt divided by 30 is 0, jump to S82, otherwise jump to S83;

[0036] S82, using the infeasible solution repair strategy to generate the offspring population OP 2 ;

[0037] S83,P 2Use the binary genetic algorithm to generate NP / 2 offspring individuals, denoted as OP 2 .

[0038] Further, the specific steps of S82 are:

[0039] S821, Random from P 1 Select NP / 2 individuals from the , denoted as ST 1 ;

[0040] S822, if the optimized network model is MDS, jump to S823; if the optimized network model is NCUA, jump to S828;

[0041] S823, set counter i=1;

[0042] S824, for ST 1 The i-th individual in is denoted as ST 1,i ;

[0043] S825, for any j∈{1,2,…,|V|}, if ST 1,i The node corresponding to the j-th dimension variable of is selected or the neighbor node of the node corresponding to the j-th dimension is selected, then ST 1 The j-th dimension variable does not change; otherwise, randomly select a node from the j-th dimension corresponding node and its neighbors, and put ST 1 The dimension corresponding to the newly selected node in is set to 1;

[0044] S826, update counter i=i+1;

[0045] S827, if i>NP, then output ST 1 , otherwise jump to S824;

[0046] S828, set counter i=1;

[0047] S829, for ST 1 The i-th individual in is denoted as ST 1,i , set j = 1;

[0048] S8210, for any j∈{1,2,…,|E|}, if ST 1,i If neither of the two nodes connected by the jth edge in the corresponding PGIN is selected, a node is randomly selected and its corresponding dimension is set to 1;

[0049] S8211, update counter i=i+1;

[0050] S8212, if i>NP, then output ST 1 , otherwise jump to S829.

[0051] The beneficial effects of the present invention are:

[0052] (1) Model the disease signal detection and drug target identification problems as multi-task optimization problems, and improve the algorithm's solution performance on both problems by transferring knowledge between the two tasks;

[0053] (2)T 1 The two single-objective optimal solutions constructed in improve the diversity of the algorithm on this task;

[0054] (3)T 2 The infeasible solution repair strategy set in T 1 to T 2 effectiveness of knowledge transfer. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flow chart of the method of the present invention;

[0056] Figure 2 It is a comparison chart of evaluation indicators of early warning scores of the present invention and the unconstrained multi-objective algorithm;

[0057] Figure 3 It is a comparison chart of the evaluation index of the area under the curve of the present invention and the constrained multi-objective algorithm;

[0058] Figure 4 It is a comparison chart of evaluation indicators of early warning scores of the present invention and the multi-objective multi-task algorithm;

[0059] Figure 5 It is a comparison chart of the evaluation index of the area under the curve of the present invention and the multi-objective multi-task algorithm. DETAILED DESCRIPTION

[0060] The present invention will be further described below in conjunction with the accompanying drawings.

[0061] S1, the disease signal detection task and drug target identification task are constructed as a multi-objective and multi-task optimization problem, the formula is:

[0062] {x 1 ,x 2 ,…,x k}=arg min(F 1 (x 1 ),F 2 (x 2 ),…,F k (x k ))

[0063] sG 1 (x 1 ),G 2 (x2 ),…,G k (x k )

[0064] x k ={x k,1 ,x k,2 ,…,x k,NP},k∈{1,2,…,K}

[0065] Where: k is the number of optimization tasks; F k is the objective function of the kth optimization task, G k is the constraint condition of the kth optimization task, x k represents the decision variables of the kth optimization task, which contains NP individuals.

[0066] The disease signal detection task T 1 The formula is:

[0067]

[0068] Where: f 1 (x) represents the minimum number of nodes, f 2 (x) represents the minimization form of maximizing the number of prior nodes, V represents the number of nodes in PGIN, and V p represents the number of prior nodes in PGIN, x 1,i is a 0 / 1 variable. If the i-th node is selected, then x i is equal to 1, otherwise it is equal to 0; l i is an indicator variable. If the i-th node is a priori node, then l i is equal to 1, otherwise it is equal to 0.

[0069] Drug target identification task T 2 The formula is:

[0070]

[0071] In the minimum dominating set model MDS, represents the index of the neighbor node of the ith node; if the ith node and the jth node are connected by an edge, then the jth node is called the neighbor of the ith node; in the nonlinear control model NCUA of undirected network algorithms, |E| represents the number of edges in PGIN, e j represents the jth edge in E; when the selected node is j When connected, e j is equal to 1, otherwise it is equal to 0;

[0072] S2, randomly initialize (NP-2) individuals as P 11 .

[0073] S3, generate two single-objective optimal solutions as P 12 ; The specific steps include:

[0074] S31, generates a |V|-dimensional row vector of all zeros, denoted as A, which is the optimal solution to the first objective.

[0075] S32, generates a |V|-dimensional row vector of all zeros, denoted as B, where all dimensional values ​​corresponding to the prior nodes are set to 1, which is the optimal solution for the second objective.

[0076] S4, Merge P 11 and P 12 , as T 1 The initial population P 1 , and at T 1 Evaluation P 1 .

[0077] S5, randomly generate three populations of size NP, respectively as T 2 The three initial populations P 2 , P 21 and P 22 , and at T 2 Three populations were evaluated above; among them, P 2 For optimizing two objectives and constraints, P 21 For optimizing the first objective and constraint, P 22 Used to optimize the second objective and constraints.

[0078] S6, set the algebraic counter Cnt equal to 1.

[0079] S7, P 1 Use the binary genetic algorithm to generate NP / 2 offspring individuals, denoted as OP 1 , and at T 1 Review OP 1 .

[0080] S8,P 2 Produce offspring, recorded as OP 2 ; The specific steps include:

[0081] S81, if the remainder when Cnt is divided by 30 is 0, jump to S82, otherwise jump to S83.

[0082] S82, using the infeasible solution repair strategy to generate the offspring population OP 2 ; The specific steps are:

[0083] S821, Random from P 1 Select NP / 2 individuals from the , denoted as ST 1 .

[0084] S822, if the optimized network model is MDS, jump to S823; if the optimized network model is NCUA, jump to S828.

[0085] S823, set counter i=1.

[0086] S824, for ST 1 The i-th individual in is denoted as ST 1,i .

[0087] S825, for any j∈{1,2,…,|V|}, if ST 1,i The node corresponding to the j-th dimension variable of is selected or the neighbor node of the node corresponding to the j-th dimension is selected, then ST 1 The j-th dimension variable does not change; otherwise, randomly select a node from the j-th dimension corresponding node and its neighbors, and put ST 1 The dimension corresponding to the newly selected node in is set to 1.

[0088] S826, update counter i=i+1.

[0089] S827, if i>NP, then output ST 1 , otherwise jump to S824.

[0090] S828, set counter i=1;

[0091] S829, for ST 1 The i-th individual in is denoted as ST 1,i , set j=1.

[0092] S8210, for any j∈{1,2,…,|E|}, if ST 1,i If neither of the two nodes connected by the j-th edge in the corresponding PGIN is selected, a node is randomly selected and its corresponding dimension is set to 1.

[0093] S8211, update counter i=i+1.

[0094] S8212, if i>NP, then output ST 1 , otherwise jump to S829.

[0095] S83,P 2 Use the binary genetic algorithm to generate NP / 2 offspring individuals, denoted as OP 2 .

[0096] S9,P 21 and P 22 Use the binary genetic algorithm to generate NP / 4 offspring individuals, denoted as OP21 and OP 22 .

[0097] S10, P 1 Use the unconstrained non-dominated selection strategy to select from the set {P 1 , OP 1} select NP individuals as the new P 1 .

[0098] S11, in T 2 Review OP 2 , OP 21 and OP 22 .

[0099] S12, P 2 Using the epsilon method from the set {P 2 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 2 .

[0100] S13, P 21 Using the epsilon method from the set {P 21 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 21 , where all individuals' f 2 The value is set to 0.

[0101] S14, P 22 Using the epsilon method from the set {P 22 , OP 2 , OP 21 , OP 22} select NP individuals as the new P 22 , where all individuals' f 1 The value is set to 0.

[0102] S15, update the algebraic counter Cnt=Cnt+1.

[0103] S16: If the maximum number of evaluations is reached, the algorithm stops and outputs P 1 and P 2 ; Otherwise, jump to S7.

[0104] In order to verify the effectiveness of the present invention, an experiment was performed. The experimental parameters were set as follows: the population size NP was set to 100, the maximum number of evaluations for the first task was 50,000, and the maximum number of evaluations for the second task was 100,000. The experimental data were breast cancer data BRCA and lung cancer data LUND, which contained 112 and 106 patient data, respectively. There are two basic PGIN models, MDS and NCUA.

[0105] The comparative experiment consists of three parts:

[0106] In the first part, the optimization method of the present invention is compared with the unconstrained multi-objective algorithm to verify its effectiveness in T 1 The comparison algorithms are DEAGNG, MOEAD, MSEA, DAEA and PREA.

[0107] The comparison results are as follows Figure 2 As shown, the evaluation index is the warning score F p The larger the value, the better the algorithm can identify the critical state of the disease. As can be seen from the figure, the multi-task multi-objective network biomarker recognition method MMUR adopted by the present invention has achieved significantly better F on both data and models. p value, proving its effectiveness.

[0108] In the second part, the optimization method of the present invention is compared with the constrained multi-objective algorithm to verify its effectiveness in T 2 The comparison algorithms are CCMO, cDPEA, ICMA, IMTCMO, and LSCV_MCEA. The comparison results are as follows Figure 3 As shown in the figure, the evaluation index is the area under the curve (AUC), which evaluates the proportion of drug targets found by the algorithm in the database. The larger the AUC, the better the algorithm performance. Figure 3 It can be seen that MMUR achieved the largest AUC value in all cases, proving its effectiveness.

[0109] In the third part, the optimization method of the present invention is compared with the multi-objective multi-task algorithm to verify its effectiveness on the proposed multi-task model. The comparison algorithms are EMT-ET, MO-EMEA, MO-MFEA, and MO-MFEA-II. The comparison results are shown in Figure 2. Figure 4 and Figure 5 As shown in the figure, the evaluation indicators are warning score F p and area under the curve AUC. The results show that compared with the multi-objective multi-task comparison algorithm, the proposed method MMUR can achieve effective knowledge transfer and improve the effectiveness of the two tasks.

Claims

1. A multi-objective and multi-task optimization method for identifying individual patient network biomarkers, characterized in that: The following steps are involved: S1, the disease signal detection task and drug target identification task are constructed as a multi-objective and multi-task optimization problem, the formula is: {x1,x2,…,x k }=arg min(F1(x1),F2(x2),…,F k ( x k )) s.t.G1(x1),G2(x2),…,G k (x k ) x k ={x k,1 ,x k,2 ,…,x k,NP },k∈{1,2,…,K} Where: k is the number of optimization tasks; F k is the objective function of the kth optimization task, G k is the constraint condition of the kth optimization task, x k represents the decision variable of the kth optimization task, which contains NP individuals; The formula for the disease signal detection task T1 is: Where: f1(x) represents the minimization of the number of nodes; f2(x) represents the minimization form of maximizing the number of prior nodes; V represents the number of nodes in PGIN; V p represents the number of prior nodes in PGIN; x 1,i is a 0 / 1 variable. If the i-th node is selected, then x i is equal to 1, otherwise it is equal to 0; l i is an indicator variable. If the i-th node is a priori node, then l i is equal to 1, otherwise it is equal to 0. The formula for drug target identification task T2 is: In the minimum dominating set model (MDS), represents the index of the neighbor node of the ith node; if the ith node and the jth node are connected by an edge, then the jth node is called the neighbor of the ith node; in the Nonlinear Control Model of Undirected Network Algorithms (NCUA), |E| represents the number of edges in PGIN, e j represents the jth edge in E; when the selected node is j When connected, e j is equal to 1, otherwise it is equal to 0; S2, randomly initialize (NP-2) individuals as P 11 ; S3, generate two single-objective optimal solutions as P 12 ; S4, Merge P 11 and P 12 , as the initial population P1 of T1, and evaluate P1 on T1; S5 randomly generates three populations of size NP, which are used as the three initial populations P2, P 21 and P 22 , and evaluate three populations on T2; P2 is used to optimize two objectives and constraints, P 21 For optimizing the first objective and constraint, P 22 Used to optimize the second objective and constraints; S6, setting the algebraic counter Cnt equal to 1; S7, P1 uses binary genetic algorithm to generate NP / 2 offspring individuals, recorded as OP1, and evaluates OP1 on T1; S8, P2 produces offspring, recorded as OP2; S9,P 21 and P 22 Use the binary genetic algorithm to generate NP / 4 offspring individuals, denoted as OP 21 and OP 22 ; S10, P1 uses an unconstrained non-dominated selection strategy to select NP individuals from the set {P1, OP1} as the new P1; S11, evaluate OP2 on T2, OP 21 and OP 22 ; S12, P2 uses the epsilon method to select from the set {P2, OP2, OP 21 , OP 22 }Select NP individuals as the new P2; S13, P 21 Using the epsilon method from the set {P 21 ,OP2,OP 21 , OP 22 } select NP individuals as the new P 21 , where the f2 values ​​of all individuals are set to 0; S14, P 22 Using the epsilon method from the set {P 22 ,OP2,OP 21 , OP 22 } select NP individuals as the new P 22 , where the f1 value of all individuals is set to 0; S15, Cnt=Cnt+1; S16, if the maximum number of evaluations is reached, the algorithm stops and outputs P1 and P2; otherwise, jump to S7.

2. The multi-objective and multi-task optimization method for identifying individual patient network biomarkers according to claim 1, characterized in that: The specific steps of S3 include: S31, generates a |V|-dimensional row vector of all zeros, denoted as A, which is the optimal solution to the first objective; S32, generates a |V|-dimensional row vector of all zeros, denoted as B, where all dimensional values ​​corresponding to the prior nodes are set to 1, which is the optimal solution for the second objective.

3. The multi-objective and multi-task optimization method for identifying individual patient network biomarkers according to claim 1, characterized in that: The specific steps of S8 are: S81, if the remainder of Cnt divided by 30 is 0, jump to S82, otherwise jump to S83; S82, using the infeasible solution repair strategy to generate the offspring population OP2; S83, P2 uses a binary genetic algorithm to produce NP / 2 offspring individuals, denoted as OP2.

4. The multi-objective and multi-task optimization method for identifying individual patient network biomarkers according to claim 3, characterized in that: The specific steps of S82 are: S821, randomly select NP / 2 individuals from P1, denoted as ST1; S822, if the optimized network model is MDS, jump to S823; if the optimized network model is NCUA, jump to S828; S823, set counter i=1; S824, for the i-th individual in ST1, denoted as ST 1,i ; S825, for any j∈{1,2,…,|V|}, if ST 1,i If the node corresponding to the j-th dimension variable of is selected or the neighbor node of the node corresponding to the j-th dimension is selected, the j-th dimension variable of ST1 does not change; otherwise, a node is randomly selected from the node corresponding to the j-th dimension and its neighbors, and the dimension of ST1 corresponding to the newly selected node is set to 1; S826, update counter i=i+1; S827, if i>NP, output ST1, otherwise jump to S824; S828, set counter i=1; S829, for the i-th individual in ST1, denoted as ST 1,i , set j = 1; S8210, for any j∈{1,2,…,|E|}, if ST 1,i If neither of the two nodes connected by the jth edge in the corresponding PGIN is selected, a node is randomly selected and its corresponding dimension is set to 1; S8211, update counter i=i+1; S8212, if i>NP, output ST1, otherwise jump to S829.

Citation Information

Patent Citations

  • Multi-modal optimization method for detecting dynamic network biomarkers of individual cancer patient

    CN114628031A

  • Constraint multi-objective optimization method for detecting drug target of individual cancer patient

    CN116189758A

  • Drug design method based on two-stage evolution multi-task optimization

    CN116994673A

  • Multi-modal medical data fusion modeling method and device based on multi-task cascading

    CN117093948A

  • Constrained multi-modal multi-objective evolution method for identifying cancer individualized multi-modal drug targets

    CN117789830A