A proximal policy optimization algorithm combined with clinical prior reward 18 F-fdg pet / ct liver cancer kinetic parameter estimation method

By combining a proximal strategy optimization algorithm with clinical prior rewards, the data sparsity problem in the estimation of hepatocellular carcinoma dynamic parameters by 18F-FDG PET/CT was solved, achieving more accurate parameter estimation and capture of physiological information, supporting the auxiliary diagnosis of hepatocellular carcinoma.

CN119048480BActive Publication Date: 2025-12-19KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411195976.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-12-19
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

Existing reinforcement learning algorithms suffer from data sparsity in estimating the dynamic parameters of hepatocellular carcinoma using 18F-FDG PET/CT, which limits the agent's exploration capabilities and lacks a reward function that incorporates prior clinical knowledge, resulting in inaccurate parameter estimation.

Method used

By combining a proximal strategy optimization algorithm with a clinical prior reward function, and by constructing a reversible dual-input three-compartment model and an intelligent agent interaction environment, a reward function was designed to guide parameter estimation. The proximal strategy optimization algorithm was then used to estimate the dynamic parameters of 18F-FDG PET/CT liver cancer.

Benefits of technology

It improves the accuracy of parameter estimation, captures more pharmacokinetic and physiological information, and the estimated pharmacokinetic parameters show statistical differences between hepatocellular carcinoma tissue and normal liver tissue, which is in line with clinical expectations and supports the auxiliary diagnosis of hepatocellular carcinoma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048480B_ABST
    Figure CN119048480B_ABST
Patent Text Reader

Abstract

The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori reward 18 The present application relates to a proximal strategy optimization algorithm combined with clinical priori
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention provides a proximate strategy optimization algorithm that incorporates clinical prior rewards. 18 The F-FDG PET / CT method for estimating hepatokinetic parameters belongs to the field of deep reinforcement learning and pharmacokinetic parameter estimation. Background Technology

[0002] Liver cancer is a malignant tumor originating from hepatocytes and bile duct cells, and is classified as primary or secondary. Hepatocellular carcinoma (HCC) is the most common primary malignant tumor of the liver. Early-stage liver cancer often presents with few or no symptoms, while mid-to-late-stage liver cancer exhibits a variety of symptoms, such as pain in the liver area, abdominal distension, loss of appetite, fatigue, weight loss, progressive hepatomegaly, or an upper abdominal mass. Some patients may experience low-grade fever, jaundice, diarrhea, and upper gastrointestinal bleeding. Liver cancer is predicted to be the sixth most common cancer worldwide.

[0003] Recent studies have shown the potential of reinforcement learning in pharmacokinetic parameter estimation. Existing research has employed temporal difference (TD-learning) and state-action-reward-state-action (SARSA) algorithms for parameter estimation in kinetic models. However, these reinforcement learning algorithms require maintaining a Q-table to store the value of each state-action pair. When the state or action space is large, the size of the Q-table becomes difficult to manage and update, leading to data sparsity and limiting the agent's exploratory capabilities.

[0004] Technically, Proximal Policy Optimization (PPO) is a deep reinforcement learning algorithm proposed in recent years and has been validated in other parameter estimation fields, such as industrial parameter estimation and optimization. This method uses neural networks to approximate the Q-value and other policy and value functions, effectively alleviating the data sparsity problem of Q-table-based reinforcement learning algorithms. However, no research has yet applied the Proximal Policy Optimization algorithm to... 18 Estimation of dynamic parameters of hepatocellular carcinoma using F-FDG PET / CT; therefore, this invention proposes to use an optimization algorithm with a proximal strategy for estimation. 18 This paper proposes parameter estimation for the dynamics of liver cancer using F-FDG PET / CT and suggests incorporating prior medical knowledge into the reward function. This approach aims to improve the accuracy of parameter estimation for the intelligent agent, aligning it with clinical expectations and providing more precise guidance for the agent's optimization. 18 F-FDG means 18F-Fluorodeoxyglucose, PET / CT is positron emission tomography / computed tomography. SUMMARY

[0005] In view of the above defects or deficiencies in the prior art, and in view of the above research gap on liver cancer, the present application proposes a 18 F-FDG PET / CT liver cancer kinetic parameter estimation method, which realizes the proximal strategy optimization algorithm combined with clinical priori reward for 18 F-FDG PET / CT liver cancer kinetic parameter estimation method, which realizes the proximal strategy optimization algorithm combined with clinical priori reward for

[0006] The technical solution of the present application is: a proximal strategy optimization algorithm combined with clinical priori reward 18 F-FDG PET / CT liver cancer kinetic parameter estimation method, which realizes the proximal strategy optimization algorithm combined with clinical priori reward for

[0007] Step 1, obtaining clinical liver PET / CT images, including liver 5-minute dynamic PET / CT scanning and 60-minute conventional static PET / CT scanning, capturing the time course of F-FDG tracer uptake of HCC patients. 18 F-FDG tracer uptake time course.

[0008] Step 2, manually drawing a region of interest (ROI), including hepatocellular carcinoma, normal liver tissue, artery and portal vein, adjusting the ROI layer by layer, obtaining a time-concentration activity curve (TAC) composed of each frame maximum SUV (SUVmax) and its data.

[0009] Step 3, dividing the TAC data obtained in Step 2 into data sets, wherein the training set accounts for 29%, and the test set accounts for 71%;

[0010] Step 4, constructing a mathematical model of a reversible double-input three-compartment model, including 18 F-FDG uptake rate constant K1 from blood to liver tissue cells, 18 F-FDG return rate constant k2(1 / min) from liver tissue cells to blood, F-FDG in liver tissue, 18 F-FDG is phosphorylated to F-FDG 6-phosphate by hexokinase, 18 F-FDG 6-phosphate rate constant k3, dephosphorylation rate constant k4, aortic blood supply fraction fa, and the proportion of measured blood volume vb.

[0011] Step 5, build an intelligent agent interaction environment combined with a reversible double-input three-chamber model, including defining the state of the dynamic parameters, the action distribution and the action space, designing the reward function, etc.; among them, the reward function part is improved according to the clinical prior knowledge of the parameters, including the upper and lower limits of the parameters obtained from previous studies and the expected parameter range of individual parameters, to improve the design of the reward function in the intelligent agent interaction environment.

[0012] Step 6, build the overall framework of the Proximal Policy Optimization (PPO) algorithm, including the policy network (Actor network), the value function network (Critic network), the Generalized advantage estimation (GAE) algorithm for calculating the advantage function, and the clipping ratio of the new and old strategies, etc. Hyperparameters.

[0013] Step 7, use the Proximal Policy Optimization algorithm combined with medical prior rewards to estimate the parameters of the atrial model and analyze the perfusion and metabolic parameters of hepatocellular carcinoma tissue and normal liver tissue.

[0014] Specifically, the specific steps of Step 1 are as follows:

[0015] Step 1.1. Perform a low-dose CT scan of the liver (120kV, 100mA) on one bed; then quickly manually inject 18 F-FDG (5.5MBq / kg) in 2mL of 0.9% saline, in order to make the tracer enter the heart and lung circulation faster, flush with 20mL of 0.9% saline at a flow rate of 2mL / s; after completing the injection, perform a 5-minute liver PET scan. The conventional static PET / CT scan is performed about 60 minutes after the F-FDG bolus, first perform a total of 11 bed positions of whole body CT scan (120kV, 200mA) from the top of the skull to the proximal thigh, then perform 1 minute PET scan on each bed position. 18

[0016] Step 1.2. After scanning, use the short-term PET acquisition protocol, and use the Ordered Subset Expectation Maximization (OSEM) algorithm to reconstruct the PET image, the first minute of data is reconstructed into 12 frames with an interval of 5 seconds, the last four minutes of data is reconstructed into 4 frames with an interval of 60 seconds, and the 60th minute selects 1 frame of liver static PET image. A total of 17 frames of PET images are obtained and fused and registered with the CT image.

[0017] ​Specifically, in Step 2, if the patient's lesion has 18 F-FDG uptake is not obvious, then the ROI is drawn relative to the conventional CT image. The ROI of the abdominal aorta and portal vein is located at about two-thirds of the cross section of the blood vessels, and in order to compare the HCC tumor with the surrounding non-tumor liver tissue, the corresponding ROI is drawn in the non-tumor liver tissue, and the interference of the internal blood vessels of the liver on the SUV is avoided as much as possible during the delineation process. The ROI is copied to the PET / CT fusion image to obtain a time-concentration activity curve (TAC) composed of the maximum SUV (SUVmax) of each frame.

[0018] Specifically, in Step 3, 5 clinical liver cancer data are randomly selected as the training set, and the remaining 12 clinical liver cancer data are used as the test set.

[0019] Specifically, in Step 4, the mathematical form of the reversible double-input three-chamber model can be described by a set of ordinary differential equations:

[0020]

[0021] where t is the time variable, C(t) is the tissue concentration function, and the calculation is performed according to the formula C(t) = C f (t) + C p (t), and its expression is [c(t1), c(t2), …, c(t k )] T , and k is the total number of frames of the PET scan protocol. The matrix form of the model system can be expressed as:

[0022]

[0023] Solving the ordinary differential equation set gives:

[0024]

[0025] For the target region in the PET image, the model not only considers the tracer activity within the tissue within the region, but also considers the proportion of tracer concentration in the blood. Therefore, the parameter blood volume factor v b is used to weight the tissue concentration C(t) and the blood input concentration C i (t), and finally the total output concentration C t (t) of the target region is obtained:

[0026] C t (t) = v b × C i (t) + (1-v b ) × C(t)

[0027] Specifically, in the intelligent agent interaction environment in Step 5, the parameter state is defined as:

[0028] s = [k1, k2, k3, k4, f a , v b ]

[0029] The intelligent agent action distribution is defined as a 12-dimensional discrete space:

[0030] set a = [k1+, k1-, k2+, k2-, k3+, k3-, k4+, k4-, f a +, f a -, v b +, v b -]

[0031] In the setting of the reward function, a reward function combining the clinical prior knowledge of the kinetic parameters is proposed, and its calculation formula is:

[0032]

[0033]

[0034] r t = λ1r1+ λ2r2+ λ3r3+ λ4r4

[0035] Where r1~r4 are multiple reward signals, λ1~λ4 are the weights of the reward signals, r t is the reward value obtained at this moment, and t represents the time.

[0036] Specifically, the proximal policy optimization algorithm based on clipping (Clip) is used in Step 6. First, the generalized advantage estimation (GAE) algorithm is used to calculate the advantage value, and its formula is as follows:

[0037]

[0038] Where γ is the reward discount coefficient, λ is a constant between 0 and 1, is the calculation method of the GAE advantage function, r t+l is the single-step total reward at the t+l moment, l is the number of steps experienced, V represents the state value function, and s t represents the state at t moment. Then PPO updates the Actor and Critic networks multiple times. For the Critic network, the loss function can be calculated as:

[0039]

[0040] In the formula, r t′ represents the reward at time t', t represents the current time, V represents the state value function, phi represents the parameter of the value function network, s t represents the state at time t. For the Actor network, the target function can be calculated in the following manner:

[0041]

[0042] wherein theta represents the model parameter, represents the expectation of the target function, epsilon represents the clipping coefficient, pi θ (a t |s t ) represents a new policy function, represents an old policy function, represents the ratio of the new and old probability policies.

[0043] Specifically, the description analysis and statistical analysis performed in Step 7 include but are not limited to root-mean-square error (RMSE), area under the curve (AUC), and T-test (P<0.05 is considered to have statistical difference). In addition, the differences in kinetic parameters K1, k2, k3, k4, fa, and vb between tumor and normal liver tissues are analyzed to evaluate their clinical significance.

[0044] The beneficial effects of the present application are that the present application combines the proximal policy optimization algorithm of the prior clinical reward 18 The F-FDG PET / CT liver cancer kinetic parameter estimation method improves the accuracy of parameter estimation, and the reward function set in combination with prior medical knowledge can better guide the intelligent agent to perform parameter estimation and parameter optimization, capture more physiological information in pharmacokinetics, and estimate pharmacokinetic parameters with statistical differences between hepatocellular carcinoma tissues and normal liver tissues, thereby effectively evaluating 18 The F-FDG tracer has significantly different transport and phosphorylation rates in different tissues, providing technical support for hepatocellular carcinoma diagnosis. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 The flowchart in the present application;

[0046] Figure 2 The PET, CT, PET / CT images and ROI in dynamic PET / CT imaging used in the present application; A: ROI in the PET image; B: ROI in the CT image; C: ROI in the PET / CT image; D: ROI in the dynamic PET / CT imaging;

[0047] Figure 3 The time-density activity curve (TAC) is obtained from the ROI region of the PET / CT image used in this invention, consisting of the maximum SUV (SUVmax) of each frame.

[0048] Figure 4 This refers to the dual-input reversible three-compartment model used in this invention;

[0049] Figure 5 This is a diagram showing the model structure of the 18F-FDG PET / CT liver cancer dynamic parameter estimation method based on the proximal strategy optimization algorithm that combines clinical prior rewards, as proposed in this invention.

[0050] Figure 6 These are the ROC curves and AUC values ​​of the parameter results obtained from the algorithm model proposed in this invention. Here, PRPPO represents the method of this invention; NLLS represents the nonlinear least squares method. Detailed Implementation

[0051] Example: Figures 1-6 As shown, a proximal policy optimization algorithm combining clinical prior rewards is presented. 18 Methods for estimating dynamic parameters of hepatocellular carcinoma using F-FDG PET / CT include the following:

[0052] Experimental materials for this embodiment: 14 patients underwent PET / CT scans and metabolic PET / CT scans; patient data were provided by the Southern Pet Center of Southern Medical University Southern Hospital, and written informed consent was obtained from all patients; this study was approved by the Yunnan Provincial People's Hospital University Review Committee, with the approval number KHLL2022-KY189. This embodiment included 14 patients (13 males and 1 female), aged 31–78 years; 17 cases were diagnosed with liver cancer, with the longest axis diameter of the tumor ranging from 6.5 ± 3.6 cm (1.9–15.0 cm).

[0053] Initial PET / CT scan: After the patient fasted for more than 6 hours and their blood glucose stabilized, a liver CT scan (120 kV, 100 mAs) was performed; 5.5 MBq / kg was administered intravenously. 18After F-FDG, a 5-minute dynamic PET scan was performed in the liver region, supplemented by a 1-minute static PET (60 minutes). The dynamic PET data was divided into 17 frames: 12 frames at 5 s, 4 frames at 60 s, and 1 frame at 60 min, reconstructed using the standard ordered subset expectation maximization algorithm. The regions of interest (ROIs) in the CT image were manually delineated for each patient, including hepatocellular carcinoma, normal liver tissue, the aorta, and the portal vein (especially the extrahepatic part). After image fusion, the regions of interest were copied to the PET / CT image to generate time activity curves (TACs) and extract the maximum standard uptake values (SUVmax) from each frame of image. Figure 2 The CT image and the corresponding position of the ROI drawn in a frame of PET / CT image are shown, where the black-drawn area is the outline of hepatocellular carcinoma, the yellow-drawn area is the position of the aorta, the red color is the position of the hepatic portal vein, and the green-drawn area is normal liver tissue.

[0054] The proximal strategy optimization algorithm of the present embodiment combined with a clinical priori reward 18 The specific steps of the F-FDG PET / CT liver cancer kinetic parameter estimation method are as follows:

[0055] Step 1, obtain the clinical liver PET / CT image, including 5-minute dynamic PET / CT scan of the liver and 60-minute conventional static PET / CT scan, capture the time course of F-FDG tracer uptake of HCC patients. 18 F-FDG tracer uptake.

[0056] Step 2, manually draw the region of interest (ROI), including hepatocellular carcinoma, normal liver tissue, artery and portal vein, adjust the ROI layer by layer, obtain the time-concentration activity curve (TAC) composed of the maximum SUV (SUVmax) of each frame and its data.

[0057] Step 3, divide the TAC data obtained in Step 2 into data sets, where the training set accounts for 29% and the test set accounts for 71%;

[0058] Step 4, construct a mathematical model of a reversible double-input three-compartment model, including 18 the rate constant K1 of F-FDG uptake from blood to liver tissue cells, 18 the rate constant k2 (1 / min) of F-FDG return from liver tissue cells to blood, the volume of distribution Vd of F-FDG in liver tissue, 18 F-FDG is phosphorylated to 18The rate constant k3 of F-FDG 6-phosphate, the rate constant k4 of dephosphorylation, the fraction of aortic blood supply fa, and the proportion of the measured blood volume vb.

[0059] Step 5, build an intelligent agent interaction environment combined with a reversible double-input three-chamber model, including defining the state of kinetic parameters, action distribution and action space, designing reward functions, etc.; wherein the reward function part is improved according to the clinical prior knowledge of the parameters, including the upper and lower limits of the parameters obtained from previous studies and the expected parameter range of individual parameters, to improve the reward function in the intelligent agent interaction environment.

[0060] Step 6, build the overall framework of the Proximal Policy Optimization (PPO) algorithm, including the policy network (Actor network), the value function network (Critic network), the generalized advantage estimation (GAE) algorithm for calculating the advantage function, and the clipping ratio of the new and old strategies, etc. Hyperparameters.

[0061] Step 7, use the Proximal Policy Optimization algorithm combined with medical prior rewards to estimate the parameters of the atrioventricular model and analyze the perfusion and metabolic parameters of hepatocellular carcinoma tissue and normal liver tissue, statistical analysis and other clinical analysis.

[0062] Specifically, the specific steps of Step 1 are as follows:

[0063] Step 1.1. Perform a low-dose CT scan (120kV, 100mA) of the liver at one bed position; then quickly manually administer 18 F-FDG (5.5MBq / kg) in 2mL of 0.9% saline to make the tracer enter the cardiopulmonary circulation faster, flush with 20mL of 0.9% saline at a flow rate of 2mL / s; after completing the injection, perform a 5-minute liver PET scan. The conventional static PET / CT scan is performed about 60 minutes after 18 F-FDG injection, first perform a total of 11 bed positions of whole-body CT scan (120kV, 200mA) from the top of the skull to the proximal thigh, then perform 1-minute PET scan on each bed position.

[0064] Step 1.2. After scanning, short-term PET acquisition protocol is used, and PET image reconstruction is performed using standard ordered subset expectation maximization (OSEM) algorithm. The first minute of data is reconstructed into 12 frames with an interval of 5 seconds, and the last four minutes of data is reconstructed into 4 frames with an interval of 60 seconds. One frame of liver static PET image is selected at the 60th minute. A total of 17 frames of PET images are obtained and fused and registered with the CT images.

[0065] Specifically, in Step 2, if the patient's lesion has 18 F-FDG uptake is not obvious, the ROI is drawn relative to the conventional CT image. The ROI of the abdominal aorta and portal vein is located at about two-thirds of the vessel cross section, and in order to compare the HCC tumor with the surrounding non-tumor liver tissue, a corresponding ROI is drawn in the non-tumor liver tissue, and the drawing process avoids the interference of internal blood vessels in the liver on SUV as much as possible. The ROI is copied to the PET / CT fusion image to obtain a time-concentration activity curve (TAC) composed of each frame maximum SUV (SUVmax), and the results are shown in Figure 3 .

[0066] Specifically, in Step 3, 5 clinical liver cancer data are randomly selected as the training set, and the remaining 12 clinical liver cancer data are used as the test set.

[0067] Specifically, as shown in Figure 4 , the mathematical form of the reversible double-input three-compartment model in Step 4 can be described by a set of ordinary differential equations:

[0068]

[0069] where t is the time variable, C(t) is the tissue concentration function, and the calculation is performed according to the formula C(t) = C f (t) + C p (t). Its expression is [c(t1), c(t2), …, c(t k )] T , and k is the total number of frames of the PET scan protocol. The matrix form of the model system can be expressed as:

[0070]

[0071] Solving the ordinary differential equation set gives:

[0072]

[0073] For the target region in the PET image, the model not only considers the tracer activity within the tissue in the region, but also considers the proportion of tracer concentration in the blood. Therefore, the parameter blood volume fraction v b is introduced to weight the tissue concentration C(t) and the blood input concentration C i (t) to obtain the total output concentration C t (t) of the target region:

[0074] C t (t)=v b ×C i (t)+(1-v b )×C(t)

[0075] Specifically, in the intelligent agent interaction environment in Step 5, the parameter state is defined as:

[0076] s=[K1,k2,k3,k4,f a ,v b ]

[0077] The action distribution of the intelligent agent is defined as a 12-dimensional discrete space:

[0078] set a =[K1+,K1-,k2+,k2-,k3+,k3-,k4+,k4-,f a +,f a -,v b +,v b -]

[0079] In the setting of the reward function, a reward function combining the clinical prior knowledge of the kinetic parameters is proposed, and its calculation formula is:

[0080]

[0081] r t =λ1r1+λ2r2+λ3r3+λ4r4

[0082] Where r1~r4 are multiple reward signals, λ1~λ4 are the weights of the reward signals, r t is the reward value obtained at this moment, and t represents the moment.

[0083] Specifically, the proximal policy optimization algorithm based on clipping (Clip) is used in Step 6. First, the generalized advantage estimation (GAE) algorithm is used to calculate the advantage value, and its formula is as follows:

[0084]

[0085] where γ is a reward discount factor, λ is a constant between 0 and 1, is the calculation method of GAE advantage function, r t+l is the total reward of a single step at the t+l moment, l is the number of steps experienced, V represents the state value function, s t represents the state at t moment. Then PPO updates the Actor and Critic networks multiple times. For the Critic network, the loss function can be calculated as:

[0086]

[0087] where r t′ represents the reward at t' moment, t represents the current moment, V represents the state value function, φ represents the parameter of the value function network, s t represents the state at t moment. For the Actor network, the objective function can be calculated as:

[0088]

[0089] where θ represents the model parameter, represents the expectation of the objective function, ε represents the clipping coefficient, π θ (a t |s t ) represents the new policy function, represents the old policy function, represents the ratio of the new and old probability policies.

[0090] Specifically, the descriptive analysis and statistical analysis performed in Step 7 include but are not limited to root-mean-square error (RMSE), area under the curve (AUC), and T-test (P<0.05 is considered statistically different). In addition, the differences in kinetic parameters K1, k2, k3, k4, fa, and vb between tumor and normal liver tissues are analyzed to evaluate their clinical significance.

[0091] Experiment 1:

[0092] Comparison experiment with traditional algorithm. The data of 12 patients with hepatocellular carcinoma were subjected to traditional nonlinear least square (NLLS) experiment, and the parameter estimation results and experimental evaluation indexes are shown in Table 1:

[0093] Table 1 Statistical analysis of parameter estimation results of NLLS method

[0094]

[0095] * represents the average value of normal liver tissue and HCC, which does not conform to the physiological expectation

[0096] As shown in Table 1, in the NLLS optimization method, for hepatocellular carcinoma, the obtained kinetic parameters K1 (1.4899 ± 0.0003 vs. 1.3925 ± 0.1914, P < 0.05), fa (0.9279 ± 0.1340 vs. 0.3095 ± 0.2378, P < 0.001) are significantly higher than those of normal liver tissue. k2 (0.8486 ± 0.4080 vs. 1.2169 ± 0.3561, P = 0.0125) is significantly lower than that of normal liver tissue, but does not conform to the physiological expectation. For k3 (0.0736 ± 0.0874 vs. 0.0198 ± 0.0431, P = 0.2025) and vb (0.0999 ± 0.0001 vs. 0.0929 ± 0.0201, P = 0.0948), although HCC is higher than normal liver tissue, there is no significant difference. For k4 (0.0669 ± 0.0347 vs. 0.0860 ± 0.0001, P = 0.1340), HCC is lower than normal liver tissue, but there is no significant difference. In general, the traditional NLLS optimization method can only effectively distinguish HCC and normal liver tissue through parameters K1 and fa.

[0097] Experiment 2:

[0098] According to the experimental procedure, test on the same batch of 12 HCC patient data, the specific parameter estimation results and experimental evaluation index comparison are shown in Table 2:

[0099] Table 2 Experimental results of the application

[0100]

[0101] As shown in Table 2, in the proximal strategy optimization algorithm combined with clinical prior reward in the application 18In the F-FDG PET / CT liver cancer kinetic parameter estimation method, for hepatocellular carcinoma, the obtained kinetic parameters K1 (1.1491±0.1620 vs. 0.9658±0.0849, P<0.001), k2 (1.4342±0.0888 vs. 1.2763±0.0846, P<0.001), k3 (0.1875±0.0040 vs. 0.1817±0.0026, P<0.001), fa (0.8251±0.2035 vs. 0.4046±0.1789, P<0.001), and vb (0.0487±0.0045 vs. 0.0425±0.0031, P<0.001) are significantly higher than those of normal liver tissue. The significant increase of K1 and k3 indicates that the glucose transporter and hexokinase activity in hepatocellular carcinoma are increased, and k2 indicates that the two regions of hepatocellular carcinoma and normal liver tissue have different glucose transport rates. 18 The different clearance rates of F-FDG transport back to blood, the significant increase of fa is consistent with the hypothesis that the double input is the actual blood supply mode of most HCC tumors, and the significant increase of vb indicates that the blood supply intensity of hepatocellular carcinoma is significantly higher than that of normal liver tissue. The k4 parameter of hepatocellular carcinoma is significantly lower than that of normal liver tissue, indicating that the G6P activity is increased, and all the parameter results obtained by the algorithm model of the application are consistent with the expected clinical performance.

[0102] Overall, the proximal strategy optimization algorithm combined with the clinical priori reward 18 The pharmacokinetic parameter results obtained by the F-FDG PET / CT liver cancer kinetic parameter estimation method are more accurate than the traditional NLLS algorithm, more consistent with the physiological characteristics and clinical performance expectations, and the overall higher AUC value and lower fitting error RMSE value are obtained, as shown in Figure 6 It can be seen that the method of the application can effectively perform 18 F-FDG PET / CT liver cancer kinetic parameter estimation, and can better evaluate the characteristics of hepatocellular carcinoma.

[0103] The specific embodiments of the application are described in detail above in combination with the drawings, but the application is not limited to the above embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the purpose of the application.

Claims

1. A method of combining a proximal policy optimization algorithm with clinical prior rewards 18 A method for estimating F-FDG PET / CT liver cancer kinetic parameters, characterized in that, Specifically comprising the following steps: Step 1, Obtain clinical liver PET / CT images, capture the time course of F-FDG tracer uptake in HCC patients 18 F-FDG tracer uptake in HCC patients Step 2, manually outline the region of interest (ROI), adjust the ROI layer by layer, obtain the time-concentration activity curve (TAC) composed of the maximum SUV of each frame and its data; Step 3, divide the TAC data obtained in Step 2 into data sets; Step 4, construct a mathematical model of a reversible double-input three-chamber model; Step 5, construct an intelligent agent interactive environment combining the reversible double-input three-chamber model, including defining the state of the dynamic parameters, the action distribution and the action space, and designing the reward function; wherein the reward function part is improved according to the parameter clinical prior knowledge, including the upper and lower limits of the parameters obtained from previous studies and the expected parameter range of individual parameters, to design the reward function in the intelligent agent interactive environment; Step 6, build the overall framework of the proximal policy optimization (PPO) algorithm, including the policy network, the value function network, the generalized advantage estimation algorithm for calculating the advantage function, and the clipping ratio hyperparameter of the new and old policies; Step 7, use the proximal policy optimization algorithm combined with medical prior rewards to estimate the parameters of the atrial and ventricular model and analyze the perfusion and metabolic parameters of hepatocellular carcinoma tissue and normal liver tissue, statistical analysis and other clinical analysis; The mathematical form of the reversible double-input three-chamber model in Step 4 is described by a system of ordinary differential equations: ; ; where is the time variable, is the tissue concentration function, calculated at time according to the formula is performed, whose expression is , is the total number of frames of the PET scan protocol, the matrix form of the model system is given by ; Solving the system of ordinary differential equations gives: ; For a target region in a PET image, the model considers not only the tissue- in activity within the region, but also the proportion of tracer concentration in the blood; thus, the parameter blood volume fraction is used to weight the tissue concentration and the blood input concentration to give the total output concentration of the target region : ; In the agent interactive environment in Step 5, the parameter state is defined as: ; The action distribution of the agent is defined as a 12-dimensional discrete space: ; In the setting of the reward function, a reward function combining the clinical prior knowledge of the dynamic parameters is proposed, and its calculation formula is: ; ; ; ; ; wherein, is a plurality of reward signals, is a weight of the reward signal, denotes a time instant, is a reward value obtained at the time instant.

2. The proximal policy optimization algorithm with clinical prior reward according to claim 1, 18 F-FDG PET / CT liver cancer kinetic parameter estimation method, characterized in that: The specific steps of Step 1 are as follows: Step 1.

1. A low-dose CT scan of the liver is performed with one bed position; this is followed by a rapid manual injection of 18 F-FDG, in order to allow the tracer to enter the cardio-pulmonary circulation more rapidly, a flush of 20 mL of 0.9% saline at a flow rate of 2 mL / s is performed; after completion of the injection, a 5-minute liver PET scan is performed; a conventional static PET / CT scan is performed 18 F-FDG push, approximately 60 minutes later, a total of 11 bed positions are performed from the vertex of the skull to the proximal thigh, followed by a 1-minute PET scan on each bed position; Step 1.

2. After scanning, use the short-term PET acquisition protocol, use the standard ordered subset expectation maximization (OSEM) algorithm for PET image reconstruction, the first minute of data is reconstructed into 12 frames with an interval of 5 seconds, the last four minutes of data is reconstructed into 4 frames with an interval of 60 seconds, and the 60th minute selects 1 frame of liver static PET image, a total of 17 frames of PET image are obtained, and are fused and registered with the CT image.

3. The proximal policy optimization algorithm with clinical prior reward according to claim 1, 18 A method for estimating liver cancer kinetic parameters of F-FDG PET / CT, characterized in that: In Step 2, if the patient's lesion has 18 If the F-FDG uptake is not obvious, the ROI is drawn relative to the conventional CT image; the ROI of the abdominal aorta and portal vein is located at about two-thirds of the cross section of the blood vessels, and in order to compare the HCC tumor with the surrounding non-tumor liver tissue, the corresponding ROI is drawn in the non-tumor liver tissue, and the interference of the internal blood vessels of the liver on the SUV is avoided as much as possible during the drawing process; the ROI is copied to the PET / CT fusion image to obtain a time-concentration activity curve TAC composed of the maximum SUV of each frame.

4. The proximal policy optimization algorithm with clinical prior reward according to claim 1, 18 A method for estimating liver cancer kinetic parameters of F-FDG PET / CT, characterized in that: In Step 2, the manually outlined region of interest includes hepatocellular carcinoma, normal liver tissue, artery and portal vein.

5. The proximal policy optimization algorithm with clinical prior reward according to claim 1, 18 A method for estimating liver cancer kinetic parameters of F-FDG PET / CT, characterized in that: In Step 6, the proximal policy optimization algorithm based on clipping is used, first the generalized advantage estimation algorithm is used to calculate the advantage value, the formula is as follows: ; in, As a reward discount factor, A constant between 0 and 1 The method for calculating the GAE dominance function. For the first Total reward per step at any given moment The number of steps taken. Represents the state value function. express The state at any given moment; Then PPO updates the Actor and Critic networks multiple times, for the Critic network, the loss function is calculated as: ; wherein represents a reward at a time instant, represents a current time instant, represents a state value function, represents a parameter of the value function network, represents a state at a time instant; For the Actor network, the target function is calculated as: ; wherein, denotes a model parameter, denotes an expectation of the objective function, denotes a clipping coefficient, denotes a new policy function, denotes an old policy function, denotes a new-old probability policy ratio.

6. The proximal strategy optimization algorithm combining clinical prior rewards as described in claim 1. 18 The method for estimating dynamic parameters of hepatocellular carcinoma using F-FDG PET / CT is characterized by: The descriptive analysis and statistical analysis performed in Step 7 include root mean square error RMSE, area under the curve AUC, and T test, P < 0.05 is considered to have statistical difference; in addition, it also includes the analysis of kinetic parameters K1, k2, k3, k4, f a , v b The differences between tumor and normal liver tissues were analyzed to evaluate their clinical significance.

Citation Information

Patent Citations

  • Medical image segmentation method for introducing priori knowledge based on reward function

    CN115187571A

  • 18F-FDG PET / CT liver cancer kinetic parameter estimation method of Advantageous Actor-Critic algorithm

    CN118378663A