Liver cancer occurrence risk detection method and system
By acquiring and simulating the data on serum alpha-fetoprotein AFP and intrahepatic nodules image characteristics of liver cancer patients, and using differential operations to quantify their independent and interactive contributions, the problem that cannot be accurately quantified in the existing technology is solved, and the accuracy of liver cancer risk assessment and the guiding principle of clinical decision-making is improved.
Patent Information
- Application Number
- CN202511022045.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing liver cancer risk assessment methods cannot accurately quantify the independent and interactive contributions of each risk indicator characteristic in complex clinical situations, resulting in difficulty in clinical decision-making.
By obtaining original medical data containing image characteristics of serum alpha-fetoprotein AFP and intrahepatic nodules, a variety of simulation data were generated using a preset risk assessment model and performing differential operations to quantify the independent and interactive risk contribution of each risk indicator feature.
It has achieved a refined analysis of the risk drivers of liver cancer, provided more explanatory and guiding clinical decision-making support, and improved the accuracy and reliability of risk assessment.
Smart Images

Figure CN120527005A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of medical data analysis and disease risk assessment, and specifically, to a method and system for detecting the risk of liver cancer. Background Art
[0002] In the field of medical diagnostics, particularly in liver disease risk assessment, automated risk assessment programs have been widely used to identify high-risk patients. These programs typically provide a contribution analysis of various risk characteristics to aid clinical decision-making. However, in certain complex clinical scenarios, existing contribution assessment methods may not provide sufficiently accurate or explanatory information, leading to clinical decision-making dilemmas.
[0003] For example, for a patient who has been determined to be at high risk for liver cancer, their electronic medical records may show a significant increase in serum alpha-fetoprotein (AFP) and be identified as the main source of risk contribution by the existing risk assessment system. However, if the patient also has a history of long-term and effective antiviral treatment, resulting in their hepatitis B virus DNA (HBV DNA) load being below the detection limit for a long time, the pathophysiological basis of the increased AFP (such as viral replication activity) no longer exists. At this point, the risk significance of AFP becomes unclear, and its contribution may be overestimated by existing models, because these models are usually trained on a general population database containing a large amount of viral activity data and cannot recognize the impact of the special background of "suppressed viral load" on the explanatory power of AFP risk. The increase in AFP may represent a carcinogenic mechanism independent of the virus, or it may simply be a manifestation of nonspecific hyperplasia in the context of cirrhosis, which is difficult to distinguish with existing methods.
[0004] At the same time, the patient's imaging report may show an intrahepatic nodule with atypical features, such as being classified as "moderately likely to be hepatocellular carcinoma" (such as LI-RADS LR-3). Although the independent risk contribution of this nodule may be quantified as low in the current system, its presence in a liver environment where the virus is effectively suppressed and should be stable is itself of great clinical significance.
[0005] The above scenario reveals a core contradiction: the risk significance of the serum marker (AFP), which numerically contributes the most, becomes uncertain due to the patient's successful treatment history; whereas, the presence of radiographic nodules, which numerically contribute less, appears abnormal in a specific context. A deeper issue lies in the fact that the true source of risk may not be the independent effect of a single marker, but rather the interactive effect resulting from the "coexistence" of these markers in a specific clinical context. Existing methods can calculate the independent contribution of each feature, but cannot identify and quantify the incremental risk of this "conditional dependency." In other words, they cannot answer how much risk is contributed independently by AFP, and how much risk is amplified by the presence of AFP, which increases the risk of nodules.
[0006] This dilemma of being unable to quantify and interpret the data directly leads to clinical decision paralysis. Unable to accurately isolate and quantify the independent contributions of each risk indicator feature and the interactive contributions between them, the diagnosis and treatment team struggles to make a clear choice between different treatment pathways. Summary of the Invention
[0007] The present application provides a method and system for detecting the risk of liver cancer, which has the advantage of being able to accurately quantify the independent risk contribution of each risk indicator feature and the interactive risk contribution between them, thereby providing more explanatory and guiding information for complex clinical decision-making.
[0008] In one aspect, the present application provides a method for detecting the risk of liver cancer, comprising: Obtaining original medical data including a first risk indication feature and a second risk indication feature, and processing the original medical data according to a preset risk assessment model to generate a first risk value; wherein the first risk indication feature is a serum feature including serum alpha-fetoprotein (AFP); and the second risk indication feature is an imaging feature of the intrahepatic nodule morphology; Modifying a first risk-indicating feature in the original medical data to a first preset state to generate first simulated medical data, and processing the first simulated medical data according to the preset risk assessment model to generate a second risk value; wherein the first preset state indicates that the serum characteristic is at an upper limit of a normal range; modifying the second risk-indicating feature in the original medical data to a second preset state to generate second simulated medical data, and processing the second simulated medical data according to the preset risk assessment model to generate a third risk value; wherein the second preset state indicates that the image feature has no abnormal description; modifying the first risk indication feature to the first preset state, and modifying the second risk indication feature to the second preset state to generate third simulated medical data, and processing the third simulated medical data according to the preset risk assessment model to generate a fourth risk value; Based on the first risk value, the second risk value, the third risk value and the fourth risk value, a differential operation is performed to determine an independent risk contribution of the first risk indication feature, an independent risk contribution of the second risk indication feature, and an interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0009] Optionally, the differential operation is performed as follows: IR1=R3-R4 IR2= R2-R4 IR12=R1-R2-R3+R4 Among them, R1 is the first risk value, R2 is the second risk value, R3 is the third risk value, R4 is the fourth risk value, IR1 is the independent risk contribution of the first risk indication feature, IR2 is the independent risk contribution of the second risk indication feature, and IR12 is the interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0010] Optionally, the first preset state is determined by the following steps, specifically including: In the original medical data, the second risk indication feature is kept unchanged and only the value of the first risk indication feature is changed to generate test data; Processing the test data using the preset risk assessment model to obtain a test risk value corresponding to a first risk indication feature value in each test data; Determining a trigger threshold according to a correspondence between the value of the first risk indication feature and the test risk value, and determining a first preset state based on the trigger threshold; Among them, the trigger threshold is the inflection point position where the risk value in the corresponding relationship changes from linear change to nonlinear change; the first preset state is set to a calibration state, and the value of the calibration state is close to the trigger threshold and does not meet the trigger condition defined by the trigger threshold.
[0011] Optionally, the step of keeping the second risk indication feature unchanged and only changing the value of the first risk indication feature in the original medical data to generate test data includes: Performing a unidirectional scan within a preset value range to generate a first subset of test data, wherein, in the unidirectional scan, the value of the first risk indication feature changes in one direction; Performing a reverse scan within the preset value range to generate a second subset of test data, wherein a direction of change of the value of the first risk indication feature in the reverse scan is opposite to a direction of change of the value of the first risk indication feature in the unidirectional scan; Combining the first subset test data and the second subset test data to form the test data; The test data is processed according to the preset risk assessment model to obtain the test risk value, where the test risk value includes a first response path corresponding to the unidirectional scan and a second response path corresponding to the reverse scan.
[0012] Optionally, the step of determining a trigger threshold according to a correspondence between a value of the first risk indication feature and the test risk value includes: Identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining a value of the first risk indication feature corresponding to the location as a first candidate threshold; Identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining a value of the first risk indication feature corresponding to the location as a second candidate threshold; The smaller one between the first candidate threshold and the second candidate threshold is determined as the trigger threshold.
[0013] Optionally, the steps of identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining a value of the first risk indication feature corresponding to the location as a first candidate threshold, and identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining a value of the first risk indication feature corresponding to the location as a second candidate threshold, include: During the process of performing the unidirectional scan to generate the first subset test data and performing the reverse scan to generate the second subset test data, maintaining the second risk indication feature as a state recorded in the original medical data; Processing the first subset of test data and the second subset of test data according to the preset risk assessment model to obtain the first response path and the second response path; A location where the risk value changes nonlinearly is identified from the first response path to determine the first candidate threshold, and a location where the risk value changes nonlinearly is identified from the second response path to determine the second candidate threshold.
[0014] Optionally, the step of identifying a position where the risk value undergoes a nonlinear change from the first response path to determine the first candidate threshold, and identifying a position where the risk value undergoes a nonlinear change from the second response path to determine the second candidate threshold includes: For any one of the first response path and the second response path, perform the following steps: Traversing the data points in the response path, and defining, for each data point, a pre-change data segment consisting of a group of adjacent data points before the data point, and a post-change data segment consisting of a group of adjacent data points after the data point; Calculating a first statistic based on the risk value within the pre-change data segment defined for each data point, and calculating a second statistic based on the risk value within the post-change data segment; determining a variation range of each data point based on the first statistic and the second statistic calculated for the data point; The data point with the largest variation amplitude is determined as the location where the risk value in any response path undergoes nonlinear variation.
[0015] Optionally, the step of calculating the first statistic based on the risk value in the pre-change data segment defined for each data point, and calculating the second statistic based on the risk value in the post-change data segment, comprises: For any data segment of the pre-change data segment and the post-change data segment, perform the following steps: Determine the distribution characteristics of risk values within the data segment; According to the distribution characteristics, a corresponding calculation rule is selected for the data segment from a preset calculation rule set; The calculation rule selected for the data segment is used to calculate and obtain the first statistic or the second statistic corresponding to the data segment.
[0016] Optionally, the step of determining the distribution characteristics of the risk value in the data segment includes: Calculate the dispersion index of the risk value within the data segment; The dispersion index is calculated and determined as the distribution characteristic of the risk value in the data segment.
[0017] In another aspect, the present application provides a liver cancer risk detection system, comprising: an original medical data acquisition module, configured to acquire original medical data including a first risk indication feature and a second risk indication feature, and process the original medical data according to a preset risk assessment model to generate a first risk value; a first single-factor risk assessment module, configured to generate first simulated medical data based on the original medical data by modifying a first risk indication feature in the original medical data to a first preset state, and process the first simulated medical data according to the preset risk assessment model to generate a second risk value; a second single-factor risk assessment module, configured to generate second simulated medical data based on the original medical data by modifying a second risk indication feature in the original medical data to a second preset state, and process the second simulated medical data according to the preset risk assessment model to generate a third risk value; a background risk assessment module for generating, based on the original medical data, third simulated medical data by modifying the first risk indication feature to the first preset state and the second risk indication feature to the second preset state, and processing the third simulated medical data according to the preset risk assessment model to generate a fourth risk value; a contribution quantification calculation module that performs a differential operation based on the first risk value, the second risk value, the third risk value, and the fourth risk value to determine an independent risk contribution of the first risk indication feature, an independent risk contribution of the second risk indication feature, and an interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0018] The present application provides a method and system for detecting the risk of liver cancer, which introduces differential operations to quantify the independent contribution and interactive contribution of each risk indicator feature, effectively solving the problem in existing technologies that is unable to accurately analyze the source of risk in complex clinical scenarios. This allows for accurate quantification of the independent risk contribution of each risk indicator feature and the interactive risk contribution between them, providing more explanatory and guiding information for complex clinical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 Schematic diagram of a method for detecting the risk of liver cancer in an embodiment is shown in FIG. Figure 2 Schematic diagram of a module configuration of a liver cancer risk detection system in an embodiment is shown in FIG.
[0021] Figure numerals: 100, liver cancer risk detection system; 10, original medical data acquisition module; 20, first single factor risk assessment module; 30, second single factor risk assessment module; 40, background risk assessment module; 50, contribution quantification calculation module. DETAILED DESCRIPTION
[0022] The technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The components of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.
[0023] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0024] Traditional liver cancer risk assessment methods struggle to effectively distinguish and quantify the independent risk contributions of different risk-indicating features (such as serum alpha-fetoprotein (AFP) and intrahepatic nodule imaging characteristics) when processing complex clinical data, particularly when patients are receiving long-term effective antiviral therapy, which suppresses key pathogenic factors. Furthermore, they are unable to quantify the additive interactive risk arising from the coexistence of these features in specific clinical settings. This leads to an ambiguous interpretation of a patient's true risk profile, which in turn impacts the accuracy and timeliness of subsequent clinical decisions.
[0025] If these issues are not addressed, clinical decision-making will be paralyzed. Unable to quantify and separate the independent contribution of AFP from the interactive contribution of AFP and nodules, the diagnostic and treatment team cannot determine whether the primary risk is an unknown systemic factor represented by AFP, or a nodule that becomes suspicious against the backdrop of abnormal serology. This will directly lead to the inability to make a clear choice between two distinct treatment paths: conducting more extensive etiology screening and adjusting systemic treatment regimens, or immediately performing a puncture biopsy and experimental ablation of the nodule. This will delay the patient's diagnosis and treatment, potentially leading to disease progression and negatively impacting the patient's prognosis.
[0026] like Figure 1 The figure shows a schematic diagram of a method for detecting the risk of liver cancer. The method proposed in this application includes: S10, obtaining original medical data including a first risk indication feature and a second risk indication feature, processing the original medical data according to a preset risk assessment model, and generating a first risk value; wherein, the first risk indication feature is a serum feature including serum alpha-fetoprotein AFP; the first risk indication feature is an imaging feature of the intrahepatic nodule morphology.
[0027] Among them, the preset risk assessment model refers to a mathematical or statistical model used to assess the risk of liver cancer. It can be implemented using a machine learning model, a statistical regression model, or a rule model based on expert knowledge. It can perform quantitative analysis on medical data to generate a risk value.
[0028] S20, modifying the first risk indication feature in the original medical data to a first preset state to generate first simulated medical data, processing the first simulated medical data according to the preset risk assessment model to generate a second risk value; wherein, the first preset state represents that the serum characteristic is at the upper limit of the normal range.
[0029] Among them, the first preset state refers to the specific baseline state to which the first risk indication characteristic (serum characteristic) is adjusted, specifically the upper limit of the normal range of the serum characteristic, which can be determined based on clinical guidelines, statistical analysis or model training data to simulate the risk level when the serum characteristic is in a non-abnormal state.
[0030] S30, modifying the second risk indication feature in the original medical data to a second preset state to generate second simulated medical data, processing the second simulated medical data according to the preset risk assessment model to generate a third risk value; wherein, the second preset state represents that the image feature has no abnormal description.
[0031] Among them, the second preset state is a specific baseline state to which the second risk indication feature (imaging feature) is adjusted, specifically, the imaging feature has no abnormal description, which can be defined according to imaging diagnostic standards or expert consensus to simulate the risk level when the imaging feature is in a non-abnormal state.
[0032] S40, modify the first risk indication feature to the first preset state, and modify the second risk indication feature to the second preset state to generate third simulated medical data, process the third simulated medical data according to the preset risk assessment model to generate a fourth risk value.
[0033] S50: Perform a differential operation based on the first risk value, the second risk value, the third risk value, and the fourth risk value to determine an independent risk contribution of the first risk indication feature, an independent risk contribution of the second risk indication feature, and an interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0034] The core innovation of this application is that by simulating the preset states of different risk indicator features and combining them with sophisticated differential operations, it is possible to accurately separate and quantify the independent risk contribution of each risk indicator feature and the interactive risk contribution between them, thereby achieving the effect of detailed and accurate analysis of the driving factors of liver cancer risk in a complex clinical context, thereby providing data support for clinical decision-making.
[0035] First, the system acquires raw medical data containing a first risk-indicating feature (serum alpha-fetoprotein (AFP)) and a second risk-indicating feature (imaging features of intrahepatic nodule morphology) and inputs this data into a pre-set risk assessment model to generate a first risk value. This first risk value represents the patient's overall risk level, taking into account all known risk factors, and serves as the benchmark for all subsequent risk contribution analyses. Based on this, the system modifies the first risk-indicating feature in the raw medical data to a first pre-set state, generating the first simulated medical data.
[0036] This simulated data is then re-entered into the pre-set risk assessment model to generate a second risk value. This second risk value reflects the risk level attributable to remaining risk factors (including the second risk-indicating feature and background factors) after the first risk-indicating feature is normalized. Similarly, the system modifies the second risk-indicating feature in the original medical data to a second pre-set state, generating second simulated medical data. This data is processed by the pre-set risk assessment model to generate a third risk value. The third risk value reflects the risk level attributable to remaining risk factors (including the first risk-indicating feature and background factors) after the second risk-indicating feature is normalized. Furthermore, to determine pure background risk, the system modifies the first risk-indicating feature in the original medical data to the first pre-set state and the second risk-indicating feature to the second pre-set state, generating third simulated medical data. This data is processed by the pre-set risk assessment model to generate a fourth risk value. This fourth risk value represents the risk level attributable to other background factors after both primary risk-indicating features have been normalized or eliminated.
[0037] Ultimately, the system performs a differential calculation based on these four key risk values (first, second, third, and fourth). Through ingenious mathematical combinations, it accurately isolates the independent risk contribution of the first risk indicator feature, the independent risk contribution of the second risk indicator feature, and, more critically, the interactive risk contribution between the first and second risk indicator features. The logic behind this differential calculation is that by comparing the risk changes under different feature combinations, it reveals the additional risk increments generated by each feature acting alone and by their interactions.
[0038] For example, raw medical data can be obtained from a hospital's electronic medical record system. This data is typically stored in a structured or semi-structured format. The pre-set risk assessment model can be a pre-trained deep learning model, such as one based on a convolutional neural network or a recurrent neural network. This model has been trained on a large number of medical data from patients with and without liver cancer, and can output a risk probability value based on the input medical features. Specifically, after obtaining a patient's raw medical data, for example, including a serum alpha-fetoprotein (AFP) value of 150 ng / mL and an imaging description of an intrahepatic nodule as "LR-3," this raw data is first input into the pre-trained deep learning model to obtain a first risk value, such as 0.75. Next, to generate the first simulated medical data, the system modifies the serum alpha-fetoprotein (AFP) value in the raw data to a first preset state, such as 20 ng / mL, the upper limit of the normal range, while maintaining all other features unchanged. This modified data is then input into the deep learning model again to obtain a second risk value, such as 0.40. Similarly, to generate the second simulated medical data, the system modifies the image description of the intrahepatic nodule in the original data to a second preset state, for example, "no abnormal description," while keeping all other features unchanged. This modified data is input into the deep learning model to obtain a third risk value, for example, 0.50. Furthermore, to generate the third simulated medical data, the system modifies the serum alpha-fetoprotein (AFP) value in the original data to 20 ng / mL and simultaneously modifies the image description of the intrahepatic nodule to "no abnormal description." This modified data is input into the deep learning model to obtain a fourth risk value, for example, 0.20. Finally, the system performs a differential operation on these risk values. For example, a preset formula can be used to calculate the independent risk contribution of the first risk-indicating feature (AFP), the independent risk contribution of the second risk-indicating feature (the intrahepatic nodule), and the interactive risk contribution between the two. These calculations can be performed on a central processing unit or graphics processing unit and implemented using a software program. These quantitative results can then be presented to clinicians in the form of a report to assist in diagnostic and treatment decisions.
[0039] Through the above technical solutions, this application can effectively solve the problem that existing liver cancer risk assessment methods cannot accurately distinguish and quantify the independent contributions and interactive contributions of different risk indicator features in a specific clinical context. This application systematically simulates the preset states of different risk factors and combines differential operations to achieve a refined analysis of the risk drivers of liver cancer. This enables clinicians to clearly identify which risk factors act independently and which factors have synergistic or amplifying effects, thereby avoiding misinterpretation of a single indicator and improving the accuracy and reliability of risk assessment. This quantitative analysis capability provides data support for the development of personalized diagnosis and treatment plans, helps optimize clinical decision-making processes, and improves patient management levels.
[0040] In some embodiments, the differential operation is performed as follows: IR1=R3-R4 IR2= R2-R4 IR12=R1-R2-R3+R4 Among them, R1 is the first risk value, R2 is the second risk value, R3 is the third risk value, R4 is the fourth risk value, IR1 is the independent risk contribution of the first risk indication feature, IR2 is the independent risk contribution of the second risk indication feature, and IR12 is the interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0041] Among them, R1 refers to the first risk value generated after the original medical data is processed by the preset risk assessment model, which serves as the benchmark risk value for subsequent differential operations.
[0042] R2 refers to the second risk value generated after the first simulated medical data is processed by the preset risk assessment model after only the first risk indication feature is modified to the first preset state, reflecting the risk change when only the first risk indication feature is changed.
[0043] R3 refers to the third risk value generated after the second simulated medical data is processed by the preset risk assessment model after only the second risk indication feature is modified to the second preset state, reflecting the risk change when only the second risk indication feature is changed.
[0044] R4 refers to the fourth risk value generated after the third simulated medical data is processed by the preset risk assessment model after the first risk indication feature is modified to the first preset state and the second risk indication feature is modified to the second preset state, reflecting the background risk when both risk indication features are changed.
[0045] IR1 is the independent risk contribution of the first risk indicator characteristic, which is used to quantify the independent impact of the first risk indicator characteristic on the total risk. IR2 is the independent risk contribution of the second risk indicator characteristic, which is used to quantify the independent impact of the second risk indicator characteristic on the total risk. IR12 is the interactive risk contribution of the first and second risk indicator characteristics, which is used to quantify the nonlinear risk change caused by the combined effect of the two risk indicator characteristics.
[0046] The present application provides a specific differential operation formula to achieve quantification of the independent risk contribution of the first risk indication feature and the second risk indication feature, as well as the interactive risk contribution therebetween.
[0047] Specifically, the independent risk contribution IR1 of the first risk indicator characteristic is calculated by subtracting the fourth risk value R4 from the third risk value R3. This can isolate the incremental risk caused solely by changes in the first risk indicator characteristic. Similarly, the independent risk contribution IR2 of the second risk indicator characteristic is calculated by subtracting the fourth risk value R4 from the second risk value R2, thereby quantifying the incremental risk caused solely by changes in the second risk indicator characteristic.
[0048] Furthermore, the interactive risk contribution IR12 of the first risk indicator feature and the second risk indicator feature is calculated by subtracting the second risk value R2 and the third risk value R3 from the first risk value R1, and then adding back the fourth risk value R4. This can reveal the nonlinear risk changes caused by the two risk indicator features acting together, that is, the additional risk or risk reduction brought about by their mutual influence. It is precisely because of the introduction of these formulas that after obtaining the first risk value, second risk value, third risk value and fourth risk value corresponding to the original medical data, the first simulated medical data, the second simulated medical data and the third simulated medical data, it is possible to accurately decompose the total risk, identify the independent effects of each risk factor and the complex interactions between them, and thus solve the problem of how to accurately perform differential operations to quantify risk contributions.
[0049] For example, suppose that after a patient's original medical data is processed by a preset risk assessment model, the first risk value R1 generated is 0.8. When only the first risk indication feature (such as serum alpha-fetoprotein AFP) in the original medical data is modified to a first preset state (such as AFP at the upper limit of the normal range) to generate the first simulated medical data, and after processing by the preset risk assessment model, the second risk value R2 generated is 0.5. When only the second risk indication feature (such as the morphology of the intrahepatic nodule) in the original medical data is modified to a second preset state (such as no abnormal description) to generate the second simulated medical data, and after processing by the preset risk assessment model, the third risk value R3 generated is 0.6. When the first risk indication feature is modified to the first preset state and the second risk indication feature is modified to the second preset state to generate the third simulated medical data, and after processing by the preset risk assessment model, the fourth risk value R4 generated is 0.2. Based on these risk values, a differential operation can be performed according to the following formula: The independent risk contribution of the first risk indicator characteristic is IR1 = R3 - R4 = 0.6 - 0.2 = 0.4.
[0050] The independent risk contribution of the second risk indicator characteristic is IR2 = R2 - R4 = 0.5 - 0.2 = 0.3.
[0051] The interactive risk contribution of the first risk indicator feature and the second risk indicator feature is IR12 = R1 - R2 - R3 +R4 = 0.8 - 0.5 - 0.6 + 0.2 = -0.1.
[0052] Through the above calculations, we can obtain the specific values of each risk contribution. For example, the first risk indicator feature independently contributes 0.4, the second risk indicator feature independently contributes 0.3, and the interaction contribution between the two is -0.1. This shows that in a specific context, the joint effect of these two features may lead to nonlinear changes in risk and even reduce the total risk to some extent.
[0053] The above technical solution provides a clear set of differential calculation formulas that can accurately quantify the independent risk contribution of the first risk indicator feature, the independent risk contribution of the second risk indicator feature, and the interactive risk contribution of the first and second risk indicator features. This enables more in-depth and accurate analysis of complex risk factors, helping to identify whether there are synergistic or antagonistic effects between different risk factors in specific contexts, thereby providing more refined data support for clinical decision-making.
[0054] In some embodiments, the first preset state is determined by the following steps, specifically including: In the original medical data, the second risk indication feature is kept unchanged and only the value of the first risk indication feature is changed to generate test data; Processing the test data using the preset risk assessment model to obtain a test risk value corresponding to a first risk indication feature value in each test data; Determining a trigger threshold according to a correspondence between the value of the first risk indication feature and the test risk value, and determining a first preset state based on the trigger threshold; Among them, the trigger threshold is the inflection point position where the risk value in the corresponding relationship changes from linear change to nonlinear change; the first preset state is set to a calibration state, and the value of the calibration state is close to the trigger threshold and does not meet the trigger condition defined by the trigger threshold.
[0055] Among them, test data refers to a series of simulated data generated on the basis of original medical data by keeping the numerical value of the second risk indication feature unchanged and only systematically changing the numerical value of the first risk indication feature, which is used to observe the impact of the first risk indication feature on the risk assessment results in isolation while controlling other variables.
[0056] The test risk value refers to the risk assessment result obtained for each specific value of the first risk indicator feature in each test data after processing the test data using a preset risk assessment model. These risk values collectively depict the functional relationship between the change in the value of the first risk indicator feature and the risk assessment result.
[0057] The correspondence relationship refers to the mapping between the numerical value of the first risk indicator feature and the test risk value obtained after processing using the preset risk assessment model. This relationship can be expressed as a curve, chart, or mathematical function to reveal how the first risk indicator feature affects the risk assessment results.
[0058] The trigger threshold is the inflection point in the relationship between the value of the first risk indicator feature and the test risk value, where the risk value changes from a linear to a nonlinear change. This inflection point marks the critical point where the impact of the first risk indicator feature on the risk value changes significantly. For example, after exceeding this point, the risk value may show an accelerating upward or downward trend.
[0059] The calibration state refers to a specific value that is determined to be the first preset state. The calibration state value is set to be close to the trigger threshold, but its value itself does not meet the trigger condition defined by the trigger threshold. This means that the calibration state is a safe or baseline value used as a reference point in subsequent risk assessments to avoid misidentifying normal or low-risk situations as high-risk.
[0060] This application solves the problem of assessment bias that may exist in the preset risk assessment model under specific circumstances by systematically determining the first preset state.
[0061] Specifically, the present application first generates a series of test data by maintaining the value of the second risk indicator feature in the original medical data and changing only the value of the first risk indicator feature. This operation aims to examine the impact of the first risk indicator feature on the risk assessment results in isolation while controlling for other variables.
[0062] Subsequently, these test data are processed using a preset risk assessment model to obtain a test risk value corresponding to the numerical value of the first risk indicator feature in each test data. Through this process, a clear correspondence is established between the numerical value of the first risk indicator feature and the test risk value, which reveals how the first risk indicator feature affects the risk assessment results. Furthermore, based on this correspondence, the inflection point where the risk value changes from linear to nonlinear change is identified, and this inflection point is determined as the trigger threshold. This trigger threshold marks the critical point at which the impact of the first risk indicator feature on the risk value changes significantly.
[0063] Finally, the first preset state is set to a calibration state, and its value is set to be close to the trigger threshold, but does not meet the trigger condition defined by the trigger threshold. This means that the calibration state is a safe value slightly lower than the critical point of a sharp increase in risk, and can be used as a calibration benchmark for the first risk indication feature. In this way, the present application can dynamically and accurately determine the first preset state based on the actual response characteristics of the existing model, avoiding the deviation that may be caused by a simple preset. When this precisely determined first preset state is used to generate the first simulated medical data and the third simulated medical data, it can provide a more accurate benchmark for the subsequent differential calculation of the independent risk contribution and the interactive risk contribution of the first risk indication feature, thereby making the final risk contribution quantification result more clinically instructive, and effectively solving the problem mentioned in the background technology that in a specific clinical context, the existing model may overestimate or misjudge the risk contribution assessment of serological indicators (such as FP).
[0064] In some embodiments, the specific process of determining the first preset state can be performed as follows. For example, assume that the first risk-indicating feature is the serum alpha-fetoprotein (α-fetoprotein) value, and the second risk-indicating feature is the imaging description of an intrahepatic nodule. First, various medical indicators of the patient can be extracted from a set of original medical data. Based on this, the imaging description of the patient's intrahepatic nodule remains unchanged, and only the α-fetoprotein (α-fetoprotein) value is systematically varied. For example, starting from a low value (e.g., 5 ng / mL) and gradually increasing in preset steps (e.g., 5 ng / mL) to a higher value (e.g., 500 ng / mL), a series of test data is generated. Each test data set contains a specific α-fetoprotein (α-fetoprotein) value and the original imaging description of the intrahepatic nodule. These generated test data sets are then input into an existing liver cancer risk assessment model for processing. This preset risk assessment model can be a deep learning-based neural network model that has been trained with a large amount of clinical data and can output a liver cancer risk value based on the input medical characteristics. For each input test data set, the model outputs a corresponding test risk value. This yields a series of pairs of FP values and corresponding test risk values. These pairs can be plotted as a curve to visually demonstrate the impact of FP value changes on risk values. Furthermore, this curve can be analyzed to determine the trigger threshold.
[0065] For example, mathematical methods such as piecewise linear regression or spline interpolation can be used to identify the inflection point where the slope of the curve changes significantly. When the FP value is low, the risk value may show a slow linear increase; however, when the FP value reaches a certain critical point, the risk value may begin to rise sharply, showing a nonlinear growth. This turning point from linear to nonlinear change is the trigger threshold. For example, if it is found that the risk value begins to rise rapidly when the FP value exceeds 100 ng / mL, then 100 ng / mL can be determined as the trigger threshold. Finally, the first preset state is set to the calibration state. The value of this calibration state is set to be close to the determined trigger threshold, but its value itself does not meet the trigger condition defined by the trigger threshold. For example, if the trigger threshold is 100 ng / mL and the trigger condition is "greater than or equal to 100 ng / mL," the value of the calibration state can be set to 99 ng / mL. In this way, when the FP value needs to be set to the "upper limit of the normal range" in subsequent risk assessments, this 99 ng / mL can be used as the first preset state, thereby ensuring that the preset value can reflect the higher normal level of FP and will not be misjudged as high risk by the existing model, providing a more accurate and reasonable benchmark point for subsequent differential operations.
[0066] Through the above technical solution, the present application can dynamically and accurately determine the first preset state based on the actual response characteristics of the preset risk assessment model. This avoids the evaluation bias that may be caused by the use of fixed or empirical thresholds, so that the first preset state can more accurately characterize the "normal" or "benchmark" level of the first risk indication feature in the model. Therefore, when the first preset state is applied to the subsequent quantitative calculation of risk contribution, it can provide a more reasonable and reliable reference benchmark, thereby improving the accuracy of liver cancer risk assessment, especially in complex situations where the value of the first risk indication feature is increased but its clinical significance is unclear, it can effectively distinguish the influence of different risk factors and provide more accurate data support for clinical decision-making.
[0067] In some embodiments, the step of generating test data by keeping the second risk indication feature unchanged and only changing the value of the first risk indication feature in the original medical data includes: Performing a unidirectional scan within a preset value range to generate a first subset of test data, wherein, in the unidirectional scan, the value of the first risk indication feature changes in one direction; Performing a reverse scan within the preset value range to generate a second subset of test data, wherein a direction of change of the value of the first risk indication feature in the reverse scan is opposite to a direction of change of the value of the first risk indication feature in the unidirectional scan; Combining the first subset test data and the second subset test data to form the test data; The test data is processed according to the preset risk assessment model to obtain the test risk value, where the test risk value includes a first response path corresponding to the unidirectional scan and a second response path corresponding to the reverse scan.
[0068] The preset numerical range refers to the numerical interval defined by the first risk indicator feature during testing, which can be determined based on historical data statistics, expert experience, or model sensitivity analysis, to ensure that the test data covers the effective variation range of the risk indicator feature; Among them, unidirectional scanning refers to the process in which the value of the first risk indication feature gradually increases or decreases in a fixed direction (for example, from small to large or from large to small) within a preset numerical range. It can be achieved by increasing in equal steps, increasing in unequal steps, or increasing based on a specific function curve, etc., to obtain the changing trend of the risk value in a single direction.
[0069] Among them, reverse scanning refers to the process in which the value of the first risk indication feature gradually changes in the opposite direction of the unidirectional scanning within a preset numerical range. It can be achieved by adopting a decreasing or increasing method with the same step size as the unidirectional scanning but in the opposite direction, and is used to obtain the changing trend of the risk value in the opposite direction and detect possible hysteresis or asymmetric effects.
[0070] Among them, the first subset test data refers to a set of simulated data points for risk assessment generated by a unidirectional scanning method. It can be constructed on the basis of the original medical data by only modifying the value of the first risk indication feature and keeping other features unchanged. It can reflect the impact of the unidirectional change of the first risk indication feature on the risk value.
[0071] Among them, the second subset test data refers to a series of simulated data point sets for risk assessment generated by reverse scanning. It can be constructed on the basis of the original medical data by only modifying the value of the first risk indication feature and keeping other features unchanged, which can reflect the impact of the reverse change of the first risk indication feature on the risk value.
[0072] Among them, the first response path refers to the trend curve or data sequence in which the test risk value obtained after the preset risk assessment model processes the first subset of test data changes unidirectionally with the first risk indication feature. It can be represented by a set of data points, a function curve or a chart, which can intuitively show the change pattern of the risk value under unidirectional scanning.
[0073] Among them, the second response path refers to the trend curve or data sequence in which the test risk value obtained after the preset risk assessment model processes the second subset of test data changes in the opposite direction of the first risk indication feature. It can be represented by a set of data points, a function curve or a chart, which can intuitively show the change pattern of the risk value under reverse scanning.
[0074] For example, when determining the first preset state of serum alpha-fetoprotein (FP) as the first risk-indicating feature, test data can be generated and a test risk value obtained in the following manner. First, a preset range of serum alpha-fetoprotein (FP) values is set, for example, from 5 ng / mL to 500 ng / mL. Then, a unidirectional scan is performed, for example, starting from 5 ng / mL and increasing in steps of 10 ng / mL, to generate a series of first subset test data points, for example, 5 ng / mL, 15 ng / mL, 25 ng / mL, and so on, up to 500 ng / mL. These data points are combined with other unchanged features in the original medical data (e.g., imaging features of intrahepatic nodule morphology) to form simulated medical data. Next, this first subset test data is input into a preset risk assessment model for processing, resulting in a test risk value corresponding to each serum alpha-fetoprotein (FP) value. These risk values constitute the first response path. Subsequently, a reverse scan is performed, for example, starting at 500 ng / mL and decreasing in steps of 10 ng / mL, to generate a second subset of test data, for example, 500 ng / mL, 490 ng / mL, 480 ng / mL, and so on, down to 5 ng / mL. Similarly, these data points are combined with other features of the original medical data, which remain unchanged, to form simulated medical data. This second subset of test data is then input into a pre-defined risk assessment model for processing, generating a test risk value corresponding to each serum alpha-fetoprotein (α-fetoprotein) value. These risk values constitute the second response path. Finally, the first subset of test data generated by the unidirectional scan and the second subset of test data generated by the reverse scan are combined to form a complete test data set. This complete test data set is processed by the pre-defined risk assessment model, and the resulting test risk value will include both the first and second response paths. For example, the first response path may show the trend of the risk value when the FP changes from low to high, while the second response path may show the trend of the risk value when the FP changes from high to low. By comparing these two paths, we can more comprehensively observe the sensitivity of the risk value to FP changes, such as whether there is hysteresis or asymmetric response, thereby providing more reliable data support for the subsequent precise determination of the trigger threshold.
[0075] The above technical solution generates test data by performing unidirectional and reverse scans within a preset numerical range, and obtains a test risk value encompassing the first and second response paths. This more completely covers the variations of the first risk indicator characteristic across the entire preset numerical range, and captures the differences in the response of the risk value in different directions of variation. This accurately reflects the complete trend of the risk value as it changes with the first risk indicator characteristic, providing a more comprehensive and reliable data foundation for the subsequent determination of the trigger threshold, significantly improving the accuracy of the determination of the first preset state.
[0076] In some embodiments, the step of determining a trigger threshold according to the correspondence between the value of the first risk indication feature and the test risk value includes: Identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining a value of the first risk indication feature corresponding to the location as a first candidate threshold; Identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining a value of the first risk indication feature corresponding to the location as a second candidate threshold; The smaller one between the first candidate threshold and the second candidate threshold is determined as the trigger threshold.
[0077] Among them, the location where the risk value undergoes nonlinear change refers to the point where the rate of change of the risk value changes significantly when the risk value changes with the first risk indication characteristic value, for example, the inflection point where a gentle change suddenly turns into a sharp rise or fall. It can be identified by mathematical analysis methods, such as by calculating the second derivative of the risk value change rate or by analyzing the sudden change of the risk value slope through a sliding window, to accurately capture the key critical points of the risk state transition.
[0078] The first candidate threshold refers to the value of the first risk indication feature corresponding to the position of the nonlinear change of the risk value identified from the first response path, which serves as a reference point for determining the final trigger threshold.
[0079] The second candidate threshold refers to the value of the first risk indication feature corresponding to the position of the nonlinear change of the risk value identified from the second response path, which serves as a reference point for determining the final trigger threshold.
[0080] The trigger threshold refers to the critical value that is ultimately determined to define the transition of risk status or trigger specific conditions. Its purpose is to provide an accurate and robust basis for risk judgment.
[0081] This application determines the trigger threshold by conducting an in-depth analysis of the correspondence between the values of the first risk indicator feature and the test risk value, specifically focusing on locations where the risk value undergoes nonlinear changes. Specifically, after obtaining a first response path generated by a unidirectional scan and a second response path generated by a reverse scan, the application identifies key locations within each path where the risk value undergoes nonlinear changes. These locations represent the inflection points where the risk value transitions from linear to nonlinear change, and are critical points where a sudden change in the risk state may occur. The values of the first risk indicator feature corresponding to the locations of nonlinear changes identified in the first response path are determined as the first candidate threshold, while the values of the first risk indicator feature corresponding to the locations of nonlinear changes identified in the second response path are determined as the second candidate threshold. Ultimately, by comparing these two candidate thresholds, the smaller value is selected as the final trigger threshold. This combination of bidirectional scanning and dual-path analysis effectively captures the varying characteristics of the risk value across different scanning directions, avoiding the potential bias introduced by unidirectional scanning. By selecting a smaller value as the trigger threshold, the risk assessment is conservative, reducing the probability of falsely misjudging a low risk, and thus improving the accuracy and reliability of the trigger threshold determination. This method, combined with the previous steps of generating response paths through one-way scanning and reverse scanning, forms a more complete threshold determination mechanism, which can more accurately locate the critical point of risk transformation. This two-way verification strategy is particularly important when the risk model response is complex or there is a lag effect.
[0082] For example, assume that in liver cancer risk detection, the first risk indicator characteristic is the serum alpha-fetoprotein (AFP) value. Within a preset value range, a series of test data is generated by unidirectional scanning (e.g., increasing from low AFP values to high AFP values), and corresponding test risk values are obtained using a preset risk assessment model, thereby forming a first response path. Simultaneously, another series of test data is generated by reverse scanning (e.g., decreasing from high AFP values to low AFP values), and corresponding test risk values are obtained, forming a second response path. Specifically, to identify locations where nonlinear changes in risk values occur, the data points in the first response path can be analyzed. For example, a sliding window method can be used to calculate the slope or rate of change of the risk value within a certain range before and after each data point. When the absolute value of the slope or rate of change increases significantly and abruptly near a data point, that point can be identified as a location of nonlinear change. The AFP value corresponding to that location is recorded as the first candidate threshold value. Similarly, the same analysis process is performed on the data points in the second response path to identify locations where nonlinear changes in risk values occur, and the corresponding AFP value is recorded as the second candidate threshold value. For example, if the first response path analysis determines a first candidate threshold of 120 ng / mL, and the second response path analysis determines a second candidate threshold of 115 ng / mL, then according to the present application, the smaller of these two values, 115 ng / mL, is determined as the final trigger threshold. This trigger threshold can be used to subsequently determine the first preset state. For example, the first preset state can be set to a value slightly below 115 ng / mL to represent the upper limit of the normal range of AFP, thereby providing a calibration benchmark for risk assessment. This approach ensures that the critical point of risk change can be robustly captured under different scanning directions and a more conservative threshold is selected to improve the accuracy of risk assessment.
[0083] The above technical solution accurately determines the trigger threshold based on the corresponding relationship between the numerical value of the first risk indicator feature and the test risk value. By identifying the locations where the risk value undergoes nonlinear changes from the first response path generated by the unidirectional scan and the second response path generated by the reverse scan, and combining the results of the two paths to select the smaller numerical value as the trigger threshold, the bias that may be introduced by a single scanning direction can be effectively eliminated, improving the accuracy and robustness of the trigger threshold determination. This allows for the precise identification of critical points in risk state transitions when the risk value exhibits complex nonlinear changes, providing a more reliable basis for subsequent risk assessment and determination of the preset state.
[0084] In some embodiments, the steps of identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining the value of the first risk indication feature corresponding to the location as a first candidate threshold, and identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining the value of the first risk indication feature corresponding to the location as a second candidate threshold, include: During the process of performing the unidirectional scan to generate the first subset test data and performing the reverse scan to generate the second subset test data, maintaining the second risk indication feature as a state recorded in the original medical data; Processing the first subset of test data and the second subset of test data according to the preset risk assessment model to obtain the first response path and the second response path; A location where the risk value changes nonlinearly is identified from the first response path to determine the first candidate threshold, and a location where the risk value changes nonlinearly is identified from the second response path to determine the second candidate threshold.
[0085] Among them, maintaining the second risk indication feature as the state recorded in the original medical data means that when generating test data, in addition to the first risk indication feature of the target analysis, the values of other non-target analysis risk indication features (such as the second risk indication feature) are fixed to their actual recorded values in the original medical data. Specifically, during the data simulation or generation process, the value of the second risk indication feature is locked so that it does not change with the change of the first risk indication feature, so as to isolate and eliminate the influence of other risk indication features on the change of risk value, ensure that the observed risk value change is only caused by the change of the first risk indication feature, thereby improving the accuracy and reliability of the analysis.
[0086] In some preferred embodiments, the present application is specifically implemented as follows. Assume that in the original medical data, the first risk indication feature is serum alpha-fetoprotein AFP, and the second risk indication feature is the imaging feature of the intrahepatic nodule morphology. When determining the trigger threshold of AFP, it is necessary to generate a series of test data to observe the impact of AFP changes on the risk value. Specifically, in the process of performing a one-way scan to generate the first subset of test data, and performing a reverse scan to generate the second subset of test data, the system will first read the imaging feature value of the patient recorded in the original medical data. For example, if the imaging feature in the original data is described as "LR-3", then when generating all test data, the value of the imaging feature will always be fixed to "LR-3" and will not change with the change of the AFP value. Subsequently, the system will process the first subset of test data and the second subset of test data in which the AFP values change and the imaging features are fixed according to the preset liver cancer risk assessment model, thereby obtaining two risk response paths. For example, the first response path may show that when AFP gradually increases from a low value, the risk value begins to change from a slow linear increase to a rapid nonlinear increase at a certain point (e.g., AFP = 100 ng / mL). The second response path may show that when AFP gradually decreases from a high value, the risk value changes from a rapid nonlinear decrease to a slow linear decrease at another point (e.g., AFP = 95 ng / mL). The system will identify the locations where the nonlinear changes in the risk value occur in each of these two paths. For example, the AFP value corresponding to the inflection point of the first response path is determined as the first candidate threshold value of 100 ng / mL, and the AFP value corresponding to the inflection point of the second response path is determined as the second candidate threshold value of 95 ng / mL. In this way, when analyzing the impact of AFP on the risk value, the interference caused by imaging features is eliminated, making the identified threshold more accurate.
[0087] Through the above technical solution, during the process of generating test data, by maintaining the second risk indicator feature in the state recorded in the original medical data, the interference of other risk indicator features on the risk assessment of the target risk indicator feature (the first risk indicator feature) is effectively isolated. This enables the obtained first response path and second response path to more accurately reflect the true relationship between the numerical change of the first risk indicator feature and the risk value, thereby more accurately identifying the location where the risk value undergoes nonlinear changes. Ultimately, this improves the accuracy of determining the first and second candidate thresholds, providing a more reliable basis for the subsequent determination of the final trigger threshold, thereby improving the accuracy and reliability of the entire liver cancer risk detection method.
[0088] In some embodiments, the steps of identifying a location where the risk value undergoes a nonlinear change in the first response path to determine the first candidate threshold, and identifying a location where the risk value undergoes a nonlinear change in the second response path to determine the second candidate threshold, include: For any one of the first response path and the second response path, perform the following steps: Traversing the data points in the response path, and defining, for each data point, a pre-change data segment consisting of a group of adjacent data points before the data point, and a post-change data segment consisting of a group of adjacent data points after the data point; Calculating a first statistic based on the risk value within the pre-change data segment defined for each data point, and calculating a second statistic based on the risk value within the post-change data segment; determining a variation range of each data point based on the first statistic and the second statistic calculated for the data point; The data point with the largest variation amplitude is determined as the location where the risk value in any response path undergoes nonlinear variation.
[0089] The pre-change data segment refers to a data set consisting of a series of continuous adjacent data points before a specific data point in the response path, and is used to capture the changing trend or stable state of the risk value before the specific data point.
[0090] The post-change data segment refers to a data set consisting of a series of continuous adjacent data points following a specific data point in the response path, and is used to capture the changing trend or stable state of the risk value after the specific data point.
[0091] The first statistic refers to a quantitative indicator calculated based on the risk value in the pre-change data segment. It can be the average, median, maximum, minimum, standard deviation, slope of the risk value in the data segment, or other statistical measures that can reflect the overall risk level or change trend of the data segment. It is used to generally characterize the characteristics of the risk value before the data point.
[0092] The second statistic refers to a quantitative indicator calculated based on the risk value in the post-change data segment. It can be the average, median, maximum, minimum, standard deviation, slope of the risk value in the data segment, or other statistical measures that can reflect the overall risk level or change trend of the data segment. It is used to generally characterize the characteristics of the risk value after the data point.
[0093] The amplitude of change refers to an indicator used to quantify the degree of change in the risk value before and after a certain data point. It is a value calculated by comparing the first statistic and the second statistic of the data point. For example, it can be the absolute difference, relative difference, ratio between the two, or the difference calculated based on a specific function. It can intuitively reflect the severity of the change in the risk value at the data point.
[0094] Specifically, when identifying the locations where the risk values in the first response path and the second response path undergo nonlinear changes, the following operations can be performed for any response path, such as the first response path. First, the system will traverse each data point in the first response path. For each data point, such as data point A, the system will define a pre-change data segment and a post-change data segment. The pre-change data segment can be composed of a preset number of adjacent data points before data point A. Similarly, the post-change data segment can be composed of a preset number of adjacent data points after data point A. Then, the system will calculate the first statistic based on the risk value in the pre-change data segment. For example, the average value of all risk values in the pre-change data segment can be calculated, or the slope of its linear regression line can be calculated.
[0095] Simultaneously, the system calculates a second statistic based on the risk values within the post-change data segment. For example, it similarly calculates the average of all risk values within the post-change data segment, or the slope of its linear regression line. Subsequently, the system determines the magnitude of change for data point A based on the calculated first and second statistics. For example, if the statistic is an average, the magnitude of change can be calculated as the absolute difference between the first and second statistics; if the statistic is a slope, the magnitude of change can be calculated as the difference between the two slopes. This calculation process is repeated for each data point in the response path, determining the magnitude of change for each data point. Finally, the system compares the magnitudes of change for all data points and identifies the data point with the largest magnitude of change as the location in the first response path where the risk value undergoes a nonlinear change. The data point at this location is considered to have undergone the most significant nonlinear change in risk value, and thus serves as the basis for determining the first candidate threshold. For the second response path, the same method is used to identify the location where the risk value undergoes a nonlinear change, thereby determining the second candidate threshold.
[0096] The above technical solution enables precise identification of key locations where nonlinear changes in risk values occur by quantitatively analyzing the risk value trends before and after each data point in the response path. This approach avoids the inaccuracies that can arise from simple threshold determination or local difference comparisons, effectively addressing the challenges of numerous data points and complex change patterns in the response path. Consequently, the first and second candidate thresholds can be determined more accurately, improving the accuracy of subsequent trigger threshold determination and, in turn, the reliability of liver cancer risk detection.
[0097] In some embodiments, the step of calculating the first statistic based on the risk value in the pre-change data segment defined for each data point, and calculating the second statistic based on the risk value in the post-change data segment, comprises: For any data segment of the pre-change data segment and the post-change data segment, perform the following steps: Determine the distribution characteristics of risk values within the data segment; According to the distribution characteristics, a corresponding calculation rule is selected for the data segment from a preset calculation rule set; The calculation rule selected for the data segment is used to calculate and obtain the first statistic or the second statistic corresponding to the data segment.
[0098] Among them, the distribution characteristics of risk values refer to the arrangement rules and central trends of risk values on the numerical axis within the data segment. They can be described by statistical methods, for example, by calculating indicators such as skewness, kurtosis, mean, median, mode, variance, standard deviation, and interquartile range of the data, or by visually presenting them through drawing histograms, kernel density estimation plots, and other visual methods, to reveal the inherent structure and statistical laws of risk values within the data segment, and provide a basis for the subsequent selection of appropriate statistical calculation methods.
[0099] The preset calculation rule set is a set of predefined algorithms or formulas for calculating the statistics of data segments. It can include a variety of statistical methods, such as arithmetic mean, weighted mean, median, mode, standard deviation, variance, interquartile range, geometric mean, harmonic mean, etc., which are used to provide a variety of statistical calculation options to adapt to data segments with different distribution characteristics and ensure the accuracy and applicability of statistical calculations.
[0100] The corresponding calculation rule is a specific calculation method that is selected from a preset set of calculation rules based on the distribution characteristics of the risk value in the data segment and that can most accurately reflect the characteristics of the data segment. It can be selected based on the distribution characteristics. For example, for data with an approximately normal distribution, the mean and standard deviation can be selected; for data with a skewed distribution or containing outliers, the median and interquartile range can be selected to ensure that the calculated statistics can represent the actual situation of the data segment to the greatest extent possible and avoid bias introduced due to improper calculation methods.
[0101] In some embodiments, the step of determining the distribution characteristics of the risk value within the data segment includes: Calculate the dispersion index of the risk value within the data segment; The dispersion index is calculated and determined as the distribution characteristic of the risk value in the data segment.
[0102] Among them, the dispersion index refers to a quantitative value used to measure the degree of dispersion or fluctuation between each value in the data set. It can be achieved using statistics such as variance, standard deviation, range or interquartile range, and is used to quantify the fluctuation range and central tendency of the risk value within the data segment.
[0103] This application quantifies the fluctuation of the risk value within a data segment by first calculating the dispersion index of the risk value within the data segment, such as the variance or standard deviation. The magnitude of the dispersion index directly reflects the degree of concentration or dispersion of the risk value within the data segment. Subsequently, the calculated dispersion index is directly used as the distribution characteristic of the risk value within the data segment. This method makes the judgment of the risk value distribution direct and quantifiable. When the dispersion index is high, it indicates that the risk value distribution is relatively dispersed, and there may be outliers or large fluctuations; when the dispersion index is low, it indicates that the risk value distribution is relatively concentrated, and the differences between data points are small.
[0104] The above technical solution effectively determines the distribution characteristics of risk values within a data segment. By calculating the dispersion index and using it as the distribution characteristic, it provides a clear basis for selecting appropriate statistical calculation rules from a set of preset calculation rules. This allows the calculation of statistics to better adapt to data segments with different risk value distributions, thereby improving the accuracy and robustness of identifying locations where risk values undergo nonlinear changes.
[0105] On the other hand, Figure 2 As shown, a liver cancer risk detection system is exemplarily shown. This application further proposes a liver cancer risk detection system 100, which includes: The original medical data acquisition module 10 is configured to acquire original medical data including a first risk indication feature and a second risk indication feature, and process the original medical data according to a preset risk assessment model to generate a first risk value; a first single-factor risk assessment module 20 for generating, based on the original medical data, first simulated medical data by modifying a first risk-indicating feature in the original medical data to a first preset state, and processing the first simulated medical data according to the preset risk assessment model to generate a second risk value; a second single-factor risk assessment module 30 for generating, based on the original medical data, second simulated medical data by modifying a second risk-indicating feature in the original medical data to a second preset state, and processing the second simulated medical data according to the preset risk assessment model to generate a third risk value; a background risk assessment module 40 for generating, based on the original medical data, third simulated medical data by modifying the first risk indication feature to the first preset state and the second risk indication feature to the second preset state, and processing the third simulated medical data according to the preset risk assessment model to generate a fourth risk value; The contribution quantification calculation module 50 performs a differential operation based on the first risk value, the second risk value, the third risk value and the fourth risk value to determine the independent risk contribution of the first risk indication feature, the independent risk contribution of the second risk indication feature, and the interactive risk contribution of the first risk indication feature and the second risk indication feature.
[0106] Through the above technical solution, the present application provides a liver cancer risk detection system. The system realizes the detection and evaluation of liver cancer risk through modular design, and solves the problem of lack of physical system support and difficulty in implementing liver cancer risk detection methods in the existing technology. The system can obtain original medical data and perform preliminary risk assessment, and quantify the independent contribution of each risk indicator feature and the interactive contribution between them by simulating the presence or absence of different risk factors. This enables clinicians to understand the source and composition of the risk, thereby providing data support for the formulation of diagnosis and treatment plans, avoiding the decision-making dilemma caused by unclear contribution of risk factors, and improving the accuracy of liver cancer risk assessment and clinical guidance value.
[0107] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for detecting the risk of liver cancer, characterized in that: include: Obtaining original medical data including a first risk indication feature and a second risk indication feature, and processing the original medical data according to a preset risk assessment model to generate a first risk value; wherein the first risk indication feature is a serum feature including serum alpha-fetoprotein (AFP); and the second risk indication feature is an imaging feature of the intrahepatic nodule morphology; Modifying a first risk-indicating feature in the original medical data to a first preset state to generate first simulated medical data, and processing the first simulated medical data according to the preset risk assessment model to generate a second risk value; wherein the first preset state indicates that the serum characteristic is at an upper limit of a normal range; Modifying the second risk indication feature in the original medical data to a second preset state to generate second simulated medical data, and processing the second simulated medical data according to the preset risk assessment model to generate a third risk value; wherein the second preset state indicates that the image feature has no abnormal description; modifying the first risk indication feature to the first preset state, and modifying the second risk indication feature to the second preset state to generate third simulated medical data, and processing the third simulated medical data according to the preset risk assessment model to generate a fourth risk value; Based on the first risk value, the second risk value, the third risk value and the fourth risk value, a differential operation is performed to determine an independent risk contribution of the first risk indication feature, an independent risk contribution of the second risk indication feature, and an interactive risk contribution of the first risk indication feature and the second risk indication feature.
2. The method for detecting the risk of liver cancer according to claim 1, wherein The operating formula of the difference operation is as follows: IR1=R3-R4 IR2= R2-R4 IR12=R1-R2-R3+R4 Among them, R1 is the first risk value, R2 is the second risk value, R3 is the third risk value, R4 is the fourth risk value, IR1 is the independent risk contribution of the first risk indication feature, IR2 is the independent risk contribution of the second risk indication feature, and IR12 is the interactive risk contribution of the first risk indication feature and the second risk indication feature.
3. The method for detecting the risk of liver cancer according to claim 1, wherein The first preset state is determined by the following steps, specifically including: In the original medical data, the second risk indication feature is kept unchanged and only the value of the first risk indication feature is changed to generate test data; Processing the test data using the preset risk assessment model to obtain a test risk value corresponding to a first risk indication feature value in each test data; Determining a trigger threshold according to a correspondence between the value of the first risk indication feature and the test risk value, and determining a first preset state based on the trigger threshold; Among them, the trigger threshold is the inflection point position where the risk value in the corresponding relationship changes from linear change to nonlinear change; the first preset state is set to a calibration state, and the value of the calibration state is close to the trigger threshold and does not meet the trigger condition defined by the trigger threshold.
4. The method for detecting the risk of liver cancer according to claim 3, wherein: The test data is generated by keeping the second risk indication feature unchanged in the original medical data and only changing the value of the first risk indication feature; The step of processing the test data using the preset risk assessment model to obtain a test risk value corresponding to the first risk indication feature value in each test data includes: Performing a unidirectional scan within a preset value range to generate a first subset of test data, wherein, in the unidirectional scan, the value of the first risk indication feature changes in one direction; Performing a reverse scan within the preset value range to generate a second subset of test data, wherein a direction of change of the value of the first risk indication feature in the reverse scan is opposite to a direction of change of the value of the first risk indication feature in the unidirectional scan; Combining the first subset test data and the second subset test data to form the test data; The test data is processed according to the preset risk assessment model to obtain the test risk value, where the test risk value includes a first response path corresponding to the unidirectional scan and a second response path corresponding to the reverse scan.
5. The method for detecting the risk of liver cancer according to claim 3, wherein: The step of determining a trigger threshold according to the correspondence between the value of the first risk indication feature and the test risk value includes: Identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining a value of the first risk indication feature corresponding to the location as a first candidate threshold; Identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining a value of the first risk indication feature corresponding to the location as a second candidate threshold; The smaller one between the first candidate threshold and the second candidate threshold is determined as the trigger threshold.
6. The method for detecting the risk of liver cancer according to claim 5, wherein: The steps of identifying a location where the risk value undergoes a nonlinear change from the first response path, and determining the value of the first risk indication feature corresponding to the location as a first candidate threshold, and identifying a location where the risk value undergoes a nonlinear change from the second response path, and determining the value of the first risk indication feature corresponding to the location as a second candidate threshold, include: During the process of performing the unidirectional scan to generate the first subset test data and performing the reverse scan to generate the second subset test data, maintaining the second risk indication feature as a state recorded in the original medical data; Processing the first subset of test data and the second subset of test data according to the preset risk assessment model to obtain the first response path and the second response path; A location where the risk value changes nonlinearly is identified from the first response path to determine the first candidate threshold, and a location where the risk value changes nonlinearly is identified from the second response path to determine the second candidate threshold.
7. The method for detecting the risk of liver cancer according to claim 6, wherein: The steps of identifying a location where the risk value undergoes a nonlinear change from the first response path to determine the first candidate threshold, and identifying a location where the risk value undergoes a nonlinear change from the second response path to determine the second candidate threshold, include: For any one of the first response path and the second response path, perform the following steps: Traversing the data points in the response path, and defining, for each data point, a pre-change data segment consisting of a group of adjacent data points before the data point, and a post-change data segment consisting of a group of adjacent data points after the data point; Calculating a first statistic based on the risk value within the pre-change data segment defined for each data point, and calculating a second statistic based on the risk value within the post-change data segment; determining a variation range of each data point based on the first statistic and the second statistic calculated for the data point; The data point with the largest variation amplitude is determined as the location where the risk value in any response path undergoes nonlinear variation.
8. The method for detecting the risk of liver cancer according to claim 7, wherein: The step of calculating the first statistic based on the risk value in the pre-change data segment defined for each data point, and calculating the second statistic based on the risk value in the post-change data segment, comprises: For any data segment of the pre-change data segment and the post-change data segment, perform the following steps: Determine the distribution characteristics of the risk value within the data segment; According to the distribution characteristics, a corresponding calculation rule is selected for the data segment from a preset calculation rule set; The calculation rule selected for the data segment is used to calculate and obtain the first statistic or the second statistic corresponding to the data segment.
9. The method for detecting the risk of liver cancer according to claim 8, wherein The step of determining the distribution characteristics of the risk value within the data segment includes: Calculate the dispersion index of the risk value within the data segment; The dispersion index is calculated and determined as the distribution characteristic of the risk value in the data segment.
10. A liver cancer risk detection system, characterized in that: The system includes: an original medical data acquisition module, configured to acquire original medical data including a first risk indication feature and a second risk indication feature, and process the original medical data according to a preset risk assessment model to generate a first risk value; a first single-factor risk assessment module, configured to generate first simulated medical data based on the original medical data by modifying a first risk indication feature in the original medical data to a first preset state, and process the first simulated medical data according to the preset risk assessment model to generate a second risk value; a second single-factor risk assessment module, configured to generate second simulated medical data based on the original medical data by modifying a second risk indication feature in the original medical data to a second preset state, and process the second simulated medical data according to the preset risk assessment model to generate a third risk value; a background risk assessment module for generating, based on the original medical data, third simulated medical data by modifying the first risk indication feature to the first preset state and the second risk indication feature to the second preset state, and processing the third simulated medical data according to the preset risk assessment model to generate a fourth risk value; a contribution quantification calculation module that performs a differential operation based on the first risk value, the second risk value, the third risk value, and the fourth risk value to determine an independent risk contribution of the first risk indication feature, an independent risk contribution of the second risk indication feature, and an interactive risk contribution of the first risk indication feature and the second risk indication feature.
Citation Information
Patent Citations
Construction method and application of diabetic nephropathy risk prediction model
CN114220540A
Health risk evaluation method, equipment, medium and product under coexistence of multiple combined effects
CN118534068A
Risk assessment system for chronic diseases
CN119207806A
Premature delivery prediction method and device, electronic equipment and readable storage medium
CN119446551A