Pancreatic cancer CT image prediction method and system based on iterative self-evolution

By constructing an iterative and self-evolving CT image prediction method for pancreatic cancer, and utilizing disease-specific data for fine-tuning and interactive graphic reasoning, the problems of uninterpretability and missed diagnosis in pancreatic cancer diagnosis were solved, achieving high-precision and low-cost intelligent assisted diagnosis.

CN121998972APending Publication Date: 2026-05-08QINGDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO UNIV
Filing Date
2026-03-20
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for pancreatic cancer diagnosis suffer from problems such as uninterpretable and insidious "black box" predictions, missed diagnoses of small lesions, and a scarcity of high-quality inference data, leading to diagnostic difficulties and a high rate of misdiagnosis.

Method used

A pancreatic cancer CT image prediction method based on iterative self-evolution is constructed. By introducing the FastSAM vision tool for fine-tuning disease-specific data and an interactive image and text reasoning generation architecture, combined with result consistency verification, a closed-loop training framework is formed to achieve autonomous learning of the model and improve diagnostic accuracy.

Benefits of technology

It significantly improves the detection rate and clinical interpretability of occult lesions, reduces the cost of manual annotation, and provides an intelligent assisted diagnosis and treatment solution that combines high accuracy and continuous learning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998972A_ABST
    Figure CN121998972A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image analysis and artificial intelligence auxiliary diagnosis, and more specifically relates to a pancreatic cancer CT image prediction method and system based on iterative self-evolution. The method comprises the following steps: S1, constructing a segmentation data set and a fine tuning visual tool of the pancreatic cancer special disease; s2, utilizing a pre-trained vision-language large model to generate interactive image-text reasoning; s3, reasoning data screening and training set construction based on result consistency; s4, carrying out model iteration self-evolution training based on the screened data; and S5, performing deployment and reasoning output of a final diagnosis model after multiple iterations. A FastSAM tool for special disease fine tuning and an interactive reasoning architecture are introduced, and the focus detection rate and interpretability are improved through active operation; a self-training mechanism is constructed, a closed-loop framework is formed, and dependence on manual annotation is reduced; after deployment, a visual diagnosis report is generated, and an intelligent diagnosis and treatment scheme which is high in precision, high in interpretation and capable of continuously learning is provided for clinic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image analysis and artificial intelligence-assisted diagnosis technology, and more specifically, relates to a method and system for predicting pancreatic cancer CT images based on iterative self-evolution. Background Technology

[0002] Pancreatic cancer has an insidious onset, rapid progression, and extremely high mortality rate. Early pancreatic cancer lesions often appear on CT images only as isodense or micro-low-density nodules within the pancreatic parenchyma, or as very subtle signs of pancreatic duct truncation and dilation, making them easily confused with other benign diseases, leading to difficulties in early diagnosis and a high misdiagnosis rate. With the digitization of medical imaging technology, hospitals have accumulated massive amounts of abdominal CT scan data, containing a wealth of untapped pathological morphological features. Existing technologies utilize artificial intelligence to deeply analyze this complex imaging information, achieving precise localization, characterization, and logical reasoning diagnosis of pancreatic cancer, initially improving early patient survival rates. Traditional non-reasoning medical imaging models (such as lesion classifiers based on dedicated CNN or ViT architectures) typically use sliding window mechanisms or local cropping based on lesion bounding boxes for training, concentrating all computational power on local pathological textures. While the simple Qwen3-VL-8B model supports high-resolution dynamic input when receiving complete panoramic abdominal CT images, this generates thousands of massive visual tokens. In a global view encompassing the liver, gastrointestinal tract, spleen, and abundant adipose tissue, a 1.5cm early-stage occult pancreatic micronodule occupies only a very small number of tokens. Under the global self-attention mechanism of the Transformer, the model's attention weights are easily dispersed and diluted by the vast and complex background tissue. Without active local focusing intervention, large models are inherently less sensitive to a massive number of tiny lesions than traditional classifiers that specifically fit local pixels.

[0003] Chinese patent document CN121101470A discloses a method for identifying the degree of danger by generating cancer growth videos using graph convolutional networks (GCN) and generative adversarial networks (GAN) based on prostate magnetic resonance (MRI) scan images and PSA results.

[0004] Chinese patent document CN120954689A discloses a method for segmenting and diagnosing prostate ultrasound videos using a multimodal large model (MedSAM2+CLIP).

[0005] Existing technologies and their representative improvements still have shortcomings.

[0006] Although Chinese patent document CN121101470A constructs a graph structure of "antigen test - magnetic resonance imaging", the edge features are simply defined as "test consistency", without deeply integrating clinical information (such as symptoms) and imaging details (such as lesion morphology and enhancement characteristics). The fusion dimension is single and cannot support complex reasoning. It relies on generative adversarial networks to generate cancer simulation growth videos, but does not verify the consistency between the simulation videos and the development of real lesions. Moreover, the long short-term neural network model only judges the degree of risk based on the temporal features of the video, without combining step-by-step reasoning logic, and the quantification of the degree of risk lacks clinical interpretability. It only screens subjects who need further examination by "antigen test risk value > preset threshold", without introducing preliminary screening at the imaging level. This may miss early cases with normal antigen test but abnormal imaging, which is not conducive to early cancer screening.

[0007] Chinese patent document CN120954689A relies on a two-stage process of "screening-alignment training" without designing a progressively refined reasoning mechanism. It cannot strengthen the reliability of reasoning for key areas (such as lesion details) through a step-by-step logic of "location-analysis-verification," and its ability to identify small lesions and lesions with blurred boundaries is weak. Although it uses similarity screening to compress irrelevant information, the sliding window sampling (fixed 8 frames) and clustering to select key frames lack dynamic adaptability, which is insufficient for capturing dynamic changes of lesions in ultrasound videos. Furthermore, it does not introduce an iterative data optimization mechanism, making it difficult to continuously improve data quality.

[0008] In summary, there is an urgent need to provide a method and system for predicting pancreatic cancer CT images based on iterative self-evolution, in order to solve the problems existing in the current technology. Summary of the Invention

[0009] The technical problems this invention aims to solve are: addressing the "black box" prediction uninterpretability issue in traditional deep learning models for pancreatic cancer diagnosis, and the issues of missed diagnoses and hallucinations caused by the lack of active visual interaction capabilities in existing general-purpose multimodal large models when dealing with occult and small pancreatic lesions. Furthermore, addressing the extreme scarcity of high-quality inference annotation data in the medical field, this invention constructs a closed-loop self-evolutionary training framework, utilizes fine-tuned specialized visual tools to endow the model with precise lesion capture capabilities, and automatically filters high-confidence inference data through result consistency verification. This enables continuous performance iteration and improvement of the model with low manual costs, ultimately providing clinicians with an intelligent assisted diagnosis and treatment solution for pancreatic cancer that combines expert-level diagnostic accuracy, complete process transparency, and continuous learning capabilities.

[0010] The technical problem to be solved by this invention is achieved by the following technical solution:

[0011] A method for predicting pancreatic cancer using CT images based on iterative self-evolution includes the following steps:

[0012] S1. Constructing a segmentation dataset and fine-tuning visual tools specifically for pancreatic cancer;

[0013] S2. Generate interactive text and image reasoning using a pre-trained visual-language large model;

[0014] S3. Data filtering and training set construction based on result consistency;

[0015] S4. Iterative self-evolutionary training of the model based on filtered data;

[0016] S5. Deployment and inference output of the final diagnostic model after multiple iterations.

[0017] This technical solution, by introducing the FastSAM visual tool finely tuned with disease-specific data and constructing an interactive image-text reasoning generation architecture, realizes a transformation from passive observation to active interaction in visual reasoning mode. Addressing the clinical diagnostic challenges of the occult nature and small lesions in pancreatic cancer, the model can actively execute a "planning-segmentation-focusing" sequence, transforming blurred global images into local visual evidence with clear physical support. The image-text interleaved reasoning method not only provides pixel-level image evidence for each diagnostic conclusion, effectively avoiding the visual illusion problem of large multimodal models, but also achieves white-box diagnostic logic through a visualized operational process, significantly improving the detection rate and clinical interpretability of occult lesions.

[0018] Meanwhile, by constructing a self-training mechanism based on result consistency verification, the bottleneck of scarce high-quality medical inference data has been overcome. The model's predicted conclusions are dynamically compared with the gold standard for pathology, automatically selecting logically consistent high-quality samples from massive inference paths. This forms a closed-loop training framework of "generation-selection-evolution," enabling the model to autonomously learn optimal tool invocation strategies and medical inference logic based on existing labeled data. With each iteration, the model's diagnostic performance and generalization ability exhibit a spiral improvement. While ensuring expert-level diagnostic accuracy, the reliance on manual annotation resources is significantly reduced, achieving continuous self-evolution of the model at low manual costs. This provides clinicians with an intelligent assisted diagnosis and treatment solution that combines high accuracy, strong interpretability, and continuous learning capabilities.

[0019] In addition, the pancreatic cancer CT image prediction method based on iterative self-evolution proposed according to the present invention may also have the following additional technical features:

[0020] According to a preferred embodiment of the present invention, the construction of the pancreatic cancer-specific segmentation dataset and the fine-tuning of the visual tools in step S1 includes the following specific steps:

[0021] S11. Construct a segmentation dataset specifically for pancreatic cancer:

[0022] Input CT image data and corresponding expert-annotated masks , The categories included are pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts, forming an image segmentation dataset;

[0023] S12, Building a fine-tuned visual tool for pancreatic cancer:

[0024] We selected the pre-trained FastSAM as the basic vision tool model and constructed a hybrid loss function that includes Dice loss and cross-entropy loss. By fine-tuning the parameters of FastSAM, a fine-tuned visual tool specifically for extracting pancreatic anatomy is obtained, where the objective function for parameter fine-tuning is expressed as:

[0025] (1)

[0026] In the formula: This represents the optimal visual tool model parameters obtained after fine-tuning. Let these be the model parameter variables to be optimized. This represents the operation of finding parameters that minimize the loss function; This is expressed as the expected value calculation for all sample pairs within the dataset; This is represented as the prediction output function of the visual segmentation model; It is represented as a hybrid segmentation loss function used to measure the difference between the predicted mask and the real mask.

[0027] This technical solution constructs a disease-specific dataset and fine-tunes the model, enabling precise segmentation of pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts. This provides doctors with clear and accurate anatomical information to aid in diagnosis. By utilizing a pre-trained model and a hybrid loss function for fine-tuning, the cost of de novo training is reduced, specialized tools are quickly obtained, and the efficiency of pancreatic cancer-related research is improved. Fine-tuning based on a general visual model retains a certain degree of universality, making it adaptable to different pancreatic imaging data and facilitating widespread application.

[0028] According to a preferred embodiment of the present invention, the step S2 of generating interactive text-image reasoning using a pre-trained large visual-language model includes the following specific steps:

[0029] S21. Utilize a pre-trained large-scale visual-language model For CT image data without labeled inference chains, multi-step inference is generated: the model follows a cyclical mechanism of "planning-execution-observation-description," based on historical inference paths. Generate visual operation instructions Then, the finely tuned visual tools were invoked. Execute the instructions to generate visual evidence images. ;

[0030] S23, Combination Generate textual evidence for the current step The entire chain of reasoning The generation process can be represented as maximizing the joint probability:

[0031] (2)

[0032] In the formula: This refers to generating a complete inference chain given image input and model parameters. The joint probability; This represents an unlabeled CT image that was entered. These are represented as parameters of a large vision-language model. This is represented as the visual tool obtained through fine-tuning in S1; This represents the total number of steps in the reasoning process; Represented as in the first The visual operation instructions generated step by step; Represented as the first The historical reasoning path before the step includes ; Indicated as according to instructions and images The generated visual evidence image; Represented as a text description generated based on visual evidence; It is represented as a conditional generation probability function.

[0033] This technical solution eliminates the need for manual annotation of the inference chain, allowing the model to automatically generate multi-step inference, saving labor costs and improving inference efficiency. By following a specific cyclic mechanism and based on joint probability optimization, the inference chain becomes logically coherent and reasonable, improving inference quality. Combined with visual and linguistic information, it can deeply mine the potential information in CT image data, providing richer evidence for medical diagnosis and other purposes.

[0034] According to a preferred embodiment of the present invention, the inference data filtering and training set construction based on result consistency in step S3 includes the following specific steps:

[0035] Define the diagnostic prediction label ultimately derived from the inference chain as The true pathological label corresponding to the case is ;

[0036] Define indicator functions Used to filter effective inference paths and construct a self-evolving training dataset. :

[0037] (3)

[0038] In the formula: This represents the high-quality dataset selected for subsequent self-evolutionary training. C represents the original CT image; C represents the inference chain generated autonomously by the model. This indicates the actual pathological diagnosis label of the case; Represented as the model through the inference chain The predicted diagnostic conclusions drawn; Represented as an indicator function, when the condition within the parentheses... If the prediction result is consistent with the gold standard, the value is 1 and the sample is retained; otherwise, the value is 0 and it is discarded as noise data.

[0039] This technical solution effectively removes noisy data and retains high-quality samples consistent with the real situation, making the training data more representative and reliable. High-quality data helps the model learn more accurate disease characteristics and diagnostic logic, improving the accuracy and stability of diagnostic prediction. The constructed self-evolving training set can be continuously updated with model inference, allowing the model to continuously adapt to new data and enhance its generalization ability and adaptability.

[0040] According to a preferred embodiment of the present invention, the iterative self-evolutionary training of the model based on the selected data in step S4 includes the following specific steps:

[0041] S41. Selected high-quality datasets Supervised fine-tuning techniques are used to optimize and update the large-scale vision-language model:

[0042] Define the self-evolutionary loss function To maximize the log-likelihood probability of generating the correct inference chain C, the formula is as follows:

[0043] (4)

[0044] In the formula: This is expressed as the objective loss function for self-evolutionary training; Represented as the parameters of the large visual-language model to be updated; This represents the high-quality training dataset constructed in S3; Represented as image-inference chain sample pairs in the dataset; This is expressed as the total number of steps in the reasoning chain C; Represented as the first inference chain C A fragment of textual evidence; Represented as generation Previous historical reasoning path; Represented as the corresponding visual evidence image; Represented as the logarithmic probability of generating the target text sequence;

[0045] S42. Update the parameters using the gradient descent algorithm. After completing one round of training, use the updated model as the new base model for the next round of iteration.

[0046] This technical solution uses iterative training to allow the model to continuously learn correct reasoning patterns, generating reasoning chains that better reflect reality and improving the accuracy of reasoning tasks such as diagnosis. Each iteration uses an updated model as a foundation, enabling the model to better adapt to different data and enhance its ability to handle various complex situations. Training based on selected high-quality data fully leverages the value of the data, avoids noise interference, and improves training efficiency and model performance.

[0047] According to a preferred embodiment of the present invention, the deployment and inference output of the final diagnostic model after multiple iterations in S5 includes the following specific steps:

[0048] S51, will pass through The final model, evolved through multiple iterations, is deployed in clinical settings; during the inference phase, the model is input with the CT images to be diagnosed. Automatically plan and generate optimal inference paths and diagnostic results :

[0049] (5)

[0050] In the formula: This represents the final diagnostic prediction result; This is represented as the optimal visual reasoning path; This refers to CT images of newly diagnosed patients imported into the clinical database. This represents the final large model parameters after multiple rounds of iterative self-evolutionary training. This represents a fine-tuned dedicated visual tool; argmax represents the combination of diagnostic conclusions and reasoning processes that maximizes the joint posterior probability.

[0051] S52, The reasoning output includes key visual evidence images. and corresponding text parsing A complete diagnostic report enables white-box diagnostics with full-process visualization.

[0052] This technical solution undergoes multiple iterations to make the model more accurate, generating reasonable reasoning paths and accurate diagnostic results, assisting doctors in reducing misdiagnosis; it outputs a complete report containing key visual evidence and text analysis, achieving white-box diagnosis and enabling doctors to understand the diagnostic basis; the model automatically plans reasoning paths and outputs results, saving doctors time and speeding up the diagnostic process.

[0053] According to a preferred embodiment of the present invention, the prediction method further includes the following instructions or judgments:

[0054] Operation planning instructions: Invoke FastSAM to segment the overall outline of the pancreas;

[0055] Active focusing command: Magnify the uncinate process region of the pancreatic head and perform density analysis;

[0056] Instructions for verifying indirect signs: Measure the diameter of the main pancreatic duct and common bile duct;

[0057] Logical synthesis judgment: Determine whether there is a low-density poorly vascularized mass, double tube sign, painless jaundice, and significantly elevated CA19-9;

[0058] Conclusion interpretation: Determine whether there is a strong indication of pancreatic head cancer and whether endoscopic ultrasound biopsy should be performed.

[0059] This technical solution undergoes multiple iterations to make the model more accurate, generating reasonable reasoning paths and accurate diagnostic results, assisting doctors in reducing misdiagnosis; it outputs a complete report containing key visual evidence and text analysis, achieving white-box diagnosis and enabling doctors to understand the diagnostic basis; the model automatically plans reasoning paths and outputs results, saving doctors time and speeding up the diagnostic process.

[0060] According to a preferred embodiment of the present invention, the prediction method further includes input data preprocessing, comprising the following specific steps:

[0061] The input data consisted of enhanced CTDICOM sequences containing pancreatic phase images showing low-density lesions in the uncinate process of the pancreatic head and dilation of the distal main pancreatic duct, as well as the patient's clinical information, including age, sex, chief complaint, and laboratory test results.

[0062] This technical solution integrates imaging and clinical information to analyze the condition from multiple dimensions, reducing misjudgments caused by single pieces of information and improving the accuracy of pancreatic head cancer prediction. A rich variety of input data exposes the model to more case characteristics, improving its adaptability to different situations and enhancing its generalization ability. It also provides doctors with comprehensive data references, helping them to better understand the condition and assisting in developing more reasonable diagnostic and treatment plans.

[0063] According to a preferred embodiment of the present invention, the prediction method further includes model performance evaluation, comprising the following specific steps:

[0064] The performance of the model was evaluated by the receiver operating characteristic ROC curve and AUC value. The AUC value of the improved inference model was significantly better than that of the traditional non-inference model and the unimproved inference model.

[0065] In this technical solution, the AUC value is significantly improved, indicating that the improved model has a strong ability to distinguish pancreatic head cancer and can more reliably determine whether a patient has the disease. By comparing the AUC values ​​of different models, the model with better performance can be selected intuitively, providing a better tool for clinical diagnosis.

[0066] In another aspect of the invention, the invention also provides a pancreatic cancer CT image prediction system based on iterative self-evolution.

[0067] A pancreatic cancer CT image prediction system based on iterative self-evolution includes the following modules:

[0068] Segmentation Dataset Module: Collects CT image data of pancreatic cancer, invites medical experts to annotate pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts to form an image segmentation dataset;

[0069] Visual tool fine-tuning module: Select the pre-trained FastSAM model, use the pancreatic cancer disease segmentation dataset, construct a Dice and cross-entropy hybrid loss function for parameter fine-tuning, and obtain a dedicated pancreatic anatomy extraction tool;

[0070] Interactive graphic reasoning module: Combining a pre-trained visual-language large model with fine-tuned visual tools, it generates visual operation instructions based on historical reasoning paths, executes instructions to generate visual evidence images and text descriptions, forming a complete reasoning chain;

[0071] Inference data filtering module: Defines diagnostic prediction labels and true pathology labels, filters effective inference paths through result consistency verification, constructs a self-evolutionary training dataset, retains samples with consistent predictions, and removes noisy data;

[0072] Model Iterative Self-Evolution Training Module: Utilizing high-quality datasets, supervised fine-tuning techniques are employed to optimize the large-scale vision-language model. A self-evolutionary loss function is defined to maximize the log-likelihood probability of generating correct inference chains. Gradient descent is used to update parameters, and iterative training is conducted to improve model performance.

[0073] Diagnostic report output module: Deploy the final model after multiple iterations to the clinical end, input the CT image to be diagnosed, automatically generate the optimal reasoning path and diagnostic results, and output a complete white-box diagnostic report including key visual evidence images and text parsing.

[0074] In this technical solution, multi-module collaboration and self-evolutionary training enable the model to better adapt to complex cases and improve the accuracy of pancreatic cancer diagnosis; output a complete white-box report containing key evidence and analysis, allowing doctors to understand the diagnostic basis and increasing diagnostic credibility; and continuously improve model performance and maintain the system's advanced nature through screening high-quality data and iterative training.

[0075] The one or more technical solutions provided by this invention have the following advantages compared with the prior art:

[0076] (1) By introducing FastSAM visual tools for fine-tuning disease-specific data and interactive graphic reasoning generation architecture, the model can actively perform "planning-segmentation-focusing" operations, transforming blurry images into local visual evidence. The graphic and text interleaved reasoning method provides pixel-level image evidence, effectively avoiding visual illusion problems and significantly improving the detection rate and clinical interpretability of hidden lesions.

[0077] (2) Construct a self-training mechanism based on result consistency verification, automatically filter high-quality samples, form a closed-loop training framework of "generation-filtering-evolution", enable the model to learn the optimal reasoning logic autonomously, reduce the dependence on manual annotation resources, and achieve continuous self-evolution with low manual cost.

[0078] (3) After the model is deployed, it can automatically generate a complete diagnostic report containing key visual evidence and text analysis, realize full-process visualized white-box diagnosis, help doctors understand the diagnostic basis, speed up the diagnostic process, and provide clinical intelligent auxiliary diagnosis and treatment solutions with high precision, strong interpretability and continuous learning ability. Attached Figure Description

[0079] Figure 1 This is a flowchart of the method of the present invention.

[0080] Figure 2 This is a comparison chart of the model results of the present invention.

[0081] Figure 3 This is the overall flowchart of the present invention.

[0082] Figure 4 This is the CT image and segmentation mask of the present invention. Detailed Implementation

[0083] Example 1

[0084] Please refer to Figure 1 and Figure 3 This embodiment provides a method for predicting pancreatic cancer CT images based on iterative self-evolution, including the following steps:

[0085] S1. Constructing a segmentation dataset specifically for pancreatic cancer and fine-tuning visual tools:

[0086] First, a CT image segmentation dataset containing the pancreas and its surrounding key anatomical structures (such as the duodenum and splenic vein) is constructed. Let the input CT image be... The corresponding expert-annotated mask is (Including categories such as pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts). A pre-trained FastSAM (Fast Segment Anything Mode) model was selected as the basic vision tool model, and its parameters are denoted as... To adapt the general segmentation model to the segmentation task of occult pancreatic cancer lesions, a hybrid loss function incorporating Dice loss and cross-entropy loss is constructed. Fine-tuning of the parameters of FastSAM is performed. The objective function for fine-tuning is expressed as:

[0087] (1)

[0088] In the formula: This represents the optimal visual tool model parameters obtained after fine-tuning. Let these be the model parameter variables to be optimized. This represents the operation of finding parameters that minimize the loss function; This is expressed as the expected value calculation for all sample pairs within the dataset; This is represented as the prediction output function of the visual segmentation model; It is represented as a hybrid segmentation loss function used to measure the difference between the predicted mask and the real mask.

[0089] S2. Generate interactive text-image reasoning using a pre-trained large-scale visual-language model:

[0090] Using a pre-trained large-scale visual-language model (Qwen3-VL-8B) as the inference engine, its parameters are denoted as follows. For pancreatic cancer CT images without labeled thought chains... The model follows a cyclical mechanism of "planning-execution-observation-description" to generate multi-step inference. In the first... In step-by-step reasoning, the model first bases its reasoning on contextual history. Generate visual operation instructions (e.g., "segmenting the pancreatic head region"), then invoke the finely tuned visual tools. Execute this command to generate a visual evidence image. Subsequently, the model was combined with Generate textual evidence for the current step The entire chain of reasoning The generation process can be formalized as maximizing the joint probability:

[0091] (2)

[0092] In the formula: This refers to generating a complete inference chain given image input and model parameters. The joint probability; This represents an unlabeled CT image that was entered. These are represented as parameters of a large vision-language model. This is represented as the visual tool obtained through fine-tuning in S1; This represents the total number of steps in the reasoning process; Represented as in the first The visual operation instructions generated step by step; Represented as the first The historical reasoning path before the step includes ; Indicated as according to instructions and images The generated visual evidence image; Represented as a text description generated based on visual evidence; It is represented as a conditional generation probability function.

[0093] S3. Data filtering and training set construction based on result consistency:

[0094] To obtain high-quality inference data without human intervention, a result consistency verification mechanism is introduced. Let the inference chain generated in step S2 be... The final exported diagnostic prediction label is The ground truth label for this case is... Define indicator functions Used to filter effective inference paths and construct a self-evolving training dataset.

[0095] (3)

[0096] In the formula: This represents the high-quality dataset selected for subsequent self-evolutionary training. C represents the original CT image; C represents the inference chain generated autonomously by the model. This represents the actual pathological diagnosis label of the case; Represented as the model through the inference chain The predicted diagnostic conclusions drawn; Represented as an indicator function, when the condition within the parentheses... If the prediction result is consistent with the gold standard, the value is 1 and the sample is retained; otherwise, the value is 0 and it is discarded as noise data.

[0097] S4. Iterative self-evolutionary training of the model based on filtered data:

[0098] High-quality datasets selected in step S3 Supervised fine-tuning (SFT) is employed to optimize and update the large-scale vision-language model. A self-evolutionary loss function is defined. To maximize the generation of correct inference chains The log-likelihood probability is calculated using the following formula:

[0099] (4)

[0100] In the formula: This is expressed as the objective loss function for self-evolutionary training; Represented as the parameters of the large visual-language model to be updated; This represents the high-quality training dataset constructed in S3; Represented as image-inference chain sample pairs in the dataset; This is expressed as the total number of steps in the reasoning chain C; Represented as the first inference chain C A fragment of textual evidence; Represented as generation Previous historical reasoning path; Represented as the corresponding visual evidence image; Represented as the logarithmic probability of generating the target text sequence;

[0101] S5. Deployment and inference output of the final diagnostic model after multiple iterations:

[0102] Will pass The final model, evolved through multiple iterations, is deployed clinically. During the inference phase, the model is input with the CT images to be diagnosed. Automatically plan and generate optimal inference paths and diagnostic results

[0103] (5)

[0104] In the formula: This represents the final diagnostic prediction result; This is represented as the optimal visual reasoning path; This refers to CT images of newly diagnosed patients imported into the clinical database. This represents the final large model parameters after multiple rounds of iterative self-evolutionary training. This represents a fine-tuned dedicated visual tool; argmax represents the combination of diagnostic conclusions and reasoning processes that maximizes the joint posterior probability.

[0105] Example 2

[0106] Please refer to Figure 1 and Figure 3 Based on Example 1, this example provides a pancreatic cancer CT image prediction system based on iterative self-evolution, including the following modules:

[0107] Segmentation Dataset Module: Collects CT image data of pancreatic cancer, invites medical experts to annotate pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts to form an image segmentation dataset;

[0108] Visual tool fine-tuning module: Select the pre-trained FastSAM model, use the pancreatic cancer disease segmentation dataset, construct a Dice and cross-entropy hybrid loss function for parameter fine-tuning, and obtain a dedicated pancreatic anatomy extraction tool;

[0109] Interactive graphic reasoning module: Combining a pre-trained visual-language large model with fine-tuned visual tools, it generates visual operation instructions based on historical reasoning paths, executes instructions to generate visual evidence images and text descriptions, forming a complete reasoning chain;

[0110] Inference data filtering module: Defines diagnostic prediction labels and true pathology labels, filters effective inference paths through result consistency verification, constructs a self-evolutionary training dataset, retains samples with consistent predictions, and removes noisy data;

[0111] Model Iterative Self-Evolution Training Module: Utilizing high-quality datasets, supervised fine-tuning techniques are employed to optimize the large-scale vision-language model. A self-evolutionary loss function is defined to maximize the log-likelihood probability of generating correct inference chains. Gradient descent is used to update parameters, and iterative training is conducted to improve model performance.

[0112] Diagnostic report output module: Deploy the final model after multiple iterations to the clinical end, input the CT image to be diagnosed, automatically generate the optimal reasoning path and diagnostic results, and output a complete white-box diagnostic report including key visual evidence images and text parsing.

[0113] In this technical solution, multi-module collaboration and self-evolutionary training enable the model to better adapt to complex cases and improve the accuracy of pancreatic cancer diagnosis; output a complete white-box report containing key evidence and analysis, allowing doctors to understand the diagnostic basis and increasing diagnostic credibility; and continuously improve model performance and maintain the system's advanced nature through screening high-quality data and iterative training.

[0114] Example 3

[0115] Based on Examples 1 and 2, the effects of the present invention are analyzed in conjunction with actual cases.

[0116] Assume the operation is performed according to the following instructions:

[0117] [Inference Input]

[0118] System prompt: You are a senior pancreatic imaging expert. Please use visual tools to carefully analyze the following patient's abdominal CT scan images. You need to actively locate the pancreatic region, check for any abnormal lesions or indirect signs (such as pancreatic duct dilatation), and provide a final diagnosis based on clinical information.

[0119] Input: [CT Imaging]: Enhanced CT DICOM sequence including the pancreatic phase. (Note: The imaging shows a hypodepth lesion of approximately 1.5 cm in the uncinate process of the pancreatic head, accompanied by complete dilation of the distal main pancreatic duct, presenting as a "double duct sign".) [Clinical Information]: Male, 68 years old. Chief complaint: Dull pain and discomfort in the upper abdomen for 3 months, with recent onset of jaundice of the skin and sclera. Laboratory tests: Elevated CA19-9 (1200 U / mL), elevated total bilirubin.

[0120] Based on the above information, please independently plan a visual observation path, predict whether the patient has a risk of pancreatic cancer, and output a detailed reasoning process and final conclusion.

[0121] [True Value of Data]

[0122] There is a risk of pancreatic cancer (pancreatic head cancer).

[0123] [Unimproved model inference results]

[0124] No risk of pancreatic cancer (or suspected chronic pancreatitis)

[0125] 1. Visual observation: No obvious mass or space-occupying lesion was found across the entire image. The pancreas appears slightly full, but its outline is still clear.

[0126] 2. Feature Analysis: Mild dilation of the pancreatic duct was noted, which may be related to ductal stenosis caused by chronic inflammation. Combined with the patient's upper abdominal pain symptoms and the absence of obvious low-density lesions, a diagnosis of chronic pancreatitis is more likely.

[0127] 3. Conclusion: Imaging findings were atypical; follow-up observation is recommended. (Analysis of shortcomings: The model adopted a passive global observation mode, failing to actively magnify the pancreatic head region, thus missing a poorly vascularized lesion of only 1.5 cm and incorrectly interpreting the "double-tube sign" as inflammatory changes, leading to a missed diagnosis.)

[0128] [Based on the final model inference results]

[0129] Risk of pancreatic cancer

[0130] 1. Operation Plan: [Instruction: Call FastSAM to segment the overall outline of the pancreas]. Atrophy was observed in the pancreatic body and tail, but the pancreatic head had an irregular shape.

[0131] 2. Active Focusing: [Command: Zoom in on the uncinate process region of the pancreatic head and perform density analysis]. In the locally enhanced image, a low-density area with a diameter of approximately 1.5 cm was found. After contrast enhancement, the enhancement level was lower than that of the surrounding normal pancreatic parenchyma (characteristic of poor blood supply), suggesting a high probability of malignancy.

[0132] 3. Indirect sign verification: [Instruction: Measure the diameter of the main pancreatic duct and common bile duct]. The results showed that the diameter of the main pancreatic duct was dilated to 8 mm, and the common bile duct was dilated to 12 mm, forming a typical "double duct sign". This is usually caused by compression from a mass in the pancreatic head.

[0133] 4. Logical synthesis: Combining the imaging findings of "low-density poorly vascularized mass" and "double duct sign" with the clinical findings of "painless jaundice" and "significantly elevated CA19-9", it is consistent with the typical manifestations of pancreatic head ductal adenocarcinoma.

[0134] 5. Conclusion: Highly suggestive of pancreatic head cancer; endoscopic ultrasound-guided fine-needle aspiration biopsy (EUS-FNA) is recommended immediately.

[0135] [Non-reasoning model]

[0136] Output only: Risk of pancreatic cancer

[0137] Disadvantages: 1. "Black box" prediction, lacking localization ability. Non-inference models (such as ResNet-based binary classifiers) can only provide a risk probability and cannot pinpoint the specific location of the lesion on the CT image. After seeing a high-risk warning, doctors still need to spend a lot of time manually searching for that tiny lesion among hundreds of slides, which does not significantly reduce the burden on doctors in interpreting images.

[0138] 2. Highly dependent on large-scale pixel-level annotation. Such models typically require training with thousands of CT scans manually drawn layer by layer by doctors, resulting in extremely high data preparation costs. In contrast, the self-evolving model of this invention can learn on its own using a large amount of data with only diagnostic labels, greatly reducing the data threshold.

[0139] Therefore, this invention introduces specialized visual tools for consistency verification of results. For example, in the [based on the final model inference result], firstly, through [instruction: call FastSAM to segment the overall outline of the pancreas], it observes that the pancreatic body and tail are atrophied, but the pancreatic head has an irregular shape; then [instruction: magnify the uncinate process region of the pancreatic head and perform density analysis], key features such as low-density areas are discovered; next, [instruction: measure the diameter of the main pancreatic duct and common bile duct] verifies indirect signs; finally, logical synthesis is performed to arrive at a conclusion that strongly suggests pancreatic head cancer. Through these operations, the inference ability of the large model is successfully anchored to real pixel features, achieving a performance reversal and breakthrough. In contrast, the [non-inference model] only outputs the risk of pancreatic cancer, which has obvious shortcomings. On the one hand, it is a "black box" prediction, lacking localization ability. After seeing the high-risk warning, doctors still need to spend a lot of time manually searching for tiny lesions in hundreds of slides, which does not significantly reduce the burden of reading slides; on the other hand, it is highly dependent on large-scale pixel-level annotation, usually requiring thousands of CT data that have been manually drawn layer by layer by doctors for training, resulting in extremely high data preparation costs. The self-evolutionary model of this invention can learn itself using a large amount of data with only diagnostic labels, which greatly reduces the data threshold and has significant advantages in the field of pancreatic cancer diagnosis.

[0140] Combined with appendix Figure 2 This paper presents a performance comparison of different models in the task of pancreatic cancer CT diagnosis, specifically demonstrating the receiver operating characteristic (ROC) curve of the inference model based on visual thought chain and self-evolution mechanism proposed in this invention on the task of pancreatic cancer CT image prediction. Accurate diagnosis of pancreatic cancer is crucial in clinical diagnostic scenarios, and the diagnostic capabilities of different models vary significantly. This curve visually illustrates the advantages and disadvantages of each model.

[0141] The orange curve represents the "improved inference model" that utilizes the core improvement strategy of this invention. Through fine-tuning visual tools and self-evolutionary training, its AUC value reaches as high as 0.968. This superior performance significantly outperforms the traditional "non-inference model" (blue curve, AUC=0.934), fully demonstrating that this method, while offering an interpretable implementation process, also surpasses mainstream deep learning baselines in diagnostic accuracy. In practical applications of pancreatic cancer diagnosis, higher diagnostic accuracy means more accurate identification of the patient's condition, providing a reliable basis for subsequent treatment.

[0142] Crucially, the "unimproved inference model" (green curve, such as a lesion classifier based on a dedicated CNN or ViT architecture), which lacks dedicated visual tools and self-evolutionary screening mechanisms, performed the worst, with an AUC of only 0.829, even lower than the non-inference model. Considering the specific diagnostic case mentioned earlier, in the [Inference Input], the system prompted a senior pancreatic imaging expert to carefully analyze the patient's abdominal CT scan images using visual tools. These images included enhanced CT DICOM sequences of the pancreatic phase and contained clear clinical information. However, the [Unimproved Model Inference Result] incorrectly concluded there was no risk of pancreatic cancer (or suspected chronic pancreatitis). Its visual observation adopted a passive, global observation mode, failing to actively magnify the pancreatic head region, thus missing a poorly vascularized lesion of only 1.5 cm and mistakenly interpreting the "double tube sign" as inflammatory changes, leading to a missed diagnosis. This significant difference strongly confirms the technical contribution of this invention: that general-purpose visual-language models are highly susceptible to hallucinations and performance degradation when lacking refined visual guidance.

[0143] Example 4 Based on Examples 1 and 2, the effectiveness of FastSAM as a visual tool for lesion segmentation is analyzed using actual cases.

[0144] Combined: 4079 cases from a disease-specific dataset, such as Figure 4 As shown in the image, FastSAM segmentation results on a disease-specific dataset demonstrate that in clinical imaging, pancreatic tumors on plain CT scans, lacking contrast enhancement, exhibit almost isodense density compared to the surrounding normal pancreatic parenchyma, making their precise boundaries extremely difficult to discern with the naked eye and traditional global visual models. The system of this invention demonstrates the following effectiveness in handling this challenging case. Based on the actual segmentation results, the finely tuned FastSAM exhibits the following three major breakthroughs.

[0145] 1. Breaking through the bottleneck of extremely low contrast in plain CT scans, achieving "creation from nothing" level extraction of occult lesions, from... Figure 4 The original plain CT scan on the left side shows that the gray-level difference between the lesion area and the surrounding normal tissue is negligible. If an untuned open-source FastSAM or a traditional global segmentation network is used, it is highly likely to result in missed segmentation or large-area adhesion. However, the FastSAM in this invention, fine-tuned using a "Dice + cross-entropy hybrid loss function" and tens of thousands of disease-specific slices, deeply internalizes the subtle texture and mass effect characteristics specific to pancreatic cancer. For example... Figure 4 As shown in the segmentation mask image on the right, FastSAM successfully and accurately captured the tumor entity against an isodense background, generating a high-confidence lesion mask.

[0146] 2. The "large model command guidance + local focusing" method completely eliminates interference from complex anatomical backgrounds. The pancreas is surrounded by complex organs such as the duodenum and superior mesenteric artery and vein, which are easily confused with pancreatic lesions on plain CT scans. From Figure 4 As can be seen from the segmentation mask on the right, FastSAM's segmentation boundaries are extremely clean, without any "bleeding" or missegmentation into the surrounding gastrointestinal tract or blood vessels. This fully demonstrates the core advantage of the "large model interactive active guidance" in step S2 of this invention: the large model first defines the region of interest (ROI) through spatial instructions (e.g., "focusing on the pancreatic head region"), and then FastSAM performs bottom-up pixel-level fitting only within this local area. This collaborative mechanism of "first defining the general direction, then refining the details" exponentially increases the segmentation tool's anti-interference capability.

[0147] 3. Accurately depicting the "edge infiltration" morphology provides irreplaceable physical evidence for large-scale model reasoning; careful observation is crucial. Figure 4 The outline of the segmentation mask on the right reveals that FastSAM not only located the lesion but also accurately reconstructed the typical geometric features of pancreatic cancer: "irregular shape and spiculated infiltrative edges." This is precisely the ultimate purpose of using FastSAM as a "visual tool" in this invention: it transforms the originally blurry global image into local visual evidence with clear boundaries and well-defined shapes. Large language models (such as Qwen3-VL) are able to generate the correct thought chain (CoT) in subsequent steps based on this precise lesion morphology provided by FastSAM, which "detects irregular low-density lesions with unclear boundaries, highly suggestive of malignancy," thus breaking the model illusion from a physical mechanism perspective.

[0148] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, and are not intended to limit the specific implementation of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A method for predicting pancreatic cancer using CT images based on iterative self-evolution, characterized in that, Includes the following steps: S1. Constructing a segmentation dataset and fine-tuning visual tools specifically for pancreatic cancer; S2. Generate interactive text and image reasoning using a pre-trained visual-language large model; S3. Data filtering and training set construction based on result consistency; S4. Iterative self-evolutionary training of the model based on filtered data; S5. Deployment and inference output of the final diagnostic model after multiple iterations.

2. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 1, characterized in that: The construction of the pancreatic cancer-specific segmentation dataset and the fine-tuning of visual tools in S1 include the following specific steps: S11. Construct a segmentation dataset specifically for pancreatic cancer: Input CT image data and corresponding expert-annotated masks , The categories included are pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts, forming an image segmentation dataset; S12, Building a fine-tuned visual tool for pancreatic cancer: We selected the pre-trained FastSAM as the basic vision tool model and constructed a hybrid loss function that includes Dice loss and cross-entropy loss. By fine-tuning the parameters of FastSAM, a fine-tuned visual tool specifically for extracting pancreatic anatomical structures is obtained, where the objective function for parameter fine-tuning is expressed as: (1) In the formula: This represents the optimal visual tool model parameters obtained after fine-tuning. Let these be the model parameter variables to be optimized. This represents the operation of finding parameters that minimize the loss function; This is expressed as the expected value calculation for all sample pairs within the dataset; This is represented as the prediction output function of the visual segmentation model; It is represented as a hybrid segmentation loss function used to measure the difference between the predicted mask and the real mask.

3. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 2, characterized in that: The S2 method, which utilizes a pre-trained large-scale visual-language model to generate interactive text-image reasoning, includes the following specific steps: S21. Utilize a pre-trained large-scale visual-language model For CT image data without labeled inference chains, multi-step inference is generated: the model follows a cyclical mechanism of "planning-execution-observation-description," based on historical inference paths. Generate visual operation instructions Then, the finely tuned visual tools were invoked. Execute the instructions to generate visual evidence images. ; S23, Combination Generate textual evidence for the current step The entire chain of reasoning The generation process can be represented as maximizing the joint probability: (2) In the formula: This refers to generating a complete inference chain given image input and model parameters. The joint probability; This represents an unlabeled CT image that was entered. These are represented as parameters of a large vision-language model. This is represented as the visual tool obtained through fine-tuning in S1; This represents the total number of steps in the reasoning process; Represented as in the first The visual operation instructions generated step by step; Represented as the first The historical reasoning path before the step includes ; Indicated as according to instructions and images The generated visual evidence image; Represented as a text description generated based on visual evidence; It is represented as a conditional generation probability function.

4. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 3, characterized in that: The inference data filtering and training set construction based on result consistency in S3 includes the following specific steps: Define the diagnostic prediction label ultimately derived from the inference chain as The true pathological label corresponding to the case is ; Define indicator functions Used to filter effective inference paths and construct a self-evolving training dataset. : (3) In the formula: This represents the high-quality dataset selected for subsequent self-evolutionary training. C represents the original CT image; C represents the inference chain generated autonomously by the model. This indicates the actual pathological diagnosis label of the case; Represented as the model through the inference chain The predicted diagnostic conclusions drawn; Represented as an indicator function, when the condition within the parentheses... If the prediction result is consistent with the gold standard, the value is 1 and the sample is retained; otherwise, the value is 0 and it is discarded as noise data.

5. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 4, characterized in that: The iterative self-evolutionary training of the model based on the selected data in S4 includes the following specific steps: S41. Selected high-quality datasets Supervised fine-tuning techniques are used to optimize and update the large-scale vision-language model: Define the self-evolutionary loss function To maximize the log-likelihood probability of generating the correct inference chain C, the formula is as follows: (4) In the formula: This is expressed as the objective loss function for self-evolutionary training; Represented as the parameters of the large visual-language model to be updated; This represents the high-quality training dataset constructed in S3; Represented as image-inference chain sample pairs in the dataset; This is expressed as the total number of steps in the reasoning chain C; Represented as the first inference chain C A fragment of textual evidence; Represented as generation Previous historical reasoning path; Represented as the corresponding visual evidence image; Represented as the logarithmic probability of generating the target text sequence; S42. Update the parameters using the gradient descent algorithm. After completing one round of training, use the updated model as the new base model for the next round of iteration.

6. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 5, characterized in that: The deployment and inference output of the final diagnostic model after multiple iterations of S5 includes the following specific steps: S51, will pass through The final model, evolved through multiple iterations, is deployed in clinical settings; during the inference phase, the model is input with the CT images to be diagnosed. Automatically plan and generate optimal inference paths and diagnostic results : (5) In the formula: This represents the final diagnostic prediction result; This is represented as the optimal visual reasoning path; This refers to CT images of newly diagnosed patients imported into the clinical database. This represents the final large model parameters after multiple rounds of iterative self-evolutionary training. This represents a fine-tuned dedicated visual tool; argmax represents the combination of diagnostic conclusions and reasoning processes that maximizes the joint posterior probability. S52, The reasoning output includes key visual evidence images. and corresponding text parsing A complete diagnostic report enables white-box diagnostics with full-process visualization.

7. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 1, characterized in that: The prediction method also includes the following instructions or judgments: Operation planning instructions: Invoke FastSAM to segment the overall outline of the pancreas; Active focusing command: Magnify the uncinate process region of the pancreatic head and perform density analysis; Instructions for verifying indirect signs: Measure the diameter of the main pancreatic duct and common bile duct; Logical synthesis judgment: Determine whether there is a low-density poorly vascularized mass, double tube sign, painless jaundice, and significantly elevated CA19-9; Conclusion interpretation: Determine whether there is a strong indication of pancreatic head cancer and whether endoscopic ultrasound biopsy should be performed.

8. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 1, characterized in that: The prediction method also includes input data preprocessing, comprising the following specific steps: The input data consisted of enhanced CTDICOM sequences containing pancreatic phase images showing low-density lesions in the uncinate process of the pancreatic head and dilation of the distal main pancreatic duct, as well as the patient's clinical information, including age, sex, chief complaint, and laboratory test results.

9. The pancreatic cancer CT image prediction method based on iterative self-evolution as described in claim 1, characterized in that: The prediction method also includes model performance evaluation, comprising the following specific steps: The performance of the model was evaluated by the receiver operating characteristic ROC curve and AUC value. The AUC value of the improved inference model was significantly better than that of the traditional non-inference model and the unimproved inference model.

10. A pancreatic cancer CT image prediction system based on iterative self-evolution, employing the pancreatic cancer CT image prediction method based on iterative self-evolution as described in any one of claims 1-9, characterized in that: Includes the following modules: Segmentation Dataset Module: Collects CT image data of pancreatic cancer, invites medical experts to annotate pancreatic parenchyma, tumor lesions, and dilated pancreatic ducts to form an image segmentation dataset; Visual tool fine-tuning module: Select the pre-trained FastSAM model, use the pancreatic cancer disease segmentation dataset, construct a Dice and cross-entropy hybrid loss function for parameter fine-tuning, and obtain a dedicated pancreatic anatomy extraction tool; Interactive graphic reasoning module: Combining a pre-trained visual-language large model with fine-tuned visual tools, it generates visual operation instructions based on historical reasoning paths, executes instructions to generate visual evidence images and text descriptions, forming a complete reasoning chain; Inference data filtering module: Defines diagnostic prediction labels and true pathology labels, filters effective inference paths through result consistency verification, constructs a self-evolutionary training dataset, retains samples with consistent predictions, and removes noisy data; Model Iterative Self-Evolution Training Module: Utilizing high-quality datasets, supervised fine-tuning techniques are employed to optimize the large-scale vision-language model. A self-evolutionary loss function is defined to maximize the log-likelihood probability of generating correct inference chains. Gradient descent is used to update parameters, and iterative training is conducted to improve model performance. Diagnostic report output module: Deploy the final model after multiple iterations to the clinical end, input the CT image to be diagnosed, automatically generate the optimal reasoning path and diagnostic results, and output a complete white-box diagnostic report including key visual evidence images and text parsing.

Citation Information

Patent Citations

  • Prostate cancer diagnosis method based on multimodal large model prompt learning mechanism

    CN120954689A

  • Cancer risk level distinguishing method and system

    CN121101470A