NK / T cell lymphoma intelligent analysis method based on large model agent

CN122617880APending Publication Date: 2026-08-21BEIJING TONGREN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611095079.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]本发明提供一种基于大模型智能体的NK/T细胞淋巴瘤智能分析方法,解决ENKTCL诊断过程高度依赖临床专家经验,存在耗时长,差异化大,难以实现标准化和规模化推广的问题,能对ENKTCL的早期精准诊断提供及时、准确的智能辅助,对于改善患者生存率和制定个体化治疗方案具有重要意义,实现ENKTCL的多模态联合推理与辅助决策

Benefits of technology

[0038]This invention provides an intelligent analysis method for NK/T-cell lymphoma based on a large-scale intelligent model. It unifies and collaboratively fuses multimodal data, including MRI images, image reports, clinical structured variables, and a medical knowledge base. It intelligently segments the patient's MRI images using a visual segmentation model and performs knowledge retrieval and context construction using the medical knowledge base. Finally, it achieves explicit thought chain reasoning based on a large language model. This addresses the problems of the NK/T-cell lymphoma (ENTCL) diagnosis process being highly dependent on clinical expert experience, resulting in long processing times, significant variability, and difficulties in standardization and large-scale implementation. It provides timely and accurate intelligent assistance for the early and precise diagnosis of ENKTCL, which is of great significance for improving patient survival rates and developing individualized treatment plans, enabling multimodal joint reasoning and decision support for ENKTCL.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122617880A_ABST
    Figure CN122617880A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical image analysis, and specifically discloses an NK / T cell lymphoma intelligent analysis method based on a large model agent, which comprises the following steps: uniformly modeling and cooperatively fusing different modal data to form a multi-modal patient information model; constructing a visual segmentation model based on a DINOv3 backbone network to intelligently segment MRI medical images of a patient; constructing an ENKTCL special medical knowledge base to perform knowledge retrieval and context construction; using a large language model to establish a chain reasoning mechanism to realize multi-modal joint reasoning and auxiliary decision of ENKTCL, construct an iterative correction mechanism based on prognosis feedback, and form an internal iterative correction mechanism; calculating uncertainty and risk values at key stages of image segmentation, forward reasoning and prognosis feedback, and determining whether to trigger doctor intervention review and correction to introduce an adaptive external doctor supervision mechanism. The application can provide accurate auxiliary decision basis for treatment of ENKTCL.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of medical image analysis, and more specifically, to an intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent. Background Technology

[0002] Extranodal NK / T-cell lymphoma (ENKTCL) is a highly aggressive malignant lymphoma originating from natural killer cells (NK cells) or cytotoxic T cells. Due to its high malignancy and rapid progression, early and accurate diagnosis of ENKTCL is of paramount importance in clinical practice. Currently, the clinical diagnosis of ENKTCL relies primarily on a comprehensive assessment of multiple sources of information, including medical imaging examinations, pathological histological analysis, immunohistochemical detection, laboratory hematological markers, and clinical symptoms. Imaging examinations provide information on lesion location, extent, and invasion; pathological sections and immunophenotypic analysis confirm the nature of the tumor; and hematological and clinical markers assist in disease staging and prognostic assessment. Because these information sources are complex and heterogeneous, with strong correlations and complementarities between different modalities, a multidisciplinary team of experts, including radiologists, pathologists, and hematologists, is typically required. Therefore, the diagnostic process heavily relies on the experience of clinical experts and lacks a unified and scalable intelligent assistance mechanism. This not only results in a lengthy process but also significant differences in experience among different doctors, leading to insufficient stability of diagnostic results and hindering standardization and large-scale implementation. Therefore, it is of great significance to explore how to achieve multimodal joint reasoning and decision support using ENKTCL by integrating medical imaging, pathology reports, structured clinical data, and medical knowledge bases. Summary of the Invention

[0003] This invention provides an intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent, which solves the problems of the ENKTCL diagnosis process being highly dependent on the clinical expert experience, time-consuming, highly variable, and difficult to standardize and scale up. It can provide timely and accurate intelligent assistance for the early and accurate diagnosis of ENKTCL, which is of great significance for improving patient survival rate and developing individualized treatment plans, and realizes multimodal joint reasoning and auxiliary decision-making for ENKTCL.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for intelligent analysis of NK / T cell lymphoma based on a large model intelligent agent includes:

[0006] The system acquires patients' MRI medical images, image reports, clinical structured variables, and knowledge base retrieval results to perform unified modeling and collaborative fusion of data from different modalities, forming a multimodal patient information model.

[0007] A visual segmentation model based on the DINOv3 backbone network was constructed to intelligently segment the patient's MRI medical images to identify and extract lesion regions, lesion boundaries, and lesion spatial location information, and use them as structured imaging evidence.

[0008] We constructed a dedicated medical knowledge base for ENKTCL to perform knowledge retrieval and context construction based on the patient's corresponding multimodal information and structured imaging evidence.

[0009] By using a large language model to establish a chain reasoning mechanism, the forward reasoning process of ENKTCL diagnosis is decomposed into evidence summary, clinical hypothesis generation, treatment planning, and outcome output.

[0010] Preferred options also include:

[0011] An iterative correction mechanism based on prognostic feedback is constructed to predict the possible prognostic outcomes of the current treatment recommendations and input the prediction results back into the reasoning process to dynamically correct the original reasoning path based on the prognostic feedback, thus forming an internal iterative correction mechanism.

[0012] Preferred options also include:

[0013] Uncertainty and risk values ​​are calculated for key stages of image segmentation, forward inference, and prognostic feedback, and the appropriate risk thresholds are used to determine whether to trigger physician intervention for review and correction, thereby introducing an adaptive external physician supervision mechanism.

[0014] Preferably, the step of uniformly modeling and collaboratively fusing data from different modalities to form a multimodal patient information model includes:

[0015] The patient's MRI medical images, image reports, clinical structured variables, and knowledge base search results are used as input. The patient input is defined as: ,in, This indicates information about MRI images and lesion areas. This indicates a structured image report. Represents the patient's clinical structured variables. This indicates the results of a knowledge base search.

[0016] By using a modality-specific encoder to learn structured representations of various types of data, information from different sources is transformed into corresponding modality feature representations that can be uniformly processed for subsequent reasoning;

[0017] After obtaining the features of each modality, a multimodal fusion function is used to integrate them to construct a unified semantic representation at the patient level, which simultaneously reflects the patient's lesion morphology, imaging manifestations, clinical status, and medical knowledge-related information.

[0018] Preferably, the intelligent segmentation of the patient's MRI medical images includes:

[0019] Extracting deep visual features using the DINOv3 visual backbone network: ,in, denoted as the visual feature extraction network, where f represents the visual features obtained from the input MRI image;

[0020] The obtained visual features are input into the segmentation head network, and the segmentation head network outputs the corresponding lesion region probability map, and then the lesion region is segmented according to the lesion region probability map.

[0021] Preferably, the step of inputting the obtained visual features into the segmentation head network and outputting the corresponding lesion region probability map through the segmentation head network includes:

[0022] According to the formula: This yields a lesion prediction mask, where each pixel value in the prediction mask represents the probability that the corresponding location belongs to the lesion region. Represents the segmentation head network. This represents the lesion prediction mask, where H and W represent the image height and width, respectively.

[0023] Preferably, the knowledge retrieval and context construction based on the patient's corresponding multimodal information and structural human imaging evidence includes:

[0024] The knowledge documents are preprocessed in a structured manner and then vectorized using a text embedding model, so that all knowledge vectors are stored in a vector database to support efficient similarity retrieval in the future.

[0025] After receiving multimodal patient information, a query vector is generated from the patient information, and a patient-specific query function is automatically constructed. A similarity search is performed through the vector database to obtain a knowledge set related to the patient's status.

[0026] The patient fusion representation is integrated with the retrieved knowledge set using a context fusion function.

[0027] Preferably, the method of establishing a chain reasoning mechanism using a large language model includes:

[0028] Using the formula: The system reasons from the knowledge set related to the patient's condition to generate a summary of clinical evidence based on the patient's contextual information. This represents the final context provided to the large language model. This represents the summary of clinical evidence. Represents the inference function of the large language model in the evidence summarization stage;

[0029] Intermediate clinical hypotheses are generated based on the results of summarizing clinical evidence to obtain descriptions of disease risk, lesion invasion, clinical staging tendencies, and potential treatment directions.

[0030] Using a structured output function, initial treatment recommendations are generated based on the obtained clinical hypotheses. The recommendations include treatment strategies, treatment regimens, and suggestions for the use of PD-1 inhibitors.

[0031] Preferably, the construction of the iterative correction mechanism based on prognostic feedback includes:

[0032] A prognostic prediction head is set up using a large language model to predict patient prognostic outcomes based on the current treatment plan and patient information;

[0033] The prognostic prediction results are re-inputted into the inference module to update the current inference state. New treatment suggestions are generated based on the updated inference state. After multiple rounds of iteration, a stable and converged final recommendation result is obtained.

[0034] Preferably, the calculation of uncertainty and risk values ​​for the key stages of image segmentation, forward inference, and prognostic feedback includes:

[0035] In the image segmentation stage, uncertainty indicators are used... Assess the reliability of lesion boundaries, among which, Let represent the predicted probability at pixel i, and N represent the total number of pixels;

[0036] In the forward reasoning phase, through the inference phase indicators To evaluate the inference confidence of large language models, among which... This represents the confidence distribution of the model for candidate recommendation results in the final inference state;

[0037] In the prognostic feedback phase, the feedback phase indicators are used. The study assesses the consistency between treatment recommendations and prognostic predictions, where D(·) represents a measure of inconsistency between treatment recommendations and predicted prognosis. Indicates the first In round-robin reasoning, the predicted prognosis corresponding to the current treatment recommendation. This represents the final recommendation result that has converged stably.

[0038] This invention provides an intelligent analysis method for NK / T-cell lymphoma based on a large-scale intelligent model. It unifies and collaboratively fuses multimodal data, including MRI images, image reports, clinical structured variables, and a medical knowledge base. It intelligently segments the patient's MRI images using a visual segmentation model and performs knowledge retrieval and context construction using the medical knowledge base. Finally, it achieves explicit thought chain reasoning based on a large language model. This addresses the problems of the NK / T-cell lymphoma (ENTCL) diagnosis process being highly dependent on clinical expert experience, resulting in long processing times, significant variability, and difficulties in standardization and large-scale implementation. It provides timely and accurate intelligent assistance for the early and precise diagnosis of ENKTCL, which is of great significance for improving patient survival rates and developing individualized treatment plans, enabling multimodal joint reasoning and decision support for ENKTCL. Attached Figure Description

[0039] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below.

[0040] Figure 1 This is a schematic diagram of an intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent provided by the present invention.

[0041] Figure 2 This is a flowchart illustrating an intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent according to an embodiment of the present invention. Detailed Implementation

[0042] To enable those skilled in the art to better understand the embodiments of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and implementation methods.

[0043] To address the current challenges of ENKTCL diagnosis, which heavily relies on clinical expert experience, is time-consuming, exhibits significant variability, and is difficult to standardize and scale up, this invention provides an intelligent analysis method for NK / T-cell lymphoma based on a large-scale intelligent model. This method solves the problems of ENKTCL diagnosis being highly dependent on clinical expert experience, time-consuming, exhibiting significant variability, and being difficult to standardize and scale up. It can provide timely and accurate intelligent assistance for the early and precise diagnosis of ENKTCL, which is of great significance for improving patient survival rates and developing individualized treatment plans, and realizes multimodal joint reasoning and auxiliary decision-making for ENKTCL.

[0044] like Figure 1 and Figure 2 As shown, a method for intelligent analysis of NK / T cell lymphoma based on a large model intelligent agent includes:

[0045] S1: Obtain the patient's MRI medical images, image reports, clinical structured variables, and knowledge base retrieval results to perform unified modeling and collaborative fusion of data from different modalities, forming a multimodal patient information model;

[0046] S2: Construct a visual segmentation model based on the DINOv3 backbone network to intelligently segment the patient's MRI medical images to identify and extract lesion regions, lesion boundaries and lesion spatial location information, and use them as structured image evidence.

[0047] S3: Construct a dedicated medical knowledge base for ENKTCL to perform knowledge retrieval and context construction based on the patient's corresponding multimodal information and structured imaging evidence;

[0048] S4: Utilize a large language model to establish a chain reasoning mechanism, decompose the forward reasoning process of ENKTCL diagnosis into evidence summary, clinical hypothesis generation, treatment planning, and outcome output.

[0049] Specifically, a collaborative diagnostic framework for ENKTCL is constructed using a Large Language Model (LLM) as the core inference engine. This framework integrates medical images, pathology reports, structured clinical data, and a medical knowledge base to achieve multimodal joint reasoning and decision support for ENKTCL. The large model agent includes a multimodal patient information model, a visual segmentation model, and a large language model. This method constructs an agent for complex ENKTCL diagnostic scenarios, capable of simultaneously supporting the joint analysis of multimodal data such as MRI images, image reports, structured clinical indicators, and a medical knowledge base. It also implements explicit chain-of-thought reasoning based on the large language model. Furthermore, it can dynamically retrieve relevant clinical guidelines, expert consensus, and medical literature based on the patient's current state, enabling knowledge-enhanced clinical decision-making and significantly improving the completeness, interpretability, and medical consistency of complex case diagnoses.

[0050] Furthermore, the unified modeling and collaborative fusion of data from different modalities to form a multimodal patient information model includes:

[0051] The patient's MRI medical images, image reports, clinical structured variables, and knowledge base search results are used as input. The patient input is defined as: ,in, This indicates information about MRI images and lesion areas. This indicates a structured image report. Represents the patient's clinical structured variables. This indicates the results of a knowledge base search.

[0052] By using a modality-specific encoder to learn structured representations of various types of data, information from different sources is transformed into corresponding modality feature representations that can be uniformly processed for subsequent reasoning.

[0053] After obtaining the features of each modality, a multimodal fusion function is used to integrate them to construct a unified semantic representation at the patient level, which simultaneously reflects the patient's lesion morphology, imaging manifestations, clinical status, and medical knowledge-related information.

[0054] Specifically, since ENKTCL clinical diagnosis typically relies on multi-source heterogeneous information such as MRI images, image reports, clinical structured indicators, laboratory test results, and external medical knowledge, unified modeling and collaborative fusion of data from different modalities are necessary. Simultaneously, due to significant differences in data structure and semantic representation between different modalities, it is necessary to convert information from different sources into feature representations that can be uniformly processed by subsequent inference modules. This is often achieved through modality-specific encoders to learn structured representations for various types of data. The corresponding modality encoding process is represented as follows: ;in, This represents the encoding function for the corresponding mode. This represents the feature representation of the corresponding mode.

[0055] For imaging data, a visual coding network is used to extract information such as the spatial location, morphological boundaries, and extent of invasion of lesions; for imaging reports and knowledge base text, a medical text coding model is used to extract semantic features; for structured clinical variables, a structured feature mapping method is used to convert blood indicators, staging information, and other clinical variables into a unified vector representation.

[0056] After obtaining the features of each modality, a multimodal fusion function is further used to integrate the above features and construct a unified semantic representation at the patient level: ;in, Represents a multimodal fusion function. This represents the fused comprehensive patient representation. This representation simultaneously reflects the patient's lesion morphology, imaging findings, clinical status, and related medical knowledge, providing a complete contextual basis for subsequent intelligent reasoning. Through the aforementioned multimodal modeling approach, this method enables the synergistic expression of different medical evidence, thereby improving the completeness and stability of complex ENKTCL case analysis.

[0057] Furthermore, the intelligent segmentation of the patient's MRI medical images includes:

[0058] Extracting deep visual features using the DINOv3 visual backbone network: ,in, denoted as the visual feature extraction network, where f represents the visual features obtained from the input MRI image.

[0059] The obtained visual features are input into the segmentation head network, and the segmentation head network outputs the corresponding lesion region probability map, and then the lesion region is segmented according to the lesion region probability map.

[0060] Furthermore, the step of inputting the obtained visual features into the segmentation head network and outputting the corresponding lesion region probability map through the segmentation head network includes:

[0061] According to the formula: This yields a lesion prediction mask, where each pixel value in the prediction mask represents the probability that the corresponding location belongs to the lesion region. Represents the segmentation head network. This represents the lesion prediction mask, where H and W represent the image height and width, respectively.

[0062] Specifically, ENKTCL lesions in medical imaging typically exhibit irregular boundaries, complex invasive ranges, and significant differences in local tissue structure. Relying solely on image reports or manual annotations is easily influenced by subjective experience. Therefore, visual segmentation models can be used to automatically identify lesion regions, transforming unstructured images into structured imaging evidence that can be used for reasoning.

[0063] Given an input MRI image: ,in, , and These represent the image height, width, and number of channels, respectively. First, deep visual features are extracted using the DINOv3 visual backbone network: ,in, This represents a visual feature extraction network. This represents the visual features obtained from the input MRI image, which can characterize the differences between the lesion area and the surrounding tissue, providing a basis for subsequent pixel-level lesion prediction.

[0064] Subsequently, the system outputs a probability map of the lesion region through the segmentation head network: ;in, Represents the segmentation head network. This represents the lesion prediction mask. Each pixel value in the prediction mask represents the probability that the corresponding location belongs to the lesion region.

[0065] To simultaneously improve the accuracy of lesion region overlap and pixel-level classification, Dice Loss and binary cross-entropy loss can be used to jointly optimize model parameters. ;

[0066] in, and The Dice loss is the loss balance coefficient. It is used to optimize the overlap between the predicted lesion region and the actual labeled region, and is defined as follows: ;

[0067] Binary cross-entropy loss is used to improve the classification accuracy of each pixel, and it is defined as follows: ;in, Indicates the first The real label of each pixel Indicates the first The predicted probability of each pixel. To prevent small constants from causing numerical instability, after the model is trained, the system can automatically output information on the lesion area, lesion boundary, and lesion spatial location, and input this information as structured imaging evidence into the subsequent reasoning system. This enables the imaging evidence to participate in clinical reasoning in an interpretable and verifiable manner.

[0068] Furthermore, the knowledge retrieval and context construction based on the patient's corresponding multimodal information and structural human imaging evidence includes:

[0069] The knowledge documents are preprocessed in a structured manner and then vectorized using a text embedding model, so that all knowledge vectors are stored in a vector database to support efficient similarity retrieval in the future.

[0070] After receiving multimodal patient information, a query vector is generated from the patient information, and a patient-specific query function is automatically constructed. A similarity search is performed through the vector database to obtain a knowledge set related to the patient's status.

[0071] The patient fusion representation is integrated with the retrieved knowledge set using a context fusion function.

[0072] Specifically, since ENKTCL diagnosis and treatment heavily rely on clinical guidelines, expert consensus, and the latest medical literature, relying solely on the knowledge stored in the parameters of the large language model itself cannot guarantee the timeliness and accuracy of the content. Furthermore, to reduce the illusion problem of large language models in medical scenarios and enhance the reliability of reasoning, a dynamic knowledge enhancement and context construction mechanism is further constructed. Therefore, a dedicated medical knowledge base for ENKTCL is built, and knowledge content related to the current case is dynamically retrieved during each patient reasoning process.

[0073] First, the knowledge documents undergo structured preprocessing, and then are vectorized using a text embedding model: ;

[0074] in, Represents knowledge entries. This represents a text embedding function. This represents the vector representation of knowledge entries. All knowledge vectors are stored in a vector database to support efficient similarity retrieval later.

[0075] After receiving multimodal patient information, the system automatically constructs patient-specific queries based on the patient data: ;

[0076] in, This indicates the query builder function. This represents a query vector generated from patient information. The system then performs a similarity search through the vector database to obtain the set of knowledge most relevant to the patient's condition. ,in, Indicates the first result obtained from the search One related knowledge item.

[0077] Then, the patient fusion representation is integrated with the retrieved knowledge using a contextual fusion function: ,in, This represents the final context provided to the large language model. This represents a context-building function. Through this knowledge enhancement mechanism, large language models can perform reasoning under the joint constraints of patient-specific data and external medical knowledge, thereby reducing the risk of hallucinations and improving the medical consistency of diagnostic and treatment recommendations.

[0078] Furthermore, the establishment of a chain reasoning mechanism using a large language model includes:

[0079] Using the formula: The system reasons from the knowledge set related to the patient's condition to generate a summary of clinical evidence based on the patient's contextual information. This represents the final context provided to the large language model. This represents the summary of clinical evidence. Represents the inference function of the large language model in the evidence summarization stage;

[0080] Based on the summary of the clinical evidence, intermediate clinical hypotheses are generated to describe disease risk, lesion invasion, clinical staging tendency, and potential treatment directions.

[0081] Using a structured output function, initial treatment recommendations are generated based on the obtained clinical hypotheses. The recommendations include treatment strategies, treatment regimens, and suggestions for the use of PD-1 inhibitors.

[0082] Specifically, traditional models typically derive the final result directly from the input data, lacking intermediate reasoning processes, making it difficult to judge the reasonableness of the model's conclusions. By employing a chain-reasoning mechanism, large language models can simulate the step-by-step analytical process of clinicians, enabling explicit reasoning for complex clinical decisions.

[0083] First, a summary of clinical evidence is generated based on the patient's contextual information: ,in, This represents the summary of clinical evidence. This represents the inference function of the large language model during the evidence summarization phase.

[0084] Subsequently, based on the above evidence, an intermediate clinical hypothesis was generated: ,in, This indicates a clinical hypothesis based on evidence, which can be used to describe disease risk, lesion invasion, clinical staging tendencies, and potential treatment directions.

[0085] Then, initial treatment recommendations are generated based on clinical assumptions: ,in, This represents a structured output function. This indicates the initial recommendation result.

[0086] The recommendation results are expressed as follows: ,in, Indicates treatment strategy, Indicates the treatment plan, This indicates recommendations for the use of PD-1 inhibitors. Through explicit chained reasoning, this method breaks down the ENKTCL diagnostic reasoning process into consecutive steps such as evidence summarization, clinical hypothesis generation, treatment planning, and outcome output. It retains interpretable intermediate states at each step and fully records the reasoning basis, intermediate conclusions, and decision-making logic, thereby improving the system's interpretability and clinical credibility.

[0087] The method also includes:

[0088] S5: Construct an iterative correction mechanism based on prognostic feedback to predict the possible prognostic outcomes of the current treatment recommendations and input the prediction results back into the reasoning process to dynamically correct the original reasoning path based on prognostic feedback, thus forming an internal iterative correction mechanism.

[0089] Furthermore, the construction of the iterative correction mechanism based on prognostic feedback includes:

[0090] A prognostic prediction head is set up using a large language model to predict patient prognostic outcomes based on the current treatment plan and patient information;

[0091] The prognostic prediction results are re-inputted into the inference module to update the current inference state. New treatment suggestions are generated based on the updated inference state. After multiple rounds of iteration, a stable and converged final recommendation result is obtained.

[0092] Specifically, to avoid the error accumulation problem caused by the unidirectional reasoning of traditional large language models, an iterative correction mechanism based on prognostic feedback is introduced. This mechanism not only generates treatment suggestions but also predicts the possible prognostic outcomes of the current treatment suggestions and feeds these predictions back into the reasoning process. This allows for dynamic correction of the original reasoning path based on prognostic feedback. The reasoning process is no longer limited to a one-time output but forms a closed-loop reasoning process that includes evaluation, feedback, and updates.

[0093] First, predict the patient's prognosis based on the current treatment plan and patient information: ;in, Indicates the first In round-robin reasoning, the predicted prognosis corresponding to the current treatment recommendation. This indicates the prognostic prediction module.

[0094] Subsequently, the prognostic prediction results are re-entered into the inference module to update the current inference state: ;

[0095] And generate new treatment suggestions based on the updated inference state: ;

[0096] After multiple iterations, the system obtains a stable and converged final recommendation result: ,in, This indicates the iteration round number. Through this feedback loop mechanism, the method can automatically detect inconsistencies between the current treatment recommendation and the prognosis, and dynamically correct the reasoning path, thereby improving the stability of the reasoning and the rationality of the treatment plan.

[0097] This method is based on a closed-loop iterative optimization mechanism using prognostic feedback. By predicting the prognosis of the current treatment plan and feeding the prediction results back into the reasoning process, it achieves dynamic correction of intermediate reasoning paths and treatment recommendations. This mechanism can effectively avoid the problem of early error propagation in the unidirectional reasoning process of traditional large language models, thereby improving the stability of system reasoning, the rationality of treatment recommendations, and robustness in complex clinical scenarios.

[0098] The method also includes:

[0099] S6: Calculate the uncertainty and risk values ​​for key stages of image segmentation, forward inference, and prognostic feedback, and determine whether to trigger physician intervention review and correction based on the corresponding risk threshold, so as to introduce an adaptive external physician supervision mechanism.

[0100] Furthermore, the calculation of uncertainty and risk values ​​for the key stages of image segmentation, forward inference, and prognostic feedback includes:

[0101] In the image segmentation stage, uncertainty indicators are used... Assess the reliability of lesion boundaries, among which, Let represent the predicted probability at pixel i, and N represent the total number of pixels.

[0102] In the forward reasoning phase, through the inference phase indicators To evaluate the inference confidence of large language models, among which... This represents the confidence distribution of the model for candidate recommendation results in the final inference state.

[0103] In the prognostic feedback phase, the feedback phase indicators are used. The consistency between treatment recommendations and prognostic predictions is assessed, where D(·) represents the measure of inconsistency between treatment recommendations and prognostic predictions.

[0104] Specifically, since ENKTCL is a high-risk malignant tumor, the reasoning process must have full-process reliability control. Therefore, a dynamic risk assessment and physician collaboration mechanism was further designed. This differs from simply providing an overall confidence level at the final output stage and deciding whether to trigger physician intervention based on the stage risk.

[0105] For any stage Calculate the uncertainty and risk value separately:

[0106] ;

[0107] ;

[0108] in, Indicates stage uncertainty score, This indicates the stage of risk assessment. The decision to trigger physician intervention is based on the risk threshold.

[0109] ;

[0110] in, This indicates that a doctor has been called in. This indicates the risk threshold for a given stage. This mechanism allows doctors to participate in the review process only when necessary, thereby reducing their workload while ensuring safety.

[0111] During the lesion segmentation phase, the system assesses the reliability of lesion boundaries using the following uncertainty indicators:

[0112] ;

[0113] in, Represents pixels The predicted probability at that location. This represents the total number of pixels. If the confidence level of the segmented region is low, it triggers the doctor to review and correct the lesion boundaries, missed areas, or mis-segmented areas.

[0114] During the inference phase, the inference confidence of the large language model is evaluated using the following metrics:

[0115] ;

[0116] in, This represents the confidence distribution of the model's candidate recommendations in the final inference state. When the inference confidence is insufficient, it triggers doctors to conduct targeted reviews of intermediate inference conclusions, risk stratification, and treatment recommendations.

[0117] During the prognostic feedback phase, the consistency between treatment recommendations and prognostic predictions is assessed using the following indicators:

[0118] ;

[0119] in, This measures the discrepancy between treatment recommendations and predicted prognosis. When a significant conflict is detected between the recommended treatment and the prognostic outcome, physicians can intervene to adjust the treatment approach or provide higher-level clinical guidance.

[0120] This method employs a dynamic risk assessment mechanism throughout the entire process, evaluating system uncertainty and risk levels at the lesion segmentation, clinical reasoning, and prognostic feedback stages. Based on the risk status, it dynamically determines whether to trigger physician intervention, thus introducing an adaptive external physician supervision mechanism. This mechanism guides physicians to conduct targeted review and correction at key high-risk points, thereby achieving efficient human-machine collaborative diagnosis. While reducing physician workload, it significantly improves system reliability and clinical safety.

[0121] This method first segments lesions in MRI images and extracts structured image features; then it integrates image, pathology, clinical indicators, and knowledge base information to construct a multimodal context for the patient; next, a large language model performs step-by-step clinical reasoning and dynamically corrects the results based on prognostic predictions; finally, the system determines whether physician intervention is needed through a risk assessment mechanism, and through internal iterative correction mechanisms and external physician supervision mechanisms, it can build a reliable medical intelligent system with "self-correction + human-machine collaboration" capabilities, achieving highly reliable collaborative diagnosis.

[0122] In one embodiment, a systematic test and validation were conducted on 568 real clinical cases of ENKTCL collected from a hospital over the past five years. The experimental data encompassed multimodal data including MRI images, pathology reports, clinical structured indicators, and treatment and prognostic information, and was rigorously labeled and reviewed by clinical experts. The experimental results are shown in Table 1.

[0123]

[0124] Indicator meaning:

[0125] ZS: Zero-Shot evaluation; FS: Few-Shot evaluation;

[0126] Mean Accuracy: The average of the Accuracy field in the recommended treatment plan output.

[0127] mean Macro F1: The output field of the recommended treatment plan is the average of Macro F1.

[0128] Strict Accuracy: The percentage of recommended treatment plans where all fields are completely correct;

[0129] Top 3 Hit Rate: The percentage of at least one completely correct treatment option among the recommended and alternative treatment options (3 in total);

[0130] BLEU-4 and ROUGE-L are matching indicators for treatment protocol descriptions. BLEU-4 measures whether the expression is consistent, and ROUGE-L measures whether the content is covered.

[0131] Experimental results show that our proposed method significantly outperforms existing baseline methods and traditional large language model methods across all metrics. Specifically, it achieves a Mean Accuracy score of 0.82, a Mean Macro F1 score of 0.78, a Strict Accuracy score of 0.70, and a Top 3 Hit Rate score of 0.90, all significantly better than existing models such as Qwen3-VL, LLaVA3.2, and Baichuan. Furthermore, our method achieves scores of 0.53 and 0.75 on the BLEU-4 and ROUGE-L metrics, respectively, related to the quality of treatment plan text generation. This indicates that our method not only improves the accuracy of treatment plan prediction but also generates treatment recommendations that are more consistent with clinical expression norms and medical semantics.

[0132] As can be seen, this invention provides an intelligent analysis method for NK / T-cell lymphoma based on a large-scale intelligent agent model. It unifies and collaboratively integrates multimodal data such as MRI images, image reports, clinical structured variables, and medical knowledge bases. It intelligently segments the patient's MRI medical images using a visual segmentation model and utilizes the medical knowledge base for knowledge retrieval and context construction. Furthermore, it achieves explicit thought chain reasoning based on a large language model. This addresses the problems of the NK / T-cell lymphoma (ENKTCL) diagnosis process being highly dependent on clinical expert experience, resulting in long processing times, significant variability, and difficulties in standardization and large-scale implementation. It provides timely and accurate intelligent assistance for the early and precise diagnosis of ENKTCL, which is of great significance for improving patient survival rates and developing individualized treatment plans, realizing multimodal joint reasoning and decision support for ENKTCL.

[0133] The structure, features, and effects of the present invention have been described in detail above with reference to the embodiments shown in the figures. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, shall be within the protection scope of the present invention as long as they do not exceed the spirit covered by the specification and figures.

Claims

1. A method for intelligent analysis of NK / T-cell lymphoma based on a large-scale intelligent agent model, characterized in that, include: The system acquires patients' MRI medical images, image reports, clinical structured variables, and knowledge base retrieval results to perform unified modeling and collaborative fusion of data from different modalities, forming a multimodal patient information model. A visual segmentation model based on the DINOv3 backbone network was constructed to intelligently segment the patient's MRI medical images to identify and extract lesion regions, lesion boundaries, and lesion spatial location information, and use them as structured imaging evidence. We constructed a dedicated medical knowledge base for ENKTCL to perform knowledge retrieval and context construction based on the patient's corresponding multimodal information and structured imaging evidence. By using a large language model to establish a chain reasoning mechanism, the forward reasoning process of ENKTCL diagnosis is decomposed into evidence summary, clinical hypothesis generation, treatment planning, and outcome output.

2. The intelligent analysis method for NK / T cell lymphoma based on a large-scale intelligent agent according to claim 1, characterized in that, Also includes: An iterative correction mechanism based on prognostic feedback is constructed to predict the possible prognostic outcomes of the current treatment recommendations and input the prediction results back into the reasoning process to dynamically correct the original reasoning path based on the prognostic feedback, thus forming an internal iterative correction mechanism.

3. The intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent according to claim 2, characterized in that, Also includes: Uncertainty and risk values ​​are calculated for key stages of image segmentation, forward inference, and prognostic feedback, and the appropriate risk thresholds are used to determine whether to trigger physician intervention for review and correction, thereby introducing an adaptive external physician supervision mechanism.

4. The intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent according to claim 3, characterized in that, The process of unified modeling and collaborative fusion of data from different modalities to form a multimodal patient information model includes: The patient's MRI medical images, image reports, clinical structured variables, and knowledge base search results are used as input. The patient input is defined as: ,in, This indicates information about MRI images and lesion areas. This indicates a structured image report. Represents the patient's clinical structured variables. This indicates the results of a knowledge base search. By using a modality-specific encoder to learn structured representations of various types of data, information from different sources is transformed into corresponding modality feature representations that can be uniformly processed for subsequent reasoning; After obtaining the features of each modality, a multimodal fusion function is used to integrate them to construct a unified semantic representation at the patient level, which simultaneously reflects the patient's lesion morphology, imaging manifestations, clinical status, and medical knowledge-related information.

5. The intelligent analysis method for NK / T cell lymphoma based on a large model intelligent agent according to claim 4, characterized in that, The intelligent segmentation of the patient's MRI medical images includes: Deep visual features are extracted using a DINOv3 visual backbone network: wherein, denotes a visual feature extraction network, and f denotes visual features derived from the input MRI image; The obtained visual features are input into the segmentation head network, and the segmentation head network outputs the corresponding lesion region probability map, and then the lesion region is segmented according to the lesion region probability map.

6. The NK / T cell lymphoma intelligent analysis method based on a large model agent according to claim 5, characterized in that, The process of inputting the obtained visual features into the segmentation head network and outputting the corresponding lesion region probability map through the segmentation head network includes: According to the formula: , the lesion prediction mask is obtained, and each pixel value in the prediction mask represents the probability that the corresponding position belongs to the lesion area, wherein, represents the segmentation head network, represents the lesion prediction mask, H and W represent the height and width of the image respectively.

7. The NK / T cell lymphoma intelligent analysis method based on a large model agent according to claim 6, characterized in that, The knowledge retrieval and context construction based on the patient's corresponding multimodal information and structural human imaging evidence includes: The knowledge documents are preprocessed in a structured manner and then vectorized using a text embedding model, so that all knowledge vectors are stored in a vector database to support efficient similarity retrieval in the future. After receiving multimodal patient information, a query vector is generated from the patient information, and a patient-specific query function is automatically constructed. A similarity search is performed through the vector database to obtain a knowledge set related to the patient's status. The patient fusion representation is integrated with the retrieved knowledge set using a context fusion function.

8. The NK / T cell lymphoma intelligent analysis method based on a large model agent according to claim 7, characterized in that, The chain reasoning mechanism established using a large language model includes: Using the formula: , the knowledge set related to the patient state is inferred to generate a clinical evidence summary according to the patient context information, wherein, represents the final context provided to the large language model, represents the clinical evidence summary result, represents the inference function of the large language model in the evidence summary stage; Intermediate clinical hypotheses are generated based on the results of summarizing clinical evidence to obtain descriptions of disease risk, lesion invasion, clinical staging tendencies, and potential treatment directions. Using a structured output function, initial treatment recommendations are generated based on the obtained clinical hypotheses. The recommendations include treatment strategies, treatment regimens, and suggestions for the use of PD-1 inhibitors.

9. The NK / T cell lymphoma intelligent analysis method based on a large model agent according to claim 8, characterized in that, The construction of the iterative correction mechanism based on prognostic feedback includes: A prognostic prediction head is set up using a large language model to predict patient prognostic outcomes based on the current treatment plan and patient information; The prognostic prediction results are re-inputted into the inference module to update the current inference state. New treatment suggestions are generated based on the updated inference state. After multiple rounds of iteration, a stable and converged final recommendation result is obtained.

10. The NK / T cell lymphoma intelligent analysis method based on a large model agent according to claim 9, wherein, The calculation of uncertainty and risk values ​​for the key stages of image segmentation, forward inference, and prognostic feedback includes: In the image segmentation phase, the lesion boundary reliability is evaluated by an image segmentation phase uncertainty indicator , wherein represents the prediction probability at pixel i, and N represents the total number of pixels. In the forward inference phase, the inference confidence of the large language model is evaluated through the inference phase indicator , wherein represents the confidence distribution of the model on the candidate recommended result in the final inference state; In the prognosis feedback phase, the consistency between the treatment recommendation and the prognosis prediction is evaluated by the feedback phase indicator , where D(·) denotes the inconsistency measure between the treatment recommendation and the predicted prognosis, denotes the predicted prognosis result corresponding to the current treatment recommendation in the th round of reasoning, denotes the final recommended result of stable convergence.