Cholangiocarcinoma ct image prediction method and system based on momentum attention and large model verification
By employing momentum attention and large model validation, the problem of inconsistent visual feature extraction and reasoning logic in CT image diagnosis of cholangiocarcinoma was solved, achieving efficient and reliable image diagnosis and report generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE AFFILIATED HOSPITAL OF QINGDAO UNIV
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115991A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging technology, and in particular to a method and system for predicting cholangiocarcinoma CT images based on momentum attention and large model validation. Background Technology
[0002] Upper abdominal CT imaging is an important tool for clinical screening and disease assessment of cholangiocarcinoma. However, current technologies lack effective momentum mechanisms to guide attention allocation when extracting visual features from cholangiocarcinoma CT images. This makes it difficult to dynamically and accurately update the focus of attention on the images, and the mining and utilization of high-dimensional visual features are insufficient. Consequently, the extraction accuracy of key signs related to cholangiocarcinoma in the images is low, and the correlation between visual features and diagnostic reasoning is weak, which directly affects the effectiveness of subsequent diagnostic analysis.
[0003] Current intelligent diagnostic methods for cholangiocarcinoma CT images lack a systematic iterative reasoning logic, have insufficient depth and coherence in image-text fusion, and lack professional logical verification of the reasoning chain and preliminary diagnostic conclusions. This easily leads to mismatches between the reasoning basis and the diagnostic results, resulting in insufficient credibility of the diagnostic conclusions. At the same time, the generation of diagnostic reports lacks a standardized compilation process, and the efficiency of the connection between each link is low. The accuracy and smoothness of the overall diagnostic process need to be improved. Therefore, how to improve the overall efficiency of intelligent prediction of cholangiocarcinoma CT images has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a method and system for predicting cholangiocarcinoma CT images based on momentum attention and large model validation, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation, comprising:
[0006] S1. Based on the pre-trained visual-language large model, visual features are extracted from the upper abdominal non-enhanced CT image data of the patient to be diagnosed in order to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and the historical momentum attention map and initial inference text sequence of the visual-language large model are initialized.
[0007] S2. Based on the historical momentum attention map and the initial inference text sequence, the attention focus of the high-dimensional visual embedding vector is updated using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed.
[0008] S3. Based on the first momentum attention map, extract the first key image patch set from the upper abdominal non-enhanced CT image, and perform image-text interleaving fusion of the first key image patch set with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed.
[0009] S4. Iterate through S2 to S3 until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated.
[0010] S5. Based on the preset plain text logic verifier, the correlation between the reasoning chain and the preliminary diagnostic conclusion is evaluated to obtain the logical consistency score of the patient to be diagnosed.
[0011] S6. Compile the logical consistency scores to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.
[0012] In a preferred embodiment, the pre-trained visual-language large model extracts visual features from the unenhanced CT images of the upper abdomen of the patient to be diagnosed, constructing a high-dimensional visual embedding vector for the patient, and initializes the historical momentum attention map and initial inference text sequence of the visual-language large model, including:
[0013] The raw tomographic scan data of the upper abdomen of the patient to be diagnosed is obtained, and the window width and window level are adjusted on the raw tomographic scan data to obtain the grayscale tomographic image sequence of the patient to be diagnosed.
[0014] The grayscale tomographic image sequence is input into a pre-trained visual-language large model for encoding and mapping to obtain the local visual feature vector of the patient to be diagnosed.
[0015] Feature aggregation is performed on the local visual feature vectors to obtain the high-dimensional visual embedding vector of the patient to be diagnosed.
[0016] A momentum attention buffer is constructed for the visual-language large model to record historical observation paths, and the momentum attention buffer is cleared to obtain the historical momentum attention map of the visual-language large model;
[0017] Start symbols are implanted into the text generator of the visual-language big model to obtain the initial inference text sequence of the visual-language big model.
[0018] In a preferred embodiment, the step of updating the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed includes:
[0019] The anatomical structure pointing information in the initial inference text sequence is parsed to extract the region of interest descriptive words of the anatomical structure pointing information;
[0020] Based on the descriptive words of the region of interest, a preliminary attention weight allocation is performed on the high-dimensional visual embedding vector to obtain the initial attention distribution map of the patient to be diagnosed.
[0021] The attention weight distribution of the last observation path in the historical momentum-attention map is used as the momentum-inertia guiding term for the patient to be diagnosed.
[0022] By fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector, a corrected attention distribution map of the patient to be diagnosed is obtained.
[0023] The modified attention distribution map is normalized to obtain the first momentum attention map of the patient to be diagnosed.
[0024] In a preferred embodiment, the step of fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector to obtain the corrected attention distribution map of the patient to be diagnosed includes:
[0025] The momentum decay coefficient of the historical momentum attention map is determined based on the semantic complexity in the initial inference text sequence.
[0026] The momentum inertia guidance term and the momentum decay coefficient are weighted and fused to obtain the historical attention contribution value of the patient to be diagnosed.
[0027] The initial attention distribution map is multiplied element-wise with the complementary coefficient of the momentum decay coefficient to obtain the current attention contribution value of the patient to be diagnosed.
[0028] The historical attention contribution value and the current attention contribution value are superimposed, and the weights of the attention focus of the high-dimensional visual embedding vector are reorganized to obtain the attention fusion result of the patient to be diagnosed.
[0029] Based on the attention fusion results, the initial attention distribution map is recalibrated to obtain the corrected attention distribution map of the patient to be diagnosed.
[0030] In a preferred embodiment, the step of extracting a first key image patch set from the upper abdominal non-contrast CT image based on the first momentum-attention map, and then performing image-text interleaving fusion of the first key image patch set with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed includes:
[0031] Obtain the attention weight distribution of each image region in the first momentum attention map;
[0032] Based on the attention weight distribution, the importance of all image regions in the non-contrast CT image of the upper abdomen is ranked to obtain the candidate key regions of the patient to be diagnosed.
[0033] Local image patches corresponding to the candidate key regions are cropped from the non-contrast CT images of the upper abdomen, and the size of the local image patches is standardized to obtain the first key image patch set of the patient to be diagnosed.
[0034] The first set of key image patches is inserted as visual evidence markers into the designated positions of the initial inference text sequence to construct the image-text interleaved input sequence of the patient to be diagnosed.
[0035] The image-text interleaved input sequence is input into the vision-language large model for text decoding to obtain the first inference text of the patient to be diagnosed. The calculation formula for the first inference text is as follows:
[0036] ;
[0037] In the formula, The first inference text, For the aforementioned large visual-language model, For global visual embedding vectors, This is a selection function that sorts and selects the top K regions based on attention weights. The original tomographic scan data, This is the first momentum attention map. A historical reasoning text sequence, These are preset model weight parameters. In a preferred embodiment, the iterations S2 to S3 are performed until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated, including:
[0038] The first inference text is used as the text generation result of the current iteration round, and the first momentum attention map is used as the attention update result of the current iteration round;
[0039] The attention update result of the current iteration round is used as the historical momentum attention graph when the attention focus update is performed next, and the text generation result of the current iteration round is appended to the historical inference text sequence as the historical inference text sequence when the attention focus update is performed next.
[0040] Based on the historical momentum attention map and the historical sequence of inference text, the attention focus update operation is performed again to obtain the next iteration momentum attention map of the patient to be diagnosed.
[0041] Based on the momentum attention map of the next iteration round, the key image patch set extraction and image-text interleaving fusion operation are performed again to obtain the next iteration round inference text of the patient to be diagnosed.
[0042] The momentum attention map generated in each iteration is used as the historical momentum attention map when performing attention focus update in subsequent iterations. The inference text generated in each iteration is appended to the historical inference text sequence. The attention focus update operation and the key image patch set extraction and image-text interleaving operation are repeatedly performed until the visual-language large model output terminator is detected and the repetition stops.
[0043] The reasoning texts generated in all repeated rounds are concatenated sequentially to obtain the reasoning thought chain of the patient to be diagnosed. The preliminary diagnostic conclusion of the patient to be diagnosed is then parsed from the reasoning text output in the last repeated round. The calculation formula for the preliminary diagnostic conclusion is as follows:
[0044] , ;
[0045] In the formula, For the aforementioned reasoning chain, This is the preliminary diagnostic conclusion. The model, given the global visual embedding vector The first key image patch set and the aforementioned historical reasoning text sequence Under the condition of predicting and generating the current word unit The conditional probability, To find the reasoning path and conclusion combination that maximizes the joint probability, The total length of the generated sequence. In a preferred embodiment, the step of evaluating the correlation between the reasoning chain and the preliminary diagnostic conclusion based on a preset plain text logic checker to obtain the logical consistency score of the patient to be diagnosed includes:
[0046] The morphological descriptor sequence and the logical analysis statement sequence in the reasoning chain are concatenated to form the verification text of the patient to be diagnosed.
[0047] The preliminary diagnostic conclusion and the text to be verified are input together into a preset plain text logic verifier;
[0048] The radiographic evidence set of the patient to be diagnosed is constructed by parsing the radiographic feature descriptive words in the text to be verified using the plain text logic verifier.
[0049] The diagnostic reference words in the preliminary diagnostic conclusion are parsed by the plain text logic validator to extract the diagnostic conclusion reference of the patient to be diagnosed.
[0050] The number of imaging features described in the set of imaging evidence that support the diagnostic conclusion is counted.
[0051] The number of descriptive words of the imaging features is correlated with the imaging evidence set to obtain the logical consistency score of the patient to be diagnosed.
[0052] In a preferred embodiment, the logical consistency score is calculated using the following formula:
[0053] ;
[0054] In the formula, The logical consistency score is given. The total number of descriptive terms for imaging features in the aforementioned set of imaging evidence. For the set of radiological evidence, the first A descriptive term for an imaging feature, The diagnostic conclusion extracted from the preliminary diagnostic conclusion points to... This is a pre-defined semantic similarity calculation function within the plain text logic validator.
[0055] In a preferred embodiment, the step of compiling reports from the logical consistency scores to obtain a target intelligent assisted diagnostic report for the patient to be diagnosed includes:
[0056] The logical consistency score is compared with a preset logical consistency threshold to obtain the verification status of the patient to be diagnosed.
[0057] When the verification status is "verification passed", extract the key morphological description statements in the reasoning chain;
[0058] The logical consistency score, the preliminary diagnostic conclusion, and the key morphological description statements are structured and filled in to obtain the initial draft of the diagnostic report for the patient to be diagnosed.
[0059] Extract the CT image layer identifiers corresponding to the first key image patch set from the reasoning chain, and embed the CT image layer identifiers as visual evidence annotations into the designated position of the initial draft of the diagnostic report;
[0060] The initial draft of the diagnostic report is formatted to obtain the target intelligent assisted diagnostic report for the patient to be diagnosed.
[0061] To address the aforementioned problems, this invention also provides a CT image prediction system for cholangiocarcinoma based on momentum attention and large model validation, the system comprising:
[0062] The visual encoding and initialization module is used to extract visual features from the non-enhanced CT image data of the upper abdomen of the patient to be diagnosed based on the pre-trained visual-language large model, so as to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and initialize the historical momentum attention map and initial inference text sequence of the visual-language large model.
[0063] The momentum attention update module is used to update the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence, so as to obtain the first momentum attention map of the patient to be diagnosed.
[0064] The image-text interleaving reasoning module is used to extract the first key image patch set from the upper abdominal non-enhanced CT image based on the first momentum attention map, and to perform image-text interleaving fusion of the first key image patch set with the initial reasoning text sequence to obtain the first reasoning text of the patient to be diagnosed.
[0065] The iterative generation module is used to iteratively execute the momentum attention update module to the graph-text interleaved reasoning module until the reasoning thought chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated.
[0066] The logic verification module is used to evaluate the correlation between the reasoning chain and the preliminary diagnostic conclusion based on a preset plain text logic verifier, and obtain the logic consistency score of the patient to be diagnosed.
[0067] The report generation module is used to compile the logical consistency score into a report to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] 1. This invention relies on a pre-trained visual-language large model to construct high-dimensional visual embedding vectors for non-enhanced CT images of upper abdominal bile duct cancer. Through a momentum mechanism, it achieves dynamic guidance and updating of attention focus, which can accurately locate and extract key image patch sets in the images. Combined with image-text interleaving, it achieves deep integration of visual features and textual reasoning. At the same time, through continuous attention updates and iterative generation of image-text fusion, it effectively improves the accuracy of CT image feature mining, gives the diagnostic reasoning process a clear logical hierarchy, and provides solid visual image evidence to support the preliminary diagnostic conclusion.
[0070] 2. This invention uses a pre-set plain text logic verifier to quantitatively evaluate the correlation between the reasoning chain and the preliminary diagnostic conclusion, obtaining a logical consistency score. This verifies the rationality of the diagnostic conclusion from a textual logic perspective, effectively ensuring the credibility of the diagnostic results. Simultaneously, based on this score, standardized diagnostic reports are compiled, embedding CT image layer markers as visual evidence annotations into the reports. The diagnostic information is then structured and formatted, improving the standardization and completeness of the intelligent assisted diagnostic report generation. Furthermore, it significantly improves the overall efficiency of CT image prediction for cholangiocarcinoma, providing efficient and reliable intelligent auxiliary support for the clinical imaging diagnosis of cholangiocarcinoma. Attached Figure Description
[0071] Figure 1 This is a flowchart illustrating a method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation, provided in an embodiment of the present invention.
[0072] Figure 2 A functional block diagram of a cholangiocarcinoma CT image prediction system based on momentum attention and large model validation provided in an embodiment of the present invention;
[0073] Figure 3 A comparison chart of model results for a cholangiocarcinoma CT image prediction system based on momentum attention and large model validation provided in an embodiment of the present invention; Detailed Implementation
[0074] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings.
[0075] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0076] This application provides a method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0077] Reference Figure 1 The diagram shown is a flowchart illustrating a method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation according to an embodiment of the present invention. In this embodiment, the method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation includes:
[0078] S1. Based on the pre-trained visual-language large model, visual features are extracted from the upper abdominal non-enhanced CT image data of the patient to be diagnosed in order to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and the historical momentum attention map and initial inference text sequence of the visual-language large model are initialized.
[0079] In this embodiment of the invention, the pre-trained visual-language large model extracts visual features from the unenhanced CT images of the upper abdomen of the patient to be diagnosed, constructs a high-dimensional visual embedding vector of the patient, and initializes the historical momentum attention map and initial inference text sequence of the visual-language large model, including:
[0080] The raw tomographic scan data of the upper abdomen of the patient to be diagnosed is obtained, and the window width and window level are adjusted on the raw tomographic scan data to obtain the grayscale tomographic image sequence of the patient to be diagnosed.
[0081] The grayscale tomographic image sequence is input into a pre-trained visual-language large model for encoding and mapping to obtain the local visual feature vector of the patient to be diagnosed.
[0082] Feature aggregation is performed on the local visual feature vectors to obtain the high-dimensional visual embedding vector of the patient to be diagnosed.
[0083] A momentum attention buffer is constructed for the visual-language large model to record historical observation paths, and the momentum attention buffer is cleared to obtain the historical momentum attention map of the visual-language large model;
[0084] Start symbols are implanted into the text generator of the visual-language big model to obtain the initial inference text sequence of the visual-language big model.
[0085] The original tomographic scan data of the upper abdomen non-contrast CT images of the patient to be diagnosed were retrieved from the medical imaging equipment. This data is a matrix of CT values containing pixels. For the image analysis needs of the upper abdominal bile duct related tissues, the window width was set to 300HU to 400HU and the window level was set to 40HU to 60HU. The CT value of each pixel in the original tomographic scan data was mapped and adjusted. CT values exceeding the upper limit of the window width were mapped to a gray value of 255, and CT values below the lower limit of the window width were mapped to a gray value of 0. CT values within the window width range were converted to corresponding gray values between 0 and 255 according to a linear ratio. After the gray value conversion of all pixels was completed, the converted grayscale images were arranged in order according to the slice thickness of the original tomographic scan to form a grayscale tomographic image sequence of the patient to be diagnosed.
[0086] The individual grayscale tomographic images are sequentially input into the visual coding layer of the vision-language large model according to the layer order of the grayscale tomographic image sequence. The visual coding layer extracts spatial features from the pixel matrix of each grayscale tomographic image through a preset convolutional kernel. The convolutional kernel slides across the pixel matrix with fixed row and column spacing to extract local spatial features such as texture, contour, and grayscale distribution at different spatial locations of each image. The extracted local spatial features of each image are then transformed into structured vectorization to form local feature sub-vectors corresponding to each image. Finally, the local feature sub-vectors of all individual images are combined in an orderly manner according to the layer order of the grayscale tomographic image sequence to form the local visual feature vector of the patient to be diagnosed.
[0087] The local visual feature vectors are aggregated by channel-dimensional stitching. The local feature sub-vectors corresponding to each grayscale tomographic image in the local visual feature vectors are stitched continuously and orderly according to the feature channel dimensions preset by the visual-language big model. During the stitching process, the spatial position information and hierarchical information of each local feature sub-vector are completely preserved. The dimensionality of the stitched feature vector is then increased to form a feature representation vector with a feature dimension higher than that of the local visual feature vector. This feature representation vector is the high-dimensional visual embedding vector of the patient to be diagnosed.
[0088] A matrix-like storage area is constructed as a momentum attention cache for the large vision-language model. The number of rows and columns of this cache corresponds one-to-one with the feature dimensions of the high-dimensional visual embedding vector. Each storage unit in the cache is used to record the attention weight information of the corresponding feature dimension in the high-dimensional visual embedding vector. The initial values of all storage units in this matrix-like storage area are set to 0 to complete the zeroing operation of the momentum attention cache. The zeroed momentum attention cache is the historical momentum attention map of the large vision-language model.
[0089] At the starting position of text sequence generation in the visual-language large model text generator, a unique text generation trigger character set after the model has completed pre-training is implanted as the start character. This start character is a fixed single character and is only used to trigger the inference text generation process of the text generator. The text sequence containing only this start character is the initial inference text sequence of the visual-language large model.
[0090] The beneficial effects are as follows: the above implementation process establishes a standardized and reproducible execution standard for visual feature extraction and model initialization; the adjustment of fixed window width and window level ensures the feature effectiveness of grayscale tomographic image sequences; the local visual feature vectors obtained by convolution extraction and ordered combination can completely preserve the spatial and tonal information of CT images; the feature aggregation method of channel dimension splicing gives high-dimensional visual embedding vectors more comprehensive feature representation capabilities; the historical momentum attention map formed after the momentum attention buffer matching the dimension of the high-dimensional visual embedding vector is cleared can provide a unified and accurate initial benchmark for subsequent attention focus updates; the initial inference text sequence formed by implanting a dedicated start symbol can effectively trigger the subsequent text inference process; the operation of the entire process revolves around the technical requirements of subsequent momentum attention updates and image-text interleaved inference, ensuring that the generated vectors and the initial state of the model can accurately match subsequent operations without deviations in feature dimension or format, thus ensuring the consistency and stability of the overall technical solution.
[0091] S2. Based on the historical momentum attention map and the initial inference text sequence, the attention focus of the high-dimensional visual embedding vector is updated using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed.
[0092] In this embodiment of the invention, the step of updating the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed includes:
[0093] The anatomical structure pointing information in the initial inference text sequence is parsed to extract the region of interest descriptive words of the anatomical structure pointing information;
[0094] Based on the descriptive words of the region of interest, a preliminary attention weight allocation is performed on the high-dimensional visual embedding vector to obtain the initial attention distribution map of the patient to be diagnosed.
[0095] The attention weight distribution of the last observation path in the historical momentum-attention map is used as the momentum-inertia guiding term for the patient to be diagnosed.
[0096] By fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector, a corrected attention distribution map of the patient to be diagnosed is obtained.
[0097] The modified attention distribution map is normalized to obtain the first momentum attention map of the patient to be diagnosed.
[0098] The process of fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector to obtain the corrected attention distribution map of the patient to be diagnosed includes:
[0099] The momentum decay coefficient of the historical momentum attention map is determined based on the semantic complexity in the initial inference text sequence.
[0100] The momentum inertia guidance term and the momentum decay coefficient are weighted and fused to obtain the historical attention contribution value of the patient to be diagnosed.
[0101] The initial attention distribution map is multiplied element-wise with the complementary coefficient of the momentum decay coefficient to obtain the current attention contribution value of the patient to be diagnosed.
[0102] The historical attention contribution value and the current attention contribution value are superimposed, and the weights of the attention focus of the high-dimensional visual embedding vector are reorganized to obtain the attention fusion result of the patient to be diagnosed.
[0103] Based on the attention fusion results, the initial attention distribution map is recalibrated to obtain the corrected attention distribution map of the patient to be diagnosed.
[0104] The initial inference text sequence is split into independent words and phrases according to Chinese semantic segmentation rules. A pre-stored vocabulary of anatomical structures of the upper abdominal biliary system is called. This vocabulary contains the names of all anatomical structures related to the diagnosis of cholangiocarcinoma, such as the hilar bile duct, intrahepatic bile duct, and common bile duct. The segmented words and phrases are matched precisely with this vocabulary one by one. Only the completely matching content is retained as anatomical structure pointing information, and the unmatched content is directly removed. Then, the core words representing the specific observation area are extracted from the successfully matched anatomical structure pointing information. These core words are the descriptive words of the region of interest of the anatomical structure pointing information.
[0105] A weight allocation matrix is constructed that corresponds one-to-one with the feature dimensions of the high-dimensional visual embedding vector. Each matrix unit uniquely corresponds to a feature dimension of the high-dimensional visual embedding vector. The descriptive words of the region of interest are mapped to the corresponding anatomical region feature dimensions of the biliary system in the high-dimensional visual embedding vector. A basic attention weight of 0.8 is assigned to the matrix units corresponding to the feature dimensions of the successfully mapped feature dimensions, and a basic attention weight of 0.2 is assigned to the matrix units corresponding to the feature dimensions of other anatomical regions that are not mapped to the descriptive words of the region of interest. The weight allocation matrix formed after the weight allocation is completed is the initial attention distribution map of the patient to be diagnosed.
[0106] The full attention weight distribution data generated during the last attention focus update operation is directly retrieved from the historical momentum attention map. This data is a matrix data that matches the feature dimensions of the high-dimensional visual embedding vector and contains the attention weight values corresponding to each feature dimension during the last observation. The full attention weight distribution data is completely extracted and its original values and matrix structure are retained. The extracted data is the momentum inertia guidance term of the patient to be diagnosed.
[0107] The semantic complexity of the initial inference text sequence is quantified by counting the total number of words in the initial inference text sequence. It is preset that a word count of 5 or less is low semantic complexity, a word count of more than 5 and less than or equal to 10 is medium semantic complexity, and a word count of more than 10 is high semantic complexity. At the same time, a momentum decay coefficient of 0.7 is preset for low semantic complexity, 0.5 for medium semantic complexity, and 0.3 for high semantic complexity. The corresponding semantic complexity level is matched according to the total number of words counted, and then a preset fixed value is extracted according to the level. This value is the momentum decay coefficient of the historical momentum attention map.
[0108] The matrix weight data of the momentum-inertia guidance term and the determined momentum decay coefficient value are retrieved. The weight value of each unit in the momentum-inertia guidance term matrix is multiplied independently with the momentum decay coefficient. The result of multiplying each unit is calculated one by one. The values of all units after the operation and the original matrix structure are retained. The new matrix data obtained after the operation is the historical attention contribution value of the patient to be diagnosed.
[0109] The complementary coefficient of the momentum decay coefficient is calculated by numerical subtraction. The calculation method is to subtract the determined momentum decay coefficient from the numerical value 1. The difference is the complementary coefficient. The matrix weight data of the initial attention distribution map is retrieved. The weight value of each unit in the initial attention distribution map matrix is multiplied independently with the complementary coefficient. The result of multiplying each unit is calculated one by one. The values of all units after the operation and the original matrix structure are retained. The new matrix data obtained after the operation is the current attention contribution value of the patient to be diagnosed.
[0110] The matrix data of historical attention contribution values and current attention contribution values are retrieved. The two matrices have the same structure and dimensions. The weight values of the corresponding units in the two matrices are added one by one to calculate the sum of the values at each corresponding position to obtain the fused basic weight matrix. Based on this basic weight matrix, the attention focus weights of each feature dimension in the high-dimensional visual embedding vector are reorganized and combined. The fused weight values of each feature dimension are retained and the correspondence between the feature dimension and the matrix unit is not changed. The reorganized matrix weight data is the attention fusion result of the patient to be diagnosed.
[0111] The matrix data of the attention fusion result and the initial attention distribution map are retrieved. The weight value of each unit in the attention fusion result matrix is used as the new calibration benchmark value to replace the original weight value of the corresponding unit in the initial attention distribution map. The value replacement of all units is completed one by one while retaining the original matrix structure and feature dimension correspondence of the initial attention distribution map. Only the weight value of each unit is updated. The new matrix data obtained after all units have been replaced is the corrected attention distribution map of the patient to be diagnosed.
[0112] First, calculate the sum of the weight values of all units in the corrected attention distribution map matrix. Then, divide the weight value of each unit in the matrix by the sum and calculate the quotient value of each unit one by one. The quotient value of each unit is within the range of 0 to 1. Keep the normalized values of all units after calculation and the original matrix structure. The new matrix data obtained after the normalization operation is the first momentum attention map of the patient to be diagnosed.
[0113] The beneficial effects are that this implementation process establishes a standardized and reproducible operational system for generating the first momentum-attention map. It achieves accurate extraction of descriptive words for the region of interest through a pre-set biliary anatomical vocabulary, providing a clear anatomical basis for attention weight allocation. The fixed-value basic attention weight allocation method ensures that the initial attention distribution map is free from subjective bias. Complete extraction of the weight distribution of historical momentum-attention maps guarantees the effective transmission of historical observation path information. The semantic complexity judgment standard based on the number of words and the pre-set attenuation coefficient value ensures that the determination of the momentum attenuation coefficient has a quantitative basis and a unique result. The generation of historical and current attention contribution values follows unified operational rules, ensuring that the calculation results are accurate and verifiable. The generation of attention fusion results is also consistent. By adding corresponding positional values and recombining weights, the comprehensive guidance of historical and current attention is fully preserved. Feature recalibration is completed through one-to-one position replacement, allowing the corrected attention distribution map to accurately reflect the fused attention guidance. Normalization processing ensures that the weight values of the first momentum attention map are in a unified range of 0 to 1, facilitating subsequent attention weight analysis and key image region selection. Each step of the entire process has clear execution standards and fixed processing methods, with no black-box operations or subjective adjustments, ensuring the consistency and accuracy of the first momentum attention map generation results. It can effectively guide the update of the momentum mechanism of the attention focus of the high-dimensional visual embedding vector, providing a reliable basis for the accurate extraction of subsequent key image patch sets.
[0114] S3. Based on the first momentum attention map, extract the first key image patch set from the upper abdominal non-enhanced CT image, and perform image-text interleaving fusion of the first key image patch set with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed.
[0115] In this embodiment of the invention, the step of extracting the first key image patch set from the upper abdominal non-enhanced CT image based on the first momentum attention map, and performing image-text interleaving fusion of the first key image patch set with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed includes:
[0116] Obtain the attention weight distribution of each image region in the first momentum attention map;
[0117] Based on the attention weight distribution, the importance of all image regions in the non-contrast CT image of the upper abdomen is ranked to obtain the candidate key regions of the patient to be diagnosed.
[0118] Local image patches corresponding to the candidate key regions are cropped from the non-contrast CT images of the upper abdomen, and the size of the local image patches is standardized to obtain the first key image patch set of the patient to be diagnosed.
[0119] The first set of key image patches is inserted as visual evidence markers into the designated positions of the initial inference text sequence to construct the image-text interleaved input sequence of the patient to be diagnosed.
[0120] The image-text interleaved input sequence is input into the vision-language large model for text decoding to obtain the first inference text of the patient to be diagnosed. The calculation formula for the first inference text is as follows:
[0121] ;
[0122] In the formula, The first inference text, For the aforementioned large visual-language model, For global visual embedding vectors, This is a selection function that sorts and selects the top K regions based on attention weights. The original tomographic scan data, This is the first momentum attention map. A historical reasoning text sequence, These are the preset model weight parameters.
[0123] The matrix weight data stored in the first momentum-attention map is extracted. Each cell of this matrix forms a unique one-to-one mapping with each independent image region of the upper abdominal non-contrast CT image. The weight values of all cells in the matrix are directly retrieved, while the positional mapping information between each weight value and the corresponding image region is fully preserved. The specific weight value corresponding to each image region is fully recorded, thus obtaining the attention weight distribution of each image region in the first momentum-attention map. For each image region of the upper abdominal non-contrast CT image, its specific weight value in the attention weight distribution is matched one by one. All image regions are arranged sequentially according to a fixed order from high to low weight values. The top 20 image regions in the sorting result are selected as the candidate range of key regions. The spatial location information and corresponding weight values of these 20 image regions are fully preserved, thus obtaining the candidate key regions of the patient to be diagnosed.
[0124] Based on the precise spatial boundary coordinates of each region in the candidate key regions, each candidate key region is precisely cropped in the grayscale tomographic image sequence of the upper abdominal non-contrast CT image, separating the independent local image blocks corresponding to each candidate key region. During the cropping process, the original pixel arrangement, grayscale value distribution and morphological features of the local image blocks are completely preserved. Then, all the cropped local image blocks are uniformly adjusted to a fixed size of 512×512 pixels. During the adjustment, a proportional scaling method is used to maintain the original shape of the image blocks and avoid morphological distortion caused by stretching or compression. After scaling, a set of local image blocks with uniform size is obtained, thus obtaining the first key image patch set of the patient to be diagnosed. In the initial inference text sequence, the fixed insertion position of the visual evidence marker is the first character position after the sequence start character. The first key image patch set is embedded into this preset position as a whole visual evidence marker. During the embedding process, the original character order and text structure of the initial inference text sequence are completely preserved. At the same time, a one-to-one association mapping is established between the visual features of the first key image patch set and the semantic information of the initial inference text sequence. The combined sequence containing the visual evidence marker and the inference text formed after the embedding is completed is the image-text interleaved input sequence of the patient to be diagnosed.
[0125] The image-text interleaved input sequence is converted into a combined input stream of visual feature tensors and text character streams that the model can recognize, according to the input specifications of the pre-trained visual-language large model. This combined input stream is then completely fed into the text decoding layer of the visual-language large model. The text decoding layer first extracts the imaging morphological features corresponding to the visual evidence markers in the input stream, and then combines them with the basic semantic information of the initial inference text sequence. Following the clinical semantic logic of cholangiocarcinoma CT image diagnosis, it decodes and generates text word by word, finally outputting a continuous and complete text that conforms to the clinical diagnostic expression specifications. This text is the first inference text for the patient to be diagnosed. The parameters in its generation calculation formula all come from the above implementation process and related steps such as visual feature extraction and model initialization. The global visual embedding vector is a high-dimensional visual embedding vector formed by adjusting the window width and window level, encoding mapping, and feature aggregation of the original tomographic scan data. It is the core representation result of the model on the visual features of CT images. The pre-trained visual-language large model adapted to abdominal CT image analysis is the core carrier for realizing the deep integration of CT image visual features and diagnostic text reasoning. A selection function was set based on clinical practice and model training results for the diagnosis of cholangiocarcinoma using CT imaging, which is used to select key image regions according to attention weights. The raw tomographic scan data collected by medical imaging equipment is the original data foundation for the entire diagnostic reasoning and analysis; The first momentum attention map obtained after the momentum mechanism-guided update is the core basis for the model to locate key lesion areas in CT images; The initial inference text sequence forms the basic text context for the model to generate its first inference text. The preset model weight parameters, obtained after general pre-training and specific fine-tuning of the visual-language large model, provide core support for model feature mapping and text generation. The first reasoning text is the core calculation output of the formula, providing the basic text content for subsequent iterative operations.
[0126] The calculation formula for the first inference text is the core logical expression of the first inference text generation process. It systematically integrates the core products of steps such as visual feature extraction, momentum attention map updating, key image region selection, and image-text interleaving. It clarifies the input relationships of each product in the first inference text generation process, and forms a precise correspondence with the implementation content of the first key image patch set extraction, image-text interleaving input sequence construction, and first inference text generation. The formula presents the process of image-text interleaving to generate inference text through standardized computational logic, giving the deep fusion of visual image features and text inference sequences a fixed execution basis. This ensures that the generation of the first inference text is always based on the key visual features of CT images, while the semantic basis of historical inference text sequences guarantees the semantic coherence of the inference text. Furthermore, the fixed computational logic of the formula enables the standardization and reproducibility of the generation process of the first reasoning text, avoiding subjective bias in the generation process. It also lays a unified computational logic foundation for subsequent iterations to perform attention focus updates and image-text interleaving operations to generate a complete reasoning thought chain. It effectively connects the core technical links of momentum attention mechanism and image-text interleaving reasoning, and promotes the orderly development of the entire cholangiocarcinoma CT image prediction process.
[0127] The beneficial effects include establishing a standardized and reproducible operational system for the extraction of the first key image patch set and the generation of the first inference text. The extraction of attention weight distribution combined with a one-to-one mapping relationship ensures accurate correspondence with image regions. The selection of a fixed number of candidate key regions ensures effective capture of key imaging features while avoiding redundant extraction of invalid regions. The proportionally scaled and standardized size gives the first key image patch set a unified feature dimension, which can accurately match the input requirements of the visual-language large model. The preset embedding method makes the structure of the image-text interleaved input sequence standardized and achieves accurate association between visual features and text sequences. The text decoding process combines visual evidence and diagnostic semantic logic, giving the first inference text solid imaging evidence support. There is no subjective adjustment or black-box operation. Each step has clear execution standards and processing basis, ensuring the accuracy of the first key image patch set extraction and the rationality of the first inference text generation. It effectively realizes the deep integration of visual features and text reasoning, laying a reliable foundation for subsequent iterations to generate reasoning thought chains.
[0128] S4. Iterate through S2 to S3 until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated.
[0129] In this embodiment of the invention, the iterative execution of S2 to S3 until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated includes:
[0130] The first inference text is used as the text generation result of the current iteration round, and the first momentum attention map is used as the attention update result of the current iteration round;
[0131] The attention update result of the current iteration round is used as the historical momentum attention graph when the attention focus update is performed next, and the text generation result of the current iteration round is appended to the historical inference text sequence as the historical inference text sequence when the attention focus update is performed next.
[0132] Based on the historical momentum attention map and the historical sequence of inference text, the attention focus update operation is performed again to obtain the next iteration momentum attention map of the patient to be diagnosed.
[0133] Based on the momentum attention map of the next iteration round, the key image patch set extraction and image-text interleaving fusion operation are performed again to obtain the next iteration round inference text of the patient to be diagnosed.
[0134] The momentum attention map generated in each iteration is used as the historical momentum attention map when performing attention focus update in subsequent iterations. The inference text generated in each iteration is appended to the historical inference text sequence. The attention focus update operation and the key image patch set extraction and image-text interleaving operation are repeatedly performed until the visual-language large model output terminator is detected and the repetition stops.
[0135] The reasoning texts generated in all repeated rounds are concatenated sequentially to obtain the reasoning thought chain of the patient to be diagnosed. The preliminary diagnostic conclusion of the patient to be diagnosed is then parsed from the reasoning text output in the last repeated round. The calculation formula for the preliminary diagnostic conclusion is as follows:
[0136] , ;
[0137] In the formula, For the aforementioned reasoning chain, This is the preliminary diagnostic conclusion. The model, given the global visual embedding vector The first key image patch set and the aforementioned historical reasoning text sequence Under the condition of predicting and generating the current word unit The conditional probability, To find the reasoning path and conclusion combination that maximizes the joint probability, The total length of the generated sequence is defined as follows: The first inference text is completely copied to the text sub-region of the current iteration round result storage area, preserving its original character order, text structure, and semantic expression. Simultaneously, the matrix weight data and image region mapping information of the first momentum attention map are completely copied to the attention sub-region of this storage area. This designates the first inference text as the text generation result of the current iteration round and the first momentum attention map as the attention update result of the current iteration round. The original stored data of the historical momentum attention map is cleared. The matrix weight data and image region mapping information of the current iteration round attention update result are completely copied to the storage address of the historical momentum attention map, making it the historical momentum attention map for the next attention focus update operation. Simultaneously, the complete text content of the current iteration round text generation result is concatenated as a character stream to the end of the historical inference text sequence storage, preserving the temporal relationship between historical and new content, making it the historical inference text sequence for the next attention focus update operation.
[0138] Following the complete operational process of generating the initial momentum attention map, based on the updated historical momentum attention map and the historical sequence of inference text, the anatomical structure pointing information in the historical sequence of inference text is first parsed and the descriptive words of the region of interest are extracted. Then, based on the descriptive words of the region of interest, the high-dimensional visual embedding vector is preliminarily weighted to obtain the initial attention distribution map. Next, the momentum inertia guiding term is extracted and fused with the initial attention distribution map. After smoothing correction, the corrected attention distribution map is obtained. Finally, the corrected attention distribution map is normalized. The matrix weight data obtained after processing is the momentum attention map for the next iteration of the patient to be diagnosed. Following the complete operational process of extracting the key image patch set and completing the image-text interleaving fusion, based on the momentum attention map of the next iteration, the attention weight distribution of each image region is first obtained. Then, the importance of all image regions in the CT image is ranked to obtain candidate key regions. After cropping and size standardization, a new key image patch set is obtained. This patch set is then inserted as a visual evidence marker into a specified position in the reasoning text history sequence to construct an image-text interleaving input sequence. This input sequence is then input into the visual-language large model to complete text decoding. The complete text output after decoding, which conforms to the diagnostic description specification of cholangiocarcinoma CT images, is the reasoning text for the next iteration of the patient to be diagnosed.
[0139] The momentum attention map generated in each iteration is used as the historical momentum attention map for subsequent iterations to perform the attention focus update operation. The inference text generated in each iteration is concatenated in character stream form to the latest storage end of the historical inference text sequence. The attention focus update operation, key image patch set extraction and image-text interleaving operation are continuously repeated. After each generation of inference text, the output content of the visual-language large model is detected character by character. When the output content contains the exclusive terminator set in the model pre-training stage, all repeated operations are stopped immediately. Following the sequential execution of the iterations, all reasoning texts generated in repeated rounds are continuously concatenated. During the concatenation process, the original content, semantic expression, and generation order of each reasoning text are fully preserved without any content deletion or order adjustment. The complete text sequence formed after concatenation is the reasoning thought chain of the patient to be diagnosed. The reasoning text output from the last repeated round is retrieved, and a pre-stored vocabulary library for cholangiocarcinoma CT image diagnostic conclusions is called. The reasoning text is segmented word by word according to Chinese semantic segmentation rules and accurately matched with the vocabulary library one by one. The core diagnostic terms that are successfully matched are extracted and combined with the corresponding imaging feature descriptions in the reasoning text to form a complete diagnostic conclusion statement. This statement is the preliminary diagnostic conclusion for the patient to be diagnosed.
[0140] The reasoning chain generated during the above iteration process Preliminary Diagnostic Conclusion The relevant calculation parameters all originate from this process and previous steps such as visual feature extraction and key image patch set extraction. The reasoning chain is a complete text sequence containing morphological description and logical analysis formed by continuously splicing the reasoning text generated in each round according to the order of generation during the model's iterative execution of attention focus update, key image patch set extraction and image-text interleaving. It is a complete semantic record of the model's step-by-step completion of the CT image diagnosis reasoning for cholangiocarcinoma. The preliminary diagnostic conclusion is a complete diagnostic conclusion formed by matching and extracting core diagnostic terms from the pre-stored vocabulary library of cholangiocarcinoma CT image diagnostic conclusions from the inference text output by the last iteration of the model, and combining it with the description of imaging features. It is the final diagnostic result obtained by the model after completing iterative inference. The conditional probability is derived from the pre-trained visual-language large model after fine-tuning with CT imaging diagnostic corpus and image data for cholangiocarcinoma. Given relevant visual features and text sequences, the probability value of the model predicting the generation of the current word is calculated by the clinical diagnostic semantic reasoning logic of the model's text decoding layer. It reflects the degree of matching between the generated current word and visual evidence and historical text semantics. The global visual embedding vector is a high-dimensional visual embedding vector formed by adjusting the window width and window level, encoding and mapping, and aggregating the original tomographic scan data of the non-enhanced CT images of the upper abdomen of the patient to be diagnosed. It is the core representation result of the model on the visual features of CT images. The first set of key image patches is derived from the set of local image patches extracted based on the first momentum attention map. It is the core visual evidence of lesions extracted by the model from CT images. The initial state is the historical reasoning text sequence after the implantation of the exclusive start symbol. In subsequent iterations, the reasoning text of each round is appended to the end of the sequence in chronological order, providing continuous semantic context support for the model to generate the current word. To find the reasoning path and conclusion combination that maximizes the joint probability, the operation is set based on the clinical logic of CT image diagnosis of cholangiocarcinoma and the optimal verification results of model training. After calculating the joint probability of all possible combinations and comparing the numerical values, the combination with the largest value is selected as the output result. The total length of the generated sequence is the total number of reasoning text words generated in all iterations from the initial reasoning text sequence until the specific terminator is detected. It is a quantitative indicator representing the overall length of the reasoning thought chain.
[0141] This calculation formula is the core logical expression of the reasoning chain and the generation process of the preliminary diagnostic conclusion. It systematically integrates the products of all core steps mentioned above, including the construction of high-dimensional visual embedding vectors, the first momentum attention map update, the first key image patch set extraction, iterative attention update, and image-text fusion reasoning. It clarifies the core role and correlation of each product in the reasoning chain and the generation process of the preliminary diagnostic conclusion, and precisely corresponds to the implementation content of the iterative generation of the reasoning chain and the preliminary diagnostic conclusion. The multiplication operation in the formula accumulates the conditional probabilities of the generated lexical units in each iteration, reflecting the overall semantic rationality and visual evidence matching degree of the entire reasoning path. The operation ensures that the selected reasoning chain and preliminary diagnostic conclusion are the best match among all possible combinations with global visual features, key image patch sets, and historical reasoning text semantics. This provides rigorous probabilistic logic support for the generation of the reasoning chain and ensures that the preliminary diagnostic conclusion is based on a complete reasoning path rather than a single word judgment. Simultaneously, the fixed computational logic of the formula standardizes and reproducibly enables the generation of the reasoning chain and preliminary diagnostic conclusion, avoiding subjective biases and effectively connecting the various technical stages of iterative reasoning. This guarantees the coherence and progression of the diagnostic reasoning process, providing the preliminary diagnostic conclusion with solid CT image visual evidence and progressive logical reasoning support. This provides a complete and standardized reasoning basis and diagnostic result for subsequent logical consistency verification.
[0142] The beneficial effects include establishing a standardized and reproducible iterative operation system for generating reasoning chains and preliminary diagnostic conclusions. Each iteration is based on the previous attention update results and reasoning text, ensuring the continuity and progression of attention focus updates and diagnostic reasoning. The repeated execution process remains completely consistent with the initial generation, ensuring the uniformity and accuracy of results during iteration. Using a dedicated terminator from the visual-language large model as the criterion for stopping iterations ensures the integrity of the reasoning chain while avoiding invalid iterations and the generation of redundant reasoning content. The method of piecing together reasoning text according to the iteration sequence fully preserves the logical hierarchy of the entire diagnostic reasoning, giving the reasoning chain a clear clinical reasoning framework. The preliminary diagnostic conclusion is analyzed based on a dedicated diagnostic conclusion vocabulary, ensuring the standardization and professionalism of the diagnostic conclusion. The entire iterative process is free from subjective adjustments and black-box operations; each step has clear execution standards and processing basis, providing the preliminary diagnostic conclusion with progressively advancing imaging evidence and logical reasoning support, effectively improving the credibility and clinical reference value of the diagnostic conclusion.
[0143] S5. Based on the preset plain text logic verifier, the correlation between the reasoning chain and the preliminary diagnostic conclusion is evaluated to obtain the logical consistency score of the patient to be diagnosed.
[0144] In this embodiment of the invention, the step of evaluating the correlation between the reasoning chain and the preliminary diagnostic conclusion based on a preset plain text logic verifier to obtain the logical consistency score of the patient to be diagnosed includes:
[0145] The morphological descriptor sequence and the logical analysis statement sequence in the reasoning chain are concatenated to form the verification text of the patient to be diagnosed.
[0146] The preliminary diagnostic conclusion and the text to be verified are input together into a preset plain text logic verifier;
[0147] The radiographic evidence set of the patient to be diagnosed is constructed by parsing the radiographic feature descriptive words in the text to be verified using the plain text logic verifier.
[0148] The diagnostic reference words in the preliminary diagnostic conclusion are parsed by the plain text logic validator to extract the diagnostic conclusion reference of the patient to be diagnosed.
[0149] The number of imaging features described in the set of imaging evidence that support the diagnostic conclusion is counted.
[0150] The number of descriptive words of the imaging features is correlated with the imaging evidence set to obtain the logical consistency score of the patient to be diagnosed.
[0151] The formula for calculating the logical consistency score is as follows:
[0152] ;
[0153] In the formula, The logical consistency score is given. The total number of descriptive terms for imaging features in the aforementioned set of imaging evidence. For the set of radiological evidence, the first A descriptive term for an imaging feature, The diagnostic conclusion extracted from the preliminary diagnostic conclusion points to... This is a pre-defined semantic similarity calculation function within the plain text logic validator.
[0154] All words characterizing the morphological features of non-enhanced CT images of the upper abdomen are extracted from the reasoning thought chain and arranged sequentially according to their appearance in the reasoning thought chain to form a morphological descriptive word sequence. Then, all statements used for the logical analysis of bile duct cancer imaging diagnosis are extracted from the reasoning thought chain and arranged sequentially according to their reasoning time in the reasoning thought chain to form a logical analysis statement sequence. The morphological descriptive word sequence is concatenated at the beginning of the logical analysis statement sequence in a continuous string of characters. During the concatenation process, the original content and inherent time sequence of the two sequences are completely preserved without any content deletion or order adjustment. The resulting complete continuous text is the verification text for the patient to be diagnosed.
[0155] The text to be verified is completely imported into the text input port of the preset plain text logic validator in the form of a character stream. Then, the preliminary diagnosis conclusion is completely imported into the diagnosis conclusion input port of the plain text logic validator in the form of a character stream. The two input contents are independently stored in different dedicated storage areas inside the plain text logic validator, and the original text structure and character order are preserved throughout the process without any content conversion. This completes the operation of inputting the preliminary diagnosis conclusion and the text to be verified into the preset plain text logic validator.
[0156] The plain text logic validator calls upon its internally stored lexicon of CT imaging features related to cholangiocarcinoma. This lexicon, constructed according to clinical imaging diagnostic guidelines for cholangiocarcinoma, includes all imaging features related to the diagnosis of cholangiocarcinoma, such as bile duct dilatation, bile duct wall thickening, intraductal soft tissue nodules, and bile duct stenosis. The text to be validated is segmented word-by-word according to Chinese semantic segmentation rules. Each segmented word is precisely matched against this imaging feature lexicon. Successfully matched words are the imaging feature descriptions in the text to be validated. All successfully matched imaging feature descriptions are then organized in chronological order of their appearance in the text to form a set of unique words. This set constitutes the imaging evidence set for the patient to be diagnosed. The plain text logic validator performs a word-by-word, unique counting of this imaging evidence set, and the specific numerical value obtained is used in the logical consistency score calculation formula. Simultaneously, each word in the image evidence set is assigned a unique numerical identifier starting from 1, based on the original occurrence sequence of the image feature descriptive words in the text to be verified. The specific radiographic descriptive terms corresponding to each numerical identifier are in the formula. .
[0157] The plain text logic validator calls an internally stored lexicon of cholangiocarcinoma diagnostic terms. This lexicon is constructed based on the clinical diagnostic criteria for cholangiocarcinoma and includes all diagnostic terms related to cholangiocarcinoma, such as intrahepatic cholangiocarcinoma, extrahepatic cholangiocarcinoma, common bile duct cancer, and benign lesions of the bile duct. The preliminary diagnostic conclusion is segmented word-by-word according to Chinese semantic segmentation rules. All segmented words are then precisely matched against this diagnostic terminology. Successfully matched words are the diagnostic terms in the preliminary diagnostic conclusion. If a single successfully matched term is used, it is directly designated as the diagnostic conclusion. If multiple terms are matched, the term with the highest priority is selected according to the validator's preset diagnostic terminology priority rules. The final determined term is the diagnostic conclusion for the patient to be diagnosed. This diagnostic conclusion is the term in the logical consistency score calculation formula. The core computation layer of the plain text logic validator also incorporates a pre-defined semantic similarity calculation function. This function is a dedicated semantic matching algorithm based on the semantic association features of bile duct cancer imaging diagnosis. It uses a corpus of imaging signs and diagnostic conclusions of bile duct cancer clinical diagnosis as training data. The calculation process involves converting the two input words into semantic feature vectors, calculating the spatial similarity between the two semantic feature vectors, and finally converting the result into a value between 0 and 1. The higher the value, the stronger the semantic association between the two words.
[0158] The plain text logic validator calls an internally stored association table between diagnostic conclusions and supporting imaging features. This association table is constructed based on the clinical diagnostic guidelines and imaging diagnostic standards for cholangiocarcinoma, clearly defining all supporting imaging feature descriptors corresponding to each diagnostic conclusion. Each imaging feature descriptor in the imaging evidence set is compared one by one with the currently extracted diagnostic conclusion in this association table, determining whether each descriptor is a supporting imaging feature descriptor for that diagnostic conclusion. After the word-by-word judgment is completed, all the imaging feature descriptors determined to be supportive are counted one by one, and the specific value obtained in the final statistics is recorded. This value is the number of imaging feature descriptors in the imaging evidence set that support the diagnostic conclusion.
[0159] The plain text logic validator performs correlation quantification calculations based on the logic consistency score calculation formula, first by... Calculate each of the radiological evidence sets separately and The semantic similarity is calculated, and then the results of all semantic similarity calculations are summed. Finally, the sum is divided by the total number of sign descriptive words in the imaging evidence set. The mean value is obtained, which is the logical consistency score of the patient to be diagnosed. This value is a quantitative indicator representing the degree of logical connection between the reasoning chain and the preliminary diagnostic conclusion. The numerical range is... The output range remains consistent between 0 and 1, which intuitively reflects the logical matching degree between the two.
[0160] The beneficial effects are that this implementation process establishes a standardized and reproducible operational system for generating logical consistency scores. Based on the dedicated vocabulary and association table constructed from clinical diagnostic guidelines and standards for cholangiocarcinoma, the parsing of imaging feature descriptive terms and the extraction of diagnostic indications have clear clinical basis. The orderly splicing of the text to be verified completely preserves the morphological features and logical analysis information in the reasoning chain, ensuring the integrity of the verification basis. The non-repeating imaging evidence set makes the evidence foundation for verification more standardized. Accurate statistics of the number of supporting features are achieved through one-by-one comparison of the association table, providing a solid numerical foundation for formula calculation. The core calculation formula is deeply integrated with the verification process of the plain text logic verifier, using the imaging evidence set and diagnostic conclusions as the calculation basis, and through the verification of all... and The method of calculating semantic similarity and then summing and averaging comprehensively reflects the overall support of all sign descriptive words for the diagnostic conclusion, avoiding the assessment bias caused by a single sign. At the same time, it transforms abstract logical relationships into standardized values between 0 and 1, giving the logical consistency score a clear quantitative judgment standard. Its calculation logic is also consistent with the principle of comprehensively judging the diagnostic conclusion based on multiple imaging signs in the clinical diagnosis of cholangiocarcinoma, ensuring the clinical rationality and scientific nature of the quantitative assessment results. The generated logical consistency score can objectively and accurately reflect the degree of matching between the preliminary diagnostic conclusion and the reasoning evidence, providing a scientific and reliable logical verification basis for the generation of subsequent diagnostic reports, and effectively improving the credibility and logical rationality of the CT image prediction diagnosis conclusion of cholangiocarcinoma.
[0161] S6. Compile the logical consistency scores to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.
[0162] In this embodiment of the invention, the step of compiling a report on the logical consistency score to obtain a target intelligent assisted diagnostic report for the patient to be diagnosed includes:
[0163] The logical consistency score is compared with a preset logical consistency threshold to obtain the verification status of the patient to be diagnosed.
[0164] When the verification status is "verification passed", extract the key morphological description statements in the reasoning chain;
[0165] The logical consistency score, the preliminary diagnostic conclusion, and the key morphological description statements are structured and filled in to obtain the initial draft of the diagnostic report for the patient to be diagnosed.
[0166] Extract the CT image layer identifiers corresponding to the first key image patch set from the reasoning chain, and embed the CT image layer identifiers as visual evidence annotations into the designated position of the initial draft of the diagnostic report;
[0167] The initial draft of the diagnostic report is formatted to obtain the target intelligent assisted diagnostic report for the patient to be diagnosed.
[0168] The logical consistency score of the patient to be diagnosed is retrieved and compared with the preset logical consistency threshold in the plain text logical verifier. This threshold is set to a fixed value of 0.7 based on clinical practice data of cholangiocarcinoma CT image diagnosis and the pre-training verification results of the visual-language large model. The specific value of the logical consistency score is directly compared with this fixed threshold. If the logical consistency score is greater than or equal to 0.7, the verification status is determined to be passed. If the logical consistency score is less than 0.7, the verification status is determined to be failed. The clear result obtained after the determination is the verification status of the patient to be diagnosed.
[0169] Assuming the verification status is deemed successful, the pre-stored core morphological feature vocabulary of cholangiocarcinoma CT images is retrieved. This vocabulary contains all core morphological descriptive terms related to the diagnosis of cholangiocarcinoma. All statements in the reasoning chain are semantically broken down sentence by sentence. Each sentence is matched with the core morphological feature vocabulary. If a sentence contains any word from the vocabulary, it is identified as a key morphological descriptive statement. The statement is extracted sentence by sentence according to its original appearance order in the reasoning chain, preserving the original content and semantic expression of the sentence without any deletions or modifications.
[0170] The pre-set intelligent assisted diagnosis structured report template for cholangiocarcinoma CT images is retrieved. This template is a fixed-format text template, including dedicated fields for logical consistency score, preliminary diagnosis conclusion, and key morphological description statements. The three fields are independently partitioned in the template and have a fixed arrangement order. The logical consistency score of the patient to be diagnosed is accurately filled into the corresponding field, the preliminary diagnosis conclusion is completely filled into its dedicated field, and the extracted key morphological description statements are filled into the corresponding fields in the original extraction sequence. All content retains the original expression without modification. The text generated after all the content is accurately filled in is the first draft of the diagnosis report for the patient to be diagnosed.
[0171] Extract the CT image tomographic layer number corresponding to the first key image patch set from the text content of the reasoning chain. This number is the inherent layer identifier of the original tomographic scan of the upper abdominal non-enhanced CT image. Directly extract this number as the CT image layer identifier. In the structured template of the initial draft of the diagnostic report, the preset fixed insertion position for visual evidence annotation is the blank area immediately below the key morphological description statement field. Convert the extracted CT image layer identifier according to the fixed format of "Visual evidence corresponds to CT layer: X". Embed the converted content completely into the preset fixed insertion position without changing the original layout and expression of the initial draft of the report.
[0172] Following the standard format requirements for clinical medical diagnostic reports, the initial draft of the diagnostic report underwent a uniform formatting process. The specific formatting rules were as follows: the font was uniformly set to SimSun, size 12; the line spacing was uniformly set to 1.5 times; all paragraphs were formatted with a first-line indent of 2 characters; key morphological descriptions were marked with square bullet points; the logical consistency score and preliminary diagnostic conclusion were centered; and visual evidence annotations were left-aligned with a blank line above the previous line. The text format, paragraph layout, and symbol annotations of the initial draft diagnostic report were adjusted one by one according to these fixed rules. After the adjustments were completed, the position and accuracy of each content were checked again. Once it was confirmed that there were no deviations, the resulting standardized text became the target intelligent assisted diagnostic report for the patient to be diagnosed.
[0173] The beneficial effects of this implementation process are that it establishes a standardized and reproducible operating system for generating intelligent assisted diagnostic reports. The fixed logical consistency threshold set based on clinical and experimental data provides a clear and unified basis for determining the verification status, ensuring the rigor of the generated diagnostic reports. Key phrases are extracted through matching with a core morphological feature vocabulary, ensuring that the core content of the initial draft diagnostic report has a solid morphological basis in imaging. The pre-set structured report template makes the logical consistency score, preliminary diagnostic conclusions, and key phrase filling more standardized. Fixed-position embedding of CT image layer markers achieves precise correspondence between textual diagnostic evidence and visual image evidence, making the report's evidence support more complete. The formatting process following clinical standards ensures that the final generated intelligent assisted diagnostic report conforms to clinical medical usage guidelines. The entire process is standardized and the steps are clear. The generated target report is complete, logically clear, and formatted according to standards. It retains the core information of cholangiocarcinoma CT image diagnosis while achieving standardized generation of intelligent assisted diagnostic reports, effectively improving the generation efficiency and clinical practicality of cholangiocarcinoma CT image prediction diagnostic reports, and providing standardized and reliable intelligent assisted support for clinical cholangiocarcinoma imaging diagnosis.
[0174] like Figure 2The diagram shown is a functional block diagram of a cholangiocarcinoma CT image prediction system based on momentum attention and large model validation, provided by an embodiment of the present invention.
[0175] The cholangiocarcinoma CT image prediction system 100 based on momentum attention and large model validation described in this invention can be installed in an electronic device. Depending on the functions implemented, the cholangiocarcinoma CT image prediction system 100 may include a visual encoding and initialization module 101, a momentum attention update module 102, a text-image interleaving reasoning module 103, an iterative generation module 104, a logic validation module 105, and a report generation module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0176] In this embodiment, the functions of each module / unit are as follows:
[0177] The visual encoding and initialization module 101 is used to extract visual features from the non-enhanced CT image data of the upper abdomen of the patient to be diagnosed based on the pre-trained visual-language large model, so as to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and initialize the historical momentum attention map and initial inference text sequence of the visual-language large model.
[0178] The momentum attention update module 102 is used to update the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence, so as to obtain the first momentum attention map of the patient to be diagnosed.
[0179] The image-text interleaving reasoning module 103 is used to extract the first key image patch set in the upper abdominal non-enhanced CT image based on the first momentum attention map, and to perform image-text interleaving fusion of the first key image patch set with the initial reasoning text sequence to obtain the first reasoning text of the patient to be diagnosed.
[0180] The iterative generation module 104 is used to iteratively execute the momentum attention update module to the graph-text interleaved reasoning module until the reasoning thought chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated.
[0181] The logic verification module 105 is used to evaluate the correlation between the reasoning chain and the preliminary diagnostic conclusion based on a preset plain text logic verifier, and obtain the logic consistency score of the patient to be diagnosed.
[0182] The report generation module 106 is used to compile the logical consistency score into a report to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.
[0183] like Figure 3 The figure shown is a comparison of model results for a cholangiocarcinoma CT image prediction system based on momentum attention and large model validation, provided by an embodiment of the present invention.
[0184] To verify the performance advantages of the proposed cholangiocarcinoma CT image prediction method based on momentum attention and large model validation, a comparative experiment was conducted in this embodiment. The experimental data came from 1000 cases of upper abdominal non-enhanced CT images (including 500 positive cases and 500 negative cases of cholangiocarcinoma). The comparison objects included a traditional non-inference model (ResNet-50 classifier) and an unimproved inference model (a large visual-language model without the introduction of momentum mechanism and logical validation). The receiver operating characteristic (ROC) curve and the area under the curve (AUC) were used as evaluation indicators. The experimental results are as follows: Figure 3 As shown.
[0185] Figure 3 This is a comparison of the ROC curves of the momentum attention and logic verification-based inference model with other models. In the figure, the orange curve represents the improved inference model described in this invention, the blue curve represents the traditional non-inference model, and the green curve represents the unimproved inference model. Figure 3 As can be seen, the AUC value of the improved inference model of this invention is as high as 0.975, while the AUC value of the traditional non-inference model is 0.941, and the AUC value of the unimproved inference model is only 0.831. These experimental results fully demonstrate that this invention, by introducing a momentum mechanism to achieve dynamic guidance and updating of the attention focus, effectively solves the problems of attention drift and discontinuous lesion tracking existing in the unimproved inference model. Simultaneously, combined with the secondary verification of the plain text logic checker, it eliminates logically inconsistent erroneous predictions, significantly reducing the risk of medical hallucinations. This allows the model to not only maintain high accuracy in the cholangiocarcinoma CT image prediction task but also possess superior stability and reliability. Its overall performance is significantly better than that of the traditional non-inference model and the unimproved inference model, further verifying the advanced nature and practicality of the technical solution of this invention.
[0186] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0187] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0188] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0189] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0190] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation, characterized in that, The method includes: S1. Based on the pre-trained visual-language large model, visual features are extracted from the upper abdominal non-enhanced CT image data of the patient to be diagnosed in order to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and the historical momentum attention map and initial inference text sequence of the visual-language large model are initialized. S2. Based on the historical momentum attention map and the initial inference text sequence, the attention focus of the high-dimensional visual embedding vector is updated using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed. S3. Based on the first momentum attention map, extract the first key image patch set from the upper abdominal non-enhanced CT image, and perform image-text interleaving fusion of the first key image patch set with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed. S4. Iterate through S2 to S3 until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated. S5. Based on the preset plain text logic verifier, the correlation between the reasoning chain and the preliminary diagnostic conclusion is evaluated to obtain the logical consistency score of the patient to be diagnosed. S6. Compile the logical consistency scores to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.
2. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 1, characterized in that, The pre-trained visual-language large model extracts visual features from the unenhanced CT images of the upper abdomen of the patient to be diagnosed, constructing a high-dimensional visual embedding vector for the patient, and initializes the historical momentum attention map and initial inference text sequence of the visual-language large model, including: The raw tomographic scan data of the upper abdomen of the patient to be diagnosed is obtained, and the window width and window level are adjusted on the raw tomographic scan data to obtain the grayscale tomographic image sequence of the patient to be diagnosed. The grayscale tomographic image sequence is input into a pre-trained visual-language large model for encoding and mapping to obtain the local visual feature vector of the patient to be diagnosed. Feature aggregation is performed on the local visual feature vectors to obtain the high-dimensional visual embedding vector of the patient to be diagnosed. A momentum attention buffer is constructed for the visual-language large model to record historical observation paths, and the momentum attention buffer is cleared to obtain the historical momentum attention map of the visual-language large model; Start symbols are implanted into the text generator of the visual-language big model to obtain the initial inference text sequence of the visual-language big model.
3. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 1, characterized in that, The step of updating the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence using a momentum mechanism to obtain the first momentum attention map of the patient to be diagnosed includes: The anatomical structure pointing information in the initial inference text sequence is parsed to extract the region of interest descriptive words of the anatomical structure pointing information; Based on the descriptive words of the region of interest, a preliminary attention weight allocation is performed on the high-dimensional visual embedding vector to obtain the initial attention distribution map of the patient to be diagnosed. The attention weight distribution of the last observation path in the historical momentum-attention map is used as the momentum-inertia guiding term for the patient to be diagnosed. By fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector, a corrected attention distribution map of the patient to be diagnosed is obtained. The modified attention distribution map is normalized to obtain the first momentum attention map of the patient to be diagnosed.
4. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 3, characterized in that, The process of fusing the initial attention distribution map with the momentum-inertia guidance term and smoothing the attention focus of the high-dimensional visual embedding vector to obtain the corrected attention distribution map of the patient to be diagnosed includes: The momentum decay coefficient of the historical momentum attention map is determined based on the semantic complexity in the initial inference text sequence. The momentum inertia guidance term and the momentum decay coefficient are weighted and fused to obtain the historical attention contribution value of the patient to be diagnosed. The initial attention distribution map is multiplied element-wise with the complementary coefficient of the momentum decay coefficient to obtain the current attention contribution value of the patient to be diagnosed. The historical attention contribution value and the current attention contribution value are superimposed, and the weights of the attention focus of the high-dimensional visual embedding vector are reorganized to obtain the attention fusion result of the patient to be diagnosed. Based on the attention fusion results, the initial attention distribution map is recalibrated to obtain the corrected attention distribution map of the patient to be diagnosed.
5. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 2, characterized in that, The first key image patch set is extracted from the upper abdominal non-contrast CT image based on the first momentum attention map, and the first key image patch set is interleaved with the initial inference text sequence to obtain the first inference text of the patient to be diagnosed, including: Obtain the attention weight distribution of each image region in the first momentum attention map; Based on the attention weight distribution, the importance of all image regions in the non-contrast CT image of the upper abdomen is ranked to obtain the candidate key regions of the patient to be diagnosed. Local image patches corresponding to the candidate key regions are cropped from the non-contrast CT images of the upper abdomen, and the size of the local image patches is standardized to obtain the first key image patch set of the patient to be diagnosed. The first set of key image patches is inserted as visual evidence markers into the designated positions of the initial inference text sequence to construct the image-text interleaved input sequence of the patient to be diagnosed. The image-text interleaved input sequence is input into the vision-language large model for text decoding to obtain the first inference text of the patient to be diagnosed. The calculation formula for the first inference text is as follows: ; In the formula, The first inference text, For the aforementioned large visual-language model, For global visual embedding vectors, This is a selection function that sorts and selects the top K regions based on attention weights. The original tomographic scan data, This is the first momentum attention map. A historical reasoning text sequence, These are the preset model weight parameters.
6. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 5, characterized in that, The iterations S2 to S3 are performed until the reasoning chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated, including: The first inference text is used as the text generation result of the current iteration round, and the first momentum attention map is used as the attention update result of the current iteration round; The attention update result of the current iteration round is used as the historical momentum attention graph when the attention focus update is performed next, and the text generation result of the current iteration round is appended to the historical inference text sequence as the historical inference text sequence when the attention focus update is performed next. Based on the historical momentum attention map and the historical sequence of inference text, the attention focus update operation is performed again to obtain the next iteration momentum attention map of the patient to be diagnosed. Based on the momentum attention map of the next iteration round, the key image patch set extraction and image-text interleaving fusion operation are performed again to obtain the next iteration round inference text of the patient to be diagnosed. The momentum attention map generated in each iteration is used as the historical momentum attention map when performing attention focus update in subsequent iterations. The inference text generated in each iteration is appended to the historical inference text sequence. The attention focus update operation and the key image patch set extraction and image-text interleaving operation are repeatedly performed until the visual-language large model output terminator is detected and the repetition stops. The reasoning texts generated in all repeated rounds are concatenated sequentially to obtain the reasoning thought chain of the patient to be diagnosed. The preliminary diagnostic conclusion of the patient to be diagnosed is then parsed from the reasoning text output in the last repeated round. The calculation formula for the preliminary diagnostic conclusion is as follows: , ; In the formula, For the aforementioned reasoning chain, This is the preliminary diagnostic conclusion. The model, given the global visual embedding vector The first key image patch set and the aforementioned historical reasoning text sequence Under the condition of predicting and generating the current word unit The conditional probability, To find the reasoning path and conclusion combination that maximizes the joint probability, This represents the total length of the generated sequence.
7. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 1, characterized in that, The process involves using a preset plain text logic checker to assess the correlation between the reasoning chain and the preliminary diagnostic conclusion, thereby obtaining a logical consistency score for the patient to be diagnosed. This includes: The morphological descriptor sequence and the logical analysis statement sequence in the reasoning chain are concatenated to form the verification text of the patient to be diagnosed. The preliminary diagnostic conclusion and the text to be verified are input together into a preset plain text logic verifier; The radiographic evidence set of the patient to be diagnosed is constructed by parsing the radiographic feature descriptive words in the text to be verified using the plain text logic verifier. The diagnostic reference words in the preliminary diagnostic conclusion are parsed by the plain text logic validator to extract the diagnostic conclusion reference of the patient to be diagnosed. The number of imaging features described in the set of imaging evidence that support the diagnostic conclusion is counted. The number of descriptive words of the imaging features is correlated with the imaging evidence set to obtain the logical consistency score of the patient to be diagnosed.
8. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 7, characterized in that, The formula for calculating the logical consistency score is as follows: ; In the formula, The logical consistency score is given. The total number of descriptive terms for imaging features in the aforementioned set of imaging evidence. For the set of radiological evidence, the first A descriptive term for an imaging feature, The diagnostic conclusion extracted from the preliminary diagnostic conclusion points to... This is a pre-defined semantic similarity calculation function within the plain text logic validator.
9. The method for predicting cholangiocarcinoma CT images based on momentum attention and large model validation as described in claim 1, characterized in that, The process of compiling reports based on the logical consistency scores to obtain the target intelligent assisted diagnostic report for the patient to be diagnosed includes: The logical consistency score is compared with a preset logical consistency threshold to obtain the verification status of the patient to be diagnosed. When the verification status is "verification passed", extract the key morphological description statements in the reasoning chain; The logical consistency score, the preliminary diagnostic conclusion, and the key morphological description statements are structured and filled in to obtain the initial draft of the diagnostic report for the patient to be diagnosed. Extract the CT image layer identifiers corresponding to the first key image patch set from the reasoning chain, and embed the CT image layer identifiers as visual evidence annotations into the designated position of the initial draft of the diagnostic report; The initial draft of the diagnostic report is formatted to obtain the target intelligent assisted diagnostic report for the patient to be diagnosed.
10. A CT image prediction system for cholangiocarcinoma based on momentum attention and large model validation, characterized in that, The system for implementing the cholangiocarcinoma CT image prediction method based on momentum attention and large model validation as described in claim 1 includes: The visual encoding and initialization module is used to extract visual features from the non-enhanced CT image data of the upper abdomen of the patient to be diagnosed based on the pre-trained visual-language large model, so as to construct the high-dimensional visual embedding vector of the patient to be diagnosed, and initialize the historical momentum attention map and initial inference text sequence of the visual-language large model. The momentum attention update module is used to update the attention focus of the high-dimensional visual embedding vector based on the historical momentum attention map and the initial inference text sequence, so as to obtain the first momentum attention map of the patient to be diagnosed. The image-text interleaving reasoning module is used to extract the first key image patch set from the upper abdominal non-enhanced CT image based on the first momentum attention map, and to perform image-text interleaving fusion of the first key image patch set with the initial reasoning text sequence to obtain the first reasoning text of the patient to be diagnosed. The iterative generation module is used to iteratively execute the momentum attention update module to the graph-text interleaved reasoning module until the reasoning thought chain and preliminary diagnostic conclusion of the patient to be diagnosed are generated. The logic verification module is used to evaluate the correlation between the reasoning chain and the preliminary diagnostic conclusion based on a preset plain text logic verifier, and obtain the logic consistency score of the patient to be diagnosed. The report generation module is used to compile the logical consistency score into a report to obtain the target intelligent assisted diagnosis report for the patient to be diagnosed.