Image analysis methods, devices, media, and products based on global spatial perception and multi-step self-verifying inference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-14
AI Technical Summary
一方面,DCLDs很罕见,样本获取困难、标注成本高,端到端监督学习模型容易出现过拟合,跨中心泛化能力不足;另一方面,目前基于三维体数据的模型计算代价高,对病灶空间分布的表达能力有限
本申请提供了一种基于全局空间感知和多步自验证推理的影像分析方法、设备、介质及产品,通过获取目标患者的胸部CT体数据和临床元数据;对胸部CT体数据分别进行肺区域分割和囊状病灶分割,并按照预设规则进行融合,得到解剖标签图;基于解剖标签图进行全局空间感知,并结合临床元数据基于预设多模态大模型,确定初始诊断后验概率。该过程能够在显著降低多模态大模型输入上下文长度的同时,保留了DCLDs鉴别诊断所必需的空间分布信息,克服了直接输入完整CT体数据所导致的上下文爆炸问题。基于初始诊断后验概率进行多步自验证推理,以检索代表性二维证据切片并逐轮更新诊断后验概率。本申请通过引入多步自验证推理机制,使得在每一轮观察到新的二维证据切片后都对当前诊断结论进行支持或修正,从而降低由临床先验偏置或局部证据不充分所带来的诊断幻觉,提升诊断稳定性。根据更新后的诊断后验概率和预设置信度阈值输出分析结果,以提高临床应用安全性。由此,本申请可提升辅助诊断的准确性和稳定性。
Smart Images

Figure CN122575674A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing, and in particular to an image analysis method, device, medium, and product based on global spatial perception and multi-step self-verifying reasoning. Background Technology
[0002] In recent years, deep learning methods have made some progress in lung image analysis, but they still have significant limitations in the context of DCLDs (Diffuse Cystic Lung Diseases). On the one hand, DCLDs are rare, making sample acquisition difficult and annotation costs high. End-to-end supervised learning models are prone to overfitting and have insufficient cross-center generalization ability. On the other hand, current models based on three-dimensional volume data are computationally expensive and have limited ability to represent the spatial distribution of lesions.
[0003] Multimodal large models possess strong zero-shot and few-shot reasoning capabilities, offering new possibilities for the diagnosis of complex and rare diseases. However, current multimodal large models still face the problem of limited context length when processing chest CT volume data. A complete CT scan often contains hundreds of slices, and directly inputting them into the model can lead to context explosion. Simply sparsely sampling a few slices can easily result in the loss of information about the overall spatial distribution of lesions, leading to problems such as unstable reasoning, severe diagnostic illusions, and a lack of interpretability in the conclusions.
[0004] Therefore, there is an urgent need to propose a DCLDs-assisted diagnostic method that can reduce the complexity of 3D CT input without sacrificing spatial distribution information, and guide multimodal large models to perform step-by-step verification reasoning around key evidence, so as to improve its accuracy, stability and interpretability. Summary of the Invention
[0005] The purpose of this application is to provide an image analysis method, device, medium, and product based on global spatial perception and multi-step self-verifying reasoning, which can improve the accuracy and stability of assisted diagnosis.
[0006] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides an image analysis method based on global spatial perception and multi-step self-verifying reasoning, which is used to assist in the diagnosis of DCLDs; the method includes: Acquire chest CT scan data and clinical metadata of the target patient; The chest CT scan data are segmented into lung regions and cystic lesions, and then fused according to preset rules to obtain an anatomical label map. Global spatial perception is performed based on the anatomical label map, and the initial posterior probability of diagnosis is determined by combining the clinical metadata with a preset multimodal large model. Multi-step self-verification reasoning is performed based on the initial diagnostic posterior probability to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round. The analysis results are output based on the updated diagnostic posterior probability and the preset reliability threshold; the analysis results include auxiliary diagnostic results, evidence images and corresponding explanatory information.
[0007] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image analysis method based on global spatial awareness and multi-step self-verification reasoning described above.
[0008] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image analysis method based on global spatial awareness and multi-step self-verification reasoning described above.
[0009] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the image analysis method based on global spatial awareness and multi-step self-verification reasoning described above.
[0010] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides an image analysis method, device, medium, and product based on global spatial awareness and multi-step self-verifying reasoning. It acquires chest CT scan data and clinical metadata from a target patient; segments the chest CT scan data into lung regions and cystic lesions, and fuses them according to preset rules to obtain an anatomical label map; performs global spatial awareness based on the anatomical label map, and combines it with clinical metadata to determine the initial diagnostic posterior probability based on a preset multimodal large model. This process significantly reduces the input context length of the multimodal large model while retaining the spatial distribution information necessary for differential diagnosis of DCLDs, overcoming the context explosion problem caused by directly inputting complete CT scan data. Multi-step self-verifying reasoning is performed based on the initial diagnostic posterior probability to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round. By introducing a multi-step self-verifying reasoning mechanism, this application supports or corrects the current diagnostic conclusion after observing new two-dimensional evidence slices in each round, thereby reducing diagnostic illusions caused by clinical prior bias or insufficient local evidence and improving diagnostic stability. Analysis results are output based on the updated diagnostic posterior probability and a preset reliability threshold to improve the safety of clinical applications. Therefore, this application can improve the accuracy and stability of auxiliary diagnosis. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 The flowchart shows an image analysis method based on global spatial awareness and multi-step self-verifying reasoning. Figure 2 This is a flowchart illustrating the specific operation of an image analysis method based on global spatial perception and multi-step self-verification reasoning. Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] DCLDs include multiple subtypes, which differ significantly in their pathological mechanisms. However, they can all present as multiple pulmonary cysts on chest CT, and the cysts are highly similar in morphology, distribution range, and cyst wall thickness, making differential diagnosis difficult.
[0015] Currently, the imaging diagnosis of DCLDs typically relies on radiologists to simultaneously synthesize information at two scales: firstly, the three-dimensional spatial distribution characteristics of cysts throughout the lungs, such as whether the cysts are predominantly located in the upper or lower lungs, whether they are near the subpleural or mediastinal regions, and whether they are diffusely distributed; secondly, the fine-grained morphological characteristics on local sections, such as the size, shape, wall thickness, and relationship of the cysts to the bronchovascular bundles. Relying solely on a single two-dimensional section is insufficient to fully reflect the spatial distribution of the lesions, while relying solely on a rough three-dimensional representation is insufficient for verifying local morphology. Therefore, accurate diagnosis of DCLDs requires both a comprehensive understanding of the global spatial distribution and accurate local morphological discrimination.
[0016] The purpose of this application is to address the difficulties in CT differential diagnosis of diffuse cystic lung diseases (DCLDs), the inability of current multimodal large models to directly process complete three-dimensional volume data, and the tendency to generate diagnostic hallucinations. By fusing the segmentation results of the lung and cystic lesions and constructing a compressed three-dimensional spatial representation, the application first provides global distribution perception for the multimodal large model. Then, based on the initial diagnostic hypothesis, it actively retrieves representative two-dimensional evidence slices and performs multiple rounds of local morphological verification and diagnostic probability updates, thereby achieving high accuracy, strong interpretability, and low hallucination rate in the auxiliary diagnosis of DCLDs.
[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] In one exemplary embodiment, an image analysis method based on global spatial awareness and multi-step self-verifying reasoning is provided, which is used to assist in the diagnosis of DCLDs.
[0019] This application takes chest CT data and clinical metadata as input. First, it uses a lung segmentation model and a cystic lesion segmentation model to obtain anatomical label maps. Then, it compresses the spatial distribution of three-dimensional lesions into a three-dimensional rendering map with a layer index, which is used by a multimodal large model to form an initial diagnostic posterior. Then, based on the current diagnostic hypothesis, it automatically selects the axial slice with the most discriminative value for local evidence verification, and updates the posterior probability of the disease category round by round until the confidence level reaches a threshold or an uncertain result is output. Finally, it outputs the disease category, confidence level, and corresponding evidence chain.
[0020] like Figure 1 As shown, the method includes: Step 100: Obtain chest CT data and clinical metadata of the target patient.
[0021] This includes acquiring the target patient's chest CT scan data and clinical metadata, specifically including: Acquire the initial chest CT volume data and corresponding clinical metadata of the target patient; the initial chest CT volume data retains the slice order information in the DICOM (Digital Imaging and Communications in Medicine) sequence.
[0022] The initial chest CT volume data is preprocessed to obtain chest CT volume data; the preprocessing includes volume data resampling, window width and window level normalization, slice order correction and size adjustment.
[0023] Step 200: Perform lung region segmentation and cystic lesion segmentation on the chest CT data, and fuse them according to preset rules to obtain an anatomical label map.
[0024] Specifically, the chest CT scan data were segmented into lung regions and cystic lesions, and then fused according to preset rules to obtain anatomical label maps, including: Input chest CT volume data into the lung segmentation model to obtain a lung region mask. Furthermore, chest CT scan data is input into the cystic lesion segmentation model to obtain a cystic lesion mask. Among them, the lung segmentation model is trained based on chest CT volume data with known lung region masks; the cystic lesion segmentation model is trained based on chest CT volume data with known cystic lesion masks; the lung segmentation model adopts a three-dimensional medical image segmentation model based on convolutional neural network, Transformer network or a combination of convolutional neural network and Transformer network; the cystic lesion segmentation model adopts the nnU-Net v2 architecture.
[0025] The lung region mask and the cystic lesion mask are fused according to preset rules to generate an anatomical label map. The mathematical expression corresponding to the anatomical label image is: .
[0026] in, This is a tag-level fusion operation.
[0027] Step 300: Perform global spatial perception based on the anatomical label map, and combine clinical metadata with a preset multimodal large model to determine the initial posterior probability of diagnosis.
[0028] This includes global spatial perception based on anatomical label maps, and determining the initial posterior probability of diagnosis based on a pre-set multimodal large model in conjunction with clinical metadata, including: The Marching Cubes algorithm was used to extract the surface of the anatomical label map, generating surface meshes for the lung region and cystic lesions, respectively.
[0029] A 3D visualization method based on mesh simplification and off-screen rendering is used to render the surface mesh of the lung region as transparent or wireframe, and the surface mesh of cystic lesions as solid, thereby obtaining a 3D visualization image that compressively expresses the distribution characteristics of cysts throughout the lung. .
[0030] 3D visualization images Together with clinical metadata, the data is input into a pre-defined multimodal large model to obtain the initial posterior probability of diagnosis for a pre-defined set of DCLD categories. ; Initial diagnosis posterior probability The corresponding mathematical expression is: .
[0031] in, For pre-defined multimodal large models; This is clinical metadata.
[0032] Step 400: Perform multi-step self-verification inference based on the initial diagnostic posterior probability to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round.
[0033] The process involves multi-step self-verification inference based on the initial diagnostic posterior probability to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round, including: In 3D visualization images An index axis is added that is consistent with the initial chest CT volume data slice order to establish a mapping relationship between the three-dimensional spatial position of the three-dimensional visualization image and the two-dimensional slice retrieval.
[0034] Based on the initial diagnostic posterior probability and mapping relationship, multi-step self-verifying inference is performed to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round; in any round: Based on the diagnostic posterior probability of the current round and the accumulated reasoning history, a slice retrieval strategy oriented towards diagnostic hypotheses is generated.
[0035] Based on the layer index set and the mapping relationship, extract the corresponding two-dimensional evidence slice set. Among them, the layer index set It is obtained from the output of a multimodal large model; layer index set Each level index in the table is used to locate axis slices with predefined identification value; level index set The size does not exceed the preset slice budget .
[0036] The reasoning history, the two-dimensional evidence slice set, and the slice retrieval strategy for diagnostic hypotheses are input into a pre-defined multimodal large model for multi-step self-verifying reasoning analysis, and the updated diagnostic posterior probability for the current round is output. .
[0037] Determine whether the iteration stopping conditions are met; the iteration stopping conditions include the diagnostic confidence level being no less than the preset confidence threshold and the current round reaching the preset maximum number of rounds; the diagnostic confidence level is determined based on the updated diagnostic posterior probability.
[0038] If the diagnostic confidence is not less than the preset confidence threshold, the iteration is terminated, and the updated diagnostic posterior probability in the current round is determined as the updated diagnostic posterior probability.
[0039] If the diagnostic confidence is less than the preset confidence threshold and the current round has not reached the preset maximum number of rounds, then the updated diagnostic posterior probability of the current round is determined as the diagnostic posterior probability of the next round, and the next round is updated to the current round, and the process returns to the step of "generating a slice retrieval strategy for diagnostic hypotheses based on the diagnostic posterior probability of the current round and the accumulated reasoning history".
[0040] If the current round reaches the preset maximum number of rounds, but still does not meet the preset confidence threshold, then output an uncertain result and stop iterating.
[0041] The mathematical expression for diagnostic confidence is: .
[0042] in, For diagnostic confidence; This represents the updated diagnostic posterior probability. For DCLDs, it is a set of categories; This refers to the category number of DCLDs in the category set of DCLDs.
[0043] Step 500: Output the analysis results based on the updated diagnostic posterior probability and the preset reliability threshold. The analysis results include auxiliary diagnostic results, evidence images, and corresponding explanatory information. Evidence images include three-dimensional visualization images corresponding to the auxiliary diagnostic results and two-dimensional evidence slices.
[0044] If the diagnostic confidence level is not less than a preset confidence threshold, the auxiliary diagnostic result is determined based on the updated diagnostic posterior probability; the auxiliary diagnostic result The corresponding expression is: .
[0045] Alternatively: In cases where the output is uncertain, review the evidence based on a chain of evidence that performs multi-step self-verifying reasoning.
[0046] Specifically, the method mentioned in this application includes the following steps: Step 1: Obtain chest CT scan data and patient clinical metadata.
[0047] Step 1.1: Collect chest CT data of the patient and retain the slice order information in the DICOM sequence.
[0048] Step 1.2: Obtain clinical metadata corresponding to the chest CT data. Clinical metadata includes age, gender, smoking history, pneumothorax history, family history, kidney disease history, and other information related to the differential diagnosis of DCLDs.
[0049] Step 1.3: Preprocess the chest CT volume data, including volume data resampling, window width and window level normalization, slice order correction, and size adjustment to match the requirements of subsequent segmentation models and large model inputs.
[0050] Step 2: Segment the lung regions and cystic lesions in the chest CT data and fuse them to obtain an anatomical label map.
[0051] Step 2.1: Input the preprocessed chest CT data into the lung segmentation model to obtain the lung region mask. .
[0052] Step 2.2: Input the preprocessed chest CT data into the cystic lesion segmentation model to obtain the cystic lesion mask. .
[0053] Step 2.3: Fuse the lung region mask and the cystic lesion mask according to preset rules to generate an anatomical label map. Its mathematical representation is: .
[0054] in, This indicates a label-level fusion operation used to simultaneously represent the spatial location of lung boundaries and cystic lesions in a unified coordinate system.
[0055] The lung segmentation model in step 2 can be a 3D medical image segmentation model based on convolutional neural networks, Transformer networks, or a combination thereof; the cystic lesion segmentation model can be a 3D segmentation model trained on DCLDs data. Preferably, the cystic lesion segmentation model adopts the nnU-Net v2 architecture, and the lung segmentation model adopts an automatic segmentation model capable of outputting a mask of the lung parenchyma region.
[0056] Step 3: Perform global spatial perception based on the anatomical label map, generate a three-dimensional spatial compressed representation, and obtain the initial diagnostic posterior probability.
[0057] Step 3.1: Perform surface extraction on the anatomical label map obtained in Step 2 to generate surface meshes for the lung region and cystic lesions, respectively.
[0058] Step 3.2: Render the surface mesh of the lung region as transparent or wireframe, and render the surface mesh of the cystic lesions as solid to obtain a three-dimensional visualization image that compresses and expresses the distribution characteristics of cysts throughout the lung. .
[0059] Step 3.3: Add a layer index axis to the 3D visualization image that is consistent with the original CT layer order, so as to establish a mapping relationship between the 3D spatial location and the subsequent 2D slice retrieval.
[0060] Step 3.4: Visualize the 3D image Together with clinical metadata, the data is input into a multimodal large model to obtain the initial posterior probability of diagnosis for a predefined set of DCLD categories. It is represented as: .
[0061] in, Represents a multimodal large model. This represents clinical metadata.
[0062] The surface extraction in step 3 can employ the Marching Cubes algorithm, and the rendering can utilize a 3D visualization method based on mesh simplification and off-screen rendering. Preferably, the lung surface is displayed with a semi-transparent green wireframe, and the surface of cystic lesions is displayed as a red solid to enhance the identification of cyst distribution.
[0063] Step 4: Perform multi-step self-verification inference based on the initial diagnostic posterior probability, retrieve representative two-dimensional evidence slices, and update the diagnostic posterior probability round by round.
[0064] Step 4.1: Generate a slice retrieval strategy based on the diagnostic posterior probability of the current round and the accumulated reasoning history.
[0065] Step 4.2: Output a sparse layer index set from the multimodal large model. Each level index is used to locate the most valuable axis slice, and the size of the set of level indexes does not exceed a preset slice budget. .
[0066] Step 4.3: Extract the corresponding two-dimensional evidence slice set from the original CT volume data according to the slice index set. .
[0067] Step 4.4: Input the current reasoning history, the set of two-dimensional evidence slices, and the current diagnostic hypothesis back into the multimodal large model. This model will perform self-validation analysis based on features such as cyst size, cyst wall thickness, morphology, distribution area, and relationship with the bronchovascular bundle, and output the updated posterior diagnostic probability. That is, the first The updated diagnostic posterior probability corresponding to each round.
[0068] Step 4.5: Calculate the diagnostic confidence level for the current round. : .
[0069] in, This is a set of DCLDs categories.
[0070] Step 4.6, if Greater than or equal to the preset reliability threshold If, then terminate the iteration; if Less than the preset confidence threshold And the number of iterations has not reached the maximum number of rounds. If the maximum number of rounds is reached, then continue with the next round of slice retrieval and self-verification reasoning; If the threshold is still not met, an uncertain result will be output.
[0071] The slide retrieval strategy in step 4 is adapted to the current diagnostic assumptions. For suspected BHD cases, slides from the lower lung, subpleural, and mediastinal regions should be prioritized; for suspected LAM cases, representative slides from diffusely distributed areas in both lungs should be prioritized; for suspected PLCH cases, slides from the dominant distribution areas in the upper lungs should be prioritized; and for suspected LIP cases, slides from areas with fewer cysts and clues of interstitial changes should be prioritized.
[0072] The reasoning history in step 4 includes at least: the initial diagnostic posterior probabilities of previous rounds, the retrieved slice index, the analyzed two-dimensional evidence slices, the corresponding morphological descriptions, and the updated posterior probabilities of each round.
[0073] Multimodal large models are not limited to a single manufacturer or a single architecture; they can be visual language foundational models with image input and text reasoning capabilities.
[0074] Step 5: Output the final diagnosis result, evidence image and corresponding interpretation based on the updated diagnostic posterior probability and the preset reliability threshold.
[0075] Step 5.1, if in the... During round iteration, satisfy Then output the final diagnosis category. : .
[0076] Step 5.2: Output the three-dimensional visualization image, two-dimensional evidence slices, diagnostic confidence score, and diagnostic interpretation generated by the multimodal large model corresponding to the final diagnosis.
[0077] Step 5.3: When the preset reliability threshold is not met, output an "uncertain" conclusion and the retrieved evidence chain for further review by the doctor.
[0078] The benefits of this application are: This application compresses the three-dimensional spatial information of lung and cystic lesions into a single three-dimensional visualization image, which significantly reduces the input context length of multimodal large models while retaining the spatial distribution information necessary for the differential diagnosis of DCLDs, thus overcoming the context explosion problem caused by directly inputting complete CT volume data.
[0079] This application does not simply perform uniform sampling of CT slices, but actively retrieves two-dimensional evidence slices with the highest discriminative value based on the current diagnostic hypothesis, enabling the model to conduct targeted local morphological analysis around key lesion areas, which significantly improves diagnostic efficiency and effectiveness.
[0080] This application introduces a multi-step self-verifying reasoning mechanism, which enables the model to support or revise the current diagnostic conclusion after observing new two-dimensional evidence in each round, thereby reducing diagnostic illusions caused by clinical prior bias or insufficient local evidence and improving diagnostic stability.
[0081] This application can output a traceable chain of evidence consisting of "global three-dimensional distribution image + local two-dimensional evidence slice + diagnostic interpretation + confidence level", which conforms to the clinical doctors' working habits of "first the whole, then the part, and then make a comprehensive judgment" and has strong interpretability and practical value.
[0082] This application sets a confidence threshold and a maximum number of iterations, which can output uncertain conclusions when there is insufficient evidence or obvious conflict, thus avoiding the model from forcibly giving incorrect diagnoses and improving the safety of clinical applications.
[0083] According to the embodiments of this application, high diagnostic accuracy and good cross-center generalization ability can be obtained on multi-center DCLDs data, indicating that this method is suitable for auxiliary diagnosis of complex rare diseases.
[0084] like Figure 2 As shown, this embodiment takes the four-category DCLDs diagnostic task of chest CT body data as an example. The category set is BHD, LAM, PLCH and LIP. The whole process includes five stages: data acquisition, mask segmentation and fusion, global spatial perception, multi-step self-verification reasoning and result output.
[0085] 1. Data acquisition.
[0086] 1.1. Acquire volumetric data of the patient's chest CT scan, preserving the DICOM slice order. Volumetric data typically contains hundreds of axial slices.
[0087] 1.2. Collect patient clinical metadata. This may include age, gender, smoking history, history of pneumothorax, family history of cancer, etc., among which this clinical information can serve as important auxiliary clues for initial inference in a multimodal large model.
[0088] 1.3. Preprocessing of CT volume data. As an example, the voxel spacing can be uniformly resampled, and the image grayscale can be normalized under lung window conditions to meet the input requirements of subsequent 3D segmentation models and multimodal large models.
[0089] 2. Mask segmentation and fusion.
[0090] 2.1. Input the preprocessed chest CT data into the cystic lesion segmentation model and output the cyst mask. In a specific case, the cystic lesion segmentation model can be trained and inferred using the 3D full-resolution configuration of nnU-Net v2.
[0091] 2.2. Input the same chest CT scan data into the lung segmentation model and output a lung region mask. The lung segmentation model can utilize existing automatic lung segmentation networks.
[0092] 2.3. Integrate the lung region mask and the cystic lesion mask into a unified label map. The lung region is used to provide organ boundaries and spatial references, while the cystic lesion region is used to highlight lesion distribution information.
[0093] 3. Global spatial awareness.
[0094] 3.1. Label-based graph Extract the 3D surface mesh of the lungs and cystic lesions. As an example, the Marching Cubes algorithm can be used to complete the surface reconstruction, and the rendering burden can be reduced by simplifying the mesh.
[0095] 3.2. Perform 3D rendering. As an example, the lung surface can be displayed as a semi-transparent wireframe, and cystic lesions can be displayed as solid structures, thus showing the overall distribution, aggregation trend, and size differences of cysts throughout the lung in a single image.
[0096] 3.3. Attach a vertical index axis next to the 3D rendered image that matches the original CT slice number to establish the correspondence between the 3D spatial location and the 2D axial slice.
[0097] 3.4. Input the above 3D rendering and clinical metadata into the multimodal large model to generate the initial posterior probability of diagnosis. In one specific embodiment, the multimodal large model can output a probability ranking of the four DCLDs categories and provide initial diagnostic hypotheses and their rationale.
[0098] 4. Multi-step self-verifying reasoning.
[0099] 4.1. Multimodal large models generate retrieval strategies based on the category with the highest posterior probability. For example, when the model initially suspects BHD, it will prioritize viewing representative axial sections of the lower lung, subpleural, or mediastinal regions to verify signs such as thin-walled, flat, or lenticulous cysts.
[0100] 4.2. Output the hierarchical index set according to the retrieval strategy. In one specific case, the slice budget is set to... That is, a maximum of 4 representative axial slices can be extracted in a single round.
[0101] 4.3. Extract the two-dimensional evidence slices corresponding to the index from the original CT body data, and send the slices and the current reasoning history into the multimodal large model.
[0102] 4.4. Multimodal large models perform local morphological verification on two-dimensional evidence slices, including but not limited to: whether the cyst wall is thin or thick, whether the cyst is round or irregular, whether the cyst is diffusely distributed, whether it is located in the subpleural region, and whether it is adjacent to the bronchovascular bundle, etc.
[0103] 4.5. The multimodal large model updates the posterior probabilities of each DCLD category based on the validation results, obtaining... And calculate the current confidence level. . Figure 2 In Indicates the first The updated diagnostic posterior probability corresponding to each round.
[0104] 4.6. In this embodiment, the maximum number of iteration rounds can be set. Confidence threshold .when The system outputs the final diagnosis; if the threshold is not met after two rounds, an uncertain result is output.
[0105] 5. Output results.
[0106] 5.1. Output the final diagnosis category.
[0107] 5.2. Output a 3D rendering, the retrieved 2D evidence slices, the changes in posterior probabilities in each round, and the explanatory text generated by the model, forming a traceable chain of diagnostic evidence.
[0108] Taking a patient with clinical information such as "55-year-old female, no history of smoking, history of pneumothorax, and family history of renal cancer" as an example, when using the basic multimodal large model directly, the model is easily influenced by clinical priors and tends to give a preliminary diagnosis of BHD.
[0109] This application first utilizes a global spatial perception module to render the spatial relationship between the patient's lungs and cystic lesions into a single 3D image, and combines this with clinical metadata to obtain an initial diagnostic posterior. Subsequently, in a multi-step self-validating inference phase, the model actively retrieves several representative axial slices based on the initial diagnostic hypothesis. Through analysis of local evidence, the model discovers numerous diffusely distributed, approximately round, thin-walled cysts in both lungs, whose distribution pattern is inconsistent with the typical lower lung and subpleural predominant distribution of BHD. Based on this, the model progressively revises the diagnostic hypothesis and ultimately updates the diagnosis to LAM (Lung Atrophy) with a high confidence level.
[0110] This embodiment illustrates that this application can not only output the final diagnostic result, but also provide a complete chain of evidence for "why the initial hypothesis was rejected and why the alternative diagnosis was supported", effectively reducing the illusion phenomenon of multimodal large models in the diagnosis of rare diseases.
[0111] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores image analysis data based on global spatial awareness and multi-step self-verifying inference. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an image analysis method based on global spatial awareness and multi-step self-verifying inference.
[0112] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0113] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0114] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0115] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0117] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0118] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logic devices, etc., and are not limited to these.
[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0120] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An image analysis method based on global spatial perception and multi-step self-verifying reasoning, characterized in that, The image analysis method based on global spatial perception and multi-step self-verifying reasoning is used to assist in the diagnosis of DCLDs; the method includes: Acquire chest CT scan data and clinical metadata of the target patient; The chest CT scan data are segmented into lung regions and cystic lesions, and then fused according to preset rules to obtain an anatomical label map. Global spatial perception is performed based on the anatomical label map, and the initial posterior probability of diagnosis is determined by combining the clinical metadata with a preset multimodal large model. Multi-step self-verification reasoning is performed based on the initial diagnostic posterior probability to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round. The analysis results are output based on the updated diagnostic posterior probability and the preset reliability threshold; the analysis results include auxiliary diagnostic results, evidence images and corresponding explanatory information.
2. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 1, characterized in that, Obtain chest CT scan data and clinical metadata of the target patient, specifically including: Acquire the initial chest CT volume data and corresponding clinical metadata of the target patient; the initial chest CT volume data retains the slice order information in the DICOM sequence; The initial chest CT volume data is preprocessed to obtain chest CT volume data; the preprocessing includes volume data resampling, window width and window level normalization, slice order correction and size adjustment.
3. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 1, characterized in that, The chest CT scan data are segmented into lung regions and cystic lesions, and then fused according to preset rules to obtain an anatomical label map, including: The chest CT scan data is input into the lung segmentation model to obtain a lung region mask. Furthermore, the chest CT data is input into the cystic lesion segmentation model to obtain a cystic lesion mask. The lung segmentation model is trained on chest CT scans with known lung region masks; the cystic lesion segmentation model is trained on chest CT scans with known cystic lesion masks; the lung segmentation model uses a 3D medical image segmentation model based on convolutional neural networks, Transformer networks, or a combination of convolutional neural networks and Transformer networks; the cystic lesion segmentation model uses the nnU-Net v2 architecture. The lung region mask and the cystic lesion mask are fused according to preset rules to generate an anatomical label map. The mathematical expression corresponding to the anatomical label image is: ; in, This is a tag-level fusion operation.
4. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 1, characterized in that, Global spatial awareness is performed based on the anatomical label map, and the initial posterior probability of diagnosis is determined based on a preset multimodal large model in conjunction with the clinical metadata, including: The Marching Cubes algorithm was used to extract the surface of the anatomical label map, generating surface meshes for the lung region and cystic lesions, respectively. A 3D visualization method based on mesh simplification and off-screen rendering is used to render the surface mesh of the lung region as transparent or wireframe, and the surface mesh of cystic lesions as solid, thereby obtaining a 3D visualization image that compressively expresses the distribution characteristics of cysts throughout the lung. ; The three-dimensional visualization image The clinical metadata is input together with a pre-defined multimodal large model to obtain the initial posterior probability of diagnosis for a pre-defined set of DCLDs categories. The initial diagnostic posterior probability The corresponding mathematical expression is: ; in, For pre-defined multimodal large models; This is clinical metadata.
5. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 4, characterized in that, Based on the initial diagnostic posterior probability, multi-step self-verifying inference is performed to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round, including: In 3D visualization images An index axis with the same layer order as the initial chest CT body data is added to establish a mapping relationship between the three-dimensional spatial position of the three-dimensional visualization image and the two-dimensional slice retrieval. Based on the initial diagnostic posterior probability and the mapping relationship, multi-step self-verification reasoning is performed to retrieve representative two-dimensional evidence slices and update the diagnostic posterior probability round by round. In any round: Based on the diagnostic posterior probability of the current round and the accumulated reasoning history, a slice retrieval strategy oriented towards diagnostic hypotheses is generated. Based on the layer index set and the mapping relationship, extract the corresponding two-dimensional evidence slice set. ; wherein, the layer index set It is obtained from the output of a multimodal large model; the layer index set Each layer index in the set is used to locate an axial slice with a predetermined identification value; the layer index set The size does not exceed the preset slice budget ; The reasoning history, the two-dimensional evidence slice set, and the slice retrieval strategy for diagnostic hypotheses are input into a pre-defined multimodal large model for multi-step self-verifying reasoning analysis, and the updated diagnostic posterior probability for the current round is output. ; Determine whether the iteration stopping condition is met; the iteration stopping condition includes a diagnostic confidence level not less than a preset confidence threshold and the current round reaching a preset maximum number of rounds; the diagnostic confidence level is determined based on the updated diagnostic posterior probability; If the diagnostic confidence is not less than the preset confidence threshold, the iteration is terminated, and the updated diagnostic posterior probability in the current round is determined as the updated diagnostic posterior probability. If the diagnostic confidence is less than the preset confidence threshold and the current round has not reached the preset maximum number of rounds, then the updated diagnostic posterior probability of the current round is determined as the diagnostic posterior probability of the next round, and the next round is updated to the current round, and the process returns to the step of "generating a slice retrieval strategy for diagnostic hypotheses based on the diagnostic posterior probability of the current round and the accumulated reasoning history". If the current round reaches the preset maximum number of rounds, but still does not meet the preset confidence threshold, then output an uncertain result and stop iterating.
6. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 5, characterized in that, The mathematical expression corresponding to the diagnostic confidence level is: ; in, For diagnostic confidence; This represents the updated diagnostic posterior probability. For DCLDs, it is a set of categories; This refers to the category number of DCLDs in the category set of DCLDs.
7. The image analysis method based on global spatial perception and multi-step self-verifying reasoning according to claim 6, characterized in that, The analysis results are output based on the updated diagnostic posterior probability and the preset reliability threshold, specifically including: If the diagnostic confidence level is not less than a preset confidence threshold, an auxiliary diagnostic result is determined based on the updated diagnostic posterior probability; the auxiliary diagnostic result... The corresponding expression is: ; or: In cases where the output is uncertain, a review is performed based on a chain of evidence that involves multi-step self-verifying reasoning. Evidence images include three-dimensional visualizations corresponding to the auxiliary diagnostic results and two-dimensional evidence slices.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the image analysis method based on global spatial awareness and multi-step self-verification reasoning as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the image analysis method based on global spatial awareness and multi-step self-verification reasoning as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the image analysis method based on global spatial awareness and multi-step self-verification reasoning as described in any one of claims 1-7.