Multi-source data-based multi-modal fusion processing method, case report generation method and system
By employing a multimodal fusion processing method for multi-source data, combined with frequency-domain guided micro-sign diffusion enhancement and vascular topology quantization, the problems of micro-sign fidelity and reporting reliability in low-dose CT images were solved, enabling efficient stratified diagnosis of invasive lung adenocarcinoma and reliable case report generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN AGRI UNIV
- Filing Date
- 2026-03-12
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for low-dose CT image processing suffer from problems such as high false positives, difficulty in stratifying the invasiveness of lung adenocarcinoma, difficulty in quantifying vascular bundles, and unreliable report generation, especially in terms of the fidelity of microsignatures and the traceability of evidence.
A multi-modal fusion processing method based on multi-source data is adopted. Through a frequency-domain guided micro-sign diffusion enhancement model and a three-dimensional wavelet high-frequency evidence channel, combined with vascular topology quantification indicators and structured knowledge graphs, deep fusion of imaging features and clinical data is carried out, and a closed-loop verification mechanism is used to generate case reports.
It improves the stability of distinguishing between MIA and IAC, enhances the sensitivity to IAC identification, ensures the reliability and credibility of report descriptions, reduces the risk of medical disputes, and improves the credibility and adaptability of the system.
Smart Images

Figure CN121839160B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to a multimodal fusion processing method based on multi-source data, a case report generation method and system. Background Technology
[0002] With the widespread use of low-dose spiral CT in lung cancer screening, the detection rate of pulmonary nodules has significantly improved, but clinical diagnosis still faces the following technical challenges:
[0003] (1) High false positives and excessive intervention: Many benign nodules are similar in shape and density to early malignant lesions. Relying solely on the appearance of images or traditional computer-aided detection can easily lead to misjudgment, resulting in unnecessary follow-up, puncture or surgery.
[0004] (2) Difficulty in stratifying invasive lung adenocarcinoma: For ground-glass opacities or subsolid nodules, the pathological progression usually follows a continuous spectrum of AAH / AIS / MIA / IAC. Accurate preoperative differentiation between MIA and IAC directly affects the surgical approach and prognostic assessment. However, IAC-related signs (such as microspiculation, superficial lobulation, subtle pleural traction, microvascular convergence, etc.) are often weak and fragmented, and are easily weakened during deep network sampling and feature smoothing, leading to unstable stratification.
[0005] (3) Key signs such as vascular convergence lack computability and quantification: Existing methods often use "vascular convergence / contraction" as a qualitative description or coarse-grained feature, lacking topological quantitative modeling of the vascular structure around the nodule, making it difficult to form a stable biomarker that can distinguish IAC.
[0006] (4) Report generation may be illusory and evidence may be untraceable: End-to-end text generation may result in the description of signs that are not present in the images or omission of key evidence; once the key evidence in the report is unreliable, it will directly mislead the conclusion of infiltration stratification, making it difficult for the system to be implemented in clinical practice.
[0007] Furthermore, existing technologies often employ common denoising algorithms (such as traditional filtering or ordinary GAN / DDPM) when processing low-dose or noisy CT images. However, the ordinary diffusion model (DDPM) tends to smooth high-frequency information during denoising, which can easily lead to the loss of key micro-signs for assessing the invasiveness of lung adenocarcinoma (such as fine spiculations and small vascular branches); or produce non-existent texture "illusions" when generating enhanced images. Simple preprocessing overlay cannot resolve the contradiction of "desensitizing while preserving the fidelity of subtle pathological features." Summary of the Invention
[0008] The purpose of this invention is to provide a multimodal fusion processing method, a case report generation method and system based on multi-source data, so as to solve the above-mentioned technical problems in chest CT image data processing and report generation.
[0009] This invention is achieved through the following technical solution:
[0010] The multimodal fusion processing method based on multi-source data includes the following steps:
[0011] Acquire chest CT imaging data and multidimensional clinical structured data;
[0012] A frequency-domain guided micro-sign diffusion enhancement model is constructed. Wavelet decomposition is performed on the ROI of lung nodules extracted from CT images to obtain high-frequency subband coefficients. The high-frequency subband coefficients are used to guide the inverse denoising process of the micro-sign diffusion enhancement model to obtain the high-frequency evidence tensor of micro-sign enhancement. Based on the high-frequency evidence tensor, an evidence scalar that can reflect the intensity of lung nodule micro-signs is obtained.
[0013] Based on the lung nodule mask extracted from CT images, vascular topology quantification indicators were obtained.
[0014] A one-dimensional vector is obtained by concatenating multidimensional clinical structured data, vascular topology quantification indicators, and evidence scalars. Deep visual features are obtained by fusing deep features extracted from CT images with high-frequency evidence tensors. The one-dimensional vector and deep visual features are used as inputs to the cross-attention mechanism of the multimodal fusion model to perform fusion processing on CT images, high-frequency evidence tensors, vascular topology quantification indicators, and multidimensional clinical structured data.
[0015] In some embodiments, the micro-feature diffusion enhancement model introduces a frequency domain cross-attention module in each downsampling and upsampling layer of the U-Net;
[0016] The high-frequency subband coefficient set is mapped to a high-dimensional embedding vector. The frequency domain cross-attention model uses the intermediate layer features of U-Net as the query and the high-dimensional embedding vector as the key and value to perform multi-head cross-attention operation to generate a high-frequency evidence tensor.
[0017] In some embodiments, the step of generating a high-frequency evidence tensor enhanced with micro-features includes:
[0018] Select multichannel volumetric data that includes at least the lung window channel and the mediastinal window channel in the 3D ROI of lung nodules;
[0019] Three-dimensional discrete wavelet decomposition is performed on multi-channel volume data to obtain a set of high-frequency sub-band coefficients;
[0020] Based on the lung nodule mask, a boundary shell region is constructed at its boundary, and an outer ring region is constructed based on the outward expansion / inward contraction parameters. The high-frequency sub-band coefficient set is masked and screened according to the boundary shell region and the outer ring region to obtain the mask high-frequency coefficient set.
[0021] The high-frequency coefficient set of the mask is stacked according to direction and scale, and then channel normalization and convolution mapping are performed to obtain a high-frequency evidence tensor with a uniform number of channels.
[0022] In some embodiments, the evidence scalar includes:
[0023] High-frequency energy This is used to measure the degree of irregularity at the edge of a nodule; it is expressed as:
[0024] ;
[0025] in, The boundary shell region of the nodule, p is the voxel coordinate, ||HF enhanced (p)|| represents the norm of the high-frequency evidence tensor value at voxel coordinate p; The larger the value, the rougher the edge and the higher the erosiveness;
[0026] Directional anisotropy Used to measure the directionality of burrs; represented as:
[0027] ;
[0028] in, Indicates standard deviation, These represent the energy values of wavelet subbands in different directions;
[0029] Boundary sharpness , represented as This is used to calculate the mean gradient magnitude of the high-frequency evidence tensor at the mask boundary, where... Represents the gradient operator. This represents the masking layer for lung nodules, and Mean() represents calculating the mean.
[0030] In some embodiments, the method further includes a step of preprocessing the acquired chest CT image data, including:
[0031] The steps for fusing deep features extracted from chest CT images with high-frequency evidence tensors to obtain deep visual features include:
[0032] High-frequency evidence tensor Performing a 1×1 convolution mapping yields deep features extracted from CT images. Dimensional alignment ;
[0033] Channel attention gating vectors are generated using the Sigmoid activation function, as follows: Where W() represents the learnable weight matrix, This indicates a feature concatenation operation. This represents the Sigmoid activation function;
[0034] Based on the following residual gating formula, and By fusing the data, deep visual features are obtained. ; indicates as:
[0035] ;
[0036] in, This indicates element-wise multiplication.
[0037] On the other hand, the present invention also provides a method for generating case reports, comprising the following steps:
[0038] Construct a three-layer structured knowledge graph of "nodules-signs-pathology";
[0039] A multimodal fusion processing method based on multi-source data is adopted to output structured quantitative data, including infiltration risk probability and symptom confidence score;
[0040] Based on a three-layer structured knowledge graph and structured quantitative data, a large language model is used to generate case reports.
[0041] In some embodiments, the three-layer structured knowledge graph includes a lesion layer, a sign layer, and a diagnostic layer. The lesion layer includes a unique ID for the lung nodule, a three-dimensional ROI mask, and its location in the world coordinate system. The sign layer includes deep semantic features describing the lung nodule and imaging features strongly correlated with the risk of malignancy. The diagnostic layer includes benign / malignant probability assessment conclusions, infiltrative grade prediction results, Lung-RADS classification suggestions, and follow-up strategies.
[0042] In some embodiments, the method further includes a step of verifying the consistency of the generated case reports, including:
[0043] Based on image features processed by the micro-feature diffusion enhancement model, an object detection algorithm is used to identify entities and relationships, and a visual scene graph is constructed. ,in Represents visual nodes. Indicates visual edge;
[0044] The description of the generated report is parsed to construct a text semantic graph. ,in Represents a text node. Represents text edges;
[0045] Calculation will Transform into The minimum required operating cost is used to calculate the logical consistency score. The overall consistency score is calculated based on the logical consistency score and the visual existence consistency score. The overall consistency score is then used to verify the consistency of the report.
[0046] In some embodiments, during consistency verification, when the corresponding attributes in the visual scene graph and the text semantic graph are determined to conflict, an error correction operation is triggered, and the visual scene graph is corrected. The corresponding attributes are output as mandatory constraints to the large language model, and a new report is generated.
[0047] On the other hand, the present invention also provides a pathology report generation system for performing the aforementioned case report generation method, comprising:
[0048] The data acquisition and processing module is used to acquire and process chest CT image data and multidimensional clinical structured data;
[0049] The ROI localization and segmentation module for lung nodules is used for multi-scale feature extraction from chest CT images.
[0050] The high-frequency evidence tensor construction and fusion module is used to generate high-frequency evidence tensors with enhanced micro-signs based on the extracted ROI of lung nodules, and to obtain evidence scalars based on the high-frequency evidence tensors.
[0051] The vascular topology quantification module is used to obtain vascular topology quantification indicators based on lung nodule masks.
[0052] The fusion reasoning module is used to concatenate multidimensional clinical structured data, vascular topology quantification indicators and evidence scalars to obtain a one-dimensional vector, fuse deep features and high-frequency evidence tensors to obtain deep visual features, and perform multi-head cross-attention operation on the one-dimensional vector and deep visual features to perform fusion processing on CT images, high-frequency evidence tensors, vascular topology quantification indicators and multidimensional clinical structured data.
[0053] The report generation module is used to generate case reports based on the structured quantitative data output by the three-layer structured knowledge graph and the fusion reasoning module;
[0054] The consistency verification and closed-loop error correction module is used to perform consistency verification and error correction on the generated case reports.
[0055] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0056] This invention deeply integrates tumor markers (CEA, CYFRA21-1, etc.), inflammatory markers (NLR), and radiomics features. It employs a three-dimensional wavelet high-frequency evidence channel to perform fidelity-preserving and explicit modeling of high-frequency micro-signs such as microspiculations and shallow lobulations, reducing the loss of invasive evidence caused by deep feature smoothing and improving the stability of distinguishing between MIA and IAC. Based on "frequency-domain guided diffusion enhancement," it uses wavelet high-frequency coefficients as the "anchor" of the diffusion enhancement model, achieving targeted enhancement of key MIA / IAC differentiation features such as microspiculations and shallow lobulations while removing CT noise. This enables deep fusion of multi-source data and significantly improves the upper limit of stratified diagnosis.
[0057] This invention models the topological structure of the microvessels in the annulus surrounding the nodule and introduces quantitative indicators such as tortuosity, fractal dimension, and convergence, transforming the qualitative "vascular clustering / convergence" into a calculable marker of invasive risk, thereby improving the sensitivity of IAC identification.
[0058] This invention employs a closed-loop verification and evidence anchoring mechanism to ensure that every sentence in the report is verifiable, achieving "fact-checking" of medical documents and significantly reducing the risk of medical disputes caused by reporting errors. Through structured slot report generation and cross-modal reverse verification, it reduces report illusion and achieves "what you see is what you get" evidence traceability; in cases of consistency conflicts, it triggers closed-loop error correction or uncertainty gating output, reducing the risk of stratified misleading and enhancing clinical credibility.
[0059] This invention transforms fuzzy consistency checks into precise topological comparisons through "visual-text" dual-map matching, fundamentally eliminating the logical illusion of large-scale model-generated reports. Employing structured reasoning paths and visualized image evidence chains enhances doctors' trust in the AI system and improves collaborative efficiency. Combined with a continuous learning mechanism based on doctor feedback, it enhances adaptability to long-tail cases and distribution drift, increasing the system's sustainable deployment value. Attached Figure Description
[0060] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 This is a flowchart of multimodal fusion processing and case report generation in an embodiment of the present invention.
[0062] Figure 2 This is a flowchart of consistency verification and closed-loop error correction in an embodiment of the present invention. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0064] Reference Figure 1 The present invention relates to a multimodal fusion processing method and a case report generation method based on multi-source data, comprising the following steps:
[0065] S1. Multi-source data acquisition and processing
[0066] Acquire chest CT imaging data and multidimensional structured clinical data of the patient, including:
[0067] High-dimensional image data: thin-slice high-resolution CT (HRCT) sequences in DICOM format, with a slice thickness ≤1.25mm;
[0068] Multidimensional clinical structured data, including:
[0069] Demographic characteristics: such as age, sex, height, weight, BMI;
[0070] High-risk factors include smoking index (number of cigarettes per year), duration of smoking cessation, and history of occupational exposure.
[0071] Past medical history: such as a detailed history of respiratory diseases (COPD, tuberculosis, asthma), and a history of tumors (primary site and treatment).
[0072] Laboratory test indicators include: carcinoembryonic antigen (CEA), carbohydrate antigens (CA125 / CA153 / CA724), cytokeratin 19 fragment (CYFRA21-1), neuron-specific enolase (NSE), progastrin-releasing peptide (ProGRP), squamous cell carcinoma antigen (SCC), and the neutrophil / lymphocyte ratio (NLR), which reflects the systemic inflammatory state.
[0073] Deep learning-based artifact detection (breathing artifacts, metal artifacts) and slice thickness verification are performed on images. Statistical missing value imputation (KNN) and outlier handling are performed on clinical data to ensure the integrity and reliability of the input data.
[0074] S2, Image Standardization and Multi-Window Multi-Scale Feature Enhancement Processing
[0075] S21. Spatial standardization: Resample the images to isotropic voxels (e.g., uniformly 1mm×1mm×1mm) to eliminate geometric distortion caused by differences in scanning parameters of different CT equipment.
[0076] S22. Multi-channel window mapping; generates channel data corresponding to different tissue characteristics, such as lung window (W:1500, L:-600) for observing the edges, internal texture and ground-glass components of nodules in the lung parenchyma; mediastinal window (W:350, L:40) for observing the enlargement of mediastinal lymph nodes, the density of calcifications and the true enhancement degree of solid components; bone window for assisting in determining whether there is rib destruction or bone metastasis.
[0077] S23. Multi-scale pyramid construction: Construct a feature pyramid network (FPN) to cover full-scale detection from tiny miliary nodules (<3mm) to giant masses (>30mm) to prevent missing tiny lesions.
[0078] S3. Frequency-domain guided micro-sign diffusion enhancement processing and construction of high-frequency evidence tensors
[0079] By designing a micro-feature diffusion enhancement model based on a frequency-conditional residual diffusion network, we can solve the problem of smooth loss of aggressive micro-features of lung adenocarcinoma (such as micro-spicules and shallow lobulation) caused by traditional noise reduction or feature extraction processes.
[0080] This frequency-domain conditional residual diffusion network does not generate entirely new images, but focuses on restoring and enhancing high-frequency details that are pathologically significant.
[0081] The frequency domain conditional residual diffusion network adopts an improved denoising network based on the U-Net architecture. ;include:
[0082] Input layer; receives two inputs, namely noisy latent space feature maps. (Originated from the original ROI after encoder compression and noise addition), and frequency domain conditional embedding vector. .
[0083] The frequency domain conditional embedding layer is used to perform the following operations:
[0084] Perform three-dimensional discrete wavelet decomposition (3D-DWT) on the three-dimensional ROI of the nodule to extract the set of high-frequency subband coefficients. ,in As a scale, For direction;
[0085] High-frequency subband coefficient set Mapped to a high-dimensional embedding vector using a frequency encoder. This is used to guide the micro-feature diffusion enhancement model; the frequency domain encoder here is a custom shallow neural network module (e.g., several fully connected or convolutional layers) used to map wavelet coefficients into vectors that match the dimensions of the main network.
[0086] Frequency cross-attention modules are introduced into each downsampling and upsampling layer of U-Net. These modules use the intermediate layer features of U-Net as the query, and... For the Key and Value, the forced frequency domain conditional residual diffusion network focuses on the real high-frequency texture locations during the denoising process, preventing the model from generating texture "illusions" out of thin air.
[0087] The frequency domain cross-attention module is based on standard cross-attention. The input is the intermediate layer features of U-Net as the query and the frequency domain conditional embedding vector. For the Key and Value, the output is an intermediate layer feature map that incorporates frequency domain information. This output will be passed to the next layer of U-Net.
[0088] The output layer outputs the high-frequency residuals predicted by the network for reconstructing high-fidelity pathological images from noisy images, instead of directly outputting pixel values. This residual learning strategy significantly reduces the difficulty of model training and focuses on the recovery of fine structures.
[0089] The micro-feature diffusion enhancement model employs a conditional DDPM algorithm, which, based on the standard prediction strategy of diffusion model (DDPM), predicts noise or residuals instead of directly regressing pixels. By predicting "high-frequency residuals," the training difficulty is reduced, allowing the model to focus on recovering fine structures (burrs, lobes) rather than wasting resources on reconstructing the overall contour.
[0090] The micro-feature diffusion enhancement model here is a latent diffusion model (LDM) that integrates frequency domain conditional constraints. The difference between this model and the traditional unconditional diffusion model is that the traditional diffusion model tends to generate textures that conform to statistical laws but may not be realistic when denoising. In contrast, the micro-feature diffusion enhancement model of this invention uses real wavelet coefficients as strong constraints to ensure that the generated textures are consistent with the real pathological structures in terms of spatial location and frequency characteristics.
[0091] After the inverse processing of the micro-feature diffusion enhancement model, the enhanced high-frequency feature distribution is output. This distribution is then restored by the decoder to obtain the micro-feature enhanced high-frequency evidence tensor. This high-frequency evidence tensor contains minute signs of invasion (such as microvascular ends and micropleural traction) that are smoothed out by traditional algorithms, providing precise hierarchical input for subsequent multimodal fusion processing. It also serves as a data source for various micro-signs in the visual atlas construction during the consistency judgment step.
[0092] Inverse diffusion enhancement using a frequency domain conditional residual diffusion network utilizes a pre-trained improved denoising network. From pure Gaussian noise The process begins with T-step iterations to gradually remove noise and inject high-frequency details. Specific steps include:
[0093] S311. Initialize random noise , This represents a standard normal distribution with a mean of 0 and a covariance of I.
[0094] S312, proceed The iterative sampling gradually decreases from t=T to t=1 (or 0). In each sampling step, wavelet high-frequency coefficients are used. The direction of noise reduction is guided as follows:
[0095] ;
[0096] in, The features of the previous time step after denoising are represented. For predefined noise variance scheduling parameters, This represents the noise residual predicted by the neural network. The added random perturbation is used to increase the diversity of the generated data.
[0097] This formula represents the denoising network. exist Guided by this principle, the edges of pathological micro-signs are preserved while irrelevant random noise is suppressed.
[0098] S313. When t=0, the enhanced latent features are obtained, and the final high-frequency evidence tensor is obtained through decoding. Mathematically, this tensor is represented as a multi-channel matrix of the same size as the ROI of the original image, but it is enriched with high-frequency responses such as spikes and lobes in the channel dimension.
[0099] In steps S311-S313, a high-frequency evidence tensor with enhanced micro-features is generated using a frequency-domain conditional residual diffusion network. The process was explained.
[0100] The following section discusses high-frequency evidence tensors. The explicit quantization process is described, including the following steps:
[0101] S321. In the three-dimensional ROI of the nodule, select multi-channel volume data X that includes at least the lung window channel and the mediastinal window channel;
[0102] S322. Perform three-dimensional discrete wavelet decomposition on X, with a decomposition level of J (e.g., J=1~3), to obtain a set of coefficients {C{j,b}} for low-frequency subbands and several high-frequency subbands, where... Indicates scale. This indicates directional combination subbands, which are high-frequency components pointing in different directions such as LH, HL, and HH;
[0103] S323. Construct a boundary shell region of thickness t at its boundary based on the nodule mask. And construct the outer ring zone region based on the outward expansion / inward contraction parameters. (Can be shared with or set independently of the vascular ring); the high-frequency sub-band coefficient is set according to... and Mask constraints and filtering are performed to obtain the set of high-frequency coefficients of the mask. ;
[0104] S324, will Stacked according to direction and scale, and then subjected to channel normalization and convolution mapping to obtain a high-frequency evidence tensor with a uniform number of channels. .
[0105] Based on the enhanced high-frequency evidence tensor The following three physical quantities are calculated to address the problem of "difficulty in quantifying vascular bundles"; the high-frequency evidence tensor is then used. The input is fed into the feature quantization module to obtain an evidence scalar that reflects the intensity of micro-features, including:
[0106] 1) High-frequency energy : Used to measure the degree of irregularity of the nodule's edge; expressed as:
[0107] ;
[0108] in, denoted as the boundary shell region of the nodule, where p is the voxel coordinate. The larger the value, the rougher the edge and the higher the erosiveness. ||HF enhanced (p)|| represents the norm of the high-frequency enhancement tensor value at voxel coordinate p (usually the L2 norm, i.e., the square root of the sum of squares), representing the texture intensity at that point.
[0109] 2) Directional anisotropy : Used to measure the directionality of burrs (burrs are usually radial); represented as:
[0110] ;
[0111] in, Indicates standard deviation, These represent the energy values of wavelet subbands in different directions (horizontal, diagonal, and vertical). The standard deviation of the energy in each subband (horizontal, vertical, and diagonal) is calculated using this expression. A larger standard deviation indicates a clear directionality in the texture (such as radial spikes); a deviation close to 0 indicates uniform noise or smooth edges.
[0112] 3) Boundary sharpness :
[0113] ;
[0114] in, This represents the gradient operator, which calculates the rate of change (derivative) of a tensor in space. The expression represents a lung nodule mask, and Mean() represents the mean value. This expression is used to calculate the mean gradient magnitude of the high-frequency evidence tensor at the mask boundary.
[0115] The three evidence scalars obtained from the calculation ) spliced into an evidence vector This is used for subsequent fusion operations.
[0116] S4, Microvascular Topology Quantization of Lung Nodules
[0117] Vascular clustering is key evidence for identifying intravascular coagulation (IAC). By calculating vascular topology quantification indicators and using them as input for multimodal fusion, the sensitivity of IAC identification can be improved.
[0118] Vascular topology quantification indicators include tortuosity index, fractal dimension, and convergence score.
[0119] S5. Construct a three-layer structured knowledge graph (Schema) of "nodules-signs-pathology".
[0120] The test results are mapped to a structured schema conforming to international medical standards (such as Lung-RADS and Fleischner guidelines) as an intermediate inference representation; including:
[0121] 1) Lesion layer (Entity): contains a unique ID for the nodule, a detailed 3D ROI mask, and its location in the world coordinate system;
[0122] 2) Attributes Layer: This layer includes deep semantic features describing the nodules, covering imaging features strongly correlated with malignancy risk, including:
[0123] Edge features: lobulation (shallow lobulation / deep lobulation), spiculation (long spiculation / short spiculation), boundary clarity (clear / blurred);
[0124] Internal features: Vacuole sign, air bronchogram (including truncation / torsion / dilation classification), honeycomb pattern, calcification pattern (central / popcorn / layered / diffuse).
[0125] Peripheral features: pleural indentation (including linear / curtain-like), vessel convergence, satellite foci, halo sign;
[0126] 3) Diagnosis: Benign / malignant probability assessment, invasiveness grading prediction (AIS / MIA / IAC), Lung-RADS classification recommendations and follow-up strategies.
[0127] S6. Multimodal Fusion Reasoning and Immersiveness Stratification Assessment
[0128] A multimodal fusion network is used to fuse raw image features, multidimensional clinical structured data, and enhanced high-frequency evidence tensors to output quantified risk probability distributions and auxiliary assessment parameters for stratified infiltration assessment, providing decision support for physicians' diagnosis.
[0129] The inputs for the fusion process in this step include:
[0130] Raw visual stream: Deep features extracted from the raw CT images by the backbone network (such as ResNet or DenseNet CNN network) in step S2. .
[0131] Enhanced Evidence Flow: The high-frequency evidence tensor output in step S3 The tensor is processed by a separate lightweight convolutional encoder to extract feature maps specifically characterizing micro-glitch and shallow lobes. This design ensures that the model (a separate lightweight convolutional encoder) does not lose the minute details recovered in step S3 due to downsampling of the deep network.
[0132] Numerical indicator stream: includes the multidimensional clinical structured data obtained in step S1 and the high-frequency evidence scalar obtained in step S3. ) and vascular topology quantification indicators; these data are normalized and mapped to a one-dimensional vector. .
[0133] The multimodal fusion model employs an asymmetric cross-attention fusion network to perform the following operations:
[0134] Feature alignment: and Concatenate along the channel dimension to form deep visual features that include both macroscopic shape and microscopic texture. .
[0135] Cross-modal interaction: Utilizing a cross-attention mechanism to map a one-dimensional vector. Generate query vectors with deep visual features Generate a key-value matrix; represented as:
[0136] ;
[0137] This mechanism enables a more realistic simulation of the reasoning process; for example, when the input displays "abnormally high CEA index" (Query), the multimodal fusion model adaptively assigns higher weights to "solid components" or "spiculated regions" (Key / Value) in visual features. When vascular topology quantification indicators (such as convergence scores) show abnormalities, the multimodal fusion model adaptively assigns weights to deep visual features through a cross-attention mechanism. The corresponding spatial region has a higher weight.
[0138] This dual input mechanism of "tensor + scalar" ensures that the deep learning model can both see the enhanced image details and make decisions using explicit physical quantification indicators, thus guaranteeing the credibility of the fused output results.
[0139] Through the fusion processing of a multimodal fusion network, the following structured quantized data can be output:
[0140] 1) Infiltration risk probability distribution: Output a normalized probability vector. , respectively, represent the radiographic phenotypes of the nodules as atypical adenomatous hyperplasia, adenocarcinoma in situ, minimally invasive adenocarcinoma, and invasive adenocarcinoma.
[0141] 2) Sign Confidence Score: Outputs a quantitative score (0-1) for key malignant signs, such as "Microvascular Cluster Confidence: 0.85" and "Pleural Traction Confidence: 0.92". This score data can be used in subsequent report generation and report consistency verification steps.
[0142] The multimodal fusion processing outputs not only hierarchical probabilities but also generates a structured schema graph containing the reasoning basis. This structured schema graph directly defines the node and edge attributes of the "visual scene graph" and serves as the "ground truth" for subsequent logical consistency verification.
[0143] This invention deeply integrates tumor markers (CEA, CYFRA21-1, etc.), inflammatory markers (NLR), and radiomics features. It employs a three-dimensional wavelet high-frequency evidence channel to perform fidelity-preserving and explicit modeling of high-frequency micro-signs such as microspiculations and shallow lobulations, reducing the loss of invasive evidence caused by deep feature smoothing and improving the stability of distinguishing between MIA and IAC. Based on "frequency-domain guided diffusion enhancement," it uses wavelet high-frequency coefficients as the "anchor" of the diffusion enhancement model, achieving targeted enhancement of key MIA / IAC differentiation signs such as microspiculations and shallow lobulations while removing CT noise, significantly improving the upper limit of stratified diagnosis.
[0144] S7. Generate a two-part report that is schema-constrained and evidence-anchored.
[0145] Based on the reasoning conclusions derived from schema structure and multimodal fusion, a standardized report is generated using a large language model, including:
[0146] 1) Objective Findings: Using slot filling and restricted decoding technology, descriptions are generated strictly based on the detected attributes. For example: "A mixed ground-glass nodule (type) is visible in the apical segment of the right upper lobe (location), with a size of approximately 15mm × 12mm (size), a solid component of approximately 6mm, and lobulation and short spiculations visible at the margin (marginal signs), with traction on the adjacent pleura (peripheral signs)."
[0147] 2) Impression: Provides Lung-RADS classification and treatment recommendations based on the stratification results. For example: "Lung-RADS 4B, early-stage lung adenocarcinoma (highly likely microinvasive adenocarcinoma MIA), multidisciplinary consultation or PET-CT scan recommended."
[0148] Evidence Anchoring: Each key description in the report (such as "vascular clustering sign" or "lobulation") is embedded with an interactive hyperlink. Clicking on it will take you to the corresponding ROI area in the image viewer and highlight it, achieving "what you see is what you get".
[0149] S8. Report Consistency Verification and Closed-Loop Error Correction
[0150] This approach employs Neuro-Symbolic Dual-Graph Verification, moving beyond black-box similarity scoring to construct an interpretable graph matching mechanism. This addresses the "illusion" problem (discrepancies between report descriptions and actual imagery) that arises when large models generate reports. Unlike traditional text similarity comparison verification methods, this approach references... Figure 2 This invention constructs a mathematically meaningful graph topology comparison, and the consistency verification process includes:
[0151] S81. Constructing the Visual Scene Graph (VSG)
[0152] Based on the image features (high-frequency evidence tensor) after diffusion enhancement processing in step S3 ), identify entities and relationships through object detection algorithms, and construct a visual scene graph. ,in Represents visual nodes. The visual boundary is indicated. The entities here refer to the nodule itself, thymus, blood vessels, etc., while the relationships refer to anatomical positional relationships such as traction (between the nodule and the pleura), truncation (between the nodule and the blood vessels), and close contact.
[0153] The visual scene diagram is an undirected or directed graph with anatomical topological significance, constructed using nodules, pleura, blood vessels, and surrounding lung tissue as nodes and traction, adhesion, passage, and truncation as edges.
[0154] For example, visual nodes : {Node_1: Nodule (attribute: ground glass), Node_2: Edge (attribute: microburrs, source: high-frequency evidence tensor enhanced by micro-signatures in step S3} )};
[0155] Visual edge : {Node_1 --[has]→Node_2}.
[0156] S82. Construct a Text Semantic Graph (TSG)
[0157] Using NLP dependency parsing, a semantic graph was obtained from the description "...a ground-glass nodule is visible with smooth edges..." in the generated report. ,in Represents a text node. Represents a text edge.
[0158] For example, text nodes {Node_A: Nodule (attribute: Ground Glass), Node_B: Edge (attribute: Smooth)};
[0159] Text edges : {Node_A --[has] →Node_B}.
[0160] S83. Calculate the graph edit distance and perform conflict determination.
[0161] Calculation will Transform into Minimum required operating cost (graph edit distance GED). The calculation process is as follows:
[0162] S831, Matching Nodes: Node_1 matches Node_A (the attribute "frosted glass" is the same, cost 0);
[0163] S832, Conflict Detection: Node_2 (micro-burrs) and Node_B (smooth) attributes are mutually exclusive;
[0164] S833, Editing Operation: Convert the attribute "Smooth" to "Microburr". Assume the attribute replacement cost is 1.
[0165] Calculated logical consistency score , The maximum number of nodes in the two graphs (used for normalization); the closer the consistency score is to 1, the higher the consistency between the image and the report; for example, in the above example, the consistency score drops significantly (e.g., below the threshold of 0.8) due to the existence of key attribute conflicts.
[0166] S84, Closed-loop error correction
[0167] When a conflict is detected between textual evidence "attribute: smooth" and visual evidence "attribute: micro-burrs," a correction operation is triggered. The correct triples <edge, have, micro-spicules> are used as prompt constraints, output to the large language model, and the report is regenerated.
[0168] The regenerated report was revised to read "...a ground glass nodule is visible, and the S3 reinforcement layer shows visible microburrs at the edges...".
[0169] The consistency verification operation can backtrack to verify whether the generated report is faithful to the image evidence extracted in steps S3 and S6, and then decide whether to trigger the error correction mechanism to output the final report. This step-by-step progression of "signal (S3) → semantics (S5) → symbol (S8)" ensures the quality and credibility of the generated report.
[0170] S9. Model-in-the-Loop Learning and Optimization
[0171] By deploying a human-machine collaboration interface, feedback from doctors on automatically generated results can be collected, including:
[0172] Boundary correction: Doctors manually adjust the segmentation boundaries of the nodules;
[0173] Conclusion correction: Doctors have revised the predicted "IAC" to "MIA", or "malignant" to "inflammatory".
[0174] These corrections are then transformed into preference data pairs (Rejected vs. Preferred). The Direct Preference Optimization (DPO) algorithm is used to periodically optimize the model, enhancing its ability to interpret long-tail, difficult cases (such as cystic lung cancer and occult metastases), thus achieving continuous model optimization.
[0175] By employing a closed-loop verification and evidence anchoring mechanism, every sentence in the report is guaranteed to be verifiable, achieving "fact-checking" of medical documents and significantly reducing the risk of medical disputes caused by reporting errors. Through structured slot report generation and cross-modal reverse verification, the illusion of reporting is reduced, and "what you see is what you get" evidence traceability is achieved. In cases of consistency conflicts, closed-loop error correction or uncertainty gating output is triggered, reducing the risk of stratified misleading and enhancing clinical credibility.
[0176] By employing a dual-map matching approach combining visual and textual data, fuzzy consistency checks are transformed into precise topological comparisons, fundamentally eliminating the logical illusions inherent in reports generated by large models. The use of structured reasoning paths and visualized image evidence chains enhances doctors' trust in the AI system and improves collaborative efficiency. Combined with a continuous learning mechanism based on doctor feedback, the system's adaptability to long-tail cases and distribution drift is enhanced, increasing its sustainable deployment value.
[0177] The method of the present invention will be described in detail below with reference to specific embodiments.
[0178] S01. Acquisition and processing of multi-source data, including:
[0179] Imaging data: Chest thin-section CT DICOM sequence, slice thickness not exceeding 1.25 mm;
[0180] Multidimensional clinical structured data: patient demographic information (age, sex, BMI, etc.), risk factors (smoking index, etc.), laboratory indicators (CEA, CYFRA21-1, NLR, etc.) and past medical history, etc.
[0181] Perform slice thickness verification, reconstruction kernel consistency check, and artifact detection on CT images;
[0182] Perform missing value processing (e.g., mean / interpolation / model imputation) and outlier pruning on multidimensional clinical structured data;
[0183] Samples that do not meet the quality control threshold are marked as "uncertain" and entered into the manual review queue.
[0184] S02, Image standardization and candidate lesion localization and segmentation, including:
[0185] Resampling: Resample CT images to isotropic voxels, for example, 1mm×1mm×1mm;
[0186] Window mapping: Generate multi-channel input tensors such as lung window and mediastinal window;
[0187] Candidate lesion localization: A three-dimensional detection network or segmentation network is used to obtain candidate ROIs of nodules and their three-dimensional nodule masks;
[0188] Circumferential definition: The peripheral circumferential region (r2>r1) is formed by expanding the nodule mask outward by r2 and contracting it inward by r1, and is used for subsequent vascular topology quantification and calculation of peripheral signs.
[0189] S03. Construction and fusion processing of high-frequency evidence tensors, including:
[0190] 1) Input definition: Obtain the 3D ROI volume data X of the nodule and the nodule mask M; where X contains at least the data of the lung window channel and the mediastinal window channel, and complete resampling and intensity normalization.
[0191] 2) Construction of the boundary shell and peripheral circumference:
[0192] Morphological expansion and erosion are performed on mask M to construct the boundary shell region: Ω bd = Dilate(M, t) \ Erode(M, t) , where t is the shell thickness parameter, and Dilate() and Erode() represent dilation (making the region larger) and erosion (making the region smaller) in morphological operations, respectively.
[0193] Construct the outer ring zone region: Ωring = Dilate(M, r2) \ Dilate(M, r1) ,in r2 > r1 ; ΩRings are used to capture high-frequency peripheral evidence such as the outer edge texture of shallow lobes and subtle pleural traction.
[0194] 3) Three-dimensional discrete wavelet decomposition and high-frequency subband selection
[0195] Perform three-dimensional discrete wavelet decomposition on X, with a decomposition level of J, to obtain a set of coefficients {C{j,b}} for low-frequency subbands and several high-frequency subbands; where b is the identifier of the directional combination subband (e.g., a high-frequency combination subband along the x / y / z direction).
[0196] Mask constraint is used: = C{j,b} ⊙ (I{Ω bd} ∪ I{Ωring}), retaining only the high-frequency coefficients within the boundary shell and outer ring zone; This indicates element-wise multiplication (Hadamard Product), and I{} represents an indicator function or mask; when a pixel is in the region Ω When the value is within the specified range, it is 1; otherwise, it is 0. This function is used to extract coefficients from a specific region.
[0197] Directional energy normalization is adopted: the energy of each directional subband is calculated and scale normalized to avoid fusion bias caused by the difference in amplitude of different scale coefficients.
[0198] 4) High-frequency evidence tensor Build
[0199] After mask constraint Stacking the data according to "scale j—direction b—window level channel" yields the initial high-frequency tensor HF_0; performing channel alignment mapping (e.g., 1×1×1 convolution or equivalent linear mapping) on HF_0 yields the high-frequency evidence tensor. ,make The number of channels matches the number of characteristic channels in the backbone network.
[0200] 5) Explicit Evidence Quantification:
[0201] To transform hidden textural features into interpretable pathological evidence, this step utilizes the enhanced high-frequency evidence tensor (HF). enhanced Calculate the following evidence vectors It directly serves the critical distinction between MIA and IAC, including:
[0202] High-frequency energy ( ): For boundary shell regions With peripheral ring The energy norm is calculated using high-frequency wavelet coefficients within the ring. The energy difference between the inside and outside of the ring band is calculated to quantify the microtexture roughness of the lesion edge and its outward invasion tendency.
[0203] Directional anisotropy ( Skewness or variance is calculated based on the energy distribution of subbands C{j,b} in different directions such as horizontal, vertical, and depth. This index is used to measure the degree of directional aggregation of spiculated radial texture (malignant nodules usually have specific radial anisotropy, while inflammatory nodules tend to be isotropic).
[0204] Boundary sharpness ( :exist Internal computation The edge response or gradient magnitude statistics of the reconstructed features. This indicator is used to characterize the subtle undulations and discontinuous changes of the boundary, and can keenly capture the "solid micro-wetting in the frosted glass composition".
[0205] 6) High-frequency evidence injection and feature fusion (Fusion Strategy), used to inject the generated high-frequency information HF is "injected" into the extracted deep semantic features In the middle, forming ;include:
[0206] A high-frequency evidence injection module is constructed to inject deep semantic features extracted from the backbone network. With high-frequency evidence tensor Perform adaptive fusion; including:
[0207] Feature mapping: for Perform a 1×1 convolution mapping to obtain the result with Dimensional alignment ;
[0208] Gated attention mechanism: Using the Sigmoid activation function to generate channel attention gating vectors. ; where W() represents the learnable weight matrix (convolutional kernel or fully connected layer weights). This indicates a feature concatenation operation. This represents the Sigmoid activation function (mapping values between 0 and 1 as gating weights). This gating mechanism can automatically identify and assign higher weights to feature channels that are highly correlated with microspurs, shallow lobulation, and pleural traction.
[0209] Fusion output formula: final fusion features The residual gating formula is obtained from the following:
[0210] ;
[0211] in, This represents element-wise multiplication (Hadamard Product). This formula ensures that deep features are infused with selected, pathologically significant high-frequency details while preserving high-level semantics.
[0212] 7) Evidence heatmap output and subsequent verification interface: From or Generate an evidence heatmap H_hf (e.g., channel weighted summation and normalization), and use H_hf as the retrieval object for report "evidence anchoring" and reverse visual verification, so that descriptions in the report such as "microburrs / shallow lobes / traction" can be traced back to the corresponding level and spatial region, thereby supporting closed-loop error correction.
[0213] S04, Peripheral microvascular topology quantification, including:
[0214] 1) Vessel extraction: Vessels are segmented within the ring region to obtain a binary mask of the vessels;
[0215] 2) Topology graph construction: The blood vessel mask skeleton is extracted / centerline is extracted, and a graph G=(V,E) is constructed, where V is the set of bifurcation points / endpoints and E is the set of blood vessel segments;
[0216] 3) Indicator calculation, including:
[0217] Tortuosity index: calculated based on the curve length of the vessel segment and the end-to-end distance;
[0218] Fractal dimension: estimated by box counting or equivalent algorithms for the vascular skeleton;
[0219] Convergence score: The convergence score is obtained by statistically analyzing the number of vessel segments pointing towards the center of the nodule, the distribution of their angles, and the density of their branches, and then normalizing the results.
[0220] 4) The above indicators are used as biomarkers of infiltration risk and are input into the multimodal fusion network model along with image depth features and high-frequency evidence.
[0221] By modeling the topological structure of the microvessels in the annulus surrounding the nodule and introducing quantitative indicators such as tortuosity, fractal dimension, and convergence, the qualitative "vascular clustering / convergence" is transformed into a calculable marker of invasive risk, thereby improving the sensitivity of IAC identification.
[0222] S05, Multimodal Fusion Reasoning and Immersive Layered Output
[0223] 1) Multi-source feature vector construction: The input to the multimodal fusion network model consists of three parts, forming a complete "image-topology-clinical" evidence chain; including:
[0224] Deep visual features ( : A fused feature map from the S03 output, including enhancements to high-frequency micro-features;
[0225] Vascular topology index vector ( ): Includes vascular fractal dimension, tortuosity index and centripetal convergence score calculated by S04, used to quantify the impact of nodules on angiogenesis in the surrounding microenvironment;
[0226] Clinical structured data feature vector It includes embedded encoded tumor markers (such as CEA, CYFRA21-1), inflammatory markers (NLR), and high-risk past medical history characteristics.
[0227] 2) Fusion Architecture Design: A hierarchical fusion strategy is adopted; firstly, clinical structured data is converted into clinical feature vectors through a multilayer perceptron (MLP). ,Will , This is concatenated with the evidence metric scalar to form a one-dimensional query vector that enters the attention mechanism, along with deep visual features that serve as key / value pairs. Multi-head cross-attention is performed. This asymmetric cross-attention mechanism allows the system to dynamically focus on the invasive edge of the nodule based on the degree of vascular malformation. This architecture outperforms simple concatenation because it allows the model to dynamically focus on specific regions in the image based on the clinical context.
[0228] 3) Fusion output, including:
[0229] The probabilities of benign and malignant nodules (Pbenign, Pmalignant)
[0230] The probability distribution of the stratified lung adenocarcinoma spectrum is as follows: p(AAH), p(AIS), p(MIA), and p(IAC), which correspond to the probabilities of atypical adenomatous hyperplasia, adenocarcinoma in situ, microinvasive adenocarcinoma, and invasive adenocarcinoma, respectively.
[0231] 4) The stratification conclusion can adopt the maximum a posteriori or threshold strategy, for example, when P(IAC)≥θ IAC If the output is IAC bias, then output MIA / AIS, etc., θ IAC This is the IAC probability threshold.
[0232] S06. Report generation and reverse verification closed-loop error correction, including:
[0233] 1) Structured slots: Nodule location, size, density type, solid component, edge features, surrounding features, and vascular topology summary are written into the schema slots;
[0234] 2) Report generation: Generate Findings and Impression based on slot filling, restricting generation to only reference slots and traceable evidence fields;
[0235] 3) Logical consistency check: Check the rule consistency of "Findings - Conclusion"; for example, if the description is typical hamartoma signs, a high-risk stratification should not be output, and calculate the logical consistency score S log ;
[0236] 4) Closed-loop error correction: When the comprehensive consistency score S = w1·S vis + w2·S log vis <T, trigger:
[0237] Error correction and regeneration (generate again after updating slots or reducing the confidence of the conclusion); or
[0238] Uncertainty gated output (prompt manual review), and record the sample into the feedback pool.
[0239] Among them, S vis represents the visual existence consistency score (whether the words in the report exist in the figure), S log represents the logical consistency score (check whether the reasoning is logical), and w1, w2 are weight coefficients.
[0240] S07. Doctor in the loop and continuous learning, including:
[0241] Doctors correct the segmentation boundary, key sign slots, and stratification conclusions;
[0242] Form preference pairs or supervised samples from "model output - doctor correction" for periodic updating of model parameters;
[0243] Set a higher sampling weight for long-tailed cases to improve the system's adaptability to rare subtypes / atypical morphologies.
[0244] Comparison and ablation experiments
[0245] Compare under the same dataset and evaluation protocol, specifically as follows:
[0246] Baseline A: Only use image depth features for stratification;
[0247] Scenario B: Baseline A + multi-dimensional clinical structured data fusion;
[0248] Scenario C: Scenario B + high-frequency evidence channel;
[0249] Scenario D: Scenario C + vascular topology quantification;
[0250] Scenario E: Scenario D + reverse verification closed-loop error correction;
[0251] Evaluation metrics include: AUC of MIA vs IAC, IAC recall rate, and hallucination rate / consistency score for reporting key symptoms.
[0252] Experimental results show that the present invention significantly improves the ability of the invention to distinguish between MIA and IAC compared to baseline A, and reduces reported hallucinations and inconsistent output.
[0253] On the other hand, the present invention also provides a pathology report generation system for performing the aforementioned case report generation method, comprising:
[0254] The data acquisition and processing module is used to acquire and process chest CT image data and multidimensional clinical structured data;
[0255] The ROI localization and segmentation module for lung nodules is used for multi-scale feature extraction from chest CT images.
[0256] The high-frequency evidence tensor construction and fusion module is used to generate high-frequency evidence tensors with enhanced micro-signs based on the extracted ROI of lung nodules, and to obtain evidence scalars based on the high-frequency evidence tensors.
[0257] The vascular topology quantification module is used to obtain vascular topology quantification indicators based on lung nodule masks.
[0258] The fusion reasoning module is used to concatenate multidimensional clinical structured data, vascular topology quantification indicators and evidence scalars to obtain a one-dimensional vector, fuse deep features and high-frequency evidence tensors to obtain deep visual features, and perform multi-head cross-attention operation on the one-dimensional vector and deep visual features to perform fusion processing on CT images, high-frequency evidence tensors, vascular topology quantification indicators and multidimensional clinical structured data.
[0259] The report generation module is used to generate case reports based on the structured quantitative data output by the three-layer structured knowledge graph and the fusion reasoning module;
[0260] The consistency verification and closed-loop error correction module is used to perform consistency verification and error correction on the generated case reports.
[0261] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A multimodal fusion processing method based on multi-source data, characterized in that, Includes the following steps: Acquire chest CT imaging data and multidimensional clinical structured data; A frequency-domain guided micro-sign diffusion enhancement model is constructed. Wavelet decomposition is performed on the ROI of lung nodules extracted from CT images to obtain high-frequency subband coefficients. The high-frequency subband coefficients are used to guide the inverse denoising process of the micro-sign diffusion enhancement model to obtain the high-frequency evidence tensor of micro-sign enhancement. Based on the high-frequency evidence tensor, an evidence scalar that can reflect the intensity of lung nodule micro-signs is obtained. ROI represents the region of interest. Based on the lung nodule mask extracted from CT images, vascular topology quantification indicators were obtained. A one-dimensional vector is obtained by concatenating multidimensional clinical structured data, vascular topology quantification indicators, and evidence scalars. Deep visual features are obtained by fusing deep features extracted from CT images with high-frequency evidence tensors. The one-dimensional vector and deep visual features are used as inputs to the cross-attention mechanism of the multimodal fusion model to perform fusion processing on CT images, high-frequency evidence tensors, vascular topology quantification indicators, and multidimensional clinical structured data. The micro-feature diffusion enhancement model introduces a frequency domain cross-attention module into each downsampling and upsampling layer of the U-Net; The high-frequency subband coefficient set is mapped to a high-dimensional embedding vector. The frequency domain cross-attention model uses the intermediate layer features of U-Net as the query and the high-dimensional embedding vector as the key and value to perform multi-head cross-attention operation to generate a high-frequency evidence tensor. The evidence scalar includes: High-frequency energy This is used to measure the degree of irregularity at the edge of a nodule; it is expressed as: ; in, The boundary shell region of the nodule, p is the voxel coordinate, ||HF enhanced (p)|| represents the norm of the high-frequency evidence tensor value at voxel coordinate p; The larger the value, the rougher the edge and the higher the erosiveness; Directional anisotropy Used to measure the directionality of burrs; represented as: ; in, Indicates standard deviation, These represent the energy values of wavelet subbands in different directions; Boundary sharpness , represented as This is used to calculate the mean gradient magnitude of the high-frequency evidence tensor at the mask boundary, where... Represents the gradient operator. Indicates a pulmonary nodule mask. This indicates calculating the mean. The steps for fusing deep features extracted from chest CT images with high-frequency evidence tensors to obtain deep visual features include: High-frequency evidence tensor Performing a 1×1 convolution mapping yields deep features extracted from CT images. Dimensional alignment ; Channel attention gating vectors are generated using the Sigmoid activation function, as follows: Where W() represents the learnable weight matrix, This indicates a feature concatenation operation. This represents the Sigmoid activation function; Based on the following residual gating formula, and By fusing the data, deep visual features are obtained. ; indicates as: ; in, This indicates element-wise multiplication.
2. The multimodal fusion processing method based on multi-source data according to claim 1, characterized in that, The steps for generating a high-frequency evidence tensor enhanced with micro-features include: Select multichannel volumetric data that includes at least the lung window channel and the mediastinal window channel in the 3D ROI of lung nodules; Three-dimensional discrete wavelet decomposition is performed on multi-channel volume data to obtain a set of high-frequency sub-band coefficients; Based on the lung nodule mask, a boundary shell region is constructed at its boundary, and an outer ring region is constructed based on the outward expansion / inward contraction parameters. The high-frequency sub-band coefficient set is masked and screened according to the boundary shell region and the outer ring region to obtain the mask high-frequency coefficient set. The high-frequency coefficient set of the mask is stacked according to direction and scale, and then channel normalization and convolution mapping are performed to obtain a high-frequency evidence tensor with a uniform number of channels.
3. A method for generating case reports, characterized in that, Includes the following steps: Construct a three-layer structured knowledge graph of "nodules-signs-pathology"; The multimodal fusion processing method based on multi-source data described in claim 1 or 2 is used to output structured quantitative data, including infiltration risk probability and symptom confidence score. Based on a three-layer structured knowledge graph and structured quantitative data, a large language model is used to generate case reports.
4. The case report generation method according to claim 3, characterized in that, The three-layer structured knowledge graph includes a lesion layer, a sign layer, and a diagnosis layer. The lesion layer includes a unique ID for lung nodules, a three-dimensional ROI mask, and a world coordinate system location. The sign layer includes deep semantic features describing lung nodules and imaging signs strongly correlated with malignancy risk. The diagnosis layer includes benign / malignant probability assessment conclusions, infiltrative grade prediction results, Lung-RADS classification suggestions, and follow-up strategies.
5. The case report generation method according to claim 3, characterized in that, It also includes a step of consistency verification of the generated case reports, including: Based on image features processed by the micro-feature diffusion enhancement model, an object detection algorithm is used to identify entities and relationships, and a visual scene graph is constructed. ,in Represents visual nodes. Indicates visual edge; The description of the generated report is parsed to construct a text semantic graph. ,in Represents a text node. Represents text edges; Calculation will Transform into The minimum operating cost required is used to calculate the logical consistency score. The overall consistency score is calculated based on the logical consistency score and the visual existence consistency score. The overall consistency score is then used to verify the consistency of the report. When the comprehensive consistency score S = w1·S vis + w2·S log <T, trigger: error correction and regeneration, update the slot or reduce the conclusion confidence and then generate; or uncertainty gating output, prompt manual review, and record the sample into the feedback pool; where S vis represents the visual existence consistency score, that is, whether the words in the report exist in the figure, and S log represents the logical consistency score, that is, check whether the reasoning is logical, and w1, w2 are weight coefficients.
6. The case report generation method according to claim 5, characterized in that, In the consistency verification, when the corresponding attributes in the visual scene graph and the text semantic graph are determined to be in conflict, an error correction operation is triggered, and the corresponding attributes in the visual scene graph are output as mandatory constraints to the large language model to regenerate the report.
7. A pathology report generation system, characterized in that, The method for generating a case report according to any one of claims 3-6 includes: The data acquisition and processing module is used to acquire and process chest CT image data and multidimensional clinical structured data; The ROI localization and segmentation module for lung nodules is used for multi-scale feature extraction from chest CT images. The high-frequency evidence tensor construction module is used to generate micro-sign-enhanced high-frequency evidence tensors based on the extracted lung nodule ROIs, and to obtain evidence scalars based on the high-frequency evidence tensors. The vascular topology quantification module is used to obtain vascular topology quantification indicators based on lung nodule masks. The fusion reasoning module is used to concatenate multidimensional clinical structured data, vascular topology quantification indicators and evidence scalars to obtain a one-dimensional vector, fuse deep features and high-frequency evidence tensors to obtain deep visual features, and perform multi-head cross-attention operation on the one-dimensional vector and deep visual features to perform fusion processing on CT images, high-frequency evidence tensors, vascular topology quantification indicators and multidimensional clinical structured data. The report generation module is used to generate case reports based on the structured quantitative data output by the three-layer structured knowledge graph and the fusion reasoning module; The consistency verification and closed-loop error correction module is used to perform consistency verification and error correction on the generated case reports.