A medical image recognition and processing system and method based on multi-modal image fusion

Through parallel neural networks, the accuracy and risk assessment of multimodal image data are solved through the extraction and fusing of tracheal tissue structure characteristics and biological metabolic characteristics, and combined with typical case feature library and AR technology, and efficient medical image recognition processing is achieved.

CN119887773BActive Publication Date: 2025-07-18上海万怡医学科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510376458.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing multimodal image fusion method is complex and has large errors, ignoring the deep correlation and complementarity between different modal features, resulting in limited accuracy of the fusion result and lack of effective risk assessment methods.

Method used

The parallel neural network extracts the structural characteristics of the trachea and biological metabolic characteristics, and fuses them to generate a comprehensive image feature map, compares it with the typical case feature library to generate risk assessment values, and combines AR visualization technology to make correlation.

Benefits of technology

It improves the analysis accuracy and efficiency of multimodal image data, realizes automated processing and rapid diagnosis, can comprehensively evaluate the patient's condition and potential risks, and continuously optimize through adaptive mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887773B_ABST
    Figure CN119887773B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical image recognition and processing system and method based on multi-modal image fusion, belonging to the technical field of image processing. The system includes obtaining medical images and performing correction to obtain an image data set; inputting the image data set into a parallel neural network to extract tracheal tissue structure features and biological metabolism features, and performing fusion to generate a comprehensive image feature map; comparing the fusion features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ, and associating it with the corresponding organ in the preset AR visualization image. By extracting tracheal tissue structure features and biological metabolism features through a parallel neural network, performing fusion to generate a comprehensive image feature map, comparing it with the typical case feature library to generate a risk assessment value, and finally associating it with the AR visualization image, the present invention realizes the effective fusion and utilization of multi-modal image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a medical image recognition and processing system and method based on multi-modal image fusion. Background Art

[0002] With the rapid development of medical imaging technology, doctors can obtain more and more types and modalities of medical imaging data, such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), PET (Positron Emission Tomography), etc. These different-modal imaging data provide rich information about the internal structure and function of the human body, and are crucial for the early diagnosis of diseases, the formulation of treatment plans, and the evaluation of treatment effects.

[0003] With the development of computer technology and artificial intelligence, multi-modal image fusion technology has gradually become a hot research direction in the field of medical image analysis. Existing multi-modal image fusion methods often require complex preprocessing steps, which not only increase the computational complexity but may also introduce additional errors. In addition, existing fusion methods often focus on simple feature stitching or weighted averaging, ignoring the deep associations and complementarities between different-modal features, resulting in limited accuracy of the fusion results. At the same time, how to effectively perform risk assessment on the fused features is also an urgent problem to be solved. Summary of the Invention

[0004] To solve the above problems, the present invention provides a medical image recognition and processing system and method based on multi-modal image fusion. By using a parallel neural network to extract tracheal tissue structure features and biological metabolism features, and fusing them to generate a comprehensive image feature map, then comparing it with a typical case feature library to generate a risk assessment value, and finally associating it with an AR visualization image, the effective fusion and utilization of multi-modal imaging data are realized.

[0005] The above object can be achieved by the following solutions:

[0006] A medical image recognition and processing method based on multi-modal image fusion, comprising: obtaining at least two types of medical images from different devices and performing correction to obtain an image data set; inputting the image data set into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features; fusing the tracheal tissue structure features and the biological metabolism features to obtain fused features, and using the fused features of each organ to generate a comprehensive image feature map; comparing the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ; associating the risk assessment value with the corresponding organ in a preset AR visualization image, so as to obtain the corresponding risk assessment value through the corresponding organ in the AR visualization image.

[0007] Optionally, obtaining at least two types of medical images from different devices and performing correction to obtain an image dataset includes: performing temporal alignment on the medical images from different devices to obtain an initialized image dataset; mapping the image coordinate systems in the initialized image dataset to a unified human anatomical coordinate system, and outputting spatially aligned multi-modal image data; performing noise reduction processing on the multi-modal image data using a preset non-linear filtering algorithm, and outputting noise-reduced image data; performing normalization calibration on the gray values of the noise-reduced image data based on device parameter differences to generate an image dataset with a unified quantization range.

[0008] Optionally, inputting the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features includes: using any one of the image data in the image dataset as a reference, performing voxel interpolation processing on each image data in the image dataset; performing bilinear interpolation compensation on the mechanically distorted areas generated during the mapping process of the image data after voxel interpolation processing to obtain a final image dataset.

[0009] Optionally, inputting the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features includes: inputting the final image dataset into a 3D U-Net network with dilated convolutions to output tracheal tissue structure features; inputting the final image dataset into a multi-scale feature pyramid network to output biological metabolism features.

[0010] Optionally, fusing the tracheal tissue structure features and the biological metabolism features to obtain fusion features includes: obtaining the structural segmentation accuracy score of the current region, calculating the structural weight value of the tracheal tissue structure features, for the structural weight value , there is

[0011] ,

[0012] wherein, is an exponential function, is the structural segmentation accuracy score of the current region, is an adjustment coefficient; obtaining the activity standard deviation of the biological metabolism features of the current region and the activity standard deviation of the biological metabolism features within the entire image range, calculating the metabolic weight value of the biological metabolism features, for the metabolic weight value , there is

[0013] ,

[0014] wherein, is the activity standard deviation of the biological metabolism features of the current region, is the standard deviation of the activity of the biological metabolic characteristics within the entire image range, is the maximum value function;

[0015] Using the structure weight value and the metabolic weight value, a fusion feature is calculated. For the fusion feature , there is

[0016] ,

[0017] In the formula, is the tracheal tissue structure feature, is the biological metabolic feature, is the outer product of tensors.

[0018] Optionally, the generating of the comprehensive image feature map using the fusion features of each organ includes: spatially registering the fusion feature with the standard organ boundaries in a preset human anatomical atlas database; obtaining the metabolic activity values of each registered organ region and sorting them from large to small, selecting the metabolic activity value ranked first, and extracting the metabolic activity quantification index; obtaining the average structure density of each registered organ region and extracting the structure density quantification index; using the metabolic activity quantification index and the structure density quantification index to generate a feature index table with three-dimensional spatial coding to obtain the comprehensive image feature map.

[0019] Optionally, the comparing of the fusion features of different organs in the comprehensive image feature map with the features in a preset typical case feature library includes: retrieving and matching the pathological template corresponding to the organ according to the organ code in the feature index table of the comprehensive image feature map in the typical case feature library; obtaining the metabolic activity quantification index and the structure density quantification index of the corresponding organ, and calculating the similarity index with the pathological template. For the similarity index , there is

[0020] ,

[0021] In the formula, is the structure density quantification index of the th organ, is the metabolic activity quantification index of the th organ.

[0022] Optionally, the generating of the risk assessment value for the corresponding organ includes: determining whether the similarity index is greater than a preset threshold; if the similarity index is greater than the preset threshold, outputting the risk assessment value as 0; if the similarity index is less than the preset threshold, marking the pathological template as a similar case; and outputting the clinical manifestation probability of the similar case through a preset Bayesian network to obtain the risk assessment value for the corresponding organ.

[0023] Optionally, the method further includes: obtaining the true pathological result of the corresponding organ, evaluating according to the true pathological result to obtain a true evaluation value; calculating the difference value between the true evaluation value and the risk evaluation value of the corresponding organ; if the difference value is greater than a preset warning value, reversely correcting the weight parameter of the corresponding pathological template in the typical case feature library.

[0024] Based on the same inventive concept, the present invention further provides a medical image recognition and processing system based on multi-modal image fusion. The system includes: an image data acquisition module, configured to acquire at least two types of medical images from different devices and perform correction to obtain an image data set; a feature extraction module, configured to input the image data set into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features; a feature map generation module, configured to fuse the tracheal tissue structure features and the biological metabolism features to obtain fusion features, and generate a comprehensive image feature map by using the fusion features of each organ; a risk assessment module, configured to compare the fusion features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ; a visualization module, configured to associate the risk assessment value with the corresponding organ in a preset AR visualization image.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] 1. The present invention extracts tracheal tissue structure features and biological metabolism features through a parallel neural network and performs deep fusion, fully considering the complementarity between different modality images, thereby improving the accuracy of image analysis; traditional multi-modal image fusion methods often focus on simple feature splicing or weighted averaging, ignoring the deep-level association between features, while the present invention realizes the optimal cooperation of multi-modal image features through an adaptive weighting mechanism.

[0027] 2. The present invention compares the comprehensive image feature map with the typical case feature library to generate a risk assessment value for the corresponding organ; this method not only considers image features but also combines clinical pathological data, and can more comprehensively evaluate the patient's condition and potential risks; in addition, through the reverse correction mechanism, the system can continuously learn and optimize to improve the accuracy of risk assessment.

[0028] 3. The present invention realizes the automated processing and analysis of multi-modal image data, greatly shortening the diagnosis time; through parallel processing and deep learning technologies, the system can quickly and accurately extract image features, generate risk assessment values, and feedback them to doctors in a timely manner, thereby improving the efficiency of the entire diagnosis process.

[0029] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the specification, claims as well as the drawings. Brief Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 It is a schematic flowchart of a medical image recognition and processing method based on multimodal image fusion according to an embodiment of the present invention.

[0032] Figure 2 It is a schematic flowchart of medical image correction according to an embodiment of the present invention.

[0033] Figure 3 It is a schematic structural diagram of a parallel neural network according to an embodiment of the present invention.

[0034] Figure 4 It is a schematic structural diagram of a medical image recognition and processing system based on multimodal image fusion according to an embodiment of the present invention. Detailed Embodiments

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0036] Refer to Figure 1 , an embodiment of the present invention provides a medical image recognition and processing method based on multimodal image fusion. Tracheal tissue structure features and biological metabolism features are extracted through a parallel neural network, and then fused to generate a comprehensive image feature map. Then, it is compared with a typical case feature library to generate a risk assessment value, and finally associated with an AR visualization image, realizing the effective fusion and utilization of multimodal image data.

[0037] The method of this embodiment specifically includes:

[0038] Obtain at least two types of medical images from different devices and perform correction to obtain an image dataset;

[0039] Specifically, at least two types of medical images are obtained from different devices, and these images can include but are not limited to CT (Computed Tomography), MRI (Magnetic Resonance Imaging), PET (Positron Emission Tomography), etc. These image data reflect the structural and functional information of different tissues and organs in the human body and are the basis for subsequent processing and analysis.

[0040] Input the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features;

[0041] Specifically, after obtaining the image data, these data will be corrected to obtain an image dataset. The correction process includes steps such as temporal alignment, spatial alignment, noise reduction processing, and gray value normalization calibration. Temporal alignment is to solve the problem of asynchronous scanning times of different devices. By extracting the scan timestamps in the DICOM file headers, a multi-modal image temporal sequence is established, and a time interpolation method based on a sliding window is used to compensate for the temporal misalignment caused by respiratory motion. Spatial alignment is to map the coordinate systems of different images to a unified human anatomical coordinate system to ensure the spatial consistency of the image data. Noise reduction processing is to use a preset non-linear filtering algorithm to perform noise reduction on the image data to improve the image quality. Gray value normalization calibration is to calibrate the gray values of the image data based on device parameter differences to ensure the comparability of the gray values of the image data scanned by different devices.

[0042] Specifically, the corrected image dataset is then input into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features. This parallel neural network can include a 3D U-Net network with dilated convolutions and a multi-scale feature pyramid network, etc. The 3D U-Net network with dilated convolutions is used to extract tracheal tissue structure features, and it can accurately capture structural information such as the four-level bifurcation morphology of the trachea. The multi-scale feature pyramid network is used to extract biological metabolism features, and it can effectively analyze the metabolic activity information in the image data.

[0043] Fuse the tracheal tissue structure features and the biological metabolism features to obtain fused features, and generate a comprehensive image feature map using the fused features of each organ;

[0044] Compare the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ;

[0045] Specifically, the extracted tracheal tissue structure features and biological metabolism features are then fused to obtain fused features. The fusion process takes into account the weights of different features. By calculating the structural weight value and the metabolic weight value, the tracheal tissue structure features and the biological metabolism features are weighted and fused. The structural weight value is calculated based on the segmentation accuracy score of the tracheal tissue structure, and the metabolic weight value is calculated based on the standard deviation of the activity of the biological metabolism features. The fused features of each organ are used to generate a comprehensive image feature map, and the fused features of different organs in the comprehensive image feature map are compared with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ. These risk assessment values can help doctors better understand the patient's condition and potential risks.

[0046] Associate the risk assessment value with the corresponding organ in the preset AR visualization image, so as to obtain the corresponding risk assessment value through the corresponding organ in the AR visualization image.

[0047] AR refers to the technology of superimposing virtual images onto the real world through Augmented Reality (AR for short). The AR technology captures the position and angle of the organ image through a camera and performs stereoscopic image processing, and superimposes the virtual image of the organ onto the real world for interaction. This technology combines the virtual world and the real world, and can improve the visual stereoscopic experience and interactivity of the organ. Exemplarily, assume that the system obtains a patient's CT image and PET image. First, the system will perform temporal alignment and spatial alignment on these two images to ensure their consistency in time and space. Then, the corrected image data set is input into a parallel neural network to extract the tracheal tissue structure features and the biological metabolism features in the PET image. Next, these features are fused to obtain fused features. Finally, the fused features are used to generate a comprehensive image feature map and compared with the features in the typical case feature library to generate a risk assessment value for the corresponding organ. Assume that there is abnormal metabolic activity in the patient's lungs, and the system may give a relatively high risk assessment value, prompting the doctor to further examine and treat.

[0048] Optionally, as Figure 2 shown, the obtaining of at least two types of medical images from different devices and the correction to obtain an image data set includes:

[0049] Perform temporal alignment on the medical images of different devices to obtain an initial image data set;

[0050] Specifically, by extracting the scan timestamp in the DICOM file header, a multi-modal image temporal sequence is established. When the time interval between the CT image and the PET image exceeds the device sampling period, a time interpolation method based on a sliding window is used to compensate for the temporal misalignment caused by respiratory motion.

[0051] Exemplarily, for the synchronization problem of chest CT and dynamic PET, the chest CT is 0.5 seconds per layer and the dynamic PET is 2 seconds per frame; the R wave of the electrocardiogram can be used as a reference point to perform gated reconstruction on the PET projection data within the same cardiac cycle to ensure that the CT thin-slice scan and the PET metabolic image are in the same respiratory phase.

[0052] Map the image coordinate system in the initialized image dataset to a unified human anatomical coordinate system, and output spatially aligned multimodal image data;

[0053] Specifically, an initial coordinate system is established with the anterior superior iliac spine - sternal notch - external occipital protuberance, and each modality image is mapped to the standard MNI152 space template using a 3D affine transformation matrix, with the feature point registration error controlled within 1.2 mm.

[0054] Exemplarily, for the PET / CT fusion system, the iterative closest point (ICP) algorithm based on feature points is adopted. First, the SIFT feature descriptors of each modality image are extracted; then, the stable matching point pairs are screened through the RANSAC algorithm; finally, the optimal rigid body transformation parameters are calculated to make the Hausdorff distance of the registered liver contour ≤ 3 mm.

[0055] Perform noise reduction processing on the multimodal image data using a preset non-linear filtering algorithm, and output noise-reduced image data;

[0056] Specifically, an improved filter based on anisotropic diffusion is constructed, and the diffusion coefficient is determined by the formula where is an exponential function, is the gradient threshold, is the gray value of a certain pixel point in the image, and the gradient threshold is dynamically adjusted based on the image signal-to-noise ratio to suppress random noise while retaining the edge features of the organ.

[0057] Normalize and calibrate the gray values of the noise-reduced image data based on the equipment parameter differences to generate an image dataset with a unified quantization range.

[0058] Specifically, establish the mapping relationship between the Hounsfield unit and the standardized uptake value (SUV), generate a calibration curve under different scanning protocols through bicubic spline interpolation, and ensure that the CT value of bone tissue is stable in the range of 100 - 300 HU, and the SUV value of the tumor is scaled to the range of [0, 5].

[0059] This embodiment realizes the precise registration of multimodal images through the collaborative optimization of four dimensions. The time dimension solves the phase difference caused by physiological movement, the space dimension ensures the anatomical structure consistency, the signal dimension eliminates equipment-related noise, and the intensity dimension unifies the quantization standard.

[0060] Optionally, as Figure 3 shown, the inputting of the image data set into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features includes:

[0061] Taking any image data in the image data set as a reference, performing voxel interpolation processing on each image data in the image data set;

[0062] Specifically, taking CT thin-layer scanning as the reference coordinate system, with a CT thin-layer thickness of 0.625 mm, performing asymmetric interpolation for MRI T2WI and PET, with an MRI T2WI layer thickness of 3 mm and a PET layer thickness of 4.25 mm, parsing the PixelSpacing and ImagePositionPatient parameters in the DICOM metadata, calculating the Z-axis voxel resampling position based on Euclidean space, and automatically inserting a transition layer when the distance between consecutive sagittal plane tomograms exceeds 5 mm.

[0063] Performing bilinear interpolation compensation on the mechanically distorted regions generated during the mapping process for the image data after voxel interpolation processing to obtain the final image data set.

[0064] Specifically, identifying the geometric distortion regions caused by magnetic resonance gradient non-linearity, preloading the field strength mapping file provided by the equipment manufacturer, such as the B0 field topology map of 3T MRI; calculating the deviation matrix between the theoretical gradient field and the actual measured value; applying piecewise linear compensation in the phase encoding direction, where the correction coefficient for the central FOV (within a diameter of 30 cm) is 0.95 - 1.05, and high-order polynomial fitting is used for the edge FOV (outside a diameter of 50 cm).

[0065] Exemplarily, for the AX T1 sequence of GE Signa 7T MRI, a 3.8 mm butterfly-shaped distortion is detected in the anterior cranial fossa region; applying a convolution kernel in the frequency domain for anisotropic smoothing; restoring the anatomical contour through bicubic interpolation, reducing the Hausdorff distance at the edge of the brainstem from 4.2 mm to 0.7 mm.

[0066] Specifically, the present invention can also perform multi-modal interpolation compensation, such as using dynamic compensation for the bidirectional deformation caused by respiratory motion. Constructing a deformation vector field using the lung field surface reconstructed by 4D-CT, performing deformation field-driven interpolation on the PET metabolic activity map, and confirming that the displacement error of the diaphragm vertex is ≤2 mm through feature matching. For example, in 4D-PET / CT fusion of lung cancer, taking the end-expiratory CT as the reference coordinate system; calculating the time-phase signal distribution of each projection angle of PET using the back-projection algorithm; mapping the PET data of 10 respiratory phases to the CT reference system through non-rigid registration to eliminate 75% of the respiratory artifacts.

[0067] This embodiment can ensure data consistency through a three - level interpolation system. The first - level geometric interpolation solves the slice - thickness difference, the second - level non - linear correction eliminates the inherent distortion of the device, and the third - level spatio - temporal compensation deals with physiological motion interference, which can significantly improve the accuracy of subsequent pathological analysis.

[0068] Optionally, as Figure 3 shown, inputting the image data set into a preset parallel neural network and extracting tracheal tissue structure features and biological metabolism features includes:

[0069] Inputting the final image data set into a 3D U - Net network with dilated convolutions to output tracheal tissue structure features;

[0070] Specifically, at the 3rd - 5th layers of the encoder, dilatation rates of d = 2 / 4 / 6 are successively adopted to achieve multi - scale receptive field coverage from 3.5 cm³ to 8.7 cm³, which can accurately capture the morphology of the four - level tracheal bifurcation; a channel attention mechanism is introduced in the skip connection to dynamically adjust the weights of the feature maps according to a preset formula; a ternary composite loss of trachea main - bronchus segmentation plus terminal bronchiole segmentation plus boundary constraint term is adopted.

[0071] Inputting the final image data set into a multi - scale feature pyramid network to output biological metabolism features.

[0072] Specifically, the basic backbone network adopts an improved 3D ResNet - 50 architecture, adjusts the stride of stage4 to 1 to retain high - resolution feature maps; the pyramid is constructed by atrous spatial pyramid pooling (ASPP) to generate feature maps with four atrous rates of 1 / 4 / 6 / 24, and multi - scale features are fused by element - wise addition instead of concatenation, and a feature calibration module is embedded after the pooling layer; a dual - path output head is set at the end of the network to respectively predict the SUVmax distribution map and the metabolic heterogeneity index (HI) to obtain biological metabolism features.

[0073] Specifically, this embodiment realizes decoupled feature extraction through a dual - depth network. The 3D U - Net focuses on depicting millimeter - level anatomical structures, and the feature pyramid focuses on metabolic heterogeneity analysis. The two networks share a pre - processing layer but are independently trained.

[0074] Optionally, fusing the tracheal tissue structure features and the biological metabolism features to obtain fused features includes:

[0075] Obtaining the structural segmentation accuracy score of the current region, calculating the structural weight value of the tracheal tissue structure features, and for the structural weight value , there is

[0076] ,

[0077] In the formula, is an exponential function, is the structural segmentation accuracy score of the current region, is the adjustment coefficient;

[0078] Specifically, the recognition accuracy of the tracheal structure in the current region is mapped through an S-shaped function. The value of the adjustment coefficient in the formula comes from the statistical results of clinical trials, making the slope of the weight change in the interval [0.6, 0.9] the steepest, sensitively capturing the topological errors of bronchi below the fourth level; for extremely low segmentation accuracy, such as <0.4 region, soft threshold truncation is adopted to avoid noise amplification; when calculating the weight, 3×3×3 voxel neighborhood smoothing filtering is applied to prevent local mutations of the weight value.

[0079] Exemplarily, assume that the segmentation result of a certain bronchus segment on CT shows excellent main bronchus segmentation effect, =0.93, then the calculated weight is 0.977, close to the full score of 1; the small branches are generally segmented due to limited CT resolution, =0.57, then the weight is 0.801, still usable but with reduced influence.

[0080] Obtain the active standard deviation of the biological metabolism characteristics in the current region and the active standard deviation of the biological metabolism characteristics within the entire image range, and calculate the metabolic weight value of the biological metabolism characteristics. For the metabolic weight value , there is

[0081] ,

[0082] In the formula, is the active standard deviation of the biological metabolism characteristics in the current region, is the active standard deviation of the biological metabolism characteristics within the entire image range, is the maximum value function;

[0083] Specifically, the weight is automatically adjusted according to the degree of abnormality of the metabolic activity in the PET image. The more abnormal the region, the higher the attention; calculate the fluctuation range of the metabolic value in the local region. The larger the value, the more uneven the metabolism; take the fluctuation of the whole-body background metabolism as the benchmark, and give a high weight when the local region is significantly higher than the whole; when the whole-body metabolism is extremely uniform, such as in a low-metabolism state, the denominator should at least retain 0.1 to prevent the calculation result from being too large.

[0084] Specifically, the 3σ principle is used to define the ROI metabolic abnormal region, and only the voxels exceeding the mean value of the whole image ±3σ are involved in the calculation of the activity standard deviation; the maximum value function is taken for the denominator term to ensure the stability of the weight calculation in the low-metabolism scenario, such as the brain gray matter background, and avoid the risk of division by zero caused by the uniform metabolism of the whole image; the Pearson correlation between the metabolic standard deviation and the malignancy degree is verified by stratified sampling.

[0085] Exemplarily, the metabolism in the tumor area of the lung cancer lesion fluctuates violently, with a local standard deviation of 2.1 and a relatively stable overall standard deviation of 0.65 for the whole body metabolism. The calculated weight is 3.23, highlighting the lesion; the metabolism of the calcified nodule hardly changes, with a standard deviation of 0.08 and an overall standard deviation of 0.52. The calculated weight is 0.15, which is almost negligible.

[0086] Using the structural weight value and the metabolic weight value, a fusion feature is calculated. For the fusion feature , there is

[0087] ,

[0088] In the formula, is the tracheal tissue structure feature, is the biological metabolism feature, is the tensor outer product.

[0089] Specifically, the two features of anatomical structure and metabolic activity are deeply combined to discover hidden associations; the structural feature is regarded as a digital bar with a length of 32, and the metabolic feature is regarded as a digital bar with a width of 16; at each position, the two digital bars are arranged into a table of 32 rows × 16 columns, which is equivalent to a feature relationship table; each grid of the table reflects the combination relationship of the two features, such as the combination of sharp structural edges and increased metabolism; a smart filter (1×1×1 convolution) is used to compress the 32×16 table into 256 key features.

[0090] Specifically, this embodiment realizes an adaptive weighting mechanism based on the credibility of anatomical segmentation, a statistical-driven modeling of metabolic heterogeneity, a high-order tensor operation to capture cross-modal coupling effects, and a deep adaptation of clinical parameters and mathematical models, and finally achieves the optimal cooperation of multi-modal imaging features.

[0091] Optionally, the generating of the comprehensive image feature map using the fusion features of each organ includes:

[0092] Performing spatial registration of the fusion feature with the standard organ boundary in the preset human anatomical atlas database;

[0093] Specifically, spatially register the fused features with the standard organ boundaries in the preset human anatomical atlas database to ensure that each fused feature can be accurately corresponding to a specific organ on the human anatomical structure. This step is crucial because it establishes the correspondence between the fused features and the anatomical structure, providing the basis for subsequent risk assessment.

[0094] Obtain the metabolic activity values of each registered organ region and sort them from large to small. Select the metabolic activity value ranked first and extract the metabolic activity quantification index.

[0095] Specifically, obtain the metabolic activity values of each registered organ region. These values reflect the biological metabolic state of the organ and are important indicators for evaluating the organ health status. Sort these metabolic activity values from large to small, select the metabolic activity value ranked first, and extract the metabolic activity quantification index. This step highlights the most metabolically active organ region, helping to quickly identify organs that may be abnormal.

[0096] Obtain the average structural density of each registered organ region and extract the structural density quantification index.

[0097] Specifically, obtain the average structural density of each registered organ region. This index reflects the tightness of the organ tissue structure and is also an important basis for evaluating the organ health status. Based on these average structural densities, extract the structural density quantification index.

[0098] Generate a feature index table with three-dimensional spatial encoding using the metabolic activity quantification index and the structural density quantification index to obtain a comprehensive image feature atlas.

[0099] Specifically, generate a feature index table with three-dimensional spatial encoding using the metabolic activity quantification index and the structural density quantification index. This feature index table integrates the fused features, metabolic activity quantification index, and structural density quantification index of each organ, forming a comprehensive image feature atlas. This atlas not only contains rich organ information but also has three-dimensional spatial encoding, facilitating subsequent risk assessment and visualization processing.

[0100] Exemplarily, assume that the multimodal imaging data of a patient has undergone the previous processing steps to obtain fused features. Next, these fused features are spatially registered with a pre-set human anatomical atlas database. After registration, it is found that the metabolic activity value in the liver region is the highest. Therefore, the metabolic activity value of the liver is selected as the metabolic activity quantification index. At the same time, we also calculated the average structural density of the liver region as the structural density quantification index. Finally, these indexes are integrated into the feature index table to obtain the comprehensive imaging feature map of this patient. This map clearly shows the metabolic activity and tissue structure characteristics of the liver, providing strong support for subsequent risk assessment.

[0101] Optionally, the comparison of the fused features of different organs in the comprehensive imaging feature map with the features in the pre-set typical case feature library includes:

[0102] Retrieve the pathological template of the corresponding organ that matches in the typical case feature library according to the organ code in the feature index table of the comprehensive imaging feature map;

[0103] Specifically, assign a unique code to each organ, such as LUNG001 for the lung; quickly retrieve the typical cases of the same type of organ in the database according to the code, such as templates for benign tumors / malignant tumors / inflammation, etc.; determine the target location through the coordinate information in the code, such as the specific area of the tumor in the upper lobe of the right lung.

[0104] Obtain the metabolic activity quantification index and the structural density quantification index of the corresponding organ, and calculate the similarity index with the pathological template. For the similarity index , there is

[0105] ,

[0106] In the formula, is the structural density quantification index of the th organ, is the metabolic activity quantification index of the th organ.

[0107] Specifically, compare the density value of the current organ with the reference value recorded in the case library to evaluate the difference in metabolic activity; only when both are close to the reference value can the score be high to avoid misjudgment by a single index; the smaller the similarity index, the greater the difference from the pathological template, and the lower the possibility of abnormality.

[0108] Optionally, the generation of the risk assessment value for the corresponding organ includes:

[0109] Judge whether the similarity index is greater than a pre-set threshold;

[0110] If the similarity index is greater than the preset threshold, the output risk assessment value is 0;

[0111] Specifically, it is determined whether the similarity index of the corresponding organ in the comprehensive image feature map is greater than the preset threshold. This step is the key to risk assessment and determines whether further risk calculation is required. If the similarity index is greater than the preset threshold, it indicates that the health status of the organ has a relatively low similarity with the pathological template in the typical case feature library and is in a normal or low-risk state. Therefore, the output risk assessment value is directly set to 0.

[0112] If the similarity index is less than the preset threshold, the pathological template is marked as a similar case;

[0113] Through a preset Bayesian network, the probability of the clinical manifestations of the similar case is output to obtain the risk assessment value of the corresponding organ.

[0114] Specifically, the Bayesian network combines medical prior knowledge with current data to calculate probabilities. The input layer is the metabolic activity and structural density of the current patient's organ; the hidden layer is the pathological results of each similar case, for example, lung cancer is 1 and tuberculosis is 0; the output layer is the posterior probability that the current case is malignant; for each confirmed real diagnosis, the conditional probability table between nodes is automatically adjusted.

[0115] Exemplarily, it is assumed that the patient data matches similar case 01 and similar case 28. Among them, the malignant probability of similar case 01 is 85%, the CT density is 145, and the PET metabolism is 22.3; the malignant probability of similar case 28 is 72%, the CT density is 130, and the PET metabolism is 18.7; after weighted calculation, the density of the current patient's organ is 138 and the metabolism is 19.5, thereby obtaining a comprehensive risk probability of 79.3%.

[0116] Optionally, the method further includes:

[0117] Obtain the real pathological result of the corresponding organ and perform an assessment based on the real pathological result to obtain a real assessment value;

[0118] Specifically, a feedback loop between the artificial intelligence output and the final diagnosis is established; the pathological report of the surgical specimen, the result of the puncture biopsy, etc. are used as data calibration anchor points; the warning value is set to 15%, and when the difference is greater than 15%, the database is updated; the template related to the misjudged case is preferentially corrected. For example, when a malignant case is misjudged as a benign case, the weight of the corresponding template is increased; a sandbox test mechanism is adopted, such as the corrected parameters need to be verified by historical cases before they take effect.

[0119] Calculate the difference value between the real assessment value and the risk assessment value of the corresponding organ;

[0120] If the difference value is greater than a preset warning value, the weight parameter of the corresponding pathological template in the typical case feature library is corrected inversely.

[0121] Specifically, mark each organ to be corrected according to the pathological report, such as the true label: benign / malignant / other; use the Manhattan distance to calculate the absolute deviation between the predicted risk value and the true value of the system; trace back the Top5 similar templates called during the comparison of this case inversely; impose a penalty factor on the weight of the wrong template; add a new annotation to the affected template, such as "V2.1_Reason for correction: Optimized after misjudging pneumonia as mucinous adenocarcinoma".

[0122] Exemplarily, assume that the system has performed a risk assessment on a patient's lungs and given a risk assessment value. However, through actual diagnosis, it is found that there are relatively serious lesions in the patient's lungs, which are significantly different from the risk assessment value. At this time, the difference value between the true assessment value and the risk assessment value will be calculated, and it will be judged whether this difference value is greater than the preset warning value. If the difference value is greater than the warning value, the weight parameter of the pathological template related to the lung lesion in the typical case feature library will be corrected inversely, so as to more accurately reflect the actual pathological situation in future risk assessments. Through such a self-optimization mechanism, the system can continuously improve the accuracy and reliability of its risk assessment and provide more accurate and useful diagnostic assistance information for doctors.

[0123] Based on the same inventive concept, as Figure 4 shown, the present invention also provides a medical image recognition and processing system based on multi-modal image fusion, and the system includes:

[0124] An image data acquisition module, configured to obtain at least two types of medical images from different devices and perform correction to obtain an image data set;

[0125] A feature extraction module, configured to input the image data set into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features;

[0126] A feature map generation module, configured to fuse the tracheal tissue structure features and the biological metabolism features to obtain fused features, and generate a comprehensive image feature map by using the fused features of each organ;

[0127] A risk assessment module, configured to compare the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ;

[0128] A visualization module, configured to associate the risk assessment value with the corresponding organ in a preset AR visualization image.

[0129] It should be noted that the electrical connections between the above-mentioned units do not necessarily represent direct connections of the circuits. Indirect connection methods, as long as they can achieve the purpose of the present invention, can be applied to the embodiments of the present invention. The above are only exemplary embodiments of the present invention and should not be used to limit the scope of the present invention.

[0130] That is, any equivalent changes and modifications made in accordance with the teachings of the present invention still fall within the scope covered by the present invention. Those skilled in the art will easily think of other implementation schemes of the present invention after considering the specification and the disclosure of the practical truth. This application aims to cover any variations, uses or adaptive changes of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not recorded in the present invention.

Claims

1. A medical image recognition and processing method based on multi-modal image fusion, characterized in that, The method includes: Obtaining at least two types of medical images from different devices and performing correction to obtain an image dataset; Inputting the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features; Fusing the tracheal tissue structure features and the biological metabolism features to obtain fused features, and generating a comprehensive image feature map by using the fused features of each organ; Comparing the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ; Associating the risk assessment value with the corresponding organ in a preset AR visualization image, so as to obtain the corresponding risk assessment value through the corresponding organ in the AR visualization image; The generating a comprehensive image feature map by using the fused features of each organ includes: Performing spatial registration on the fused features and the standard organ boundaries in a preset human anatomy atlas database; Obtaining the metabolic activity values of each registered organ region, sorting them from large to small, selecting the metabolic activity value ranked first, and extracting a metabolic activity quantification index; Obtaining the average structural density of each registered organ region and extracting a structural density quantification index; Using the metabolic activity quantification index and the structural density quantification index to generate a feature index table with three-dimensional space encoding to obtain a comprehensive image feature map.

2. The medical image recognition and processing method based on multi-modal image fusion according to claim 1, wherein The obtaining at least two types of medical images from different devices and performing correction to obtain an image dataset includes: Performing temporal alignment on the medical images of different devices to obtain an initialized image dataset; Mapping the image coordinate system in the initialized image dataset to a unified human anatomy coordinate system, and outputting spatially aligned multimodal image data; Performing noise reduction processing on the multimodal image data by using a preset non-linear filtering algorithm, and outputting noise-reduced image data; Performing normalization calibration on the gray values of the noise-reduced image data based on device parameter differences to generate an image dataset with a unified quantization range.

3. A medical image recognition and processing method based on multi-modal image fusion according to claim 1, characterized in that, The inputting the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features includes: Taking any image data in the image dataset as a reference and performing voxel interpolation processing on each image data in the image dataset; Performing bilinear interpolation compensation on the mechanically distorted regions generated during the mapping process of the image data after voxel interpolation processing to obtain a final image dataset.

4. A medical image recognition and processing method based on multi-modal image fusion according to claim 1, characterized in that, The inputting the image dataset into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features includes: Inputting the final image dataset into a 3D U-Net network with dilated convolution to output tracheal tissue structure features; Inputting the final image dataset into a multi-scale feature pyramid network to output biological metabolism features.

5. A medical image recognition and processing method based on multi-modal image fusion according to claim 1, characterized in that, The fusing the tracheal tissue structure features and the biological metabolism features to obtain fused features includes: Obtain the structural segmentation accuracy score of the current region, calculate the structural weight value of the tracheal tissue structure feature, and for the structural weight value , there is , In the formula, is an exponential function, is the structural segmentation accuracy score of the current area, is the adjustment coefficient; Obtain the activity standard deviation of the biological metabolic characteristics in the current region and the activity standard deviation of the biological metabolic characteristics within the entire map range, and calculate the metabolic weight value of the biological metabolic characteristics. For the metabolic weight value , there is , In the formula, is the active standard deviation of the biological metabolism characteristics of the current area, is the active standard deviation of the biological metabolism characteristics within the entire map range, is the maximum value function; Using the structural weight value and the metabolic weight value, a fused feature is calculated. For the fused feature , there is , In the formula, is the tracheal tissue structure feature, is the biological metabolism feature, is the tensor outer product.

6. A medical image recognition and processing method based on multi-modal image fusion according to claim 1, characterized in that The comparing the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library includes: Retrieve the pathological template corresponding to the organ in the typical case feature library according to the organ code in the feature index table of the comprehensive image feature map; Obtain the quantified index of the metabolic activity and the quantified index of the structural density of the corresponding organ, and calculate the similarity index with the pathological template. For the similarity index , there is , In the formula, is the structural density quantization index of the th organ, is the metabolic activity quantization index of the th organ.

7. A medical image recognition and processing method based on multi-modal image fusion according to claim 6, characterized in that, The generation of the risk assessment value for the corresponding organ includes: Determine whether the similarity index is greater than a preset threshold; If the similarity index is greater than the preset threshold, output the risk assessment value as 0; If the similarity index is less than the preset threshold, mark the pathological template as a similar case; Output the clinical manifestation probability of the similar case through a preset Bayesian network to obtain the risk assessment value for the corresponding organ.

8. A medical image recognition and processing method based on multi-modal image fusion according to claim 1, characterized in that, The method further includes: Obtain the true pathological result of the corresponding organ and perform an evaluation based on the true pathological result to obtain the true evaluation value; Calculate the difference value between the true evaluation value and the risk assessment value of the corresponding organ; If the difference value is greater than the preset warning value, reversely correct the weight parameter of the corresponding pathological template in the typical case feature library.

9. A medical image recognition and processing system based on multimodal image fusion, which is applied to a medical image recognition and processing method based on multimodal image fusion according to any one of claims 1-8, and is characterized in that, The system includes: An image data acquisition module for obtaining at least two types of medical images from different devices and performing correction to obtain an image data set; A feature extraction module for inputting the image data set into a preset parallel neural network to extract tracheal tissue structure features and biological metabolism features; A feature map generation module for fusing the tracheal tissue structure features and the biological metabolism features to obtain fused features, and generating a comprehensive image feature map using the fused features of each organ; A risk assessment module for comparing the fused features of different organs in the comprehensive image feature map with the features in a preset typical case feature library to generate a risk assessment value for the corresponding organ; A visualization module for associating the risk assessment value with the corresponding organ in a preset AR visualization image.

Citation Information

Patent Citations

  • Augmented reality visualization method, system and device for vessel anatomical structure and medium

    CN118280528A

  • Knowledge mining method and system for tumor field

    CN119673479A