Image information interpretation method based on medical image multi-modal feature fusion
By separating robust structural features and local features in multimodal medical images and constructing a registration error correlation map for reorganization and fusion, the problem of misjudgment of local features caused by registration errors in traditional methods is solved, and the accuracy and reliability of image interpretation are improved.
Patent Information
- Application Number
- CN202511087133.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional multimodal medical image fusion methods suffer from misjudgment of local features due to registration errors, which affects diagnostic accuracy, and are computationally intensive and lack generalization capabilities.
By extracting structural features that are robust to registration errors and separating them from local features, a registration error correlation map is constructed. Invariant features are used to compensate for the unreliability of variable features, and recombined and fused, the images are finally input into the interpretation model to obtain image information.
It improves the accuracy and reliability of medical image interpretation, solves the problem of misjudgment of local features caused by registration errors, and reduces computational complexity.
Smart Images

Figure CN120599429A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image interpretation, and in particular to an image information interpretation method based on multimodal feature fusion of medical images. Background Art
[0002] In medical imaging diagnosis, doctors often need to simultaneously observe images from different imaging modalities (such as CT, MRI, PET, and ultrasound) in order to leverage the complementary information from each modality to make a comprehensive judgment. Traditional multimodal image fusion methods generally follow a "registration first, fusion second" approach: first, rigid or non-rigid registration algorithms are used to precisely align images from different modalities at the pixel level, and then features are superimposed or combined on corresponding pixels. However, in real clinical scenarios, due to factors such as differences in imaging device parameters, scanning time intervals, changes in patient position, and physiological tissue movement or deformation, registration errors between different modalities are often difficult to completely eliminate.
[0003] To address these issues, existing technologies typically employ higher-precision registration algorithms or introduce complex deformation correction models to minimize errors. However, these methods place extremely high demands on image quality and algorithm accuracy, are computationally intensive, and are susceptible to noise, artifacts, and low resolution. This results in insufficient generalization across different devices, locations, or patients, hindering their widespread clinical adoption. Summary of the Invention
[0004] The purpose of the present invention is to provide an image information interpretation method based on multimodal feature fusion of medical images, which solves the technical problem of local feature misjudgment caused by registration errors in traditional multimodal fusion.
[0005] Image information interpretation methods based on multimodal feature fusion of medical images include: Acquiring medical image data of at least two modalities of a target object, wherein the medical image data of at least two modalities have registration errors; Performing feature extraction on medical image data of at least two modalities to obtain a feature set corresponding to each modality, wherein the feature set includes structural features and local features that are robust to registration errors; Perform feature decomposition on the feature set to separate the invariant features that are unrelated to the registration error and the variable features that are affected by the registration error; Construct a registration error correlation graph, which uses invariant features as nodes and spatial deviations between variable features as edge weights to characterize the error correlation relationship between different modal features. Based on the correlation graph, the invariant features and variable features are recombined and fused. In the registration error area, the unreliability of local features is compensated by structural features to obtain fused features. The fused features are input into the preset interpretation model to obtain the image information interpretation results of the target object.
[0006] Furthermore, feature decomposition is achieved by analyzing the sensitivity of features to registration errors, including: Screen out the features that remain stable when the registration error changes as invariant features. The invariant features include the overall morphological features of the organs and the relative positional relationship features between different organs. Screen out features that change significantly with the registration error as variable features, including pixel distribution features and subtle texture features in local areas; During the feature decomposition process, the sensitivity level of the feature is determined by comparing the consistency of the same feature under different registration error levels.
[0007] Furthermore, the construction of the registration error correlation map includes: Core nodes are selected from the structural features of each modality. Core nodes are landmark anatomical features that are robust to registration errors, including bone contours and the direction of main blood vessels. The three-dimensional coordinate difference of the local features corresponding to different modes is calculated as the spatial deviation, and the spatial deviation is used as the basis for calculating the edge weight. The edge weight is positively correlated with the deviation size. A multi-scale association structure consisting of organ-level coarse scale and tissue-level fine scale is constructed. The organ-level coarse scale is formed by connecting the corresponding organ structured feature nodes, and its edge weight reflects the overall registration deviation. The tissue-level fine scale is formed by connecting the local features of the substructure within the organ with the structured feature nodes, and its edge weight reflects the local registration error. The multi-scale association structure is combined to form a registration error association graph.
[0008] Furthermore, the screening of core nodes needs to give priority to anatomically iconic invariant features; the setting of edge weights needs to highlight feature associations with significant errors, that is, when the spatial deviation between variable features increases, the edge weight value increases accordingly to strengthen the error representation; in the multi-scale association structure, the coarse-scale association at the organ level is reflected as the alignment deviation of the overall position of the organ corresponding to different modalities, and the fine-scale association at the tissue level is reflected as the alignment error of the local area of the substructure inside the organ.
[0009] Furthermore, in the recombination and fusion step, the judgment of the unreliability of local features specifically includes: Calculating the spatial matching degree between the local feature and the surrounding structured features. The matching degree is evaluated by the consistency of the spatial distribution patterns of the local feature and the structured features, including whether the spatial coordinates of the local feature fall within the anatomically reasonable range of the structured features. When the spatial matching degree is lower than a first preset threshold, the local feature is marked as a potentially unreliable feature. For the marked potentially unreliable features, the registration error level of the region in which they are located is combined to make another judgment. If the potentially unreliable feature is located in a high error region and is significantly different from the corresponding local features of other modalities in the region, it is confirmed as an unreliable feature. For local features that are confirmed to be unreliable, go back to the feature extraction stage, re-extract the structural features of the area where the local feature is located and compare them. If N consecutive times, N ≥ 2 comparison results show that the spatial matching degree is lower than the first preset threshold, then the unreliability of the local feature is locked; The compensation method includes building anatomical constraint rules based on the spatial distribution law of structural features in unreliable areas, and screening, correcting or replacing unreliable local features based on the rules to make the integration results conform to anatomical logic; Repeat the above judgment process until the proportion of unreliable features is lower than the preset proportion of the total feature quantity, or when the recognition results of unreliable features in two consecutive judgment results are consistent, the judgment process is stopped.
[0010] Furthermore, the restructuring and integration includes regional implementation, specifically: Based on the distribution range of edge weights in the registration error correlation graph, a weight limit is set, and the area where the edge weight value exceeds the limit is divided into the area with large error, and the rest are the area with small error; For areas with smaller errors, more detailed information of local features is retained, and the matching degree between local features and structural features is calculated using structural features as an auxiliary reference. When the matching degree is lower than the first matching threshold, the local feature is marked as a feature to be verified, and the process returns to the feature extraction step to re-extract the local feature and calculate the matching degree again until the matching degree is no lower than the first matching threshold or the preset number of extractions is reached; For areas with large errors, the matching degree between local features and structured features is calculated based on the structured features. Only local features with a matching degree not lower than the second matching threshold are retained. Local features with a matching degree lower than the second matching threshold are marked as features to be optimized. They are adjusted and optimized based on the structured features and the matching degree is recalculated until the matching degree is no lower than the second matching threshold or the preset number of optimizations is reached. The fusion results of the areas with smaller errors and larger errors are checked as a whole, and the consistency of the fusion features is judged by the spatial continuity of the structured features. If there are incoherent features, return to the feature retention step of the corresponding area and reprocess until the overall fusion features meet the consistency requirements.
[0011] Furthermore, structured features that are robust to registration errors include: The organ contour features determined based on anatomical atlas can reflect the overall shape and boundary range of the organ; The connection relationship characteristics between different anatomical structures, such as the attachment relationship between blood vessels and organs, and the distribution characteristics of tissue spaces.
[0012] Furthermore, local features include: The grayscale distribution characteristics of a local area in an image can reflect the difference in tissue density in that area; The regular arrangement characteristics of pixels in a local area can reflect the microstructure of the tissue.
[0013] Furthermore, when medical image data of some modalities is missing, the method further includes: Based on the structural features of the acquired modality, key information related to the structural features in the missing modality is inferred; The inferred key information is combined with the acquired feature set to perform feature decomposition and reorganization fusion to ensure that the fusion process is not significantly affected by modality loss.
[0014] Furthermore, it also includes: In the area with smaller errors, when local features conflict with structural features, the probability of occurrence of the conflicting features in the clinical case database is calculated. If it is lower than the rare threshold, the structural features are used to correct the local features. In areas with large errors, cross-modal verification is performed on local features that match the structured features. When the response value difference exceeds the tolerance threshold, the structured features are used to guide the reconstruction of local features. After regional fusion, the fused features are divided into high, medium and low reliability levels according to the matching stability between local features and structural features; Processing according to anatomical constraint rules: Highly reliable features are directly retained, moderately reliable features are retained or modified after verification with adjacent structural features, and low-reliability features are reconstructed according to the topology of structural features; The above process is repeated until the difference of the fused features of two consecutive iterations is lower than the convergence threshold and the matching degree with the anatomical template exceeds the threshold.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention extracts structured features (anatomical structure topology features, organ contour features) that are robust to registration errors and combines them with local features, decomposing the features into invariant features (with structured features as the core) and variable features (including local features), constructing a registration error correlation map to achieve precise fusion. Finally, the unreliability of local features is compensated by structured features in the registration error area, which is beneficial to solving the problem of local feature misjudgment caused by registration errors in traditional multimodal fusion and improving the accuracy of medical image interpretation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the framework structure of the method of the present invention. DETAILED DESCRIPTION
[0017] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] Registration error refers to the spatial positional deviation or inconsistency that persists after registration (spatial alignment) of images of different modalities (such as CT, MRI, and PET) in multimodal medical image analysis. This error can prevent the corresponding positions of the same anatomical structure in images of different modalities from completely aligning, thereby affecting the accuracy of subsequent feature extraction, fusion, and image interpretation. For example, registration errors can directly lead to spatial misalignment of local features (such as pixel grayscale and texture), making these features unreliable during multimodal fusion: Traditional multimodal feature fusion methods typically directly fuse features from the original image or without distinguishing the effects of errors. This can easily lead to misjudgment of local features (such as the grayscale of the lesion edge) due to registration deviation, affecting diagnostic accuracy. For example, if there is a spatial offset between the highlighted area of a tumor in CT and the enhanced area in MRI, direct fusion may obscure the true lesion boundary.
[0019] For the above questions, please refer to Figure 1 This application provides an image information interpretation method based on multimodal feature fusion of medical images, including: S1. Acquire medical image data of at least two modalities of a target object, where the medical image data of at least two modalities have registration errors; S2. performing feature extraction on medical image data of at least two modalities to obtain a feature set corresponding to each modality, wherein the feature set includes structural features and local features that are robust to registration errors; Among them, registration error robustness refers to the characteristic that the feature can still maintain stability and reliability when there is registration error in the medical image (i.e., spatial alignment deviation between images of different modalities), and is not easily distorted or failed due to error interference.
[0020] Specifically, when medical images of different modalities are inaccurately aligned in spatial position (registration error) due to equipment differences, tissue deformation, scanning time interval, etc., features that are robust to registration errors can still accurately reflect the key information in the image in the presence of such deviations, and will not lose their representational significance due to spatial misalignment of local areas.
[0021] S3. Decompose the feature set to separate the invariant features that are not related to the registration error and the variable features that are affected by the registration error; S4. Construct a registration error correlation graph. The correlation graph uses invariant features (structural features) as nodes and the spatial deviations between variable features (local features) as edge weights to characterize the error correlation relationship between different modal features. S5. Based on the correlation graph, the invariant features and variable features are recombined and fused. In the registration error area, the unreliability of local features is compensated by the structural features to obtain the fused features. S6. Input the fusion features into a preset interpretation model to obtain the image information interpretation result of the target object.
[0022] Among them, medical image multimodality refers to the use of equipment with different imaging principles to obtain multiple types of medical images of the same target object; The topological characteristics of anatomical structures refer to the characteristics that describe the connection relationships and spatial distribution between human organs and tissues.
[0023] Local features are features that focus on the details of local areas of the image.
[0024] The pixel grayscale gradient feature is a feature that measures the rate of change of the grayscale values of adjacent pixels in an image.
[0025] Texture features are the distribution pattern characteristics of pixel grayscale in a local area of an image. The working principle of this application is as follows: First, at least two modal image data of the target object (such as the patient's chest or brain) is obtained using medical imaging equipment (such as a CT machine or MRI scanner). Due to factors such as imaging parameters (such as slice thickness and scanning angle) and patient positional micro-movement, there may be deviations in the spatial coordinates of the same anatomical structure (such as the lung lobe or brain gray matter) in the different modal images, which is called registration error. For example, in chest CT and PET images, the lung edge may have a positional offset of 1-3 pixels; Secondly, features are extracted by combining deep learning models (such as U-Net) with traditional computer vision algorithms: Structured features: anatomical topological features (such as the connection relationship between vascular branches) and organ contour features (such as the morphological contour of the liver) are extracted through edge detection and topological analysis. These features are less affected by local spatial offsets and are robust to registration errors. Local features: Gray-level co-occurrence matrix is used to calculate pixel gray gradient features (such as the rate of change between light and dark in the lesion area) and texture features (such as the thickness of the texture inside the tumor). These features are strongly correlated with spatial position and are easily affected by registration errors. Next, the sensitivity of the features to the registration error is classified: Invariant features: With structural features as the core, such as the topological connection pattern of brain sulci and gyri, the overall structural relationship remains stable even if there are local alignment deviations. Variable features: mainly include local features, such as the grayscale gradient of the bone edge. If the registration offset exceeds 1 pixel, the value of the feature may change significantly.
[0026] Using invariant features (structured features) as nodes (e.g., "left upper lobe contour" and "pulmonary artery branch topology"), the spatial deviation of variable features (local features) in different modalities is calculated (e.g., the deviation distance between the grayscale gradient of the tumor edge in CT and the corresponding position in MRI). This deviation value is used as the edge weight to construct an association graph. For example, the edge weight between the "tumor contour (structured feature)" node is the spatial offset of the tumor edge texture features in CT and MRI, which intuitively reflects the error correlation between features. Finally, in areas with small registration errors, the invariant features are directly fused with the variable features. In areas with large registration errors (e.g., areas where edge weights are greater than a preset threshold), the variable features are corrected using the invariant features as a benchmark. For example, the liver contour (structural features) is used to calibrate the grayscale feature position of the liver lesion in MRI to compensate for the unreliability of local features caused by offset, ultimately generating fused features. The fused features are input into a preset interpretation model (such as a convolutional neural network (CNN) or a clinical diagnostic rule model). The model classifies and identifies the features (such as distinguishing between benign nodules and malignant tumors) and outputs the image information interpretation results (such as "a malignant tumor with a diameter of 1.2 cm is present in the left upper lobe of the lung with clear boundaries").
[0027] The innovation of this application lies in extracting structured features (anatomical structure topology features, organ contour features) that are robust to registration errors and combining them with local features, decomposing the features into invariant features (with structured features as the core) and variable features (including local features), constructing a registration error correlation map to achieve precise fusion, and finally compensating for the unreliability of local features in the registration error area through structured features, which is conducive to solving the problem of local feature misjudgment caused by registration errors in traditional multimodal fusion, and is conducive to improving the accuracy of medical image interpretation.
[0028] In some embodiments of the present application, when fusion features in multimodal medical images are used directly for subsequent analysis without distinguishing between features, some features may be significantly affected by registration errors, leading to distortion in the fusion results and compromising image interpretation accuracy. For example, subtle local texture features may be effective when registration errors are small, but will deviate from the true picture when the errors increase. If treated the same as stable organ morphology features, this can interfere with diagnostic judgment.
[0029] In this regard, the present application further proposes to screen out features that remain stable when the registration error changes as invariant features, including the overall morphological features of the organs and the relative positional relationship features between different organs; screen out features that change significantly with the change of the registration error as variable features, including the pixel distribution features and subtle texture features of the local area; in the feature decomposition process, the sensitivity level of the feature is determined by comparing the consistency of the same feature under different registration error levels.
[0030] The working principle of this invention is that the overall morphological features of organs, such as the general outline of the liver or the overall shape of the heart, remain unchanged by small changes in registration error. Furthermore, the relative positional relationships between organs, such as the relative position of the kidneys to the spine or the proximity of the lungs to the heart, remain stable even in the presence of registration error. By observing the performance of these features at different levels of registration error, it was found that their numerical fluctuations were minimal, and therefore they were classified as invariant features.
[0031] Pixel distribution characteristics in a local area, such as the grayscale value distribution of pixels at the edge of a tumor, will be disrupted when the registration error changes, and the distribution characteristics will change significantly. Subtle texture features, such as the gray and white texture details of local brain tissue, are extremely sensitive to spatial position. A slight registration error will cause the texture features to be completely distorted. These features will experience significant fluctuations in their characteristic values when the registration error changes, and are therefore identified as variable features. To determine the sensitivity of a feature, we first manually set different registration error levels (e.g., 0 pixels, 1 pixel, 3 pixels, 5 pixels, etc.). We then calculate the consistency of the feature values for the same feature at these error levels (e.g., using cosine similarity or mean squared error). If the consistency is high (e.g., cosine similarity greater than 0.9), the feature has a low sensitivity level and is classified as an invariant feature. If the consistency is low (e.g., cosine similarity less than 0.5), the feature has a high sensitivity level and is classified as a variable feature. Features between these two levels can be set to a medium sensitivity level, and the processing method is determined based on actual needs. For example, by comparing the consistency of the overall morphological features of the liver under different registration errors, it was found that it always remained above 0.9, while the consistency of the local pixel distribution features of the tumor dropped to 0.4 at an error of 3 pixels, thus clarifying the sensitivity levels of both. In this way, it can help solve the problem of fusion result distortion caused by different sensitivity of features to registration errors. By accurately distinguishing invariant features from variable features, it provides a reliable basis for subsequent feature recombination and fusion based on error correlation maps, ensuring that stable and effective features can still be extracted for image interpretation in the presence of registration errors, which is conducive to improving the accuracy and reliability of multimodal medical image interpretation.
[0032] In some embodiments of the present application, when constructing a registration error correlation map, the traditional single-scale structure has difficulty in taking into account the correlation between global and local registration errors, resulting in insufficient comprehensiveness of feature association and affecting the accuracy of subsequent feature fusion.
[0033] In this regard, this application further proposes to select core nodes from the structured features of each modality. The core nodes are selected from iconic anatomical features that are robust to registration errors, such as bone contours (such as the overall shape of the skull and spine) and the direction of major blood vessels (such as the distribution path of the aorta and basilar artery). These features can remain stable when the registration error changes and can serve as the basic anchor points of the association graph; The spatial deviation is calculated by calculating the difference in the 3D coordinates of the local features corresponding to different modalities. This deviation is then converted into an edge weight through normalization, and the edge weight increases as the deviation increases. For example, when the 3D coordinate difference of the local features at the edge of the tumor in CT and MRI is 2mm, the edge weight is set to 0.2; when the difference increases to 5mm, the edge weight is increased to 0.5, which directly reflects the degree of registration deviation of the local features. Constructing a multi-scale association structure: The coarse organ-level scale is formed by connecting the structured feature nodes of the corresponding organs, such as connecting the structured feature nodes of the liver and gallbladder. The edge weight is calculated based on the overall registration deviation of the two organs, reflecting the overall registration accuracy of the region. The fine tissue-level scale is formed by connecting the local features of the substructures within the organ (such as the subtle texture of the intrahepatic bile duct) with the structured feature nodes of the corresponding organ. The edge weight is determined based on the spatial deviation of the local features, reflecting the registration error of the subtle structure. Combining the two scale structures, a registration error association map covering different levels is formed. In summary, this helps to solve the problem that traditional correlation maps do not fully express registration errors. By reflecting the overall and local errors through multi-scale structural hierarchies, it provides a more accurate error correlation basis for subsequent feature reconstruction, which is conducive to improving the adaptability of multimodal feature fusion to registration errors.
[0034] In some embodiments of the present application, in the construction of the registration error correlation map, if the core node screening lacks anatomical landmarks, the edge weights cannot highlight the significant correlation of errors, and the multi-scale structure does not clearly define the hierarchical deviation differences, it will cause the correlation map to be difficult to accurately reflect the true distribution of the registration error, affecting the targetedness of subsequent feature fusion.
[0035] In this regard, the present application further proposes that the screening of core nodes should give priority to anatomically iconic invariant features; the setting of edge weights should highlight feature associations with significant errors, that is, when the spatial deviation between variable features increases, the edge weight value increases accordingly to strengthen the error representation; in the multi-scale association structure, the coarse-scale association at the organ level is reflected as the alignment deviation of the overall position of the organ corresponding to different modalities, and the fine-scale association at the tissue level is reflected as the alignment error of the local area of the substructure inside the organ.
[0036] It should be understood that iconic invariant features have a clear and stable positioning role in the human body structure, such as the sella turcica structure of the skull (an important anatomical landmark for brain imaging) and the acetabulum contour of the pelvis (the positioning reference for pelvic imaging). Their shape and position are minimally affected by the registration error in different modal images, and can provide reliable spatial anchors for the association map, ensuring the consistency of cross-modal feature associations.
[0037] It should be explained that when the spatial deviation of the local features of the tumor edge increases from 2 mm to 8 mm, the edge weight is linearly increased from 0.2 to 0.8, making the feature associations with larger errors more prominent in the graph, making it easier to focus on correcting the features of these high-error areas during subsequent fusion.
[0038] Furthermore, the spatial offset of the overall liver contour in CT and MRI is reflected by the edge weights connecting the liver's structural feature nodes (e.g., an offset of 5 mm corresponds to a weight of 0.5), which intuitively demonstrates the overall registration accuracy at the organ level; Furthermore, the positional deviation of the subtle texture of the intrahepatic portal vein branches in different modalities is represented by the edge weights of the substructure local features and the liver structured feature nodes (e.g., a 1mm offset corresponds to a weight of 0.1), accurately capturing the error details at the microscopic level.
[0039] This method solves the problem of fuzzy error representation and unclear hierarchy in traditional correlation graphs. By accurately screening core nodes, dynamically adjusting edge weights, and distinguishing multi-scale deviations, the registration error correlation graph can more comprehensively and meticulously reflect the error relationship between features, providing an accurate error reference basis for subsequent feature recombination and fusion, and improving the reliability of multimodal medical image interpretation.
[0040] In some embodiments of the present application, in multimodal medical image feature fusion, local features are susceptible to registration errors and become unreliable. If not promptly identified and processed, the fused features may deviate from anatomical logic, affecting the accuracy of image interpretation. For example, registration errors may cause local texture features of a tumor to be mismatched with normal tissue areas. Direct use of these features can lead to misidentification of lesions.
[0041] In this regard, the present application further proposes calculating the spatial matching degree between local features and surrounding structured features. The matching degree is evaluated by the consistency of the spatial distribution patterns of local features and structured features, including whether the spatial coordinates of the local features fall within the anatomically reasonable range of the structured features. When the spatial matching degree is lower than a first preset threshold, the local feature is marked as a potentially unreliable feature. It should be understood that the coordinates of the local lung feature should fall within the lung contour (structured feature). If it exceeds, the spatial matching degree is low. When the spatial matching degree is lower than the first preset threshold (such as 0.6), it indicates that the spatial correlation between the local feature and the surrounding structured features is abnormal, and it is marked as a potentially unreliable feature.
[0042] For the marked potentially unreliable features, the registration error level of the region in which they are located (high / low error region) is combined to make another judgment. If the potentially unreliable feature is located in a high error region and is significantly different from the corresponding local features of other modalities in the region (the difference value exceeds a second preset threshold), it is confirmed as an unreliable feature; Secondly, high-error regions refer to areas where the registration error exceeds a preset value (e.g., 3 pixels), while low-error regions are the opposite. If a potentially unreliable feature is located in a high-error region and is significantly different from the corresponding local features of other modalities in that region (the difference value exceeds a second preset threshold, such as a grayscale mean difference of more than 50 for texture features), it is confirmed as an unreliable feature. For example, if the density feature of a local area in CT differs significantly from the signal feature of the corresponding area in MRI and is located in a high-error region, the feature will be confirmed as unreliable.
[0043] For local features that are confirmed to be unreliable, the system goes back to the feature extraction stage and re-extracts the structural features of the area where the local feature is located and compares them. If the spatial matching results of N consecutive comparisons (N ≥ 2) show that the degree of spatial matching is lower than the first preset threshold, the local feature is determined to be unreliable. If a local feature is found to fall outside the liver contour after two consecutive extractions, the feature can be determined to be unreliable. Compensation methods include: constructing anatomical constraint rules based on the spatial distribution patterns of structural features in unreliable areas (for example, local features within an organ boundary should conform to the distribution patterns of the organ's tissue characteristics), and screening, correcting, or replacing unreliable local features based on these rules to make the integration results conform to anatomical logic; The above judgment process is repeated until the proportion of unreliable features falls below a preset ratio (e.g., 5%) of the total features, or when the unreliable features are identified in two consecutive judgments, the judgment process is stopped. This ensures that the interference of unreliable features is eliminated as much as possible while avoiding the waste of resources caused by excessive judgment. The innovation of this application is that this method helps solve the problem of local features being unreliable but not identified and processed due to registration errors. By locking unreliable features through multiple rounds of strict judgment and compensating them based on anatomical rules, it ensures that the fusion features conform to anatomical logic, improves the accuracy and reliability of multimodal medical image interpretation, and provides a more reliable basis for disease diagnosis.
[0044] In some embodiments of the present application, during the multimodal medical image feature recombination and fusion process, applying the same fusion strategy to regions with different registration errors can result in loss of local feature details in regions with smaller errors, or unreliable local features in regions with larger errors interfering with the fusion results. For example, in normal tissue regions with smaller errors, over-reliance on structural features may overlook valuable local lesion features; whereas in complex structural regions with larger errors, retaining too many low-matching local features can disrupt the consistency of the fused features.
[0045] In this regard, the present application further proposes setting a weight limit (e.g., setting an edge weight of 0.5 as the limit) based on the distribution range of edge weights in the registration error correlation graph. Regions with edge weight values exceeding this limit are classified as regions with large errors (e.g., the complex region where the tumor and surrounding tissue meet), and the rest as regions with small errors (e.g., normal muscle tissue regions). This division can clarify the focus of feature fusion in different regions and provide a basis for subsequent differentiated processing. For areas with smaller errors, since the registration error has less impact on local features, more detailed information of local features is retained, and structural features are used as an auxiliary reference. The matching degree between local features and structural features is calculated, and when the matching degree is lower than the first matching threshold (such as 0.7), the local feature is marked as a feature to be verified. For example, in an area with smaller errors in normal brain tissue, if the matching degree between a local grayscale feature and the brain tissue contour feature is only 0.6, which is lower than the first matching threshold, it is marked as a feature to be verified, and the process returns to the feature extraction step to re-extract the local feature (such as adjusting the parameters of the feature extraction algorithm) and calculate the matching degree again until the matching degree is no lower than the first matching threshold or the preset number of extractions (such as 3 times) is reached. This process ensures that the local features in areas with smaller errors retain details and match the structural features.
[0046] For areas with large errors, the registration error has a significant impact on local features, so the structured features are used as the main basis. Similarly, the matching degree between local features and structured features is calculated, and only local features with a matching degree not lower than the second matching threshold (such as 0.8, which is higher than the first matching threshold, because the matching degree requirements for areas with large errors are more stringent) are retained. Local features with a matching degree lower than the second matching threshold are marked as features to be optimized, and they are adjusted and optimized based on the structured features (such as correcting the spatial coordinates of local features according to organ contour features, or adjusting the attribute values of local features according to the distribution law of tissue characteristics), and then the matching degree is recalculated until the matching degree is not lower than the second matching threshold or the preset number of optimization times is reached (such as 5 times). For example, in the area with large errors at the edge of the lung tumor, the matching degree between a local texture feature and the lung contour feature is low. The spatial position is corrected based on the lung contour and the matching degree is recalculated to ensure that the retained local features are reliable; The fusion results of the areas with smaller errors and the areas with larger errors are checked as a whole, and the consistency of the fusion features is judged by the spatial continuity of the structural features (such as the consistency of the blood vessel direction and the integrity of the organ contour). If there are incoherent features (such as a segment of blood vessel features being broken in the fusion results of the two areas), return to the feature retention step of the corresponding area and reprocess (such as re-extracting local features in the area with smaller errors and re-optimizing local features in the area with larger errors) until the overall fusion features meet the consistency requirements and form a complete fusion feature that conforms to anatomical logic; This regional reorganization and fusion method is conducive to solving the problem of inappropriate feature fusion strategy in different registration error areas. By adopting differentiated feature processing methods according to the size of the error, it not only retains the local feature details in the area with smaller errors, but also uses structured features to compensate for the unreliability of local features in the area with larger errors. At the same time, the consistency of the fusion features is ensured through overall verification, which improves the quality of multimodal medical image feature fusion and provides a more accurate and reliable fusion feature basis for subsequent image information interpretation.
[0047] In some embodiments of the present application, when extracting structural features from multimodal medical images, if the extraction of structural features relies on precise spatial alignment between the modalities, registration errors can lead to feature extraction failure or distortion. For example, if spatial alignment of CT and MRI images is required to determine the liver contour, registration errors can directly cause misalignment in the extracted contour, thus losing the stability advantage of structural features.
[0048] In this regard, the present application further proposes organ contour features determined based on anatomical atlases, the extraction of which relies solely on identifiable anatomical structures in single-modality images. The anatomical atlas contains standard morphological parameters of various human organs (such as the average size of the liver and the range of contour curvature). There is no need to refer to other modal images during extraction. Contour fitting is performed only through the identifiable features of the organ in a single-modality image (such as CT), such as the grayscale distribution and edge gradient, combined with the atlas. For example, in a CT image, the liver contour is directly extracted by utilizing the density difference between the liver and surrounding tissues (the liver has a higher density than fat and a lower density than bone), combined with the contour morphological constraints of the liver in the anatomical atlas. This contour can reflect the overall morphology and boundary range of the organ and is not affected by the registration error of the MRI image. Connectivity features between different anatomical structures are also independently extracted from single-modality images, independent of spatial alignment between modalities. For example, in MRI images, the flow void effect (signal loss caused by blood flow) is used to identify the course of blood vessels, while simultaneously locating organ contours and directly determining the connection points between blood vessels and organs (such as where the renal artery enters the kidney). The distribution of interstitial spaces is directly identified in CT images through density differences between different tissues (such as the gaps formed by the density difference between muscle and bone), without the need for alignment with other modal images. These connectivity relationships are determined by the inherent spatial distribution of anatomical structures in single-modality images, and even if there are registration errors in other modalities, this does not affect the extraction of connectivity features within a single modality. This extraction method solves the problem of structured feature extraction relying on cross-modal spatial alignment: on the one hand, independent extraction of a single modality ensures that organ contours and structural connectivity can be stably acquired even when registration errors exist; on the other hand, based on anatomical atlases and unimodal identifiable structures, feature extraction has a clear anatomical basis, avoiding feature distortion caused by modal alignment deviations. For example, in multimodal brain image analysis, even if there is a spatial offset between CT and MRI, the cerebral cortex contours and the attachment relationship between cerebral blood vessels and brain regions can still be stably extracted from the two modalities, providing a solid foundation for subsequent use of structured features to compensate for the unreliability of local features, which is conducive to significantly improving the accuracy of multimodal feature fusion.
[0049] The present application further proposes that the local features include: The grayscale distribution characteristics of a local area in an image can reflect the difference in tissue density in that area; The regular arrangement characteristics of pixels in a local area can reflect the microstructure of the tissue.
[0050] Among them, the extraction of local features is only for the local area inside the single-modal image, and does not involve cross-modal feature correspondence.
[0051] It should be understood that local features are detailed features focused on specific small areas in medical images and can be divided into two categories: The grayscale distribution characteristics of a local area are analyzed by analyzing the size and variation of pixel grayscale values within a specific region of an image (e.g., an area a few millimeters in diameter) to reflect differences in tissue density within that region. For example, in CT images, tumor tissue typically appears as an area with higher grayscale values than normal tissue. The uniformity or dispersion of this grayscale distribution can be used to distinguish between benign and malignant tumors. Malignant tumors often have a more chaotic grayscale distribution due to uneven cell proliferation, while benign nodules have a relatively uniform grayscale distribution.
[0052] The regular arrangement of pixels in a local area refers to the spatial arrangement of pixels in a local area, which can reflect the microstructure of the tissue. For example, the pixel arrangement of gray and white matter in brain tissue in MRI images has a clear layered pattern; a disruption of this pattern may indicate cerebral infarction or demyelination. Similarly, a striped arrangement of pixels in a breast image may indicate breast hyperplasia, while a chaotic arrangement may be associated with tumor infiltration.
[0053] These local features are directly related to the pathophysiological characteristics of tissues. However, since they rely on precise spatial positions, feature misalignment or distortion is prone to occur when there are errors in multimodal image registration. Therefore, they need to be verified and corrected in combination with structured features.
[0054] In some embodiments of the present application, when fusion of multimodal medical image features occurs, if some modalities of medical image data are missing (e.g., a patient cannot undergo MRI examination due to metal implants in the body, resulting in missing MRI modality data), directly fusing the existing data will reduce the comprehensiveness of the fused features due to incomplete information, affecting the accuracy of image interpretation. For example, if MRI data that clearly shows soft tissue is missing, relying solely on CT data may not accurately identify the boundary details of the tumor; In this regard, the present application further proposes to infer key information related to the structured features in the missing modality based on the structured features of the acquired modality. The structured features of the acquired modality (such as the bone contour and the overall morphology of the organ in the CT image) have stable anatomical properties, and the key information corresponding to the same structured feature in different modalities is correlated. For example, the contour of the liver in the CT image (structural feature) is known, and combined with the signal characteristics of the liver usually presented in the MRI image (such as the medium signal of the liver parenchyma on the T2-weighted image), the key signal information related to the liver contour in the missing MRI modality can be inferred; for example, based on the direction of the pulmonary blood vessels in the CT (structural feature), key information such as the signal intensity and distribution range of the corresponding blood vessels in the MRI can be inferred. The inference process can be assisted by a pre-trained cross-modal feature mapping model. This model can learn the mapping relationship between the structured features of different modalities and key information through training with a large amount of paired multimodal data, and thus output the key information of the missing modality based on the existing structured features.
[0055] The inferred key information is combined with the acquired feature set to perform feature decomposition and recombinant fusion. The inferred key information is added to the acquired feature set as a supplementary feature for the missing modality (e.g., combining the inferred MRI key information with the CT feature set). Subsequently, following the feature decomposition process, the sensitivity of each feature in the supplemented feature set to registration error is analyzed to screen for invariant and variable features. During the recombinant fusion phase, regions with small and large errors are processed using a combination of inferred information and existing features. In regions with small errors, local feature details from the inferred key information are combined with existing local features, with structural features serving as an auxiliary reference for verification. In regions with large errors, structural features are used as the primary basis, integrating highly matched local features from the inferred information. This ensures that the fused features both supplement the missing information and conform to anatomical logic, thus ensuring that the fusion process is not significantly affected by the missing modality.
[0056] This method effectively solves the problem of incomplete feature fusion information caused by the missing data of some modal medical images. By utilizing the stability of structured features to infer the key information of the missing modalities, the comprehensiveness of the feature set is guaranteed, so that the recombined and fused features can still accurately reflect the human anatomical structure and pathological characteristics, providing a reliable basis for image interpretation, and improving the applicability and accuracy of multimodal medical image interpretation in the case of incomplete modal data.
[0057] In some embodiments of the present application, even after preliminary processing, after regional fusion of multimodal medical images, issues such as feature conflict, insufficient cross-modality consistency, and unclear reliability of fused features may still exist. For example, the tissue boundary indicated by the local grayscale gradient in an area with small errors may not match the organ contour features, or the local features retained in an area with large errors may have significantly different responses in other modalities. These can cause the fused features to deviate from anatomical facts, affecting subsequent interpretation.
[0058] In this regard, the present application further proposes that in areas with smaller errors, when the detailed information of the local feature conflicts with the auxiliary reference of the structured feature (such as the tissue boundary indicated by the local grayscale gradient does not match the organ contour feature), the probability of occurrence of the conflicting feature in the clinical case library is calculated. The clinical case library contains a large number of multimodal images and feature records of real patients. If the probability of occurrence of the conflicting feature in the case library is lower than the clinical rare threshold (such as 0.5%), it means that it does not conform to the common anatomical rules and is an abnormal feature. In this case, the local feature details are corrected based on the structured feature. For example, the local grayscale gradient shows that there is an abnormal protrusion at the edge of the liver, but it conflicts with the liver contour feature, and the protrusion rarely appears in the case library. At this time, the local feature is corrected based on the liver contour to eliminate the protrusion signal.
[0059] In areas with large errors, the retained local features that are highly matched to the structured features are further verified through cross-modal feature consistency. Specifically, the characteristic response value of the local feature corresponding to the anatomical location in the other modality is calculated (such as the signal intensity of the local density feature of a tumor in CT and the corresponding location in MRI). If the response value difference exceeds the modality tolerance threshold (set according to the imaging characteristics of different modalities, such as the density-signal difference threshold between CT and MRI), the local feature reconstruction guided by the structured feature is initiated. Based on the spatial distribution pattern of the structured feature (such as the positional relationship between the tumor and the surrounding blood vessels), a new local feature replacement is generated. For example, if the response difference of a local metabolic feature in PET and CT is too large, the metabolic feature is reconstructed based on the tumor contour and vascular distribution to make it consistent with cross-modal consistency.
[0060] After completing the regional fusion, the reliability of the fused features is graded. This classification is based on the stability of the match between the local features and the structural features (the fluctuation in the matching degree over three consecutive feature extractions). High-reliability features are those with a fluctuation less than or equal to a stability threshold (e.g., 5%), indicating a stable matching relationship with the structural features over multiple extractions. Medium-reliability features have a fluctuation between the stability threshold and the fluctuation threshold (e.g., 15%). Low-reliability features have a fluctuation greater than the fluctuation threshold, indicating an unstable matching relationship. For example, if the matching degree between the local texture and the vascular contour of the lung vessels extracted three times fluctuates by 3%, the feature is considered highly reliable; a fluctuation of 10% is considered moderately reliable; and a fluctuation of 20% is considered lowly reliable.
[0061] Based on the anatomical constraint reinforcement rules, fusion features of different reliability levels are processed: high-reliability features are directly retained due to their strong stability; medium-reliability features are supplemented with verification based on the structured features of adjacent regions (such as the contours of adjacent organs and the direction of blood vessels). If the verification shows that they match the adjacent structures (such as the continuity of the tumor edge features and the contours of the surrounding tissues), they are retained; otherwise, local corrections are made based on the structured features (such as adjusting the feature boundaries to match the contours of adjacent organs); low-reliability features are reconstructed based on the spatial topological relationship of the structured features (such as the connection pattern of vascular branches within organs and the hierarchical structure of tissues). For example, local gray matter features with large fluctuations are reconstructed based on the topological relationship of the brain sulci to ensure that the reconstructed features conform to the anatomical structure of the region.
[0062] The conflict mediation, consistency verification, and reliability grading process is repeated until the difference between the fused features in two consecutive iterations falls below a convergence threshold (e.g., 3%), and the matching degree between the fused features and the preset anatomical template features (based on a standard human anatomy model) exceeds a template threshold (e.g., 90%). The anatomical template features are the standard characteristic patterns of normal or typical lesions. A satisfactory matching degree indicates that the fused features conform to universal anatomical laws, ultimately resulting in a stable and reliable fusion result.
[0063] This process solves the problems of feature conflict, cross-modal inconsistency and reliability ambiguity after regional fusion. Through clinical probability verification, cross-modal verification, hierarchical processing and iterative optimization, the fusion features are more in line with anatomical facts, which is conducive to significantly improving the accuracy and stability of multimodal medical image interpretation, and is particularly suitable for feature fusion scenarios in complex lesion areas.
[0064] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An image information interpretation method based on multimodal feature fusion of medical images, characterized in that: include: Acquiring medical image data of at least two modalities of a target object, wherein the medical image data of at least two modalities have registration errors; Performing feature extraction on medical image data of at least two modalities to obtain a feature set corresponding to each modality, wherein the feature set includes structural features and local features that are robust to registration errors; Perform feature decomposition on the feature set to separate the invariant features that are unrelated to the registration error and the variable features that are affected by the registration error; Construct a registration error correlation graph, which uses invariant features as nodes and spatial deviations between variable features as edge weights to characterize the error correlation relationship between different modal features. Based on the correlation graph, the invariant features and variable features are recombined and fused. In the registration error area, the unreliability of local features is compensated by structural features to obtain fused features. The fused features are input into the preset interpretation model to obtain the image information interpretation results of the target object.
2. The image information interpretation method based on multimodal feature fusion of medical images according to claim 1, characterized in that: Feature decomposition is achieved by analyzing the sensitivity of features to registration errors, including: The features that remain stable when the registration error changes are selected as invariant features. The invariant features include the overall morphological features of the organs and the relative positional relationship features between different organs. Screen out features that change significantly with the registration error as variable features, including pixel distribution features and subtle texture features in local areas; During the feature decomposition process, the sensitivity level of the feature is determined by comparing the consistency of the same feature under different registration error levels.
3. The image information interpretation method based on multimodal feature fusion of medical images according to claim 1, characterized in that: The construction of the registration error correlation map includes: Core nodes are selected from the structural features of each modality. Core nodes are landmark anatomical features that are robust to registration errors, including bone contours and the direction of main blood vessels. The three-dimensional coordinate difference of the local features corresponding to different modes is calculated as the spatial deviation, and the spatial deviation is used as the basis for calculating the edge weight. The edge weight is positively correlated with the deviation size. A multi-scale association structure consisting of organ-level coarse scale and tissue-level fine scale is constructed. The organ-level coarse scale is formed by connecting the corresponding organ structured feature nodes, and its edge weight reflects the overall registration deviation. The tissue-level fine scale is formed by connecting the local features of the substructure within the organ with the structured feature nodes, and its edge weight reflects the local registration error. The multi-scale association structure is combined to form a registration error association graph.
4. The image information interpretation method based on multimodal feature fusion of medical images according to claim 1, characterized in that: The screening of core nodes should give priority to anatomically iconic invariant features; the setting of edge weights should highlight feature associations with significant errors, that is, when the spatial deviation between variable features increases, the edge weight value increases accordingly to strengthen the error representation; in the multi-scale association structure, the coarse-scale association at the organ level is reflected as the alignment deviation of the overall position of the organ corresponding to different modalities, and the fine-scale association at the tissue level is reflected as the alignment error of the local area of the substructure inside the organ.
5. The image information interpretation method based on multimodal feature fusion of medical images according to claim 4, characterized in that: In the recombination and fusion step, the judgment of the unreliability of local features specifically includes: Calculating the spatial matching degree between the local feature and the surrounding structured features. The matching degree is evaluated by the consistency of the spatial distribution patterns of the local feature and the structured features, including whether the spatial coordinates of the local feature fall within the anatomically reasonable range of the structured features. When the spatial matching degree is lower than a first preset threshold, the local feature is marked as a potentially unreliable feature. For the marked potentially unreliable features, the registration error level of the region in which they are located is combined to make another judgment. If the potentially unreliable feature is located in a high error region and is significantly different from the corresponding local features of other modalities in the region, it is confirmed as an unreliable feature. For local features that are confirmed to be unreliable, go back to the feature extraction stage, re-extract the structural features of the area where the local feature is located and compare them. If N consecutive times, N ≥ 2 comparison results show that the spatial matching degree is lower than the first preset threshold, then the unreliability of the local feature is locked; The compensation method includes building anatomical constraint rules based on the spatial distribution law of structural features in unreliable areas, and screening, correcting or replacing unreliable local features based on the rules to make the integration results conform to anatomical logic; Repeat the above judgment process until the proportion of unreliable features is lower than the preset proportion of the total feature quantity, or when the recognition results of unreliable features in two consecutive judgment results are consistent, the judgment process is stopped.
6. The image information interpretation method based on multimodal feature fusion of medical images according to claim 4, characterized in that: Reorganization and integration will be carried out in different regions, specifically: Based on the distribution range of edge weights in the registration error correlation graph, a weight limit is set, and the area where the edge weight value exceeds the limit is divided into the area with large error, and the rest are the area with small error; For areas with smaller errors, more detailed information of local features is retained, and the matching degree between local features and structural features is calculated using structural features as an auxiliary reference. When the matching degree is lower than the first matching threshold, the local feature is marked as a feature to be verified, and the process returns to the feature extraction step to re-extract the local feature and calculate the matching degree again until the matching degree is no lower than the first matching threshold or the preset number of extractions is reached; For areas with large errors, the matching degree between local features and structured features is calculated based on the structured features. Only local features with a matching degree not lower than the second matching threshold are retained. Local features with a matching degree lower than the second matching threshold are marked as features to be optimized. They are adjusted and optimized based on the structured features and the matching degree is recalculated until the matching degree is no lower than the second matching threshold or the preset number of optimizations is reached. The fusion results of the areas with smaller errors and larger errors are checked as a whole, and the consistency of the fusion features is judged by the spatial continuity of the structured features. If there are incoherent features, return to the feature retention step of the corresponding area and reprocess until the overall fusion features meet the consistency requirements.
7. The image information interpretation method based on multimodal feature fusion of medical images according to claim 3, characterized in that: Structured features that are robust to registration errors include: The organ contour features determined based on anatomical atlas can reflect the overall shape and boundary range of the organ; The connection relationship characteristics between different anatomical structures, such as the attachment relationship between blood vessels and organs, and the distribution characteristics of tissue spaces.
8. The image information interpretation method based on multimodal feature fusion of medical images according to claim 1, characterized in that: Local features include: The grayscale distribution characteristics of a local area in an image can reflect the difference in tissue density in that area; The regular arrangement characteristics of pixels in a local area can reflect the microstructure of the tissue.
9. The image information interpretation method based on multimodal feature fusion of medical images according to claim 8, characterized in that: When medical image data of some modalities is missing, the method further includes: Based on the structural features of the acquired modality, key information related to the structural features in the missing modality is inferred; The inferred key information is combined with the acquired feature set to perform feature decomposition and reorganization fusion to ensure that the fusion process is not significantly affected by modality loss.
10. The image information interpretation method based on multimodal feature fusion of medical images according to claim 1, characterized in that: Also includes: In the area with smaller errors, when local features conflict with structural features, the probability of occurrence of the conflicting features in the clinical case database is calculated. If it is lower than the rare threshold, the structural features are used to correct the local features. In areas with large errors, cross-modal verification is performed on local features that match the structured features. When the response value difference exceeds the tolerance threshold, the structured features are used to guide the reconstruction of local features. After regional fusion, the fused features are divided into high, medium and low reliability levels according to the matching stability between local features and structural features; Processing according to anatomical constraint rules: Highly reliable features are directly retained, moderately reliable features are retained or modified after verification with adjacent structural features, and low-reliability features are reconstructed according to the topology of structural features; The above process is repeated until the difference of the fused features of two consecutive iterations is lower than the convergence threshold and the matching degree with the anatomical template exceeds the threshold.
Citation Information
Cited By
Traditional Chinese medicine meridian detection method and system based on multi-modal data fusion
CN120878097A