Image fusion method, device and medium based on multi-modal images

CN115908516BActive Publication Date: 2026-08-21ZINGBOT (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211444252.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2026-08-21
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

[0003]发明人发现,尽管已有厂商推出了CTA与X光成像融合显示的功能(如GE ASSIST、Philips Vessel Navigation等),但大部分产品采用的是简单的刚性配准融合,术前的三维模型与术中的组织器官的实际的形态有较大的差异,二维与三维解剖学信息未能充分互补,导致融合影像未能带来最佳的引导效果,并且,由于术前采集CTA时和术中获取X光影像时静态和动态不完全匹配,患者姿态变化、组织器官形变不同,融合显示存在较大的误差,应用于术中导航效率较低

Benefits of technology

[0045] As can be seen from the above technical solution, this application provides an image fusion method, device and medium based on multimodal images. It performs multi-label segmentation on the acquired three-dimensional static image data through deep learning algorithms, and realizes measurement and three-dimensional reconstruction on the basis of segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The fusion display result is corrected and the adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain the fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908516B_ABST
    Figure CN115908516B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of image fusion method, equipment and medium based on multi-modal image, method includes: three-dimensional static image is carried out multi-label segmentation, according to the result of the multi-label segmentation, corresponding three-dimensional model reconstruction is carried out, and the three-dimensional model after reconstruction is obtained;The three-dimensional static image and at least one two-dimensional dynamic image are registered, corresponding deformation and displacement are determined and mapped to the three-dimensional model in real time to carry out model adaptive deformation, and the three-dimensional model after correction is obtained;The three-dimensional model after correction is superimposed to the two-dimensional dynamic image, and the fusion image is obtained;The application can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, specifically to an image fusion method, device, and medium based on multimodal images. Background Technology

[0002] In the field of vascular surgery, surgeons typically perform the procedure under the guidance of X-ray imaging, supplemented by the injection of contrast agents to obtain digital subtraction angiography. They then determine how to continue advancing and placing surgical instruments based on visual observation and experience. Usually, CTA imaging and X-ray imaging are independent of each other during the procedure.

[0003] The inventors discovered that although some manufacturers have launched CTA and X-ray imaging fusion display functions (such as GE ASSIST, Philips Vessel Navigation, etc.), most products use simple rigid registration fusion. There is a large difference between the preoperative 3D model and the actual morphology of the tissues and organs during the operation. The 2D and 3D anatomical information are not fully complementary, resulting in the fused image failing to bring the best guidance effect. Furthermore, due to the incomplete matching of static and dynamic aspects when acquiring CTA images preoperatively and when acquiring X-ray images intraoperatively, and the differences in patient posture changes and tissue and organ deformation, the fusion display has a large error, resulting in low efficiency when applied to intraoperative navigation. Summary of the Invention

[0004] To address the problems in the prior art, this application provides an image fusion method, device, and medium based on multimodal images, which can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0005] To solve at least one of the above problems, this application provides the following technical solution:

[0006] Firstly, this application provides an image fusion method based on multimodal images, including:

[0007] Multi-label segmentation is performed on the 3D static image, and the corresponding 3D model is reconstructed based on the results of the multi-label segmentation to obtain the reconstructed 3D model.

[0008] The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement, and then mapped to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model.

[0009] The corrected 3D model is superimposed on the 2D dynamic image to obtain a fused image.

[0010] Furthermore, prior to performing multi-label segmentation on the 3D static image, the following steps are included:

[0011] Determine the number of fields of view for the input 3D static image;

[0012] If there are multiple 3D static images with multiple fields of view, then the 3D static images with multiple fields of view are fused to obtain the 3D static image after image fusion.

[0013] Further, the step of image fusion of the three-dimensional static images from the multiple fields of view to obtain the fused three-dimensional static image includes:

[0014] The three-dimensional static images of the multiple fields of view are resampled according to a preset axial distance;

[0015] Based on the resampled three-dimensional static images of multiple fields of view and their respective spatial positions, image fusion is performed to obtain the three-dimensional static image after image fusion.

[0016] Further, the resampling of the three-dimensional static images of the multiple fields of view according to a preset axial distance includes:

[0017] Based on the preset axial distance and the number of voxels and corresponding physical length of the three-dimensional static images of the multiple fields of view, the three-dimensional static images of the multiple fields of view are resampled to obtain the three-dimensional static images after resampling.

[0018] Further, the step of fusing the resampled three-dimensional static images of multiple fields of view with their respective spatial positions to obtain the fused three-dimensional static image includes:

[0019] The correspondence between the three-dimensional static images of multiple fields of view is determined based on the coordinate axis positions and spatial offset information of different cross sections of the resampled three-dimensional static images.

[0020] Based on the correspondence between the three-dimensional static images of the multiple fields of view, image fusion is performed to obtain the three-dimensional static image after image fusion.

[0021] Furthermore, the step of performing multi-label segmentation on the 3D static image and reconstructing the corresponding 3D model based on the results of the multi-label segmentation to obtain the reconstructed 3D model includes:

[0022] The three-dimensional static image is segmented into multiple labels according to a preset deep learning semantic segmentation algorithm to obtain the multi-label segmentation result.

[0023] After fusing the results of the multi-label segmentation, the corresponding 3D model is reconstructed to obtain the reconstructed 3D model.

[0024] Furthermore, the step of performing multi-label segmentation on the 3D static image according to a preset deep learning semantic segmentation algorithm to obtain the multi-label segmentation result includes:

[0025] Extracting features from 3D static images;

[0026] The preset deep learning semantic segmentation algorithm is trained using a preset classifier and loss function. Then, multi-label segmentation is performed based on the deep learning semantic segmentation algorithm trained by the model and the features of the three-dimensional static image to obtain the multi-label segmentation result.

[0027] Further, the registration of the three-dimensional static image and at least one two-dimensional dynamic image to determine the corresponding deformation and displacement includes:

[0028] Convert two-dimensional dynamic images into direct digital X-ray images;

[0029] Image feature matching is performed on the direct digital X-ray image and the three-dimensional static image, and the corresponding linear transformation matrix is ​​determined based on the result of the image feature matching, wherein the linear transformation matrix contains the corresponding deformation and displacement information.

[0030] Furthermore, the step of registering the three-dimensional static image with at least one two-dimensional dynamic image to determine the corresponding deformation and displacement includes:

[0031] Feature extraction is performed on the three-dimensional static image and at least one two-dimensional dynamic image;

[0032] The corresponding deformation field is determined based on the feature extraction results, wherein the deformation field contains corresponding deformation and displacement information.

[0033] Furthermore, determining the corresponding deformation and displacement includes:

[0034] The registered image is segmented into target parts according to a preset video semantic segmentation algorithm;

[0035] Based on the segmentation results of the target region, the segmentation results of consecutive frames are compared to determine the real-time displacement and deformation information.

[0036] Further, the step of determining the corresponding deformation and displacement and mapping them to the three-dimensional model in real time to perform adaptive deformation of the model and obtain the corrected three-dimensional model includes:

[0037] The corresponding reference section in the three-dimensional model is determined based on the determined deformation and displacement;

[0038] The deformation and displacement are mapped to the three-dimensional model in real time based on the reference section and the set center line to obtain the corrected three-dimensional model.

[0039] Further, before mapping the deformation and displacement to the three-dimensional model in real time according to the reference section and the set centerline, the process includes:

[0040] Based on the results of the multi-label segmentation, the boundaries of each region of the target part are extracted, and the maximum inscribed sphere corresponding to each region boundary is obtained based on the Euclidean distance.

[0041] The intersection point of the plane containing the boundary of the region and the centerline is determined based on the center of the inscribed sphere, and the centerline is determined based on the line connecting the intersection points.

[0042] Secondly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the image fusion method based on multimodal images.

[0043] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the image fusion method based on multimodal images.

[0044] Fourthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the image fusion method based on multimodal images.

[0045] As can be seen from the above technical solution, this application provides an image fusion method, device and medium based on multimodal images. It performs multi-label segmentation on the acquired three-dimensional static image data through deep learning algorithms, and realizes measurement and three-dimensional reconstruction on the basis of segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The fusion display result is corrected and the adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain the fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is one of the flowcharts illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0048] Figure 2 This is the second flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0049] Figure 3 This is the third flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0050] Figure 4 This is the fourth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0051] Figure 5 This is the fifth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0052] Figure 6 This is the sixth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0053] Figure 7 This is the seventh flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0054] Figure 8 This is the eighth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0055] Figure 9 This is the ninth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0056] Figure 10 This is the tenth flowchart illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0057] Figure 11 This is eleventh of the flowcharts illustrating the image fusion method based on multimodal images in the embodiments of this application;

[0058] Figure 12 This is one of the schematic diagrams of multi-view three-dimensional static image fusion in a specific embodiment of this application;

[0059] Figure 13 This is the second schematic diagram of multi-view three-dimensional static image fusion in a specific embodiment of this application;

[0060] Figure 14 This is a schematic diagram of spatial position fusion in a specific embodiment of this application;

[0061] Figure 15 This is a schematic diagram of the reconstruction of a three-dimensional model in a specific embodiment of this application;

[0062] Figure 16This is a schematic diagram of rigid registration in a specific embodiment of this application;

[0063] Figure 17 This is a schematic diagram of flexible registration in a specific embodiment of this application;

[0064] Figure 18 This is a schematic diagram of a deformation displacement determination method in a specific embodiment of this application;

[0065] Figure 19 This is a schematic diagram of the model adaptive deformation method in a specific embodiment of this application;

[0066] Figure 20 This is a schematic diagram of a centerline determination method in a specific embodiment of this application;

[0067] Figure 21 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0069] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0070] The surgical medical terms used in the description of the embodiments in this application are only for the purpose of illustrating the image fusion technology solution of this application more clearly, and the protection claimed in this application is only for the technology solution itself.

[0071] In view of the problems existing in the prior art, this application provides an image fusion method, device and medium based on multimodal images. The method uses a deep learning algorithm to perform multi-label segmentation on the acquired three-dimensional static image data, and realizes measurement and three-dimensional reconstruction on the basis of segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The fusion display result is corrected and the adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain a fused image. This method can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0072] To accurately fuse and match 3D static image data and 2D dynamic image data, this application provides an embodiment of an image fusion method based on multimodal imagery, see [link to embodiment]. Figure 1 The image fusion method based on multimodal images specifically includes the following:

[0073] Step S101: Perform multi-label segmentation on the three-dimensional static image, and reconstruct the corresponding three-dimensional model based on the results of the multi-label segmentation to obtain the reconstructed three-dimensional model.

[0074] Optionally, this application can automatically perform multi-label segmentation on the acquired three-dimensional static images (such as CTA (Computed Tomography Axial Transmission) image data) using deep learning algorithms, and realize measurement and three-dimensional reconstruction based on the segmentation to form an accurate static three-dimensional model.

[0075] Step S102: Register the three-dimensional static image with at least one two-dimensional dynamic image, determine the corresponding deformation and displacement, and map them to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model.

[0076] Optionally, this application can register a three-dimensional static image with at least one two-dimensional dynamic image, for example, rigidly register CTA image data with X-ray imaging, and fuse the static three-dimensional model and the two-dimensional X-ray imaging for display.

[0077] Meanwhile, this application can also perform real-time and synchronous flexible registration of static 3D models based on changes identified in each 2D dynamic image data, transforming the 3D static model into an adaptive dynamic model.

[0078] Step S103: Superimpose the corrected 3D model onto the 2D dynamic image to obtain a fused image.

[0079] Optionally, this application can correct the fusion display result by superimposing the corrected three-dimensional model (i.e., the adaptive dynamic model) onto the two-dimensional dynamic image to obtain a fused image.

[0080] As can be seen from the above description, the image fusion method based on multimodal images provided in this application can perform multi-label segmentation on the acquired three-dimensional static image data through deep learning algorithms, and realize measurement and three-dimensional reconstruction on the basis of segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The fusion display result is corrected and the adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain a fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0081] In order to enable the judgment and fusion of multi-view 3D images, in one embodiment of the image fusion method based on multimodal images in this application, see [link to relevant documentation]. Figure 2 It can also specifically include the following:

[0082] Step S201: Determine the number of fields of view for the input 3D static image.

[0083] Step S202: If there are multiple three-dimensional static images of multiple fields of view, then perform image fusion on the three-dimensional static images of the multiple fields of view to obtain the three-dimensional static image after image fusion.

[0084] Optional, see Figure 12 This application can determine whether multiple FOV (Field of View) CTA data exist from the input CTA data. If so, the multi-FOV CTA data is first fused, and then the fused CTA data is segmented to output the segmentation result. If not, multi-label segmentation is directly performed on the current CTA data, and the segmentation result is output. The segmentation result will be used for 3D reconstruction and centerline extraction. The 3D reconstruction will output a static 3D model based on the segmentation result, and the centerline extraction will be based on the virtual centerlines of each tissue from the segmentation result.

[0085] To accurately fuse multi-view 3D images, in one embodiment of the image fusion method based on multimodal images in this application, see [link to relevant documentation]. Figure 3 It can also specifically include the following:

[0086] Step S301: Resample the three-dimensional static images of the multiple fields of view according to the preset axial distance.

[0087] Step S302: Perform image fusion based on the resampled three-dimensional static images of multiple fields of view and their respective spatial positions to obtain the three-dimensional static image after image fusion.

[0088] Optional, see Figure 13 When multiple FOV CTA data exist, this application selects the CTA with the largest FOV as a reference and fuses the remaining FOV CTAs with it one by one. During the fusion process, CTA FOV1 and CTA FOV2 are first resampled, each resampled to a size of 1 mm along each axis. After resampling, fusion is performed based on the spatial positions of the two CTAs.

[0089] In order to enable accurate resampling, one embodiment of the image fusion method based on multimodal images in this application may further include the following:

[0090] Based on the preset axial distance and the number of voxels and corresponding physical length of the three-dimensional static images of the multiple fields of view, the three-dimensional static images of the multiple fields of view are resampled to obtain the three-dimensional static images after resampling.

[0091] Understandably, for example, in different endovascular surgeries, surgeons will collect CTA data at different FOVs (fields of view) before surgery, depending on the needs. Typically, large FOV aortic CTA is used for surgical approach planning, while small FOV is used for lesion identification and measurement. To facilitate subsequent image fusion, the CTA data from different FOVs need to be fused during the preoperative procedure. This allows for the use of the fused CTA data (segmentation results) for subsequent 3D reconstruction, centerline and centroid extraction, and intraoperative image fusion.

[0092] Optionally, this application can resample CTA data of different FOVs to make the voxel size and coarseness uniform, for example, resample the physical space size of each axis of different FOV CTAs to 1mm.

[0093] See below for the specific formula:

[0094]

[0095] Where N represents the number of voxels, S represents the physical length represented by the voxels, i represents the axis of the current resampling, o represents original, i.e. the original value; and n represents new, i.e. the value after resampling.

[0096] To enable accurate fusion based on spatial location, in one embodiment of the image fusion method based on multimodal images in this application, see [link to relevant documentation]. Figure 4 It can also specifically include the following:

[0097] Step S401: Determine the correspondence between the three-dimensional static images of multiple fields of view based on the coordinate axis positions and spatial offset information of different cross sections of the resampled three-dimensional static images of multiple fields of view.

[0098] Step S402: Perform image fusion based on the correspondence between the three-dimensional static images of the multiple fields of view to obtain the three-dimensional static image after image fusion.

[0099] Optional, see Figure 14 This application may select the CTA with the largest FOV as a reference and merge other FOV CTAs into the large FOV CTA.

[0100] Understandably, although the FOVs of each CTA are different, since they are scans of the same patient performed in a single examination, there are spatial correspondences between the different FOVs. These correspondences can be determined in resampled CTAs using spatial information such as the x, y, and z axis positions and offsets of different cross-sections, allowing for direct fusion.

[0101] To accurately construct the 3D model, in one embodiment of the image fusion method based on multimodal imagery in this application, see [link to relevant documentation]. Figure 5 It can also specifically include the following:

[0102] Step S501: Perform multi-label segmentation on the 3D static image according to the preset deep learning semantic segmentation algorithm to obtain the multi-label segmentation result.

[0103] Step S502: After fusing the results of the multi-label segmentation, the corresponding three-dimensional model is reconstructed to obtain the reconstructed three-dimensional model.

[0104] Optional, see Figure 15 When multiple FOV CTA datasets exist, this application can train a segmentation task specific to each FOV CTA, using the corresponding trained deep learning semantic segmentation algorithm to perform the segmentation task and generate segmentation results. After obtaining the multi-label segmentation results of multiple FOV CTAs, this application can perform multi-FOV segmentation result fusion. Based on the obtained multi-FOV segmentation results, a static 3D model is generated through 3D reconstruction.

[0105] To achieve accurate multi-label segmentation, in one embodiment of the image fusion method based on multimodal images in this application, see [link to relevant documentation]. Figure 6 It can also specifically include the following:

[0106] Step S601: Extract the features of the three-dimensional static image.

[0107] Step S602: Train a preset deep learning semantic segmentation algorithm model based on a preset classifier and loss function, and perform multi-label segmentation based on the deep learning semantic segmentation algorithm trained by the model and the features of the three-dimensional static image to obtain the multi-label segmentation result.

[0108] Optionally, for different FOVs or for different CTA data to be segmented within the same FOV, this application can equip each of them with a deep learning-based multi-label segmentation algorithm model to perform multi-label segmentation tasks for that CTA data.

[0109] Each segmentation algorithm model includes an image feature extraction part and a classifier part. It can use mainstream 2D ​​segmentation networks such as UNet, FCN, DeepLab and 3D segmentation networks such as VNet, 3D UNet, etc. The loss function can include, but is not limited to, cross-entropy loss function, multi-label Dice loss function, etc.

[0110] Therefore, this application can selectively collect a large amount of labeled data to train the corresponding model. After training, the model can automatically perform inference and generate multi-label segmentation results when new CTA data is input. If CTAs of multiple FOVs exist, they are fused to obtain the final multi-label segmentation result, and 3D reconstruction visualization is achieved accordingly.

[0111] To enable accurate rigid registration, in one embodiment of the image fusion method based on multimodal images in this application, see [link to relevant documentation]. Figure 7 It can also specifically include the following:

[0112] Step S701: Convert the two-dimensional dynamic image into a direct digital X-ray image.

[0113] Step S702: Perform image feature matching on the direct digitized X-ray image and the three-dimensional static image, and determine the corresponding linear transformation matrix based on the result of the image feature matching, wherein the linear transformation matrix contains the corresponding deformation and displacement information.

[0114] Optional, see Figure 16 This application, without considering any displacement or deformation, aims solely to overlay and display a 3D model onto DSA data. Feature matching can employ methods including, but not limited to, SIFT and Harris feature matching.

[0115] In this process, the CTA is first used to generate a DDR (Digital Representation Matrix). The generated DDR is then matched with the DSA (Digital Subtraction Animation) image, and a linear transformation matrix is ​​calculated based on the feature matching results. The DDR is then transformed according to the obtained transformation matrix, and the segmentation results are correspondingly transformed and superimposed on the DSA image through the projection of the DDR.

[0116] To enable accurate flexible registration, in one embodiment, see [link to embodiment]. Figure 8 It can also specifically include the following:

[0117] Step S801: Extract features from the three-dimensional static image and at least one two-dimensional dynamic image.

[0118] Step S802: Determine the corresponding deformation field based on the feature extraction result, wherein the deformation field includes corresponding deformation and displacement information.

[0119] Optional, see Figure 17 The embodiments can also be used to match CTA data with dynamic DSA images, TTE images, and IVUS images to map the tissue and organ deformations identified based on DSA, TTE, and IVUS images into a three-dimensional model and perform adaptive deformation correction on the three-dimensional model.

[0120] Specifically, the model takes two modalities of images to be registered as input, extracts features, infers the deformation field, and obtains a spatial transformation based on the deformation field to transform the images to be registered. Unsupervised algorithms, including but not limited to VoxelMorph, can be used.

[0121] In order to accurately determine the deformation displacement, in one embodiment, see [reference needed]. Figure 9 It can also specifically include the following:

[0122] Step S901: Segment the target parts of the registered image according to the preset video semantic segmentation algorithm.

[0123] Step S902: Compare the segmentation results of consecutive frames based on the segmentation results of the target part to determine the real-time displacement and deformation information.

[0124] Optional, see Figure 18 The embodiment can also use dynamic video as input and output real-time target organization segmentation results. The core algorithm is a video semantic segmentation algorithm, which can make full use of temporal context information by utilizing structures including but not limited to RNN, GRU, LSTM, etc., to achieve accurate segmentation of target organizations of interest in dynamic video. After obtaining the segmentation results, the real-time displacement and deformation information of the target organization can be calculated by comparing the segmentation results of consecutive frames.

[0125] To accurately perform adaptive deformation of the model, in one embodiment of the image fusion method based on multimodal imagery, see [link to relevant documentation]. Figure 10 It can also specifically include the following:

[0126] Step S1001: Determine the corresponding reference section in the three-dimensional model based on the determined deformation and displacement.

[0127] Step S1002: Based on the reference section and the set centerline, the deformation and displacement are mapped to the three-dimensional model in real time to obtain the corrected three-dimensional model.

[0128] Optional, see Figure 19When registering dynamic DSA / TTE / IVUS video with CTA image, the reference plane for registration is a cross section of the CTA image at the corresponding position. After registration, the deformation of tissues and organs detected in the dynamic DSA / TTE / IVUS video can be directly matched with the corresponding tissue and organ model in the three-dimensional model.

[0129] Understandably, since dynamic DSA / TTE / IVUS videos only observe the dynamics of a single cross-section, it is necessary to apply displacement and deformation to the entire tissue / organ model. Therefore, it is necessary to use a pre-obtained centerline to map the transformation of a plane onto the entire tissue / organ model. That is, to find the components of the displacement and deformation vector of that plane along the plane containing each point on the centerline, and to make corresponding changes to the tissue contour boundary on that plane based on these components.

[0130] To accurately determine the centerline, in one embodiment of an image fusion method based on multimodal imagery, see [reference needed]. Figure 11 It can also specifically include the following:

[0131] Step S1101: Extract the boundaries of each region of the target part based on the results of the multi-label segmentation and obtain the maximum inscribed sphere corresponding to each region boundary based on Euclidean distance.

[0132] Step S1102: Determine the intersection point of the plane containing the boundary of the region and the center line based on the center of the inscribed sphere, and determine the center line based on the line connecting the intersection points.

[0133] Optional, see Figure 20 The purpose of centerline extraction is to extract the "skeleton" of the tissue structure of interest, so that the shape transformation of the corresponding 3D model of each tissue can be performed when the static 3D model is adaptively deformed.

[0134] Optionally, this application may first extract the contour of the segmentation results of the organization of interest. The region boundaries are traversed in a certain order, and the largest inscribed sphere is found for the current boundary using Euclidean distance. The center of this inscribed sphere is the point through which the centerline passes. After the boundary traversal is complete, the line connecting all the found inscribed sphere centers is the desired centerline.

[0135] From a hardware perspective, in order to accurately fuse and match 3D static image data and 2D dynamic image data, this application provides an embodiment of an electronic device for implementing all or part of the image fusion method based on multimodal images. The electronic device specifically includes the following components:

[0136] The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the image fusion device based on multimodal images and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the image fusion method based on multimodal images and the embodiments of the image fusion device based on multimodal images in the embodiments, the contents of which are incorporated herein, and repeated details will not be described again.

[0137] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.

[0138] In practical applications, parts of the image fusion method based on multimodal images can be executed on the electronic device side as described above, or all operations can be completed in the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.

[0139] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0140] Figure 21 This is a schematic block diagram illustrating the system configuration of an electronic device 9600 according to another embodiment of the present invention. Figure 21 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 21 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0141] In one embodiment, the image fusion method based on multimodal images can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0142] Step S101: Perform multi-label segmentation on the three-dimensional static image, and reconstruct the corresponding three-dimensional model based on the results of the multi-label segmentation to obtain the reconstructed three-dimensional model.

[0143] Step S102: Register the three-dimensional static image with at least one two-dimensional dynamic image, determine the corresponding deformation and displacement, and map them to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model.

[0144] Step S103: Superimpose the corrected 3D model onto the 2D dynamic image to obtain a fused image.

[0145] As can be seen from the above description, the electronic device provided in the embodiment performs multi-label segmentation on the acquired three-dimensional static image data through a deep learning algorithm, and realizes measurement and three-dimensional reconstruction based on the segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The result of the fusion display is corrected and the three-dimensional model after adaptive deformation is superimposed on the two-dimensional dynamic image to obtain a fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0146] In another embodiment, the image fusion apparatus based on multimodal images can be configured separately from the central processing unit 9100. For example, the image fusion apparatus based on multimodal images can be configured as a chip connected to the central processing unit 9100, and the image fusion method function based on multimodal images can be implemented through the control of the central processing unit.

[0147] like Figure 21 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 21 All components shown; in addition, the electronic device 9600 may also include Figure 21 For components not shown, please refer to existing technologies.

[0148] like Figure 21 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.

[0149] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.

[0150] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.

[0151] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.

[0152] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device's communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0153] The communication module 9110 is a transmitter / receiver 9110 that transmits and receives signals via the antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which can be the same as in a conventional mobile communication terminal.

[0154] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored sound via the speaker 9131.

[0155] Another embodiment also provides a computer-readable storage medium capable of implementing all steps of the image fusion method based on multimodal images in the above embodiments, wherein the execution subject is a server or client. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the image fusion method based on multimodal images in the above embodiments, wherein the execution subject is a server or client. For example, when the processor executes the computer program, it implements the following steps:

[0156] Step S101: Perform multi-label segmentation on the three-dimensional static image, and reconstruct the corresponding three-dimensional model based on the results of the multi-label segmentation to obtain the reconstructed three-dimensional model.

[0157] Step S102: Register the three-dimensional static image with at least one two-dimensional dynamic image, determine the corresponding deformation and displacement, and map them to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model.

[0158] Step S103: Superimpose the corrected 3D model onto the 2D dynamic image to obtain a fused image.

[0159] As can be seen from the above description, the computer-readable storage medium provided in the embodiments of the present invention performs multi-label segmentation on the acquired three-dimensional static image data through a deep learning algorithm, and realizes measurement and three-dimensional reconstruction based on the segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The result of the fusion is corrected and displayed. The adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain a fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0160] This invention also provides a computer program product capable of implementing all steps of the image fusion method based on multimodal images, where the execution subject is a server or client, as described in the above embodiments. When executed by a processor, this computer program / instruction implements the steps of the image fusion method based on multimodal images. For example, the computer program / instruction implements the following steps:

[0161] Step S101: Perform multi-label segmentation on the three-dimensional static image, and reconstruct the corresponding three-dimensional model based on the results of the multi-label segmentation to obtain the reconstructed three-dimensional model.

[0162] Step S102: Register the three-dimensional static image with at least one two-dimensional dynamic image, determine the corresponding deformation and displacement, and map them to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model.

[0163] Step S103: Superimpose the corrected 3D model onto the 2D dynamic image to obtain a fused image.

[0164] As can be seen from the above description, the computer program product provided in the embodiments of the present invention performs multi-label segmentation on the acquired three-dimensional static image data through deep learning algorithms, and realizes measurement and three-dimensional reconstruction based on the segmentation to form an accurate static three-dimensional model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement and map them to the three-dimensional model in real time for adaptive deformation of the model. The result of the fusion is corrected and displayed. The adaptively deformed three-dimensional model is superimposed on the two-dimensional dynamic image to obtain a fused image. Thus, it can accurately fuse and match three-dimensional static image data and two-dimensional dynamic image data.

[0165] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0166] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0167] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0168] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0169] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. An image fusion method based on multimodal images, characterized in that, The method includes: Multi-label segmentation is performed on the 3D static image, and the corresponding 3D model is reconstructed based on the results of the multi-label segmentation to obtain the reconstructed 3D model. The three-dimensional static image and at least one two-dimensional dynamic image are registered to determine the corresponding deformation and displacement, and then mapped to the three-dimensional model in real time to perform adaptive deformation of the model, thereby obtaining the corrected three-dimensional model. The corrected 3D model is superimposed on the 2D dynamic image to obtain a fused image; The step of performing multi-label segmentation on a 3D static image and reconstructing a corresponding 3D model based on the results of the multi-label segmentation to obtain a reconstructed 3D model includes: performing multi-label segmentation on the 3D static image according to a preset deep learning semantic segmentation algorithm to obtain multi-label segmentation results; and performing model reconstruction on the corresponding 3D model after fusing the results of the multi-label segmentation to obtain a reconstructed 3D model. The step of registering the three-dimensional static image and at least one two-dimensional dynamic image to determine the corresponding deformation and displacement includes: converting the two-dimensional dynamic image into a direct digital X-ray image; performing image feature matching on the direct digital X-ray image and the three-dimensional static image, and determining the corresponding linear transformation matrix based on the image feature matching result, wherein the linear transformation matrix contains the corresponding deformation and displacement information; the step of determining the corresponding deformation and displacement includes: segmenting the registered image into target parts according to a preset video semantic segmentation algorithm; comparing the segmentation results of consecutive frames based on the target part segmentation result to determine real-time displacement and deformation information; the step of determining the corresponding deformation and displacement and mapping it to the three-dimensional model in real time for adaptive deformation of the model to obtain a corrected three-dimensional model includes: determining the corresponding reference section in the three-dimensional model based on the determined deformation and displacement; mapping the deformation and displacement to the three-dimensional model in real time based on the reference section and a set centerline to obtain a corrected three-dimensional model; Before mapping the deformation and displacement to the three-dimensional model in real time based on the reference section and the set centerline, the process includes: extracting the boundaries of each region of the target part based on the results of the multi-label segmentation and obtaining the maximum inscribed sphere corresponding to each region boundary based on the Euclidean distance; determining the intersection point of the plane where the region boundary is located and the centerline based on the center of the inscribed sphere, and determining the centerline based on the line connecting the intersection points.

2. The image fusion method based on multimodal images according to claim 1, characterized in that, Before performing multi-label segmentation on the 3D static image, the following steps are included: Determine the number of fields of view for the input 3D static image; If there are multiple 3D static images with multiple fields of view, then the 3D static images with multiple fields of view are fused to obtain the 3D static image after image fusion.

3. The image fusion method based on multimodal images according to claim 2, characterized in that, The process of fusing the three-dimensional static images from the multiple fields of view to obtain the fused three-dimensional static image includes: The three-dimensional static images of the multiple fields of view are resampled according to a preset axial distance; Based on the resampled three-dimensional static images of multiple fields of view and their respective spatial positions, image fusion is performed to obtain the three-dimensional static image after image fusion.

4. The image fusion method based on multimodal images according to claim 3, characterized in that, The resampling of the three-dimensional static images of the multiple fields of view according to a preset axial distance includes: Based on the preset axial distance and the number of voxels and corresponding physical length of the three-dimensional static images of the multiple fields of view, the three-dimensional static images of the multiple fields of view are resampled to obtain the three-dimensional static images after resampling.

5. The image fusion method based on multimodal images according to claim 3, characterized in that, The step of fusing the resampled three-dimensional static images from multiple fields of view with their respective spatial positions to obtain the fused three-dimensional static image includes: The correspondence between the three-dimensional static images of multiple fields of view is determined based on the coordinate axis positions and spatial offset information of different cross sections of the resampled three-dimensional static images. Based on the correspondence between the three-dimensional static images of the multiple fields of view, image fusion is performed to obtain the three-dimensional static image after image fusion.

6. The image fusion method based on multimodal images according to claim 1, characterized in that, The step of performing multi-label segmentation on the 3D static image according to a preset deep learning semantic segmentation algorithm to obtain the multi-label segmentation result includes: Extracting features from 3D static images; The preset deep learning semantic segmentation algorithm is trained using a preset classifier and loss function. Then, multi-label segmentation is performed based on the deep learning semantic segmentation algorithm trained by the model and the features of the three-dimensional static image to obtain the multi-label segmentation result.

7. The image fusion method based on multimodal images according to claim 1, characterized in that, The step of registering the three-dimensional static image with at least one two-dimensional dynamic image to determine the corresponding deformation and displacement includes: Feature extraction is performed on the three-dimensional static image and at least one two-dimensional dynamic image; The corresponding deformation field is determined based on the feature extraction results, wherein the deformation field contains corresponding deformation and displacement information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the image fusion method based on multimodal images as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the image fusion method based on multimodal images as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • 2D / 3D coronary artery automatic registration method and system, and medium

    CN113935889A