Medical image processing method, system and device and storage medium
Through a dynamic registration method based on video image processing technology, the problem that the prior art cannot achieve preoperative model and intraoperative image registration when the endoscopic fixation is performed but the organ is moved is solved, and the navigation accuracy and reliability are improved.
Patent Information
- Application Number
- CN202510206318.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-10
AI Technical Summary
The existing intraoperative navigation technology cannot register preoperative models and intraoperative images when the endoscopic fixes but the organs move, resulting in reduced navigation accuracy.
Through video image processing technology, based on the front and back continuous video frames and three-dimensional models, the correspondence between video frames and three-dimensional models is determined in real time to achieve dynamic registration.
It realizes that when the endoscopic fixation and organ movement, the preoperative model and intraoperative images can still be accurately registered, improving the accuracy and reliability of navigation.
Smart Images

Figure CN120125631A_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is: 202411282643.X, the filing date of the original application is September 13, 2024, the invention title of the original application is: Medical Image Processing Method, System, Device and Storage Medium, and the entire content of the original application is incorporated herein by reference. Technical Field
[0002] This application belongs to the technical field of medical image processing, and particularly relates to a medical image processing method, system, device and storage medium. Background Art
[0003] Minimally invasive surgery has the advantages of small trauma, fast recovery, and reduced patient pain. As the "eyes" of doctors, endoscopes can effectively help doctors see the lesions clearly during minimally invasive surgery. The surgical field of view under the endoscope is limited, and many lesion blood vessels are hidden under the organ surface, which requires high capabilities and experience of doctors. Intraoperative navigation technology can map the preoperative model of the target organ during the operation and provide real-time guidance to doctors. However, some current navigation systems achieve intraoperative navigation by means of additional hardware devices. The hardware devices may include sensors such as electromagnetic sensors or IMUs. However, such sensors usually can only reflect the movement of the endoscope. When the endoscope moves, the preoperative model and the intraoperative image are rigidly registered according to the movement of the endoscope. However, when the endoscope is fixed and the organ moves, the existing intraoperative navigation technology cannot register the preoperative model and the intraoperative image. Summary of the Invention
[0004] Based on this, the embodiments of this application provide a medical image processing method, system, device and storage medium, which can perform real-time dynamic registration on video frames and three-dimensional models based on video images.
[0005] The embodiments of this application provide a medical image processing system, including:
[0006] An input module, configured to input a video image of a target tissue during surgery;
[0007] A processing module, configured to determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and register the three-dimensional model and the second video frame according to the correspondence, where the first video frame and the second video frame are consecutive video frames before and after;
[0008] A display module, configured to display the registered second video frame and the three-dimensional model in real time.
[0009] In some embodiments, the processing module is configured to determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, including:
[0010] Determine a vector field of pixel movement in the second video frame according to the first video frame and the second video frame;
[0011] Determine the updated positions of the control points in the second video frame according to the vector field;
[0012] Determine the correspondence between the second video frame and the three-dimensional model according to the updated positions of the control points and the correspondence between the first video frame and the three-dimensional model.
[0013] In some embodiments, the processing module is further configured to:
[0014] Update the weight values of each feature point relative to the control points according to the updated positions of the control points in the second video frame.
[0015] In some embodiments, the processing module determines the updated positions of the control points in the second video frame according to the vector field, including:
[0016] Map the feature points in the feature point dense area in the first video frame to the second video frame according to the vector field, and determine the updated positions of the control points in the second video frame according to the weight values of the feature points relative to the control points, where the feature point dense area is an area where the number of feature points in the second video frame is greater than a number threshold;
[0017] Map the control points in the feature point sparse area in the first video frame to the second video frame according to the vector field to obtain the updated positions of the control points in the feature point sparse area, where the feature point sparse area is an area where the number of feature points in the second video frame is less than the number threshold.
[0018] In some embodiments, the processing module determines the vector field of pixel movement in the second video frame according to the first video frame and the second video frame, including:
[0019] For any target pixel in the first video frame, search for the target pixel on the second video frame;
[0020] Determine the displacement field of the target pixel movement according to the position of the target pixel in the first video frame and the position of the target pixel in the second video frame, so as to obtain the vector field of pixel movement in the second video frame.
[0021] In some embodiments, when the first video frame is the first frame, the processing module establishes control point pairs according to the registration result, including:
[0022] Uniformly sample feature points from the first video frame to obtain initial control points;
[0023] Screen out target control points from the initial control points according to the registration result obtained by registering the three-dimensional model and the first video frame, where corresponding control points exist for the target control points in the three-dimensional model to obtain the control point pairs.
[0024] In some embodiments, the processing module is used to determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, including:
[0025] Extract feature points from the second video frame and extract feature points from the first video frame;
[0026] Match the feature points extracted from the first video frame with the feature points extracted from the second video frame to obtain a matching result;
[0027] Determine the updated positions of the control points according to the matching result;
[0028] Determine the correspondence between the second video frame and the three-dimensional model according to the updated positions and the correspondence between the first video frame and the three-dimensional model.
[0029] In some embodiments, the processing module registers the three-dimensional model and the second video frame according to the correspondence, including:
[0030] Determine the relative position relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, where the intraoperative image sensor is used to collect video images of the target tissue during the operation;
[0031] Register the three-dimensional model and the second video frame according to the elastic deformation and the relative position relationship.
[0032] In some embodiments, determining the relative position relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model includes:
[0033] Determine the relative position relationship between the target tissue and the intraoperative image sensor according to the correspondence;
[0034] Determine the displacement change information of the rigid motion of the 3D model according to the relative position relationship;
[0035] Determine the virtual force by which the 3D model changes according to the displacement change information;
[0036] Simulate the elastic deformation of each feature point in the 3D model according to the virtual force and the constraint conditions of elastic deformation to obtain the elastic deformation of each feature point of the 3D model, or;
[0037] Calculate the elastic deformation of each feature point in the 3D model by using the finite element analysis method according to the virtual force and the biomechanical model of the target tissue.
[0038] An embodiment of the present application provides a medical image processing method, including:
[0039] Obtain the video image of the target tissue during the operation;
[0040] Determine the correspondence between the second video frame and the 3D model according to the first video frame, the 3D model of the target tissue, and the second video frame, and register the 3D model with the second video frame according to the correspondence, where the first video frame and the second video frame are consecutive video frames before and after;
[0041] Real-time display the registered second video frame and the 3D model.
[0042] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above is implemented.
[0043] An embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0044] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, it causes the electronic device to execute the method described in any one of the above.
[0045] A medical image processing system provided by an embodiment of the present application inputs video images of a target tissue during surgery through an input module; a processing module determines the correspondence between a second video frame and the three-dimensional model according to a first video frame, the three-dimensional model of the target tissue, and the second video frame, and registers the three-dimensional model and the second video frame according to the correspondence, where the first video frame and the second video frame are respectively consecutive video frames before and after; a display module displays the registered second video frame and the three-dimensional model in real time, and can perform real-time dynamic registration on the video frame and the three-dimensional model based on the video image. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Hereinafter, the present application will be described in more detail according to embodiments and with reference to the drawings.
[0047] Figure 1 It is a schematic structural diagram of a medical image processing system provided by an embodiment of the present application;
[0048] Figure 2 It is a schematic implementation flowchart of a medical image processing method provided by an implementation of the present application;
[0049] Figure 3 It is a schematic flowchart of a medical image processing method provided by an embodiment of the present application;
[0050] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0051] In the drawings, the same components are denoted by the same reference numerals, and the drawings are not drawn to actual scale. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0053] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0054] If similar descriptions such as "first / second / third" appear in the application documents, the following explanations shall be added. In the following descriptions, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0056] Based on the problems in the related art, embodiments of the present application provide a medical image processing system. Each module included in the system of the embodiments of the present application, as well as each unit included in each module, can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. During implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0057] Embodiments of the present application provide a medical image processing system. Figure 1 As a schematic structural diagram of a medical image processing system provided by the embodiments of the present application, as Figure 1 shown, the medical image processing system 100 includes: an input module 101, a processing module 102, and a display module 103. The input module is used to input video images of a target tissue during surgery; the processing module is used to determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and register the three-dimensional model and the second video frame according to the correspondence. The first video frame and the second video frame are respectively consecutive video frames before and after; the display module is used to display the registered second video frame and the three-dimensional model in real time.
[0058] In the embodiments of the present application, the target organ can be an organ in the human body, and the target organ can be: the liver, the heart, the stomach, etc.
[0059] In the embodiments of the present application, during the surgical operation, the surgical operator can use the image acquisition module to collect intraoperative images of the target organ in real time, thereby obtaining video images. The image acquisition module can be an endoscope. The video images include two consecutive frames, and any two consecutive frames can be the first video frame and the second video frame.
[0060] In the embodiments of the present application, the three-dimensional model of the target tissue can be a three-dimensional model established for the target tissue before surgery.
[0061] In the embodiments of the present application, the processing module can determine the correspondence between the second video frame and the three-dimensional model according to the pixel movement vector field of the video frame, and the vector field of pixel movement can be determined based on the first video frame and the second video frame.
[0062] In some embodiments, the processing module can also use feature matching to determine the correspondence between the second video frame and the three-dimensional model.
[0063] In the embodiments of the present application, the correspondence between the second video frame and the three-dimensional model can be the correspondence between the features in the second video frame and the features in the three-dimensional model. Since the features in the second video frame correspond to the pixel points, the correspondence can also be the correspondence between the pixel points in the second video frame and the features in the three-dimensional model.
[0064] In the embodiments of the present application, registration can include: rigid registration and elastic registration.
[0065] In the embodiments of the present application, rigid registration is a method of matching two data sets so that they are aligned in the same coordinate system. When rigidly registering the second video frame and the three-dimensional model, rigid transformations such as rotation, translation, and scaling can be performed on the three-dimensional model to make the three-dimensional model and the second video frame coincide as much as possible in spatial position. This type of registration is suitable for situations where the position and shape need to remain unchanged, and it can help the doctor align the previously planned surgical plan with the anatomical structure of the actual patient.
[0066] In the embodiments of the present application, elastic registration is used to perform more precise and flexible matching of information from different data sources, taking into account more complex deformations and distortions of the target organ. Nonlinear deformations and local changes can be performed on the three-dimensional model to achieve elastic registration of the three-dimensional model and the second video frame. In the embodiments of the present application, elastic registration of the three-dimensional model planned by the doctor before surgery and the intraoperative images is to more precisely align the structures between the two to provide more accurate guidance and auxiliary information.
[0067] In the embodiments of the present application, the display module can display the registered three-dimensional model and the second video frame in real time. The display module displaying the registered three-dimensional model and the second video frame may include: superimposing and displaying the three-dimensional model and the second video frame.
[0068] A medical image processing system provided by the embodiments of the present application inputs a video image of a target tissue during surgery through an input module; a processing module determines the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and registers the three-dimensional model and the second video frame according to the correspondence. The first video frame and the second video frame are respectively two consecutive video frames; the display module displays the registered second video frame and the three-dimensional model in real time, and can perform real-time dynamic registration on the video frame and the three-dimensional model based on the video image.
[0069] In the embodiments of the present application, since the video image can be collected through an endoscope, and the registration of the video frame and the three-dimensional model can be achieved through the video image, the registration can be achieved without relying on other hardware devices (such as electromagnetic sensors or IMU and other sensors). In addition, in the case where the endoscope is fixed and the organ moves, the registration of the three-dimensional model and the video frame can also be achieved, and the movement of the target organ itself can be accurately tracked.
[0070] In some embodiments, the processing module is used to determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, including:
[0071] The processing module determines the vector field of pixel movement in the second video frame according to the first video frame and the second video frame, determines the updated position of the control points in the second video frame according to the vector field, and determines the correspondence between the second video frame and the three-dimensional model according to the updated position of the control points and the correspondence between the first video frame and the three-dimensional model.
[0072] In the embodiments of the present application, the vector field of pixel movement is used to describe the movement of pixels between two video frames. The displacement of the pixels can be calculated by comparing the position changes of the pixels in the two video frames, and the vector field of the movement can be determined through the displacement.
[0073] In the embodiments of the present application, the control points can be feature points set in the video frame.
[0074] In the embodiments of the present application, the correspondence between the first video frame and the three-dimensional model may be established based on the video frames before the first video frame. In some embodiments, if the first video frame is the first frame of the video image, the correspondence between the first video frame and the three-dimensional model may be obtained after a preliminary registration of the three-dimensional model and the first video frame.
[0075] In the embodiments of the present application, based on the updated positions of the control points and the correspondence between the first video frame and the three-dimensional model, the approximate correspondence between the second video frame and the three-dimensional model can be inferred.
[0076] In some embodiments, the processing module may determine the vector field of the pixel movement in the second video frame according to the first video frame and the second video frame, including:
[0077] For any target pixel in the first video frame, the processing module searches for the target pixel in the second video frame, and determines the displacement field of the target pixel movement according to the position of the target pixel in the first video frame and the position of the target pixel in the second video frame, so as to obtain the vector field of the pixel movement in the second video frame.
[0078] In the embodiments of the present application, the target pixel may be any pixel. In the embodiments of the present application, when searching for the target pixel in the second video frame, methods such as sliding window, template matching, and feature matching may be used for searching and matching.
[0079] Exemplarily, taking feature matching as an example, features of the target pixel in the first video frame can be extracted, and features such as color, texture, and shape can be selected. Using the selected features, pixels similar to the target pixel in the first video frame are searched in the second video frame. Similarity measurement methods (such as Euclidean distance, cosine similarity, etc.) can be used to compare the similarity between the target pixel and the candidate pixels in the second video frame. The pixel closest to the target pixel in the first video frame is found in the second video frame according to the similarity measurement, so as to search for the target pixel on the second video frame.
[0080] In the embodiments of the present application, the processing module can find the position of the target pixel in the second video frame by searching for the target pixel in the second video frame. By comparing the position differences of the target pixel in the two video frames, the displacement of the target pixel in space can be calculated to form a displacement field. Combining the displacement fields of all target pixels together, the vector field of the pixel movement in the second video frame can be obtained, and the movement of the pixels in the entire video frame can be described through the vector field.
[0081] In the embodiments of the present application, combining the displacement fields of all target pixels together to obtain the vector of pixel movement in the second video frame can be achieved in the following manner: For each target pixel, calculate the corresponding displacement vector according to the position difference between it in the first video frame and the second video frame. Divide the pixels in the video frame into grids or lattice points so as to apply the displacement information to the entire picture. According to the known displacement information of the target pixels, estimate the displacement information of the remaining pixels on the grid through an interpolation method (such as bilinear interpolation). Combine the displacement information of each pixel point into a vector field of overall pixel movement, thereby describing the movement condition of the pixels in the entire second video frame.
[0082] In the embodiments of the present application, through such a method, the motion information of a single pixel can be integrated onto the entire image, so as to more comprehensively understand the movement of the pixels in the video frame.
[0083] In some embodiments, the processing module determining the updated position of the control points in the second video frame according to the vector field may include:
[0084] The processing module maps the feature points in the feature point dense area in the first video frame to the second video frame according to the vector field, and determines the updated position of the control points in the second video frame according to the weight value of the feature points relative to the control points. The feature point dense area is the area where the number of feature points in the second video frame is greater than the number threshold.
[0085] In the embodiments of the present application, the weight value of the feature points relative to the control points is determined based on the first video frame. Feature points can be searched within the neighborhood of each control point in the first video frame with a preset radius, and the weight of the feature points corresponding to each control point is determined based on the distance between the feature points searched within the neighborhood of each control point and the corresponding control point. In the embodiments of the present application, the closer the distance, the greater the weight, and the farther the distance, the smaller the weight. The preset radius can be configured.
[0086] In the embodiments of the present application, mapping the feature points in the feature point dense area in the first video frame to the second video frame according to the vector field includes: Select some key feature points within the dense feature point area in the first video frame according to the calculated vector field of pixel movement, and determine the positions of these feature points in the second video frame by mapping each feature point to the second video frame according to the movement in the vector field.
[0087] In the embodiments of the present application, for each feature point, its weight value at this position can be determined according to its distance from surrounding control points or other features. Generally, the closer the control point is, the greater its influence on the feature point. According to the position of the feature point in the second video frame and its weight value relative to the control point, the updated position of the control point in the second video frame is determined. The information of the feature point obtained by weighting can effectively determine the position of the control point to better map and track the entire area.
[0088] In the embodiments of the present application, determining the position of the control point through the feature points in the dense area can make the accuracy of the determined position of the control point higher.
[0089] The processing module maps the control points in the sparse area of feature points in the first video frame to the second video frame according to the vector field, and obtains the updated positions of the control points in the sparse area of feature points, where the sparse area of feature points is an area in the second video frame where the number of feature points is less than the number threshold.
[0090] In the embodiments of the present application, the vector field of pixel movement can be used to map the control points in the sparse area of feature points in the first video frame to the second video frame, so as to determine the positions of these control points in the second video frame.
[0091] In the embodiments of the present application, for the sparse area, feature points may not be searched in the neighborhood of the control point, or the accuracy of the feature points cannot be guaranteed. Therefore, the control point and the feature point are decoupled, and the control point is mapped from the first image frame to the second image frame according to the vector field, and the updated position of the updated control point is directly obtained.
[0092] In the embodiments of the present application, the sparse area can be an area where the surface of the target organ is smooth or the anatomical structure is invisible under the endoscope.
[0093] In some embodiments, after the processing module registers the three-dimensional model with the second video frame according to the corresponding relationship, the processing module is further configured to update the weight value of each feature point relative to the control point according to the updated position of the control point in the second video frame.
[0094] In the embodiments of the present application, feature points can be searched in the neighborhood of each control point in the second video frame with a preset radius, and the weight of the feature point corresponding to each control point is determined based on the distance between the feature points searched in the neighborhood of each control point and the corresponding control point, so that the weight value of each feature point relative to the control point can be updated.
[0095] In some embodiments, when the first video frame is the first frame, the processing module establishes control point pairs according to the registration result, including:
[0096] The processing module uniformly samples feature points from the first video frame to obtain initial control points, and filters out target control points from the initial control points according to the registration result obtained by registering the three-dimensional model and the first video frame, so as to obtain the control point pairs, where there are corresponding control points for the target control points in the three-dimensional model.
[0097] In the embodiments of the present application, for uniform sampling of feature points in the first video frame, various feature detection algorithms (such as SIFT, SURF, etc.) can be used to identify some key feature points, so as to obtain initial control points.
[0098] In the embodiments of the present application, the three-dimensional model and the first video frame can be registered. Here, the registration can include: rigid registration and / or elastic registration. After registration, there is a corresponding relationship between the feature points in each first video frame and the feature points of the three-dimensional model.
[0099] In the embodiments of the present application, the target control points corresponding to the initial control points can be found in the registered three-dimensional model, so as to obtain the control point pairs.
[0100] In the embodiments of the present application, after obtaining the control point pairs, the control point pairs can be used for subsequent updated position tracking and deformation estimation.
[0101] In the related art, when performing position tracking, it can only rely on the anatomical features or image features of the target organ itself, with a narrow scope of application and easy to fail. It is difficult to detect feature points on the surface of some smooth target organs, and the uneven distribution of such features will affect the accuracy of intraoperative navigation. In the embodiments of the present application, by setting control points, accurate registration can still be achieved when the feature points fail.
[0102] In the related art, in order to perform registration, tracking markers are installed on the surface of the organ to achieve intraoperative tracking of the target organ. However, the method in the related art introduces additional surgical procedures and is only applicable to the tracking of rigid organs such as bones, and markers cannot be fixed on soft tissue organs such as the liver. The method provided in the embodiments of the present application performs registration by setting control points, can be independent of additional visual markers, and can achieve accurate intraoperative tracking for both rigid organs and soft tissue organs.
[0103] In some embodiments, the processing module is configured to determine the correspondence between the second video frame and the three-dimensional model based on the first video frame, the three-dimensional model of the target tissue, and the second video frame, which may include: The processing module extracts feature points from the second video frame, extracts feature points from the first video frame, matches the feature points extracted from the first video frame with the feature points extracted from the second video frame to obtain a matching result, determines the updated position of the control points according to the matching result, and determines the correspondence between the second video frame and the three-dimensional model according to the updated position and the correspondence between the first video frame and the three-dimensional model.
[0104] In the embodiments of the present application, feature points can be extracted from the first video frame and the second video frame, and the extraction of feature points in the first video frame and the second video frame can be implemented by various feature extraction algorithms such as SIFT, SURF, ORB, etc.
[0105] In the embodiments of the present application, the similarity between the feature points extracted from the first video frame and the feature points extracted from the second video frame can be calculated for matching. The greater the similarity, the better the match; the smaller the similarity, the worse the match.
[0106] In the embodiments of the present application, since the feature points include control points, the updated position of the control points can be determined from the position information of the matched feature points.
[0107] In the embodiments of the present application, the approximate correspondence between the second video frame and the three-dimensional model can be inferred according to the updated position of the control points in combination with the correspondence between the first video frame and the three-dimensional model.
[0108] In some embodiments, the processing module registers the three-dimensional model and the second video frame according to the correspondence, including:
[0109] The processing module determines the relative position relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, where the intraoperative image sensor is used to collect video images of the target tissue during the operation, and registers the three-dimensional model and the second video frame according to the elastic deformation and the relative position relationship.
[0110] In the embodiments of the present application, since the second video frame is collected by the intraoperative image sensor, the relative position relationship between the target tissue and the intraoperative image sensor can be determined through the relationship between the second video frame and the three-dimensional model.
[0111] In some embodiments, the processing module determines the relative positional relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, where the intraoperative image sensor is used to collect video images of the target tissue during surgery, and may include:
[0112] The processing module determines the relative positional relationship between the target tissue and the intraoperative image sensor according to the correspondence, determines the displacement change information of the rigid motion of the three-dimensional model according to the relative positional relationship, determines the virtual force of the change of the three-dimensional model according to the displacement change information, and simulates the elastic deformation of each feature point in the three-dimensional model according to the virtual force and the constraint conditions of the elastic deformation to obtain the elastic deformation of each feature point of the three-dimensional model, or; calculates the elastic deformation of each feature point in the three-dimensional model by using the finite element analysis method according to the virtual force and the biomechanical model of the target tissue.
[0113] In the embodiments of the present application, since the second video frame is collected by the intraoperative sensor, therefore, the relative position transformation can be solved based on the correspondence and the geometric transformation theory, so as to determine the relative positional relationship between the target tissue and the intraoperative image sensor. The geometric transformation theory may include a homography matrix, an essential matrix, etc.
[0114] In the embodiments of the present application, the displacement change information may include: changes such as the translation and rotation of the overall model.
[0115] In the embodiments of the present application, a virtual force field matching the target tissue can be designed and created according to the geometric shape and material properties of the target tissue. The virtual force field may include various types of forces, such as binding forces, spring forces, frictional forces, etc. After creating the virtual force field, the virtual force of the change of the three-dimensional model can be simulated through the virtual force field by the displacement change information.
[0116] In the embodiments of the present application, an appropriate elastic model can be determined according to the material properties of the target tissue and the positions of the feature points. This may include a linear elastic model, a nonlinear elastic model, etc., and a suitable model is selected according to the characteristics of the target tissue. In order to set the constraint conditions of each feature point during the simulation according to the virtual force and the constraint conditions of the elastic deformation, so as to simulate the external constraints and internal constraints received by the target tissue, and the simulation calculation can be performed according to the virtual force and the constraint conditions to simulate the elastic deformation of each feature point in the three-dimensional model.
[0117] In the embodiments of the present application, the biomechanical model is used to represent the mechanical properties of the target tissue. The biomechanical model may include: the elastic model and the constitutive model of the material. Then, the elastic deformation of each feature point in the three-dimensional model is calculated by using the finite element analysis method according to the virtual force and the biomechanical model of the target tissue.
[0118] In the embodiments of the present application, the three-dimensional model and the second video frame can be rigidly registered according to the relative position key, and the three-dimensional model and the second video frame can be elastically registered according to the elastic deformation.
[0119] The method provided by the embodiments of the present application can achieve real-time registration of both the rigid and elastic motions of the target organ.
[0120] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application.
[0121] Based on the foregoing medical image processing system, the embodiments of the application provide a medical image processing method that can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), and medical imaging devices. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.
[0122] The functions implemented by the medical image processing method provided by the embodiments of the present application can be realized by the processor of the electronic device calling the program code, where the program code can be stored in a computer storage medium.
[0123] The embodiments of the present application provide a medical image processing method, Figure 2 which is a schematic flowchart of the implementation of a medical image processing method provided for the implementation of the present application. As Figure 2 shown, the medical image processing method includes:
[0124] Step S1, obtaining a video image of the target tissue during the operation;
[0125] Step S2, determine the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and register the three-dimensional model and the second video frame according to the correspondence. The first video frame and the second video frame are consecutive video frames before and after respectively;
[0126] Step S3, display the registered second video frame and the three-dimensional model in real time.
[0127] A medical image processing method provided by an embodiment of the present application obtains video images of a target tissue during surgery; determines the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and registers the three-dimensional model and the second video frame according to the correspondence. The first video frame and the second video frame are consecutive video frames before and after respectively; displays the registered second video frame and the three-dimensional model in real time, and can perform real-time dynamic registration on the video frame and the three-dimensional model based on the video image.
[0128] In some embodiments, the determining the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame includes:
[0129] Determine the vector field of pixel movement in the second video frame according to the first video frame and the second video frame;
[0130] Determine the updated positions of the control points in the second video frame according to the vector field;
[0131] Determine the correspondence between the second video frame and the three-dimensional model according to the updated positions of the control points and the correspondence between the first video frame and the three-dimensional model.
[0132] In some embodiments, the method further includes:
[0133] Update the weight values of each feature point relative to the control point according to the updated positions of the control points in the second video frame.
[0134] In some embodiments, the determining the updated positions of the control points in the second video frame according to the vector field includes:
[0135] Map the feature points in the feature point dense area in the first video frame to the second video frame according to the vector field, and determine the updated positions of the control points in the second video frame according to the weight values of the feature points relative to the control points. The feature point dense area is an area where the number of feature points in the second video frame is greater than a threshold number;
[0136] Mapping the control points in the sparse region of feature points in the first video frame to the second video frame according to the vector field to obtain the updated positions of the control points in the sparse region of feature points, where the sparse region of feature points is the region in the second video frame where the number of feature points is less than the number threshold.
[0137] In some embodiments, determining the vector field of pixel movement in the second video frame according to the first video frame and the second video frame includes:
[0138] Searching for any target pixel in the first video frame on the second video frame;
[0139] Determining the displacement field of the target pixel movement according to the position of the target pixel in the first video frame and the position of the target pixel in the second video frame to obtain the vector field of pixel movement in the second video frame.
[0140] In some embodiments, when the first video frame is the first frame, establishing control point pairs according to the registration result includes:
[0141] Performing uniform sampling of feature points from the first video frame to obtain initial control points;
[0142] Screening out target control points from the initial control points according to the registration result obtained by registering the three-dimensional model and the first video frame to obtain the control point pairs, where there are corresponding control points for the target control points in the three-dimensional model.
[0143] In some embodiments, determining the correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame includes:
[0144] Extracting feature points from the second video frame and extracting feature points from the first video frame;
[0145] Matching the feature points extracted from the first video frame with the feature points extracted from the second video frame to obtain a matching result;
[0146] Determining the updated positions of the control points according to the matching result;
[0147] Determining the correspondence between the second video frame and the three-dimensional model according to the updated positions and the correspondence between the first video frame and the three-dimensional model.
[0148] In some embodiments, registering the three-dimensional model and the second video frame according to the correspondence includes:
[0149] Determine the relative positional relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, where the intraoperative image sensor is used to collect video images of the target tissue during the operation;
[0150] Register the three-dimensional model and the second video frame according to the elastic deformation and the relative positional relationship.
[0151] In some embodiments, the determining the relative positional relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model includes:
[0152] Determine the relative positional relationship between the target tissue and the intraoperative image sensor according to the correspondence;
[0153] Determine the displacement change information of the rigid motion of the three-dimensional model according to the relative positional relationship;
[0154] Determine the virtual force of the change of the three-dimensional model according to the displacement change information;
[0155] Simulate the elastic deformation of each feature point in the three-dimensional model according to the virtual force and the constraint conditions of the elastic deformation to obtain the elastic deformation of each feature point of the three-dimensional model, or;
[0156] Calculate the elastic deformation of each feature point in the three-dimensional model by using the finite element analysis method according to the virtual force and the biomechanical model of the target tissue.
[0157] For the working process of the above method, reference can be made to the corresponding process in the foregoing system embodiment, which will not be elaborated here.
[0158] Based on the medical image processing method provided in the foregoing embodiments, the embodiments of the present application further provide a medical image processing method, Figure 3 For the flowchart of a medical image processing method provided by the embodiments of the present application, as Figure 3 shown, it includes:
[0159] Step S11, perform initial registration on the three-dimensional model and the intraoperative image frame.
[0160] In the embodiments of the present application, initial rigid registration can be achieved by finding corresponding key features between the three-dimensional model and the intraoperative image frames. Different anatomical features are defined for different target organs. Deep learning networks can be used to extract corresponding anatomical key points / key lines on the three-dimensional model and the intraoperative image frames respectively. The corresponding key anatomical structures, points, and their projection positions in the two-dimensional images are calculated to obtain the result of the initial registration. The result of the registration includes the correspondence between the image frame and the three-dimensional model.
[0161] Step S12, initialization of control points.
[0162] In the embodiments of the present application, uniformly sampled points can be set on the intraoperative image frame at a fixed density, and the uniformly sampled points are screened according to the correspondence obtained from the initial registration. The uniformly sampled points within the target organ are used as control points.
[0163] General feature points are extracted on the intraoperative image frame, including but not limited to anatomical feature points of the target organ, image corner points, etc. Search is performed in the neighborhood of the control points with a fixed radius, and the control points are coupled with adjacent feature points according to information such as distance, that is, the weights of the feature points in the neighborhood relative to the control points are calculated.
[0164] Step S13, update of the positions of the control points.
[0165] In the embodiments of the present application, when there is movement within the field of view, that is, when the first video frame becomes the second video frame, the positions of the existing control points are updated.
[0166] In the embodiments of the present application, the frame rate of the endoscopic surgery video is generally above 30Hz, and the update speed is very fast. Therefore, the pixel points of the image can be regarded as not changing in terms of color, brightness, etc. between two adjacent frames. When the organ moves within the field of view, it will cause the corresponding pixel points on the video frame to move. The vector field of pixel movement can reflect the movement speed of the organ. Therefore, a local search method can be used to calculate the vector field of pixel movement. For any target pixel on the first video frame, search for the pixel point on the second video frame within a preset-sized neighborhood to obtain the correspondence of the same pixel point on the front and back video frames. After completing the search for all pixel points, the vector field of pixel movement can be obtained. After determining the vector field of pixel movement, the updated positions of the control points can be determined.
[0167] In the embodiments of the present application, for the dense area of feature points, the positions of the control points are determined by the feature points in the neighborhood. The feature points can be mapped from the first image frame to the second image frame according to the vector field, and the updated positions of the control points can be obtained using the coupling relationship.
[0168] In the embodiments of the present application, for the sparse feature point region, the sparse feature point region may include: a region where the surface of the target organ is smooth or a region where the anatomical structure is invisible under the endoscope, and the solution for tracking based on anatomical features will fail. In the embodiments of the present application, for the sparse feature point region, feature points may not be searched in the neighborhood of the control points, or the accuracy of the feature points cannot be guaranteed. Therefore, the control points and the feature points are decoupled. According to the vector field, the control points are mapped from the first image frame to the second image frame, and the updated position of the control points is directly obtained. The first image frame may be the nth image frame, and the second image frame may be the (n + 1)th image frame.
[0169] In the embodiments of the present application, after determining the updated position of the control points, the coupling relationship of the control points can be updated: general feature points are re-extracted on the second image frame, and the coupling relationship between the control points and the feature points is recalculated.
[0170] Step S14, perform dynamic registration.
[0171] In the embodiments of the present application, after determining the updated position of the control points, the feature points on the three-dimensional model maintain the same corresponding relationship with the feature points of the first video frame, and the updated position can be used to perform dynamic registration on the three-dimensional model and the second video frame.
[0172] In the embodiments of the present application, dynamic registration includes: rigid registration and elastic registration.
[0173] In the embodiments of the present application, after determining the updated position of the control points, the corresponding points in the new three-dimensional model and the second video frame are obtained. The Perspective-n-Point (PnP) is calculated using the updated points to obtain the relative position relationship between the new target tissue and the endoscope, and the preoperative model is driven to perform rigid motion to achieve rigid registration.
[0174] In the embodiments of the present application, after the control points are updated, the displacement field of all control points of the (N + 1)th frame relative to the Nth frame can be obtained. The position-based dynamics method PBD is used to simulate the elastic deformation of the model, and distance constraints and volume constraints are applied to drive the preoperative model to undergo elastic deformation, thereby achieving elastic registration.
[0175] The method provided by the embodiments of the present application uses initial registration to obtain the initial corresponding relationship of the target tissue on the preoperative model and the intraoperative image, and maps the control points on the intraoperative image to the preoperative model to drive the three-dimensional model to deform.
[0176] In the embodiments of the present application, control points are formed by uniform sampling on the image frame. When the characteristic anatomical structure of the target organ is invisible or the feature fails, the control points are decoupled from the anatomical features and can independently drive the deformation of the three-dimensional model. Moreover, the density of the sampling points can be flexibly adjusted according to different accuracy requirements to achieve high-precision and high-efficiency dynamic tracking.
[0177] The method provided by the embodiments of the present application establishes the corresponding relationship of pixels between the front and back video frames when the image changes, maps the preoperative and intraoperative corresponding relationship at the previous moment to the current moment, updates the control points, and drives the model to obtain a new registration result, which can achieve both rigid registration and elastic registration. This method quickly obtains a new corresponding relationship based on temporal information, realizes real-time dynamic registration, and can track different forms of organ movements in real time.
[0178] In some embodiments, when setting the control points, the control points are screened. By screening the control points, the movement of complex organ tissues around the target tissue can be avoided from affecting the tracking of the target tissue. Therefore, it can also be achieved through the segmentation of the target tissue. The segmentation of the target tissue can adopt traditional methods or can be performed using a deep learning network. The objects to be segmented include but are not limited to the target organ, surrounding tissue organs, instruments, etc.
[0179] In some embodiments, when updating the position of the control points, the calculation of the vector field is to obtain the corresponding feature points on the subsequent frame. Therefore, the matching of feature points can also be achieved by extracting features in both the front and back frames and then comparing them. The matching of feature points can adopt traditional methods, such as calculating the feature descriptors and then comparing them, or can directly use deep learning networks such as superglue to achieve accurate matching.
[0180] In some embodiments, during dynamic registration, the model can be physically simulated using Position-Based Dynamics (PBD). The advantage of physical simulation is fast calculation speed and good real-time performance. To improve the accuracy of elastic registration, biomechanical modeling can be performed on the target tissue, and materials of the target organ can be simulated using non-linear elastic models, poroelastic models, etc., and the elastic deformation of the organ driven by the control points can be calculated using finite elements.
[0181] Figure 4 The structural schematic diagram of the electronic device provided by the embodiments of the present application is as Figure 4 shown. The electronic device 3 of this embodiment may include: at least one processor 30 ( Figure 4 only one processor 30 is shown in Figure 2 ), a memory 31, and a computer program 32 stored in the memory 31 and executable on at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above method embodiments are implemented, such as Figure 1 the steps S1 to S3 in the embodiment shown. Or, when the processor 30 executes the computer program 32, the functions of each module / unit in the above system embodiments are implemented, such as Figure 1 the functions of the modules 101 to 103 shown.
[0182] Exemplarily, the computer program 32 can be divided into one or more modules / units. One or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program 32 instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.
[0183] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program 32, and when the computer program 32 is executed by the processor 30, the steps in the above-mentioned various method embodiments can be implemented.
[0184] The embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned various method embodiments when executed.
[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can be completed by instructing relevant hardware through the computer program 32. The computer program 32 can be stored in a computer-readable storage medium. When the computer program 32 is executed by the processor 30, the steps in the above-mentioned various method embodiments can be implemented. Among them, the computer program 32 includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the terminal, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0186] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0187] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0188] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0189] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0190] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A medical image processing system, characterized in that: include: An input module, used for inputting a video image of a target tissue during surgery; A processing module, used for determining a correspondence between the second video frame and the three-dimensional model according to a first video frame, the three-dimensional model of the target tissue and a second video frame, and registering the three-dimensional model with the second video frame according to the correspondence, wherein the first video frame and the second video frame are respectively consecutive video frames, and the processing module registers the three-dimensional model with the second video frame according to the correspondence, including: determining a relative position relationship between the target tissue and an intraoperative image sensor, and elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, wherein the intraoperative image sensor is used to collect video images of the target tissue during surgery; and registering the three-dimensional model with the second video frame according to the elastic deformation and the relative position relationship; The display module is used to display the registered second video frame and the three-dimensional model in real time.
2. The medical image processing system according to claim 1, characterized in that: The processing module is used to determine the corresponding relationship between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, including: Determine a vector field of pixel movement in the second video frame according to the first video frame and the second video frame; determining an updated position of a control point in a second video frame according to the vector field; The correspondence relationship between the second video frame and the three-dimensional model is determined according to the updated position of the control point and the correspondence relationship between the first video frame and the three-dimensional model.
3. The medical image processing system according to claim 2, characterized in that: The processing module determines a vector field of pixel movement in the second video frame according to the first video frame and the second video frame, including: For any target pixel in the first video frame, searching for the target pixel in the second video frame; According to the position of the target pixel in the first video frame and the position of the target pixel in the second video frame, a displacement field of the target pixel movement is determined to obtain a vector field of pixel movement in the second video frame.
4. The medical image processing system according to claim 2, characterized in that: When the first video frame is a first frame, the processing module establishes a control point pair according to the registration result, including: Uniformly sampling feature points from the first video frame to obtain initial control points; A target control point is screened out from the initial control points according to a registration result obtained after registering the three-dimensional model with the first video frame to obtain the control point pair, wherein the target control point has a corresponding control point in the three-dimensional model.
5. The medical image processing system according to claim 1, characterized in that: The processing module is used to determine the corresponding relationship between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, including: Extracting feature points from the second video frame, and extracting feature points from the first video frame; Matching the feature points extracted from the first video frame with the feature points extracted from the second video frame to obtain a matching result; Determine an update position of the control point according to the matching result; The correspondence relationship between the second video frame and the three-dimensional model is determined according to the update position, the correspondence relationship between the first video frame and the three-dimensional model.
6. The medical image processing system according to claim 1, characterized in that: The determining, according to the correspondence between the second video frame and the three-dimensional model, the relative position relationship between the target tissue and the intraoperative image sensor and the elastic deformation of each feature point of the three-dimensional model includes: Determining a relative position relationship between the target tissue and the intraoperative image sensor according to the corresponding relationship; Determining displacement change information of rigid motion of the three-dimensional model according to the relative position relationship; Determining a virtual force causing a change in the three-dimensional model according to the displacement change information; The elastic deformation of each feature point in the three-dimensional model is simulated according to the constraints of the virtual force and elastic deformation to obtain the elastic deformation of each feature point in the three-dimensional model, or the elastic deformation of each feature point in the three-dimensional model is calculated using a physical simulation method based on the virtual force and the biomechanical model of the target tissue.
7. The medical image processing system according to claim 6, characterized in that: The method of calculating the elastic deformation of each feature point in the three-dimensional model by using a physical simulation method according to the virtual force and the biomechanical model of the target tissue includes: The elastic deformation of each characteristic point in the three-dimensional model is calculated using a finite element analysis method according to the virtual force and the biomechanical model of the target tissue.
8. A medical image processing method, characterized in that: include: Acquire video images of target tissue during surgery; Determine a correspondence between the second video frame and the three-dimensional model according to the first video frame, the three-dimensional model of the target tissue, and the second video frame, and align the three-dimensional model with the second video frame according to the correspondence, wherein the first video frame and the second video frame are respectively consecutive video frames, and aligning the three-dimensional model with the second video frame according to the correspondence, including: determining a relative position relationship between the target tissue and an intraoperative image sensor, and elastic deformation of each feature point of the three-dimensional model according to the correspondence between the second video frame and the three-dimensional model, wherein the intraoperative image sensor is used to collect video images of the target tissue during surgery; and aligning the three-dimensional model with the second video frame according to the elastic deformation and the relative position relationship; The registered second video frame and the three-dimensional model are displayed in real time.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to claim 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to claim 8 is implemented.