A scene reconstruction method for pulmonary bronchoscopic surgical robots

Through the combination of U-Net network and HAPCG algorithm, the problem of three-dimensional reconstruction in pulmonary bronchoscopic surgical robots is solved, and high-accuracy and low-cost three-dimensional reconstruction of pulmonary bronchos is achieved to adapt to the needs of narrow scenarios.

CN115222878BActive Publication Date: 2025-08-29ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210691216.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2025-08-29
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to reconstruct the three-dimensional lung bronchus in pulmonary bronchoscopic surgical robots, especially the monocular depth estimation method lacks accuracy and calculation efficiency. Excessive sensor size will increase human injury, and the application value of learning-based methods is not high.

Method used

U-Net network filtering is used to enhance the endoscopic image features, combine the HAPCG algorithm to extract feature points and match them, and estimate the depth through triangulation measurement method to achieve three-dimensional reconstruction, avoiding the use of additional sensors.

Benefits of technology

It provides a three-dimensional reconstruction method of pulmonary bronchus with high accuracy and robustness, reduces costs, avoids harm to the human body, and adapts to pulmonary bronchoscopy stenosis scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222878B_ABST
    Figure CN115222878B_ABST
Patent Text Reader

Abstract

The present invention discloses a scene reconstruction method applied to a pulmonary bronchoscopic surgical robot, belonging to the field of pulmonary bronchoscopic surgical robots. An endoscopic image sequence of the pulmonary bronchi is obtained and filtered; feature points are extracted from the filtered image sequence and feature points of adjacent frames are matched; the first frame image perspective is used as the initialization posture, and the posture of the second frame image is calculated based on the feature points and matching relationship of the first two frames; based on the feature points, matching relationship and posture, a triangulation measurement method is used to estimate the depth of the matched feature points to obtain an initial point cloud; based on the position of the feature points of the t-th frame image in the point cloud, the posture of the t-th frame image is obtained, t≥3; based on the feature points, matching relationship and posture of the t-th frame and the previous frame image, a triangulation measurement method is used to estimate the depth of the matched feature points, and the point cloud is updated; finally, a three-dimensional reconstruction image of the local pulmonary bronchi is obtained. The present invention does not require additional sensor information assistance, and has high reconstruction accuracy and good robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of pulmonary bronchoscopic surgical robots, to the realization of the functions of the perception part of the robot, and in particular to a scene reconstruction method applied to the pulmonary bronchoscopic surgical robot. Background Art

[0002] Traditional bronchoscopic navigation techniques are mainly divided into the following categories: Radical EBUS (radical endobronchial ultrasonography), EMN (electromagnetic navigation) and VBN (virtual bronchoscopic navigation), each with its own advantages and disadvantages.

[0003] Radical EBUS: The radial EBUS probe provides ultrasound images of the tissue surrounding the tip. The system can determine the location of peripheral lung lesions by comparing EBUS images of normal lung tissue and cancerous tissue. However, EBUS-GS cannot perform biopsies under visual control. The EBUS probe must be removed before the biopsy tool is introduced from the GS into the working channel, which adds certain difficulties to biopsy.

[0004] EMN: SuperDimension (Medtronic) is an EMN system that can be used in conjunction with an EBUS probe to improve the diagnosis of lung cancer. However, the cost of the EMN navigation system itself limits the development of this technology and its frequency of use in routine bronchoscopy.

[0005] Image-based bronchoscopic navigation (VBN) can achieve three-dimensional reconstruction of the chest cavity and lung nodules. Before a bronchoscopic biopsy, the location of peripheral lung lesions is marked, and the bronchoscope's trajectory is precalculated and displayed in real time on the image during the biopsy. Despite the ease of obtaining reconstructed information and the simplicity of the underlying methods, VBN is not widely used. This is partly because the accuracy of CT image reconstruction still needs to be improved, and CT image slice thickness should be less than 1 mm. Thicker CT slices, respiratory tracheal deformation, and intraluminal secretions shorten the length of the visually reconstructed bronchial tree, resulting in a loss of peripheral lung information and lowering diagnostic yield. Furthermore, 3D reconstruction of real-time biopsy probe images faces numerous challenges. Limited by the size of the lung bronchi, binocular depth cameras are often difficult to use, making it difficult to obtain environmental depth information. Monocular cameras, on the other hand, face the challenge of limited and uniform features in the lung bronchi. Therefore, research on the extraction and matching of feature points in the 3D reconstruction of lung bronchi is valuable.

[0006] In the case of lung bronchi, visual information is very scarce, with few and simple features. Below the primary bronchus, there is typically a long lumen. However, as the endoscope at the end of the bronchoscope rotates, the resulting image often only shows the tube wall in a single color, without distinct features. Even images with an elliptical lumen outline only show a single, minimal, and simple feature, making feature extraction extremely difficult.

[0007] Typically, 3D reconstruction methods use a binocular camera, which can directly determine the depth of the environment. Once the pose is determined, 3D reconstruction can be performed. However, the working environment of a lung bronchoscopy robot is narrow, and excessively large sensors would increase harm to the human body, so the use of binocular cameras is not permitted. Depth estimation can only be performed using monocular images, that is, RGB images, which presents certain challenges in ensuring accuracy and computational efficiency.

[0008] Currently, some researchers are actively exploring deep learning-based monocular depth estimation methods for medical applications. These methods can obtain depth estimates from a single RGB image through the principle of style transfer. However, this method lacks strong mathematical basis and is similar to a black box approach. Currently, there are no open-source, proven learning-based monocular depth estimation methods for the lungs and bronchi. While learning-based methods can be considered innovative, their application value is limited. Summary of the Invention

[0009] In response to the difficulties in the three-dimensional reconstruction technology of lung bronchi by bronchoscopic surgical robots in the existing technology, the present invention proposes a scene reconstruction method applied to lung bronchoscopic surgical robots, and provides a complete three-dimensional reconstruction solution for the scene. This technology is closely related to the positioning and navigation functions of the bronchoscope.

[0010] The present invention adopts the following technical solutions:

[0011] A scene reconstruction method for a pulmonary bronchoscopic surgical robot comprises the following steps:

[0012] Step 1: Acquire an endoscopic image sequence of the lung bronchi, filter the endoscopic image, enhance the lumen features of the trachea on the endoscopic image, and obtain an enhanced image sequence;

[0013] Step 2: extract feature points from the enhanced image sequence and perform feature point matching on adjacent frame images to obtain a matching relationship between adjacent frame images;

[0014] Step 3: Using the viewing angle of the first endoscopic image as the initial pose, the pose of the second endoscopic image relative to the first endoscopic image is calculated based on the feature points and matching relationship of the first two endoscopic images; based on the feature points, matching relationship, and pose, a triangulation method is used to estimate the depth of the matching feature points to obtain an initial point cloud;

[0015] Step 4: According to the position of the feature points of the endoscopic image of the t-th frame in the point cloud, the pose of the endoscopic image of the t-th frame is obtained, where t≥3; according to the feature points, matching relationship, and pose of the endoscopic images of the t-th frame and the t-1-th frame, the depth of the matched feature points is estimated using the triangulation measurement method, and the point cloud is updated;

[0016] Step 5: Obtain a 3D reconstruction of the local lung bronchi based on the final point cloud.

[0017] Furthermore, in step 1, a U-Net network is used to filter the endoscopic image, and the U-Net network includes two parts: downsampling and upsampling;

[0018] The downsampling is specifically as follows: after every two 3×3 convolution operations, a maximum pooling operation with a stride of 2 is performed as a downsampling operation, and each downsampling process uses a ReLU activation function;

[0019] The upsampling is specifically as follows: first, a 2×2 convolution operation is performed to double the number of channels, and then it is spliced ​​with the feature map generated by the corresponding downsampling process, and then two 3×3 convolution operations are performed as an upsampling, and each upsampling process uses the ReLU activation function;

[0020] The output of the last upsampling is passed through a 1×1 convolutional layer to obtain a filtered endoscopic image, which is used to enhance the tracheal lumen contour and internal features on the endoscopic image.

[0021] Furthermore, the feature extraction and matching algorithm used in step 2 is the HAPCG algorithm.

[0022] Furthermore, the step 3 is specifically as follows:

[0023] Step 3.1, solve the 2D-2D pose using the eight-point method

[0024] For a pair of matching points of the first two frames of endoscopic images, defined on the normalized coordinate plane, we have X = [u, v, 1] T and X1=[u1,v1,1] T Two pixels:

[0025] According to the epipolar geometry constraints, we have:

[0026]

[0027] in, Represents a 3×3 constraint matrix, where u1 and v1 are the coordinates of pixel point X1, and u and v are the coordinates of pixel point X;

[0028] Written in linear form:

[0029] [u1u,u1v,u1,v1u,v1v,v1,u,v,1]·e=E1·e=0

[0030] Where e represents the 9×1 constraint matrix, E1 is the local essential matrix;

[0031] Similarly, select the remaining seven pairs of matching points and express them in the same way as above; integrate the expression results of the eight pairs of matching points into a linear equation and solve to obtain the essential matrix

[0032] Perform singular value decomposition on the essential matrix E:

[0033] E=U∑V T

[0034] Calculate the solution:

[0035]

[0036]

[0037] Among them, U and V represent orthogonal matrices, Σ represents diagonal matrices; R1 and R2 are rotation matrices, and t1 and t2 are translation matrices; Indicates rotation along the z axis The rotation matrix of Indicates rotation along the z axis The rotation matrix of

[0038] From the above four solutions (R1, t1), (R1, t2), (R2, t1), (R2, t2), select the only solution that meets the actual situation as the pose estimation result (R, t);

[0039] Step 3.2: Triangulation to estimate depth

[0040] According to the pose estimation result (R, t), the relationship between pixel points X and X1 is obtained:

[0041] s1X1=sRX+t

[0042] Convert the above formula into:

[0043] s1X1^X1=0=sX1^RX+X1^t

[0044] Where s1 is the depth information of pixel X1, and s is the depth information of pixel X. If we regard the formula as an equation about depth, we can directly solve s1 and s.

[0045] Step 3.3, repeat steps 3.1 and 3.2, traverse all matching points, and obtain the initial point cloud based on the depth of the matched feature points.

[0046] Furthermore, in step 4, the pnp algorithm is used to calculate the pose of the t-th frame of the endoscopic image, t≥3.

[0047] Furthermore, in step 4, before each point cloud update, the estimated point cloud is first projected onto different imaging planes, and the Euclidean distance is calculated with the feature points originally corresponding to the plane. If the distance is greater than the threshold, the point is considered to be poorly estimated and is removed from the point cloud.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention proposes a feature point extraction and matching method for pulmonary bronchoscopic endoscopic images, namely, a feature point extraction and matching algorithm based on HAPCG supplemented by U-Net filtering method. HAPCG enhances the image's invariance to illumination and contrast from the perspective of texture, while U-Net filtering enhances the lumen content features of bronchoscope images from the perspective of content. Compared with traditional ORB and SIFT algorithms, this method is more suitable for pulmonary bronchoscopic scenarios and provides a good foundation for three-dimensional reconstruction of the lung bronchi.

[0050] 2. The pulmonary bronchoscopy robot 3D reconstruction method proposed in the present invention can directly obtain the 3D model of the scene from the 2D image given a sequence of images, without the need for additional sensor information assistance, and has high reconstruction accuracy and good robustness.

[0051] 3. The present invention directly uses RGB images for pose estimation to achieve three-dimensional reconstruction, without the need for expensive sensors such as NDI, thus saving costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 1 is a flow chart of a scene reconstruction method applied to a pulmonary bronchoscopic surgical robot, according to an embodiment of the present invention;

[0053] Figure 2 This is a schematic diagram of the three-dimensional scene reconstruction effect proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The present invention is further described below with reference to the accompanying drawings and embodiments. The accompanying drawings are merely schematic illustrations of the present invention. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0055] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to actual circumstances.

[0056] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by a person having ordinary skills in the technical field to which this application belongs. In this application, words such as "one", "a", "the", "these" and the like do not indicate a limit on quantity, and they may be singular or plural. The terms "include", "comprising", "having" and any variations thereof used in this application are intended to cover non-exclusive inclusions; words such as "connected", "connected", "coupled" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0057] Essentially, the 3D reconstruction problem of lung bronchoscopes can also be viewed as an SFM (Structure from Motion) problem. This paper obtains sequential images of the bronchoscope robot's terminal motion using a monocular endoscope, designs a special method to extract features from these images and perform feature matching on adjacent images, calculates relative poses, and achieves depth estimation through triangulation to complete 3D scene reconstruction.

[0058] like Figure 1 As shown, the scene reconstruction method applied to the pulmonary bronchoscopic surgical robot of the present invention mainly includes two parts: (1) feature point extraction and matching, and (2) pose estimation and scene reconstruction.

[0059] In this embodiment, feature point extraction and matching include:

[0060] (1.1) Position, scale and orientation invariant feature transform (histogram of absolute phase consistency gradients, HAPCG).

[0061] The main purpose of applying the algorithm based on the HAPCG feature descriptor is to enhance the feature points based on the texture characteristics of the endoscopic image. The HAPCG algorithm was originally designed for feature point extraction and matching between heterogeneous remote sensing images. This scenario has problems such as significant image illumination differences, large contrast differences, and nonlinear radiation distortion, and the algorithm can solve these problems well. Similarly, analyzing the data collected by bronchoendoscopy, it can be found that bronchoscopic images are affected by the light provided by the endoscope. The contrast and illumination differences between different images are large, and a feature extraction algorithm with illumination and contrast invariance is also required. It can be seen that the HAPCG algorithm is equally applicable to bronchoendoscopic scenarios.

[0062] During HAPCG algorithm execution, two adjacent image frames are first subjected to nonlinear diffusion through anisotropic filtering to obtain the maximum and minimum directional matrices of phase consistency. Based on this, anisotropic weighted moment equations are established, and anisotropic weighted moment maps are calculated. Next, the absolute phase consistency directional gradient formula is derived by expanding the absolute phase consistency model. A histogram of absolute phase consistency gradients (HAPCG) is then defined based on a template in logarithmic polar coordinates. Finally, the Euclidean example method is used as a matching metric to identify feature points with the same name, and a fast sample consensus method is employed to remove false matches, thereby achieving detection and registration of feature points.

[0063] This algorithm has shown better results than traditional scale-invariant feature transform (SIFT), position-scale-orientation-invariant feature transform (PSO-SIFT), and Log-Gabor histogram descriptor (LGHD). In the HAPCG algorithm, anisotropic filtering is a nonlinear filtering method, and the edge information of the image is well preserved, making the subsequent feature point extraction more satisfactory and making up for the lack of image feature points. The phase consistency method is used to extract features in the frequency domain, which can well extract the boundaries and corners of the image from the maximum phase of the superposition of the Fourier harmonic components of the image, and is not affected by the amplitude of the image amplitude spectrum, so as to achieve good invariance to image illumination and contrast. The HAPCG algorithm has good feature extraction and matching effects in the pulmonary bronchoscopy scene, providing a solid foundation for three-dimensional scene reconstruction.

[0064] (1.2) Content segmentation based on U-Net network;

[0065] The images provided by bronchoscopes are typically simple, consisting only of the tube wall and the circular or elliptical outline of the trachea. However, while the tube wall occupies the majority of the image area in a single frame, the feature points that require extraction and matching are often concentrated only on the tracheal lumen and interior. Therefore, the present invention adds a filter to focus the entire image on the tracheal lumen and interior, essentially enhancing the primary image content. This can also be considered as adding a filtering operation to the image.

[0066] In this embodiment, U-Net is selected as the filtering method. This is a network architecture based on convolutional neural networks, which is usually used for medical image segmentation tasks and is suitable for lung bronchoscopy scenarios. Data sets in the field of medical images are difficult to obtain, and the amount of data is usually relatively small. Therefore, this network makes good use of data enhancement methods and can achieve satisfactory results using only limited labeled data. The U-Net network includes a series of downsampling steps to extract image content, and a symmetrical upsampling process to well locate the content information at the original image size. During the experiment, the network was able to greatly reduce the workload of manual labeling, while accurately extracting the content of the image, that is, the position of the tube wall, providing good input for subsequent work.

[0067] The U-Net network structure used in this embodiment includes two parts: downsampling and upsampling;

[0068] The downsampling is specifically as follows: after every two 3×3 convolution operations, a maximum pooling operation with a stride of 2 is performed as a downsampling operation, and each downsampling process uses a ReLU activation function;

[0069] The upsampling is specifically as follows: first, a 2×2 convolution operation is performed to double the number of channels, and then it is spliced ​​with the feature map generated by the corresponding downsampling process, and then two 3×3 convolution operations are performed as an upsampling, and each upsampling process uses the ReLU activation function;

[0070] The output of the last upsampling is passed through a 1×1 convolutional layer to obtain a filtered endoscopic image, which is used to enhance the tracheal lumen contour and internal features on the endoscopic image.

[0071] In this embodiment, pose estimation and scene reconstruction include:

[0072] (2.1) The eight-point pair method is used to solve the 2D-2D pose.

[0073] After the previous steps, some paired feature points can be obtained. Based on these results, the relative motion of adjacent frames of pulmonary bronchoscope endoscope images can be calculated. Taking a pair of matching points as an example, assuming that the matching is correct, then the matching point pair x and x1 is the projection of the same spatial point p on two different imaging planes with c0 and c1 as the optical center, that is, and intersect at p. For the convenience of description, some terms are defined to describe the set relationship between them, as shown in Table 1 below:

[0074] Table 1 Description of the geometric relationships involved in epipolar geometry constraints

[0075]

[0076]

[0077] Add epipolar geometry constraints, the formula is as follows:

[0078]

[0079] Among them, K is the camera intrinsic parameter matrix, t is the translation matrix, and R is the rotation matrix;

[0080] Its geometric meaning is that p, c0, and c1 are coplanar. The middle part is recorded as two matrices: the basic matrix F and the essential matrix E. The difference between E and F is only the camera intrinsic parameter matrix K, which further simplifies the epipolar geometry constraint:

[0081] E=t^R,F=K -T EK -1 ,

[0082] Therefore, camera pose estimation is divided into the following two steps:

[0083] The first step is to find E or F based on the pixel position of the paired points.

[0084] The second step is to solve R,t based on E or F.

[0085] E = t^R has six degrees of freedom. This is because, based on spatial relationships, the translation matrix and rotation matrix each have three degrees of freedom. Taking into account scale uncertainty in real-world situations, eliminating one degree of freedom still leaves five. This means that the matrix E can be uniquely determined with at least five pairs of points. However, it is worth noting that due to the inherent nonlinearity of E, its estimation is relatively difficult. Therefore, in practical research and applications, eight pairs of points are often used to solve E.

[0086] Consider a pair of matching points, defined on the normalized coordinate plane, with X = [u, v, 1] T and X1=[u1,v1,1]T Two pixels:

[0087] According to the epipolar geometry constraint, we have

[0088]

[0089] in, Represents a 3×3 constraint matrix, where u1 and v1 are the coordinates of pixel X1, and u and v are the coordinates of pixel X;

[0090] Written in linear form with respect to e

[0091] [u1u,u1v,u1,v1u,v1v,v1,u,v,1]·e=E1·e=0

[0092] Where e represents the 9×1 constraint matrix and E1 is the local essential matrix.

[0093] Similarly, the remaining seven pairs of points can be expressed in the same way. By integrating all eight points into a linear equation, the essential matrix can be solved. Then, we need to perform SVD decomposition on E to obtain a non-unique pose estimation solution:

[0094] E=U∑V T

[0095]

[0096]

[0097] Among them, U and V represent orthogonal matrices, Σ represents diagonal matrices; R1 and R2 are rotation matrices, and t1 and t2 are translation matrices; Indicates rotation along the z axis The rotation matrix of Indicates rotation along the z axis The rotation matrix of the four solutions (R1, t1), (R1, t2), (R2, t1), and (R2, t2) satisfies the constraints, meaning the depth of the point is positive in the camera coordinate system. Substituting all the solutions into the original equation and determining the depth of the point yields the correct solution.

[0098] (2.2) Depth estimation using triangulation

[0099] The 3D depth of pixels cannot be directly obtained from a single 2D image from a monocular camera. To address this problem, triangulation is often used to calculate the depth of pixels.

[0100] According to the above epipolar geometry constraints, and Intersect at point p. However, noise often exists in feature point extraction and matching, making this difficult to achieve and the two lines difficult to compare in three-dimensional space. To solve this problem, the least squares method can usually be used for approximation.

[0101] Assume that X, X1 are the normalized coordinates of two feature points, satisfying

[0102] s1X1=sRX+t

[0103] Given R and t, we want to solve the depths s1 and s of two feature points. Geometrically, Find the coordinate point of the three-dimensional coordinate system so that its projection position is close to x1. For example, when solving for s, multiply both sides of the above equation by X1^ on the left to obtain:

[0104] s1X1^X1=0=sX1^RX+X1^t

[0105] The formula can be viewed as an equation about depth, with one side being zero, and one of the depths can be solved directly.

[0106] (2.3) PnP algorithm solves 3D-2D pose

[0107] During the initialization phase, an initial 3D point cloud is generated. Successive 2D images then need to be matched to this initial point cloud. PnP is a method used to solve this problem. Compared to the eight-point method used for 2D-2D pose estimation, PnP requires only three point pairs for 3D-2D pose estimation. Since the endoscope is a monocular camera, the 3D positions of the feature points are determined by triangulation.

[0108] At this point, for each frame, we can obtain its feature points and their matching relationships with the feature points of adjacent frames. For the first two frames, we use the eight-point method to calculate the initial pose and triangulate the feature point depth. For each subsequent frame, we can use the 3D points and their corresponding 2D projections onto that frame to determine the pose of that frame. This yields a pose sequence from which we can perform 3D reconstruction of the feature points across multiple frames.

[0109] During the reconstruction process, the three-dimensional structure is optimized using the idea of ​​Bundle Adjustment (BA). The so-called BA refers to extracting the optimal 3D model and camera parameters from the visual graphics. If the calculation results and related parameters are adjusted so that the light bundle where the feature points are located converges to the optical center of the camera, it is called BA. In this embodiment, the specific approach is to project the estimated three-dimensional feature points onto different imaging planes respectively, and calculate the Euclidean distance with the original corresponding feature points on the plane. If the distance is greater than the threshold, it is considered that the point is poorly estimated and is removed from the three-dimensional point cloud.

[0110] (2.4) Scene reconstruction

[0111] In this embodiment, MATLAB is used to reconstruct the scene, and the steps are as follows:

[0112] Input the image sequence acquired by endoscope {I t},t=1,2,…n;

[0113] Calculate the feature descriptor of each image {D t},t=1,2,…n;

[0114] Calculate the matching relationship between adjacent frame images {M t}, t=1,2,…n-1;

[0115] Calculate the pose of I2 relative to I1 and use triangulation to estimate the depth of feature points to obtain the initial point cloud p1;

[0116] Using M t , calculate I t+1 to p t-1 PnP relationship, estimate the three-dimensional coordinates p of the feature point in the image t+1 (t=2,..n);

[0117] For all p t (t=1,2,..n-1) perform BA optimization to obtain the final 3D reconstruction image.

[0118] In this embodiment, Figure 2 The reconstruction results of 7 endoscopic images are shown, with small reconstruction error, high accuracy and good robustness.

[0119] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A scene reconstruction method for a pulmonary bronchoscopic surgical robot, characterized in that: The following steps are involved: Step 1: Acquire an endoscopic image sequence of the lung bronchi, filter the endoscopic image, enhance the lumen features of the trachea on the endoscopic image, and obtain an enhanced image sequence; Step 2: extract feature points from the enhanced image sequence and perform feature point matching on adjacent frame images to obtain a matching relationship between adjacent frame images; Step 3: Using the viewing angle of the first endoscopic image as the initial pose, the pose of the second endoscopic image relative to the first endoscopic image is calculated based on the feature points and matching relationship of the first two endoscopic images; based on the feature points, matching relationship, and pose, a triangulation method is used to estimate the depth of the matching feature points to obtain an initial point cloud; The step 3 is specifically as follows: Step 3.1, solve the 2D-2D pose using the eight-point method For a pair of matching points of the first two frames of endoscopic images, defined on the normalized coordinate plane, we have X = [u, v, 1] T and X1=[u1,v1,1] T Two pixels: According to the epipolar geometry constraints, we have: in, Represents a 3×3 constraint matrix, where u1 and v1 are the coordinates of pixel point X1, and u and v are the coordinates of pixel point X; Written in linear form: [u1u,u1v,u1,v1u,v1v,v1,u,v,1]·e=E1·e=0 Where e represents the 9×1 constraint matrix, E1 is the local essential matrix; Similarly, select the remaining seven pairs of matching points and express them in the same way as above; integrate the expression results of the eight pairs of matching points into a linear equation and solve to obtain the essential matrix Perform singular value decomposition on the essential matrix E: E=U∑V T Calculate the solution: Among them, U and V represent orthogonal matrices, Σ represents diagonal matrices; R1 and R2 are rotation matrices, and t1 and t2 are translation matrices; Indicates rotation along the z axis The rotation matrix of Indicates rotation along the z axis The rotation matrix of From the above four solutions (R1, t1), (R1, t2), (R2, t1), (R2, t2), select the only solution that meets the actual situation as the pose estimation result (R, t); Step 3.2: Triangulation to estimate depth According to the pose estimation result (R, t), the relationship between pixel points X and X1 is obtained: s1X1=sRX+t Convert the above formula into: s1X1^X1=0=sX1^RX+X1^t Where s1 is the depth information of pixel X1, and s is the depth information of pixel X. If we regard the formula as an equation about depth, we can directly solve s1 and s. Step 3.3, repeat steps 3.1 and 3.2, traverse all matching points, and obtain the initial point cloud based on the depth of the matched feature points; Step 4: According to the position of the feature points of the endoscopic image of the t-th frame in the point cloud, the pose of the endoscopic image of the t-th frame is obtained, where t≥3; according to the feature points, matching relationship, and pose of the endoscopic images of the t-th frame and the t-1-th frame, the depth of the matched feature points is estimated using the triangulation measurement method, and the point cloud is updated; Step 5: Obtain a 3D reconstruction of the local lung bronchi based on the final point cloud.

2. The scene reconstruction method for a pulmonary bronchoscopic surgical robot according to claim 1, characterized in that: In step 1, the endoscopic image is filtered using a U-Net network, wherein the U-Net network includes two parts: downsampling and upsampling; The downsampling is specifically as follows: after every two 3×3 convolution operations, a maximum pooling operation with a stride of 2 is performed as a downsampling operation, and each downsampling process uses a ReLU activation function; The upsampling is specifically as follows: first, a 2×2 convolution operation is performed to double the number of channels, and then it is spliced ​​with the feature map generated by the corresponding downsampling process, and then two 3×3 convolution operations are performed as an upsampling, and each upsampling process uses the ReLU activation function; The output of the last upsampling is passed through a 1×1 convolutional layer to obtain a filtered endoscopic image, which is used to enhance the tracheal lumen contour and internal features on the endoscopic image.

3. The scene reconstruction method for a pulmonary bronchoscopic surgical robot according to claim 1, characterized in that: The feature extraction and matching algorithm used in step 2 is the HAPCG algorithm.

4. The scene reconstruction method for a pulmonary bronchoscopic surgical robot according to claim 1, characterized in that: In step 4, the pnp algorithm is used to calculate the pose of the t-th frame of the endoscope image, t≥3.

5. The scene reconstruction method for a pulmonary bronchoscopic surgical robot according to claim 1, characterized in that: In step 4, before each point cloud update, the estimated point cloud is first projected onto different imaging planes, and the Euclidean distance is calculated with the feature points originally corresponding to the plane. If the distance is greater than the threshold, the point is considered to be poorly estimated and is removed from the point cloud.

Citation Information

Patent Citations

  • Robot scene self-adaptive pose estimation method based on RGB-D camera

    CN110223348A

  • Three-dimensional reconstruction method and device for monocular endoscope image and terminal equipment

    CN111145238A