Surgical navigation system space registration method based on feature point detection
Through the method based on feature point detection, the image space and patient space are matched by artificial markers divided into two parts A and B, which solves the traumatic and operational complexity of the existing surgical navigation system, and achieves efficient and accurate spatial registration.
Patent Information
- Application Number
- CN202510498120.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-05
AI Technical Summary
The spatial registration method of existing surgical navigation systems has problems of traumaticity, operational complexity and reliability of feature point acquisition, which affects registration efficiency and accuracy.
A method based on feature point detection is adopted, artificial markers divided into two parts A and B are used to match the image space and the patient space. The three-dimensional coordinates of the marker are identified by the detection network and binocular camera, which simplifies the operation steps and improves the recognition accuracy.
It reduces the risk of trauma to patients, simplifies the operating steps of doctors, improves surgical efficiency and navigation accuracy, and reduces artificial errors.
Smart Images

Figure CN120420083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-aided technology, and in particular to a spatial registration method for a surgical navigation system based on feature point detection. Background Art
[0002] Surgical navigation systems are designed to provide surgeons with real-time, accurate information about the surgical site, assisting them in performing surgical procedures with greater precision, reducing surgical risks, and improving surgical success rates. Spatial registration, a core component of surgical navigation systems, establishes an accurate correspondence between a patient's preoperative medical imaging data and the actual anatomy during surgery. This allows surgeons to accurately locate lesions and avoid critical structures such as nerves and blood vessels during surgery, based on preoperative planning.
[0003] Currently, the spatial registration methods for surgical navigation are mainly divided into two categories: one is spatial registration based on markers, and the other is spatial registration based on surface information.
[0004] The marker point registration method has been widely used in clinical surgical navigation. It identifies feature points in preoperative images and registers them with these feature points identified and tracked during surgery. The types of marker points mainly include bone implanted markers, artificial markers, and physiological markers. Bone implanted markers have the highest accuracy, but require the injection of instruments in advance, which will cause additional trauma and pain to the patient. Artificial markers have high accuracy and are simple to operate. They do not cause additional trauma to the patient and are the most commonly used markers. However, during the registration process, the doctor needs to manually select the marker points for identification, which increases the complexity of the registration process. Physiological markers use significant feature points on the human body as identifiers, such as the tip of the nose, the corners of the eyes, etc., but due to the errors in manual identification and the insufficient or insignificant number of feature points in many parts that cannot meet the registration requirements, they are less used.
[0005] Surface-based spatial registration methods utilize surface data from the patient in two spaces to create a transformation relationship. This method uses a handheld scanner to obtain a set of surface points. The surgeon first clicks on several characteristic anatomical points on the patient's face in both the 3D image and the actual surgical space for a rough match. Surface data is then collected for a precise match, which adds several steps for the surgeon.
[0006] In summary, the problems with existing registration methods include: (1) In terms of trauma, some methods, such as bone-implanted markers, can cause additional trauma to patients and increase the risk of infection. (2) In terms of operational complexity, both the manual selection of artificial markers and physiological markers and methods based on surface information greatly increase the burden on doctors, easily lead to human errors, and affect registration efficiency and accuracy. Overall, existing methods need to be improved in terms of trauma, operational convenience, and reliability of feature point acquisition to enhance the quality and effectiveness of surgical navigation space registration. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings and defects of existing technologies by providing a spatial registration method for a surgical navigation system based on feature point detection. This method utilizes artificial markers with unique structural features, divided into two parts, A and B, to match image space with patient space. This method effectively improves recognition accuracy and stability, eliminates manual operation steps, shortens preoperative preparation time, and effectively enhances overall surgical efficiency and navigation accuracy.
[0008] The present invention is achieved in that:
[0009] A spatial registration method for a surgical navigation system based on feature point detection comprises the following steps:
[0010] S1. Before surgery, attach at least four markers A that can be visualized on CT to the patient's skin, preprocess the patient's preoperative CT image data, and annotate the markers A;
[0011] S2. Using the marker A detection dataset to train the marker A detection network, and obtain the weight parameters of the trained detection network;
[0012] S3. The detection network processes the pre-processed preoperative CT image data based on the weight parameters to obtain the three-dimensional coordinates of the marker in the image space;
[0013] S4. During surgery, the binocular camera captures an image of a marker consisting of marker B attached to marker A, analyzes images of the same marker from different viewpoints, and obtains the three-dimensional coordinates of the marker in the patient's space.
[0014] S5. Determine the corresponding marker points in the image space and the patient space according to the three-dimensional coordinates of the marker in the image space and the patient space;
[0015] S6. Map the image space to the patient space based on the corresponding marker points in the image space and the patient space to complete the spatial registration.
[0016] Preferably, the marker A is a disposable ECG electrode sheet with a metal cylinder in the center, and the marker B is a cylinder that forms a detachable connection with the metal cylinder of the marker A through a slot in the center thereof. The marker B consists of an inner black circular part and an outer white annular part.
[0017] Preferably, the detection network includes:
[0018] A candidate detection model is used to obtain the location information, diameter, and anchor box target score of candidate markers based on patient imaging data;
[0019] The classification and screening model is used to classify candidate markers based on patient image data and the output data of the candidate detection model to determine the probability that they are true markers and obtain the three-dimensional coordinates of the target markers.
[0020] Preferably, the detection network is improved based on Faster RCNN, and the candidate detection model includes a CNN encoder, a feature enhancement module, and a CNN decoder from the input side to the output side, and the feature enhancement module includes two parallel feature enhancement units;
[0021] The CNN encoder includes a first encoding module, a second encoding module, a third encoding module and a fourth encoding module connected in sequence, and the CNN decoder includes a first decoding module, a second decoding module and a third decoding module connected in sequence, wherein the second encoding module is connected to the third decoding module through a feature enhancement unit, and the third encoding module is connected to the second decoding module through another feature enhancement unit.
[0022] Preferably, the classification and screening model includes, from the input side to the output side, a three-dimensional image extraction unit, a convolution unit, a feature decoding module, an average pooling unit, and a fully connected layer unit; the feature decoding module is composed of three groups of feature decoding units connected in sequence, and each group of feature decoding units is composed of a maximum pooling unit and a feature enhancement unit connected.
[0023] Preferably, the feature enhancement unit includes three convolutional layers arranged in parallel and a residual connection: the feature maps obtained by the three convolutional layers are spliced, and then passed through a convolutional layer, and the output obtained is residually connected with the original input as the output of the feature enhancement unit.
[0024] Preferably, in step S4, the image of the marker formed by the marker B fastened to the marker A captured by the binocular camera is processed by combining the internal and external parameters of the binocular camera obtained by calibration and using a continuous binary threshold method and binocular stereo vision principle.
[0025] Preferably, in step S5, a marker matching algorithm is used to process the three-dimensional coordinates of the markers in the image space and the patient space, obtain matching indexes of the two sets of point pairs, and determine corresponding marker points in the image space and the patient space.
[0026] Preferably, in step S6, an input alignment algorithm is used to calculate the transformation matrix between the image space and the patient space based on the two sets of matched point pairs. Based on the transformation matrix, the image space is mapped to the patient space to complete the spatial registration, provide the surgical navigation system with the required spatial position correspondence, and assist doctors in performing surgical operations.
[0027] The artificial marker of the present invention is pasted on the surface of the patient's skin. In the patient space, the center point of the marker can be identified using the continuous binary threshold method and binocular stereo vision principle; in the image space, the center point of the marker is obtained using a deep learning detection method, and then registration is completed, thereby effectively realizing spatial registration; avoiding the infection risks caused by invasive operations, creating extremely favorable conditions for the patient's rapid and smooth recovery after surgery, and greatly improving the patient's experience and safety during surgical treatment.
[0028] The detection network of this invention, trained with extensive data, accurately extracts marker features and obtains marker coordinates in image space. The binocular camera, combined with calibrated internal and external parameters, acquires marker coordinates in patient space. Both methods offer high precision and significantly simplify the registration process, reducing physician workload. They also effectively reduce errors caused by human factors, improving the efficiency and accuracy of spatial registration.
[0029] This invention utilizes a marker matching algorithm that analyzes the three-dimensional coordinates of markers in both image and patient space to accurately calculate the matching index between the two sets of point pairs. This process eliminates the need for the physician to manually select point pairs in a sequential order, significantly reducing the number of steps required. The resulting matching index establishes a more reliable and accurate spatial position correspondence, providing high-quality data support for the subsequent calculation of the transformation matrix using the closest point iterative algorithm, ultimately achieving precise mapping from image space to patient space. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a flow chart of the spatial registration method of a surgical navigation system based on feature point detection of the present invention.
[0031] Figure 2 Schematic diagram of the candidate detection network structure of the artificial marker A detection network of the present invention.
[0032] Figure 3 Schematic diagram of the classification and screening network structure of the artificial marker A detection network of the present invention.
[0033] Figure 4 Schematic diagram of the structure of the feature enhancement unit (FAU) of the present invention.
[0034] Figure 5 This is a schematic diagram of the overall artificial marker of the present invention attached to the surface of the patient's skin.
[0035] Figure 6 This is a schematic diagram of the overall artificial marker of the present invention after being separated into marker A and marker B. DETAILED DESCRIPTION
[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0037] See also Figure 1 As shown, in an exemplary embodiment of the present application, the spatial registration method of the surgical navigation system based on feature point detection is implemented by the following steps:
[0038] S1. Before surgery, attach at least four markers A that can be visualized on CT to the patient's skin, preprocess the patient's preoperative CT image data, and annotate the markers A;
[0039] S2. Using the marker A detection dataset to train the marker A detection network, and obtain the weight parameters of the trained detection network;
[0040] S3. The detection network processes the pre-processed preoperative CT image data based on the weight parameters to obtain the three-dimensional coordinates of the marker in the image space;
[0041] S4. During surgery, the binocular camera captures an image of a marker consisting of marker B attached to marker A, analyzes images of the same marker from different viewpoints, and obtains the three-dimensional coordinates of the marker in the patient's space.
[0042] S5. Determine the corresponding marker points in the image space and the patient space according to the three-dimensional coordinates of the marker in the image space and the patient space;
[0043] S6. Map the image space to the patient space based on the corresponding marker points in the image space and the patient space to complete the spatial registration.
[0044] For example, the marker A is a disposable ECG electrode sheet with a metal cylinder in the center, which is composed of a metal cylinder, a non-woven lining and a solid gel. The solid gel can be attached to the surface of human skin, and the metal cylinder can be clearly visualized in CT (Computed Tomography). Figure 5 、 Figure 6As shown. For example, 4-6 artificial markers A are attached to each patient's skin surface, although other numbers are possible. For example, after collecting preoperative CT images and performing data preprocessing, the doctor can use 3D Slicer software to annotate markers A to obtain the true three-dimensional coordinate labels of markers A.
[0045] Among them, in order to train the detection network, it is necessary to pre-construct the marker A detection data set and divide the data set into a training set and a test set in proportion (such as 8:2). The required detection network and network parameters are obtained through training and testing, which can then be used to process preoperative CT images to determine the three-dimensional coordinates of the marker in the image space.
[0046] Exemplarily, in step S1, the data preprocessing may be to use a Gaussian filter to smooth the image to remove image noise, resample the image, and normalize the data, such as unifying the voxel spacing of the image data to 1mm×1mm×1mm and the number of slices to 512×512×512.
[0047] For example, in step S2, the weight parameters of the detection network, including the learning rate, number of iterations, batch size, etc., are used to fully train the detection network model. After the training is completed, the weight parameters of the trained detection network model are obtained and used for subsequent processing of the preoperative CT image to determine the three-dimensional coordinates of the marker in the image space.
[0048] Exemplarily, in step S2, the detection network is improved based on Faster RCNN, including a candidate detection model and a classification screening model. The candidate detection stage module is used for the location information, diameter and anchor box target score of the candidate marker. The classification screening model classifies the candidate marker and obtains the probability that the candidate marker is a true marker. Through these two steps, the three-dimensional coordinates of the target marker are finally obtained.
[0049] like Figure 2 As shown, exemplarily, the detection network is improved based on Faster RCNN, and the candidate detection model includes a CNN encoder, a feature enhancement module, and a CNN decoder from the input side to the output side, and the feature enhancement module includes two parallel feature enhancement units;
[0050] The CNN encoder includes a first encoding module, a second encoding module, a third encoding module and a fourth encoding module connected in sequence, and the CNN decoder includes a first decoding module, a second decoding module and a third decoding module connected in sequence, wherein the second encoding module is connected to the third decoding module through a feature enhancement unit, and the third encoding module is connected to the second decoding module through another feature enhancement unit.
[0051] like Figure 2 As shown in the figure, the input of the candidate detection model is patient image data, which is passed through the CNN encoder and the output is a high-level semantic feature map. The candidate detection model contains two skip connections. The output of the second encoding module (Encoder Block2) of the CNN encoder passes through a feature enhancement unit (FAU) and is sent to the third decoding module (Decoder Block3) together with the output of the second decoding module (DecoderBlock2). The output of the third encoding module (Encoder Block3) of the CNN encoder passes through another feature enhancement unit (FAU) and is sent to the second decoding module (Decoder Block2) together with the output of the first decoding module (DecoderBlock1). The anchor mechanism is applied to the feature map of the last layer, and the output (output1) of the candidate detection model is finally obtained, which is a set of object region proposals. Each object region proposal has a 5*1 vector representing the location information, diameter and anchor box score of the candidate marker.
[0052] Due to GPU memory limitations, the input image size is 128×128×128. Assuming the original input dimensions are c×w×h×d (e.g., 1×128×128×128), the CNN encoder has four modules, each consisting of two 3×3×3 convolutional layers and a 3D max pooling layer. The convolutional layers have a kernel size of 3×3×3, a stride of 1, and a padding of 1. The pooling layer has a pooling window size of 2×2×2, a stride of 2, and a padding of 0. Local feature extraction is used to capture detailed information at different scales in the image, resulting in a feature map of c1×w / 16×h / 16×d / 16 (e.g., 64×8×8×8).
[0053] In the embodiment of the present application, the structure of the feature enhancement unit (FAU) is as follows Figure 4 As shown in the figure, it includes three convolutional layers arranged in parallel and a residual connection: the feature maps obtained by the three convolutional layers are spliced, and then passed through another convolutional layer. The output is residually connected with the original input as the output of the feature enhancement unit.
[0054] When the feature enhancement unit (FAU) processes data, it obtains feature maps δ, θ, and β through 1×1×1, 3×3×3, and 5×5×5 convolutions, respectively. The stride of the 1×1×1 convolution branch is 1 and the padding is 0, the stride of the 3×3×3 convolution branch is 1 and the padding is 1, and the stride of the 5×5×5 convolution branch is 1 and the padding is 2. δ, θ, and β are then concatenated (Concat) and then subjected to a 1×1×1 convolution to output data K, so that the size and number of channels of the feature map remain unchanged. In addition, a residual connection is added to alleviate the gradient vanishing problem. The initial input data X and the convolution output data K are connected to obtain the final output Y. The size of the feature map obtained by the feature enhancement unit (FAU) remains unchanged. The specific process is as follows: the input data is X and the final output data is Y:
[0055]
[0056] In formula (1), Indicates the Concat operation.
[0057] In the embodiment of this application, Figure 2 As shown, among the three decoder modules included in the decoder, DecoderBlock1 includes an upsampling layer; Decoder Block2 includes Concat splicing, an upsampling layer, and a convolution layer. The Concat is used to splice the output of Decoder Block1 with the output of Encoder Block3. The convolution uses a 3×3×3 convolution kernel, with stride and padding both of 1, to adjust the number of channels, resulting in c2×w / 4×h / 4×d / 4 (such as 32×32×32×32). The Decoder Block3 includes Concat splicing and a convolution layer. The Concat splices the output of DecoderBlock2 with the output of Encoder Block2. The convolution uses a 3×3×3 convolution kernel, with stride and padding both of 1, to adjust the number of channels. The upsampling of the first two decoder modules uses an interpolation algorithm with a scale factor of 2, which doubles the feature map. After the decoder, the final output (output1) is c3×w / 4×h / 4×d / 4 (such as 15×32×32×32).
[0058] In the embodiment of the present application, the candidate detection model adopts a multi-task loss function, which is defined as:
[0059] L=L cls +L loc ,
[0060] In formula (2), y i is the label of the i-th category, p i is the probability of the i-th class predicted by the model, and x is the difference between the predicted value and the true value.
[0061] In the candidate detection model, data enhancement is performed on positive samples by left-right flipping and random rotation to alleviate the overfitting problem. For example, the candidate detection model is trained for 150 rounds, and the batch size parameter (batch size) is set to 32.
[0062] To better balance the convergence speed and accuracy of the candidate detection models during training, a dynamic learning rate adjustment strategy was adopted. The initial learning rate was 0.01 to accelerate convergence. After 60 epochs, it was reduced to 0.001, enabling a more refined search in the parameter space. After 120 epochs, it was reduced to 0.0001 to avoid model oscillation near the optimal solution due to excessive learning rates, ultimately allowing the model to converge to a more optimal state.
[0063] In this application, the classification and screening model is centered on the three-dimensional coordinates of the candidate markers obtained by the candidate detection model, extracts a three-dimensional image block from the preprocessed CT image, such as a size of 24×24×24, as the input of the module, and then processes it together with the output of the candidate detection model as input.
[0064] like Figure 3 As shown, exemplarily, in an embodiment of the present application, the classification and screening model includes, from the input side to the output side, a three-dimensional image extraction unit, a convolution unit, a feature decoding module, an average pooling unit, and a fully connected layer unit; the feature decoding module is composed of three groups of feature decoding units connected in sequence, and each group of feature decoding units is composed of a maximum pooling unit and a feature enhancement unit connected.
[0065] When processing input data, the classification and screening model first performs a conventional 3×3×3 convolution, followed by three max pooling layers and a feature augmentation unit (FAU). The pooling window is set to 2×2×2, the stride is 2, and the padding is 0. Finally, the average pooling layer and the fully connected layer are used to obtain the output P(output2), which represents the probability that the candidate marker is a true marker. In addition, a dropout layer with a probability of 0.2 is added after each max pooling layer and the fully connected layer to prevent overfitting.
[0066] The convolutional and pooling layers in this network are both three-dimensional. Compared to two-dimensional representations, they can extract richer and more complete three-dimensional image features, which helps the network learn the characteristic information of the markers. This classification network fully utilizes the strong representational capabilities of deep hierarchical networks. However, as depth increases, planar networks are prone to gradient vanishing or gradient exploding problems. The residual connections in the Feature Augmentation Unit (FAU) effectively enhance the feature propagation information flow, which can alleviate the vanishing gradient problem.
[0067] In the embodiment of the present application, the classification screening model uses binary cross entropy error to calculate the classification loss, which is defined as: L(z) = -[ylog(σ(z)) + (1-y)log(1-σ(z))] (3)
[0068] In formula (3), y is the true label, z is the output of the model, and σ(z) is the Sigmoid function, which is used to convert the value to (0, 1) and calculate the probability value of the two classifications.
[0069] The classification and screening model was trained for 100 epochs with a batch size of 32. During training, the learning rate for the classification and screening model was initialized to 0.01, then to 0.001 after 10 epochs, and to 0.001 after 50 epochs. Batch normalization was applied to improve the regularization capability of the model.
[0070] In this application, both models of the detection network adopt stochastic gradient descent with momentum of 0.9.
[0071] In the embodiment of the present application, in step S3, after the preoperative CT image data enters the detection network, the detection network first preprocesses and crops it to obtain an image block of 128×128×128; then the candidate detection model and classification screening model are processed, and the confidence level is set to 0.99. Finally, the marker point set M in the image space is obtained. 1, m 2, …m n},n≥4.
[0072] For example, in the embodiment of the present application, in step S4, the marker B is fastened to the marker A on the patient's skin surface to form a complete marker, such as Figure 5 、 Figure 6 When obtaining the three-dimensional coordinates of the marker in the patient space, the internal and external parameters of the binocular camera obtained through calibration can be combined, and the continuous binary threshold method and binocular stereo vision principle can be used to analyze the images of the same marker taken by the binocular camera from different perspectives to obtain the three-dimensional coordinates of the marker in the patient space;
[0073] like Figure 5 、 Figure 6As shown, marker B can be 3D printed and consists of a white ring and a black dot, both 5mm thick. The white ring has an inner diameter of 12mm and an outer diameter of 20mm; the black dot has a diameter of 12mm and a cylindrical groove at the bottom, which can tightly fit with marker A.
[0074] Exemplarily, the specific steps of analyzing the images of the same marker captured by the binocular camera from different perspectives using the continuous binary threshold method and the binocular stereo vision principle to obtain the three-dimensional coordinates of the marker in the patient space may include:
[0075] (1) The continuous binary threshold method is used to binarize the video streams of the left and right cameras of the binocular camera, and the circularity and convexity are calculated to identify the area close to the circle and detect the markers in the left and right camera images, as shown in the following formula:
[0076]
[0077] In formula (4), S represents the square of the blob area, C represents the perimeter of the blob area, and H represents the convex square of the blob area.
[0078] (2) Gaussian similarity algorithm is used to perform stereo matching between the markers detected by the left and right cameras, combining epipolar constraints and similarity constraints to find the corresponding marker point pairs in the left and right camera images.
[0079] (3) Apply the stereo triangulation algorithm to the best matching pair to calculate the three-dimensional coordinate position of each marker and obtain the marker point set C = {c1, c2, c3, c4} in the patient space.
[0080] Through the above methods or steps, the three-dimensional coordinates of the marker in the patient space can be obtained.
[0081] Exemplarily, in an embodiment of the present application, in step S5, a marker matching algorithm is used to process the three-dimensional coordinates of the markers in the image space and the patient space, and the matching indexes of the two sets of point pairs are obtained to determine the corresponding marker points in the image space and the patient space.
[0082] For example, the marker matching algorithm can be implemented by calculating the distance between each pair of points and the angle between each three points in the two point sets. The distance calculation formula between each pair of points is as follows:
[0083] d ij =||m (i) -m (j) ||, d kl =||c (k) -c (l) || (5)
[0084] In formula (5), d ij represents the Euclidean distance between the i-th and j-th points detected in the image space, d kl represents the Euclidean distance between the kth and lth points detected in the patient space.
[0085] The formula for calculating the angle between three points is as follows:
[0086] θ ijk =∠(m (i) ,m (j) ,m (k) ),θ klm =∠(c (k) ,c (l) ,c (m) ) (6)
[0087] In formula (6), θ ijk represents the angle formed by the i-th, j-th and k-th points detected in the image space, θ klm represents the angle formed by the k-th point, l-th point, and m-th point detected in the patient space.
[0088] By comparing the distances and angles, the best match between the points in the image space and the patient space can be obtained, and the new point set in the image space is M′={m′1,m′2,m′3,m′4}, and the new point set in the patient space is C′={c′1,c′2,c′3,c′4}, which lays the foundation for the calculation of the spatial transformation matrix in the next step.
[0089] In an embodiment of the present application, in step S6, an alignment algorithm is used to process the two sets of matched point pairs input, calculate the transformation matrix of the two spaces, map the image space to the patient space, complete the spatial registration, provide the surgical navigation system with accurate spatial correspondence, and assist doctors in performing surgical operations.
[0090] Exemplarily, the specific steps of the alignment algorithm are as follows:
[0091] (1) Calculate the centroid of the two point sets of the image space point set and the patient space point set; the method is as follows:
[0092]
[0093] (2) Translation correction:
[0094] Move the center of mass of the two point sets to the same point, that is, there is only a rotation transformation relationship, and obtain the new point set after the two point sets are transformed; the expression is as follows:
[0095] M″={m1′-sm ,m2′-s m ,m3′-s m ,m4′-s m}, C″={c1′-s c ,c2′-s c ,c3′-s c ,c4′-s c}(8)
[0096] In formula (9), M″ represents the new point set obtained by transforming the points of the image space point set M′, C″ represents the new point set obtained by transforming the points of the image space point set C′, s c represents the centroid of the patient space point set C′, s m Represents the centroid of the obtained image space point set M′.
[0097] (3) Calculate the rotation matrix:
[0098] Based on M″ and C″, the singular value decomposition method is used to calculate the best fitting rotation matrix R.
[0099] H=M″ T C″, U, Σ, V=svd(H), R=V T U T (9)
[0100] (4) Calculate the translation vector t:
[0101] The centroid of the original target point set C′ is subtracted from the centroid of the rotated M′ by Rs. m calculate.
[0102] t=s c -Rs m (10)
[0103] (5) Based on the translation vector t, the transformation matrix R is expressed as a homogeneous coordinate matrix T; as shown below:
[0104]
[0105] In formula (11), R represents the obtained rotation matrix, and t represents the obtained translation vector.
[0106] Finally, the image space is mapped to the patient space through the homogeneous coordinate matrix T to complete the spatial registration, provide accurate spatial correspondence for the surgical navigation system, and assist doctors in performing precise surgical operations.
[0107] The basic principles, main features and advantages of the present invention are shown and described above. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments and that the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention.
[0108] The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are therefore intended to be embraced therein.
[0109] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A spatial registration method for a surgical navigation system based on feature point detection, characterized in that: The following steps are involved: S1. Before surgery, attach at least four markers A that can be visualized on CT to the patient's skin, preprocess the patient's preoperative CT image data, and annotate the markers A; S2. Using the marker A detection dataset to train the marker A detection network, and obtain the weight parameters of the trained detection network; S3. The detection network processes the pre-processed preoperative CT image data based on the weight parameters to obtain the three-dimensional coordinates of the marker in the image space; S4. During surgery, the binocular camera captures an image of a marker consisting of marker B attached to marker A, analyzes images of the same marker from different viewpoints, and obtains the three-dimensional coordinates of the marker in the patient's space. S5. Determine the corresponding marker points in the image space and the patient space according to the three-dimensional coordinates of the marker in the image space and the patient space; S6. Map the image space to the patient space based on the corresponding marker points in the image space and the patient space to complete the spatial registration.
2. The spatial registration method for a surgical navigation system based on feature point detection according to claim 1, characterized in that: The marker A is a disposable ECG electrode with a metal cylinder in the center. The marker B is a cylinder that forms a detachable connection with the metal cylinder of the marker A through a slot in the center. The marker B consists of an inner black circular part and an outer white ring part.
3. The spatial registration method for a surgical navigation system based on feature point detection according to claim 1, characterized in that: The detection network includes: A candidate detection model is used to obtain the location information, diameter, and anchor box target score of candidate markers based on patient imaging data; The classification and screening model is used to classify candidate markers based on patient image data and the output data of the candidate detection model to determine the probability that they are true markers and obtain the three-dimensional coordinates of the target markers.
4. The spatial registration method for a surgical navigation system based on feature point detection according to claim 3, characterized in that: The detection network is improved based on Faster RCNN. The candidate detection model includes a CNN encoder, a feature enhancement module, and a CNN decoder from the input side to the output side. The feature enhancement module includes two parallel feature enhancement units. The CNN encoder includes a first encoding module, a second encoding module, a third encoding module and a fourth encoding module connected in sequence, and the CNN decoder includes a first decoding module, a second decoding module and a third decoding module connected in sequence, wherein the second encoding module is connected to the third decoding module through a feature enhancement unit, and the third encoding module is connected to the second decoding module through another feature enhancement unit.
5. The spatial registration method for a surgical navigation system based on feature point detection according to claim 3, characterized in that: The classification and screening model includes, from the input side to the output side, a three-dimensional image extraction unit, a convolution unit, a feature decoding module, an average pooling unit, and a fully connected layer unit; the feature decoding module is composed of three groups of feature decoding units connected in sequence, and each group of feature decoding units is composed of a maximum pooling unit and a feature enhancement unit connected.
6. The spatial registration method for a surgical navigation system based on feature point detection according to claim 4 or 5, characterized in that: The feature enhancement unit includes three convolutional layers arranged in parallel and a residual connection: the feature maps obtained by the three convolutional layers are spliced, and then passed through a convolutional layer, and the output obtained is residually connected with the original input as the output of the feature enhancement unit.
7. The spatial registration method for a surgical navigation system based on feature point detection according to claim 1, characterized in that: In step S4, the image of the marker formed by the marker B fastened on the marker A captured by the binocular camera is processed by using the continuous binary threshold method and the binocular stereo vision principle in combination with the internal and external parameters of the binocular camera obtained through calibration.
8. The spatial registration method for a surgical navigation system based on feature point detection according to claim 1, characterized in that: In step S5, the three-dimensional coordinates of the markers in the image space and the patient space are processed using a marker matching algorithm to obtain matching indexes of the two sets of point pairs and determine corresponding marker points in the image space and the patient space.
9. The spatial registration method for a surgical navigation system based on feature point detection according to claim 1, characterized in that: In step S6, the input alignment algorithm is used to calculate the transformation matrix between the image space and the patient space based on the two sets of matched point pairs. Based on the transformation matrix, the image space is mapped to the patient space to complete the spatial registration, provide the required spatial position correspondence for the surgical navigation system, and assist doctors in performing surgical operations.