A method and system for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks
By using a neural network method based on VGG networks, real-time acquisition and annotation of recurrent laryngeal nerve videos were performed, and parameters were dynamically adjusted. This solved the problem of low accuracy in real-time identification and dynamic tracking of the recurrent laryngeal nerve, and achieved stable real-time identification and tracking of the recurrent laryngeal nerve.
Patent Information
- Application Number
- CN202511019088.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing methods for real-time identification and dynamic tracking of the recurrent laryngeal nerve have low accuracy. Traditional intraoperative neurophysiological monitoring techniques lack sensitivity. Deep learning models struggle to capture the temporal correlation between consecutive frames in dynamic video scenes. Single-lens video segmentation techniques cannot achieve stable localization between consecutive frames.
A neural network-based approach based on VGG networks was adopted. By acquiring intraoperative video of recurrent laryngeal nerve surgery in real time, performing image frame binarization annotation, constructing a neural network for feature extraction, and aligning edge features, semantic features, and specific features of the recurrent laryngeal nerve, the neural network parameters were dynamically adjusted to achieve real-time identification and dynamic tracking of the recurrent laryngeal nerve.
It improves the accuracy of real-time identification and dynamic tracking of the recurrent laryngeal nerve, enabling rapid focus on the current morphology of the patient's recurrent laryngeal nerve, adapting to dynamic changes during surgery, and achieving stable real-time identification and tracking.
Smart Images

Figure CN120526290B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, specifically relating to a method and system for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks. Background Technology
[0002] The recurrent laryngeal nerve (RLN) is easily damaged during thyroid surgery due to direct cutting, thermal injury, or traction. Traditional intraoperative neurophysiological monitoring (IONM) technology is not only insufficient in sensitivity, but also cannot locate the nerve position in real time.
[0003] While the da Vinci robotic surgical system can improve operational precision, it relies on the surgeon's experience and suffers from problems such as surgeon fatigue and time-consuming nerve localization.
[0004] Current deep learning models lack generalization ability and struggle to adapt to dynamic intraoperative scenarios. For example, while U-Net (U-shaped Network) performs well in medical static image segmentation, its encoder-decoder structure struggles to capture the temporal correlation between consecutive frames in dynamic video scenes, leading to jitter in segmentation results during tracking (e.g., Dice coefficient fluctuations exceeding ±0.15). Fully Convolutional Networks (FCNs) lack effective modeling of multi-scale contextual information, resulting in a significant drop in segmentation accuracy when intraoperative lighting changes or tissue deformation occurs (IOU values drop from 0.85 to below 0.60; IOU stands for Intersection over Union). Mask R-CNN (MaskRegion-Based Convolutional Neural Network), while suitable for instance segmentation, suffers from insufficient real-time performance due to its two-stage detection mechanism (single-frame processing time > 2 seconds), failing to meet intraoperative real-time requirements.
[0005] Current single-lens video segmentation techniques cannot effectively integrate with intraoperative target tracking, failing to achieve stable localization between consecutive frames. For example, tracking algorithms based on Siamese Networks rely on fixed template matching, resulting in tracking drift errors exceeding 2mm when intraoperative neural morphology changes or occlusion occur. While DeepLab (a semantic segmentation series of models) expands the receptive field through dilated convolutions, it lacks sufficient modeling of inter-frame correlations for dynamic targets (such as the intraoperatively moving recurrent laryngeal nerve), leading to inconsistencies between segmentation and tracking results. Traditional KCF (Kernelized Correlation Filters) relies on manual feature extraction, exhibiting insufficient sensitivity to neural texture in complex surgical contexts, resulting in a false detection rate as high as 30%.
[0006] In summary, current methods for real-time identification and dynamic tracking of the recurrent laryngeal nerve suffer from low accuracy. Summary of the Invention
[0007] The purpose of this invention is to provide a method and system for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks, in order to solve the problem of low accuracy in the real-time identification and dynamic tracking of the recurrent laryngeal nerve in the prior art.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on a neural network, comprising the following steps:
[0010] Real-time acquisition of intraoperative video of recurrent laryngeal nerve surgery, and conversion of the real-time acquired intraoperative video of recurrent laryngeal nerve surgery into image frames;
[0011] Binarize and annotate the image frames to obtain several frames of annotated recurrent laryngeal nerve region images;
[0012] A neural network was constructed, and features were extracted from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0013] Select a set number of frame-annotated images of the recurrent laryngeal nerve region, align edge features and semantic features with specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters.
[0014] Based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters.
[0015] The recurrent laryngeal nerve is identified and dynamically tracked in real time using the final neural network parameters in the intraoperative video of the recurrent laryngeal nerve, and the real-time identification and dynamic tracking results of the recurrent laryngeal nerve are obtained.
[0016] A further improvement of the present invention is that the neural network is a VGG (Visual Geometry Group) network.
[0017] A further improvement of this invention lies in the construction of a neural network, which is used to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve, specifically including:
[0018] Data augmentation was performed on several frames of labeled recurrent laryngeal nerve region images to obtain several frames of augmented recurrent laryngeal nerve region images;
[0019] A training set was constructed using several frames of enhanced images of the recurrent laryngeal nerve region.
[0020] A neural network is constructed, and it is pre-trained using the ImageNet and Davis datasets to obtain a pre-trained neural network.
[0021] The pre-trained neural network is fine-tuned based on the constructed training set to obtain the fine-tuned neural network.
[0022] By using a fine-tuned neural network, feature extraction was performed on several frames of enhanced images of the recurrent laryngeal nerve region to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0023] A further improvement of this invention lies in the fact that the feature extraction of several frames of enhanced recurrent laryngeal nerve region images using a fine-tuned neural network yields edge features, semantic features, and specific features of the recurrent laryngeal nerve, specifically including:
[0024] Freeze the parameters of the first set convolutional group of the fine-tuned neural network to obtain edge features and semantic features;
[0025] Further adjustments to the parameters of the remaining convolutional groups in the fine-tuned neural network yielded specific features of the recurrent laryngeal nerve.
[0026] A further improvement of this invention is that the loss function used during the pre-training of the neural network is the Dice loss function, the expression of which is:
[0027]
[0028] in, Represents the Dice loss function. and They represent the first The predicted and actual values of each pixel. This represents the total number of pixels. Represents the smoothing factor. This indicates the number of pixels.
[0029] A further improvement of this invention lies in that the initial neural network parameters are dynamically adjusted based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames to obtain the final neural network parameters, specifically including:
[0030] The second set of convolutional groups of the fine-tuned neural network is used to extract features from the laryngeal nerve region image of a set number of frame annotations to obtain specific features of the recurrent laryngeal nerve.
[0031] Store specific features of the recurrent laryngeal nerve as a reference set;
[0032] A fine-tuned neural network was used to extract features from real-time acquired images of the recurrent laryngeal nerve region to obtain intraoperative video stream features of the recurrent laryngeal nerve.
[0033] Real-time calculation of semantic similarity between intraoperative video stream features and reference set during recurrent laryngeal nerve surgery;
[0034] When the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters.
[0035] A further improvement of this invention is that, when the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted using the momentum accumulation gradient descent method to obtain the final neural network parameters. The specific steps of the momentum accumulation gradient descent method include:
[0036] The formula for calculating the final parameter change of the initial neural network after the last adjustment is as follows:
[0037]
[0038] in, This represents the final parameter change of the initial neural network after the last adjustment. This represents the change in parameters of the initial neural network after the last adjustment. Represents the Dice loss function. Represents the gradient. Momentum factor;
[0039] The parameters of the initial neural network after the update are obtained using the following formula:
[0040]
[0041] in, and These represent the parameters of the neural network before and after initial adjustment, respectively. This represents the final parameter change after the initial neural network adjustment. This is the learning rate.
[0042] In a second aspect, the present invention provides a real-time identification and dynamic tracking system for the recurrent laryngeal nerve based on a neural network, including a data acquisition module, a binarization annotation module, a feature extraction module, an initial network parameter acquisition module, a final network parameter acquisition module, and a real-time identification and dynamic tracking module for the recurrent laryngeal nerve.
[0043] The data acquisition module is used to collect intraoperative video of recurrent laryngeal nerve surgery in real time and convert the real-time collected intraoperative video of recurrent laryngeal nerve surgery into image frames;
[0044] The binarization annotation module is used to perform binarization annotation on image frames to obtain several frames of annotated recurrent laryngeal nerve region images;
[0045] The feature extraction module is used to construct a neural network, and uses the neural network to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features and specific features of the recurrent laryngeal nerve.
[0046] The initial network parameter acquisition module is used to select a set number of frame-annotated recurrent laryngeal nerve region images, align edge features and semantic features with specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters.
[0047] The final network parameter acquisition module is used to dynamically adjust the initial neural network parameters based on the semantic similarity between a set number of frame-annotated recurrent laryngeal nerve region images and real-time acquired recurrent laryngeal nerve region image frames, and obtain the final neural network parameters.
[0048] The recurrent laryngeal nerve real-time identification and dynamic tracking module is used to identify and dynamically track the recurrent laryngeal nerve in real-time using the final neural network parameters in the real-time acquired intraoperative video of the recurrent laryngeal nerve, and obtain the real-time identification result and dynamic tracking result of the recurrent laryngeal nerve.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] This invention is an improved version. Compared with existing methods for real-time identification and dynamic tracking of the recurrent laryngeal nerve, this invention selects a set number of frame-annotated images of the recurrent laryngeal nerve region, aligns edge features and semantic features with specific features of the recurrent laryngeal nerve, and adjusts the parameters of the neural network to obtain initial neural network parameters. This operation allows the neural network to quickly focus on the morphology of the current patient's recurrent laryngeal nerve. On the other hand, this invention dynamically adjusts the initial neural network parameters based on the semantic similarity between the set number of frame-annotated images of the recurrent laryngeal nerve region and real-time acquired image frames of the recurrent laryngeal nerve region to obtain final neural network parameters. This operation allows the neural network to better capture the semantic features of the recurrent laryngeal nerve region images. When encountering recurrent laryngeal nerve region images that are semantically similar but slightly different, the neural network can dynamically adjust its parameters according to the degree of semantic similarity, thereby accurately identifying and dynamically tracking the recurrent laryngeal nerve in real time. This effectively solves the problem of low accuracy in real-time identification and dynamic tracking of the recurrent laryngeal nerve in existing technologies. Attached Figure Description
[0051] Figure 1 This is a flowchart of the real-time identification and dynamic tracking method for the recurrent laryngeal nerve based on neural networks according to the present invention;
[0052] Figure 2 This is a schematic diagram of the real-time identification and dynamic tracking system for the recurrent laryngeal nerve based on neural networks according to the present invention;
[0053] Figure 3 This is a structural diagram of the VGG network after fine-tuning according to the present invention. Detailed Implementation
[0054] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.
[0055] This invention proposes a neural network-based real-time identification and dynamic tracking method for the recurrent laryngeal nerve. It selects a predetermined number of labeled frames of the recurrent laryngeal nerve region image, aligns edge and semantic features with specific features of the recurrent laryngeal nerve, and adjusts the neural network parameters to obtain initial neural network parameters. Based on the semantic similarity between the labeled frames and real-time acquired frames of the recurrent laryngeal nerve region image, the initial neural network parameters are dynamically adjusted to obtain final neural network parameters. These final neural network parameters are then used to perform real-time identification and dynamic tracking of the recurrent laryngeal nerve in real-time acquired intraoperative video, yielding real-time identification and dynamic tracking results. Compared to existing technologies, this invention effectively solves the problem of low accuracy in real-time identification and dynamic tracking of the recurrent laryngeal nerve in existing technologies.
[0056] Example 1:
[0057] The flowchart of the real-time identification and dynamic tracking method of the recurrent laryngeal nerve based on neural networks of this invention is as follows: Figure 1 As shown, the real-time identification and dynamic tracking method for the recurrent laryngeal nerve based on neural networks of the present invention includes the following steps:
[0058] S1. Real-time acquisition of intraoperative video of recurrent laryngeal nerve surgery, and conversion of the real-time acquired intraoperative video of recurrent laryngeal nerve surgery into image frames;
[0059] S2. Binarize and annotate the image frames to obtain several annotated images of the recurrent laryngeal nerve region;
[0060] S3. Construct a neural network and use the neural network to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve;
[0061] S4. Select a set number of frame-annotated images of the recurrent laryngeal nerve region, align the edge features and semantic features with the specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters.
[0062] S5. Based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters.
[0063] S6. Using the final neural network parameters, the recurrent laryngeal nerve in the intraoperative video of the recurrent laryngeal nerve is identified and dynamically tracked in real time to obtain the real-time identification results and dynamic tracking results of the recurrent laryngeal nerve.
[0064] Example 2:
[0065] A schematic diagram of the real-time identification and dynamic tracking system for the recurrent laryngeal nerve based on neural networks of this invention is shown below. Figure 2 As shown, the recurrent laryngeal nerve real-time identification and dynamic tracking system based on neural networks of the present invention includes a data acquisition module, a binarization annotation module, a feature extraction module, an initial network parameter acquisition module, a final network parameter acquisition module, and a recurrent laryngeal nerve real-time identification and dynamic tracking module.
[0066] The data acquisition module is used to collect intraoperative video of recurrent laryngeal nerve surgery in real time and convert the real-time collected intraoperative video of recurrent laryngeal nerve surgery into image frames.
[0067] The binarization annotation module is used to perform binarization annotation on image frames, resulting in several frames of annotated images of the recurrent laryngeal nerve region.
[0068] The feature extraction module is used to construct a neural network. The neural network is used to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0069] The initial network parameter acquisition module is used to select a set number of frame-annotated images of the recurrent laryngeal nerve region, align edge features and semantic features with specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters.
[0070] The final network parameter acquisition module is used to dynamically adjust the initial neural network parameters based on the semantic similarity between a set number of frame-annotated recurrent laryngeal nerve region images and real-time acquired recurrent laryngeal nerve region image frames, and obtain the final neural network parameters.
[0071] The recurrent laryngeal nerve real-time identification and dynamic tracking module is used to identify and dynamically track the recurrent laryngeal nerve in real-time using the final neural network parameters acquired in the intraoperative video of the recurrent laryngeal nerve, and obtain the real-time identification results and dynamic tracking results of the recurrent laryngeal nerve.
[0072] Example 3:
[0073] The present invention provides a method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks, comprising the following steps:
[0074] S1. Real-time acquisition of intraoperative video of recurrent laryngeal nerve surgery, and conversion of the real-time acquired intraoperative video of recurrent laryngeal nerve surgery into image frames.
[0075] First, real-time video of the recurrent laryngeal nerve surgery was acquired (spatial resolution 1080p / frame rate 30fps), and then the real-time video of the recurrent laryngeal nerve surgery was converted into image frames.
[0076] S2. Binarize and annotate the image frames to obtain several frames of annotated images of the recurrent laryngeal nerve region.
[0077] S3. Construct a neural network and use the neural network to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0078] This step constructs a neural network (in this embodiment, the neural network is a VGG network), and uses the neural network to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features (e.g., the boundary between the nerve perineurium and surrounding adipose tissue), semantic features, and specific features of the recurrent laryngeal nerve (e.g., the spatial topological relationship between the recurrent laryngeal nerve and the inferior thyroid artery). Specifically, this includes:
[0079] A. Perform data augmentation on several frames of labeled recurrent laryngeal nerve region images to obtain several frames of augmented recurrent laryngeal nerve region images;
[0080] B. Construct a training set using several frames of enhanced images of the recurrent laryngeal nerve region;
[0081] C. Construct a neural network, and pre-train the neural network using the ImageNet dataset and the Davis dataset to obtain a pre-trained neural network;
[0082] D. Fine-tune the pre-trained neural network based on the constructed training set to obtain the fine-tuned neural network;
[0083] E. Using a fine-tuned neural network, feature extraction is performed on several frames of enhanced recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0084] In step A, rotation, translation, scaling, and flip transformations are used to perform data augmentation on several frames of labeled recurrent laryngeal nerve region images. The rotation, translation, scaling, and flip transformations are performed in no particular order.
[0085] In step D, the loss function used for pre-training the VGG network is the Dice loss function, and the expression for the Dice loss function is:
[0086]
[0087] in, Represents the Dice loss function. and They represent the first The predicted and actual values of each pixel. This represents the total number of pixels. This represents the smoothing factor; in this embodiment, the smoothing factor is... Values To prevent the denominator from being zero in the Dice loss function expression, This indicates the number of pixels.
[0088] The structure of the fine-tuned VGG network in this embodiment is as follows: Figure 3 As shown, the fine-tuned VGG network consists of 5 convolutional groups. The first convolutional group consists of two layers (Conv1-1 and Conv1-2), the second convolutional group consists of two layers (Conv2-1 and Conv2-2), the third convolutional group consists of three layers (Conv3-1, Conv3-2, and Conv3-3), the fourth convolutional group consists of three layers (Conv4-1, Conv4-2, and Conv4-3), and the fifth convolutional group consists of three layers (Conv5-1, Conv5-2, and Conv5-3).
[0089] In step E, a fine-tuned neural network (VGG network) is used to extract features from several frames of enhanced recurrent laryngeal nerve region images, obtaining edge features, semantic features, and specific features of the recurrent laryngeal nerve, specifically including:
[0090] a. Freeze the parameters (e.g., weights and biases of the convolutional layers) of the first set convolutional group (first, second, and third convolutional groups) of the fine-tuned neural network (VGG network) to obtain edge features and semantic features;
[0091] b. Further adjust the parameters of the remaining convolutional groups (the fourth and fifth convolutional groups) of the fine-tuned neural network (VGG network) to obtain specific features of the recurrent laryngeal nerve.
[0092] In steps a and b, the first set of convolutional groups and the remaining convolutional groups can be adjusted according to actual needs.
[0093] In this step, the Dice coefficient and cross-union ratio are used to evaluate the performance of the fine-tuned neural network. The evaluation results of the Dice coefficient and cross-union ratio are shown in Table 1.
[0094] Table 1. Evaluation results of Dice coefficient and crossover ratio.
[0095]
[0096] As can be seen from the evaluation results in Table 1, the evaluation metrics (Dice coefficient and crossover ratio) of the fine-tuned VGG network are higher than those of the commonly used fully convolutional neural network. Therefore, this implementation chooses to use the fine-tuned VGG network to extract features from several frames of enhanced recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve.
[0097] S4. Select a set number of frame-annotated images of the recurrent laryngeal nerve region, align the edge features and semantic features with the specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters.
[0098] This step specifically selects two labeled images of the recurrent laryngeal nerve region; the number can be adjusted according to actual needs.
[0099] This step specifically employs a feature space adaptive method to align edge features and semantic features with specific features of the recurrent laryngeal nerve. Specifically, by further adjusting the intermediate layer parameters of the fine-tuned VGG network (the fourth and fifth convolutional groups), the fine-tuned VGG network can quickly focus on the current patient's recurrent laryngeal nerve morphology and better extract the specific features of the current patient's recurrent laryngeal nerve.
[0100] S5. Based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters.
[0101] The specific steps involved in obtaining the final neural network parameters include:
[0102] A. The second set of convolutional groups (the fifth convolutional group) of the fine-tuned neural network (VGG network) is used to extract features from the laryngeal nerve region images of a set number of frames to obtain specific features of the recurrent laryngeal nerve.
[0103] B. Store specific features of the recurrent laryngeal nerve as a reference set. The formula for calculating the reference set is:
[0104]
[0105] in, Represents a reference set. It represents the set of real numbers.
[0106] C. The finely tuned neural network is used to extract features from the real-time acquired recurrent laryngeal nerve region image frames to obtain the intraoperative video stream features of the recurrent laryngeal nerve.
[0107] D. Real-time calculation of the semantic similarity between the video stream features during recurrent laryngeal nerve surgery and the reference set. The formula for calculating the semantic similarity between the video stream features during recurrent laryngeal nerve surgery and the reference set is as follows:
[0108]
[0109] in, Indicates calculation and semantic similarity, Indicates calculation and The inner product, Indicates will The result is normalized to the range of 0-1. Features of the video stream during recurrent laryngeal nerve surgery. Represents a reference set. The temperature coefficient is the temperature coefficient in this embodiment. The value is 0.07.
[0110] E. When the semantic similarity is less than a set threshold (the threshold is set to 0.7 in this embodiment), the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters.
[0111] Step E specifically employs the momentum accumulation gradient descent method to dynamically adjust the initial neural network parameters, obtaining the final neural network parameters. The specific steps of the momentum accumulation gradient descent method include:
[0112] The final parameter change after the last update on the fine-tuned VGG network is obtained using the following formula:
[0113]
[0114] in, This represents the final parameter change after the last update (also called adjustment) on the fine-tuned VGG network. This indicates the amount of parameter change on the fine-tuned VGG network after the last update. Represents the Dice loss function. Represents the gradient. The momentum factor is used in this embodiment. The value is 0.9.
[0115] The parameters of the fine-tuned VGG network after the update are obtained using the following formula:
[0116]
[0117] in, and These represent the parameters of the VGG network before and after the update, respectively. This represents the final parameter change after the fine-tuned VGG network update. The learning rate is the learning rate used in this embodiment. The value is 0.001.
[0118] S6. Using the final neural network parameters, the recurrent laryngeal nerve in the intraoperative video of the recurrent laryngeal nerve is identified and dynamically tracked in real time to obtain the real-time identification results and dynamic tracking results of the recurrent laryngeal nerve.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks, characterized in that, Includes the following steps: Real-time acquisition of intraoperative video of recurrent laryngeal nerve surgery, and conversion of the real-time acquired intraoperative video of recurrent laryngeal nerve surgery into image frames; Binarize and annotate the image frames to obtain several frames of annotated recurrent laryngeal nerve region images; A neural network was constructed, and features were extracted from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve. The neural network is constructed to extract features from several frames of labeled recurrent laryngeal nerve region images, obtaining edge features, semantic features, and specific features of the recurrent laryngeal nerve, specifically including: Data augmentation was performed on several frames of labeled recurrent laryngeal nerve region images to obtain several frames of augmented recurrent laryngeal nerve region images; A training set was constructed using several frames of enhanced images of the recurrent laryngeal nerve region. A neural network is constructed, and it is pre-trained using the ImageNet and Davis datasets to obtain a pre-trained neural network. The pre-trained neural network is fine-tuned based on the constructed training set to obtain the fine-tuned neural network. By using a fine-tuned neural network, feature extraction was performed on several frames of enhanced images of the recurrent laryngeal nerve region to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve. Select a set number of frame-annotated images of the recurrent laryngeal nerve region, align edge features and semantic features with specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters. Based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters. The initial neural network parameters are dynamically adjusted based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames to obtain the final neural network parameters. Specifically, this includes: The second set of convolutional groups of the fine-tuned neural network is used to extract features from the laryngeal nerve region image of a set number of frame annotations to obtain specific features of the recurrent laryngeal nerve. Store specific features of the recurrent laryngeal nerve as a reference set; A fine-tuned neural network was used to extract features from real-time acquired images of the recurrent laryngeal nerve region to obtain intraoperative video stream features of the recurrent laryngeal nerve. Real-time calculation of semantic similarity between intraoperative video stream features and reference set during recurrent laryngeal nerve surgery; When the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters. When the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted using the momentum accumulation gradient descent method to obtain the final neural network parameters. The specific steps of the momentum accumulation gradient descent method include: The formula for calculating the final parameter change of the initial neural network after the last adjustment is as follows: in, This represents the final parameter change of the initial neural network after the last adjustment. This represents the change in parameters of the initial neural network after the last adjustment. Represents the Dice loss function. Represents the gradient. Momentum factor; The initial parameters of the adjusted neural network are obtained using the following formula: in, and These represent the parameters of the neural network before and after initial adjustment, respectively. This represents the final parameter change after the initial neural network adjustment. The learning rate; The recurrent laryngeal nerve is identified and dynamically tracked in real time using the final neural network parameters in the intraoperative video of the recurrent laryngeal nerve, and the real-time identification and dynamic tracking results of the recurrent laryngeal nerve are obtained.
2. The method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks according to claim 1, characterized in that, The neural network is a VGG network.
3. The method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on a neural network according to claim 1, wherein the step of using a fine-tuned neural network to extract features from several frames of enhanced recurrent laryngeal nerve region images to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve specifically includes: Freeze the parameters of the first set convolutional group of the fine-tuned neural network to obtain edge features and semantic features; Further adjustments to the parameters of the remaining convolutional groups in the fine-tuned neural network yielded specific features of the recurrent laryngeal nerve.
4. The method for real-time identification and dynamic tracking of the recurrent laryngeal nerve based on neural networks according to claim 1, characterized in that, The loss function used when pre-training a neural network is the Dice loss function, and its expression is: in, Represents the Dice loss function. and They represent the first The predicted and actual values of each pixel. This represents the total number of pixels. Represents the smoothing factor. This indicates the number of pixels.
5. A real-time identification and dynamic tracking system for the recurrent laryngeal nerve based on a neural network, characterized in that, It includes a data acquisition module, a binarization annotation module, a feature extraction module, an initial network parameter acquisition module, a final network parameter acquisition module, and a real-time recurrent laryngeal nerve recognition and dynamic tracking module; The data acquisition module is used to collect intraoperative video of recurrent laryngeal nerve surgery in real time and convert the real-time collected intraoperative video of recurrent laryngeal nerve surgery into image frames; The binarization annotation module is used to perform binarization annotation on image frames to obtain several frames of annotated recurrent laryngeal nerve region images; The feature extraction module is used to construct a neural network, and uses the neural network to extract features from several frames of labeled recurrent laryngeal nerve region images to obtain edge features, semantic features and specific features of the recurrent laryngeal nerve. The neural network is constructed to extract features from several frames of labeled recurrent laryngeal nerve region images, obtaining edge features, semantic features, and specific features of the recurrent laryngeal nerve, specifically including: Data augmentation was performed on several frames of labeled recurrent laryngeal nerve region images to obtain several frames of augmented recurrent laryngeal nerve region images; A training set was constructed using several frames of enhanced images of the recurrent laryngeal nerve region. A neural network is constructed, and it is pre-trained using the ImageNet and Davis datasets to obtain a pre-trained neural network. The pre-trained neural network is fine-tuned based on the constructed training set to obtain the fine-tuned neural network. By using a fine-tuned neural network, feature extraction was performed on several frames of enhanced images of the recurrent laryngeal nerve region to obtain edge features, semantic features, and specific features of the recurrent laryngeal nerve. The initial network parameter acquisition module is used to select a set number of frame-annotated recurrent laryngeal nerve region images, align edge features and semantic features with specific features of the recurrent laryngeal nerve, adjust the parameters of the neural network, and obtain the initial neural network parameters. The final network parameter acquisition module is used to dynamically adjust the initial neural network parameters based on the semantic similarity between a set number of frame-annotated recurrent laryngeal nerve region images and real-time acquired recurrent laryngeal nerve region image frames, and obtain the final neural network parameters. The initial neural network parameters are dynamically adjusted based on the semantic similarity between the recurrent laryngeal nerve region image with a set number of frame annotations and the real-time acquired recurrent laryngeal nerve region image frames to obtain the final neural network parameters. Specifically, this includes: The second set of convolutional groups of the fine-tuned neural network is used to extract features from the laryngeal nerve region image of a set number of frame annotations to obtain specific features of the recurrent laryngeal nerve. Store specific features of the recurrent laryngeal nerve as a reference set; A fine-tuned neural network was used to extract features from real-time acquired images of the recurrent laryngeal nerve region to obtain intraoperative video stream features of the recurrent laryngeal nerve. Real-time calculation of semantic similarity between intraoperative video stream features and reference set during recurrent laryngeal nerve surgery; When the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted to obtain the final neural network parameters. When the semantic similarity is less than a set threshold, the initial neural network parameters are dynamically adjusted using the momentum accumulation gradient descent method to obtain the final neural network parameters. The specific steps of the momentum accumulation gradient descent method include: The formula for calculating the final parameter change of the initial neural network after the last adjustment is as follows: in, This represents the final parameter change of the initial neural network after the last adjustment. This represents the change in parameters of the initial neural network after the last adjustment. Represents the Dice loss function. Represents the gradient. Momentum factor; The initial parameters of the adjusted neural network are obtained using the following formula: in, and These represent the parameters of the neural network before and after initial adjustment, respectively. This represents the final parameter change after the initial neural network adjustment. The learning rate; The recurrent laryngeal nerve real-time identification and dynamic tracking module is used to identify and dynamically track the recurrent laryngeal nerve in real-time using the final neural network parameters in the real-time acquired intraoperative video of the recurrent laryngeal nerve, and obtain the real-time identification result and dynamic tracking result of the recurrent laryngeal nerve.
6. The real-time identification and dynamic tracking system for the recurrent laryngeal nerve based on a neural network according to claim 5, characterized in that, The neural network is a VGG network.
Citation Information
Patent Citations
Method and system for identifying and tracking recurrent laryngeal nerves in endoscopic thyroid surgery
CN117710324A
Visual location identification method and system based on image enhancement and scene semantic optimization
CN118196484A