Real-time navigation method for glaucoma operation based on deep learning

Through deep learning multi-task model, the surgical stage and device anatomy structure in glaucoma surgery are identified in real time, and the problems of navigation delay and misjudgment in the prior art are solved, providing low-latency and high-precision intraoperative decision-making support, improving surgical safety and standardization level.

CN120241245APending Publication Date: 2025-07-04ZHONGSHAN OPHTHALMIC CENT SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357117.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing AI-assisted technology cannot handle the interaction timing between instruments and tissues in dynamic surgical videos during glaucoma surgery, resulting in high navigation delays, high misjudgment rates, and lack of real-time dynamic intraoperative decision support tools. Relying on high-performance computing devices cannot achieve low-latency inference on conventional hardware, and lacks deep coordination with microscope optical systems.

Method used

A multi-task model based on deep learning, including step recognition model and target tracking model, obtain video streams through surgical microscopy, utilize Transformer architecture and PIDNet's real-time semantic segmentation network, identify surgical stages and instrument anatomy in real time, and trigger voice prompts to generate navigation video streams.

Benefits of technology

It achieves low-latency and high-precision intraoperative decision-making support, improves the accuracy and safety of surgical navigation, shortens the doctor's learning curve, and improves the level of operation standardization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120241245A_ABST
    Figure CN120241245A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and medical assistance, in particular to a glaucoma operation real-time navigation method based on deep learning, and the method mainly comprises the steps: inputting to-be-processed data into a pre-trained multi-task deep learning model; the step recognition model is constructed based on a Transform architecture and is used for extracting spatial and temporal characteristics of continuous frames in the to-be-processed data and outputting an operation stage corresponding to the current frame according to the extracted characteristics and a preset operation step definition; the target tracking model is constructed based on a real-time semantic segmentation network of PIDNet, and is used for identifying and tracking an instrument and an anatomical structure based on a single-frame image in the to-be-processed data, and outputting a target identification result; and triggering a predefined voice prompt according to the operation stage corresponding to the current frame and the target recognition result. According to the method, the problem of multi-task real-time collaborative navigation in a dynamic operation scene is solved, and low-delay and high-precision intra-operation decision support is provided for medical staff.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision and medical assistance technologies, and particularly to a real-time navigation method for glaucoma surgery based on deep learning. Background Art

[0002] As the leading cause of irreversible blindness globally, the mainstream treatment for glaucoma relies on minimally invasive surgeries such as goniotomy, but precise localization of submillimeter structures such as the trabecular meshwork is required during the operation. In the surgical videos of glaucoma, the variability is much greater compared to the images or video results from radiography or endoscopy. This variability is not only manifested in aspects such as microscope focusing, magnification, image quality, background noise, and color deviation, but also in the interaction between surgical instruments and anatomical structures. Inappropriate interactions between instruments can lead to insufficient exposure or partial visibility of key structures, thus affecting the subsequent surgical process. Most existing AI-assisted technologies achieve anatomical segmentation based on static images, but they cannot handle the interaction timing between instruments and tissues in dynamic surgical videos, resulting in high navigation latency and high misjudgment rates. In addition, the lack of ophthalmic surgery training resources and the absence of operation standardization further exacerbate the steep learning curve problem for novice doctors, and there is an urgent clinical need for real-time dynamic intraoperative decision support tools.

[0003] The current application of AI algorithms in ophthalmic surgery has significant limitations. Firstly, at the data level, there is a lack of multi-center dynamic video datasets with standardized annotations. Existing research relies on manually intercepting key frames, which are difficult to cover intraoperative sudden scenarios (such as bleeding, instrument occlusion). For example, the recognition of TM (the goniotomy step in the operation) under a gonioscope is limited to static gonioscope photos. Secondly, at the technical level, traditional models only support single tasks, such as step recognition or instrument tracking, and rely on preset time thresholds to divide stages, unable to adapt to intraoperative operation fluctuations. Moreover, the clinical adaptability of existing systems is insufficient. Most solutions rely on high-performance computing devices and cannot achieve low-latency inference on conventional operating room hardware (such as devices with 4GB video memory), and there is a lack of in-depth coordination with the microscope optical system, such as insufficient annotation overlay accuracy and failure to optimize detection using the focusing signal.

[0004] Therefore, there is an urgent need to develop a dynamic navigation system for glaucoma surgery that takes into account real-time performance, multi-task accuracy, and hardware lightweight to improve the safety and standardization level of glaucoma surgery.

[0005] Application Content

[0006] This application provides a real-time navigation method for glaucoma surgery based on deep learning, which solves the problem of multi-task real-time collaborative navigation in dynamic surgical scenarios and provides low-latency and high-precision intraoperative decision support for medical staff.

[0007] To achieve the above object, the present application adopts the following technical solutions:

[0008] In a first aspect, the present application provides a real-time navigation method for glaucoma surgery based on deep learning, including:

[0009] Real-time obtain a continuous video stream through the interface of the operating microscope to generate data to be processed;

[0010] Input the data to be processed into a pre-trained multi-task deep learning model, and the multi-task deep learning model includes a step recognition model and a target tracking model;

[0011] The step recognition model is constructed based on the Transformer architecture, and is used to extract the spatio-temporal features of consecutive frames in the data to be processed, and output the corresponding surgical stage of the current frame according to the extracted features and the predefined surgical step definition;

[0012] The target tracking model is constructed based on the real-time semantic segmentation network of PIDNet, and is used to recognize and track instruments and anatomical structures based on single-frame images in the data to be processed, and output target recognition results;

[0013] According to the surgical stage corresponding to the current frame and the target recognition result, trigger a predefined voice prompt;

[0014] Overlay the information of the instrument and the anatomical structure in the video frames of the continuous video stream to obtain a navigation video stream;

[0015] Output the navigation video stream to a monitor in real time.

[0016] In a preferred example of the present application, it can be further set that when pre-training the multi-task deep learning model, it includes:

[0017] Obtain the historical continuous video stream of the operating microscope;

[0018] According to the predefined surgical stage, label the start frame and end frame of each surgical stage in the historical continuous video stream to form a time series label;

[0019] Extract frame images from the historical continuous video stream at a frame interval of 50:1;

[0020] Detect whether there are anatomical structures and instruments in each frame image, and if so, determine it as a valid frame;

[0021] Label the anatomical structures and instruments in the valid frames;

[0022] Input the labeled historical continuous video stream into the multi-task deep learning model.

[0023] In a preferred example of the present application, it can be further set that the target tracking model and the step recognition model process the data to be processed in parallel.

[0024] In a preferred example of the present application, it can be further set that the annotation of the anatomical structure and the instrument in the valid frame includes:

[0025] Annotating the anatomical structure and the instrument in the valid frame based on the annotation of the Rosetta platform;

[0026] When annotating, use polygon labels to annotate the anatomical structure.

[0027] In a preferred example of the present application, it can be further set that when pre-training the multi-task deep learning model, it further includes:

[0028] Use the AdamW optimizer to optimize the weight learning of the step recognition model;

[0029] Use the SGD optimizer to optimize the weight learning of the target tracking model.

[0030] In a preferred example of the present application, it can be further set that it further includes:

[0031] Receive a voice command and turn on or off the voice prompt according to the voice command.

[0032] In a second aspect, the present application provides a real-time navigation device for glaucoma surgery based on deep learning, and the device includes:

[0033] A data acquisition module, configured to obtain a continuous video stream in real time through the interface of a surgical microscope and generate data to be processed;

[0034] A step recognition module, configured to input the data to be processed into a pre-trained multi-task deep learning model, and the multi-task deep learning model includes a step recognition model and a target tracking model; the step recognition model is constructed based on the Transformer architecture and is used to extract the spatio-temporal features of consecutive frames in the data to be processed, and output the surgical stage corresponding to the current frame according to the extracted features and the predefined surgical step definition;

[0035] A target tracking module, configured to the target tracking model is constructed based on the real-time semantic segmentation network of PIDNet and is used to identify and track instruments and anatomical structures based on a single-frame image in the data to be processed and output a target recognition result;

[0036] A voice prompt module, configured to trigger a predefined voice prompt according to the surgical stage corresponding to the current frame and the target recognition result;

[0037] A display module, configured to superimpose information of the instrument and the anatomical structure on video frames of the continuous video stream to obtain a navigation video stream; and output the navigation video stream to a monitor in real time.

[0038] In a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the real-time navigation method for glaucoma surgery based on deep learning as described in any one of the above are implemented.

[0039] In a fourth aspect, the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the real-time navigation method for glaucoma surgery based on deep learning as described in any one of the above is implemented.

[0040] In a fifth aspect, the present application provides a computer program product, including computer instructions, which implement the steps of the real-time navigation method for glaucoma surgery based on deep learning as described in any one of the above when executed by a processor.

[0041] In summary, compared with the prior art, the beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0042] By constructing the multi-task deep learning model, which integrates temporal analysis and spatial analysis, real-time intraoperative identification of surgical step stages, such as steps like goniotomy and viscoelastic injection, and precise tracking of key anatomical structures and surgical instruments are achieved, solving the problem of inaccurate navigation results caused by only analyzing a single dimension and improving the accuracy of surgical navigation.

[0043] By analyzing the real-time video stream provided by the operating microscope to obtain dynamic information of the surgical process, it is possible to more precisely locate tiny eye structures that are prone to confusion, such as the trabecular meshwork, and trigger voice navigation prompts at specific steps, realizing real-time, efficient, and precise assistance for doctors in localizing eye structures and reducing the risk of complications caused by anatomical variations or operation deviations.

[0044] By training the multi-task deep learning model, the generalization ability of the algorithm in real surgical scenarios is effectively improved, solving the technical problems of traditional glaucoma surgery relying on subjective experience and lacking objective navigation, significantly shortening the doctor's learning curve, and improving the surgical safety and operation standardization level. Description of the Drawings

[0045] Figure 1 It is a flowchart of a real-time navigation method for glaucoma surgery based on deep learning provided by an embodiment of the present application.

[0046] Figure 2This is a model architecture diagram of a real-time navigation method for glaucoma surgery based on deep learning provided by an embodiment of the present application.

[0047] Figure 3 This is a module diagram of a real-time navigation device for glaucoma surgery based on deep learning provided by an embodiment of the present application. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0049] In an embodiment of the present application, a real-time navigation method for glaucoma surgery based on deep learning is provided, which is applied to glaucoma surgery. By continuously identifying fine eye structures such as the trabecular meshwork in the surgical video and triggering a reminder according to the analysis structure, please refer to Figure 1 As shown, the method includes:

[0050] S100: Real-time obtain a continuous video stream through the interface of the surgical microscope to generate data to be processed.

[0051] Specifically, obtain the video data of the AG OPMI Lumera 700 surgical microscope (Carl Zeiss Meditec AG, Jena, Germany) and the AIAVS-ST12 surgical video recording platform (AIIST Co., Ltd., Suzhou, Zhejiang, China) through a wired or wireless data interface. The data is a continuous video stream with a resolution of 1920×1080 and 25 frames per second.

[0052] S200: Input the data to be processed into a pre-trained multi-task deep learning model, and the multi-task deep learning model includes a step recognition model and a target tracking model.

[0053] Specifically, the step recognition model and the multi-task deep learning model can process the data to be processed simultaneously, and the structure of the model is as Figure 2 shown.

[0054] S300: The step recognition model is constructed based on the Transformer architecture, and is used to extract the spatio-temporal features of consecutive frames in the data to be processed, and output the surgical stage corresponding to the current frame according to the extracted features and the predefined surgical step definitions.

[0055] Specifically, the step recognition model is based on the End-to-End Long-form Online Action Detection (E2E-LOAD) network of the Transformer architecture. It has excellent performance and code availability. This network uses a lightweight flow buffer to extract and cache the spatial representation of incoming frames. It integrates two branches to simulate historical and current information. The long-term compression (LC) branch compresses the temporal resolution of earlier and longer features, dynamically leveraging the coarse-scale historical information in long-term memory, while the short-term modeling (SM) branch focuses on short windows of transient frames, carefully modeling the recent context. Finally, the representations of these two branches are fused together by the long-short-term fusion (LSF) module to predict the latest frame. While retaining key fine-scale information, the long-term history is effectively compressed. During model execution, video frames in the data to be processed are sampled at a frequency of 24 frames per second (FPS), and the video is structured into chunks of 6 frames. Decisions occur at the chunk level, allowing surgical steps to be evaluated every 0.25 seconds. Regarding the transformer unit, the number of heads is set to 1, and the hidden units are set to 1024 dimensions. When training this model, to facilitate model weight learning, the AdamW optimizer is used, with an initial learning rate of 0.001, a weight decay rate of 0.05, and a momentum value of 0.9. The training batch size is 8, and it ends after 50 epochs. The size of the input frame is adjusted to 320×256 pixels and randomly cropped to 224×224.

[0056] The object tracking model is based on the real-time semantic segmentation network of PIDNet, which identifies anatomical structures such as TM and iris and instruments such as the KDB knife in a single frame.

[0057] S400: The object tracking model is constructed based on the real-time semantic segmentation network of PIDNet, and is used to identify and track instruments and anatomical structures based on single-frame images in the data to be processed, and output the object recognition result.

[0058] Specifically, each frame of the current video stream is sequentially segmented by PIDNet (Proportional-Integral-Derivative Controller network) to generate a segmentation map of the current frame. The previous frames are stored and organized into a sequence as the input of E2E-LOAD (End-to-End Long-form Online Action Detection) to integrate spatio-temporal information. The spatial embedding is combined with short-term and long-term memories to jointly predict the surgical steps of the frame. Subsequently, the intraoperative surgical guidance tools tailored for different surgical steps will be activated.

[0059] Then, a real-time semantic segmentation network named PIDNet is used to construct the target tracking model. PIDNet is an innovative architecture fusion that combines a Proportional-Integral-Derivative (PID) controller and a Convolutional Neural Network (CNN). It integrates three branches for analyzing detail information, context information, and boundary information respectively. The training process of the target tracking model goes through 150 epochs, and the batch processed contains 16 images. Among them, the SGD optimizer (SGD optimizer) is adopted, with an initial learning rate of 0.001, a weight decay rate of 0.0005, and a momentum value of 0.9. During training, the polynomial decay strategy (polynomial decay strategy) is used to gradually reduce the learning rate, with a decay coefficient of 0.9. During the training process, the size of the input image is adjusted to 1536×1024 pixels, while the model evaluation is carried out at the original resolution of 1920×1080.

[0060] The above model analysis code is based on Python 3.8, the OpenCV library, and the PyTorch framework.

[0061] S500: Trigger a predefined voice prompt according to the surgical stage corresponding to the current frame and the target recognition result.

[0062] Specifically, after obtaining the output results of the classification of surgical instruments, anatomical structures, and surgical steps of the multi-task deep learning model, it will trigger the activation of a surgical guidance tool that is consistent with the ongoing surgical step. The intraoperative guidance tool continuously and automatically monitors the positioning of key surgical instruments in relevant steps in the background. According to the target recognition results and surgical stages, audio prompts are used to activate tools tailored for each specific step and issue prompts through real-time audio cues. The purpose of this audio prompt is to notify the operator without interrupting the microscopic observation.

[0063] Among them, a specific Critical Operating Area (COA) is defined for each step, representing the operating area near the key anatomical structure that is most likely to affect the surgical prognosis. When the surgical instrument area (such as a keratome) intersects with the critical operating area (such as the limbus), the critical operating area will serve as a warning range for intraoperative navigation and emit a sound prompt. The intersection range is determined by the common edge of the two areas on the grayscale image through an algorithm. The limbus area is designated as the COA for surgeries such as corneal incision, carbachol injection, OVDs injection, and OVDs irrigation / aspiration. Whenever the relevant instrument area (such as a lacrimal duct cannula, toothed forceps, etc.) intersects with the COA of the corneal edge area, the intraoperative guidance tool will trigger an external speaker to emit a low-frequency (once per second) "beep" sound.

[0064] In addition, the minimum external connecting rectangle of the predicted TM area is calculated, and this rectangular area is marked as the COA in specific steps of gonioscopy observation and GT (highlighted in light red on the user interface). Usually, this COA area covers key anatomical structures such as scleral spur, ciliary body band, and the attachment point of the iris root. When a surgical instrument is detected in the field of view of the gonioscopy prism, a slow-frequency (once per second) "beep" chord will be activated. When the instrument enters the COA, a faster-frequency (twice per second) chord sound "beep" will indicate the proximity of the target. The "ding" chord sound will stop, and if the instrument moves out of the gonioscopy prism or the instrument is not visible in the field of view, a "ding" chord sound will be prompted.

[0065] S600: Superimpose the information of the instrument and the anatomical structure on the video frames of the continuous video stream to obtain a navigation video stream; output the navigation video stream to the monitor in real time.

[0066] Specifically, the processed image information will be superimposed on the output of the original video frame and displayed on the user interface of a third-party monitor to enhance visualization and educational guidance for novices. The surgical guidance function will not be activated during the idle phase or when no instrument is detected in the field of view.

[0067] In this embodiment, by constructing the multi-task deep learning model, the fusion of temporal analysis and spatial analysis is realized, and the intraoperative real-time identification of surgical step stages, such as steps like goniosynechialysis and viscoelastic injection, as well as the precise tracking of key anatomical structures and surgical instruments, are achieved. The problem of inaccurate navigation results caused by only analyzing a single dimension is solved, and the accuracy of surgical navigation is improved. By combining the real-time video stream provided by the operating microscope for analysis, voice navigation prompts are triggered at specific steps to assist the doctor in accurately positioning the tiny and easily confused eye structures, reducing the risk of complications caused by anatomical variations or operation deviations. By training the multi-task deep learning model, the generalization ability of the algorithm in the real surgical scenario is effectively improved, the technical problems of traditional glaucoma surgery relying on subjective experience and lacking objective navigation are solved, the doctor's learning curve is significantly shortened, and the surgical safety and operation standardization level are improved.

[0068] In some embodiments, when pre-training the multi-task deep learning model, it includes:

[0069] Obtain the historical continuous video stream of the operating microscope;

[0070] According to the predefined surgical stages, label the start frame and end frame of each surgical stage in the historical continuous video stream to form time series labels;

[0071] Extract frame images from the historical continuous video stream at a frame interval of 50:1;

[0072] Detect whether there are anatomical structures and instruments in each frame image. If so, determine it as a valid frame;

[0073] Label the anatomical structures and instruments in the valid frames;

[0074] Input the labeled historical continuous video stream into the multi-task deep learning model.

[0075] Furthermore, the labeling of the anatomical structures and instruments in the valid frames includes:

[0076] Based on the Rosetta platform annotation, label the anatomical structures and instruments in the valid frames;

[0077] When labeling, use polygon labels to label the anatomical structures.

[0078] In specific implementation, first, the historical continuous video stream is obtained, and then each frame in the video stream is annotated to determine the time periods for annotating the start and end frames of each step. For each segment of video in Task 1, the annotator watches the surgical video in real time and fills in a comma-separated file using a drop-down menu containing predefined operations, and finally summarizes the start and end frame sequences of the steps. The GT is divided into eight recognizable steps: corneal incision, carbamylcholine injection, injection of ophthalmic viscoelastic devices (OVD), gonioscopic observation, GT, OVD irrigation / aspiration, wound closure, and the intervening idle phase. Each surgical step is precisely defined to ensure accurate annotation of the start and end times. Among them, since goniosynechialysis must be performed under gonioscopic observation, the step of "GT" occurs within the time span of the step of "direct gonioscopic observation", and the other video steps are independently distributed without the possibility of overlap or intersection. According to the time continuity of each segment of video, combined with the detailed analysis of the start and end frames of each step, each frame is clearly classified into one of the above eight steps.

[0079] Then, the surgical instruments and key ocular anatomical structures that may be used in each surgical step are listed, and a comprehensive annotation method is developed, aiming to accurately track and identify surgical instruments and important anatomical structures in order to provide real-time guidance during the operation. Considering that only a relatively fixed number of frames during the operation contain clinically significant instruments or structures, after randomly selecting unedited complete videos, the research further compiled a downsampling script using Python to extract frames from the unprocessed complete videos at a ratio of 50:1 for further fine annotation. According to the previous observation, in the images captured during the operation, only a limited number depict clinically relevant instruments or structures, and many images are considered unnecessary for annotation due to lack of substantial content. Therefore, this process requires classifying video frames into two categories: "valid" and "invalid", and only the frames classified as "valid" can enter the subsequent detailed annotation stage. Among them, the key to classification lies in the presence of important instruments or key anatomical landmarks. In the following situations, the following frames lacking substantial content will be regarded as "invalid": blurring caused by improper focusing, obstruction by key anatomical structures or surgical instruments, periodic corneal instillation obstructing the field of view, and the idle time spent by the surgeon adjusting the microscope or replacing instruments.

[0080] The data annotation of the target tracking model is carried out on the Rosetta platform (Stardust Inc, San Francisco, USA). The image sequence includes downsampled frames extracted from randomly selected videos, which are uploaded to the Rosetta platform and annotated frame by frame for valid / invalid classification. For each valid frame in the training data of the target tracking model, all relevant anatomical structures that may affect the surgical prognosis within each valid box are finely annotated with polygon labels and undergo multiple rounds of careful proofreading. These structures include but are not limited to the corneal margin, iris (observed under gonioscopy), TM (observed under gonioscopy), and hemorrhage (observed under gonioscopy). In addition, various commonly used ophthalmic instruments and tools involved in incising the TM during the surgical procedure will also be annotated, such as the Kahook Dual Blade (KDB, New World Medical, Rancho Cucamonga, CA), Tanito microhook (TMH, Inami & Co, Ltd, Tokyo, Japan), and a self-made 25G bent needle. During the proofreading review process, if the polygon annotation is not fitting, inappropriate, or the label is incorrect, the annotation result will be returned for modification and proofread again until it conforms to the annotation consensus.

[0081] In this embodiment, by combining the glaucoma surgical procedure and characteristics to annotate and screen the training dataset, the accuracy of the multi-task deep learning model is improved.

[0082] In some embodiments, when pre-training the multi-task deep learning model, it further includes:

[0083] Using the AdamW optimizer to optimize the weight learning of the step recognition model;

[0084] Using the SGD optimizer to optimize the weight learning of the target tracking model.

[0085] In this embodiment, the accuracy and generalization ability of the multi-task deep learning model are improved.

[0086] In some embodiments, it further includes:

[0087] Receiving a voice command to turn on or off the voice prompt according to the voice command.

[0088] During specific implementation, through the existing intelligent voice recognition function, it interacts with medical staff in language to control the turning on and off of the voice prompt.

[0089] In this embodiment, doctors can be reminded when navigation is needed and muted when it is not, avoiding unnecessary distraction of doctors' attention during surgery and improving the user experience.

[0090] This application also provides a real-time navigation device for glaucoma surgery based on deep learning. Please refer to Figure 3 as shown, the device includes:

[0091] A data acquisition module 100, configured to obtain a continuous video stream in real time through the interface of a surgical microscope and generate data to be processed;

[0092] A step recognition module 200, configured to input the data to be processed into a pre-trained multi-task deep learning model, and the multi-task deep learning model includes a step recognition model and a target tracking model; the step recognition model is constructed based on the Transformer architecture, and is configured to extract the spatio-temporal features of consecutive frames in the data to be processed, and output the corresponding surgical stage of the current frame according to the extracted features and the predefined surgical step definition;

[0093] A target tracking module 300, configured to construct the target tracking model based on the real-time semantic segmentation network of PIDNet, and is configured to identify and track instruments and anatomical structures based on a single-frame image in the data to be processed, and output a target recognition result;

[0094] A voice prompt module 400, configured to trigger a predefined voice prompt according to the surgical stage corresponding to the current frame and the target recognition result;

[0095] A display module 500, configured to superimpose the information of the instrument and the anatomical structure on the video frame of the continuous video stream to obtain a navigation video stream; and output the navigation video stream to a monitor in real time.

[0096] The function implementation of each module in the above real-time navigation device for glaucoma surgery based on deep learning corresponds to the steps in the above embodiments of the real-time navigation method for glaucoma surgery based on deep learning, and its functions and implementation processes will not be elaborated here one by one.

[0097] This application also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the real-time navigation method for glaucoma surgery based on deep learning as described in any of the above embodiments.

[0098] The present application also provides a computer-readable storage medium, on which a program is stored. Herein, the computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of a real-time navigation method for glaucoma surgery based on deep learning in the foregoing text, which will not be elaborated herein.

[0099] The application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the real-time navigation method for glaucoma surgery based on deep learning as described in any of the foregoing embodiments.

[0100] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the foregoing embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the foregoing methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memories (ROMs), programmable ROMs (PROMs), electrically programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), or flash memories. Volatile memories may include random access memories (RAMs) or external cache memories. By way of illustration and not limitation, RAMs are available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0101] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A real-time navigation method for glaucoma surgery based on deep learning, characterized in that Including: Real-time acquisition of a continuous video stream through the interface of a surgical microscope to generate data to be processed; Inputting the data to be processed into a pre-trained multi-task deep learning model, where the multi-task deep learning model includes a step recognition model and an object tracking model; The step recognition model is constructed based on the Transformer architecture, used to extract the spatio-temporal features of consecutive frames in the data to be processed, and according to the extracted features and the predefined surgical step definitions, output the surgical stage corresponding to the current frame; The object tracking model is constructed based on the real-time semantic segmentation network of PIDNet, used to identify and track instruments and anatomical structures based on single-frame images in the data to be processed, and output the object recognition result; Trigger a predefined voice prompt according to the surgical stage corresponding to the current frame and the object recognition result; Overlay the information of the instrument and the anatomical structure on the video frames of the continuous video stream to obtain a navigation video stream; Output the navigation video stream to a monitor in real time.

2. The real-time navigation method for glaucoma surgery based on deep learning according to claim 1, wherein When pre-training the multi-task deep learning model, it includes: Obtaining the historical continuous video stream of the surgical microscope; According to the predefined surgical stages, mark the start frames and end frames of each surgical stage in the historical continuous video stream to form time series labels; Extract frame images from the historical continuous video stream at a frame interval of 50:1; Detect whether there are anatomical structures and instruments in each frame image, and if so, determine it as a valid frame; Label the anatomical structures and instruments in the valid frames; Input the labeled historical continuous video stream into the multi-task deep learning model.

3. The real-time navigation method for glaucoma surgery based on deep learning according to claim 2, wherein, The object tracking model and the step recognition model process the data to be processed in parallel.

4. The real-time navigation method for glaucoma surgery based on deep learning according to claim 3, wherein, The labeling of the anatomical structures and instruments in the valid frames includes: Label the anatomical structures and instruments in the valid frames based on the labeling of the Rosetta platform; When labeling, use polygon labels to label the anatomical structures.

5. The real-time navigation method for glaucoma surgery based on deep learning according to claim 2, wherein When pre-training the multi-task deep learning model, it also includes: Use the AdamW optimizer to optimize the weight learning of the step recognition model; Use the SGD optimizer to optimize the weight learning of the object tracking model.

6. The real-time navigation method for glaucoma surgery based on deep learning according to claim 2, wherein It also includes: Receive a voice command and turn on or off the voice prompt according to the voice command.

7. A real-time navigation device for glaucoma surgery based on deep learning, characterized in that, Including: A data acquisition module for real-time acquisition of a continuous video stream through the interface of a surgical microscope to generate data to be processed; A step recognition module for inputting the data to be processed into a pre-trained multi-task deep learning model, where the multi-task deep learning model includes a step recognition model and an object tracking model; the step recognition model is constructed based on the Transformer architecture, used to extract the spatio-temporal features of consecutive frames in the data to be processed, and according to the extracted features and the predefined surgical step definitions, output the surgical stage corresponding to the current frame; An object tracking module for the object tracking model is constructed based on the real-time semantic segmentation network of PIDNet, used to identify and track instruments and anatomical structures based on single-frame images in the data to be processed, and output the object recognition result; A voice prompt module, configured to trigger a predefined voice prompt according to the surgical stage corresponding to the current frame and the target recognition result; A display module, configured to superimpose information of the instrument and the anatomical structure on a video frame of the continuous video stream to obtain a navigation video stream; and output the navigation video stream to a monitor in real time.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the real-time navigation method for glaucoma surgery based on deep learning according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, A program is stored on the computer-readable storage medium, and when the program is executed by a processor, the real-time navigation method for glaucoma surgery based on deep learning according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising computer instructions, characterized in that, When the computer instruction is executed by a processor, the steps of the real-time navigation method for glaucoma surgery based on deep learning according to claims 1 to 6 are implemented.