Image following type surgery annotation system
Patent Information
- Application Number
- PCT/JP2025/007040
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-02
AI Technical Summary
Existing image-tracking surgery systems fail to effectively display annotations in accordance with camera movement, lacking mechanisms to adjust and maintain the position and scale of annotations relative to the video as the camera moves.
An annotation display device that includes a video playback unit and an annotation display unit, utilizing optical flow tracking to automatically move and scale annotations based on camera movement, with mechanisms to pause video playback during annotation input to ensure precise placement and scaling.
Enables annotations to follow camera movement seamlessly, maintaining their position and scale relative to the video, enhancing the review experience by reducing distractions and improving annotation clarity.
Smart Images

Figure JP2025007040_02102025_PF_FP_ABST
Abstract
Description
Image-tracking surgical annotation system
[0001] RELATED APPLICATIONS This application claims priority to Japanese Patent Application No. 2024-035539, entitled "Image-Tracking Surgery Annotation System," filed on March 8, 2024, the disclosure of which is incorporated herein by reference in its entirety. The disclosure of this application relates to an image-tracking surgery annotation system.
[0002] International Publication No. WO 2018 / 163600 (Patent Document 1) is a document disclosing background technology in this technical field. This publication states, "Fig. 6 is a diagram showing another example of a display screen of a terminal on which surgical data is displayed. In the example shown in Fig. 6, the display screen 421 is divided into six areas. The largest area 422 displays the user's endoscopic video, similar to the example shown in Fig. 5. However, unlike the example shown in Fig. 5, the example shown in Fig. 6 differs from the example shown in Fig. 5 in that area 422 displays the endoscopic video on which annotation data is superimposed. By displaying such superimposed annotation data, it is expected that the review effect will be enhanced when the doctor receiving the training later reviews the endoscopic video of his or her own surgery" (see paragraph
[0128] ).
[0003] International Publication No. 2018 / 163600
[0004] The above-mentioned Patent Document 1 describes displaying annotation data superimposed on an endoscopic image. However, there is no mention of moving the displayed annotation data in accordance with camera movement. The disclosure of the present application has been made in consideration of such circumstances, and provides a mechanism for displaying annotations in accordance with camera movement.
[0005] In order to solve the above problems, for example, the configuration described in the claims is adopted. The present application includes a plurality of means for solving the above problems, and one example thereof is an annotation display device including a video playback unit that plays back a video captured by an endoscopic camera, and an annotation display unit that performs at least one of movement, enlargement, and reduction of an annotation superimposed on the video and input by a user based on the movement of the endoscopic camera.
[0006] According to the disclosure of the present application, it is possible to provide a mechanism for displaying annotations in accordance with camera movement. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments.
[0007] FIG. 1 shows an example of the configuration of a surgery support system 100. FIG. 2 shows an example of the configuration of an annotation display device 107. FIG. 3 shows an example of a video playback flow. FIG. 4 shows an example of a tracking processing flow. FIG. 5 shows an example of an annotation display flow. FIG. 6 shows an example of a menu screen. FIG. 7 shows an example of an annotation screen. FIG. 8 shows another example of an annotation screen. FIG. 9 shows another example of an annotation screen. FIG. 10 shows another example of an annotation display flow. FIG. 11 shows an example of an image before and after image processing. FIG. 12 shows an example of a tracking processing flow.
[0008] 1. Examples Hereinafter, examples of the present invention will be described with reference to the drawings. 1-1. Overview The annotation display device according to this example has the following features: (1) It is possible to use a pen or mouse to overlay and depict lines and characters (examples of annotations) on endoscopic surgery images.
[0009] (2) The rendered annotations can be automatically moved and scaled to follow the movement of the video. This maintains the size and positional relationship of the rendered image to the video. (3) Optical flow tracking is used to make the annotations follow the movement of the video, but sampling the entire screen reduces the influence of surgical instruments.
[0010] (4) When drawing annotations, if the system continues to track the surgical video, it will be impossible to draw lines or letters as intended. Therefore, the system constantly checks whether an annotation is currently being drawn, and pauses the playback of the surgical video while drawing. At the same time, tracking is seemingly stopped, allowing the user to draw lines and letters as intended.
[0011] The tracking process itself continues in the background, and once you finish drawing, the image switches to the current time. By continuing the tracking process in the background, the annotation is moved to the correct position and the zoomed-in / zoomed display is maintained.
[0012] (5) The color, thickness, and transparency of the lines drawn can be changed. (6) Annotations can be made remotely on images sent in real time from the operating room.
[0013] 1 shows an example of the configuration of a surgery assistance system 100 according to this embodiment. The surgery assistance system 100 shown in the figure includes an endoscopic camera 101, an encoder 102, a decoder 103, and a monitor 104 as equipment on the side where surgery is performed. The surgery assistance system 100 also includes a decoder 105, a capture device 106, an annotation display device 107, and an encoder 108 as equipment on the side where annotations are added to the surgery video.
[0014] An endoscopic camera 101 on the surgical side outputs video data representing surgical video to an encoder 102. The encoder 102 encodes the input video data and transmits it to a decoder 105 on the annotation side via a wired or wireless network 109.
[0015] The decoder 105 decodes the received encoded data and outputs the generated video data to the capture device 106. The capture device 106 outputs the input video data to the annotation display device 107. The annotation display device 107 displays the input video data on a display. The annotation display device 107 also displays annotations superimposed on the surgical video in response to pen input or mouse operation. The annotation display device 107 outputs the data of the surgical video on which the annotations are superimposed to the encoder 108. The encoder 108 encodes the input video data and transmits it to the decoder 103 on the surgery side via a network 109.
[0016] The decoder 103 decodes the received encoded data and outputs the generated video data to the monitor 104. The monitor 104 displays the input video data. The displayed video data is a surgical video on which annotations are superimposed.
[0017] Next, the annotation display device 107 will be described with reference to Fig. 2. Fig. 2 shows an example of the configuration of the annotation display device 107.
[0018] The annotation display device 107 may be, for example, a mobile device such as a smartphone, tablet, mobile phone, or personal digital assistant (PDA), or a wearable device such as glasses, a wristwatch, or clothing. The annotation display device 107 may also be a stationary or mobile computer, or a server located on the cloud or a network. The annotation display device 107 may also function as a VR (Virtual Reality) device, an AR (Augmented Reality) device, or an MR (Mixed Reality) device. Alternatively, the annotation display device 107 may be a combination of multiple of these devices. For example, a combination of one smartphone and one wearable device may logically function as a single device. Other information processing devices may also be used.
[0019] The annotation display device 107 includes a processor 203 that executes an operating system, applications, programs, etc., a main storage device 201 such as RAM (Random Access Memory), an auxiliary storage device 202 such as an IC card, hard disk drive, SSD (Solid State Drive), or flash memory, a communication control unit 206 such as a network card, wireless communication module, or mobile communication module, an input device 204 such as a touch panel, keyboard, mouse, pen input, voice input, or input based on motion detection captured by a camera unit, and an output device 205 such as a monitor or display. Note that the output device 205 may also be a device or terminal that transmits information to be output to an external monitor, display, printer, or other device.
[0020] The main memory device 201 stores various programs, applications, etc. (modules), and the processor 203 executes these programs and applications to realize the various functional elements of the overall system. These modules may be implemented in hardware, such as by integration. Each module may be an independent program or application, or may be implemented as a subprogram or function within a single integrated program or application.
[0021] In this specification, each module is described as a subject that performs processing, but in reality, the processor 203 that processes various programs, applications, etc. (modules) executes the processing.
[0022] Various databases (DBs) are stored in the auxiliary storage device 202. A "database" is a functional element (storage unit) that stores a set of data so that it can accommodate any data manipulation (e.g., extraction, addition, deletion, overwriting, etc.) from the processor 203 or an external computer. The method for implementing the database is not limited, and may be, for example, a database management system, spreadsheet software, or a text file such as XML or JSON.
[0023] The main memory device 201 of the annotation display device 107 stores programs and applications such as a UI module 210, a frame acquisition module 211, a video playback module 212, a tracking module 213, and an annotation display module 214. The processor 203 executes these programs and applications to realize the various functional elements of the annotation display device 107. Each module will be described below.
[0024] The UI module 210 receives instructions from the user and displays an application startup screen and an annotation screen, which will be described later.
[0025] When a user instructs playback of a moving image, the frame acquisition module 211 sequentially acquires frames that make up the moving image and provides them to the moving image playback module 212 and tracking module 213 .
[0026] When a user instructs the video playback module 212 to play a specified video, the video playback module 212 plays the specified video. The videos to be played include videos captured by the endoscopic camera 101. The videos to be played also include streaming videos and videos stored in the auxiliary storage device 202.
[0027] The video playback module 212 stops the video playback while the user is inputting annotations, allowing the user to draw annotations without being distracted by the video being played.
[0028] After the annotation is input, the video playback module 212 starts playing the video by skipping images for the duration of the annotation input, thereby avoiding display delays when the video is displayed in real time.
[0029] In this specification, annotations refer to characters or graphics added to an image. These annotations are input to indicate display elements in the image or to provide supplemental explanations for the display elements in the image.
[0030] When a user instructs playback of a video, the tracking module 213 calculates the movement of the camera that captured the video by measuring the optical flow of the video.
[0031] More specifically, this module calculates camera movement in the following steps: (1) Of the previous and latest frames that make up the video, multiple points (in other words, dots) are placed in the previous frame; (2) Optical flow measurement is performed to estimate where the multiple points have moved in the latest frame; (3) The median of the movement amounts of each point is calculated and obtained as the up / down / left / right movement amounts of the camera; and (4) The median of the change in distance between each point is calculated and obtained as the forward / backward movement amounts of the camera.
[0032] In step (1) of this procedure, the module places multiple points over the entire previous frame, i.e., the module samples the entire frame, which allows the module to reduce the influence of surgical instrument movement on the calculations compared to when sampling over a portion of the frame.
[0033] In addition, when performing processes (3) and (4), the module performs forward-backward error checking. Specifically, the module performs the following processes: (5) For each point whose position in the latest frame has been estimated, optical flow measurement is performed to estimate the position in the previous frame. (6) Points whose estimated position differs from their pre-positioned position in the previous frame by a predetermined value or more are excluded from the calculation of the movement amount. This improves the accuracy of estimating camera movement.
[0034] In addition, when performing steps (3) and (4), this module also performs normalized cross correction (NCC). Specifically, this module excludes points where the similarity (similarity of pixel values) between the surrounding pixels in the previous frame and the surrounding pixels in the latest frame is below a predetermined value from the movement amount calculation. This improves the accuracy of camera movement estimation.
[0035] When the user inputs an annotation to be superimposed on the video being played, the annotation display module 214 performs at least one of moving, enlarging, and reducing the annotation based on the movement of the camera that captured the video.
[0036] Furthermore, this module accumulates values calculated sequentially by the tracking module 213 while the annotation is being input, and corrects the annotation based on the accumulated value after the annotation is input. This allows the annotation to be displayed at the correct position and scale in the latest frame, even if the displayed image jumps to the latest frame from the user's viewpoint after the annotation is input.
[0037] Next, we will explain the auxiliary storage device 202. The auxiliary storage device 202 stores video data 220. The stored video data 220 includes data of videos captured by a camera (for example, video data of an endoscopic surgery).
[0038] 1-3. Operation Next, the operation of the annotation display device 107 will be described. First, the operation at the time of application startup will be described. When the user starts an application, the UI module 210 of the annotation display device 107 displays a startup menu screen. Figure 6 shows an example of this menu screen.
[0039] The menu screen 600 shown in the figure has a pull-down menu 601 for specifying the method of obtaining the video data, an input field 602 for specifying the URL or file name of the video data, an input field 603 for specifying the resolution of the video, a microphone setting area 604 for setting the microphone, a speaker setting area 605 for setting the speaker, and a play button 612.
[0040] Of these, the pull-down menu 601 allows the selection of the following data acquisition methods: (1) Use the capture device 106; (2) Specify video data stored in the auxiliary storage device 202; (3) Specify the URL of video data distributed via the network 109. One of these data acquisition methods is selected. Note that, for example, RTSP (Real Time Streaming Protocol) is used to distribute video via the network 109.
[0041] In the input field 602, the URL of the video data distributed via the network 109 or the file name of the video is input.
[0042] In the input field 603, the resolution of the video to be acquired from the capture device 106 is specified.
[0043] The microphone setting area 604 includes a pull-down menu 606 for selecting a device to be used for the microphone, a pull-down menu 607 for selecting the output destination of the microphone sound, and an indicator 608 for adjusting the microphone volume.
[0044] The speaker setting area 605 includes a pull-down menu 609 for selecting a device to be used for the speaker, a pull-down menu 610 for selecting a sound source, and an indicator 611 for adjusting the speaker volume.
[0045] The play button 612 is a button for instructing playback of a video. When this button is selected, playback of the video begins based on settings specified by the user. When the video is being played, the UI module 210 displays an annotation screen. Figure 7 shows an example of this annotation screen.
[0046] The annotation screen 700 shown in the same figure has a pull-down menu 701 for selecting either pen mode or eraser mode, a button 702 for specifying the pen or eraser size, a button 703 for specifying the line color, a button 704 for specifying the line opacity, a button 705 for specifying the time until resuming from a paused state after the drawing operation is completed, a button 706 for erasing all drawn annotations at once, a button 707 for stopping video playback and returning to the menu screen, a seek bar 708 for specifying the playback position, and a video playback area 709.
[0047] Of these, the pull-down menu 701 allows the user to select either pen mode or eraser mode. By selecting pen mode, the user can draw annotations superimposed on the video being played. By selecting eraser mode, the user can erase part or all of the drawn annotations.
[0048] When a pen tablet is used as the input device 204, the mode may be automatically switched, such as pen mode when the pen tip is used to draw and eraser mode when the pen tip is used to draw.
[0049] Button 702 is a button for specifying the size of the pen or eraser. In this example screen, four sizes can be specified: "5", "17", "30", and "40", but this number of sizes may be changed as desired.
[0050] A button 703 is a button for specifying the color of the line. In this example screen, seven colors can be specified, but this number may be changed arbitrarily.
[0051] A button 704 is a button for specifying the opacity of the line. In this example screen, two opacity levels can be specified: "0.5" and "1", but the number of levels may be changed as desired.
[0052] Button 705 is a button for specifying the time from the paused state until resuming after the drawing operation is completed. After the annotation is input, the video playback module 212 starts playing the video after the time specified by this button has elapsed. Note that in this example screen, three types of time (seconds) can be specified: "0.25", "0.50", and "1.00", but the number of types may be changed as desired.
[0053] A button 706 is a button for erasing all drawn lines at once. When this button is selected, the annotation display module 214 erases all of the annotations currently being displayed.
[0054] A button 707 is used to stop the video playback and return to the menu screen. When this button is selected, the UI module 210 displays a menu screen such as the one shown in FIG.
[0055] The seek bar 708 is a button for specifying a playback position of the video. The video playback module 212 plays the video from the position specified by the seek bar 708.
[0056] The video currently being played is displayed in the video playback area 709. In this example screen, video of endoscopic surgery is displayed. Surgical instruments 710 and 711 are shown in the displayed video. An annotation 712 is also superimposed on the displayed video. The annotation 712 is a substantially circular line.
[0057] Next, the video playback process will be described with reference to Fig. 3. Fig. 3 shows an example of a video playback flow. Flow 300 shown in Fig. 3 is executed when a user instructs playback of a video.
[0058] When a video is played back, the frame acquisition module 211 sequentially acquires frames that make up the video and provides them to the video playback module 212 (not shown). The processing of this frame acquisition module 211 stops when the video playback stops, but continues while annotations are being input.
[0059] The video playback module 212 acquires the latest frame from the frame acquisition module 211 (step 301). After acquiring the frame, the video playback module 212 updates the display image based on the acquired frame (step 302).
[0060] After updating the image, the module determines whether or not annotation input has been detected (step 303). If the result of this determination is that annotation input has not been detected (NO in step 303), the module skips step 304 and proceeds to step 305. On the other hand, if the result of this determination is that annotation input has been detected (YES in step 303), the module then determines whether or not completion of annotation input has been detected (step 304). If the result of this determination is that completion of annotation input has not been detected (NO in step 304), the module executes step 304 again and waits until annotation input is complete. On the other hand, if the result of this determination is that completion of annotation input has been detected (YES in step 304), the module proceeds to step 305.
[0061] In step 305, the module determines whether or not an instruction to end video playback has been issued. If the result of this determination is that an instruction to end video playback has been issued (YES in step 305), the module ends this flow. On the other hand, if the result of this determination is that an instruction to end video playback has not been issued (NO in step 305), the module returns to step 301 and acquires the latest frame again. This concludes the description of video playback flow 300.
[0062] In the video playback flow 300 described above, updating of the displayed image stops while an annotation is being input. That is, playback of the video stops from the user's viewpoint. This allows the user to draw annotations without being distracted by the video being played.
[0063] Next, the tracking process will be described with reference to Fig. 4. Fig. 4 shows an example of a tracking process flow. Flow 400 shown in the figure is a process for tracking the movement of a camera that captured a video being played back. This flow 400 is executed when a user instructs playback of the video.
[0064] When a video is being played, the frame acquisition module 211 sequentially acquires frames that make up the video and provides them to the tracking module 213 (not shown). The processing of this frame acquisition module 211 stops when the video playback stops, but continues while annotations are being input.
[0065] The tracking module 213 acquires the latest frame from the frame acquisition module 211 (step 401). If the acquired frame is the first frame (YES in step 402), the video playback module 212 executes step 401 again to acquire the next latest frame. On the other hand, if the acquired frame is not the first frame (NO in step 402), the tracking module 213 executes processing to calculate the camera motion.
[0066] Specifically, the module sets the range of the tracking target by enclosing the target in the latest frame and the immediately preceding frame with a rectangle (step 403). In this embodiment, the module sets the entire frame as the tracking target.
[0067] Next, the module places a predetermined number of points at equal intervals within a rectangle in the previous frame and performs optical flow measurement using the Lucas-Kanade algorithm between the previous frame and the latest frame (step 404). Next, the module performs forward-backward error checking (step 405) and NCC (step 406). Based on these calculations, the module measures where the previous point moved in the latest frame. The module obtains the median of the movement of all points as the up / down / left / right movement of the camera, and obtains the median of the change in distance between points as the forward / backward movement of the camera (in other words, the amount of scaling). The module then outputs the obtained movement and scaling to the annotation display module 214 (step 407). The output movement is, for example, represented as a vector (i.e., direction and magnitude).
[0068] A further note about the above forward-backward error check and NCC: If the optical flow measurement in step 404 is performed correctly, the point positions measured when the comparison frames are reversed (from the latest frame to the immediately preceding frame) should be close to the original point positions. In forward-backward error check, the difference between the point positions from the latest frame to the immediately preceding frame and the originally placed point positions is considered to be an error, and point positions calculated in step 404 with an error equal to or greater than the median are excluded as not being tracked correctly.
[0069] The pixels surrounding the point before and after the movement should be similar. NCC calculates the similarity for each point, and excludes points whose similarity is lower than the median value as they have not been correctly tracked.
[0070] After outputting the amount of movement, etc., the tracking module 213 determines whether an instruction to end video playback has been issued (step 408). If the result of this determination is that an instruction to end video playback has been issued (YES in step 408), the module ends this flow. On the other hand, if the result of this determination is that an instruction to end video playback has not been issued (NO in step 408), the module returns to step 401 and acquires the latest frame again. This concludes the explanation of the tracking processing flow 400.
[0071] In the tracking process flow 400 described above, the entire frame is tracked. Additionally, this flow uses medians as the movement and scaling amounts. Therefore, this flow can reduce the influence of the movement of surgical instruments that are partially visible in the camera image when calculating the camera movement.
[0072] Next, the annotation display process will be described with reference to Fig. 5. Fig. 5 shows an example of an annotation display flow. A flow 500 shown in the figure is executed when an annotation input is detected.
[0073] First, when the annotation display module 214 detects the input of an annotation, it starts accumulating each of the movement amount and the scaling amount sequentially output from the tracking module 213 (step 501). Next, the annotation display module 214 determines whether or not the completion of the annotation input has been detected (step 502). If the result of this determination is that the completion of the annotation input has not been detected, the module executes step 502 again and waits until the annotation input is complete. During this time, the video playback module 212 stops the playback of the video from the user's viewpoint. Meanwhile, the tracking module 213 continues the tracking process.
[0074] If the result of the determination in step 502 above indicates that the annotation input has been completed (YES in step 502), the module corrects the input and displayed annotation based on the cumulative values of the movement amount and the scaling amount (step 503). This correction ensures that the annotation is displayed at the position and scale it should be in the latest frame, even if the displayed image jumps to the latest frame from the user's viewpoint. In other words, the movement and scale of the annotation follow the movement of the camera that captured the video being played. As a result, the position and size of the annotation relative to the captured object are maintained.
[0075] Next, the module determines whether the input annotation has been deleted (step 504). If the result of this determination is that the input annotation has been deleted (YES in step 504), the module ends the accumulation of the movement amount and the scaling amount, and initializes the accumulated values (step 505). The module then ends this flow. On the other hand, if the result of the determination in step 504 is that the input annotation has not been deleted (NO in step 504), the module returns to step 503.
[0076] In step 503, the module corrects the annotation being displayed based on the latest accumulated values of the movement and scaling. This correction ensures that the movement and scale of the annotation follow the movement of the camera that captured the video being played. As a result, the position and size of the annotation relative to the captured object are maintained.
[0077] Here, the correction of annotations will be described with reference to Figures 8 and 9. Figures 8 and 9 show other examples of annotation screens.
[0078] 8, the annotation 801 is shown as an annotation. The annotation 801 has been reduced in size compared to the annotation 712 shown in FIG. 7 as the endoscopic camera 101 has moved backward.
[0079] On the other hand, in the screen 900 shown in Fig. 9, the reference numeral 901 indicates an annotation. As the endoscopic camera 101 moves diagonally forward to the right, the annotation 901 is enlarged and moved leftward compared to the annotation 712 shown in Fig. 7.
[0080] After executing step 503, the module again executes step 504. The annotation display flow 500 has been described above.
[0081] The annotation display flow 500 described above allows the movement and scale of the annotation to follow the movement of the camera that captured the video being played. In addition, this flow allows the annotation to be displayed at the position and scale that it should be in the latest frame, even if the displayed image jumps to the latest frame from the user's viewpoint after the annotation is input.
[0082] 2. Modifications The above embodiment may be modified as follows: The following modifications may be combined with each other.
[0083] (1) Application of the Present Invention The above-described embodiment assumes a case in which endoscopic surgery (or endoscopic examination) is remotely supported, but the use case of the present invention is not limited to this. The present invention can also be applied to other cases as long as annotations are drawn on video captured by a camera. Note that the camera assumed in the present invention is a camera whose shooting range can be changed, such as a movable camera or a pan-tilt-zoom camera.
[0084] When the present invention is used in a situation other than endoscopic surgery, the camera used does not necessarily have to be an endoscopic camera.
[0085] (2) Annotation Control In the above embodiment, both movement and scaling are performed in accordance with the camera movement. However, it is not necessary to perform both. Only one of movement and scaling may be performed in accordance with the camera movement.
[0086] (3) Camera Tracking Method In the above embodiment, a different tracking method may be used to track the camera movement. For example, an object tracking algorithm implemented in an extension module of OpenCV (Open Source Computer Vision Library) may be used.
[0087] (4) Tracking Target Within a Frame In the tracking process flow 400 of the above embodiment, the entire frame is the tracking target. This is to reduce the influence of the movement of the surgical instrument, which is partially visible in the camera image, when calculating the camera movement. However, in a situation where the influence of partial display elements is small to begin with, the entire frame does not necessarily need to be the tracking target. For example, part of the frame may be excluded from the tracking target.
[0088] (5) System Configuration In the above embodiment, the encoder 102, decoder 103, and monitor 104 are each separate (see FIG. 1), but they do not necessarily have to be separate. Two or more of these devices may be configured as an integrated unit. For example, a single PC may have the functions of all of these devices. Similarly, two or more of the decoder 105, capture device 106, annotation display device 107, and encoder 108 may be configured as an integrated unit.
[0089] In the above embodiment, since it is assumed that annotation will be performed from a remote location, the endoscopic camera 101 and the monitor 104 are connected to the annotation display device 107 via the network 109. However, the location where annotation is performed is not limited to a remote location. For example, if annotation is performed near an operating room, the endoscopic camera 101 and the monitor 104 may be directly connected to the annotation display device 107.
[0090] (6) Three-Dimensional Annotations In the above embodiment, it is assumed that two-dimensional images are displayed on the annotation display device 107. However, a three-dimensional model may be generated from video data, and the generated three-dimensional model may be displayed stereoscopically to the user. In this case, the user may add three-dimensional annotations to the three-dimensionally displayed three-dimensional model. The tracking module 213 may then calculate the camera movement by tracking the three-dimensional model, and control the movement, scaling, and rotation of the annotation based on the calculated movement.
[0091] (7) Modification of Annotation Display Flow The annotation display flow may be executed as described below. Figure 10 shows another example of the annotation display flow. Flow 1000 shown in the figure is executed when an annotation input is detected.
[0092] First, when the annotation display module 214 detects the input of an annotation, it starts accumulating each of the movement amount and the scaling amount sequentially output from the tracking module 213 (step 1001). Next, the annotation display module 214 determines whether or not the completion of the annotation input has been detected (step 1002). If the result of this determination is that the completion of the annotation input has not been detected, the module executes step 1002 again and waits until the annotation input is complete. During this time, the video playback module 212 stops the playback of the video from the user's viewpoint. Meanwhile, the tracking module 213 continues the tracking process.
[0093] As a result of the determination in step 1002, if the completion of annotation input is detected (YES in step 1002), the module ends the accumulation of the movement amount and the enlargement / reduction amount (step 1003).
[0094] Next, the module corrects the input and displayed annotation based on the cumulative values of the movement and scaling amounts (step 1004). This correction ensures that the annotation is displayed at the correct position and scale for the latest frame, even if the displayed image jumps to the latest frame from the user's viewpoint. In other words, the annotation's movement and scale follow the movement of the camera that captured the video being played. As a result, the annotation's position and size relative to the captured object are maintained.
[0095] Next, the module determines whether the input annotation has been deleted (step 1005). If the result of this determination is that the input annotation has been deleted (YES in step 1005), the module ends this flow. On the other hand, if the result of this determination is that the input annotation has not been deleted (NO in step 1005), the module proceeds to step 1006.
[0096] In step 1006, the module acquires the latest movement amount and scaling amount from the tracking module 213. Then, the module corrects the annotation being displayed based on the acquired movement amount, etc. (step 1007). This correction causes the movement and scale of the annotation to follow the movement of the camera that captured the video being played. As a result, the position and size of the annotation relative to the captured object are maintained.
[0097] After executing step 1007, the module again executes step 1005. The annotation display flow 1000 has been described above.
[0098] (8) Tracking Target in Frame (Part 2) In the above tracking process flow 400, the entire frame is the tracking target. This is to reduce the influence of the movement of the surgical instrument. However, if the influence of the movement of the surgical instrument is small, it is not necessary to track the entire frame. Therefore, if the influence of the movement of the surgical instrument is small, only a part of the frame may be the tracking target, and if the influence of the movement of the surgical instrument is large, the entire frame may be the tracking target. Such a modified example will be described below.
[0099] In this modification, the tracking module 213 duplicates and binarizes the frame acquired from the frame acquisition module 211. In this process, the module converts pixels with saturation equal to or greater than a threshold to white, and pixels with saturation less than the threshold to black. The threshold referenced in this conversion is set in advance so that pixels representing internal organs are converted to white and pixels representing surgical instruments are converted to black. To perform tracking processing, a logical product is performed between the grayscale image of the frame acquired from the frame acquisition module 211 and the binarized image, and pixels representing the internal organs in the grayscale image are retained.
[0100] Generally, internal organs are displayed in reddish-brown or pink, while surgical instruments are displayed in silver. Therefore, the former have a higher saturation than the latter. Therefore, in this modified example, the binarization process of the frame is performed based on saturation.
[0101] 11A and 11B show examples of images before and after image processing. Fig. 11A shows the image before image processing, and Fig. 11B shows the grayscale image after logical product processing. In the image after logical product processing (i.e., Fig. 11B), the internal organs 1101 are mostly white, while the surgical instruments 1102 are mostly black.
[0102] Next, the tracking module 213 sets a tracking area in the processed image. The set tracking area is a partial area to be tracked. This tracking area is a square area of a predetermined size. This tracking area is set in the vicinity of the input annotation. More specifically, this area is set so that the center of the input annotation or its center of gravity overlaps with the center of the area. This makes it possible to track the area in the vicinity of the annotation. FIG. 11B shows a tracking area 1103 as an example of this tracking area. Note that the center of the annotation referred to here is, for example, the center of the circumscribing rectangle of the annotation.
[0103] Next, the tracking module 213 calculates the proportion of black pixels in the set tracking area. The tracking module 213 then determines whether the calculated proportion is equal to or greater than a predetermined value (e.g., 40%). If the result of this determination is that the calculated proportion is not equal to or greater than the predetermined value, the tracking module 213 determines that the influence of the surgical instrument movement is small, and therefore tracks the set tracking area. On the other hand, if the result of this determination is that the calculated proportion is equal to or greater than the predetermined value, the influence of the surgical instrument movement is large, and therefore the tracking module 213 tracks the entire grayscale image.
[0104] After determining the tracking target, the tracking module 213 arranges a predetermined number of points at equal intervals on the tracking target and performs optical flow measurement using the Lucas-Kanade algorithm between the tracking target and subsequent frames. Next, the tracking module 213 performs forward-backward error check and NCC, and outputs the camera movement amount and zoom amount to the annotation display module 214. The optical flow measurement, forward-backward error check, and NCC performed at this time have been described in the above embodiments, so their description will be omitted.
[0105] Next, the tracking process according to this modification will be described with reference to Fig. 12. Fig. 12 shows an example of a tracking process flow according to this modification. Flow 1200 shown in the figure is a process for tracking the movement of a camera that captured a video being played back. Flow 1200 is executed for each annotation after the annotation is input.
[0106] First, the tracking module 213 acquires the latest frame from the frame acquisition module 211 (step 1201). The tracking module 213 then copies the acquired frame and binarizes and grayscales it (step 1202). During binarization, the module converts pixels with saturation equal to or greater than a threshold to white and pixels with saturation less than the threshold to black. The threshold referenced during this conversion is preset so that pixels representing internal organs are converted to white and pixels representing surgical instruments are converted to black.
[0107] Next, if the frame acquired in step 1201 is the first frame (YES in step 1203), the tracking module 213 executes step 1201 again to acquire the next latest frame. On the other hand, if the acquired frame is not the first frame (NO in step 1203), the tracking module 213 sets a tracking region in the previous grayscaled frame (step 1204).
[0108] The tracking area set here is a partial area to be tracked. This tracking area is a square area of a predetermined size. Furthermore, this tracking area is set so that the center of the area overlaps with the center of the target annotation or its center of gravity. Note that the center of the annotation here refers to, for example, the center of the circumscribing rectangle of the annotation.
[0109] Next, the tracking module 213 calculates the proportion of black pixels in the set tracking area (step 1205).The tracking module 213 then determines whether the calculated proportion is equal to or greater than a predetermined value (e.g., 40%) (step 1206).If the result of this determination is that the calculated proportion is not equal to or greater than the predetermined value (NO in step 1206), the tracking module 213 determines that the influence of the surgical instrument movement is small, and therefore tracks the set tracking area (step 1207).On the other hand, if the result of this determination is that the calculated proportion is equal to or greater than the predetermined value (YES in step 1206), the tracking module 213 determines that the influence of the surgical instrument movement is large, and therefore tracks the entire previous frame (step 1208).
[0110] Next, the tracking module 213 places a predetermined number of points at equal intervals on the tracking target in the previous frame and performs optical flow measurement using the Lucas-Kanade algorithm between the previous frame and the latest frame (step 1209). Next, the tracking module 213 performs forward-backward error checking (step 1210) and NCC (step 1211). Based on these calculations, the tracking module 213 measures where the previous point moved in the latest frame. The tracking module 213 obtains the median of the movement of all points as the up / down / left / right movement of the camera, and obtains the median of the change in distance between points as the forward / backward movement of the camera (in other words, the amount of scaling). The tracking module 213 then outputs the obtained movement and scaling to the annotation display module 214 (step 1212). The output movement is, for example, represented as a vector (i.e., direction and magnitude).
[0111] A further note about the above forward-backward error check and NCC: If the optical flow measurement in step 1209 is performed correctly, the point positions measured when the comparison frames are reversed (from the latest frame to the immediately preceding frame) should be close to the original point positions. In forward-backward error check, the difference between the point positions from the latest frame to the immediately preceding frame and the originally placed point positions is considered to be an error, and point positions calculated in step 1209 with an error equal to or greater than the median are excluded as not being tracked correctly.
[0112] The pixels surrounding the point before and after the movement should be similar. NCC calculates the similarity for each point, and excludes points whose similarity is lower than the median value as they have not been correctly tracked.
[0113] After outputting the amount of movement, etc., the tracking module 213 determines whether or not an instruction to end video playback has been issued (step 1213). If the result of this determination is that an instruction to end video playback has been issued (YES in step 1213), the tracking module 213 ends this flow. On the other hand, if the result of this determination is that an instruction to end video playback has not been issued (NO in step 1213), the tracking module 213 returns to step 1201 and acquires the latest frame again. This concludes the explanation of the tracking processing flow 1200.
[0114] In the tracking process flow 1200 described above, the tracking module 213 places multiple points throughout the entire previous frame when the proportion of pixels representing a surgical instrument in a partial region (i.e., a tracking region) set in the previous frame near the annotation is equal to or greater than a threshold. On the other hand, when the proportion of pixels in the partial region is less than the threshold, the tracking module 213 places multiple points only in the partial region. As a result, when the influence of the surgical instrument movement is small, only a portion of the frame is tracked, and when the influence of the surgical instrument movement is large, the entire frame is tracked.
[0115] The shape and size of the tracking area in this modification may be changed as appropriate. The position of the tracking area may be in the vicinity of the annotation, and the center of the tracking area does not necessarily have to coincide with the center of the annotation or its center of gravity.
[0116] In addition, the binarization process may be omitted in this modified example. That is, a tracking area may be set in the color image, and the tracking target may be set according to the proportion of pixels representing the surgical instrument in that area. In this case, a predetermined range of pixel values is set in advance to identify the pixels representing the surgical instrument. Then, if the proportion of pixels belonging to that range is equal to or greater than a predetermined value, multiple points are placed throughout the entire previous frame. If the proportion of pixels belonging to that range is less than the predetermined value, multiple points are placed only in the tracking area.
[0117] (8) Others The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0118] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0119] Furthermore, the control lines and information lines shown are those considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. The above-mentioned embodiments disclose at least the configurations described in the claims.
[0120] 100...surgery support system, 101...endoscopic camera, 102...encoder, 103...decoder, 104...monitor, 105...decoder, 106...capture device, 107...annotation display device, 108...encoder, 109...network, 201...main memory device, 202...auxiliary memory device, 203...processor, 204...input device, 205...output device, 206...communication control unit, 210...UI module, 211...frame acquisition module, 212...video playback module, 213...tracking module, 214...annotation display module, 220...video data, 600...menu screen, 601...pull-down menu, 602...input field, 603...input field, 604...microphone setting area area, 605...speaker setting area, 606...pull-down menu, 607...pull-down menu, 608...indicator, 609...pull-down menu, 610...pull-down menu, 611...indicator, 612...play button, 700...annotation screen, 701...pull-down menu, 702...button, 703...button, 704...button, 705...button, 706...button, 707...button, 708...seek bar, 709...video playback area, 710...surgical instrument, 711...surgical instrument, 712...annotation, 800...annotation screen, 801...annotation, 900...annotation screen, 901...annotation, 1101...internal organ, 1102...surgical instrument, 1103...tracking area
Claims
1. An annotation display device comprising: a video playback unit that plays video captured by an endoscopic camera; and an annotation display unit that performs at least one of movement, enlargement, and reduction of annotations input by a user and superimposed on the video, based on the movement of the endoscopic camera.
2. The annotation display device according to claim 1, wherein the video playback unit stops playback of the video while the annotation is being input.
3. The annotation display device according to claim 2, wherein the video playback unit, after inputting the annotation, starts playing the video by skipping images for a period of time corresponding to the input time of the annotation.
4. The annotation display device according to claim 1, further comprising a tracking unit that performs optical flow measurement on the video to calculate the movement of the endoscopic camera.
5. The annotation display device according to claim 4, characterized in that the annotation display unit accumulates values calculated sequentially by the tracking unit while the annotation is being input, and corrects the annotation based on the accumulated values after the annotation is input.
6. The annotation display device of claim 4, characterized in that the tracking unit: places a plurality of points in the immediately preceding frame of the immediately preceding frame and the latest frame that constitute the video; performs optical flow measurement to estimate where the plurality of points have moved in the latest frame; calculates the median of the movement of each point and obtains it as the up / down / left / right movement of the endoscopic camera; and calculates the median of the change in distance between each point and obtains it as the forward / backward movement of the endoscopic camera.
7. The annotation display device according to claim 6, wherein the tracking unit arranges the plurality of points over the entire immediately preceding frame.
8. The annotation display device according to claim 6, characterized in that the tracking unit performs optical flow measurement for each point whose position in the latest frame has been estimated, to estimate its position in the immediately preceding frame, and excludes points for which the error between the estimated position and a position previously placed in the immediately preceding frame is equal to or greater than a predetermined value from the calculation of the amount of movement.
9. The annotation display device according to claim 6, wherein the tracking unit excludes points where the similarity between the surrounding pixels in the immediately preceding frame and the surrounding pixels in the latest frame is equal to or less than a predetermined value from the calculation of the movement amount.
10. The annotation display device described in claim 7, characterized in that the tracking unit places the multiple points over the entire previous frame when the proportion of pixels representing a surgical instrument in a partial area set in the previous frame near the annotation is equal to or greater than a threshold, and places the multiple points only in the partial area when the proportion of pixels in the partial area is not equal to or greater than the threshold.
11. An annotation display method executed by a computer, comprising: a video playback step of playing back a video captured by an endoscopic camera; and an annotation display step of performing at least one of movement, enlargement, and reduction on an annotation input by a user and superimposed on the video, based on the movement of the endoscopic camera.
12. The annotation display method according to claim 11, wherein the video playback step stops playback of the video while the annotation is being input.
13. The annotation display method according to claim 12, wherein the video playback step starts playback of the video after inputting the annotation by skipping images for a period of time equivalent to the input time of the annotation.
14. The annotation display method according to claim 11, further comprising a tracking step of calculating the movement of the endoscopic camera by performing optical flow measurement on the video.
15. The annotation display method according to claim 14, wherein the annotation display step accumulates values calculated sequentially by the tracking step while the annotation is being input, and corrects the annotation based on the accumulated values after the annotation is input.
16. The annotation display method according to claim 14, characterized in that the tracking step comprises the steps of: placing a plurality of points in the immediately preceding frame of the latest frame that constitutes the video; measuring optical flow to estimate where the plurality of points have moved in the latest frame; calculating the median of the amount of movement of each point and obtaining it as the amount of up / down and left / right movement of the endoscopic camera; and calculating the median of the change in distance between each point and obtaining it as the amount of forward / backward movement of the endoscopic camera.
17. The annotation display method according to claim 16, wherein said tracking step arranges said plurality of points over the entire immediately preceding frame.
18. The annotation display method according to claim 16, characterized in that the tracking step comprises: a step of performing optical flow measurement for each point whose position in the latest frame has been estimated, to estimate its position in the immediately preceding frame; and a step of excluding from the calculation of the movement amount any point whose estimated position has an error of a predetermined value or more from its position previously positioned in the immediately preceding frame.
19. The annotation display method according to claim 16, characterized in that the tracking step includes a step of excluding from the calculation of the movement amount any point where the similarity between the surrounding pixels in the immediately preceding frame and the surrounding pixels in the latest frame is equal to or less than a predetermined value.
20. The annotation display method of claim 17, characterized in that the tracking step places the multiple points over the entire immediately preceding frame if the proportion of pixels representing a surgical instrument in a partial region set in the immediately preceding frame near the annotation is equal to or greater than a threshold, and places the multiple points only in the partial region if the proportion of pixels in the partial region is not equal to or greater than the threshold.
21. A program for causing a computer to execute the annotation display method according to any one of claims 11 to 19.