Method and apparatus for real-time tracking of camera configurations with automatic keyframe-based management - Patents.com

JP2025500146A5Pending Publication Date: 2025-11-26スーガー +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024532229
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-15
Filing Date
2022-12-14
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Existing methods for automatic keyframe-based management in laparoscopic surgery struggle with maintaining real-time camera tracking, especially in unexamined zones or areas with changing appearances, leading to loss of camera placement and inability to display augmented elements accurately.

Method used

A method and apparatus for real-time camera tracking that automatically manages keyframes by selecting new keyframes based on multi-criteria evaluation of image similarity and sharpness, using a buffer memory to maintain tracking even in unexamined zones, without manual intervention.

Benefits of technology

Ensures reliable and real-time tracking of camera placement, allowing accurate display of augmented reality elements by dynamically updating keyframes, reducing user dependence and system resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for tracking in real time the positioning of a camera relative to a 3D model comprising a keyframe base comprising a first execution thread making it possible to estimate the current positioning of the camera from selected keyframes, characterized in that the method further comprises the steps of analysing the quality of the estimation of the camera position from images stored in a buffer memory, analysing the sharpness of the images stored in the buffer memory from data representative of the camera movement, and adding the images stored in the buffer memory as new keyframes to the keyframe base if the quality of the estimation of the camera position from the stored images and the sharpness of the stored images meet predetermined criteria.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and apparatus for tracking the placement of a camera relative to a 3D model in real time, in particular for augmented reality applications, in particular for endoscopic surgery, in particular for laparoscopic surgery for the purposes of robotized procedures or operations. More specifically, the present invention relates to a method and apparatus that allows automatic management of a keyframe base, in particular automatic addition of keyframes during real-time tracking. [Background technology]

[0002] Laparoscopy is a medical technique for visual medical examination of the inside of a patient's body using an endoscope, or more specifically a laparoscope, which is used to view the abdominal and / or pelvic cavities of a patient. An endoscope generally includes a light source and a means for capturing the light, such as, for example, optical fibers and / or a video capture device.

[0003] During laparoscopic surgery, the laparoscope allows direct or remote viewing of the abdominal cavity, allowing observation of the surgical site and direct intervention with surgical instruments. This surgical technique has the advantage of not requiring a large opening in the abdominal wall (in contrast to open surgery), making it a minimally invasive technique.

[0004] Similarly, minimally invasive surgical procedures using an endoscope can be performed in the thoracic cavity (thoracoscopy) or in the pelvic cavity. The terms endoscopic surgery or surgical endoscopy are commonly used.

[0005] Recent technological advances have evolved laparoscopy from simply allowing the medical practitioner to view an image of the area being operated on, towards an augmented display where additional information can be displayed on a screen to accompany the observed image to assist the medical practitioner during surgery.

[0006] In particular, computer vision techniques are used on images acquired in real time by laparoscopy to provide additional information by means of augmented reality. For example, hidden structures in an organ, such as a tumor, can be displayed in the image. In particular, it may be desirable to display surgical parts (e.g., incisions) in the image of the organ. More generally, the term computer-guided procedure or surgery is used in this context.

[0007] The camera geometry, i.e. its position and orientation in a given coordinate system, is calculated in real time relative to a 3D model of the environment through which the camera is moving, so that images captured by the camera can display augmented elements from a pre-computed pre-operative augmented reality model. This preoperative augmented reality model is typically created from preoperative or intraoperative images from CT, MRI, US, or other modalities used in radiology, presumably registered beforehand to the 3D model, regardless of whether these are 2D or 3D.

[0008] In particular, the camera position is calculated from a keyframe base, which includes the position of each keyframe in a repository of a 3D model of the surgical environment, and the base operates by comparing the current image with the keyframes of the keyframe base.

[0009] This method allows to determine the placement of the camera in known zones of the environment, but in the absence of a key frame that meets the criteria of similarity with the current image, or due to changes in the appearance of the imaged organ or cavity, in particular due to movement caused by blood flow, color changes and / or texture changes, it is not possible to estimate the placement of the camera in unexamined zones, thus losing track of the placement of the camera and making it impossible to display the augmentation elements in the images captured by the camera.

[0010] Solutions are proposed for the management of keyframes in zones that are not inspected or that change appearance over time.

[0011] The first solution, used especially in laparoscopic surgery, is to manually add keyframes to the database to complement the 3D model at different viewing angles. These additions are made according to user-specific visual criteria that estimate the relevance of images to add to the keyframe base.

[0012] This solution is limited, especially within the context of surgery, because it requires frequent user intervention via dedicated human-machine interfaces (Man-Machine Interface (MMI) touch screen, keyboard, mouse, etc.), which has a detrimental effect on the user experience. Moreover, the effectiveness of this solution is highly dependent on the surgeon responsible for the surgery.

[0013] Other solutions have attempted to automate the selection of new keyframes in an attempt to relieve the user of the manual selection task. In particular, these proposed solutions involve automatic updating of the keyframe base through the use of different selection criteria.

[0014] However, these solutions are not optimal for use in a surgical context: in particular, the management of the database is complex and the automatic selection takes a lot of physical resources (such as processing, graphic or storage resources) and is not suitable for real-time use in a surgical context.

[0015] Therefore, the inventors have sought to improve the automatic management of the keyframe base by defining meaningful and effective criteria for the selection of new keyframes without compromising the overall performance of the system. Summary of the Invention [Problem to be solved by the invention]

[0016] Object of the invention The present invention aims to provide a method and apparatus for real-time tracking of a camera's placement relative to a 3D model, which enables automatic keyframe-based management of the 3D model.

[0017] SUMMARY OF THE PRESENT EMBODIMENTS The present invention aims to provide a tracking method and device that can better take into account the camera arrangement and the clarity of images added on a key frame basis.

[0018] SUMMARY OF THE PRESENT EMBODIMENT The present invention aims to provide a tracking method and apparatus that allows a user to reduce the need to add new key frames.

[0019] The present invention aims to provide a tracking method and device that allows managing the keyframe base without affecting the processing of the real-time video stream from the camera. [Means for solving the problem]

[0020] Detailed Description of the Invention To achieve this, the invention provides a method for tracking in real time the positioning of a camera relative to a 3D model, said 3D model comprising a keyframe base and a positioning base associated with each keyframe, each keyframe being characterized in particular by a reference point, said method comprising a first execution thread, said first execution thread comprising: receiving a video stream from the camera comprising a plurality of images; - mapping a reference point of the last received image of the video stream, called the current image, with a reference point of a key frame of the 3D model by a mapping algorithm; estimating the current position of the camera from the key frame, called the selected key frame, that has the most reference points in common with the current image by estimating the deviation between the position of the camera associated with the selected key frame from the position base and the current position of the camera associated with the current image, The method further comprises: incrementing a alignment loss counter if the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold; or storing the current image of the video stream in a buffer memory if a number of common reference points between the selected key frame and the current image is greater than a predetermined threshold; Furthermore, if the placement loss counter is greater than or equal to a predetermined threshold, the method further comprises a second execution thread: The second thread of execution analysing the quality of the estimation of the position of the camera from the images stored in the buffer memory; analyzing the clarity of the images stored in the buffer memory from data representative of the camera movement; adding the image stored in the buffer memory as a new key frame to the key frame base if the quality of the estimation of the position of the camera from the stored image and the sharpness of the stored image meet a predetermined criterion; and if not, removing the image stored in the buffer memory.

[0021] The tracking method according to the invention allows an automatic keyframe-based management for tracking of the camera configuration, by performing a multiple criterion selection of new keyframes from the images of the video stream coming from the cameras. In this way, the keyframe-based management is performed without manual intervention, thus relieving the user from this task. The possibility of adding keyframes during the tracking method allows the tracking to be performed reliably regardless of the viewing angle, especially in initially unknown viewing angles of the 3D model, especially in zones never seen by the camera or whose appearance changes over time.

[0022] The 3D model is generated upstream of the tracking phase by an image-based representation method, for example a method of obtaining structure from motion called the Structure-from-Motion (SFM) method. The images obtained by this method can constitute the keyframe-based static part and cannot be changed during the execution of the tracking method. The 3D model is supplemented by new keyframes via the tracking method according to the invention, which form the keyframe-based dynamic part that can be modified. The configuration base is also updated each time a new keyframe is added to add the estimated configuration of this new keyframe.

[0023] This automation of management allows for improved repeatability of tracking of camera placements by making the rendering process independent of the user.

[0024] In particular, the criteria of mapping between the current image and keyframe-based keyframes of the 3D model, the criteria of analyzing the quality of the estimation of the camera positioning, and the criteria of image sharpness from data representing the camera movement allow an optimal selection of new keyframes compatible with real-time use.

[0025] The principle is to use an old image, a previous current image placed in the buffer memory that satisfies these selection criteria, to improve the tracking of the current image of the video stream when the camera alignment is lost due to the absence of a keyframe close enough to the current image captured by the camera. The use of this old image is believed to be able to satisfy the mapping criteria and allows tracking of the camera alignment, and allows more certainty to have as a reference an image of the video stream against which the current tracking can be adjusted.

[0026] The 3D model thus supplemented allows adding augmented reality objects to the display of the video stream, adjusted according to the known camera positioning due to the effectiveness of the tracking method. For example, for laparoscopy applications, a 3D modeling of an organ can be displayed in augmented reality on top of the real organ as it is captured by the camera and displayed on the display screen. Keyframe-based updates ensure that the camera positioning is correctly tracked so that the display of the 3D modeling is permanently adjusted to the image of the real organ. The 3D modeling can be supplemented by invisible elements on the video stream, such as the presence of a tumor.

[0027] According to the present invention, the first and second threads of execution are preferably executed in parallel.

[0028] According to this aspect of the invention, the execution threads, also referred to as processes or threads, do not affect each other, in particular the keyframe-based management managed by the second execution thread, referred to as the keyframe-based management thread, does not affect the performance of the real-time tracking of the camera configuration managed by the first execution thread, referred to as the tracking thread.

[0029] According to the invention, preferably, said predetermined threshold value is greater than or equal to 50. According to this aspect of the invention, the predetermined threshold value can ensure that the image stored in the buffer memory has a large number of reference points with the key frames to allow good tracking of the camera alignment between these two images. In another variant of the invention, the predetermined threshold value can be smaller than 50, at the risk of a decrease in the quality of the alignment estimation when the threshold value is low.

[0030] According to the invention, preferably, the step of analysing the quality of the location estimation comprises at least one of the following sub-steps: a sub-step of calculating the spatial coverage of reference points; and a sub-step of calculating an error in the estimation of the estimated location between the reference points observed in the images stored in the buffer memory and the reference points calculated from the selected image and location estimation.

[0031] According to this aspect of the invention, these sub-steps allow the quality of the camera alignment estimation to be reliable for the selection of the image stored in the buffer memory as a new keyframe. The quality of the alignment estimation is more precisely defined as the quality of the parameters calculated from the alignment estimation for the mapping of the images stored in the buffer memory and the reference points in the nearest keyframe. In particular, the alignment estimation error allows the quality of this alignment estimation to be determined by ensuring that the estimation error of the alignment parameters remains small with respect to the known alignment assigned to the keyframe-based keyframe. The calculation of the alignment estimation error is calculated by a calculated reprojection between the reference points observed in the current image and the reference points predicted by the transfer and alignment estimation from the keyframe.

[0032] The calculation of the estimated error of the configuration can be, for example, a calculation of the mean squared error of the configuration or the square root of the mean squared error of the configuration.

[0033] According to the invention, preferably, the step of analysing the sharpness of the image comprises a sub-step of calculating a movement of a reference point for calculation of the velocity and / or acceleration and / or jolt of said reference point.

[0034] According to this aspect of the invention, image sharpness is taken into consideration when selecting images stored in a buffer memory as new key frames based on camera motion, with the goal of eliminating images that contain insufficient information to be integrated into the key frame base.

[0035] A jolt, also called a jerk or hyperacceleration, is also the derivative of acceleration with respect to time. Acceleration itself is the derivative of velocity with respect to time, and velocity is the derivative of position with respect to time.

[0036] According to the invention, preferably the keyframe base comprises a static part with predefined static keyframes and a dynamic part for storing new keyframes added in a step of adding images stored in a buffer memory to said keyframe base.

[0037] According to this aspect of the invention, preservation of static keyframes for keyframe-based static portions allows for limited derivatives on the positioning estimate, in particular the lack of static keyframes introduces the risk of drifting away from the positioning estimate, i.e., a gradual degradation in the positioning estimate from the keyframes.

[0038] According to the invention, preferably the method comprises the step of measuring, for each key frame, the duration since the last mapping of said key frame using images of the video stream, and if the number of said key frames is greater than or equal to a predetermined limit number, the step of adding said images stored in the buffer memory to said key frame base comprises the sub-step of removing the key frame having the longest duration since the last mapping.

[0039] According to this aspect of the invention, the setting of the keyframe limit in the keyframe base makes it possible to ensure that the mapping of keyframes with the video stream does not affect too strongly the computation time of the device executing the method, in particular to maintain the real-time aspect of the tracking method. If the keyframe limit is reached and the keyframes are removed in the static part of the keyframe base, i.e. the keyframes in the dynamic part of the keyframe base, in particular the keyframes that were previously added as the tracking method progresses, these keyframes are not removed. In other words, the management of the keyframe base is only effected in the dynamic part of the keyframe base with the keyframes that can be removed as necessary, and the management of the keyframe base does not affect the static part of the keyframe base with the static images that cannot be removed by the tracking method.

[0040] This limit number of key frames can be, for example, between 20 and 50 for the application of tracking the position of an endoscope in laparoscopic surgery, and in particular about 30 images as a suitable compromise between reasonable computation time and robustness of tracking the position of the camera. The limit number of key frames is selected as a compromise between the accuracy and robustness of the tracking method on the one hand, and the computation speed and the number of images processed per second on the other hand. The limit number of key frames depends on the equipment used and its supply means.

[0041] According to the invention there is preferably provided an apparatus for tracking in real time the placement of a camera relative to a 3D model, comprising: The 3D model comprises a keyframe base and a base of placements associated with each keyframe, each keyframe being characterized in particular by a reference point, the device comprising a first tracking module, the first tracking module comprising: a sub-module for receiving a video stream from said camera, said video stream comprising a plurality of images; a sub-module for mapping a reference point of a last received image of the video stream, called the current image, with a reference point of a key frame of the 3D model by means of a mapping algorithm; a sub-module for estimating a current configuration of the camera from a selected keyframe from the configuration base, referred to as a selected keyframe, which has the largest number of reference points in common with the current image, by estimating a shift between the configuration of the camera associated with the selected keyframe and the current configuration of the camera associated with the current image, The apparatus further comprises a module for managing the keyframe base, the module comprising: a sub-module comprising an alignment loss counter that is incremented if the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold; a sub-module for transferring the current image of the video stream to a buffer memory of the device if a number of common reference points between the selected key frame and the current image is higher than a predetermined threshold; a submodule for comparing a value of the placement loss counter with a predetermined value; a sub-module for analyzing the quality of the estimation of the camera's position from the images stored in the buffer memory; a sub-module for analyzing the clarity of the images stored in the buffer memory from data representative of the camera movement; and a sub-module for managing the buffer memory configured to add a stored image into the keyframe base if the quality of the estimation of the position of the camera from the stored images and the sharpness of the stored images meet predetermined criteria, and further configured to remove the image stored in the buffer memory if they do not meet predetermined criteria.

[0042] A module may comprise a computing device, such as, for example, a computer, a group of computing devices, an electronic component, or a group of electronic components, or, for example, a computer program, a group of computer programs, a library of computer programs, or the functionality of a computer program executed by a computing device, such as a computer, a group of computing devices, an electronic component, or a group of electronic components.

[0043] Advantageously, the tracking device according to the invention is adapted to implement the tracking method according to the invention.

[0044] Advantageously, the tracking method according to the invention is implemented by a tracking device according to the invention.

[0045] The present invention relates to a computer program product for tracking in real time the positioning of a camera with respect to a 3D model, said 3D model comprising a keyframe base and a positioning base associated with each keyframe, each keyframe being characterized in particular by a reference point, said computer program product comprising program code instructions for executing: When the computer program product is executed on a computer, the method steps of the computer program product include a first thread of execution; The first execution thread receiving a video stream from the camera comprising a plurality of images; - mapping a reference point of the last received image of the video stream, called the current image, with a reference point of a key frame of the 3D model by a mapping algorithm; estimating the current position of the camera from a keyframe, called a selected keyframe, that has the largest number of reference points in common with the current image by estimating a shift between a position of the camera associated with a keyframe selected from the position base and the current position of the camera associated with the current image, The method comprises: incrementing a alignment loss counter if the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold; or storing the current image of the video stream in a buffer memory if a number of common reference points between the selected key frame and the current image is greater than a predetermined threshold; The method further comprises: a second execution thread when the placement loss counter is greater than or equal to a predetermined value; The second execution thread analyzing the quality of the estimation of the camera position from the images stored in the buffer memory; analyzing the clarity of the images stored in the buffer memory from data representative of the camera movement; adding the image stored in the buffer memory to the keyframe base if the quality of the estimation of the position of the camera from the stored image and the clarity of the stored image meet predetermined criteria, and further removing the image stored in the buffer memory if they do not meet.

[0046] Advantageously, the computer program product according to the invention is adapted to implement the tracking method according to the invention.

[0047] Advantageously, the tracking method according to the invention is implemented by a computer program product according to the invention.

[0048] The present invention relates to an endoscopic imaging system comprising an endoscope configured to capture a video stream, a tracking device according to the present invention configured to track the position of the endoscope, and further a display screen configured to display images acquired by the endoscope and additional information provided by the processing unit according to the position of the endoscope relative to a 3D model.

[0049] The system is preferably used for laparoscopic, thoracoscopic or pelvicoscopic imaging.

[0050] The invention also relates to a tracking method, a tracking device and an endoscope system characterized by any or all of the above or below mentioned features in any combination.

[0051] List of Figures Other objects, features and advantages of the present invention will become apparent upon reading the following description, given purely in a non-limiting manner, with reference to the accompanying drawings, in which: [Brief description of the drawings]

[0052] [Figure 1] FIG. 1 is a schematic diagram of a laparoscopic imaging system according to one embodiment of the present invention. [Diagram 2] FIG. 2 is a schematic diagram of the steps of a tracking method according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0053] In the figures, for purposes of illustration and clarity, scale and proportions are not strictly respected.

[0054] Furthermore, identical, similar or analogous elements are designated with the same reference numerals in all figures.

[0055] 1 shows a schematic representation of a laparoscopic imaging system 10 according to an embodiment of the present invention. The purpose of the system is to be able to acquire and output images taken within a cavity 50 in a patient's body, in this case within the patient's abdominal cavity (i.e. the abdominal cavity 50), in particular within the scope of laparoscopic procedures, such as for example laparoscopic surgery, which is directed to the operation of a target organ 52.

[0056] To achieve this, the system 10 comprises a tracking device according to an embodiment of the present invention, which in particular comprises an endoscope 12 configured to acquire images of the patient's abdominal cavity 50. The endoscope 12 is placed within the patient's abdominal cavity 50 by means of a trocar 14 that allows the endoscope to pass through the abdominal wall 56. Endoscopes used in laparoscopic surgery are currently referred to as laparoscopes.

[0057] The tracking device comprises several modules enabling the implementation of the method according to the invention, which are in this case grouped together in a processing unit 16. The processing unit 16 is a computer or electronic platform equipped with a processor, for example a processor that contributes to the processing of the images of the method according to the invention, or a general-purpose processor that is arranged, among other functions, to execute program instructions for carrying out the steps of the method according to the invention.

[0058] Images acquired from the endoscope 12 are displayed on a display screen 18 intended for medical personnel. The images acquired may be augmented, such as with additional information added by a laparoscopic imaging system, which may come from a tracking device or other devices.

[0059] A tracking method 100 according to an embodiment of the present invention comprises several steps, which are illustrated with reference to FIG.

[0060] The method 100 allows real-time tracking of the positioning of the camera relative to a 3D model. The 3D model is provided with a keyframe base that allows the synthesis of a modeled scene to occur, which corresponds to the real scene in which the camera is moving. Each keyframe is characterized in particular by reference points that can effect a mapping between the virtual scene of the 3D model and the real scene captured by the camera, and by these reference points the positioning of the camera can be logically guided.

[0061] The 3D model comprises a geometry base that associates to each keyframe the geometry of the camera during the capture of that keyframe, which allows the keyframes to be associated with each other by their geometry represented by common markers, and the estimation of changes in geometry is performed by estimating translations and / or rotations according to the common markers.

[0062] The method comprises two execution threads, specifically a first execution thread 110, that can be executed in parallel.

[0063] A first execution thread 110 comprises a step 112 of receiving a video stream from a camera, which comprises a number of images, the camera being, for example, an endoscope imaging a cavity of a patient.

[0064] The first execution thread 110 then comprises a step 114 of mapping, by means of a mapping algorithm, the reference points of the last received image of the video stream, called the current image, with the reference points of the keyframes of the 3D model, which step makes it possible to match the mapping of the current image with the mapping of the images of the keyframes of the 3D model.

[0065] The first execution thread 110 then comprises a step 116 of estimating the current configuration of the camera from the keyframe, called the selected keyframe, which comprises the largest number of reference points in common with the current image. Due to the mapping between the current image and the keyframe, it is possible to determine that the camera has a configuration close to the configuration of the camera that allowed to obtain the selected keyframe, and therefore, by estimating from the configuration base the shift between the configuration of the camera associated with the selected keyframe and the current configuration of the camera associated with the current image, the configuration of the camera during the taking of the current image can be deduced therefrom. In particular, the configuration is estimated by estimating parameters of the configuration, for example the rotation and translation parameters of the current configuration expressed according to the reference configuration, i.e. in this context the configuration associated with the selected keyframe.

[0066] In practice, the comparison of the current image with the keyframe-based keyframes is achieved by pre-processing the image (e.g., extracting color channels depending on the application, applying a Gaussian blur, re-dimensionalizing the image, etc.) from which reference points can be easily extracted, for example by using a scale-invariant feature transform, SIFT, type detector. The reference points are then compared by a comparison method, for example by a mapping algorithm, in particular the Brute Force Matching (BFT) algorithm.

[0067] The alignment estimate is then derived from the most relevant images, i.e. the images selected, for example, by a robust estimator, in particular of the PnP RANSAC type using the Perspective-n-point projection, PnP, in particular the iterative parameter estimation method RANSAC for RANdom SAmple Consensus.

[0068] If the position estimate is reliable enough, the tracking device can augment the imagery in the video stream with augmented reality.

[0069] The method comprises steps, which may or may not be included in the first execution thread, based on a step 120 of comparing the number of common reference points between the selected keyframe and the current image. These steps may be performed from the mapping between the current image and the keyframe obtained in the mapping step 114.

[0070] If the number of common reference points between the selected keyframe and the current image is higher than a predetermined threshold (preferably between 20 and 100, for example 50), the method comprises a step 122 of storing the current image in a buffer memory 200. This step makes it possible to process the current image outside the first execution thread 110, which allows the image to be held while the video stream is provided to a unit for processing new images captured by the camera (sic).

[0071] If the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold, the method comprises a step 124 of incrementing a alignment loss counter 202 is performed. This counter makes it possible to detect a loss of camera alignment when the keyframe-by-keyframe cannot estimate the camera alignment. When the camera alignment is lost, it is no longer possible to display additional information by augmented reality.

[0072] If the placement loss counter is higher than or equal to the predetermined value, a second execution thread 150 of the method is executed.

[0073] This second execution thread 150 may be used in particular when the zone has not been previously captured by the camera, or when the zone has only been captured infrequently by the camera, or The system is configured to manage the keyframe base, and in particular to add new images to the keyframe base, if a zone changes in appearance over time and for which the 3D model does not have enough keyframes, these images can improve camera tracking in such zones.

[0074] The method includes a step 152 of analyzing the quality of the estimation of the camera geometry from the images stored in the buffer memory.

[0075] In particular, the step 152 of analyzing the quality of the estimation of the location comprises at least one of the following sub-steps: a sub-step of calculating the spatial coverage of the reference points; A substep of calculating an estimated error of the alignment, also called reprojection error, for example by calculating the mean squared error of the alignment or the square root of the mean squared error of the alignment. This error can be quantified during the estimation of the above constellation, in particular during the application of a PnP type algorithm, in particular PnPRANSAC: the mean squared error is particularly referred to as MSE, and the square root of the mean squared error is particularly referred to as RMSE.

[0076] The second execution thread 150 then comprises a step 154 ​​of analysing the sharpness of the images stored in the buffer memory from the data representative of the camera movement. The sharpness of the images is then estimated, for example from the data representative of the camera position, movement, speed, acceleration and / or rapid movement, such estimate enabling high quality images to be stored so as to be added on a keyframe basis. In particular, the step 154 ​​of analysing the sharpness of the images comprises a sub-step of calculating the movement of a reference point, in order to calculate the speed and / or acceleration and / or jitter of the reference point relative to the reference point of the selected keyframe. The analysis of the sharpness of the images stored in the buffer memory can be complemented by calculating the variance of the Laplacian to validate the analysis.

[0077] If the quality of the estimation of the camera's position from the stored images and the quality of the estimation of the sharpness of the stored images comply with predetermined criteria verified in a positioning and sharpness verification step 160, the second execution thread 150 comprises a step 156 of adding the image stored in the buffer memory as a new key frame to the key frame base, and otherwise a step 158 of removing the image stored in the buffer memory.

[0078] When the position tracking is lost due to lack of a key frame close enough to the current image of the video stream captured by the camera, the adding step can use images stored in the buffer memory that are images from the past, and the tracking of the next position of the camera can be improved by adding the images from the past to the key frame base. The position base is also updated to store the position information associated with the new key frame.

[0079] For reasons related to the overall performance of the method, the tracking method 100 may also comprise a step 164 of measuring, for each keyframe, the duration since the last mapping of that keyframe using images of the video stream, and if the number of keyframes is greater than or equal to a predetermined limit, the step of adding the images stored in the buffer memory to the keyframe base comprises a substep 166 of removing the keyframe with the longest duration since the last mapping.

[0080] The invention is not limited to the described embodiments, but is particularly applicable to any type of endoscopic imaging system, for example within the scope of endoscopy in the thoracic or pelvic cavity, and more generally to any system using augmented reality and in which it is necessary to track the position of a camera in order to display a 3D model on the actual video stream captured by the camera.

Claims

1. A method for tracking in real time the positioning of a camera relative to a 3D model, the 3D model comprising a keyframe base and a positioning base associated with each keyframe, each keyframe being characterized in particular by a reference point, the method comprising a first execution thread (110), the first execution thread (110) comprising: receiving (112) a video stream from the camera comprising a plurality of images; - mapping a reference point of the last received image of the video stream, called the current image, with a reference point of a key frame of the 3D model by a mapping algorithm; and estimating (116) the current position of the camera from a keyframe, called the selected keyframe, that has the most reference points in common with the current image by estimating the deviation between the position of the camera associated with the selected keyframe from a position base and the current position of the camera associated with the current image, The method further comprises: Incrementing (124) a placement loss counter if the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold; or storing (122) the current image of the video stream in a buffer memory if the number of common reference points between the selected keyframe and the current image is greater than a predetermined threshold, Furthermore, if the placement loss counter is greater than or equal to a predetermined threshold, the method further comprises a second execution thread (150); The second execution thread (150) analyzing (152) the quality of the estimation of the camera's position from the images stored in the buffer memory; analyzing (154) the clarity of the images stored in the buffer memory from the data representing the camera movement; adding (156) the image stored in the buffer memory as a new keyframe to the keyframe base if the quality of the estimation of the camera's position from the stored image and the clarity of the stored image meet predetermined criteria; and if there is no match, removing (158) the image stored in said buffer memory.

2. The method of claim 1 , wherein the first and second execution threads execute in parallel.

3. The tracking method of claim 1 , wherein the predetermined threshold is greater than or equal to 50.

4. The tracking method, wherein the step of analyzing the quality of the location estimate comprises the following substeps: a sub-step of calculating the spatial coverage of the reference points; and calculating an error in an estimated location estimation between reference points observed in the images stored in the buffer memory and reference points calculated from the selected images and location estimation.

5. 2. The tracking method according to claim 1, wherein the step of analyzing the sharpness of the image comprises a sub-step of calculating the movement of a reference point for calculation of the speed and / or acceleration and / or sudden movement of the reference point.

6. 2. The tracking method of claim 1, wherein the keyframe base comprises a static portion comprising predefined static keyframes, and a dynamic portion storing new keyframes added in the step of adding images stored in the buffer memory to the keyframe base.

7. 2. The tracking method of claim 1, comprising a step of measuring, for each key frame, a duration since the last mapping of the key frame using images of the video stream, and if the number of key frames is greater than or equal to a predetermined limit number, the step of adding images stored in the buffer memory to the key frame base comprises a substep of removing the key frame having the longest duration since the last mapping.

8. 1. An apparatus for tracking the placement of a camera relative to a 3D model in real time, comprising: The 3D model comprises a keyframe base and a base of placements associated with each keyframe, each keyframe being characterized in particular by a reference point, and the device comprises a first tracking module, the first tracking module comprising: a sub-module for receiving a video stream from said camera, said video stream comprising a plurality of images; a sub-module for mapping a reference point of the last received image of said video stream, called the current image, with a reference point of a key frame of said 3D model by means of a mapping algorithm; a sub-module for estimating the current position of the camera from a keyframe, called the selected keyframe, that has the largest number of reference points in common with the current image by estimating the deviation between the position of the camera associated with a keyframe selected from a position base and the current position of the camera associated with the current image, The apparatus further comprises a module for managing the keyframe base, the module comprising: a sub-module comprising an alignment loss counter that is incremented if the number of common reference points between a selected keyframe and the current image is less than or equal to a predetermined threshold; a sub-module for transferring the current image of the video stream to a buffer memory of the device if the number of common reference points between the selected key frame and the current image is higher than a predetermined threshold; a sub-module for comparing the value of the placement loss counter with a predetermined value; a sub-module for analyzing the quality of the estimation of the camera's position from the images stored in the buffer memory; a sub-module for analyzing the clarity of the images stored in the buffer memory from the data representing the camera movement; configured to add a stored image into the keyframe base if the quality of the estimation of the camera's position from the stored image and the clarity of the stored image meet predetermined criteria; and a sub-module for managing said buffer memory, configured to remove said images stored in said buffer memory if they do not match.

9. 1. A computer program product for tracking in real time the positioning of a camera relative to a 3D model, said 3D model comprising a keyframe base and a positioning base associated with each keyframe, each keyframe being characterized in particular by a reference point, said computer program product comprising program code instructions for execution, When the computer program product is executed on a computer, the method steps include: The first execution thread receiving a video stream from the camera comprising a plurality of images; - mapping a reference point of the last received image of said video stream, called the current image, with a reference point of a key frame of said 3D model by a mapping algorithm; and estimating the current position of the camera from a keyframe, called the selected keyframe, that has the largest number of reference points in common with the current image by estimating the deviation between the position of the camera associated with a keyframe selected from the position base and the current position of the camera associated with the current image, The method comprises: incrementing a placement loss counter if the number of common reference points between the selected keyframe and the current image is less than or equal to a predetermined threshold; or storing the current image of the video stream in a buffer memory if the number of common reference points between the selected keyframe and the current image is greater than a predetermined threshold; The method further comprises: a second execution thread if the placement loss counter is greater than or equal to a predetermined value; The second execution thread analyzing the quality of the estimation of the camera's position from the images stored in the buffer memory; analyzing the clarity of the images stored in the buffer memory from the data representing the camera movement; adding an image stored in the buffer memory to the keyframe base if the quality of the estimation of the camera's position from the stored image and the clarity of the stored image meet predetermined criteria, and further removing the image stored in the buffer memory if they do not meet predetermined criteria.

10. 1. An endoscopic imaging system comprising: an endoscope configured to capture a video stream; 9. A tracking device according to claim 8 configured to track the placement of the endoscope; and a display screen configured to display images acquired by the endoscope and additional information provided by the processing unit according to the positioning of the endoscope relative to a 3D model.