Method and system for producing annotated arthroscopic video streams

By generating annotated arthroscopic video streams that correct distortion and train a keypoint identifying model, the method improves landmark identification in arthroscopic procedures, reducing surgical risks through enhanced visual aids.

WO2025226200A1PCT designated stage Publication Date: 2025-10-30AUGMENTED KNEE ARTHROSCOPY AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2025/050359
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2025-04-17
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing arthroscopic images are distorted due to physical constraints of the arthroscope, making it difficult for surgeons to accurately identify landmarks or keypoints during procedures like ACLR surgery, leading to a high risk of surgical failure.

Method used

The method produces annotated arthroscopic video streams by correcting lens projection distortion using a precomputed distortion model, reconstructing a 3D representation of the joint, and training a computer-implemented keypoint identifying model to augment the video stream with graphical information.

Benefits of technology

This approach assists surgeons by providing accurate visual aids during arthroscopy, enhancing the precision of landmark identification and reducing the risk of surgical complications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025050359_30102025_PF_FP_ABST
    Figure SE2025050359_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for producing an annotated arthroscopic video stream from video frames of a training arthroscope video stream (140). The annotated arthroscopic video stream and the training video stream (140) are then used as input data for training a computer-implemented keypoint identifying model (150). The so-trained computer-implemented keypoint identifying model (150) can then be used to identify keypoints representing points of interest in or in vicinity of a joint in video frames of an arthroscopic video stream. The arthroscopic video stream is augmented by annotating pixels in video frames of the arthroscopic video stream corresponding to the keypoints. The augmentation of the arthroscopic video stream can be used as a visual aid during arthroscopy, such as during anterior cruciate ligament reconstruction surgery.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHOD AND SYSTEM FOR PRODUCING ANNOTATED ARTHROSCOPIC VIDEO STREAMS TECHNICAL FIELDThe present invention generally relates to arthroscopy and in particular to methods and a system forproducing an annotated arthroscopic video stream useful in assisting surgeons, for instance, during anterior cruciate ligament (ACL) reconstruction surgery, but it is applicable to all arthroscopic ligament reconstructions as well as assessing and evaluating any joint including meniscus and cartilage injuries. BACKGROUND Anterior cruciate ligament reconstruction (ACLR) surgery is a surgical tissue graft replacement of the ACL, located in the knee, to restore its function after an injury. The torn ligament can either be removed from the knee (most common), or preserved (where the graft is passed inside the preserved ruptured native ligament, so called remnant preserving) before reconstruction through an arthroscopic procedure. In ACLR surgery, a new ACL is made from a graft of replacement tissue from a portion of the patient's own hamstring, quadriceps or patellar tendon or an allograft from a human organ donor. ACLR surgery is performed using minimally invasive arthroscopic techniques, in which a combination of fiber optics, small incisions and small instruments are used. ACLR is a difficult surgical procedure requiring an experienced surgeon to have a good outcome. Typically, 7-10 % of ACLRs require revision due to failed surgery. One of the main difficulties in ACLR is to identify landmarks or keypoints with a high precision during surgery. This is hard since arthroscopic images are distorted due to physical constraints of the equipment. This is schematically shown in Fig.1 showing a difference between the perceived midpoint and the actual midpoint of the medial wall of the lateral femoral condyle in an arthroscopic image. Hence, the surgeons need to be trained to correctlyinterpret the arthroscopic images. A majority of the revisions of ACLR are caused by poor placement ofthe femoral tunnel with a misplacement of more than a couple of mm is associated with a high risk offailure. U.S. Patent no.10,499,996 discloses computer-assisted procedures of surgery that target rigid, non- deformable anatomical parts. Fiducial markers are attached to instruments and anatomy of interest, with each fiducial marker having a printed known pattern for detection and unique identification in images acquired by a free-moving camera, and a geometry that enables estimating its rotation and translation with respect to the camera using solely image processing techniques. There is still a need for a technology that can assist surgeons in interpreting arthroscopic images, such as in connection with ACLR surgery. SUMMARYIt is a general objective to produce annotated arthroscopic video streams.It is a particular objective to train, by the annotated arthroscopic video streams, a computer-implemented keypoint identifying model to identify keypoints in arthroscopic video streams. It is another particular objective to use such a trained computer-implemented keypoint identifying model as a tool in augmenting video frames of an arthroscopic video stream with graphical information as a visual aid during arthroscopy. These and other objectives are met by embodiments disclosed herein. The present invention is defined in the independent claims. Further embodiments of the invention are defined in the dependent claims. An aspect of the invention relates to a method for producing an annotated arthroscopic video stream. The method comprises identifying, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame. The method also comprises transforming, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model. The method further comprises undistorting, for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames. The method additionally comprises reconstructing a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream. The method further comprises selecting at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream and producing an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. Another aspect of the invention relates to a method for training a computer-implemented keypoint identifying model. The method comprises producing an annotated arthroscopic video stream according to above from a training arthroscopic video stream of video frames. The method also comprises training the computer-implemented keypoint identifying model to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint. A further aspect of the invention relates to a method for visual aid in arthroscopy. The method comprises identifying at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on a computer-implemented keypoint identifying model trained according to above and augmenting the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint. A further aspect of the invention relates to a non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to identify, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame, transform, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model, undistort, for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames, reconstruct a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream, select at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. Yet another aspect of the invention relates to a system for producing an annotated arthroscopic video stream. The system comprises at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to identify, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame, transform, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model, undistort, for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames, reconstruct a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the trainingarthroscopic video stream, select at least one keypoint in or in vicinity of the joint in a video frame of thetraining arthroscopic video stream and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. The present invention produces an annotated arthroscopic video stream from video frames of a trainingarthroscopic video stream. The annotated arthroscopic video stream can then be used to train acomputer-implemented keypoint identifying model in identifying keypoints in video frames of an arthroscopic video stream. The trained computer-implemented keypoint identifying model can thereby be used to identify such keypoints in video frames of an arthroscopic video stream allowing augmentation of the arthroscopic video stream with graphical information that can be used as visual aid assisting surgeons during arthroscopy. BRIEF DESCRIPTION OF THE DRAWINGS The embodiments, together with further objects and advantages thereof, may best be understood by making reference to the following description taken together with the accompanying drawings, in which: Fig.1 is an arthroscope image of the medial wall of the lateral femoral condyle; Fig.2 schematically illustrates transformation of scope into canonical position; Fig.3 schematically illustrates training data for the machine learning task of aligning a reference mesh over the medial wall of the lateral femoral condyle in an arthroscopic image; Fig.4 schematically illustrates graphics to augment an arthroscopic video frame (4A) and an arthroscopic video frame with augmented graphics (4B); Fig.5 is a flow chart illustrating a method for producing an annotated arthroscopic video stream according to an embodiment; Fig.6 is a flow chart illustrating an embodiment of the identifying step in Fig.5; Fig.7 is a flow chart illustrating an embodiment of the iteratively estimating step in Fig.6; Fig.8 is a flow chart illustrating additional, optional steps of the method in Fig.5 according to various embodiments; Fig.9 is a flow chart illustrating an additional, optional step of the method in Fig.5 according to an embodiment; Fig.10 is a flow chart illustrating an additional, optional step of the method in Fig.5 according to an embodiment; Fig.11 is a flow chart illustrating additional, optional steps of the method in Fig.5 according to various embodiments; Fig.12 is a flow chart illustrating additional, optional steps of the method in Fig.11 according to various embodiments;Fig. 13 is a schematic illustration of a system for producing an annotated arthroscopic video streamaccording to an embodiment;Fig. 14 is a schematic illustration of a device configured to produce an annotated arthroscopic videostream according to an embodiment; and Fig.15 illustrates an embodiment of drawing graphics objects to an arthroscopic camera feed. (15A) Screen shot of arthroscopic camera feed. (15B) Graphics objects drawn onto an arthroscopic camera feed. (15C) Annotation of graphics objects from (15B). DETAILED DESCRIPTION The present invention generally relates to arthroscopy and in particular to methods and a system for producing an annotated arthroscopic video stream useful in assisting surgeons, for instance, during anterior cruciate ligament (ACL) reconstruction surgery, but it is applicable to all arthroscopic ligament reconstructions as well as assessing and evaluating any joint including meniscus and cartilage injuries.ACL reconstruction (ACLR) surgery is performed using minimally invasive arthroscopic techniques.Arthroscopic images are, however, distorted due to physical constraints of the arthroscope, which means that perceived landmarks or keypoints in the arthroscopic images may actually differ from the actual landmarks or keypoints. This is shown in Fig.1 illustrating the difference between the perceived midpoint and the actual midpoint of the medial wall of the lateral femoral condyle in an arthroscopic image. Successful outcome of ACLR surgery thereby requires trained and experienced surgeon that can handle such a distortion in the arthroscopic images. The present invention allows generation of annotated arthroscopic video streams that can assist surgeons during arthroscopy, such as ACLR surgery. This means that the invention can provide visualaid in the form of augmented reality (AR) during arthroscopy by rendering graphics into the arthroscopicvideo stream from the arthroscope in real-time, thereby facilitating measurements and recognition of areas of interest in the video stream. In more detail, the invention produces an annotated arthroscopic video stream, which can be used to train a computer-implemented keypoint identifying model to identify keypoints (points of interest) in video frames of an arthroscopic video stream. The so-trained computer-implemented keypoint identifying model can then be used for augmenting an arthroscopic video stream and thereby provide visual aid during arthroscopy. Keypoint, as used herein, refers to a landmark or reference marker in video frames of an arthroscopic video stream. Keypoint is also referred to as point or area of interest in the art. Illustrative, but non-limiting, examples of such keypoints in the context of ACLR could be midpoint of the medial wall of the lateral femoral condyle, points at lateral intercondylar ridge (LIR), bifurcate ridge (BR), femoral notch roof, over-the-top position (OTP), posterior notch outlet, points at the cartilage boundary, lateral and medial tibialeminences, medial and lateral tibial plateau articular cartridge borders, ACL ridge, ACL tubercle, anterolateral fossa, retro-eminence ridge, anterior horn of the lateral meniscus, anterior intermeniscal ligament, native ACL footprint on both femur and tibia, etc. Arthroscopy, also referred to as arthroscopic or keyhole surgery, is a minimally invasive surgical procedure on a joint, in which an examination and treatment of damage is performed using an arthroscope, i.e., an endoscope that is inserted into the joint through a small incision. Endoscopes, including arthroscopes, are inspection instruments comprising an image sensor, optical lens, light source and mechanical device. Endoscopes use tubes, which are typically only a few millimeters thick to transfer illumination in one direction and high-resolution images in real-time in the other direction, resulting inminimally invasive surgeries. Accordingly, an arthroscopic video stream as used herein represents avideo stream recorded or generated using an arthroscope. Fig.5 is a flow chart illustrating a method for producing an annotated arthroscopic video stream according to an embodiment. The method comprises identifying, in step S1 and for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame. A next step S2 comprises transforming, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model. The method also comprises undistorting, in step S3 and for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames. The method also comprises reconstructing, in step S4, a three-dimensional (3D) representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream. A next step S5 comprises selecting at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream. The method further comprises producing, in step S6, an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. Steps S1 and S3 of the method shown in Fig.5 are performed to generate a stream of canonically positioned and undistorted video frames from the input arthroscopic video stream referred to as trainingarthroscopic video stream herein. These steps S1 and S3 are therefore performed for each video frame of the training arthroscopic video stream that should be transformed and undistorted, which is schematically illustrated by the loop L1 in Fig.5. The method of Fig.5 starts by identifying the lens projection of the arthroscope within the video frame in step S1 and then transforms that identified position into a canonical position in step S2. This canonical position is typically a position selected during calibration of the arthroscope and thereby has a lens projection distortion that is defined by a precomputed distortion model. This distortion model is typically computed during arthroscope calibration and applies for the selected canonical position. The distortion model thereby corrects for distortion caused by physical elements in the lens not being perfectly aligned. The transformation of video frames in step S2 thereby transforms the lens projection position into the canonical position, at which the distortion model is known and valid, to thereby use the precomputed distortion model to undistort the lens projection of the arthroscope in step S3. The result is, thus, a stream of canonically positioned and undistorted video frames. As an alternative to transforming the lens projection position into the canonical position in step S2, the distortion model known and valid at thecanonical position could be transformed to be applicable at the lens projection position.Steps S1 to S3 of Fig.5 thereby similarity transform the video stream from the arthroscope with respect to scale, shift and rotation of the lens projection. This is typically referred to as a similarity transformation. In an embodiment, a random sampling consensus (RANSAC) style algorithm is used for the transformation. RANSAC is an iterative method to estimate parameters of a mathematical model from a set of observed data that contains outliers, when outliers are to be accorded no influence on the values of the estimates. Therefore, it can be interpreted as an outlier detection method. It is a non-deterministic algorithm in the sense that it produces a reasonable result only with a certain probability, with this probability increasing as more iterations are allowed. RANSAC uses repeated random sub-sampling. A basic assumption is that the data consists of “inliers”, i.e., data whose distribution can be explained by some set of model parameters, though may be subject to noise, and “outliers”, which are data that do not fit the model. The outliers can come, for example, from extreme values of the noise or from erroneous measurements or incorrect hypotheses about the interpretation of data. RANSAC also assumes that,given a (usually small) set of inliers, there exists a procedure, which can estimate the parameters of amodel that optimally explains or fits this data. In an embodiment, Canny edge detection can be used to extract candidate points that indicate the edgeof the scope circle in the training arthroscope video stream. Scope circle as used herein represents the circle or outer diameter of the arthroscope. Canny edge detection is an edge detection operation that uses a multi-stage algorithm to detect edges in images. For instance, a random subset is sampled fromthe candidate points, making use of inverse distance weighting and likelihood given recent scopelocations to improve sampling efficiency. A candidate scope location is computed from the sampled subset. Inliers are selected among edge points, yielding a new collection of candidates, to which a new scope location can be estimated. Thus, scope localization is improved iteratively, in standard RANSAC fashion. Once the scope is localized, video frames can be scaled, shifted and rotated into the canonical position, and thereafter the lens projection is undistorted with respect to the precomputed distortion model. These steps of the method are done both to stabilize live arthroscopic video streams, and to preprocess collected data. The video manipulations are preferably stored, so that they can later bereversed. Lens rotation as part of the similarity transformation can be determined according to variousembodiments. For instance, the rotation can be determined by locating a bulge on the circle. This can be done as part of the process described above. Hence, the position of a “circle with a bulge” can be estimated instead of a circle. This can be done by transforming the border of the circle into a flat strip and matching against a similarity filter. In an embodiment, step S1 of Fig.5 is performed as shown in Fig.6. In such an embodiment, candidate pixels indicating an edge of the arthroscope within a video frame are extracted in step S20 for each video frame of the training arthroscopic video stream. A next step S21 comprises iteratively estimating, for each video frame of the training arthroscopic video stream, the position of the lens projection of the arthroscope based on the candidate pixels. In an embodiment, step S20 in Fig.6 comprises applying, for each video frame of the training arthroscopic video stream, Canny edge detection on the video frame to extract the candidate pixels. In an embodiment, step S21 in Fig.6 comprises iteratively estimating, for each video frame of the training arthroscopic video stream, the position of the lens projection of the arthroscope by RANSAC. Such an embodiment is shown in Fig.7 and comprises sampling candidate pixels into a candidate set in step S22, fitting a model for the lens projection of the arthroscope to the candidate set in step S23 and identifying a consensus set of candidate pixels that fit the model for the lens projection of the arthroscope and adding the identified consensus set of candidate pixels to the candidate set in step S24. Steps S22 to S24 are then repeated at least once by sampling the candidate pixels from the consensus set, which is schematically illustrated by the loop L2 in Fig.7. The Canny edge detection applied in step S20 and the RANSAC process used in steps S22 to S24thereby identify the position of the lens projection and where the localization is improved iteratively.Fig.8 is a flow chart illustrating an embodiment of step S2 in Fig.5. In this embodiment, step S2 comprises determining, in step S30 and for each video frame of the training arthroscopic video stream, the scaling, shift and / or rotation that transform i) the identified position of the lens projection of the arthroscope to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope. A next step S31 comprises scaling, shifting and / or rotating, for each video frame of the training arthroscopic video stream, i) the identified position of the lens projection of the arthroscope to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope.. Fig.2 schematically illustrates the lens projection identified in step S1 in Fig.5, or in the illustrative embodiments of Figs.6 and 7, in hatching and the canonical position shown in full line. The transformation in step S2, or in the illustrative embodiment of Fig.8, thereby transforms the hatched circle to correspondto the full line circle or vice versa and where the operations, such as scaling, shifting and / or rotating,needed to match the two circles in size and position, i.e., align or overlay the circles in the video frame, correspond to the transformation used to transform the identified position of the lens projection of the arthroscope into the canonical position. In an embodiment, the method also comprises S32 as shown in Fig.8, which comprises storing, for each video frame of the training arthroscopic video stream, information of the scaling, shift and / or rotation determined in step S31 in a memory. In an embodiment, annotation of video streams from arthroscopy can be provided using the photogrammetric technique structure from motion (SfM), which allows annotated target points to be tracked through the video streams. Generally, SfM is a photogrammetric range imaging process of estimating 3D structures of a scene from a set of 2D images that may be coupled with local motion signals. Given a sequence of canonically positioned and undistorted video frames from a video stream of the arthroscope, SfM can be used to recreate a 3D geometry in the form of a mesh of triangles or a point cloud representing the interior of the joint as well as camera movement relative to this 3D geometry. In order for SfM to be successful, the arthroscopic video stream is preferably preprocessed as discussed in the foregoing in connection with steps S1 to S3 in Fig.5. Points in a 3D space can then be associated with pixels and points in time in the training arthroscopic video stream. This allows annotating keypoints in a single video frame of the training arthroscopic video stream and translating these keypoints to the entire video stream. Other approaches for this step include temporal interpolation and local filter-based point tracking. Based on manual annotations in the training arthroscopic video streams, and the solved arthroscope or camera movement and geometry, for instance a 15-point “skeleton” mesh can be positioned to cover, for example, the medial wall of the lateral femoral condyle, in a half disc whose perimeter coincides with the cartilage boundary as shown in Fig.3. The positions of the nodes in the mesh are spread out on thesurface to indicate certain geometric features. For example, one of the points indicates the half-way pointfrom the near end to the far end. It is also possible to annotate the position of the tip of a marking instrument, such as a Steadman / Chondral pick or an arthroscopic probe when visible, as well as the position of a pilot hole. The pilot hole is an indentation made on the surface of the condyle. This indentation is subsequently used to “lock” the drill to the condyle by attaching the tip of the drill to the indentation before activating the drill. In an embodiment, step S4 of Fig.5 comprises reconstructing the 3D representation by applying SfM on the stream of the undistorted video frames with canonically positioned lens projections. In a particular embodiment, this step S4 comprises reconstructing the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections to reconstruct a 3D representation in the form of a mesh of triangles or a point cloud representing the interior of the joint. Fig.3 illustrates an example of such an approach with a mesh of triangles consisting of 15 points or nodes covering the medial wall of the lateral femoral condyle in a half disc with perimeter coinciding with the cartilage boundary. An advantage of such a mesh of triangles is that the nodes of the triangles, i.e., the corners or vertices, could be spread out on the surface to indicate certain geometric features. Hence, in an embodiment, step S6 in Fig.5 comprises generating the annotated arthroscopic video stream defining, for each video frame of the annotated arthroscopic video stream, a mesh covering the medial wall of the lateral condyle of femur with a perimeter coinciding with the cartilage boundary. An alternative to such a mesh of triangles could be a point cloud spanning the interior of the joint. Fig.9 illustrates a flow chart of an additional step of the method in Fig.5 according to an embodiment. The method continues from step S4 in Fig.5. A next step S40 comprises estimating, for each canonically positioned and undistorted video frame of the stream of canonically positioned and undistorted video frames, arthroscope movement relative to the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections. In an embodiment, the movement of the arthroscope relative to the reference geometry (3D representation) is solved at the same time as the 3D representation itself by applying SfM. The reconstructed 3D geometry generated, such as, by using SfM in the form of, for instance, a mesh of triangles or point cloud can then be used to associate keypoints in the 3D geometry with pixels and pointsin time. This allows annotating keypoints in a single video frame and translating such keypoints identifiedin one video frame into other video frames of the video stream. Thus, in an embodiment, identifying a respective position of at least one keypoint in step S6 in Fig.5 comprises identifying the respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on information of the scaling, shifting and / or rotation and the precomputed distortion model. In an embodiment, the method of Fig.5 may comprise an additional step S7 as shown in Fig.10. In such an embodiment, the method continues from step S6 in Fig.5. Step S7 comprises annotating a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur in video frames of the annotated arthroscopic video stream, in which the pilot hole is visible. Thus, in this embodiment, the position of the pilot hole is annotated, i.e., indicated, in those video frames where the pilot hole is visible. Furthermore, the annotation in step S7 may additionally, or alternatively, involve annotating the position of the tip of any marking instrument if visible in video frames. Such marking instruments include microfracture picks, such as Steadman picks or chondral picks, and arthroscopic probes. Hence, in an embodiment, step S7 comprises annotating a position of a marking instrument, such as a chondral pick or an arthroscopic probe, in video frames of the annotated arthroscopic video stream, in which the tip of the marking instrument is visible. The above-described embodiments thereby produce an annotated arthroscopic video stream that can be used as training data for training a computer-implemented (CI) keypoint identifying model. An aspect of the invention therefore relates to a method for training a computer-implemented keypoint identifying model. The method comprises producing an annotated arthroscopic video stream from a training arthroscopic video stream of video frames as described herein. The method also comprises, see step S8 in Fig.11, training the computer-implemented keypoint identifying model to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream. The at least one keypoint represents a respective point of interest in or in vicinity of a joint. The computer-implemented keypoint identifying model is thereby trained using the annotated arthroscopic video stream as training data. The annotated arthroscopic video stream thereby comprisesvideo frames together with annotations of at least one keypoint in the video frames, which are therebyused as training data for the computer-implemented keypoint identifying model. The computer- implemented keypoint identifying model is trained by the training data to recognize and identify keypoints in video frames of an input (non-annotated) arthroscopic video stream. In other words, once trained, the computer-implemented keypoint identifying model receives video frames of an arthroscopic video stream as input and outputs keypoints identified in the video frames. The training data thereby preferably comprises the video frames and the point annotations placed in the annotated arthroscopic video stream as described in the foregoing. The annotations could, for instance, be in the form of 3D annotations that are re-projected on the video frames using information of the positions of the arthroscope relative to the 3D model, which were resolved with SfM and the position of the lens projection relative to the video frames. The computer-implemented keypoint identifying model may be implemented according to various embodiments. For instance, the keypoint identifying model is a computer-implemented keypoint identifying model and could be in the form of a machine learning (ML) model. Generally, ML algorithms build a mathematical model based on training data, e.g., arthroscopic video stream annotated according to the invention, in order to make predictions or decisions without being explicitly programmed to do so. There are various types of ML algorithms that differ in their approach, the type of data they input and output, and the type of task or problem that they are intended to solve. Illustrative, but non-limiting, examples of such ML algorithms include supervised learning algorithms, unsupervised learning algorithms, semi-supervised learning algorithms, reinforcement learning algorithms, self-learning algorithms, feature learning algorithms, sparse dictionary learning algorithms, anomaly detection algorithms, and association rule learning algorithms. Performing machine learning involves creating a model, which is trained on training data and can then process additional data to make predictions or decisions. Various types of ML models could be used according to the embodiments, including, but not-limited to artificial neural networks, decision trees, support vector machines, regression analysis, Bayesian networks and Genetic algorithms. Furthermore, deep learning, also known as deep structured learning, is a ML method based on artificial neural networks with representation learning. Learning can be supervised, semi-supervised or unsupervised. Deep learning architectures, such as deep neural networks, deep belief networks, recurrent neural networks and convolutional neural networks, could be used to train and implement the computer-implemented keypoint identifying model. "Deep" in deep learning comes from the use of multiple layers in the network. Deep learning is concerned with an unbounded number of layers of bounded size, which permits practical application and optimized implementation, while retaining theoretical universality under mild conditions. In deep learning the layers are also permitted to be heterogeneous and to deviate widely from biologically informed connectionist models, for the sake of efficiency, trainability and understandability. Hence, in an embodiment, step S8 in Fig. 11 comprises training a computer-implemented keypoint identifying model implemented by a keypoint regression network to identify the at least one keypoint in video frames of the arthroscopic video stream based on the annotated arthroscopic video stream and the video frames of the training arthroscopic video stream. In a particular embodiment, the computer-implemented keypoint identifying model is implemented by a keypoint region-based convolutional neural network (R-CNN). R-CNN is a type of deep learning architecture used for object detection. R-CNN typically operates by extracting region proposals from an input image, such as video frame, compute CNN features and then classify the regions based on the computed CNN features. There are various R-CNN models that could be used according to embodiments including original R-CNN, fast R-CNN, faster R-CNN, mask R-CNN and mesh R-CNN. Given an input image, the original R-CNN begins by applying a mechanism called Selective Search to extract regions of interest (ROI), where each ROI is a rectangle that may represent the boundary of anobject in the image. After that, each ROI is fed through a neural network to produce output features. Foreach ROI’s output features, a collection of support-vector machine classifiers is used to determine whattype of object (if any) is contained within the ROI. While the original R-CNN independently computed theneural network features on each of as many as two thousand ROIs, fast R-CNN runs the neural network once on the whole image. At the end of the neural network is a novel method called ROIPooling, which slices out each ROI from the neural network's output tensor, reshapes it, and classifies it. As in the original R-CNN, the fast R-CNN uses Selective Search to generate its region proposals. While fast R-CNN used Selective Search to generate ROIs, faster R-CNN integrates the ROI generation into the neural network itself. While the original, fast and faster versions of R-CNN focused on object detection, mask R-CNN adds instance segmentation. Mask R-CNN also replaced ROIPooling with a new method called ROIAlign, which can represent fractions of a pixel. Mesh R-CNN adds the ability to generate a 3D mesh from a 2D image. The invention allows an artificial neural network to be trained to recognize target points (keypoints) in arthroscopic video streams, using the aforementioned annotated arthroscopic video streams as trainingdata. With the 3D geometric information produced as described herein, the artificial neural networks canbe trained for computer vision tasks. Of particular interest are keypoint regression networks, such as a keypoint R-CNN. This means that the neural network outputs the positions within the video frame, which correspond to the points of interest (keypoints) within the joint. With enough such keypoints, we have a well-defined 3D pose of the joint relative to the camera or arthroscope. Although keypoint regression networks, such as keypoint R-CNN, is a suitable example for implementing the computer-implemented keypoint identifying model, the embodiments are not limited thereto. Also other types of neural networks or ML models that can be trained to identify keypoints in video frames of an arthroscopic video stream could be used according to the embodiments. The computer-implemented keypoint identifying model trained according to the embodiments can then be used for visual aid in arthroscopy. For instance, the computer-implemented keypoint identifying model can be used to detect points of interest (keypoints) in video streams from arthroscopes during surgery. The detected points can then be used to position graphic objects in the arthroscopic video stream and where these graphic objects can be used as visual aid for the surgeon during surgery. A further aspect of the invention relates to a method for visual aid in arthroscopy. The method comprises identifying, see step S9 in Fig.11, at least one keypoint representing a respective point of interest in orin vicinity of a joint in video frames of an arthroscopic video stream based on a computer-implementedkeypoint identifying model trained according to the invention, see step S8. The method also comprises augmenting, in step S10, the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint. The invention uses the artificial neural network to detect points of interest (keypoints) in an arthroscopic video stream from an arthroscope. The detected points of interest can then be used to position graphicobjects in the arthroscopic video stream. In an embodiment, the neural network outputs the estimatedpixel positions of, for instance, the nodes in a 15 point skeleton mesh for a given video frame in the arthroscopic video stream, and optionally estimated pixel positions of the tip of a marking instrument (if visible), and / or optionally estimated pixel positions of the pilot hole (if visible). The estimated pixel positions are optionally, but preferably filtered in time to account for predictive noise, using, for example, a Kalman filter. The estimated pixel positions of the tip of the marking instrument and / or the pilot hole may be compared to the keypoint positions in order to translate the pixel positions to “condyle coordinates”. Using the pixel coordinates of the nodes in the skeleton mesh, the shift, scale and / or rotation of the arthroscope, as well as the camera or arthroscope intrinsics and distortion model, the perspective-n-point problem can be solved to find 3D positions of the pixels relative to the camera or arthroscope. Optionally, another filtering can be done on the 3D coordinates using a state space model. This may be accomplished by an extended Kalman filter, or another state space model relative to the reprojection error of the 3D coordinates.In an embodiment, multiple, i.e., at least two, keypoints are identified in the video frames in step S9 basedon the computer-implemented keypoint identifying model. In an embodiment, step S9 in Fig.11 comprises estimating pixel positions of points in a mesh in the video frames of the arthroscopic video stream based on the computer-implemented keypoint identifying model trained in step S8. The mesh could, for instance, be a 15-point skeleton mesh as an illustrative, but non-limiting, example of a mesh that could be estimated in step S9. In a general embodiment, the mesh is a N-point skeleton mesh, wherein N≥3. In an embodiment, the method comprises an additional step S50 as shown in Fig.12. This step S50 comprises filtering the pixel positions in time to suppress predictive noise. The filtering in step S50 is then preferably performed using a Kalman filter. In an embodiment, the method also comprises the additional steps S51 to S53 shown in Fig.12. In such an embodiment, step S51 comprises identifying, for each video frame of the arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame. A next step S52 comprises determining, for each video frame of the arthroscopic video stream, a transformation between the identified position of the lens projection of the arthroscope and a canonical position having a lens projection distortion defined by a precomputed distortion model.3D coordinates corresponding to the atleast one keypoint are then estimated in step S53 relative to the arthroscope based on the pixel positions,information of the transformation and the precomputed distortion model. Steps S51 to S53 shown in Fig.12 can be implemented without the filtering step S50. Fig. 12 illustrates an additional, optional step S54 of the method by filtering the 3D coordinates corresponding to the at least one keypoint using a state space model. The filtering done in step S54 is preferably performed using an extended Kalman filter. The arthroscopic video stream could be augmented in step S10 with various types of graphic objects. Figs.4A and 4B illustrate an example of such an embodiment, with Fig.4A showing an overlay preview and Fig.4B showing the overlay in an arthroscopic video stream. In this illustrative example, the graphicobjects comprise an outline (A) of the footprint, level curves (C) on the footprint and the contour of avirtual drill hole of a pre-configured drill diameter or radius (B). In Fig.4B several such contours of virtualdrill holes are shown, which could represent different drill diameters, such as 8, 9 and 11 mm. The level curves (C) could then indicate the scale, such as in the form of a pre-configured percentage of the distance from the far cartilage border, or an actual distance. The scale of the graphic objects can be fixed given a known distance between two points on the footprint. This distance can, for instance, be specified by the user from manual measurements, or estimated from diagnostic images, such as magneticresonance imaging (MRI) images or computed tomography (CT) images. It is also possible to fix the scale by calibrating the 3D model relative to a tool, such as a marking instrument or an arthroscopic probe, that is inserted into the portal and placed tangent to the surface of the condyle. In an embodiment, step S10 of Fig.11 comprises annotating pixels in video frames of the arthroscopic video streams corresponding to the at least one keypoint selected from the group consisting of a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur, a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, a position of the arthroscope relative to the condyle, a position of a drill hole, and level curves representing distances to the cartilage border. In a particular embodiment, step S10 comprises annotating pixels in video frames of the arthroscopic video stream defining at least one graphic object positioned, such as centered, relative to a pilot hole representing an indentation made in a surface of the lateral condyle of femur or relative to a drill hole in the lateral condyle of femur. For instance, the video frames of the arthroscopic video stream could be augmented or annotated with circles (B) shown in Figs.4A and 4B centered relative to the pilot hole or the hole drilled in the indentation defined by the pilot hole, i.e., the drill hole. Alternatively, such circles (B) could be centered at a pre- selected position or at the tip of hook as illustrative, but non-limiting, examples. In an embodiment, step S10 of Fig.11 comprises annotating pixels in video frames of the arthroscopic video stream with at least one graphic object selected from the group consisting of contour of a virtual drill hole of a pre-configured drill diameter or radius, a level curve, a footprint of at least a portion of a surface of the condyle, and any combination thereof. The augmentation of the arthroscopic video stream is preferably done in real-time on a video stream as generated by an arthroscope. This means that the user, such as a surgeon, is presented with graphical objects as visual aid in the arthroscopic video stream as it is presented in real-time to the user.The so-presented graphical information in the augmented arthroscopic video stream thereby assists thesurgeon by providing visual aids allowing identifying of keypoints or structures within arthroscopic video stream allowing the surgeon to correctly identify and interpret video stream even with distortion caused by the arthroscope. The augmentation of the arthroscopic video stream does not necessarily have to be done in real-time. The augmentation can alternatively be performed offline on an already recorded arthroscopic video stream. Herein follows an example of creating an annotated arthroscopic video stream. Let an SfM solution for acollection of video frames ^ , ^ = 1, ⋯ , ^ and ^ ∈ ^^ ℕ consist of a collection of 3D points ^^ ∈ ℝ , ^ =1, ⋯ , ^, ^ ∈ ℕ, and a sequence of projection functions where ^^ projectspoints in the 3D space to video frame ^^ . A projection function (transformation) may be composed of anaffine transformation in ℝ^, a projective transformation, a distortion adjustment, such as radial distortion,and camera intrinsic matrix transformation. For example, suppose that the 3D solution has beenannotated with 3 points ^^, ^^, ^^ ∈ ℝ^; let ^ denote the curve along the boundary between the articularcartilage and the femoral condyle and define 1. ^^ as the point on ^ furthest from the ground in standing position;2. ^^ as the point on ^ nearest to the ground in standing position; and3. ^^ as the point on ^ furthest in total from the previous two points.The same procedure can be carried out with any collection of anatomical landmarks.We annotate the video stream by transforming the 3D points of the anatomical landmarks ^^, ^^, ^^ intoeach of the video frames. The annotations may consist of triplets of numbers (^^^, ^^^ , ^^^) ∈ [0,1]^where (^^^^, ^^^^) = ^^(^^) denotes the position of keypoint ^ in video frame ^^ of width ^ andheight ^, and where ^^^ denotes the visibility of keypoint ^ in the video frame ^^ . Define the boundingbox ^^ = (^^, ^^ , ^^ , ℎ^), where (^^, ^^ ) denote the pixel coordinates of the center and (^^ , ℎ^) denotethe width and height of the bounding box.To achieve more precise visibility data, it is preferred that the visibility flag is set to zero when the pointis obstructed or occluded, e.g., by surgical tools or by organic matter. This may be done by iteratingthrough the sequence of video frames and flagging points as obstructed or occluded by clicking in agraphical user interface with a mouse.A neural network can be trained given an arthroscopic video stream annotated as described above. Inmore detail, given an annotated arthroscopic video stream consisting of a sequence of video frames^^, ⋯ , ^^ and keypoint annotations (^^^, ^^^ , ^^^) ∈ [0,1]^, ^ = 1, ⋯ , ^, ^ = 1, ⋯ , ^, where ^ ∈ ℕis the number of keypoints. For ^ = 1, ⋯ , ^ define the matrix representing a training data label corresponding to training data image or video frame ^^ . Let denote a convolutional neural network (Bishop, Pattern recognition andmachine learning, Springer, 2006; Szeliski. Computer vision: algorithms and applications. SpringerNature, 2022) with parameters ^, which maps 3 channel images (such as red, green, blude (RGB)images) of height ^ and width ^ to ^ triples of numbers corresponding to the keypoint annotations. Letℒ denote a loss function with respect to the training data. For simplicity, we can set it to be the coordinate-wise mean squared error denoting by ^∙^^F the squared Frobenius norm on ℝ^×^. Alternatively, select a more suitable loss functioninvolving box loss for the bounding boxes ^^ , and keypoint loss according to KeypointRCNN (Kaiming etal., Mask R-CNN, In Proceedings of the 2017 IEEE international conference on Computer Vision (ICCV),pages 2961–2969, 2017). The neural network is trained by minimizing the loss over all possibleparameters ^, denoting the optimal parameters ^∗ = argmin^ℒ(^), which is approximatednumerically using a suitable gradient based optimizer, such as stochastic gradient descent (SGD).Graphics objects can be drawn on an arthroscopic video stream as described below. Given a neuralnetwork NN^∗: ℝ^×^×^ → ℝ^×^ trained according to above, which trained neural network NN^∗outputs predicted locations of a predefined collection of ^ anatomical landmarks, consider a stream ofvideo frames ^^^ , ^ = 1, ⋯ , ^^. This stream of video frames is distinct from the video stream(s) used totrain the neural network, and can, for instance, come from a live video stream from an arthroscopiccamera during surgery or an offline video stream. See Fig. 15A for an example of a captured videostream, also referred to as video or camera feed herein. There are multiple alternatives for using theoutput of the neural network to draw graphics on the video stream. Some examples are listed here, andthey can be used separately or combined to produce a useful navigation system.Example 1: Draw keypointsIf the keypoints represent important anatomical landmarks useful for navigating the joint, such as theintercondular ridge or the bifurcate ridge, the predicted location of the keypoints can be drawn on to thevideo frame, e.g., with small circles at each predicted keypoint with sufficiently high visibility.Example 2: Render a 3D model of a femurSuppose we have a 3D model of a femur (e.g., CT scan of the patient), and suppose that this 3D modelhas been annotated with the locations corresponding to the keypoints in the training data used to trainthe neural network. If we have a large enough number of keypoints ^, we can solve a perspective-n-point problem (PnP, see Szeliski. Computer vision: algorithms and applications. Springer Nature, 2022)in order to find a perspective that rotates the 3D model into the video stream. This lets us render the 3Dmodel on top of the video stream. However, this rendered 3D model may obstruct the view of the jointeven if rendered with transparency, which is typically not desirable.Example 3: Render graphics relative to a 3D model of a femurInstead of drawing the 3D model directly on the video stream in Example 2, we can have parametrizeddrawing on the 3D model that is rendered. For example, given a target percentage for portal placementvertically on the condular surface (which is usually around 30%), we can compute the height ℎmodel ofthe condyle surface in the 3D model, and find the sphere centred at the furthest point of the cartilageborder with radius equal to the target percentage of the condular height ℎmodel . Intersecting this spherewith the 3D model gives us a curve that can be rendered on top of the video stream.Example 4: Projecting a fixed drill contourSuppose the surgeon has a preferred drill position on the femur, and this position is known in thecanonical femur coordinates. This position can be entered by selecting a point on the reference 3D modelin a graphical user interface. The desired drill position can be indicated in the video stream with a marker(e.g., a small circle) by projecting the 3D position into the video stream similarly to Example 3.If we also have fixed the scale we are able to draw the drill contour resulting from positioning the drill atthe desired position. The scale can, for instance, be fixed in two ways:1. we have a CT model of the specimen from the video stream; or2. we have a reference CT model with some known measurement, and have taken thecorresponding measurement from the specimen in the video stream, allowing us to rescale the referenceCT model.In order for the contour to be accurate, we preferably know beforehand the diameter of the drill and the3D direction of the drill portal position. This information fixes both the angle of incidence affecting theellipticity of the contour and the direction of the major axis of the elliptical contour. A more accuratecontour can be computed by intersecting the 3D model of the femur with a cylinder of radius equal to thatof the drill which passes through the desired drill position and the drill portal position.Knowing the exact location of the drill portal position relative to the condyle can be difficult in practice but,assuming that the surgeon uses a standard three portal technique (Cohen and Fu, Three-portal techniquefor anterior cruciate ligament reconstruction: use of a central medial portal. Arthroscopy: The Journal ofArthroscopic & Related Surgery, 23(3):325.e1–5, 2007), the position can be approximated with apredefined position relative to the keypoints, which is adjusted to the fixed 3D model scale.Example 5: Projecting a drill contour relative to a tracked pointWe can choose to center the drill contour around a tracked point. The point in question can be one of theoutput points from the neural network, provided that the point represents a suitable drilling position. Wemay extend the training by, for each video frame ^^ , letting the first ^ − 1 keypoints represent anatomicallandmarks and letting (^^^, ^^^ , ^^^) represent the pixel location and visibility of the tip of a surgical tool,such as a hook or a probe. If there are no surgical tools in the training data, they can be added digitallyusing photo editing software.The neural network outputs pixel coordinates, so centring the drill hole contour around a pixel positionrequires that the pixel coordinates are transformed to a location on the 3D model. This is a straight-forward computation since the same perspective information that lets us draw the 3D model onto thescreen can be used to map pixel coordinates to rays in the 3D space of the 3D model of the femur.Centring the drill contour around a surgical tool lets the surgeon interact with the system by moving thetool to a candidate drill hole position and seeing a contour of the resulting drill hole immediately, beforedrilling into the femur and causing irreversible anatomical changes.Example 6: Drawing a perspective view of the 3D model with additional graphicsMost of the graphical objects described above have fixed coordinates relative to a 3D model. Instead ofprojecting the graphical objects to the screen and overlaying them to the video stream, they can berendered to a separate image from a separate perspective that is suitable to get an overview of theanatomy. Call this separate image the perspective view. The perspective view can be drawn to the displayon top of the video stream, positioning it to the side where it does not obstruct the arthroscope view. Wecan indicate the camera location in the perspective view by rendering a 3D model of an arthroscope atthis location. The field of view can be displayed by rendering the 3D model of the femur with a directionallight centred at the camera location pointing in the viewing direction.Fig. 13 is a schematic illustration of a system 1 for producing an annotated arthroscopic video streamaccording to an embodiment. The system 1 comprises a computing device 100, such as computer, comprising a memory 120 configured to store a computer-implemented keypoint identifying model 150 and, at least temporarily, store one or more arthroscopic video streams 140. The computing device 100 in Fig.13 has been shown with a single memory 120. The embodiments are, however, not limited thereto. In clear contrast, the computing device 100 could comprise or be, wirelessly or with wire, connected tomultiple memories 120, such as a memory system of multiple memories. The computing device 100 alsocomprises a processor 110. The computing device 100 in Fig.13 has been shown with a single processor130. The embodiments are, however, not limited thereto. In clear contrast, the computing device 100 could comprise or be, wirelessly or with wire, connected to multiple processors 130, such as one or moreprocessing circuitries. The computing device 100 further comprises a general input and output (I / O) unit130 configured to communicate with external devices, such as an arthroscope 160 and / or a display screen (not shown). The I / O unit 130 could represent a transmitter and receiver, or transceiver, configured to conduct wireless communication. Alternatively, or in addition, the I / O unit 130 could be configured to conduct wired communication and may then, for instance, comprise one or more input and / or output ports. Generally, in the training step, recorded arthroscopic video streams 140 are stored on a memory 120 of the system 1. These arthroscopic video streams 140 are then annotated as disclosed herein, such as with lens tracking information, geometric information from Structure-from-Motion (SfM) and manual annotations. The latter may include annotations of events, such as tools being brought into the field of view as well as the position of the tip of a hook, when present. As a consequence annotated arthroscopic video streams are produced that are used as training data for the computer-implemented keypoint identifying model 150. During use, when the system 1 runs against an arthroscopic video stream, e.g., a live or “on-line” arthroscopic video stream, the video frames of such an arthroscopic video stream are received by the I / O unit 130 and stored in the memory 120. These video frames are then input to trained computer- implemented keypoint identifying model 150 to thereby determine positions of graphics object, i.e., augment the video frames of the arthroscopic video stream. The system 1 as shown in Fig.13 could comprise one computing device 100 used to both provide the annotated arthroscopic video streams and train the CI keypoint identifying model 150 and then use the trained CI keypoint identifying model 150 for visual aid in arthroscopy. In another embodiment, the system100 comprises a first computing device to train the CI keypoint identifying model 150 based on providedannotated arthroscopic video streams. The trained CI keypoint identifying model 150 is then used by a second computing device to augment arthroscopic video streams obtained from an arthroscope 160. In such an embodiment, the first computing device could operate “off-line” to train the CI keypoint identifying model 150, whereas the second computing device operates “on-line” at the hospital or medical facility to assist in arthroscopy in real-time. An aspect of the invention therefore relates to a system 1 for producing an annotated arthroscopic video stream. The system 1 comprises at least one processor 110 and at least one memory 120 storing instructions that, when executed by the at least one processor 110, cause the at least one processor 110 to identify, for each video frame of a training arthroscopic video stream 140, a position of the lensprojection of an arthroscope 160 within the video frame. The at least one processor 110 is also causedto transform, for each video frame of the training arthroscopic video stream 140, the identified position ofthe lens projection of the arthroscope 160 into a canonical position having a lens projection distortiondefined by a precomputed distortion model 170. The at least one processor 110 is further caused to undistort, for each video frame of the training arthroscopic video stream 140, the lens projection of thearthroscope 160 based on the precomputed distortion model 170 to obtain a stream of canonicallypositioned and undistorted video frames. The at least one processor 110 is also caused to reconstruct a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream 140. The at least one processor 110 is further caused to select at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream 140 and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream 140 based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream 140. In an embodiment, the at least one processor 110 is caused to extract, for each video frame of the training arthroscopic video stream 140, candidate pixels indicating an edge of the arthroscope 160 within the video frame and iteratively estimate, for each video frame of the training arthroscopic video stream 140, the position of the lens projection of the arthroscope 160 based on the candidate pixels. In an embodiment, the at least one processor 110 is caused to apply, for each video frame of the training arthroscopic video stream 140, Canny edge detection on the video frame to extract the candidate pixels. In an embodiment, the at least one processor 110 is caused to iteratively estimate, for each video frame of the training arthroscopic video stream 140, the position of the lens projection of the arthroscope 160 by RANSAC by sampling candidate pixels into a candidate set, fitting a model for the lens projection of the arthroscope 160 to the candidate set, identifying a consensus set of candidate pixels that fit the model for the lens projection of the arthroscope 160 and adding the identified consensus set of candidate pixels to the candidate set and repeating the sampling, fitting and identifying operations at least once by sampling the candidate pixels from the consensus set. In an embodiment, the at least one processor 110 is caused to scale, shift and / or rotate, for each video frame of the training arthroscopic video stream 140, i) the identified position of the lens projection of the arthroscope 160 to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope 160. The at least one processor 110 is also caused to determine, for each video frame of the training arthroscopic video stream, the scaling, shift and / or rotation that transforms i) the identified position of the lens projection of the arthroscope 160 to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope 160. In an embodiment, the at least one processor 110 is caused to store, for each video frame of the training arthroscopic video stream 140, information of the scaling, shift and / or rotation in a memory 120. In an embodiment, the at least one processor 110 is caused to identify the respective position of the at least one keypoint in other video frames of the training arthroscopic video stream 140 based on information of the scaling, shifting and / or rotation and the precomputed distortion model 170. In an embodiment, the at least one processor 110 is caused to reconstruct the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections. In an embodiment, the at least one processor 110 is caused to reconstruct the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections to reconstruct a 3D representation in the form of a mesh of triangles or a point cloud representing the interior of the joint. In an embodiment, the at least one processor 110 is caused to estimate, for each canonically positioned and undistorted video frame of the stream of canonically positioned and undistorted video frames,arthroscope movement relative to the 3D representation by applying SfM on the stream of undistortedvideo frames with canonically positioned lens projections. In an embodiment, the at least one processor 110 is caused to generate the annotated arthroscopic video stream defining, for each video frame of the annotated arthroscopic video stream, a mesh covering themedial wall of the lateral condyle of femur within a perimeter coinciding with the cartilage boundary.In an embodiment, the at least one processor 110 is caused to annotate a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, in video frames of the annotated arthroscopic video stream, in which the tip of the marking instrument is visible. In an embodiment, the at least one processor 110 is caused to annotate a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur in video frames of the annotated arthroscopic video stream, in which the pilot hole is visible. In an embodiment, the at least one memory 120 stores a computer-implemented keypoint identifying model 150 and instructions that, when executed by the at least one processor 110, cause the at least one processor 110 to train the computer-implemented keypoint identifying model 150 to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint. In an embodiment, the at least one processor 110 is caused to train the computer-implemented keypoint identifying model 150 implemented by a keypoint regression network, preferably a keypoint R-CNN, to identify the at least one keypoint in video frames of the arthroscopic video stream based on the annotated arthroscopic video stream. In an embodiment, the at least one memory 120 stores instructions that, when executed by the at least one processor 110, cause the at least one processor 110 to identify at least one keypoint representing arespective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream basedon the trained computer-implemented keypoint identifying model 150 and augment the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint. In an embodiment, the at least one processor 110 is caused to estimate pixel positions of points in a mesh in the video frames of the arthroscopic video stream based on the trained computer-implemented keypoint identifying model 150. In an embodiment, the at least one processor 110 is caused to filter the pixel positions in time to suppress predictive noise, preferably using a Kalman filter. In an embodiment, the at least one processor 110 is caused to identify, for each video frame of the arthroscopic video stream, a position of the lens projection of the arthroscope 160 within the video frame.The at least one processor 110 is also caused to determine, for each video frame of the arthroscopicvideo stream, a transformation between the identified position of the lens projection of the arthroscope 160 and a canonical position having a lens projection distortion defined by a precomputed distortion model 170 and estimate the 3D coordinates of the pixel positions relative to the arthroscope 160 based on the pixel positions, information of the transformation and the precomputed distortion model 170. In an embodiment, the at least one processor 110 is caused to filter the 3D coordinates using a state space model, preferably an extended Kalman filter. In an embodiment, the at least one processor 110 is caused to annotate pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint selected from the group consisting of a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur, a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, a position of the arthroscope relative to the condyle, a position of a drill hole, and level curves representing distances to the cartilage border. In an embodiment, the at least one processor 110 is caused to annotate pixels in video frames of the arthroscopic video stream with at least one graphic object selected from the group consisting of contour of a virtual drill hole of a pre-configured drill diameter or radius, a level curve, a footprint of at least aportion of a surface of the condyle, and any combination thereof.In an embodiment, the system 1 also comprises the arthroscope 160. Fig.14 is a schematic block diagram of a device 200 configured to produce an annotated arthroscopic video stream according to an embodiment, such as computer. The device 200 comprises a processor210 and a memory 220 interconnected to each other to enable normal software execution. An I / O unit230 is preferably connected to the processor 210 and / or the memory 220 to enable reception of video frames of an arthroscopic video stream from an arthroscope. The term processor should be interpreted in a general sense as any circuitry, system or device capable of executing program code or computer program instructions to perform a particular processing, determining or computing task. The processing circuitry including one or more processors 110 is, thus, configured to perform, when executing a computer program 240, well-defined processing tasks such as those described herein. The processor 210 does not have to be dedicated to only execute the above-described steps, functions, procedure and / or blocks, but may also execute other tasks. In an embodiment, the computer program 240 comprises instructions, which when executed by a processor 210, cause the processor 210 to identify, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame, transform, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model, undistort, for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames, reconstruct a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream, select at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream, and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. In an embodiment, the computer program 240 comprises instructions, which when executed by a processor 210, cause the processor 210 to train a computer-implemented keypoint identifying model to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint. In an embodiment, the computer program 240 comprises instructions, which when executed by a processor 210, cause the processor 210 to identify at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on the trainedcomputer-implemented keypoint identifying model and augment the arthroscopic video stream byannotating pixels in video frames of the arthroscopic video streams corresponding to the at least one keypoint. The various embodiments discussed in the foregoing for the method in connection with Figs.5 to 12 and the system 1 in connection with Fig.13 also applies to the computer program embodiments. The proposed technology also provides a non-transitory computer-readable storage medium 250 comprising the computer program 240. By way of example, the software or computer program 240 may be realized as a computer program product, which is normally carried or stored on the non-transitorycomputer-readable medium 250, in particular a non-volatile medium. The non-transitory computer-readable medium 250 may include one or more removable or non-removable memory devices including, but not limited to a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc (CD), a Digital Versatile Disc (DVD), a Blu-ray disc, a Universal Serial Bus (USB) memory, a Hard Disk Drive (HDD) storage device, a flash memory, a magnetic tape, or any other conventional memory device. The computer program 240 may, thus, be loaded into the operating memory 220 of the computer for execution by the processor 210 thereof. An aspect of the invention relates to a non-transitory computer-readable medium 250 storing instructions 240 that, when executed by at least one processor 210, cause the at least one processor 210 to identify, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope within the video frame, transform, for each video frame of the training arthroscopic video stream, the identified position of the lens projection of the arthroscope into a canonical position having a lens projection distortion defined by a precomputed distortion model, undistort, for each video frame of the training arthroscopic video stream, the lens projection of the arthroscope based on the precomputed distortion model to obtain a stream of canonically positioned and undistorted video frames, reconstruct a 3D representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream, select at least one keypoint in or in vicinity of the joint in a video frame of thetraining arthroscopic video stream, and produce an annotated arthroscopic video stream by identifying arespective position of the at least one keypoint in other video frames of the training arthroscopic video stream based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream. In an embodiment, the non-transitory computer-readable medium 250 further stores instructions 240 that, when executed by the at least one processor 210, cause the at least one processor 210 to train a computer-implemented keypoint identifying model to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint. In an embodiment, the non-transitory computer-readable medium 250 further stores instructions 240 that, when executed by the at least one processor 210, cause the at least one processor 210 to identify at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on the trained computer-implemented keypoint identifying model, and augment the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video streams corresponding to the at least one keypoint. The various embodiments discussed in the foregoing for the method in connection with Figs.5 to 12 and the system 1 in connection with Fig.13 also applies to the non-transitory computer-readable medium embodiments. The embodiments described above are to be understood as a few illustrative examples of the present invention. It will be understood by those skilled in the art that various modifications, combinations and changes may be made to the embodiments without departing from the scope of the present invention. In particular, different part solutions in the different embodiments can be combined in other configurations, where technically possible. The scope of the present invention is, however, defined by the appended claims.

Claims

CLAIMS1. A method for producing an annotated arthroscopic video stream, the method comprising:identifying (S1), for each video frame of a training arthroscopic video stream (140), a position of the lens projection of an arthroscope (160) within the video frame; transforming (S2), for each video frame of the training arthroscopic video stream (140), the identified position of the lens projection of the arthroscope (160) into a canonical position having a lens projection distortion defined by a precomputed distortion model (170); undistorting (S3), for each video frame of the training arthroscopic video stream (140), the lens projection of the arthroscope (160) based on the precomputed distortion model (170) to obtain a stream of canonically positioned and undistorted video frames; reconstructing (S4) a three-dimensional (3D) representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream (140); selecting (S5) at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream (140); and producing (S6) an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream (140) based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream (140).

2. The method according to claim 1, wherein identifying (S1) the position of the lens projectioncomprises: extracting (S20), for each video frame of the training arthroscopic video stream (140), candidate pixels indicating an edge of the arthroscope (160) within the video frame; and iteratively estimating (S21), for each video frame of the training arthroscopic video stream (140), the position of the lens projection of the arthroscope (160) based on the candidate pixels.

3. The method according to claim 2, wherein extracting (S20) candidate pixels comprises applying,for each video frame of the training arthroscopic video stream (140), Canny edge detection on the video frame to extract the candidate pixels.

4. The method according to claim 2 or 3, wherein iteratively estimating (S21) the position comprisesiteratively estimating, for each video frame of the training arthroscopic video stream (140), the position of the lens projection of the arthroscope by random sample consensus (RANSAC) by:sampling (S22) candidate pixels into a candidate set; fitting (S23) a model for the lens projection of the arthroscope (160) to the candidate set; identifying (S24) a consensus set of candidate pixels that fit the model for the lens projection of the arthroscope (160) and adding the identified consensus set of candidate pixels to the candidate set; and repeating (L2) the sampling, fitting and identifying operations at least once by sampling the candidate pixels from the consensus set.

5. The method according to any one of claims 1 to 4, wherein transforming (S2) the identified positionof the lens projection of the arthroscope (160) comprises: determining (S30), for each video frame of the training arthroscopic video stream (140), a scaling, shift and / or rotation that transforms i) the identified position of the lens projection of the arthroscope (160)to correspond to the canonical position or ii) the canonical position to correspond to the identified positionof the lens projection of the arthroscope (160); and scaling, shifting and / or rotating (S31), for each video frame of the training arthroscopic video stream (140), i) the identified position of the lens projection of the arthroscope (160) to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope (160).

6. The method according to claim 5, further comprising storing (S32), for each video frame of thetraining arthroscopic video stream (140), information of the scaling, shift and / or rotation in a memory (120).

7. The method according to claim 5 or 6, wherein identifying a respective position of the at least onekeypoint comprises identifying the respective position of the at least one keypoint in other video frames of the training arthroscopic video stream (140) based on information of the scaling, shifting and / or rotation and the precomputed distortion model (170).

8. The method according to any one of claims 1 to 7, wherein reconstructing (S4) the 3Drepresentation comprises reconstructing (S4) the 3D representation by applying structure from motion (SfM) on the stream of undistorted video frames with canonically positioned lens projections.

9. The method according to claim 8, wherein reconstructing (S4) the 3D representation comprisesreconstructing (S4) the 3D representation by applying SfM on the stream of undistorted video frames withcanonically positioned lens projections to reconstruct a 3D representation in the form of a mesh of triangles or a point cloud representing the interior of the joint.

10. The method according to claim 8 or 9, further comprising estimating (S40), for each canonicallypositioned and undistorted video frame of the stream of canonically positioned and undistorted video frames, arthroscope movement relative to the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections.

11. The method according to any one of claims 1 to 10, wherein producing (S6) the annotatedarthroscopic video stream comprises generating the annotated arthroscopic video stream defining, for each video frame of the annotated arthroscopic video stream, a mesh covering the medial wall of the lateral condyle of femur within a perimeter coinciding with the cartilage boundary.

12. The method according to any one of claims 1 to 11, further comprising annotating (S7) a positionof a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, in video frames of the annotated arthroscopic video stream, in which the tip of the marking instrument is visible.

13. The method according to any one of claims 1 to 12, further comprising annotating (S7) a positionof a pilot hole representing an indentation made in a surface of the lateral condyle of femur in video frames of the annotated arthroscopic video stream, in which the pilot hole is visible.

14. A method for training a computer-implemented keypoint identifying model (150), the methodcomprising: producing an annotated arthroscopic video stream according to any one of claims 1 to 13 from a training arthroscopic video stream (140) of video frames; and training (S8) the computer-implemented keypoint identifying model (150) to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint.

15. The method according to claim 14, wherein training (S8) the computer-implemented keypointidentifying model (150) comprises training (S8) a computer-implemented keypoint identifying model (150)implemented by a keypoint regression network, preferably a keypoint region-based convolutional neuralnetwork (R-CNN), to identify the at least one keypoint in video frames of the arthroscopic video streambased on the annotated arthroscopic video stream.

16. A method for visual aid in arthroscopy, the method comprising:identifying (S9) at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on a computer-implemented keypoint identifying model (150) trained according to claim 14 or 15; and augmenting (S10) the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint.

17. The method according to claim 16, wherein identifying (S9) at least one keypoint comprisesestimating pixel positions of points in a mesh in the video frames of the arthroscopic video stream based on the computer-implemented keypoint identifying model (150) trained according to claim 14 or 15.

18. The method according to claim 17, further comprises filtering (S50) the pixel positions in time tosuppress predictive noise, preferably using a Kalman filter.

19. The method according to claim 17 or 18, further comprisingidentifying (S51), for each video frame of the arthroscopic video stream, a position of the lens projection of an arthroscope (160) within the video frame; determining (S52), for each video frame of the arthroscopic video stream, a transformation between the identified position of the lens projection of the arthroscope (160) and a canonical position having a lens projection distortion defined by a precomputed distortion model (170); and estimating (S53) 3D coordinates corresponding to the at least one keypoint relative to thearthroscope (160) based on the pixel positions, information of the transformation and the precomputed distortion model (170).

20. The method according to claim 19, further comprising filtering (S54) the 3D coordinates using astate space model, preferably an extended Kalman filter.

21. The method according to any one of claims 16 to 20, wherein augmenting (S10) the arthroscopicvideo stream comprises annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint selected from the group consisting of a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur, a position of a tip of a marking instrument,preferably a chondral pick or an arthroscopic probe, a position of the arthroscope relative to the condyle, a position of a drill hole, and level curves representing distances to the cartilage border.

22. The method according to any one of claims 16 to 21, wherein augmenting (S10) the arthroscopicvideo stream comprises annotating pixels in video frames of the arthroscopic video stream with at least one graphic object selected from the group consisting of contour of a virtual drill hole of a pre-configured drill diameter, a level curve, a footprint of at least a portion of a surface of the condyle and any combination thereof.

23. A non-transitory computer-readable medium (250) storing instructions (240) that, when executedby at least one processor (210), cause the at least one processor (210) to: identify, for each video frame of a training arthroscopic video stream, a position of the lens projection of an arthroscope (160) within the video frame; transform, for each video frame of the training arthroscopic video stream (140), the identified position of the lens projection of the arthroscope (160) into a canonical position having a lens projection distortion defined by a precomputed distortion model (170); undistort, for each video frame of the training arthroscopic video stream (140), the lens projection of the arthroscope (160) based on the precomputed distortion model (170) to obtain a stream of canonically positioned and undistorted video frames; reconstruct a three-dimensional (3D) representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream (140); select at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream (140); and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream (140) based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream (140).

24. The non-transitory computer-readable medium according to claim 23, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to: extract, for each video frame of the training arthroscopic video stream (140), candidate pixels indicating an edge of the arthroscope (160) within the video frame; anditeratively estimate, for each video frame of the training arthroscopic video stream (140), the position of the lens projection of the arthroscope (160) based on the candidate pixels.

25. The non-transitory computer-readable medium according to claim 24, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to apply, for each video frame of the training arthroscopic video stream (140), Canny edge detection on the video frame to extract the candidate pixels.

26. The non-transitory computer-readable medium according to claim 24 or 25, further storinginstructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to iteratively estimate, for each video frame of the training arthroscopic video stream (140), the position of the lens projection of the arthroscope (160) by random sample consensus (RANSAC) by: sampling candidate pixels into a candidate set; fitting a model for the lens projection of the arthroscope (160) to the candidate set; identifying a consensus set of candidate pixels that fit the model for the lens projection of the arthroscope (160) and adding the identified consensus set of candidate pixels to the candidate set; and repeating the sampling, fitting and identifying operations at least once by sampling the candidate pixels from the consensus set.

27. The non-transitory computer-readable medium according to any one of claims 23 to 26, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to: determine, for each video frame of the training arthroscopic video stream (140), a scaling, shift and / or rotation that transforms i) the identified position of the lens projection of the arthroscope (160) to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope (160); and scale, shift and / or rotate, for each video frame of the training arthroscopic video stream (140), i) the identified position of the lens projection of the arthroscope (160) to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope (160).

28. The non-transitory computer-readable medium according to claim 27, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) tostore, for each video frame of the training arthroscopic video stream (140), information of the scaling, shift and / or rotation in a memory (220).

29. The non-transitory computer-readable medium according to claim 27 or 28, further storinginstructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to identify the respective position of the at least one keypoint in other video frames of thetraining arthroscopic video stream (140) based on information of the scaling, shifting and / or rotation andthe precomputed distortion model (170).

30. The non-transitory computer-readable medium according to any one of claims 23 to 29, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least oneprocessor (210) to reconstruct the 3D representation by applying structure form motion (SfM) on thestream of undistorted video frames with canonically positioned lens projections.

31. The non-transitory computer-readable medium according to claim 30, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to reconstruct the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections to reconstruct a 3D representation in the form of a mesh oftriangles or a point cloud representing the interior of the joint.

32. The non-transitory computer-readable medium according to claim 30 or 31, further storinginstructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to estimate, for each canonically positioned and undistorted video frame of the stream of canonically positioned and undistorted video frames, arthroscope movement relative to the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections.

33. The non-transitory computer-readable medium according to any one of claims 23 to 32, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to generate the annotated arthroscopic video stream defining, for each video frame of the annotated arthroscopic video stream, a mesh covering the medial wall of the lateral condyle of femur within a perimeter coinciding with the cartilage boundary.

34. The non-transitory computer-readable medium according to any one of claims 23 to 33, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to annotate a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, in video frames of the annotated arthroscopic video stream, in which the tip of the marking instrument is visible.

35. The non-transitory computer-readable medium according to any one of claims 23 to 34, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least oneprocessor (210) to annotate a position of a pilot hole representing an indentation made in a surface ofthe lateral condyle of femur in video frames of the annotated arthroscopic video stream, in which the pilot hole is visible.

36. The non-transitory computer-readable medium according to any one of claims 23 to 35, furtherstoring instructions (240) that, when executed by the at least one processor (210), cause the at least one processor (210) to train a computer-implemented keypoint identifying model to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint.

37. The non-transitory computer-readable medium according to claim 36, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) totrain the computer-implemented keypoint identifying model (150) implemented by a keypoint regressionnetwork, preferably a keypoint R-CNN, to identify the at least one keypoint in video frames of the arthroscopic video stream based on the annotated arthroscopic video stream.

38. The non-transitory computer-readable medium according to claim 36 or 37, further storinginstructions (240) that, when executed by the at least one processor (210), cause the at least one processor (210) to: identify at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on the trained computer-implemented keypoint identifying model (150); and augment the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video streams corresponding to the at least one keypoint.

39. The non-transitory computer-readable medium according to claim 38, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to estimate pixel positions of points in a mesh in the video frames of the arthroscopic video stream based on the trained computer-implemented keypoint identifying model (150).

40. The non-transitory computer-readable medium according to claim 39, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to filter the pixel positions in time to suppress predictive noise, preferably using a Kalman filter.

41. The non-transitory computer-readable medium according to claim 39 or 40, further storinginstructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to: identify, for each video frame of the arthroscopic video stream, a position of the lens projection ofthe arthroscope (160) within the video frame;determine, for each video frame of the arthroscopic video stream, a transformation between theidentified position of the lens projection of the arthroscope (160) and a canonical position having a lensprojection distortion defined by a precomputed distortion model (170); andestimate the 3D coordinates of the pixel positions relative to the arthroscope (160) based on thepixel positions, information of the transformation and the precomputed distortion model (170).

42. The non-transitory computer-readable medium according to claim 41, further storing instructions(24) that, when executed by the at least one processor (210), cause the at least one processor (210) to filter the 3D coordinates using a state space model, preferably an extended Kalman filter.

43. The non-transitory computer-readable medium according to any one of claims 38 to 42, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least one processor (210) to annotate pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint selected from the group consisting of a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur, a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, a position of the arthroscope relative to the condyle, a position of a drill hole, and level curves representing distances to the cartilage border.

44. The non-transitory computer-readable medium according to any one of claims 38 to 43, furtherstoring instructions (24) that, when executed by the at least one processor (210), cause the at least oneprocessor (210) to annotate pixels in video frames of the arthroscopic video stream with at least one graphic object selected from the group consisting of contour of a virtual drill hole of a pre-configured drilldiameter or radius, a level curve, a footprint of at least a portion of a surface of the condyle, and anycombination thereof.

45. A system (1) for producing an annotated arthroscopic video stream, the system (1) comprising:at least one processor (110); and at least one memory (120) storing instructions that, when executed by the at least one processor (110), cause the at least one processor (110) to: identify, for each video frame of a training arthroscopic video stream (140), a position of the lens projection of an arthroscope (160) within the video frame; transform, for each video frame of the training arthroscopic video stream (140), the identified position of the lens projection of the arthroscope (160) into a canonical position having a lens projection distortion defined by a precomputed distortion model (170); undistort, for each video frame of the training arthroscopic video stream (140), the lensprojection of the arthroscope (160) based on the precomputed distortion model (170) to obtain a streamof canonically positioned and undistorted video frames; reconstruct a three-dimensional (3D) representation of an interior of a joint based on the stream of canonically positioned and undistorted video frames by associating points within the 3D representation to pixels in video frames of the training arthroscopic video stream (140); select at least one keypoint in or in vicinity of the joint in a video frame of the training arthroscopic video stream (140); and produce an annotated arthroscopic video stream by identifying a respective position of the at least one keypoint in other video frames of the training arthroscopic video stream (140) based on the association between points within the 3D representation and pixels in the video frames of the training arthroscopic video stream (140).

46. The system according to claim 45, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to: extract, for each video frame of the training arthroscopic video stream (140), candidate pixelsindicating an edge of the arthroscope (160) within the video frame; anditeratively estimate, for each video frame of the training arthroscopic video stream (140), theposition of the lens projection of the arthroscope (160) based on the candidate pixels.

47. The system according to claim 46, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to apply, for each video frame of the training arthroscopic video stream (140), Canny edge detection on the video frame to extract the candidate pixels.

48. The system according to claim 46 or 47, wherein the at least one memory (120) stores instructionsthat, when executed by the at least one processor (110), cause the at least one processor (110) to iteratively estimate, for each video frame of the training arthroscopic video stream 140, the position of thelens projection of the arthroscope 160 by random sample consensus (RANSAC) by:sampling candidate pixels into a candidate set; fitting a model for the lens projection of the arthroscope (160) to the candidate set; identifying a consensus set of candidate pixels that fit the model for the lens projection of thearthroscope (160) and adding the identified consensus set of candidate pixels to the candidate set; andrepeating the sampling, fitting and identifying operations at least once by sampling the candidate pixels from the consensus set.

49. The system according to any one of claims 45 to 48, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor (110) to: determine, for each video frame of the training arthroscopic video stream, a scaling, shift and / or rotation that transforms i) the identified position of the lens projection of the arthroscope (160) to correspond to the canonical position or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope (160); and scale, shift and / or rotate, for each video frame of the training arthroscopic video stream (140), i)the identified position of the lens projection of the arthroscope (160) to correspond to the canonicalposition or ii) the canonical position to correspond to the identified position of the lens projection of the arthroscope (160).

50. The system according to claim 49, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to store, for each video frame of the training arthroscopic video stream (140), information of the scaling, shift and / orrotation in the at least one memory (120).

51. The system according to claim 49 or 50, wherein the at least one memory (120) stores instructionsthat, when executed by the at least one processor (110), cause the at least one processor (110) to identify the respective position of the at least one keypoint in other video frames of the training arthroscopic videostream (140) based on information of the scaling, shifting and / or rotation and the precomputed distortionmodel (170).

52. The system according to any one of claims 45 to 51, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor(110) to reconstruct the 3D representation by applying structure from motion (SfM) on the stream ofundistorted video frames with canonically positioned lens projections.

53. The system according to claim 52, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to reconstruct the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections to reconstruct a 3D representation in the form of a mesh of triangles or a point cloud representing the interior of the joint.

54. The system according to claim 52 or 53, wherein the at least one memory (120) stores instructionsthat, when executed by the at least one processor (110), cause the at least one processor (110) to estimate, for each canonically positioned and undistorted video frame of the stream of canonically positioned and undistorted video frames, arthroscope movement relative to the 3D representation by applying SfM on the stream of undistorted video frames with canonically positioned lens projections.

55. The system according to any one of claims 45 to 54, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor (110) to generate the annotated arthroscopic video stream defining, for each video frame of the annotated arthroscopic video stream, a mesh covering the medial wall of the lateral condyle of femur within a perimeter coinciding with the cartilage boundary.

56. The system according to any one of claims 45 to 55, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor (110) to annotate a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, in video frames of the annotated arthroscopic video stream, in which the tip of the marking instrument is visible.

57. The system according to any one of claims 45 to 56, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor (110) to annotate a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur in video frames of the annotated arthroscopic video stream, in which the pilot hole is visible.

58. The system according to any one of claims 45 to 57, wherein the at least one memory (120) storesa computer-implemented keypoint identifying model (150) and instructions that, when executed by the at least one processor (110), cause the at least one processor (110) to train the computer-implemented keypoint identifying model (150) to identify at least one keypoint in video frames of an arthroscopic video stream based on the annotated arthroscopic video stream, wherein the at least one keypoint represents a respective point of interest in or in vicinity of a joint.

59. The system according to claim 58, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to train thecomputer-implemented keypoint identifying model (150) implemented by a keypoint regression network,preferably a keypoint R-CNN, to identify the at least one keypoint in video frames of the arthroscopic video stream based on the annotated arthroscopic video stream.

60. The system according to claim 58 or 59, wherein the at least one memory (120) stores instructionsthat, when executed by the at least one processor (110), cause the at least one processor (110) to: identify at least one keypoint representing a respective point of interest in or in vicinity of a joint in video frames of an arthroscopic video stream based on the trained computer-implemented keypoint identifying model (150); and augment the arthroscopic video stream by annotating pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint.

61. The system according to claim 60, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) estimate pixelpositions of points in a mesh in the video frames of the arthroscopic video stream based on the trained computer-implemented keypoint identifying model (150).

62. The system according to claim 61, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to filter the pixel positions in time to suppress predictive noise, preferably using a Kalman filter.

63. The system according to claim 61 or 62, wherein the at least one memory (120) stores instructionsthat, when executed by the at least one processor (110), cause the at least one processor (110) to: identify, for each video frame of the arthroscopic video stream, a position of the lens projection ofthe arthroscope (160) within the video frame;determine, for each video frame of the arthroscopic video stream, a transformation between theidentified position of the lens projection of the arthroscope (160) and a canonical position having a lensprojection distortion defined by a precomputed distortion model (170); andestimate the 3D coordinates of the pixel positions relative to the arthroscope (160) based on thepixel positions, information of the transformation and the precomputed distortion model (170).

64. The system according to claim 63, wherein the at least one memory (120) stores instructions that,when executed by the at least one processor (110), cause the at least one processor (110) to filter the 3D coordinates using a state space model, preferably an extended Kalman filter.

65. The system according to any one of claims 60 to 64, wherein the at least one memory (120) storesinstructions that, when executed by the at least one processor (110), cause the at least one processor (110) to annotate pixels in video frames of the arthroscopic video stream corresponding to the at least one keypoint selected from the group consisting of a position of a pilot hole representing an indentation made in a surface of the lateral condyle of femur, a position of a tip of a marking instrument, preferably a chondral pick or an arthroscopic probe, a position of the arthroscope relative to the condyle, a position of a drill hole, and level curves representing distances to the cartilage border.

66. The system according to any one of claims 60 to 65, wherein the at least one memory (120)stores instructions that, when executed by the at least one processor (110), cause the at least one processor (110) to annotate pixels in video frames of the arthroscopic video stream with at least one graphic object selected from the group consisting of contour of a virtual drill hole of a pre-configured drill diameter or radius, a level curve, a footprint of at least a portion of a surface of the condyle, and any combination thereof.

67. The system according to any one of claims 45 to 66, further comprising the arthroscope (160).

Citation Information

Patent Citations

  • Methods and systems for computer-aided surgery using intra-operative video acquired by a free moving camera

    US10499996B2

  • Methods for Autoregistration of Arthroscopic Video Images to Preoperative Models and Devices Thereof

    US20230200928A1

  • Probes, systems, and methods for computer-assisted landmark or fiducial placement in medical images

    US20230263573A1