Image processing systems and methods

A markerless image processing system using RGB-D cameras and neural networks addresses the lack of quantitative analysis in patellar tracking by employing synthetic data augmentation, achieving precise and dynamic tracking in surgical environments.

WO2026096769A1PCT designated stage Publication Date: 2026-05-07SMITH & NEPHEW INC +2
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SMITH & NEPHEW INC
Filing Date
2025-10-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current methods for measuring patellar tracking in total knee arthroplasty lack quantitative and dynamic analyses, relying on markers or fiducials that are not robust in surgical environments with factors like bleeding and surgical lighting, and lack sufficient training data for depth imaging in surgical scenarios.

Method used

A markerless image processing system using a combination of RGB-D cameras and neural networks for robust region of interest localization and segmentation, trained with synthetic data augmentation to handle occlusions and variations in surgical scenes, employing techniques like domain randomization and data augmentation to enhance network performance.

Benefits of technology

Provides accurate and dynamic tracking of patellar movement, overcoming environmental challenges with high precision and reducing the need for physical training data, enabling effective surgical navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025053344_07052026_PF_FP_ABST
    Figure US2025053344_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Examples relate to an image process system for markerless patella-femoral joint identification, the system comprising: a. a machine learning interface to a multiclass classification deep learning model; the machine learning interface being arranged to receive an input vector; the input vector comprising at least one, or both, of: image information and depth information associated with a patella-femoral joint; the multiclass classification deep learning model being trained to generate multiclass classification data; the multiclass classification data comprising at least: i. semantically segmented image data comprising at least one mask corresponding to a respective at least one member (distal end of femur, proximal end of tibia, patella) of the knee joint one of which being the patella, and b. a machine learning output interface of the multiclass classification deep learning model; the machine learning output interface being arranged to output the multiclass classification data.
Need to check novelty before this filing date? Find Prior Art

Description

IMAGE PROCESSING SYSTEMS AND METHODSTECHNICAL FIELD

[0001] The present application generally relates to image processing systems and methods such as, for example, image processing systems and methods for markerlessly identifying body parts using Artificial Intelligence (Al).BACKGROUND

[0002] Patellar maltracking and instability can play a role in patellofemoral pain and produce poor outcomes including anterior knee pain in total knee arthroplasty (TKA) Patellar tracking is defined as the movement of the patella relative to the patellofemoral joint through flexion and extension and can be influenced by patellar tilt, subluxation or complete dislocation. The quest for “optimal” patellar tracking continues as it can be affected by many factors such as surgical technique, implant selection, and / or resurfaced and unresurfaced patella. Most of the methods used to measure patellar tracking consist of computed tomography, nuclear magnetic resonance imaging, an infrared tracking system, and a fluorescence capture system. However, abnormalities and variations in patellar tracking remain unclear due to the lack of quantitative and dynamic analyses focused on patella tracking currently studied.

[0003] US20190388159A1 describes a method for selecting a properly sized patellar implant using a surgical system for patellar tracking. The surgical system uses markers to collect patella and bone range of motion. CN101484085B also uses tracking attachments to the patella and femur to help determine the medial-lateral position of the patella implant during a Total Knee Replacement (TKR). US8571637B2 uses an apparatus to track the patella comprising a frame fixed in relation to the patella along with markers that are detectable by a surgical navigation system through the patella’s range of motion. US20210315640A1 uses a computing system to determine the position of the posterior apex of the patella by characterizing the anterior geometry of the patella. The computing system receives location information of the patella relative to the knee joint and provides recommendations to improve patella-femoral response. However, the foregoing use markers or fiducials for tracking bodily parts.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] By way of illustration, examples of image processing systems, methods and machine-readable storage for markerless tracking of body parts will now be described, with reference to the accompanying drawings, in which:

[0005] FIG. 1 shows a view of an image processing system according to examples;

[0006] FIG. 2 depicts a view of a region of interest neural network for identifying a region of interest within an image clipping according to examples;

[0007] FIG. 3 illustrates shows a view of the architecture of the region of interest neural network according to examples;

[0008] FIG. 4 depicts a view of an image segmentation neural network according to examples;

[0009] FIG. 5A shows a view of the architecture of the image segmentation neural network according to examples;

[0010] FIG. 5B shows a view of segmentation images according to examples;

[0011] FIG. 5C shows a view of an overall view of processing according to examples;

[0012] FIG. 5D shows a comparison of a marker based registration process and a markerless based registration process according to examples;

[0013] FIG. 6 illustrates a view of a flowchart for generating ground truth images according to examples;

[0014] FIG. 7A shows a view of a flowchart for image segmentation according to examples;

[0015] FIG. 7B depicts a view of a flowchart for determining a region of interest according to examples;

[0016] FIG. 7C is a view of a flowchart for producing labelled ground truth data according to examples;

[0017] FIG. 8 illustrates a view of a computer assisted surgical system using image processing according to examples; and

[0018] FIGs. 9 to 13 depict views of machine-readable storage according to examples.DETAILED DESCRIPTION

[0019] FIG. 1 shows a view 100 of an image processing system 102 according to examples. The image processing system 102 comprises an RGB Depth (RGB D) camera node 104, a region of interest localisation network node 106, and an image segmentation neural network node 108 arranged sequentially. The image processing system 102 can also comprise a visualisation node 110 also arranged sequentially.

[0020] The RGB D camera node 104 is arranged to receive RGB D image data 112 from an RGB D camera 114. The RGB D image data 112 can comprise an RGB image, which is an example of an original image, and a depth image, which is an example of an original depth image comprising depth data. The reason for using both colour and depth information is twofold. Firstly, depth imaging is still not as developed and stable as RGB imaging. Therefore, in a depth image, there is a risk that a considerable number of pixels might not have valid values (i.e. their positions cannot be measured), which makes it difficult to globally localise the position of the knee according to such a ‘broken’ depth image. Furthermore, there is an exponentially increasing measurement noise with respect to distance that also diminishes the usability of the whole depth image. Secondly, RGB imaging is able to provide a robust Region of Interest (Rol) estimate from the whole scene, but in the smaller scope of the surgical site, the high brightness of the surgical light may compromise the colour features that are useful for segmentation. Moreover, bleeding at the surgical site can also complicate the colour conditions, whereas depth imaging is hardly affected by bleeding. Therefore, a stable RGB stream ensures robust target localisation in the full scope of captures, while the depth data ensure fine segmentation, as the depth data are less impacted by bleeding and surgical lighting. Accordingly, a combination of RGB and Depth data results in an improved segmented image containing any detected predetermined bodily structures. Bodily structures can also be referred to an anatomical structures.

[0021] In the example depicted in figure 1 , the RGB D image data 112 comprises RGB images together with depth data. The depth data provides an indication of the distance of a given point of an object from the RGB D camera 114. The object in the present example is a knee 116. The knee 116 bears an open incision 118 that reveals the femur 120 andthe tibia 122 in the view shown. However, examples can be realised in which the incision reveals the patella. Examples can be realised that reveal the patella, but that leave the retinaculum intact. A depth image is a map describing the spatial geometry of the scene. Like RGB images, a depth image is also a matrix of pixels, each of which contains three values. Instead of representing colour, the values of each pixel in a depth image are the x, y and z coordinates of that point relative to the depth camera. As depth images and RGB images share the same data structure, the architecture of artificial neural networks that perform well on RGB images can also be utilised for depth image processing. However, very few studies apply depth imaging to surgical scenarios, and so there are no labelled datasets of surgical depth images available for segmentation training. Collecting and labelling a large number of training data is highly tedious and time-consuming in complex surgical scenes, (e.g., where the target is surrounded by blood and tissues). Therefore, to address these shortcomings, examples can be realised in which a large dataset with occlusion instances can be created to train a network that works within an intraoperative scenario. To expand a given training data in a fast and efficient way, examples can generate synthetic data via randomized scenes using a modular procedural pipeline, such as, for example, BlenderProc, on an opensource rendering platform, such as, for example, Blender. Existing real data containing no occlusion instances can be augmented with synthetic RGBD data containing various simulated target interactions. Advantageously, by utilising both 2D RGB images and 3D point clouds converted from depth frames, the neural networks of examples successfully learn to be robust to occlusions from synthetic data only, which also generalises well to future generations of RGB-D cameras and knee targets. Domain randomisation can be beneficial in overcoming the simulation-to-reality gap in RGB data. Therefore, examples can be realised in which the training data was processed or augmented to alter, randomly or systematically, the scene during image generation in respect of a set of factors. Examples can be realised in which the set of factors comprises one or more than one of the following taken jointly and severally in any and all permutations: (a) the type (point or surface) and strength of lighting, (b) the room background, which contains arbitrary extrusions and objects loaded from the Ikea assembly dataset as distractors, the materials of the wall, floor and loaded objects are randomly sampled from a large public material database, ambientCG, and (c) the material of skin and exposed bone, by blending a random texture with a random RGBcolour. Examples simulated foreground occlusion using 3D models of human hands and surgical tools, which were prepared and imported as foreground distractors. These objects are randomly positioned and orientated within the camera’s line-of-sight of the exposed knee joint to simulate partial target occlusion. The fingertip or tooltip, defined as the origin of local object coordinates, can optionally contact and be translated onto the exposed anatomy. The material of these objects is also altered using a random texture blending method.

[0022] Aso, depth sampling noise arising from pixel location offsets due to quantised disparity can be added, which is challenging given that the material- and illuminationdependent interaction cannot be physically simulated. Furthermore, dropout density is subject to specific camera properties like spatial sampling resolution and the baseline distance between projector and receiver. For each image-generation session with a settled scene, at least 20 captures are taken with random camera poses. The viewpoint is controlled to be 0.8-1 m away from the target to replicate a typical physical working distance. The sampling intrinsic parameters and resolution are set to the physical values of the SpryTrack300 camera calibrated by a standard routine. The visibility of the exposed femur, tibia and patella is checked for each sampling pose to ensure a meaningful capture. The simulation is repeated to produce at least 10,000 randomised synthetic RGB D images together with respective (automatically) labelled masks.

[0023] Examples can be realised in which the synthetic training data can be further improved for better network performance without the need for network retraining. An imported knee model is currently considered as a rigid body with a fixed surgical exposure. By modelling the skin part as a non-rigid body controlled by respective nodes, various extents of skin exposure could be included in the synthetic images. Including more target geometries in a simulation enriches the generated data.

[0024] The RGB D camera node 104 is arranged to input the RGB D image data 112, in the form of an input vector (not shown), to the region of interest localisation neural network node 106. Commercial depth cameras, e.g. Atracsys SpryTrack 300, Intel RealSense™ normally have a wide field of view that captures a large portion of the environment. Therefore, only a small part of the image is associated with the target bodily structure. To decrease the size of the segmentation network and potentially improve its accuracy, examples can be realised in which a localisation network is used that utilises the RGBinformation to estimate the ROI position, and relative to which the depth image can be cropped to remove most of the background.

[0025] The region of interest localisation neural network node 106 is arranged to provide access to a trained region of interest neural network 124. The trained region of interest neural network 124 is configured to clip, or otherwise identify, a region of interest 126 from the RGB frame according to which the aligned depth frame is cropped and resampled into a 3D point cloud. In the example depicted, the region of interest 126 encompasses the region of the knee 116 bearing the open incision 118.

[0026] The training data for the region of interest localisation neural network 124 comprises a set of RGB images and a set of corresponding segmented images, that is, a set of mask images comprising masks corresponding to a feature or features of interest in the RGB images, as well as a set of depth images and a set of corresponding segmented depth images, that is, a set of depth mask images comprising masks corresponding to the same feature or features of interest in the depth images.

[0027] Although the region of interest neural network 124 has been depicted in figure 1 as being accessible to the region of interest localisation neural network node 106, examples are not limited to such an arrangement. Examples can be realised in which the region of interest localisation neural network 124 forms part of the region of interest localisation neural network node 106 and, therefore, part of the overall system 102.

[0028] The region of interest localisation network node 106 is arranged to forward clipped image data (not shown) associated with the region of interest 126 of the RGB D image data 112 to the image segmentation neural network node 108. The image segmentation neural network node 108 is arranged to input the clipped image data, in the form of an input vector (not shown), to an image segmentation neural network 128 that produces a number of segmentation outputs (Nf). The segmentation outputs can be used, for instance, for pose estimation and / or pose registration. The number of segmentation outputs (Nf) can depend upon the number of features to be identified or segmented within the input image. For instance, in a simple case, an example can be realised in which the segmented image of the neural network 128 is arranged to identify markers associated with bodily structures or to identify the bodily structures per se. The image segmentation neural network 128 is configured to segment the clipped image data to identify a set ofbodily structures within the RGB D image data 112. The image segmentation neural network 128 is an example of a multiclass classification deep learning model. Given that high speed segmentation is desirable for real-time tracking, only a certain number of resampled points (N) can be taken by the segmentation network. Consequently, ROI cropping ensures a high target occupation rate Nf / N. Furthermore, if the camera moves towards or away from the target pose, or the network is redeployed to a newer version of the camera with a considerably different focal length, a fixed cropping size may fail to cover the whole target dimensions. Therefore, examples can be realised in which ROI cropping can be dynamic in size. Examples can be realised in which such dynamically sized ROI cropping supports ensuring a nearly constant value for the target occupancy rate Nf / N.

[0029] Examples can be realised in which the set of bodily structures comprises a set of bodily structures associated with the knee 116. However, examples are not limited to a set of bodily structures associated with the knee. Examples can be realised in which the set of bodily structures can be any selected bodily structures subject to the ROI network and / or the image segmentation neural network being trained with suitable original image : segmentation image pairs. For instance, the set of bodily structures can comprise at least one, or more than one, of the following taken jointly and severally in any and all permutations: a patella (not shown), the femur 120 and the tibia 122. The image segmentation neural network node 108 is arranged to forward data (not shown) associated with the clipped image data and the set of bodily structures / masks are predicted from the cloud for point-wise segmentation to be displayed within the visualisation node 110.

[0030] The visualisation node 110 is arranged to output, on a display 130, a set of images 132 derived from the RGB D image data 112. Examples can be realised in which the set of images 132 comprises at least one image. Examples can be realised in which the set of images 132 comprises a plurality of images. In the example depicted in figure 1 , the set of images 132 comprises a first image 134 and a second image 136. The first image 134 has a respective segment 138 identified by the image segmentation neural network 128. The respective segment 138 is associated with the femur 120. The second image 136 also has a respective segment 140 identified by the image segmentation neural network 128. The respective segment 140 is associated with the tibia 120.

[0031] Although the image segmentation neural network 128 has been shown in figure 1 as being accessible to the image segmentation node 108, examples are not limited to such an arrangement. Examples can be realised in which the image segmentation neural network 128 forms part of the image segmentation neural network node 108 and, therefore, part of the overall system 102.

[0032] Examples can be realised in which the system 102 additionally comprises, or has access to, a registration node 142. The registration node 142 is arranged to maintain a model (not shown) of the knee 116 based upon the set of segmented images 132.

[0033] To assess the accuracy of the registration between the segmented femur, tibia and patella masks and corresponding reference models, a SpyTrack300 optical tracker (Atracsys LLC, Switzerland) was used to obtain gold standard reference measurements, by tracking two reference frames rigidly attached to a handpiece (M1), distal femur, proximal tibia and patella (M2) respectively. The SpyTrack300 camera consisted of an RGB-D and IR sensor in a common reference frame, which negates needing to attach an additional external optical marker to the camera. Examples can be realised in which the optical tracker was considered as the world frame W. After the entire exposed femur, tibia and patella surfaces were scanned using a digitized probe from a NAVIO system (Smith & Nephew Inc., Pittsburgh), a point cloud of the femur surface Pf in the femur marker frame was defined, which would be used to label the femur points in the depth images. The depth camera was rigidly attached to the arm of a surgical lighting system either directly above or oblique to the specimen at a height of approximately 1m. The camera was then used to simultaneously capture depth and RGB images of a cadaveric knee joint at approximately 50Hz. A robotic handpiece of the NAVIO™ surgical system (Smith & Nephew inc.) has a rigidly attached reference frame with optical markers Mi for reliable tool tracking. The handpiece probe can be used to digitise the surface of the femur, tibia and patella by collecting point clouds and then creating a series of custom digitised models “reference model” using bone morphing 3D statistical shape modelling software. The “landmark" tracked in the depth camera frame D can then be transformed to the static world space W at a time t, as the global landmark position. A ground truth pose, for example, for a femur, can thus be expressed as:

[0035] where ^T(t) andarethe optically tracked poses of the reference frames, is the initial femur pose registered in the marker frame, andis the relative static pose between the reference frame of the RGB D camera and the rigid marker Mi. These two matrices are unaffected by the movement of the handpiece and thus only need to be calibrated once. Examples can be realised in which more than 2,000 depth images of the knee are collected, together with the same number of RGB images. The points belonging to the femur, tibia and patella in the depth images can be labelled by finding the matching points of Pc. The Pcpoints are a set of depth points that are transformed into the camera frame for each depth image. Examples can be realised in which the Pcpoints comprise one or more than one of the following taken jointly and severally in any and all permutations: a set of femur depth points, a set of tibia depth points, and a set of patella points. However, due to errors existing in the camera-marker calibration, the transformed reference points, Pcdo not perfectly overlap the anatomical features in the corresponding depth images. Therefore, a standard iterative closest point (ICP) algorithm can be used to align Pcwith the depth images, such that the overlapped points in the depth image can be labelled as the surface points belonging to a femur, a tibia or a patella for segmentation training. The ICP algorithm principle is reflected in the following: Given a reference point set P and a data point set Q (at a provided initial estimate R, T), the ICP algorithm finds the corresponding nearest point in P for each point in Q to form a matching point pair. Then, the ICP algorithm uses the sum of the Euclidean distances of all matching point pairs as the value of the error objective function error, utilizes singular value decomposition (SVD) to find R and T to minimize the error, rotates Q according to R and T, and finds the corresponding point pairs again. Further details on the ICP algorithm can be found in, for example, Besl, Paul J.; N.D. McKay (1992). "A Method for Registration of 3-D Shapes". IEEE Transactions on Pattern Analysis and Machine Intelligence. 14 (2): 239-256.

[0036] Having labelled the points in the depth images, the dataset is augmented in order to improve training performance. First, a depth image can be cropped to a square shape of predetermined dimensions nx n. Examples can be realised in which the predetermined size is 160 x 160 around the centre of labelled points, with the cropping centre being recorded as the label of the corresponding RGB image for ROI localisation training.Subsequently, each image in the set of depth images is flipped to increase the size of the dataset. In the depth images, pixel position (row, column) and pixel value (x, y, z) are correlated because a pixel in the image is projected from a physical point according to its spatial position. Thus, depth images cannot be augmented by simply flipping the pixel positions, which means that the pixel values also need to be modified. According to the coordinate system of the depth camera, the x-axis points to the right and the y-axis points downwards, so the x values of the pixels change sign if the depth image is flipped horizontally, and y values change sign if the flip is vertical. In this way, a left knee geometry can be produced from a right knee geometry and vice versa. Rotation is also used to augment the dataset, and for the same reasons as above, the points in the depth image are rotated around the z-axis of the depth camera (pointing forward) by ±90° before the pixels were rotated (anti-)clockwise. For each depth image, the pixel values are multiplied by a random scalar (arbitrarily set between 0.9 and 1.1) to represent different knee sizes. After data augmentation, a dataset of over 10,000 labelled depth images can be obtained, which was shuffled and divided into three groups with the ratio of 6:2:2 for network training, validation and testing.

[0037] Examples of the image segmentation neural network 128 can be realised using, for instance, UNet as will be described below with reference to figures 5A and 5B. Examples of the region of interest neural network 124 can be realised using, for instance, Alexnet or ROINet. Although examples can be realised in which the segmentation network uses UNet, other CNNs can be used such as, for example, VGG, more particularly, VGGNet-19 or VGG16..

[0038] FIG. 2 depicts a view 200 of the region of interest neural network node 106 for image clipping according to examples. The region of interest neural network 124 is trained using a set of training data 202. The set of training data 202 comprises supervised learning training data in the form of a number of images 204 to 208; each of which depicts a knee 116 bearing open incision 118 showing the set of bodily structures as described above. The images 204 to 208, or at least a subset thereof, have been annotated, or otherwise labelled, to indicate which pixels are associated with which respective anatomical structures of the set of bodily structures. The set of bodily structures of interest can be defined using a bounding box 209. The bounding box defines the region of interest (Rol).The set of bodily structures can comprise one or more than one bodily structure. Therefore, the bounding box can be associated with a single bodily structure or with multiple bodily structures. The bounding box will have a respective set of characteristics. The set of respective characteristics can comprise, for example, height and width dimensions of the bounding box and position data relating to the bounding box, such as, for instance, the coordinates of the centre, or some other reference, of the bounding box. For instance, a first image 204 is shown as having been annotated with the bounding box 209 to identify or contain the open incision showing the femur and the tibia. Similarly, the set of images 202 can comprise images that have been annotated with respective bounding boxes to identify, for example, the femur. Still further, the set of images 202 can comprise images that have been annotated with respective bounding boxes to identify, for example, the patella. Examples can be realised in which the images in the set 202 have been annotated with respective bounding boxes associated with at least one or more than one of the following taken jointly and severally in any and all permutations: the incision, the knee joint, the tibia, the femur, and the patella. The region of interest neural network 124 is trained to produce the above-described clipped images or at least to produce data defining one or more than one respective bounding box containing at least one or more than one anatomical feature of interest. In the example shown in figure 2, the region of interest neural network 124 is arranged to produce a set 210 of clipped images. The set of clipped images 210 is derived from the above described input vector, which is depicted as input vector 214. The input vector 214 comprises the RGB D image data set 112. The RGB D image data set 112 comprises the RGB colour planes 216 to 220 together with corresponding depth data 222. Although examples have been described in which RGB D data is used to determine the ROI and / or to produce clipped or cropped images, examples are not limited to such an arrangement. Examples can be realised in which RGB data, without the depth data, is used to determine the ROI and / or to produce the clipped or cropped images.

[0039] Therefore, the set of images 210 comprises clipped RGB images 224 to 228 that are derived from respective colour plane images 216 to 220. The region of interest neural network 124 can also produce clipped depth data 230 from the depth data 222 of the input vector 214 such that the set of images 210 comprises both clipped RGB images 224 to 228 as well as the clipped images 230.

[0040] FIG. 3 shows a view 300 of the architecture 302 of the region of interest neural network 124 according to examples. The architecture 302 comprises a plurality of convolutional layers for extracting features from the RGB image. The RGB input image or first feature map 304 is subjected to a first convolutional operation via a first convolutional operator followed by a first batch normalisation operation 308. The first convolutional operator is an 11x11 kernel with a stride length of 4, which results in a second feature map 310 having dimensions 88x158 with 32 channels.

[0041] The second feature map 310 is subjected to a MaxPooling operator 312 having dimensions 3x3 with a stride of 2, which results in a third feature map 314.

[0042] The third feature map 314 is subjected to a second convolutional operation via a second convolutional operator followed by a second batch normalisation operation 316. The second convolutional operator is a 5x5 kernel with a stride length of 1, which results in a fourth feature map 318 having dimensions of 43x78 and 32 channels.

[0043] The fourth feature map 318 is subjected to a second MaxPooling operator 320 having dimensions 3x3 with a stride length of 2, which results in a fifth feature map 322 having dimensions 21x38 with 64 channels.

[0044] The fifth feature map 322 is subjected to a third convolutional operation via a third convolutional operator followed by a third batch normalisation operation 324. The third convolutional operator is a 3x3 kernel with a stride length of 1 , which results in a sixth feature map 326 having dimensions of 21x38 and 64 channels.

[0045] The sixth feature map 326 is subjected to a fourth convolutional operation via a fourth convolutional operator followed by a fourth batch normalisation operation 328. The fourth convolutional operator is a 3x3 kernel with a stride length of 1, which results in a seventh feature map 330 having dimensions of 21x38 and 64 channels.

[0046] The seventh feature map 330 is subjected to a fifth convolutional operation via a fifth convolutional operator followed by a fifth batch normalisation operation 332. The fifth convolutional operator is a 3x3 kernel with a stride length of 1 , which results in an eighth feature map 334 having dimensions of 21x38 and 32 channels.

[0047] The eighth feature map 334 is subjected to sixth parallel convolutional operation via a pair of convolutional operators followed by a respective pair of batch normalisationoperations 336. The pair of convolutional operators comprise a 1x1 kernel with a stride length of 1 and a 3x3 kernel with a stride length of 2, which results in a ninth feature map 338 having dimensions of 11x19 and 64 channels.

[0048] The sixth feature map 326 is also subjected to a further pair of convolution operations via a respective pair of convolutional operators 340. A first convolutional operator of the pair 340 has 3x3 kernel with a stride length of 1 , which results in a single channel that is subjected to a respective activation function, which can be the sigmoid activation function. A second convolutional operator of the pair 340 has a 3x3 kernel, with a stride length of 3 which results in 3 channels.

[0049] The ninth feature map 338 is also subjected to a still further pair of convolution operations via a respective pair of convolutional operators 342. A first convolutional operator of the pair 342 has 3x3 kernel with a stride length of 1 , which results in a single channel that is subjected to a respective activation function, which can be the sigmoid activation function. A second convolutional operator of the pair 342 has a 3x3 kernel, with a stride length of 3 which results in 3 channels.

[0050] The outputs of the two pairs of convolutional operators 340 and 342 are subjected to a softmax operation via a softmax function 344. The softmax function produces a number, M, of bounding boxes each having an associated probability, f , where 0 < / j- < 1, of containing a respective target object of a set of possible target objects. The set of possible target objects comprises the femur, the tibia, the patella, or background. Each bounding box has respective dimensions and a respective centre. Examples can be realised in which M=21x38+11x19.

[0051] The softmax function outputs 345 are used to position a set of up to M bounding boxes on the original input image 304 to create a final feature map 346 containing bounding boxes surrounding respective target objects each having a respective probability or confidence measure with which the target has been identified using Non Max Suppression to select the best or most likely bounding box of a plurality of identified bounding boxes for each respective target object.

[0052] An alternative ROI Neural Network Architecture can be realised according to the following table.

[0053] It will be appreciated that a 1 * 1 convolutional layer is used to compress the feature map to one channel and then normalise it instead of using fully connected layers at the end of the network for classification, which ensures that the spatial information of the features are preserved. The value of each element in the compressed map represents the probability of the pixel in that position belonging to the ROI. The compressed map is then multiplied elementwise by a pre-defined position weight map to calculate the ROI position. The position weight map has the same size as the final feature map, and each cell in the map has two values representing the relative position of that cell in terms of the row and the column of that cell.

[0054] It can be appreciated from the table above, that input images are cropped to a predetermined size of (360 x 640) to fit the region of interest neural network 124. The image data set is augmented to increase the amount of image data. Augmenting the data set can comprise one or multiple operations such as any of the following taken jointly and severally in any and all permutations: the images can be flipped vertically, flipped horizontally, have the brightness adjusted or have the saturation adjusted to enlarge the dataset. Furthermore, batch normalisation is used after each convolutional layer to facilitate network training.

[0055] Unlike ReLU, which gates inputs by their sign, the Gaussian Error Linear Unit (GELU) activation function weights inputs by their percentiles. This means that GELU allows small negative values when the input is less than zero, providing a richer gradient for backpropagation. GELU is often described as a smoother version of ReLU

[0056] In an alternative embodiment, the ROINet 124 can be modified by adding two midlayer auxiliaries and a multi-box loss function. Accordingly, with a similar design to Alexnet, examples can be realised in which the first five convolutional layers extract feature maps with shrinking sizes from the input RGB image. Similar to a Single Shot Multibox Detector (SSD), examples can be realised in which M multi-scale feature maps are taken from different layers and convolved by 3x3 kernels to produce M bounding boxes with a probability for the presence of the target in the box (0<f<1). Each bounding box c = [c1 , c2, c3, c4] is uniquely defined by the x and y offset of upper-left and lower-right corners relative to the default box coordinates of [-0.5, -0.5, 0.5, 0.5], The overall Mx(4 + 1) predictions are processed by a non-maximum suppression to decide the best ROI box.

[0057] For the ROI localisation neural network 124, the loss function is defined as the mean of the squared pixel errors where m and n are the number of rows and columns of the input image and,and the label, ltj are the prediction and the corresponding label of the pixel row at row i and column j. To prevent overfitting, dropout and weight regularisation can be applied during the ROI localisation neural network training. After training, the test dataset is used to test the performance of ROI localisation neural network 124 on images previously unseen by the network 124. The mean of the squared pixel error distance between the predicted ROI position and the label is 5.2 (SD: 4.3) pixels. The lossfunction of the ROI localisation neural network 124 was defined as the mean of the squared pixel errors:

[0059] where m, n are the numbers of rows and columns of the input image,are the prediction and the corresponding label of the pixel at row i and column j. Due to the vanishing gradient problem caused by the sigmoid activation in the last layer, it is taxing to train the segmentation network if the parameters are poorly initialised. The activation function introduces non-linearities to the neural network that allow it to capture complex patterns and relationships in the input data. Consequently, examples can be realised in which one or more than one activation function is implemented that aligns with the specific task and data characteristics. The elected activation function influences the effectiveness of deep learning models, as it influences learning capacity, stability, and computational efficiency. Examples can be realised in which, the Rectified Linear Unit (ReLU) activation function is used due to its simplicity, efficiency, and effectiveness in various applications. The ReLu activation function is used in the last layer and the network is pre-trained for a number of epochs to initialise the network parameters. The ReLU activation function is then changed to sigmoid to compute the segmentation map needed (values range between 0 and 1). The segmentation network is trained for at least 250 epochs before the validation error stops decreasing. Different metrics are available to test the image segmentation accuracy, however, most are designed for binary classification (0 or 1) problems. The values in the current generated segmentation map represent the probability of pixels belonging to each class (femur, tibia and patella), which are between 0 and 1.

[0060] Pixel accuracy (PA) is a metric calculating the ratio of pixels that are correctly classified to all pixels:TP+TN

[0061] PA =TP+TN+FP+FN

[0062] where TP, TN, FP and FN represent the pixel counts for true positive (label: 1, prediction: 1), true negative, (label: 0, prediction: 0), false positive, (label: 0 , prediction: 1) and false negative (label: 1, prediction: 0) respectively. Given prediction values between0 and 1 , a threshold is set to judge if the prediction is positive (prediction > threshold) or negative (prediction < threshold), and derive a weighted pixel accuracy from:- p(TP)+q(TN)

[0063] weighted PA = p(TP)+q(TN)+p(FP)+q(FN)

[0064] where p(... ) is the contribution of each pixel in that area to the count in the prediction value of the pixel rather than 1, and q(... ) means that the contribution is 1 minus the prediction value.

[0065] Although examples can be realised using pixel accuracy as measure of the pixels that are correctly classified, such a measure bears a risk of being misleading when the target area is too small compared to the background because the measure can be biased by a large TN. Therefore, examples can be realised in which a further metric is used for image segmentation. The further metric is intersection over union (IOU), which does not account for TN, IOU is given by:TP

[0066] IOU =TP+FP+FN

[0067] Examples can be realised that use a weighted IOU that is given by:

[0068] WeightedJOU

[0069] The weighted IOU calculates a weighted pixel count based upon the generated pixel values, where p() means that the contribution of each pixel in that area to the count is the prediction value of that pixel rather than 1 , and q() means that the contribution is 1 minus the prediction value.

[0070] FIG. 4 depicts a view 400 of the image segmentation neural network node 108 according to examples. The generated depth maps can lack a realistic depth dropout. Fortunately, the 3D point cloud representation of depth data is less vulnerable to such sampling artefacts compared to 2D depth maps. Network-learned features should be similar in both real and synthetic domains to ensure knowledge transfer; they should also be robust to camera sampling properties so that the trained network is camera-agnostic. Consequently, examples can be realised in which the image segmentation neural network is arranged to learn from the 3D point cloud representation rather than the 2D depth maps.

[0071] The image segmentation neural network node 108 comprises, or at least has access to, the image segmentation neural network 128. In the example depicted in figure 4, the image segmentation neural network 128 is shown as forming part of the image segmentation neural network node 108, as an alternative to the arrangement shown in figure 1 in which the image segmentation neural network 128 was accessible to the node 108.

[0072] The image segmentation neural network 128 is trained using a respective set of training data 402 comprising images. For example, the set of training data is randomly divided into training and validation sets according to a predetermined ratio. Examples can be realised in which the predetermined ratio is 8:2. Although examples have been described using a ratio of 8:2, examples are not limited to such an arrangement. Examples can be realised in which some other ratio is used instead.

[0073] The set of training data 402 comprises supervised learning training data in the form of a number of pairs images 404 to 408 and 404’ to 408’; each pair of images comprises an initial image depicting a knee 116 bearing the open incision 118 showing the set of bodily structures described above and a corresponding segmented image. The corresponding segmented images have been annotated, or otherwise labelled, to distinguish between pixels associated with respective scene features and pixels associated with one or more than one anatomical feature of interest. For instance, a first image 404 of an image pair 404 and 404’ shows a leg 414 with an open incision 416 depicting the femur 418, tibia 420, upper background portion 422 and lower background portion 424. A second corresponding image 404’ comprises the following regions of annotated pixels: a leg region 414’ with an open incision region 416’ depicting a femur region 418’, a tibia region 420’, an upper background portion region 422’ and a lower background portion 424’ region that respectively correspond the leg 414, the open incision 416, the femur 418, the tibia 420, the upper background portion 422 and the lower background portion 424 of the first image 404. Each region is known as a mask or segment. Therefore, the second image 404’ is also known as a mask image or a segmented image. Similarly, the set of training data 402 will comprise image pairs that also include an original image and an annotated or segmented image to identify, for example, the femur. Still further, the set of images 402 can also comprise images thathave been annotated to identify, for example, the patella. Although the second image 404’ has been shown as comprising multiple masks of segments, examples are not limited to such an arrangement. Examples can be realised in which each mask in the second image 404’ is actually annotated on a respective image such that each original image : segmented image pair is directed to depicting and annotating a set of anatomical features comprising a single anatomical feature, as opposed to the multiple anatomical features of the second image 404’.

[0074] The image segmentation neural network 128 is trained using the set of training data 402 to produce output data 412. The output data 412 comprises segmented image data. The segmented image data identifies a set of segments or regions of an image that identify, or are otherwise associated with, a set of bodily structures. The set of segments can comprise one segment, or multiple segments, that identify, or that are otherwise associated with, one bodily structure or multiple respective bodily structures. In the example depicted in figure 4, the output data 412 comprises the pair of images 132 comprising the first image 134 showing the segmented region 138 corresponding to the femur and the second image 136 showing the segmented region 140 corresponding to the tibia. The image segmentation neural network node 108 is arranged to receive the set of images 210 comprising the clipped images 224 to 228 that are derived from the respective colour plane images 216 to 220, together with the clipped depth data 230, as the input vector 410. After the localisation and segmentation networks are trained with satisfactory accuracy, they were used to process RGB and depth images from the depth camera directly. The localisation neural network 124 provides the position of the surgical site based on the RGB image, which is then used to crop the corresponding depth image to the required size. The cropped depth image is then fed into the image segmentation neural network 128 to remove surrounding tissues and obtain a clean surface of a set of bodily features. The set of bodily features can comprise one or more than one of the following taken in any and all permutations: a femur, a tibia and a patella, which are similar or comparable to the surface that the surgeon would map out manually using a digitising probe. Once the target bone surfaces have been obtained from the depth image, the pose of the set of bodily features such as, for example, the femur, the tibia and the patella, can be computed by comparing the acquired surface with a reference model relating to that set of bodily structures using the registration node 142.

[0075] The ICP algorithm is an effective and widely used algorithm for precise surface matching in surgical registration. However, the standard ICP requires a large number of iterations to achieve satisfactory convergence. In order to reduce the number of iterations, thus reducing convergence time, a more efficient variant of ICP, the point-to-plane ICP algorithm, is used to realise examples. Once initialised with a rough estimate of a limb pose with respect to the camera system, the ICP algorithm searches for corresponding points for each point in the segmented depth image and the reference model; the former representing a subset of the latter in terms of the features available for matching. The ICP implementation computes the best pose between these points and uses it to estimate a better candidate pose, which is then applied to the depth image to search for better correspondences, until monotonic convergence to a minimum. Since the segmentation process attributes to each point a probability of belonging to the bone, such a probability is used as a weight of that point when calculating total point-to-plane errors. Consequently, points with higher probability give a larger contribution to the ICP pose estimation, which helps further improve the registration accuracy.

[0076] FIG. 5A shows a view 500 of the architecture 502 of the image segmentation neural network 128 according to examples. The image segmentation neural network 128 comprises an encoder 504 and a decoder 506. The encoder 502 receives image data 508. The image data 508 is an example of the above described RGB image data. The received image data 508 has predetermined height, nH, and predetermined width, nw, dimensions in terms of pixels, and nccolour space channels. In the example depicted, the predetermined height and the predetermined width are the same. The predetermined height is 128 pixels. The predetermined width is 128 pixels. Although the architecture 502 depicted in figure 5A has a predetermined height of 128 pixels and a predetermined width of 128 pixels, examples are not limited to such an arrangement. Examples can be realised in which the aspect ratio of the RGB image is not 1 :1. The received image data 308 comprises 3 channels, as indicated by the numeral “3” above the layer. The format used in describing the data produced will be nHxnwxn , where nHis the height of an image or feature map, nwis the width of an image or feature map, and ncis the number of channels of an image or feature map. Although the example depicted in figure 5A uses 3 channels, examples are not limited to such an arrangement. Examples can be realised in which a set of channels is provided to the image segmentation neural network 128. The set ofchannels can comprise one channel or multiple channels. For instance, the image segmentation neural network 128 can be arranged to receive the depth data, that is, a depth image, either alone or together with one or more than one additional channel. Examples can be realised in which the image segmentation neural network can be arranged to receive 4 channels. The 4 channels can comprise the RGB channels and a Depth channel, that is, the RGBD output from the above mentioned camera(s). The 4 channels can be cropped according to respective regions of interest.

[0077] Table 2: Image Segmentation Neural Network Architecture

[0078] Table 2 above describes, in tabular form, the, or an example of, image segmentation neural network architecture 502. The received data is, or the received channels are, subjected to a convolution operation using a convolutional operator 510 of a predetermined size. In the example depicted, the convolutional operator comprises a matrix of the predetermined size. Examples can be realised in which the predetermined size is 7x7. The values of the convolutional operator 510 will have been learnt or derived from training the image segmentation neural network. The convolutional operator has a stride of 1.

[0079] Applying the convolution operation results in a second set of image data or feature map 512. The second set of image data 512 has dimensions 32x128x128. The second set of image data 512 is subject to a second convolutional operation using a second convolutional operator 514. The second convolutional operator has predetermined dimensions. In the example depicted, the predetermined dimensions are 3x3. The values of the operator will have been derived or learnt as a consequence of training the image segmentation neural network 128. The stride length is 1. Applying the second convolutional operator to the second set of image data 512 results in a third set of image data or feature map 516. The third set of image data 516 has dimensions 128x128x32.

[0080] A pooling operator 518 is applied to the third set of image data 516, which results in a fourth set of image data or feature map 520. The fourth set of image data 520 has smaller dimensions compared to the third set of image data 516, that is, the fourth set of image data 520 is a down-sampled version of the third set of image data 516. The extent of the reduction is governed by the pooling operator, in particular, by the dimensions of the pooling operator together with the stride of the pooling operator. The example depicted uses 2x2 max pooling with a stride of 2, which results in the fourth set of image data 520 having dimensions 64x64x32, that is, there are 32 channels each comprising a respective 64x64 feature map.

[0081] The fourth set of image data 520 is subjected to a third convolution operation using a third convolutional operator 522. The third convolutional operator 522 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the third convolutional operator 522 is the same as the second convolutional operator 514. Applying the third convolutional operator 522 to thefourth set of image data 520 results in a fifth set of image data or fifth feature map 524. The fifth set of image data 524 has dimensions 64x64x64.

[0082] The fifth set of image data 524 is subjected to a fourth convolution operation using a fourth convolutional operator 526. The fourth convolutional operator 526 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the fourth convolutional operator 526 is the same as the second convolutional operator 514. Applying the fourth convolutional operator 526 to the fifth set of image data 524 results in a sixth set of image data or sixth feature 528. The sixth set of image data 528 has dimensions 64x64x64.

[0083] A second pooling operator 530 is applied to the sixth set of image data 528, which results in a seventh set of image data or feature map 532. The seventh set of image data 532 has smaller dimensions compared to the sixth set of image data 528, that is, the seventh set of image data 532 is a down-sampled version of the sixth set of image data 528. The extent of the reduction is governed by the pooling operator, in particular, by the dimensions of the pooling operator and the stride. The example depicted uses 2x2 max pooling with a stride of 2, which results in the seventh set of image data 532 having dimensions 32x32x64.

[0084] The seventh set of image data 532 is subject to fifth convolution operation using a fifth convolutional operator 534. The fifth convolutional operator 534 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the fifth convolutional operator 534 is the same as the second convolutional operator 514. Applying the fifth convolutional operator 534 to the seventh set of image data 532 results in an eighth set of image data or an eighth feature map 536. The eighth set of image data 536 has dimensions 32x32x128.

[0085] The eighth set of image data 536 is subjected to a sixth convolution operation using a sixth convolutional operator 538. The sixth convolutional operator 538 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the sixth convolutional operator 538 is the same as the second convolutional operator 514. Applying the sixth convolutional operator 538 to the eighth set of image data 536 results in a ninth set of image data or ninth feature map 540. The ninth set of image data 540 has dimensions 32x32x128.

[0086] A third pooling operator 542 is applied to the ninth set of image data 540, which results in a tenth set of image data or tenth feature map 544. The tenth set of image data 544 has smaller dimensions compared to the ninth set of image data 540, that is, the tenth set of image data 544 is a down-sampled version of the ninth set of image data 540. The extent of the reduction is governed by the pooling operator, in particular, by the dimensions of the pooling operator and the stride. The example depicted uses 2x2 max pooling, which results in the tenth set of image data 544 having dimensions 16x16x128.

[0087] The tenth set of image data 544 is subjected to a seventh convolution operation using a seventh convolutional operator 546. The seventh convolutional operator 546 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the seventh convolutional operator 546 is the same as the second convolutional operator 514. Applying the seventh convolutional operator 546 to the tenth set of image data 544 results in an 11thset of image data or an 11thfeature map 548. The 11thset of image data 548 has dimensions 16x16x256.

[0088] The eleventh set of image data 548 is subjected to an eighth convolution operation using an eighth convolutional operator 550. The eighth convolutional operator 550 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 3x3. Examples can be realised in which the eighth convolutional operator 550 is the same as the second convolutional operator 514. Applying the eighth convolutional operator 550 to the eleventh set of image data 548 results in a twelfth set of image data or twelfth feature map 552. The twelfth set of image data 552 has dimensions 16x16x256.

[0089] Each convolutional layer is subject to a respective activation function. Examples can be realised, as will be appreciated from table 2 above, in which the activation functions are GeLu activation functions.

[0090] Next the operations performed by the decoder 506 will be described. The decoder 506 generates segmentation masks from or corresponding to the features identified by the encoder 504 in the image data 508. The decoder 506 operates in reverse to the encoder 504, but also uses concatenated images from the encoder 504 to preserve identified features to generate segmentation masks from features extracted by the encoder 504.

[0091] A deconvolution operation using a deconvolution operator 554 is applied to the twelfth set of image data 552. The deconvolution operator 554 has predetermined dimensions. The predetermined dimensions are 2x2. The deconvolution operator is learnt or derived during training the image segmentation neural network 502. Applying the deconvolutional operator 554 to the twelfth set of image data 552 results in a thirteenth set of image data or thirteenth feature map 556. The thirteenth set of image data 556 is an upscaled version of the twelfth set of image data 552. To preserve identified or extracted features, a copy of the ninth set of image data 540 is concatenated, in a concatenation operation 541 , with the thirteenth set of image data 556. The dimensions of the thirteenth set of concatenated image data 556 are 32x32x256.

[0092] The thirteenth set of concatenated image data 556 is subjected to a convolution operation using a ninth convolutional operator 558, which results in a fourteenth set of image data or fourteenth feature map 560. The fourteenth set of image data 560 has dimensions 32x32x128. The ninth convolutional operator 558 is the same as the second convolutional operator 514.

[0093] The fourteenth set of image data 560 is subjected to a tenth convolution operation using a tenth convolutional operator 562 to produce a fifteenth set of image data or fifteenth feature map 564. The tenth convolutional operator 562 has dimensions of 3x3. The tenth convolutional operator 562 is the same as the second convolutional operator 514.

[0094] The fifteenth set of image data 564 is subjected to a second deconvolution operation using a second deconvolutional operator 566. The second deconvolutional operator 566 upscales the fifteenth set of image data 564 to produce a sixteenth set of image data or sixteenth feature map 568. The sixteenth set of image data 568 is concatenated, using a second concatenation operation 567, with the sixth set of image data 528. The resulting sixteenth set of concatenated image data 568 has dimensions of 64x64x128.

[0095] The sixteenth set of concatenated image data 568 is subjected to a convolution operation using an eleventh convolutional operator 570. The eleventh convolutional operator 570 has predetermined dimensions. The predetermined dimensions, in the example, shown are 3x3. The eleventh convolutional operator 570 is the same as thesecond convolutional operator 514. Subjecting the sixteenth set of concatenated image data 568 to the eleventh convolutional operator 570 results in a seventeenth set of image data or seventeenth feature map 572. The seventeenth set of image data 572 has dimensions 64x64x64.

[0096] The seventeenth set of image data 572 is subjected to a twelfth convolutional operation using a twelfth convolutional operator 574. The twelfth convolutional operator 574 is a 3x3 convolutional operator. Subjecting the seventeenth set of image data 572 to the twelfth convolutional operator 574 results in an eighteenth set of image data or eighteenth feature map 576. The eighteenth feature map has dimensions 64x64x32.

[0097] The eighteenth set of image data 576 is subjected to a third deconvolution operation using a third deconvolutional operator 578. The third deconvolutional operator 578 upscales the eighteenth set of image data 576 to produce the nineteenth set of image data or nineteenth feature map 580. The nineteenth set of image data 580 has dimensions of 128x128x32.

[0098] The nineteenth set of image data 580 is concatenated, using a third concatenation operation 581 , with a copy of the third set of image data 516 to create a nineteenth set of concatenated image data. The nineteenth set of concatenated image data has dimensions of 128x128x64.

[0099] The nineteenth set of concatenated image data is subjected to a convolution operation using a thirteenth convolutional operator 582. The thirteenth convolutional operator 582 is a 3x3 convolutional operator. Subjecting the nineteenth set of concatenated image data 580 to the thirteenth convolutional operator 582 results in a twentieth set of image data or twentieth feature map 584. The twentieth set of image data 584 has dimensions 128x128x32.

[0100] The twentieth set of image data 584 is subjected to a convolution operation using a fourteenth convolutional operator 586. The fourteenth convolutional operator 586 is a 3x3 convolutional operator. Subjecting the twentieth set of image data 584 to the fourteenth convolutional operator 586 results in a twenty-first set of image data or twenty- first feature map 588. The twenty-first set of image data 588 has dimensions 128x128x32.

[0101] The twenty-first set of image data 588 is subjected to a convolution operation using a fifteenth convolutional operator 590. Subjecting the twenty-first set of image data 588 to the fifteenth convolutional operator 590 results in a set of segmented image data 592. The fifteenth convolutional operator 590 has predetermined dimensions. Examples can be realised in which the predetermined dimensions are 1x1 , such as depicted in figure 5.

[0102] The set of segmented image data 592 comprises at least one image containing at least one mask segmenting, or otherwise identifying, at least one feature of the image. Segmenting, or otherwise identifying, the at least one feature of the image comprises distinguishing between pixels associated with the at least one feature from pixels not associated with the at least one feature. Examples can be realised in which the at least one feature of the image is considered to be a foreground feature, and the pixels corresponding to the at least one feature are considered to be foreground pixels, whereas pixels not associated with, that is, not corresponding to, the at least one feature are considered to be background pixels associated with a background image.

[0103] Examples can be realised in which the set of segmented image data 592 is subjected to, or has been subjected to, an associated softmax function that gives a probability of a number of possible probabilities, each one corresponding to the probability that a segmentation image in the set of segmented image data 592 corresponds to a segmented image associated with a respective bodily part.

[0104] FIG. 5B shows a view 500B of a set 502B of potential outputs of the image segmentation network 128. The set 502B of potential outputs of the image segmentation network 128 is an example of multiclass classification data. In the examples depicted, the set 502B comprises four segmentation images 504B to 510B. The set 502B is an example of the above described set of segmented image data 592. Each segmentation image 504B to 510B corresponds to either a classified or respective bodily part or the background. For instance, assuming that the image segmentation network 128 is arranged to produce segmentation images corresponding to any of the set {femur, tibia, patella, background}, the outputs of the image segmentation network 128 can comprise one or more than one of the following segmentation images taken jointly or severally in any and all permutations: a femur segmentation image 504B, a tibia segmentation image506B, a patella segmentation image 508B and a background segmentation image 51 OB. The femur segmentation image 504B comprises a mask 512B of pixels corresponding to the distal femur 418. The tibia segmentation image 506B comprises a mask 514B of pixels corresponding to the tibia 420. The patella segmentation image 508B comprises a mask 516B of pixels corresponding to the patella. The background segmentation image 51 OB comprises a mask 518B of pixels corresponding to the background. In the background segmentation image 51 OB depicted, the background comprises pixels other than pixels related to the leg. However, examples can be realised in which the background segmentation image 51 OB comprises the mask 518B of pixels corresponding to pixels other than pixels relating to the femur, tibia and patella. In essence, the background segmentation image 51 OB can comprise pixels that are the logical NOT of the union, that is, the logical OR, of the other segmentation images in the set 502B of segmentation images. Each pixel of an output segmentation image has an associated probability. If the associated probability is greater than a predetermined threshold, the pixel is deemed to relate to an anatomical feature of interest. For example, if the probability of a pixel is greater than 0.8, the pixel is assumed to relate to a corresponding anatomical feature, otherwise the pixel is assumed to relate to either a different anatomical feature not associated with a particular class or to relate to the background.

[0105] Therefore, the image segmentation neural network 128 can be realised as a multiclass classifier that outputs multiclass classification data in the form a segmentation image according to a classified or otherwise detected bodily part.

[0106] FIG. 5C shows a view 500C of the overall process 502C of producing a segmented image 504C from an RGB image 506C and a depth image 508C. The RGB image 506C is an example of an original image. The depth image 508C is an example of an original depth image. The RGB image 506C is processed by the region of interest neural network node 106 to identify a region of interest 512C. The output of the region of interest neural network 106 can comprise an image 510C depicting the region of interest 512C.

[0107] Alternatively, or additionally, data describing or otherwise defining the identified region of interest 512C can be output. As an example, the data describing or otherwise defining the region of interest 512C can comprise data defining a box or other closedboundary shape. In the example depicted in figure 5C, the region of interest 512C is shown as a cropped image that is derived from the image 510C containing the region of interest 512C. The region of interest 512C is used to crop or clip the depth image 508C to produce a cropped or clipped depth image 514C. The cropped or clipped depth image is processed by the image segmentation neural network node 128 to produce the segmented image 504C. In the example depicted, the segmented image 504C identifies the femur 516C. Within the context of the examples, to crop or clip an image comprises at least one, or more than one, of the following: (1) creating a subset of data from given initial data, (2) identifying within given initial data, a subset of data, (3) extracting a subset of data from given initial data. Applying the foregoing, it can be appreciated that:

[0108] (a) discarding image data of the given initial image 510C to retain the subset of image data within the bounding box creates the cropped or clipped regional of interest image data 512C, and / or

[0109] (b) copying data from within, or otherwise creating associated with, data within the bounding box within the given initial image 510C creates a subset of date 512C,

[0110] defines the subset of data via the bounding box, or other closed form shape, relative the given initial image 510C.

[0111] Examples can be realised in which both the RGB image 506C and the depth image 508C are fed into the region of interest localization neural network 106 to produce region of interest image 512C. In such a case, the region of interest localization neural network 106 will have first been trained using RGB image: annotated segmented image pairs and depth image data: annotated segmentation depth image pairs.

[0112] Furthermore, examples can be realised in which the region of interest image 512C and the cropped depth image 514C are fed into the image segmentation neural network 128. In such a case, the image segmentation neural network 128 will have been trained using the cropped RGB image: annotated segmented image pairs and cropped depth image: annotated segment image pairs.

[0113] Referring to figure 5D, there is shown a view 500D of a side by side comparison between a marker based registration process 502D and a markerless based registration process 504D.

[0114] Referring to the marker based registration process 502D, a camera 503D is used to capture a view of an exposed knee joint 506D. At 508D, pins are inserted into the exposed knee joint for supporting one or more than one marker array assembly. The one or more marker array assemblies are mounted to the pins at 510D. At 512D, verification checkpoint placement is performed, which verifies the placement of the marker assemblies. At 514D, patient landmark collection is performed to collect data relating to anatomical landmarks associated with the patient. The landmark collection can be performed using a probe with an attached marker. The data representing the landmark collection can comprise 3D point cloud data. At 516D, registration of the collected landmark data with a model of the patient anatomy is performed that, in turn, supports registration, at 518D, of one or more than one anatomical feature of the patient. In the present example, the one or more than one anatomical feature comprises at least one or more than one of the following taken jointly and severally in any and all permutations: patella, femur and tibia. At 520, robotic surgery is performed.

[0115] In contrast, referring to the markerless based registration process 504D, an RGB image 522D of an exposed knee joint 524D is captured; the exposed knee joint having an intact retinaculum. A depth image 526D is captured. Capturing the depth image 526D can be performed simultaneously with capturing the RGB image 522D. The depth image comprises a 3D point cloud of the exposed knee joint 524D. Knee localisation is performed to produce a region of interest 528D. The region of interest 528D in the present example is defined using a rectangular bounding box. The region of interest 528D is used to crop the depth image 526D to produce a reduced depth image 530D that is focused or otherwise centred on 3D point cloud data corresponding to the region of interest 528D. The amount of RGB image data to be processed can be reduced by creating a coarse detection image 532D comprising the RGB data within the region of interest 528D. The coarse detection image 532D can be used to produce the reduced depth image 530D. A mask 534D is created corresponding to anatomical features derived from the 3D point cloud and the region of interest and / or the coarse detection image 532D. The pixels within the mask can corresponding to one or more than one anatomical feature. The anatomical features can comprise at least one or more than one of the following taken in any and all permutations: the patella, the femur and the tibia. The mask 534D can be used to label the RGB image 522D to create ground truth data for training the segmentation network128. Using multiple instances of the ground truth data, the segmentation network 128 is trained to produce segmented images that identify, or correspond to, anatomical structures of the exposed knee 524D. In the example depicted in figure 5D, there is shown a segmented image 536D showing the patella 538D and a segmented image 540D showing the femur 542D and tibia 544D. An I PC registration process can be applied to register the segmented images 536D and 540D with at least one, or both, of: the cropped depth image 530D and corresponding anatomical models 546D and 548D of the patella and femur and tibia respectively. Finally, RGB output images 550D and 552D can be generated that identify the anatomical features 542D and 544D on corresponding RGB images.

[0116] FIG. 6 illustrates a view 600 of a flowchart 602 for generating ground truth images according to examples. The images 404 to 408 in the set of images 402 are examples of ground truth images. At 602, a plurality of images is generated. Each image in the plurality of images shows a respective knee 116 bearing a respective open incision 118 that reveals one or more than one respective bodily structure of a respective set of bodily structures. The open incision 118 is arranged to leave the retinaculum intact.

[0117] The plurality of images, or at least a subset of the plurality of images, is annotated to identify bodily structures at 604 to create a set of ground truth images. Examples can be realised in which each pixel corresponding to a bodily structure is labelled, or otherwise identified as, an appropriate bodily structure.

[0118] At 606, foreground occlusion images are generated as indicated above. To evaluate the performance of the trained segmentation network in the real world, over 1000 RGB-D images were captured with the camera, during which the target knee joint was partially occluded by hands or surgical tools. When hands or tools occluded the target surface, the registration between digitised surfaces and unsegmented captures became highly unreliable. To ensure correct annotation under target occlusion, pairwise captures are utilized; namely, a set of pairwise images comprising frame 1 and frame 2. A first image, frame 1 , of the target was first captured with no surface occlusion or contact, then labelled by ICP-based point matching. Subsequently, without moving the camera or target, a second image, frame 2, was captured for the target, such as the knee surface, while being partially occluded by a hand in a purple-coloured surgical glove or a tool to simplifythe segmentation process, as follows. The femur mask labelled in frame 1 was applied to frame 2’s RGB frame to segment an Rol, which was then converted to hue saturation and value (HSV) format, and filtered by a band-pass hue filter in the purple colour range to identify the pixels that belong to the foreground. The ground truth femur, tibia and patella pixels for frame 2 were computed by subtracting the femur, tibia and patella pixels in frame 1 from the detected foreground pixels in frame 2. The ground truth Rol box was computed as the smallest rectangle that covers all ground truth femur, tibia and patella pixels. Regardless of hand occlusion, tool manipulation, capturing perspective and human presentation, the network properly pays attention to the exposed femur. Regardless of hand occlusion, tool manipulation, capturing perspective and human presentation, the region of interest localisation neural network 124 properly pays attention to the exposed femur. If the intersection over union (loU) between the predicted Rol and the ground truth Rol is higher than 0.5, the prediction is regarded as successful. The overall accuracy is presented by the success rate of predictions over the entire test dataset. The trained localisation network 124 is also tested as a reference for comparison. The predicted Rol is regarded as the box drawn around the inferred target location, with the same size as the ground truth Rol box. Depending on the ground truth label (positive: is femur, tibia or femur; otherwise, negative) and the correctness of the prediction (true: prediction matches ground truth; otherwise false), the N points can be classified as true positive (TP), truenegative (TN), false positive (FP) and false-negative (FN). To avoid the bias arising from a large number of TN predictions for background points, the segmentation accuracy is defined as the loU score in each frame:TP

[0119] loU =TP+FP+FN

[0120] FIG. 7A shows a view 700A of a flowchart 702 for image segmentation according to examples. The RGB D image data 112 is received at 704 by the region of interest neural network node 106. The RGB-D image data 112 is cropped at 706 by the region of interest neural network 124. The cropped RGB-D image data is formed, at 708, into the input vector 410 for input to the image segmentation neural network node 108 and sent, at 710, for processing by the image segmentation neural network 128 to produce the output data 412. The output data 412, received at 712, is a segmented view of the RGB image data. The output data 412 can be, or is, passed to thevisualisation node 110 for display on the display 130 at 714. The image segmentation neural network (ISNN) receives the input vector at 716. The ISNN processes the received input vector at 718 to produce the output data 412. The output data is returned at 720.

[0121] FIG. 7B shows a view 700B of a flowchart 702B for cropping the depth data according to an example. The RGB data is received for processing at 704B. The region of interest 126 is determined at 706B using the ROI network 124. At 708B, the region of interest 126 can be used to crop the RGB data so that the segmentation neural network has less information to process in creating the set of segmented images 132. The depth data from the RGB-D camera is received at 710B, and the depth data is cropped to the region of interest at 712B before being processed by the segmentation neural network 108 to produce the set of segmented images 132.

[0122] FIG. 7C shows a view 700C of a flowchart 702C for generating training data for the examples described herein. The training data is used to train the system 102 to identify a patella, or to derive data associated with a patella such as, for example, a segmentation image comprising a mask associated with the patella. The markers described below are tracked using an tracking camera. The tracking camera can use IR tracking to capture the orientation and position of a marker within a frame of reference, that is, within tracking camera world coordinates. The orientation and position of a marker is known as a pose.

[0123] At 704C, a marker, such as a passive or active marker, coupled to, or positioned with respect to, the patella is identified, and the position and orientation of the patella marker are determined. The patella is revealed due to a corresponding incision. The marker defines a respective patella pose, Mpt. The corresponding incision is such that the retinaculum is intact. Therefore, the incision is a partial arthrotomy exposing the patella contained within the retinaculum.

[0124] At 706C, a patella point cloud, Ppt :is created comprising a set of points associated with the patella. The set of points associated with the patella can comprise at least one or more of the following taken jointly and severally in any and all combinations: a set of points defining, or associated with, a surface of the patella, a set of points defining landmarks of the patella, a set of points associated with the perimeteror extremities of the patella. The point cloud can be created by tracking the position and orientation of a probe bearing a respective marker that defines a respective probe pose, Mp, using an IR camera.

[0125] At 708C, RGB-D images of the patella are captured using the RGB-D camera 114. The RGB-D camera has an associated marker that defines a respective camera pose, Mc,

[0126] Any movement of the patella during RGB-D image capture at 708C can be noted at 710C, that is, the RGB-D images and the patella poses can be simultaneously monitored or captured.

[0127] At 712C, the RGB-D image data and the patella point cloud data are aligned. The RGB-D images data and the patella point cloud data can be aligned using, for example, an Iterative Closest Point (ICP) algorithm.

[0128] Having the patella point cloud data, Ppt, aligned with at least one, or both, of: the RGB image data and the depth image data allows at least one, or both, of: the RGB image data and the depth image data to be labelled as patella or not patella.

[0129] Accordingly, the foregoing example supports capturing depth images together with corresponding RGB images, as well as the corresponding camera poses, Mc, and patella poses, Mpt, in which at least one, or both, of: the RGB image data and the depth image data are labelled using the patella point cloud as being part of the patella or not. The labelling can be realised by transforming the patella point cloud, Ppt, into a camera point cloud, Pc, within the camera frame of reference by

[0130] Pc= T^T^xT^xPpt

[0131] where is a transformation that maps the patella pose, Mpt, into a frameof reference of the IR camera,is a transformation that maps the camera frame of reference into the frame of reference of the IR camera, andrepresents the inverse transformation of (T )-1,and TMCis atransformation that maps the camera pose, Mc, into the camera frame of reference.

[0132] Accordingly, a mapping can be made between patella point cloud data, Ppt, and RGB-D image data by mapping the patella point cloud data, Ppt, into the RGB-D camera reference frame. The mapping of patella point cloud data, Ppt, into the camera frame of reference allows at least one, or both, of: the RGB image data and the depth image data to be labelled accurately.

[0133] The accuracy of the mapping can be improved using, for example, an Iterative Closest Point algorithm to align the patella point cloud data, Ppt, expressed in the camera frame of reference, Pc, with the at least one, or both, of: the RGB image data and the depth image data.

[0134] Although the example described with reference to figure 7C relates to a patella, examples are not limited to such an anatomical part. Examples can be realised in which other anatomical parts are exposed such as, for example, at least one or more of: the femur or tibia, or some other joint, or anatomical feature or features.

[0135] Referring to FIG. 8, there is shown a view 800 of a Computer Assisted Surgery System (CASS) 802 according to an example. The CASS 802 can be made responsive to the above image segmentation to control one or more than one aspect of the CASS 802. In the example depicted, the CASS 802 is arranged to aid surgeons in performing orthopaedic surgical procedures such as, for example, a knee arthroplasty (e.g., total knee arthroplasty (TKA)) or total hip arthroplasty (THA). An Effector Platform 804 positions surgical tools relative to a patient during surgery. For example, for a knee surgery, the Effector Platform 804 may include an End Effector 804B that holds surgical tools or instruments during their use. Effector Platform 105 can include a Limb Positioner 804C for positioning the patient’s limbs during surgery. Resection Equipment (not shown in FIG. 8) performs bone or tissue resection using, for example, mechanical, ultrasonic, or laser techniques. Effector Platform 804 can also include a cutting guide or jig 804D that is used to guide saws or drills used to resect tissue during surgery. Such cutting guides 804D can be formed integrally as part of the Effector Platform 804 or Robotic Arm 804A, or cutting guides can be separate structures that can be matingly and / or removably attached to the Effector Platform 804 or Robotic Arm 804A.

[0136] The CASS 802 comprises a Tracking System 806 that uses one or more sensors to collect real-time position data to locate the patient’s anatomy and surgical instruments. Any suitable tracking system can be used for tracking surgical objects and patient anatomy in the surgical theatre. For example, a combination of infrared (I R) and visible light cameras can be used in an array. Such a Tracking System 806 can use the EMR retro- reflected from any of the retro-reflectors described and / or claimed herein to determine real-time position data that locates at least one, or both, of the patient’s anatomy and surgical instruments. The Tracking System 806 is an example of the abovedescribed cameras such as, for example, the RGB-D camera 114 and the camera 505D.

[0137] Accordingly, the CASS 802 shown in FIG. 8 depicts a number of retroreflectors. The retro- reflectors are examples of marker assemblies. The retro-reflectors can be placed on objects or body parts to be tracked or for which respective positions are to be determined. For example, a first retro- reflector 814 is situated on the robot arm 804A. Knowing the position of the retro-reflector 814 can allow, for exam. , the position of the actuator 816 of the robot arm 804A to be determined. A second retro-reflector 818 is placed on the handheld tool 804B to allow the position of the handheld tool to be determined and / or tracked in 3D space. A third retro-reflector 820 can be situated relative to the jig 804D. At least a fourth retro-reflector 822 can be used to determine not only the position of the jig 804D, but also the attitude in 3D space of the jig 804D. A sixth retroreflector 824 can be placed on the Limb Positioner 804C to assist in determining the position of a respective distal actuator 826 for holding a limb.

[0138] Although the CASS 802 has been described with reference to a set of retroreflectors comprising six retro-reflectors, examples are not limited thereto. Examples can be realised in which such a set of retro-reflectors comprises one or more than one retroreflector to suit the needs of the operation to be performed. Still further, the deployment of the retro-reflectors can be realised other than in relation to the robot arm 804A, the handheld tool 804B, the jig 804D and the limb positioner 804C.

[0139] The registration process that registers the CASS 802 to the relevant anatomy of the patient can also involve the use of anatomical landmarks, such as landmarks on a bone or cartilage. For example, the CASS 802 can include a 3D model of the relevant bone or joint and the surgeon can intraoperatively collect data regarding the location ofbony landmarks on the patient’s actual bone using a probe that is connected to the CASS. Alternatively, the CASS 802 can construct a 3D model of the bone or joint without preoperative image data by using location data of bony landmarks and the bone surface that are collected by the surgeon using a CASS probe or other means.

[0140] A Tissue Navigation System (not shown in figure 8) provides the surgeon with intraoperative, real-time visualization for the patient’s bone, cartilage, muscle, nervous, and / or vascular tissues surrounding the surgical area.

[0141] The CASS 802 comprises a Display 808 to provide graphical user interfaces (GUIs) that display images collected by the Tissue Navigation System as well other information relevant to the surgery to a surgeon or other operating threatre staff 828. For example, the Display 808 overlays image information collected from various modalities (e.g., CT, MRI, 8-ray, fluorescent, ultrasound, etc.) collected pre-operatively or intra- operatively to give the surgeon various views of the patient’s anatomy as well as real-time conditions. A Surgical Computer 810 provides control instructions to various components of the CASS 802, collects data from those components, and provides general processing for various data needed during surgery. The surgical computer 810 can be used to realise, or to provide access to, the system 102. In the example depicted in figure e. 8, the surgeon 828 is shown as wearing protective eye-ear 830.

[0142] Examples can be realised in the form of machine-instructions. The machineinstructions described herein can be stored using respective machine-readable storage. The machine-instructions can be arranged to realise any of the examples described herein. The machine-instructions can be realised as either hardware, software, or a combination of hardware and software. If the machine-instructions are realised as software, the machine-instructions can be processed by an interpreter, processed by a compiler and executed by a processor or given effect in some other way such as, for example, being realised as an FPGA or ASIC. The terms circuitry and logic are examples of hardware, software or a combination of hardware and software.

[0143] Therefore, referring to figure 9, there is shown a view 900 of machine-readable storage 902 storing machine-instructions 904 for generating the ground truth images according to examples. The machine-instructions 904 can be processed by one or more than one processor 906. The machine-instructions 904 comprise:

[0144] Instructions 908 to generate the randomised scenes;

[0145] Instruction 910 to generate the above described ground truth images; and

[0146] Instructions 912 to generate foreground occlusion images.

[0147] Furthermore, referring to figure 10, there is shown a view 1000 of machine- readable storage 1002 storing machine-instructions 1004 for processing by a processor 1006. The machine-instructions 1004 comprise:

[0148] Instructions 1008 to receive or otherwise access RGB-D data or images;

[0149] Instructions 1010 to crop RGB-D data and create a Rol using region of interest localisation neural network 124;

[0150] Instructions 1012 to form the cropped RGB-D data into an input vector for the image segmentation neural network 128;

[0151] Instructions 1014 to send or submit the input vector to the image segmentation neural network 128;

[0152] Instructions 1016 to receive or otherwise access the segmented image data; and

[0153] Instructions 1018 to output the segmented image data for further processing.

[0154] Furthermore, referring to figure 11 , there is shown a view 1100 of machine- readable storage 1102 storing machine-instructions 1104 for processing by a processor 1106. The machine-instructions 1104 comprise:

[0155] Instructions 1108 to receive the input vector;

[0156] Instructions 1110 to process the input vector using the image segmentation neural network; and

[0157] Instructions 1112 to output the segmented image data.

[0158] Referring to figure 12, there is shown a view 1200 of machine-readable storage 1202 storing machine-instructions 1204 for processing by a processor 1206 for cropping depth data according to an example. The machine-instructions 1204 comprise:

[0159] Instructions 1208 to receive RGB image data;

[0160] Instructions 1210 to determine the region of interest 126 using the ROI network 124; the region of interest 126 can be used to crop the RGB data so that the segmentation neural network has less information to process in creating the set of segmented images 132;

[0161] Instructions 1212 to receive the depth image data; and

[0162] Instructions 1214 to crop the depth image data according to the region of interest 126.

[0163] Referring to figure 13, there is shown a view 1300 of machine-readable storage 1302 storing machine-instructions 1304 for processing by a processor 1306 for generating training data for the examples described herein with reference to, for example, figure 7C. The training data is used to train the system 102 to identify a patella, or to derive data associated with a patella such as, for example, a segmentation image comprising a mask associated with the patella. The markers described below are tracked using an tracking camera. The tracking camera can use IR tracking to capture the orientation and position of a marker within a frame of reference, that is, within tracking camera world coordinates. The machine-instructions 1304 comprise:

[0164] Instructions 1306 to identify, and determine the position and orientation of a marker, such as a passive or active marker, coupled to, or positioned with respect to, the patella The patella is revealed due to a corresponding incision. The marker defines a respective patella pose, Mpt. The corresponding incision is such that the retinaculum is intact. Therefore, the incision is a partial arthrotomy exposing the patella contained within the retinaculum;

[0165] Instructions 1308 to create a patella point cloud, Ppt, comprising a set of points associated with the patella. The set of points associated with the patella can comprise at least one or more of the following taken jointly and severally in any and all combinations: a set of points defining, or associated with, a surface of the patella, a set of points defining landmarks of the patella, a set of points associated with the perimeter or extremities of the patella. The point cloud can be created by tracking the position and orientation of a probe bearing a respective marker that defines a respective probe pose, Mp, using an IR camera;

[0166] Instructions 1310 to capture RGB-D images of the patella using the RGB-D camera 114. The RGB-D camera has an associated marker that defines a respective camera pose, Mc-,

[0167] Instructions 1312 to note any movement of the patella during RGB-D image capture, that is, the RGB-D images and the patella poses can be simultaneously monitored or captured;

[0168] Instructions 1314 to align the RGB-D image data and the patella point cloud data. Aligning the RGB-D image data and the patella point cloud data. The RGB-D images data and the patella point cloud data can be aligned using, for example, an Iterative Closest Point (ICP) algorithm; and

[0169] Instructions 1316 to label at least one, or both, of: the RGB image data and the depth image data as patella or not patella using the patella point cloud data, Ppt, aligned with at least one, or both, of: the RGB image data and the depth image data.

[0170] Accordingly, the foregoing example supports capturing depth images together with corresponding RGB images, as well as the corresponding camera poses, Mc, and patella poses, Mpt, in which at least one, or both, of: the RGB image data and the depth image data are labelled using the patella point cloud as being part of the patella or not. The labelling can be realised by transforming the patella point cloud, Ppt, into a camera point cloud, Pc, within the camera frame of reference by Pc= TMcx(TMRc)~1xTl]RtxPptwhere TM is a transformation that maps the patella pose, Mpt, into a frame of reference of the IR camera, T^Ris a transformation that maps the camera frame of reference into the frame of reference of the IR camera, andrepresents the inverse transformation of (T )- 1, and T^cis a transformation that maps the camera pose, Mc, into the camera frame of reference.

[0171] Accordingly, a mapping can be made between patella point cloud data, Ppt, and RGB-D image data by mapping the patella point cloud data, Ppt, into the RGB-D camera reference frame. The mapping of patella point cloud data, Ppt, into the camera frame of reference allows at least one, or both, of: the RGB image data and the depth image data to be labelled accurately.

[0172] The accuracy of the mapping can be improved using, for example, an Iterative Closest Point algorithm to align the patella point cloud data, Ppt, expressed in the camera frame of reference, Pc, with at least one, or both, of: the RGB image data and the depth image data.

[0173] Suitably, examples can be realised in which the machine-instructions 1314 to align the RGB-D image data and the patella point cloud data comprise machineinstructions (not shown) to implement an Iterative Closest Point algorithm.

[0174] Further examples can be realised according to the following clauses.

[0175] Clause 1 : A method for generating a segmented image of a bodily structure; the method comprising:

[0176] determining a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to define the region of interest as comprising any detected predetermined bodily structures within the image; and

[0177] generating a segmented image comprising one or more masks associated with the detected predetermined bodily structures using a second neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth data being defined by the region of interest.

[0178] Clause 2: The method of clause 1 in which the generating a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment depth data of a depth image; the depth data being defined by the region of interest comprises:

[0179] producing a cropped depth image of the detected bodily structures from a depth image comprising depth data associated with the scene containing the detected bodily structures; and

[0180] generating the segmented image comprising the one or more masks associated with the detected bodily structures using a second neural network trained to segment the cropped depth image.

[0181] Clause 3: The method of either of clauses 1 to 2, in which the depth image is derived from a 3D point cloud.

[0182] Clause 4: The method of any of clauses 1 to 3, in which the image of the scene is a colour image, optionally, an RGB image.

[0183] Clause 5: The method of any preceding clause, in which the depth data and the image of the scene are derived from an RGB-D camera.

[0184] Clause 6: A method for generating a segmented image of bodily structures; the method comprising:

[0185] generating a segmented image comprising one or more masks respectively associated with one or more than one detected predetermined bodily structure using a neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth image being derived from a scene comprising at least one bodily structure of the predetermined bodily structures.

[0186] Clause 7: The method of clause 6, in which the depth data of the depth image is associated with a 3D point cloud.

[0187] Clause 8: The method of either of clauses 6 and 7, in which the depth data of the depth image is a 2D map of depth data derived from the 3D point cloud or the 3D point cloud per se.

[0188] Clause 9: A method for generating a segmented image of bodily structures; the method comprising:

[0189] determining a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to detect the bodily structures within the image;

[0190] producing a cropped depth image containing the bodily structures from a depth image comprising depth data associated with the scene containing the bodily structures; and

[0191] generating a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

[0192] Clause 10: Machine-readable storage storing machine-instruction for generating a segmented image of bodily structures; the machine-instructions comprising:

[0193] instructions to determine a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to define the region of interest as comprising any detected predetermined bodily structures within the image;

[0194] instructions to generate a segmented image comprising one or more masks associated with the detected predetermined bodily structures using a second neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth data being defined by the region of interest.

[0195] Clause 11 : The machine-readable storage of clause 10, in which the instructions to generate a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment depth data of a depth image, the depth data being defined by the region of interest, comprises:

[0196] instructions to produce a cropped depth image of bodily structures from a depth image comprising depth data associated with the scene containing bodily structures;

[0197] instructions to generate the segmented image comprising the one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

[0198] Clause 12: The machine-readable storage of any of clauses 10 to 11 , in which the depth image is derived from a 3D point cloud.

[0199] Clause 13: The machine-readable storage of any of clauses 10 to 12, in which the image of the scene is a colour image, optionally, an RGB image.

[0200] Clause 14: The machine-readable storage of any of preceding clause, in which the depth data and the image of the scene are derived from an RGB-D camera.

[0201] Clause 15: Machine-readable storage storing machine-instructions for generating a segmented image of bodily structures; the machine-instructions comprising:

[0202] instructions to generate a segmented image comprising one or more masks respectively associated with one or more than one detected predetermined bodily structure using a neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth image being derived from a scene comprising at least one bodily structure of the predetermined bodily structures.

[0203] Clause 16: The machine-readable storage of clause 15, in which the depth data of the depth image is associated with a 3D point cloud.

[0204] Clause 17: The machine-readable storage of either of clauses 6 and 7, in which the depth data of the depth image is a 2D map of depth data derived from the 3D point cloud, or the 3D point cloud per se.

[0205] Clause 18: Machine-readable storage storing instructions to generate a segmented image of bodily structures; the machine-instructions comprising:

[0206] instructions to determine a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to detect the bodily structures within the image;

[0207] instructions to produce a cropped depth image of bodily structures from a depth image comprising depth data associated with the scene containing bodily structures; and

[0208] instructions to generate a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

[0209] Clause 19: An image processing system for generating a segmented image of bodily structures; the image processing system comprising:

[0210] circuitry to determine a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to define the region of interest as comprising any detected predetermined bodily structures within the image; and

[0211] circuitry to generate a segmented image comprising one or more masks associated with the detected predetermined bodily structures using a second neuralnetwork trained to segment depth data of a depth image according to predetermined bodily structures; the depth data being defined by the region of interest.

[0212] Clause 20: The image processing system of clause 19, in which the circuitry to generate a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment depth data of a depth image, the depth data being defined by the region of interest, comprises:

[0213] circuitry to produce a cropped depth image of bodily structures from a depth image comprising depth data associated with the scene containing bodily structures;

[0214] circuitry to generate the segmented image comprising the one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

[0215] Clause 21 : The image processing system of any of clauses 19 to 20, in which the depth image is derived from a 3D point cloud.

[0216] Clause 22: The image processing system of any of clauses 19 to 21, in which the image of the scene is a colour image, optionally, an RGB image.

[0217] Clause 23: The image processing system of any of clauses 19 to 22, in which the depth data and the image of the scene are derived from an RGB-D camera.

[0218] Clause 24: An image processing system for generating a segmented image of bodily structures; the image processing system comprising:

[0219] logic for generating a segmented image comprising one or more masks respectively associated with one or more than one detected predetermined bodily structure using a neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth image being derived from a scene comprising at least one bodily structure of the predetermined bodily structures.

[0220] Clause 25: The image processing system of clause 24, in which the depth data of the depth image is associated with a 3D point cloud.

[0221] Clause 26: The image processing system of either of clauses 24 and 25, in which the depth data of the depth image is a 2D map of depth data derived from the 3D point cloud, or the 3D point cloud per se.

[0222] Clause 27: An image processing system for generating a segmented image of bodily structures; the image processing system comprising a processor configured to:

[0223] determine a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to detect the bodily structures within the image;

[0224] produce a cropped depth image of bodily structures from a depth image comprising depth data associated with the scene containing bodily structures; and

[0225] generate a segmented image comprising one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

[0226] Clause 28: An image processing system for markerless patella-femoral joint identification, the system comprising:

[0227] a machine learning interface to a multiclass classification deep learning model; the machine learning interface being arranged to receive an input vector; the input vector comprising at least one, or both, of: image information (RGB) and depth information associated with a patella-femoral joint; the multiclass classification deep learning model being trained to generate multiclass classification data; the multiclass classification data comprising at least:

[0228] segmented image data comprising at least one mask corresponding to a respective at least one member of the knee joint one of which being the patella, and

[0229] a machine learning output interface of the multiclass classification deep learning model; the machine learning output interface being arranged to output the multiclass classification data.

[0230] Clause 29: The system of clause 28, in which the multiclass classification data comprises at least one or more than one of: a segmented image associated with a first member, such as, for example, the femur, of the at least one member, a segmented image associated with a second member such as, for example, the tibia, of the at least one member, a segmented image associated with a third member such as, for example, patella, of the at least one member, a segmented image associated with a backgroundassociated with at least one member, taken jointly and severally in any and all permutations.

[0231] Clause 30: The system of either of clause 28 and 29, comprising the multiclass classification deep learning model.

[0232] Clause 31 : The system of any of clauses 28 to 30, comprising a region of interest neural network trained to generate the image information from an original image associated with an image camera.

[0233] Clause 32: The system of clause 31 , in which the image information comprises a cropped image derived from the original image comprising the at least one member; the cropped image comprising the at least one member.

[0234] Clause 33: The system of either of clauses 31 and 32, in which the region of interest neural network is also trained to generate the image information using the depth data associated with a depth camera.

[0235] Clause 34: The system of any of clauses 28 to 33, in which the depth information comprises depth data associated with the at least one member derived from an original depth image associated with the at least one member.

[0236] Clause 35: The system of clause 34, in which the depth information comprising the depth data associated with the at least one member is derived a cropped version of the original depth image associated with the at least one member.

[0237] Clause 36: The system of any of clauses 28 to 35, comprising at least one display for displaying the multiclass classification data.

[0238] Clause 37: The system of any of clauses 28 to 36, comprising a camera for generating at least one, or both, of: the original image and the original depth image.

Claims

CLAIMS1. An image processing system for markerless patella-femoral joint identification, the system comprising: a. a machine learning interface to a multiclass classification deep learning model; the machine learning interface being arranged to receive an input vector; the input vector comprising at least one, or both, of: image information and depth information associated with a patella-femoral joint; the multiclass classification deep learning model being trained to generate multiclass classification data; the multiclass classification data comprising at least: i. segmented image data comprising at least one mask corresponding to a respective at least one member of the knee joint one of which being the patella, and b. a machine learning output interface of the multiclass classification deep learning model; the machine learning output interface being arranged to output the multiclass classification data.

2. The system of claim 1 , in which the multiclass classification data comprises at least one or more than one of: a. a segmented image associated with a first member of the at least one member, b. a segmented image associated with a second member of the at least one member, c. a segmented image associated with a third member of the at least one member, d. a segmented image associated with a background associated with at least one member, taken jointly and severally in any and all permutations.

3. The system of claim 1 , comprising the multiclass classification deep learning model.

4. The system of claim 1 , comprising a region of interest neural network trained to generate the image information from an original image associated with an image camera.

5. The system of claim 4, in which the image information comprises a cropped image derived from the original image comprising the at least one member; the cropped image comprising the at least one member.

6. The system of either of claims 4 and 5, in which the region of interest neural network is also trained to generate the image information using the depth data associated with a depth camera.

7. The system of claim 1 , in which the depth information comprises depth data associated with the at least one member derived from an original depth image associated with the at least one member.

8. The system of claim 7, in which the depth information comprising the depth data associated with the at least one member is derived a cropped version of the original depth image associated with the at least one member.

9. The system of claim 1 , comprising at least one display for displaying the multiclass classification data.

10. The system of claim 1 , comprising a camera for generating at least one, or both, of: the original image and the original depth image.11 . A method for generating a segmented image of bodily structures; the method comprising: a. determining a region of interest from an image of a scene containing the bodily structures; said determining using a first neural network trained to define the region of interest as comprising any detected predetermined bodily structures within the image; b. generating a segmented image comprising one or more masks associated with the detected predetermined bodily structures using a second neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth data being defined by the region of interest.

12. The method of claim 11 in which the generating a segmented image comprising one or more masks associated with the bodily structures using a second neuralnetwork trained to segment depth data of a depth image; the depth data being defined by the region of interest comprises: a. producing a cropped depth image of bodily structures from a depth image comprising depth data associated with the scene containing bodily structures; b. generating the segmented image comprising the one or more masks associated with the bodily structures using a second neural network trained to segment the cropped depth image.

13. The method of any of either of claims 11 to 12, in which the depth image is derived from a 3D point cloud.

14. The method of claim 11, in which the image of the scene is a colour image.

15. The method of claim 11, in which the depth data and the image of the scene are derived from an RGB-D camera.

16. Machine-readable storage storing machine-instructions for generating a segmented image of bodily structures; the machine instructions comprising: a. Instructions for generating a segmented image comprising one or more masks respectively associated with one or more than one detected predetermined bodily structure using a neural network trained to segment depth data of a depth image according to predetermined bodily structures; the depth image being derived from a scene comprising at least one bodily structure of the predetermined bodily structures.

17. The machine-readable storage of claim 16 in which the depth data of the depth image is associated with a 3D point cloud.

18. The machine-readable storage of either of claims 16 and 17, in which the depth data of the depth image is a 2D map of depth data derived from the 3D point cloud.

Citation Information

Patent Citations

  • Patellar tracking

    CN101484085B

  • Patella Tracking

    US20190388159A1

  • Patella tracking method and system

    US20210315640A1

  • Patella tracking method and apparatus for use in surgical navigation

    US8571637B2