Apparatus and method for enhancing image of eye during medical procedure
By acquiring eye images and location information, and using microscopes and neural networks for image difference recognition, the problem of insufficient information in existing ophthalmic surgical systems is solved, achieving real-time augmented reality image enhancement effects and improving the accuracy and efficiency of surgery.
Patent Information
- Application Number
- CN202480044392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-10
- Filing Date
- 2024-07-10
- Publication Date
- 2026-02-10
AI Technical Summary
Existing ophthalmic surgery systems only provide general and/or static information, lack augmented reality support, and cannot effectively enhance the information support for surgeons during ophthalmic surgery.
By acquiring a first image of the eye and related positional information, combining the first and second images for difference recognition, determining second positional information, and providing this information to the user along with the second image, augmented reality display of the image is achieved. This system can be implemented in hardware or software, utilizing a microscope and camera to acquire images, and employing neural networks for image alignment and enhancement.
It enables accurate image enhancement during ophthalmic surgery, providing real-time, dynamic augmented reality information support, and improving the accuracy and efficiency of surgeons' operations.
Smart Images

Figure CN121511048A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to apparatus and methods for enhancing images of the eye during medical procedures. Additionally, a microscope for image enhancement is disclosed. Background Technology
[0002] Eye surgery is a medical procedure that treats a variety of eye conditions, including cataracts, glaucoma, refractive errors, and other vision-related problems. Intraocular lens (IOL) surgery is a common type of eye surgery that involves replacing the eye's natural lens with an artificial one. This procedure is often performed to treat cataracts, which occur when the natural lens becomes cloudy and causes visual impairment.
[0003] Eye surgeries sometimes utilize augmented reality (AR) information, typically displayed to the surgeon on a separate monitor. Furthermore, preoperative information is used to support the surgeon during the procedure. However, existing systems only provide general and / or static information. Improvements are desired. Summary of the Invention
[0004] The purpose of this disclosure is to improve ophthalmic medical procedures with enhanced information.
[0005] This objective is achieved through the disclosed embodiments, which are specifically defined by the subject matter of the independent claims. The dependent claims provide further embodiments. The following summary and description also disclose aspects and embodiments thereof, which provide additional features and advantages.
[0006] The first aspect of this disclosure relates to a device for enhancing images of the eye during medical procedures.
[0007] Configured as:
[0008] - Obtain the first image of the eye;
[0009] - Obtain first location information associated with the first image, wherein the first location information is related to at least one of the state of the eye or a medical procedure to be performed at the eye;
[0010] - Obtain a second image of the eye during a medical procedure;
[0011] - Determine the second location information based on the first image, the first location information, and the second image;
[0012] - Provide the second location information to the user along with the second image.
[0013] Devices for enhancing images of the eye can be implemented in hardware and / or software. In particular, such devices can be implemented solely in software and configured to operate on general-purpose computers and / or specialized hardware.
[0014] Medical procedures can include a series of actions or steps taken by a medical professional to diagnose, treat, or prevent a medical condition. The range of medical procedures can be from simple and / or minimally invasive procedures (such as physical examinations) to more complex and invasive procedures (such as surgery). Medical procedures can involve the use of a variety of medical tools and techniques, including diagnostic tests, imaging equipment, and / or surgical instruments. They can be performed in a wide variety of settings, including hospitals, clinics, and outpatient facilities. An example of a medical procedure is cataract surgery, which involves removing the cloudy lens from the eye and replacing it with an artificial lens implant (such as an IOL). Another example of a medical procedure is glaucoma surgery, which involves lowering the intraocular pressure within the eye, for example, through a drainage implant.
[0015] The first image can be obtained before the medical procedure. For example, the first image can be taken during an examination procedure in preparation for the medical procedure. Alternatively or alternatively, the first image can be obtained during another medical procedure (specifically, during another surgery). The device according to the first aspect can receive the first image from an external source. Alternatively or alternatively, the device can acquire the first image from an external source (e.g., shared memory). Furthermore, the device can autonomously generate the first image. In this case, the device can include a camera. For example, the device can be a microscope with a camera.
[0016] Within the first image of the eye, biometric features of the eye can be identified. These biometric features can be, for example, a portion of the iris or a portion of the pupil. These two organs have their own distinct patterns and structures. Furthermore, biometric features can be or include portions of one or more blood vessels, such as those in the sclera.
[0017] The first location information may relate to the medical condition of the eye, such as one or more astigmatic axes. Additionally or alternatively, the first location information may relate to information about where examinations and / or medical treatments should be performed (e.g., incision point, incision path, and / or IOL placement). Location information may include location, distance, region, perimeter, orientation, and / or direction. Additionally or alternatively, the first location information may also relate to gradients, particularly gradients with respect to location or time.
[0018] The second image acquired during a medical procedure may be obtained using the same or different equipment as the first image. Specifically, the second image may be obtained using a microscope or a camera attached to a microscope. The second image (and the first image) may capture the entire eye or only one or more parts of the eye. The device according to the first aspect may receive and / or acquire the second image from an external source. Furthermore, the device may autonomously generate the second image. In this case, the device may include a camera. For example, the device may be a microscope equipped with a camera.
[0019] The second positional information may include first positional information aligned with the second image. For example, the astigmatic axis of the first positional information identified as the first image is converted into second positional information to fit the second image. This can be done by taking information from the first image that is also present in the second image and determining the differences between the two images. Based on this difference, the second positional information is then determined as the first positional information that is converted into the background of the second image. Determining the differences can be difficult because the second image may differ from the first image in terms of viewpoint, resolution, detail, and / or noise, for example, by the light from the microscope or camera that captured the second image. Therefore, the information from the first image and the second image used to determine the differences in the size and orientation of the eye in the two images can be multifaceted. For example, it may include positional or chromatic information of certain parts of the eye (e.g., blood vessels, eyelids, and / or structures in the pupil). Alternatively or alternatively, the differences between the two images can be calculated using a normalized artificial scaling present in both the first and second images.
[0020] The second location information, together with the second image, can be provided in a format such that it can be displayed to the user as augmented reality. The second location information can be provided as augmented information within the second image. Alternatively or separately, the second location information can be provided separately from the second image so that the two types of information can be displayed independently of each other. In an embodiment, the device can be configured to adapt a first image based on the second image and display the adapted first image together with the second image to the user. Advantageously, the second image can be enhanced with the first image, for example, by overlay and / or by picture-in-picture arrangement.
[0021] Advantageously, accurate enhancement of images captured during medical procedures is feasible using the equipment according to the first aspect.
[0022] The first aspect of the implementation plan involves a device.
[0023] Configured as:
[0024] - Determine at least one of the scaling difference, rotation difference, or translation difference between the first image and the second image; and
[0025] - Determine second location information on at least one of the identified differences.
[0026] Geometric differences can be determined by visual comparison between the first and second images. Visual comparison can be based on positional and / or chromatic information within the images. For example, visual comparison can be based on biometric structures within the eye that can be identified in both the first and second images. This could be, for example, the trajectory of the eyelid, blood vessels within the eye, or invariant forms within the pupil. Advantageously, alignment based on the first positional information of geometrical differences provides efficient enhancement of the second image.
[0027] The first aspect of the implementation plan involves a device.
[0028] It is configured to determine the second location information based on non-iterative difference recognition.
[0029] A rough fit can provide a fast and computationally efficient determination of starting values for optimizing alignment for the second image.
[0030] The first aspect of the implementation plan involves a device.
[0031] Configured to determine the second location information based on the following:
[0032] - The seating position of the person performing the medical procedure; and / or - Whether the eye is the left or right eye.
[0033] Seating position and / or eye position can be specifically provided by the surgeon or other personnel as manual input. This allows for, advantageously, more accurate alignment.
[0034] The first aspect of the implementation plan involves a device.
[0035] Configured to adapt the second location information based on the following:
[0036] - Scaling of the eye in the first image relative to the eye in the second image based on biometric features of the eye;
[0037] - Based on whether the eye is the left or right eye in the first image, the rotation of the eye relative to the eye in the second image;
[0038] - Translation of the eye in the first image relative to the eye in the second image, based on the pupil center of the eye.
[0039] Scaling-based adaptation can be performed, for example, based on scaling-related features of a first and a second image, which can be identified as representing the same feature. This allows for the identification of key structures in the first and second images, such as eyelids. Based on the identified key structures, the first image can be scaled to match the second image. This method can also be applied to matching rotations and / or translations of the first and second images. Based on this, more accurate alignment is possible.
[0040] The first aspect of the implementation plan involves a device.
[0041] It is configured to determine the second location information based on iterative difference identification.
[0042] Iterative methods can be used to improve the estimation of geometric differences between a first and a second image. Iterative methods can include deterministic optimization and / or statistical optimization. They can also include gradient descent algorithms and / or machine learning algorithms. Advantageously, accurate alignment can be obtained based on iterative methods.
[0043] The first aspect of the implementation plan involves a device.
[0044] It is configured to perform iterative difference recognition based on non-iterative difference recognition.
[0045] Non-iterative difference identification can provide starting values for iterative difference identification. Specifically, a coarse but computationally efficient difference identification can be used to obtain one or more starting values for computationally more complex but also more precise iterative alignment. In this way, complex iterative alignment will also be performed efficiently, as it only processes one or more small differences to obtain secondary positional information.
[0046] The first aspect of the implementation plan involves a device.
[0047] Configured as:
[0048] - Obtain a third image during a medical procedure;
[0049] - Determine the third location information based on the second location information and the third image;
[0050] - Provide the user with third location information along with a third image.
[0051] The third image can capture the entire eye or only one or more parts of the eye. Through the third image, further changes in the condition that may occur during the medical procedure can be considered. Such changes may, for example, be related to incisions and / or lens placement. The third image can also be the next image in an image stream provided by a camera that also provides the second image. The third image may also be followed by other images. Alignment of positional information with the third image can be achieved in the same manner as alignment of positional information with the second image.
[0052] Advantageously, augmentation information can be aligned with each frame of the video stream during medical procedures. If the differences between images are small during the medical procedure, difference recognition can be based on accurate iterative alignment. Thus, accurate alignment of location information with images can be provided in real time during the medical procedure.
[0053] The first aspect of the implementation plan involves a device.
[0054] Configured as:
[0055] - Determining the second position information based on both non-iterative and iterative methods; and
[0056] - The third position information is determined based on the iterative method.
[0057] The first aspect of the implementation plan involves a device.
[0058] It is configured to determine at least one of a first location information, a second location information, or a third location information through a neural network (particularly a convolutional neural network and / or a regressive neural network).
[0059] Precise alignment can be achieved using neural networks, particularly with high accuracy in calculating second and / or third positional information. Neural networks can include standard neural networks such as VGG and / or ResNET. VGG is a convolutional neural network architecture characterized by the use of small 3x3 convolutional filters and deep stacking of convolutional layers. ResNET is a deep convolutional neural network architecture that uses residual connections to train very deep neural networks (up to hundreds of layers deep). In particular, convolutional neural networks can be used for efficient image processing. To determine continuous variables such as scaling, translation, and rotation, fully connected single-layer or multi-layer regression neural networks can be used. Regression neural networks can be feedforward neural networks, such as multilayer perceptrons. Regression neural networks can be specifically configured to receive their inputs from convolutional neural networks. This allows for the calculation of geometric differences with high accuracy.
[0060] The first aspect of the implementation plan involves a device.
[0061] The neural network is configured to determine at least one of the following:
[0062] - Bounding box;
[0063] - Angle or a value related to the angle.
[0064] A bounding box can be a rectangular box that encloses a portion of the eye and / or a region of interest within the eye in an image. For example, a bounding box could enclose the lens and / or incision traces. By identifying and locating objects using bounding boxes, computer algorithms can be trained to recognize and classify objects in a supervised learning environment. The bounding box does not need to be rectangular and can have other shapes such as squares, circles, trapezoids, and / or polygons. Angles or angle-related values can be values in, for example, degrees, radians, or gradients. Furthermore, angle-related values can be trigonometric function values, such as sine and / or cosine values.
[0065] The first aspect of the implementation plan involves a device.
[0066] Configured to determine at least one of a first location information, a second location information, or a third location information based on optimization, wherein the optimization includes at least one of the following:
[0067] - Pixel-wise loss function; or
[0068] - Perceptual loss function.
[0069] A pixel-wise loss function measures the difference between two images by comparing the values of individual pixels in each image (e.g., a first image and a second image). In a pixel-wise loss function, the differences between pixels in the two images are calculated, and then the differences of all pixels are aggregated to obtain a scalar value representing the overall difference between the images. A pixel-wise loss function may include mean squared error and / or mean absolute error.
[0070] The perceptual loss function is designed to measure the perceptual similarity between two images (e.g., between a first image and a second image). The perceptual loss function may, for example, include a measure of content loss. Content loss is based on feature representations. Feature representations capture high-level information about the content of an image, such as edges, textures, and objects. The output of a convolutional neural network can be used as a feature representation. The content loss function can measure the distance between the feature representations of two images, for example, by using mean squared error. Alternatively or alternatively, the perceptual loss function may include a style loss function that measures the similarity of texture and color patterns between two images.
[0071] The first aspect of the implementation plan involves a device.
[0072] The neural network was trained manually, at least in part.
[0073] Hand-labeled bounding boxes can be inserted into the training images. The difference between the estimated bounding boxes and the hand-labeled bounding boxes can then be calculated and used as the loss value for training the network. This can be specifically based on the cross-union loss function. Alternatively or alternatively, other parameters used to align different images can be based on self-supervised learning, such as rotational differences and / or translational differences.
[0074] A second aspect of this disclosure relates to a method for enhancing images of the eye during medical procedures.
[0075] Use the following steps:
[0076] - Obtain the first image of the eye before medical procedures;
[0077] - Prior to a medical procedure, first location information related to an image is obtained, wherein the first location information is related to at least one of the state of the eye or the medical procedure to be performed at the eye;
[0078] - Obtain a second image of the eye during a medical procedure;
[0079] - Determine the second location information based on the first image, the first location information, and the second image;
[0080] - Provide the second location information to the user along with the second image.
[0081] A third aspect of this disclosure relates to a microscope.
[0082] Configured as:
[0083] - Perform according to the method in the second aspect; or
[0084] - Interact with a device based on one of the first aspects. Attached Figure Description
[0085] Further advantages and features arise from the following embodiments, some of which are illustrated in the accompanying drawings. The drawings are not always to scale. The dimensions of the various features may be enlarged or reduced accordingly, particularly for clarity of description. Therefore, they are shown, and in part illustrated, as follows:
[0086] Figure 1 - According to the embodiments of this disclosure, the positional information of the first image and the second image are aligned;
[0087] Figure 2 - A flowchart for generating second position information within a second image, according to an embodiment of this disclosure;
[0088] Figure 3- Two neural network systems for identifying rotational displacement and pupil region according to embodiments of this disclosure;
[0089] Figure 4 - A runtime graph for feature recognition in two consecutive images according to an embodiment of this disclosure;
[0090] Figure 5 - A microscope with enhanced image rendering according to an embodiment of this disclosure.
[0091] In the following description, reference is made to the accompanying drawings, which form part of this disclosure and illustrate specific aspects that may help to understand this disclosure. The same reference numerals indicate the same or at least functionally or structurally similar features.
[0092] Generally, the disclosure of the described methods also applies to the corresponding device (or apparatus) performing the method or to a corresponding system comprising one or more devices, and vice versa. For example, if specific method steps are described, the corresponding device may include features for performing the described method steps, even if such features are not explicitly described or represented in the figures. On the other hand, if a specific device is described, for example, based on functional units, the corresponding method may include one or more steps for performing the described function, even if such steps are not explicitly described or represented in the figures. Similarly, corresponding device features or features for performing specific method steps may be provided for the system. Unless otherwise expressly stated, features of the various exemplary aspects and embodiments described above or below may be combined. Detailed Implementation
[0093] Figure 1 The figure illustrates the alignment of position information of a first image 100a and a second image 100b, which are outputs of a device according to the first aspect of this disclosure.
[0094] The first image 100a provides a visual perspective of the eye, including the eyelid 102, sclera 104, blood vessels 105, iris 106, and pupil 108. Location and chromaticity information of the detected features of the eye are stored as first location information. The first image 100a was captured during the patient's preoperative evaluation. In this case, the patient has cataracts. During the preoperative scan, two astigmatic axes 110a and 112a were identified and correlated with the first image 100a. This information is also stored as first location information. Furthermore, during the preoperative scan, the visual axis 114a, which connects the center of the retinal region of highest visual acuity (fovea) to the center of corneal curvature, was also identified and additionally assigned to the first location information. The identified information will be important to the surgeon in subsequent surgery. During the surgery, the naturally deteriorating lens behind the pupil of the patient's eye will be removed, and an artificial lens (IOL) will be implanted.
[0095] The second image 100b shows the eye from a slightly different perspective. In this perspective, the eyelid 102 is not visible. Furthermore, the second image 100b is shifted by approximately 90° relative to the first image 100b. Only the sclera 104, iris 106, and pupil 108 are visible. The second image 100b was captured at the start of the surgery.
[0096] Based on the position and chromaticity information of blood vessels 105, scaling, rotation, and translation differences between the first and second images are determined. (Alternatively or alternatively, any features common to both images, such as structures in the iris, scleral pigmentation, etc., can be used.) Based on the determined difference information, astigmatic axes 110b and 112b and visual axis 114b are determined in the second image. As can be seen in the figure, the aligned astigmatic axes 110b and 112b are shifted by approximately 90°. Furthermore, various parts of the eye 104, 106, and 108 are also determined in the second image. As can be seen in the second image, the pupil 108 is larger relative to the iris 106. Based on the determined alignment, positional information (specifically, astigmatic axes 110b, 112b, and visual axis 114b) can be provided as enhancement information in the second image. Thus, the surgeon receives crucial information for the success of the medical procedure.
[0097] Figure 2 The figure illustrates a flowchart of a method 200 for generating second position information within a second image according to an embodiment of the present disclosure.
[0098] In the first step 210, a first image of the eye is obtained, for example... Figure 1 The first image 100a is shown in the diagram. The first image can be passively received and / or actively acquired by the computing device performing the method. Furthermore, the first image can be acquired during preoperative examinations prior to medical procedures (such as surgery). The first image can be captured by any imaging device.
[0099] In the second step 220, first location information related to the first image is obtained. The first location information is related to at least one of the state of the eye and / or the medical procedure to be performed on the eye. Therefore, the first location information may include the astigmatic axis, the circumference of the pupil and / or sclera, an indication of the area in the eye related to the medical procedure, the incision point and / or incision trajectory, or the orientation in which the IOL should be located.
[0100] In the third step 230, a second image of the eye is captured. This step is performed during a medical procedure, for example, using a microscope used in the medical procedure. The second image is different from the first image. Therefore, the first positional information cannot be used within the background of the second image.
[0101] Therefore, in step 240, second location information is determined. The second location information is determined based on information provided by the first and second images. This information may, for example, be a biometric feature. If this feature is present in both the first and second images, it can be used to determine the geometric and chromatic differences between the first and second images. Once the geometric and / or chromatic transformations between the first and second images are determined, the second location information is determined as the first location information transformed into the background of the second image.
[0102] In step 240 of method 200, the determination of the second position information may include multiple sub-steps. In an embodiment of method 200, the first position information is aligned with the second position information in the second image in two sub-steps.
[0103] In the first sub-step 242, coarse alignment is performed. This defines initial values for scaling difference, rotation difference, and translation difference between the first and second images. The scaling difference is based on the size of the iris in both images. Iris size can be detected automatically. The rotation difference is determined based on detections of the left and right eyes (OD / OS) in both images. This information can be provided manually by a medical professional. The translation difference is determined based on the identified center of the detected pupil and / or the location based on other biometric features. This can also be automated.
[0104] Fine alignment is performed in the second sub-step 244. Fine alignment 244 is performed using iterative optimization. Minimization of the perceptual loss is then performed. The perceptual loss function is used in conjunction with a neural network for image alignment. The perceptual loss function is configured to measure the perceptual similarity between two first and second images. Alternatively, a pixel-wise loss function can also be implemented as an alternative to the perceptual loss function.
[0105] After determining the second location information, it is provided to the user along with the second image in step 250 for display. Therefore, the second location information is combined with the second image to provide an enhanced second image. The enhanced second image can be a single, separate image.
[0106] Alternatively or alternatively, the enhanced second image can be part of an image stream (i.e., a video stream). In this case, third positional information of a third image, which is a successor image to the second image, is determined. Furthermore, the third positional information is determined based on the second image, the third image, and the second positional information. This is accomplished in the same manner as determining the second positional information of the second image. In this way, other positional information of other images can also be provided, thereby achieving an enhanced video stream with continuously updated positional information regarding the state of the eye and / or regarding medical procedures.
[0107] Figure 3The figure illustrates a system with two neural networks as part of an embodiment of the device according to the first aspect. A first neural network 300b is configured to determine the pupil center. A second neural network 300a is configured to determine rotational displacement between two images (e.g., between a first image and a second image).
[0108] The first neural network 300a and the second neural network 300b are convolutional neural networks. They are trained using a perceptual loss function. The perceptual loss function can be implemented using a pre-trained convolutional neural network (CNN). A pre-trained CNN can be configured as a feature extractor to extract high-level features from an image.
[0109] A perceptual loss function typically consists of two main components: content loss and style loss. Content loss measures the difference between high-level features of the first and second images, while style loss measures the difference in low-level features (such as texture and color) between the two images. To implement the perceptual loss function, a pre-trained CNN (such as VGG-16 or ResNet) is chosen. Features are then extracted from the intermediate layers of the network. The content loss can then be computed by comparing the feature maps of the first and second images at specific layers of the network. The style loss can be computed by comparing the correlation between the feature maps of the first and second images across multiple layers of the network. Once the content and style losses are computed, they are combined to form the overall perceptual loss function. This perceptual loss function can be used to train the neural network. As an alternative to the perceptual loss function, a pixel-wise loss function can be implemented.
[0110] The first neural network 300a and the second neural network 300b are based on the VGG16 standard architecture of convolutional neural networks. However, unlike the standard VGG16 architecture, neural networks 300a and 300b include 13 convolutional layers and a single fully connected layer for the final regression analysis (the standard VGG16 includes 13 convolutional layers and 3 fully connected layers for the classifier).
[0111] The structure of the first neural network 300a is explained below. The entry layer includes two convolutional layers 302a, each with a size of 224x224x64. The first hidden layer includes a max-pooling layer 304a, which downsamples the initial image to a size of 112x112x128 using a 3x3 perceptual field of view. The first hidden layer also includes two convolutional layers 302b. The second hidden layer performs downsampling to 56x56x256 using the max-pooling layer 304b, followed by three convolutional layers 302c. The third hidden layer performs downsampling to 28x28x512 using the max-pooling layer 304c, followed by three convolutional layers 302d. The fourth hidden layer performs downsampling to 14x14x512 using the max-pooling layer 304d, followed by three convolutional layers 302e. The final max-pooling layer 304e performs further downsampling to 7x7x512. This layer is connected to a single 8x1x1 fully connected layer 306 for regression analysis.
[0112] The first neural network is trained on discrete frames of the surgical video as input and with corresponding bounding boxes of the iris and pupil as labels. Labels are collected from human-labeled data. During training, the first neural network is optimized via supervised backpropagation to maximize the intersection-over-union (IoU) ratio between the predicted bounding boxes of each surgical frame and the human-collected labels.
[0113] Now let's explain the operation of the first neural network. Following... Figure 2 Steps 242 and 244 align the first surgical image with the preoperative image. Subsequently, based on the alignment with the preoperative image, the first neural network 300a infers the center of the pupil 301 of the eye, deriving information about the top of the pupil, the left edge of the pupil, the maximum height of the pupil, and the maximum width of the pupil. The first neural network can also be used to infer the iris of the eye in the same manner. Therefore, the neural network can also be used in real-time on continuous images during medical procedures to provide continuous centering and / or scaling of the images (in which case no further alignment is performed).
[0114] The second neural network 300b has a similar structure. It includes 13 convolutional layers 312a, 312b, 312c, 312d, and 312e, which are transformed by 5 max-pooling layers 314a, 314b, 314c, 314d, and 314e. However, for regression analysis, the second neural network 300b uses 2x1x1 fully connected layers 316 instead of 8x1x1 fully connected layers.
[0115] The second neural network is also trained using discrete frames from the surgical video. Unlike the first neural network, the second neural network is trained using a semi-supervised method as follows: A copy of each surgical frame is generated, rotated by a randomly determined angle θ311. The original unrotated image and the copied rotated image are fed as input to the network. The network is then optimized via backpropagation to minimize the average difference between the predicted angle θ and its true value. In this way, human labels are not required.
[0116] The operation of the second neural network 300b will now be explained. A centered image (i.e., an image corrected by parameters obtained from the first neural network 300a) is used to infer the translational shift and scaling between the two images. Subsequently, based on the alignment of the translation and scaling, the second neural network 300b infers the angle θ or trigonometric function values (cos(θ), sin(θ)) of the image or the eye within the image. Therefore, the neural network can also be used in real-time on successive images during medical procedures to provide continuous rotational alignment of the image and / or the eye (in which case no further alignment is performed).
[0117] Figure 4 A runtime diagram of feature recognition in two consecutive images according to an embodiment of this disclosure is provided. This diagram is related to... Figure 3 An example of real-time alignment of related explanations. A first image 402 is provided to a machine learning algorithm. The first image includes first positional information 402a, 402b associated with two astigmatic axes. In a first step 404, the center of the pupil (or iris) depicted in the first image 402 is predicted. This can be achieved, for example, by... Figure 3 The convolutional neural network 300a is used to complete this. As a result, information about the center of the eye is obtained in step 406.
[0118] Simultaneously, the second image 412, captured after the first image, also undergoes center detection 414, producing a centered image 416. The angular difference is inferred based on the centered first and second images, for example, through a neural network 300b. Once the angle is inferred, the astigmatic axes 420a, 420b can be correlated with the second image as second positional information to obtain an enhanced second image 420.
[0119] Generally, in this method, the prediction for each frame is relative to the previous frame. The acquired information can be propagated back to the initial image, such as the image used for preoperative registration. This method can be used to enhance video streams during medical procedures and provide surgeons with crucial updated information related to the state of the eye or the medical procedures being performed.
[0120] Some implementations involve a microscope that includes... Figures 1 to 4One or more related systems described in the text. Alternatively, the microscope can be... Figures 1 to 4 It is part of or connected to one or more related descriptions of the system. Figure 5 A schematic diagram of a system 500 configured to perform the methods described herein is shown. System 500 includes a microscope 510 and a computer system 520. The microscope 510 is configured to capture images and is connected to the computer system 520. The computer system 520 is configured to perform at least a portion of the methods described herein. The computer system 520 may be configured to execute machine learning algorithms. The computer system 520 and the microscope 510 may be separate entities, or they may be integrated together in a common housing. The computer system 520 may be part of the central processing system of the microscope 510 and / or the computer system 520 may be part of a sub-component of the microscope 510, such as a sensor, actuator, camera, or illumination unit of the microscope 510.
[0121] Computer system 520 may be a local computer device (e.g., a personal computer, laptop, tablet, or mobile phone) having one or more processors and one or more storage devices, or it may be a distributed computer system (e.g., a cloud computing system with one or more processors and one or more storage devices distributed in various locations, such as local clients and / or one or more remote server farms and / or data centers). Computer system 520 may include any circuitry or combination of circuitry. In one embodiment, computer system 520 may include one or more processors, which may be of any type. As used herein, a processor may refer to any type of computing circuitry, such as, but not limited to, microprocessors, microcontrollers, complex instruction set computing (CISC) microprocessors, reduced instruction set computing (RISC) microprocessors, very long instruction word (VLIW) microprocessors, graphics processors, digital signal processors (DSPs), multi-core processors, field-programmable gate arrays (FPGAs), computing circuitry for microscopes or microscope components (e.g., cameras), or any other type of processor or processing circuitry. Other types of circuitry that may be included in computer system 520 may be custom circuitry, application-specific integrated circuits (ASICs), etc., such as one or more circuits (e.g., communication circuitry) used in wireless devices such as mobile phones, tablets, laptops, two-way radios, and similar electronic systems. Computer system 520 may include one or more storage devices, which may include one or more storage elements suitable for a particular application, such as main memory in the form of random access memory (RAM), one or more hard disk drives, and / or one or more drives that process removable media (such as optical discs (CDs), flash memory cards, digital video discs (DVDs), etc.). Computer system 520 may also include a display device, one or more speakers, and a keyboard and / or controller, which may include a mouse, trackball, touchscreen, voice recognition device, or any other device that allows a system user to input and receive information from computer system 520.
[0122] Some or all of the method steps can be performed by (or using) hardware devices, such as, for example, processors, microprocessors, programmable computers, or electronic circuits. In some embodiments, such devices can perform one or more of the most important method steps.
[0123] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be performed using a non-transient storage medium, such as a digital storage medium like a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, storing electronically readable control signals that cooperate (or are capable of cooperating with) a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0124] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0125] Typically, embodiments of the present invention can be implemented as a computer program product having program code operable to perform one of the methods when the computer program product is run on a computer. The program code may, for example, be stored on a machine-readable medium. Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein. In other words, therefore, an embodiment of the present invention is a computer program having program code for performing one of the methods described herein when the computer program is run on a computer.
[0126] Therefore, another embodiment of the invention is a storage medium (or data carrier, or computer-readable medium) including a computer program stored thereon for performing one of the methods described herein when executed by a processor. Data carriers, digital storage media, or recording media are typically tangible and / or non-transitional. Another embodiment of the invention is an apparatus as described herein, including a processor and a storage medium.
[0127] Therefore, another embodiment of the invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0128] Other embodiments include processing devices, such as computers or programmable logic devices, configured or adapted to perform one of the methods described herein.
[0129] Another implementation includes a computer on which a computer program for performing one of the methods described herein is installed.
[0130] Another embodiment of the invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.
[0131] In some embodiments, programmable logic devices (e.g., field-programmable gate arrays) may be used to perform some or all of the functions of the methods described herein. In some embodiments, field-programmable gate arrays may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0132] As used herein, the term “and / or” includes any and all combinations of one or more of the related listed items and may be abbreviated as “ / ”.
[0133] Although some aspects are described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a box or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of an item or feature of the corresponding box or corresponding apparatus.
[0134] Implementation can be based on the use of machine learning models or algorithms. Machine learning can refer to algorithms and statistical models that a computer system can use to perform a specific task without explicit instructions, relying instead on models and inference. For example, in machine learning, data transformations inferred from the analysis of historical and / or training data can be used instead of rule-based data transformations. For example, machine learning models or algorithms can be used to analyze the content of images. To use a machine learning model to analyze the content of an image, the model can be trained using training images as input and training content information as output. By training the machine learning model with a large number of training images and / or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine learning model “learns” to recognize the content of images, thus enabling it to identify image content not included in the training data. The same principle can be applied to other categories of sensor data: by training the machine learning model using training sensor data and the desired output, the model “learns” the transformation between sensor data and output, which can be used to provide outputs based on non-training sensor data provided to the machine learning model. The provided data (e.g., sensor data, metadata, and / or image data) can be preprocessed to obtain feature vectors, which are used as inputs to the machine learning model.
[0135] Machine learning models can be trained using training input data. The example detailed above uses a training method called "supervised learning." In supervised learning, a machine learning model is trained using multiple training samples, where each sample can include multiple input data values and multiple desired output values; that is, each training sample is associated with a desired output value. By specifying both the training samples and the desired output values, the machine learning model "learns" which output value to provide based on input samples similar to those provided during training. Besides supervised learning, semi-supervised learning can also be used. In semi-supervised learning, some training samples lack corresponding desired output values. Supervised learning can be based on supervised learning algorithms (such as classification algorithms, regression algorithms, or similarity learning algorithms). Classification algorithms are used when the output is restricted to a finite set of values (categorical variables), where the input is classified into one of a finite set of values. Regression algorithms are used when the output can have any numerical value (within a certain range). Similarity learning algorithms can be similar to classification and regression algorithms, but are based on learning from examples using a similarity function that measures the similarity or relevance between two objects. In addition to supervised or semi-supervised learning, unsupervised learning can also be used to train machine learning models. In unsupervised learning, input data can be provided (only), and unsupervised learning algorithms can be used to find structure in the input data (e.g., by grouping or clustering the input data to find commonalities). Clustering is the process of assigning input data, which includes multiple input values, into subsets (clusters) such that input values within the same cluster are similar according to one or more (predefined) similarity criteria, while input values included in other clusters are different.
[0136] Reinforcement learning is a third type of machine learning algorithm. In other words, reinforcement learning can be used to train machine learning models. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Rewards are calculated based on the actions taken. Reinforcement learning is based on training one or more software agents to select actions, thereby increasing the cumulative reward, making the software agents better at a given task (as demonstrated by the increase in rewards).
[0137] Furthermore, certain techniques can be applied to some machine learning algorithms. For example, feature learning can be used. In other words, machine learning models can be trained, at least in part, using feature learning, and / or machine learning algorithms can include feature learning components. Feature learning algorithms, also known as representation learning algorithms, can preserve information from their inputs but can also transform them in a way that makes them useful, often as a preprocessing step before performing classification or prediction. For example, feature learning can be based on principal component analysis or cluster analysis.
[0138] In some examples, anomaly detection (i.e., outlier detection) can be used, the purpose of which is to provide identification of suspicious input values by making them significantly different from most of the input or training data. In other words, machine learning models can be trained using anomaly detection at least in part, and / or machine learning algorithms can include anomaly detection components.
[0139] In some examples, machine learning algorithms can use decision trees as predictive models. In other words, machine learning models can be based on decision trees. In a decision tree, observations about an item (e.g., a set of input values) are represented by branches of the decision tree, and the output value corresponding to that item is represented by the leaves of the decision tree. Decision trees can support both discrete and continuous values as output values. If discrete values are used, the decision tree can be approximated as a classification tree; if continuous values are used, it can be approximated as a regression tree.
[0140] Association rules are another technique that can be used in machine learning algorithms. In other words, a machine learning model can be based on one or more association rules. Association rules are created by identifying relationships between variables in a large amount of data. Machine learning algorithms can identify and / or utilize one or more relationship rules that represent knowledge derived from the data. For example, these rules can be used to store, manipulate, or apply knowledge.
[0141] Machine learning algorithms are typically based on machine learning models. In other words, the term "machine learning algorithm" can broadly refer to a set of instructions that can be used to create, train, or use machine learning models. The term "machine learning model" can broadly refer to a data structure and / or set of rules that represent (e.g., based on training performed by a machine learning algorithm). In implementations, using a machine learning algorithm may mean using an underlying machine learning model (or multiple underlying machine learning models). Using a machine learning model may mean that the machine learning model and / or the data structure / rule set used as the machine learning model were trained by a machine learning algorithm.
[0142] For example, a machine learning model can be an artificial neural network (ANN). An ANN is a system inspired by biological neural networks, such as those found in the retina or brain. An ANN consists of multiple interconnected nodes and multiple connections between nodes, known as edges. There are typically three types of nodes: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can transmit information from one node to another. The output of a node can be defined as a (non-linear) function of its inputs (e.g., the sum of its inputs). The node's inputs can be used in the function based on the "weights" of the edges or the nodes that provide the inputs. The weights of the nodes and / or edges can be adjusted during the learning process. In other words, training an artificial neural network can involve adjusting the weights of the nodes and / or edges of the artificial neural network to achieve the desired output for a given input.
[0143] Alternatively, the machine learning model can be a Support Vector Machine (SVM), a Random Forest model, or a Gradient Boosting model. A Support Vector Machine (SVM), also known as a Support Vector Network, is a supervised learning model with associated learning algorithms that can be used to analyze data (e.g., in classification or regression analysis). An SVM can be trained by providing inputs with multiple training input values belonging to one of two categories. An SVM can be trained to assign new input values to one of the two categories. Alternatively, the machine learning model can be a Bayesian Network, a probabilistic directed acyclic graph model. A Bayesian Network can use a directed acyclic graph to represent a set of random variables and their conditional dependencies. Alternatively, the machine learning model can be based on a genetic algorithm, a search algorithm and heuristic technique that mimics the process of natural selection.
[0144] List of reference numerals
[0145] 100a First Image
[0146] 100b Second Image
[0147] 102 Eyelids
[0148] 104 Sclera
[0149] 105 blood vessels
[0150] 106 Iris
[0151] 108 pupils
[0152] 110a Astigmatic axis
[0153] 110b aligned astigmatic axis
[0154] 112a Astigmatic axis
[0155] 112b aligned astigmatic axis
[0156] 114a line of sight
[0157] 114b aligned visual axis
[0158] 200 methods
[0159] 210 Obtain the first image
[0160] 220 Obtain first location information
[0161] 230 Obtain the second image
[0162] 240 Determine the second location information
[0163] 242 Rough alignment
[0164] 244 Fine Alignment
[0165] 250 Display of second location information and second image
[0166] CNN for 300a Angle Difference Detection
[0167] CNN detected by 300b center
[0168] 301. Pupil size
[0169] 302a Convolutional layer at the entry layer
[0170] 302b First hidden layer of convolutional layer
[0171] 302c Second hidden layer convolutional layer
[0172] 302d convolutional layer with third hidden layer
[0173] 302e Convolutional layer with third hidden layer
[0174] 304a First hidden layer max pooling layer
[0175] 304b Second hidden layer max pooling layer
[0176] 304c Second hidden layer max pooling layer
[0177] 304d second hidden layer max pooling layer
[0178] 304e The final max pooling layer
[0179] 306 Fully Connected Layer
[0180] 311 Rotational Difference
[0181] 312a Convolutional Layer
[0182] 312b convolutional layer
[0183] 312c convolutional layer
[0184] 312d convolutional layer
[0185] 312e convolutional layer
[0186] 314a Max Pooling Layer
[0187] 314b max pooling layer
[0188] 314c max pooling layer
[0189] 314d max pooling layer
[0190] 314e max pooling layer
[0191] 316 Fully Connected Layer
[0192] 402 First Image
[0193] 402a First location information
[0194] 402b First Location Information
[0195] 404 Iris Detection
[0196] 406 Center Testing
[0197] 412 Second Image
[0198] 414 Center Testing
[0199] 416 Centralized Images
[0200] 420 Enhanced Second Image
[0201] 420a Second Position Information
[0202] 420b Second Position Information
[0203] 500 Microscope System
[0204] 510 Microscope
[0205] 520 Computer System
Claims
1. A device for enhancing images of the eye during medical procedures. Configured as: - Obtain the first image (100a) of the eye (102, 106, 108); - Obtain first location information (110a, 112a, 114a) associated with the first image (100a), wherein the first location information is associated with at least one of the state of the eye or a medical procedure to be performed at the eye; - A second image (100b) of the eye (106, 108) is obtained during the medical procedure; - Determine second location information (110b, 112b, 114b) based on the first image, the first location information, and the second image; - Provide the second location information to the user along with the second image.
2. The device according to the preceding claims, Configured as: - Determine at least one of the scaling difference, rotation difference, or translation difference between the first image and the second image; and - Determine the second location information (110b, 112b, 114b) on at least one of the identified differences.
3. The device according to any one of the preceding claims, It is configured to determine the second location information (110b, 112b, 114b) based on non-iterative difference recognition.
4. The device according to any one of the preceding claims, The second location information (110b, 112b, 114b) is configured to be determined based on the following: - The seating position of the person performing the medical procedure; and / or - Is the eye described the left or right eye? 5. The device according to any one of the preceding claims, It is configured to adapt the second location information (110b, 112b, 114b) based on the following: - Scaling of the eye in the first image relative to the eye in the second image based on the biometric features of the eye; - The rotation of the eye in the first image relative to the eye in the second image, based on whether the eye is the left or right eye; - Translation of the eye in the first image relative to the eye in the second image, based on the pupil center of the eye.
6. The device according to any one of the preceding claims, It is configured to determine the second location information based on iterative difference recognition (244).
7. The device according to any one of the preceding claims, It is configured to perform the iterative difference (244) identification based on the non-iterative difference identification (242).
8. The device according to any one of the preceding claims, Configured as: - A third image (412) is obtained during the medical procedure; - Based on the second location information (402a, 402b) and based on the third image, determine the third location information (420a, 420b); - Provide the third location information to the user along with the third image.
9. The device according to the preceding claims, Configured as: - The second location information (110b, 112b, 114b) is determined based on both non-iterative and iterative methods; and - The third location information is determined based on an iterative method.
10. The device according to any one of the preceding two claims, It is configured to determine at least one of the first location information, the second location information, or the third location information through a neural network, wherein the neural network is in particular a convolutional neural network and / or a regressive neural network.
11. The device according to any one of the preceding three claims, The neural network is configured to determine at least one of the following: - Bounding box; - Angle or an angle-related value (311).
12. The device according to any one of the preceding four claims, Configured to determine at least one of the first location information, the second location information, or the third location information based on optimization, wherein the optimization includes at least one of the following: - Pixel-wise loss function; or - Perceptual loss function.
13. The device according to any one of the preceding five claims, The neural network was at least partially trained manually.
14. A computer-implemented method for enhancing images of the eye during medical procedures, comprising the following steps: - Obtain first images (100a) of the eye (102, 106, 108) before the medical procedure; - Prior to the medical procedure, first location information (110a, 112a, 114a) related to the image is obtained, wherein the first location information is related to at least one of the state of the eye or the medical procedure to be performed on the eye; - Obtain a second image (100b) of the eye during the medical procedure; - Determine second location information (110b, 112b, 114b) based on the first image, the first location information, and the second image; - Provide the second location information to the user along with the second image.
15. A microscope, Configured as: - Perform the method according to the preceding claims; or - To interact with the device according to any one of the preceding claims.