Noise reduction in retina image

A computer-implemented method aligns and averages retinal images based on offset thresholds to suppress noise and preserve texture, addressing the challenge of noise in retinal imaging and improving image quality for clinical use.

JP2025100457AActive Publication Date: 2025-07-03OPTOS PLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024221715
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-18
Publication Date
2025-07-03
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Retinal imaging techniques face challenges in suppressing noise while preserving valuable low-contrast anatomical features and texture due to low signal-to-noise ratio and inherent noise in images, leading to difficulty in discriminating boundaries and reducing contrast.

Method used

A computer-implemented method processes a sequence of retinal images by aligning and averaging selected images based on offset thresholds to generate a noise-reduced image that retains texture, using techniques like convolutional neural networks for further noise filtering.

Benefits of technology

The method effectively suppresses noise while maintaining anatomical texture, enhancing image quality for clinical assessment and improving the performance of AI noise-removal algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100457000001_ABST
    Figure 2025100457000001_ABST
Patent Text Reader

Abstract

To generate an image that holds a texture related to the retina from which noise is removed.SOLUTION: A computer mounting method which processes an image sequence in one area of the retina and generates an averaging image in the area includes: determining each offset between a reference image and each comparison image with respect to each set of a reference image selected from the image sequence and each of comparison images being the remaining images in the sequence; comparing each offset with threshold to determine whether the offset is smaller than the threshold; selecting each comparison image in each set whose offset is determined to be smaller than the threshold; and using the selected comparison image so as to generate an averaging image in the area. The threshold is determined so that an averaging image shows more texture characteristics in the area of the retina than a reference averaging image generated from the image in the image sequence.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary embodiments of the present specification generally relate to techniques for suppressing noise in images of the retina of an eye, and more particularly, to processing a sequence of images of a region of the retina to generate a noise-reduced image of the region.

Background Art

[0002] Retinal imaging using various different modalities is widely used to detect pathological changes in the retina of an eye and physiological changes associated with aging. In all cases, for such changes to be detected properly, the retinal image is required to be in focus, have as little noise as possible, and have as few artifacts as possible. However, for patient safety, retinal imaging devices generally use low-power illumination sources, which may result in a lower signal-to-noise ratio (SNR) compared to imaging of other tissues. Additionally, noise is inherent in some imaging techniques. For example, optical coherence tomography (OCT) images tend to contain speckle noise, resulting in reduced contrast and difficulty in discriminating boundaries between strongly scattering structures. Fundus autofluorescence (FAF) images result from the detection of low-intensity levels of fluorescence and often have a low SNR due to a combination of photon noise and amplifier (readout) noise.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Noise can be suppressed by adapting the image capture process, for example, by using an averaging technique that records a plurality of image frames, aligns them with each other, and averages the aligned images to generate a noise removal image in which noise from the component images is averaged to a low level. Also, noise can be suppressed by post-capture image processing techniques that use, for example, a convolutional neural network (CNN) or other supervised machine learning algorithm. The CNN or other supervised machine learning algorithm is trained using the averaged (noise removal) image to remove noise from the captured retinal image. However, in both of these approaches, valuable low-contrast anatomical features, including texture related to the retina, tend to be suppressed or missing in the final noise removal image. Therefore, it is desirable to provide a technique for generating a noise-removed retinal image that retains texture related to the retina.

Means for Solving the Problems

[0004] According to a first exemplary aspect of the present specification, a computer-implemented method is provided for processing an image sequence of a region of the retina of an eye to generate an averaged image of the region. The computer-implemented method includes, for each combination of a reference image selected from the image sequence and each comparison image that is an image from the remaining images of the sequence, determining each offset between the reference image and each comparison image; comparing each determined offset with an offset threshold to determine whether the offset is less than the offset threshold; selecting each comparison image in each combination for which it is determined that the offset between the reference image and each comparison image is less than the offset threshold; and using the selected comparison images to generate an averaged image of the region. The offset threshold is defined such that when the image sequence includes at least one image offset by more than the threshold from the reference image and images each offset by less than the threshold from the reference image, the averaged image shows more texture indicating the structure of the first region of the retina than a reference averaged image generated from the images in the image sequence.

[0005] According to a second exemplary aspect of the present specification, a computer-implemented method is provided for training a machine learning algorithm to filter noise from a retinal image. The computer-implemented method includes generating ground truth training data by processing each sequence of a plurality of sequences of retinal images to generate a respective averaged retinal image, each averaged retinal image being generated according to the computer-implemented method of the first exemplary aspect; generating training input data by selecting a respective image from each of the image sequences; and using the ground truth training data and the training input data to train a machine learning algorithm to filter noise from a retinal image.

[0006] Also, according to a third exemplary aspect of the present specification, there is provided a computer program including computer-readable instructions that, when executed by at least one processor, cause the at least one processor to execute at least one method of the aforementioned first exemplary aspect or second exemplary aspect. The computer program may be stored in a non-transitory computer-readable storage medium (e.g., a computer hard disk, CD, etc.) or may be carried by a computer-readable signal.

[0007] Also, according to a fourth exemplary aspect of the present specification, there is provided a data processing device arranged to process an image sequence of a region of the retina of an eye to generate an averaged image of the region. The data processing device includes at least one processor and at least one memory storing computer-readable instructions. The computer-readable instructions, when executed by the at least one processor, cause the at least one processor to determine, for each combination of a reference image selected from the image sequence and each comparison image that is an image from the remaining images of the sequence, each offset between the reference image and each comparison image; compare each determined offset with an offset threshold to determine whether the offset is less than the offset threshold; select each comparison image in each combination for which it is determined that the offset between the reference image and each comparison image is less than the offset threshold; and use the selected comparison images to generate an averaged image of the region. The offset threshold is defined such that when the image sequence includes at least one image offset from the reference image by more than the threshold and each image offset from the reference image by less than the threshold, the averaged image shows more texture indicative of the structure of the first region of the retina than a reference averaged image generated from the images in the image sequence.

[0008] Also, according to a fifth exemplary aspect of the present specification, an ophthalmic imaging system is provided that includes an ophthalmic imaging device arranged to acquire an image sequence of a region of the retina of an eye, and a data processing device according to a fourth exemplary aspect arranged to process the image sequence acquired by the ophthalmic imaging device to generate an averaged image of the region of the retina.

[0009] Next, exemplary embodiments will be described in detail by way of non-limiting examples only, with reference to the accompanying drawings. Unless otherwise specified, like reference numerals indicate the same or functionally similar elements in other drawings.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

[0011] [First Exemplary Embodiment] FIG. 1 is a schematic diagram of an ophthalmic imaging system 100 according to an exemplary embodiment, the system including an ophthalmic imaging device 10 arranged to acquire an image sequence 20 of a common region 30 of the retina 40 of the eye 50. The ophthalmic imaging system 100 further includes a data processing device 60 arranged to process the images from the image sequence 20 to generate an averaged image 70 of the region 30 of the retina 40.

[0012] As schematically shown in FIG. 2, the data processing device 60 is arranged to generate an averaged image 70 by first selecting images from the image sequence 20 according to one or more predefined criteria. These images are first the reference image I Ref selected from the image sequence 20 and each image from the sequence 20 being considered for selection (referred to herein as comparison image I Comp ). For each combination of the reference image I Ref and the respective comparison image I Comp , it may be selected by determining the respective offset between the reference image I Comp and the respective comparison image I Ref . Each determined offset is compared with an offset threshold value, and it may be determined whether the offset is smaller than the offset threshold value. As an example, one or more predefined criteria may include the criterion that the magnitude t = |t| of the determined translational offset t between the comparison image I Comp and the reference image I Ref is smaller than a predetermined translational offset threshold value T. Additionally or alternatively, one or more predefined criteria may include the criterion that the magnitude of the determined rotational offset (i.e., relative rotation) Δφ between the comparison image I Comp and the reference image I Ref is smaller than a predetermined rotational offset threshold value Θ. The selected comparison images 25 are aligned with each other by the data processing device 60, and the data processing device 60 then, for example, calculates the average of the aligned images such that the respective pixel values of each pixel in the resulting averaged image 70 are the average (e.g., arithmetic mean) of the pixel values of the corresponding pixels at the corresponding positions in the aligned selected images 25 that correspond to the respective common positions within the region 30 of the retina 40. Alternatively, the data processing device 60 may calculate the average of the aligned images such that the respective pixel values of each pixel in the resulting averaged image 70 are the weighted average of the pixel values of the corresponding pixels at the corresponding positions in the aligned selected images 25 that correspond to the respective common positions within the region 30 of the retina 40.

[0013] In this process (an example of which will be described later), when the image sequence 20 includes at least one image offset by an offset greater than a threshold from the reference image I Ref and an image offset by each offset less than the threshold from the reference image I Ref , the resulting averaged image 70 shows more texture than the reference averaged image 75 generated from some or all of the images of the image sequence 20 without these images being selected in the manner described herein by one or more predefined criteria. In this context, the texture is the anatomical texture related to the retina and shows the anatomical structure in the region 30 of the retina 40. For example, as in this exemplary embodiment, when the image acquired by the ophthalmic imaging device 10 is a fundus autofluorescence (FAF) image, the structure can be defined by the spatial distribution of the phosphors throughout the region 30 of the retina 40. As another example, when the image acquired by the ophthalmic imaging device 10 is an OCT image, the structure can include the physical structure of one or more layers of the retina 40 in the region 30 of the retina 40. As a further example, when the image acquired by the ophthalmic imaging device 10 is a reflectance image (e.g., a red, blue, or green reflectance image) of the region 30 of the retina 40, the structure may include the upper surface of the retina in the region 30 such that the texture is the physical texture of the surface reflecting the topography of the surface.

[0014] Furthermore, the averaged image 70 tends to have a higher SNR than a single image from the image sequence 20. This is because the selected images 25 are averaged, partially canceling out the noise components within the selected images 25 with each other. Thus, the averaged image 70 generated by the data processing apparatus 60 of the exemplary embodiment is a noise-removed image of the region 30 of the retina 40, and retains texture better than a conventional noise-removed image (e.g., image 75) generated by averaging images within the sequence 20 that are not selected according to one or more criteria described herein. Thus, the image processing techniques described herein can generate a noise-removed and anatomically accurate retinal image in which important clinical data is preserved. Such noise-removed images are valuable not only to clinicians for assessing the health of the retina 40, but also for training an AI noise-removal algorithm that suppresses noise while reducing or avoiding texture loss in noise-removed images, which often occurs with the use of conventionally trained artificial intelligence (AI) noise-removal algorithms, as will be described in detail below.

[0015] In order to suppress noise while retaining the texture of the averaged image 70, in the processing of the image sequence 20, it is necessary to consider the following classes of imperfections in the image acquisition process and their effects on the acquired images. First, the reference image I Ref and the comparison image I CompIt may be associated by an affine transformation or a non - affine transformation (e.g., translation, rotation, or distortion in one or more directions), which can occur as a result of movement of the retina in a direction perpendicular to the imaging axis during image acquisition, rotation of the retina around the imaging axis, or movement of the retina due to a change in the line of sight of the eye. Second, the acquired image may have non - affine distortion in it. This can be caused by eye movements such as rapid horizontal saccades that occur during image capture. Third, in some of the images, there may be one or more local regions or patches of micro - distortion (micro - displacement), which can be caused by variations in the imaging system (e.g., a beam - scanning system including a polygon mirror or other scanning elements) or resulting optical - path variations between captures of different images. The optical - path variations can be caused by, for example, deviation from scanning linearity and / or mirror imperfection among other factors.

[0016] The ophthalmic imaging device 10 may be any kind of FAF imaging device well - known to those skilled in the art, such as a fundus camera, a confocal scanning laser ophthalmoscope (SLO), or an ultra - wide - field imaging device such as Optos® Daytona®, as in this exemplary embodiment, and is arranged to capture a sequence of FAF images of the same region of the retina 30. The FAF imaging device often uses short - to medium - wavelength visible - light excitation to collect emissions within a predefined spectral band (e.g., within a wavelength range of 500 - 750 nm) to form a luminance map that reflects the distribution of lipofuscin, which is the dominant fluorescence present in the retinal pigment epithelium (RPE). However, other excitation wavelengths such as near - infrared can be used to detect additional fluorescent pigments such as melanin. FAF is useful for the evaluation of various diseases involving the retina and RPE.

[0017] However, the form of the ophthalmic imaging device 10 is not so limited. In other exemplary embodiments, the ophthalmic imaging device 10 may take the form of, for example, an optical coherence tomography (OCT) imaging device such as a swept-source OCT (SS-OCT) imaging device or a spectral-domain OCT (SD-OCT) imaging device arranged to acquire OCT images of the region 30 of the retina 40, such as B-scans, C-scans, and / or En-face OCT images. In yet other exemplary embodiments, the ophthalmic imaging device 10 may include, for example, a scanning laser ophthalmoscope (SLO), a fundus camera, etc. arranged to capture reflectance images (e.g., red, green, or blue images) of the region 30 of the retina 40.

[0018] In FIG. 1, the ophthalmic imaging device 10 is shown acquiring a sequence of 10 images, but this number is shown by way of example only. The number of images included in the image sequence 20 is not limited to 10 and may be more or less than this. Also, note that the region 30 is a part of the retina 40 within the field of view (FoV) of the ophthalmic imaging device 10 and may be smaller than the portion of the retina 40 that extends into the FoV. As will be described in more detail below, the region 30 may be a segment of a larger region of the retina 40 where the ophthalmic imaging device 10 may be positioned to capture an image, and each image spans the FoV of the ophthalmic imaging device 10. Thus, each image of the image sequence 20 may be a segment (e.g., a rectangular image tile) of a respective larger retinal image of the retina 40 captured by the ophthalmic imaging device 10.

[0019] The data processing device 60 can be provided in any suitable form, such as, for example, a programmable signal processing hardware 300 of the type schematically shown in FIG. 3. The programmable signal processing device 300 includes a communication interface (I / F) 310 for receiving the image sequence 20 from the ophthalmic imaging device 10 and outputting the averaged image 70. The signal processing hardware 300 further includes a processor (for example, a central processing unit, CPU and / or a graphics processing unit, GPU) 320, a working memory 330 (for example, a random access memory), and an instruction store 340 that stores a computer program 345 including computer-readable instructions that, when executed by the processor 320, cause the processor 320 to execute the various functions of the data processing device 60 described herein. The working memory 330 stores a reference image I Ref , a comparison image I Comp1 ~I Comp9Stores information used by the processor 320 during the execution of the computer program 345, such as the calculated image offset, the image offset threshold, the selected comparison image 25, the determined similarity, the similarity threshold, and the results of various intermediate processes described herein. The instruction store 340 can include a ROM (e.g., in the form of an electrically erasable programmable read-only memory (EEPROM) or flash memory) in which computer-readable instructions are pre-stored. Alternatively, the instruction store 340 may include a RAM or a similar type of memory, and the computer-readable instructions of the computer program 345 can be input from a non-transitory computer-readable storage medium 350 in the form of a CD-ROM, DVD-ROM, etc., or a computer-readable signal 360 that conveys the computer-readable instructions, such as a computer program product. In any case, when the computer program 345 is executed by the processor 320, it causes the processor 320 to execute the functions of the data processing device 60 described herein. Thus, the data processing device 60 of the exemplary embodiment may include a computer processor 320 (or two or more similar processors) and a memory 340 (or two or more similar memories) that stores computer-readable instructions that, when executed by the processor(s), cause the processor(s) to process the image sequence 20 and generate an averaged image 70 of the region 30 of the retina 40 as described herein.

[0020] However, it should be noted that the data processing device 60 may alternatively be implemented as non-programmable hardware such as an ASIC, an FPGA, or other dedicated integrated circuits that execute the functions of the data processing device 60 described herein, or a combination of such non-programmable hardware and programmable hardware as described with reference to FIG. 3. The data processing device 60 may be provided as an independent product or as part of the ophthalmic imaging system 100.

[0021] Next, a process in which the data processing apparatus 60 of the present exemplary embodiment processes the image sequence 20 to generate the averaged image 70 of the region 30 of the retina 40 will be described with reference to FIG. 4.

[0022] In process S10 of FIG. 4, the processor 320 of the data processing apparatus 60 determines, for each combination of a reference image I selected from the image sequence 20 Ref and each comparison image I from the remaining images of the image sequence 20 Comp the respective offset between the reference image I Ref and each comparison image I Comp . Each offset may include a rotation of the comparison image I in the combination with respect to the reference image I in the combination Ref and / or a translation of the comparison image I in the combination with respect to the reference image I in the combination Comp . Ref Comp Ref Ref

[0023] The reference image I Ref can be selected from the image sequence 20 in a variety of different ways. The reference image I Ref may be, for example, an image that appears at a predetermined position in the image sequence 20, such as the beginning, end, or preferably the middle position of the image sequence 20 (e.g., the middle of the sequence). Alternatively, the reference image I Ref may be randomly selected from the images within the image sequence 20. However, in order to more reliably realize as many selection images 25 as possible and thereby make the noise suppression in the averaged image 70 (derived from the selection images 25) more effective, the reference image I Ref is preferably selected from the images within the image sequence 20 as the image for which the number of remaining images in the sequence offset by an amount less than a predetermined offset threshold from the reference image I Ref is the largest (or equal thereto). The reference image I RefFor example, for any image selected from the image sequence 20 to function as a base image, the respective translational offsets between the base image and each of the remaining images in the image sequence 20 can be determined, for example, by determining the positions of the respective peaks in the cross-correlation calculated between the base image and each of the remaining images. The translational offsets determined in this way can be represented as a set of points in a 2D plot. For each point in this plot, the respective number of the remaining points within a circle of radius T centered on that point is determined. When this is done for all the points in the plot, the point with the most (or equal) number of the remaining points within the circle centered on that point is selected. Then, from the image sequence, the image represented by the point whose offset from the base image was selected is selected as the reference image I Ref is selected as.

[0024] In this way, for each of the remaining images (i.e., images other than the base image) in the sequence 20, the processor 320 determines the respective translational offsets between that image and all the other images in the sequence, compares the determined translational offsets with the offset threshold T, identifies all the other images in the sequence whose translational offsets from the image are less than the translational threshold T, counts the identified images, and determines the respective counts by determining each count. Next, the processor 320 uses the determined counts to identify, among the remaining images in the sequence 20, the image with the highest (or equal) determined count as the reference image I Ref is identified as.

[0025] As in this exemplary embodiment, in the process S10 of FIG. 4, the processor 320, for each combination of the reference image I Ref and each comparison image I from some or all of the remaining images in the image sequence 20 Comp the respective translational offsets between the reference image I Ref and each comparison image I Comp and, between the reference image I Ref and each comparison image IComp Each rotational offset between them may be determined. However, in other exemplary embodiments, the processor 320 may use the reference image I Ref and each comparison image I from the remaining images of the image sequence 20 Comp For each combination of the reference image I Ref and each comparison image I Comp either the translational offset between the reference image I Ref and each comparison image I Comp or the rotational offset between the reference image I

[0026] The translational and rotational offsets between images may be determined using various techniques as described herein or well known to those skilled in the art. For example, if the relative rotation between the images within the image sequence 20 can be ignored, the translational offset t between the reference image I Ref and each comparison image I Comp may be determined by calculating the cross-correlation between the reference image I Ref and the comparison image I Comp and the position of the peak of the calculated cross-correlation is used to determine the respective translational offset t. Due to the self-similarity of the images within the image sequence 20, a relatively wide peak may occur in the calculated cross-correlation. Therefore, the images may be pre-processed (e.g., using an edge detector) before cross-correlating to sharpen the peak and enable a more accurate determination of the translational offset t.

[0027] Alternatively, the translational offset t between the reference image I Ref and each comparison image I Comp may first be determined by calculating the normalized cross-power spectrum of the reference image I Ref and the comparison image I Comp which may be expressed as F Ref F Comp * / |F Ref F Comp * | where F Refis the discrete Fourier transform of the reference image I Ref and F Comp * is the complex conjugate of the discrete Fourier transform of the comparison image I Comp . Next, the inverse Fourier transform of the normalized cross-power spectrum is calculated to provide a delta function having a maximum value at (Δx, Δy) where the comparison image I Comp is shifted by Δx pixels in the x-axis direction with respect to the reference image I Ref and by Δy pixels in the y-axis direction with respect to the reference image I Ref .

[0028] The above cross-correlation approach can be extended to allow for relative rotation between images in the image sequence 20 by determining each translational offset t between each of a plurality of rotated versions of the reference image I Ref and the comparison image I Comp . This can be done by calculating each cross-correlation between the reference image I Ref and each rotated version of the comparison image I Comp . As an example, and not by way of limitation, each comparison image I Comp is rotated between -6 degrees and +6 degrees in 0.1 degree increments (regardless of the presence or absence of the above preprocessing), generating 120 rotated versions of each comparison image I Comp , each of which is cross-correlated with the reference image I Comp along with the original (unrotated) comparison image I Ref . The -6 degrees to +6 degrees angular range and 0.1 degree increment are shown by way of example, and other ranges and / or angular increments can also be used. The rotation (if any) applied to the image cross-correlated with the reference image I Ref is taken to produce the cross-correlation having the highest peak among the 121 calculated cross-correlations, and may provide an indication of the rotational offset between the reference image I Ref and the comparison image I Comp . On the other hand, the position of the highest cross-correlation peak may be taken to provide an indication of the translational offset t between the reference image I Ref and the comparison image I Comp .

[0029] Reference image I Ref and comparison image I Comp The rotational offset Δφ and translational offset t between them may alternatively be calculated in the Fourier domain. This approach relies on two properties of the Fourier transform. First, the frequency spectrum is always centered at the origin regardless of the shift of the underlying image. Second, rotation of an image always results in a corresponding rotation of its Fourier transform. In another approach, the reference image I Ref and comparison image I Comp are converted to polar coordinates. The image I(x, y) is converted to I(r, θ), and the discrete Fourier transform F(u, v) of the image becomes F(ρ, φ), and as a result,

[0030]

Number

[0031] I Comp If the angle β is given such that I Ref is rotated clockwise by β from I Comp (r, θ) = I Ref (r, θ + β). This means the following.

[0032]

Number

[0033] The above equation is equivalent to the following equation.

[0034]

Number

[0035] Therefore, the following can be said.

[0036]

Number

[0037] These two characteristics can separate the rotational and translational components of image alignment and solve each independently. For a set of images that are shifted and rotated relative to each other, the rotational shift between the images can be separated by taking their Fourier transforms. The magnitude of the frequency components generated by this Fourier transform is treated as if it were the pixel values of a new set of images that require only rotational alignment. When rotational correction is required, it is applied to the underlying image for alignment. Next, the translational parameters are determined by the same algorithm. Thus, for the reference image I Ref and each comparison image I Comp the respective rotational offset Δφ between them can be determined by calculating the inverse Fourier transform of the normalized cross-power spectrum of the polar transformation between the reference image I Ref and each comparison image I Comp and. Further details of this approach are described in the paper by J. Luce et al., "Medical Image Registration Using the Fourier Transform", International Journal of Medical Physics, Clinical Engineering and Radiation Oncology, 2024.3, pp49-55 (February 2014), the content of which is hereby incorporated by reference in its entirety into this specification.

[0038] In the process S10 of FIG. 4, to determine the offset between the reference image I Ref and the comparison image I Comp other intensity-based algorithms may alternatively be used, including those based on similarity measures other than cross-correlation, such as mutual information, sum of squared intensity differences, or ratio image uniformity (among others).

[0039] Referring to FIG. 4 again, in process S20, the processor 320 compares each determined offset with an offset threshold value to determine whether the offset is smaller than the offset threshold value. As in this exemplary embodiment, the processor 320 may compare each translation offset t with a translation offset threshold value T to determine whether the translation offset t is smaller than the translation offset threshold value T. As in this exemplary embodiment, the processor 320 may compare each rotation offset Δφ with a rotation offset threshold value Θ to determine whether the rotation offset Δφ is smaller than the rotation offset threshold value Θ. In other exemplary embodiments, the reference image I Ref and each comparison image I Comp between each translation offset, or the reference image I Ref and each comparison image I Comp between each rotation offset, if any of the combinations of the reference image I Ref and each comparison image I Comp from the remaining images of the image sequence 20 is determined by the processor 320 in process S10 of FIG. 4, the processor 320 may optionally determine in process S20 of FIG. 4 whether either the translation offset t is smaller than the translation offset threshold value T or the rotation offset Δφ is smaller than the rotation offset threshold value Θ. The setting of the offset threshold values T and Θ will be described below.

[0040] In process S30 of FIG. 4, the processor 320 selects each comparison image I Ref in each combination in which it is determined that each offset between the reference image I Comp and each comparison image I Comp is smaller than the offset threshold value. For any combination in which it is determined that each offset between the reference image I Ref and each comparison image I Comp is not smaller than the offset threshold value, the processor 320 does not select the comparison image I Comp in that combination.

[0041] As in this exemplary embodiment, in process S20 of FIG. 4, in each combination where each translational offset t is determined to be less than the translational offset threshold T and each rotational offset Δφ is determined to be less than the rotational offset threshold Θ, each comparison image I Comp may be selected. In other exemplary embodiments, for each combination between the reference image I Ref and each comparison image I Comp , or for each combination between the reference image I Ref and each comparison image I Comp , if any of the translational offsets or rotational offsets between them is determined by the processor 320 (in process S10 of FIG. 4), then in process S20 of FIG. 4, the processor 320 may, in each combination where each translational offset t is determined to be less than the translational offset threshold T or each rotational offset Δφ is determined to be less than the rotational offset threshold Θ, select each comparison image I Ref from the reference image I Comp and the remaining images of the image sequence 20. Comp may be selected.

[0042] In process S40 of FIG. 4, the processor 320 generates an averaged image 70 of the region 30 of the retina 40 using the selected comparison image 25. As in this exemplary embodiment, the processor 320 uses the selected comparison image 25 (and optionally the reference image I Ref as well) to generate the selected image 25 (and optionally the reference image I RefBy aligning them with each other, an averaged image 70 of region 30 is generated. For example, for each pixel value of each pixel in the averaged image 70, it is the average (e.g., arithmetic mean) of the pixel values of the pixels at the corresponding positions in the aligned images corresponding to the respective common positions within region 30 of the retina 40. The aligned images are averaged so that the averaged image 70 of the first region 30 is generated. Alternatively, the processor 320 may average the aligned images such that, for example, each pixel value of each pixel in the averaged image 70 is the weighted average of the pixel values of the pixels at the corresponding positions in the aligned selected image 25 corresponding to the respective common positions within region 30 of the retina 40, to generate the averaged image 70 of the first region 30.

[0043] The offset thresholds T and Θ are such that when the image sequence 20 includes an image offset by a translational offset t greater than the translational offset threshold T and / or a rotational offset Δφ greater than the rotational offset threshold Θ from the reference image I Ref and an image with translational and rotational offsets from the reference image I Ref each less than T and Θ respectively, the averaged image 70 shows more texture indicating the structure of region 30 of the retina 40 than the reference averaged image 75 generated from the images within the image sequence 20.

[0044] The images compared in process S20 of FIG. 4 often differ due to non - affine transformations that tend to produce increasing deformations as the offset (regardless of translation or rotation) between the images increases. Thus, by selecting the comparison image I Ref offset by less than the offset threshold from the reference image I Comp a higher correlation between the contents of the images, and thus an improved retention of the texture of the averaged image 70, can be achieved. Any affine frame offset (typically translation and / or rotation) of the selected image 25 can be effectively corrected by the image alignment performed in process S40 of FIG. 4 so that the aligned images are highly correlated.

[0045] The amount of texture in the averaged image 70 can be quantified using an algorithm as described in the paper by R. Bergman et al., "Detection of Textured Areas in Images Using a Disorganization Indicator Based on Component Counts", J. Electronic Imaging, 17.043003, 2008. The entire content of this paper is hereby incorporated by reference in its entirety. The texture detector presented in this paper is based on the intuition that the texture of natural images is "disordered". The measure used to detect texture examines the structure of local regions of the image. This structural approach enables the detection of both structured and unstructured textures at multiple scales. Furthermore, it differentiates edges from textures and also textures from noise. The results of the automatic detection are shown in the above paper to be consistent with the human classification of the corresponding image regions. The amount of texture in the averaged image 70 and the reference averaged image 75 can be compared by comparing the regions of these images designated as "texture" by the algorithm.

[0046] The inventors have found that in the process S40 of FIG. 4, when interpolation is used to align the selected images 25 with respect to each other, at least a part of the texture present in the images acquired by the ophthalmic imaging device 10 tends to be removed, and without interpolation, that is, without interpolating between the pixel values of any of the selected images 25 to align the selected images 25, more of the original texture can be retained by the averaged image 70. Avoiding interpolation in the process S40 not only improves the retention of this clinical data in the averaged image 70, but also significantly improves the effectiveness of the machine learning (ML) noise removal algorithm in differentiating texture features from noise when the machine learning (ML) noise removal algorithm is trained using the averaged image 70 generated without interpolation.

[0047] Thus, as in this exemplary embodiment, the selected comparison image 25 can be used by the processor 320 to generate the averaged image 70 of the region 30 by first aligning the selected comparison images 25 with each other. In this process, one of the selected images 25 is selected as the base image, and each of the remaining selected images is sequentially aligned with the base image to generate a set of aligned images. When one of the remaining images is aligned with the base image, the pixel values of the first image of the pair are redistributed according to the geometric transformation between the image coordinate systems of the pair of images, i.e., the pixel values of the first image are maintained but reassigned to different pixels of the first image according to the geometric transformation required to align the images, and no interpolation is performed. Next, the processor 320 generates the averaged image 70 of the region 30, for example, by averaging the aligned images such that the respective pixel values of each pixel of the averaged image 70 are the average (e.g., arithmetic mean) of the pixel values of the corresponding pixels in the aligned selected images 25 corresponding to the respective common positions within the region 30 of the retina 40. Alternatively, the processor 320 may generate the averaged image 70 by averaging the aligned images such that the respective pixel values of each pixel in the averaged image 70 are the weighted average of the pixel values of the corresponding pixels in the aligned selected images 25 corresponding to the respective common positions within the region 30 of the retina 40.

[0048] When interpolation is avoided in the process S40 of FIG. 4, each geometric transformation between the image coordinate systems of each pair of images of the selected comparison image 25 consists of (i) a respective first translation by a respective first integer number of pixels along a first pixel arrangement direction in which the pixels of the selected comparison image 25 are arranged, and / or (ii) a respective second translation by a respective second integer number of pixels along a second pixel arrangement direction in which the pixels of the selected comparison image 25 are arranged, the second direction being orthogonal to the first direction. Therefore, in the pair of images to be aligned, neither image rotates. When this constraint is used in the alignment process, the rotation offset threshold Θ used in the process S20 of FIG. 4 requires that the image rotation required to compensate for the rotation offset Δφ between a pair of images is small enough to cause a displacement at the edge of the image frame that is smaller than the resolution of the ophthalmic imaging device 10. For example, in the case of a FAF imaging device, since the resolution is usually limited by the point spread function of the device, the rotation offset threshold Θ is set such that the rotation required for alignment that causes a displacement smaller than the width of the point spread function along the edge of the image frame is less than Θ (in this case, in S30 of FIG. 4, the comparison image of the pair of images will be selected), while the rotation required for alignment that causes a displacement larger than the width of the point spread function along the edge of the image frame is greater than Θ (in this case, in S30 of FIG. 4, the comparison image of the pair of images will not be selected). As an example, when the size of the images in the sequence is 512×512 pixels and the width of the point spread function of the region 30 is 1.5 pixels as shown in FIG. 5, setting Θ = 0.3 degrees results in a maximum displacement of tan(0.3)×256 = 1.3 pixels due to the rotation for alignment (at the edge of the image frame), which is smaller than the width of the point spread function.

[0049] FIG. 6B shows an averaged image obtained by averaging an exemplary FAF image shown in FIG. 6A and another FAF image rotated off by 4.5 degrees from the image of FIG. 6A. For comparison, FIG. 6C shows another averaged image obtained by averaging the FAF image shown in FIG. 6A and another FAF image having a negligible rotational offset from the image of FIG. 6A. As can be understood from the comparison between FIGS. 6B and 6C, the averaged image of FIG. 6C shows anatomical details of the emphasized region R that are not present in the averaged image of FIG. 6B.

[0050] In some cases, the images within the sequence 20 may have a high proportion of in-frame distortion and / or local micro-distortion regions of the type described above (regions with minute displacements between the imaged retinal features). In such cases, in the process S30 of FIG. 4, a relatively low proportion of the images within the image sequence 20 may be selected for averaging, and an averaged image 70 may be generated. As a result, noise suppression in the averaged image 70 may not be effective. In such cases, as schematically shown in FIG. 7, it may be beneficial to adapt the process of FIG. 4 by including a preprocessing (before S10) of generating each image of the image sequence 20 by segmenting each image of the second image sequence 15 of the second region of the retina 40. Thereby, the image of the region 30 becomes the segment 17 of each image of the second image sequence 15. Thus, the image sequence 20 may be a sequence of image segments of the second image sequence 15 acquired by the ophthalmic imaging device 10 that images a larger region of the retina 40 including the region 30. Each image segment 17 has a respective position within each image of the second image sequence 15, and that position is the same as the respective positions of all other image segments within each image of the second image sequence 15. Limiting the subsequent image processing operations in the process of FIG. 4 to the segments of the second image sequence 15 increases the proportion of image segments selected in process S30 for averaging to generate the averaged image 70. As a result, compared to the case where the images of the second image sequence 15 are processed as a whole according to the process described above with reference to FIG. 4, the noise in the averaged image 70 may be reduced. The averaged image 70 based on such selection of image segments has less noise, but since it only covers a part of the region of the retina 40 covered by the images of the second image sequence 15, a plurality of other image segments 17 at the corresponding positions of the second image sequence 15 can be processed in the same way to generate additional averaged images that collectively cover a wider range of the second region of the retina 40.

[0051] Each image of the second image sequence 15 may similarly be segmented into a one- or two-dimensional array of rectangular image tiles (e.g., a two-dimensional array of square image tiles as shown in FIG. 7), and each set of image tiles at corresponding positions from the segmented images may be processed according to the process described above with reference to FIG. 4 to generate corresponding averaged image tiles. The averaged image tiles may then be combined to generate an averaged image of the second region of the retina, which can exhibit various levels of noise suppression among the component averaged image tiles. The averaged image tiles are also useful for training a machine learning noise removal algorithm that more effectively distinguishes noise from retinal structure, and thus, as will be described later, a noise removal retinal image can be generated that retains more of the texture present in the originally acquired image. The two-dimensional orthogonal array of square image segments 17 in FIG. 7 is for illustration only, and it should be noted that the image segments 17 need not be of the same size or shape, nor need they be in a regular array. [Second Exemplary Embodiment]

[0052] FIG. 8 is a flowchart showing a method according to a second exemplary embodiment in which a data processing apparatus 60 can process an image sequence 20 to generate an averaged image 70 of a region 30.

[0053] The processes S10, S20, and S40 in FIG. 8 are the same as the processes in FIG. 4, and a repeated description here is omitted. The method of the second exemplary embodiment differs from the method of the first exemplary embodiment in that it includes additional processes S15 and S25, and a modified form of process S30 (labeled S30' in FIG. 8), and these processes will be described next.

[0054] In process S15 of FIG. 8, the processor 320 determines, for each combination of a reference image I Ref and a respective comparison image I Comp when they are aligned with respect to each other using respective offsets (plural possible), e.g., respective translational offsets t and respective rotational offsets Δφ of the first exemplary embodiment described above, the reference image IRef and each comparison image I Comp to determine the respective similarity therebetween. The reference image I in each combination Ref and each comparison image I Comp The respective translational offsets t between and are, as in the first exemplary embodiment, the reference image I Ref and each comparison image I Comp may be determined by calculating the cross-correlation between. In this case, when aligned with each other using the respective translational offsets t, the reference image I Ref and each comparison image I Comp The respective similarity between and may be determined by determining the maximum value of the calculated cross-correlation. Note that in some exemplary embodiments, the order of processes S10 and S15 may be reversed. Therefore, when the processor 320 calculates the cross-correlation between the reference image I Ref and each comparison image I Comp the processor 320 may determine the maximum value of the calculated cross-correlation before determining the translational (x-y) offset corresponding to that maximum value. When aligned with each other using the respective offset(s), the reference image I Ref and each comparison image I Comp The respective similarity between and need not be derived from the calculated cross-correlation between the images, and instead, for example, may be obtained by calculating the sum of squared differences or mutual information based on the images.

[0055] In process S25 of FIG. 7, the processor 320 compares each similarity determined in process S15 with a first similarity threshold value, and determines whether the determined similarity is greater than the first similarity threshold value. Note that in some exemplary embodiments, the order of processes S20 and S25 may be reversed.

[0056] In process S30' of FIG. 8, the processor 320 is the reference image I Ref and each comparison image I CompFor each combination determined to have, during that time, respective offsets smaller than the offset threshold and respective similarities greater than the first similarity threshold when aligned with each other using the respective offsets (i.e., each combination satisfying the condition of S30 in FIG. 4), each comparison image I Comp is selected. The latter additional selection criterion helps to avoid a comparison image I having non-affine distortion within the frame Comp and / or a comparison image I having minute distortion (i.e., the misalignment of retinal features is limited to one or more local regions) Comp from being selected for averaging to generate the averaged image 70. Such images tend to have a relatively low correlation with the reference image I Ref even when the offset(s) between them and the reference image I Ref is / are determined to be smaller than the offset threshold(s). Therefore, the additional processes S15 and S25 in FIG. 8 and the correction process S30' may help to avoid loss of valuable clinical data such as texture in the averaged image 70. The first similarity threshold can be set by trial and error while observing the effect of adjusting the first similarity threshold with respect to the amount of texture present in the averaged image 70

[0057] FIG. 9A shows an exemplary FAF image acquired by the ophthalmic imaging device 10. For comparison, FIG. 9B shows an example of the averaged image 70 obtained by averaging 10 FAF images including the FAF image of FIG. 9A, and the 10 FAF images are selected from the FAF image sequence according to the method described herein with reference to FIG. 8. FIG. 9B shows that the averaged image has less noise than the FAF image of FIG. 9A while retaining most of the texture of the FAF image of FIG. 9A (which is most prominent in the bright gray regions of the image).

[0058] Referring again to FIG. 8, in process S25, the processor 320 may also compare each determined similarity with a second similarity threshold to determine whether the determined similarity is less than the second similarity threshold, where the second similarity threshold is greater than the first similarity threshold. In this case, the processor 320 may use the reference image I Ref and each comparison image I Comp has a respective offset less than the offset threshold therebetween, and for each combination determined to have a similarity greater than the first similarity threshold and less than the second similarity threshold when the reference image I Ref and each comparison image I Comp are aligned with each other using their respective offsets, each comparison image I CompBy selecting, a modified version of process S30' may be executed. Thus, the comparison images may be selected for averaging to generate the averaged image 70 using the same criteria as in the first exemplary embodiment and an additional criterion that the similarity is within a predefined range of values, i.e., between a first similarity threshold and a second similarity threshold. (Determined by comparison with an appropriate value of the second similarity threshold set by trial and error) Comparison images that result in an unduly high determined similarity, as well as comparison images that result in a determined similarity that is not high enough (determined by comparison with an appropriate value of the first similarity threshold set by trial and error) are excluded from the selection, whereby the present inventors have found that the retention of texture in the averaged image 70 is improved. Without wishing to be bound by theory, the reasons for the above may be understood as follows. Each imaging region on the retina has a certain amount of information, and each single image captures only a part of that information space due to factors such as subtle changes in illumination, changes in the scanning system (subtle changes in the line start position and related optical paths, etc.), and the way light scatters from the layers of the retina and other complex effects. Statistically, images of regions with slightly different information content tend to generate an averaged image that overall reassembles a larger proportion of the total amount of available information compared to the individual images. Put another way, very highly correlated images with exactly the same information content do not generate an averaged image with more information than the individual images, but will still contain noise. [Third Exemplary Embodiment]

[0059] FIG. 10 is a flowchart showing a method according to a third exemplary embodiment in which a data processing apparatus 60 trains a machine learning (ML) algorithm to filter noise from a retinal image. The ML algorithm may be any type of supervised ML algorithm known to those skilled in the art that is suitable for removing one or more different types of noise from a FAF image or other type of retinal image when trained to perform this task using labeled training data. The ML algorithm may consist of a convolutional neural network (CNN), as in this exemplary embodiment. The other paper by R.S. Thakur et al., "Image De-Noising With Machine Learning: A Review", IEEE Access, Vol. 9, 2021, p93338-93363 (the content of which is incorporated herein by reference in its entirety) describes various state-of-the-art machine learning-based image denoisers such as dictionary learning models, convolutional neural networks, and adversarial generative networks for a wide range of noises such as Gaussian noise, impulse noise, Poisson noise, mixed noise, real-world noise, etc. The ML algorithm may be any of the various types described in this paper or other ones known to those skilled in the art.

[0060] In process S100 of FIG. 10, the processor 320 of the data processing apparatus 60 generates ground truth training target data by processing each sequence 20 of a plurality of sequences of retinal images as described in any of the foregoing exemplary embodiments or variations thereof to generate respective averaged retinal images.

[0061] FIG. 11 is a flowchart summarizing a method by which the processor 320 generates each of the averaged retinal images from each sequence of retinal images in the third exemplary embodiment.

[0062] In process S110 of FIG. 11, the processor 320 uses a reference image I Ref and each comparison image I that is an image from the remaining images in sequence 20Comp For each combination, the reference image I Ref and each comparison image I Comp to determine each offset therebetween.

[0063] In the process S120 of FIG. 11, the processor 320 compares each determined offset with an offset threshold value to determine whether the offset is smaller than the offset threshold value.

[0064] In the process S130 of FIG. 11, the processor 320 Ref and each comparison image I Comp In each combination where it is determined that the offset between and each comparison image I is smaller than the offset threshold value, each comparison image I Comp is selected.

[0065] In the process S140 of FIG. 11, the processor 320 generates an averaged image 70 for each region 30 using the selected comparison image 25.

[0066] As described above, the offset threshold value is such that when the sequence 20 includes at least one image offset by an amount greater than the threshold from the reference image I Ref and an image offset by an amount smaller than the threshold from the reference image I Ref the averaged image 70 is defined to show more texture indicating the structure of the first region 30 of the retina 40 than the reference averaged image generated from the images within the image sequence 20.

[0067] Referring again to FIG. 10, in the process S200, the processor 320 generates training input data by selecting each image from each of the image sequences.

[0068] Next, in the process S300 of FIG. 10, the processor 320 trains an ML algorithm to filter noise from the retinal image using well-known techniques to those skilled in the art using the ground truth training target data and the training input data.

[0069] Some of the exemplary embodiments described above are summarized in clauses E1 to E14 numbered below.

[0070] E1 A data processing device 60 arranged to process a first image sequence 20 of a region 30 of the retina 40 of an eye 50 to generate an averaged image 70 of the region 30, the data processing device 60 including at least one processor 320 and at least one memory 340 storing computer-readable instructions, the computer-readable instructions, when executed by the at least one processor 320, causing the at least one processor 320 to a reference image I selected from the first image sequence 20 Ref and each respective comparison image I which is an image from the remaining images of the first image sequence 20 Comp for each combination of the reference image I Ref and each respective comparison image I Comp to determine each respective offset between the reference image I Comp and each respective comparison image I; to compare each determined offset with an offset threshold to determine whether the offset is less than the offset threshold; in each combination where it is determined that the respective offset between the reference image I Ref and each respective comparison image I Comp is less than the offset threshold, to select each respective comparison image I Comp ; to use the selected comparison images 25 to generate an averaged image 70 of the region 30, wherein the offset threshold is defined such that when the first image sequence 20 includes at least one image offset by more than the threshold from the reference image I Ref and images offset by less than the threshold from the reference image I Ref the averaged image 70 shows more texture indicative of the structure of the first region 30 of the retina 40 than a reference averaged image generated from images within the first image sequence 20. The data processing device 60.

[0071] E2 Reference image I Ref and each comparison image I Comp Each offset determined for each combination includes a translational offset, and when the computer-readable instructions are executed by at least one processor 320, cause the at least one processor 320 to compare each translational offset with a translational offset threshold to determine whether the translational offset is less than the translational offset threshold, and thereby compare each determined offset with the offset threshold to determine whether the offset is less than the offset threshold; In each combination in which each translational offset is determined to be less than the translational offset threshold, each comparison image I Comp is selected, so that in each combination in which each offset between the reference image I Ref and each comparison image I Comp is determined to be less than the offset threshold, each comparison image I Comp is selected, The data processing device 60 of E1.

[0072] E3 When the computer-readable instructions are executed by at least one processor 320, cause the at least one processor 320 to, for each combination, calculate the mutual correlation between the reference image I Ref and each comparison image I Comp using; and Reference image I Ref and each comparison image I Comp calculate the inverse Fourier transform of the normalized cross-power spectrum calculated using; and Reference image I Ref and each comparison image I Comp Thereby determining by any one of, the data processing device 60 of E2.

[0073] E4 Reference image I Refand each comparison image I Comp Each offset determined for each combination includes a rotation offset, When the computer-readable instructions are executed by at least one processor 320, cause the at least one processor 320 to, compare each determined offset with an offset threshold by comparing each rotation offset with a rotation offset threshold to determine whether the rotation offset is less than the rotation offset threshold, thereby determining whether the offset is less than the offset threshold; In each combination where each rotation offset is determined to be less than the rotation offset threshold, each comparison image I Comp is selected, thereby selecting each comparison image I Ref in each combination where each offset between the reference image I Comp and each comparison image I Comp is determined to be less than the offset threshold, any one of the data processing devices 60 from E1 to E3.

[0074] E5 When the computer-readable instructions are executed by at least one processor 320, cause the at least one processor 320 to, for the reference image I in each combination Ref and each comparison image I Comp calculate the cross-correlation using the rotated version of each comparison image I ; and Comp calculate the inverse Fourier transform of the normalized cross-power spectrum calculated using the polar transformation of the reference image I and each comparison image, Ref and cause the determination to be made by any one of the data processing device 60 of E4.

[0075] E6 When the computer-readable instructions are executed by at least one processor 320, cause the at least one processor 320 to, Aligning the selected comparison images 25 with respect to each other, wherein aligning each pair of the selected comparison images 25 includes redistributing the pixel values of one of the pair of images according to each geometric transformation between the image coordinate systems of the pair of images; Generating an averaged image 70 of the first region by averaging the aligned images; A data processing device 60 of any one of E1 to E5 that generates an averaged image of the first region using the selected comparison images 25.

[0076] E7 Each geometric transformation between the image coordinate systems of the images of each pair of the selected comparison images 25 is Each first translation by a respective first integer number of pixels along the first pixel array direction in which the pixels of the selected comparison image 25 are arranged; and Each second translation by a respective second integer number of pixels along the second pixel array direction in which the pixels of the selected comparison image 25 are arranged, A data processing device 60 of E6 comprising at least one of the above.

[0077] E8 When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is further caused to For each combination of the reference image I Ref and each respective comparison image I Comp determine the respective similarity between the reference image I and each respective comparison image I when they are aligned with respect to each other using each respective offset; Ref and each respective comparison image I Comp ; Compare each determined similarity with a first similarity threshold to determine whether the determined similarity is greater than the first similarity threshold, When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is caused to, for the reference image I Ref and each respective comparison image I Compin each combination determined to have each offset smaller than the offset threshold during that time and each similarity larger than the first similarity threshold when aligned with each other using each offset, each comparison image I Comp including causing selection of any one of data processing apparatuses 60 from E1 to E7.

[0078] E9 When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is further caused to compare each determined similarity with a second similarity threshold to determine whether the determined similarity is smaller than the second similarity threshold, the second similarity threshold being larger than the first similarity threshold, When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is caused to, with respect to the reference image I Ref and each comparison image I Comp in each combination determined to have each offset smaller than the offset threshold during that time and each similarity larger than the first similarity threshold and smaller than the second similarity threshold when aligned with each other using each offset, each comparison image I Comp including causing selection of the data processing apparatus 60 of E8.

[0079] E10 When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is caused to, for each combination, the reference image I Ref and each comparison image I Comp to determine each translational offset between them by calculating a cross-correlation using the reference image I Ref and each comparison image I Comp and, when aligned with each other using each translational offset, the reference image I Ref and each comparison image IComp Determine the similarity between each of them by determining the maximum value of the calculated cross-correlation. A data processing device 60 of E8 when subordinate to E2, and E9 when subordinate to E2.

[0080] E11 When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is further caused to generate each image of the first image sequence 20 by segmenting each image of the second image sequence in the second region of the retina 40 such that the image of the first region 30 is a segment of each image of the second image sequence. A data processing device 60 of any one of E1 to E10.

[0081] E12 A data processing device 60 of any one of E1 to E11, wherein the first image sequence 20 includes an autofluorescence image sequence of the first region 30 of the retina 40 of the eye 50.

[0082] E13 When the computer-readable instructions are executed by at least one processor 320, the at least one processor 320 is further caused to train a machine learning algorithm for filtering noise from a retinal image, and further By processing each sequence 20 of a plurality of sequences of retinal images to generate respective averaged retinal images, ground truth training target data is generated, and each of the averaged retinal images is Reference image I Ref And each comparison image I which is an image from the remaining images in the sequence 20 Comp For each combination of Ref Reference image I Comp And each respective comparison image I Determine each determined offset compared to an offset threshold to determine if the offset is less than the offset threshold; Reference image I Ref And each respective comparison image I CompIn each combination where each offset between them is determined to be less than the offset threshold, select each comparison image I Comp ; Use the selected comparison image 25 to generate each averaged image 70 of region 30, generated by, The offset threshold is such that when sequence 20 includes at least one image offset by more than the threshold from reference image I Ref and each image offset by less than the threshold from reference image I Ref the averaged image 70 is defined to show more texture indicating the structure of the first region 30 of retina 40 than the reference averaged image generated from the images within image sequence 20; Generate training input data by selecting each image from each of the image sequences; Use the ground truth training target data and the training input data to train a machine learning algorithm for filtering noise from the retinal image, Any of the data processing devices 60 from E1 to E12.

[0083] E14 An ophthalmic imaging device 10 arranged to acquire an image sequence 20 of a region 30 of a retina 40 of an eye 50; Any of the data processing devices 60 from E1 to E13 arranged to process the image sequence 20 acquired by the ophthalmic imaging device 10 to generate an averaged image 70 of the region 30 of the retina 40, An ophthalmic imaging system 100 including.

[0084] In the foregoing description, exemplary aspects have been described with reference to several exemplary embodiments. Accordingly, this specification is not limiting and should be regarded as exemplary. Similarly, the figures in the drawings highlighting the features and advantages of the exemplary embodiments are presented for illustrative purposes only. The architecture of the exemplary embodiments is flexible and configurable in various ways so that it can be utilized in ways other than those shown in the accompanying drawings.

[0085] Some aspects of the examples presented in this specification may be provided as one or more programs, such as a computer program, or software, containing instructions or instruction sequences that are contained in or stored on a manufactured article such as a machine-accessible medium or machine-readable medium, instruction store, or computer-readable storage device, each of which may be non-transitory. Programs or instructions on a non-transitory machine-accessible medium, machine-readable medium, instruction store, or computer-readable storage device can be used to program a computer system or other electronic device. Machine or computer-readable media, instruction stores, and storage devices may include, but are not limited to, floppy (registered trademark) disks, optical disks, and magneto-optical disks, or other types of media / machine-readable media / instruction stores / storage devices suitable for storing or transmitting electronic instructions. The techniques described herein are not limited to any particular software configuration. They are applicable to any computer environment or processing environment. As used herein, the terms "computer-readable," "machine-accessible medium," "machine-readable medium," "instruction store," and "computer-readable storage device" shall include any medium that can store, encode, or transmit instructions or instruction sequences for execution by a machine, computer, or computer processor, and cause the machine / computer / computer processor to perform any one of the methods described herein. Further, in the art, it is common to consider software in some form or another (e.g., program, procedure, process, application, module, unit, logic, etc.) as that which causes an action to occur or a result to be caused. Such expressions are merely shorthand expressions indicating that the execution of software by a processing system causes a processor to perform an action and result in a consequence.

[0086] Also, some or all of the functions of the data processing device 60 may be implemented by preparing an application-specific integrated circuit, a field-programmable gate array, or by interconnecting a suitable network of conventional component circuits.

[0087] A computer program product can be provided in the form of one or more storage media, instruction stores (s), or storage devices (s), which store or hold instructions that can be used for controlling a computer or a computer processor or for causing a computer or a computer processor to execute any of the procedures of the exemplary embodiments described herein. The storage media / instruction store / storage device may include, for example, but are not limited to, optical disks, ROM, RAM, EPROM, EEPROM, DRAM, VRAM, flash memory, flash cards, magnetic cards, optical cards, nanosystems, molecular memory integrated circuits, RAID, remote data storage / archive / warehousing, and / or other types of devices suitable for storing instructions and / or data.

[0088] Some implementations are software stored in any one or more computer-readable media, instruction stores (s), or storage devices (s), the software controlling the system's hardware and enabling the system or microprocessor to interact with a human user or other entity using the results of the exemplary embodiments described herein. Such software may include, but is not limited to, device drivers, operating systems, user applications. Ultimately, these computer-readable media or storage devices (s) further include software for performing the exemplary aspects of the present invention as described above.

[0089] The programming and / or software of the system includes software modules for performing the procedures described herein. In some exemplary embodiments herein, a module includes software, and in other exemplary embodiments herein, a module includes hardware, or a combination of hardware and software.

[0090] The various exemplary embodiments of the present invention have been described above. It should be understood that these are merely examples and do not limit the present invention in any way. It will be apparent to those skilled in the relevant technical field(s) that various changes can be made to the shape and details. Therefore, the present invention should not be limited to any of the above-described exemplary embodiments, but should be defined only by the following claims and their equivalents.

[0091] Furthermore, the purpose of the abstract is to enable the Patent Office, the general public, especially scientists, engineers, and those skilled in the art who are not familiar with patent terms and legal terms, to quickly judge the nature and essence of the technical disclosure of this application at a glance. The abstract is not intended to limit the scope of the exemplary embodiments presented in this specification in any way. Also, it should be understood that the procedures described in the claims do not have to be executed in the presented order.

[0092] This specification includes details of many specific embodiments, but these should not be construed as limiting any invention or the scope of the claims. Instead, they should be construed as descriptions of features specific to the particular embodiments described in this specification. The specific features described in separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Furthermore, a feature may be described above as operating in a particular combination and may even be initially claimed as such. However, one or more features from the claimed combination may, in some cases, be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0093] Depending on the situation, multitasking or parallel processing may be advantageous. Furthermore, the separation of the various components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it will be understood that the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.

[0094] Although several exemplary embodiments and aspects have been described, the above are exemplary and it is clear that they are presented for purposes of example. In particular, many of the examples presented herein include specific combinations of apparatus or software elements, but those elements can be combined in other ways to achieve the same purpose. The acts, elements, and features discussed only in connection with one embodiment are not intended to be excluded from a similar role in other embodiments or other combinations of embodiments.

Claims

1. A computer-implemented method for processing a first image sequence (20) of a first region (30) of a retina (40) of an eye (50) to generate an averaged image (70) of the first region (30), comprising: A reference image (I selected from the first image sequence (20) Ref ), and for each combination with each comparison image (I which is an image from the remaining images of the first image sequence (20) Comp ), determining each offset (t; Δφ) between the reference image (I Ref ) and each of the comparison images (I Comp ); comparing each determined offset (t; Δφ) with an offset threshold (T; Θ) to determine whether the offset (t; Δφ) is less than the offset threshold (T; Θ); the reference image (I Ref ), and in each combination in which each of the offsets (t) between the reference image (I Comp ) and each of the comparison images (I Comp ) is determined to be smaller than the offset threshold (T; Θ), selecting each of the comparison images (I ) using the selected comparison image (25) to generate the averaged image (70) of the first region (30). The offset threshold (T; Θ) is such that when the first image sequence (20) includes at least one image offset by an offset (t) greater than the threshold (T) from the reference image (I Ref ) and an image offset by each offset (t) smaller than the threshold (T) from the reference image (I Ref ), the averaged image (70) is defined to show more texture indicating the structure of the first region (30) of the retina (40) than a reference averaged image (75) generated from the images within the first image sequence (20). A computer-implemented method.

2. The reference image (I Ref ) and each of the respective comparison images (I Comp ) for each combination of the respective offsets (t; Δφ) determined for each combination include a translational offset (t), The comparing includes comparing each translational offset (t) with a translational offset threshold (T) to determine whether the translational offset (t) is less than the translational offset threshold (T). Said selecting includes, in each combination determined such that each respective translational offset (t) is smaller than the translational offset threshold value (T), selecting each respective comparison image (I Comp ). The computer-implemented method according to Claim 1.

3. The respective translational offsets (t) between the reference image (I Ref ) and the respective comparison images (I Comp ) in each combination are Calculating the cross-correlation using the reference image (I Ref ), and the respective comparison images (I Comp ), and The inverse Fourier transform of the normalized cross-power spectrum calculated using the reference image (I Ref ) and each of the comparison images (I Comp ) is calculated, The computer-implemented method according to Claim 2, as determined by any one of the above.

4. The reference image (I Ref ) and each of the respective comparison images (I Comp ), each of the respective offsets (t; Δφ) determined for each combination, includes a rotational offset (Δφ), The comparing includes comparing each rotational offset (Δφ) with a rotational offset threshold (Θ) to determine whether the rotational offset (Δφ) is less than the rotational offset threshold (Θ). Said selecting includes, in each combination determined such that each of said respective rotational offsets (Δφ) is smaller than said rotational offset threshold value (Θ), selecting said respective comparison images (I Comp ). The computer-implemented method according to any one of Claims 1 to 3.

5. The respective rotational offsets (Δφ) between the reference image (I Ref ), and the respective comparison images (I Comp ) in each combination are Calculating the cross-correlation using the rotated versions of the respective comparison images (I Comp ), and The inverse Fourier transform of the normalized cross-power spectrum calculated using the polar transformation of the reference image (I Ref ) and each of the comparison images (I Comp ) is calculated, The computer-implemented method according to Claim 4, as determined by any one of the above.

6. The selected comparison image (25) is aligning the selected comparison images (25) with respect to each other, wherein aligning each pair of the selected comparison images (25) includes redistributing pixel values of one of the images of the pair according to a respective geometric transformation between the image coordinate systems of the images of the pair; generating the averaged image (70) of the first region (30) by averaging the aligned images; The computer-implemented method according to any one of the preceding claims, used to generate the averaged image (70) of the first region (30).

7. The respective geometric transformation between the image coordinate systems of the images of each pair of the selected comparison images (25) is each first translation by a respective first integer number of pixels along a first pixel array direction in which pixels of the selected comparison image (25) are arranged, and Each second translation, for each second integer number of pixels, along a second pixel array direction in which the pixels of the selected comparison image (25) are arranged, The computer-implemented method according to claim 6, comprising at least one of: **Claim 8** The reference image (I Ref ), for each combination of the reference image (I Comp ) and each of the comparison images (I Ref ), determining the similarity of the reference image (I Comp ) and each of the comparison images (I Comp ) when they are aligned with each other using the respective offsets, and Comparing each determined similarity with a first similarity threshold to determine whether the determined similarity is greater than the first similarity threshold; further comprising Said selecting includes, in each combination determined to have, between the reference image (I Ref ), and each of the comparison images (I Comp ), respective offsets (t; Δφ) smaller than the offset threshold (T; Θ) therebetween, and respective similarities greater than the first similarity threshold when the respective comparison images are aligned with each other using the respective offsets, selecting each of the comparison images (I Comp ). The computer-implemented method according to any one of the preceding claims. **Claim 9** To determine whether each of the determined similarities is less than a second similarity threshold, further comprising comparing each of the determined similarities with the second similarity threshold, wherein the second similarity threshold is greater than the first similarity threshold, and the selecting comprises, for each combination determined to have, between the reference image (I Ref ), and each of the comparison images (I Comp ), a respective offset (t; Δφ) less than the offset threshold (T; Θ), and a respective similarity greater than the first similarity threshold and less than the second similarity threshold when the respective offsets are used to align the comparison images with respect to each other, selecting each of the comparison images (I Comp ). The computer-implemented method according to claim 8. **Claim 10** The respective translational offsets (t) between the reference image (I Ref ) and the respective comparison images (I Comp ) in each combination are determined by calculating the cross-correlation using the reference image (I Ref ) and the respective comparison images (I Comp ). When aligned with respect to each other using the respective translational offsets (t), the similarity between the reference image (I Ref ) and each of the comparison images (I Comp ) is determined by determining the maximum value of the calculated cross-correlation, The computer-implemented method according to claim 8 when dependent on claim 2, and according to claim 9 when dependent on claim 2. **Claim 11** Generating each image of the first image sequence (20) by segmenting each image of a second image sequence of a second region of the retina (40) such that the image of the first region (30) is a segment of each image of the second image sequence, the computer-implemented method according to any one of the preceding claims. **Claim 12** The computer-implemented method according to any one of the preceding claims, wherein the first image sequence (20) comprises an autofluorescence image sequence of the first region (30) of the retina (40) of the eye (50). **Claim 13** A computer-implemented method for training a machine learning algorithm for filtering noise from retinal images, comprising: Generating ground truth training target data by processing each sequence of a plurality of sequences of retinal images to generate a respective averaged retinal image, each averaged retinal image being generated according to the computer-implemented method according to any one of the preceding claims (S100); Generating training input data by selecting a respective image from each of the image sequences (S200); Using the ground truth training target data and the training input data to train the machine learning algorithm for filtering noise from retinal images (S300). A computer-implemented method comprising. **Claim 14** A computer program (345) comprising computer-readable instructions which, when executed by at least one processor (320), cause the at least one processor (320) to execute the method according to any one of the preceding claims.

15. A data processing apparatus (60) arranged to process an image sequence (20) of a region (30) of a retina (40) of an eye (50) to generate an averaged image (70) of the region (30), the data processing apparatus (60) comprising at least one processor (320) and at least one memory (340) storing computer-readable instructions, the computer-readable instructions, when executed by the at least one processor (320), causing the at least one processor (320) to The reference image (I) selected from the image sequence Ref ), and for each combination with each comparison image (I Comp ) that is an image from the remaining images of the image sequence (20), determine each offset (t; Δφ) between the reference image (I Ref ) and each of the comparison images (I Comp ), compare each determined offset (t; Δφ) with an offset threshold (T; Θ) to determine whether the offset (t; Δφ) is less than the offset threshold (T; Θ), The reference image (I Ref ), and in each combination where it is determined that each offset (T; Θ) between the reference image and each of the comparison images (I Comp ) is smaller than the offset threshold (T; Θ), select each of the comparison images (I Comp ). use the selected comparison image (25) to generate the averaged image (70) of the region (30), The offset threshold (T; Θ) is such that when the image sequence (20) includes at least one image offset by an offset (t) greater than the threshold (T) from the reference image (I Ref ) and an image offset by each offset (t) less than the threshold (T) from the reference image (I Ref ), the averaged image (70) is defined to show more texture indicating the structure of the region (30) of the retina (40) than a reference averaged image (75) generated from the images within the image sequence (20). a data processing apparatus (60).

Citation Information

Patent Citations

  • A speckle denoising method for OCT imaging based on conditional generation antagonistic network is proposed

    CN109345469A

  • Image processing apparatus, image processing method, and program

    JP2019080724A

  • Image fusion method and device, and terminal device

    JP2019503614A

  • OCT image transformation, system for denoising ophthalmic images, and neural network therefor

    JP2022520415A

  • Image processing device, image processing method, image processing program, and recording medium storing said program

    WO2014103501A1