Retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling

By constructing a dynamic eyeball model and a camera array model, and combining the SuperPoint and SuperGlue algorithms, the problems of applicability and inter-frame continuity in retinal image registration are solved, achieving high-precision retinal image sequence registration and synthetic image generation, which is suitable for method evaluation and deep learning.

CN122115518APending Publication Date: 2026-05-29NANKAI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANKAI UNIV
Filing Date
2026-02-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies have limitations in retinal image registration, being applicable only to the registration of a pair of images and exhibiting poor inter-frame continuity.

Method used

A retinal image sequence registration method based on a longitudinal three-dimensional fundus examination scene is constructed. A dynamic eyeball model and a camera array model are used, and the SuperPoint and SuperGlue algorithms are combined for key point extraction and matching. The particle swarm optimization algorithm is used for optimization to achieve joint registration of frame to reference and frame to frame.

Benefits of technology

It achieves more comprehensive retinal image sequence registration, improves registration accuracy and inter-frame continuity, and generates realistic synthetic retinal image sequences, which are suitable for method evaluation and deep learning model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115518A_ABST
    Figure CN122115518A_ABST
Patent Text Reader

Abstract

The application provides a retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling, and belongs to the field of retinal image registration. The method develops a more comprehensive and more universal transformation model, which is used for global registration of retinal image sequences in longitudinal research. Specifically, according to the process of imaging the retina by a fundus camera and the change of the shape of the eyeball in the natural growth or disease progression process, a 3D space-time model simulating the longitudinal fundus examination scene is constructed. The patent uses a joint registration strategy of frame-to-reference and frame-to-frame, realizes a retinal image sequence registration framework (RISeR) based on the 3D space-time model, and aligns all images in the sequence to the same longitudinal 3D fundus examination scene by using a matching key point between the retinal images as a medium and adopting a pose estimation and optimization algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of retinal image registration, and more specifically to a method for retinal image sequence registration based on longitudinal three-dimensional fundus examination scene modeling. Background Technology

[0002] Retinal image registration works on a pair of retinal images, aiming to align the two images in space. Patent CN 117314982 A provides a retinal image registration method for myopia development. The proposed method, RIRMD, constructs a three-dimensional spatial model for a pair of retinal images showing myopia progression based on the process of capturing retinal images with a fundus camera and the changes in eyeball shape during myopia development. Figure 1 As shown. In this model, the world coordinate system ( x w - y w - z w The origin of ) is located at the center of the eyeball (c w The reference camera (c0) is used to acquire reference images ( F The test camera (c1) is used to acquire test images. F 1), where the pose of the reference camera is fixed, and the relative pose between the test camera and the reference camera represents the change in eye pose when acquiring the reference image and the test image. Furthermore, the reference eye ( E 0) and test eyeballs ( E 1) Corresponding to the reference image and the test image respectively, they approximately simulate the changes in eye shape caused by the development of myopia. RIRMD achieves registration on the constructed 3D spatial model using corresponding points between retinal image pairs as the medium. The process is as follows: First, SuperPoint and SuperGlue are used to detect, describe and match key points in the retinal image pairs; then, the initial pose of the test camera is estimated using the RANSAC-based PnP algorithm; finally, the 3D spatial model is optimized using the Particle Swarm Optimization (PSO) algorithm, and the test image is registered accordingly.

[0003] However, RIRMD has the following drawbacks and shortcomings: 1) The three-dimensional spatial model constructed in RIRMD only considers the changes in eyeball shape caused by myopia development, and its applicability is limited; 2) The three-dimensional spatial model constructed in RIRMD can only be used for the registration of a pair of retinal images; 3) RIRMD is designed for the application requirement of registering a pair of retinal images; 4) When using RIRMD to register a retinal image sequence in a pairwise registration manner, the inter-frame continuity of the final sequence registration result is poor, that is, the registration error between adjacent frames is large. Summary of the Invention

[0004] The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling provided by this invention is an improvement and extension of the background technique method RIRMD. The goal is to further develop a more comprehensive and universal transformation model for global registration of retinal image sequences in longitudinal studies.

[0005] To solve the above problems, the present invention adopts the following technical solution: A retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling. This method, based on the RISeR method, constructs a three-dimensional spatiotemporal model for the longitudinal three-dimensional fundus examination scene. The three-dimensional spatiotemporal model includes a dynamic eyeball model and a camera array model. The RISeR method includes the following steps: S1: Input retinal image sequence { F 0, F 1, ..., F n}, F 0 is the selected reference image in the sequence; S2: Calculate the reference image F The intrinsic parameter matrix K0 and the extrinsic parameter matrix {R0, t0} of 0; S3: Reference image F 0 and remaining images F 1. Perform key point extraction and matching, and perform pairwise registration, using frame-to-reference and frame-to-frame strategies to register the remaining images in the sequence; S4: After registration, output the corresponding model parameters in the longitudinal three-dimensional fundus examination scene; S5: Register and transform all remaining images to the viewpoint of the reference image; S6: Output the registered retinal image sequence { F 0, F 1→0 , ..., F n→0}

[0006] Furthermore, the dynamic eyeball model consists of eyeballs of different shapes but with the same pose, using a world coordinate system. origin Multiple ellipsoids centered on The system is designed to simulate the shape changes that may occur in the eye between each examination; the camera array model consists of multiple cameras in different poses. Composition, where corresponding to the reference image F 0 reference camera C Pose 0 It is fixed and known that the camera pose change is used to simulate the inevitable pose change of the eye between each examination.

[0007] Furthermore, in the dynamic eyeball model, all eyeballs are approximated by ellipsoids, and the three orthogonal semi-axis between each ellipsoid... The lengths of the ellipsoids differ, and optimizations are performed separately for each. Regarding pose, it is assumed that the pose parameters of all the ellipsoids are completely identical. It is also assumed that the size of the ellipsoid decreases continuously as the three orthogonal semi-axis shortens, eventually becoming an infinitesimal ellipsoid located at the center of the original ellipsoid. The line connecting corresponding points between two ellipsoids is approximately the line connecting a point on the surface of the original ellipsoid to its center point, with the center point of the dynamic eyeball model as the reference point. Starting from the eyeball Points on the curved surface of the fundus The ray forms and interacts with the eyeball. The intersection of the fundus surfaces That is, the corresponding point after mapping: ; in, and It is a proportionality coefficient, and .

[0008] Furthermore, the camera array model includes n+1 cameras. Corresponding to the reference image and the remaining image Given a sequence of retinal images, the parameters of the camera array model consist of the following two parts: 1) the poses of all remaining cameras except the reference camera; and ;2) Fourth-order radial distortion coefficients for all cameras, .

[0009] Furthermore, the dynamic eyeball model and the camera array model are connected by a mapping between three-dimensional retinal points and two-dimensional image points, which includes two opposite processes: 1) mapping from three-dimensional retinal points to two-dimensional image points, which is essentially the process of imaging the fundus retina with a camera; 2) mapping from two-dimensional image points to three-dimensional retinal points, which is equivalent to the process of light rays emitted from the camera intersecting with the fundus.

[0010] Furthermore, in S3, the reference image... F 0 and remaining images F 1. Perform pairwise registration: 1) Keypoint extraction and matching are performed using the SuperPoint and SuperGlue algorithms; 2) Calculate the remaining image F The intrinsic parameter matrix K1 of 1; 3) Initialize the reference image F 0 and remaining images F 1 corresponds to the unknown model parameters; 4) Optimize the reference image F 0 and remaining images F The model parameters corresponding to 1.

[0011] Furthermore, the RISeR framework involves two different types of optimization: 1) optimization under a single constraint from frame to reference, used for registering sequences other than the reference image. F The first image outside of 0 F 1; 2) Optimization under joint frame-to-reference and frame-to-frame constraints, used for further registration of subsequent images in the sequence. Both optimization methods use the particle swarm optimization algorithm. p Particle iteration g The optimal candidate solution is found in the given search space.

[0012] Furthermore, in S6, a synthetic retinal image sequence The generation mainly includes the following steps: 1) Use the given retinal image as the reference image. F 0, based on which the relevant parameters of the reference camera and reference eyeball in the 3D spatiotemporal model are set. ; 2) Use appropriate strategies to set the parameters of the remaining cameras in the camera array model and the remaining eyes in the dynamic eye model. Ultimately, a complete three-dimensional spatiotemporal model is constructed to simulate the longitudinal fundus examination scenario; 3) In this scenario, synthesize a sequence of retinal images. It is generated by mapping image points between a reference camera and a residual camera, bilinear interpolation, and optional color space transformation, which simulates changes in image color, brightness, and contrast caused by variations in camera, camera settings, or environment.

[0013] Compared with the prior art, the present invention has the following advantages and effects: Thanks to the three-dimensional spatiotemporal simulation of longitudinal fundus examination scenarios in this patent, a large number of realistic synthetic retinal image sequences can be generated given only a single retinal image. This is potentially applicable to method evaluation, deep learning-based model training, and data augmentation. To demonstrate the effectiveness of the proposed method, a comprehensive experimental evaluation was conducted on two publicly available retinal image pair registration datasets, two newly constructed retinal image sequence registration datasets (one real and one synthetic), and one publicly available retinal image sequence registration dataset. The results show that this patent achieves state-of-the-art registration accuracy, reliability, and applicability.

[0014] (1) Based on the process of retinal imaging by fundus camera and the changes in eyeball shape during natural growth or disease progression, a three-dimensional spatiotemporal model simulating longitudinal fundus examination is constructed: ① The three-dimensional spatiotemporal model in this patent can be used to register a retinal image sequence, while the three-dimensional spatial model in RIRMD can only be used to register a pair of retinal images; ② The eyeball shape change model in this patent is more generalized, simulating the overall eyeball shape change caused by natural growth or disease progression, while the myopia development model in RIRMD only simulates the shape change of the posterior hemisphere of the eyeball caused by myopia development. (2) This patent also designs a joint registration strategy of frame-to-reference and frame-to-frame, realizing a global registration framework for retinal image sequences based on a three-dimensional spatiotemporal model. This is not found in the background technology RIRMD, and it is also one of the key technologies for this patent method to achieve high accuracy and good inter-frame continuity of retinal image sequence global registration results. The framework uses the matching key points between retinal images as a medium and uses pose estimation and optimization algorithms to align all images in a sequence to the same vertical 3D fundus examination scene. Attached Figure Description

[0015] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the description of the specific embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the three-dimensional spatial model of RIRMD in the background technology; Figure 2 This is a schematic diagram of the three-dimensional spatiotemporal model constructed in the RISeR method of this invention; Figure 3 This is a schematic diagram illustrating how the present invention solves for the parameters used to calculate the corresponding camera intrinsic parameter matrix based on a given retinal image; Figure 4 This is a schematic diagram of the mapping process between three-dimensional retinal points and two-dimensional image points in this invention; Figure 5 This is a flowchart of the RISeR framework used in this invention. Figure 6 This is a schematic diagram illustrating the generation of five different types of synthetic retinal image sequences according to the present invention; Figure 7 This is a schematic diagram of a retinal image sequence with labeled control points in RISeR-real; Figure 8 This is a schematic diagram of a retinal image sequence with labeled control points in RISeR-syn; Figure 9 This is a schematic diagram of labeled control points that exist only in the common retinal region between two adjacent frames and not in the reference frame; Figure 10 This is a schematic diagram of the registration results of RISeR and REMPE on the RISeR-real dataset. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] This patent is an improvement and extension of the RIRMD mentioned in the background art, with the goal of further developing a more comprehensive and general transformation model for global registration of retinal image sequences in longitudinal studies. The objectives of this patent include two aspects: (1) constructing a more comprehensive and general retinal image transformation model; and (2) achieving accurate registration of retinal image sequences with state-of-the-art registration accuracy and inter-frame continuity. Thanks to the three-dimensional spatiotemporal simulation of longitudinal fundus examination scenarios in this patent, the proposed method can generate a large number of realistic synthetic retinal image sequences given only a single retinal image. This may be applicable to application areas such as method evaluation, deep learning-based model training, and data augmentation.

[0019] Specifically, this invention discloses a retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling. The method RISeR constructs a three-dimensional spatiotemporal model for the longitudinal three-dimensional fundus examination scene. The three-dimensional spatiotemporal model includes a dynamic eyeball model and a camera array model. The RISeR method includes the following steps: S1: Input retinal image sequence {F0, F1, ..., F...} n}, F0 is the selected reference image in the sequence; S2: Calculate the reference image F The intrinsic parameter matrix K0 and the extrinsic parameter matrix {R0, t0} of 0; S3: Reference image F 0 and remaining images F 1. Perform key point extraction and matching, and perform pairwise registration, using frame-to-reference and frame-to-frame strategies to register the remaining images in the sequence; S4: After registration, output the corresponding model parameters in the longitudinal three-dimensional fundus examination scene; S5: Register and transform all remaining images to the viewpoint of the reference image; S6: Output the registered retinal image sequence { F 0, F 1→0 , ..., F n→0}

[0020] The RISeR method will be described in detail below.

[0021] 1. Longitudinal 3D Funduscopy Scene Modeling in RISeR Based on the process of retinal imaging by a fundus camera and the changes in eyeball shape during natural growth or disease progression, the method RISeR proposed in this patent is a longitudinal retinal image sequence. A three-dimensional spatiotemporal model was constructed, specifically as follows: Figure 2 As shown. This three-dimensional spatiotemporal model simulates a longitudinal fundus examination scenario, mainly including the following two important components: (1) a dynamic eyeball model, consisting of different shapes but the same pose, in a world coordinate system. origin Multiple ellipsoids centered on (2) A camera array model, consisting of multiple cameras with different poses, used to simulate the shape changes that may occur between each examination; Composition, where corresponding to the reference image Reference camera position It is fixed and known; the changes in camera pose are used to simulate the inevitable changes in eye pose between each examination.

[0022] A. Dynamic eyeball model This patent approximates a dynamic eyeball model by simulating the changes in the lengths of the three axes of the eye as they change with natural growth or disease progression. Specifically, such as... Figure 2 As shown, this patent uses n+1 eyeballs. , corresponding to the reference image and the remaining images For shape, all eyeballs are approximated by ellipsoids, with three orthogonal semi-axes between each ellipsoid. Since their lengths differ, they can be optimized separately. Regarding pose, the pose parameters of all ellipsoids are completely identical. The centers of these ellipsoids are located at the origin of the world coordinate system. The orientation of these ellipsoids is determined by three orthogonal semi-axises. Around the world coordinate axis Angle of rotation This indicates that the final equation for the dynamic eyeball model obtained by this patent is as follows:

[0023] in It is the eyeball Three-dimensional points on the surface As can be seen, the parameters in the dynamic eye model consist of the following two parts: 1) the shape of all ellipsoids. ;2) The pose of each ellipsoid .

[0024] Furthermore, this patent establishes a novel mapping scheme that maps corresponding points between different fundus surfaces to the same fundus surface, facilitating further alignment constraints on key matching points in three-dimensional space. Specifically, this patent employs the concept of limits, assuming that the size of the ellipsoid continuously decreases as the three orthogonal semi-axis shortens, eventually becoming an infinitesimal ellipsoid located at the center of the original ellipsoid. Then, the line connecting corresponding points between these two ellipsoids can be approximated as the line connecting a point on the surface of the original ellipsoid to its center point. Since all ellipsoids in the dynamic eyeball model share the same center point, this patent uses this center point as the starting point, approximating the intersection points of light rays emitted from this point on several fundus surfaces as a set of mutually mappable corresponding points. For example... Figure 2 As shown, the specific mapping process can be described as follows: using the center point of the dynamic eyeball model... Starting from the eyeball Points on the curved surface of the fundus A ray is formed, as shown in equation (2), and this ray interacts with the eyeball. The intersection of the fundus surface (1) That is, the corresponding point after mapping.

[0025]

[0026] in, and It is a proportionality coefficient, and .

[0027] B. Camera Array Model Determining the camera in a 3D scene typically involves two elements: (1) an intrinsic parameter matrix that is only related to the camera itself; and (2) an extrinsic parameter matrix that reflects the camera pose. Considering that this patent transforms changes in eye pose into changes in camera pose, and that the camera used in each fundus examination may be different, this patent approximates the camera array model in a longitudinal 3D fundus examination scene by setting the intrinsic and extrinsic parameter matrices of the camera used in each fundus examination.

[0028] Specifically, such as Figure 2 As shown, this patent uses n+1 cameras. , corresponding to the reference image and the remaining images Regarding the intrinsic parameter matrix, given a sequence of retinal images, the intrinsic parameter matrix for each corresponding camera can be directly calculated using equations (3) and (4). .

[0029]

[0030] in, and A camera that measures pixels. The focal lengths were all approximated using the method proposed by Hernandez-Matas et al., as shown in equation (4); Figure 3 As shown, It is an image The coordinates of the center point in the pixel coordinate system can be calculated as the image. Half of its width and height.

[0031]

[0032] Among them, such as Figure 3 As shown, It is an image The radius of the retinal region shown in the image. and These are the distance from the lens to the cornea and the field of view, respectively, which can be obtained from a fundus camera. The information can be found in the device specifications (if unknown, the specifications of commonly used fundus cameras can be used). This is a rough estimate of the radius of the spherical eye model. .

[0033] Regarding the extrinsic parameter matrix, it consists of two parts: a rotation matrix R and a translation vector t, which determine the camera pose in the scene. By using equation (5), this patent will perform a longitudinal three-dimensional fundus examination of the scene corresponding to the reference image. Reference camera position Set to be fixed and known. Remaining cameras in the array. pose relative to reference camera Adjustments and optimizations were made, and the specific form is shown in equation (6).

[0034]

[0035]

[0036] in, It's a camera. The angle of rotation relative to the world coordinate system, The origin of the world coordinate system is at the camera. The position in the camera coordinate system.

[0037] Furthermore, fundus cameras typically exhibit some distortion during retinal imaging, which is actually part of the inherent parameters of cameras used in practice. In previous studies, this patent determined that the basic form of fundus camera distortion is fourth-order radial distortion. Therefore, this patent adds a fourth-order radial distortion model to each camera in the array, as shown in Equation (7).

[0038]

[0039] in, and These are points on the normalized image plane before and after distortion. It's a camera. The fourth-order radial distortion coefficient reflects the severity of the distortion.

[0040] Finally, this patent concludes that, given a sequence of retinal images, the parameters of the camera array model consist of the following two parts: 1) the poses of all remaining cameras except the reference camera; and ;2) Fourth-order radial distortion coefficients for all cameras, .

[0041] C. Mapping between three-dimensional retinal points and two-dimensional image points The mapping between 3D retinal points and 2D image points serves as a bridge connecting the dynamic eye model and the camera array model. It involves two opposing processes: 1) mapping from 3D retinal points to 2D image points; and 2) mapping from 2D image points back to 3D retinal points. The former is essentially the process of imaging the retina with a camera, while the latter is equivalent to the process of light rays emitted from the camera intersecting with the retina. The specific process is as follows: Figure 4 As shown.

[0042] like Figure 4 As shown in red, the mapping process from three-dimensional retinal points to two-dimensional image points is described in detail below: 1) Using the extrinsic parameter matrix of a fundus camera Transform the coordinates of the three-dimensional retina point from the world coordinate system Transform to camera coordinate system ,Right now ; 2) Transform to normalized coordinate system The normalized image plane in the middle is obtained ,Right now ; 3) Use formula (7) to calculate the distortion after adding distortion. coordinates ; 4) Using the intrinsic parameter matrix K of the fundus camera to... Transform to pixel coordinate system u - v In the process, obtain two-dimensional image points ,Right now .

[0043] like Figure 4 The green portion shows the mapping process from two-dimensional image points to three-dimensional retinal points, described in detail below: 1) For two-dimensional image points... Distortion correction is performed to obtain distortion-free two-dimensional image points. 2) Connecting two-dimensional image points and the center point of the camera This forms a ray pointing towards the fundus (Equation (8)). The intersection of this ray with the posterior hemisphere of the ellipsoid representing the fundus is the corresponding three-dimensional retinal point. The posterior hemisphere of the fundus is shown in equation (1).

[0044]

[0045] in, The coefficients are obtained by combining equations (8) and (1). It is the camera's projection matrix.

[0046] 2. Implementation of the RISeR framework The method RISeR proposed in this patent uses corresponding points between images as a medium to optimize the parameters of a constructed three-dimensional spatiotemporal model under joint alignment constraints of frame-to-reference and frame-to-frame, ultimately obtaining a longitudinal three-dimensional fundus examination scene for registering a given retinal image sequence. Figure 5 The workflow of the RISeR framework is summarized. The key steps in this workflow will be described in detail in the following sections.

[0047] A. Keypoint Extraction and Matching This step directly employs existing, commonly used methods to extract and match keypoints between retinal images. Specifically, the SuperPoint algorithm proposed by DeTone et al. is used to detect and describe keypoints in retinal images, and then the SuperGlue algorithm proposed by Sarlin et al. is used to match keypoint sets extracted from different retinal images. It is important to note that this patent does not perform any additional network structure design or model weight fine-tuning for SuperPoint and SuperGlue; it only selects and sets some of their hyperparameters. Specifically, this patent sets the confidence threshold for keypoint detection in SuperPoint to 0.001, selects the "outdoor" type weight file loaded in SuperGlue, and uses the default values ​​for the remaining hyperparameters.

[0048] B. Initialization of the three-dimensional spatiotemporal model Given a sequence of retinal images The unknown parameters in the corresponding three-dimensional spatiotemporal model consist of the following four parts: 1) the shape of all ellipsoids in the dynamic eyeball model; ;2) The pose of the dynamic eye model, 3) The poses of all remaining cameras in the camera array model, excluding the reference camera. ,Right now and ;4) The fourth-order radial distortion coefficients of all cameras in the camera array model, .

[0049] The above unknown parameters are initialized as follows: 1) Based on the Navarro eye model, the initial shape of each eyeball is set to a radius of 12. mm A sphere; 2) All eyeballs have the same pose, and the initial values ​​are set to have no rotation relative to the world coordinate system; 3) The poses of all remaining cameras are initialized by solving the Perspective-n-Point (PnP) problem within the RANSAC framework. The PnP problem refers to estimating the camera pose {R, t} using a set of 2D-3D corresponding points and the camera's intrinsic parameter matrix K. Here, the 2D points are the matched keypoints in the retinal image corresponding to the target camera. Their corresponding 3D points are obtained by mapping the matched keypoints in the reference image to the fundus surface of the initial eyeball; 4) The fourth-order radial distortion coefficients of all cameras are initialized using the calibration results from the previous work [2] of this patent.

[0050] C. Optimization of the three-dimensional spatiotemporal model like Figure 5 As shown, the RISeR framework involves two different types of optimization: 1) optimization under a single constraint on frame-to-reference; and 2) optimization under joint constraints on frame-to-reference and frame-to-frame. The former is used for registration sequences excluding the reference image. F The first image outside of 0 F 1, while the latter is used for further registration of subsequent images in the sequence. They all use the Particle Swarm Optimization (PSO) algorithm to achieve... p Particle iteration g The optimal candidate solution is found in the given search space.

[0051] For the first type of optimization, such as Figure 2 As shown, the calculation method of its objective function is described as follows: 1) F 0 and F Matched in 1 Mapping the key points onto their respective fundus surfaces yields... Group corresponding points ; 2) Put p t Mapping to q t On the curved surface of the fundus, we obtained ; 3) Calculate according to formula (9) Euclidean distance ; 4) Sort in ascending order to get The objective function is to minimize the sum of the first 80% of the distances (to reduce the images of mismatched keypoints), as shown in Equation (10).

[0052]

[0053]

[0054] in, and Representing reference images F 0 and remaining images F The corresponding model parameters to be optimized are 1.

[0055] For the second type of optimization, such as Figure 2 As shown, the calculation method of its objective function is described as follows: 1) F 0 and F n Matched Mapping the key points onto their respective fundus surfaces yields... Group corresponding points ; 2) G t Mapping to q t On the curved surface of the fundus, we obtained ; 3) Calculate the three-dimensional space according to equation (11) Euclidean distance ; 4) F n-1 and F n Matched Mapping the key points onto their respective fundus surfaces yields... Group corresponding points ; 5) G t Mapping to p t On the curved surface of the fundus, we obtained ; 6) Calculate the three-dimensional space according to equation (12). Euclidean distance ; 7) and Sort in ascending order to get and ; 8) Calculate the sum of the first 80% of the values ​​in each of the two groups, and then add a certain weight to these two sums (to balance the two parts) as the objective function to be minimized, as shown in equation (13).

[0056]

[0057]

[0058]

[0059] in, Represents the remaining imageF n The corresponding model parameters to be optimized.

[0060] Additionally, it should be noted that after each optimization, a subset of unknown parameters in the 3D spatiotemporal model are determined and used as fixed known values ​​in subsequent optimizations. This process continues until all n optimizations are completed, at which point all unknown parameters in the 3D spatiotemporal model are determined, forming a longitudinal fundus examination scenario simulating the imaging process of a given retinal image sequence. In this scenario, the retinal image to be registered in the sequence is mapped to the viewpoint of the reference image to obtain multiple sets of two-dimensional floating-point coordinates and corresponding RGB values. Then, bilinear interpolation is used to generate the registered retinal image sequence.

[0061] D. Generation of synthetic retinal image sequences Because of its three-dimensional spatiotemporal simulation of longitudinal fundus examination scenarios, the RISeR framework can generate a large number of realistic synthetic retinal image sequences from a single input image through a series of physically meaningful transformations. This has significant potential applications in areas such as registration method evaluation, deep learning-based model training, and data augmentation. Specifically, a synthetic retinal image sequence... The generation mainly includes the following steps: 1) Use the given retinal image as the reference image. F 0, based on which the relevant parameters of the reference camera and reference eyeball in the 3D spatiotemporal model are set. ; 2) Use certain strategies to set the parameters of the remaining cameras in the camera array model and the remaining eyes in the dynamic eye model. Ultimately, a complete three-dimensional spatiotemporal model is constructed to simulate the longitudinal fundus examination scenario; 3) In this scenario, synthesize a sequence of retinal images. It is generated by mapping image points between a reference camera and a residual camera, bilinear interpolation, and optional color space transformation (to simulate changes in image color, brightness, contrast, etc. due to changes in camera, camera settings, or environment).

[0062] As shown in Table 1, based on the strategies adopted when setting the parameters of the remaining camera and eyeball in the 3D spatiotemporal model, and whether color space transformation is performed when generating images, five different types of longitudinal fundus examination scenarios can be constructed: 1) P Simulates a series of retinal images captured as the eye moves during the same fundus photography session; 2) PC Simulates a series of retinal images captured from an eye without any shape changes using the same fundus camera during different examinations; 3) PCI Simulates a series of retinal images captured from an eye without any shape changes using different fundus cameras during different examinations; 4) PCS Simulates a series of retinal images captured from an eye with shape changes using the same fundus camera during different examinations; 5) PCSI Simulates a series of retinal images captured from an eye with shape changes using different fundus cameras during different examinations.

[0063] Using a retinal image as input, five different types of synthetic retinal image sequences are generated based on the above scenario, as shown in the following figures. Figure 6 As shown.

[0064] Table 1. Construction of five different types of longitudinal funduscopy examination scenarios

[0065] The model built using this method will be validated using a specific dataset in the following section.

[0066] 1. Dataset Five datasets were used in the experiments of this invention. The first two are publicly available retinal image pair registration datasets, including the Fundus Image Registration (FIRE) dataset shared by Hernandez-Matas et al. and the Fundus Image Myopia Development (FIMD) dataset shared in previous work of this patent. The next two are newly created retinal image sequence registration datasets, including the real dataset (RISeR-real) and the synthetic dataset (RISeR-syn). The last is a publicly available retinal image registration dataset containing the progression of retinopathy of preterm birth (ROP), called COph100, shared by Hu et al.

[0067] FIRE contains 134 pairs of retinal images, divided into three categories: 1) 71 pairs of images with large overlapping areas and no anatomical changes were classified as... S Class 2) 49 pairs of images with small overlapping areas and no anatomical changes were classified as P Class; 3) 14 pairs of images with large overlapping areas and significant anatomical changes were classified as A These anatomical changes are caused by the progression of retinal disease, mainly manifested as increased vascular tortuosity and the appearance of microaneurysms. Ten pairs of widely distributed and evenly distributed ground truth points (also called control points) were manually annotated for each pair of retinal images as a medium for quantitatively assessing registration accuracy. However, it should be noted that... P The labeling of the 6th pair of control points in the 37th pair of retinal images was incorrect, therefore this patent uniformly excluded this pair of control points in the evaluation of subsequent experiments.

[0068] FIMD consists of 70 pairs of retinal images, each showing significant anatomical changes due to myopia progression. Similar to FIRE, each pair of images was manually annotated with 12 control points. Furthermore, the long time interval between image acquisitions and the use of different fundus cameras resulted in variations in resolution, brightness, contrast, and visual field between the two images.

[0069] RISeR-real contains 10 longitudinal retinal image sequences, each from one of ten different eyes. Each sequence consists of five or six retinal images taken at one- or two-year intervals using different fundus cameras. Furthermore, each retinal image sequence shows no other fundus diseases besides obvious myopia onset and progression. Figure 7 As shown, this patent manually marks 12 sets of control points for each retinal image sequence.

[0070] RISeR-syn is generated using the method described in Section 2-D of this patent. Specifically, this patent selects 10 retinal images from different eyes, using one of them as a given reference image each time, to generate 5 different types of synthetic retinal image sequences, for a total of 50 sequences, each consisting of 6 synthetic retinal images. Compared to RISeR-real, which is time-consuming and prone to errors when manually annotating control points, RISeR-syn can easily obtain densely distributed and precisely corresponding control points during generation, thus achieving a more objective and effective registration evaluation. Figure 8 This displays all the control points labeled in a sequence, covering the common retinal area of ​​all images. Furthermore, assuming the first image in the sequence is used as the reference frame, such as... Figure 9 As shown, this patent also marks control points that exist only in the common retinal region between two adjacent frames but not in the reference frame, in order to more comprehensively evaluate the inter-frame continuity of the sequence registration results.

[0071] COph100 is a retinal image registration dataset specifically focused on advancements in ROP (Retinal Optical Process). This dataset consists of 100 eyes, each with 2 to 9 retinal images collected at different examination times. These images show significant anatomical changes due to ROP advancements. Similarly, to evaluate the registration results, all retinal images for each eye were manually annotated with 10 sets of control points, primarily located around vascular crossovers. Finally, it should be noted that to better evaluate the sequence registration performance of the proposed method, this patent selected 24 eyes from COph100 that underwent 4 or more subsequent examinations for experiments. This portion of the data is named COph100-45689 to distinguish it from the complete dataset containing all eyes.

[0072] 2. Experimental Setup The RISeR framework was implemented using PyTorch and OpenCV. All RISeR-related registration experiments were conducted on an Intel(R) Core(TM) i5-12500 CPU @ 3.00 GHz with 16.0 GB of RAM. The PSO algorithm used was the official implementation pyswarms.single.GlobalBestPSO, and its hyperparameters were set as follows in the experiments: (1) Number of particles p = 10000, number of iterations g = 300; (2) Cognitive parameters c 1 = 1.6, social parameter c 2 = 1.4, inertial parameter w = 0.6, with update strategies of "nonlin_mod", "lin_variation", and "exp_decay" respectively; (3) The search spaces for eye shape parameters, eye pose parameters, camera rotation parameters, camera translation parameters, and camera distortion parameters are 4.0 and 4.0, respectively. mm 2.0 rad 1.0 rad 2.0 mm The maximum search step size is 0.1, and the maximum search step size is 1.0. mm 0.01 rad 0.01 rad 0.1 mm and 0.01; (4) The position and velocity adjustment strategies for out-of-bounds particles are “nearest” and “zero”, respectively.

[0073] Finally, it should be noted that this patent does not set any additional optimization termination criteria; optimization will only stop when the preset maximum number of iterations is reached.

[0074] 3. Experimental results on FIRE and FIMD 1) Evaluation index: As shown in Equation (14), each retinal image pair Registration error e By applying registration transformation through calculation Then mark the control points. Average Euclidean distance Received, among which This indicates the number of labeled control point pairs. Based on this, this patent further employs the following two methods to process all retinal image pairs in the dataset. Registration error Perform statistical analysis and use it as an indicator to evaluate the registration results of the algorithm on this dataset.

[0075]

[0076] First, this patent calculates The average value is used to obtain the average registration error. It represents the overall registration accuracy of a method on a dataset.

[0077]

[0078] in N It represents the number of retinal image pairs in the dataset.

[0079] Secondly, this patent statistically analyzes different error thresholds. (1-25 pixels) The registration success rate is calculated, and then the area under the curve (AUC) is calculated to evaluate the overall registration success rate of a method on the dataset.

[0080]

[0081] Wherein is the indicator function.

[0082] 2) Registration Methods: This patent selects four classic methods (GDB-ICP, SURF-PIIFD-RPM, GFEMR, REMPE) and three newly proposed advanced methods (SuperRetina, GeoFormer, RIRMD) for comparison. This patent uses publicly available implementations and retains default parameter settings to perform the registration experiments. For SuperRetina and GeoFormer, this patent uses model weights shared by the authors and runs them on an NVIDIA GeForce RTX 4090 GPU. The remaining five comparison methods are run on the same device as RISeR.

[0083] 3) Results: Table 2 records the average registration error of all registration methods on the FIRE and FIMD datasets. And AUC. As seen in this patent, the proposed method RISeR achieves the best registration results on both datasets and has significant advantages. Furthermore, only the proposed method RISeR achieves the best or second-best results in all columns of Table 2, further demonstrating that the proposed method RISeR also possesses very good stability and robustness. Finally, it is worth noting that the image pairs in FIMD and FIRE-A respectively contain anatomical changes caused by myopia progression and retinal disease progression, and the proposed method RISeR achieves the best registration results on both datasets. This initially indicates that RISeR has good applicability for registration of retinal images with anatomical changes caused by myopia progression or non-myopia fundus disease progression.

[0084] Table 2 Evaluation results on the FIRE and FIMD datasets

[0085] 4. Experimental results on RISeR-real 1) Evaluation metrics: For a retinal image sequence By estimating a set of transformations All remaining images All were registered to the selected reference image. From this perspective. In order to more comprehensively evaluate The registration results, this patent utilizes Group label control points The following two indicators are calculated. First, as shown in equation (17), this patent calculates... Control points after transformation and The corresponding control points Average registration error between This metric reflects the performance of an algorithm. Overall registration accuracy.

[0086]

[0087] Secondly, as shown in equation (18), this patent calculates Control points transformed between all adjacent frames Average drift distance This indicator reflects the registration result. Inter-frame continuity. The smaller the value of this index, the better the consistency of the registration results, and the smoother and more natural the switching between adjacent frames, that is, the better the continuity.

[0088]

[0089] 2) Registration Method: Existing global registration methods have limited research on image sequence registration. A common approach is to directly apply pairwise registration methods multiple times, independently registering the remaining images to the selected reference image. The few global registration methods specifically designed for image sequences are also difficult to compare because their implementations are not publicly available. Therefore, this patent selects the most competitive contrast method SuperRetina and highly relevant methods (REMPE, RIRMD) on the FIMD dataset, and performs registration on the RISeR-real dataset using the aforementioned common approach. Furthermore, to explore the effectiveness of the frame-to-reference and frame-to-frame joint registration strategies employed by the proposed RISeR method, this patent also conducted experiments on an ablation version of RISeR (RISeR-F2R) using a single frame-to-reference registration strategy.

[0090] 3) Results: The experimental results are shown in Table 3, which records the average registration error of the proposed method (RISeR), its ablation version (RISeR-F2R), and the comparison methods (REMPE, SuperRetina, RIRMD) on the RISeR-real dataset. and average drift distance As can be seen from this patent, RISeR achieves the highest registration accuracy and the best inter-frame continuity, with a significant improvement over other comparative methods. This demonstrates the superior performance of the proposed method in retinal image sequence registration tasks. Furthermore, this patent compares the ablation experimental results of RISeR and RISeR-F2R, finding that applying frame-to-reference and frame-to-frame joint alignment constraints can significantly improve the inter-frame continuity of retinal image sequence registration results and help the transformation model more accurately simulate fundus examination scenarios, thereby obtaining better registration results.

[0091] Table 3 Evaluation results on the RISeR-real dataset

[0092] Finally, this patent qualitatively compares the registration results of RISeR and REMPE on the RISeR-real dataset. For example... Figure 10 As shown, the left two columns belong to RISeR, and the right two columns belong to REMPE. The first three rows are the registration results from the current frame to the reference frame, and the last three rows are the registration results from the current frame to the previous frame. Generally speaking, for blood vessels that pass through multiple image blocks, the smoother and more natural the connection between adjacent image blocks, the better the registration effect. This patent shows that, compared with REMPE, the proposed method RISeR can achieve better registration results in both coarser main vessel regions and finer branch vessel regions.

[0093] 5. Experimental results on RISeR-syn 1) Evaluation index: In addition to the average registration error calculated by equation (17) and the average drift distance calculated by equation (18) ,like Figure 9 As shown, this patent also uses only those existing in Two adjacent frames The common retinal region between them, but not in the reference frame In the retinal region Additional control points for the group Calculate the average alignment deviation The specific calculation process is shown in equation (19), where This represents the length of set J. It's important to note that this is not... There are such control points between each pair of adjacent frames in the data, therefore J is... A subset of. In fact, this indicator is more than The inter-frame continuity of sequence registration results is evaluated more rigorously, therefore its value is typically greater than [value missing]. The value of .

[0094]

[0095] 2) Registration method: Since REMPE performs relatively poorly on the RISeR-real dataset, this method is not used. The other registration methods are the same as those used in the experiments on RISeR-real.

[0096] 3) Results: Table 4 records the average registration errors of the proposed method (RISeR), its ablation version (RISeR-F2R), and comparative methods (SuperRetina, RIRMD) on the RISeR-syn dataset. Average drift distance and average alignment deviation The bold text in the table represents the best results. As can be seen from this patent, RISeR still achieves the highest registration accuracy and the best inter-frame continuity, with significant advantages. Furthermore, RISeR performs well in five significantly different simulated fundus camera scenarios (…). P , PC , PCI , PCS , PCSI The best sequence registration results were achieved in both methods, further demonstrating that the method also has good stability and robustness. Similarly, the analysis of ablation experiments of RISeR and RISeR-F2R in this patent shows that applying joint alignment constraints of frame-to-reference and frame-to-frame can significantly improve the inter-frame continuity of the sequence registration results.

[0097] Table 4 Evaluation results on the RISeR-syn dataset

[0098] 6. Experimental results on COph100-45689 1) Evaluation metrics: Due to the high registration difficulty of certain retinal image sequences in COph100-45689, some methods may encounter registration failures, i.e., registration errors are particularly large (>50 pixels) or cannot output registration results at all. Therefore, this patent first counts the number of failed registration sequences. Then, this patent uses equations (17) and (18) to calculate the average registration error of successfully registered retinal image sequences. and average drift distance .

[0099] 2) Registration method: The same comparison method as the experiment on RISeR-real (REMPE, SuperRetina, RIRMD) was used.

[0100] 3) Results: The experimental results are shown in Table 5. It can be seen that RISeR not only achieves zero-failure registration but also achieves the highest registration accuracy and the best inter-frame continuity, demonstrating significant advantages. This further proves that the proposed method is also well-suited for the progression of non-myopic fundus diseases.

[0101] Table 5 Evaluation results on the COph100-45689 dataset

[0102] This patent proposes a retinal image sequence registration method for longitudinal studies, comprising two key elements: a three-dimensional spatiotemporal model and a joint registration strategy. The former simulates a longitudinal fundus examination scenario by constructing a dynamic eyeball model and a camera array model. Benefiting from this, the patent can efficiently generate a large number of realistic synthetic retinal image sequences. The latter specifically considers the inter-frame continuity, which is also of concern in sequence registration tasks, and develops an algorithmic framework for registering retinal image sequences based on a joint alignment constraint strategy of frame-to-reference and frame-to-frame. Experimental results show that this method achieves state-of-the-art registration accuracy, reliability, and applicability on two publicly available retinal image pair registration datasets, two newly constructed retinal image sequence registration datasets, and one publicly available retinal image sequence registration dataset.

[0103] The present invention has been described in detail above through embodiments, but the content is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the patent coverage of the present invention.

Claims

1. A retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling, characterized in that: The method, based on the RISeR method, constructs a three-dimensional spatiotemporal model for a longitudinal three-dimensional fundus examination scene. The three-dimensional spatiotemporal model includes a dynamic eyeball model and a camera array model. The RISeR method includes the following steps: S1: Input retinal image sequence { F 0, F 1, ..., F n }, F 0 is the selected reference image in the sequence; S2: Calculate the reference image F The intrinsic parameter matrix K0 and the extrinsic parameter matrix {R0, t0} of 0; S3: Reference image F 0 and remaining images F 1. Perform key point extraction and matching, and perform pairwise registration, using frame-to-reference and frame-to-frame strategies to register the remaining images in the sequence; S4: After registration, output the corresponding model parameters in the longitudinal three-dimensional fundus examination scene; S5: Register and transform all remaining images to the viewpoint of the reference image; S6: Output the registered retinal image sequence { F 0, F 1→0 , ..., F n→0 } 2. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 1, characterized in that: The dynamic eyeball model consists of different shapes but the same pose, using a world coordinate system. origin Multiple ellipsoids centered on The system is designed to simulate the shape changes that may occur in the eye between each examination; the camera array model consists of multiple cameras in different poses. Composition, where corresponding to the reference image F 0 reference camera C Pose 0 It is fixed and known that the camera pose change is used to simulate the inevitable pose change of the eye between each examination.

3. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 2, characterized in that: In the dynamic eyeball model, all eyeballs are approximated by ellipsoids, and the three orthogonal semi-axis of each ellipsoid are defined. The lengths of the ellipsoids differ, and optimizations are performed separately for each. Regarding pose, it is assumed that the pose parameters of all the ellipsoids are completely identical. It is also assumed that the size of the ellipsoid decreases continuously as the three orthogonal semi-axis shortens, eventually becoming an infinitesimal ellipsoid located at the center of the original ellipsoid. The line connecting corresponding points between two ellipsoids is approximately the line connecting a point on the surface of the original ellipsoid to its center point, with the center point of the dynamic eyeball model as the reference point. Starting from the eyeball Points on the curved surface of the fundus The ray forms and interacts with the eyeball. The intersection of the fundus surfaces That is, the corresponding point after mapping: ; in, and It is a proportionality coefficient, and .

4. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 2, characterized in that: The camera array model contains n+1 cameras. Corresponding to the reference image and the remaining image Given a sequence of retinal images, the parameters of the camera array model consist of the following two parts: 1) the poses of all remaining cameras except the reference camera; and ;2) Fourth-order radial distortion coefficients for all cameras, .

5. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 1, characterized in that: The dynamic eye model and camera array model are connected by a mapping between three-dimensional retinal points and two-dimensional image points. This includes two opposite processes: 1) mapping from three-dimensional retinal points to two-dimensional image points, which is essentially the process of imaging the fundus retina with a camera; 2) mapping from two-dimensional image points to three-dimensional retinal points, which is equivalent to the process of light rays emitted from the camera intersecting with the fundus.

6. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 1, characterized in that: In S3, the reference image F 0 and remaining images F 1. Perform pairwise registration: 1) Keypoint extraction and matching are performed using the SuperPoint and SuperGlue algorithms; 2) Calculate the remaining image F The intrinsic parameter matrix K1 of 1; 3) Initialize the reference image F 0 and remaining images F 1 corresponds to the unknown model parameters; 4) Optimize the reference image F 0 and remaining images F The model parameters corresponding to 1.

7. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 1, characterized in that: The RISeR framework involves two different types of optimization: 1) Optimization under a single frame-to-reference constraint, used for registration sequences excluding the reference image. F The first image outside of 0 F 1; 2) Optimization under joint frame-to-reference and frame-to-frame constraints, used for further registration of subsequent images in the sequence. Both optimization methods use the particle swarm optimization algorithm. p Particle iteration g The optimal candidate solution is found in the given search space.

8. The retinal image sequence registration method based on longitudinal three-dimensional fundus examination scene modeling according to claim 1, characterized in that: In S6, a synthetic retinal image sequence The generation mainly includes the following steps: 1) Use the given retinal image as the reference image. F 0, based on which the relevant parameters of the reference camera and reference eyeball in the 3D spatiotemporal model are set. ; 2) Use appropriate strategies to set the parameters of the remaining cameras in the camera array model and the remaining eyes in the dynamic eye model. Ultimately, a complete three-dimensional spatiotemporal model is constructed to simulate the longitudinal fundus examination scenario; 3) In this scenario, synthesize a sequence of retinal images. It is generated by mapping image points between a reference camera and a residual camera, bilinear interpolation, and optional color space transformation, which simulates changes in image color, brightness, and contrast caused by variations in camera, camera settings, or environment.