Multi-mode fundus imaging system based on FPGA and real-time eye movement compensation method

The multimodal fundus imaging system implemented by FPGA adopts a dual-wavelength parallel scanning beam architecture and hardware-level eye movement compensation, which solves the feedback lag effect of the multimodal fundus imaging system, realizes real-time synchronization and high-precision fusion of cSLO and OCT data streams, and improves image quality and repeatability of follow-up scans.

CN122004745APending Publication Date: 2026-05-12Gaoshi Innovation Technology Co., Ltd.
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Gaoshi Innovation Technology Co., Ltd.
Filing Date
2026-01-19
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing multimodal fundus imaging systems suffer from feedback lag due to their serial processing architecture, making it impossible to complete eye movement compensation before the start of the next imaging scan frame or scan block. This prevents deterministic synchronization and real-time fusion display of cSLO and OCT data streams, and makes it difficult to guarantee accurate automatic reproduction of longitudinal follow-up scan positions.

Method used

A multimodal fundus imaging system based on FPGA is adopted. Through a dual-wavelength parallel scanning beam architecture and the synchronization and control mechanism inside the FPGA, eyeball displacement is calculated in real time and hardware-level compensation is performed. Combined with a parallel processing pipeline, stable synchronization and high-precision spatial alignment of multimodal imaging data are achieved.

Benefits of technology

It significantly improves the spatial consistency of multimodal images, reduces motion artifacts, enhances the repeatability and comparability of follow-up scans, simplifies the system structure, reduces power consumption and cost, and supports flexible system updates and expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122004745A_ABST
    Figure CN122004745A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode fundus imaging system based on an FPGA and a real-time eye movement compensation method, and belongs to the field of medical imaging equipment. The system comprises an optical imaging and tracking module, an FPGA processing control module, a storage module and a system control and display unit. A unified frame-level clock management unit is arranged in the FPGA module and is used for synchronously scheduling scanning of the tracking light beam and the imaging light beam; eyeball displacement parameters are solved in real time through a hardware feature matching unit, and initial coordinates of cSLO and SD-OCT scanning galvanometers are corrected in a feedforward mode according to the parameters before each imaging frame or scanning block starts, so that frame-level eye movement compensation is achieved; meanwhile, preprocessing, displacement-based multi-frame alignment accumulation and image registration and fusion are performed on cSLO and SD-OCT data streams in parallel in a hardware pipeline. According to the invention, real-time synchronization, motion artifact suppression and high-precision fusion of multi-modal data are realized through full-hardware processing, and the imaging quality, the system response speed and the repeatability of follow-up scanning are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of medical imaging equipment, and particularly relates to a multi-modal fundus imaging system based on FPGA and a real-time eye movement compensation method. BACKGROUND

[0002] Multi-modal fundus imaging technology is an important tool for modern ophthalmic diagnosis. It integrates the advantages of different imaging modalities to provide more comprehensive information for screening, diagnosis and follow-up of retinal and choroidal diseases. Among them, confocal scanning laser ophthalmoscope (cSLO) can provide high-contrast two-dimensional images of the retinal surface, while spectral domain optical coherence tomography (SD-OCT) can obtain three-dimensional high-resolution tomographic structure of each layer of the retina. Combining the wide-area high-speed imaging capability of cSLO with the depth resolution capability of SD-OCT can comprehensively evaluate the morphological and functional changes of diseases such as glaucoma, age-related macular degeneration, and diabetic retinopathy.

[0003] However, high-quality multi-modal fusion imaging faces many technical challenges, and the core difficulty comes from the physiological movement of the eyeball. Even with well-coordinated patients, there are unavoidable micro-tremor, drift and micro-saccades. These movements occur on the frame or scan block time scale and can cause the following problems: (1) scanning offset and motion artifacts: in a single scan (especially a time-consuming three-dimensional OCT volume scan), eye movement can cause scanning line misalignment or distortion, resulting in trailing, broken or repeated artifacts, which significantly reduces image quality; (2) multi-modal image registration error: even if cSLO and OCT images are collected synchronously, due to the lack of a unified motion compensation reference, subsequent software registration still has spatial misalignment, affecting the accuracy of image fusion and diagnostic accuracy. (3) Insufficient consistency of follow-up scans: in long-term monitoring of diseases, the alignment accuracy of scan images at different time points directly affects the quantitative evaluation of lesion changes, and traditional manual or software post-processing registration methods are inefficient and have limited repeatability.

[0004] To address eye-tracking interference, existing technologies mainly employ two approaches: one is to use software algorithms for post-processing registration and artifact correction, but this approach suffers from drawbacks such as large processing delays, inability to guide imaging in real time, and limited ability to correct large-scale motions; the other approach is to introduce active eye-tracking technology at the hardware level. However, existing hardware solutions still have room for improvement: (1) Insufficient real-time performance and determinism: Tracking, compensation, and multimodal image processing usually rely on general-purpose processors or distributed hardware modules, which are prone to performance bottlenecks under high frame rates and large data streams; (2) Limited real-time fusion display and frame-level compensation capabilities: Existing systems cannot achieve frame-level feedforward compensation and real-time image fusion during scanning; (3) Low system integration: Functions such as motion tracking, frame-level image denoising, multi-frame cumulative averaging, and automatic follow-up alignment (AutoRescan) are not deeply integrated and accelerated in parallel at the hardware level, resulting in increased power consumption, complexity, and cost.

[0005] The aforementioned shortcomings stem from the inherently serial processing architecture of existing systems. The acquisition, processing, compensation command generation, and feedback of eye-tracking signals inevitably involve delays, resulting in a severe "feedback lag effect." Therefore, a novel technical solution is urgently needed to achieve frame-level multimodal data synchronization, eye-tracking compensation, and real-time image processing on underlying hardware such as FPGAs, thereby comprehensively improving the real-time performance, stability, fusion accuracy, and clinical application efficiency of multimodal fundus imaging. Summary of the Invention

[0006] [Technical Issues] The technical problem to be solved by the present invention is to overcome the "feedback lag effect" caused by the serial processing architecture of the existing multimodal fundus imaging system, which is specifically manifested in: 1) the inability to complete eye movement compensation before the start of the next imaging scan frame or scan block; 2) the inability to achieve deterministic synchronization and real-time fusion display of cSLO and OCT data streams; and 3) the difficulty in ensuring accurate automatic reproduction of the longitudinal follow-up scan position.

[0007] [Technical Solution] To address the aforementioned technical problems, this invention provides a multimodal fundus imaging system based on FPGA. This system employs a dual-wavelength parallel scanning beam architecture, including a near-infrared reference tracking beam for continuously acquiring a reference field of view and an imaging beam for performing confocal laser scanning fundus imaging (cSLO), and an imaging beam for spectral-domain optical coherence tomography (SD-OCT). The reference tracking beam and the cSLO imaging beam share a first set of scanning mirrors for scanning modulation, while the SD-OCT imaging beam uses a separate second set of scanning mirrors. The output optical paths of the two sets of scanning mirrors are coaxially propagated at the rear end via a beam combiner and splitter, thereby ensuring the consistency of the scanning field of view of the three imaging beams on the examined fundus plane.

[0008] An internal synchronization and control mechanism, with scan frames or scan blocks as the smallest granularity, is established within the FPGA. Hardware feature matching and displacement calculation are performed on the continuous reference images acquired by the reference tracking beam. Before the start of each imaging scan frame, the starting coordinates and scanning trajectory of the imaging scan beam are updated in real time to compensate for low-frequency displacement errors caused by eye drift or saccades. For high-frequency motion components such as microsaccades, suppression is achieved through subsequent multi-frame cumulative averaging and image registration strategies based on displacement alignment. Through the aforementioned dual-beam parallel co-path architecture and FPGA hardware-level parallel processing pipeline, this invention achieves stable synchronization and high-precision spatial alignment of multimodal imaging data, significantly improves the spatial consistency of multimodal images, effectively reduces motion artifacts, and enhances the repeatability and comparability of follow-up scans.

[0009] The system includes: The optical imaging and tracking module includes: a confocal laser scanning fundus imaging unit for generating and utilizing a first set of scanning galvanometers to scan a common-path propagating imaging beam and an eye movement tracking beam; and a spectral-domain optical coherence tomography unit for generating and utilizing a second set of scanning galvanometers to scan an independent tomographic scanning beam. The FPGA processing control module has a built-in unified frame-level clock management unit and integrates a parallel processing pipeline built from hardware logic resources. The pipeline includes: a real-time eye-tracking signal processing unit, a real-time compensation control unit, a parallel logic unit group, and a real-time image registration and automatic rescanning unit. The FPGA processing control module is connected to the optical imaging and tracking module; The storage module is connected to the FPGA processing and control module; in, The frame-level clock management unit is used to generate a unified timing reference for the system and assign frame identifiers with a unified time reference to the reference image frames acquired by the eye movement tracking beam and the data streams acquired by the imaging beam and the tomographic scanning beam, so that the acquired data of different modalities are correlated in time through the frame identifiers. The real-time eye movement signal processing unit is used to receive continuous reference image frames acquired by the eye movement tracking beam, calculate the displacement parameters of the eyeball relative to the reference position in real time through an image feature matching algorithm implemented by hardware logic, and store the displacement parameters and their corresponding frame identifiers in the storage module. The real-time compensation control unit is used to read the displacement parameters from the storage module before the start of the scanning cycle of each imaging frame or scanning block, and convert them into starting coordinate correction instructions for the first set of scanning galvanometers and the second set of scanning galvanometers, so as to realize feedforward eye movement compensation. The parallel logic unit group is used to receive the acquisition data streams from the imaging beam and the tomographic scanning beam in parallel, and to call the corresponding displacement parameters from the storage module based on the frame identifier associated with the data stream. In the hardware pipeline, the acquisition data of the tomographic scanning beam is processed by multi-frame alignment and accumulation based on the displacement parameters, while the acquisition data of the imaging beam is processed by parallel preprocessing. The real-time image registration and automatic rescanning unit is used to register and fuse the data processed by the parallel logic unit group, output the fused image, and realize the automatic rescanning function in the follow-up scanning mode. The data transmission and state switching between the real-time eye-tracking signal processing unit, the real-time compensation control unit, the parallel logic unit group, and the real-time image registration and automatic rescanning unit are driven by the hardware timing signals generated by the frame-level clock management unit, forming a closed-loop processing and control path in full hardware.

[0010] Optionally, the parallel logic unit group's processing of the tomographic scan beam acquisition data includes: performing a fast Fourier transform through hardware logic to reconstruct depth information, and using the displacement parameters, performing spatial position correction and signal accumulation on the reconstruction results of multiple scans from the same anatomical location through a hardware accumulator.

[0011] Optionally, in follow-up scanning mode, the real-time image registration and automatic rescanning unit uses a block average absolute difference algorithm to perform feature matching between the currently acquired image and the baseline image pre-stored in the storage module, and dynamically adjusts the coordinate system of subsequent scans according to the matching results to achieve automatic rescanning.

[0012] Optionally, the real-time compensation control unit calculates the correction amount of the scanning galvanometer driving voltage based on the displacement parameters, and the correction amount is injected into the scanning control signal before the start of the next imaging scanning frame via a digital-to-analog converter.

[0013] Optionally, it also includes a system control and display unit, which communicates with the FPGA processing control module to send scanning protocols and receive and display the fused image.

[0014] This invention also provides a real-time image fusion and eye-motion compensation method for multimodal fundus imaging, applied to the FPGA-based multimodal fundus imaging system described above, the method comprising: The unified timing generated by the frame-level clock management unit based on the FPGA processing control module synchronously starts eye movement tracking scanning, confocal laser scanning fundus imaging, and spectral domain optical coherence tomography, and assigns frame identifiers to the data acquired by each channel. In the real-time eye movement signal processing unit of the FPGA processing control module, hardware-level feature matching is performed on the image sequence acquired in real time by tracking and scanning to calculate the current eye displacement parameters, and the displacement parameters and corresponding frame identifiers are stored. Before the data acquisition of each imaging frame or scan block begins, the real-time compensation control unit of the FPGA processing control module reads the stored displacement parameters, converts them into scan coordinate correction values, and feeds them forward to the galvanometer control system of confocal laser scanning fundus imaging and spectral domain optical coherence tomography. During the imaging data acquisition process, the parallel logic unit group of the FPGA processing control module performs the following in parallel: real-time preprocessing of confocal laser scanning fundus imaging data, and multi-frame spatial alignment and accumulation of spectral domain optical coherence tomography data based on the read displacement parameters. Based on the association of the frame identifier, the processed confocal laser scanning fundus imaging image and the spectral domain optical coherence tomography image are registered and fused in the real-time image registration and automatic rescanning unit of the FPGA processing control module.

[0015] Optionally, the hardware-level feature matching employs a cross-correlation algorithm to calculate the cross-correlation function between the real-time acquired reference image frame and the pre-stored baseline reference image frame to solve for the displacement parameters of the eyeball; the multi-frame spatial alignment and accumulation includes: for multiple A-scan lines at the same anatomical location, after position correction using the displacement parameters, the average is accumulated in real time in a hardware accumulator.

[0016] Optionally, in follow-up scanning mode, the method further includes: using a block average absolute difference algorithm to perform feature matching between the currently acquired image and a pre-stored baseline image, and adjusting the global coordinates of subsequent scans based on the matching results to achieve automatic rescanning.

[0017] Optionally, the steps of real-time eye-tracking signal processing, real-time compensation control, parallel logic processing, and image registration and rescanning are driven by hardware timing signals inside the FPGA and executed in a closed loop within a unified hardware pipeline.

[0018] [Beneficial Effects] First, addressing the issue of not being able to complete eye movement compensation before the start of imaging scanning, this invention utilizes the collaboration of a real-time eye movement signal calculation unit and a real-time compensation control unit to rapidly calculate displacement and feedforward to correct scan coordinates based on hardware logic before the start of each imaging frame or scan block cycle. This fully hardware-based closed-loop feedback and frame-level feedforward compensation mechanism reduces processing latency to the microsecond level, thereby completing compensation at the physical start of scanning and significantly suppressing image artifacts caused by eye movements from the source.

[0019] Secondly, addressing the issue of uncertain synchronization and real-time fusion display of multimodal data streams, this invention assigns a unified time-reference frame identifier to all data through a frame-level clock management unit, achieving forced alignment between cSLO and OCT acquisition times. Combined with the real-time registration and fusion unit in the hardware pipeline, pixel-level spatial registration and image fusion can be completed during acquisition, controlling the lateral registration error of multimodal images to within 30 micrometers, and enabling real-time display of the fusion results. This effectively solves the data misalignment and display lag caused by time asynchrony.

[0020] Finally, addressing the issue of difficulty in accurately reproducing retinal locations during longitudinal follow-up scans, this invention activates an automatic rescan function in follow-up mode. By comparing the current image with a pre-stored baseline image using a hardware feature matching algorithm and dynamically adjusting the global scanning coordinate system, high-precision and highly repeatable automatic localization of the same anatomical location on the retina is achieved, greatly improving the reliability and quantitative assessment efficiency of long-term lesion monitoring.

[0021] Furthermore, this invention utilizes parallel logic unit groups to perform multi-frame alignment and accumulation of tomographic scan data based on displacement parameters, thereby improving the image signal-to-noise ratio in real time during acquisition. Simultaneously, all core processing functions are integrated into a single FPGA chip, achieving a high degree of unity between high system performance, low latency, low power consumption, and low cost.

[0022] In summary, this invention integrates the entire workflow—eye-tracking computation, scanning control, image preprocessing, registration and fusion—on a single FPGA chip, replacing the complex architecture of traditional multi-processor or host computer software processing. This highly integrated hardware design not only simplifies the system structure and reduces power consumption and cost, but also further lowers the implementation threshold by reusing existing imaging hardware, while improving the determinism and reliability of system operation. Furthermore, based on the reconfigurable hardware characteristics of FPGAs, the core processing pipeline of this invention can be flexibly updated, configured, or expanded using hardware description languages. The system easily integrates new imaging modalities, scanning protocols, or advanced processing algorithms, adapting to continuous upgrades in clinical needs and technology, effectively protecting R&D investment and accelerating product iteration. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 The system hardware structure framework diagram provided for this invention; Figure 2A timing diagram illustrating the synchronous acquisition of multimodal data provided by this invention; Figure 3 A block diagram of the logic module for the FPGA internal parallel real-time processing pipeline provided by the present invention. Figure 4 This is a schematic diagram illustrating the working principle of real-time eye movement compensation and scan correction provided by the present invention. Figure 5 A flowchart of the system workflow provided by this invention; Figure 6 This is an imaging schematic diagram provided by the present invention. Detailed Implementation

[0025] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. In the specific implementation of the present invention, the optical components are customized according to the requirements of imaging field of view, resolution, and aberration control, while other components and chips are commercially available general-purpose products, selected based on meeting the performance indicators of each module.

[0026] This embodiment provides a multimodal fundus imaging system based on FPGA. See [link to documentation]. Figure 1 The system mainly includes an optical imaging and tracking module, an FPGA processing and control module, a storage module, and a system control and display unit.

[0027] The optical imaging and tracking module, serving as the system's data acquisition front-end, is designed based on confocal scanning laser fundus imaging (cSLO) technology and spectral domain optical coherence tomography (SD-OCT) technology, and deeply integrates active eye-tracking functionality based on hardware multiplexing. This module aims to achieve multimodal, synchronized data acquisition of fundus structures.

[0028] Regarding the light source system, this module integrates multiple light sources. For confocal laser scanning of the fundus, it is equipped with multi-wavelength laser diodes to support different imaging modes, providing 486nm blue light, 518nm green light, and 786nm and 815nm near-infrared light, respectively for blue light reflection, infrared reflection, or angiography imaging. The reference beam used for eye movement tracking directly reuses the near-infrared light source in the cSLO optical path. For spectral-domain optical coherence tomography, an independent broadband light source with a center wavelength of approximately 870 nm is used, specifically for acquiring depth tomographic information.

[0029] In terms of the scanning and optical path system, its core architecture lies in the coordination of two sets of scanning mirrors and the beam combining of the back-end optical path. The reference tracking beam and the cSLO imaging beam are multiplexed using the first set of scanning mirrors for synchronous scanning modulation. The SD-OCT imaging beam is modulated using a separate second set of scanning mirrors. Subsequently, the beams emitted from the two sets of mirrors are spatially combined by a dichroic mirror at a position behind the scanning mirrors and in front of the imaging lens, achieving common path propagation with consistent visual axes, and finally focusing together on the same area of ​​the retina through the objective lens. The dichroic mirror is selective for beams of different wavelengths: it has high transmittance (>90%) for the SD-OCT beam with a center wavelength of approximately 870nm, while it has high reflectivity (>90%) for the cSLO imaging beam and eye-tracking beam with wavelengths in the range of 780nm to 820nm. This design ensures strict consistency of the spatial coordinates of all imaging and tracking beams while achieving initial spectral separation.

[0030] In terms of the detection system, the cSLO imaging signal and the eye-tracking light signal, after returning via the same common optical path, are first separated by a dichroic mirror. Specifically, the 486nm blue light reflection signal, the 518nm green light reflection signal, and the 815nm near-infrared light signal used for tracking are directed to different high-speed photodetectors for independent photoelectric conversion and acquisition. The SD-OCT optical path is equipped with an independent high-speed spectrometer, the core of which is a CMOS linear array sensor used to acquire interference spectral signals. Data output from all detection channels is transmitted in real time to the FPGA processing and control module, providing raw input for subsequent synchronous processing and control.

[0031] The FPGA processing and control module, serving as the real-time processing and control hub of this invention, is implemented using a high-performance FPGA chip such as the Xilinx Kintex-7 series. This module integrates a high-precision clock source and clock management circuitry, forming a unified frame-level clock management unit serving the entire system, providing a precise time reference for all data acquisition, processing flows, and synchronization triggering. This module connects to the various scanning galvanometer controllers and detectors in the optical imaging and tracking module via a high-speed digital interface for sending and receiving control commands and acquiring data; simultaneously, it communicates with the system control and display unit via high-speed buses such as Thunderbolt-to-PCIe interfaces to achieve the interaction of commands and image data.

[0032] The core of this module lies in its highly parallelized hardware processing pipeline. This pipeline uses scan frames or scan blocks as the basic processing granularity, coordinating multiple dedicated hardware logic units to work together, forming a complete hardware processing chain from reference tracking and displacement calculation to multimodal image processing and fusion, thereby ensuring the system's high real-time performance, low latency, and multimodal synchronous processing capabilities.

[0033] Specifically, the production line mainly includes the following functional units, see Figure 2 and Figure 3 : (1) Real-time eye movement signal processing unit: This unit directly processes the continuous reference image frames returned by the reference tracking beam that is fully multiplexed with the cSLO imaging optical path. Through a hardware-implemented feature matching and correlation calculation module, it compares the real-time acquired reference frames with the reference frames pre-stored in the storage module, and calculates the translational displacement parameters of the eyeball in the two-dimensional plane in real time. and And can be extended to obtain rotational components. In one specific embodiment, feature matching is achieved by calculating the cross-correlation function between a real-time acquired reference image frame and a pre-stored baseline image frame. By searching for the peak position of this cross-correlation function, the translation of the eyeball in the two-dimensional plane can be determined. , By comparing the cross-correlation values ​​calculated at different preset rotation angles, the calculation of the rotation component of the eyeball can be further expanded. This process is continuously executed within the FPGA in a frame-level pipeline manner, providing real-time and reliable input for subsequent scan coordinate updates.

[0034] (2) Real-time compensation control unit: This unit receives the displacement parameters output by the real-time eye movement signal calculation unit, and before the start of the data acquisition cycle of each imaging scan frame or each scan block, calculates the correction amount for the starting coordinates and scanning trajectory of the first group of cSLO scanning mirrors and the second group of SD-OCT scanning mirrors based on the current eye displacement, and generates corresponding control commands. This unit calculates the correction amount based on the calculated eye displacement parameters ( , , Based on the preset mapping relationship between the scanning coordinate system and the tracking coordinate system, the voltage correction amount required to drive the scanning galvanometer is calculated. The specific calculation is as follows:

[0035]

[0036] in, , This is the original mirror driving voltage; , The mapping coefficient from pixel to voltage; To track the installation offset angle between the coordinate system and the physical coordinate system of the galvanometer.

[0037] This correction amount will be injected into the scan control signal via a digital-to-analog converter (DAC) before the start of the next imaging frame or scan block, thereby achieving feedforward adjustment of the scan start point and trajectory. This frame-level or scan block-level control enables adaptive feedforward compensation for low-frequency motion components such as eye drift and slow saccades, correcting the scan path from the source.

[0038] (3) Parallel Logic Unit Group: This unit is responsible for parallel real-time preprocessing and enhancement of the cSLO and SD-OCT imaging data streams. For the cSLO image data stream, this unit performs preprocessing operations such as flat field correction and contrast enhancement. For the SD-OCT signal data stream, the processing flow includes: first, performing spectral preprocessing and fast Fourier transform on the original interferometric spectral data to reconstruct the A-scan signal representing depth information; then, using the displacement parameters obtained from eye tracking, aligning the spatial positions of multiple A-scan signals from the same anatomical location, and only signals with displacement amounts less than a preset tolerance range are considered validly aligned (in this embodiment, the preset tolerance range is set to ±64 pixels); finally, completing the multi-frame accumulation (8 frames in this embodiment) and averaging of the aligned signals in the hardware accumulator, the averaging can be done using a weighted average method, linearly weighted according to the relevant peak confidence of each frame signal. This process achieves hardware-level noise suppression and signal-to-noise ratio improvement, and all processing is completed in real time in the data acquisition pipeline.

[0039] (4) Real-time image registration and automatic rescanning unit: This unit is responsible for the final data synthesis and advanced application functions. Firstly, it realizes real-time registration and fusion: the preprocessed cSLO image stream and the en-face image stream reconstructed by the B-scan of SD-OCT are spatially registered at the frame level, and a multimodal fused image is generated through a pixel-level fusion algorithm. This fusion algorithm is based on the mapping relationship between the cSLO image coordinates and the OCT scan coordinates, and spatially fuses the intensity information of cSLO and the depth information of OCT. The specific fusion process is as follows: Map cSLO pixel coordinates to OCT B-scan horizontal coordinates: ; ②Add the cSLO intensity to the OCT: .

[0040] in, The horizontal pixel coordinates in the SD-OCT B-scan corresponding to the cSLO pixel; cSLO represents the horizontal pixel coordinates of a pixel in the image along the scanning direction. These are pixel-scale mapping coefficients used to compensate for the differences between cSLO and SD-OCT in terms of scanning field of view, resolution, and galvanometer scanning amplitude. This is the coordinate offset, used to compensate for the fixed system offset error caused by the optical path installation, scanning start angle, etc., between the two imaging modes; The pixel intensity values ​​at the horizontal coordinate x and depth coordinate z of the fused multimodal image; Raw intensity values ​​of SD-OCT imaging at the horizontal x and depth z. The cSLO pixel intensity value corresponding to the horizontal coordinate of this OCT; These are OCT image weighting coefficients, used to control the proportion of OCT structural information in the fusion result; cSLO image weight coefficients are used to control the proportion of cSLO reflection intensity information in the fusion result, and satisfy the following conditions: .

[0041] Secondly, it supports automatic rescanning: This unit is activated during follow-up scans of the same patient. It performs hardware feature matching again with the currently acquired reference tracking frame or cSLO image and the baseline reference image from previous examinations stored in the storage module, calculating the overall coordinate transformation relationship (translation and rotation) between the two examinations. This embodiment preferably uses the block mean absolute difference (SAD) matching algorithm for feature matching, and the criterion for successful matching is that the matching score SAD value must be lower than 80% of a set threshold. If three consecutive matching attempts fail, the host computer UI will issue a "Please maintain eye position" prompt. Based on this relationship, the unit dynamically adjusts the global coordinate system of subsequent scan frames, enabling the current scan to automatically and accurately locate the retinal anatomical region that is exactly the same as in the historical scans, thereby ensuring high repeatability and consistency of longitudinal follow-up imaging.

[0042] To ensure the spatiotemporal consistency of multimodal data for accurate processing, the FPGA processing control module assigns a frame identifier based on the aforementioned unified clock to each set of acquired data (including eye movement parameters, cSLO image data blocks, and SD-OCT raw spectral data). Subsequent processing units use this frame identifier to associate, synchronize, and match data belonging to the same time or anatomical location, thereby ensuring the correct alignment and fusion of displacement parameters and imaging data in the hardware pipeline.

[0043] Through the coordinated operation of the aforementioned modular pipeline, the FPGA processing control module realizes the entire process of eye-tracking signal processing, scanning feedforward compensation, real-time image preprocessing, multi-frame alignment and noise reduction, high-precision registration and fusion, and automatic rescan positioning at the hardware level, without relying on host computer software for post-processing, thus ensuring the determinism and extremely low latency of the system.

[0044] The storage module is directly connected to the FPGA processing and control module, serving as the system's data storage and caching hub. This module stores, caches, and manages various types of data generated during system operation, providing data support for real-time processing and long-term monitoring.

[0045] Specifically, storage modules typically consist of two levels: (1) Cache: Used for real-time temporary data storage. Its main functions include: temporarily storing eye movement parameters calculated in real time by the FPGA processing and control module. , and optional The cache caches raw data from the optical imaging and tracking modules, such as raw interferometric spectral data from SD-OCT and raw image data from cSLO; and temporarily stores intermediate results during the FPGA's internal pipeline processing. The low latency of the cache ensures fast read and write access to displacement parameters by the FPGA, as well as real-time data flow during pipeline processing.

[0046] (2) Non-volatile memory: Used to store important data and information stably for a long time. Its stored contents mainly include: multimodal fusion images finally generated and output by the FPGA processing control module; tracking logs and detailed displacement data generated during eye tracking; scanning parameter configuration information used by the system; and patient baseline reference images pre-stored in follow-up scanning mode to support automatic rescanning function. Non-volatile memory ensures the long-term traceability of patient data, historical images and system settings.

[0047] Through this hierarchical storage architecture, the storage module efficiently supports the core operation of the FPGA processing and control module: on the one hand, the high-speed cache meets the real-time processing requirements for fast data access, such as allowing the FPGA to read displacement parameters in real time for feedforward compensation, or temporarily storing intermediate data for use by downstream units in the pipeline; on the other hand, the non-volatile memory enables data persistence, not only saving the final diagnostic images and operation records, but also providing the necessary baseline image data foundation for feature matching and coordinate adjustment in the follow-up scanning mode described in the claims. The design of the entire storage module ensures the system's comprehensive data needs from real-time control to long-term data management.

[0048] The system control and display unit, which serves as the human-computer interaction and advanced management hub of this invention, is typically implemented by an industrial control computer that runs operating systems such as Windows and is equipped with a graphical user interface.

[0049] The core function of this unit is to communicate and collaborate with the FPGA processing and control module to complete system control, image display, and data management tasks. Specifically, its functions mainly include: (1) User control and command issuance: A user-friendly graphical user interface is provided for operators (such as doctors) to enter patient information, select scanning protocols, and set parameters. Commands issued by the operator, especially control commands with high real-time requirements, such as starting scanning, setting scanning timing, and updating frame-level compensation parameters, are sent to the FPGA processing control module through a predefined application programming interface (API), thereby driving the hardware system to execute precise scanning and processing procedures.

[0050] (2) Image reception and real-time display: The unit receives the final multimodal fundus image stream after hardware pipeline processing and fusion in the FPGA processing control module in real time. This unit is responsible for rendering and displaying these image data on the graphical interface in real time, so that the operator can observe the imaging effect and diagnostic information immediately.

[0051] (3) Patient Data and System Management: Responsible for managing the patient database, including storing and retrieving basic patient information, historical examination records, corresponding baseline reference images, and fused image results of the current examination. In addition, this unit can also perform some non-real-time large-scale data analysis or post-processing tasks, but its core positioning is to ensure the stability and efficiency of real-time control and display, while offloading computationally intensive real-time processing tasks to FPGA hardware.

[0052] Through this division of labor, the system control and display unit and the FPGA processing and control module form a collaborative two-tier architecture: the host computer unit focuses on flexible human-computer interaction, protocol management, image display, and long-term data management; while the FPGA module focuses on deterministic, low-latency real-time signal acquisition, eye-tracking compensation, high-speed image processing, and fusion. This architecture leverages the real-time processing advantages of the FPGA while retaining the user-friendliness and data management convenience of a general-purpose computer.

[0053] This embodiment combines Figure 4 , Figure 5 and Figure 6 The invention further demonstrates how the system operates and its value through a complete workflow and a specific clinical example.

[0054] See Figure 5 The workflow of this system embodies the real-time image fusion and eye-tracking compensation method of this invention, including the following steps: S1: Based on the unified timing generated by the frame-level clock management unit built into the FPGA processing control module, eye movement tracking scanning, cSLO imaging scanning and SD-OCT imaging scanning are started synchronously, and frame identifiers with a unified time reference are assigned to the data acquired by each channel.

[0055] S2: In the real-time eye-tracking signal processing unit of the FPGA processing and control module, hardware-level feature matching is performed on the image sequence acquired in real time by the tracking scan to calculate the displacement parameters of the current eyeball relative to the reference position. , The displacement parameters and their corresponding frame identifiers are stored in the storage module.

[0056] S3: Before the start of the scanning cycle of each imaging frame or scanning block, the real-time compensation control unit of the FPGA processing control module reads the stored displacement parameters from the storage module, converts them into starting coordinate correction instructions for the first and second sets of scanning galvanometers, and feeds them forward to the galvanometer control system of cSLO and SD-OCT to realize feedforward eye-tracking compensation, and then starts the imaging data acquisition of that frame or scanning block.

[0057] S4: During the imaging data acquisition process, the parallel logic unit group of the FPGA processing control module performs the following in parallel: real-time preprocessing of cSLO imaging data, and multi-frame spatial alignment and accumulation of SD-OCT data based on displacement parameters read from the storage module.

[0058] S5: Based on the association of frame identifiers, in the real-time image registration and automatic rescanning unit of the FPGA processing control module, the processed cSLO image and the SD-OCT en-face image are registered and fused to generate a multimodal fused image.

[0059] S6: Output the fused image and send it to the system control and display unit for real-time display, and repeat steps S2 to S6 until all scans are completed.

[0060] See Figure 4 This illustrates the working principle of real-time eye-tracking compensation and scan correction: the system acquires the current position of the eyeball through real-time eye tracking, and the FPGA calculates its displacement relative to a pre-stored reference position; at the start of the next imaging scan frame, the starting coordinates of the imaging scan beam are pre-adjusted accordingly, ensuring that the scan area still aligns with and covers the retinal target area corresponding to the reference position, thereby physically compensating for the offset caused by eye movement. To illustrate its clinical application, a follow-up of a patient with diabetic retinopathy is considered. During the initial examination, the system established a high-quality baseline reference image of the macula. Several months later, during a follow-up examination in follow-up mode: (1) The system detects slight changes in eye position through real-time tracking and feature matching.

[0061] (2) Before each SD-OCT B scan, the FPGA fine-tunes the starting point of the galvanometer to ensure that each line falls in the same position as the baseline.

[0062] (3) Simultaneously perform alignment and accumulation of multiple A-scans to obtain a clear OCT image.

[0063] (4) The cSLO image and the OCT en-face image are precisely fused within the FPGA.

[0064] (5) such as Figure 6 As shown, the display shows a fused image that closely overlaps with the baseline, clearly demonstrating the retinal vascular network and deep structures.

[0065] By comparing images from two examinations, doctors can accurately assess changes in microaneurysms, bleeding points, or retinal thickness, providing reliable and intuitive imaging evidence for disease monitoring and treatment effectiveness evaluation.

[0066] In practical implementation, through the above-described hardware architecture, processing pipeline, and workflow, the system of this invention can achieve the following key performance indicators: (1) Synchronization and tracking performance: The frame-level synchronization delay between the cSLO and SD-OCT modules does not exceed 50 microseconds. The FPGA-based reference tracking displacement calculation latency is no more than 1 millisecond (ms), and the eye tracking accuracy can reach 50 micrometers. Within 50 Hz; the axial resolution of the SD-OCT module is not less than 50 Hz. The cSLO module's imaging frame rate is no less than 100 frames per second (fps).

[0067] (2) Image processing and fusion performance: Using a frame-level image registration algorithm accelerated by FPGA hardware, the registration error between cSLO images and SD-OCT tomographic data is no greater than 30%. Through multi-frame alignment and cumulative noise reduction processing based on displacement parameters, motion artifacts can be effectively suppressed, keeping the image artifact rate below 5%, while improving image contrast by no less than 15%. In follow-up scanning mode, the system can achieve high repeatability imaging, with a spatial repeatability positioning error of no more than 20. .

[0068] In summary, this invention systematically solves the core problems of multimodal fundus imaging in terms of synchronization, real-time performance, spatial consistency, and follow-up repeatability through a fully hardware-based processing and frame-level compensation control architecture. It provides a high-performance, highly integrated, and feasible advanced imaging solution with significant clinical value and market potential.

[0069] Those skilled in the art will understand that the hardware configuration, dataset, hyperparameters, etc. in the above embodiments can be adjusted according to the actual application scenario, and these adjustments should all be included within the protection scope defined by the claims of this invention.

[0070] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multimodal fundus imaging system based on FPGA, characterized in that, include: The optical imaging and tracking module includes: a confocal laser scanning fundus imaging unit for generating and utilizing a first set of scanning galvanometers to scan a common-path propagating imaging beam and an eye movement tracking beam; and a spectral-domain optical coherence tomography unit for generating and utilizing a second set of scanning galvanometers to scan an independent tomographic scanning beam. The FPGA processing control module has a built-in unified frame-level clock management unit and integrates a parallel processing pipeline built from hardware logic resources. The pipeline includes: a real-time eye-tracking signal processing unit, a real-time compensation control unit, a parallel logic unit group, and a real-time image registration and automatic rescanning unit. The FPGA processing control module is connected to the optical imaging and tracking module. The storage module is connected to the FPGA processing and control module; in, The frame-level clock management unit is used to generate a unified timing reference for the system and assign frame identifiers with a unified time reference to the reference image frames acquired by the eye movement tracking beam and the data streams acquired by the imaging beam and the tomographic scanning beam, so that the acquired data of different modalities are correlated in time through the frame identifiers. The real-time eye movement signal processing unit is used to receive continuous reference image frames acquired by the eye movement tracking beam, calculate the displacement parameters of the eyeball relative to the reference position in real time through an image feature matching algorithm implemented by hardware logic, and store the displacement parameters and their corresponding frame identifiers in the storage module. The real-time compensation control unit is used to read the displacement parameters from the storage module before the start of the scanning cycle of each imaging frame or scanning block, and convert them into starting coordinate correction instructions for the first set of scanning galvanometers and the second set of scanning galvanometers, so as to realize feedforward eye movement compensation. The parallel logic unit group is used to receive the acquisition data streams from the imaging beam and the tomographic scanning beam in parallel, and to call the corresponding displacement parameters from the storage module based on the frame identifier associated with the data stream. In the hardware pipeline, the acquisition data of the tomographic scanning beam is processed by multi-frame alignment and accumulation based on the displacement parameters, while the acquisition data of the imaging beam is processed by parallel preprocessing. The real-time image registration and automatic rescanning unit is used to register and fuse the data processed by the parallel logic unit group, output the fused image, and realize the automatic rescanning function in the follow-up scanning mode. The data transmission and state switching between the real-time eye-tracking signal processing unit, the real-time compensation control unit, the parallel logic unit group, and the real-time image registration and automatic rescanning unit are driven by the hardware timing signals generated by the frame-level clock management unit, forming a closed-loop processing and control path in full hardware.

2. The FPGA-based multimodal fundus imaging system according to claim 1, characterized in that, The parallel logic unit group processes the data acquired by the tomographic scanning beam by: performing a fast Fourier transform through hardware logic to reconstruct depth information, and using the displacement parameters, performing spatial position correction and signal accumulation on the reconstruction results of multiple scans from the same anatomical location through a hardware accumulator.

3. The FPGA-based multimodal fundus imaging system according to claim 1, characterized in that, In follow-up scanning mode, the real-time image registration and automatic rescanning unit uses a block average absolute difference algorithm to perform feature matching between the currently acquired image and the baseline image pre-stored in the storage module, and dynamically adjusts the coordinate system of subsequent scans according to the matching results to achieve automatic rescanning.

4. The FPGA-based multimodal fundus imaging system according to claim 1, characterized in that, The real-time compensation control unit calculates the correction amount of the scanning mirror drive voltage based on the displacement parameters, and the correction amount is injected into the scanning control signal before the start of the next imaging scanning frame through a digital-to-analog converter.

5. The FPGA-based multimodal fundus imaging system according to claim 1, characterized in that, It also includes a system control and display unit, which communicates with the FPGA processing control module to send scanning protocols and receive and display the fused image.

6. A method for real-time image fusion and eye-motion compensation in multimodal fundus imaging, characterized in that, The method, applied to the FPGA-based multimodal fundus imaging system as described in any one of claims 1 to 5, comprises: The unified timing generated by the frame-level clock management unit based on the FPGA processing control module synchronously starts eye movement tracking scanning, confocal laser scanning fundus imaging, and spectral domain optical coherence tomography, and assigns frame identifiers to the data acquired by each channel. In the real-time eye movement signal processing unit of the FPGA processing control module, hardware-level feature matching is performed on the image sequence acquired in real time by tracking and scanning to calculate the current eye displacement parameters, and the displacement parameters and corresponding frame identifiers are stored. Before the data acquisition of each imaging frame or scan block begins, the real-time compensation control unit of the FPGA processing control module reads the stored displacement parameters, converts them into scan coordinate correction values, and feeds them forward to the galvanometer control system of confocal laser scanning fundus imaging and spectral domain optical coherence tomography. During the imaging data acquisition process, the parallel logic unit group of the FPGA processing control module performs the following in parallel: real-time preprocessing of confocal laser scanning fundus imaging data, and multi-frame spatial alignment and accumulation of spectral domain optical coherence tomography data based on the read displacement parameters. Based on the association of the frame identifier, the processed confocal laser scanning fundus imaging image and the spectral domain optical coherence tomography image are registered and fused in the real-time image registration and automatic rescanning unit of the FPGA processing control module.

7. The method for real-time image fusion and eye movement compensation in multimodal fundus imaging according to claim 6, characterized in that, The hardware-level feature matching uses a cross-correlation algorithm to calculate the cross-correlation function between the real-time acquired reference image frame and the pre-stored baseline reference image frame in order to solve the displacement parameters of the eyeball. The multi-frame spatial alignment and accumulation includes: for multiple A-scan lines at the same anatomical location, after position correction using the displacement parameters, averaging is accumulated in real time in a hardware accumulator.

8. The method for real-time image fusion and eye movement compensation in multimodal fundus imaging according to claim 6, characterized in that, In follow-up scanning mode, the method further includes: using a block average absolute difference algorithm to perform feature matching between the currently acquired image and a pre-stored baseline image, and adjusting the global coordinates of subsequent scans based on the matching results to achieve automatic rescanning.

9. The method for real-time image fusion and eye movement compensation in multimodal fundus imaging according to claim 6, characterized in that, The steps of real-time eye-tracking signal processing, real-time compensation control, parallel logic processing, and image registration and rescanning are driven by hardware timing signals inside the FPGA and executed in a closed loop within a unified hardware pipeline.