Distributed accelerated computing architecture and method based on face confocal surface topography recovery

The distributed accelerated computing architecture for surface topography restoration utilizes multiple computing engines with row-region division of labor and independent DDR memory to achieve sliced ​​processing and parallel computing, solving the problem of insufficient computing throughput in existing technologies and realizing efficient image reconstruction and fault recovery capabilities.

CN120598762BActive Publication Date: 2026-05-15HUBEI CUGUANG 3D SENSING TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUBEI CUGUANG 3D SENSING TECH CO LTD
Filing Date
2025-05-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing heterogeneous computing architectures are difficult to scale linearly at high resolutions or multi-layer depths, and cannot independently optimize and partition different processing stages, resulting in insufficient computing throughput.

Method used

A distributed accelerated computing architecture based on surface confocal topography restoration is adopted. Multiple computing engines are divided into row regions to achieve slice processing and parallel computing. Each engine is equipped with independent DDR memory to perform row region segmentation processing of image sequences and forwards the computing results through the GTX interface. Finally, the first computing engine summarizes and transmits the results to the host computer.

Benefits of technology

It significantly improves computing throughput, shortens overall reconstruction time, reduces bus pressure, supports continuous online processing of high frame rate and high resolution image sequences, and can quickly recover in the event of local failures, achieving high bandwidth and low latency computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598762B_ABST
    Figure CN120598762B_ABST
Patent Text Reader

Abstract

The application provides a distributed acceleration computing architecture and method based on a surface confocal surface topography recovery, relates to the technical field of surface confocal surface hardware design, and comprises a plurality of serially connected computing engines, wherein the first computing engine in the computing engines receives image sequence data through a CXP interface, forwards the image sequence data to the remaining computing engines through a GTX high-speed interface, and each computing engine is configured with a DDR memory; the computing engines perform row region segmentation processing on the image sequence in a preset division mode, and each computing engine is responsible for processing a specific row region of the image; after the computing engines complete optical tomographic demodulation calculation and peak positioning calculation required for surface confocal surface topography recovery, the calculation results of the other computing engines in the computing engines are forwarded to the first computing engine through the GTX interface, and the first computing engine transmits the summary result to an upper computer through a fiber interface. The application helps to improve the computing throughput of the acceleration computing architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hardware design technology for confocal surfaces, and in particular to a distributed accelerated computing architecture and method based on confocal surface topography recovery. Background Technology

[0002] With the profound development of science and technology, various industries have achieved unprecedented progress, and advancements in data processing technology have been a significant contributor to this development. Fields such as geology, atmospheric science, medical imaging, and virtual reality each have their own characteristics, and their data can be two-dimensional, three-dimensional, scalar, or vector-based. What they have in common is the sheer volume of data, not only in the large number of data nodes but also in the wide range of dimensions and complexities presented. Extracting key data from this massive dataset and presenting it clearly graphically places high demands on visualization technology.

[0003] Chinese patent CN110765064B discloses an edge image processing system and method with a heterogeneous computing architecture. The system includes an FPGA-implemented network interface card (NIC), an FPGA-implemented system control unit, a GPU, a PCIe switch, and system storage. The NIC, system control unit, and GPU are interconnected via the PCIe switch. A NIC array is configured, with the number of NICs depending on the network packet processing capability and GPU computing power. The system storage is used to store information within the system. However, the NIC array size in this solution is solely determined by the network packet processing capability and GPU computing power. The system storage only serves as an information cache and cannot perform independent optimization partitioning for different processing stages, making linear scaling at higher resolutions or deeper layers difficult. Therefore, providing a distributed accelerated computing architecture and method based on planar confocal surface topography restoration is essential to improve the computational throughput of the accelerated computing architecture. Summary of the Invention

[0004] In view of this, this invention proposes a distributed accelerated computing architecture and method based on surface confocal topography restoration. By dividing the work among multiple computing engines by row region, slice processing and parallel computing are achieved, significantly reducing the overall reconstruction time. Each engine runs optical tomography demodulation and peak localization simultaneously, and the overall throughput increases almost linearly with the number of engines, thereby improving the computational throughput of the accelerated computing architecture.

[0005] This invention provides a distributed accelerated computing architecture based on surface confocal topography recovery, comprising multiple serially connected computing engines, wherein...

[0006] The first computing engine in the computing engine receives image sequence data through the CXP interface and forwards the image sequence data to the remaining computing engines through the GTX high-speed interface. Each computing engine is equipped with DDR memory to store the image data, maximum value map and intermediate results allocated to the corresponding computing engine. The computing engines perform row region segmentation processing on the image sequence according to a preset division of labor, and each computing engine is responsible for processing a specific row region of the image.

[0007] After the computing engine completes the optical tomography demodulation calculation and peak positioning calculation required for surface confocal topography restoration, the calculation results of other computing engines in the computing engine are forwarded to the first computing engine through the GTX interface, and the first computing engine transmits the summarized results to the host computer through the fiber optic interface.

[0008] Based on the above technical solutions, preferably, the plurality of computing engines includes a first computing engine, a second computing engine, a third computing engine, and a fourth computing engine connected in series, and the first computing engine, the second computing engine, the third computing engine, and the fourth computing engine have the same chip and memory parameters.

[0009] Based on the above technical solutions, preferably, when the first computing engine receives image sequence data, it saves the initial row to the first preset row of each frame of the image sequence data into the DDR memory corresponding to the first computing engine, and calculates the first maximum value map corresponding to the initial row to the first preset row of the current image, and saves the first maximum value map into the DDR memory corresponding to the first computing engine.

[0010] When receiving image sequence data, the second computing engine saves the first preset row to the second preset row of each frame of the image sequence data into the DDR memory corresponding to the second computing engine, and calculates the second maximum value map corresponding to the first preset row to the second preset row of the current image, and saves the second maximum value map into the DDR memory corresponding to the second computing engine.

[0011] When receiving image sequence data, the third computing engine saves the second preset row to the third preset row of each frame of the image sequence data into the DDR memory corresponding to the third computing engine, and calculates the third maximum value map corresponding to the second preset row to the third preset row of the current image, and saves the third maximum value map into the DDR memory corresponding to the third computing engine.

[0012] When receiving image sequence data, the fourth computing engine saves the third preset row to the fourth preset row of each frame of the image sequence data into the DDR memory corresponding to the fourth computing engine, calculates the fourth maximum value map corresponding to the third preset row to the fourth preset row of the current image, and saves the fourth maximum value map into the DDR memory corresponding to the fourth computing engine.

[0013] More preferably, the DDR memory of each computing engine is divided into a first region, a second region, and a third region. The first region is used to store the original image segments, the second region is used to store the corresponding maximum value image, and the third region is used to store intermediate results.

[0014] More preferably, the GTX interface between the computing engines includes a pair of transceivers for forwarding image sequence data and reverse transmission of centroid results.

[0015] More preferably, the process of generating the image sequence data includes:

[0016] The sample is illuminated with spatially modulated structured light from an incoherent structured illumination source via a beam splitter and a microscope objective;

[0017] And by driving the objective lens or sample stage, it is positioned along the optical axis at predetermined steps Δz to z1…z1…z2. n The detectors at each location acquire corresponding structured raw images to form a structured image stack.

[0018] An optical tomography demodulation algorithm is executed on the structured image stack to obtain optical tomographic image stacks for each axial plane.

[0019] Based on the tomographic image stack, an axial response curve is generated for each pixel along the optical axis. A peak localization algorithm is applied to the response curve to determine the optimal depth-of-focus peak position for each pixel, so as to output image sequence data.

[0020] More preferably, the first computing engine, the second computing engine, the third computing engine, and the fourth computing engine are all FPGA chips, and the latency of the first computing engine, the second computing engine, the third computing engine, and the fourth computing engine in receiving the image sequence data is less than 1ms.

[0021] Furthermore, the processing flow of each computing engine includes:

[0022] Calculate and store the maximum value map of the corresponding allocation region based on the original image of the engine storage allocation region;

[0023] The centroids of the original image and the maximum value image stored in the allocated region are extracted using an improved centroid fitting function, and the centroid extraction results are forwarded to the first calculation engine.

[0024] More preferably, the expression for the improved centroid fitting function is:

[0025]

[0026] Where μ represents the peak position obtained from the fitting, argmin represents the optimization operator, which is used to return the independent variable that minimizes the objective function, a represents the Gaussian function amplitude, e represents the natural constant, and z i δ represents the coordinate of the i-th axial position. r Let b represent the standard deviation of the Gaussian function, and g represent the center position of the Gaussian function. norm (z i () represents the intensity value measured at the coordinates of the i-th axial position on the axial response curve, z i ∈FWHM indicates that data points with a peak intensity greater than half that of the axial response curve are selected as valid data points for fitting.

[0027] A second aspect of this application provides a distributed accelerated computing method based on surface confocal topography recovery, comprising the following steps:

[0028] Image sequence data is transmitted from the area confocal sensor to the first computing engine via the CXP interface;

[0029] The first computing engine stores the preset rows of each frame of the image sequence data into the DDR memory corresponding to the first computing engine, calculates the maximum pixel value corresponding to the preset row, and forwards the remaining rows of each frame of the image to the remaining computing engine via the GTX interface, so that the remaining computing engine stores the corresponding row segments and the maximum value image respectively.

[0030] After all computing engines generate the corresponding maximum value map, each computing engine reads the original image and the corresponding maximum value map from the corresponding DDR memory and intermediate cache, and uses the improved centroid fitting function to extract the centroid.

[0031] The centroid extraction results are reversed by each computing engine and aggregated to the first computing engine via the GTX interface. The first computing engine then sends the centroid extraction results to the host computer via the fiber optic interface.

[0032] The distributed accelerated computing architecture and method based on surface confocal topography restoration provided by this invention have the following advantages over existing technologies:

[0033] (1) By dividing the work by row area through multiple computing engines, the processing and parallel computing are realized, which greatly shortens the overall reconstruction time. Each engine runs optical tomography demodulation and peak positioning at the same time. The overall throughput increases almost linearly with the number of engines. Each engine is equipped with independent DDR memory to store image segments, maximum value maps and intermediate results locally, reducing the data transfer overhead across modules. The intermediate results are only aggregated through the GTX interface when necessary, which significantly reduces the bus pressure. The first engine uses CXP to receive high-speed camera data, uses GTX to forward it with low latency in the link, and uses the fiber optic interface to summarize the data to the host computer in real time, ensuring high bandwidth and low latency at the end. It supports continuous online processing of high frame rate and high resolution image sequences. At the same time, local failure of each engine will not cause the entire link to collapse, but only affects the corresponding row area. It can be quickly recovered through hot standby or fault switching, so that resources can be more finely load balanced and utilized, and the computing throughput of the accelerated computing architecture is improved.

[0034] (2) By significantly suppressing scattering and background through structured illumination, the contrast of the optical cross section brought about by the modulation frequency is improved. The fine stepping of Δz in the optical axis direction ensures the sampling density and forms a high-quality structured image stack. Optical tomography demodulation is performed on the image stack, which can effectively remove out-of-focus signals and restore the true optical tomographic map of each axial plane. The continuous z-layer is sliced ​​so that the resolution of each depth layer is close to the limit of the system point spread function. Based on the response curve of each pixel along the z direction, a high-precision peak interpolation / fitting algorithm is used to determine the position of the focal depth peak. At the same time, the end-to-end process from the structured original image to the tomographic map and then to the focal depth map can directly produce the 3D surface height map of each frame. Without additional deconvolution or global optimization, high signal-to-noise ratio and high depth resolution surface morphology data can be obtained. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 A schematic diagram of a distributed accelerated computing architecture based on surface confocal topography recovery provided by the present invention;

[0037] Figure 2 A schematic diagram illustrating the surface morphology restoration process for a confocal surface provided by the present invention;

[0038] Figure 3 A schematic diagram of the storage process of DDR memory provided by the present invention. Detailed Implementation

[0039] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0040] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.

[0041] refer to Figure 1 This invention provides a distributed accelerated computing architecture based on surface confocal topography recovery, comprising multiple serially connected computing engines, wherein...

[0042] The first computing engine in the computing engine receives image sequence data through the CXP interface and forwards the image sequence data to the remaining computing engines through the GTX high-speed interface. Each computing engine is equipped with DDR memory to store the image data, maximum value map and intermediate results allocated to the corresponding computing engine. The computing engines perform row region segmentation processing on the image sequence according to a preset division of labor, and each computing engine is responsible for processing a specific row region of the image.

[0043] After the computing engine completes the optical tomography demodulation calculation and peak positioning calculation required for surface confocal topography restoration, the calculation results of other computing engines in the computing engine are forwarded to the first computing engine through the GTX interface, and the first computing engine transmits the summarized results to the host computer through the fiber optic interface.

[0044] Each computing engine's DDR memory is divided into three regions: Region 1 stores the original image segments, Region 2 stores the corresponding maximum value images, and Region 3 stores intermediate results. The GTX interface between the computing engines includes a pair of transceivers, which are used for forwarding image sequence data and reverse transmission of centroid results.

[0045] The processing flow of each computing engine includes: calculating and storing the maximum value map of the corresponding allocation region based on the original image of the allocation region stored in the engine; performing centroid extraction calculation on the original image and maximum value map of the allocation region stored based on the improved centroid fitting function, and forwarding the centroid extraction result to the first computing engine.

[0046] In one example, the first computing engine receives the entire frame image sequence through the CXP interface or the upper-level network, while the other computing engines receive the row segment pixel data forwarded by the previous engine through the GTX interface. Each engine writes the pixels into its own DDR4 raw image area according to the preset row segment index.

[0047] While writing the original image, the FPGA's internal hardware logic / pipeline unit performs a "maximum value" operation on each scan line (or each pixel along the time sequence). The calculated maximum intensity value is written to the maximum value map area in DDR4, with each pixel occupying a corresponding 29 bits or a predetermined width. After storage, the original image area and the maximum value map area are prepared for the next stage of reading via the DMA controller (FDMA256-bit). The controller reads the original image segment (I(x,y)) and the corresponding pixel's maximum value map M(x,y) from DDR4, and both are entered into the centroid calculation unit through a dedicated intermediate cache (A2 area). The calculation process is completed entirely in parallel within the FPGA's DSP unit and LUT, and the result is the centroid coordinates (x,y on a two-dimensional plane or z value in three dimensions) for each assigned pixel. The centroid position of each pixel is packaged into a small data frame and stored in the centroid result buffer. The FPGA controller sends all the calculated centroid results of its engine to the first computing engine through the GTX interface (reverse channel). After receiving the segmented results from the other computing engines, the first computing engine merges them with its own results and sends them to the host computer or control host through the fiber optic interface. The DDR4 storage area is pre-divided into three parts: the original image area, the maximum value area, and the centroid result buffer, and read and write operations do not conflict with each other. The FDMA / GTX high-bandwidth channel ensures that data transmission and pipeline calculations at each stage are parallel and overlapping. After multiple engines are connected in series, horizontal segmentation and parallelism at the row level are achieved, which greatly improves the overall frame rate and real-time performance.

[0048] Furthermore, the expression for the improved centroid fitting function is:

[0049]

[0050] Where μ represents the peak position obtained from the fitting, argmin represents the optimization operator, which is used to return the independent variable that minimizes the objective function, a represents the Gaussian function amplitude, e represents the natural constant, and z i δ represents the coordinate of the i-th axial position. rLet b represent the standard deviation of the Gaussian function, and g represent the center position of the Gaussian function. norm (z i () represents the intensity value measured at the coordinates of the i-th axial position on the axial response curve, z i ∈FWHM indicates that data points with a peak intensity greater than half that of the axial response curve are selected as valid data points for fitting.

[0051] In one example, the process of generating image sequence data includes: illuminating the sample with spatially modulated structured light from an incoherent structured illumination source via a beam splitter and a microscope objective; and positioning the sample along the optical axis at predetermined steps Δz to z1…z1 using a driving objective or sample stage. n The detectors at each location acquire corresponding structured raw images to form a structured image stack. Optical tomography demodulation algorithms are applied to the structured image stack to obtain optical tomographic image stacks for each axial plane. Based on the tomographic image stacks, an axial response curve is generated for each pixel along the optical axis. A peak localization algorithm is applied to the response curves to determine the optimal depth-of-focus peak position for each pixel, so as to output image sequence data.

[0052] In this embodiment, structured illumination significantly suppresses scattering and background, enhances the optical cross-sectional contrast brought about by the modulation frequency, and fine-grained Δz stepping along the optical axis ensures sampling density, forming a high-quality structured image stack. Optical tomography (e.g., carrier demodulation + filtering) is performed on the image stack to effectively eliminate defocus signals and restore the true optical tomographic maps of each axial plane. Slicing the continuous z-layers makes the resolution of each depth layer approach the system's point spread function limit. Based on the response curve of each pixel along the z-direction, a high-precision peak interpolation / fitting algorithm is used to determine the depth-of-focus peak position. Simultaneously, the end-to-end process from the original structured image → tomographic map → depth-of-focus map directly produces a 3D surface height map for each frame, obtaining high signal-to-noise ratio and high depth-resolution surface morphology data without additional deconvolution or global optimization. Incoherent light sources reduce the risk of phototoxicity / photobleaching, making them suitable for living or sensitive samples. The pipelined processing of structured illumination + tomography + peak localization can be efficiently implemented on FPGA / ASIC, meeting the requirements of high frame rate online detection.

[0053] Furthermore, such as Figure 2 As shown, an incoherent structured illumination source (e.g., a projection grating or digital micromirror array) generates a stripe / lattice light field with known spatial frequencies and phases. The beam is incident on the microscope objective via a beam splitter, forming a high-contrast structured illumination pattern on the sample surface. The driving objective or sample stage is moved sequentially along the optical axis to a series of discrete positions z1, z2, ..., z along a predetermined step size Δz (typically 0.1–1 μm). nAfter stabilizing the structured illumination at each zi, a detector (area confocal camera or high-speed CMOS sensor) is triggered to acquire a frame of the original structured image I. i (x,y), optionally perform k accumulations or averages on each frame to improve the signal-to-noise ratio. Repeat the above steps until z1 to z2 are completed. n Full depth-of-focus scanning yields a structured image stack {I1(x,y),I2(x,y),…,I...} n (x,y)}.

[0054] The structured image stack is input into the optical tomography demodulation module. Common methods include phase retrieval algorithms, Fourier reconstruction, or lock-in amplification techniques. Each layer I... i In (x,y), the intensity is separated and filtered using a known modulation frequency f0 to extract the optical tomographic information O of the focal depth plane. i (x,y). The resulting tomographic image stack {O1,O2,…,O n This accurately reflects the scattering / emission intensity distribution on each depth-of-focus plane. For each pixel (x, y) in the stack, its position in O1…O is extracted in depth-of-focus order. n The intensity values ​​at each point form a discrete axial response curve R_{x,y}(z)=[O1(x,y),O2(x,y),…,O n (x,y)]. To eliminate exposure differences between different levels, each curve can be normalized: I_norm(zi)=O i (x,y) / maxk{O k (x,y)}.

[0055] For each normalized response curve R_{x,y}(z), a peak localization algorithm, such as an improved centroid fitting function, is applied to obtain μ(x,y), which is the optimal depth-of-focus position of pixel (x,y) in the axial direction. The set of μ(x,y) across the entire field is then used to construct a depth map, which is the final sample surface topography map. Simultaneously, the O values ​​of each layer are preserved. i (x,y) or the response curve R_{x,y}(z) can be used as additional sequence data to meet the needs of 3D reconstruction, post-processing or quality assessment.

[0056] The generated raw image stack, tomographic image stack, and depth map can be sent to the FPGA computing engine or host computer at high speed via CXP / fiber optic link for further analysis and visualization. On the FPGA, the above steps can be processed in real time with the help of parallel pipelines and DDR cache partitions, ensuring continuous data output at high frame rates.

[0057] The system comprises multiple computing engines, including a first computing engine, a second computing engine, a third computing engine, and a fourth computing engine connected in series, all of which share identical chip and memory parameters. All four computing engines are FPGA chips, and their latency in receiving image sequence data is less than 1ms.

[0058] Furthermore, when receiving image sequence data, the first computing engine saves the initial row to the first preset row of each frame of the image sequence data into the DDR memory corresponding to the first computing engine, and calculates the first maximum value map corresponding to the initial row to the first preset row of the current image, and saves the first maximum value map into the DDR memory corresponding to the first computing engine.

[0059] When receiving image sequence data, the second computing engine saves the first preset row to the second preset row of each frame of the image sequence data into the DDR memory corresponding to the second computing engine, and calculates the second maximum value map corresponding to the first preset row to the second preset row of the current image, and saves the second maximum value map into the DDR memory corresponding to the second computing engine.

[0060] When receiving image sequence data, the third computing engine saves the second preset row to the third preset row of each frame of the image sequence data into the DDR memory corresponding to the third computing engine, and calculates the third maximum value map corresponding to the second preset row to the third preset row of the current image, and saves the third maximum value map into the DDR memory corresponding to the third computing engine.

[0061] When receiving image sequence data, the fourth computing engine saves the third to fourth preset rows of each frame of the image sequence data into the DDR memory corresponding to the fourth computing engine, and calculates the fourth maximum value map corresponding to the third to fourth preset rows of the current image, and saves the fourth maximum value map into the DDR memory corresponding to the fourth computing engine.

[0062] In this embodiment, each engine only stores the line segment and corresponding frame data it is responsible for, resulting in smaller and more predictable DDR space usage. This eliminates redundant storage of the entire frame image by each engine, reducing memory requirements and power consumption. During frame stream access, each engine can calculate the maximum value map of its segment in parallel with data reception, achieving overlap between reception and preprocessing, reducing overall pipeline latency. Furthermore, each locally generated maximum value map only needs to retain key grayscale peaks, eliminating the need for full frame backhaul, significantly reducing link transmission volume. Only in the later stages is the maximum value map of each segment (size ≈ 5-10% of the original image) summarized, rather than the entire frame. This causes a sharp drop in GTX / fiber link load, ensuring sufficient margin for critical data such as peak positioning and fiber output. The released bandwidth can then be used to support higher frame rates or larger resolution inputs, improving overall system resilience. The first engine promptly intercepts the frame header and processes the first segment, while the last segment is forwarded sequentially. Each engine forms a three-stage pipeline of "reception → local calculation → forwarding," with single-frame end-to-end latency controllable within milliseconds, effectively smoothing data jitter and sudden loads, and ensuring stable output. When adding the N+1th segment, it is only necessary to insert the corresponding engine at the end of the link and adjust the preset line number, without reconstructing the original storage or calculation logic; if a segment engine fails, it only affects the generation of the maximum value graph of its segment, which can be supplemented or downgraded by adjacent engines in real time, and the system as a whole remains online.

[0063] In one example, when the first computing engine receives the image sequence, it saves rows 0 to 1 / 4 of each image to its DDR memory and calculates its maximum value, saving the corresponding maximum value image to DDR memory. The second computing engine performs the operations for rows 1 / 4 to 1 / 2; the third computing engine performs the operations for rows 1 / 2 to 3 / 4; and the fourth computing engine performs the operations for rows 3 / 4 to the last. If performance requirements increase later, the number of computing engines and operations can be increased using this method.

[0064] In this embodiment, multiple computing engines are used to perform tasks according to row regions, achieving fragmented processing and parallel computing, which significantly shortens the overall reconstruction time. Each engine runs optical tomography demodulation and peak positioning simultaneously, and the overall throughput increases almost linearly with the number of engines. Each engine is equipped with independent DDR memory to store image segments, maximum value maps, and intermediate results locally, reducing data transfer overhead across modules. Furthermore, intermediate results are only aggregated through the GTX interface when necessary, significantly reducing bus pressure. The first engine uses CXP to receive high-speed camera data, uses GTX for low-latency forwarding within the link, and uses the fiber optic interface to aggregate data to the host computer in real time, ensuring high bandwidth and low latency end-to-end. This supports continuous online processing of high frame rate and high resolution image sequences. At the same time, local failures of each engine will not cause the entire link to collapse, but only affect the corresponding row region. It can be quickly recovered through hot standby or fault switching, enabling more refined load balancing and utilization of resources, and improving the computing throughput of the accelerated computing architecture.

[0065] Please see Figure 3 , Figure 3 This diagram illustrates the storage process of DDR memory. A confocal camera or high-speed CMOS sensor is connected via CXP / fiber optic cable, supporting continuous transmission of ≥10Gbps. The receiving logic writes line-by-line (or pixel-by-pixel) streaming data into the FPGA's line buffer and packages it into DDR-A via the AXI-4FDMA controller. Frame synchronization signals (Start-of-Frame / End-of-Frame) drive the read / write timing of subsequent stages.

[0066] exist Figure 3 In DDR-A, the channel width per pixel is 256 bits, the original image area is 8 bits × number of pixels, the ΣA2 middle area is 50–68 bits, and the bit width required to accumulate the square intensity of each depth of focus is used; in DDR-B, the channel width per pixel is 512 bits, the maximum value image area is 29 bits, and the ΣA2i middle area is stored in two segments, with some results being 0–49 bits (accumulated in online streaming) and the complete result being 0–77 bits (finally accumulated after combining with the maximum value image).

[0067] Stage 1 Image Preprocessing (i.e., ΣA2 Calculation): The data flow is from the DDR-A original image area → A2 calculation pipeline → DDR-A ΣA2 area. Each incoming pixel I(x,y,zi) (8 bits) is fed into the multiplier in parallel to calculate I. 2 (16 bits); will I 2 The result is accumulated with the previous level (50–68 bits wide) to output a new ΣA2; the entire pipeline depth can reach 5–10 stages, supporting one pixel throughput per cycle at a clock speed of 250MHz. It runs in parallel with stage 2 and alternates between “writing the second half of ΣA2” and “reading the first half of ΣA2” through 256-bit AXIFDMA.

[0068] Phase 2: Maximum value extraction and partial accumulation of ΣA2i: The data flow is from DDR-A ΣA2 area → Max comparison unit → DDR-B maximum value map area, while simultaneously sending ΣA2 data to the ΣA2i partial accumulation unit. Each row of pixel sequence undergoes pipelined comparison (several comparators in parallel), ultimately outputting a 29-bit intensity peak value M(x,y); after the peak value is generated, it is written back to DDR-B; ΣA2 and M(x,y) are received synchronously, and intensity weighted accumulation (0–49 bits) begins to be compared; transmission: a 256-bit FDMA channel is used again, with the Max result and partial ΣA2i being interleaved and transmitted on the same bus.

[0069] Phase 3 ΣA2i complete accumulation and centroid calculation preparation: The data flow is as follows: DDR-A original image area + DDR-B maximum value image area + DDR-B ΣA2i partial area → ΣA2i complete accumulation unit → DDR-B ΣA2i full area. The original I(x,y,zi), M(x,y) and partial ΣA2i are concatenated, and weighted or normalized by pixel to expand to 77 bits. The data is written back to DDR-B to prepare complete data for the centroid stage.

[0070] Each stage uses an independent FDMA channel (256-bit for Stage 1 / 2, 512-bit for Stage 3), ensuring no read / write conflicts and true pipelining. Partitioned DDR memory and deep pipeline guarantee write-to-use functionality, with latency controlled to <100 cycles. The FPGA's internal BRAM and DSP resources are compactly multiplexed, supporting multi-channel, multi-resolution parallel processing.

[0071] Based on the above architecture, this application discloses a distributed accelerated computing method for surface topography restoration based on surface confocal surface, including the following steps:

[0072] Step S1: Image sequence data is transmitted from the area confocal sensor to the first computing engine via the CXP interface;

[0073] Step S2: The first computing engine stores the preset row of each frame of the image sequence data into the DDR memory corresponding to the first computing engine, calculates the maximum pixel value corresponding to the preset row, and forwards the remaining rows of each frame of the image to the remaining computing engine through the GTX interface, so that the remaining computing engine stores the corresponding row segments and the maximum value image respectively.

[0074] Step S3: After all computing engines generate the corresponding maximum value map, each computing engine reads the original image and the corresponding maximum value map from the corresponding DDR memory and intermediate cache, and uses the improved centroid fitting function to extract the centroid.

[0075] Step S4: The centroid extraction results are reversed by each computing engine and aggregated to the first computing engine via the GTX interface. The first computing engine then sends the centroid extraction results to the host computer via the fiber optic interface.

[0076] In this embodiment, steps S1 and S2 employ a three-stage pipeline of receiving → local maximum / minimum value calculation → forwarding. Each engine can execute in parallel with data receiving and forwarding, fully overlapping I / O and computation, significantly reducing end-to-end latency. Multiple computing engines work collaboratively in rows, and the overall reconstruction throughput is approximately N × the processing rate of a single engine, demonstrating a significant linear acceleration effect.

[0077] In step S2, the maximum pixel value map of the preset row of each frame is extracted in real time and stored in the local DDR. There is no need to send back the whole frame, which reduces the amount of data across the engine / host computer. Only the necessary original row segments and centroid extraction results are forwarded, so that the GTX / fiber link is kept at a very low load and more bandwidth is released to support higher frame rates or larger resolutions.

[0078] In step S3, an improved centroid fitting function is used to optimize the noise and asymmetry characteristics of the surface confocal response curve, which accelerates convergence, suppresses defocus pseudo-peaks, improves the centroid positioning accuracy from sub-pixel to sub-nanometer level, enhances the anti-interference ability of the extraction results, and improves the depth resolution and reliability of subsequent 3D topography reconstruction.

[0079] In step S4, data is aggregated via reverse GTX to the first engine and then directly output to the host computer via fiber optic cable. The end-to-end total latency is controllable to the millisecond level, meeting the requirements of industrial online inspection or high-speed imaging for scientific research. The host computer only needs to receive the final centroid data, simplifying the backend storage and post-processing workflow. Each engine segment is pre-parameterized, allowing for plug-and-play operation. A single point of failure only affects the corresponding segment, and the system can degrade to a lower level or be supplemented by an adjacent engine, ensuring uninterrupted online operation. Multi-level parallel / serial hybrid scheduling achieves the optimal balance between high throughput and high hardware utilization.

[0080] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A distributed accelerated computing system based on surface confocal topography recovery, characterized in that, It includes multiple cascaded computing engines, among which, The first computing engine in the computing engine receives image sequence data through the CXP interface and forwards the image sequence data to the remaining computing engines through the GTX high-speed interface. Each computing engine saves the preset row of the received image sequence data to the DDR memory corresponding to each computing engine. Based on optical tomography demodulation calculation, the maximum pixel value in the preset row of each frame of the image sequence data is obtained. The peak position corresponding to the maximum pixel value is obtained by peak positioning calculation. The maximum pixel value and the peak position are saved in the DDR memory corresponding to each computing engine. The maximum pixel value represents the maximum value of the intensity values ​​of all pixels in each frame of the image. After the first computing engine completes the optical tomography demodulation calculation and peak positioning calculation required for surface confocal topography restoration, the maximum pixel value and peak position calculated by the remaining computing engines in the computing engine are forwarded to the first computing engine through the GTX interface, and the first computing engine transmits the summarized results to the host computer through the optical fiber interface. Peak location is calculated using an improved centroid fitting function, wherein the improved centroid fitting function is expressed as: ; in, This indicates the peak position obtained from the fitting. This represents an optimization operator, which returns the independent variable that minimizes the objective function. a This represents the amplitude of the Gaussian function. e Represents the natural constant. z i δ represents the coordinate of the i-th axial position. r The standard deviation of the Gaussian function is represented by... b Indicates the center position of the Gaussian function. This represents the intensity value measured at the coordinates of the i-th axial position on the axial response curve. This means selecting data points with a peak intensity greater than half that of the axial response curve as valid data points for fitting.

2. The distributed accelerated computing system based on surface confocal topography recovery as described in claim 1, characterized in that, The plurality of serially connected computing engines include a first computing engine, a second computing engine, a third computing engine, and a fourth computing engine connected in series, and the first computing engine, the second computing engine, the third computing engine, and the fourth computing engine have the same chip and memory parameters.

3. The distributed accelerated computing system based on surface confocal topography recovery as described in claim 2, characterized in that, When the first computing engine receives image sequence data, it saves the initial row to the first preset row of each frame of the image sequence data into the DDR memory corresponding to the first computing engine, and calculates the first maximum pixel value corresponding to the initial row to the first preset row of the current image, and saves the first maximum pixel value into the DDR memory corresponding to the first computing engine. When receiving image sequence data, the second computing engine saves the first preset row to the second preset row of each frame of the image sequence data into the DDR memory corresponding to the second computing engine, and calculates the second maximum pixel value corresponding to the first preset row to the second preset row of the current image, and saves the second maximum pixel value into the DDR memory corresponding to the second computing engine. When receiving image sequence data, the third computing engine saves the second preset row to the third preset row of each frame of the image sequence data into the DDR memory corresponding to the third computing engine, and calculates the third maximum pixel value corresponding to the second preset row to the third preset row of the current image, and saves the third maximum pixel value into the DDR memory corresponding to the third computing engine. When receiving image sequence data, the fourth computing engine saves the third preset row to the fourth preset row of each frame of the image sequence data into the DDR memory corresponding to the fourth computing engine, calculates the fourth maximum pixel value corresponding to the third preset row to the fourth preset row of the current image, and saves the fourth maximum pixel value into the DDR memory corresponding to the fourth computing engine.

4. The distributed accelerated computing system based on surface confocal topography recovery as described in claim 1, characterized in that, Each computing engine's DDR memory is divided into a first region, a second region, and a third region. The first region is used to store segmented image sequence data, the second region is used to store the corresponding maximum pixel value, and the third region is used to store intermediate results.

5. The distributed accelerated computing system based on surface confocal topography recovery as described in claim 1, characterized in that, The GTX interface between the computing engines includes a pair of transceivers for forwarding image sequence data and reverse transmission of centroid results.

6. The distributed accelerated computing system based on surface confocal topography recovery as described in claim 2, characterized in that, The first computing engine, the second computing engine, the third computing engine, and the fourth computing engine are all FPGA chips, and the latency of the first computing engine, the second computing engine, the third computing engine, and the fourth computing engine in receiving the image sequence data is less than 1ms.