Parallel four-neighborhood real-time multi-light-spot center detection method based on FPGA
Through the parallel four-neighborhood real-time multi-spot center detection method of FPGA, combined with image preprocessing and connectivity domain center detection, the problems of high hardware resource occupation and insufficient computing accuracy are solved, and efficient and real-time multi-spot center detection is achieved.
Patent Information
- Application Number
- CN202510641312.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing multi-spot detection algorithm based on FPGA has problems such as high hardware resource occupation, loss of calculation accuracy and insufficient real-time performance, and it is difficult to meet the real-time processing needs of high frame rate and large-resolution images.
The parallel four-neighborhood real-time multi-spot center detection method based on FPGA is adopted, including two-level median filtering, adaptive binarization and morphological processing modules. Combined with the four-neighborhood center detection module, through image preprocessing and connectivity domain center detection, the pipeline architecture is used to realize label allocation, equivalence table merging and center calculation of center.
It reduces hardware resource usage, improves computing accuracy and real-time performance, realizes accurate positioning of subpixel-level spot centers, and adapts to the real-time detection needs of high-dynamic scenarios.
Smart Images

Figure CN120467178A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical measurement and laser radar, and relates to a parallel four-neighborhood real-time multi-spot center detection method based on FPGA. Background Art
[0002] Multi-spot center detection technology is a core technology in optical measurement, lidar, and national defense science and technology. Its core goal is to quickly and accurately extract the sub-pixel center coordinates of multiple light spots. In high-dynamic scenarios such as space laser communications and industrial precision inspection, the ability to locate the center of the light spot in real time directly affects the system's measurement accuracy and response speed. Traditional methods rely primarily on general-purpose processing devices such as CPUs or ARM architectures, implementing light spot detection through software algorithms. However, such methods have inherent flaws: general-purpose processors lack hardware-level parallel computing capabilities, resulting in large processing delays and high power consumption, making it difficult to meet the real-time processing requirements of high frame rates and large resolution images.
[0003] In recent years, field-programmable gate arrays (FPGAs) have become the preferred solution for real-time image processing due to their parallel computing advantages and customizable hardware architecture. However, existing FPGA-based multi-spot detection algorithms still face significant challenges: First, connected domain labeling algorithms require a large amount of storage resources to cache intermediate results, resulting in high hardware resource utilization. Second, existing methods suffer from computational precision loss during label merging and coordinate accumulation, affecting the sub-pixel accuracy of centroid positioning. Third, the algorithm relies on multiple full-image scans, making it difficult to pipeline frame processing and data output, limiting the system's real-time performance.
[0004] Therefore, there is an urgent need for a multi-spot center detection method that can balance low resource usage, high computational accuracy and real-time performance to resolve the contradiction between hardware resource efficiency, algorithm accuracy and processing timeliness in existing FPGA implementation solutions. Summary of the Invention
[0005] In view of this, the object of the present invention is to provide a parallel four-neighborhood real-time multi-spot center detection method based on FPGA.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A parallel four-neighborhood real-time multi-spot center detection method based on FPGA includes two core functional modules: image preprocessing and connected domain center detection module. The image preprocessing module includes a two-stage median filter submodule, an adaptive binarization submodule and a morphological processing submodule.
[0008] Furthermore, the two-stage median filter submodule uses multiple synchronous FIFOs to cache multiple lines of image data. The 3×3 median filter template uses two synchronous FIFOs to cache the first two lines of image data, and the 5×5 median filter template uses four synchronous FIFOs to cache the first four lines of data, which respectively form 3×3 and 5×5 image data matrices with the image data of the current line. The matrix window is slid to the right once per clock cycle, and the median of the matrix elements is calculated as the current pixel to realize the median filtering function.
[0009] Furthermore, the adaptive binarization submodule uses the highest grayscale value of the previous frame as a reference, sets the binarization threshold of the current frame to 1 / 2 or 1 / 4 of the reference value, outputs 0 for grayscale values less than the threshold, and outputs 1 for grayscale values greater than the threshold.
[0010] Furthermore, the morphological processing module uses a 3×3 erosion template and a 3×3 expansion template respectively. The expansion and erosion templates have the same implementation code, and the expansion and erosion functions can be achieved by setting different thresholds.
[0011] Furthermore, the four-neighborhood center detection algorithm module includes the following steps:
[0012] Step 1: Input binary image data;
[0013] Step 2: Filter the target pixel and check its neighborhood labels;
[0014] Step 3: Assign new labels or merge equivalent labels;
[0015] Step 4: Accumulate the coordinates and pixel values of the pixels corresponding to the labels;
[0016] Step 5: Determine whether the current frame has been processed;
[0017] Step 6: Parse the equivalence table and calculate the centroid.
[0018] Furthermore, step 2 determines for each pixel in the input video data whether its pixel value is greater than 0. If the pixel value is 0, the pixel is considered to belong to the background and does not participate in subsequent calculations; if the pixel value is greater than 0, the subsequent connected component marking step is entered.
[0019] Furthermore, step 3 uses a dual-port RAM to construct an equivalence table, takes the address of the RAM as the node value, and the value under the corresponding address as the label value, then assigns the label values of the current pixel and its four neighboring pixels (left, right, top, and top-right) to the label1, label2, label3, label4, and label5 registers, and uses combinational logic to find the minimum non-zero value of these five labels. In the first clock cycle, the timing logic is used to assign the minimum label value to label1_r1 as the main label; in the second clock cycle, the label with a label value of 0 is marked separately to avoid invalid merging, and label_r1 is assigned to label_r2 as the main label; in the third clock cycle, the label node points to the main label to complete the equivalent merge.
[0020] Furthermore, step 4 uses three dual-port RAMs to store the accumulated values of the horizontal coordinates, vertical coordinates, and pixel values corresponding to different label values. Since the RAM has a delay of one clock cycle for the read data, it is impossible to read, modify, or write the value of the corresponding address in the RAM within one clock cycle. Therefore, two registers are defined to cache the rewritten data for one and two beats respectively, forming a two-stage pipeline structure to realize continuous read, modify, and write operations on the same address of the RAM.
[0021] Furthermore, the step 6 of parsing the equivalence table is to traverse the equivalence table RAM by address, find out the labels with equivalent relationships, and calculate the center of the connected domain under the equivalent labels. Three distributed RAMs are used to store the horizontal coordinates, vertical coordinates and pixel values of the pixels with equivalent labels and accumulate them. The spot centroid algorithm formula is used for each connected domain:
[0022]
[0023] Where X and Y represent the center coordinates, N and M represent the value ranges of the horizontal and vertical coordinates of the current connected domain, x and y represent the coordinate values of the current pixel, and I(x, y) represents the grayscale value of the current pixel coordinate. In the binarized image, I(x, y) = 1, so this item can be ignored.
[0024] A detection system for implementing the method, comprising:
[0025] An image acquisition module configured to receive an 8-bit grayscale image signal and output it via a parallel data bus;
[0026] The preprocessing module includes a cascaded 5×5 median filter unit, a 3×3 median filter unit, an adaptive binarization unit, and a morphological processing unit, where:
[0027] The 5×5 median filter unit buffers the first 4 lines of image data through 4 synchronous FIFOs.
[0028] The 3×3 median filter unit buffers the first two lines of image data through two synchronous FIFOs.
[0029] The morphological processing unit uses a 3×3 sliding window and is configured to perform a dilation or erosion operation;
[0030] The real-time processing module receives the binary image signal output by the pre-processing module and includes:
[0031] Four neighborhood label analysis units are configured to assign labels to target pixels and merge equivalence relations.
[0032] a centroid calculation unit configured to parse the tag data and calculate the coordinates of the center of the light spot;
[0033] Output interface module, outputs centroid coordinate data through AXI-Stream protocol;
[0034] The system uses a global synchronous clock and asynchronous reset signal, and each module is connected through a pipeline architecture to meet the following requirements:
[0035] The line buffer delay between the image acquisition module and the preprocessing module is 2 clock cycles.
[0036] The label allocation and equivalence table merging operations of the real-time processing module take up 3 clock cycles.
[0037] The output delay of the center of mass coordinates does not exceed 1 frame period.
[0038] The beneficial effects of the present invention are:
[0039] (1) A four-neighborhood connected domain labeling algorithm combined with a two-level pipeline accumulation mechanism is used to implement label merging and coordinate calculation through on-chip storage, which greatly reduces external storage dependence and hardware resource usage, allowing the system to run efficiently on a low-cost FPGA platform.
[0040] (2) Based on pixel-level pipeline processing and intra-frame parallel computing, a seamless connection between spot detection and centroid output is achieved. After one frame of image transmission is completed, the center positioning of all connected domains can be completed without a second scan, meeting the microsecond-level real-time response requirements in high-dynamic scenes.
[0041] (3) The influence of environmental interference on the light spot contour is suppressed through the collaborative denoising of two-level median filtering and adaptive binarization. At the same time, the dynamic merging strategy of the equivalence table is combined with the sub-pixel centroid algorithm to ensure the accurate extraction of the edge of the connected domain and the sub-pixel positioning of the center coordinates.
[0042] (4) The morphological processing module can flexibly configure the corrosion and expansion parameters to adapt to the spot morphology correction in different noise environments; the adaptive threshold mechanism dynamically adjusts the segmentation threshold according to the grayscale characteristics of the previous frame to improve the detection stability under complex lighting conditions.
[0043] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0045] Figure 1 This is the system structure diagram of the FPGA-based parallel four-neighborhood real-time multi-spot center detection method;
[0046] Figure 2 This is the module connection diagram of the FPGA-based parallel four-neighborhood real-time multi-spot center detection method;
[0047] Figure 3 Generate schematics for an FPGA-based pixel matrix sliding window;
[0048] Figure 4 This is the flow chart of the parallel four-neighborhood center detection algorithm based on FPGA;
[0049] Figure 5 Construct an operational schematic for the equivalent table RAM in the FPGA-based parallel four-neighborhood center detection algorithm;
[0050] Figure 6 This is a schematic diagram of the read, modify, and write operations of the pixel coordinate accumulation RAM in the FPGA-based parallel four-neighborhood center detection algorithm;
[0051] Figure 7 This is a comparison chart of the simulation results of the connected domain center detection algorithm between MATLAB and VIVADO;
[0052] Figure 8 This is the timing diagram of the key signals inside the FPGA-based parallel four-neighborhood center detection algorithm module. DETAILED DESCRIPTION
[0053] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0054] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0055] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0056] This paper proposes a parallel, four-neighborhood, real-time, multi-spot center detection method based on an FPGA. Compared to traditional hardware-connected domain multi-spot algorithms, this method reduces hardware resource utilization and improves computational accuracy and real-time performance. Therefore, this paper is of great significance in scenarios with stringent requirements for real-time performance, hardware resources, and computational accuracy.
[0057] like Figure 1 As shown, a parallel four-neighborhood real-time multi-spot center detection method based on FPGA includes: a two-stage median filter submodule, an adaptive binarization submodule, a morphological processing submodule, and a parallel four-neighborhood center algorithm module.
[0058] Each module is based on Figure 2The modules are connected in a 3D manner, with each module sharing the clock and reset interfaces. An 8-bit grayscale video signal is used as input and connected to the 5×5 median filter module. The output of the 5×5 median filter is then connected to the input of the 3×3 median filter module to perform noise reduction. The output of the 3×3 median filter module is then connected to the input of the adaptive binarization module to perform binarization. The output of the adaptive binarization module, along with a pre-set threshold, is then connected to the input of the morphological processing module to perform morphological processing. Finally, the output of the morphological processing module is connected to the input of the four-neighborhood center algorithm module to perform multi-spot center detection.
[0059] Furthermore, the median filter module and the dilation / erosion module will construct a key sliding window, and the construction process of the window is as follows: Figure 3 As shown, FIFO-2 and FIFO-1 form a chain cache structure. The data read in the current row is first cached in FIFO-2. After reading a row of data, the next row of data is read. At the same time, the data is written into FIFO-2 and the data is read from FIFO-2 and written into FIFO-1. When the third row of data is read, FIFO-2 and FIFO-1 cache the image data of the previous row and the row before that respectively. At the same time, data is read from FIFO-2 and FIFO-1. Together with the data of the current row, a 3×3 sliding pixel window can be formed. By medianizing the pixel values in the window and taking the median as the current pixel value, the median filtering function can be implemented.
[0060] Repeat the above operation to obtain a 3×3 sliding pixel window. After accumulating all the pixels in the window, compare it with the threshold set in advance. If the accumulated value of the pixels in the window is greater than the threshold, the current pixel value is set to 1; if it is less than the threshold, the current pixel value is set to 0, thus realizing the functions of dilation and erosion.
[0061] Furthermore, after completing the two-stage median filtering, binarization and morphological processing operations, the obtained binarized spot image is used as input to execute the parallel four-neighborhood center detection algorithm. The algorithm flow is as follows: Figure 4 As shown, the following steps are included:
[0062] Step 1: Input binary image data;
[0063] Step 2: Filter the target pixel and check its neighborhood labels;
[0064] Step 3: Assign new labels or merge equivalent labels;
[0065] Step 4: Accumulate the coordinates and pixel values of the pixels corresponding to the labels;
[0066] Step 5: Determine whether the current frame has been processed;
[0067] Step 6: Parse the equivalence table and calculate the centroid.
[0068] Furthermore, the workflow of step 3 is as follows Figure 5 As shown in the figure, in the process of labeling pixels, pixels with a value of 1 are treated as target pixels, and pixels with a value of 0 are not processed. To label the target pixel, you first need to declare a register type variable label_cnt. label_cnt records the current label value starting from 1. When encountering the target pixel, check whether there are already labeled pixels on the left, upper left, upper right, and upper right sides of the target pixel. If there are already labeled pixels, the one with the smallest label value in these four neighborhoods is selected as the label of the current target pixel; if there are no labeled pixels around the target pixel, the value of label_cnt is used as the label value of the target pixel, and label_cnt+1 is assigned to label_cnt.
[0069] While performing the labeling operation, the labels need to be merged. Use a dual-port RAM to build an equivalence table RAM. The address of the RAM is used as the root label, and the value of the RAM is used as the label equivalent to the corresponding address. The pixel label value of the pixel coordinate (2, 2) is 1, and the pixel label value of the neighboring pixel on the left is 2. Therefore, the value corresponding to address 2 of the equivalence table RAM is assigned to 1, indicating that the pixels corresponding to label value 1 and label value 2 are in the same connected domain.
[0070] Furthermore, step 4 accumulates the horizontal coordinate value, vertical coordinate value and pixel value of pixels with the same label by constructing three dual-port RAMs, such as Figure 6 As shown, a pixel value accumulation RAM is constructed. When the video signal is transmitted to the pixel coordinate position (4, 3), 7 pixels with a label value of 1 have appeared in front of the pixel point. Therefore, the value of pixel value accumulation RAM address 1 at this time is 7. The value 7 at this address is read out in the first clock cycle, and the value is added by 1 in the second clock cycle to replace the value at pixel value accumulation RAM address 1. Similarly, the X-axis coordinate accumulation RAM and the Y-axis coordinate accumulation value RAM can be constructed respectively to store the horizontal and vertical coordinate accumulation values of pixels with the same label value.
[0071] Furthermore, after completing the transmission of a frame of image, step 6 traverses the equivalence table RAM in step 3 according to the address. During the traversal process, if the address and value of the equivalence table RAM are different, the value of the equivalence table RAM is used as the main label, the address of the equivalence table RAM is used as the sub-label, and the main label and the sub-label are used as the addresses of the pixel, horizontal and vertical coordinate cumulative value RAM in step 4. The value under the address corresponding to the main label is read as the main value, and the value under the address corresponding to the sub-label is read as the sub-value. The sub-value is accumulated to the main value and written back to the address corresponding to the main label. Then, the value under the address corresponding to the sub-label is set to zero, and the above operation is repeated to finally obtain the pixel, horizontal and vertical coordinate cumulative value RAM after label equivalent merging. The addresses corresponding to the values in these RAMs that are not 0 represent the connected domain labels; if the address and value of the equivalence table RAM are the same, no operation is performed.
[0072] Finally, the values of the accumulated values of pixels, horizontal and vertical coordinates in RAM at different addresses are substituted into the above centroid algorithm to obtain the center coordinates of different connected domains.
[0073] Furthermore, VIVADO was used to simulate the parallel four-neighborhood center detection algorithm. 16 64×64 binary images were continuously input, and each image had two connected domains. The calculated coordinates of the two connected domain centers were recorded. The same 16 images were processed using the MATLAB (R2022b) system function [L,n]=bwlabel(BW,conn). The bwlabel function parameter description is shown in Table 1. The coordinates of the two connected domain centers were also calculated. The difference between the two coordinates and the coordinates obtained by VIVADO simulation was used as the reference value, and the absolute value was taken to obtain the following: Figure 7 As shown in the comparison results, the absolute errors of the four center coordinate values obtained by the FPGA parallel four-neighborhood center detection algorithm and the four center coordinate values obtained by the MATLAB function are all between 0 and 0.02 pixels. Therefore, the parallel four-neighborhood center detection algorithm can meet the requirements of high-precision calculation.
[0074] Table 1
[0075] Parameter name illustrate L Label Matrix n Number of connected objects BW binary image conn Pixel connectivity (default 8-connected)
[0076] Furthermore, the time delay analysis of the parallel four-neighborhood center detection algorithm is carried out. In the simulation, the key signals inside the parallel four-neighborhood center detection algorithm module are as follows: Figure 8As shown in the figure, the description of each signal is shown in Table 2. In theory, the next frame of data can be received when the global RAM initialization signal is pulled low. Therefore, we can use the period from the input frame synchronization signal being pulled low to the global RAM initialization signal being pulled low to represent the algorithm's runtime. If the system uses a 50MHz clock, that is, a clock cycle of 20ns, the number of clock cycles from the input frame synchronization signal being pulled low to the global RAM initialization signal being pulled low is approximately 544. Therefore, the algorithm runtime is 544 × 20 = 10880ns. This time is only related to the FPGA system clock and internal RAM depth, and is independent of the frame rate or resolution of the input video. It can meet the real-time processing requirements of high frame rate and high resolution images.
[0077] Table 2
[0078] Signal name illustrate clk clock signal pre_img_vsync Input frame synchronization signal post_img_vsync Output frame synchronization signal find_ET Lookup equivalence table signal find_ET_cnt[7:0] Lookup equivalence table counter output_flag Connected domain center coordinate output signal output_flag_cnt[4:0] Connected region center coordinate output counter init_RAM Global RAM initialization signal init_RAM_cnt[7:0] Global RAM initialization counter
[0079] Furthermore, the functional modules involved in the present invention were placed and routed in VIVADO. The FPGA selected was the xc7z020clg400-2 chip, which is a mid-to-low-end SOC+FPGA heterogeneous chip from Xilinx. After checking the hardware resource usage after placement and routing, as shown in Table 3, the utilization rate of key resources was mostly below 10%. Therefore, the algorithm can be implemented using a lower-end FPGA, effectively reducing device costs.
[0080] Table 3
[0081] Resource Name Number of users Total resources Utilization rate (%) LUT 6486 53200 12.19 LUTRAM 1020 17400 5.86 FF 1997 106400 1.88 BRAM 4 140 2.88 BUFG 1 32 3.13
[0082] In summary, the present invention implements a parallel four-neighborhood center detection algorithm through FPGA, which has the advantages of low processing delay, low resource usage and accurate results.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A parallel four-neighborhood real-time multi-spot center detection method based on FPGA, characterized by: The following steps are involved: The input image is noise-filtered using a two-stage median filter submodule, where the first stage uses a 5×5 filter template and the second stage uses a 3×3 filter template; Converting the filtered image into a binary image through an adaptive binarization submodule, wherein the adaptive binarization submodule adopts a dynamic threshold generation mechanism; Perform edge optimization processing on binary images through morphological processing submodule; Multi-spot center detection is performed through the four-neighborhood center detection algorithm module, which includes: Filter the target pixel and check its neighborhood label status, Assign new labels or merge equivalent labels based on the neighboring label status, The pixel coordinates corresponding to the labels are accumulated and stored through the two-stage pipeline architecture. After the image frame processing is completed, the equivalence relationship is resolved and the coordinates of the center of mass of the light spot are calculated.
2. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The two-stage median filter submodule uses a first-in-first-out memory (FIFO) to construct a sliding window, where: The 5×5 filter template uses 4 synchronous FIFOs to cache the first 4 lines of data. The 3×3 filter template uses two synchronous FIFOs to cache the first two lines of data. The sliding window is shifted right by one bit in each clock cycle to calculate the median.
3. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The threshold generation method of the adaptive binarization submodule is: T k =α·max(G k-1 ) Among them, T k represents the threshold of the kth frame, α is the proportional coefficient of the (0,1) interval, G k-1 represents the highest grayscale value of the k-1th frame, When the pixel value is greater than T k Output 1 when , otherwise output 0.
4. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The morphological processing submodule adopts a configurable template structure and is implemented by setting different operation thresholds: When the accumulated value of pixels in the template is greater than the first threshold, a dilation operation is performed. When the accumulated value of pixels in the template is less than the second threshold, the corrosion operation is performed. The template structure uses a 3×3 neighborhood window.
5. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The equivalent label merging method of the four-neighborhood center detection algorithm module includes: Constructing dual-port random access memory (RAM) storage equivalence, Compare the current pixel with its left, right, above, and right-upper neighbor labels, Determine the main label and update the equivalence table through a three-stage pipeline structure. The third-level pipeline points the equivalent label node to the main label.
6. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The two-stage pipeline architecture includes: Three independent dual-port RAMs store the accumulated value of the horizontal coordinate, the accumulated value of the vertical coordinate and the pixel count respectively. Set up two registers to build a read-modify-write pipeline, Each clock cycle completes the address read → value update → write back operation.
7. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The centroid coordinate calculation formula is: Where (x i ,y i ) is the coordinate of the i-th pixel in the connected domain, N is the total number of pixels in the connected domain, The calculation is started within 1 clock cycle after the image frame processing is completed.
8. The FPGA-based parallel four-neighborhood real-time multi-spot center detection method according to claim 1, characterized in that: The method, when implemented in FPGA, includes: Under the condition of clock frequency 50MHz, The processing delay of each frame is t = (544 ± 10) clock cycles. Hardware resource utilization is less than 15%, The LUT resource occupancy rate is ≤12.2%, and the BRAM resource occupancy rate is ≤2.9%.
9. A detection system for implementing the method according to any one of claims 1 to 8, characterized in that: include: An image acquisition module configured to receive an 8-bit grayscale image signal and output it via a parallel data bus; The preprocessing module includes a cascaded 5×5 median filter unit, a 3×3 median filter unit, an adaptive binarization unit, and a morphological processing unit, where: The 5×5 median filter unit buffers the first 4 lines of image data through 4 synchronous FIFOs. The 3×3 median filter unit buffers the first two lines of image data through two synchronous FIFOs. The morphological processing unit uses a 3×3 sliding window and is configured to perform a dilation or erosion operation; The real-time processing module receives the binary image signal output by the pre-processing module and includes: Four neighborhood label analysis units are configured to assign labels to target pixels and merge equivalence relations. a centroid calculation unit configured to parse the tag data and calculate the coordinates of the center of the light spot; Output interface module, outputs centroid coordinate data through AXI-Stream protocol; The system uses a global synchronous clock and asynchronous reset signal, and each module is connected through a pipeline architecture to meet the following requirements: The line buffer delay between the image acquisition module and the preprocessing module is 2 clock cycles. The label allocation and equivalence table merging operations of the real-time processing module take up 3 clock cycles. The output delay of the center of mass coordinates does not exceed 1 frame period.
Citation Information
Patent Citations
Image processing apparatus, image forming apparatus, image reading apparatus and image processing method
CN101064009A
FPGA-oriented high-parallelism light spot segmentation method
CN112330611A
Light spot center positioning method and system optimized by using convolutional neural network
CN116823941A
High-precision laser spot center detection method based on FPGA
CN118261980A
ZYNQ-based real-time multi-target detection system
CN119136063A