Single field multi-depth target focusing method and system for single continuous focusing scan

CN122802789APending Publication Date: 2026-09-22SUZHOU MINGJIAN IMAGING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611241524.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0008]本发明提出了单次连续调焦扫描的单视野多深度目标对焦方法及系统,旨在解决现有技术中存在的在单一固定光学视野内面对深度差超出单次成像景深范围的多个目标时须多次串行往复调焦导致检测效率低下的技术问题

Benefits of technology

[0039]第一,由于采用单次、单向连续调焦扫描代替多次串行往复搜索,对焦耗时恒定为单次扫描时间而与目标数量无关,从根本上突破了光学景深对多深度目标同时检测的效率瓶颈。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802789A_ABST
    Figure CN122802789A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of optical imaging and auto-focusing, and particularly relates to a single-continuous-focusing-scan single-view multi-depth target focusing method and system. The method selects N target regions with axial depth differences in a fixed view, generates corresponding pixel masks, controls an imaging system to perform single one-way continuous focusing scan, simultaneously collects an image sequence, calculates a definition evaluation value in parallel by operating each image with each mask to extract local data and obtaining a definition curve of each target region, extracts a peak corresponding optimal focal plane position or image frame, and respectively outputs each target sharpest image. The application can complete multi-depth target focusing in parallel by only one scan, the focusing time consumption is constant, mechanical return error is eliminated, and the detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical imaging and autofocus technology, and in particular to a single-field-of-view multi-depth target focusing method and system for single-segment continuous focusing scan. Background Technology

[0002] In the fields of optical imaging and autofocus technology, when multiple target features at different physical depths exist within the same fixed optical field of view, a single focal plane cannot simultaneously achieve sharp imaging of all targets due to the limited optical depth of field of the imaging system. Traditional autofocus methods typically employ a serial reciprocating search strategy: an independent autofocus process is performed for each target area, driving the focusing mechanism to the approximate depth range of the target, using search strategies such as hill-climbing algorithms to find the sharpness peak, and then driving the mechanism to the depth range of the next target to search again.

[0003] This traditional method has the following significant technical drawbacks:

[0004] First, the detection efficiency is low. Whenever there are N targets at different depths in the field of view, N independent focus searches must be performed, and the total time is the sum of the time spent on each focus search, which severely restricts the throughput of the production line.

[0005] Second, significant mechanical errors occur. Frequent motor starts, stops, accelerations, decelerations, and reversals inevitably introduce mechanical backlash errors. The micro-vibrations caused by each start and stop require time to decay, not only increasing time overhead but also reducing the physical extraction accuracy of the focal point. In solutions using liquid lenses, frequent voltage reversals can also induce thin-film oscillations at the fluid interface.

[0006] Third, there is a significant waste of computing power. Traditional methods typically process the entire image when calculating image sharpness, while the actual effective area of ​​interest only occupies a small portion of the field of view. When sharpness needs to be calculated separately for multiple target areas, the repeated calculation of the entire image results in a large waste of computing power, making it difficult to match the real-time output frame rate of high-speed cameras.

[0007] To address the aforementioned issues, some improvements have been proposed in the prior art. For example, multiple pre-focus points can be set to perform surface fitting to estimate the focal plane position of other fields of view. However, this method is only applicable to focal plane prediction between fields of view and cannot solve the problem of independent focusing on multiple targets at different depths within the same field of view. Summary of the Invention

[0008] This invention proposes a single-field-of-view multi-depth target focusing method and system for single-segment continuous focusing scanning, aiming to solve the technical problem of low detection efficiency caused by the need for multiple serial reciprocating focusing within a single fixed optical field of view when facing multiple targets with depth differences exceeding the depth of field range of a single imaging.

[0009] In a first aspect, the present invention provides a single-field-of-view multi-depth target focusing method for a single continuous focusing scan, comprising the following steps:

[0010] S1. Within the same fixed optical field of view of the imaging system, N target regions are selected, where N≥1; the N target regions have a depth difference along an axis parallel to the optical axis of the imaging system, and the maximum depth difference is greater than the single optical depth of field of the imaging system.

[0011] S2, control the imaging system to perform single, unidirectional continuous movement and focusing within a preset axial range, so that the focal plane of the imaging system continuously sweeps across the physical depth of N target areas.

[0012] S3, generate N pixel masks corresponding to N target regions; during the continuous movement of the focal plane, acquire M frame image sequences, M≥N;

[0013] S4, each frame of the acquired image is processed with N pixel masks to extract N local image data; the sharpness evaluation value of each of the local image data is calculated in multiple independent calculation channels to obtain N sharpness evaluation curves;

[0014] S5, extract the optimal focal plane position or optimal image frame corresponding to the peak value of each of the aforementioned sharpness evaluation curves, and output the local sharpest image of each target region.

[0015] Furthermore, the single, unidirectional continuous focusing movement in step S2 is achieved through any of the following methods:

[0016] Drive the entire optical lens or microscope objective to move unidirectionally along the optical axis;

[0017] Alternatively, keep the optical lens fixed and drive the image sensor to move unidirectionally along the optical axis;

[0018] Alternatively, a monotonically increasing or monotonically decreasing drive signal can be applied to the autofocus liquid lens in the optical path to change the curvature of the fluid interface inside the autofocus liquid lens.

[0019] Furthermore, the target region in step S1 has the shape of a polygon, a circle, an annulus, or an irregular closed curve; the N target regions are spatially distributed independently or partially overlap each other on the image plane.

[0020] Furthermore, the pixel mask in step S3 is a binary mask matrix, wherein pixels within the target area are marked as valid pixels, and pixels outside the target area are marked as background pixels.

[0021] Further, the operation in step S4 is a Boolean AND operation or a multiply-add operation; when calculating the sharpness evaluation value, the sharpness operator operation is performed only on the valid pixels marked by the pixel mask, and background pixels are skipped; the sharpness operator is a Laplacian operator, a Tenengrad gradient function, or a variance function.

[0022] Furthermore, the acquisition of the image sequence in step S3 is controlled by a hardware trigger signal: triggered by axial physical displacement at equal intervals, or triggered by equal time intervals, or triggered by a preset trigger sequence inside the camera; the axial physical displacement at equal intervals is realized based on the feedback pulse of the grating ruler or encoder.

[0023] Furthermore, the multiple independent computing channels in step S4 are multiple computing cores of a graphics processor, multiple parallel logic units of a field-programmable gate array, or multiple thread processing units of a central processing unit.

[0024] Furthermore, when extracting the peak value in step S5, curve fitting is performed on the discrete data points of the peak frame and its adjacent frames of the sharpness evaluation curve to calculate the theoretical optimal focal plane position at the sub-frame level.

[0025] The curve fitting uses Gaussian curve fitting:

[0026] ;

[0027] in This refers to the axial position of the focal plane. The sharpness rating is... To obtain the theoretically optimal focal plane position through fitting, Let be the standard deviation of the Gaussian distribution. For amplitude coefficient, For bias terms;

[0028] Alternatively, a quadratic parabola can be used for fitting:

[0029] ;

[0030] in , , The fitting coefficients are given, and the theoretical optimal focal plane position is... .

[0031] Furthermore, it also includes step S6: based on the optimal image frames of each target region extracted in step S5 and the pixel masks generated in step S3, the effective pixels of each target region in its corresponding optimal image frame are extracted, and pixel-level stitching is performed according to the spatial coordinates of each pixel mask to generate a composite panoramic depth image.

[0032] In a second aspect, the present invention provides a single-field-of-view multi-depth target focusing system for single-segment continuous focusing scanning, the system being used to perform the method, comprising:

[0033] The image acquisition module is used to acquire image sequences of the object under test within the same fixed optical field of view;

[0034] The region planning module is used to set N target regions with depth differences in the field of view and generate N corresponding independent pixel masks;

[0035] The continuous focusing execution module is used to control the focal plane to perform a single, unidirectional continuous displacement within a preset depth range;

[0036] The parallel processing module is used to synchronously receive the image sequence during the operation of the continuous focusing execution module, calculate the sharpness curve of each target region in parallel based on the pixel mask, and extract the extreme values.

[0037] The image output module is used to extract the clearest local image of each target region based on the extreme values.

[0038] The technical effects of the disclosed technical solution are as follows:

[0039] First, by using a single, unidirectional continuous focusing scan instead of multiple serial reciprocating searches, the focusing time is constant at the time of a single scan and is independent of the number of targets, fundamentally breaking through the efficiency bottleneck of optical depth of field for simultaneous detection of targets at multiple depths.

[0040] Secondly, by adopting a single, unidirectional, continuous, and smooth focusing motion, frequent acceleration and deceleration and mechanical start-stop during the reciprocating optimization process are avoided, eliminating mechanical return errors and liquid film oscillations, thus improving the accuracy of focal point extraction.

[0041] Third, since independent pixel masks corresponding to each target area are generated and calculations are performed only on the effective pixels covered by the mask during the sharpness calculation, the computational load per frame is compressed from the full image size to the sum of the areas of each target area. Combined with parallel processing of multiple independent calculation channels, the sharpness calculation of multiple targets can match the real-time output frame rate of high-speed cameras.

[0042] Fourth, by performing Gaussian or parabolic fitting on the discrete sharpness evaluation curve during peak extraction, the theoretical optimal focal plane position at the sub-frame level can be obtained by interpolation between discrete sampling frames, which significantly improves the focal plane positioning accuracy.

[0043] Fifth, by stitching together effective pixels according to spatial coordinates based on the best frame and pixel mask of each target region, a clear composite panoramic depth image with multiple depth features can be generated without the need for complex global optimization algorithms. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating the single-field-of-view, multi-depth target focusing method for single-segment continuous focusing scanning proposed in an embodiment of the present invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] To address the technical problem mentioned in the background art, where multiple serial repetitive focusing is required within a single fixed optical field of view to detect multiple targets with depth differences exceeding the depth of field of a single imaging operation, resulting in low detection efficiency, this invention provides a single-field-of-view multi-depth target focusing method using continuous focusing scanning. This embodiment takes the detection scenario of advanced semiconductor packaging as an example. The sample under test is an advanced packaging module, and under the same high-magnification microscopic field of view, there are three key structures with extremely large height differences: the bottom substrate trace (depth defined as 0 μm), the middle silicon interposer (depth + 80 μm), and the top copper pillar solder joint (depth + 150 μm). The imaging system uses a 5x microscope objective, with a single optical depth of field of 20 μm, which is much smaller than the maximum depth difference of 150 μm.

[0047] The imaging system includes: an industrial area scan camera (global shutter image sensor), a 5x microscope objective, a Z-axis precision positioning platform (stepper motor with closed-loop control of a grating ruler, the grating ruler resolution being better than 0.1μm), and a host computer (including an image acquisition card and a multi-core central processing unit). The camera is connected to the host computer via Camera Link or USB 3.0 interface, and the Z-axis platform driver is connected to the host computer via RS232 or Ethernet. The differential signal from the grating ruler is simultaneously fed back to both the driver and the host computer.

[0048] In addition to being used for closed-loop position control of the motor, the feedback pulses from the grating ruler are also introduced into the external trigger input port of the camera. The grating ruler outputs a TTL level pulse every 10μm of axial displacement, which is directly used as the frame trigger signal for the camera. The host computer sends the target position and uniform motion speed commands to the driver through the motion control card, and the driver drives the motor to move at a uniform speed under closed-loop control.

[0049] If the camera supports the internal trigger sequence function, the inter-frame delay parameter sequence can be predefined inside the camera. The host computer or slave computer control board only needs to send a single external trigger start signal, and the camera will automatically complete the acquisition of multiple frames according to the predefined delay sequence, without the need for external hardware to generate an independent trigger pulse for each frame.

[0050] The driving target of the Z-axis positioning platform is to move the entire microscope objective and camera assembly along the optical axis. In other implementation scenarios: if the optical lens is a large, heavy-duty telecentric lens that is inconvenient to move, the lens can be kept fixed, and the image sensor can be mounted on a piezoelectric nano-positioning stage. By driving the sensor to move along the optical axis, the image distance can be changed, which is equivalent to continuous movement of the object-side focal plane. If the system optical path integrates an autofocus liquid lens, all mechanical parts can be kept stationary, and a monotonically increasing analog voltage command (such as a constant increase from 35V to 55V) can be sent to the host computer's analog-to-digital converter card to drive the continuous change of the curvature of the fluid interface inside the liquid lens, thereby achieving optical zoom scanning without mechanical movement.

[0051] refer to Figure 1 Specifically, it includes the following steps:

[0052] Step S1: Single-field-of-view, multi-depth target region selection. Within the same fixed optical field of view of the imaging system, N=3 target regions are selected on a preview image captured at the initial focusing position: Region A is the bottom substrate trace area (rectangular, depth 0μm), Region B is the middle silicon interposer area (ring-shaped, depth +80μm), and Region C is the top copper pillar solder joint area (circular, depth +150μm). The three target regions are arranged sequentially along the Z-axis. The maximum depth difference between Region A and Region C is 150μm, significantly greater than the 20μm depth of field. The spatial coordinates of the three regions on the image plane are as follows: Region A is located in the upper left quadrant, Region B is located slightly to the right of the center, and Region C is located in the lower right quadrant. They are independent and do not overlap. In other implementation scenarios, each target region can be of any shape and partial overlap is allowed.

[0053] Step S2, single continuous focusing scan.

[0054] The host computer sends control commands to the Z-axis platform driver via the motion control card: the starting Z-axis position is -20μm (20μm below the depth of region A), the ending position is +180μm (30μm above the depth of region C), and the total travel is 200μm; the platform's movement speed is 250μm / s, and the acceleration is 500μm / s² (only used for acceleration and deceleration during the initial and final stages; the majority of the movement is at a constant speed). The driver drives the motor to move the microscope objective from the starting position to the ending position in a single, unidirectional, continuous, and uniform motion.

[0055] The grating ruler outputs a TTL level pulse every 10μm of axial displacement, simultaneously serving two purposes: one is fed into the motor driver as a position correction reference, and the other is sent directly to the camera's external trigger input port as a hardware trigger signal. The camera is set to rising edge trigger acquisition mode. Theoretically, 20 triggers are performed within a total travel of 200μm, with actual acquisition of M=20 frames (from -20μm to +180μm, uniformly sampled at 10μm intervals).

[0056] Key motion parameters: The host computer commands do not include any mid-course reversal or pause commands. The driver speed loop and position loop parameters are tuned to ensure that the actual speed fluctuation is less than ±1% of the set value. Acceleration and deceleration segments exist only within a very short distance (approximately 2μm) at the beginning and end, with a constant speed segment of 196μm in between. This motion mode fundamentally avoids the mechanical backlash and micro-vibration caused by multiple starts and stops and repeated reversals in traditional hill climbing methods.

[0057] In addition to the aforementioned equal-interval triggering, image acquisition can also be triggered at equal time intervals (using the lower-level computer's clock signal to trigger the camera at fixed time intervals), or by acquiring images according to a preset trigger sequence within the camera. In the preset trigger sequence mode, after receiving a single external trigger signal, the camera automatically completes continuous multi-frame acquisition according to a predefined inter-frame delay sequence (e.g., alternating between 3ms, 5ms, 3ms, 5ms, etc.), reducing the requirements for the real-time response capability of the upper-level computer.

[0058] Step S3: Mask generation and image sequence acquisition.

[0059] After selecting three target regions in step S1, the host computer region planning module generates three independent binary pixel mask matrices in memory, denoted as follows: , , The size is consistent with the resolution of the image sensor (e.g., 2048×2048 pixels). In the diagram, the pixel positions inside region A (rectangle) are marked as logic 1 (valid pixels), and the rest are marked as logic 0 (background pixels). In the diagram, only the area inside the annular region is marked as 1. In the diagram, only the area inside the circular region is marked as 1. If there are overlapping regions (such as a macroscopic observation area covering a microscopic local area), multiple masks within the overlapping region are all marked as 1, and subsequent calculations are performed independently.

[0060] During the continuous movement of the focal plane, the camera acquires images in 10μm increments based on the hardware trigger signal of the grating ruler. Starting from Z=-20μm, frame indices are acquired sequentially: t=1 (Z=-20μm), t=2 (Z=-10μm), t=3 (Z=0μm, region A is roughly in focus)... up to t=20 (Z=+180μm). All image frames are temporarily stored in a circular buffer in the host computer's memory in the order of acquisition, and are processed in a pipelined manner for subsequent processing.

[0061] Step S4, resolution calculation.

[0062] The core of this step lies in the parallel computation of local sharpness under mask constraints. After receiving each frame of image, the host computer's parallel processing module (such as an 8-core CPU or a graphics processor equipped with thousands of computing cores) immediately starts multiple independent computing channels.

[0063] With the first Frame Image Taking a size of H×W pixels as an example, for three target areas, the parallel processing module simultaneously starts three independent computing threads (or three computing core groups in the graphics processor):

[0064] Thread 1: Will and Perform pixel-level multiplication and addition operations to obtain local image data. , The pixel value corresponding to position 0 is set to zero, and only the effective pixel at position 1 is retained. Then, the sharpness operator is applied within the effective pixel area. Region A is the substrate trace, and the texture is mainly composed of line edges, so the Tenengrad gradient function is selected.

[0065] The Tenengrad function is defined as follows: It calculates the horizontal and vertical gradients by convolving the image with the Sobel operator.

[0066] ;

[0067] ;

[0068] ;

[0069] ;

[0070] The frame is Tenengrad sharpness rating within the effective area:

[0071] ;

[0072] in for The set of all valid pixel locations marked as 1. The summation range is strictly limited to... Inside the mask, pixels outside the mask are not involved in the calculation and are skipped directly. Region A has an area of ​​approximately 100,000 pixels, and the entire image has approximately 4 million pixels. The computational cost per frame is only 2.5% of that of the entire image.

[0073] Thread 2: and Performing multiplication and addition yields Region B is a ring-shaped area of ​​the silicon interposer, with a smooth surface texture but exhibiting subtle periodic structures; a variance function is chosen for this region. Effective pixel area. Mean grayscale value of inner pixels:

[0074] ;

[0075] in for Total number of effective pixels. Variance sharpness evaluation value:

[0076] ;

[0077] Thread 3: Will and Performing multiplication and addition yields .area For the circular region of the copper pillar solder joint with a distinct circular boundary, the Laplacian operator is chosen. The Laplacian operator's response to the second derivative of the image produces a strong zero-crossing response at the edges, making it suitable for detecting the sharpness of the solder joint boundary. Discretized convolution kernel:

[0078] ;

[0079] The sharpness evaluation value is the sum of the absolute values ​​of the Laplacian response within the effective pixel area:

[0080] ;

[0081] The three threads operate in complete parallelism with no data dependencies. After the first frame is calculated, they are output separately. , , As the scan continues, the host computer will sequentially... , , Store the data in three independent arrays to form three separate entries for each frame number. A sharpness evaluation curve showing changes in (or Z-axis position).

[0082] The parallel computing architecture described above is also applicable to graphics processing units (GPUs) and field-programmable gate arrays (FPGAs): each computing core on a GPU can simultaneously handle computational tasks for different masks; different logic blocks on an FPGA can execute convolution and accumulation operations for each mask in a parallel pipeline. If hardware resources are limited, each region can also be computed sequentially in a serial manner. Since only the mask region is computed, the total computational burden of the serial approach is still much smaller than that of traditional full-image processing.

[0083] Step S5: Extreme value extraction and subframe-level depth fitting. After scanning, three sharpness evaluation curves (t=1,…,20) are obtained, with the Z-axis position mapped to the frame number as follows: Z(t) = -20 + (t-1)×10 μm.

[0084] The peak of curve A is approximately at t=3 (Z=0μm). Since the discrete sampling interval is 10μm, the actual focal point may be located between two sampling frames (e.g., Z=+5μm). Therefore, only the discrete peak frame is taken with a precision of 10μm. To this end, curve fitting is performed on the discrete data points of the peak frame and its adjacent frames (t=2, 3, 4, corresponding to Z=-10μm, 0μm, +10μm) to calculate the theoretically optimal focal plane position at the sub-frame level.

[0085] Gaussian curve fitting was used:

[0086] ;

[0087] in The axial position of the focal plane (independent variable). For clarity rating (dependent variable). This is the theoretically optimal focal plane position (the axis of symmetry of the Gaussian curve). Standard deviation For amplitude coefficient, This is a bias term.

[0088] The least squares method is used. Select the peak frame and one to two frames of data to its left and right (e.g., t=2, 3, 4), substitute them into the above formula, and transform the Gaussian function into a quadratic function through logarithmic transformation:

[0089] ;

[0090] Expand on the topic The quadratic polynomial:

[0091] ;

[0092] set up , , Transform into linear regression:

[0093] ;

[0094] in , , The regression coefficients are obtained as follows: .

[0095] The alternative solution uses a quadratic parabola fit:

[0096] ;

[0097] The vertex of the parabola (the theoretically optimal focal plane) is = -B / (2A). Parabolic fitting is suitable for scenarios where the sharpness curve is approximately symmetrical near the peak and the dispersion step size is uniform.

[0098] After Gaussian fitting, the theoretical optimal focal plane for region A is... = +0.3μm (relative to the nominal depth of 0μm, the error is 0.3μm, much smaller than the 10μm sampling interval). Similarly, = +80.2μm, = +149.7μm. The submicron-level Z-axis physical coordinates obtained from the fitting can be directly output to the subsequent three-dimensional phase shift measurement algorithm as accurate Z-axis initial prior data.

[0099] At the same time, the best image frame for each target region is locked according to the peak frame number (or fitted μ value): region A is t=3 (Z≈0μm), region B is t=11 (Z≈80μm), and region C is t=18 (Z≈150μm). The local image data of the corresponding target regions are extracted and cached independently.

[0100] In the preferred embodiment, the image output module generates a composite panoramic depth image based on the optimal frame number and corresponding pixel mask for each target region.

[0101] Create a blank composite image of the same size as the original image. All pixels are initialized to 0.

[0102] Region A: Read 3 frames of image data at time t=3. Extract effective pixels ,according to Assigning effective pixel space coordinates to Corresponding position.

[0103] Region B: Read frames t=11. ,extract Assigned to Corresponding positions (if the coordinates do not overlap, they are directly covered; if they overlap, they can be averaged by weight or the maximum sharpness value can be taken. In this embodiment, there is no overlap).

[0104] Region C: Read images at t=18 frames ,extract Assigned to Corresponding position.

[0105] final The image can simultaneously contain clearly focused bottom substrate traces, middle silicon interposers, and top copper pillar solder joints, without relying on complex global wavelet transforms or Laplacian pyramid fusion in traditional extended depth-of-field algorithms. It avoids a large amount of full-image iterative calculations and edge artifacts, and can be generated in milliseconds.

[0106] In this embodiment, the total scanning distance is 200μm, the scanning speed is 250μm / s, and including the acceleration and deceleration time at both ends (approximately 4ms, negligible), the total focusing time is 0.8 seconds. Traditional serial search methods perform independent hill-climbing searches for each of the three targets, taking approximately 0.8 seconds per search, for a total of 2.4 seconds for three searches, resulting in a 3-fold efficiency improvement. If there are five targets in the field of view, the traditional method requires 4.0 seconds, while this invention still takes 0.8 seconds, representing a 5-fold efficiency improvement.

[0107] The Z-axis platform performs only one unidirectional continuous uniform motion throughout the entire focusing process. Acceleration exists only at the beginning and end moments, while the speed remains constant in between. This fundamentally avoids the mechanical backlash (±1~2μm) and micro-vibration waiting (approximately 50ms decay for each vibration, totaling 150ms for three vibrations) caused by multiple start-stop and direction-changing operations. Equally spaced grating ruler hardware triggers ensure sampling position accuracy better than 0.1μm, and combined with Gaussian fitting, sub-micron level focal plane positioning accuracy is achieved.

[0108] The total effective pixel area of ​​the three target regions is approximately 7.5% of the entire image (Area A 2.5% + Area B 2.5% + Area C 2.5%). The computational cost per frame is only 7.5% of that of the entire image. With the help of three-thread parallel processing, the time for single-frame sharpness calculation is controlled within 3ms. Matching the camera's highest acquisition frame rate (approximately 333 frames / second, frame interval 3ms), it achieves real-time pipeline processing of simultaneous acquisition, calculation, and extraction. All focus results are obtained instantly after acquisition ends.

[0109] This invention can be widely applied in technical fields such as advanced semiconductor packaging inspection, microscopic imaging, industrial vision inspection, medical imaging, and automated optical inspection, where multiple targets at different depths need to be focused on within the same field of view. Focusing efficiency increases proportionally with the number of targets, demonstrating high industrial practical value and broad market application prospects.

[0110] Based on the same inventive concept, embodiments of the present invention also provide a single-field-of-view multi-depth target focusing system for single-segment continuous focusing scanning, the system being used to execute the method, including:

[0111] The image acquisition module is used to acquire image sequences of the object under test within the same fixed optical field of view;

[0112] The region planning module is used to set N target regions with depth differences in the field of view and generate N corresponding independent pixel masks;

[0113] The continuous focusing execution module is used to control the focal plane to perform a single, unidirectional continuous displacement within a preset depth range;

[0114] The parallel processing module is used to synchronously receive the image sequence during the operation of the continuous focusing execution module, calculate the sharpness curve of each target region in parallel based on the pixel mask, and extract the extreme values.

[0115] The image output module is used to output the best focal plane image of each target region based on the extreme value extraction results.

[0116] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A single-field-of-view, multi-depth target focusing method using a single continuous focusing scan, characterized in that, Includes the following steps: S1, within the same fixed optical field of view of the imaging system, select N target regions, N≥1; the N target regions have a depth difference along an axis parallel to the optical axis of the imaging system, and the maximum depth difference is greater than the single optical depth of field of the imaging system. S2, control the imaging system to perform single, unidirectional continuous movement and focusing within a preset axial range, so that the focal plane of the imaging system continuously sweeps across the physical depth of N target areas. S3, generate N pixel masks corresponding to N target regions; during the continuous movement of the focal plane, acquire M frame image sequences, M≥1; S4, each frame of the acquired image is processed with N pixel masks to extract N local image data; the sharpness evaluation value of each of the local image data is calculated in multiple independent calculation channels to obtain N sharpness evaluation curves; S5, extract the optimal focal plane position or optimal image frame corresponding to the peak value of each of the aforementioned sharpness evaluation curves, and output the local sharpest image of each target region.

2. The method according to claim 1, characterized in that, The single, unidirectional continuous focusing movement in step S2 is achieved through any of the following methods: Drive the entire optical lens or microscope objective to move unidirectionally along the optical axis; Alternatively, keep the optical lens fixed and drive the image sensor to move unidirectionally along the optical axis; Alternatively, a monotonically increasing or monotonically decreasing drive signal can be applied to the autofocus liquid lens in the optical path to change the curvature of the fluid interface inside the autofocus liquid lens.

3. The method according to claim 1, characterized in that, The target region in step S1 has the shape of a polygon, a circle, an annulus, or an irregular closed curve; the N target regions are spatially independent or partially overlapping in the image plane.

4. The method according to claim 1, characterized in that, The pixel mask mentioned in step S3 is a binary mask matrix, in which pixels within the target area are marked as valid pixels, and pixels outside the target area are marked as background pixels.

5. The method according to claim 1, characterized in that, The operation described in step S4 is a Boolean AND operation or a multiply-add operation; when calculating the sharpness evaluation value, the sharpness operator operation is performed only on the valid pixels marked by the pixel mask, and background pixels are skipped; the sharpness operator is a Laplacian operator, a Tenengrad gradient function, or a variance function.

6. The method according to claim 1, characterized in that, In step S3, the acquisition of the image sequence is controlled by a hardware trigger signal: triggered by axial physical displacement at equal intervals, or triggered by equal time intervals, or triggered by a preset trigger sequence inside the camera; the axial physical displacement at equal intervals is realized based on the feedback pulse of the grating ruler or encoder.

7. The method according to claim 1, characterized in that, The multiple independent computing channels in step S4 are multiple computing cores of a graphics processor, multiple parallel logic units of a field-programmable gate array, or multiple thread processing units of a central processing unit.

8. The method according to claim 1, characterized in that, When extracting the peak value in step S5, curve fitting is performed on the discrete data points of the peak frame and its adjacent frames of the sharpness evaluation curve to calculate the theoretical optimal focal plane position at the sub-frame level. The curve fitting uses Gaussian curve fitting: ; in This refers to the axial position of the focal plane. The sharpness rating is... To obtain the theoretically optimal focal plane position through fitting, Let be the standard deviation of the Gaussian distribution. For amplitude coefficient, For bias terms; Alternatively, a quadratic parabola can be used for fitting: ; in , , The fitting coefficients are given, and the theoretical optimal focal plane position is... .

9. The method according to claim 8, characterized in that, It also includes step S6: based on the optimal image frames of each target region extracted in step S5 and the pixel masks generated in step S3, the effective pixels of each target region in its corresponding optimal image frame are extracted, and pixel-level stitching is performed according to the spatial coordinates of each pixel mask to generate a composite panoramic depth image.

10. A single-field-of-view, multi-depth target focusing system for single-pass continuous focusing scanning, characterized in that, The system is used to perform the method according to any one of claims 1-9, comprising: The image acquisition module is used to acquire image sequences of the object under test within the same fixed optical field of view; The region planning module is used to set N target regions with depth differences in the field of view and generate N corresponding independent pixel masks; The continuous focusing execution module is used to control the focal plane to perform a single, unidirectional continuous displacement within a preset depth range; The parallel processing module is used to synchronously receive the image sequence during the operation of the continuous focusing execution module, calculate the sharpness curve of each target region in parallel based on the pixel mask, and extract the extreme values. The image output module is used to output the clearest local image of each target region based on the extreme value extraction results.