Chip positioning detection method and system based on image analysis
Through the joint frequency domain-space augmentation algorithm, multi-scale deformation feature fusion network, deep learning correction model and dynamic self-calibration system of geometric prior constraints, the defects of chip positioning detection technology in terms of robustness, adaptability and stability are solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510306336.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing chip positioning detection technology has significant defects in image preprocessing robustness, multi-scale deformation adaptability, physical accuracy guarantee and long-term system stability, and it is difficult to meet the requirements of rapid development.
The frequency domain-space combined enhancement algorithm, multi-scale deformation feature fusion network (MS-DFN), deep learning correction model with geometric prior constraints, and dynamic self-calibration system are used to improve detection accuracy, efficiency and industrial scenario adaptability.
Through these technical means, the accuracy and efficiency of chip positioning detection are improved, the adaptability to multi-scale deformation and physical constraints is improved, and the long-term operation stability of the system is improved.
Smart Images

Figure CN120219340A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a chip positioning detection method and system based on image analysis. Background Art
[0002] Currently, in the field of chip detection both internationally and domestically, visual servo positioning technology is both a hot topic and a difficult point. Although foreign countries have developed finished products such as chip detection machines and die bonders, it is very difficult for their detection positioning accuracy and positioning speed to meet the requirements of high-speed development.
[0003] In the field of semiconductor manufacturing and packaging in China, the accurate position detection of chips is the core link to ensure the quality of subsequent bonding, soldering, and packaging processes. Traditional chip positioning detection technologies mainly rely on machine vision and image processing algorithms, but there are the following key bottlenecks in practical applications:
[0004] (1) Insufficient robustness in the image preprocessing link
[0005] When local overexposure or shadow occlusion occurs on the chip surface due to the light reflection characteristics of the material, it is difficult for traditional methods to retain complete edge and texture information, resulting in incorrect subsequent positioning feature extraction. In addition, a single spatial domain filtering algorithm (such as median filtering) is prone to cause detail blurring and cannot balance the requirements of noise suppression and feature retention.
[0006] (2) Adaptability defects in multi-scale and deformation scenarios
[0007] Current mainstream methods use template matching (such as the normalized cross-correlation algorithm) or convolutional neural networks (CNNs) with fixed structures for feature matching. However, when the size of the chip changes due to packaging process differences (such as from 2mm×2mm to 10mm×10mm), thermal expansion deformation, or tilted placement, the calculation efficiency and accuracy of traditional template matching drop sharply, while a single-scale CNN model is difficult to effectively capture multi-scale features and has insufficient generalization ability for minor deformations.
[0008] (3) Conflict between positioning accuracy and physical constraints
[0009] Although deep learning models based on pure data-driven can improve the detection speed, they lack prior knowledge constraints on the geometric rules of chips. For example, physical characteristics such as the symmetry of chip pins and the topological relationship between the center point and the boundary are not effectively utilized, resulting in the model output may deviate from the actual physical position. In addition, existing sub-pixel edge detection algorithms (such as the gray moment method) are prone to drift errors under complex background interference and are difficult to meet the micron-level positioning requirements.
[0010] (4) Stability problems in the long-term operation of the system
[0011] Existing inspection systems rely on initial calibration parameters, but during long-term operation, industrial cameras will produce posture deviations due to mechanical vibration or temperature changes, causing the coordinate transformation matrix to gradually fail (the cumulative error can reach ±0.5 pixels). Existing solutions require regular shutdowns for calibration or manual intervention corrections, which seriously reduces production line efficiency and cannot meet the high-continuity process requirements of semiconductor manufacturing.
[0012] In summary, the existing chip positioning detection technology has significant defects in image preprocessing robustness, multi-scale deformation adaptability, physical accuracy assurance and system long-term stability. The present invention systematically solves the above problems through the frequency domain-spatial domain joint enhancement algorithm, multi-scale deformation feature fusion network (MS-DFN), deep learning correction model of geometric prior constraints and dynamic self-calibration system (DSCS), and improves detection accuracy, efficiency and industrial scene adaptability. Summary of the invention
[0013] The purpose of the present invention is to provide a chip positioning detection method and system based on image analysis, which solves the existing problems of insufficient detection accuracy and low detection efficiency through a frequency domain-spatial domain joint enhancement algorithm, a multi-scale deformation feature fusion network (MS-DFN), a deep learning correction model of geometric prior constraints, and a dynamic self-calibration system.
[0014] In order to solve the above technical problems, the present invention is achieved through the following technical solutions:
[0015] The present invention is a chip positioning detection method based on image analysis, comprising the following steps:
[0016] Step S1: The circular polarized light source adjusts the light on the workstation according to the illumination compensation template generated in real time by the illumination estimation model;
[0017] Step S2: Collect data information of the chip on the workstation through a high-resolution linear array CMOS camera and a single-line laser radar;
[0018] Step S3: obtaining characteristic corner points at the endpoints of the chip point cloud, and judging and identifying breakpoints therein;
[0019] Step S4: split the chip point cloud into several parts by breakpoints, and extract the point cloud line segment features by using a windowed straight line fitting method;
[0020] Step S5: construct a lightweight convolutional network, extract chip contours and solder joint features of different scales through dilated convolution, introduce a deformable convolution module, dynamically adjust the receptive field, and adaptively match the chip deformation;
[0021] Step S6: Input the extracted point cloud line segment features into the constructed lightweight convolutional network, embed the chip design rules at the output of the neural network, and correct the predicted coordinates through nonlinear optimization;
[0022] Step S7: Move the chip to the corrected predicted coordinates by the robotic arm.
[0023] As a preferred technical solution, in the step S1, the specific process of the illumination compensation template generated in real time according to the illumination estimation model is as follows:
[0024] Step S11: Decompose the input image into an illumination component and a reflection component, where the illumination component and the reflection component are the low-frequency component and the high-frequency component respectively;
[0025] S(x,y) = L(x,y) × R(x,y);
[0026] In the formula, L(x,y) represents the incident light image, the illumination component of the ambient light, and R(x,y) represents the reflection image. The incident light irradiates the reflecting object, and through the reflection of the object, the reflected light enters the camera and finally forms an image; L(x,y) represents the reflection image that the camera can receive.
[0027] Step S12: Convert the linear RGB space to the logarithmic domain to simplify the multiplication operation into an addition operation:
[0028] logS(x,y) = logL(x,y) + logR(x,y);
[0029] Step S13: Perform convolution calculation on the original image using a large-scale Gaussian kernel to obtain an initial illumination estimation, and use a separable Gaussian kernel to accelerate the calculation;
[0030] Step S14: Dynamically correct the illumination component; during correction, overexposure can be performed locally to fill the illumination mutation area, or the gradient information of the original image can be combined to retain the edge smoothness of the illumination component;
[0031] Step S15: Normalize the illumination component to [0,1] and apply adaptive gamma correction. The specific formula is as follows:
[0032]
[0033] In the formula, L compensated (x,y) is the image after gamma correction, and γ represents the dynamic adjustment factor of the overall brightness distribution of the image. γ < 1 in the dark area and γ > 1 in the bright area;
[0034] Step S16: Enhance the detail contrast by inversely deriving the reflection component. The formula is as follows:
[0035] R(x,y) = exp(logS(x,y) - logL compensated (x,y));
[0036] Step S17: Perform illumination estimation on the low-resolution image and then upsample it to the original size to accelerate the construction of the illumination compensation template;
[0037] Step S18: Based on the inter-frame difference method, judge the sudden change of the illumination at the workstation and trigger re-estimation. At the same time, the value of σ needs to be adjusted according to the complexity of the image content. The higher the complexity of the image content, the larger the value of σ.
[0038] As a preferred technical solution, in step S2, after the illumination compensation template adjusts the light of the workstation, preprocess the data collected by the high-resolution line array CMOS camera and the single-line lidar. The high-resolution line array CMOS camera obtains the point cloud data of the chip, and after two voxel filters to remove redundant data, the preprocessed chip point cloud data is obtained. And measure the initial position of the chip through the single-line lidar, and the initial position is used for scan matching of the point cloud data; among them, the first voxel filter divides the original point cloud by a voxel grid, and retains one representative point in each voxel; the second voxel filter reduces the voxel size and performs secondary downsampling on the local high-density area.
[0039] As a preferred technical solution, in step S2, when the single-line lidar determines that the chip is not properly placed, it means that the chip point cloud data collected by the high-resolution line array CMOS camera is tilted, then correct the chip point cloud data. The specific correction process is as follows:
[0040] Step S21: Perform pixel-level coordinate transformation on each image according to the camera calibration parameters;
[0041] Step S22: Map the image coordinate system to the object space coordinate system through the coordinate transformation matrix;
[0042] Step S23: Extract the corresponding feature points through the SIFT / SURF algorithm, use the RANSAC algorithm to remove the mismatched points, and establish the spatial relationship between high-precision images.
[0043] As a preferred technical solution, in step S3, split the chip point cloud into multiple parts, select a certain number of chip points with larger curvature as feature corner points for each part, and calculate the curvature of each part; the curvature calculation formula corresponding to the chip point cloud is as follows:
[0044]
[0045] In the formula, c is the curvature corresponding to the chip point cloud, S is the chip point cloud set, represents the spatial coordinate of the current i-th chip point;
[0046] When the curvature is greater than the threshold, the point cloud point is taken as a feature corner point; the break points of the feature corner points are distinguished by the distance difference between the feature corner points and the adjacent single-line lidar to segment the point cloud.
[0047] As a preferred technical solution, the break points divide the chip point cloud into several parts, and linear fitting is respectively performed on the chip point cloud to obtain candidate feature line segments, and the distance between the chip points and the candidate feature line segments is calculated. If it exceeds the set value, windowed linear fitting is performed; if it does not exceed the set value, the feature line segment is the final fitted line segment.
[0048] As a preferred technical solution, the windowed linear fitting obtains the linear parameters of the chip, and the specific process is as follows:
[0049] Step S41: Initialize the window size to n, and set the starting position of each piece of point cloud as P1 i , the end position is
[0050] Step S42: Determine whether the current window has traversed the entire point cloud set; if the traversal is completed, the program is terminated; otherwise, proceed to step S43;
[0051] Step S43: Perform LM linear fitting on the point cloud points under the window to generate a line L, and calculate the distance d from each point cloud point to L; if d is greater than the maximum distance, execute step S44; otherwise execute step S45;
[0052] Step S44: It means that a line cannot be successfully fitted under this window. Move the starting position of the window by +1, and at the same time initialize the window size, and return to step S42;
[0053] Step S45: It means that a line is successfully fitted under this window. Set the window size to n + 1, and calculate the distance d from the newly added point cloud point to the current L; if d is less than the maximum distance, repeat this step, otherwise enter step S46;
[0054] Step S46: Re-perform linear fitting on the point cloud point set under the window to generate L′, and calculate the distance d from the current point cloud point to L′; if d is greater than the maximum distance, enter step S47, otherwise return to step S45;
[0055] Step S47: Record the parameters of the line segment L, and at the same time move the starting position of the window to the original end position, initialize the window size, and return to step S42.
[0056] As a preferred technical solution, the windowed linear fitting obtains the linear parameters of the chip, and the specific process is as follows:
[0057] Step S41: Initialize the window size to n, and set the starting position of each piece of point cloud as P1i , the end position is
[0058] Step S42: Determine whether the current window has traversed the entire point cloud set; if the traversal is completed, terminate the program; otherwise, proceed to step S43;
[0059] Step S43: Perform LM linear fitting on the point cloud points under the window to generate a straight line L, and calculate the distance d from each point cloud point to L; if d is greater than the maximum distance, execute step S44; otherwise, execute step S45;
[0060] Step S44: It indicates that a straight line cannot be successfully fitted under this window. Increase the starting position of the window by 1, initialize the window size at the same time, and return to step S42;
[0061] Step S45: It indicates that a straight line has been successfully fitted under this window. Set the window size to n + 1, and calculate the distance d from the newly added point cloud point to the current L; if d is less than the maximum distance, repeat this step, otherwise enter step S46;
[0062] Step S46: Re-perform linear fitting on the point cloud point set under the window to generate L′, and calculate the distance d from the current point cloud point to L′; if d is greater than the maximum distance, enter step S47, otherwise return to step S45;
[0063] Step S47: Record the parameters of the line segment L, move the starting position of the window to the original end position at the same time, initialize the window size, and return to step S42.
[0064] As a preferred technical solution, in step S5, the specific implementation process is as follows:
[0065] Step S51: Scale the chip images of different sizes to a unified fixed size through adaptive scaling, retain the original aspect ratio and fill the edges to avoid distortion caused by direct scaling;
[0066] Step S52: Adopt three-branch parallel dilated convolution to extract the chip contour and solder joint features of different sizes, covering the cross-scale features of chip sizes from 2 mm to 10 mm, and avoid missing small targets by traditional single-branch networks;
[0067] The three branches include:
[0068] Branch 1: Dilatation rate = 1 (conventional convolution), receptive field = 3×3, capturing fine solder joints;
[0069] Branch 2: Dilatation rate = 3, receptive field = 7×7, extracting the main body contour of the chip;
[0070] Branch 3: Dilatation rate = 6, receptive field = 13×13, adapting to large-size packages;
[0071] Step S53: Predict the offset on the feature map, dynamically adjust the convolutional kernel sampling position, and adapt to the chip deformation;
[0072] When predicting the offset, dynamically adjust the convolutional kernel sampling position. The formula is:
[0073]
[0074] In the formula, y(p) is the value of the output feature map at position p, K is the convolutional kernel size, W k is the weight parameter of the k-th convolutional kernel, x(·) is the pixel value of the input feature map at the specified position, p is the central position coordinate of the currently calculated output feature map, p k is the fixed offset of the k-th sampling point in the standard convolution, and Δp k is the dynamic offset of the k-th sampling point in the deformable convolution;
[0075] Step S54: Apply L2 regularization to the offset Δp to prevent excessive deformation from deviating from physical rationality, and apply random affine transformation to the chip image during training to force the network to learn deformation invariance;
[0076] Step S55: Generate the channel weight vector through global average pooling and fully connected layers for cross-scale feature weighted fusion;
[0077] In Step S55, the weight generates the channel weight vector through global average pooling (GAP) and fully connected layers:
[0078] s = σ(W2 × δ(W1 × GAP(F)));
[0079] In the formula, s is the channel weight vector, W1 ∈ R C / r×C , W2 ∈ R C×C / r is the weight matrix of the fully connected layer, δ is the ReLU activation function, σ is the Sigmoid activation function, GAP(F) is the global average pooling of the input feature map F, and the compression ratio r = 16;
[0080] Multiply the weight vector and the original feature channel by channel:
[0081] F weighted = s × F;
[0082] The specific fusion formula is as follows:
[0083] F funsion = α × F low + β × F mid + χ × F high ;
[0084] In the formula, F weightedis the weighted feature, where α, β, and χ are dynamically generated by SE-Net, and F low , F mid , F high are the low-level feature map, the middle-level feature map, and the high-level feature map respectively, and F funsion is the fused feature map.
[0085] As a preferred technical solution, in step S6, the working process of the lightweight convolutional network is as follows:
[0086] Step S61: Use a lightweight U-Net or CenterNet, input the extracted point cloud line segment features, and output the key point heat map; the heat map is generated by using a Gaussian kernel function to label the real coordinates to generate a probability distribution heat map (such as a resolution of 512×512); the loss function uses Focal Loss + mean square error (MSE) of the heat map to enhance the sensitivity to small targets;
[0087] Step S62: Construct geometric constraint modeling, construct an objective function, and minimize the deviation between the predicted coordinates and the geometric rules;
[0088] The geometric prior rules are as follows:
[0089] Symmetry constraint: The pin pitch in the same row / column is equal;
[0090] Center point alignment: The connecting line of the centers of all pins needs to pass through the geometric center of the chip;
[0091] Angle constraint: The pin arrangement direction is parallel to the chip edge;
[0092] The calculation formula for minimizing the deviation between the predicted coordinates and the geometric rules is as follows:
[0093]
[0094] In the formula, Δx and Δy are the coordinate correction amounts, λ is the balance between the position and the angle error, and (x i,rule , y i,rule ) is the theoretical coordinate under the geometric rule constraint;
[0095] Step S63: Perform sub-pixel edge point calculation and B-spline curve fitting;
[0096] The calculation formula for sub-pixel edge points is as follows:
[0097]
[0098] In the formula, θ is the edge direction angle, Z 11 , Z 20 are the real parts of the Zernike moments; ρ is the sub-pixel offset, and R is the ROI radius;
[0099] When fitting the B-spline curve, optimize the control point p i , so that the curve approximates the sub-pixel points. The specific formula is as follows:
[0100]
[0101] In the formula, p i is the control point of the B-spline curve, M is the number of data points, B(t j ) is the coordinate of the B-spline curve at t j , and (x j , y j ) is the coordinate of the measured edge point;
[0102] Step S64: Fuse the deep learning prediction, geometric correction, and sub-pixel fitting results according to the confidence level.
[0103] The present invention is a chip positioning and detection system based on image analysis, including an operation station, an optical acquisition module located above the station, an industrial computer, and a robotic arm;
[0104] The optical acquisition module includes a high-resolution linear array CMOS camera, a single-line lidar, and a circularly polarized light source; the high-resolution linear array CMOS camera is installed on an inclined bracket for image acquisition of the chip; the single-line lidar is used to measure the distance to any point on the chip; the circularly polarized light source is used to provide the light environment for shooting and suppress the metal reflection on the chip surface;
[0105] The industrial computer includes a front-end FPGA and a back-end GPU server; the front-end FPGA is used to perform image preprocessing and feature extraction in real time; the back-end GPU server is built-in with a lightweight convolutional network, a chip design rule customization module, moment sub-pixel edge detection, and a dynamic self-calibration system; the multi-scale deformation feature fusion network is used for; the lightweight convolutional network is used to extract chip contour and solder joint features of different scales through dilated convolution; the chip design rule customization module is used to formulate the design rules of the chip; the moment sub-pixel edge detection is used to locate the chip through the Zernike moment sub-pixel edge detection algorithm in combination with B-spline curve fitting;
[0106] The robotic arm is used to move the chip to the corrected predicted coordinates.
[0107] The present invention has the following beneficial effects:
[0108] (1) The present invention improves the detection accuracy, efficiency, and industrial scenario adaptability through the frequency domain-spatial domain joint enhancement algorithm, the multi-scale deformation feature fusion network, the deep learning correction model with geometric prior constraints, and the dynamic self-calibration system;
[0109] (2) The present invention adjusts the light on the workstation through the light compensation template generated in real time by the circularly polarized light source according to the light estimation model, and can intelligently adjust the brightness of the circularly polarized light source according to the actual light environment on the workstation, improving the acquisition efficiency of the high-resolution linear array CMOS camera and the clarity of the acquired images.
[0110] (3) Through the cascaded design of dilated convolution and deformable convolution, the present invention uses dilated convolution to solve the scale difference and deformable convolution to compensate for geometric deformation, forming a complementary dynamic feature fusion mechanism, avoiding manual setting of multi-scale weights, and realizing adaptive fusion lightweight deployment optimization through SE-Net.
[0111] (4) By complementing data-driven and physical rules, the present invention solves the accuracy ceiling of pure learning models, reduces the dependence on the amount of training data using geometric priors, adapts to chips of different package types, and the optimization process conforms to industrial design rules, facilitating debugging and acceptance by engineers.
[0112] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0114] Figure 1 It is a flowchart of a chip positioning and detection method based on image analysis according to the present invention;
[0115] Figure 2 It is a schematic structural diagram of a chip positioning and detection system based on image analysis according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0116] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0117] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0118] In order to make the purpose, technical solutions and advantages of the present application clearer, the following is combined with the attached Figure 1-2Examples are provided to further elaborate on this application. It should be understood that the specific examples described herein are for the purpose of explaining this application and not for limiting it.
[0119] Example 1
[0120] Please refer to Figure 1 As shown, the present invention is a chip positioning detection method based on image analysis, including the following steps:
[0121] Step S1: The circularly polarized light source adjusts the light on the workstation according to the light compensation template generated in real time by the light estimation model;
[0122] Step S2: Collect data information of the chip on the workstation through a high-resolution linear array CMOS camera and a single-line lidar;
[0123] Step S3: Obtain the characteristic corner points among the end points of the chip point cloud, and judge and identify the break points among the characteristic corner points;
[0124] Step S4: The break points divide the chip point cloud into several parts, and the windowed linear fitting method is used to extract the point cloud line segment features;
[0125] Step S5: Construct a lightweight convolutional network, extract the chip contour and solder joint features of different scales through dilated convolution, introduce a deformable convolution module, dynamically adjust the receptive field, and adaptively match the chip deformation;
[0126] Step S6: Input the extracted point cloud line segment features into the constructed lightweight convolutional network, embed the chip design rules at the output end of the neural network, and correct the predicted coordinates through non-linear optimization;
[0127] Step S7: Move the chip to the corrected predicted coordinates through a robotic arm.
[0128] In step S1, the specific process of the light compensation template generated in real time according to the light estimation model is as follows:
[0129] Step S11: Decompose the input image into a light component and a reflection component, and the light component and the reflection component are the low-frequency component and the high-frequency component;
[0130] S(x,y) = L(x,y) × R(x,y);
[0131] In the formula, L(x,y) represents the incident light image, the illumination component of the ambient light, and R(x,y) represents the reflection image. The incident light irradiates on the reflecting object, and through the reflection of the object, the reflected light enters the camera and finally forms an image; L(x,y) represents the reflection image that the camera can receive.
[0132] Step S12: Convert the linear RGB space to the logarithmic domain to simplify multiplication to an addition operation:
[0133] logS(x,y) = logL(x,y) + logR(x,y);
[0134] Step S13: Perform convolution calculation on the original image using a large-scale Gaussian kernel to obtain an initial illumination estimate, and use a separable Gaussian kernel to accelerate the calculation;
[0135] Step S14: Dynamically correct the illumination component; during correction, overexposure can be performed locally to fill the illumination mutation area, or the gradient information of the original image can be combined to retain the edge smoothness of the illumination component;
[0136] Step S15: Normalize the illumination component to [0,1] and apply adaptive gamma correction. The specific formula is as follows:
[0137]
[0138] In the formula, L compensated (x,y) is the image after gamma correction, γ represents the dynamic adjustment factor of the overall brightness distribution of the image, γ < 1 in the dark area, and γ > 1 in the bright area;
[0139] Step S16: Derive the reflection component by reverse deduction to enhance the detail contrast. The formula is as follows:
[0140] R(x,y) = exp(logS(x,y) - logL compensated (x,y));
[0141] Step S17: Perform illumination estimation on the low-resolution image and then upsample it to the original size to accelerate the construction of the illumination compensation template;
[0142] Step S18: Judge the illumination mutation of the workstation based on the inter-frame difference method and trigger re-estimation. At the same time, the σ value needs to be adjusted according to the complexity of the image content. The higher the complexity of the image content, the larger the σ value.
[0143] In Step S2, after the illumination compensation template adjusts the light of the workstation, preprocess the data collected by the high-resolution line array CMOS camera and the single-line lidar. The high-resolution line array CMOS camera obtains the point cloud data of the chip, and after two voxel filters to remove redundant data, the preprocessed chip point cloud data is obtained. And measure the initial position of the chip through the single-line lidar. The initial position is used for scan matching of the point cloud data;
[0144] The first voxel filter is through 0.2mm 3The voxel grid divides the original point cloud in space, and a representative point is retained within each voxel. The first voxel filtering can quickly remove redundant data in dense areas, reducing the data volume while retaining the overall morphological features of the chip surface. When setting it, attention should be paid to:
[0145] The voxel size should be slightly smaller than the size of the key structures on the chip surface (such as pin pitch) to avoid losing details due to excessive merging. If there is obvious noise in the point cloud, statistical filtering within the voxel (such as removing points with an average distance exceeding the threshold) can be combined for synchronous denoising.
[0146] In the second voxel filtering, the voxel size is reduced to less than 0.05 mm 3 The voxel grid performs secondary downsampling on local high-density areas, further optimizing based on the point cloud after the first filtering. The second voxel filtering mainly performs secondary downsampling on local high-density areas (chip edges) to eliminate microscopic redundancy caused by too high sensor accuracy.
[0147] In step S2, when the single-line lidar determines that the chip is not properly positioned, it means that the high-resolution line-array CMOS camera captures tilted chip point cloud data. Then, the point cloud data of the chip is corrected. The specific correction process is as follows:
[0148] Step S21: According to the camera calibration parameters (such as focal length, principal point offset, radial / tangential distortion coefficients), perform pixel-level coordinate transformation on each image to eliminate barrel or pincushion distortion;
[0149] Step S22: Map the image coordinate system to the object space coordinate system through a coordinate transformation matrix (including attitude angle parameters);
[0150] Step S23: Extract corresponding feature points through the SIFT / SURF algorithm, use the RANSAC algorithm to eliminate mismatched points, and establish the spatial relationship between high-precision images.
[0151] In step S3, the chip point cloud is split into multiple parts. For each part, a certain number of chip points with a relatively large curvature (greater than 90°) are selected as feature corner points, and the curvature of each part is calculated. The curvature calculation formula corresponding to the chip point cloud is as follows:
[0152]
[0153] In the formula, c is the curvature corresponding to the chip point cloud, S is the set of chip point clouds, represents the spatial coordinate of the current i-th chip point;
[0154] When the curvature is greater than the threshold, the point cloud point is taken as a feature corner point. The point cloud is segmented by the distance difference between the feature corner point and the adjacent single-line lidar to distinguish the breakpoints of the feature corner points.
[0155] The breakpoint divides the chip point cloud into several parts, respectively performs linear fitting on the chip point cloud to obtain candidate feature line segments, calculates the distance between the chip points and the candidate feature line segments, and if it exceeds the set value, windowed linear fitting is performed; if it does not exceed the set value, the feature line segment is the final fitted line segment.
[0156] In step S5, the specific implementation process is as follows:
[0157] Step S51: Scale the chip images of different sizes into a unified fixed size through adaptive size scaling, retain the original aspect ratio and fill the edges to avoid distortion caused by direct scaling.
[0158] Step S52: Adopt three-branch parallel atrous convolution to extract the chip contour and solder joint features of different sizes, covering the cross-scale features of chip sizes from 2mm to 10mm, and avoid missing small targets by traditional single-branch networks.
[0159] The three branches include:
[0160] Branch 1: Atrous rate = 1 (conventional convolution), receptive field = 3×3, capturing fine solder joints.
[0161] Branch 2: Atrous rate = 3, receptive field = 7×7, extracting the main body contour of the chip.
[0162] Branch 3: Atrous rate = 6, receptive field = 13×13, adapting to large-size packages.
[0163] Step S53: Predict the offset on the feature map, dynamically adjust the sampling position of the convolution kernel, and adapt to chip deformation.
[0164] When predicting the offset, the sampling position of the convolution kernel is dynamically adjusted, and the formula is:
[0165]
[0166] In the formula, y(p) is the value of the output feature map at position p, K is the size of the convolution kernel, W k is the weight parameter of the k-th convolution kernel, x(·) is the pixel value of the input feature map at the specified position, p is the center position coordinate of the currently calculated output feature map, p k is the fixed offset of the k-th sampling point in the standard convolution, and Δp k is the dynamic offset of the k-th sampling point in the deformable convolution.
[0167] Step S54: Apply L2 regularization to the offset Δp to prevent excessive deformation from deviating from physical rationality, and apply random affine transformation to the chip images during training to force the network to learn deformation invariance.
[0168] Step S55: Generate a channel weight vector through global average pooling and a fully connected layer, and perform cross-scale feature weighted fusion;
[0169] In step S55, the weights generate a channel weight vector through global average pooling (GAP) and a fully connected layer:
[0170] s = σ(W2 × δ(W1 × GAP(F)));
[0171] In the formula, s is the channel weight vector, W1 ∈ R C / r×C , W2 ∈ R C×C / r is the weight matrix of the fully connected layer, δ is the ReLU activation function, σ is the Sigmoid activation function, GAP(F) is the global average pooling of the input feature map F, and the compression ratio r = 16;
[0172] Multiply the weight vector and the original features channel by channel:
[0173] F weighted = s × F;
[0174] The specific fusion formula is as follows:
[0175] F funsion = α × F low + β × F mid + χ × F high ;
[0176] In the formula, F weighted is the weighted feature, α, β, χ are dynamically generated by SE-Net, F low , F mid , F high are the low-level feature map, the middle-level feature map, and the high-level feature map respectively, and F funsion is the fused feature map.
[0177] In step S6, the working process of the lightweight convolutional network is as follows:
[0178] Step S61: Adopt a lightweight U-Net or CenterNet, input the extracted point cloud line segment features, and output a key point heat map; the heat map is generated by using a Gaussian kernel function to label the real coordinates to generate a probability distribution heat map (such as a resolution of 512×512); the loss function uses Focal Loss + mean square error (MSE) of the heat map to enhance the sensitivity to small targets;
[0179] Step S62: Construct geometric constraint modeling, construct an objective function, and minimize the deviation between the predicted coordinates and the geometric rules;
[0180] The geometric prior rules are as follows:
[0181] Symmetry constraint: The pin pitch in the same row / column is equal;
[0182] Center point alignment: The center connection lines of all pins need to pass through the geometric center of the chip;
[0183] Angle constraint: The pin arrangement direction is parallel to the chip edge;
[0184] The formula for minimizing the deviation between the predicted coordinates and the geometric rules is as follows:
[0185]
[0186] In the formula, Δx and Δy are the coordinate correction amounts, λ is the balance position and angle error, and (x i,rule , y i,rule ) are the theoretical coordinates under the geometric rule constraints;
[0187] Step S63: Perform sub-pixel edge point calculation and B-spline curve fitting;
[0188] The formula for sub-pixel edge points is as follows:
[0189]
[0190] In the formula, θ is the edge direction angle, Z 11 , Z 20 are the real parts of the Zernike moments; ρ is the sub-pixel offset, and R is the ROI radius;
[0191] When performing B-spline curve fitting, optimize the control point p i , so that the curve approximates the sub-pixel points. The specific formula is as follows:
[0192]
[0193] In the formula, p i is the control point of the B-spline curve, M is the number of data points, B(t j ) is the coordinate of the B-spline curve at t j , and (x j , y j ) are the measured edge point coordinates;
[0194] Step S64: Fuse the deep learning prediction, geometric correction, and sub-pixel fitting results according to the confidence level.
[0195] Example 2
[0196] Refer to Figure 2 As shown, the present invention is a chip positioning and detection system based on image analysis, which can be used to execute the method content of Example 1 of the present invention, including: an operation station, an optical acquisition module located above the station, an industrial control computer, and a robotic arm;
[0197] The optical acquisition module includes a high-resolution linear array CMOS camera, a single-line lidar, and a circularly polarized light source; the high-resolution linear array CMOS camera is installed on an inclined bracket for image acquisition of the chip; the single-line lidar is used to measure the distance to any point on the chip; the circularly polarized light source is used to provide the light environment for shooting and suppress the metal reflection on the chip surface;
[0198] The industrial control computer includes a front-end FPGA and a rear-end GPU server; the front-end FPGA is used to perform image preprocessing and feature extraction in real time; the rear-end GPU server is built-in with a lightweight convolutional network, a chip design rule customization module, moment sub-pixel edge detection, and a dynamic self-calibration system; the multi-scale deformation feature fusion network is used for; the lightweight convolutional network is used to extract chip contour and solder joint features of different scales through dilated convolution; the chip design rule customization module is used to formulate the design rules of the chip; the moment sub-pixel edge detection is used to locate the chip through the Zernike moment sub-pixel edge detection algorithm and combined with B-spline curve fitting;
[0199] The robotic arm is used to move the chip to the corrected predicted coordinates.
[0200] It should be noted that in the above system embodiments, the various units included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for easy distinction from each other and do not limit the protection scope of the present invention.
[0201] In addition, those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium.
[0202] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not elaborate on all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A chip positioning detection method based on image analysis, characterized in that: The steps include: Step S1: The circular polarized light source adjusts the light on the workstation according to the illumination compensation template generated in real time by the illumination estimation model; Step S2: Collect data information of the chip on the workstation through a high-resolution linear array CMOS camera and a single-line laser radar; Step S3: obtaining characteristic corner points at the endpoints of the chip point cloud, and judging and identifying breakpoints therein; Step S4: split the chip point cloud into several parts by breakpoints, and extract the point cloud line segment features by using a windowed straight line fitting method; Step S5: construct a lightweight convolutional network, extract chip contours and solder joint features of different scales through dilated convolution, introduce a deformable convolution module, dynamically adjust the receptive field, and adaptively match the chip deformation; Step S6: Input the extracted point cloud line segment features into the constructed lightweight convolutional network, embed the chip design rules at the output of the neural network, and correct the predicted coordinates through nonlinear optimization; Step S7: Move the chip to the corrected predicted coordinates by the robotic arm.
2. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S1, the specific process of generating the illumination compensation template in real time according to the illumination estimation model is as follows: Step S11: decomposing the input image into an illumination component and a reflection component; Step S12: converting the linear RGB space to the logarithmic domain; Step S13: Use a large-scale Gaussian kernel to perform convolution calculation on the original image to obtain an initial illumination estimate; Step S14: dynamically correct the illumination component; Step S15: normalize the illumination component to [0, 1] and apply adaptive gamma correction; Step S16: deriving the reflected component by reverse derivation; Step S17: performing illumination estimation on the low-resolution image and then upsampling it to the original size; Step S18: Determine the sudden change of illumination at the workstation based on the inter-frame difference method, and trigger re-estimation.
3. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S2, after the illumination compensation template adjusts the light of the workstation, the data collected by the high-resolution linear array CMOS camera and the single-line laser radar are preprocessed, the high-resolution linear array CMOS camera obtains the point cloud data of the chip, and the redundant data is eliminated through two voxel filtering to obtain the preprocessed chip point cloud data, and the initial position of the chip is measured by the single-line laser radar, and the initial position is used for scanning and matching the point cloud data; wherein, the first voxel filtering spatially divides the original point cloud through the voxel grid, and retains a representative point in each voxel; the second voxel filtering reduces the voxel size and performs secondary downsampling on the local high-density area.
4. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S2, when the single-line laser radar determines that the chip is not aligned, it means that the chip point cloud data collected by the high-resolution linear array CMOS camera is tilted, and the chip point cloud data is corrected. The specific correction process is as follows: Step S21: performing pixel-level coordinate transformation on each image according to camera calibration parameters; Step S22: Mapping the image coordinate system to the object space coordinate system through a coordinate conversion matrix; Step S23: extracting feature points with the same name through SIFT / SURF algorithm, eliminating mismatched points using RANSAC algorithm, and establishing high-precision spatial relationship between images.
5. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S3, the chip point cloud is split into multiple parts, and the curvature of each part is calculated; when the curvature is greater than a threshold, the point cloud point is used as a feature corner point; The point cloud is segmented by distinguishing the breakpoints of feature corner points through the distance difference between the feature corner points and the adjacent single-line lidar.
6. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S4, the breakpoints divide the chip point cloud into several parts, and straight line fitting is performed on each of the chip point clouds to obtain candidate feature line segments, and the distance between the chip point and the candidate feature line segments is calculated. If it exceeds the set value, windowed straight line fitting is performed; if it does not exceed the set value, the feature line segment is the final fitting line segment.
7. The chip positioning detection method based on image analysis according to claim 6, characterized in that: The windowed straight line fitting obtains the straight line parameters of the chip, and the specific process is as follows: Step S41: Initialize the window size to n, and set the starting position of each point cloud to P1 i The end point is Step S42: determine whether the current window has traversed the entire point cloud set; if the traversal is complete, terminate the program; otherwise, proceed to step S43; Step S43: Perform LM straight line fitting on the point cloud points under the window to generate a straight line L, and calculate the distance d from each point cloud point to L; if d is greater than the maximum distance, execute step S44; Otherwise, execute step S45; Step S44: It indicates that a straight line cannot be successfully fitted under the window, the window start position is increased by 1, the window size is initialized, and the process returns to step S42; Step S45: It indicates that a straight line is successfully fitted under the window, the window size is set to n+1, and the distance d between the newly added point cloud point and the current L is calculated; If d is less than the maximum distance, repeat this step, otherwise go to step S46; Step S46: re-perform straight line fitting on the point cloud point set under the window to generate L′, and calculate the distance d from the current point cloud point to L′; if d is greater than the maximum distance, proceed to step S47, otherwise return to step S45; Step S47: record the parameters of line segment L, move the window start position to the original end position, initialize the window size, and return to step S42.
8. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S5, the specific implementation process is as follows: Step S51: scaling chip images of different sizes to a uniform fixed size through adaptive scaling, retaining the original aspect ratio and filling the edges; Step S52: using three-branch parallel dilated convolution to extract chip outlines and solder joint features of different sizes; Step S53: predicting the offset on the feature map and dynamically adjusting the convolution kernel sampling position; Step S54: applying L2 regularization to the offset Δp, and applying a random affine transformation to the chip image during training; Step S55: Generate a channel weight vector through global average pooling and a fully connected layer to perform cross-scale feature weighted fusion.
9. The chip positioning detection method based on image analysis according to claim 1, characterized in that: In step S6, the workflow of the lightweight convolutional network is as follows: Step S61: using a lightweight U-Net or CenterNet, inputting the extracted point cloud line segment features, and outputting a key point heat map; Step S62: construct geometric constraint modeling, construct an objective function, and minimize the deviation between the predicted coordinates and the geometric rules; Step S63: performing sub-pixel edge point calculation and B-spline curve fitting; Step S64: Fusion the deep learning prediction, geometric correction, and sub-pixel fitting results according to the confidence level.
10. A chip positioning detection system based on image analysis, characterized in that: It includes an operating station, an optical acquisition module located above the station, an industrial computer and a robotic arm; The optical acquisition module includes a high-resolution linear array CMOS camera, a single-line laser radar and a circular polarized light source; the high-resolution linear array CMOS camera is installed on a tilt bracket to complete the image acquisition of the chip; the single-line laser radar is used to measure the distance to any point on the chip; the circular polarized light source is used to provide a light environment for shooting and suppress metal reflections on the chip surface; The industrial computer includes a front-end FPGA and a back-end GPU server; the front-end FPGA is used to perform image preprocessing and feature extraction in real time; the back-end GPU server is equipped with a lightweight convolutional network, a chip design rule customization module, a rectangular sub-pixel edge detection and a dynamic self-calibration system; the multi-scale deformation feature fusion network is used; the lightweight convolutional network is used to extract chip contours and solder joint features of different scales through hole convolution; the chip design rule customization module is used to formulate chip design rules; the rectangular sub-pixel edge detection is used to locate the chip through the Zernike rectangular sub-pixel edge detection algorithm combined with B-spline curve fitting; The robot arm is used to move the chip to the corrected predicted coordinates.