Dense ultrasound micro-column segmentation and counting method and system with cross-scale feature interaction

CN122473616BActive Publication Date: 2026-09-29广州多浦乐电子科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610926256.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-29
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

[0011]有鉴于此,本发明的目的在于提供一种跨尺度特征交互的密集超声微柱分割计数方法和系统,通过构建背景抑制特征提取、跨尺度加权融合、可学习非线性尺度编码及联合损失优化,实现密集微柱高精度分割与计数;以解决超声C扫描图像中密集小尺寸微柱因特征微弱、尺度差异大、边缘模糊、背景干扰强而难以精确分割计数的技术难题

Benefits of technology

(1)特征保持方面:采用背景抑制层级特征提取网络BS Hiera提取多个不同层级的微柱特征图,在多次下采样过程中有效保持小尺度微柱目标的深度特征与边缘响应,避免特征丢失;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473616B_ABST
    Figure CN122473616B_ABST
Patent Text Reader

Abstract

The application discloses a dense ultrasonic micro-column segmentation and counting method and system for cross-scale feature interaction, and aims at the problem that dense micro-columns in an ultrasonic C-scan image are difficult to segment due to small size, close arrangement and large scale span, and constructs a segmentation and counting model: first, multi-level features are extracted through a background suppression feature extraction module BS Hiera, a Sobel edge extraction branch and a spatial context branch based on a KAN linear layer are used in parallel structure in the background suppression layer, the micro-column edge is strengthened and the background interference is suppressed; then, adaptive bidirectional weighted fusion is performed on the features of each level through a cross-scale feature weighted fusion network CSWF, and then the features are enhanced through a feature enhancement module; a KAN network is used to replace an MLP to construct a scale-specific query encoder, and modeling of multi-scale nonlinear changes is performed; finally, end-to-end optimization is performed in combination with a Focal loss and a counting loss. The application can effectively improve the segmentation precision and counting accuracy of dense small target micro-columns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ultrasonic nondestructive testing image analysis and processing technology, specifically a dense ultrasonic micropillar segmentation and counting method with cross-scale feature interaction. Background Technology

[0002] In the field of industrial nondestructive testing and microstructure quality assessment, accurate geometric parameter measurement and quantity statistics of microstructures such as micropores and micropillars inside workpieces are a core requirement. Taking the heat dissipation substrate of IGBT (Insulated Gate Bipolar Transistor) power modules as an example, its interior is densely packed with a large array of micropillars to enhance heat dissipation efficiency. The permeability of these micropillars directly determines the heat dissipation performance and operational reliability of the module. Ultrasonic C-scan imaging technology is widely used to detect the integrity of such micropillar structures due to its sensitivity to changes in the acoustic impedance within materials. In actual testing, the grayscale value of the ultrasonic C-scan image characterizes the energy distribution of the ultrasonic echo: low grayscale areas usually correspond to through-hole areas where ultrasonic waves can penetrate normally, while high grayscale areas indicate the presence of the micropillar itself or defects caused by foreign matter blockage. However, with the continuous advancement of microfabrication technology, the diameter of micropillars has become as small as 0.8 to 1.2 mm, with an arrangement density of hundreds per square centimeter. In ultrasonic C-scan images, the effective pixel area corresponding to a single micropillar is usually only about 10 × 10 pixels, which is a typical scenario of dense small targets. Accurate segmentation and counting of such images presents significant technical challenges.

[0003] For the segmentation and counting of dense small targets, existing technologies mainly include the following two categories of methods.

[0004] The first category comprises methods based on traditional image processing, such as thresholding, morphological operations, connected component analysis, and watershed algorithms. These methods typically rely on manually designed shallow features such as grayscale, gradients, or textures. They can achieve certain processing results when the target size is large and sparsely distributed. However, when faced with densely packed small targets with extremely weak pixel-level features, these methods struggle to extract stable and discriminative foreground information. Furthermore, the inherent speckle noise in ultrasonic images, edge blurring caused by beam diffusion, and grayscale inhomogeneity caused by workpiece surface curvature further interfere with the accuracy of edge or region-growing-based segmentation algorithms, leading to severe oversegmentation or undersegmentation. In some complex scenarios, researchers have proposed multi-scale segmentation methods based on minimum spanning trees or fractal network evolution, or coarse-to-fine segmentation strategies by constructing image pyramids, to segment multi-scale targets. However, when dealing with targets with extremely large scale spans, these methods still struggle to accurately locate the contours of small-scale targets while maintaining the consistency of large-scale target regions. Moreover, their segmentation results are often highly sensitive to the selection of merging criteria or scale parameters, lacking robustness.

[0005] The second category comprises methods based on general deep learning models. In recent years, instance segmentation models such as Mask R-CNN and SOLO, and object detection models such as the YOLO series and Faster R-CNN, have been widely applied to natural image segmentation and counting tasks. However, directly applying these models to ultrasonic micropillar segmentation presents significant drawbacks. First, the feature extraction networks in most deep learning segmentation models increase the receptive field and extract high-level semantic features through multiple convolutions and downsampling operations. However, in this process, the spatial details of small targets are almost completely lost, leaving only a blurry response in the deep feature map, making it impossible to effectively locate and segment small-sized micropillars. Although existing multi-scale feature fusion structures (such as Feature Pyramid Networks, FPN) can fuse high-level semantics and shallow details to some extent, they typically employ a unidirectional or simple fusion path, lacking a refined interaction mechanism for small target features, making it difficult to solve the problem of effective coordination between shallow edge noise and deep semantic information. Secondly, in the densely packed scenarios mentioned above, the target detection boxes overlap significantly. Traditional non-maximum suppression (NMS) operations are prone to incorrectly suppressing adjacent real targets, while the error accumulation effect of bounding box-based counting methods is obvious, further amplifying the counting bias.

[0006] To address the aforementioned feature loss issue, some existing techniques enhance the representational power of key features by introducing attention mechanisms during feature extraction. For example, multi-head self-attention mechanisms can capture global contextual information and are widely used to improve the model's sensitivity to target regions. However, in dense, small-target scenarios, relying solely on self-attention mechanisms is insufficient to effectively distinguish micropillar targets from background noise, especially in areas where target edges are blurred with the background. In such cases, the model often lacks accurate perception of boundary information, leading to incomplete segmentation contours. Furthermore, some studies have attempted to introduce edge detection branches into the segmentation network to assist in target boundary localization. However, these approaches mostly treat edge detection as an independent branch, resulting in insufficient fusion with the main segmentation network's features and failing to achieve effective synergistic enhancement of edge information and deep features.

[0007] On the other hand, in the perception and encoding of multi-scale targets, existing methods typically rely on scale-specific query encoders to generate query vectors corresponding to different target sizes. These encoders often employ traditional multilayer perceptron (MLP) structures, which are limited by fixed activation functions and linear weight combinations, making it difficult to effectively model complex nonlinear scale changes caused by large target size spans. When targets with vastly different scales coexist in a scene, MLP-based encoders are prone to insufficient fitting ability for extreme-scale targets, limiting the model's overall performance in multi-scale perception.

[0008] Furthermore, in terms of model training and optimization, existing segmentation methods typically employ standard cross-entropy loss or Dice loss for optimization. For dense micropillar segmentation tasks, the number of background pixels (matrix material regions) and foreground pixels (micropillar regions) is severely imbalanced, and the proportion of pixels at the micropillar edges is extremely low. Standard loss functions struggle to provide sufficient gradient attention to hard-to-classify pixels such as edges, causing the model to tend to ignore edge details during training. This results in overly smooth segmentation boundaries, affecting the separation accuracy of adhered targets. Although existing research has mitigated the class imbalance problem by introducing hard sample mining strategies such as Focal loss, how to simultaneously ensure global counting accuracy and local boundary integrity in dense small target segmentation tasks remains a challenge that current loss function designs have not adequately addressed.

[0009] In summary, existing technologies generally face the following core problems when dealing with the ultrasonic image segmentation and counting tasks of dense small-sized micropillars: (1) Severe loss of small target features: Multiple downsampling causes the spatial details and edge information of small-sized micropillars to disappear in deep features, and existing feature extraction networks lack background suppression and edge preservation mechanisms for dense small targets; (2) Inefficient cross-scale feature fusion: The feature interaction between shallow details and deep semantics is insufficient, and the traditional FPN unidirectional fusion structure cannot effectively recover the edge details of small targets; (3) Limited multi-scale perception capability: The scale encoder based on MLP is difficult to accurately model the nonlinear change relationship of cross-scale targets, which limits the generalization ability of the model in large-scale difference scenarios; (4) Insufficient boundary segmentation accuracy: Existing loss functions are difficult to focus on hard-to-classify pixels such as edges under extremely unbalanced class conditions, resulting in blurred segmentation boundaries of adhered micropillars and affecting counting accuracy.

[0010] Therefore, there is an urgent need in this field for a novel dense micropillar segmentation and counting method that can effectively solve the above problems, so as to achieve high-precision segmentation and reliable counting of dense small target micropillars in ultrasound C-scan images. Summary of the Invention

[0011] In view of this, the purpose of this invention is to provide a method and system for segmenting and counting dense ultrasound micropillars with cross-scale feature interaction. By constructing background suppression feature extraction, cross-scale weighted fusion, learnable nonlinear scale coding and joint loss optimization, high-precision segmentation and counting of dense micropillars can be achieved. This solves the technical problem that it is difficult to accurately segment and count dense small-sized micropillars in ultrasound C-scan images due to weak features, large scale differences, blurred edges and strong background interference.

[0012] To achieve the above objectives, the present invention provides the following technical solution: This invention first proposes a dense ultrasonic micropillar segmentation and counting method with cross-scale feature interaction, comprising the following steps: Step 1: Acquire an ultrasonic C-scan image of the workpiece to be inspected, and generate a target bounding box prompt on the ultrasonic C-scan image; Step 2: Input the ultrasound C-scan image into the pre-constructed background suppression hierarchical feature extraction network BHiera to extract multiple micropillar feature maps at different levels, including shallow feature maps, middle feature maps, and deep feature maps; Step 3: Input the feature maps of different levels into the scale-specific query encoder to generate query description vectors of the corresponding scales and output multiple scale-specific feature maps containing dense information. Step 4: Input the multiple scale-specific feature maps into the pre-constructed cross-scale feature weighted fusion network CSWF to perform adaptive fusion of deep and shallow features, and output the fused multi-scale features; Step 5: Input the fused multi-scale features and the target bounding box prompt information into the mask decoding network to generate an instance segmentation mask for each micropillar, and output the segmentation count result based on the instance segmentation mask.

[0013] Furthermore, in step two, the background suppression hierarchical feature extraction network (BS Hiera) includes multiple BS Hiera Blocks connected in sequence; each BS Hiera Block includes a multi-head self-attention module (MHA), a background suppression layer (BS Layer), skip connections, and an MLP layer connected in sequence. The background suppression layer (BS Layer) includes a target boundary extraction branch and a spatial context branch set in parallel, as well as a feature fusion unit; The target boundary extraction branch is used to extract edge features in the H and W directions of the input features using the Sobel operator, and to perform weighted enhancement on the edge features to output enhanced edge features; The spatial context branch is used to extract spatial context feature weights using global average pooling and nonlinear transformation based on KAN linear layers, and to perform weighted fusion of input features according to the spatial context feature weights to suppress background interference and output background suppression features. The feature fusion unit is used to fuse the enhanced edge features, the background suppression features and the input features, and output the background suppression layer (BS Layer) features after adjusting the channels through a 1×1 convolution.

[0014] Furthermore, the method for weighted enhancement of the edge features by the target boundary extraction branch is as follows:

[0015]

[0016] in: Input features; Input features The feature gradient of the c-th channel is in the H direction; Input features The feature gradient of the c-th channel in the W direction; Input features Features of the c-th channel; and These are the Sobel gradient operator kernels for the H and W directions, respectively; This represents the gradient magnitude. This represents the batch size; The number of input feature channels; The height of the feature map; The width of the feature map; For the gradient magnitude Perform a 1×1 convolution operation to obtain the enhanced edge features:

[0017] in: To enhance edge features; This represents a 1×1 convolution operation.

[0018] Furthermore, the method for extracting the spatial context feature weights by the spatial context branch is as follows:

[0019] in: The output characteristics of the average pooling operation; Input features; and These are the height and width of the feature map, respectively; Nonlinear transformation using KAN linear layers:

[0020]

[0021] in: For the spline nonlinear function of the KAN linear layer; For linear spline basis functions of the KAN linear layer; These are learnable control point parameters; The nonlinear transformation output features of the KAN linear layer; This indicates that the KAN linear layer is affected by the input features. After performing nonlinear transformation, the first One output dimension; The feature dimension represents the input feature. Calculate the spatial context feature weights:

[0022] Weight the input features:

[0023] in: Background suppression features; The boundary weights of the micro-pillars.

[0024] Furthermore, in step four, the cross-scale feature weighted fusion network CSWF includes a feature alignment module, a dual-path bidirectional feature fusion module DBFF, and a feature enhancement module based on structural reparameterization. The feature alignment module is used to downsample and align the channel dimensions of the shallow feature map and the middle feature map to adapt to cross-scale feature fusion. The dual-path bidirectional feature fusion module DBFF is used to receive aligned shallow features, mid-level features, and deep features, calculate the fusion weight of cross-layer features, and perform weighted fusion of the shallow features, mid-level features, and deep features according to the fusion weight to realize cross-layer feature interaction. The feature enhancement module based on structural reparameterization is used to enhance the features of each layer after interaction, and then splice the enhanced features of each layer to output the fused multi-scale features.

[0025] Furthermore, the dual-path bidirectional feature fusion module DBFF includes a first side branch, a second side branch, and an intermediate branch; The first side branch is used to retain the feature information of the first input feature; the second side branch is used to retain the feature information of the second input feature; the middle branch is used to learn the weights of the concatenated first input feature and second input feature, and output the first fusion weight corresponding to the first input feature and the second fusion weight corresponding to the second input feature. The output features of the dual-path bidirectional feature fusion module DBFF are:

[0026] in: The output features of the dual-path bidirectional feature fusion module DBFF; and These are two input features; This represents a 1×1 convolution operation; and These are the first fusion weight and the second fusion weight, respectively, and:

[0027] in: This is a global average pooling operation; Calculated for MLP; For the Softmax function; .

[0028] Furthermore, the scale-specific query encoder is constructed using a KAN network, where each weight parameter is replaced by a learnable one-dimensional spline function to model the nonlinear scale variation relationship of multi-scale targets; the mask decoding network uses the SAM2 Mask Decode network.

[0029] Furthermore, training is performed using a total loss function jointly optimized by Focal loss and counting loss, wherein the total loss function is:

[0030] in: For count loss; Loss due to counting small targets; Focal loss; Weights for counting small targets.

[0031] Furthermore, in step five, the method for outputting the segmentation count result based on the instance segmentation mask is as follows: The pixel area of ​​the instance segmentation mask for each micropillar is counted and compared with the theoretical pixel area of ​​the micropillar's designed aperture to calculate the single-aperture throughput. The total number of micropillar vias is calculated based on the number of segmented masks described in the example. Based on the single-hole through-hole ratio and / or the total number of through-holes in the micropillars, output the through-hole ratio statistics, blockage defect location results, and / or manufacturing pass rate determination results of the workpiece.

[0032] The present invention also proposes a system for implementing the dense ultrasonic micropillar segmentation and counting method for cross-scale feature interaction as described above, comprising: The image acquisition module is used to acquire an ultrasonic C-scan image of the workpiece to be inspected, and generate a target box prompt information on the ultrasonic C-scan image; The feature extraction module is used to input the ultrasound C-scan image into the pre-constructed background suppression hierarchical feature extraction network BS Hiera to extract multiple micropillar feature maps at different levels, including shallow feature maps, middle feature maps and deep feature maps; The scale encoding module is used to input the feature maps of the multiple different levels into the scale-specific query encoder, generate query description vectors of the corresponding scale, and output multiple scale-specific feature maps containing dense information. The cross-scale fusion module is used to input the multiple scale-specific feature maps into a pre-constructed cross-scale feature weighted fusion network (CSWF) to perform adaptive fusion of deep and shallow features and output the fused multi-scale features. The segmentation counting module is used to input the fused multi-scale features and the target bounding box prompt information into the mask decoding network, generate an instance segmentation mask for each micropillar, and output the segmentation counting result based on the instance segmentation mask.

[0033] The beneficial effects of this invention are as follows: The present invention provides a dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction, which has the following technical advantages: (1) Feature preservation: The background suppression hierarchical feature extraction network BS Hiera is used to extract micro-pillar feature maps at multiple different levels. During multiple downsampling processes, the depth features and edge responses of small-scale micro-pillar targets are effectively preserved, avoiding feature loss. (2) Feature fusion: The cross-scale feature weighted fusion network CSWF is adopted to achieve adaptive collaborative enhancement of shallow details and deep semantics, which overcomes the problem of insufficient recovery of edge details of small targets by traditional one-way fusion strategy; (3) Scale perception: A scale-specific query encoder is used to decompose and model targets of different sizes, which improves the accuracy of multi-scale target perception. (4) In terms of segmentation accuracy: the micro-pillar instance segmentation mask is generated by the mask decoding network and the segmentation counting result is directly output, which ensures the end-to-end segmentation accuracy and counting accuracy.

[0034] In summary, this invention effectively solves the technical problem that dense, small-sized micropillars are difficult to segment precisely due to their weak features, large scale differences, and blurred edges.

[0035] The present invention also has the following technical effects: (1) In the background suppression hierarchical feature extraction network BS Hiera, the background suppression layer BS Layer adopts a parallel dual-branch structure of Sobel edge detection and KAN linear layer spatial attention, which respectively strengthens the micro-pillar edge response and suppresses background interference. The skip connection preserves the integrity of the original features, enabling the network to accurately locate the micro-pillar boundary in densely arranged scenes, reduce false background detection, and significantly improve the recall rate of small-sized micro-pillars. (2) In the cross-scale feature weighted fusion network CSWF, the dual-path bidirectional feature fusion module DBFF realizes the adaptive weighted fusion of shallow details and deep semantics through bidirectional interaction, which overcomes the defect that the edge information of small targets cannot be recovered in the traditional FPN unidirectional fusion; the feature enhancement module realizes the equivalent transformation between training multi-branch enhancement and inference single-branch, which ensures the segmentation accuracy without increasing the inference overhead. (3) In the KAN linear layer, a learnable spline function is used to replace the fixed activation function of the MLP, so as to accurately model the nonlinear scale change relationship of multi-scale targets with fewer parameters and improve the perception accuracy of extreme scale targets. (4) Joint loss optimization: Focal loss and counting loss are complementary. Counting loss ensures the correct global quantity response, while Focal loss drives the model to focus on hard-to-classify pixels such as edges, effectively solving the boundary blurring problem of densely connected micropillars. (5) By calculating the single-hole through-hole rate, judging the blockage defect and evaluating the manufacturing qualification rate, the segmentation results are directly converted into industrial quality inspection indicators, which have clear engineering application value. Attached Figure Description

[0036] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 This is a network architecture diagram of a cross-scale segmentation counting model; Figure 2 Here is a structural diagram of the BS Hiera Block; Figure 3 Here is a structural diagram of the background suppression layer (BS Layer); Figure 4 Structure diagram of the cross-scale feature weighted fusion network CSWF; Figure 5 This is a flowchart of the dense ultrasonic micropillar segmentation and counting method with cross-scale feature interaction according to the present invention. Figure 6 The first image shows the visualization results; (a), (b), and (c) are the visualization results of the segmentation and counting of micropillars in the sample workpiece. Detailed Implementation

[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0038] Ultrasonic C-scan imaging is based on the propagation and reflection characteristics of ultrasound in materials. When ultrasound encounters an interface with different acoustic impedances (such as the interface between a micropillar pore and the substrate material), reflection and transmission occur. Pore regions allow ultrasound to penetrate normally, resulting in weaker echo energy collected by the receiving probe, corresponding to high grayscale values ​​in the C-scan image. Blocked regions (micropillar bodies or foreign matter fillings) obstruct ultrasound propagation, resulting in stronger echo energy collected by the receiving probe, corresponding to low grayscale values ​​in the C-scan image. This embodiment utilizes this physical characteristic to propose a cross-scale feature interaction-based method for segmenting and counting dense small-target micropillars. This method can automatically identify and segment micropillar pore regions in C-scan images, achieving high-precision counting and defect detection. First, a background suppression hierarchical feature extraction network (BS Hiera) is designed. Through a background suppression layer and a multi-head self-attention mechanism, the depth features and edge information of dense micropillar targets are enhanced, effectively avoiding feature loss in small-scale micropillars. Second, a cross-scale feature weighted fusion network (CSWF) is constructed, employing a bidirectional feature interaction and adaptive weighted fusion strategy to achieve synergistic enhancement of shallow details and deep semantics. Then, a learnable nonlinear scaling code based on the KAN (Kolmogorov-Arnold Network) is introduced to replace the traditional MLP (Multilayer Perceptron) query encoder. Leveraging the multi-scale decomposition capability and superior neural scaling law of the KAN network, the model's perception accuracy for micropillars of different scales is significantly improved. Finally, end-to-end optimization is performed by combining Focal loss and counting loss, forcing the model to focus on hard-to-classify pixels such as edges, thus improving the segmentation boundaries of adhered micropillars.

[0039] This embodiment uses an ultrasonic C-scan image of the micropillar region of a certain type of IGBT cooling plate as an example to provide a detailed description of the specific implementation of the dense ultrasonic micropillar segmentation and counting method with cross-scale feature interaction.

[0040] 1. Construct a cross-scale segmentation counting model.

[0041] The network architecture for constructing the dense ultrasonic micropillar segmentation and counting method for implementing the cross-scale feature interaction of this invention is as follows: Figure 1As shown. In this embodiment, the cross-scale segmentation counting model mainly consists of an input image, target bounding box prompts, a background suppression layer feature extraction network (BS Hiera, Hiera with backgroundsuppression Layer, BS Hiera), a scale-specific query encoder, a cross-scale feature weighted fusion network (CSWF, CSWF), and a mask decoding network. The overall process is as follows: First, an ultrasonic C-scan image of the micropillar region of the IGBT cooling plate is acquired. The image grayscale value represents the distribution of ultrasonic echo energy: low grayscale values ​​(close to 0) correspond to the micropillar through-hole region where ultrasonic energy can penetrate normally, and high grayscale values ​​(close to 255) correspond to the region where there is a micropillar body or blockage defect at this location. By using a rectangular bounding box annotation method, the micropillar array region to be detected is selected on the C-scan image to generate target bounding box prompt information, which serves as the spatial prior guidance information for subsequent networks. Then, the Background Suppression Hierarchical Feature Extraction Network (BS Hiera) is used to obtain rich depth features and edge information of the dense micropillar ultrasound image, avoiding feature loss during the feature extraction process. The output includes three layers of features: a deep feature map S4, a middle feature map S3, and a shallow feature map S2, representing the edge details, local structural information, and global semantic information of the micropillar, respectively. Next, the deep feature map S4, the middle feature map S3, and the shallow feature map S2 are each passed through a scale-specific query encoder to obtain multiple scale-specific feature maps containing dense information. These feature maps are then fused using a cross-scale feature weighting fusion network (CSWF) with depth adaptive fusion to obtain the fused multi-scale features. Finally, the fused multi-scale features are passed through a cue-guided mask decoding network to output the final counting and segmentation result.

[0042] (1) Background suppression hierarchical feature extraction network BS Hiera.

[0043] Specifically, the Background Suppression Hierarchical Feature Extraction Network (BS Hiera) comprises multiple sequentially connected BS Hiera Blocks. Each BS Hiera Block includes a multi-head self-attention module (MHA), a background suppression layer (BS Layer), skip connections, and an MLP layer, all connected sequentially. For example... Figure 2The diagram shows the structure of the BS Hiera Block designed in this embodiment. The BS Hiera Block first slices the input features, fuses positional encoding information, stabilizes the feature distribution using Layer Norm, and calculates the Q, K, and V matrices. Then, a multi-head self-attention module (MHA) is used to divide the Q, K, and V matrices into multiple attention heads. Within each attention head, the matching score between the Q and K matrices is calculated. The matrix is ​​then converted into a weighted distribution using the Softmax activation function, and this weighted distribution is used to sum the V matrix. The MHA-weighted features are then output through linear projection. Simultaneously, a background suppression layer (BS Layer) is used to further focus the output features, ensuring the network accurately focuses on dense micropillar targets. The BS Layer utilizes a concatenated structure of average pooling (Avgpool), KAN linear layers (KAN linear), and sigmoid to capture multi-scale micropillar features caused by size differences and focus on local information, maintaining high recognition robustness. By integrating Sobel edge detection and boundary-guided gating, richer micropillar edge features are extracted and fused to enhance the features of dense micropillar targets. This results in more accurate boundary contours, significantly reduced interference from complex backgrounds, decreased segmentation errors, and improved integrity and accuracy. Finally, after feature fusion by skip connections between the output features of the background suppression layer (BS Layer) and the input features, an MLP is used to integrate the nonlinear micropillar boundary features, outputting the target micropillar features extracted by a BS Hiera Block layer.

[0044] Specifically, such as Figure 3 As shown, the Background Suppression Layer (BS Layer) includes a target boundary extraction branch and a spatial context branch set in parallel, as well as a feature fusion unit. The target boundary extraction branch uses the Sobel operator to extract edge features in the H and W directions of the input features, and then weights and enhances these edge features to output enhanced edge features. The spatial context branch uses global average pooling and a nonlinear transformation based on a KAN linear layer to extract spatial context feature weights, and then weights and fuses the input features according to these spatial context feature weights to suppress background interference, outputting background suppression features. The feature fusion unit fuses the enhanced edge features, the background suppression features, and the input features, and then outputs the BS Layer features after adjusting the channels through a 1×1 convolution.

[0045] Specifically, the background suppression layer (BS Layer) first takes the features as input. It is divided into two branches: pillar boundary extraction and spatial context, which are used to enhance pillar boundary features and suppress background interference, respectively. Pillar boundary extraction The branch uses the Sobel operator to extract edge features in the H and W directions respectively, and then uses a 1×1 convolution to scale the features. Spatial context The branch uses AvgPool, KAN linear layers, and Sigmoid to compute spatial context feature weights. The KAN linear layer performs a non-linear transformation on each input feature using a learnable spline function, while global pooling compresses the spatial dimension into channel response vectors, replacing spatial context location information with global semantics for the KAN linear layer to extract high-level features. Sigmoid computes the weight vector for this branch and fuses it with the input. Finally, the outputs of the two branches and the input features are combined. Fusion, output features of the background suppression layer (BS Layer) .

[0046] Specifically, the method for weighted enhancement of the edge features by the target boundary extraction branch is as follows.

[0047]

[0048]

[0049] in: Input features; Input features The feature gradient of the c-th channel is in the H direction; Input features The feature gradient of the c-th channel in the W direction; Input features Features of the c-th channel; and These are the Sobel gradient operator kernels for the H and W directions, respectively; This represents the gradient magnitude. This represents the batch size; The number of input feature channels; The height of the feature map; This represents the width of the feature map.

[0050] For the gradient magnitude Perform a 1×1 convolution operation to obtain the enhanced edge features:

[0051] in: To enhance edge features; This represents a 1×1 convolution operation.

[0052] Specifically, the method for extracting the spatial context feature weights using the spatial context branch is as follows.

[0053]

[0054] in: The output characteristics of the average pooling operation; Input features; and These represent the height and width of the feature map, respectively.

[0055] Nonlinear transformation using KAN linear layers:

[0056]

[0057] in: For the spline nonlinear function of the KAN linear layer; For linear spline basis functions of the KAN linear layer; These are learnable control point parameters; The nonlinear transformation output features of the KAN linear layer; This indicates that the KAN linear layer is affected by the input features. After performing nonlinear transformation, the first One output dimension; This represents the feature dimension of the input features.

[0058] Calculate the spatial context feature weights:

[0059] Weight the input features:

[0060] in: Background suppression features; The boundary weights of the micro-pillars.

[0061] (2) Cross-scale feature weighted fusion network CSWF.

[0062] In this embodiment, the cross-scale feature weighted fusion network (CSWF) includes a feature alignment module, a dual-path bidirectional feature fusion module (DBFF), and a feature enhancement module based on structural reparameterization. Specifically, the feature alignment module is used to downsample and align the channel dimensions of the shallow feature map S2 and the mid-level feature map S3 to adapt to cross-scale feature fusion; the dual-path bidirectional feature fusion module (DBFF) is used to receive the aligned shallow, mid-level, and deep features, calculate the fusion weights of the cross-layer features, and perform weighted fusion of the shallow, mid-level, and deep features according to the fusion weights to achieve cross-layer feature interaction; the feature enhancement module based on structural reparameterization is used to enhance the features of each layer after interaction, and then concatenate the enhanced features of each layer to output the fused multi-scale features.

[0063] To address the feature weakening and interference caused by the varying sizes of micropillars, as well as the loss of details and feature sparsity during feature sampling and calculation, this embodiment is designed as follows: Figure 4 The cross-scale feature weighted fusion network (CSWF) module is shown. This module, through cross-level bidirectional feature interaction and adaptive feature selection mechanisms, simultaneously enhances shallow detail information and deep texture features during multi-level feature fusion, significantly improving the model's segmentation and counting accuracy for various micropillars in complex backgrounds.

[0064] The cross-scale feature weighted fusion network CSWF takes deep feature map S4, mid-level feature map S3, and shallow feature map S2 as input, such as... Figure 4 As shown, firstly, the deep feature map S4 and the mid-level feature map S3 are downsampled and aligned with the channel dimension features using 3×3 convolutions to adapt to the subsequent cross-scale feature fusion process. Then, the aligned shallow and deep features are respectively processed by the designed dual-branch feature fusion module DBFF (Dual Branch Feature Fusion) for feature interaction, and then the interacted features are respectively enhanced by a feature enhancement module based on structure reparameterization. Finally, the multi-scale features are output after being concatenated by the stepwise processing of the three-layer feature enhancement module. The design of the dual-branch feature fusion module DBFF can flexibly handle multiple input features and adaptively fuse them. Taking the deep feature map S4 and the mid-level feature map S3 as input features as an example, the dual-branch feature fusion module DBFF first uses 1×1 Conv to adjust the channel features of each feature, and then divides them into three branches including the first side branch, the second side branch, and the middle branch. The features of the first side branch and the second side branch are respectively and The first branch preserves the feature information of the original input, while the second branch learns weights for the concatenated features to adaptively obtain the fusion weights for the two branches. Finally, these weights are fused with the features from the two branches using corresponding weighted feature fusion, thus achieving cross-layer feature interaction and fusion. This process allows deep semantic information to guide the selection of shallow edge features, and shallow edge information to supplement the minute aperture details of deep features, effectively solving the problem that minute aperture edge information cannot be effectively recovered from deep semantics in traditional FPN unidirectional fusion. That is, in this embodiment, the input feature map... For the deep feature map S4, input features The adjustment number S3 is for the middle layer. Of course, in some other embodiments, the shallow feature map S2 and the middle feature map S3 can be used as two input features respectively, or the deep feature map S4 and the shallow feature map S2 can be used as two input features respectively, which will not be elaborated further.

[0065] Specifically, in this embodiment, the output features of the dual-path bidirectional feature fusion module DBFF are:

[0066] in: The output features of the dual-path bidirectional feature fusion module DBFF; and These are two input features; This represents a 1×1 convolution operation; and These are the first fusion weight and the second fusion weight, respectively, and:

[0067] in: This is a global average pooling operation; Calculated for MLP; For the Softmax function; .

[0068] (3) Scale-specific query encoder.

[0069] To address the significant scale variations in dense micropillar segmentation and counting tasks, the Scale-Specific Query Encoder (SME) decomposes the complex cross-scale modeling task into several intra-scale subtasks by generating multiple dedicated query descriptions corresponding to different object sizes in parallel. Each encoder is only responsible for capturing and aggregating the corresponding micropillar features within its preset scale range, thus avoiding the missed detections (small-sized micropillars) and boundary confusion (large-sized micropillars) caused by the fixed receptive field of traditional single-scale models. Furthermore, this module collaborates with cross-attention and deformable attention mechanisms to guide the model in generating more refined instance segmentation masks and more accurate density maps. However, MLP-based SMEs are limited by fixed activation functions and linear weight structures, making it difficult to effectively model multi-scale variations caused by targets of different sizes, thus limiting their localization and counting accuracy in dense scenes with uneven scale distribution. Given that KAN networks possess inherent multi-scale decomposition capabilities and superior neural scaling, enabling stronger nonlinear fitting with fewer parameters, this embodiment uses KAN networks instead of MLPs to construct a scale-specific query encoder. Specifically, in this embodiment, the scale-specific query encoder is constructed using a KAN network, where each weight parameter is replaced by a learnable one-dimensional spline function to model the nonlinear scale variation relationship of multi-scale targets, thereby significantly improving the model's overall performance in multi-scale perception.

[0070] (4) Loss function The cross-scale segmentation counting model uses counting loss to globally constrain the quantity response of micropillar targets. However, in dense micropillar segmentation scenarios, the proportion of edge pixels is extremely low, and the counting loss is insufficient to give enough gradient attention to hard-to-classify edge samples, resulting in insufficient segmentation accuracy of dense small target micropillar boundaries. To address this, this embodiment introduces Focal loss, which dynamically reduces the loss weight of easily classified samples, forcing the model to focus its optimization on hard-to-classify pixels such as edges. It compares the pixel-by-pixel predicted probability with the true label, reduces the loss contribution to flat regions with high confidence, and increases the loss contribution to edge regions with low confidence, supplemented by a balancing factor to alleviate the imbalance between positive and negative samples. Focal loss and the original counting loss are jointly optimized and complement each other, where the counting loss ensures the correct overall response of the micropillar, and Focal loss drives fine boundary segmentation, thereby effectively improving the integrity of the edges of small vias in ultrasound C-scan images without increasing the computational burden. Specifically, this embodiment uses a total loss function jointly optimized by Focal loss and counting loss for training. The total loss function of the cross-scale segmentation counting model is:

[0071] in: For count loss; Loss due to counting small targets; Focal loss; Weights for counting small targets.

[0072] The backbone network of the Background Suppression Hierarchical Feature Extraction Network (BS Hiera) extracts relatively rich depth and edge features of dense micropillars, effectively avoiding the loss of small-scale micropillar features during extraction and improving the recall rate of small-sized micropillar targets. The Cross-Scale Feature Weighted Fusion Network (CSWF) further enhances the edge response of micropillar targets, suppressing interference from non-target regions (such as matrix material and background noise), making the micropillar contour information clearer and reducing the probability of missed detections in densely arranged scenes. Learnable nonlinear scale encoding, through the introduction of learnable gating and bi-branch interaction mechanisms, achieves adaptive weighted fusion of features at adjacent scales, effectively suppressing shallow high-frequency noise and significantly enhancing the edge response of small-sized micropillars, solving the key problems of easy loss of small target features and blurred edges in dense micropillar scenes. Finally, the introduced Focal loss allows the model to focus optimization on difficult-to-classify micropillar pixels such as boundaries during training, solving the problem of missed counts caused by the smoothness and severe adhesion of dense micropillar features.

[0073] 2. Dense ultrasonic micropillar segmentation and counting with cross-scale feature interaction.

[0074] 2.1 A dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction.

[0075] like Figure 5 As shown, the dense ultrasound micropillar segmentation and counting method with cross-scale feature interaction in this embodiment includes the following steps: Step 1: Obtain an ultrasonic C-scan image of the workpiece to be inspected, and generate a target box prompt on the ultrasonic C-scan image.

[0076] Specifically, ultrasonic C-scan images of the micropillar region of the IGBT cooling plate are acquired. The image grayscale values ​​characterize the distribution of ultrasonic echo energy: low grayscale values ​​(close to 0) correspond to micropillar through-hole areas where ultrasonic energy can penetrate normally, while high grayscale values ​​(close to 255) correspond to areas where there are micropillar bodies or blockage defects. Using a rectangular bounding box annotation method, the micropillar array region to be detected is selected on the C-scan image, generating target bounding box coordinates as spatial prior guidance information for subsequent network processing.

[0077] Step 2: Input the ultrasound C-scan image into the pre-constructed background suppression hierarchical feature extraction network BHiera to extract multiple micropillar feature maps at different levels, including shallow feature maps, intermediate feature maps, and deep feature maps.

[0078] Specifically, the input image is fed into the Background Suppression Hiera (BS Hiera) network to extract the depth features and edge information of the micropillars, outputting three different layers of feature maps: shallow feature map S2, mid-layer feature map S3, and deep feature map S4, representing the edge details, local structural information, and global semantic information of the micropillars, respectively. The specific process is as follows: (1) Slicing and positional encoding fusion: The input features are sliced, positional encoding information is fused, the feature distribution is stabilized by LayerNorm (layer normalization), and the Q (query), K (key), and V (value) matrices are calculated.

[0079] (2) Multi-head self-attention MHA calculation: The Q, K, and V matrices are divided into multiple attention heads. The matching score between the Q and K matrices is calculated in each attention head. The K matrix is ​​calculated as a weight distribution through the Softmax activation function. The V matrix is ​​weighted and summed using these weights. The weighted features of the multi-head self-attention MHA are output through linear projection.

[0080] (3) Background Suppression Layer (BS Layer) Processing: The output features of the Multi-Head Self-Attention (MHA) layer are fed into the Background Suppression Layer (BS Layer), which is divided into two parallel branches. Target Boundary Extraction Branch: The Sobel operator is used to extract edge features in the H and W directions respectively, and then 1×1 convolution is used for scale adjustment. The Sigmoid function is used to calculate the micro-pillar boundary weights to weight and strengthen the edge features. Spatial Context Branch: The spatial dimension is compressed into a channel response vector using AvgPool (global average pooling). Nonlinear feature extraction is performed through a KAN linear layer (a linear layer based on learnable spline functions). The spatial context feature weights are calculated using Sigmoid and weighted and fused with the input to suppress background interference. Finally, the outputs of the two branches are fused with the input features to output the final features of the Background Suppression Layer (BS Layer).

[0081] (4) Skip connection and MLP integration: The output features of the background suppression layer (BS Layer) are fused with the original input features by skip connection, and then the MLP is used to integrate the nonlinear micro-pillar boundary features. The target micro-pillar features extracted by the BS HieraBlock layer are output, and finally the three-level feature maps from shallow to deep are obtained, namely: shallow feature map S2, middle feature map S3 and deep feature map S4.

[0082] Step 3: Input the feature maps of different levels into the scale-specific query encoder to generate query description vectors of the corresponding scales and output multiple scale-specific feature maps containing dense information.

[0083] The feature maps of three layers—shallow feature map S2, mid-level feature map S3, and deep feature map S4—are fed into a Scale-Specific Query Encoder. The Scale-Specific Query Encoder generates multiple dedicated query description vectors corresponding to different target sizes in parallel, decomposing the cross-scale modeling task into several intra-scale sub-tasks. Each encoder is only responsible for capturing and aggregating the corresponding target features within its preset scale range. This embodiment uses a KAN network instead of a traditional MLP to construct the Scale-Specific Query Encoder. Utilizing the mechanism in the KAN network where each weight parameter is replaced by a learnable one-dimensional spline function, it fully maps the nonlinear relationships between multiple features, improving the model's perception accuracy at different scale locations. The encoded output contains feature maps with more accurate and dense information for subsequent cross-scale fusion.

[0084] Step 4: Input the multiple scale-specific feature maps into the pre-constructed cross-scale feature weighted fusion network CSWF to perform adaptive fusion of deep and shallow features, and output the fused multi-scale features.

[0085] The shallow feature map S2, the middle feature map S3, and the deep feature map S4 output from step three are fed into the cross-scale feature weighted fusion network CSWF to perform adaptive fusion of shallow and deep features. The specific process is as follows.

[0086] (1) Feature alignment: Use 3×3 convolution to downsample and align the mid-level feature map S3 and the deep feature map S4 with the channel dimension features to adapt to subsequent cross-scale feature fusion.

[0087] (2) DBFF Bidirectional Feature Fusion: The aligned shallow and deep features are respectively processed by the dual-path bidirectional feature fusion module DBFF for feature interaction. The dual-path bidirectional feature fusion module DBFF uses 1×1 convolution to adjust the channels of each feature, dividing it into two side branches (including the first side branch and the second side branch, which retain the original feature information) and a middle branch (learning the fusion weights of the two side branches). The fusion weights are adaptively calculated through global average pooling, MLP calculation and Softmax function, and the weights are weighted and fused with the features of the two side branches to realize cross-layer feature interaction.

[0088] (3) Feature enhancement module based on structure reparameterization: In the specific example, the feature enhancement module based on structure reparameterization adopts RepC3. The interactive features are enhanced by the feature enhancement module, and the three-layer feature enhancement module is used to process them step by step.

[0089] (4) Multi-scale feature splicing output: The features output by the three-layer feature enhancement module are spliced ​​to output the fused multi-scale features, which simultaneously enhance shallow detail information and deep semantic texture features.

[0090] Step 5: Input the fused multi-scale features and the target bounding box prompt information into the mask decoding network to generate an instance segmentation mask for each micropillar, and output the segmentation count result based on the instance segmentation mask.

[0091] The multi-scale features fused in step four and the target bounding box cue information obtained in step one are fed into a cue-guided mask decoding network to output an instance segmentation mask for each micropillar. As a preferred implementation, this embodiment uses the SAM2 Mask decoder as the cue-guided mask decoding network. The SAM2 mask decoder uses the target bounding box cue information as a spatial prior to guide the decoder to focus on the target micropillar array region, avoiding missegmentation of irrelevant background regions in the image. For each identified micropillar via instance, its pixel area is calculated and compared with the theoretical pixel area of ​​the designed aperture of the micropillar to calculate the single-hole permeability.

[0092] in: To divide the area of ​​the mask pixels; Calculate the pixel area corresponding to the aperture in CAD design.

[0093] Ultimately, as Figure 6 As shown, the total number of micropillar through holes is obtained by counting the number of output masks. Combined with the single hole through hole rate calculation results, quantitative detection results are output for workpiece through hole rate statistics, blockage defect location and manufacturing qualification rate determination.

[0094] Furthermore, this embodiment employs a training strategy that combines Focal loss and counting loss for joint optimization. The counting loss globally constrains the response of the micropillar targets, ensuring the overall correctness of the micropillar response. The counting weight for small targets is set to 0.3. Focal loss dynamically reduces the loss weight of easily classified samples, forcing the model to focus its optimization efforts on hard-to-classify pixels such as edges. It reduces the loss contribution to flat regions with high confidence and increases the loss contribution to edge regions with low confidence. A balancing factor is also used to alleviate the imbalance between positive and negative samples, driving fine segmentation of the dense micropillar boundaries.

[0095] 2.2 Dense Ultrasonic Micropillar Segmentation and Counting System with Cross-Scale Feature Interaction

[0096] This embodiment also proposes a dense ultrasound micropillar segmentation and counting system with cross-scale feature interaction, used to implement the aforementioned dense ultrasound micropillar segmentation and counting system method with cross-scale feature interaction. Specifically, the dense ultrasound micropillar segmentation and counting system with cross-scale feature interaction in this embodiment includes: an image acquisition module, a feature extraction module, a scale encoding module, a cross-scale fusion module, and a segmentation and counting module. Specifically, the image acquisition module acquires an ultrasonic C-scan image of the workpiece to be inspected and generates a target bounding box prompt on the ultrasonic C-scan image; the feature extraction module inputs the ultrasonic C-scan image into a pre-constructed background suppression hierarchical feature extraction network (BS Hiera) to extract multiple micropillar feature maps at different levels, including shallow, medium, and deep feature maps; the scale encoding module inputs the multiple feature maps at different levels into a scale-specific query encoder to generate query description vectors for the corresponding scales and outputs multiple scale-specific feature maps containing dense information; the cross-scale fusion module inputs the multiple scale-specific feature maps into a pre-constructed cross-scale feature weighted fusion network (CSWF) to perform adaptive fusion of shallow and deep features and outputs the fused multi-scale features; the segmentation counting module inputs the fused multi-scale features and the target bounding box prompt into a mask decoding network to generate an instance segmentation mask for each micropillar and outputs the segmentation counting result based on the instance segmentation mask.

[0097] 3. Technical effects.

[0098] (1) This embodiment proposes a background suppression hierarchical feature extraction network BS Hiera based on background suppression. On the basis of the Hiera hierarchical structure, a background suppression layer BS Layer is introduced. Through two parallel branches, Sobel edge detection and KAN linear layer learnable spline spatial attention, the response of micro-pillar target edge contour and the suppression of background interference are strengthened respectively. This effectively solves the problem that small-scale target features disappear during multiple downsampling processes in traditional backbone networks, and significantly improves the recall rate of small targets in dense scenes.

[0099] (2) A cross-scale feature weighted fusion network (CSWF) is designed. The dual-path bidirectional feature fusion module (DBFF) adaptively weights and fuses shallow detail features and deep semantic features. After fusion, the feature enhancement module realizes the equivalent transformation between multi-branch feature enhancement in the training stage and single-branch feature enhancement in the inference stage. This structure solves the problem that edge information of small targets cannot be effectively recovered from deep semantics in the traditional unidirectional fusion of FPN.

[0100] (3) Introducing the KAN network into a scale-specific query encoder to replace the traditional MLP. By using the learnable spline activation function of each connection edge in the KAN network to replace the fixed activation function in the MLP, the network can model the nonlinear scale variation relationship of multi-scale targets with fewer parameters. The local controllability of the spline function allows different scale intervals to be adjusted by independent control point parameters, avoiding the problem of insufficient fitting ability of traditional MLP to extreme scale targets due to fixed activation functions.

[0101] (4) Construct a joint optimization strategy of Focal loss and counting loss. The counting loss constrains the global micropillar number response, while the Focal loss drives the model to focus on the boundary region of the sticky target by dynamically reducing the loss weight of easily classified background pixels and increasing the loss weight of hard-to-classify pixels such as edges. The two losses complement each other and solve the problems of undersegmentation and undercounting caused by boundary ambiguity in dense target segmentation.

[0102] The micropillar through-hole segmentation and counting results output in this embodiment can be directly used in the following industrial application scenarios: (1) Workpiece through hole ratio statistics: count the number of masks for all micro-pillar through holes, calculate the overall through hole ratio of the workpiece, and determine whether it meets the design requirements (usually the through hole ratio is required to be ≥95%). (2) Blockage defect location: For micropillars whose segmentation mask area is less than 70% of the designed aperture area, they are identified as blockage defects, and a list of defect coordinates is output for subsequent rework. (3) Manufacturing qualification rate determination: Combine the through hole rate and defect distribution density to determine the manufacturing qualification rate of the batch of workpieces, and provide data support for production process optimization.

[0103] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.

Claims

1. A dense ultrasonic micropillar segmentation and counting method with cross-scale feature interaction, characterized in that: Includes the following steps: Step 1: Acquire an ultrasonic C-scan image of the workpiece to be inspected, and generate a target bounding box prompt on the ultrasonic C-scan image; Step 2: Input the ultrasound C-scan image into the pre-constructed background suppression hierarchical feature extraction network BHiera to extract multiple micropillar feature maps at different levels, including shallow feature maps, middle feature maps, and deep feature maps; Step 3: Input the multiple micro-pillar feature maps of different levels into the scale-specific query encoder to generate query description vectors of the corresponding scales and output multiple scale-specific feature maps containing dense information. Step 4: Input the multiple scale-specific feature maps into the pre-constructed cross-scale feature weighted fusion network CSWF to perform adaptive fusion of deep and shallow features, and output the fused multi-scale features; Step 5: Input the fused multi-scale features and the target bounding box prompt information into the mask decoding network to generate an instance segmentation mask for each micropillar, and output the segmentation count result based on the instance segmentation mask; In step two, the background suppression hierarchical feature extraction network BS Hiera includes multiple BS HieraBlocks connected in sequence; each BS Hiera Block includes a multi-head self-attention module MHA, a background suppression layer BSLayer, skip connections, and an MLP layer connected in sequence. The background suppression layer (BS Layer) includes a target boundary extraction branch and a spatial context branch set in parallel, as well as a feature fusion unit; The target boundary extraction branch is used to extract edge features in the H and W directions of the input features using the Sobel operator, and to perform weighted enhancement on the edge features to output enhanced edge features; The spatial context branch is used to extract spatial context feature weights using global average pooling and nonlinear transformation based on KAN linear layers, and to perform weighted fusion of input features according to the spatial context feature weights to suppress background interference and output background suppression features. The feature fusion unit is used to fuse the enhanced edge features, the background suppression features, and the input features, and output the background suppression layer (BS Layer) features after adjusting the channels through a 1×1 convolution. The method for weighted enhancement of the edge features in the target boundary extraction branch is as follows: in: For input features; Input features The feature gradient of the c-th channel is in the H direction; Input features The feature gradient of the c-th channel in the W direction; Input features Features of the c-th channel; and These are the Sobel gradient operator kernels for the H and W directions, respectively; This represents the gradient magnitude. This represents the batch size; The number of input feature channels; The height of the feature map; The width of the feature map; For the gradient magnitude Perform a 1×1 convolution operation to obtain the enhanced edge features: in: To enhance edge features; This represents a 1×1 convolution operation.

2. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 1, characterized in that: The method for extracting the spatial context feature weights in the spatial context branch is as follows: in: The output characteristics of the average pooling operation; For input features; and These are the height and width of the feature map, respectively; Nonlinear transformation using KAN linear layers: in: For the spline nonlinear function of the KAN linear layer; For linear spline basis functions of the KAN linear layer; These are learnable control point parameters; The nonlinear transformation output features of the KAN linear layer; This indicates that the KAN linear layer is affected by the input features. After performing nonlinear transformation, the first One output dimension; The feature dimension represents the input feature. Calculate the spatial context feature weights: Weight the input features: in: Background suppression features; The boundary weights of the micro-pillars.

3. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 1, characterized in that: In step four, the cross-scale feature weighted fusion network CSWF includes a feature alignment module, a dual-path bidirectional feature fusion module DBFF, and a feature enhancement module based on structural reparameterization. The feature alignment module is used to downsample and align the channel dimensions of the shallow feature map and the middle feature map to adapt to cross-scale feature fusion. The dual-path bidirectional feature fusion module DBFF is used to receive aligned shallow features, mid-level features, and deep features, calculate the fusion weight of cross-layer features, and perform weighted fusion of the shallow features, mid-level features, and deep features according to the fusion weight to realize cross-layer feature interaction. The feature enhancement module based on structural reparameterization is used to enhance the features of each layer after interaction, and then splice the enhanced features of each layer to output the fused multi-scale features.

4. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 3, characterized in that: The dual-path bidirectional feature fusion module DBFF includes a first side branch, a second side branch, and a middle branch; The first side branch is used to retain the feature information of the first input feature; The second side branch is used to retain the feature information of the second input feature; The intermediate branch is used to learn the weights of the concatenated first input feature and second input feature, and output the first fusion weight corresponding to the first input feature and the second fusion weight corresponding to the second input feature; The output features of the dual-path bidirectional feature fusion module DBFF are: in: The output features of the dual-path bidirectional feature fusion module DBFF; and These are two input features; This represents a 1×1 convolution operation; and These are the first fusion weight and the second fusion weight, respectively, and: in: This is a global average pooling operation; Calculated for MLP; For the Softmax function; .

5. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 1, characterized in that: The scale-specific query encoder is constructed using a KAN network, where each weight parameter is replaced by a learnable one-dimensional spline function to model the nonlinear scale variation relationship of multi-scale targets; the mask decoding network uses a SAM2Mask Decode network.

6. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 1, characterized in that: The training process employs a total loss function jointly optimized by Focal loss and counting loss. The total loss function is as follows: in: For count loss; Loss due to counting small targets; Focal loss; Weights for counting small targets.

7. The dense ultrasonic micropillar segmentation and counting method based on cross-scale feature interaction according to claim 1, characterized in that: In step five, the method for outputting the segmentation count result based on the instance segmentation mask is as follows: The pixel area of ​​the instance segmentation mask for each micropillar is counted and compared with the theoretical pixel area of ​​the micropillar's designed aperture to calculate the single-aperture throughput. The total number of micropillar vias is calculated based on the number of segmented masks described in the example. Based on the single-hole through-hole ratio and / or the total number of through-holes in the micropillars, output the through-hole ratio statistics, blockage defect location results, and / or manufacturing pass rate determination results of the workpiece.

8. A system for implementing the dense ultrasonic micropillar segmentation and counting method for cross-scale feature interaction as described in any one of claims 1-7, characterized in that: include: The image acquisition module is used to acquire an ultrasonic C-scan image of the workpiece to be inspected, and generate a target box prompt information on the ultrasonic C-scan image; The feature extraction module is used to input the ultrasound C-scan image into a pre-constructed background suppression hierarchical feature extraction network BS Hiera to extract multiple micropillar feature maps at different levels, including shallow feature maps, middle feature maps and deep feature maps. The scale encoding module is used to input the multiple micro-pillar feature maps of different levels into the scale-specific query encoder, generate query description vectors of corresponding scales, and output multiple scale-specific feature maps containing dense information. The cross-scale fusion module is used to input the multiple scale-specific feature maps into a pre-constructed cross-scale feature weighted fusion network (CSWF) to perform adaptive fusion of deep and shallow features and output the fused multi-scale features. The segmentation counting module is used to input the fused multi-scale features and the target bounding box prompt information into the mask decoding network, generate an instance segmentation mask for each micropillar, and output the segmentation counting result based on the instance segmentation mask.

Citation Information

Patent Citations

  • Automatic image segmentation technology based on convolutional neural network

    CN120807537A

  • Small target detection system and method based on multi-modal feature fusion and fine-grained modeling

    CN122200227A