Silicon wafer surface defect detection method and device based on multi-scale feature fusion, and medium

Through the multi-scale feature fusion silicon wafer surface defect detection method, cross-level polymerization processing and multi-scale fusion technology are used to solve the problems of low detection accuracy and insufficient generalization ability in the existing technology, and efficient and accurate detection of silicon wafer surface defects is achieved.

CN120471918AActive Publication Date: 2025-08-12ZHEJIANG UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510969275.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-12
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

The prior art has the loss of micro defect feature information, flooding of redundant information, and lack of effective fusion of information of different scales and levels in the detection of surface defects of silicon wafers, resulting in low detection accuracy and insufficient generalization ability.

Method used

The silicon wafer surface defect detection method based on multi-scale feature fusion is adopted, and feature extraction and fusion is performed through cross-level aggregation processing. Multiple feature maps with different resolutions are obtained using the backbone network, neck network and detection head to realize the retention and integration of multi-scale information.

Benefits of technology

It improves the accuracy of surface defect detection of silicon wafers, ensures timely retention and fusion of multi-level information in the feature extraction stage, enhances the model's learning ability of multi-dimensional intrinsic feature information, and improves the ability to detect multi-scale defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471918A_ABST
    Figure CN120471918A_ABST
Patent Text Reader

Abstract

The invention discloses a silicon wafer surface defect detection method and device based on multi-scale feature fusion, and a medium. The method comprises the following steps: obtaining a to-be-detected silicon wafer surface image and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model comprises a backbone network, a neck network and a detection head which are connected in sequence; inputting the silicon wafer surface image into a silicon wafer surface defect detection model; wherein a backbone network in the silicon wafer surface defect detection model is used for performing feature extraction on the silicon wafer surface image through cross-level aggregation processing to obtain at least four first effective feature maps with different resolutions; the neck network is used for performing feature fusion on the first effective feature maps through cross-level aggregation processing to obtain at least four second effective feature maps with different resolutions; the detection head is used for detecting each second effective feature pattern; and obtaining the position and the type of the silicon wafer surface defect output by the silicon wafer surface defect detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of defect detection, and in particular to a method, equipment, and medium for detecting surface defects of silicon wafers based on multi-scale feature fusion. Background Art

[0002] In modern semiconductor manufacturing, silicon wafers, the core carriers of integrated circuits and microelectronic devices, have a surface quality that directly determines the performance, reliability, and lifespan of these electronic products. Any tiny defect can trigger a chain reaction in the subsequent manufacturing process, leading to decreased device performance, increased failure rates, and even the scrapping of entire batches of products. This can cause significant economic losses and waste of resources, and in severe cases, can even lead to major accidents. Therefore, timely and accurate detection of silicon wafer surface defects is crucial during the manufacturing process.

[0003] However, current defect detection on silicon wafer surfaces relies primarily on manual identification methods, which requires extensive inspection experience and a strong sense of responsibility. This leads to drawbacks such as low inspection accuracy, long inspection times, and a significant labor investment. Therefore, an efficient, high-precision automated inspection method is urgently needed to quickly and accurately identify defects on the surface of silicon wafers.

[0004] With the continuous development of deep learning and machine vision technologies, neural network-based target detection algorithms can now be used to automatically detect defects on silicon wafer surfaces. For example, Chinese patent CN119359702A discloses a YOLOv8-based silicon wafer surface defect detection method. This method extracts features from input images using a backbone network, fuses features using a Neck structure, and identifies target defects in the image using a Head structure, completing the detection process.

[0005] However, this method of using a neural network-based target detection algorithm to detect silicon wafer surface defects has the following problems, which in turn prevents it from achieving good detection performance:

[0006] (1) In the target detection algorithm, as the neural network deepens, it is necessary to continuously downsample the feature map to extract the required features, which will cause the feature information of small defects to be lost during the downsampling process. At the same time, the introduction of other redundant information will also overwhelm the feature information of small defects, leading to information confusion.

[0007] (2) In the feature extraction process, convolution operations are performed only with fixed-size convolution kernels, focusing only on the extraction of local information. There is a lack of mining of long-distance contextual semantic information rich in large-scale defects, resulting in insufficient perception of deep intrinsic features.

[0008] (3) The lack of effective fusion of information at different scales and levels makes it difficult to simultaneously detect defects of different sizes and shapes, and thus the generalization ability is insufficient when detecting defects with large differences in size, shape, and distribution. Summary of the Invention

[0009] In response to the shortcomings of the existing technology, the present invention provides a silicon wafer surface defect detection method, equipment, and medium based on multi-scale feature fusion to improve the detection accuracy of silicon wafer surface defects.

[0010] In a first aspect, an embodiment of the present invention provides a method for detecting surface defects of a silicon wafer based on multi-scale feature fusion, the method comprising the following steps:

[0011] Obtaining a surface image of a silicon wafer to be inspected and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model includes a backbone network, a neck network, and a detection head connected in sequence;

[0012] The silicon wafer surface image is input into the silicon wafer surface defect detection model; wherein the backbone network in the silicon wafer surface defect detection model is used to extract features from the silicon wafer surface image through cross-level aggregation processing to obtain at least four first effective feature maps with different resolutions; the neck network is used to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least four second effective feature maps with different resolutions; the detection head is used to detect each second effective feature map separately;

[0013] The locations and types of silicon wafer surface defects output by the silicon wafer surface defect detection model are obtained.

[0014] In a second aspect, an embodiment of the present invention provides an electronic device comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned silicon wafer surface defect detection method based on multi-scale feature fusion.

[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned silicon wafer surface defect detection method based on multi-scale feature fusion.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned silicon wafer surface defect detection method based on multi-scale feature fusion.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] The present invention provides a silicon wafer surface defect detection method based on multi-scale feature fusion, by inputting a silicon wafer surface image into a silicon wafer surface defect detection model; wherein, the backbone network in the silicon wafer surface defect detection model is used to perform feature extraction on the silicon wafer surface image through cross-level aggregation processing to obtain at least 4 first effective feature maps with different resolutions; the neck network is used to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least 4 second effective feature maps with different resolutions; the detection head is used to detect each second effective feature map separately; the present invention uses cross-level aggregation processing to enable the feature map to obtain and integrate more information of different dimensions and abstract levels in the perception domain, ensuring that multi-level information can be retained and fused in time at each stage of feature extraction, promoting the flow of multi-scale features, deepening the model's learning ability of the multi-dimensional intrinsic feature information of defects, and enabling the feature maps at each stage to establish rich multi-scale information dependence, effectively solving the problem of multi-scale feature information being lost in the model, and improving the detection accuracy of silicon wafer surface defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A flow chart of a silicon wafer surface defect detection method based on enhanced multi-scale feature fusion provided by an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of a silicon wafer surface defect detection model provided by an embodiment of the present invention;

[0022] Figure 3 A flowchart of cross-level aggregation processing provided by an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of cross-level aggregation processing provided by an embodiment of the present invention;

[0024] Figure 5 A flowchart of multi-scale fusion processing provided by an embodiment of the present invention;

[0025] Figure 6 A schematic diagram of multi-scale fusion processing provided by an embodiment of the present invention;

[0026] Figure 7 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0028] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0029] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0030] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.

[0031] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting silicon wafer surface defects based on multi-scale feature fusion, the method comprising the following steps:

[0032] Step S1, obtaining a silicon wafer surface image to be detected and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model includes a backbone network, a neck network, a detection head ( Figure 2 ).

[0033] Step S2, inputting the silicon wafer surface image into the silicon wafer surface defect detection model; wherein, the backbone network in the silicon wafer surface defect detection model is used to extract features of the silicon wafer surface image through cross-level aggregation processing to obtain at least 4 first effective feature maps with different resolutions; the neck network is used to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least 4 second effective feature maps with different resolutions; the detection head is used to detect each second effective feature map separately.

[0034] Among them, such as Figure 3 As shown in Figure 2, the cross-level aggregation process includes:

[0035] Extract features from the input feature map and divide it into the first feature map X1 and the second feature map X2;

[0036] Performing three-fold cascade multi-scale fusion processing on the first feature map X1 to obtain a third intermediate feature map; including: performing multi-scale fusion processing on the first feature map X1 to obtain a first intermediate feature map; performing multi-scale fusion processing on the first intermediate feature map to obtain a second intermediate feature map; performing multi-scale fusion processing on the second intermediate feature map to obtain a third intermediate feature map;

[0037] Concatenate the first feature map X1, the first intermediate feature map, the second intermediate feature map, the third intermediate feature map, and the second feature map X2 to obtain the third feature map X3;

[0038] The third feature map X3 is convolved to obtain an output feature map.

[0039] Specifically, if Figure 4 As shown, in this example, the cross-level aggregation processing process specifically includes:

[0040] The input feature map X is extracted using a standard CBS module consisting of a 1×1 convolution operation, batch normalization, and SiLU activation function. It is then split along the channel dimension into the first feature map X1 and the second feature map X2. The expression is as follows:

[0041]

[0042] Where, Indicates separation operation along the channel, Represents a standard CBS module with 1×1 convolution operation, batch normalization and SiLU activation function; the first feature map As the main part, multi-dimensional information is collected and integrated, and the second feature map It is used as the expression of the stable feature of the residual part.

[0043] Performing three-fold cascade multi-scale fusion processing on the first feature map X1 to obtain a third intermediate feature map; including: performing multi-scale fusion processing on the first feature map X1 to comprehensively capture and fuse rich multi-scale information to obtain a first intermediate feature map; performing multi-scale fusion processing on the first intermediate feature map to extract higher-level abstract multi-scale features to obtain a second intermediate feature map; performing multi-scale fusion processing on the second intermediate feature map to obtain a third intermediate feature map;

[0044] The first feature map X1, the first intermediate feature map, the second intermediate feature map, the third intermediate feature map, and the second feature map X2 are concatenated to achieve cross-dimensional information fusion, thereby enriching the expression and cardinality of the features and obtaining the third feature map X3; the expression is as follows:

[0045]

[0046] Where, and represent the concatenation operation and enhanced multi-scale fusion operation respectively.

[0047] The third feature map X3 is processed using a standard CBS module consisting of a 1×1 convolution operation, batch normalization, and SiLU activation function to obtain the output feature map ; The expression is as follows:

[0048]

[0049] Where, Represents a standard CBS module with 1×1 convolution operation, batch normalization and SiLU activation function.

[0050] It should be noted that feature maps at different levels have different representations of multi-scale information. This example uses three cascades of multi-scale fusion processing to gradually perceive and integrate contextual semantic information, further deepening the model's ability to learn multi-dimensional feature information. This allows feature maps at each stage to establish rich multi-scale information dependencies, thereby constructing multi-scale feature representations of the input feature map at different levels of abstraction.

[0051] Among them, such as Figure 5 As shown in Figure 2, the multi-scale fusion processing process includes:

[0052] The input feature map M is split into 4 branches; the first branch, the second branch, and the third branch are processed by the first convolution, the second convolution, and the third convolution with different convolution kernel sizes respectively to obtain the first scale feature map M1, the second scale feature map M2, and the third scale feature map M3; the fourth branch obtains the fourth scale feature map M4 through the identity mapping; the first scale feature map M1, the second scale feature map M2, the third scale feature map M3, and the fourth scale feature map M4 are spliced to obtain the first fusion feature map M f ;

[0053] Concatenate the input feature map M and the first fusion feature map M f , get the second fusion feature map M s ;

[0054] For the second fusion feature map M s Feature extraction is performed through dilated convolution, and the feature map obtained by dilated convolution is combined with the second fused feature map M s Splice and get the third fusion feature map M l ;

[0055] The third fusion feature map M l After layer normalization, convolution, and GELU activation function processing, it is concatenated with the input feature map M to obtain the multi-scale fusion feature map M'.

[0056] Specifically, if Figure 6 As shown, in this example, the multi-scale fusion processing process specifically includes:

[0057] The input feature map M is split into four branches in proportion in the channel dimension; the first branch is subjected to a depth-separable convolution with a convolution kernel size of 3×3 to extract features, and the first-scale feature map M1 is obtained; the second branch is subjected to a depth-separable convolution with a convolution kernel size of 9×1 to extract features, and the second-scale feature map M2 is obtained; the third branch is subjected to a depth-separable convolution with a convolution kernel size of 1×9 to extract features, and the third-scale feature map M3 is obtained; the fourth branch retains the original features through identity mapping to obtain the fourth-scale feature map M4; the first-scale feature map M1, the second-scale feature map M2, the third-scale feature map M3, and the fourth-scale feature map M4 are spliced to obtain the first fusion feature map M f ; The expression is as follows:

[0058]

[0059] Where, Represents the connection operation between channels, Represents a depthwise separable convolution operation.

[0060] The first fusion feature map M fIt combines rich feature information from different directions and scales, concatenates the input feature map M and the first fused feature map M f , get the second fusion feature map M s ; The expression is as follows:

[0061]

[0062] Where, Represents a pixel-by-pixel matrix addition operation.

[0063] For the second fusion feature map M s Feature extraction is performed through dilated convolution with a kernel size of 7×7 and a separation rate of 3 to capture more high-level semantic and contextual information, thereby significantly enhancing the network's ability to perceive global information; the feature map obtained by dilated convolution is combined with the second fused feature map M s Perform element-by-element addition to obtain the third fusion feature map M with a large receptive field l ; The expression is as follows:

[0064]

[0065] Where, Represents a depthwise atrous convolution operation.

[0066] The third fusion feature map M l After layer normalization, convolution, and GELU activation function processing, it is spliced with the input feature map M. This not only integrates multi-scale feature information, but also effectively alleviates the problems of gradient disappearance and network degradation, thereby obtaining a multi-scale fusion feature map M'; the expression is as follows:

[0067]

[0068] Where, represents the GELU activation function, Presentation layer standardization.

[0069] It should be noted that this example utilizes large kernel convolutions with multiple shapes to provide the model with a wide, all-round, and multi-angle receptive field. It re-establishes long-range feature associations in both the horizontal and vertical directions, comprehensively aggregating effective contextual information to associate multi-scale defect features, and enhancing sensitivity to defects with large length-width differences (such as long scratches and cracks on silicon wafers), enabling the model to "see" more broadly. Small kernel convolutions are used to efficiently extract more complex spatial representations and more detailed local representations for the model, enabling deep perception and interaction of feature information, enhancing perception of subtle defects (such as fine particles and subtle dirt on the silicon wafer surface), and enabling the model to "see" more deeply and finely. Multi-scale fusion processing enhances the richer representation and broader patterns of feature information in multi-scale space, enabling the model to go beyond texture features when detecting defects and gain a comprehensive understanding of overall shape information.

[0070] Exemplarily, the size of the input feature map is 448×448×3. After passing through the backbone network based on cross-level aggregation operation, four first effective feature maps with different resolutions are obtained, whose sizes are 112×112×128, 56×56×256, 28×28×512, and 14×14×512 respectively.

[0071] Furthermore, in this example, the feature pyramid network is reconstructed with cross-level aggregation processing operations as the core. Using upsampling and concatenation operations, the deep semantic information and discriminant features contained in the low-resolution feature map are transferred to the high-resolution feature map, and the full fusion of multi-scale features is achieved with the help of cross-level aggregation processing operations. Next, the high-resolution feature map is processed with 3×3 convolution, batch normalization, and SiLU activation function to implement downsampling operations, and the precise defect location information and local details therein are fed back to the low-resolution feature map. This allows the detail advantages of low-level features to be perfectly combined with the semantic expression of high-level features, thereby constructing a detection network that performs well in both defect recognition and positioning.

[0072] It is worth noting that under the cross-level aggregation processing, the feature pyramid network can maintain full perception of small-scale defects (such as particles, protrusions, etc.) and large-scale defects (such as scratches, cracks, etc.) on the wafer at all stages of feature extraction, and also ensure the continuous flow and propagation of discernible information of multi-scale defects of silicon wafers in the detection network, avoiding the loss of small-scale defect features such as particles in low-resolution feature maps and the loss of large-scale defect features such as scratches in high-resolution feature maps. At the same time, it can obtain and integrate more information of different dimensions and abstract levels in the perception domain, thereby improving the model's generalization ability for multiple scales.

[0073] Ultimately, the feature pyramid network can output at least four second effective feature maps with different resolutions, providing strong multi-scale support for subsequent defect detection tasks. At the same time, this example uses a decoupled detection head to predict each second effective feature map of different resolutions generated by the feature pyramid network. Each detection head independently processes position information and category information, and sets different weights for different categories of defects based on the different detection heads. That is, high-resolution feature maps are highly sensitive to tiny defects such as particles and protrusions, and thus have a larger weight coefficient when predicting such defects; low-resolution feature maps are highly sensitive to large-area defects such as scratches and cracks, and thus have a larger weight coefficient when predicting such defects, effectively improving detection accuracy.

[0074] Step S3: Obtain the location and type of silicon wafer surface defects output by the silicon wafer surface defect detection model.

[0075] Furthermore, the training process of the silicon wafer surface defect detection model includes:

[0076] Step S100 , obtaining a silicon wafer surface defect image, and marking the locations and types of defects in the silicon wafer surface defect image to obtain a silicon wafer surface defect dataset.

[0077] Specifically, in this example, under appropriate lighting conditions, an industrial camera is used to capture and acquire high-quality silicon wafer surface images. Silicon wafer surface images without defects are manually removed from the captured images, and silicon wafer surface images containing defects are retained. The LabelImg annotation tool is used to annotate square frames at the defect locations in the silicon wafer surface images, so that the annotation frames can just cover the complete form of the silicon wafer surface defects, ensuring that the annotation frames neither miss the defective areas nor include the redundant background areas within the rectangular frames. The annotated rectangular frame labels are saved in .txt format to the Annotations folder, and the corresponding silicon wafer surface defect images are saved to the Images folder, thereby forming a silicon wafer surface defect dataset.

[0078] Step S200: construct a joint loss function, where the joint loss function is a weighted sum of the CIoU loss function and the Varifocal loss function.

[0079] In this example, the CIoU loss function was introduced for bounding box regression. This not only considers the overlap between the predicted and true bounding boxes, but also comprehensively factors in the center point distance and aspect ratio, ensuring more accurate positioning. Furthermore, the Varifocal loss function was used to determine defect categories. This loss function balances prediction confidence with classification accuracy, better adapting to uneven sample distribution.

[0080] Step S300 , training a silicon wafer surface defect detection model based on a joint loss function using a silicon wafer surface defect dataset.

[0081] Specifically, in this example, the silicon wafer surface defect dataset constructed in step S100 was randomly divided into a 7:3 ratio to obtain corresponding training and test sets. The silicon wafer surface defect detection model was trained using the training set. During this process, the loss and accuracy changes were observed based on its performance on the test set. The hyperparameters of the silicon wafer surface defect detection model were continuously adjusted until the silicon wafer surface defect detection model reached a fully converged state, and the weights that achieved the best detection performance were selected.

[0082] The optimal model weights obtained through training are deployed as the core parameters of the final silicon wafer surface defect detection model, and are used to accurately identify, locate, and classify silicon wafer surface defects, thereby realizing automated defect detection.

[0083] In summary, the present invention provides a silicon wafer surface defect detection method based on multi-scale feature fusion, by inputting a silicon wafer surface image into a silicon wafer surface defect detection model; wherein, the backbone network in the silicon wafer surface defect detection model is used to perform feature extraction on the silicon wafer surface image through cross-level aggregation processing to obtain at least 4 first effective feature maps with different resolutions; the neck network is used to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least 4 second effective feature maps with different resolutions; the detection head is used to detect each second effective feature map separately; the present invention uses cross-level aggregation processing to enable the feature map to obtain and integrate more information of different dimensions and abstract levels in the perception domain, ensuring that multi-level information can be retained and fused in time at each stage of feature extraction, promoting the flow of multi-scale features, deepening the model's learning ability of the multi-dimensional intrinsic feature information of defects, and enabling the feature maps at each stage to establish rich multi-scale information dependence, effectively solving the problem of multi-scale feature information being lost in the model, and improving the detection accuracy of silicon wafer surface defects.

[0084] On the other hand, an embodiment of the present invention provides a silicon wafer surface defect detection system based on multi-scale feature fusion, the system comprising:

[0085] A detection preparation module is used to obtain a surface image of a silicon wafer to be detected and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model includes a backbone network, a neck network, and a detection head connected in sequence;

[0086] The target detection module is configured to input the silicon wafer surface image into the silicon wafer surface defect detection model; wherein the backbone network in the silicon wafer surface defect detection model is configured to perform feature extraction on the silicon wafer surface image through cross-level aggregation processing to obtain at least four first effective feature maps with different resolutions; the neck network is configured to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least four second effective feature maps with different resolutions; and the detection head is configured to perform detection on each second effective feature map separately;

[0087] The result acquisition module is used to obtain the position and type of silicon wafer surface defects output by the silicon wafer surface defect detection model.

[0088] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0089] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0090] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the silicon wafer surface defect detection method based on multi-scale feature fusion as described above. Figure 7 As shown in FIG, a hardware structure diagram of any device with data processing capability in which the silicon wafer surface defect detection method based on multi-scale feature fusion provided by an embodiment of the present invention is located, except Figure 7 In addition to the processor, memory, and network interface shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0091] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the silicon wafer surface defect detection method based on multi-scale feature fusion as described above. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.

[0092] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only.

[0093] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A silicon wafer surface defect detection method based on multi-scale feature fusion, characterized in that: The method comprises the following steps: Obtaining a surface image of a silicon wafer to be inspected and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model includes a backbone network, a neck network, and a detection head connected in sequence; The silicon wafer surface image is input into the silicon wafer surface defect detection model; wherein the backbone network in the silicon wafer surface defect detection model is used to extract features from the silicon wafer surface image through cross-level aggregation processing to obtain at least four first effective feature maps with different resolutions; the neck network is used to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least four second effective feature maps with different resolutions; the detection head is used to detect each second effective feature map separately; The locations and types of silicon wafer surface defects output by the silicon wafer surface defect detection model are obtained.

2. The method for detecting silicon wafer surface defects based on multi-scale feature fusion according to claim 1, characterized in that: The cross-level aggregation process includes: Extract features from the input feature map and divide it into the first feature map X1 and the second feature map X2; Perform three-fold cascade multi-scale fusion processing on the first feature map X1 to obtain the third intermediate feature map; Concatenate the first feature map X1, the first intermediate feature map, the second intermediate feature map, the third intermediate feature map, and the second feature map X2 to obtain the third feature map X3; The third feature map X3 is convolved to obtain an output feature map.

3. The method for detecting silicon wafer surface defects based on multi-scale feature fusion according to claim 2, characterized in that: The process of performing three-fold cascade multi-scale fusion processing on the first feature map X1 to obtain the third intermediate feature map includes: The first feature map X1 is subjected to multi-scale fusion processing to obtain a first intermediate feature map; the first intermediate feature map is subjected to multi-scale fusion processing to obtain a second intermediate feature map; the second intermediate feature map is subjected to multi-scale fusion processing to obtain a third intermediate feature map.

4. A silicon wafer surface defect detection method based on multi-scale feature fusion according to claim 2 or 3, characterized in that: The multi-scale fusion processing process includes: The input feature map M is split into 4 branches; the first branch, the second branch, and the third branch are processed by the first convolution, the second convolution, and the third convolution with different convolution kernel sizes respectively to obtain the first scale feature map M1, the second scale feature map M2, and the third scale feature map M3; the fourth branch obtains the fourth scale feature map M4 through the identity mapping; the first scale feature map M1, the second scale feature map M2, the third scale feature map M3, and the fourth scale feature map M4 are spliced to obtain the first fusion feature map M f ; Concatenate the input feature map M and the first fusion feature map M f , get the second fusion feature map M s ; For the second fusion feature map M s Feature extraction is performed through dilated convolution, and the feature map obtained by dilated convolution is combined with the second fused feature map M s Splice and get the third fusion feature map M l ; The third fusion feature map M l After layer normalization, convolution, and GELU activation function processing, it is concatenated with the input feature map M to obtain the multi-scale fusion feature map M'.

5. The method for detecting silicon wafer surface defects based on multi-scale feature fusion according to claim 4, characterized in that: The process of obtaining the first scale feature map M1, the second scale feature map M2, the third scale feature map M3, and the fourth scale feature map M4 includes: The first branch performs feature extraction through depth-wise separable convolution with a convolution kernel size of 3×3 to obtain the first-scale feature map M1; the second branch performs feature extraction through depth-wise separable convolution with a convolution kernel size of 9×1 to obtain the second-scale feature map M2; the third branch performs feature extraction through depth-wise separable convolution with a convolution kernel size of 1×9 to obtain the third-scale feature map M3; the fourth branch retains the original features through identity mapping to obtain the fourth-scale feature map M4.

6. The method for detecting silicon wafer surface defects based on multi-scale feature fusion according to claim 1, characterized in that: The training process of the silicon wafer surface defect detection model includes: Obtain silicon wafer surface defect images, and mark the locations and types of defects in the silicon wafer surface defect images to obtain a silicon wafer surface defect dataset; Construct a joint loss function, which is a weighted sum of the CIoU loss function and the Varifocal loss function; The silicon wafer surface defect detection model is trained based on the joint loss function using the silicon wafer surface defect dataset.

7. A silicon wafer surface defect detection system based on multi-scale feature fusion, characterized in that: The system comprises: A detection preparation module is used to obtain a surface image of a silicon wafer to be detected and a pre-trained silicon wafer surface defect detection model; wherein the silicon wafer surface defect detection model includes a backbone network, a neck network, and a detection head connected in sequence; The target detection module is configured to input the silicon wafer surface image into the silicon wafer surface defect detection model; wherein the backbone network in the silicon wafer surface defect detection model is configured to perform feature extraction on the silicon wafer surface image through cross-level aggregation processing to obtain at least four first effective feature maps with different resolutions; the neck network is configured to perform feature fusion on the first effective feature map through cross-level aggregation processing to obtain at least four second effective feature maps with different resolutions; and the detection head is configured to perform detection on each second effective feature map separately; The result acquisition module is used to obtain the position and type of silicon wafer surface defects output by the silicon wafer surface defect detection model.

8. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the silicon wafer surface defect detection method based on multi-scale feature fusion as described in any one of claims 1-7 above.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the silicon wafer surface defect detection method based on multi-scale feature fusion as described in any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the silicon wafer surface defect detection method based on multi-scale feature fusion described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Wafer surface defect detection method and system based on improved YOLO network

    CN113222982A

  • Wafer defect detection method and device based on lightweight target detection model

    CN118334032A

  • Wafer defect detection method and system for IGP processing unit, and storage medium

    CN119205764A

  • Lightweight wafer defect detection method

    CN119359702A

  • Small target detection method and system based on feature fusion

    CN119418037A