Part surface defect detection method and system based on Transform model
Through the part surface defect detection method based on the Transformer model, the problem of inability to effectively identify small size defects and low detection efficiency in the prior art is solved, and high-precision and high-efficiency defect detection are achieved.
Patent Information
- Application Number
- CN202510181429.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively identify small size defects in part surface defect detection, and the detection efficiency is low, which cannot meet the real-time inspection requirements of industrial sites.
Using the part surface defect detection method based on the Transformer model, by obtaining the image to be detected and performing grayscale processing and filtering, different frequency parts of the grayscale image are extracted, frequency domain enhancement and image reconstruction are performed, and finally the processed images are input to the pre-trained Transformer model for defect recognition.
It improves the recognition accuracy and robustness of micro defects, enhances detection efficiency, and can meet the real-time detection needs of industrial sites.
Smart Images

Figure CN120107207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and in particular to a method and system for detecting surface defects of parts based on a Transformer model. Background Art
[0002] In the manufacturing process of industrial parts, effective defect detection, especially small defects such as microcracks, pores, and fine scratches caused by material properties, processing technology, processing wear or environmental factors, is of great practical significance to ensure the safe operation of equipment in subsequent practical applications and avoid corresponding economic losses.
[0003] In the prior art, surface defect detection of parts is mainly carried out through manual visual inspection or automatic detection based on image processing. Among them, although manual visual inspection is simple to operate and does not require additional hardware support, it has low detection efficiency, high fatigue, strong subjectivity, and it is difficult to ensure consistency; and it is impossible to detect tiny defects. Although automatic detection based on image processing avoids the above-mentioned defects of manual visual inspection, it can only be applied to defect detection under simple backgrounds or large-size defect detection; and when the defect contrast is low, the shape is complex or it is seriously interfered by noise, the automatic detection based on image processing often has reduced accuracy and significantly increased false detection and missed detection rates, which is difficult to meet the needs of various defect detections in industrial sites, especially tiny defect detection.
[0004] Although there are related studies that apply convolutional neural networks to defect detection on the surface of parts in order to improve detection accuracy. However, the convolutional neural network is simply combined with surface detection without considering the various defects that exist in actual applications. Specifically, first, there are still limitations in processing high-resolution images and capturing global information, so it is not possible to effectively identify tiny defects shown in the image, resulting in missed detection. Secondly, the inference process of the convolutional neural network model consumes a lot of computing power, which makes it difficult to meet the real-time detection needs of a large number of parts in industrial sites. Summary of the invention
[0005] The purpose of the present invention is to provide a part surface defect detection method and system based on the Transformer model, so as to solve the technical problems that the convolutional neural network used in the prior art for part surface defects cannot effectively detect tiny defects, has low detection efficiency, and is difficult to meet actual detection needs.
[0006] To achieve the above object, the present invention proposes the following technical solutions:
[0007] In the first aspect, the technical solution provides a method for detecting surface defects of parts based on a Transformer model, comprising:
[0008] Acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and perform grayscale processing on each image to be detected to obtain each grayscale image;
[0009] After low-pass filtering and high-pass filtering are respectively performed on each grayscale image in the row direction, two-fold downsampling is continued simultaneously to obtain the corresponding row low-frequency part and row high-frequency part respectively; after low-pass filtering and high-pass filtering are respectively performed on the low-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding LL original subband and LH original subband respectively; after low-pass filtering and high-pass filtering are respectively performed on the high-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding HL original subband and HH original subband respectively;
[0010] Among them, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction;
[0011] Introducing preset fixed gain factors into the LH original subband, the HL original subband, and the HH original subband to obtain first gain coefficients at various positions in the original subbands, calculating the sum of squares of the first gain coefficients at various positions in the original subbands as energy values of the corresponding original subbands; calculating dynamic gain factors corresponding to the corresponding original subbands based on the energy values; and performing frequency domain enhancement on the LH original subband, the HL original subband, and the HH original subband based on the corresponding fixed gain factors and the dynamic gain factors to obtain LH gain subbands, HL gain subbands, and HH gain subbands respectively;
[0012] The dynamic gain factor is Among them, α is the basic weight constant; E ref is the reference energy value; ∈ is an infinitesimal constant used to prevent the denominator from being zero; E d is the energy value;
[0013] The LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband are sequentially subjected to inverse column filtering operations and inverse row filtering operations, and then image reconstruction is performed to obtain a spatial domain image; the spatial domain image is fused with the grayscale image to obtain a target image;
[0014] The target image is subjected to patch cutting and position encoding optimization and then input into a pre-trained Transformer model to output defect recognition results on the surface of the part to be inspected.
[0015] Furthermore, the step of acquiring a plurality of images to be detected corresponding to the surface of the part to be detected includes:
[0016] Acquire a number of original images corresponding to the surface of the part to be inspected based on a high-resolution camera;
[0017] Performing format processing on each original image to obtain each intermediate image; wherein the resolution and size of each intermediate image remain consistent;
[0018] Based on the filtering algorithm, each intermediate image is randomly denoised to obtain each image to be detected.
[0019] Further, after low-pass filtering and high-pass filtering are respectively performed on each grayscale image in the row direction, two-fold downsampling is continued simultaneously to obtain the corresponding row low-frequency part and row high-frequency part respectively; after low-pass filtering and high-pass filtering are respectively performed on the low-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding LL original subband and LH original subband respectively; after low-pass filtering and high-pass filtering are respectively performed on the high-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding HL original subband and HH original subband respectively, including:
[0020] The grayscale image f(i, j) is low-pass filtered and downsampled twice in the row direction based on a low-pass filter to obtain the row low-frequency part At the same time, high-pass filtering and double downsampling are performed based on a high-pass filter in the row direction to obtain the high-frequency part of the row Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index; h low is a low-pass filter, h high is a high pass filter;
[0021] The row low-frequency part is low-pass filtered and down-sampled twice based on a low-pass filter in the column direction to obtain the LL original subband In the column direction, high-pass filtering and double downsampling are performed based on a high-pass filter to obtain the LH original subband
[0022] The high frequency part of the row is low-pass filtered and down-sampled twice based on a low-pass filter in the column direction to obtain the HL original subband And based on the high-pass filter, high-pass filtering and double downsampling are performed to obtain the HH original subband
[0023] Further, the step of performing inverse column filtering and reverse row filtering on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband, respectively, and then reconstructing the image to obtain a spatial domain image includes:
[0024] Performing size processing on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband so that their sizes meet the sampling ratio requirement;
[0025] The LL original subband after the size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the LH gain subband after the size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the HL gain subband after the size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction; the HH gain subband after the size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction;
[0026] Spatial reconstruction is performed to obtain a spatial domain image.
[0027] Furthermore, the target image is subjected to patch cutting and position encoding optimization and then input into a pre-trained Transformer model to output defect recognition results on the surface of the part to be inspected, including:
[0028] Patch cutting the target image according to a preset grid size to obtain a plurality of image blocks;
[0029] Add coordinate position coding to each tile, and add directional labels to each tile based on the defect features in each direction retained in frequency domain enhancement;
[0030] Each tile is expanded in the spatial domain or a simple linear projection is performed to obtain the corresponding feature vector, and the corresponding coordinate position code and directional label are added to each feature vector to obtain the embedding vector corresponding to each tile;
[0031] Based on the coordinate position encoding, each embedding vector is input into the Transformer model in sequence to obtain the defect recognition result under the guidance of the self-attention mechanism and directional labels.
[0032] In the second aspect, the technical solution provides a part surface defect detection system based on the Transformer model, including:
[0033] An image acquisition module is used to acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and grayscale process each image to be detected to acquire each grayscale image;
[0034] A filtering processing module is used to perform low-pass filtering and high-pass filtering on each grayscale image in the row direction, and then continue to perform two-fold downsampling to obtain the corresponding row low-frequency part and row high-frequency part respectively; perform low-pass filtering and high-pass filtering on the low-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding LL original subband and LH original subband respectively; perform low-pass filtering and high-pass filtering on the high-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding HL original subband and HH original subband respectively;
[0035] Among them, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction;
[0036] A gain processing module, used for introducing preset fixed gain factors into the LH original subband, the HL original subband, and the HH original subband to obtain a first gain coefficient at each position in each original subband, calculating the sum of squares of the first gain coefficient at each position in each original subband as an energy value of the corresponding original subband; calculating a dynamic gain factor corresponding to the corresponding original subband based on the energy value; and performing frequency domain enhancement on the LH original subband, the HL original subband, and the HH original subband based on the corresponding fixed gain factor and the dynamic gain factor to obtain a LH gain subband, a HL gain subband, and a HH gain subband respectively;
[0037] The dynamic gain factor is Among them, α is the basic weight constant; E ref is the reference energy value; ∈ is an infinitesimal constant used to prevent the denominator from being zero; E d is the energy value;
[0038] an inverse filtering processing module, configured to perform inverse column filtering operations and inverse row filtering operations on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband, respectively, and then reconstruct the image to obtain a spatial domain image; and perform a fusion operation on the spatial domain image and the grayscale image to obtain a target image;
[0039] The defect recognition module is used to perform patch cutting and position coding optimization on the target image and then input the image into a pre-trained Transformer model to output the defect recognition result of the surface of the part to be inspected.
[0040] Furthermore, the image acquisition module includes:
[0041] An acquisition unit, used for acquiring a plurality of original images corresponding to the surface of the part to be inspected based on a high-resolution camera;
[0042] A first processing unit is used to perform format processing on each original image to obtain each intermediate image; wherein the resolution and size of each intermediate image are kept consistent;
[0043] The second processing unit is used to perform random denoising on each intermediate image based on a filtering algorithm to obtain each image to be detected.
[0044] Furthermore, the filtering processing module includes:
[0045] The third processing unit is used to perform low-pass filtering and double downsampling on the grayscale image f(i, j) in the row direction based on a low-pass filter to obtain the row low-frequency part At the same time, high-pass filtering and double downsampling are performed based on a high-pass filter in the row direction to obtain the high-frequency part of the row Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index; h low is a low-pass filter, h high is a high-pass filter;
[0046] The fourth processing unit is used to perform low-pass filtering and double downsampling on the row low-frequency part in the column direction based on a low-pass filter to obtain an LL original subband In the column direction, high-pass filtering and double downsampling are performed based on a high-pass filter to obtain the LH original subband
[0047]
[0048] The fifth processing unit is used to perform low-pass filtering and double downsampling on the high-frequency part of the row in the column direction based on a low-pass filter to obtain the HL original subband And based on the high-pass filter, high-pass filtering and double downsampling are performed to obtain the HH original subband
[0049]
[0050] In a third aspect, the technical solution provides an electronic device, comprising at least one processor, wherein the processor is coupled to a memory, wherein a computer program is stored in the memory, and wherein the computer program is configured to execute the method described when executed by the processor.
[0051] In a fourth aspect, the present technical solution provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to execute the described method.
[0052] Beneficial effects:
[0053] It can be seen from the above technical solutions that the technical solution of the present invention provides a part surface defect detection method based on the Transformer model to solve the detection defects existing in various defect detection means in the prior art, especially the problem that tiny size defects cannot be effectively detected.
[0054] Specifically, first, a number of images to be detected corresponding to the surface of the part to be detected are obtained, and each image to be detected is gray-scale processed to obtain each gray-scale image. In order to effectively identify defects in gray-scale images, especially small defects, the following steps are continued: After low-pass filtering and high-pass filtering are performed on each gray-scale image in the row direction, two times downsampling is continued at the same time to obtain the corresponding row low-frequency part and row high-frequency part respectively; after low-pass filtering and high-pass filtering are performed on the low-frequency part in the column direction, two times downsampling is continued at the same time to obtain the corresponding LL original sub-band and LH original sub-band respectively; after low-pass filtering and high-pass filtering are performed on the high-frequency part in the column direction, two times downsampling is continued at the same time to obtain the corresponding HL original sub-band and HH original sub-band respectively. At this time, the overall brightness distribution and macroscopic structure will be retained through the LL original sub-band, so as to effectively obtain global information. And the LH original sub-band, the HL original sub-band and the HH original sub-band are used to record the detailed information in the horizontal, longitudinal and diagonal directions respectively. Specifically, these detailed information will contain defect features such as small cracks and fine scratches. Continuing, each original subband carrying detail information is enhanced to obtain each enhanced subband, thereby further highlighting the defect features. Specifically, considering that different industrial parts have different defect characteristics and different detection requirements, the corresponding original subbands are enhanced by solid-state gain factors and dynamic gain mechanisms based on dynamic gain factors. At this time, the enhancement operation can be adaptive to different frequency characteristics and energy distributions by adjusting the basic weight constants and energy parameters in the dynamic gain factors, while suppressing noise and background textures, ensuring that the defect features are effectively amplified. Then, the corresponding enhanced subbands and LL original subbands are spatially reconstructed and fused with the original grayscale image to obtain a target image with frequency domain enhancement characteristics. When defect recognition is performed based on the target image, considering that the Transformer model has a self-attention mechanism on the one hand, it can capture long-distance dependencies in the image and focus on global information, so it is adaptable to small and scattered defects; on the other hand, it is more flexible when processing high-resolution images and subtle structures under complex backgrounds, and can perform global modeling for defects of different scales and shapes. Then it is determined to use the Transformer model to recognize the target image. In the specific recognition process, patch cutting and position encoding optimization are also introduced to make the enhanced high-frequency features more closely integrated with the global semantic information, thereby improving the recognition accuracy and robustness of minor defects.
[0055] It should be appreciated that all combinations of the foregoing concepts, as well as additional concepts described in greater detail below, may be considered to be part of the inventive subject matter of the present disclosure, provided such concepts are not mutually inconsistent.
[0056] The foregoing and other aspects, embodiments and features of the present invention can be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the present invention, such as the features and / or beneficial effects of the exemplary embodiments, will be apparent from the following description or learned from the practice of the specific embodiments according to the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in various figures may be represented by the same reference numeral. For clarity, not every component is labeled in every figure. Embodiments of various aspects of the present invention will now be described by way of example and with reference to the accompanying drawings, in which:
[0058] Figure 1 Flow chart of the part surface defect detection method based on the Transformer model described in this embodiment;
[0059] Figure 2 This is a flow chart of acquiring an image to be detected as described in this embodiment;
[0060] Figure 3 This is a flow chart of acquiring each original sub-band according to this embodiment;
[0061] Figure 4 This is a flow chart of acquiring spatial domain images according to this embodiment;
[0062] Figure 5 A flowchart of defect identification using the Transformer model described in this embodiment;
[0063] Figure 6 is a structural block diagram of the part surface defect detection system based on the Transformer model described in this embodiment;
[0064] Figure 7 is a structural block diagram of the electronic device described in this embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meaning understood by people with general skills in the field to which the present invention belongs.
[0066] The words "first", "second" and similar words used in the specification and claims of this application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, unless the context clearly indicates otherwise, the singular form of "a", "an" or "the" and other similar words do not indicate a quantitative limitation, but rather indicate the presence of at least one. Words such as "include" or "comprise" and the like mean that the elements or objects appearing before "include" or "comprise" cover the features, wholes, steps, operations, elements and / or components listed after "include" or "comprise", and do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0067] During the manufacturing process of industrial parts, various types of tiny defects often occur on the surface due to material properties, processing technology, wear or environmental factors. Although these defects are small in size, they may cause stress concentration, fatigue cracking and performance degradation in the subsequent use of the parts, and even cause safety accidents and economic losses. Therefore, how to quickly and accurately detect and identify tiny defects on the surface of industrial parts (such as slewing bearings, bearing housings, gears, etc.) has important practical significance in the field of intelligent manufacturing and quality control. However, the existing manual visual inspection, automatic detection based on image processing, and intelligent detection based on machine learning are all unable to efficiently and accurately identify tiny defects in industrial scenarios. Based on this, the present embodiment aims to provide a part surface defect detection method based on the Transformer model to solve the above-mentioned technical defects.
[0068] The following is a detailed introduction to the part surface defect detection method based on the Transformer model described in this embodiment with reference to the accompanying drawings.
[0069] Combination Figure 1 As shown, the method described in this embodiment includes the following steps:
[0070] Step S102: Acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and perform grayscale processing on each image to be detected to obtain each grayscale image.
[0071] Specifically, in order to facilitate subsequent identification, combined Figure 2 As shown, the image to be detected is obtained in the following manner:
[0072] Step S10202: Acquire a number of original images corresponding to the surface of the part to be inspected based on a high-resolution camera.
[0073] At this time, it can be ensured that the original image has sufficient detail resolution, thereby being able to clearly display information on various types of tiny defects.
[0074] Step S10204: perform format processing on each original image to obtain each intermediate image.
[0075] Specifically, the purpose of format processing is to make the resolution and size of each intermediate image consistent, so as to facilitate subsequent unified processing.
[0076] Step S10206: Perform random denoising on each intermediate image based on an algorithm to obtain each image to be detected.
[0077] In this embodiment, the filtering algorithm is specifically a Gaussian filtering algorithm or a median filtering algorithm. The denoising process based on this step can remove invalid information in the intermediate image to avoid interference, thereby highlighting and retaining edge and detail information.
[0078] Step S104: After low-pass filtering and high-pass filtering are performed on each grayscale image in the row direction, down-sampling is continued simultaneously by a factor of two to obtain the corresponding row low-frequency part and row high-frequency part; after low-pass filtering and high-pass filtering are performed in the column direction, down-sampling is continued simultaneously by a factor of two to obtain the corresponding LL original sub-band and LH original sub-band; after low-pass filtering and high-pass filtering are performed in the column direction, down-sampling is continued simultaneously by a factor of two to obtain the corresponding HL original sub-band and HH original sub-band.
[0079] Specifically, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction. At this time, the global information carried by the image to be detected and the detail information related to defects, especially small defects, will be highlighted at the same time through each original subband.
[0080] Combination Figure 3 As shown, in this embodiment, each original sub-band is obtained in the following manner:
[0081] Step S10402, low-pass filtering and two-fold downsampling are performed on the grayscale image f(i, j) in the row direction based on a low-pass filter to obtain the row low-frequency part; at the same time, high-pass filtering and two-fold downsampling are performed in the row direction based on a high-pass filter to obtain the row high-frequency part.
[0082] Specifically, the low-frequency part is expressed as:
[0083]
[0084] The high frequency part is expressed as:
[0085]
[0086] Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index; h low is a low-pass filter, h high is a high pass filter.
[0087] Specifically, the low-pass filter is defined as:
[0088]
[0089] The low pass filter is defined as:
[0090]
[0091] Step S10404: low-pass filter and double down-sample the row low-frequency part in the column direction based on a low-pass filter to obtain the LL original sub-band; and high-pass filter and double down-sample the row low-frequency part in the column direction based on a high-pass filter to obtain the LH original sub-band.
[0092] Wherein, the LL original subband is:
[0093]
[0094] The LH original subband is:
[0095]
[0096] Step S10404, low-pass filtering and two-fold downsampling are performed on the high-frequency part of the row in the column direction based on a low-pass filter to obtain the HL original sub-band; and high-pass filtering and two-fold downsampling are performed based on a high-pass filter to obtain the HH original sub-band.
[0097] Wherein, the HL original subband is:
[0098]
[0099] The HH original subband is:
[0100]
[0101] On the basis of step S104, in order to further highlight the defect features such as micro cracks and fine scratches, continue to perform the following steps:
[0102] Step S106: respectively introduce preset fixed gain factors into the LH original sub-band, the HL original sub-band, and the HH original sub-band to obtain the first gain coefficient of each position in each original sub-band, calculate the sum of the squares of the first gain coefficients of each position in each original sub-band as the energy value of the corresponding original sub-band; calculate the dynamic gain factor corresponding to the corresponding original sub-band based on the energy value; and perform frequency domain enhancement on the LH original sub-band, the HL original sub-band, and the HH original sub-band based on the corresponding fixed gain factor and the dynamic gain factor to obtain the LH gain sub-band, the HL gain sub-band, and the HH gain sub-band accordingly.
[0103] In the specific implementation, first, define the fixed gain factor of the LH original subband as G LH , the result of using it to amplify the coefficient of the LH original subband is:
[0104] C′ LH (i, j) = C LH (i, j)×G LH ;
[0105] Among them, C LH (i, j) is the coefficient of the LH original subband;
[0106] Similarly, the fixed gain factor of the HL original subband is defined as G HL , the result of using it to amplify the coefficient of the HL original subband is:
[0107] C′ HL (i, j) = C HL (i, j)×G HL ;
[0108] Among them, C HL (i, j) is the coefficient of the original HL subband;
[0109] Define the fixed gain factor of the HH original subband as G HH , the result of using it to amplify the coefficient of the HH original subband is:
[0110] C′ HH (i, j) = C HH (i, j)×G HH ;
[0111] Among them, C HH (i, j) is the coefficient of the original subband of HH.
[0112] Secondly, in order to adapt to small defects of different sizes, a dynamic gain control mechanism is introduced to achieve adaptive frequency domain enhancement, including:
[0113] The energy value of the subband after fixed gain factor gain is calculated by the following formula:
[0114]
[0115] Among them, d∈{LH,HL,HH}.
[0116] The dynamic gain factor is calculated based on the energy value:
[0117]
[0118] Among them, α is the basic weight constant, which is used to control the overall amplification level; E ref is the reference energy, used to determine the energy level of the current subband; ∈ is an infinitesimal constant, used to prevent the denominator from being zero.
[0119] At this time, the corresponding original subband is gained based on the following gain factor to obtain the corresponding gain subband:
[0120]
[0121] in, is the coefficient of gain subband d.
[0122] In the specific implementation, based on the defect characteristics and detection requirements of different industrial parts, by adjusting the parameters α and E ref That is, it can adapt to different frequency characteristics and energy distribution to ensure that defect features are effectively amplified while suppressing noise and background texture.
[0123] Step S108, performing inverse column filtering operations and reverse row filtering operations on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband respectively to reconstruct the image to obtain a spatial domain image; performing a fusion operation on the spatial domain image and the grayscale image to obtain a target image.
[0124] Specific, combined Figure 4 As shown, the following steps are included:
[0125] Step S10802: resize the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband so that their sizes meet the sampling ratio requirement.
[0126] Step S10804: Perform inverse column filtering on each high-frequency sub-band and low-frequency sub-band to restore information in the original direction.
[0127] Specifically, the LL original subband after size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the LH gain subband after size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the HL gain subband after size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction; the HH gain subband after size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction.
[0128] In the specific implementation, first, in the row direction, the row low frequency and the row high frequency are inversely filtered. Specifically, the inverse low-pass filtering process is:
[0129]
[0130] The inverse high-pass filtering process is:
[0131]
[0132] Where, d∈{LH, HL, HH}, g low and g high is with h low and h high The corresponding inverse filter coefficients are:
[0133]
[0134] Secondly, the above operation results are subjected to inverse filtering in the column direction for the column low frequency and column high frequency. Specifically, the inverse low-pass filtering process is as follows:
[0135]
[0136] The inverse high-pass filtering process is:
[0137]
[0138] Step S10806: perform spatial reconstruction to obtain a spatial domain image:
[0139]
[0140] Specifically, in image fusion, the reconstructed spatial domain The fusion operation is performed with the original image f(i, j) to ensure that the brightness and contrast of the overall image are balanced:
[0141]
[0142] Among them, β is the fusion weight parameter, which is usually set to 0.5.
[0143] Step 110: Patch cutting and position encoding optimization are performed on the target image and then the image is input into a pre-trained Transformer model to output defect recognition results on the surface of the part to be inspected.
[0144] Combination Figure 5 As shown, identification and detection are performed through the following steps:
[0145] Step S11002: Patch the target image according to a preset grid size to obtain a plurality of image blocks.
[0146] Specifically, the target image is divided into patches of fixed size, which are divided into P×P pixel blocks, and each block can be regarded as a token of Transformer input.
[0147] Furthermore, assuming the target image size is H×W, the total size can be divided into For different industrial parts or defect sizes, the size of P can be adjusted to balance resolution and computational efficiency.
[0148] Step S11004: add coordinate position codes to each image block, and add directional labels to each image block based on defect features in each direction retained in frequency domain enhancement.
[0149] Specifically, after Patch cutting, a two-dimensional coordinate position code (x Patch ,y Patch ) is to meet the Transformer model's requirement of position information to make up for the lack of convolution translation invariance. The directional information (specifically including horizontal defect features, vertical defect features, and diagonal defect features) retained in the patch-level record and frequency domain enhancement is convenient for improving the recognition rate of subsequent model recognition.
[0150] Step S11006: Expand each image block in the spatial domain or perform a simple linear projection to obtain corresponding feature vectors, and add corresponding coordinate position codes and directional labels to each feature vector to obtain an embedding vector corresponding to each image block.
[0151] Specifically, the embedding vector can be expressed as:
[0152]
[0153] in, Indicates vector concatenation or additive fusion, Z patch is the feature representation of the image patch, POS patchIt is the position code of each patch, which is used to indicate the position of the patch in the image. Patch can be a trainable vector representing a direction label.
[0154] Step S11008: Based on the coordinate position encoding, each embedding vector is input into the Transformer model in sequence to obtain the defect recognition result under the guidance of the self-attention mechanism and the directional label.
[0155] Specifically, the series of patches processed above are embedded into the vector {z′ Patch1 , z′ Patch2 , ...} are sequentially input into the encoder of the Transformer model or the corresponding detection head to form a token sequence. During the recognition process, the Transformer model will focus on the correlation between patches through the self-attention mechanism, and under the guidance of the direction label, accurately identify the defect features after high-frequency enhancement, such as tiny cracks and fine scratches.
[0156] In summary, this embodiment combines frequency domain enhancement with the Transformer model to amplify defect details in the frequency domain and achieve global correlation characterization of the image, so as to achieve a balance between local details and overall structure in the overall detection process. In the model recognition process, image segmentation and position encoding and other processing methods are introduced to make the enhanced high-frequency features more closely integrated with the global semantic information, greatly improving the recognition accuracy of tiny defects and the robustness of the recognition process.
[0157] The above program can be run in the processor, or it can be stored in the memory (or computer-readable storage medium), which includes permanent and non-permanent, removable and non-removable media. Information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include temporary computer-readable media, such as modulated data signals and carrier waves.
[0158] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes can be implemented by different modules corresponding to different steps.
[0159] This embodiment also provides a part surface defect detection system based on the Transformer model. Figure 6 As shown, the system includes the following functional modules:
[0160] An image acquisition module is used to acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and grayscale process each image to be detected to acquire each grayscale image;
[0161] A filtering processing module is used to perform low-pass filtering and high-pass filtering on each grayscale image in the row direction, and then continue to perform two-fold downsampling to obtain the corresponding row low-frequency part and row high-frequency part respectively; perform low-pass filtering and high-pass filtering on the low-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding LL original subband and LH original subband respectively; perform low-pass filtering and high-pass filtering on the high-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding HL original subband and HH original subband respectively;
[0162] Among them, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction;
[0163] A gain processing module, used for introducing preset fixed gain factors into the LH original subband, the HL original subband, and the HH original subband to obtain a first gain coefficient at each position in each original subband, calculating the sum of squares of the first gain coefficient at each position in each original subband as an energy value of the corresponding original subband; calculating a dynamic gain factor corresponding to the corresponding original subband based on the energy value; and performing frequency domain enhancement on the LH original subband, the HL original subband, and the HH original subband based on the corresponding fixed gain factor and the dynamic gain factor to obtain a LH gain subband, a HL gain subband, and a HH gain subband respectively;
[0164] The dynamic gain factor is Among them, α is the basic weight constant; E refis the reference energy value; ∈ is an infinitesimal constant used to prevent the denominator from being zero; E d is the energy value;
[0165] an inverse filtering processing module, configured to perform inverse column filtering operations and inverse row filtering operations on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband, respectively, and then reconstruct the image to obtain a spatial domain image; and perform a fusion operation on the spatial domain image and the grayscale image to obtain a target image;
[0166] The defect recognition module is used to perform patch cutting and position coding optimization on the target image and then input the image into a pre-trained Transformer model to output the defect recognition result of the surface of the part to be inspected.
[0167] The system is built based on the method, so what has been described above will not be repeated here.
[0168] For example, the image acquisition module includes units:
[0169] The acquisition unit is used to acquire a plurality of original images corresponding to the surface of the part to be inspected based on a high-resolution camera.
[0170] The first processing unit is used to perform format processing on each original image to obtain each intermediate image; wherein the resolution and size of each intermediate image remain consistent.
[0171] The second processing unit is used to perform random denoising on each intermediate image based on a filtering algorithm to obtain each image to be detected.
[0172] For another example, the filtering processing module includes:
[0173] The third processing unit is used to perform low-pass filtering and double downsampling on the grayscale image f(i, j) in the row direction based on a low-pass filter to obtain the row low-frequency part At the same time, high-pass filtering and double downsampling are performed based on a high-pass filter in the row direction to obtain the high-frequency part of the row Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index;
[0174] h low is a low-pass filter, h high is a high-pass filter;
[0175] The fourth processing unit is used to perform low-pass filtering and double downsampling on the row low-frequency part in the column direction based on a low-pass filter to obtain an LL original subband In the column direction, high-pass filtering and double downsampling are performed based on a high-pass filter to obtain the LH original subband
[0176] The fifth processing unit is used to perform low-pass filtering and double downsampling on the high-frequency part of the row in the column direction based on a low-pass filter to obtain the HL original subband And based on the high-pass filter, high-pass filtering and double downsampling are performed to obtain the HH original subband
[0177] Correspondingly, this embodiment also provides an electronic device, such as Figure 7 As shown, the electronic device includes at least one processor, the processor is coupled to a memory, a computer program is stored in the memory, and the computer program is configured to execute the method when executed by the processor.
[0178] At the same time, the technical solution also provides a computer-readable storage medium on which a computer program is stored, and the computer program is used to execute the described method.
[0179] Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. A person with ordinary knowledge in the technical field to which the present invention belongs may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the definition of the claims.
Claims
1. A part surface defect detection method based on Transformer model, characterized in that: include: Acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and perform grayscale processing on each image to be detected to obtain each grayscale image; After low-pass filtering and high-pass filtering are respectively performed on each grayscale image in the row direction, two-fold downsampling is continued simultaneously to obtain the corresponding row low-frequency part and row high-frequency part respectively; after low-pass filtering and high-pass filtering are respectively performed on the low-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding LL original subband and LH original subband respectively; after low-pass filtering and high-pass filtering are respectively performed on the high-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding HL original subband and HH original subband respectively; Among them, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction; Introducing preset fixed gain factors into the LH original subband, the HL original subband, and the HH original subband to obtain first gain coefficients at various positions in the original subbands, calculating the sum of squares of the first gain coefficients at various positions in the original subbands as energy values of the corresponding original subbands; calculating dynamic gain factors corresponding to the corresponding original subbands based on the energy values; and performing frequency domain enhancement on the LH original subband, the HL original subband, and the HH original subband based on the corresponding fixed gain factors and the dynamic gain factors to obtain LH gain subbands, HL gain subbands, and HH gain subbands respectively; The dynamic gain factor is Among them, α is the basic weight constant; E ref is the reference energy value; ∈ is an infinitesimal constant used to prevent the denominator from being zero; E d is the energy value; The LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband are sequentially subjected to inverse column filtering operations and inverse row filtering operations, and then image reconstruction is performed to obtain a spatial domain image; the spatial domain image is fused with the grayscale image to obtain a target image; The target image is subjected to patch cutting and position encoding optimization and then input into a pre-trained Transformer model to output defect recognition results on the surface of the part to be inspected.
2. The part surface defect detection method based on the Transformer model according to claim 1 is characterized in that: The step of acquiring a plurality of images to be detected corresponding to the surface of the part to be detected comprises: Acquire a number of original images corresponding to the surface of the part to be inspected based on a high-resolution camera; Performing format processing on each original image to obtain each intermediate image; wherein the resolution and size of each intermediate image remain consistent; Based on the filtering algorithm, each intermediate image is randomly denoised to obtain each image to be detected.
3. The part surface defect detection method based on the Transformer model according to claim 1 is characterized in that: After low-pass filtering and high-pass filtering are respectively performed on each grayscale image in the row direction, two-fold downsampling is continued simultaneously to obtain the corresponding row low-frequency part and row high-frequency part respectively; after low-pass filtering and high-pass filtering are respectively performed on the low-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding LL original subband and LH original subband respectively; after low-pass filtering and high-pass filtering are respectively performed on the high-frequency part in the column direction, two-fold downsampling is continued simultaneously to obtain the corresponding HL original subband and HH original subband respectively, including: The grayscale image f(i, j) is low-pass filtered and downsampled twice in the row direction based on a low-pass filter to obtain the row low-frequency part At the same time, high-pass filtering and double downsampling are performed based on a high-pass filter in the row direction to obtain the high-frequency part of the row Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index; h low is a low-pass filter, h high is a high pass filter; The row low-frequency part is low-pass filtered and down-sampled twice based on a low-pass filter in the column direction to obtain the LL original subband In the column direction, high-pass filtering and double downsampling are performed based on a high-pass filter to obtain the LH original subband The high frequency part of the row is low-pass filtered and down-sampled twice based on a low-pass filter in the column direction to obtain the HL original subband And based on the high-pass filter, high-pass filtering and double downsampling are performed to obtain the HH original subband 4. The part surface defect detection method based on the Transformer model according to claim 1 is characterized in that: The step of performing inverse column filtering and reverse row filtering on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband in sequence and then reconstructing the image to obtain a spatial domain image comprises: Performing size processing on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband so that their sizes meet the sampling ratio requirement; The LL original subband after the size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the LH gain subband after the size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse low-pass filtering and double downsampling in the row direction; the HL gain subband after the size processing is sequentially subjected to inverse low-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction; the HH gain subband after the size processing is sequentially subjected to inverse high-pass filtering and double downsampling in the column direction, and inverse high-pass filtering and double downsampling in the row direction; Spatial reconstruction is performed to obtain a spatial domain image.
5. The part surface defect detection method based on the Transformer model according to claim 1 is characterized in that: The target image is subjected to patch cutting and position encoding optimization and then input into a pre-trained Transformer model to output defect recognition results on the surface of the part to be inspected, including: Patch cutting the target image according to a preset grid size to obtain a plurality of image blocks; Add coordinate position coding to each tile, and add directional labels to each tile based on the defect features in each direction retained in frequency domain enhancement; Each tile is expanded in the spatial domain or a simple linear projection is performed to obtain the corresponding feature vector, and the corresponding coordinate position code and directional label are added to each feature vector to obtain the embedding vector corresponding to each tile; Based on the coordinate position encoding, each embedding vector is input into the Transformer model in sequence to obtain the defect recognition result under the guidance of the self-attention mechanism and directional labels.
6. A part surface defect detection system based on Transformer model, characterized in that: include: An image acquisition module is used to acquire a plurality of images to be detected corresponding to the surface of the part to be detected, and grayscale process each image to be detected to acquire each grayscale image; A filtering processing module is used to perform low-pass filtering and high-pass filtering on each grayscale image in the row direction, and then continue to perform two-fold downsampling to obtain the corresponding row low-frequency part and row high-frequency part respectively; perform low-pass filtering and high-pass filtering on the low-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding LL original subband and LH original subband respectively; perform low-pass filtering and high-pass filtering on the high-frequency part in the column direction, and then continue to perform two-fold downsampling to obtain the corresponding HL original subband and HH original subband respectively; Among them, the LL original subband is used to store the brightness distribution and macro structure of the grayscale image, the LH original subband is used to store the horizontal features of the grayscale image, the HL original subband is used to store the vertical features of the grayscale image, and the HH original subband is used to store the features of the grayscale image in the diagonal direction; A gain processing module, used for introducing preset fixed gain factors into the LH original subband, the HL original subband, and the HH original subband to obtain a first gain coefficient at each position in each original subband, calculating the sum of squares of the first gain coefficient at each position in each original subband as an energy value of the corresponding original subband; calculating a dynamic gain factor corresponding to the corresponding original subband based on the energy value; and performing frequency domain enhancement on the LH original subband, the HL original subband, and the HH original subband based on the corresponding fixed gain factor and the dynamic gain factor to obtain a LH gain subband, a HL gain subband, and a HH gain subband respectively; The dynamic gain factor is Among them, α is the basic weight constant; E ref is the reference energy value; ∈ is an infinitesimal constant used to prevent the denominator from being zero; E d is the energy value; an inverse filtering processing module, configured to perform inverse column filtering operations and inverse row filtering operations on the LL original subband, the LH gain subband, the HL gain subband, and the HH gain subband, respectively, and then reconstruct the image to obtain a spatial domain image; and perform a fusion operation on the spatial domain image and the grayscale image to obtain a target image; The defect recognition module is used to perform patch cutting and position coding optimization on the target image and then input the image into a pre-trained Transformer model to output the defect recognition result of the surface of the part to be inspected.
7. The part surface defect detection system based on the Transformer model according to claim 6 is characterized in that: The image acquisition module comprises: An acquisition unit, used for acquiring a plurality of original images corresponding to the surface of the part to be inspected based on a high-resolution camera; A first processing unit is used to perform format processing on each original image to obtain each intermediate image; wherein the resolution and size of each intermediate image are kept consistent; The second processing unit is used to perform random denoising on each intermediate image based on a filtering algorithm to obtain each image to be detected.
8. The part surface defect detection system based on the Transformer model according to claim 6 is characterized in that: The filtering processing module comprises: The third processing unit is used to perform low-pass filtering and double downsampling on the grayscale image f(i, j) in the row direction based on a low-pass filter to obtain the row low-frequency part At the same time, high-pass filtering and double downsampling are performed based on a high-pass filter in the row direction to obtain the high-frequency part of the row Wherein, the size of the grayscale image is M×N, i∈{0, 1, ..., M-1} is a row index, j∈{0, 1, ..., N-1} represents a column index, and k is a frequency domain index; h low is a low-pass filter, h high is a high pass filter; The fourth processing unit is used to perform low-pass filtering and double downsampling on the row low-frequency part in the column direction based on a low-pass filter to obtain an LL original subband In the column direction, high-pass filtering and double downsampling are performed based on a high-pass filter to obtain the LH original subband The fifth processing unit is used to perform low-pass filtering and double downsampling on the high-frequency part of the row in the column direction based on a low-pass filter to obtain the HL original subband And based on the high-pass filter, high-pass filtering and double downsampling are performed to obtain the HH original subband 9. An electronic device, characterized in that: The method comprises at least one processor, wherein the processor is coupled to a memory, wherein a computer program is stored in the memory, and wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed by the processor.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to execute the method according to any one of claims 1 to 5.