Diamond wire saw steel ball detection method and system based on improved yov8 and medium
By improving the YOLOv8 model, introducing an illumination-invariant feature extractor and attention mechanism, optimizing the loss function, and combining the Retinex algorithm for image preprocessing, the problems of low small target detection accuracy and the influence of illumination changes in YOLOv8 in diamond wire saw steel ball detection are solved, and high-precision and stable detection effects are achieved.
Patent Information
- Application Number
- CN202510878756.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-26
AI Technical Summary
The existing YOLOv8 algorithm has problems in diamond wire saw steel ball detection, such as low small target detection accuracy, repeated detection in complex backgrounds, and detection stability affected by illumination changes. It is difficult to achieve high-precision and stable detection in industrial environments.
By improving the YOLOv8 model, introducing the illumination-invariant feature extractor LIF and the CA attention mechanism, optimizing the loss function, and combining the improved Retinex algorithm for image preprocessing, the small target detection capability is enhanced, background interference is reduced, and detection accuracy is improved by dynamically adjusting the loss weight and feature fusion.
It achieves high-precision, real-time detection of diamond wire saw beads in complex industrial environments, improves detection accuracy and stability, reduces false detection rate, and meets the real-time detection needs of industrial production.
Smart Images

Figure CN120707835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to a diamond wire saw steel ball detection method, system and medium based on improved YOLOv8. Background Art
[0002] With the advancement of machine vision technology and the accelerated digital transformation of the manufacturing industry, bead inspection, a core step in the diamond wire saw production process, has a direct impact on production efficiency and product quality. Early detection algorithms relied on manually extracting bead shape and color features, such as dilation and corrosion algorithms. This cumbersome feature extraction process significantly reduced detection accuracy in scenarios such as changing lighting and scale.
[0003] With the rise of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have become mainstream. YOLOv8, as a leading object detection algorithm, excels in detection speed and accuracy, but it still has shortcomings in detecting small objects such as beads. Because beads typically appear as small objects in images and lack distinct color and texture features, YOLOv8's feature extraction capabilities struggle to effectively capture sufficient detail when processing small objects, resulting in low detection accuracy. Furthermore, in the actual production of diamond wire saws, complex backgrounds and close proximity between beads make occlusion and dense distribution a common occurrence. When handling such complex scenes, YOLOv8 is prone to duplicate or missed detection of beads.
[0004] To address the aforementioned issues with YOLOv8 in bead detection, some scholars have proposed improved algorithms. For example, one study introduced an attention mechanism into the YOLOv8 backbone network to improve focus on small targets. However, in complex industrial environments, unstable illumination changes can severely impact the effectiveness of the attention mechanism, leading to large fluctuations in detection performance. Other studies have attempted to optimize small target detection by adjusting the structure of the detection head. However, for targets with small scale variations such as beads, simply adjusting the detection head is unlikely to significantly improve detection accuracy and may increase computational effort, impacting detection speed. Furthermore, in terms of loss function design, traditional loss functions such as CIoU do not differentiate the quality of predicted frames precisely enough when dealing with dense targets, making it difficult to effectively resolve the problem of duplicate detection.
[0005] In summary, the existing YOLOv8 and related improved algorithms face challenges in bead detection, such as low small target detection accuracy, repeated detection in complex backgrounds, and the impact of illumination changes on detection stability. A more effective detection method is urgently needed to solve these problems. Summary of the Invention
[0006] The present invention aims to at least solve the technical problems existing in the prior art, and in particular innovatively proposes a diamond wire saw steel ball detection method, system and medium based on improved yolov8.
[0007] In order to achieve the above-mentioned object of the present invention, the present invention provides a diamond wire saw steel ball detection method based on improved yolov8, comprising:
[0008] S1, collect the original image of diamond wire saw beads; use labImage label image to annotate the beads, and form training set, validation set, and test set data according to the proportion;
[0009] S2, preprocessing the image, after scaling the image, using the improved Retinex algorithm to preprocess the image;
[0010] S3 improves the yolov8 model by adding an illumination-invariant feature extractor (LIF) to the backbone. The CA attention mechanism is integrated into the C2f module. Based on the small size variation of beads, a small target detection head is added while the large-scale detection head is removed, making it suitable for detecting small targets in beads.
[0011] S4, optimize the loss function of yolov8, use the improved IoU loss function to improve the accuracy of small target detection, use the improved yolov8 model for training, and obtain the training result best.pt file.
[0012] The above technical solution is preferably, wherein S2 includes:
[0013] S2-1, use the improved RetineX algorithm to perform image enhancement preprocessing on the image to reduce the interference of the background image;
[0014] Obtain a scaled diamond steel ball image; calculate the grayscale average value V of the image and perform logarithmic transformation on smaller V and exponential transformation on larger V according to the following formula, and obtain a better quality image I1 through histogram equalization;
[0015]
[0016] Where: x and y are pixel coordinates; g(x,y) is the output pixel value; f(x,y) is the original pixel value; a, b, c are custom parameters;
[0017] When 0≤V<100, logarithmic transformation is used to stretch the low-light area and improve the detail contrast in low light;
[0018] When the illumination is medium (100 ≤ V ≤ 170), the original pixel value is directly output;
[0019] When the light intensity is 170<V≤255, the exponential function is used to reduce overexposure interference and maintain background details.
[0020] The above technical solution is preferably, wherein S2 further includes:
[0021] S2-2, performing Wiener filtering on the I1 image to obtain image I2;
[0022] The steps for performing Wiener filtering on an image are as follows:
[0023] Calculate the mean and variance of each pixel in the image within the range of M*N. The formula is as follows:
[0024]
[0025] M, N: The range of the image to be filtered by Wiener, M represents width, N represents height; x, y: represents pixel coordinates, η is the M*N range set, μ is the mean value of pixels in the M*N range, σ 2 is the pixel variance within the M*N range.
[0026] The above technical solution is preferably, wherein S2 further includes:
[0027] S2-3, calculate the average value of the local variance, the formula is as follows:
[0028]
[0029] Where Z represents the number of pixels; Represents the pixel variance of a point i within the M*N range, v 2 Indicates the average value of pixel variance within the M*N range.
[0030] The new image pixel value is calculated as follows:
[0031]
[0032] Perform logarithmic operation on the I1 and I2 images; obtain images logI1 and logI2; subtract logI2 from logI1 to obtain the logI image; and perform normalization on the logI image.
[0033] In the above technical solution, preferably, S3 includes:
[0034] S3-1, calculate the gradient images of the image in the x and y directions respectively. The formula is as follows:
[0035] f x ′(x,y)=f(x+1,y)-f(x,y);
[0036] f y′(x,y)=f(x,y+1)-f(x,y);
[0037] Convolve the gradient image with the Gaussian filter to obtain the gradient component feature information in the x and y directions. The formula is as follows:
[0038]
[0039]
[0040] is the convolution operation, F x (x,y) and F y (x, y) are the gradient component features of the image in two different directions, X and Y; G x (x,y,σ) and G y (x, y, σ) are Gaussian filters in two different directions, X and Y respectively; σ is the standard deviation of the Gaussian kernel, which controls the smoothness of the filter. The larger σ is, the better the smoothing effect is.
[0041] G x (x,y,σ) represents the derivative of G(x,y,σ) in the x direction
[0042]
[0043] G y (x,y,σ) represents the derivative of G(x,y,σ) in the y direction,
[0044]
[0045] F x (x,y) and F y Perform the inverse tangent operation on (x,y) to obtain the gradient characteristics of the image. The formula is as follows:
[0046]
[0047] The principles, steps, and network structure of the CA attention mechanism are as follows:
[0048] S3-2, CA attention mechanism decomposes the input features into horizontal and vertical feature vectors, then performs feature fusion, generates attention weights, and extracts more refined target features. The steps are as follows:
[0049] First, the input feature map is divided into two directions, width and height, and global average pooling is performed separately to obtain feature maps in the width and height directions respectively. The formula is as follows:
[0050]
[0051] in, is the global average pooling feature vector in the height direction; c is the channel index of the feature map, and h is the position index of the current operation in the height direction;
[0052] is the global average pooling feature vector in the width direction; c is the channel index of the feature map, w is the position index of the current operation in the width direction; x c (h,i) is the pixel value of the cth channel in the input feature map with coordinates (h,i); x c (j,w) is the pixel value of the cth channel in the input feature map with coordinates (j,w), where i and j represent the indexes in the height and width directions, respectively.
[0053] In the above technical solution, preferably, S3 includes:
[0054] S3-3, the feature maps of the width and height of the global receptive field are spliced together, and then they are sent to the convolution module with a shared convolution kernel of 1×1 to reduce their dimension to the original C / r. The batch normalized feature map F1 is then sent to the Sigmoid activation function to obtain a feature map f of the shape 1×(W+H)×C / r. The formula is as follows:
[0055] f=δ(F1([z h ,z w ]));
[0056] Among them, z h is the feature map in the height direction; z w is the feature map in the width direction; F1(*) is the 1*1 convolution operation; δ(*) represents the calculation result of the sigmoid function with * as the parameter,
[0057] Then, the feature map f is convolved with a kernel of 1×1 according to the original height and width to obtain feature maps with the same number of channels as the original. After the Sigmoid activation function, the attention weights g of the feature map in height and width are obtained respectively. w and the attention weight g in the width direction h , the formula is:
[0058] g w =σ(F w (f w ))
[0059] g h =σ(F h (f h ))
[0060] F w (*) and F h(*) represents the 1×1 convolution operation in the width and height directions respectively; f w and f h Represents z h and z w The feature maps in width and height directions are obtained after 1*1 convolution operation and Sigmoid activation function.
[0061] S3-4, through multiplication weighted calculation on the original feature map, the final feature map with attention weights in the width and height directions will be obtained. The formula is as follows:
[0062]
[0063] y c (i, j) is the pixel value at coordinate (i, j) in the cth channel of the final feature map with attention weight;
[0064] x c (i, j) is the pixel value at coordinate (i, j) in the c-th channel of the feature map input to the CA attention mechanism;
[0065] is the attention weight in the height direction, corresponding to the weight value at the height position i of the c-th channel;
[0066] is the attention weight in the width direction, corresponding to the weight value at the width position j of the c-th channel;
[0067] i and j represent the index in the height and width directions respectively.
[0068] The above technical solution is preferably, wherein S4 includes:
[0069] S4-1, calculate the prediction box B p With the real box B g The intersection over union (IoU)
[0070]
[0071] Through logarithmic space conversion and exponential weighting, the loss weight of small targets is strengthened.
[0072]
[0073] Among them, s is the target box area, s ref is the reference area, i.e., the average target area of the dataset, λ is the scale sensitivity coefficient, ranging from 1.0 to 2.0; the scale adaptation weight is Ω s , when s<s ref When Ω s >1, increase the loss weight, otherwise reduce the weight;
[0074] S4-2, introduces normalized center point distance and diagonal constraints to solve the gradient vanishing problem in non-overlapping scenes;
[0075]
[0076] Where: k is the distance between the center of the predicted box and the real box, c is the diagonal length of the minimum circumscribed rectangle, ω is the gradient smoothing index, ρ is the distance attenuation coefficient, α D is the constraint coefficient, R D is the gradient equalization distance penalty.
[0077] The above technical solution is preferably, wherein S4 further includes:
[0078] S4-3, constrain the aspect ratio consistency between the predicted box and the real box, and introduce an angle perception mechanism.
[0079]
[0080] Among them, w p and h p are respectively the width and height of the prediction box, w g h g are the width and height of the real frame respectively; A is the shape penalty coefficient, ranging from 0.5 to 1.0; π 2 is the square of pi, R A To penalize shape alignment, the dual constraint mechanism simultaneously optimizes the aspect ratio angle and scale difference to improve the detection accuracy of irregularly shaped objects.
[0081] An angle constraint mechanism is introduced to perform dual constraints together with the aspect ratio to improve the detection accuracy of irregularly shaped targets.
[0082] S4-4, integrate prediction confidence and dynamically adjust loss weights,
[0083]
[0084] Among them, C is the confidence score of the prediction box, Focus coefficient, ε stability coefficient, Ω C The confidence-guided weights are applied to low-confidence samples, giving priority to optimizing the localization accuracy of difficult samples.
[0085] S4-5, calculate the loss function according to the following formula
[0086]
[0087] Among them, the balance coefficient, (recommended α = 1.0, β = 0.5, γ = 0.2), is used to align the scale of RD with other loss terms; L is the improved loss function; Ω s is the scale-adaptive weight; Ω C is the confidence guide weight, IoU is the original loss function, R D is the gradient equalization distance penalty, R A Penalizes shape alignment.
[0088] The present invention also discloses a computer system, comprising:
[0089] processor;
[0090] a memory for storing processor-executable instructions;
[0091] The processor is configured to implement the application method according to any one of claims 1 to 8 when executing the executable instructions.
[0092] The present invention also discloses a computer-readable storage medium, comprising:
[0093] a memory having a computer program stored thereon;
[0094] A processor is configured to execute the program in the memory to implement the application method according to any one of claims 1 to 8.
[0095] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0096] This technology addresses the problem of detecting small targets in diamond wire saw beads during the production process. Through illumination-invariant feature extraction, attention mechanism optimization, and dynamic loss function design, it achieves high-precision, real-time detection of diamond wire saw beads in complex industrial environments.
[0097] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0099] Figure 1 It is the overall flow chart of the present invention;
[0100] Figure 2 It is the yolov8 network model structure of the present invention;
[0101] Figure 3 It is the improved network model structure of yolov8 of the present invention;
[0102] Figure 4 It is the illumination invariant feature extractor LIF structure of the present invention;
[0103] Figure 5 This is the CA attention mechanism module structure of the present invention;
[0104] Figure 6 It is the steel ball collecting device of the present invention;
[0105] Figure 7 This is a schematic diagram of obtaining the original image of the steel ball according to the present invention;
[0106] Figure 8 It is a schematic diagram of the steel ball coordinate parameter data of the present invention. DETAILED DESCRIPTION
[0107] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0108] like Figures 1 to 5 As shown, the present invention proposes a system-level optimized diamond wire saw steel bead detection method, which achieves a breakthrough improvement in detection performance through deep collaboration and interaction between various functional modules, avoids excessive dependence on a single module, and solves the problem of bead detection from the overall architecture level. In terms of system architecture design, the present invention constructs a multi-module linked bead detection system. The adaptive image preprocessing module intelligently adjusts the preprocessing parameters according to the illumination, noise and other characteristics of the input image to provide high-quality images for subsequent detection; the feature extraction module adopts a hierarchical feature extraction strategy to cross-layer fuse shallow detail features with deep semantic features, effectively making up for the defect of YOLOv8's insufficient feature extraction of small bead targets; the dynamic detection and decision module dynamically adjusts the detection threshold based on the feature map output by the feature enhancement module, combined with the target confidence and spatial position relationship, accurately identifies and eliminates duplicate detection results, and effectively solves the problem of false detection when beads are densely distributed under complex backgrounds.
[0109] The specific algorithm steps of the diamond wire saw bead detection algorithm based on the improved yolov8 are as follows:
[0110] S1, such as Figure 6 and 7 As shown in the figure, the original image of diamond wire saw beads is collected by a camera on the production line workbench; the beads are annotated using labImage label images, and training set, validation set, and test set data are formed according to the proportion;
[0111] The data is stored in the steel ball folder. There are three folders in the steel ball folder: train, val, and test, which respectively store the images and annotation files of the training set, validation set, and test set. The image files of each data set are placed in the images folder, and the annotation files are placed in the labels folder.
[0112] S2, preprocessing the image, after scaling the image, using the improved Retinex algorithm to preprocess the image;
[0113] S3 improves the yolov8 model by adding an illumination-invariant feature extractor (LIF) to the backbone. The CA attention mechanism is integrated into the C2f module. Based on the small size variation of beads, a small target detection head is added while the large-scale detection head is removed, making it suitable for detecting small targets in beads.
[0114] S4, optimize the loss function of yolov8, use the improved IoU loss function to improve the accuracy of small target detection, use the improved yolov8 model for training, and obtain the training result best.pt file.
[0115] The S2 includes:
[0116] S2-1, use the improved RetineX algorithm to perform image enhancement preprocessing on the image to reduce the interference of the background image. The specific steps are as follows:
[0117] Obtain a scaled diamond steel ball image; calculate the grayscale average value V of the image and perform logarithmic transformation on smaller V and exponential transformation on larger V according to the following formula, and obtain a better quality image I1 through histogram equalization;
[0118]
[0119] Where: x and y are pixel coordinates; g(x,y) is the output pixel value; f(x,y) is the original pixel value; a, b, c are custom parameters; Figure 8 To collect the steel ball coordinate position parameter data.
[0120] When 0≤V<100, logarithmic transformation is used to stretch the low-light area and improve the detail contrast in low light;
[0121] When the illumination is medium (100 ≤ V ≤ 170), the original pixel value is directly output;
[0122] When the light intensity is 170<V≤255, the exponential function is used to reduce overexposure interference and maintain background details.
[0123] In this embodiment, a=24, b=0.7, and d=4.
[0124] S2-2, performing Wiener filtering on the I1 image to obtain image I2;
[0125] The steps for performing Wiener filtering on an image are as follows:
[0126] Calculate the mean and variance of each pixel in the image within the range of M*N. The formula is as follows:
[0127]
[0128] M, N: The range of the image to be filtered by Wiener, M represents width, N represents height; x, y: represents pixel coordinates, η is the M*N range set, μ is the mean value of pixels in the M*N range, σ 2 is the pixel variance within the M*N range.
[0129] S2-3, calculate the average value of the local variance, the formula is as follows:
[0130]
[0131] Where Z represents the number of pixels; Represents the pixel variance of a point i within the M*N range, v 2 Indicates the average value of pixel variance within the M*N range.
[0132] The new image pixel value is calculated as follows:
[0133]
[0134] Perform logarithmic operation on the I1 and I2 images; obtain images logI1 and logI2; subtract logI2 from logI1 to obtain the logI image; and perform normalization on the logI image.
[0135] Furthermore, the principles, steps, and network structure of the S3 illumination-invariant feature extractor LIF are as follows:
[0136] The illumination-invariant feature extraction of the present invention adopts the gradient extraction algorithm in the x and y directions. The image is regarded as a discrete two-dimensional function. Each pixel in the image is regarded as a point in the function. The gradient of the point in the x and y directions is expressed as the following formula:
[0137]
[0138] The gradients in two directions are convolved with the Gaussian filter to obtain the illumination invariant feature extractor LIF.
[0139] The specific steps are as follows:
[0140] S3-1, calculate the gradient images of the image in the x and y directions respectively. The formula is as follows:
[0141] f x ′(x,y)=f(x+1,y)-f(x,y);
[0142] f y ′(x,y)=f(x,y+1)-f(x,y);
[0143] Convolve the gradient image with the Gaussian filter to obtain the gradient component feature information in the x and y directions. The formula is as follows:
[0144]
[0145] is the convolution operation, F x (x,y) and F y (x, y) are the gradient component features of the image in two different directions, X and Y; G x (x,y,σ) and G y (x, y, σ) are Gaussian filters in two different directions, X and Y respectively; σ is the standard deviation of the Gaussian kernel, which controls the smoothness of the filter. The larger σ is, the better the smoothing effect is.
[0146] G x (x,y,σ) represents the derivative of G(x,y,σ) in the x direction
[0147]
[0148] G y (x,y,σ) represents the derivative of G(x,y,σ) in the y direction
[0149]
[0150] F x (x,y) and F y Perform the inverse tangent operation on (x,y) to obtain the gradient characteristics of the image. The formula is as follows:
[0151]
[0152] The principles, steps, and network structure of the CA attention mechanism are as follows:
[0153] S3-2, CA attention mechanism decomposes the input features into horizontal and vertical feature vectors, then performs feature fusion, generates attention weights, and extracts more refined target features. The steps are as follows:
[0154] First, the input feature map is divided into two directions, width and height, and global average pooling is performed separately to obtain feature maps in the width and height directions respectively. The formula is as follows:
[0155]
[0156] in, is the global average pooling feature vector in the height direction; c is the channel index of the feature map, and h is the position index of the current operation in the height direction;
[0157] is the global average pooling feature vector in the width direction; c is the channel index of the feature map, w is the position index of the current operation in the width direction; x c (h,i) is the pixel value of the cth channel in the input feature map with coordinates (h,i); x c (j,w) is the pixel value of the cth channel in the input feature map with coordinates (j,w), where i and j represent the indexes in the height and width directions, respectively.
[0158] S3-3, the feature maps of the width and height of the global receptive field are spliced together, and then they are sent to the convolution module with a shared convolution kernel of 1×1 to reduce their dimension to the original C / r. The batch normalized feature map F1 is then sent to the Sigmoid activation function to obtain a feature map f of the shape 1×(W+H)×C / r. The formula is as follows:
[0159] f=δ(F1([z h ,z w ]))
[0160] Among them, z h is the feature map in the height direction; z w is the feature map in the width direction; F1(*) is the 1*1 convolution operation; δ(*) represents the calculation result of the sigmoid function with * as the parameter,
[0161] Then, the feature map f is convolved with a kernel of 1×1 according to the original height and width to obtain feature maps with the same number of channels as the original. After the Sigmoid activation function, the attention weights g of the feature map in height and width are obtained respectively. w and the attention weight g in the width direction h , the formula is:
[0162] g w =σ(F w (f w ))
[0163] g h =σ(F h (f h ))
[0164] F w (*) and F h(*) represents the 1×1 convolution operation in the width and height directions respectively; f w and f h Represents z h and z w The feature maps in width and height directions are obtained after 1*1 convolution operation and Sigmoid activation function.
[0165] S3-4, through multiplication weighted calculation on the original feature map, the final feature map with attention weights in the width and height directions will be obtained. The formula is as follows:
[0166]
[0167] y c (i, j) is the pixel value at coordinate (i, j) in the cth channel of the final feature map with attention weight;
[0168] x c (i, j) is the pixel value at coordinate (i, j) in the c-th channel of the feature map input to the CA attention mechanism;
[0169] is the attention weight in the height direction, corresponding to the weight value at the height position i of the c-th channel;
[0170] is the attention weight in the width direction, corresponding to the weight value at the width position j of the c-th channel;
[0171] i and j represent the index in the height and width directions respectively.
[0172] The two-level illumination suppression of pixel domain ILF and feature domain CA is connected in series. The pixel domain solves the “pixel value distortion caused directly by illumination” (such as overexposure and shadow), and the feature domain solves the “feature confusion caused indirectly by illumination” (such as overlapping features of background and target), forming hierarchical complementarity. At the same time, the output of the pixel domain (illumination-invariant image) is used as the input of the feature domain, so that the attention mechanism can locate the target more accurately. If ILF is skipped, CA will mistakenly regard the background as the target due to strong illumination interference in the original features. The present invention dynamically and collaboratively adjusts the pixel domain σ, that is, the standard deviation of the Gaussian kernel and the r parameter of the feature domain, through an algorithm based on feature entropy. The specific steps are as follows:
[0173] Calculate the entropy value of the ILF output feature map H = -∑p i log p i , reflecting the feature complexity. Where H represents the entropy value of the ILF output feature map, p i is the probability of the i-th pixel value appearing in the feature map;
[0174] S3-5, through linear mapping
[0175] Dynamically adjust the Gaussian kernel standard deviation,
[0176] To dynamically adjust the dimensionality reduction ratio, H min and H max is the preset entropy value range, where σ min is the minimum standard deviation of the Gaussian kernel, σ max is the maximum value of the Gaussian kernel standard deviation, r min is the minimum value of dimensionality reduction ratio, r max is the maximum dimensionality reduction ratio, H min Preset minimum entropy value, H max Preset the maximum entropy value,
[0177] S4, through yolov8 optimization loss function, introduces four core improvements on the original IoU loss function: scale adaptation mechanism, gradient equalization, shape constraint and confidence guidance, to enhance the attention of small objects and avoid repeated detection. The specific steps are as follows:
[0178] S4-1, calculate the prediction box B p With the real box B g Intersection over Union (IoU)
[0179]
[0180] Strengthen the loss weight of small targets through logarithmic space transformation and exponential weighting
[0181]
[0182] Among them, s is the target box area, s ref is the reference area, i.e., the average target area of the dataset, λ is the scale sensitivity coefficient, ranging from 1.0 to 2.0; the scale adaptation weight is Ω s , when s<s ref (small target), Ω s >1, increase the loss weight, otherwise reduce the weight.
[0183] S4-2, introduces normalized center point distance and diagonal constraints to solve the gradient vanishing problem in non-overlapping scenes;
[0184]
[0185] Where: k is the distance between the center of the predicted box and the real box, c is the diagonal length of the minimum circumscribed rectangle, ω is the gradient smoothing index (range 1.5-2.0), and ρ is the distance attenuation coefficient (range 0.1-0.3) α D is the constraint coefficient (range 0.5-2.0); R DThe gradient is balanced with a distance penalty. The exponential decay term is used to enhance the gradient response of distant boxes to prevent gradient vanishing.
[0186] S4-3, constrain the aspect ratio consistency between the predicted box and the real box, and introduce an angle perception mechanism.
[0187]
[0188] Among them, w p and h p are respectively the width and height of the prediction box, w g h g are the width and height of the real frame respectively; A is the shape penalty coefficient, ranging from 0.5 to 1.0; π 2 is the square of pi, R A To penalize shape alignment, the dual constraint mechanism simultaneously optimizes the aspect ratio angle and scale difference to improve the detection accuracy of irregularly shaped objects.
[0189]
[0190] Together with the aspect ratio, the dual constraint improves the detection accuracy of irregularly shaped objects.
[0191] S4-4, integrate prediction confidence and dynamically adjust loss weights,
[0192]
[0193] Among them, C is the confidence score of the predicted box (output by the model, with a value range of 0.00-1.00); Focus factor (value range is 2.0-5.0); ε Stability coefficient (value range is 0.01-0.1); Ω C The confidence-guided weights are applied to low-confidence samples, giving priority to optimizing the localization accuracy of difficult samples.
[0194] S4-5, calculate the loss function according to the following formula
[0195]
[0196] Among them, the balance coefficient, (recommended α = 1.0, β = 0.5, γ = 0.2), is used to align the scale of RD with other loss terms; L is the improved loss function; Ω s is the scale-adaptive weight; Ω C is the confidence guide weight, IoU is the original loss function, R D is the gradient equalization distance penalty, R A is the shape alignment penalty,
[0197] The scale adaptation weight Ω s and the confidence guide weight Ω C The multiplication is essentially to implement a dual dynamic weighting mechanism, solving the joint optimization problem of target scale differences and model prediction reliability in complex tasks. Scale-adaptive weights address the learning imbalance problem of targets of different scales, while confidence-guided weights focus on uncertain model predictions, alleviating the loss dominated by easy examples.
[0198] In the calculation of RD itself, α is already D The constraint coefficient is to adjust the RD value itself. If the original gradient of RD is too large (for example, the distance penalty has too strong an impact on parameter update), the balance coefficient can reduce its gradient proportionally:
[0199] Improvements to the RetineX algorithm, the attention mechanism, and the optimized yolov8 loss function allow the three to be cascaded, but each solves different recognition problems. The better the preprocessing in the front, the more accurate the feature extraction in the back. The more accurate the feature extraction, the smaller the value of the loss function and the more accurate the recognition.
[0200] Improve the RetineX algorithm for image enhancement preprocessing to reduce the interference of background images.
[0201] By connecting the two-level illumination suppression of pixel domain ILF and feature domain CA in series, the pixel domain solves the “pixel value distortion caused directly by illumination” (such as overexposure and shadow), and the feature domain solves the “feature confusion caused indirectly by illumination” (such as the overlap of background and target features), forming hierarchical complementarity.
[0202] The loss function is improved by adding logarithmic space transformation and exponential weighting to the original IoU function to strengthen the loss weight for small objects. Normalized center point distance and diagonal constraints are introduced to address the vanishing gradient in non-overlapping scenes. An angle-aware mechanism is introduced to ensure the aspect ratio consistency between the predicted and ground-truth boxes. The prediction confidence is integrated to dynamically adjust the loss weight. The improved loss function ensures more accurate recognition and avoids duplicate recognition.
[0203] Beneficial effects of the present invention: The proposed diamond wire saw bead detection method, system, and medium based on the improved YOLOv8 significantly improve the detection performance of diamond wire saw beads in complex industrial environments through multi-module collaborative optimization and algorithm innovation. This is of great significance for promoting the automation and intelligence of diamond wire saw production. The specific beneficial effects are as follows:
[0204] By introducing the illumination-invariant feature extractor (LIF) into the backbone network, combined with x- and y-directional gradient extraction and Gaussian filter convolution, it effectively captures the edge and structural features of small beaded objects, addressing the shortcomings of YOLOv8 in extracting insufficient detail in small objects. Furthermore, the CA attention mechanism, integrated into the C2f module, generates refined attention weights through horizontal and vertical feature decomposition and fusion, enhancing the model's focus on the beaded objects and avoiding background interference. Experimental data shows that using either the feature extraction module or the attention mechanism module alone increases mAP to 69.1% and 72.3%, respectively, while reducing false positive rates to 12.9% and 10.5%.
[0205] To address the dense distribution and occlusion of beads, the optimized loss function incorporates scale adaptation, gradient equalization, shape constraints, and confidence guidance. This effectively addresses the vanishing gradient problem in non-overlapping scenes, strengthens the loss weight for small objects, and simultaneously constrains the aspect ratio consistency between the predicted and ground-truth bounding boxes, improving the detection accuracy of irregularly shaped objects. The combined full-module approach achieves a mean average detection accuracy of 83.4% and a false positive rate as low as 5.1%, significantly improving detection accuracy compared to the native YOLOv8 solution.
[0206] The improved Retinex algorithm intelligently adjusts preprocessing parameters based on the average grayscale value of the image, and performs logarithmic transformation, exponential transformation, and histogram equalization on images under different lighting conditions, effectively reducing the interference of background images and solving the problem of pixel value distortion directly caused by lighting (such as overexposure and shadows). At the same time, the output of the pixel domain is used as the input of the feature domain, enabling the attention mechanism to locate the target more accurately, avoiding the situation where the background is mistakenly regarded as the target due to strong lighting interference in the original features, and solving the problem of feature confusion indirectly caused by lighting (such as overlapping features of the background and the target). Experimental data show that using the preprocessing module alone, the mAP is increased to 68.5%, and the false detection rate is reduced to 13.2, indicating that lighting preprocessing plays an important role in improving detection stability.
[0207] Through an algorithm based on feature entropy, the pixel-domain σ and feature-domain r parameters are dynamically and collaboratively adjusted, achieving hierarchical complementarity in illumination suppression between the pixel and feature domains. When the entropy value of the ILF output feature map reflects high feature complexity, the dimensionality reduction ratio is automatically adjusted, enabling the model to adaptively optimize feature extraction and attention allocation under varying lighting conditions, further improving the stability of detection performance.
[0208] This invention automates the entire process, from image acquisition, preprocessing, feature extraction, target detection, to output, eliminating the need for human intervention and significantly reducing the time and cost of manual inspection. By installing an industrial camera to capture raw bead images and using the labImage annotation tool to generate training, validation, and test data sets, combined with an improved YOLOv8 model for training and inference, it can quickly and accurately detect the number and position of beads, providing a reliable detection basis for automated diamond wire saw injection molding.
[0209] While improving the model, we deleted the large-scale detection head and added the small-target detection head to optimize the beads for the small scale variation. While ensuring the detection accuracy, we avoided unnecessary computational overhead and maintained a high detection speed to meet the real-time detection needs in industrial production.
[0210] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A diamond wire saw steel ball detection method based on improved yolov8, characterized in that: include: S1, collects the original image of diamond wire saw beads; And use labImage label images to annotate beads, and form training set, validation set, and test set data according to the proportion; S2, preprocessing the image, after scaling the image, using the improved Retinex algorithm to preprocess the image; S3 improves the yolov8 model by adding an illumination-invariant feature extractor (LIF) to the backbone. The CA attention mechanism is integrated into the C2f module. Based on the small size variation of beads, a small target detection head is added while the large-scale detection head is removed, making it suitable for detecting small targets in beads. S4, optimize the loss function of yolov8, use the improved IoU loss function to improve the accuracy of small target detection, use the improved yolov8 model for training, and obtain the training result best.pt file.
2. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 1, characterized in that, The S2 includes: S2-1, use the improved RetineX algorithm to perform image enhancement preprocessing on the image to reduce the interference of the background image; Obtain a scaled diamond steel ball image; calculate the grayscale average value V of the image and perform logarithmic transformation on smaller V and exponential transformation on larger V according to the following formula, and obtain a better quality image I1 through histogram equalization; Where: x and y are pixel coordinates; g(x,y) is the output pixel value; f(x,y) is the original pixel value; a, b, c are custom parameters; When 0≤V<100, logarithmic transformation is used to stretch the low-light area and improve the detail contrast in low light; When the illumination is medium (100 ≤ V ≤ 170), the original pixel value is directly output; When the light intensity is 170<V≤255, the exponential function is used to reduce overexposure interference and maintain background details.
3. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 2, characterized in that, Said S2 further comprises: S2-2, performing Wiener filtering on the I1 image to obtain image I2; The steps for performing Wiener filtering on an image are as follows: Calculate the mean and variance of each pixel in the image within the range of M*N. The formula is as follows: M, N: The range of the image to be filtered by Wiener, M represents width, N represents height; x, y: represents pixel coordinates, η is the M*N range set, μ is the mean value of pixels in the M*N range, σ 2 is the pixel variance within the M*N range.
4. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 3, characterized in that: Said S2 further comprises: S2-3, calculate the average value of the local variance, the formula is as follows: Where Z represents the number of pixels; Represents the pixel variance of a point i within the M*N range, v 2 Indicates the average value of pixel variance within the M*N range. The new image pixel value is calculated as follows: Perform logarithmic operation on the I1 and I2 images; obtain images logI1 and logI2; subtract logI2 from logI1 to obtain the logI image; and perform normalization on the logI image.
5. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 1, characterized in that, The S3 includes: S3-1, calculate the gradient images of the image in the x and y directions respectively. The formula is as follows: f x ′(x,y)=f(x+1,y)-f(x,y); f y ′(x,y)=f(x,y+1)-f(x,y); Convolve the gradient image with the Gaussian filter to obtain the gradient component feature information in the x and y directions. The formula is as follows: is the convolution operation, F x (x,y) and F y (x, y) are the gradient component features of the image in two different directions, X and Y; G x (x,y,σ) and G y (x, y, σ) are Gaussian filters in two different directions, X and Y respectively; σ is the standard deviation of the Gaussian kernel, which controls the smoothness of the filter. The larger σ is, the better the smoothing effect is. G x (x,y,σ) represents the derivative of G(x,y,σ) in the x direction G y (x,y,σ) represents the derivative of G(x,y,σ) in the y direction, F x (x,y) and F y Perform the inverse tangent operation on (x,y) to obtain the gradient characteristics of the image. The formula is as follows: The principles, steps, and network structure of the CA attention mechanism are as follows: S3-2, CA attention mechanism decomposes the input features into horizontal and vertical feature vectors, then performs feature fusion, generates attention weights, and extracts more refined target features. The steps are as follows: First, the input feature map is divided into two directions, width and height, and global average pooling is performed separately to obtain feature maps in the width and height directions respectively. The formula is as follows: in, is the global average pooling feature vector in the height direction; c is the channel index of the feature map, and h is the position index of the current operation in the height direction; is the global average pooling feature vector in the width direction; c is the channel index of the feature map, w is the position index of the current operation in the width direction; x c (h,i) is the pixel value of the cth channel in the input feature map with coordinates (h,i); x c (j,w) is the pixel value of the cth channel in the input feature map with coordinates (j,w), where i and j represent the indexes in the height and width directions, respectively.
6. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 5, characterized in that: The S3 includes: S3-3, the feature maps of the width and height of the global receptive field are spliced together, and then they are sent to the convolution module with a shared convolution kernel of 1×1 to reduce their dimension to the original C / r. The batch normalized feature map F1 is then sent to the Sigmoid activation function to obtain a feature map f of the shape 1×(W+H)×C / r. The formula is as follows: f=δ(F1([z h ,z w ])); Among them, z h is the feature map in the height direction; z w is the feature map in the width direction; F1(*) is the 1*1 convolution operation; δ(*) represents the calculation result of the sigmoid function with * as the parameter, Then, the feature map f is convolved with a kernel of 1×1 according to the original height and width to obtain feature maps with the same number of channels as the original. After the Sigmoid activation function, the attention weights g of the feature map in height and width are obtained respectively. w and the attention weight g in the width direction h , the formula is: g w =σ(F w (f w )) g h =σ(F h (f h )) F w (*) and F h (*) represents the 1×1 convolution operation in the width and height directions respectively; f w and f h Represents z h and z w The feature maps in width and height directions are obtained after 1*1 convolution operation and Sigmoid activation function. S3-4, through multiplication weighted calculation on the original feature map, the final feature map with attention weights in the width and height directions will be obtained. The formula is as follows: y c (i, j) is the pixel value at coordinate (i, j) in the cth channel of the final feature map with attention weight; x c (i, j) is the pixel value at coordinate (i, j) in the c-th channel of the feature map input to the CA attention mechanism; is the attention weight in the height direction, corresponding to the weight value at the height position i of the c-th channel; is the attention weight in the width direction, corresponding to the weight value at the width position j of the c-th channel; i and j represent the index in the height and width directions respectively.
7. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 1, characterized in that: The S4 includes: S4-1, calculate the prediction box B p With the real box B g Intersection over Union (IoU) Through logarithmic space conversion and exponential weighting, the loss weight of small targets is strengthened. Among them, s is the target box area, s ref is the reference area, i.e., the average target area of the dataset, λ is the scale sensitivity coefficient, ranging from 1.0 to 2.0; the scale adaptation weight is Ω s , when s<s ref When Ω s >1, increase the loss weight, otherwise reduce the weight; S4-2, introduces normalized center point distance and diagonal constraints to solve the gradient vanishing problem in non-overlapping scenes; Where: k is the distance between the center of the predicted box and the real box, c is the diagonal length of the minimum circumscribed rectangle, ω is the gradient smoothing index, ρ is the distance attenuation coefficient, α D is the constraint coefficient, R D is the gradient equalization distance penalty.
8. The diamond wire saw steel ball detection method based on improved yolov8 according to claim 7, characterized in that: Said S4 further comprises: S4-3, constrain the aspect ratio consistency between the predicted box and the real box, and introduce an angle perception mechanism. Among them, w p and h p are respectively the width and height of the prediction box, w g h g are the width and height of the real frame respectively; A is the shape penalty coefficient, ranging from 0.5 to 1.0; π 2 is the square of pi, R A To penalize shape alignment, the dual constraint mechanism simultaneously optimizes the aspect ratio angle and scale difference to improve the detection accuracy of irregularly shaped objects. and The angle constraint mechanism is introduced to perform dual constraints together with the aspect ratio to improve the detection accuracy of irregularly shaped targets. S4-4, integrate prediction confidence and dynamically adjust loss weights, Among them, C is the confidence score of the prediction box, Focus coefficient, ε stability coefficient, Ω C The confidence-guided weights are applied to low-confidence samples, giving priority to optimizing the localization accuracy of difficult samples. S4-5, calculate the loss function according to the following formula Among them, the balance coefficient, (recommended α = 1.0, β = 0.5, γ = 0.2), is used to align the scale of RD with other loss terms; L is the improved loss function; Ω s is the scale-adaptive weight; Ω C is the confidence guide weight, IoU is the original loss function, R D is the gradient equalization distance penalty, R A Penalizes shape alignment.
9. A computer system, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the application method according to any one of claims 1 to 8 when executing the executable instructions.
10. A computer-readable storage medium, characterized in that include: a memory having a computer program stored thereon; A processor is configured to execute the program in the memory to implement the application method according to any one of claims 1 to 8.
Citation Information
Cited By
Low-proportion infrared small target detection method and device
CN121767790A