Bearing surface defect automatic detection system based on image recognition

By adopting image recognition technology in bearing surface defect detection systems, including image preprocessing, feature extraction and U-Net network recognition, the problem of insufficient accuracy and robustness of defect detection in the prior art is solved, and more accurate and reliable defect recognition and severity assessment are achieved.

CN120147263AActive Publication Date: 2025-06-13SHANDONG REHE BEARING TECH CO LTD

Patent Information

Application Number
CN202510221809.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art fails to fully consider the contrast difference between the defect area and the background in the detection of bearing surface defects, resulting in the masking of small defects, low detection accuracy, poor defect classification, and a single method of severity assessment.

Method used

An automatic detection system for surface defects based on image recognition is adopted, including image acquisition and preprocessing module, feature extraction and enhancement module, defect classification and labeling module, and result output and reporting module. Local contrast and texture features are extracted through image cropping, filtering, and contrast adjustment, defects are identified using U-Net networks, and defect severity is calculated.

Benefits of technology

It improves the distinction between defect areas and background, enhances the ability to identify small defects, improves the robustness of defect classification and the accuracy of severity assessment, and meets the detection needs under complex operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147263A_ABST
    Figure CN120147263A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a bearing surface defect automatic detection system based on image recognition, and the system comprises an image collection and preprocessing module which obtains a bearing surface image, and carries out the cutting and filtering of the image, so as to remove the noise and adjust the contrast, and obtain a processed image; and carrying out gray level conversion on the processed image to generate a preprocessing result. According to the method, environmental noise interference is reduced through image cutting and filtering, the contrast ratio is adjusted, the image quality is optimized, and the distinction degree of the defect area and the background is improved. Image channel information is standardized through gray level conversion, so that the calculation precision of feature extraction is kept consistent, and the analysis result is prevented from being interfered by multi-channel data. According to the method, image hierarchies with different scales are constructed by an image pyramid method, so that the adaptability to fine cracks and large-area spalling defects is improved in the local contrast and texture feature extraction process, and the stable recognition capability to defects with different sizes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an automatic bearing surface defect detection system based on image recognition. Background Art

[0002] Image recognition is a technical field based on computer vision and pattern recognition, aiming to enable a computer to understand and analyze the content in an image. The automatic bearing surface defect detection system uses image recognition technology to detect the bearing surface to identify potential defects such as cracks, spalls, and wear.

[0003] However, the prior art does not fully consider the contrast difference between the defect area and the background, which easily causes small defects to be covered by the background information with strong contrast, reducing the detection accuracy. The defect classification method mainly relies on morphological feature matching, and has low robustness for classifying complex defects. In the case of blurred boundaries or similar morphologies, misclassification problems are likely to occur. The evaluation method of the severity only relies on simple area calculation. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the deficiencies in the prior art, and to propose an automatic bearing surface defect detection system based on image recognition.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An automatic bearing surface defect detection system based on image recognition includes:

[0006] An image acquisition and preprocessing module, which acquires the bearing surface image, crops and filters the image to remove noise and adjust the contrast, obtaining a processed image; performs gray conversion on the processed image to generate a preprocessing result;

[0007] A feature extraction and enhancement module, based on the preprocessing result, decomposes the image through image pyramid, extracts local contrast and texture features at each scale respectively, generating a feature set; adjusts the enhancement parameters according to the statistical distribution of the feature set to obtain an enhanced feature map;

[0008] A defect classification and annotation module, based on the enhanced feature map, uses a U-Net network to analyze the image, identifies crack and spall defects on the bearing surface, obtaining a classified defect result;

[0009] A result output and reporting module, based on the classified defect result, sorts out the position, type, and severity of each detected defect, obtaining a defect result.

[0010] Preferably, the step of obtaining the processed image is:

[0011] Obtain the image of the bearing surface, crop and remove the invalid areas at the edges of the image, and optimize the signal-to-noise ratio of the image using spatial domain filtering to obtain the cropped and filtered image;

[0012] According to the cropped and filtered image, calculate the contrast enhancement index, and the calculation formula is:

[0013]

[0014] where C represents the contrast enhancement index, I i,j represents the pixel gray value of the i-th row and j-th column in the cropped and filtered image, M represents the global average gray value of the cropped and filtered image, I max represents the maximum gray value in the cropped and filtered image, I min represents the minimum gray value in the cropped and filtered image, and m and n respectively represent the number of rows and columns of the image;

[0015] According to the contrast enhancement index, adjust the gray histogram distribution of the cropped and filtered image to obtain the processed image.

[0016] Preferably, the step of obtaining the preprocessing result is:

[0017] Obtain the processed image, separate the pixel values of the red channel, green channel, and blue channel, and at the same time convert to the standardized color space to obtain the image after separating the channels;

[0018] According to the image after separating the channels, calculate the gray value distribution balance degree, and the calculation formula is:

[0019]

[0020] where AG represents the gray value distribution balance degree, R represents the pixel value of the red channel in the image after separating the channels, G represents the pixel value of the green channel in the image after separating the channels, B represents the pixel value of the blue channel in the image after separating the channels, and HM represents the global average gray value of the image after separating the channels;

[0021] According to the gray value distribution balance degree, perform gray mapping on the image after separating the channels to generate the preprocessing result.

[0022] Preferably, the step of obtaining the feature set is:

[0023] Obtain the preprocessing result, decompose the image using the image pyramid, construct image levels of different scales to obtain the image after multi-scale decomposition;

[0024] According to the image after multi-scale decomposition, calculate the local contrast score, and the calculation formula is:

[0025]

[0026] Among them, CF represents the local contrast score, and P i,j represents the gray value of the pixel at the j-th position in the i-th scale of the image after multi-scale decomposition. N represents the number of local pixel blocks at the current scale, and P avg represents the average gray value of all pixels at the current scale, and P max represents the maximum gray value of the pixels at the current scale, and P min represents the minimum gray value of the pixels at the current scale. k represents the number of scales obtained after image decomposition, and z represents the total number of pixels at the current scale;

[0027] According to the local contrast score and combined with the texture distribution characteristics at different scales, the structural change situation of the image at each scale is extracted to generate a feature set.

[0028] Preferably, the steps for obtaining the enhanced feature map are as follows:

[0029] Based on the feature set, the enhancement parameter is calculated, and the calculation formula is:

[0030]

[0031] Among them, E represents the enhancement parameter, and F n represents the value of the n-th feature, σ n represents the standard deviation of the n-th feature, and μ n represents the average value of the n-th feature, and TN is the total number of features in the feature set;

[0032] According to the enhancement parameter, the intensity of each feature is adjusted to obtain the enhanced feature map.

[0033] Preferably, the steps for obtaining the classified defect result are as follows:

[0034] The U-Net network is used to process the enhanced feature map to identify and label the crack and spalling defects in the image, and the unoptimized defect recognition result is obtained;

[0035] Image post-processing is performed on the unoptimized defect recognition result, including using threshold segmentation to separate the foreground and background, and applying morphological dilation and erosion to clarify the defect boundaries to generate the classified defect result.

[0036] Preferably, the steps for obtaining the defect result are as follows:

[0037] The classified defect result is obtained, the position information of each identified defect is extracted, and the boundary shape and size in the image coordinates are analyzed to form the spatial distribution data of the defects;

[0038] Calculate the defect severity score based on the spatial distribution data of the defects. The calculation formula is as follows:

[0039]

[0040] where D represents the defect severity score, X i , Y i represents the central coordinates of the i-th defect, X c , Y c represents the centroid coordinates of all defects, FG i represents the average gray gradient of the i-th defect area, A i represents the area of the i-th defect, G med represents the median gray gradient of all defect areas, M represents the total number of detected defects;

[0041] Based on the defect severity score, organize the spatial position, type, and severity of each defect to generate defect results.

[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0043] In the present invention, environmental noise interference is reduced through image cropping and filtering, the image quality is optimized by adjusting the contrast, and the distinguishability between the defect area and the background is improved. Gray conversion standardizes the image channel information, keeps the calculation accuracy of feature extraction consistent, and avoids the interference of multi-channel data on the analysis results. The image pyramid method constructs image levels of different scales, improves the adaptability to fine cracks and large-area spalling defects during the extraction of local contrast and texture features, and enhances the stable recognition ability for defects of different sizes. Feature enhancement is based on statistical distribution adjustment, making the high-frequency features of the defect area more prominent and reducing the interference of surface texture complexity on defect recognition. The severity calculation combines the defect position, area, and boundary gradient distribution, making the impact degree of the defect more quantitative and meeting the detection requirements under various complex working conditions. The output of defect classification and numerical information enables the direct use of defect data in the maintenance and quality monitoring links, improving the efficiency and operability of the decision-making process. Description of the Drawings

[0044] Figure 1 is the system flow chart of the present invention. Detailed Embodiments

[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] Please refer to Figure 1, the present invention provides a technical solution: An automatic bearing surface defect detection system based on image recognition includes:

[0047] An image acquisition and preprocessing module, which acquires the bearing surface image, crops and filters the image to remove noise and adjust the contrast, obtaining a processed image; performs gray-scale conversion on the processed image to generate a preprocessing result;

[0048] A feature extraction and enhancement module, based on the preprocessing result, decomposes the image through image pyramid, extracts local contrast and texture features at each scale respectively, generating a feature set; adjusts the enhancement parameters according to the statistical distribution of the feature set to obtain an enhanced feature map;

[0049] A defect classification and annotation module, based on the enhanced feature map, uses the U-Net network to analyze the image, identifies crack and spalling defects on the bearing surface, obtaining a classified defect result;

[0050] A result output and reporting module, based on the classified defect result, sorts out the position, type and severity of each detected defect, obtaining a defect result.

[0051] The steps for obtaining the processed image are as follows:

[0052] Acquire the bearing surface image, crop and remove the invalid area at the edge of the image, and optimize the signal-to-noise ratio of the image using spatial domain filtering to obtain a cropped and filtered image;

[0053] According to the cropped and filtered image, calculate the contrast enhancement index, and the calculation formula is:

[0054]

[0055] where C represents the contrast enhancement index, I i,j represents the pixel gray value of the i-th row and j-th column in the cropped and filtered image, M represents the global average gray value of the cropped and filtered image, I max represents the maximum gray value in the cropped and filtered image, I min represents the minimum gray value in the cropped and filtered image, and m and n respectively represent the number of rows and columns of the image;

[0056] According to the contrast enhancement index, adjust the gray histogram distribution of the cropped and filtered image to obtain the processed image.

[0057] Specifically, based on the original images of the bearing surface collected, an industrial camera with a resolution of 1920×1080 is selected to capture 15 frames per second. The grayscale mean distribution data of the image edges is read from the original frames. By comparing the contrast of each pixel in these edge regions with the internal pixels, when the fluctuation amplitude of the continuous grayscale values at the edges is less than 5 and the brightness of the corresponding pixels does not exceed 10, this region is determined as an invalid region. After completing the judgment of all edge pixels, 15 rows are removed from the top and bottom of the image respectively, and 20 columns are removed from the left and right sides. The covered range of the remaining area is determined within a rectangle of 800×1040 in the center. Then, a spatial domain filtering process is applied to the remaining area. The size of the filtering core is determined to be 3×3 from the noise distribution map collected from the industrial site, and it is scanned pixel by pixel in combination with the convolution method. Matrix operations are performed on the values of adjacent pixels and the central pixel to obtain the spatial domain filtering output. Each pixel point is weighted and accumulated by the values of 8 surrounding pixels. Then, according to the noise signal level data collected previously, pixel points with a grayscale jitter exceeding 2 and a continuous range greater than 3 rows are further smoothed by averaging. To avoid texture detail loss in high-brightness or low-brightness regions, by tracking the noise peak distribution range recorded in the industrial site, the brightness range of 80 to 220 in the range of 0 to 255 is taken as the contrast boundary. If it exceeds 220, the original value remains unchanged. If it is lower than 80, fine-tuning is performed through local diffusion. After all pixels are processed, the cropped and filtered image is obtained.

[0058] The advantage of the formula is that by accumulating the absolute values of the differences between all pixels and the global average in the numerator and combining the comprehensive information of the variance of pixel grayscale from the mean and the difference between the maximum and minimum grayscale values in the denominator, a quantitative evaluation of the overall contrast improvement requirement is achieved, so that more targeted adjustments can be made to the histogram distribution in the subsequent stage according to this quantitative result.

[0059] I i,j The acquisition steps of

[0060] This is the pixel grayscale value of the i-th row and j-th column in the cropped and filtered image, which needs to be read one by one in row and column order from the remaining image pixel array after completing the spatial domain filtering and removing the invalid regions. To obtain this parameter, first, the edge rows and columns with obvious noise distribution are deducted from the original image collected by the industrial camera, and the known thermal interference data and grayscale jump situations are compared with the central section of the image. Only the pixels that meet the smoothness requirements will be retained. Technicians will screen out the stable pixel data most suitable for calculation by repeatedly collecting images of the same size 10 times on-site. For example, in a 5×5 image matrix, I 1,1 = 90, I 1,2 = 95, I 2,3 = 99, etc.

[0061] The steps to obtain M are as follows:

[0062] This parameter represents the global average gray value of the image after cropping and filtering. It can be obtained by summing the gray values of all valid pixels and then dividing by the total number of pixels m×n. During the acquisition process, an accumulation function can be called on an automated platform to add the gray values of each pixel one by one, and then divide the result by the total number of pixels to obtain M. In actual data recording, if some pixels are affected by occasional transient light fluctuations, the gray values of these pixels can be compromised by referring to the results of multiple on-site acquisitions in the previous step. For example, after statistically analyzing a 5×5 test image, the sum of all pixels is 2480, and the number of image pixels is 25, then M = 2480 / 25 = 99.2.

[0063] I max The steps to obtain it are as follows:

[0064] This is the maximum gray value in the image after cropping and filtering. The one with the highest value is directly found and determined from the list of all obtained valid pixels as I max . In the test images at the industrial site, all pixels will be read in row-major order first, and then the gray values of each pixel will be compared to lock and record the highest pixel value, which will be applied to the denominator part of the formula to measure the dynamic range of the gray level of the entire image. In an actual calculation example, if the highest value among the pixel values of a 5×5 test image is 115, then I max = 115.

[0065] I min The steps to obtain it are as follows:

[0066] This parameter corresponds to the minimum gray value in the image after cropping and filtering. Through a scanning method similar to that of I max , the one with the lowest value is selected from all pixel gray values as I min . In a computer program, the array index can be first positioned to the first pixel value and used as the current minimum value, and then compared with all subsequent pixels one by one. If there is a lower gray value, the minimum value is updated. Taking the above 5×5 test image as an example, the minimum pixel value may be 87, then I min = 87.

[0067] The steps to obtain m are as follows:

[0068] This parameter represents the number of rows of the image and is usually obtained by acquiring the height information of the image data. In the actual field, if the remaining number of pixel rows changes after filtering or removing the edges of the image obtained by an industrial camera, the height value of the valid area of the image needs to be read again to update m. In a 5×5 test image, 5 rows of pixels remain after cropping and filtering, then m = 5.

[0069] The steps to obtain n are:

[0070] This parameter represents the number of columns in the image, which is similar to the acquisition process of m. During on-site testing, n is obtained based on the width of the cropped image. Through row-first or column-first scanning, the software can determine the number of pixel columns actually contained in the current image. Taking a 5×5 test image as an example, if 5 columns of pixels can still be maintained, then n=5.

[0071] Calculation process:

[0072] First, read data from the pixel matrix of the 5×5 test image, as shown in the following matrix example:

[0073] 90, 95, 100, 105, 110

[0074] 92, 96, 99, 106, 112

[0075] 88, 94, 101, 103, 115

[0076] 87, 97, 98, 108, 113

[0077] 90, 92, 104, 109, 114

[0078] Add up all 25 pixel values ​​to get 2480, divide by 25 to get M = 99.2, and then traverse the absolute value of the difference between all pixels and M to get the numerator:

[0079]

[0080] Then calculate the denominator, first find the sum of the squares of the difference between the pixel and M:

[0081]

[0082] Adding the squares of all pixel differences, we get a value (omitted here for itemization) of about 1870.24, and find I max =115 and I min =87 and then calculate the denominator:

[0083]

[0084] Dividing the numerator by the denominator, we get:

[0085]

[0086] The result shows that the contrast enhancement index of the current image reaches 2.94. When combined with the subsequent histogram distribution adjustment, the incremental demand for image contrast can be determined based on the value of 2.94, and the calculation result can be used to guide a more detailed grayscale allocation process.

[0087] According to the value of the aforementioned contrast enhancement index C, read the value within the range of 2.00 to 3.50, divide the grayscale histogram of the current image into 16 levels of distribution intervals, then count the corresponding grayscale cumulative quantity for each interval one by one, and judge the relative concentration degree of pixel distribution in each interval through the statistical mean and median. When the grayscale cumulative quantity in a certain interval exceeds 500 pixels, redistribute the grayscale distribution of this interval to the adjacent interval in a small range, and combine the grayscale mean sequence collected previously during the calculation to compress the brightness values exceeding 200. The specific approach includes traversing each pixel in row-column order. When the grayscale value of the pixel is greater than or equal to 200, map it to 200 plus a random compensation amount of approximately 1 to 5. This compensation amount is obtained by observing the long-term collected brightness record table on-site, and at the same time referring to the grayscale difference array obtained from multiple previous detections. For the case where the contrast enhancement index is near 2.94, appropriately reduce the pixel accumulation in the highest grayscale interval to ensure that the distribution of each grayscale level in the overall range is relatively balanced. Then, stretch or compress each interval according to the pixel distribution density, shift the grayscale values in the low-density interval upward by a value of 5 to 10, and shift the grayscale values in the high-density interval downward by a value of 3 to 7. After all the mappings are completed, perform a global check. When it is detected that there are still excessive pixel accumulations in some intervals, continue to use the same method for refined mapping. Finally, confirm that the number of pixels in all intervals is at an acceptable level, and record the grayscale distribution after the overall mapping to obtain the processed image.

[0088] The steps for obtaining the preprocessing result are as follows:

[0089] Obtain the processed image, separate the pixel values of the red channel, green channel, and blue channel, and at the same time convert to the standardized color space to obtain the image after separating the channels;

[0090] According to the image after separating the channels, calculate the grayscale value distribution balance degree, and the calculation formula is:

[0091]

[0092] Among them, AG represents the grayscale value distribution balance degree, R represents the pixel value of the red channel in the image after separating the channels, G represents the pixel value of the green channel in the image after separating the channels, B represents the pixel value of the blue channel in the image after separating the channels, and HM represents the global average grayscale value of the image after separating the channels;

[0093] According to the grayscale value distribution balance degree, perform grayscale mapping on the image after separating the channels to generate the preprocessing result.

[0094] Specifically, based on the processed image obtained previously, the gray values of the red, green, and blue channels are read pixel by pixel, and the corresponding position indices are sequentially extracted according to the pre-recorded pixel arrangement order. Combining the exposure and gain setting data during the industrial camera acquisition, in the internal calculation, the original readings of each channel are compared with the on-site measured light compensation table. When the channel gray value exceeds 255 or is lower than 0, boundary limits are imposed according to the actually observed channel range, and the extreme out-of-limit values are corrected to an acceptable range in a linear mapping manner. The three separated channel values are combined to form the original RGB data information. Subsequently, according to the sRGB standard color conversion method, the separated red, green, and blue values are multiplied by a determined matrix to obtain the standardized tristimulus values. The specific approach includes first recording the separated channel data in a two-dimensional data table with three columns, and then applying the conversion coefficient matrix collected from the industry manual row by row to obtain a more stable color representation form. These conversion coefficient matrices are determined by actual multiple color calibration processes and are uniformly fixed in the image processing step. When the converted channel values are distributed within the range of 0 to 255, the standardized color value of the current pixel is recorded. If it exceeds this range, it is truncated in combination with the previously measured brightness distribution curve on site and slightly adjusted with the average value of adjacent pixels. The entire conversion process maintains the corresponding relationship between channels so that the pixel channel information at the same position can be retrieved during subsequent shape detection or texture recognition. After completing the channel separation and standardization transformation of all pixels, the obtained data is integrated into a new image matrix, and during the row and column traversal, it is checked whether there are extreme situations where a large number of pixels are concentrated in the same color component. If it is detected that the channel values are concentrated between 230 and 255 in some rows or columns, these pixels are corrected with reference to the simultaneously collected brightness monitoring sequence to avoid distortion problems caused by overexposed data in subsequent stages. In the final stage, a whole scan and statistics are performed on the converted image to clarify the number of rows, columns of the image, and the pixel distribution range of each channel, obtaining the image after channel separation.

[0095] The advantage of the formula is that by quantifying the mutual difference degree of the three channels and the deviation of the total channel value from the global average gray value, the balance degree of the color distribution can be comprehensively examined in the same step, thereby providing an executable numerical reference for subsequent gray mapping.

[0096] The steps for obtaining R are as follows:

[0097] This parameter represents the pixel values of the red channel in the image after channel separation. The range of all pixel values is usually between 0 and 255. During acquisition, technicians measure the lighting conditions and the exposure parameters of the industrial camera on-site and fix its gain to maintain consistency. After imaging, the red channel is scanned in row-column order, and the R value of each pixel is extracted into a list, and its coordinate position in the resolution space is recorded to avoid offsets in subsequent calculations. If it is detected that the R value of some pixels exceeds 255 or is lower than 0, the pixel will be re-estimated based on the compensation coefficient measured in advance by the industrial camera to ensure that the final R value falls within the effective range of 0 to 255. For example, if the original read value of a pixel in a test image is 270, it is corrected to the range of 230 through on-site calibration, and then 230 is stored as the R value of this pixel.

[0098] The steps for obtaining G are as follows:

[0099] This parameter represents the pixel values of the green channel in the image after channel separation. When obtaining it, the same row-column order and coordinate correspondence method as the R value are required. During the scanning of the green channel, the brightness record of the shooting site is read first, and the possible extreme values are corrected accordingly by referring to the exposure benchmark and noise distribution of the industrial camera at different times. The final range of the G value is also between 0 and 255. If the peak value of 310 appears in the on-site record for individual pixels, it will be cropped according to the sensor saturation value collected, and the pixel will be set to the saturated value of 255.

[0100] The steps for obtaining B are as follows:

[0101] This parameter represents the pixel values of the blue channel in the image after channel separation. The acquisition method is the same as that of R and G. In order to maintain the corresponding relationship with other channels, it is necessary to ensure that the row-column coordinates of the same pixel position are read for R, G, and B simultaneously, so as to correctly reflect the differences in the three channels in subsequent calculations. On-site, the blue values at the same position will be compared through multiple shootings. If the blue values measured at this position fluctuate up and down by more than 20, the lighting environment around the pixel will be queried, and the on-site brightness monitoring record will be combined to determine whether numerical correction is required. After the B value is completely obtained, it is placed within the range of 0 to 255 to avoid numerical out-of-bounds during subsequent operations.

[0102] The steps for obtaining HM are as follows:

[0103] This parameter represents the global average grayscale value of the image after separating the channels, and the pixel values of the R, G, and B channels are used for comprehensive evaluation. During actual operation, all valid pixels in the image will be traversed. First, the approximate grayscale value of each pixel is calculated, and then the approximate grayscale values of all pixels are accumulated and divided by the total number of pixels in the image to obtain HM. In a scenario with a higher resolution, to prevent shadows or bright spots from overly affecting the statistics, extreme outliers need to be removed or corrected, and then the average is calculated based on the remaining total number of data. Field technicians may use a sliding window method to calculate the local mean of several pixels around each position and then summarize it to obtain a more stable HM. In an example area of 10,000 pixels, if the sum of the approximate grayscale values of all pixels is 920,000, then HM = 920,000 / 10,000 = 92.

[0104] Calculation process:

[0105] In a certain detection, the red, green, and blue channel values of a pixel are obtained as R = 130, G = 115, and B = 100 respectively, and HM = 110 is obtained from the global statistical result. Substitute them into the formula item by item:

[0106]

[0107] First, calculate the absolute value differences:

[0108] |130 - 115| = 15, |115 - 100| = 15, |100 - 130| = 30

[0109] Then, calculate the squares and sum them up:

[0110] 15 2 +15 2 +30 2 = 225 + 225 + 900 = 1350

[0111] Divide by 3 and then take the square root:

[0112]

[0113] Then, calculate the fraction on the right side:

[0115] 130 + 115 + 100 = 345, 3 × 110 = 330, 345 - 330 = 15, |15| = 15, 15 / 3 = 5

[0116] Therefore:

[0117] AG = 21.213 + 5 = 26.213

[0118] The results show that when the value of AG is 26.213, the differences between the three channels and the degree of deviation of the total channel value from the global average gray level are both relatively obvious, indicating that the current pixel has a relatively large color deviation. If this value generally exceeds 20 in the subsequent global statistics, it means that the color balance of the entire image is low, while if the average value is generally lower than 10, it indicates that the image channels are more likely to be closer to each other.

[0119] According to the gray level distribution balance AG obtained previously, when performing gray level mapping on the image after separating channels, first check whether there is a situation where a large area of channel values are concentrated in the range of 180 to 255 or 0 to 50 in the collected three-channel matrix. If it is found that the cumulative pixel number of a single channel in the above range exceeds 30% of the total number of pixels in the image, then refer to the statistics of the on-site brightness records and light distribution data previously carried out, and proportionally compress or increase the pixel values in the over-concentrated area, map them to the average level of adjacent interval values, and compare with the recorded color reference table during the mapping process to confirm that they are consistent with the registered benchmark shooting parameters of the industrial camera. Then, re-traverse the three-channel values of each pixel in the image in the order of resolution rows and columns, assign new gray level values to each channel using the linear mapping formula, and avoid numerical mutations during mapping by performing point-by-point verification on the previously obtained bright spot distribution map. Dispersedly adjust the pixels with an original brightness greater than 220 to the range of 200 to 220, and similarly increase the pixels below 30 and make them close to the range of 40 to 50. Reduce the number of extremely bright or dark areas through this process of channel mapping. Subsequently, conduct a unified statistics on the re-obtained channel values. If it is actually found that there is still an obvious accumulation in the brightness distribution, further correction shall be carried out according to the second group of brightness monitoring sequences collected from the equipment site until the main pixel distributions of each channel tend to be centered. Finally, merge all the mapped channel combinations to generate a new image matrix, and check whether the row and column positions of each pixel correctly match the previously reserved marking information to prevent coordinate misalignment during merging, and obtain the preprocessing result.

[0120] The steps for obtaining the feature set are as follows:

[0121] Obtain the preprocessing result, decompose the image using the image pyramid, construct image levels at different scales, and obtain the multi-scale decomposed image;

[0122] According to the multi-scale decomposed image, calculate the local contrast score, and the calculation formula is:

[0123]

[0124] Among them, CF represents the local contrast score, P i,jrepresents the gray value of the pixel at the j-th position in the i-th scale after multi-scale decomposition. N represents the number of local pixel blocks at the current scale, and P avg represents the average gray value of all pixels at the current scale, and P max represents the maximum gray value of pixels at the current scale, and P min represents the minimum gray value of pixels at the current scale. k represents the number of scales obtained after image decomposition, and z represents the total number of pixels at the current scale;

[0125] According to the local contrast score and combined with the texture distribution characteristics at different scales, the structural changes of the image at each scale are extracted to generate a feature set.

[0126] Specifically, after obtaining the preprocessing result, when reading the image information corresponding to the result and performing image pyramid decomposition, first divide the obtained image into several levels in a top-down hierarchical manner, and reduce the resolution of each level at a fixed ratio. For example, on an image with an original resolution of 1920×1080, smaller-sized images such as 960×540 and 480×270 are successively formed to perform parallel observations on different scales of the same scene. After the row-column indexing of each scaled image is rearranged, these sub-images are analyzed row by row, and a preliminary reference value range is set in combination with the previously recorded pixel gray distribution. For example, the gray value is compared in the range of 0 to 255, and the brightness concentration is used as a separate observation index. If the number of pixels in a certain area of a certain-level image distributed in the high-brightness interval of 180 to 255 is significantly higher than the average level, it is necessary to reconfirm with reference to the brightness statistical results collected in the industrial field for a long time. If it is confirmed that this high-brightness interval exceeds the preset threshold of 30%, histogram equalization processing is performed on the corresponding area of this level of image, and the pixel average value and variance during the process are recorded. After the gray balance of all levels is completed, a multi-scale image hierarchy is formed. Subsequently, all levels of images are matched and corresponding pixel by pixel to ensure consistent coordinate mapping between multiple scales. For areas that do not meet pixel coherence, unreasonable pixels are checked and corrected by referring to the previous noise distribution data. Through this observation of the multi-scale decomposed image, it is possible to better compare the illumination changes in the same area at different resolutions during subsequent processing. Finally, after the summary of all hierarchical image information is completed, the specific number of the current level and the corresponding row-column dimensions of each level are confirmed and recorded to obtain the multi-scale decomposed image.

[0127] The advantage of the formula is that by accumulating the absolute differences between pixels at multiple scales and the average value of local pixel blocks in the numerator part, and combining the form of squared deviation and range in the denominator, the discreteness of the local contrast of the image can be quantified, which helps to perform fine analysis on the image structure at different scales subsequently.

[0128] P i,j is obtained as follows:

[0129] This parameter represents the gray value of the image after multi-scale decomposition at the i-th scale and the j-th pixel. During actual measurement, it is necessary to traverse each level of the image in row-column order on-site, extract and record the gray values at the corresponding coordinates. When performing image pyramid decomposition, since the resolution of each level changes, the row and column indices of the j-th pixel are not exactly the same as those of other levels, but will be uniformly summarized into an index system to ensure that the order of all pixels can correspond. The key to obtaining this parameter is to scale, filter, or perform other types of denoising on the original image collected on-site, and then arrange the pixels according to the new resolution to obtain P at each level i,j . For example, at the first scale, if the image size is 512×384, the total number of pixels is 196608. By extracting the gray values of these 196608 pixels one by one, they are named P 1,1 , P 1,2 , …, at the second scale when the image size is 256×192, the total number of pixels is 49152, and similarly, P 2,1 , P 2,2 , … can be obtained

[0130] The steps to obtain N are as follows:

[0131] This parameter represents the number of local pixel blocks at the current scale. Each level of the image will be divided into several pixel blocks according to certain rules. These blocks can be non-overlapping or partially overlapping. During on-site detection, the size and number of local blocks will be determined based on the complexity of the bearing surface. If the resolution of the image after pyramid decomposition to a certain level is 256×192, it can be divided into 16×12 = 192 local blocks, then N = 192. To determine the accurate value of N, the technical personnel will first determine the range of texture they hope to observe at this level of scale, then divide the block size according to the distribution of texture details, and finally count the total number of pixel blocks. For example, in actual detection, if the size of each block area is set to 16×16 pixels, the entire 256×192 image can be subdivided into (256 / 16)×(192 / 16) = 16×12 = 192 blocks, that is, N = 192

[0132] P avg is obtained as follows:

[0133] This parameter represents the average gray value of all pixels at the current scale. In a specific scale, the gray values of all pixels at this scale need to be accumulated and then divided by the total number of pixels to obtain an overall gray mean value. During on-site application, all the pixels retained from the image collected by the industrial camera at this resolution will be traversed first, and all the gray values will be accumulated, and then divided by the total number of pixels of this image. In order to eliminate special abnormal pixels, technicians will query the previously collected brightness distribution records, make corresponding corrections or eliminations for pixels that are too bright or too dark and have a very small number, and then perform the mean operation. If in an image of size 256×192, the gray value of each pixel is distributed in the range of 0 to 255, the staff can record the traversal result. For example, if the total accumulated value is 286720, then the average gray value P avg = 286720 / (256×192) = 286720 / 49152 ≈ 5.83.

[0134] P max is obtained as follows:

[0135] This is the maximum gray value of the pixels at the current scale. It is necessary to scan within the same resolution range and find the pixel with the highest gray value, and use it as P max . In the field, the method of row-by-row retrieval is generally used to compare the gray values of each pixel, and the maximum value is retained as the final result. For example, in an image of 256×192, after traversal, it may be found that the gray value of a certain pixel reaches 220, and 220 is used as P max . When obtaining it, it is also necessary to perform a validity check on the collected original gray value. If the readings of some pixels are 275 due to noise or out-of-range factors, it is necessary to truncate or correct them with reference to the previously measured sensor saturation value of 255. Finally, the values not exceeding 255 after correction will be included in the candidate range of P max .

[0136] P min is obtained as follows:

[0137] This parameter represents the minimum gray value of the pixels at the current scale. Similar to P max , it is determined by traversing to find the pixel with the smallest gray value. During on-site recording, an initial minimum value will be set first, and then compared with other pixels one by one. If a smaller value is found, it will be updated. Finally, the global minimum gray value is obtained after traversing all pixels. If it is detected that a certain pixel has a value of -10, technicians will correct it to 0 according to the previously obtained zero lower limit, because the gray value of the industrial camera will not be less than 0 under normal operating conditions. For a certain level of image with a total number of pixels of 49152, if the lowest value is found to be 3, then P min = 3.

[0138] The steps to obtain k are as follows:

[0139] It represents the number of scales obtained after image decomposition. A common image pyramid setting in the industry is to continuously downsample the image to generate 2 to 4 scales without affecting the macroscopic details. The specific number of scales to be selected needs to be determined according to the actual requirements of the bearing surface defect detection task. If precise positioning of relatively small cracks is required, k can be increased so that sufficient details can be retained even at a smaller resolution level. In the field, technicians will refer to the accuracy requirements, detection time consumption, and hardware resources in the previous round of detection records to determine the number of layers of the pyramid and record this number in the system. For example, in an image detection task of 1024×768, using a three-layer pyramid can make k = 3.

[0140] The steps to obtain z are as follows:

[0141] It represents the total number of pixels at the current scale and directly corresponds to the resolution of the image at this scale. If the size of a certain level of image is 256×192, then the total number of pixels is 49152, and thus z = 49152. During on-site statistics, it is necessary to traverse the actual valid pixels at this resolution and record the grayscale values. If in order to exclude a very small number of invalid pixels (such as the blanks caused by uneven illumination at the image edge), these pixels can be excluded first and then the statistics are carried out. In this way, the obtained z will be slightly smaller than the theoretical product of the resolution. However, in most cases, all the pixels of the entire image will be directly included in the statistics, making z equal to the product of the pixel rows and columns.

[0142] Calculation process:

[0143] In a detection task, through the multi-scale decomposed images obtained previously, taking k = 3 means there are three levels. Set the number of pixels at a certain level of scale as z = 49152, and determine the number of local pixel blocks as N = 192. After traversing the collected grayscale data, calculate P avg = 5.83. When detecting the extreme value, P max = 220 and P min = 3. In the numerator part, it is necessary to first subtract the grayscale value of each pixel from the average value of its corresponding local block and take the absolute value and accumulate. For example, for the i = 1 level, the jth pixel belongs to a certain local block, then subtract it from the average value of this block and take the absolute value, and then traverse all the pixels at this level and sum them up, and then accumulate the results with the second level and the third level. If the grayscale of a typical pixel is 15 and the average value of the corresponding local block is 20, then the absolute difference is 5, and the rest of the pixels are processed similarly. After the statistics are completed, assume that the total absolute difference of the three levels is A;

[0144] Then look at the denominator part. It is necessary to subtract the grayscale value of each pixel from P avg = 5.83, square and accumulate, then take the square root, and compare it with |Pmax -P min Add them up to 217. Assume that the square root of the sum of the squared grayscale differences is B, then the denominator is B + 217.

[0145] Therefore:

[0146]

[0147] After traversing all pixels, when A = 256000 and B = 850 are obtained, the final calculation can be performed:

[0148]

[0149] The result shows that under the local block observation of the three-layer pyramid, the local contrast score of this image is 240.11, which belongs to a relatively high range, indicating that there are obvious differences between bright and dark regions in the multi-scale structure of the image. If a lower contrast is required during the detection process, the brightness of the image or the local equalization range can be adjusted, or the size of the local block can be refined. If this value is below 30, it means that the overall image is relatively smooth and lacks obvious brightness contrast.

[0150] According to the local contrast score CF obtained previously, process it in combination with the texture distribution records at each level in the three-layer image, and traverse all pixels in row and column order in each level of the image. After collecting the grayscale patterns of each pixel, compare adjacent pixels in sequence. When the comparison result shows that there is a grayscale fluctuation amplitude exceeding the pre-set threshold of 10 in the same local area and the corresponding pixel proportion is higher than 30% of the area of this local block, mark this area as a texture mutation area, and then extend the scan along the surrounding to capture the edge range. Through a large number of on-site sample comparisons, it is known that these texture mutation areas often correspond to the initial traces of wear or cracks on the bearing surface. After confirmation, record the spatial coordinates of these mutation areas and the grayscale mean distribution at each scale level, and compare them in sequence according to the order from the first level to the third level. If multiple levels contain mutation phenomena in the same or adjacent pixel coordinate ranges, it is determined as a structural change feature with a relatively high probability, and centralize and summarize it to form a statistical list of the image structure changes at different scales. In this list, list the pixel positions, local grayscale distributions, and fluctuation values of the mutation areas, and at the same time mark their historical occurrence frequencies in the previous detection samples. When all statistical summaries are completed, finally output this multi-scale feature set.

[0151] The steps to obtain the enhanced feature map are as follows:

[0152] Based on the feature set, calculate the enhancement parameter, and the calculation formula is:

[0153]

[0154] Among them, E represents the enhancement parameter, and F n represents the value of the nth feature, and σ n represents the standard deviation of the nth feature, and μ n represents the average value of the nth feature, and TN is the total number of features in the feature set;

[0155] According to the enhancement parameter, the intensity of each feature is adjusted to obtain an enhanced feature map.

[0156] Specifically, the advantage of the formula is that by introducing this denominator structure, a moderate enhancement amplitude can be maintained in the range of relatively large or small feature values, and the cube root is used after accumulating the products of all features, taking into account the balance between high and low feature values.

[0157] F n is obtained as follows:

[0158] This parameter represents the value of the nth feature, and its range can cover, for example, texture directionality, texture detail intensity, local contrast gradient, etc. When obtaining data, it is necessary to check each feature in the previously obtained feature set one by one, and express each feature in a quantitative manner. Taking texture directionality as an example, the gray-scale distribution records collected on-site can be used. According to the gradient change curve of the image at different angles, specific values are statistically obtained to characterize the direction intensity of the texture. If this feature contains multiple dimensions, the data of different dimensions will be first combined into a numerical expression, and then recorded item by item in the electronic table. During the whole process, the light monitoring sequence and texture statistical records saved in the device are required. Through continuous contrast analysis and principal component extraction, redundant information is eliminated, and finally the deterministic feature value F n is obtained. In an example, 4 significant texture features have been detected, and the values are 18, 30, 22, and 35 respectively. These values correspond to the 1st, 2nd, 3rd, and 4th features in the feature set respectively, that is, F 1 = 18, F 2 = 30, F 3 = 22, F 4 = 35.

[0159] σ n is obtained as follows:

[0160] This parameter represents the standard deviation of the nth feature, which is used to measure the degree of deviation of this feature from the uniform distribution in the collected samples. In actual detection, technicians will repeat the measurement of the same type of features under multiple shootings or different working conditions, store the obtained multiple groups of values in a continuous observation sequence, and then calculate the variance of each feature through the variance formula, and then take the square root to obtain the standard deviation σ nFor example, for a certain texture directionality feature, the feature value will be repeatedly obtained at different rotational speeds, different loads, and different lighting conditions on-site. If 5 numerical values are recorded: 19, 21, 20, 20, 22, the mean of this sequence, 20.4, can be calculated first. Then, calculate the sum of the squares of the differences between each numerical value and the mean in turn, divide by the sample size to obtain the variance, and finally take the square root to get σ n In this example, the calculated variance is approximately 1.84, and its square root is approximately 1.36. Therefore, σ n = 1.36.

[0161] μ n The steps to obtain μ are as follows:

[0162] This parameter represents the average value of the nth feature, which is usually obtained by taking the arithmetic mean of the numerical values obtained for the same feature in multiple acquisitions. When obtaining it on-site, it is necessary to first record the numerical values of the same feature at different time points or in different states, and then divide the sum of all recorded numerical values by the number of samples. To ensure the stability of the numerical values, it is necessary to eliminate possible extreme abnormal readings of the sensor and repeat the measurement during periods with large light fluctuations to ensure that μ n can truly reflect the level of the feature under normal detection conditions. For example, if the numerical values of this feature obtained in 5 detections are 10.5, 10.8, 11.2, 9.9, and 10.6 respectively, the sum after accumulation is 53.0, and the total number of samples is 5, then μ n = 53.0 / 5 = 10.6.

[0163] The steps to obtain TN are as follows:

[0164] This parameter represents the total number of features in the feature set, which needs to be clearly counted after all previous feature detections, screenings, and classifications are completed. During the acquisition process, the staff will first list all recognizable and quantifiable features, number them in the project order, and then confirm whether they meet the complete acquisition and calculation conditions. For features that meet the conditions, they are regarded as valid features and included in the set. Finally, the total number of these valid features is recorded as TN. In a complete bearing surface defect detection, common feature types include aspects such as gray-scale changes, edge strength, texture direction, first-order differences, second-order differences, etc. Each aspect may be split into several dimensions. Relevant personnel will record and number all features during the acceptance process for subsequent calls. If the final confirmed number of available features is 10, then TN = 10.

[0165] Calculation process:

[0166] To give a specific calculation example, combining the example feature values, standard deviations, and mean values listed above, select TN = 4 and specify:

[0167] F 1= 18, σ 1 = 1.36, μ 1 = 10.6

[0168] F 2 = 30, σ 2 = 1.50, μ 2 = 12.0

[0169] F 3 = 22, σ 3 = 2.00, μ 3 = 11.0

[0170] F 4 = 35, σ 4 = 2.30, μ 4 = 9.8

[0171] First, calculate the denominator in the logarithmic operation:

[0172]

[0173] Then, perform the following operations for each feature:

[0174]

[0175] Perform the operations item by item. The examples are as follows:

[0176] The first item:

[0177]

[0178] The second item:

[0179]

[0180] The third item:

[0181]

[0182] The fourth item:

[0183]

[0184] Denote the sum of the four items as S:

[0185] S = 32.94 + 66.60 + 43.12 + 84.00 = 226.66

[0186] Substitute into the formula:

[0187] E = (S) 1 / 3 = (226.66) 1 / 3

[0188] E ≈ 6.08

[0189] The result shows that the enhancement parameter corresponding to the current feature set is approximately 6.08. A value greater than 5 indicates a relatively high activity level of the overall features, making it easy to enlarge the intensity difference in subsequent enhancement processing. When the value is below 2, it represents that the overall feature distribution is relatively stable. Technicians can determine the distribution strategy of the increase amplitude for each feature based on the comparison relationship between the enhancement parameter and the range of 5 to 10.

[0190] Based on the enhancement parameter E obtained previously, read the value of this parameter within the range of 5 to 10, and first retrieve the corresponding data such as local contrast, texture dispersion, and edge sharpness in all the listed feature records. Then, compare the gray range, gradient range, and spatial distribution degree of each feature one by one from the feature sequence. When it is found that the gray range of a certain feature is between 50 and 200 and the gradient range is between 10 and 30, mark it as a feature with normal intensity and do not add an obvious increase in amplitude. When it is found in the same sequence that the gradient range of some features is higher than 30 and the feature value is between 220 and 255, set the intensity increase during enhancement processing to a larger multiple. The specific method includes positioning the corresponding feature to the pixel area in the image coordinates row by row and column by column, and then performing successive addition operations on the gray values of each pixel point to ensure that the pixel range of the feature is completely covered during the calculation process. For pixel values that have exceeded the preset threshold of 255, truncate them to control within 255. For the part of the feature value in the low range, that is, less than 50, increase the brightness by multiplying the enhancement parameter by 1.2 to 1.5 times. Record the part of the feature value that is very close to the middle range in a comparison table for subsequent centralized verification and confirmation of its distribution status. After the statistics, summarize all the features, and finally adjust to obtain the enhanced feature map by comparing the gradient mean and the enhancement results within each pixel range.

[0191] The steps to obtain the classification defect results are as follows:

[0192] Use the U-Net network to process the enhanced feature map, identify and label the cracks and spalling defects in the image to obtain the unoptimized defect recognition results;

[0193] Perform image post-processing on the unoptimized defect recognition results, including using threshold segmentation to separate the foreground and background, applying morphological dilation and erosion to clarify the defect boundaries, and generating the classification defect results.

[0194] Specifically, based on the previously obtained enhanced feature maps, texture and grayscale features of each pixel and its surrounding neighborhood are mapped. When preparing for training and testing the U-Net neural network, it is necessary to first screen multiple images containing crack and spalling defects from the previously annotated sample library. All images are divided into a training group and a validation group in an 8:2 ratio. Then, the training group images are read row by row and their corresponding true mask distribution data are loaded into memory. By gradually adjusting the learning rate and batch size of the network to control the training convergence process, during training, features are matched pixel by pixel with the target labels. After one forward propagation, the pixel-level difference is calculated and the cross-entropy is used as the loss value. The error is backpropagated layer by layer and the learnable parameters of the network are updated. After hundreds of iterations, the detection results in the validation group gradually stabilize, confirming that the network has the ability to distinguish between crack and spalling defects. Subsequently, the trained network is used to perform inference on each pixel of the previously obtained enhanced feature maps. In the encoding-decoding process under the U-shaped structure, multi-scale features are extracted and the defect edges and central regions are mapped. The inferred outputs are integrated and pixels with a matching degree exceeding the specified value of 0.8 are retained. Finally, a mask result of crack and spalling defects is generated for each input image, obtaining an unoptimized defect recognition result.

[0195] According to the unoptimized defect recognition results obtained previously, the predicted type information of each pixel is read and distinguished from the background pixels. When using the threshold segmentation method to further distinguish the foreground and background, it is necessary to statistically calculate the grayscale mean of typical defect pixels in the on-site collected data and define it as the reference value of 150. If some pixels exceed this reference value during detection, they are classified as the foreground, otherwise as the background. This threshold is a value set after comparing the mean distribution of 100 defect images. Then, when performing morphological dilation on the foreground area, each pixel is verified. If it is detected that there are multiple continuous markings of adjacent pixels and the pixel count exceeds the specified number of 100, they are classified into the same area. The horizontal and vertical coordinates of the pixels are recorded and an adjacency query is performed. Then, a 3×3 structuring element is applied to these pixels for dilation. Subsequently, an erosion operation is performed on the dilation result. When eroding, if it is found that more than half of the neighborhood around any pixel still maintains the foreground label, the pixel is retained, otherwise it is converted to the background. After the morphological processing, the type information of all pixels is checked in row and column order, and the final confirmation of the processed defect area is made, summarizing the crack and spalling ranges and their corresponding pixel markings therein to generate a classified defect result.

[0196] The steps for obtaining the defect results are as follows:

[0197] Obtain the classified defect results, extract the location information of each identified defect, analyze the boundary shape and size in the image coordinates, and form the spatial distribution data of the defects;

[0198] Based on the spatial distribution data of defects, calculate the defect severity score. The calculation formula is as follows:

[0199]

[0200] where D represents the defect severity score, X i , Y i represents the center coordinates of the i-th defect, X c , Y c represents the centroid coordinates of all defects, FG i represents the average gray gradient of the i-th defect area, A i represents the area of the i-th defect, G med represents the median gray gradient of all defect areas, M represents the total number of detected defects;

[0201] According to the defect severity score, organize the spatial position, type, and severity of each defect to generate defect results.

[0202] Specifically, obtain the classified defect results, extract the position information of each identified defect, analyze the boundary shape and size in the image coordinates to form the spatial distribution data of the defects. Read the previously recorded defect recognition coordinate sequence and parse it line by line. Each record contains the starting and ending horizontal and vertical coordinates of the defect and possible arcs or fracture trends in the middle. It is necessary to first confirm whether these coordinates are within the valid area by comparing with the previously saved image resolution range. If the coordinate range exceeds the image width or height, this part of the record outside the image valid area will be excluded by referring to the original resolution upper limit set when the on-site camera was shooting. Subsequently, the selected coordinate information is associated with the corresponding defect label in a one-to-one manner. By row and column scanning, the outermost boundary of the defect on the image can be found and the number of pixels it occupies can be calculated, so as to obtain the perimeter and approximate shape of the boundary. If the number of external edge points of a certain defect exceeds the predetermined standard of 800 pixels, it will be marked as a large-range defect. This standard of 800 comes from the area ratio statistics when actually measuring 50 defect-containing images at the industrial site. By calculating the distances between all external edge points, the local width and height data of the defect can be obtained. If the aspect ratio of the defect is greater than 2 and the horizontal span exceeds 150 pixels, it can be identified as an extended defect. Record the measurement results of the boundary shapes and sizes of these defects item by item, and then match them with the previously obtained pixel coordinates after summarization to generate a complete spatial distribution data.

[0203] The advantage of the formula is that by multiplying the Euclidean distance between the defect center in the numerator part and the center of gravity by the average gray level gradient, the position of the defect in the image plane and the local gray level gradient can be comprehensively considered, and the defect area and the value of its gray level gradient deviating from the median are simultaneously introduced in the denominator part, so as to reflect the combined influence of the position dispersion and the gray level complexity in one score.

[0204] X i ,Y i The acquisition steps of are as follows:

[0205] They represent the center coordinates of the i-th defect. After completing the defect boundary tracing and determining the defect connected region in the front, the center position needs to be calculated by means of the mean value or the geometric center. At the scene, the center point can be obtained by averaging the coordinates of all pixels in the defect region, or the center combination of the outermost point and the innermost point can also be used to determine the coordinates. If the pixel set of a certain defect occupies several rows and columns in the image, then the row and column coordinate values can be accumulated pixel by pixel and then divided by the number of pixels to obtain (X i ,Y i ). In the example, for a certain defect with an area of 200, if the accumulated row and column coordinate values are (24000, 18000), then X i = 24000 / 200 = 120, Y i = 18000 / 200 = 90.

[0206] X c ,Y c The acquisition steps of are as follows:

[0207] They represent the center of gravity coordinates of all defects, which need to be calculated after summarizing the center coordinates of all defects. If a total of M defects are detected at the scene, and the center of each defect is (X i ,Y i ), then the average or weighted average (if the defect areas are different) can be taken for all X i to obtain X c , and similarly Y c can be calculated. To ensure the accuracy of the data, the technical personnel need to first confirm that these defect centers are in the same image coordinate system and exclude the errors generated by the edge part. If the centers of three defects are (120, 90), (205, 95), and (180, 150) respectively, then the global center of gravity (168.3, 111.7) can be obtained by dividing the sum of these coordinates by 3.

[0208] FG i The acquisition steps of are as follows:

[0209] This parameter represents the average gray-scale gradient of the i-th defect area. During on-site measurement, it is necessary to extract the gray-scale gradient of each pixel from the set of defect pixels that have been separated previously, and then accumulate or perform weighted statistics on this gradient to obtain the average value of the entire defect area. In industrial testing, the local gradient is obtained by taking the difference in gray-scale around each pixel. For example, at the position (x, y), the differences in the horizontal and vertical directions are taken as the gradient magnitude, and then the pixel gradients of all pixels in this defect area are accumulated and divided by the number of pixels to obtain FG. i In one example, the defect area contains 300 pixels. The technician calculates the local gradient for each pixel and obtains a total of 1800. Then FG i = 1800 / 300 = 6.

[0210] A i The steps for obtaining it are as follows:

[0211] This parameter corresponds to the area of the i-th defect, usually measured by the number of pixels covered by this defect area in the image coordinates. During on-site acquisition, if a certain defect is composed of 200 consecutive pixel points, then A i = 200. If a higher-resolution image is used, the same defect may show a larger or smaller value at the pixel level, and it needs to be matched with the previous calibration data based on physical size in order to compare the areas of other defects. When technicians perform area statistics, they will first mark the defect pixels through connected component search or boundary tracing, and then sum up the counts of these marked pixels to obtain A i . In the same image, if there are multiple defects, they are distinguished and the area values are recorded separately according to their different connected regions.

[0212] G med The steps for obtaining it are as follows:

[0213] This parameter represents the median gray-scale gradient of all defect areas. It is necessary to first summarize the FG of each defect i to form a list, and then sort the FG in the list i from smallest to largest. When there are an odd number of defects, the middle value is taken as the median. If the number of defects is even, the average of the two middle values is taken as the median. In the industrial field, it is possible to record the average gradient FG of all defects after multiple detections. When there are up to dozens of defects, the value at the middle position or the two middle positions is taken to calculate G i . If all the defect FGs med in the test are 3, 4, 6, 7, 9, and the sorted result is 3, 4, 6, 7, 9, the middle 6 is G i . med .

[0214] The steps for obtaining M are as follows:

[0215] This parameter represents the total number of detected defects, that is, the number of all true defects confirmed in the current image detection. Usually, the final defect count is obtained after morphological processing or screening of pseudo-defects on-site. If 5 connected regions are found in an image and all are confirmed to be actual defects, then M = 5.

[0216] Calculation process:

[0217] In one detection, the device identifies M = 3 defects, with the center coordinates being (X 1 , Y 1 ) = (120, 90), (X 2 , Y 2 ) = (205, 95), (X 3 , Y 3 ) = (180, 150), and the areas are A 1 = 200, A 2 = 220, A 3 = 300, and the average gray gradient FG 1 = 5, FG 2 = 8, FG 3 = 10. First, calculate the centroid coordinates (X c , Y c ) of all defects, and the simple average can be taken:

[0218]

[0219] Subtract the centroid of each defect from the center and calculate the Euclidean distance, then multiply by its FG i and accumulate:

[0220] Defect 1:

[0221]

[0222] Defect 2:

[0223]

[0224] Defect 3:

[0225]

[0226] The sum of the numerator part = about 264.6 + 322.48 + 400.9 = 988.0;

[0227] Then look at the denominator. First, calculate the FG i list of all defects, which is 5, 8, 10 and sorted as 5, 8, 10. Therefore, the median G med = 8. Then add |FG i - G med|:

[0228] Defect 1: 200 + |5 - 8| = 200 + 3 = 203;

[0229] Defect 2: 220 + |8 - 8| = 220 + 0 = 220;

[0230] Defect 3: 300 + |10 - 8| = 300 + 2 = 302;

[0231] Sum of denominator part = 203 + 220 + 302 = 725;

[0232] Final score:

[0233]

[0234] This result indicates that when the score D is approximately 1.36, it shows that the defects in the current image are at a relatively low comprehensive level in terms of position dispersion and gray gradient characteristics. If the D value exceeds 2, it means that there are defects with a more significant distribution range or a larger gradient. Users can set a reference threshold according to the actual detection task requirements and judge the size relationship between D and this threshold, so as to further distinguish the severity of the defects.

[0235] According to the aforementioned defect severity score D, read the score in the range of 0 to 3 and combine the spatial distribution data of each defect. When a defect with a score higher than 2 is detected, it is classified as a defect that requires priority attention, and its center point and boundary vertices recorded in the image coordinates are checked in turn. If it is found during the inspection that the defect area value exceeds 400 pixels and there is no obvious overlap with other surrounding defect areas, it is determined as a high-severity defect with a separate distribution, and its detailed area, gray gradient, and boundary characteristics are added to the defect entry list. If a defect with a score lower than 1 and an area less than 150 pixels is detected, it is regarded as a low-severity defect and recorded as a secondary entry. For the remaining defects with scores between 1 and 2, they are located at an intermediate level, and continue to check the concavity and convexity of their boundaries. If the irregularity degree of the boundary exceeds the previously measured threshold of 0.3, it is marked as a moderately complex-shaped defect, and the geometric characteristics of its outermost edge pixels are supplemented in the record for reference in subsequent maintenance. Finally, based on the spatial positions and score intervals of all defects, a summary is made to give the type, center coordinates, boundary shape, and severity range of each defect, and a defect result is generated.

Claims

1. An automatic bearing surface defect detection system based on image recognition, characterized in that: The system comprises: The image acquisition and preprocessing module acquires the bearing surface image, crops and filters the image to remove noise and adjust the contrast to obtain a processed image; performs grayscale conversion on the processed image to generate a preprocessing result; A feature extraction and enhancement module, based on the preprocessing result, decomposes the image through an image pyramid, extracts local contrast and texture features at each scale, generates a feature set, and adjusts the enhancement parameters according to the statistical distribution of the feature set to obtain an enhanced feature map; A defect classification and labeling module, based on the enhanced feature map, uses a U-Net network to analyze the image, identifies cracks and spalling defects on the bearing surface, and obtains a classification defect result; The result output and report module sorts out the location, type and severity of each detected defect based on the classified defect results to obtain the defect results.

2. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps of obtaining the processed image are: Acquire the bearing surface image, crop and remove the ineffective area at the edge of the image, optimize the signal-to-noise ratio of the image by using spatial domain filtering, and obtain the cropped and filtered image; According to the cropped filtered image, the contrast enhancement index is calculated, and the calculation formula is: Among them, C represents the contrast enhancement index, I i,j represents the pixel grayscale value of the i-th row and j-th column in the cropped filtered image, M represents the global average grayscale value of the cropped filtered image, and I max Represents the maximum grayscale value in the cropped filtered image, I min represents the minimum grayscale value in the cropped filtered image, m and n represent the number of rows and columns of the image respectively; According to the contrast enhancement index, the grayscale histogram distribution of the cropped and filtered image is adjusted to obtain a processed image.

3. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps for obtaining the preprocessing results are: Acquire the processed image, separate the pixel values ​​of the red channel, the green channel and the blue channel, and convert them into a standardized color space to obtain an image after channel separation; According to the image after the channel separation, the gray value distribution balance is calculated, and the calculation formula is: Among them, AG represents the gray value distribution balance, R represents the pixel value of the red channel in the image after channel separation, G represents the pixel value of the green channel in the image after channel separation, B represents the pixel value of the blue channel in the image after channel separation, and HM represents the global average gray value of the image after channel separation; According to the gray value distribution balance, gray mapping is performed on the image after channel separation to generate a preprocessing result.

4. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps of obtaining the feature set are: Obtain the preprocessing result, decompose the image using an image pyramid, construct image hierarchies of different scales, and obtain a multi-scale decomposed image; According to the multi-scale decomposed image, the local contrast score is calculated using the following formula: Where CF represents the local contrast score, P i,j represents the gray value of the jth pixel at the i-th scale of the image after multi-scale decomposition, N represents the number of local pixel blocks at the current scale, and P avg Represents the average gray value of all pixels at the current scale, P max Represents the maximum gray value of the pixel at the current scale, P min represents the minimum grayscale value of the pixel at the current scale, k represents the number of scales obtained after the image is decomposed, and z represents the total number of pixels at the current scale; According to the local contrast score and in combination with texture distribution features at different scales, the structural changes of the image at each scale are extracted to generate a feature set.

5. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps of obtaining the enhanced feature map are: Based on the feature set, the enhancement parameter is calculated using the following formula: Among them, E represents the enhancement parameter, F n represents the value of the nth feature, σ n represents the standard deviation of the nth feature, μ n represents the average value of the nth feature, and TN is the total number of features in the feature set; According to the enhancement parameters, the intensity of each feature is adjusted to obtain an enhanced feature map.

6. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps for obtaining the classification defect result are: Using a U-Net network to process the enhanced feature map, identify and mark cracks and spalling defects in the image, and obtain an unoptimized defect recognition result; The unoptimized defect recognition result is subjected to image post-processing, including using threshold segmentation to separate foreground and background, applying morphological dilation and erosion to clarify defect boundaries, and generating a classified defect result.

7. The automatic bearing surface defect detection system based on image recognition according to claim 1 is characterized in that: The steps for obtaining the defect result are: Obtaining the classified defect results, extracting the position information of each identified defect, analyzing the boundary shape and size in the image coordinates, and forming spatial distribution data of the defects; Based on the spatial distribution data of the defects, the defect severity score is calculated using the following formula: Where D represents the defect severity score, X i ,Y i represents the center coordinate of the i-th defect, X c ,Y c represents the centroid coordinates of all defects, represents the average grayscale gradient of the i-th defect area, A i represents the area of ​​the i-th defect, G med represents the median grayscale gradient of all defective areas, and M represents the total number of defects detected; According to the defect severity score, the spatial location, type and severity of each defect are sorted out to generate a defect result.

Citation Information

Patent Citations

  • circle radius measuring method based on contour merging and convex hull fitting

    CN109658391A

  • Self-supervised water surface image enhancement method and related equipment

    CN116579953A

  • Railway accessory defect detection method and system

    CN118037726A

  • System for improved image enhancement

    US20130257887A1

Cited By

  • Hole opening positioning method for automobile part machining

    CN120318330A

  • Weld defect intelligent identification and distribution visualization method and system

    CN120912617A

  • Water turbine coating abrasion detection method and system

    CN121353295A

  • Shearing behavior dynamic correction method and system based on wear state recognition

    CN121437507A

  • A shearing behavior dynamic correction method and system based on wear state identification

    CN121437507B