Complex surface damage area identification and segmentation method
By using the method of collaborative training of PP-OCRv4 and Lama model in complex surface damage detection, combined with morphological opening operations and edge detection technology, the automated and precise detection and evaluation of complex surface damage is achieved, and the problem of limited detection effect and reliance on professional and technical personnel in the existing technology is solved.
Patent Information
- Application Number
- CN202510137208.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art is difficult to effectively detect and evaluate complex surface damage, especially in workpieces with complex shapes and surface conditions. Traditional methods are highly subjective, have limited detection effects, and require professional equipment and technical personnel.
The text area positioning and image repair method based on collaborative training of PP-OCRv4 and Lama model is adopted, combined with morphological opening operations and edge detection technology, and the damage area is cut through the fixed threshold method to achieve automated, precise segmentation and evaluation of complex surface damage.
It realizes automated and precise detection and evaluation of complex surface damage, reduces dependence on professional and technical personnel, improves detection efficiency and accuracy, and can effectively identify and divide complex surface damage areas.
Smart Images

Figure CN120047466A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of surface damage location methods, and particularly relates to a method for identifying and segmenting complex surface damage areas. Background Art
[0002] In the modern industrial field, complex surface damage assessment is a crucial task. The forms of complex surface damage are diverse, including but not limited to wear, corrosion, cracks, deformation, etc. These damages not only affect the appearance of products, but also reduce the mechanical properties of products during long-term use, and may even cause serious safety accidents. Traditional surface damage detection methods mainly include visual inspection, magnetic particle inspection, penetrant inspection, ultrasonic inspection, etc. Visual inspection relies on the experience and skills of inspectors, is highly subjective, and it is difficult to detect small or hidden damages; magnetic particle inspection and penetrant inspection have high requirements for the operation of inspectors and can only detect surface opening defects; although ultrasonic inspection can detect internal defects, its detection effect on workpieces with complex shapes and surface conditions is limited, and the detection process is relatively complex, requiring professional equipment and technicians. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for identifying and segmenting complex surface damage areas, so as to provide an effective way for realizing the automatic and precise segmentation and evaluation of complex surface damage.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions.
[0005] A method for identifying and segmenting complex surface damage areas includes the following steps:
[0006] Step 1. Data collection, specifically referring to obtaining images containing complex surface damages of different degrees, where complex surface damage refers to a surface including at least one of wear, corrosion, cracks, and deformation;
[0007] Step 2. Coordination of text region location and image restoration based on the collaborative training of PP-OCRv4 and Lama models; specifically including B1 to B3:
[0008] B1. Input the preprocessed image into the PP-OCRv4 model to detect the position information of the text region in the image, and transfer to B2;
[0009] B2. According to the text position information provided by PP-OCRv4, determine the text region to be repaired, and use the Lama model to repair the detected text region, and transfer to B3;
[0010] B3. Use the PP-OCRv4 model to check the repaired area to see if there is still residual text. If there is residual text, repeat steps B1 and B2 until the text interference in the image is eliminated;
[0011] Step 3. Image information processing, specifically refers to using morphological opening operation to remove small isolated noise points and connected small parts, and using edge detection technology to capture the damage edge and retain valuable damage information;
[0012] Step 4. Damage area cutting
[0013] Adopt the fixed threshold method, set a specific gray value range to perform binary processing on the image, and determine the damaged area and the normal area.
[0014] For a further improvement or preferred implementation of the foregoing complex surface damage area recognition and segmentation method, step 1 further includes: cropping or padding the picture size to a fixed size, and dividing the entire data set into a training set and a test set according to a ratio of 8:2, where 10% of the training set data is used as the validation set.
[0015] For a further improvement or preferred implementation of the foregoing complex surface damage area recognition and segmentation method, the PP-OCRv4 model is based on a two-stage OCR module, including a text detection module and a text recognition module. The text detection module is established based on the DB algorithm, and the text recognition module is established based on the CRNN network. A text direction classifier is added between the text detection module and the text recognition module to handle text recognition in different directions.
[0016] For a further improvement or preferred implementation of the foregoing complex surface damage area recognition and segmentation method, in step 2, the Lama model directly calls the standard data set of the Lama model and uses the places2 and places challenge models for training. At the same time, by adding a custom data set, real data is obtained through actual shooting, and the model is retrained;
[0017] At the same time, a perceptual loss function is introduced. The perceptual loss function inputs the real image and the generated image into the pre-trained VGG19 network respectively to evaluate the training result of the model; The perceptual loss function can be expressed as:
[0018]
[0019] Where x is the input image, y is the target image, Fi(x) and Fi(y) respectively represent their feature representations at the i-th layer in the VGG19 network, and N is the number of feature layers.
[0020] For a further improvement or preferred implementation of the foregoing method for identifying and segmenting complex surface damage regions, step 2 further includes: using the accuracy rate P, recall rate R, and average precision AP as evaluation indicators to evaluate the performance of the PP_OCRv4 recognition model for a single-class target, and adjusting and optimizing the parameters of the PP_OCRv4 recognition model according to the evaluation results, where
[0021]
[0022] TP represents the number of correct target detections, FP represents the number of incorrect target detections, and FN represents the number of targets not detected; the area enclosed by the accuracy rate and recall rate curves is the average precision AP of the class target detection;
[0023] For multi-class targets, the mean average precision mAP is used for evaluation, where
[0024] N represents the number of types of targets in the dataset.
[0025] For a further improvement or preferred implementation of the foregoing method for identifying and segmenting complex surface damage regions, step 2 further includes continuously monitoring the changes in the loss function values and evaluation indicators of the model on the training set and validation set, plotting the numerical and indicator curves, determining whether the model has overfitting or underfitting conditions, and adjusting the training parameters or using regularization methods for optimization according to the results;
[0026] And using the mean squared error MSE and frames per second FPS to evaluate the performance of the Lama model, where
[0027] n represents the total number of pixel points in the image, I original (x i ,y i ) represents the pixel value of the original image at the coordinate (x i ,y i ), I repaired (x i ,y i ) represents the pixel value of the restored image at the coordinate (x i ,y i ); N′ represents the number of processed image frames, and T represents the total running time.
[0028] For a further improvement or preferred implementation of the foregoing method for identifying and segmenting complex surface damage areas, step 4 further includes: in order to make the calculated damage area closer to the result of machine scanning, error analysis is performed on the damage areas of the experimental results and the calculated data results, and the gray value threshold is continuously adjusted to determine the preferred threshold range that makes the calculated damage area most similar to the actual damage area.
[0029] For a further improvement or preferred implementation of the foregoing method for identifying and segmenting complex surface damage areas, in step 4, the gray value threshold of the fixed threshold method is set to 70-80.
[0030] For a further improvement or preferred implementation of the foregoing method for identifying and segmenting complex surface damage areas, in step 4, the gray value threshold of the fixed threshold method is set to 75. Description of the Drawings
[0031] Figure 1 is the system block diagram of the PP-OCRv4 model;
[0032] Figure 2 is the change curve of the loss function during the training process;
[0033] Figure 3 is the recognition result of the test set images;
[0034] Figure 4 is the change curve of the loss function of the Lama model during the training process;
[0035] Figure 5 is the processing result of the test set images;
[0036] Figure 6 is the binary output comparison chart of different methods. Detailed Implementation Modes
[0037] The present invention will be described in detail below in conjunction with specific embodiments.
[0038] Complex surface detection plays a key role in modern industrial inspection, and its damage condition is directly related to the performance and safety of industrial metals. This application provides an innovative and accurate deep learning-based method for evaluating the damage of complex surfaces, and realizes the full-process automatic evaluation from image acquisition to damage ratio calculation and result recording by comprehensively using advanced deep learning models, aiming to provide a scientific, efficient and accurate evaluation means for the maintenance and guarantee of complex surfaces.
[0039] The main step process of the complex surface damage area identification and segmentation method of this application specifically includes:
[0040] Step 1. Data collection:
[0041] Obtain images containing complex surface damages of different degrees, where the complex surface damages refer to surfaces including at least one of wear, corrosion, cracks, and deformation;
[0042] In actual implementation, images containing complex surface damages should be obtained as much as possible to ensure that the actual industrial scene images cover various damage types (such as wear, corrosion, cracks, etc.), damages of different degrees, and various text markings, so as to obtain more accurate recognition and processing results that are more in line with the actual situation of the industrial scene.
[0043] In specific implementation, the picture size should also be cropped or filled to a fixed size, and the entire dataset should be divided into a training set and a test set according to a ratio of 8:2, where 10% of the training set data is used as the validation set.
[0044] Step 2. Coordinate text region localization and image restoration based on the collaborative training of PP-OCRv4 and Lama models; specifically including B1 to B3:
[0045] B1. Input the preprocessed image into the PP-OCRv4 model to detect the position information of the text regions in the image, and transfer to B2;
[0046] B2. Based on the text position information provided by PP-OCRv4, determine the text regions that need to be restored, and use the Lama model to restore the detected text regions, and transfer to B3;
[0047] B3. Use the PP-OCRv4 model to check the restored regions to see if there are still residual texts. If there are residual texts, repeat steps B1 and B2 until the desired text removal effect is achieved to ensure that the text interference in the image is completely eliminated;
[0048] The PP-OCRv4 model in the present invention is a two-stage OCR model, including a text detection module and a text recognition module. The text detection module is established based on the DB algorithm, and the text recognition module is established based on the CRNN network. At the same time, a text direction classifier is added between the text detection module and the text recognition module to handle text recognition in different directions. The system block diagram of the PP-OCRv4 model is as Figure 1 shown. Based on the foregoing model, higher accuracy and robustness can be achieved when processing texts and characters.
[0049] As shown in Table 1, in the test results of the dataset in this application, the character recognition accuracy of the PP_OCRv4 recognition model of the present invention reached 75.45%, meeting the application requirements.
[0050] Table 1 Comparison of the results of the PP-OCRv4 model of this application with existing models
[0051] Model Model Size Hmean (%) Cpu + mk1dnn Speed (ms) PP_OCRv3 3.4 76.22 69 PP_OCRv4 4.8 79.87 67
[0052] Based on the processing method of the present application, it can effectively avoid the misrepair of non-text areas, ensure the integrity and accuracy of the restored image, and avoid the information mismatch and error accumulation problems when the traditional recognition model works.
[0053] To ensure the effectiveness of the Lama model, by directly calling the official dataset of the Lama model and using the places2 and placeschallenge models for training, and at the same time, in the present application, the method of adding a custom dataset is adopted, and real data is obtained through actual shooting to perform secondary training on the model;
[0054] To ensure that the Lama model can effectively learn and generate accurate and natural repair results, and improve the structural similarity between the real image and the repaired image, a perceptual loss function is introduced. The perceptual loss function inputs the real image and the generated image into the pre-trained VGG19 network respectively, and calculates the Euclidean distance of the activation values of each intermediate layer. The perceptual loss function can be expressed as:
[0055] where x is the input image, y is the target image, Fi(x) and Fi(y) respectively represent their feature representations at the i-th layer in the VGG19 network, and N is the number of feature layers;
[0056] In particular, in the present application, the detection accuracy is used as the evaluation index with the accuracy P, recall rate R, and average precision AP to evaluate the performance of the PP_OCRv4 recognition model for a single-class target, and the parameters of the PP_OCRv4 recognition model are adjusted and optimized according to the evaluation results, where
[0057]
[0058] TP represents the number of correct target detections, FP represents the number of incorrect target detections, and FN represents the number of targets not detected; the area enclosed by the accuracy and recall rate curves is the detection average precision AP of the class target;
[0059] For multi-class targets, the mean average precision mAP is used for evaluation, where
[0060] N represents the number of types of targets in the dataset.
[0061] During the model training process, by monitoring the loss function values and the changes of evaluation indicators of the model on the training set and the validation set, drawing the corresponding numerical and index curves, judging whether the model has overfitting or underfitting situations, and adjusting the training parameters according to the results or adopting regularization methods for optimization processing;
[0062] In this application, the mean squared error (MSE) and the number of frames processed per second (FPS) are used to evaluate the performance of the Lama model, where:
[0063]
[0064] n represents the total number of pixel points in the image, and I original (x i , y i ) represents the pixel value of the original image at the coordinate (x i , y i ), and I repaired (x i , y i ) represents the pixel value of the restored image at the coordinate (x i , y i ); N′ represents the number of frames of the processed image, and T represents the total running time.
[0065] The software and hardware facilities selected for this experiment are shown in Table 2:
[0066] Table 2 Software and Hardware Facilities for the Experiment
[0067]
[0068] All input image sizes are cropped or padded to 1002×852. The entire dataset is divided into a training set and a test set in a ratio of 8:2. Among them, 10% of the training set data is used as the validation set. After 100 rounds of training, the change curve of the loss function during the training process is as Figure 2 shown. In the figure, the abscissa represents the number of training rounds, and the ordinate represents the training loss value.
[0069] It can be seen from the figure that the loss value is the smallest in the 81st round of training. Therefore, the training result of the 81st round is used as the final training weight. After the training is completed, the test set detection results are shown in Table 3.
[0070] Table 3 Detection Results of PP-OCRv4 Based on Iterative Optimization
[0071]
[0072] It can be seen from Table 3 that the average precision (AP) of PP-OCRv4 for the three types of targets are 95.83%, 96.31%, and 92.28% respectively, and the mean average precision (mAP) reaches 94.81%. Among them, the detection precision of special symbols is the lowest, mainly because the number of pictures of special symbols in the dataset is relatively small.
[0073] Figure 3For the detection results of the test set images, the left figure is the original image, and the white area in the right figure is the area of the recognized text. It can be seen that the model can well complete the detection of text targets.
[0074] After 100 rounds of training, the change curve of the loss function of the Lama model during the training process is as Figure 4 shown. In the figure, the abscissa represents the number of training rounds, and the ordinate represents the training loss value. It can be seen from the figure that the loss value of the model rapidly drops to 0.06 in a short time, meeting the requirement of rapid convergence. In addition, the training loss gradually decreases and finally reaches a stable state, and good results can also be obtained on the validation set and the test set. The stability of the model meets the requirements. To improve the running fluency, a multi-threaded method is used for programming in this application.
[0075] Table 4 shows the experimental results on the custom dataset. The improved algorithm is superior to the original algorithm in all evaluation indicators. The MSE is reduced by 0.32%, and the FPS is increased by 24.
[0076] Table 4 Detection Results of Lama Based on Iterative Optimization
[0077] Method MSE (%) FPS Original Lama 2.44 30 Lama of This Application 2.12 54
[0078] Figure 5 For the detection results of the test set images, the left figure is the original image, and the right figure is the final image output after repairing the image after removing the text area. From the above, it can be seen that the model can well complete the redrawing work of the specified block in the complex surface image.
[0079] Step 3. Image Information Processing
[0080] Use morphological opening operation to remove small isolated noise points and connected small parts. The edge detection technology can capture the damage edges and retain valuable damage information;
[0081] By the above method, irrelevant interference information such as scratches and isolated noise points in the image is removed, while the defects that affect the damage calculation are retained, improving the image quality and ensuring the accuracy of the damage area.
[0082] As Figure 6 shown, it is a comparison chart of the binary output of different methods.
[0083] Step 4. Damage Area Cutting
[0084] Adopt the fixed threshold method, set a specific gray value range to perform binary processing on the image, and determine the damage area and the normal area;
[0085] In order to make the calculated damage area closer to the result of machine scanning, through multiple experiments and data analysis, by continuously adjusting the threshold and comparing it with the damage area obtained by machine scanning, it is determined that when the preferred threshold is set to 70-80, the calculated damage area is most similar to the conclusion given by the machine;
[0086] Verified by a large number of experiments, the fixed threshold method based on the aforementioned threshold can accurately handle complex surface damage features, effectively segment the damage area, and avoid problems such as over-segmentation and misjudgment that occur in complex surface images by other methods.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A complex surface damage area recognition and segmentation method, characterized in that: The steps include: Step 1. Data collection, specifically, obtaining images containing complex surface damage of varying degrees, where complex surface damage refers to a surface that has at least one of the following damages: wear, corrosion, cracks, and deformation; Step 2. Collaboration of text area localization and image restoration based on the co-training of PP-OCRv4 and Lama model; specifically including B1 to B3: B1. Input the preprocessed image into the PP-OCRv4 model to detect the text area location information in the image, and then go to B2; B2. Determine the text area that needs to be repaired based on the text position information provided by PP-OCRv4, and use the Lama model to repair the detected text area, and then go to B3; B3. Use the PP-OCRv4 model to check the repaired area to see if there is any residual text. If there is residual text, repeat steps B1 and B2 until the text interference in the image is eliminated. Step 3. Image information processing, specifically refers to the use of morphological opening operations to remove small isolated noise points and small connected parts, and the use of edge detection technology to capture the damage edge and retain valuable damage information; Step 4. Cutting the damaged area The fixed threshold method is used to set a specific gray value range to perform binary processing on the image and determine the damaged area and normal area.
2. The complex surface damage area identification and segmentation method according to claim 1, characterized in that: The step 1 also includes: cropping or padding the image size to a fixed size, and dividing the entire data set into a training set and a test set in a ratio of 8:2, wherein 10% of the training set data is used as a verification set.
3. The complex surface damage area identification and segmentation method according to claim 1, characterized in that: The PP-OCRv4 model is based on a two-stage OCR module, including a text detection module and a text recognition module. The text detection module is established based on the DB algorithm, and the text recognition module is established based on the CRNN network. A text direction classifier is added between the text detection module and the text recognition module to cope with text recognition in different directions.
4. The complex surface damage area identification and segmentation method according to claim 1, characterized in that: In step 2, the Lama model directly calls the Lama model standard data set and uses the places2, places challenge model for training. At the same time, a custom data set is added to obtain real data through actual shooting, and the model is trained twice. At the same time, the perceptual loss function is introduced. The perceptual loss function inputs the real image and the generated image into the pre-trained VGG19 network to evaluate the model training results; the perceptual loss function can be expressed as: Among them, x is the input image, y is the target image, Fi(x) and Fi(y) represent their feature representations at the i-th layer in the VGG19 network, and N is the number of feature layers.
5. The complex surface damage area identification and segmentation method according to claim 1, characterized in that: The step 2 also includes: using the accuracy P, the recall R and the average precision AP as evaluation indicators to evaluate the performance of the PP_OCRv4 recognition model for a single class of targets, and adjusting and optimizing the parameters of the PP_OCRv4 recognition model according to the evaluation results, wherein TP represents the number of correct target detections, FP represents the number of incorrect target detections, and FN represents the number of targets that have not been detected. The area enclosed by the accuracy and recall curves is the average precision AP of the target class. For multiple categories of targets, the average precision mAP is used for evaluation, where N represents the number of target types in the dataset.
6. The complex surface damage area identification and segmentation method according to claim 5, characterized in that: The step 2 also includes continuously monitoring the loss function value of the model on the training set and the validation set and the changes in the evaluation index, drawing the numerical value and index curve, judging whether the model is overfitting or underfitting, and adjusting the training parameters or optimizing the model by using the regularization method according to the results; The mean square error (MSE) and frames per second (FPS) are used to evaluate the performance of the Lama model. n represents the total number of pixels in the image, I original (x i ,y i ) indicates that the original image is at coordinate (x i ,y i ), I repaired (x i ,y i ) indicates that the restored image is at coordinate (x i ,y i ) at the pixel value; N′ represents the number of image frames processed, and T represents the total running time.
7. The complex surface damage area identification and segmentation method according to claim 1, characterized in that: The step 4 also includes, in order to make the calculated damage area closer to the result of machine scanning, performing error analysis by comparing the damage area of the experimental results and the calculated data results, continuously adjusting the grayscale value threshold, and determining the preferred threshold range that makes the calculated damage area most similar to the actual damage area.
8. The complex surface damage area identification and segmentation method according to claim 7, characterized in that: In step 4, the gray value threshold of the fixed threshold method is set to 70-80.
9. The complex surface damage area identification and segmentation method according to claim 7, characterized in that: In step 4, the gray value threshold of the fixed threshold method is set to 75.
Citation Information
Patent Citations
Optical element surface damage identification method based on deep learning and image processing
CN114120317A
Image restoration method, model and device based on convolution and converter hybrid network
CN116309155A
Visual inspection and evaluation method for scratch damage on surface of ceramic material
CN117218086A
Image processing method and apparatus, device, and storage medium
US20240320807A1
Cited By
Method for identifying iridovirus damage grade of micropterus salmoides based on machine vision assistance
CN120761383A