Defect Identification Method, Device, Computer Equipment and Readable Storage Medium
By acquiring and processing RGB images and depth images of the target area, fusing and enhancing the spatial weight in the image, the problems of low efficiency and safety hazards of traditional detection methods are solved, and more efficient and accurate defect recognition is achieved.
Patent Information
- Application Number
- CN202411168187.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-23
AI Technical Summary
Traditional bird's nest defect detection methods rely on manual inspection and helicopter inspection, which are costly and have problems such as safety hazards and low detection efficiency.
By acquiring the RGB image and depth image of the target area, extracting and fusing the spatial weights in the image, image enhancement is performed, and finally identifying whether there are target defects.
It improves the accuracy and efficiency of defect identification, reduces the safety risks and costs of manual inspections, and enhances the target defects in the image more significant.
Smart Images

Figure CN118887200B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and particularly to a defect recognition method, apparatus, computer device, and readable storage medium. Background Art
[0002] The bird's nest defect detection technology is of great significance in the power transmission scenario. In the maintenance and management of high-voltage transmission lines, foreign objects such as bird's nests may pose a threat to the safe operation of the transmission lines. Therefore, it is crucial to detect and remove foreign objects such as bird's nests on the transmission lines in a timely and accurate manner to ensure the stable operation of the power system.
[0003] In the power transmission system, high-voltage transmission lines are usually distributed in a vast outdoor environment with long lines and complex terrains. Birds often build nests on transmission towers and utility poles. These bird's nests may not only cause power failures but also lead to serious consequences such as fires.
[0004] Traditional detection methods mainly rely on manual inspections and helicopter patrols, which are not only costly but also have problems such as potential safety hazards to personnel and low detection efficiency. Therefore, there is an urgent need for improvement. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a defect recognition method, apparatus, computer device, and readable storage medium that can improve the accuracy of defect recognition.
[0006] In a first aspect, the present application provides a defect recognition method, which includes:
[0007] Obtain an RGB image and a depth image collected for a target area;
[0008] Extract a first spatial weight from the RGB image and a second spatial weight from the depth image;
[0009] Perform a superimposition process on the first spatial weight and the second spatial weight to obtain a fused weight;
[0010] Enhance the RGB image according to the fused weight to obtain an enhanced image;
[0011] Identify whether there is a target defect in the enhanced image and output an identification result.
[0012] In one embodiment, extracting a first spatial weight from the RGB image and a second spatial weight from the depth image includes:
[0013] Perform Gaussian denoising on the RGB image to obtain a denoised RGB image;
[0014] Perform a filling process on the depth image to obtain a filled depth image;
[0015] Extract the first spatial weight from the denoised RGB image, and extract the second spatial weight from the filled depth image.
[0016] In one embodiment, the depth image is filled to obtain a filled depth image, including:
[0017] Using the neighborhood average method or the interpolation method, the depth image is filled to obtain a filled depth image.
[0018] In one embodiment, extracting the first spatial weight from the denoised RGB image and extracting the second spatial weight from the filled depth image includes:
[0019] Perform registration processing on the denoised RGB image and the filled depth image to obtain a registered RGB image and a registered depth image;
[0020] Extract the first spatial weight from the registered RGB image and extract the second spatial weight from the registered depth image.
[0021] In one embodiment, performing registration processing on the denoised RGB image and the filled depth image to obtain a registered RGB image and a registered depth image includes:
[0022] Calibrate the parameters of the RGB sensor corresponding to the RGB image and the depth sensor corresponding to the depth image;
[0023] After parameter calibration, perform registration processing on the denoised RGB image and the filled depth image from the feature point dimension to obtain a registered RGB image and a registered depth image.
[0024] In one embodiment, according to the fusion weight, perform enhancement processing on the RGB image to obtain an enhanced image, including:
[0025] Multiply the fusion weight and the RGB image element by element to obtain an enhanced image.
[0026] In a second aspect, the present application also provides a defect recognition device, which includes:
[0027] An acquisition module, configured to acquire an RGB image and a depth image collected for a target area;
[0028] An extraction module, configured to extract the first spatial weight from the RGB image and extract the second spatial weight from the depth image;
[0029] A weight fusion module, configured to perform superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0030] An enhancement processing module, configured to perform enhancement processing on the RGB image according to the fusion weight to obtain an enhanced image;
[0031] A defect recognition module, configured to recognize whether there is a target defect in the enhanced image and output a recognition result.
[0032] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0033] Obtain the RGB image and the depth image collected for the target area;
[0034] Extract the first spatial weight in the RGB image and extract the second spatial weight in the depth image;
[0035] Perform superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0036] Perform enhancement processing on the RGB image according to the fusion weight to obtain an enhanced image;
[0037] Recognize whether there is a target defect in the enhanced image and output a recognition result.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0039] Obtain the RGB image and the depth image collected for the target area;
[0040] Extract the first spatial weight in the RGB image and extract the second spatial weight in the depth image;
[0041] Perform superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0042] Perform enhancement processing on the RGB image according to the fusion weight to obtain an enhanced image;
[0043] Recognize whether there is a target defect in the enhanced image and output a recognition result.
[0044] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0045] Obtain the RGB image and the depth image collected for the target area;
[0046] Extract the first spatial weight in the RGB image and extract the second spatial weight in the depth image;
[0047] Perform superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0048] Perform enhancement processing on the RGB image according to the fusion weight to obtain an enhanced image;
[0049] Identify whether there are target defects in the enhanced image and output the identification result.
[0050] In the above defect identification method, device, computer device and readable storage medium, the RGB image provides rich color information, which helps to identify the appearance features of the target. The depth image provides spatial structure information and can reflect the shape and depth changes of the target. By fusing the spatial weights of the RGB image and the depth image, the color and spatial structure information can be comprehensively utilized to improve the effect of image enhancement. In addition, the fusion weight can reflect the spatial distribution of important features in the RGB image and the depth image. Performing enhancement processing on the RGB image according to the fusion weight can highlight important features and suppress irrelevant or noise information. The enhanced image is more conducive to subsequent defect identification. The target defects in the enhanced image are more prominent, which helps to improve the accuracy of defect identification. Fusing depth information can make up for the difficulty of defect identification caused by factors such as illumination and shadow in the RGB image. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a schematic flowchart of a defect identification method in an embodiment;
[0053] Figure 2 It is a schematic flowchart of the steps of extracting the first spatial weight in the RGB image and extracting the second spatial weight in the depth image in an embodiment;
[0054] Figure 3 It is a schematic flowchart of the steps of performing registration processing on the denoised RGB image and the filled depth image in an embodiment;
[0055] Figure 4 It is a schematic system diagram of a defect identification system in an embodiment;
[0056] Figure 5 It is a schematic model diagram of a multi-modal detection model in an embodiment;
[0057] Figure 6 Schematic diagram of a defect recognition device in an embodiment;
[0058] Figure 7 Internal structure diagram of a computer device in an embodiment. Specific implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] In an exemplary embodiment, as Figure 1 shown, a defect recognition method is provided, including the following S101 to S105, where:
[0061] S101, Obtain an RGB image and a depth image collected for a target area.
[0062] First, it is necessary to clarify the target area to be detected, which can be an object, a scene or any area where defects need to be identified. After determining the target area, it needs to be placed within the field of view of a color camera and a depth sensor.
[0063] Optionally, use a color camera (such as a CMOS or CCD camera) to take a picture of the target area to obtain its RGB image. The RGB image will contain the color and texture information of the target area, which is an important basis for subsequent defect recognition. At the same time, use a depth sensor (such as a structured light camera, a time-of-flight camera, etc.) to scan the target area to obtain its depth image. The depth image will provide the three-dimensional shape information of the target area, including concavity, convexity, distance, etc., which is crucial for understanding the three-dimensional structure of the target area.
[0064] S102, Extract the first spatial weight in the RGB image and extract the second spatial weight in the depth image.
[0065] Specifically, preprocess the RGB image, such as denoising, grayscale conversion, etc., to improve the accuracy of subsequent processing. Then, use image processing techniques, such as edge detection, texture analysis, etc., to extract the spatial features in the RGB image. According to the importance or significance of the spatial features, assign a weight to each pixel in the RGB image to form a first spatial weight map (i.e., the first spatial weight map). This weight map reflects the importance of different regions in the RGB image on the two-dimensional plane.
[0066] Furthermore, preprocess the depth image, such as denoising, smoothing, etc., to improve its quality. Utilize the depth information to analyze the three-dimensional shape of the object surface, including concavity, convexity, distance, etc. According to the importance or significance of the three-dimensional shape, assign a weight to each pixel in the depth image to form a second spatial weight map (i.e., the second spatial weight). This weight map reflects the significance of different regions in the depth image in three-dimensional space.
[0067] S103. Perform a superposition process on the first spatial weight and the second spatial weight to obtain a fused weight.
[0068] It can be understood that after extracting the first spatial weight and the second spatial weight, they can be fused to form a more comprehensive representation of the spatial weight. The fusion can be achieved through weighted summation, maximum selection, or other fusion strategies.
[0069] Specifically, determine the superposition strategy, such as weighted summation, maximum selection, average calculation, etc. The choice of strategy depends on the application scenario, data characteristics, and expected recognition effect. For example, weighted summation is a commonly used strategy that allows different weight coefficients to be assigned to the first spatial weight and the second spatial weight to reflect their importance in the fusion process.
[0070] According to the selected superposition strategy, perform a pixel-by-pixel superposition process on the first spatial weight and the second spatial weight. Under the weighted summation strategy, multiply the first spatial weight of each pixel by the corresponding weight coefficient, then multiply the second spatial weight of each pixel by the corresponding weight coefficient, and finally add the two products to obtain the fused weight of the pixel. Optimize the fused weight, such as removing noise, smoothing, normalizing, etc., to improve its quality and usability. Normalization is an important optimization step that ensures the value of the fused weight is within an appropriate range, facilitating subsequent image enhancement and defect recognition processing. Verify the fused weight to ensure that it accurately reflects the spatial information in the RGB image and the depth image. According to the verification result, adjust the superposition strategy or weight coefficient to optimize the effect of the fused weight.
[0071] S104. Enhance the RGB image according to the fused weight to obtain an enhanced image.
[0072] Specifically, the fusion weights are applied to each pixel of the RGB image, which generally involves multiplying the color value of each pixel by the corresponding fusion weight to emphasize or suppress the importance of that pixel. The color, brightness, contrast, and other parameters of the RGB image are adjusted according to the fusion weights to enhance the key information in the image. Specifically, it includes increasing the brightness of certain regions, increasing the contrast to highlight the edges, or adjusting the color saturation to improve the visual effect. During the enhancement process, it is necessary to ensure that the enhanced image still maintains consistency with the target area. This means that the enhancement process should not introduce any information or artifacts unrelated to the target area. The enhanced image is optimized, such as removing noise, smoothing, adjusting the color balance, etc., to further improve the image quality.
[0073] S105, identify whether there is a target defect in the enhanced image and output the identification result.
[0074] Specifically, according to the characteristics of the target defect and the identification requirements, a corresponding defect identification algorithm is selected. The defect identification algorithm can be a method based on image processing, a machine learning algorithm, or a deep learning model. The selected defect identification algorithm is applied to the enhanced image to detect whether there is a target defect in the image. For example, the target defect can be a bird's nest defect on a cable line.
[0075] It can be understood that the defect identification algorithm can analyze features such as pixels, textures, and shapes in the enhanced image and match or compare them with known defect patterns. According to the output of the defect identification algorithm, it is determined whether there is a target defect in the enhanced image. If a defect is identified, the algorithm will provide relevant information such as the location, type, and size of the defect. The identification result is output in an appropriate form, such as displayed on the screen, saved to a file, or sent to other systems or users. The output result should clearly and accurately reflect the identified defect information for subsequent analysis, decision-making, or reporting.
[0076] In the above-mentioned defect identification method, the RGB image provides rich color information, which helps to identify the appearance features of the target. The depth image provides spatial structure information, which can reflect the shape and depth changes of the target. By fusing the spatial weights of the RGB image and the depth image, the color and spatial structure information can be comprehensively utilized to improve the effect of image enhancement. In addition, the fusion weights can reflect the spatial distribution of important features in the RGB image and the depth image. Enhancing the RGB image according to the fusion weights can highlight important features and suppress irrelevant or noise information. The enhanced image is more conducive to subsequent defect identification. The target defects in the enhanced image are more prominent, which helps to improve the accuracy of defect identification. Fusing depth information can make up for the difficulties in defect identification caused by factors such as lighting and shadows in the RGB image.
[0077] In an exemplary embodiment, such as Figure 2As shown, extract the first spatial weight from the RGB image and extract the second spatial weight from the depth image, including:
[0078] S201, perform Gaussian denoising on the RGB image to obtain the denoised RGB image.
[0079] It can be understood that the Gaussian filtering algorithm is used to denoise the RGB image. Gaussian filtering is famous for its smoothing effect and noise suppression ability. It can effectively reduce the interference of random noise while retaining the important features of the image. Through Gaussian denoising, the quality of the RGB image is significantly improved, and the noise in the image is effectively suppressed, creating favorable conditions for subsequent feature extraction and fusion.
[0080] S202, perform filling processing on the depth image to obtain the filled depth image.
[0081] Specifically, performing filling processing on the depth image to obtain the filled depth image includes: using the neighborhood average method or the interpolation method to perform filling processing on the depth image to obtain the filled depth image.
[0082] It can be understood that the neighborhood average method: This method fills the missing area by calculating the average value of the valid pixels around the missing pixel. The average value calculation can be based on different neighborhood sizes, such as 3x3, 5x5, etc. The specific choice depends on the size and shape of the missing area. The neighborhood average method is simple and easy to implement and is suitable for cases where the missing area is small and the surrounding pixel distribution is uniform.
[0083] Interpolation method: The interpolation method is a more complex filling method. It can calculate the estimated value of the missing pixel based on the valid pixel values around the missing pixel through a mathematical model. Commonly used interpolation methods include bilinear interpolation, cubic spline interpolation, etc. These methods can accurately estimate the value of the missing pixel according to the position of the missing pixel and the distribution of the surrounding pixels. The interpolation method is suitable for cases where the missing area is large or the surrounding pixel distribution is uneven and can provide a more accurate filling effect.
[0084] The filling process can be as follows: First, it is necessary to identify the missing or invalid areas in the depth image, which may be caused by sensor failures, object occlusions, etc. According to the size, shape and surrounding pixel distribution of the missing area, select the appropriate filling method (neighborhood average method or interpolation method). Apply the selected filling method to fill the missing area to obtain the filled depth image. Finally, it is necessary to verify the filled depth image to ensure that the filling effect meets the expectations and no new noise or artifacts are introduced.
[0085] S203, extract the first spatial weight from the denoised RGB image and extract the second spatial weight from the filled depth image.
[0086] Specifically, extract the first spatial weight from the denoised RGB image and the second spatial weight from the filled depth image, including: registering the denoised RGB image and the filled depth image to obtain the registered RGB image and the registered depth image; extracting the first spatial weight from the registered RGB image and the second spatial weight from the registered depth image.
[0087] Specifically, as Figure 3 shown, registering the denoised RGB image and the filled depth image to obtain the registered RGB image and the registered depth image includes:
[0088] S301, calibrate the parameters of the RGB sensor corresponding to the RGB image and the depth sensor corresponding to the depth image.
[0089] It can be understood that this step aims to accurately calibrate the parameters of the RGB sensor corresponding to the RGB image and the depth sensor corresponding to the depth image. The purpose of parameter calibration is to ensure that when the two sensors capture the same scene, they have consistent internal parameters (such as focal length, optical center, etc.) and external parameters (such as position, attitude, etc.). Through parameter calibration, the image spatial position deviation caused by sensor differences can be significantly reduced, providing a more accurate basis for subsequent feature point matching and spatial registration.
[0090] S302, after parameter calibration, register the denoised RGB image and the filled depth image from the feature point dimension.
[0091] Specifically, after completing the sensor parameter calibration, perform spatial registration on the denoised RGB image and the filled depth image from the feature point dimension matching algorithm. The feature point matching algorithm will identify the common feature points in the two images and calculate the spatial correspondence between them. Based on these matched feature points, the algorithm will perform a spatial transformation to make the RGB image and the depth image consistent in spatial position.
[0092] In an exemplary embodiment, enhance the RGB image according to the fusion weight to obtain an enhanced image, including: multiplying the fusion weight and the RGB image using element-wise multiplication to obtain the enhanced image.
[0093] It can be understood that the fusion weight is obtained by integrating the information of the RGB image and the depth image, and it reflects the importance or significance of different regions in the image. These weights usually exist in the same size and resolution as the RGB image for element-wise multiplication operations.
[0094] Multiply the fusion weights element-wise with the corresponding pixels of the RGB image. This means that each pixel value in the RGB image will be scaled according to its corresponding fusion weight. Regions with higher fusion weights will be more emphasized in the enhanced image, while regions with lower weights will be relatively weakened.
[0095] Through the element-wise multiplication operation, an enhanced image is obtained. In this image, important or significant regions are enhanced, while unimportant regions are relatively weakened, thereby improving the overall quality and visibility of the image.
[0096] Guided by the fusion weights, the enhanced image can highlight important or significant regions, making these regions easier to detect in subsequent defect recognition. The enhancement process can improve the overall quality of the image, increase the contrast and clarity of the image, making the defects more clearly visible. Since the enhanced image highlights important regions and improves the image quality, the accuracy and reliability of defect recognition can be improved.
[0097] In an exemplary embodiment, as Figure 4 shown, this embodiment provides a defect recognition system corresponding to a defect recognition method. The defect recognition system includes an image acquisition module, an image preprocessing module, and a multi-modal detection model. The multi-modal detection model further includes a feature extraction module, a feature fusion module, a defect detection module, and a post-processing module. The output of the multi-modal detection model is the detection result, that is, the recognition result in S105. For example, taking the defect recognition method for bird's nest detection as an example, each part of the defect recognition system will be introduced below. Among them:
[0098] (1) Image acquisition module:
[0099] The image acquisition module includes two parts: an RGB image acquisition module and a depth image acquisition module.
[0100] Among them, the RGB image acquisition module is a high-resolution image sensor that acquires bird's nest images with clear image details, and these images are finely labeled manually.
[0101] The depth image acquisition module can acquire images in two ways. On the one hand, an RGB-D camera can be mounted on a drone to acquire images, and the images acquired in this part contain RGB information and depth information. On the other hand, a monocular depth estimation technology for deep learning can be used to generate corresponding depth images from the RGB images acquired by ordinary machine patrol. In addition, the most direct way to obtain a depth image is to use a depth camera to capture the depth image.
[0102] It is understandable that when using an RGB-D camera to collect transmission line images, the RGB image provides color and texture information, and the depth image provides three-dimensional morphological information. Through synchronous acquisition, the consistency of the RGB image and the depth image in time and space can be ensured.
[0103] In this embodiment, to solve the problems of high cost of the depth camera for capturing depth images and being not conducive to popularization in practical applications, a monocular depth estimation technology based on deep learning is introduced to generate corresponding depth images. For example, the monocular depth estimation method of MiDaS can achieve unprecedented zero-shot generalization performance for eight unseen datasets from indoor and outdoor domains, thus ensuring that the generated depth images meet the requirements of practical applications.
[0104] (2) Image preprocessing module:
[0105] The image preprocessing module performs image denoising, alignment processing, normalization processing, and data augmentation operations on the RGB image and the depth image respectively to ensure the quality and consistency of the RGB-D multimodal data, thereby providing a reliable data basis for subsequent feature extraction and defect detection.
[0106] Among them, the first step of the image preprocessing module is image denoising, the purpose of which is to remove the noise generated by the sensor during the acquisition process and improve the image quality. Gaussian filtering is used for RGB image denoising to remove Gaussian noise and smooth the image; while for depth image denoising, depth filling is used, and the domain average value or interpolation method is used to fill the holes and missing data in the depth map.
[0107] Then, the RGB image and the depth image are registered in space so that their corresponding pixel points represent the same spatial position. If a depth camera is used to capture depth information, calibration of the internal parameters of the sensor (such as focal length, optical center position) and external parameters (such as rotation and translation matrices) is required. At the same time, feature points are used, for example, Scale-Invariant Feature Transform (SIFT), Accelerated Robust Features or Speeded Up Robust Features (SURF) to align the RGB image and the depth image at the pixel level.
[0108] Normalization processing is mainly used to normalize the image data to a certain range for subsequent processing and analysis. Among them, the normalization of the RGB image is as shown in Equation (1), and each channel pixel of the RGB image is normalized to the range [0, 1]; the normalization of the depth image is as shown in Equation (2):
[0109] (1)
[0110] Among them, in formula (1), I is the original image, mean represents the mean value, std represents the standard deviation, and I norm is the normalized representation of the pixel values of each channel of the RGB image.
[0111] (2)
[0112] Among them, in formula (2), D represents the original depth map, max represents the maximum value, min represents the minimum value, and D norm is the normalized representation of the depth value.
[0113] In addition, data augmentation improves the generalization ability of the model by generating diverse training samples. Among them, RGB image augmentation mainly performs data augmentation through operations such as rotation and flipping, cropping and scaling, and color jitter; depth image augmentation is performed by adding noise, deformation, etc.
[0114] (3) Multimodal detection model:
[0115] The multimodal detection model is a multimodal detection model constructed based on the YOLOv8 algorithm: The multimodal detection model is mainly based on the YOLOv8 framework and consists of four main modules, such as Figure 5 shown, which are the feature extraction module, the feature fusion module, the defect detection module, and the post-processing module; among them:
[0116] (31) Feature extraction module:
[0117] The feature extraction module includes an RGB feature extraction module and a depth map feature extraction module. Both the RGB feature extraction module and the depth map feature extraction module use the backbone network module and parameters of YOLOv8 itself. Although these two backbone networks of the RGB feature extraction module and the depth map feature extraction module have similar structures, they do not share parameters, aiming to process multimodal data separately.
[0118] Among them, the RGB feature extraction module uses an optimized backbone network (such as CSPDarknet) to process the input RGB image and extracts multi-level features through a multi-layer convolutional neural network. These features cover from low-level edge information to high-level semantic information.
[0119] The depth map feature extraction module introduces another feature extraction path to specifically process the input depth image. A similar optimized backbone network is used to ensure effective extraction of depth information features within different depth ranges.
[0120] (32) Feature fusion module:
[0121] Extract the first spatial weight from the RGB image, extract the second spatial weight from the depth image, and superimpose the first spatial weight and the second spatial weight to activate and obtain the fusion weight.
[0122] Specifically, determine the first spatial weight in the RGB image according to the RGB feature map of the RGB image, and determine the second spatial weight in the depth image according to the depth feature map of the depth image; then, multiply the fusion weight by the RGB image. This method simplifies the multi-modal detection model, effectively utilizes the valuable information in the depth image, and reduces the side effects of blind spots on the entire multi-modal detection model.
[0123] Fusion weight W S The calculation formula is shown in Equation (3):
[0124] w s = σ ( f AvgP( F RGB ); MaxP( F RGB )];f AvgP( F Depth ); MaxP( F Depth )]) (3)
[0125] Where represents the sigmoid function, and f represents a convolutional layer with a kernel size of 5 × 5. and represent the RGB feature map of the RGB image and the depth feature map of the depth image respectively, and AvgP and MaxP represent the average pooling operation and the max pooling operation respectively.
[0126] The overall architecture diagram of the feature fusion module is specifically described as follows:
[0127] First, for the depth feature map ( ), two depth feature maps with a size of H×W×1 are obtained through average merging and max merging operations, where H represents the height of the feature map and C represents the width of the feature map. Then, the feature maps of the two scales are concatenated together, and the number of channels is reduced to 1 through a standard convolutional layer.
[0128] Perform the same average pooling and max pooling operations on the RGB feature map to obtain two RGB feature maps with a size of H×W×1.
[0129] Then, connect the processed depth feature map with the two RGB feature maps, and reduce the number of channels to 1 through a standard convolutional layer.
[0130] Finally, the spatial weight W S can be obtained after activation by the Sigmoid function to obtain the fusion weight. Multiply the fusion weight W S by element-wise multiplication with the RGB feature map to obtain the RGB feature map with fused depth information.
[0131] (33) Defect detection module:
[0132] After obtaining the fused multi-modal features, the defect detection module further locates and classifies the defects through a classification and regression network, and calculates the multi-task loss function.
[0133] Among them, the classification network means that the multi-modal features pass through a series of convolutional layers and fully connected layers to process the enhanced image (i.e., the RGB feature map with fused depth information), output the class probability distribution of each candidate region, and judge whether it contains defects and its specific categories; while the regression network performs bounding box regression on the features of each candidate region, predicts its precise bounding box coordinates, and further locates the defect region.
[0134] The multi-task loss function is determined according to the weights of the classification loss and the regression loss. By adjusting the weight coefficients, the importance of the classification task and the regression task is balanced to achieve joint optimization. Among them, the classification loss uses the cross-entropy loss function to calculate the difference between the predicted class probability distribution of each candidate region and the true class label. The formula for the classification loss is shown in Equation (4):
[0135] (4)
[0136] Among them, N is the number of candidate regions, y i is the true label of the i-th candidate region, and p i is the predicted probability of the i-th candidate region.
[0137] The regression loss uses the smooth L1 loss function to calculate the bounding box regression error and calculates the difference between the predicted bounding box coordinates and the true coordinates. The formula for the regression loss is shown in Equation (5):
[0138] (5)
[0139] Among them, is the predicted bounding box coordinate of the i-th candidate region, is its true bounding box coordinate, and the smooth L1 function is defined as shown in Equation (6):
[0140] (6)
[0141] The total loss function L total is the weighted sum of the classification loss and the regression loss, and is used to jointly optimize the classification and localization performance of the model. The formula for the total loss function is shown in Equation (7):
[0142] (7)
[0143] Among them, and is the weight coefficient for the classification loss and the regression loss. By adjusting these two coefficients, the importance of the classification task and the regression task can be balanced.
[0144] (34) Post-processing module:
[0145] In the YOLOv8 model, the post-processing module is an important part to ensure the accuracy and practicality of the detection results. The post-processing tasks mainly include steps such as Non-Maximum Suppression (NMS), confidence threshold screening, and result visualization. The purpose of non-maximum suppression is to eliminate redundant detection boxes and ensure that only one best detection box is retained for each target.
[0146] First, perform confidence sorting, sort all candidate detection boxes according to the confidence scores, and the detection boxes with higher scores are processed first.
[0147] Secondly, perform iterative processing. Starting from the detection box with the highest score, perform the following operations in sequence: Mark the current detection box as the finally retained detection box; Calculate the overlap degree (Intersection over Union, IoU) between the current detection box and all the other detection boxes; Delete the other detection boxes whose overlap degree with the current detection box exceeds the set threshold (such as 0.5), and consider these boxes as duplicate detections.
[0148] Repeat the above steps until all detection boxes are processed. Subsequently, perform confidence threshold screening to filter out low-confidence detection boxes and reduce false detections and noise. Screen all candidate detection boxes according to the confidence scores. Set a confidence threshold (such as 0.5), and only retain the detection boxes with confidence scores higher than this threshold (this candidate box indicates the existence of target defects), and delete the other detection boxes. Finally, perform result visualization to display the detection results in a visual form for easy intuitive viewing and analysis by users.
[0149] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.
[0150] Based on the same inventive concept, an embodiment of the present application further provides a defect recognition device for implementing the above-mentioned defect recognition method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the defect recognition device provided below can refer to the limitations on the defect recognition method in the above text and will not be repeated here.
[0151] In an exemplary embodiment, a defect recognition device is provided, as Figure 6 shown, including an acquisition module 11, an extraction module 12, a weight fusion module 13, an enhancement processing module 14, and a defect recognition module 15; where:
[0152] The acquisition module 11 is configured to acquire an RGB image and a depth image collected for a target area;
[0153] The extraction module 12 is configured to extract a first spatial weight from the RGB image and a second spatial weight from the depth image;
[0154] The weight fusion module 13 is configured to perform a superposition process on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0155] The enhancement processing module 14 is configured to perform enhancement processing on the RGB image according to the fusion weight to obtain an enhanced image;
[0156] The defect recognition module 15 is configured to identify whether there is a target defect in the enhanced image and output an identification result.
[0157] In one of the embodiments, the extraction module 12 includes:
[0158] The first denoising sub-module is configured to perform Gaussian denoising processing on the RGB image to obtain a denoised RGB image;
[0159] The second denoising sub-module is configured to perform filling processing on the depth image to obtain a filled depth image;
[0160] The fusion sub-module is configured to extract the first spatial weight from the denoised RGB image and the second spatial weight from the filled depth image.
[0161] In one of the embodiments, the second denoising sub-module is further configured to: perform filling processing on the depth image by using the neighborhood average method or the interpolation method to obtain a filled depth image.
[0162] In one of the embodiments, the second denoising sub-module is further configured to: perform registration processing on the denoised RGB image and the filled depth image to obtain a registered RGB image and a registered depth image;
[0163] Extract the first spatial weight from the registered RGB image and extract the second spatial weight from the registered depth image.
[0164] In one embodiment, the fusion sub-module is further configured to: calibrate the parameters of the RGB sensor corresponding to the RGB image and the depth sensor corresponding to the depth image;
[0165] After parameter calibration, perform registration processing on the denoised RGB image and the filled depth image from the feature point dimension to obtain the registered RGB image and the registered depth image.
[0166] In one embodiment, the enhancement processing module is further configured to: multiply the fusion weight and the RGB image using element-wise multiplication to obtain an enhanced image.
[0167] Each module in the above defect recognition device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0168] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a communication method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0169] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0170] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0171] Obtain the RGB image and the depth image collected for the target area;
[0172] Extract the first spatial weight in the RGB image and extract the second spatial weight in the depth image;
[0173] Perform superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0174] According to the fusion weight, perform enhancement processing on the RGB image to obtain an enhanced image;
[0175] Identify whether there is a target defect in the enhanced image and output the identification result.
[0176] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0177] Obtain an RGB image and a depth image collected for a target area;
[0178] Extract a first spatial weight from the RGB image and extract a second spatial weight from the depth image;
[0179] Perform a superposition process on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0180] Enhance the RGB image according to the fusion weight to obtain an enhanced image;
[0181] Identify whether there is a target defect in the enhanced image and output the identification result.
[0182] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0183] Obtain an RGB image and a depth image collected for a target area;
[0184] Extract a first spatial weight from the RGB image and extract a second spatial weight from the depth image;
[0185] Perform a superposition process on the first spatial weight and the second spatial weight to obtain a fusion weight;
[0186] Enhance the RGB image according to the fusion weight to obtain an enhanced image;
[0187] Identify whether there is a target defect in the enhanced image and output the identification result.
[0188] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0189] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.
[0190] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A defect identification method, characterized in that: The method comprises: Obtain RGB images and depth images collected for the target area; Performing Gaussian denoising on the RGB image to obtain a denoised RGB image; Identifying missing or invalid areas in the depth image; According to the size and shape of the missing area or the invalid area, and the distribution of surrounding pixels, a neighborhood average method or an interpolation method is selected to fill the depth image to obtain a filled depth image; Extract the first spatial weight in the denoised RGB image, and extract the second spatial weight in the filled depth image; wherein the first spatial weight is determined according to the weight assigned to each pixel in the RGB image, and the second spatial weight is determined according to the weight assigned to each pixel in the depth image; and perform pixel-by-pixel superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight corresponding to each pixel, which is specifically the following formula: Ws=σ(f[AvgP(F RGB );MaxP(F RGB )];f[AvgP(F Depth );MaxP(F Depth )]) Among them, W S represents the fusion weight; σ represents the sigmoid function; f represents the convolution layer with a kernel size of 5×5; F RGB represents the RGB feature map of the RGB image; F Depth represents the depth feature map of the depth image; AvgP represents the average pooling operation; MaxP represents the maximum pooling operation; Multiply the fusion weights corresponding to each pixel by the corresponding pixels in the RGB image to obtain an enhanced image; Identify whether a target defect exists in the enhanced image and output a recognition result; Wherein, the identifying whether there is a target defect in the enhanced image includes: The enhanced image is input into a classification network and a regression network to determine the defect area in the enhanced image and the defect category corresponding to the defect area, so as to determine whether there is a target defect in the enhanced image; wherein the classification network is used to determine the category probability distribution of each candidate area in the enhanced image, so as to judge whether each candidate area contains a defect and the specific category; the regression network is used to perform bounding box regression on the features of each candidate area, and predict the bounding box coordinates of each candidate area, so as to determine the defect area.
2. The method according to claim 1, characterized in that: The extracting the first spatial weight in the denoised RGB image and extracting the second spatial weight in the filled depth image includes: Performing registration processing on the denoised RGB image and the filled depth image to obtain a registered RGB image and a registered depth image; The first spatial weight in the registered RGB image is extracted, and the second spatial weight in the registered depth image is extracted.
3. The method according to claim 2, characterized in that The registering process of the denoised RGB image and the filled depth image to obtain the registered RGB image and the registered depth image includes: Performing parameter calibration on an RGB sensor corresponding to the RGB image and a depth sensor corresponding to the depth image; After parameter calibration, the denoised RGB image and the filled depth image are registered from the feature point dimension to obtain a registered RGB image and a registered depth image.
4. The method according to claim 1, characterized in that: The identifying whether there is a target defect in the enhanced image includes: Selecting a defect recognition algorithm according to the characteristics and recognition requirements of the target defect; The defect recognition algorithm is applied to the enhanced image to detect whether the target defect exists in the enhanced image.
5. The method according to claim 4, characterized in that The defect recognition algorithm is any one of an image processing-based method, a machine learning algorithm or a deep learning model.
6. The method according to claim 1, characterized in that The target defect is a bird's nest defect.
7. A defect identification device, characterized in that: The device comprises: An acquisition module is used to acquire RGB images and depth images collected for the target area; An extraction module is used to perform Gaussian denoising on the RGB image to obtain a denoised RGB image; identify missing areas or invalid areas in the depth image; select a neighborhood average method or an interpolation method according to the size and shape of the missing area or the invalid area and the distribution of surrounding pixels to fill the depth image to obtain a filled depth image; extract a first spatial weight from the denoised RGB image and extract a second spatial weight from the filled depth image; wherein the first spatial weight is determined according to the weight assigned to each pixel in the RGB image, and the second spatial weight is determined according to the weight assigned to each pixel in the depth image; The weight fusion module is used to perform pixel-by-pixel superposition processing on the first spatial weight and the second spatial weight to obtain a fusion weight corresponding to each pixel, which is specifically the following formula: Ws=σ(f[AvgP(F RGB );MaxP(F RGB )];f[AvgP(F Depth );MaxP(F Depth )]) Among them, W S represents the fusion weight; σ represents the sigmoid function; f represents the convolution layer with a kernel size of 5×5; F RGB represents the RGB feature map of the RGB image; F Depth represents the depth feature map of the depth image; AvgP represents the average pooling operation; MaxP represents the maximum pooling operation; An enhancement processing module, used for multiplying the fusion weight corresponding to each pixel by the corresponding pixel in the RGB image to obtain an enhanced image; A defect recognition module is used to identify whether there is a target defect in the enhanced image and output a recognition result; wherein, the identification of whether there is a target defect in the enhanced image includes: inputting the enhanced image into a classification network and a regression network, determining the defect area in the enhanced image and the defect category corresponding to the defect area, so as to determine whether there is a target defect in the enhanced image; wherein, the classification network is used to determine the category probability distribution of each candidate area in the enhanced image, so as to judge whether each candidate area contains a defect and a specific category; the regression network is used to perform bounding box regression on the features of each candidate area, and predict the bounding box coordinates of each candidate area, so as to determine the defect area.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image fusion method, electronic equipment, storage medium and computer program product
CN113888452A
Defect detection method and device, computer equipment and storage medium
CN118446982A