A high-precision multi-view image recognition detection method based on a CCD camera

By employing self-supervised learning and multi-view adaptive segmentation techniques, combined with image feature enhancement and bounding box regression, the problems of viewpoint differences and noise interference in CCD detection are solved, achieving high-precision motor defect detection.

CN120279228BActive Publication Date: 2025-11-04NIDEC (SHAOGUAN) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510287963.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-11-04
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Existing CCD inspection technology suffers from problems such as viewing angle differences, lighting changes, and noise interference in motor defect detection, leading to missed detections and false detections, making it difficult to achieve high-precision defect identification.

Method used

A method combining self-supervised learning with multi-view adaptive segmentation and image feature enhancement is adopted. Noise-aware features are extracted through convolutional neural networks, and weighted average fusion and denoising are performed. Semantic segmentation is performed using an adaptive multi-view semantic segmentation network. Target region features are extracted by combining a self-attention mechanism, and view consistency matching and feature fusion are performed. Finally, defective product regions are located by bounding box regression.

Benefits of technology

It effectively removes noise interference, improves the recognition accuracy of multi-view images, ensures the accuracy of target recognition and positioning in complex environments, and enhances the accuracy and efficiency of motor defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279228B_ABST
    Figure CN120279228B_ABST
Patent Text Reader

Abstract

The application provides a high-precision multi-view image recognition detection method based on a CCD camera, which comprises the following steps: a multi-view image set is captured by a CCD camera array, noise perception features of each view image are extracted by using a convolutional neural network; semantic segmentation is performed by using an adaptive multi-view semantic segmentation network; final feature representation of a target region is constructed based on enhanced features and key features; similarity calculation is used to perform view consistency matching based on the final feature representation of the target region; and a classification module is combined to recognize a defect type based on a final target feature map, and the class and position information of a defective product are output. The application realizes noise removal of multi-view images by using a self-supervised learning method, and optimizes image features under different views by combining an adaptive semantic segmentation technology, so that accurate target recognition of multi-view images is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of high-precision multi-view image recognition, and particularly relates to a high-precision multi-view image recognition detection method based on a CCD camera. BACKGROUND

[0002] With the development of industrial automation and precision detection technology, CCD (Charge-Coupled Device) high-definition detection equipment has been widely used in the electronic manufacturing industry, especially in the defect detection of motor assemblies. Defects of the motor may manifest as missing solder, missing glue, foreign matter attachment, etc., which have a direct impact on the performance and safety of the product. Therefore, accurate and efficient defect detection is an important link to improve product quality and reduce the rate of defective products.

[0003] Although current CCD detection technology can provide high-resolution images and detect some common defects, there are still many challenges in practical applications. First, due to the complex and variable surface texture of the motor, the CCD detection equipment may not be able to fully capture all defects at certain angles, resulting in missed detection and false detection. In addition, due to environmental noise, changes in lighting, and the angle limitations of the equipment itself, the manifestations of defects in the motor may differ significantly in different images, which greatly affects the accuracy of defect detection.

[0004] Currently, traditional methods for motor defect detection mostly rely on manual feature extraction and rule-based judgment, which are inefficient and inaccurate for defect recognition in high-complexity and large-scale production lines. At the same time, although some automatic detection methods based on convolutional neural networks (CNN) have made preliminary progress, these methods are often affected by factors such as lighting, angle, and noise, and still perform poorly in detecting small defects or subtle differences.

[0005] Therefore, there is an urgent need for a new type of CCD detection technology that can improve the detection accuracy of motor defects (such as missing solder, missing glue, and foreign matter attachment) through more accurate image processing and feature optimization, and overcome problems such as angle differences, lighting changes, and noise interference. SUMMARY

[0006] The purpose of the present application is to provide a high-precision multi-view image recognition detection method based on a CCD camera, which combines self-supervised learning, multi-view adaptive segmentation, and image feature enhancement to provide an innovative multi-view image denoising and accurate recognition solution.

[0007] To achieve the above purpose, a high-precision multi-view image recognition detection method based on a CCD camera is provided, which comprises:

[0008] The multi-view image set photographed by the CCD camera array extracts noise perception features of each view image by using a convolutional neural network, fuses a global noise estimation value based on the noise perception features of each view image by weighted average, and denoises the global noise estimation value based on the convolutional neural network to obtain a denoised multi-view image set;

[0009] The denoised multi-view image set is subjected to semantic segmentation by using an adaptive multi-view semantic segmentation network to obtain a multi-view image set after semantic segmentation, wherein each image in the multi-view image set is the result of semantic segmentation of the corresponding view image, and contains semantic label information of each pixel; wherein the adaptive multi-view semantic segmentation network is used to jointly model the features of the denoised multi-view image, and realizes semantic segmentation by sharing feature learning and view feature weighting;

[0010] The target region features are adaptively enhanced to obtain enhanced features according to the semantic label information of each pixel, and the key features of the target region are extracted by using a self-attention mechanism, and the final feature representation of the target region is constructed based on the enhanced features and the key features;

[0011] Similarity calculation is used to perform view consistency matching based on the final feature representation of the target region, and feature fusion is performed based on the final feature representation of the target region, and a final target feature map is generated based on convolution operation and pooling operation;

[0012] Based on the final target feature map, the defective region is located by feature enhancement and bounding box regression, the defect type is identified in combination with a classification module, and the category and position information of the defective product are output.

[0013] Further, the noise perception features are extracted by the convolutional neural network from the gradient and local texture features of the multi-view image; the weighted average fusion includes:

[0014] The noise perception features of each view image are given different weights according to their quality, and the noise perception features are fused by weighted average to obtain the global noise estimation value;

[0015] The denoising of the global noise estimation value based on the convolutional neural network includes:

[0016] The convolutional neural network is used for denoising processing, and the denoised image is recovered from each multi-view image by removing the global noise estimation value, and the smoothness constraint of the image is introduced in the recovery process to ensure that the structure and texture of the image are preserved.

[0017] Further, the convolutional neural network adopts a self-supervised learning method to minimize the reconstruction error of the image after denoising; the loss function of the convolutional neural network includes the image reconstruction error and the noise estimation error, and the network weight is adjusted through self-supervised learning for gradually approximating the real noise.

[0018] Further, the adaptive multi-view semantic segmentation network includes a deep convolutional neural network; the semantic segmentation is performed by using the adaptive multi-view semantic segmentation network to obtain a multi-view image set after semantic segmentation, including:

[0019] Based on the similarity between each view image, the features of each view are weighted and fused to adaptively adjust the contribution of different view images to the segmentation result to obtain the weighted and fused view features.

[0020] The weighted and fused view features of each view image are adjusted by an adaptive weighting mechanism to obtain the final global features; the adaptive weighting mechanism uses the similarity weighting sum of the multi-view image and other images as the weight to weight and calculate the weighted and fused view features.

[0021] The final global features are input into the deep convolutional neural network for final semantic segmentation to obtain the multi-view image set after semantic segmentation.

[0022] Further, the adaptive enhancement is weighted and calculated based on the distance from the i-th pixel to the center of the target region and the phonetic label information to obtain the weighting coefficient of each pixel in the target region; according to the weighting coefficient of each pixel in the target region, the weight of each pixel can be dynamically adjusted according to the semantic category and position of the target to enhance the representation of the target feature; the self-attention mechanism is based on the weighting and calculation of the semantic suppression based on the weighting coefficient of each pixel in the target region and the semantic feature of the corresponding pixel to obtain the key feature; wherein the semantic suppression suppresses the background interference by calculating the contrast between the background region and the target region.

[0023] Further, the similarity calculation is used to perform view consistency matching based on the final feature representation of the target region, including:

[0024] The view consistency loss function is determined by the target features from the i-th and j-th views.

[0025] Based on the view consistency loss function, the target features from different views are optimized to maintain consistency between target features with similar appearances.

[0026] Further, the feature fusion based on the final feature representation of the target region is weighted fusion based on the feature difference measurement of the i-th view and the features of all other views to determine the fused target feature map.

[0027] Further, the feature enhancement includes high-frequency texture enhancement and edge strengthening of the target feature map, which amplifies the texture features of the missing soldering tin, missing glue coating and foreign matter attachment area.

[0028] Further, the bounding box regression is performed by the target feature map to regress the bounding box and accurately locate the defects of each motor.

[0029] Wherein, the smooth L1 loss is used to optimize the accuracy of positioning, and a spatial transformation regularization term is introduced to process the positioning error caused by the characteristics of the CCD device.

[0030] Further, the classification module introduces a fuzzy adversarial loss function as a regularization term to simulate a fuzzy scene of a low-resolution or noisy image, thereby improving the robustness of the classification module to fuzzy images.

[0031] The beneficial technical effects of the present application are at least as follows:

[0032] The present application effectively solves the problems of noise interference, view angle difference and information inconsistency in the existing multi-view image processing technology. Unlike the prior art, the present application realizes noise removal of multi-view images through a self-supervised learning method, and simultaneously combines adaptive semantic segmentation technology to optimize image features under different views, thereby realizing accurate target recognition of multi-view images. The specific invention points are as follows:

[0033] Multi-view image noise modeling and denoising self-supervised learning:

[0034] The present application constructs a multi-view image noise modeling and denoising network through a self-supervised learning mechanism. By extracting gradient information and local texture features from each view image, combining an adaptive weighted fusion strategy, accurately estimating and removing noise in multi-view images, and ensuring that key details and image structures are preserved during the denoising process. This innovation enables the denoising network to not only adapt to different noise types, but also to adaptively adjust the denoising intensity in multi-view data, thereby improving the denoising effect.

[0035] Adaptive multi-view semantic segmentation:

[0036] In multi-view image processing, directly applying traditional semantic segmentation networks often fails to utilize the complementary information from different views. To address this issue, the present invention proposes an adaptive multi-view semantic segmentation network that calculates the similarity between images from different views and adaptively adjusts the contribution of each view to the segmentation result based on the differences in the views. This method improves the accuracy of semantic segmentation, especially in cases where the view difference is large, ensuring accurate extraction of key semantic information from the image.

[0037] Target feature enhancement and key feature extraction:

[0038] For the feature representation of the target region, the present invention improves the representation ability of the target feature through adaptive feature enhancement and key feature extraction mechanism. After image semantic segmentation, the dynamic weight is given to the feature of the target region by combining the semantic information of the target region, thereby improving the discrimination of the target feature. At the same time, the semantic suppression mechanism is introduced to reduce the background interference, ensuring the accurate extraction of the target feature in complex environment, providing high-quality feature support for subsequent target recognition and positioning.

[0039] Feature fusion based on view consistency optimization:

[0040] The present invention proposes a feature fusion strategy based on view consistency optimization. By optimizing and weighting the target features from different views for consistency, it ensures that the target features can be accurately preserved in multi-view images. The introduction of view consistency matching further improves the quality of feature fusion, suppresses the interference of redundant information, and generates more accurate and efficient feature maps, improving the accuracy of target recognition and positioning. BRIEF DESCRIPTION OF DRAWINGS

[0041] The present invention is further illustrated by the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained without creative labor based on the following drawings.

[0042] Figure 1 A flowchart of a high-precision multi-view image recognition and detection method based on a CCD camera according to the present invention. DETAILED DESCRIPTION

[0043] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are only used to explain the present invention, and cannot be understood as a limitation on the present invention.

[0044] As Figure 1As shown, the embodiment of the application provides a high-precision multi-view image recognition detection method based on a CCD camera, which comprises the following steps S1-S5:

[0045] S1, a multi-view image set photographed by a CCD camera array, noise perception features of each view image are extracted by using a convolutional neural network, a global noise estimation value is fused by weighted average based on the noise perception features of each view image, and the global noise estimation value is denoised based on the convolutional neural network, so as to obtain a multi-view image set after denoising.

[0046] Specifically, the input of step S1 is a multi-view image set photographed by a CCD camera array, and it is assumed that these image sets are I1, I2, …, I n , respectively. i The i-th CCD camera in the array. Different cameras may produce noise interference due to factors such as shooting angle, light, lens distortion, etc. The target is to construct a system for modeling and denoising of these CCD camera image noise features through self-supervised learning, to eliminate noise and preserve key details of the image. Since the CCD camera array often has noise differences in multi-view images caused by different hardware conditions and environmental factors, the application designs a noise perception model combining gradient information and local texture of the image to model these noise features. Specifically, the application uses a convolutional neural network (CNN) to extract noise perception features of each view image I i . This process not only captures high-frequency noise in the image, but also considers texture distortion. The application defines the following feature extraction formula:

[0047]

[0048] , wherein, represents the gradient of the image I i , which is used to capture noise and high-frequency components in the image, represents the local texture feature of the image I i , which is usually extracted by a gray level co-occurrence matrix, and θ is a trainable parameter of the CNN. The goal of this step is to extract the latent noise features from the CCD camera images to provide basic information for the denoising stage.

[0049] Further, in the multi-view image, since the noise features of each camera view are not completely the same, the application adopts an adaptive weighted fusion strategy to calculate the global noise estimation combining the noise features of each view. Specifically, the noise features of each view are given different weights w i according to their quality (including sharpness, noise intensity, etc.), and these noise features are fused by weighted average. The global noise estimation value after fusion By the following formula:

[0050]

[0051] wherein, is the noise feature extracted from the i-th view image, w i is the adaptive weight of the i-th view, and the specific calculation method is as follows:

[0052]

[0053] The calculation method of the weight w i ensures that the image with less or clearer noise has a lower weight, and the image with stronger noise (for example, the image taken in the case of insufficient light or large lens distortion) has a higher weight, thereby ensuring the effectiveness of noise fusion.

[0054] Further, after obtaining the global noise estimate, the present application uses a convolutional neural network (CNN) for denoising processing. By removing the estimated noise from each view image I i to recover the clear image I cleaned . The denoising operation is as follows:

[0055]

[0056] wherein, I cleaned is the denoised image, is the global noise estimate.

[0057] In order to prevent the loss of image details in the denoising process, the present application introduces a regularization constraint to ensure that the structure and texture of the image are preserved. The specific loss function is as follows:

[0058]

[0059] wherein, λ is the weight of the regularization term, controlling the smoothness, is the smoothness constraint of the image, ensuring the preservation of details.

[0060] Further, in order to train the denoising network, the present application adopts a self-supervised learning method to minimize the reconstruction error of the denoised image. The loss function includes the image reconstruction error and the noise estimation error, and the network weight is adjusted through self-supervised learning to gradually approach the true noise. The training loss function is as follows:

[0061]

[0062] wherein, α is the adjustment hyperparameter, controlling the balance between the denoising error and the noise estimation error, and ∈ true is the true noise.

[0063] The final output is a set of denoised multi-view images I cleaned = {I1', I2',..., In'} n These denoised images will be used for subsequent multi-view image recognition and detection tasks.

[0064] S2, on the set of denoised multi-view images, an adaptive multi-view semantic segmentation network is used for semantic segmentation to obtain a set of multi-view images after semantic segmentation, wherein each image in the multi-view image set is the result of its corresponding view image after semantic segmentation, containing semantic label information of each pixel; wherein the adaptive multi-view semantic segmentation network is used to jointly model the features of the denoised multi-view images, and realizes semantic segmentation through shared feature learning and view feature weighting.

[0065] Specifically, since multi-view images may have significant differences in view, illumination, occlusion, etc., directly inputting denoised images into traditional semantic segmentation networks may not be able to fully utilize the view information between images. Therefore, the present application designs an adaptive multi-view semantic segmentation network that can adaptively adjust according to the features of each view image and its association with other view images. Specifically, the present application jointly models the features of multi-view images, and realizes more accurate semantic segmentation through shared feature learning and view feature weighting. The main structure of the network includes the following parts:

[0066] Further, the present application designs an adaptive feature fusion mechanism for fusing information from different view images. By calculating the similarity S ij between each view image, the features of each view are fused using this similarity. This mechanism can adaptively adjust the contribution of different view images to the segmentation result. The similarity S ij is calculated as follows:

[0067]

[0068] where cosine_similarity(I i ', I j ') is the cosine similarity between the denoised images I i ' and I j ', used to measure the similarity of two view images, and ||I i ' ||2 and ||I j ' ||2 are the L2 norms of images I i ' and I j ', respectively, used to normalize image features. The similarity measure value S ijThe features of each view image are weighted for better reflecting the similarities and differences between views.

[0069] Further, after the feature fusion, the features of each view image are adjusted by an adaptive weighting mechanism to obtain the optimal segmentation effect. The features of each view image are weighted to obtain the final global features The formula is as follows:

[0070]

[0071] wherein, is the feature of the i-th view image obtained by the view feature fusion module, w i is an adaptive weighting coefficient, which is adjusted based on image quality, view difference and image content complexity. i The calculation method of w

[0072]

[0073] wherein, is the similarity weighted sum of image I i ′ with other images, which is used to measure the importance of the image in the overall segmentation.

[0074] Through this weighting mechanism, each view image is dynamically adjusted according to its relationship and contribution with other view images, so that the multi-view information is more accurately utilized in semantic segmentation.

[0075] Further, after obtaining the globally weighted fused features , the present application inputs them into a deep convolutional neural network (CNN) for final semantic segmentation. The present application uses structures such as U-Net, which has been widely used in image segmentation tasks and can accurately segment the boundaries while preserving the details of the image. The loss function of the segmentation task is:

[0076]

[0077] wherein y k is the k-th class pixel label of the ground truth label image, is the k-th class pixel label predicted by the network, K is the number of classes, is the regularization term of image I cleaned to prevent the segmentation result from being too rough.

[0078] The final output is the multi-view image set after semantic segmentation is the result of the corresponding view image after semantic segmentation. These segmented images will provide accurate semantic information for subsequent image recognition, detection and other analysis tasks.

[0079] S3, adaptively enhance the target region features according to the semantic label information of each pixel to obtain enhanced features, and extract key features of the target region using a self-attention mechanism, and construct the final feature representation of the target region based on the enhanced features and the key features.

[0080] Specifically, the input of step 3 is the output of step 2, the image set processed by the denoising and adaptive multi-view semantic segmentation network Each image has been segmented and contains semantic label information for each pixel.

[0081] Further, the goal of this step is to improve the feature representation ability of the target based on semantic information, combined with adaptive feature enhancement and target key feature extraction mechanism, to further improve the accuracy of subsequent target recognition and detection. The core strategy is to enhance the features of the target region by assigning dynamic weights to the target region, while extracting key target features through semantic information.

[0082] Further, according to the semantic label information output by step 2, the present application adaptively enhances the features of the target region. After semantic segmentation of the image, the present application can determine the semantic class of each pixel, and enhance the expression ability of the target region by weighting the pixel features of the target region. For each pixel in the target region, the weighting coefficient w i is calculated by considering not only the distance d i of the pixel to the center of the target, but also introducing a weighting factor a i based on the target class to adapt to the differences in target features of different semantic classes. The specific calculation formula is as follows:

[0083]

[0084] where d i is the distance of the i-th pixel to the center of the target region, and s is a hyperparameter used to control the smoothness of the weight; a i is a weighting factor based on the semantic label (for example, for the background region, a i is small, and for the target region, a i is large). By combining d i and a i , the present application can dynamically adjust the weight of each pixel according to the semantic class and position of the target, and optimize the representation of the target features.

[0085] Further, after the target region feature enhancement, the application further extracts the key features of the target. In order to ensure accurate identification of the target under complex background and different perspectives, the application introduces a feature suppression mechanism based on semantic information to eliminate the interference of background information on target features. The key feature extraction process combines a self-attention mechanism, enabling the model to focus on important areas in the image. The application realizes feature extraction through the following calculation method:

[0086]

[0087] wherein, is the semantic feature of the i-th pixel, w i is the weighting coefficient of the target region (as described above), is the semantic suppression factor, which suppresses background interference by calculating the contrast between the background region and the target region. The calculation method of is as follows:

[0088]

[0089] wherein, is the feature vector of the background region, and β is a balance factor for adjusting the contrast between the target and the background features. Through this suppression mechanism, the model can focus more on the key information of the target region and suppress background interference, thereby improving the identification accuracy of the target.

[0090] Further, finally, based on the enhanced features and the extracted key features, the application constructs the final feature representation of the target. This representation combines the adaptively enhanced target region features and the key features extracted from semantic information, and weightedly fuses to obtain the final target feature representation:

[0091]

[0092] wherein, is the final feature representation of the target region, is the weighting coefficient matrix, which is dynamically adjusted based on the semantic label and the relative importance of the target region. This feature representation contains the semantic information and geometric features of the target, which can effectively support subsequent target identification and positioning.

[0093] By introducing the target feature enhancement based on semantic information and the adaptive key feature extraction mechanism, this step can effectively improve the feature representation ability of the target. In the multi-perspective and complex background of the image, the scheme of the application can automatically focus on the target region and reduce background interference through semantic suppression, and finally construct a more accurate target feature representation. This process plays a key role in improving the target detection accuracy and positioning ability.

[0094] S4, similarity calculation is adopted to perform target region-based final feature representation view consistency matching, and feature fusion is performed based on the target region-based final feature representation, and a final target feature map is generated based on convolution operation and pooling operation.

[0095] Specifically, the main goal of the present step is to combine target features under different views through view consistency optimization, perform effective feature fusion, and then generate a final high-quality feature map for subsequent target recognition and positioning. In order to achieve this goal, the present application designs a feature fusion strategy based on adaptive view consistency matching, and through optimization of multi-view feature maps, consistency and accuracy are ensured under different views.

[0096] Further, under multi-view, the appearance of the target may change, in order to ensure the consistency of the features, the present application first introduces a view consistency optimization mechanism, which calculates the feature changes of the target under different views to optimize the similarity between the features. In this process, the present application matches based on the geometric features and semantic information of the target.

[0097] The present application measures the similarity of features under different views through a view consistency loss function The specific definition is as follows:

[0098]

[0099] Wherein, and are target features from the i-th and j-th views, respectively; is the view consistency weight, which is calculated according to the view difference, and is used to adjust the matching strength of features from different views. 2 Euclidean distance metric, used to measure the difference between feature vectors. Through this loss function, the present application can optimize the target features from different views, so that the target features with similar appearances maintain consistency, thereby enhancing the accuracy of the target features.

[0100] Further, on the basis of view consistency matching, the present application then performs feature fusion to effectively synthesize target features under multiple views. Feature fusion not only retains important information under each view, but also suppresses redundant or noisy information. The present application designs an adaptive feature fusion function to weight and fuse features from different views to obtain a unified target feature map. The fusion process is as follows:

[0101]

[0102] Wherein, is the fused target feature map; σ iσ represents the weighting coefficient from the i-th viewpoint, indicating the contribution of that viewpoint to the final feature map. i The value is dynamically adjusted based on perspective consistency and feature importance. Weighting coefficient σ i The calculation method is as follows:

[0103]

[0104] Among them, D i γ is a measure of the difference between the features from the i-th viewpoint and the features from all other viewpoints, typically calculated using Euclidean distance. γ is a hyperparameter used to control the smoothness of the weighting coefficients. Through this weighted fusion strategy, this invention ensures that the most valuable viewpoint information contributes the most to the final feature map, thereby improving the target recognition capability.

[0105] Furthermore, after feature fusion, this invention further optimizes the generated feature map to ensure its superior performance in object detection tasks. By performing convolution and pooling operations on the fused feature map, this invention can extract high-level semantic features and further compress the feature dimension to generate the final target feature map. The calculation formula is as follows:

[0106]

[0107] Where: Conv(·) represents the convolution operation, used to extract high-level features; Pool(·) represents the pooling operation, used to reduce the dimensionality of the feature map and retain the most important information; This represents the splicing operation of features.

[0108] Finally, the generated target feature map This will provide accurate feature support for subsequent target detection and localization.

[0109] By optimizing viewpoint consistency and fusing features, this step effectively improves the consistency and accuracy of target features across different viewpoints. Utilizing an adaptive viewpoint consistency loss function and a weighted fusion mechanism, this invention optimizes the feature representation of targets from different viewpoints, generating a more accurate target feature map. This process is significant for multi-view target detection and localization, and can effectively improve the robustness and accuracy of the system in practical applications.

[0110] S5. Based on the final target feature map, the defective product area is located through feature enhancement and bounding box regression. Combined with the classification module, the defect type is identified, and the category and location information of the defective product are output.

[0111] Specifically, the input is the output of step S4, which is the final target feature map after viewpoint consistency optimization, feature fusion, and enhancement processing.

[0112]

[0113] Each is a feature map containing various types of defects on the motor surface (such as missing solder, missing glue, foreign matter attachment, etc.), providing a deep understanding of image content and texture. Our goal is to conduct precise defect detection and positioning based on these optimized feature maps.

[0114] Further, although the feature map of step S4 has been basically optimized, in order to further enhance the saliency of the target defect and ensure efficient recognition, we add a feature enhancement module to further amplify the features of the defects using special convolution operations and enhancement functions. Through such enhancement, we can better reveal those defects that may be partially displayed or changed due to camera perspective.

[0115] The operation of the feature enhancement module is as follows:

[0116]

[0117] where ε represents the feature enhancement operation, which enhances the features of defects such as missing solder, missing glue, and foreign matter attachment on the motor surface by using high-frequency texture enhancement filters, edge enhancement, and blur enhancement, making them more prominent and easy to identify. Specifically, is the enhanced target feature map, where the texture features of the defect area are more obvious, facilitating subsequent processing.

[0118] Further, after obtaining the enhanced feature map we input it into the target positioning module, whose core task is to locate the defective area according to the texture features in the image. Through the optimized feature map, we perform bounding box regression and accurately locate each motor defect.

[0119] For each defect (such as missing solder, missing glue, or foreign matter attachment), we not only predict its location but also calculate its actual size and orientation in the image. To this end, we use the smooth L1 loss to optimize the accuracy of positioning, and introduce an additional spatial transformation regularization term to better handle the distortion and perspective changes in the CCD image.

[0120] The positioning optimization calculation formula is as follows:

[0121]

[0122] where is the true bounding box of target i (i.e., the actual area position of missing solder, missing glue, or foreign matter attachment); is the predicted bounding box; is the localization loss, which calculates the difference between the predicted bounding box and the ground truth bounding box using a smooth L1 loss; is the spatial transformation loss, which is used to correct the localization error caused by CCD device distortion or imperfect viewing angle; λ spatial is the weight of the spatial transformation regularization term, which controls its impact on model training.

[0123] This spatial transformation loss is specifically designed for the distortion of defective areas appearing on the motor surface, which can help to handle the localization error caused by the characteristics of the CCD device (such as optical distortion, viewing angle change, etc.).

[0124] Further, once the precise positioning of the defective area is completed, the next step is to classify each defect area and determine its specific type. For example: missing solder, missing glue, foreign matter attachment, etc. To improve detection accuracy, we introduce a deep convolutional feature separation module that extracts fine-grained texture features from the positioning area and classifies them.

[0125] The formula for the classification process is as follows:

[0126]

[0127] where, is the classification module, which is used to classify the target class (missing solder, missing glue, foreign matter attachment, etc.) based on the enhanced feature map; is the predicted class label.

[0128] To deal with the possibility of blur and unclear situations in the CCD device image, we added a blur adversarial loss as a regularization term to ensure that the model can correctly classify defective products even in low-resolution or noisy images. The calculation method of the blur adversarial loss is as follows:

[0129]

[0130] where, is the standard classification loss (such as mean square error); is the blur loss based on adversarial training, which is used to simulate different degrees of blur in the image, thereby improving the model's robustness to blurred or noisy images; β is the weight hyperparameter of the blur adversarial loss.

[0131] Training and loss function design:

[0132] To ensure the overall performance of the model on positioning and classification tasks, we combine multiple loss functions such as positioning loss, classification loss, and fuzzy adversarial loss for multi-objective training. The final total loss function is as follows:

[0133]

[0134] where: and are the classification and positioning losses, respectively; is the fuzzy adversarial loss, which helps improve the recognition ability of low-quality images; is the spatial transformation loss, which solves the problems caused by device viewing angle or optical distortion; and λ1 and λ2 are weight hyperparameters of each loss.

[0135] Output and results:

[0136] The final output will be:

[0137] The predicted category of each defective area (soldering tin leakage, glue coating leakage, foreign matter attachment, etc.);

[0138] The accurate bounding box position of each defective area.

[0139] These outputs can be directly applied to CCD detection equipment to help achieve efficient and accurate automated detection and classification, thereby greatly reducing the risk of defective products and ensuring product quality.

[0140] It should be noted that the workflow described above is only illustrative and does not limit the scope of protection of the present application. In actual applications, those skilled in the art can select part or all of them to achieve the purpose of the present embodiment, and this place does not limit.

[0141] In addition, technical details not described in detail in this embodiment can be referred to the parameter operation method provided by any embodiment of the present application, which will not be repeated here.

[0142] It should be noted that in this paper, the term "includes" "contains" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or system. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of another identical element in the process, method, article or system that includes the element.

[0143] The above embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and the necessary general hardware platform, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, an optical disk), and includes a plurality of instructions for making a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) execute the methods described in various embodiments of the present application.

[0145] The above are only preferred embodiments of the present application, and do not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A high-precision multi-view image recognition and detection method based on a CCD camera, characterized in that, The method includes: A set of multi-view images captured by a CCD camera array is used to extract noise perception features of each view image using a convolutional neural network. Based on the noise perception features of each view image, a global noise estimate is fused by weighted averaging. The global noise estimate is then denoised using a convolutional neural network to obtain a denoised set of multi-view images. For the denoised multi-view image set, an adaptive multi-view semantic segmentation network is used for semantic segmentation to obtain a semantically segmented multi-view image set. Each image in the multi-view image set is the result of semantic segmentation of its corresponding view image, containing semantic label information for each pixel. The adaptive multi-view semantic segmentation network is used to jointly model the features of the denoised multi-view images, and achieves semantic segmentation through shared feature learning and view feature weighting. The target region features are adaptively enhanced based on the semantic label information of each pixel to obtain enhanced features, and the key features of the target region are extracted using a self-attention mechanism. The final feature representation of the target region is constructed based on the enhanced features and the key features. Similarity calculation is used to perform viewpoint consistency matching based on the final feature representation of the target region, and feature fusion is performed based on the final feature representation of the target region. The final target feature map is generated based on convolution and pooling operations. Based on the final target feature map, the defective product region is located through feature enhancement and bounding box regression. Combined with the classification module, the defect type is identified, and the category and location information of the defective product are output.

2. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 1, characterized in that, The noise-aware features are extracted from the gradient and local texture features of multi-view images using a convolutional neural network; the weighted average fusion includes: The noise perception features of each viewpoint image are assigned different weights according to their quality, and the noise perception features are fused by weighted averaging to obtain a global noise estimate. The denoising of the global noise estimate based on the convolutional neural network includes: Denoising is performed using a convolutional neural network, and the denoised image is recovered from each multi-view image by removing global noise estimates. At the same time, image smoothness constraints are introduced during the recovery process to ensure that the structure and texture of the image are preserved.

3. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 2, characterized in that, The convolutional neural network employs a self-supervised learning method to minimize the reconstruction error after image denoising. The loss function of the convolutional neural network includes the image reconstruction error and the noise estimation error. The network weights are adjusted through self-supervised learning to gradually approximate the real noise.

4. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 1, characterized in that, The adaptive multi-view semantic segmentation network includes a deep convolutional neural network; then, the step of using the adaptive multi-view semantic segmentation network to perform semantic segmentation to obtain a multi-view image set after semantic segmentation includes: Based on the similarity between images from each viewpoint, the features of each viewpoint are weighted and fused to adaptively adjust the contribution of different viewpoint images to the segmentation result, thus obtaining the features of the weighted fused viewpoints. The features of the weighted fused viewpoints of each viewpoint image are adjusted by an adaptive weighting mechanism to obtain the final global features. The adaptive weighting mechanism uses the similarity between the multi-viewpoint images and other images as weights to calculate the features of the weighted fused viewpoints. The final global features are input into a deep convolutional neural network for final semantic segmentation, resulting in a multi-view image set after semantic segmentation.

5. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 1, characterized in that, The adaptive enhancement is based on the first The distance from each pixel to the center of the target region and the semantic label information are weighted and calculated to obtain the weighting coefficient of each pixel in the target region. Based on the weighting coefficient of each pixel in the target region, the weight of each pixel can be dynamically adjusted according to the semantic category and location of the target to enhance the representation of the target features. The self-attention mechanism calculates key features by performing a weighted sum based on semantic suppression, which is performed on the weighted coefficient of each pixel in the target region and the semantic features of the corresponding pixel. The semantic suppression is achieved by calculating the contrast between the background region and the target region to suppress background interference.

6. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 1, characterized in that, The method of using similarity calculation to perform viewpoint consistency matching based on the final feature representation of the target region includes: Through from the and the Target features from each perspective are used to determine the perspective consistency loss function; The view consistency loss function optimizes target features from different viewpoints, ensuring consistency among target features with similar appearances.

7. A high-precision multi-view image recognition and detection method based on a CCD camera according to any one of claims 5 to 6, characterized in that, The feature fusion step based on the final feature representation of the target region is as follows: Based on the first The features from one perspective are weighted and fused with the feature difference measures of all other perspectives to obtain the fused target feature map.

8. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 7, characterized in that, The feature enhancement includes high-frequency texture enhancement and edge strengthening of the target feature map, amplifying the texture features of areas with missing solder, missing adhesive, and foreign matter attachment.

9. The high-precision multi-view image recognition and detection method based on a CCD camera according to claim 7, characterized in that, The bounding box regression uses the target feature map to perform bounding box regression and accurately locate the defects of each motor. Among them, smoothed L1 loss is used to optimize the positioning accuracy, and a spatial transformation regularization term is introduced to handle the positioning error caused by the characteristics of CCD device.

10. A high-precision multi-view image recognition and detection method based on a CCD camera according to claim 8 or 9, characterized in that, The classification module introduces a fuzzy adversarial loss function as a regularization term to simulate the fuzzy scene of low-resolution or noisy images, thereby improving the robustness of the classification module to fuzzy images.

Citation Information

Patent Citations

  • Image semantic segmentation method based on multi-task deep learning

    CN112950645A

  • Image semantic segmentation method and system based on stereoscopic perception scene

    CN117373019A