Wood surface defect detection method and device based on multi-perspective anomaly detection

The method integrates multi-view contrastive learning and reconstruction-based anomaly detection to enhance wood defect detection, addressing efficiency and precision issues in traditional methods by automating defect identification and localization without manual annotation.

CN119991677BActive Publication Date: 2025-07-15QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510474235.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has problems such as low efficiency, poor accuracy, high cost, and insufficient detection ability for small targets and similar texture defects in wood surface defect detection, especially in complex texture environments, and relying on a large amount of manual labeling data.

Method used

The multi-view anomaly detection method is adopted, combined with CMC multi-view comparison learning and reconstruction anomaly detection, and the encoder and decoder are built through the AlexNet variant network, and the Lab color space is used for self-supervised comparison learning, and a defect marking module with error fusion and adaptive threshold segmentation is built to achieve high-precision defect detection without manual annotation.

Benefits of technology

It realizes accurate detection and pixel-level positioning of tiny defects on the wood surface, reduces labeling costs, improves the robustness and generalization capabilities of the model in complex texture environments, and meets the efficient and high-precision needs of industrial quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991677B_ABST
    Figure CN119991677B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and specifically relates to a method and device for detecting wood surface defects based on multi-view anomaly detection. The method includes: obtaining a target image based on the original wood image; constructing a wood surface defect detection model including an encoder, a decoder, and a defect marking module, constructing positive and negative samples based on the target image, and training the encoder through contrastive learning using the positive and negative samples; using the output of the encoder as the input of the decoder, constructing a loss function through the mean square error and the structural similarity index to train the decoder, and the decoder outputs a reconstructed image; obtaining a residual map between the reconstructed image and the target image through the defect marking module to obtain a fused residual map, and marking the pixel regions exceeding the dynamic threshold as defects; using the trained wood surface defect detection model to detect the preprocessed wood image to be detected, realizing wood defect detection and pixel-level positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a wood surface defect detection method and device based on multi-viewing angle anomaly detection. Background Art

[0002] Wood is an indispensable key raw material in the fields of construction, furniture manufacturing and decoration. Its surface quality has a crucial impact on the safety performance, service life and commercial value of the product. There are many defects on the surface of wood, such as dead knots, cracks, dry scars, wormholes, etc. These defects vary in size and color. These defects will not only significantly weaken the mechanical properties of wood, but may also cause safety hazards during subsequent processing and use.

[0003] Traditional wood defect detection mainly relies on manual visual inspection, but manual inspection is inefficient and the results are inconsistent, especially when facing complex textures or tiny defects, which are prone to missed detection and misjudgment. This inefficient and high-cost detection method has been difficult to adapt to the needs of modern intelligent manufacturing. Therefore, the development of high-precision, automated wood surface defect detection technology has become a core issue that needs to be urgently addressed in the wood processing industry.

[0004] In order to improve detection efficiency, early studies have attempted to use automated techniques based on image processing, such as grayscale analysis or edge detection algorithms to locate defective areas. However, the complex diversity of natural wood textures and the variability of defect morphology pose a severe challenge to traditional algorithms. When defects are highly similar to wood grain in color, shape or direction, the algorithm often has difficulty in stably distinguishing between real defects and normal textures. In addition, factors such as changes in lighting conditions and surface reflections further weaken the robustness of traditional methods. Although some improvement schemes have introduced multispectral imaging or three-dimensional scanning technology, their high equipment costs and complex operating procedures limit the widespread application of these technologies in industrial scenarios.

[0005] With the continuous development of deep learning technology, the object detection model based on convolutional neural network has gradually become the mainstream solution. Single-stage detection algorithms represented by YOLO are widely used in industrial quality inspection systems due to their efficient real-time processing capabilities. YOLO divides the input image into grids and directly regresses the bounding box coordinates, category probability and confidence through a single neural network, thereby completing the object detection task in a single forward propagation. However, in the task of wood surface defect detection, such algorithms still have obvious limitations: first, the tiny defects on the wood surface account for a very low proportion in the image, and it is difficult for general detection models to effectively capture their features, resulting in a prominent problem of missed detection of small targets; second, the visual similarity between wood texture and defects can easily confuse the model, especially in densely textured areas, where normal textures are frequently mistakenly identified as defects; finally, the traditional detection framework relies on bounding boxes to locate defects, and the contour capture accuracy of irregular or subtle defects is insufficient, making it difficult to achieve accurate pixel-level segmentation.

[0006] Most of the current improvement schemes focus on optimizing the YOLO model structure, such as strengthening local feature extraction by introducing an attention mechanism, or building a multi-scale feature fusion network to improve the detection ability of small targets. However, these methods still have essential shortcomings: first, although the data enhancement strategy can expand the diversity of samples, it cannot fundamentally remove the coupling relationship between features and other information such as wood board color and ambient brightness; second, although the attention mechanism can focus on suspicious areas, it lacks the ability to learn normal texture features, resulting in poor generalization of the model when facing unseen wood species or texture types; finally, and most importantly, the existing technology is highly dependent on a large amount of accurately labeled defect data. The types of wood surface defects are complex and diverse, and professional quality inspectors are required to mark the defect boundaries. In addition, the natural variability of wood texture and the fuzziness of defect boundaries make manual labeling prone to subjective bias. This strong dependence on high-precision labeled data not only greatly increases the cost of manpower and time, but also forms a sharp contradiction with the rapid iteration production needs of the wood processing industry.

[0007] For example, Chinese Patent Document CN113989224A discloses a method for detecting defects in colored texture fabrics based on a generative adversarial network. A contrastive learning generative adversarial network model based on channel attention is constructed, and the dataset of colored texture fabrics with superimposed Gaussian noise is input into the contrastive learning generative adversarial network model based on channel attention for training to obtain a trained contrastive learning generative adversarial network model based on channel attention. The trained contrastive learning generative adversarial network model based on channel attention is used to reconstruct the colored texture fabric image to be detected, and the reconstructed image is output, and then detection is performed to locate the defect area. The present invention realizes the rapid detection and location of the defect area by calculating the residual between the colored texture fabric image to be detected and the corresponding reconstructed image, and then performing threshold segmentation and closing operation on the residual result.

[0008] Chinese Patent Document CN110969606A discloses a method and system for detecting texture surface defects. First, a multi-channel texture prior of the input texture image is extracted through a prior extraction step; then, under the guidance of the extracted prior, the texture background of the input image is accurately reconstructed through a texture reconstruction step, and more accurate texture features are encoded in the latent space of the texture reconstruction module to suppress the reconstruction of defects in the texture background; finally, through pixel-level adversarial learning, the texture reconstruction accuracy is further improved. When performing detection, only the difference operation between the input image and the reconstructed texture background image needs to be performed to detect the defects. The present invention has high detection accuracy for defects of different sizes and different contrasts on different texture surfaces.

[0009] In view of the above problems, the present invention proposes an innovative solution that combines multi-view contrastive learning with a reconstruction-based anomaly detection method. It should be noted that in the present invention, "multi-view" and "multi-perspective" are expressions of the same technical concept, both referring to methods of analyzing from different data dimensions or feature channels, and they have the same technical meaning in this context. The present invention draws on the core idea of the CMC multi-view contrastive learning method proposed by YongLong Tian et al., and deeply integrates multi-view contrastive learning technology with reconstruction-based anomaly detection technology, so as to achieve accurate detection of small target defects and texture-similar defects on the wood surface. Summary of the Invention

[0010] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a method for detecting wood surface defects based on multi-perspective anomaly detection.

[0011] The present invention also discloses a device loaded with a method for detecting wood surface defects based on multi-perspective anomaly detection.

[0012] The detailed technical solution of the present invention is as follows:

[0013] A wood surface defect detection method based on multi-view anomaly detection, the method comprising:

[0014] S1. Obtain the original wood image dataset and preprocess the original wood images therein to obtain target images, where the original wood images are RGB images and the target images are Lab images;

[0015] S2. Construct a wood surface defect detection model, including:

[0016] Construct an encoder of the wood surface defect detection model based on the AlexNet variant network, and construct positive and negative samples based on the target images, and use the positive and negative samples to train the encoder through contrastive learning;

[0017] Construct a decoder of the wood surface defect detection model based on the flipped form of the AlexNet variant network, freeze the parameters of the trained encoder, and use the output of the encoder as the input of the decoder, and construct a loss function through the mean square error MSE and the structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image;

[0018] Construct a defect marking module for error fusion and adaptive threshold segmentation, where the defect marking module is used to obtain the residual map between the reconstructed image and the target image to obtain a fused residual map, and determine whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, so as to mark the pixel regions that exceed the dynamic threshold as defects;

[0019] S3. Input the preprocessed wood image to be detected into the trained wood surface defect detection model, and after the encoder and decoder perform encoding-decoding operations in sequence, obtain a wood reconstructed image, and calculate the residual between the wood reconstructed image and the preprocessed wood image to be detected through the defect marking module, so as to realize wood defect detection and pixel-level positioning.

[0020] Preferably according to the present invention, in the S1, preprocessing the original wood images in the original wood image dataset specifically includes:

[0021] Crop the original wood images into several image blocks of 512×512 pixels, and adopt an overlapping strategy with a step size of 256 pixels to ensure texture continuity and completely cover the defect area;

[0022] Perform data augmentation operations on the cropped image blocks, including horizontal flipping, vertical flipping, and rotation;

[0023] Use the color module in the skimage library of Python to convert the augmented image blocks into Lab image blocks, that is, obtain the target images.

[0024] Preferably according to the present invention, in S2, constructing positive and negative samples based on the target image specifically includes: using the L view and the ab view of the same Lab image block in the target image as a positive sample pair, and using the L view and the ab view between different Lab image blocks as a negative sample pair.

[0025] Preferably according to the present invention, in S2, training the encoder through contrastive learning using the positive and negative samples specifically includes:

[0026] Extracting the representation vectors of the L view in the positive and negative samples respectively through the encoder and the representation vectors of the ab view , and calculating the similarity between the representation vectors and the representation vectors :

[0027] (1);

[0028] In formula (1): represents the similarity between the representation vectors corresponding to the L view and the ab view ; and respectively represent the feature representations obtained by extracting the L view and the ab view through two encoders and ; is a hyperparameter for controlling the similarity scale, and by adjusting the value of , the sensitivity of the similarity is controlled, and the smaller its value, the higher the sensitivity; is the exponential function;

[0029] Constructing a contrastive loss function, and by maximizing the similarity of the positive sample pair ( , ), minimizing the similarity between the negative sample pair ( , ) and ( , ), where i≠j, optimizing the contrastive loss function to optimize the training of the encoder;

[0030] Among them, the contrastive loss of anchoring the L view and contrasting the ab view is:

[0031] (2);

[0032] In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; is the similarity between the anchor point and the positive sample , and the sum of the similarities of negative samples; denotes taking the average over all possible sample pairs, including positive and negative samples;

[0033] The contrastive loss for anchoring the ab view and comparing with the L view is:

[0034] (3);

[0035] Then the contrastive loss function of the encoder is:

[0036] (4).

[0037] According to a preferred embodiment of the present invention, in S2, a loss function is constructed by mean squared error (MSE) and structural similarity index (SSIM) to train the decoder, specifically:

[0038] (11);

[0039] In formula (11): represents the input target image; represents the reconstructed image output by the decoder; represents the final loss function of the decoder; is the weight coefficient used to control the loss weights of MSE and SSIM, set to .

[0040] According to a preferred embodiment of the present invention, in S2, a residual map between the reconstructed image and the target image is obtained to obtain a fused residual map, specifically:

[0041] The MSE loss and SSIM loss between the reconstructed image and the target image are calculated respectively to obtain the corresponding residual maps, and then they are normalized and weighted and summed to obtain the fused residual map :

[0042] (12);

[0043] In formula (12): is the fused residual map; is the weight coefficient used to control the weights of the residual maps corresponding to the MSE loss and SSIM loss in the fused residual map; represents the residual map obtained by calculating the MSE loss; It represents the residual map obtained by calculating the SSIM loss; It represents the maximum value of the residual map obtained by calculating the MSE loss, It represents the maximum value of the residual map obtained by calculating the SSIM loss, which is used to normalize to the range of [0, 1].

[0044] Preferably according to the present invention, calculating a dynamic threshold is further included in S2, that is:

[0045] Perform Gaussian smoothing filtering on the fused residual map to eliminate isolated noise points;

[0046] Calculate the dynamic threshold , where is the mean value of the pixel residual values in the fused residual map ; is the standard deviation of the pixel residual values; is the threshold coefficient, which is used to dynamically adjust the detection sensitivity.

[0047] In another aspect of the present invention, a device for implementing a wood surface defect detection method based on multi-view anomaly detection is provided. The device includes:

[0048] A sample acquisition module, which is used to acquire the original wood image dataset and preprocess the original wood images therein to obtain target images. The original wood images are RGB images, and the target images are Lab images;

[0049] A model construction module, which is used for the wood surface defect detection model, including:

[0050] Construct an encoder of the wood surface defect detection model based on the AlexNet variant network, and construct positive and negative samples based on the target images, and use the positive and negative samples to train the encoder through contrast learning;

[0051] Construct a decoder of the wood surface defect detection model based on the flipped form of the AlexNet variant network, freeze the parameters of the trained encoder, and use the output of the encoder as the input of the decoder, and construct a loss function through the mean square error MSE and the structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image;

[0052] Construct a defect marking module for error fusion and adaptive threshold segmentation. The defect marking module is used to obtain the residual map between the reconstructed image and the target image to obtain a fused residual map, and judge whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, so as to mark the pixel regions that exceed the dynamic threshold as defects;

[0053] The defect detection module is used to detect the pre - processed wood image to be detected by using the trained wood surface defect detection model. After the encoder and decoder perform the encoding - decoding operations in sequence, a wood reconstruction image is obtained. The defect marking module calculates the residual between the wood reconstruction image and the pre - processed wood image to be detected, so as to realize wood defect detection and pixel - level positioning.

[0054] In another aspect of the present invention, an electronic device is further provided, including:

[0055] At least one processor; and

[0056] A memory, the memory stores instructions, when the instructions are executed by the at least one processor, the at least one processor is made to execute the wood surface defect detection method based on multi - perspective anomaly detection as described above.

[0057] In another aspect of the present invention, a machine - readable storage medium is further provided, which stores executable instructions, and when the instructions are executed, the machine is made to execute the wood surface defect detection method based on multi - perspective anomaly detection as described above.

[0058] Compared with the prior art, the beneficial effects of the present invention are:

[0059] (1) Fully self - supervised training reduces the annotation cost. Innovatively integrating multi - view self - supervised contrast learning and generative reconstruction network, the model training can be completed without any manual annotation, fundamentally solving the industry pain points of diverse wood defect types, strong annotation subjectivity, and high cost, and providing a technical basis for rapid cross - species migration.

[0060] (2) The ability to detect tiny defects is significantly enhanced. Through multi - view contrast learning CMC, the essential features of wood texture in the Lab color space are extracted, accurately distinguishing the subtle differences between tiny defects and normal textures, and overcoming the problem of high missed detection rate of small targets by traditional object detection models.

[0061] (3) Pixel - level high - precision defect positioning. Abandoning the traditional anchor box mechanism, adopting a residual map fusion strategy driven by reconstruction error, directly marking abnormal pixels through adaptive threshold segmentation, breaking through the accuracy limit of bounding box positioning, and being able to accurately capture irregular defects, meeting the high - precision quality inspection standards.

[0062] (4) Strong robustness under complex texture interference. Based on multi - channel color space contrast decoupling technology, separating the coupling of features and interference information, and combining bilinear interpolation to suppress edge artifacts, significantly reducing the false detection rate in scenes with dense textures or uneven illumination, and the anti - interference performance is better than the improved scheme based on the attention mechanism.

[0063] (5)High adaptability design for industrial scenarios. The encoder-decoder symmetric architecture supports modular deployment, and the pre-trained feature extraction network is compatible with different imaging conditions, such as reflection and chromatic aberration; the image block detection strategy combines a weighted fusion mechanism to meet the balance requirements of stability and efficiency in the industrial assembly line while ensuring full-image coverage. Brief Description of the Drawings

[0064] Figure 1 is a flowchart of the wood surface defect detection method based on multi-view anomaly detection according to the present invention.

[0065] Figure 2 is a schematic diagram of the original wood image obtained in Embodiment 1 of the present invention.

[0066] Figure 3 is a schematic diagram of the operation of data augmentation on the cropped original wood image in Embodiment 1 of the present invention.

[0067] Figure 4 is a comparison diagram of the original wood image in RGB color mode and the wood image in Lab color mode in Embodiment 1 of the present invention.

[0068] Figure 5 is a network structure diagram of the wood surface defect detection model in Embodiment 1 of the present invention.

[0069] Figure 6 is a schematic diagram of the positive and negative samples of the contrast learning constructed in Embodiment 1 of the present invention.

[0070] Figure 7 is a network structure diagram of the encoder of the wood surface defect detection model in Embodiment 1 of the present invention.

[0071] Figure 8 is a network structure diagram of the decoder of the wood surface defect detection model in Embodiment 1 of the present invention.

[0072] Figure 9 is a flowchart of the wood surface defect detection model performing defect detection in Embodiment 1 of the present invention. Detailed Description of the Invention

[0073] The following further describes the present disclosure in conjunction with the drawings and embodiments.

[0074] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0075] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0076] In the case of no conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0077] Aiming at the deficiencies of the prior art, the present invention provides a wood surface defect detection method based on multi-view anomaly detection, which combines the CMC multi-view contrast learning method with the reconstruction method in anomaly detection, constructs a complete network architecture, and applies it to the wood surface defect detection task.

[0078] The CMC method can be optimized between multiple views through self-supervised contrast learning, and the model is trained by optimizing the contrast loss to extract the essential features of the data. The principle of the anomaly detection method based on reconstruction is that the model first extracts the features of normal images, then reconstructs the original images, and trains the reconstruction ability of the model by optimizing the reconstruction loss. When the model reconstructs abnormal images, there will be a large deviation between the reconstruction results of the abnormal regions and the original images, and the positions of the abnormal regions can be marked by comparing with the original images. This innovation not only breaks through the limitations of traditional defect detection frameworks, but also significantly improves the accuracy of defect detection.

[0079] Specifically, contrastive multiview coding (CMC) is a method in contrastive learning. By performing contrastive learning between multiple views, it attempts to extract the representation that best conforms to the essence of the features. In coding theory, a basic idea is to learn a compressed representation that can reconstruct the original data, and this idea is reflected in autoencoders, generative models, and contrastive learning in representation learning. Among them, both autoencoders and generative models are committed to representing data points or distributions in an as lossless as possible way. However, in the present invention, lossless representation is not the most suitable feature extraction method. Because such methods will also save other information such as wood color and environmental brightness into the compressed representation, thus interfering with the representation of texture and defect features, which obviously does not conform to the design concept of this method. The CMC method retains useful information and eliminates other irrelevant information. In the design of the present invention, the encoder performs contrastive learning in the Lab color space, which can not only learn a powerful feature representation, but also improve the anti-interference ability against redundant information.

[0080] In unsupervised anomaly detection methods, compared with feature embedding methods, reconstruction methods perform better in terms of detection accuracy. Feature embedding methods usually learn the low-dimensional feature representations of normal images and try to make the feature representations of all normal images concentrated. The feature representations of abnormal images will deviate significantly from the feature distribution of normal images. In the testing stage, by calculating the deviation between the image to be tested and the normal distribution, image-level anomaly detection can be achieved. Although this method can calculate the abnormal regions with the help of feature representations and improve the detection accuracy to a certain extent, its calculation results may deviate from the actual situation, and the accuracy is difficult to meet the requirements of the present invention.

[0081] Therefore, the present invention adopts a reconstruction method to train the network to learn the ability to reconstruct normal images. During testing, the reconstructed image is compared with the original image to calculate the residual map. If the error values of the pixels in the residual map are higher than the set threshold, these pixel positions are marked as abnormal regions. In the present invention, the design of the reconstruction method can not only detect the defects on the wood surface, but also achieve pixel-level defect marking, making the detection results more accurate.

[0082] The following further describes the wood surface defect detection method and device based on multi-view anomaly detection of the present invention with specific embodiments.

[0083] Embodiment 1

[0084] Refer Figure 1 , this embodiment provides a wood surface defect detection method based on multi-view anomaly detection, and the method includes:

[0085] S1. Obtain the original wood image dataset and preprocess the original wood images therein to obtain target images. The original wood images are RGB images, and the target images are Lab images.

[0086] In this embodiment, the original wood image dataset used comes from a self-built dataset, and all wood defect pictures are taken by industrial cameras, and the pixel resolution of the pictures is mostly 1960×1080. The dataset collected about 12,200 images in total, and the wood colors, textures, and lighting brightness in each picture are different. Specifically, as Figure 2 shown, the original wood images therein are all images in RGB color mode.

[0087] The ultimate goal of the method in this embodiment is to achieve the detection and precise marking of small target defects and texture-similar defects on the wood surface, rather than simply determining whether there are abnormalities in the image. Therefore, when training the model, more refined image data should be used. In order to effectively extract the essential features of the texture and accurately distinguish normal images from defective images even on a small dataset, the method in this embodiment performs a series of preprocessing operations on the collected original wood images. In this embodiment, the preprocessing operations performed on the original wood images include image cropping, data augmentation, and color conversion, etc.

[0088] Specifically, image cropping includes cropping a complete original wood image into several image patches.

[0089] Although a complete original wood image contains the normal texture of the wood, it also contains many defects. If the original wood image is directly used for contrast learning, it is difficult to train an encoder that is sensitive to texture and defect features. Therefore, before contrast learning, the original wood image is first cropped into several image patches of 512×512 pixels, and an overlapping strategy with a step size of 256 pixels is adopted to ensure texture continuity and completely cover the defective area.

[0090] Among these cropped image patches, some only contain normal wood texture, while some contain a small number of defects, which interfere with the normal texture, resulting in a significant difference between defective image patches and normal image patches. The method in this embodiment mixes the image patches that only contain normal texture with the image patches that contain a small number of defects for contrast learning. This can not only improve the effect of contrast learning but also make a significant difference between the representation of abnormal images and normal images. In this way, the problem of false detection caused by the confusion between normal images and abnormal images can be effectively avoided in the test stage.

[0091] To achieve better results on a small dataset, the method in this embodiment performs data augmentation operations on the cropped image patches. These operations mainly include horizontal flipping, vertical flipping, and rotation, as shown specifically Figure 3 as follows.

[0092] Based on Figure 2 as shown, the work of taking wood images in an industrial environment is interfered by various factors, such as shooting angle, board angle, light intensity, and board color. These factors directly affect the quality of the dataset. Therefore, the method in this embodiment adopts a method of multi-view contrast learning in the Lab color space.

[0093] On this basis, the method of this embodiment uses the color module in the skimage library of Python to directly convert the RGB image patch into a Lab image patch. In the Lab color space, the L channel represents lightness, indicating the brightness or darkness of the color; the a channel and the b channel respectively represent the two axes of the color, the a channel represents the green - red axis, and the b channel represents the blue - yellow axis. Performing multi - view contrast learning in this color space can not only learn the features of the image but also decouple the features from other irrelevant information, thereby reducing the interference of these factors on feature extraction and enhancing the robustness of the model under different conditions.

[0094] The comparison of the same wood image patch in the RGB color mode and the Lab color mode is as Figure 4 shown.

[0095] After the above operations, the pre - processed target image is obtained, and this target image is the Lab image patch.

[0096] S2. Construct a wood surface defect detection model, including:

[0097] Construct an encoder of the wood surface defect detection model based on the AlexNet variant network, construct positive and negative samples based on the target image, and train the encoder through contrast learning using the positive and negative samples;

[0098] Construct a decoder of the wood surface defect detection model based on the flipped form of the AlexNet variant network, freeze the parameters of the trained encoder, and use the output of the encoder as the input of the decoder. Construct a loss function through the mean square error (MSE) and the structural similarity index (SSIM) to train the decoder, and the decoder outputs a reconstructed image;

[0099] Construct a defect marking module for error fusion and adaptive threshold segmentation. The defect marking module is used to obtain the residual map between the reconstructed image and the target image to obtain a fused residual map, and determine whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, and mark the pixel regions that exceed the dynamic threshold as defects.

[0100] In this embodiment, in order to efficiently extract the essential features of the wood surface image under wooden boards of different colors and different lighting conditions, the method of this embodiment adopts a variant network of AlexNet adapted to the CMC idea, and uses the image patches converted to the Lab color mode as input data for pre-training. After the pre-training is completed, the network parameters are frozen and used as the encoder module. Subsequently, referring to the principle of the reconstruction method in anomaly detection, a decoder module symmetric to the encoder structure is constructed to reconstruct the original image, and the ability of this module to reconstruct normal images is trained. In the final stage, a technique that combines the residual map and adaptive threshold segmentation is adopted to calculate the combined residual map between the reconstructed image and the original image, and the defect positions are marked accordingly, so as to achieve the defect detection of the wood surface.

[0101] The complete network structure of the wood surface defect detection model is as Figure 5 shown, where the encoder structure part includes convolution operation (Conv), batch normalization (BN), activation function (ReLU), downsampling operation (Pool), and fully connected layer (FC), and the unique operations in the decoder structure are reconstructed feature map, upsampling operation (UnPool), and transposed convolution operation (DeConv).

[0102] Specifically, first build the encoder.

[0103] In this embodiment, an encoder of a wood surface defect detection model is constructed based on a variant network of AlexNet, and then positive and negative samples are constructed based on the target image, and the encoder is trained through contrastive learning using the positive and negative samples.

[0104] The target image is a Lab image patch. In order to perform multi-view contrastive learning in the Lab color mode, it is necessary to first define the views used and the positive and negative samples. The method of this embodiment draws on the idea of the CMC method: in the Lab color space, the L channel represents brightness, and the a channel and b channel represent colors. Therefore, the method of this embodiment uses the L channel as the L view, and combines the a channel and b channel as another view, called the ab view.

[0105] As Figure 6 shown, the definition of the positive and negative samples in this embodiment is as follows: the L view and ab view of the same image patch are used as a positive sample pair, and the L view and ab view between different image patches are used as a negative sample pair. Through this multi-view contrastive learning method, the encoder can be made to ignore the interference of brightness and color differences and focus on extracting the wood board texture features and defect features.

[0106] The encoder network structure selected in this embodiment is an AlexNet variant network adapted to the CMC concept. The network is divided into two parts in the channel dimension of the feature map: one part is used to train the L view and the other part is used to train the ab view. Functionally, these can be regarded as two independent encoders, but in terms of network structure, they still belong to the same network, except that the parameters of the two parts are independent. The overall network architecture is still 5 convolutional layers and 3 fully connected layers. The convolutional layers include convolution and downsampling operations. This design not only realizes the functions of two encoders, but also does not significantly increase the number of parameters.

[0107] Since the two parts of the network perform encoding operations independently, when a complete Lab image block is input, the encoder extracts the representation vectors of the L view and the ab view respectively. and The contrast loss function maximizes the positive sample pair ( , ), while minimizing the similarity of negative sample pairs ( , )and( , ) similarity (where i≠j), continuously optimizing the contrast loss function, the model gradually brings the representation vectors between positive samples closer and pushes the representation vectors between negative samples farther away, thereby establishing a strong ability to distinguish normal textures from defect features. The network structure of the encoder is as follows Figure 7 As shown, it includes 5 convolutional layers (Conv) and 3 fully connected layers (FC).

[0108] First, calculate the similarity between the two representation vectors :

[0109] (1);

[0110] In formula (1): Indicates L view and ab view The similarity between corresponding representation vectors; and Represents L view and ab view Through two encoders and The extracted feature representation; is a hyperparameter that controls the similarity scale and can be adjusted The value of controls the sensitivity of similarity. The smaller the value, the higher the sensitivity. It is an exponential function. The exponent is taken to ensure that the output result is a positive value and can quickly amplify the difference according to the size of the similarity.

[0111] After obtaining the similarity between the representation vectors, the next step is to calculate the contrast loss of positive and negative samples and supervise the encoder training by optimizing the contrast loss; the calculation of the contrast loss is as follows:

[0112] (2);

[0113] In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; is the similarity including the anchor and the positive sample as well as negative samples The sum of similarities is normalized for the similarities of all sample pairs; represents taking the average over all possible sample pairs, including positive and negative samples, to ensure the optimization of the loss function.

[0114] The above formula (2) can be understood as anchoring the L view and calculating the contrast with the ab view. To make the effect better, the contrast loss of anchoring the ab view and comparing with the L view can be calculated again:

[0115] (3).

[0116] The above two formulas only show the loss calculation of one anchor sample. In actual application, each sample will be taken as the anchor in turn, calculate the two losses respectively, and find the average value. Finally, the sum of the two is used as the final contrast loss:

[0117] (4).

[0118] When the contrast loss value of the validation set data continuously appears below the set threshold, the training of the encoder part is completed, the encoder parameters are frozen, and finally an operation is added to concatenate the finally obtained representation vectors and into a representation vector to prepare for training the decoder.

[0119] Next, build the decoder.

[0120] Traditional autoencoders achieve the goal by compressing the input features through the encoder and then restoring the original image by the decoder. However, its encoder tries to retain all information as much as possible during the compression process, which makes the feature extraction and reconstruction results vulnerable to interference from irrelevant factors. Therefore, the method of this embodiment draws on the idea of CMC to construct the encoder to extract features, making the extracted features closer to the essence.

[0121] When designing the decoder part in this embodiment, its network structure needs to be adapted to the encoder. The flipped form of the AlexNet variant network is selected, which is symmetric with the encoder structure. Specifically, all the convolutional and downsampling operations in the encoder are replaced with transposed convolutional and upsampling operations. Each layer performs upsampling, halving the number of channels, and non-linear activation processing in sequence, and finally outputs a reconstructed image with the same size as the original image block. This design makes the reconstructed image simple and efficient. In the pre-training stage before this, the encoder has been able to extract the essential texture representation from the Lab image block, and the decoder only needs to focus on gradually restoring the original image through transposed convolutional and upsampling operations. The network structure of the decoder is as Figure 8 shown, which includes a fully connected layer (FC), a reconstructed feature layer, and 5 transposed convolutional layers (DeConv).

[0122] The transposed convolution learns the parameterized upsampling process through backpropagation. Its essence is to use a learnable convolutional kernel to slide on the low-resolution feature map and fill in the gaps, thereby gradually restoring the high-resolution features. The upsampling uses the bilinear interpolation method, which determines the new pixel value by calculating the weighted average of adjacent pixels, and the weights are determined by the relative distances between the target pixel and the surrounding 4 original pixels. Bilinear interpolation can achieve a balance between smoothness and computational efficiency, effectively suppressing checkerboard artifacts and ensuring the continuity of wood texture and the natural transition of defect boundaries.

[0123] The anomaly detection method based on reconstruction usually requires a loss function. This loss function is not only used in the training stage to guide the network to learn the ability to reconstruct the original image, but also used in the testing stage to calculate the residual map between the reconstructed image and the original image, so as to mark the anomaly positions. Commonly used reconstruction loss metrics include mean squared error MSE, cross-entropy loss, and structural similarity index SSIM. Among them, the cross-entropy loss is mainly used to process binary or probability distribution data and measure the difference between two probability distributions, so it is not suitable for use in the method of this embodiment. The mean squared error MSE is achieved by calculating the error for each sample and then taking the average. Specifically in image processing, it compares the errors of two images pixel by pixel and is suitable for application in the method of this embodiment. The calculation formula of the mean squared error MSE is:

[0124] (5);

[0125] In formula (5): represents the original image, that is, the target image; represents the reconstructed image; represents the height and width of the image size; represents the original image at the pixel position of represents the reconstructed image At this pixel position , calculate the difference in pixel intensity values at each corresponding position of the two images, take the square of the difference and sum it up, and finally find the average.

[0126] The pixel-by-pixel calculation makes the mean square error (MSE) very sensitive to pixel value differences, and it is easy to identify the location of anomalies. However, due to its oversensitivity and the neglect of the structural information of the image, when there is a slight offset between the reconstructed image and the original image, these errors caused by alignment problems will also be treated as anomalies by the mean square error (MSE). In addition, this method is difficult to detect defective areas that have changed visually but whose pixel values remain roughly the same. Bergmann et al. pointed out that methods such as the mean square error (MSE) assume that adjacent pixels are independent, but this assumption does not fully conform to the structural characteristics of actual images, resulting in their poor performance in detecting structural differences. The structural similarity index (SSIM) can improve this problem well.

[0127] The structural similarity index SSIM was proposed by Wang et al. in 2004. It aims to better evaluate the visual quality of an image by integrating these factors. Compared with traditional pixel-level error metrics, the structural similarity index SSIM pays more attention to the structural characteristics of the image.

[0128] The structural similarity index SSIM defines a distance metric to measure the similarity between two K×K image patches p and q, which takes into account the brightness similarity. (p,q), contrast similarity (p,q) and structural similarity (p, q), which is applied to the method of this embodiment to calculate , and , the formula is as follows:

[0129] (6);

[0130] In formula (6): , , ∈R is a user-defined weight factor, which is used to control the influence of brightness, contrast and structure, and is usually set to 1.

[0131] Among them, brightness similarity for:

[0132] (7);

[0133] In formula (7): , Respectively represent the original image and reconstructed image The mean value of is a small constant, usually set to 0.01, which is used to prevent the denominator from being zero.

[0134] Contrast similarity is:

[0135] (8);

[0136] In formula (8): , respectively represent the standard deviation of the original image and the reconstructed image , is a small constant, usually set to 0.03.

[0137] Structural similarity is:

[0138] (9);

[0139] In formula (9): represents the covariance between the original image and the reconstructed image .

[0140] Substituting the above three calculation formulas into the calculation formula of the structural similarity index SSIM, we can get:

[0141] (10).

[0142] When the original image and the reconstructed image are exactly the same, = 1; on the contrary, when the original image and the reconstructed image are completely different in brightness, contrast and structure, = -1; therefore, the value range of SSIM is [-1, 1].

[0143] However, SSIM is calculated based on local windows and is less sensitive to the defects of the global distribution. Moreover, there is an overall shift in the brightness distribution between the reconstructed image and the original image, and SSIM may not be able to fully reflect such problems.

[0144] Therefore, the method of this embodiment uses the weighted sum of the mean square error MSE and the structural similarity index SSIM as the final loss function to train the decoder, that is:

[0145] (11);

[0146] In formula (11): represents the final loss function of the decoder; is the weight coefficient, used to control the loss weights of the mean square error (MSE) and the structural similarity index (SSIM), and is set here to .

[0147] In this way, a balance can be achieved between pixel-level accuracy and structural consistency, thereby improving the comprehensive performance of the decoder.

[0148] Finally, a defect marking module that combines error fusion and adaptive threshold segmentation is constructed to calculate the residual map and locate defects.

[0149] After the decoder part is trained, the network can complete the entire process from the input image to the reconstructed original image. After the decoder outputs the reconstructed image, the residual map between the reconstructed image and the input target image needs to be calculated to highlight the defective areas, so as to facilitate marking the defect positions.

[0150] The method of this embodiment designs a defect marking module that combines error fusion and adaptive threshold segmentation. Specifically, the MSE and the SSIM are used to calculate the residual maps of the reconstructed image and the target image respectively, and then the two calculated residual maps are fused into one residual map. Finally, with the help of the adaptive threshold segmentation technology, defect localization can be achieved more accurately.

[0151] That is, first calculate the MSE loss and the SSIM loss between the reconstructed image and the target image respectively, calculate the residual maps, and then perform normalization processing and weighted summation to obtain the fused residual map :

[0152] (12);

[0153] In formula (12): is the fused residual map; is the weight coefficient, used to control the weights of the two residual maps in the fused residual map; represents the residual map obtained by calculating the MSE loss; represents the residual map obtained by calculating the SSIM loss; represents the maximum value of the residual map obtained by calculating the MSE loss, represents the maximum value of the residual map obtained by calculating the SSIM loss, used for normalization to the range of [0, 1].

[0154] Finally, in order to convert the pixel residual values into a binary defect mask, this embodiment adopts an adaptive threshold algorithm based on image statistics. Specifically, first perform Gaussian smoothing filtering on the fused residual map to eliminate isolated noise points; subsequently, calculate the dynamic threshold , where is The mean value of the pixel residual values is the standard deviation is the threshold coefficient, which is used to dynamically adjust the detection sensitivity The larger the value of the larger the calculated threshold is, and the stricter the detection is. Finally, the pixel regions with pixel residual values exceeding the threshold are marked as defects

[0155] Based on the above, the construction and training process of the wood surface defect detection model is completed, and then it is deployed to the actual wood surface defect detection application

[0156] S3. Input the preprocessed wood image to be detected into the trained wood surface defect detection model. After the encoder and decoder perform the encoding and decoding operations in sequence, a wood reconstruction image is obtained. The defect marking module calculates the residual between the preprocessed wood image to be detected and the wood reconstruction image to achieve wood defect detection and pixel-level positioning

[0157] This step is the application of the trained wood surface defect detection model deployed to the actual wood surface defect detection. For specific reference Figure 9 which shows a complete wood surface defect detection process

[0158] That is, first, based on the preprocessing operation in step S1, the wood image to be detected is preprocessed and converted into a Lab image, and then input into the trained wood surface defect detection model. After being encoded by its encoder, the corresponding feature vector is output to the decoder, and the decoder performs the decoding operation and outputs the wood reconstruction image. Finally, the defect marking module calculates the residual map between the preprocessed wood image to be detected and the wood reconstruction image, and determines whether there are pixel residual values exceeding the dynamic threshold. If so, the pixel regions with pixel residual values exceeding the threshold are marked as defects, thereby realizing wood defect detection and pixel-level positioning

[0159] In summary, the wood surface defect detection method based on multi-view anomaly detection of the present invention constructs an encoder module based on the CMC idea for extracting image features; at the same time, a decoder module is built according to the principle of the decoder for reconstructing the original image. By comparing the reconstructed image with the original image, the defect positions are marked, thus efficiently completing the wood defect detection task

[0160] The innovative design of the present invention organically combines the excellent feature extraction ability in the CMC method with the pixel-level anomaly detection ability of the reconstruction anomaly detection method. By decoupling the essential features of wood texture from the interference information and combining pixel-level residual analysis, the present invention realizes accurate defect detection and positioning without manual annotation, fundamentally avoiding the dependence on a large amount of annotated data in traditional methods. In addition, the encoder constructed based on the CMC idea focuses on extracting the essential features of normal texture, enabling the model to have good generalization ability when facing unseen wood species, thus providing a low-cost and highly generalizable defect detection solution for the wood processing industry.

[0161] Embodiment 2

[0162] This embodiment provides a device for implementing a wood surface defect detection method based on multi-view anomaly detection. The device includes:

[0163] A sample acquisition module, configured to acquire an original wood image dataset and preprocess the original wood images therein to obtain target images. The original wood images are RGB images, and the target images are Lab images;

[0164] A model construction module, configured to construct a wood surface defect detection model, including:

[0165] Construct an encoder of the wood surface defect detection model based on the AlexNet variant network, and construct positive and negative samples based on the target images. Use the positive and negative samples to train the encoder through contrastive learning;

[0166] Construct a decoder of the wood surface defect detection model based on the flipped form of the AlexNet variant network. Freeze the parameters of the trained encoder, and use the output of the encoder as the input of the decoder. Construct a loss function through the mean square error (MSE) and the structural similarity index (SSIM) to train the decoder, and the decoder outputs a reconstructed image;

[0167] Construct a defect marking module for error fusion and adaptive threshold segmentation. The defect marking module is configured to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and determine whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, so as to mark the pixel regions that exceed the dynamic threshold as defects;

[0168] A defect detection module, configured to use the trained wood surface defect detection model to detect the preprocessed wood image to be detected. After the encoder and the decoder perform encoding and decoding operations in sequence, a wood reconstructed image is obtained. Calculate the residual between the preprocessed wood image to be detected and the wood reconstructed image through the defect marking module to achieve wood defect detection and pixel-level positioning.

[0169] Example 3

[0170] This embodiment also provides an electronic device, including:

[0171] At least one processor; and

[0172] A memory storing instructions, which when executed by the at least one processor, cause the at least one processor to execute the wood surface defect detection method based on multi-view anomaly detection as described above.

[0173] In this embodiment, the electronic device may include but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and the like.

[0174] Example 4

[0175] This embodiment also provides a machine-readable storage medium storing executable instructions, which when executed cause the machine to execute the wood surface defect detection method based on multi-view anomaly detection as described above.

[0176] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.

[0177] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0178] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer or a cloud via a communication network.

[0179] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0180] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0181] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0183] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A wood surface defect detection method based on multi-view anomaly detection, characterized in that, The method includes: S1. Obtain the original wood image dataset and preprocess the original wood images therein to obtain target images, where the original wood images are RGB images and the target images are Lab images; S2. Construct a wood surface defect detection model, including: Construct an encoder of the wood surface defect detection model based on the AlexNet variant network, and construct positive and negative samples based on the target images. Use the positive and negative samples to train the encoder through contrastive learning; Construct a decoder of the wood surface defect detection model based on the flipped form of the AlexNet variant network. Freeze the parameters of the trained encoder, and use the output of the encoder as the input of the decoder. Construct a loss function through the mean square error (MSE) and the structural similarity index (SSIM) to train the decoder, and the decoder outputs a reconstructed image; Construct a defect marking module for error fusion and adaptive threshold segmentation. The defect marking module is used to obtain the residual map between the reconstructed image and the target image to obtain a fused residual map, and judge whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, so as to mark the pixel regions that exceed the dynamic threshold as defects; S3. Input the preprocessed wood image to be detected into the trained wood surface defect detection model. After the encoder and decoder perform the encoding-decoding operations in sequence, obtain the wood reconstructed image, and calculate the residual between the wood reconstructed image and the preprocessed wood image to be detected through the defect marking module, so as to realize wood defect detection and pixel-level positioning.

2. The method for detecting wood surface defects based on multi-view anomaly detection according to claim 1, wherein, In S1, the preprocessing of the original wood images in the original wood image dataset specifically includes: Crop the original wood images into several image blocks of 512×512 pixels, and adopt an overlapping strategy with a step size of 256 pixels to ensure texture continuity and completely cover the defect area; Perform data augmentation operations on the cropped image blocks, including horizontal flipping, vertical flipping, and rotation; Use the color module in the skimage library of Python to convert the augmented image blocks into Lab image blocks, that is, obtain the target images.

3. The wood surface defect detection method based on multi-perspective anomaly detection according to claim 1, characterized in that, In S2, constructing positive and negative samples based on the target images specifically means: using the L view and ab view of the same Lab image block in the target image as a positive sample pair, and using the L view and ab view between different Lab image blocks as a negative sample pair.

4. The method for detecting wood surface defects based on multi-view anomaly detection according to claim 3, wherein In S2, training the encoder through contrastive learning using the positive and negative samples specifically includes: Extract the feature vectors of the L view in the positive and negative samples respectively through the encoder and the feature vectors of the ab view , and calculate the similarity between the feature vectors and the feature vectors : (1); In formula (1): represents the similarity between the characterization vectors corresponding to the L view and the ab view ; and respectively represent the feature representations obtained by extracting the L view and the ab view through two encoders and ; is a hyperparameter for controlling the similarity scale, and by adjusting the value of , the sensitivity of the similarity is controlled. The smaller its value, the higher the sensitivity; is an exponential function; Construct a contrastive loss function and optimize the training of the encoder by maximizing the similarity of positive sample pairs ( , ) and minimizing the similarity between negative sample pairs ( , ) and ( , ), where i ≠ j, to optimize the contrastive loss function. Among them, the contrast loss for comparing the ab view while anchoring the L view is as follows: (2); In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; is the similarity between the anchor point and the positive sample , and negative samples is the sum of similarities; represents taking the average over all possible sample pairs, including positive and negative samples; Anchoring the ab view, the contrast loss compared with the L view is as follows: (3); Then the contrast loss function of the encoder is: (4)。 5. The wood surface defect detection method based on multi-view anomaly detection according to claim 1, characterized in that In S2, constructing a loss function through the mean square error (MSE) and the structural similarity index (SSIM) to train the decoder specifically means: (11); In formula (11): represents the input target image; represents the reconstructed image output by the decoder; represents the final loss function of the decoder; is the weight coefficient, used to control the loss weights of the mean square error MSE and the structural similarity index SSIM, and is set to .

6. The wood surface defect detection method based on multi-perspective anomaly detection according to claim 1, characterized in that In S2, obtaining the residual map between the reconstructed image and the target image to obtain a fused residual map specifically means: Calculate the MSE loss and SSIM loss between the reconstructed image and the target image respectively, obtain the corresponding residual maps respectively, and then perform normalization processing and weighted summation on them to obtain the fused residual map : (12); In formula (12): is the fused residual map; is the weight coefficient, used to control the weights of the residual maps corresponding to the MSE loss and the SSIM loss in the fused residual map; represents the residual map obtained by calculating the MSE loss; represents the residual map obtained by calculating the SSIM loss; represents the maximum value of the residual map obtained by calculating the MSE loss, represents the maximum value of the residual map obtained by calculating the SSIM loss, used for normalization to the range of [0, 1].

7. The wood surface defect detection method based on multi-view anomaly detection according to claim 6, characterized in that In S2, it also includes calculating the dynamic threshold, that is: Perform Gaussian smoothing filtering on the fusion residual map to eliminate isolated noise points; Calculate the dynamic threshold , where is the mean of the pixel residual values in the fused residual map ; is the standard deviation of the pixel residual values; is the threshold coefficient for dynamically adjusting the detection sensitivity.

8. An apparatus for implementing a wood surface defect detection method based on multi-view anomaly detection, characterized in that, The device includes: A sample acquisition module, which is used to obtain the original wood image dataset and preprocess the original wood images therein to obtain target images, where the original wood images are RGB images and the target images are Lab images; A model construction module for a wood surface defect detection model, including: Constructing an encoder of the wood surface defect detection model based on a variant network of AlexNet, constructing positive and negative samples based on the target image, and training the encoder through contrastive learning using the positive and negative samples; Constructing a decoder of the wood surface defect detection model based on a flipped form of the AlexNet variant network, freezing the parameters of the trained encoder, and using the output of the encoder as the input of the decoder. Constructing a loss function through mean square error (MSE) and structural similarity index (SSIM) to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, where the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and determine whether there are pixel residual values in the fused residual map that exceed the dynamic threshold, and mark the pixel regions that exceed the dynamic threshold as defects; A defect detection module for using the trained wood surface defect detection model to detect a preprocessed wood image to be detected. After encoding-decoding operations are sequentially performed by the encoder and the decoder, a wood reconstructed image is obtained, and the residual between the wood reconstructed image and the preprocessed wood image to be detected is calculated through the defect marking module to achieve wood defect detection and pixel-level positioning.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory that stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the wood surface defect detection method based on multi-view anomaly detection according to any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that, An executable instruction is stored on the machine-readable storage medium, and when the instruction is executed, the machine executes the wood surface defect detection method based on multi-view anomaly detection according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Texture surface defect detection method and system

    CN110969606A

  • Color texture fabric defect detection method based on generative adversarial network

    CN113989224A

  • Multi-view generation method based on comparative learning

    CN112598775A

  • Super-pixel-based flexible IC substrate color change defect detection method and device

    CN112991302A