Wood surface defect detection method and device based on multi-view anomaly detection

By combining CMC multi-view comparison learning and reconstruction anomaly detection technology, a wood surface defect detection model was constructed, which solved the problems of low efficiency and low accuracy of wood surface defect detection in the existing technology, and achieved high-precision and low-cost defect detection and positioning.

CN119991677AActive Publication Date: 2025-05-13QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2

Patent Information

Application Number
CN202510474235.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency and low accuracy in the detection of wood surface defects, especially when facing complex textures or minor defects, missed detection and misjudgment are prone to occur.

Method used

Using a method based on multi-view anomaly detection, the CMC multi-view comparison learning method is combined with the reconstruction anomaly detection technology to construct a wood surface defect detection model. The model extracts wood texture and defect features through a comparative learning and reconstruction process of the encoder and decoder, and achieves defect detection and pixel-level positioning through error fusion and adaptive threshold segmentation.

Benefits of technology

The detection accuracy of small target defects on the surface of wood and similar texture defects is significantly improved, the dependence on manual labeled data is reduced, the robustness and generalization ability of the model are improved, and the efficiency and high accuracy requirements of industrial quality inspection are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991677A_ABST
    Figure CN119991677A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a wood surface defect detection method and device based on multi-view anomaly detection, and the method comprises the steps: obtaining a target image based on an original wood image; constructing a wood surface defect detection model comprising an encoder, a decoder and a defect marking module, constructing positive and negative samples based on the target image, and training the encoder by using the positive and negative samples through contrast learning; the output of the encoder is used as the input of a decoder, a loss function is constructed through a mean square error and a structural similarity index to train the decoder, and the decoder outputs a reconstructed image; obtaining a residual image between the reconstructed image and the target image through a defect marking module to obtain a fused residual image, and marking a pixel region exceeding a dynamic threshold as a defect; and detecting the preprocessed to-be-detected wood image by using the trained wood surface defect detection model to realize wood defect detection and pixel-level positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a wood surface defect detection method and device based on multi-viewing angle anomaly detection. Background Art

[0002] Wood is an indispensable key raw material in the fields of construction, furniture manufacturing and decoration. Its surface quality has a crucial impact on the safety performance, service life and commercial value of the product. There are many defects on the surface of wood, such as dead knots, cracks, dry scars, wormholes, etc. These defects vary in size and color. These defects will not only significantly weaken the mechanical properties of wood, but may also cause safety hazards during subsequent processing and use.

[0003] Traditional wood defect detection mainly relies on manual visual inspection, but manual inspection is inefficient and the results are inconsistent, especially when facing complex textures or tiny defects, which are prone to missed detection and misjudgment. This inefficient and high-cost detection method has been difficult to adapt to the needs of modern intelligent manufacturing. Therefore, the development of high-precision, automated wood surface defect detection technology has become a core issue that needs to be urgently addressed in the wood processing industry.

[0004] In order to improve detection efficiency, early studies have attempted to use automated techniques based on image processing, such as grayscale analysis or edge detection algorithms to locate defective areas. However, the complex diversity of natural wood textures and the variability of defect morphology pose a severe challenge to traditional algorithms. When defects are highly similar to wood grain in color, shape or direction, the algorithm often has difficulty in stably distinguishing between real defects and normal textures. In addition, factors such as changes in lighting conditions and surface reflections further weaken the robustness of traditional methods. Although some improvement schemes have introduced multispectral imaging or three-dimensional scanning technology, their high equipment costs and complex operating procedures limit the widespread application of these technologies in industrial scenarios.

[0005] With the continuous development of deep learning technology, the object detection model based on convolutional neural network has gradually become the mainstream solution. Single-stage detection algorithms represented by YOLO are widely used in industrial quality inspection systems due to their efficient real-time processing capabilities. YOLO divides the input image into grids and directly regresses the bounding box coordinates, category probability and confidence through a single neural network, thereby completing the object detection task in a single forward propagation. However, in the task of wood surface defect detection, such algorithms still have obvious limitations: first, the tiny defects on the wood surface account for a very low proportion in the image, and it is difficult for general detection models to effectively capture their features, resulting in a prominent problem of missed detection of small targets; second, the visual similarity between wood texture and defects can easily confuse the model, especially in densely textured areas, where normal textures are frequently mistakenly identified as defects; finally, the traditional detection framework relies on bounding boxes to locate defects, and the contour capture accuracy of irregular or subtle defects is insufficient, making it difficult to achieve accurate pixel-level segmentation.

[0006] Most of the current improvement schemes focus on optimizing the YOLO model structure, such as strengthening local feature extraction by introducing an attention mechanism, or building a multi-scale feature fusion network to improve the detection ability of small targets. However, these methods still have essential shortcomings: first, although the data enhancement strategy can expand the diversity of samples, it cannot fundamentally remove the coupling relationship between features and other information such as wood board color and ambient brightness; second, although the attention mechanism can focus on suspicious areas, it lacks the ability to learn normal texture features, resulting in poor generalization of the model when facing unseen wood species or texture types; finally, and most importantly, the existing technology is highly dependent on a large amount of accurately labeled defect data. The types of wood surface defects are complex and diverse, and professional quality inspectors are required to mark the defect boundaries. In addition, the natural variability of wood texture and the fuzziness of defect boundaries make manual labeling prone to subjective bias. This strong dependence on high-precision labeled data not only greatly increases the cost of manpower and time, but also forms a sharp contradiction with the rapid iteration production needs of the wood processing industry.

[0007] For example, Chinese patent document CN113989224A discloses a color texture fabric defect detection method based on generative adversarial network, constructs a contrastive learning generative adversarial network model based on channel attention, inputs the color texture fabric data set after superimposing Gaussian noise into the contrastive learning generative adversarial network model based on channel attention for training, and obtains a trained contrastive learning generative adversarial network model based on channel attention; uses the trained contrastive learning generative adversarial network model based on channel attention to reconstruct the color texture fabric image to be detected, outputs the reconstructed image, and then detects to locate the defect area. The present invention calculates the residual of the color texture fabric image to be tested and the corresponding reconstructed image, and then performs threshold segmentation and closed operation processing on the residual result to achieve rapid detection and positioning of the defect area.

[0008] Chinese patent document CN110969606A discloses a texture surface defect detection method and system, which first extracts the multi-channel texture prior of the input texture image through a prior extraction step; then accurately reconstructs the texture background of the input image under the guidance of the extracted prior through a texture reconstruction step, encodes more accurate texture features in the latent space of the texture reconstruction module, and suppresses defects from being reconstructed in the texture background; finally, through pixel-level adversarial learning, the texture reconstruction accuracy is further improved. When performing detection, defects can be detected by simply performing a difference operation between the input image and the reconstructed texture background image. The present invention has a high detection accuracy for defects of different sizes and different contrasts on different texture surfaces.

[0009] In view of the above problems, the present invention proposes an innovative solution that combines multi-view contrast learning with a reconstruction-based anomaly detection method. It should be noted that "multi-view" and "multi-perspective" in the present invention are expressions of the same technical concept, both referring to methods of analyzing from different data dimensions or feature channels, and the two have the same technical meaning in this context. The present invention draws on the core idea of ​​the CMC multi-view contrast learning method proposed by YongLong Tian et al., and deeply integrates the multi-view contrast learning technology with the reconstruction-based anomaly detection technology, thereby realizing the accurate detection of small target defects and texture-similar defects on the wood surface. Summary of the invention

[0010] The present invention aims to overcome at least one defect of the above-mentioned prior art and provide a wood surface defect detection method based on multi-view anomaly detection.

[0011] The invention also discloses a device loaded with a wood surface defect detection method based on multi-viewing angle anomaly detection.

[0012] The detailed technical scheme of the present invention is as follows: A wood surface defect detection method based on multi-view anomaly detection, the method comprising: S1, obtaining an original wood image dataset and preprocessing the original wood images therein to obtain a target image, wherein the original wood image is an RGB image, and the target image is a Lab image; S2. Construct a wood surface defect detection model, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether a pixel residual value in the fused residual map exceeds a dynamic threshold, so as to mark a pixel area exceeding the dynamic threshold as a defect; S3, input the preprocessed wood image to be detected into the trained wood surface defect detection model, and after the encoder and decoder perform encoding-decoding operations in sequence, obtain a wood reconstructed image, and calculate the residual between the wood reconstructed image and the preprocessed wood image to be detected through the defect marking module to achieve wood defect detection and pixel-level positioning.

[0013] Preferably, according to the present invention, in S1, preprocessing the original wood images in the original wood image dataset specifically includes: The original wood image is cropped into a number of image blocks of 512×512 pixels, and an overlapping strategy with a step size of 256 pixels is adopted to ensure texture continuity and completely cover the defect area; Perform data augmentation operations on the cropped image blocks, including horizontal flipping, vertical flipping, and rotation; Use the color module in Python's skimage library to convert the augmented image block into a Lab image block to obtain the target image.

[0014] Preferably, according to the present invention, in S2, positive and negative samples are constructed based on the target image, specifically: the L view and ab view of the same Lab image block in the target image are used as a positive sample pair, and the L view and ab view between different Lab image blocks are used as a negative sample pair.

[0015] Preferably, according to the present invention, in S2, using the positive and negative samples to train the encoder through contrastive learning specifically includes: The encoder extracts the representation vectors of the L views in the positive and negative samples respectively and the representation vector of ab view , and calculate the characterization vector and the characterization vector The similarity between: (1); In formula (1): Indicates L view and ab view The similarity between corresponding representation vectors; and Represents L view and ab view Through two encoders and The extracted feature representation; To control the hyperparameters of the similarity scale, by adjusting The value of controls the sensitivity of similarity. The smaller the value, the higher the sensitivity. is an exponential function; Construct a contrast loss function and maximize the positive sample pair ( , ) similarity, minimize the negative sample pair ( , )and( , ), where i≠j, optimize the contrast loss function to optimize the training encoder; Among them, anchoring L view, comparing the contrast loss of ab view for: (2); In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; To include anchor points With positive samples The similarity of Negative samples The sum of similarities; It means taking the average of all possible sample pairs, including positive and negative samples; Anchoring the ab view, comparing the contrast loss of the L view for: (3); Then the contrast loss function of the encoder is: (4).

[0016] Preferably, in S2, a loss function is constructed by using mean square error MSE and structural similarity index SSIM to train the decoder, specifically: (11); In formula (11): Represents the input target image; represents the reconstructed image output by the decoder; represents the final loss function of the decoder; is the weight coefficient, which is used to control the loss weight of the mean square error MSE and the structural similarity index SSIM, and is set to .

[0017] Preferably, according to the present invention, in S2, a residual map between the reconstructed image and the target image is obtained to obtain a fused residual map, specifically: The MSE loss and SSIM loss between the reconstructed image and the target image are calculated respectively, and the corresponding residual graphs are obtained respectively, and then they are normalized and weighted summed to obtain the fused residual graph. : (12); In formula (12): is the fusion residual map; is the weight coefficient, which is used to control the weight of the residual graph corresponding to the MSE loss and SSIM loss in the fused residual graph; Represents the residual graph obtained by calculating the MSE loss; Represents the residual graph obtained by calculating the SSIM loss; Represents the maximum value of the residual graph obtained by calculating the MSE loss, Indicates the maximum value of the residual map obtained by calculating the SSIM loss, which is used to normalize to the range [0,1].

[0018] Preferably, according to the present invention, S2 further includes calculating a dynamic threshold, that is: For the fusion residual map Perform Gaussian smoothing filtering to eliminate isolated noise points; Calculating dynamic thresholds ,in, is the fusion residual graph The mean of the residual values ​​of the pixels in ; is the standard deviation of the pixel residual value; is the threshold coefficient, which is used to dynamically adjust the detection sensitivity.

[0019] In another aspect of the present invention, a device for implementing a wood surface defect detection method based on multi-view anomaly detection is provided, the device comprising: A sample acquisition module is used to acquire an original wood image data set and preprocess the original wood image therein to obtain a target image, wherein the original wood image is an RGB image and the target image is a Lab image; Model building modules for wood surface defect detection models, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether a pixel residual value in the fused residual map exceeds a dynamic threshold, so as to mark a pixel area exceeding the dynamic threshold as a defect; The defect detection module is used to detect the preprocessed wood image to be detected by using the trained wood surface defect detection model. After the encoder and the decoder perform encoding-decoding operations in sequence, a wood reconstructed image is obtained. The residual between the wood reconstructed image and the preprocessed wood image to be detected is calculated by the defect marking module to realize wood defect detection and pixel-level positioning.

[0020] In another aspect of the present invention, there is also provided an electronic device, comprising: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enable the at least one processor to execute the wood surface defect detection method based on multi-view anomaly detection as described above.

[0021] In another aspect of the present invention, a machine-readable storage medium is provided, which stores executable instructions, and when the instructions are executed, the machine executes the wood surface defect detection method based on multi-view anomaly detection as described above.

[0022] Compared with the prior art, the present invention has the following beneficial effects: (1) Fully self-supervised training reduces labeling costs. The innovative integration of multi-view self-supervised comparative learning and generative reconstruction networks allows model training to be completed without any manual labeling, fundamentally solving the industry pain points of diverse wood defect types, strong subjectivity in labeling, and high costs, and provides a technical foundation for rapid migration across tree species.

[0023] (2) The ability to detect tiny defects is significantly enhanced. Through multi-view contrast learning CMC extracts the essential characteristics of wood texture in the Lab color space, accurately distinguishes the subtle differences between tiny defects and normal texture, and overcomes the problem of high missed detection rate of small targets in traditional target detection models.

[0024] (3) Pixel-level high-precision defect positioning. Abandoning the traditional anchor frame mechanism, the residual map fusion strategy driven by reconstruction error is adopted. Abnormal pixels are directly marked through adaptive threshold segmentation, breaking through the positioning accuracy limit of the bounding box, and can accurately capture irregular defects to meet high-precision quality inspection standards.

[0025] (4) Strong robustness under complex texture interference. Based on multi-channel color space contrast decoupling technology, the coupling between feature and interference information is separated, and bilinear interpolation is combined to suppress edge artifacts. The false detection rate is greatly reduced in scenes with dense textures or uneven lighting. The anti-interference performance is better than the improved solution based on the attention mechanism.

[0026] (5) Highly adaptable design for industrial scenarios. The encoder-decoder symmetric architecture supports modular deployment, and the pre-trained feature extraction network is compatible with different imaging conditions, such as reflection and color difference. The image block detection strategy is combined with the weighted fusion mechanism to ensure full image coverage while meeting the balance between stability and efficiency of industrial pipelines. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flow chart of the wood surface defect detection method based on multi-view anomaly detection described in the present invention.

[0028] Figure 2 Schematic diagram of the original wood image obtained in Example 1 of the present invention.

[0029] Figure 3 1 is a schematic diagram of an operation of performing data augmentation on a cropped original wood image in Embodiment 1 of the present invention.

[0030] Figure 4 It is a comparison diagram of the original wood image in the RGB color mode and the wood image in the Lab color mode in Example 1 of the present invention.

[0031] Figure 5 It is a network structure diagram of the wood surface defect detection model in Example 1 of the present invention.

[0032] Figure 6Schematic diagram of positive and negative samples of contrastive learning constructed in Example 1 of the present invention.

[0033] Figure 7 4 is a network structure diagram of the encoder of the wood surface defect detection model in Example 1 of the present invention.

[0034] Figure 8 4 is a network structure diagram of the decoder of the wood surface defect detection model in Example 1 of the present invention.

[0035] Fig. 9 This is a flow chart of the wood surface defect detection model performing defect detection in Example 1 of the present invention. DETAILED DESCRIPTION

[0036] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0037] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.

[0038] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0039] In the absence of conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.

[0040] In view of the shortcomings of the prior art, the present invention provides a wood surface defect detection method based on multi-view anomaly detection, which combines the CMC multi-view contrast learning method with the reconstruction method in anomaly detection, constructs a complete network architecture, and applies it to the task of wood surface defect detection.

[0041] The CMC method can optimize between multiple views through self-supervised contrast learning, and train the model through the optimization of contrast loss to extract the essential features of the data. The principle of the reconstruction-based anomaly detection method is that the model first extracts the features of the normal image, then reconstructs the original image, and trains the reconstruction ability of the model by optimizing the reconstruction loss. When the model reconstructs the abnormal image, there will be a large deviation between the reconstruction result of the abnormal area and the original image. The location of the abnormal area can be marked by comparing with the original image. This innovation not only breaks through the limitations of the traditional defect detection framework, but also significantly improves the accuracy of defect detection.

[0042] Specifically, Contrastive Multiview Coding (CMC) is a method in contrastive learning, which attempts to extract the representation that best fits the nature of the feature by performing contrastive learning between multiple views. In coding theory, a basic idea is to learn a compressed representation that can reconstruct the original data. This idea is embodied in representation learning as autoencoders, generative models, and contrastive learning. Among them, both the autoencoder and the generative model are committed to representing data points or distributions in a lossless manner as much as possible, but in the present invention, lossless representation is not the most suitable feature extraction method. Because this type of method will also save other information such as wood color and ambient brightness into the compressed representation, thereby interfering with the representation of texture and defect features, which is obviously not in line with the design concept of this method. The CMC method retains useful information and eliminates other irrelevant information. In the design of the present invention, the encoder performs contrastive learning in the Lab color space, which can not only learn powerful feature representations, but also improve the anti-interference ability of redundant information.

[0043] Among the unsupervised anomaly detection methods, the reconstruction method performs better in detection accuracy than the feature embedding method. The feature embedding method usually learns the low-dimensional feature representation of normal images and tries to make the feature representation of all normal images concentrated. The feature representation of abnormal images will have a large deviation from the feature distribution of normal images. In the test phase, image-level anomaly detection can be achieved by calculating the deviation between the image to be tested and the normal distribution. Although this method can calculate the abnormal area with the help of feature representation and improve the detection accuracy to a certain extent, its calculation results may deviate from the actual situation, and the accuracy is difficult to meet the requirements of the present invention.

[0044] Therefore, the present invention adopts a reconstruction method to learn the ability to reconstruct normal images by training the network. During the test, the reconstructed image is compared and calculated with the original image to obtain a residual image. If the error value of the pixels in the residual image is higher than the set threshold, these pixel positions are marked as abnormal areas. In the present invention, the design of the reconstruction method can not only detect defects on the wood surface, but also realize pixel-level defect marking, making the detection result more accurate.

[0045] The wood surface defect detection method and device based on multi-view anomaly detection of the present invention are further described below in conjunction with specific embodiments.

[0046] Embodiment 1, Ginseng Figure 1 This embodiment provides a wood surface defect detection method based on multi-view anomaly detection, the method comprising: S1. Obtain an original wood image dataset and preprocess the original wood images therein to obtain a target image, wherein the original wood image is an RGB image, and the target image is a Lab image.

[0047] In this example, the original wood image dataset used is from a self-built dataset, in which all wood defect images are taken by industrial cameras, and the pixel resolution of the images is mostly 1960×1080. The dataset contains about 12,200 images in total, and the wood color, texture, and light brightness in each image are different. Figure 2 As shown, the original wood images are all images in RGB color mode.

[0048] The ultimate goal of the method of this embodiment is to detect and accurately mark small target defects and texture-similar defects on the wood surface, rather than just judging whether there are abnormalities in the image. Therefore, more sophisticated image data should be used when training the model. In order to effectively extract the essential characteristics of the texture and accurately distinguish normal images from defective images on a small data set, the method of this embodiment performs a series of preprocessing operations on the collected original wood images. In this embodiment, the preprocessing operations performed on the original wood images include image cropping, data augmentation, and color conversion.

[0049] Specifically, image cropping includes cropping a complete original wood image into several image blocks.

[0050] Although a complete original wood image contains the normal texture of wood, it also contains many defects. If the original wood image is directly used for contrastive learning, it is difficult to train an encoder that is sensitive to texture and defect features. Therefore, before contrastive learning, the original wood image is first cropped into several 512×512 pixel image blocks, and an overlapping strategy with a step size of 256 pixels is adopted to ensure texture continuity and complete coverage of defect areas.

[0051] Among these cropped image blocks, some contain only normal wood textures, while others contain a small amount of defects, which interfere with the normal texture, resulting in significant differences between the defective image blocks and the normal image blocks. The method of this embodiment mixes the image blocks containing only normal textures with the image blocks containing a small amount of defects for contrastive learning, which not only improves the effect of contrastive learning, but also makes the representation of abnormal images significantly different from that of normal images. In this way, the problem of false detection caused by the confusion between normal images and abnormal images can be effectively avoided during the test phase.

[0052] In order to achieve better results on a small data set, the method of this embodiment performs data augmentation operations on the cropped image blocks. These operations mainly include horizontal flipping, vertical flipping and rotation. The specific operations are as follows: Figure 3 shown.

[0053] based on Figure 2 As shown in FIG. 1 , the work of shooting wood images in an industrial environment is interfered by many factors, such as shooting angle, board angle, light intensity, and board color, which directly affect the quality of the data set. Therefore, the method of this embodiment adopts a method of multi-view contrast learning in Lab color space.

[0054] On this basis, the method of this embodiment uses the color module in Python's skimage library to directly convert the RGB image block into a Lab image block. In the Lab color space, the L channel represents lightness, which indicates the brightness of the color; the a channel and the b channel represent the two axes of the color, respectively, the a channel represents the green-red axis (Green-Red Axis), and the b channel represents the blue-yellow axis (Blue-Yellow Axis). Multi-view contrast learning in this color space can not only learn the features of the image, but also decouple the features from other irrelevant information, thereby reducing the interference of these factors on feature extraction and improving the robustness of the model under different conditions.

[0055] The comparison of the same wood image block in RGB color mode and Lab color mode is as follows: Figure 4 shown.

[0056] After the above operations, a preprocessed target image is obtained, which is the Lab image block.

[0057] S2. Construct a wood surface defect detection model, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; A defect marking module for error fusion and adaptive threshold segmentation is constructed, wherein the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether there is a pixel residual value in the fused residual map that exceeds a dynamic threshold, so as to mark the pixel area that exceeds the dynamic threshold as a defect.

[0058] In this embodiment, in order to efficiently extract the essential features of the wood surface image under different colors of wooden boards and different lighting conditions, the method of this embodiment adopts an AlexNet variant network adapted to the CMC concept, and pre-trains the image blocks converted to the Lab color mode as input data. After the pre-training is completed, the network parameters are frozen and used as an encoder module. Subsequently, referring to the principle of the reconstruction method in anomaly detection, a decoder module symmetrical to the encoder structure is constructed to reconstruct the original image and train the module to reconstruct the normal image. In the final stage, the fusion residual map and adaptive threshold segmentation technology are used to calculate the fusion residual map between the reconstructed image and the original image, and the defect position is marked accordingly, thereby realizing defect detection on the wood surface.

[0059] The complete network structure of the wood surface defect detection model is as follows: Figure 5 As shown in the figure, the encoder structure includes convolution operation (Conv), batch normalization (BN), activation function (ReLU), downsampling operation (Pool) and fully connected layer (FC), while the unique operations in the decoder structure include reconstructed feature map, upsampling operation (UnPool) and transposed convolution operation (DeConv).

[0060] Specifically, first build the encoder.

[0061] This embodiment constructs an encoder of a wood surface defect detection model based on an AlexNet variant network, then constructs positive and negative samples based on a target image, and uses the positive and negative samples to train the encoder through contrast learning.

[0062] The target image is a Lab image block. In order to perform multi-view contrast learning in Lab color mode, it is necessary to first define the views used and the positive and negative samples. The method of this embodiment draws on the idea of ​​the CMC method: in the Lab color space, the L channel represents brightness, and the a channel and the b channel represent color. Therefore, the method of this embodiment uses the L channel as the L view, and combines the a channel and the b channel as another view, called the ab view.

[0063] Ginseng Figure 6 As shown, the positive and negative samples of this embodiment are defined as follows: the L view and ab view of the same image block are used as positive sample pairs, and the L view and ab view between different image blocks are used as negative sample pairs. Through this multi-view contrast learning method, the encoder can ignore the interference of brightness and color differences and focus on extracting the texture features and defect features of the wood board.

[0064] The encoder network structure selected in this embodiment is an AlexNet variant network adapted to the CMC concept. The network is divided into two parts in the channel dimension of the feature map: one part is used to train the L view and the other part is used to train the ab view. Functionally, these can be regarded as two independent encoders, but in terms of network structure, they still belong to the same network, except that the parameters of the two parts are independent. The overall network architecture is still 5 convolutional layers and 3 fully connected layers. The convolutional layers include convolution and downsampling operations. This design not only realizes the functions of two encoders, but also does not significantly increase the number of parameters.

[0065] Since the two parts of the network perform encoding operations independently, when a complete Lab image block is input, the encoder extracts the representation vectors of the L view and the ab view respectively. and The contrast loss function maximizes the positive sample pair ( , ), while minimizing the similarity of negative sample pairs ( , )and( , ) similarity (where i≠j), continuously optimizing the contrast loss function, the model gradually brings the representation vectors between positive samples closer and pushes the representation vectors between negative samples farther away, thereby establishing a strong ability to distinguish normal textures from defect features. The network structure of the encoder is as follows Figure 7 As shown, it includes 5 convolutional layers (Conv) and 3 fully connected layers (FC).

[0066] First, calculate the similarity between the two representation vectors : (1); In formula (1): Indicates L view and ab view The similarity between corresponding representation vectors; and Represents L view and ab view Through two encoders and The extracted feature representation; is a hyperparameter that controls the similarity scale and can be adjusted The value of controls the sensitivity of similarity. The smaller the value, the higher the sensitivity. It is an exponential function. The exponent is taken to ensure that the output result is a positive value and can quickly amplify the difference according to the size of the similarity.

[0067] After obtaining the similarity between the representation vectors, the next step is to calculate the contrast loss of positive and negative samples and supervise the encoder training by optimizing the contrast loss; the calculation of the contrast loss is: (2); In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; To include anchor points With positive samples The similarity of Negative samples The sum of similarities is used to normalize the similarities of all sample pairs; It means taking the average of all possible sample pairs, including positive and negative samples, to ensure the optimization of the loss function.

[0068] The above formula (2) can be understood as anchoring the L view and performing a comparison calculation with the ab view. To achieve better results, the comparison loss of anchoring the ab view and comparing it with the L view can be calculated again: (3).

[0069] The above two formulas only show the loss calculation of one anchor point sample. In actual application, each sample will be used as an anchor point in turn, and the two losses will be calculated separately and the average value will be calculated. Finally, the sum of the two losses is used as the final comparison loss: (4).

[0070] When the contrast loss value of the validation set data is continuously lower than the set threshold, the encoder is partially trained, the encoder parameters are frozen, and an operation is added at the end to convert the final representation vector and Concatenate into a representation vector in preparation for training the decoder.

[0071] Next we build the decoder.

[0072] Traditional autoencoders compress input features through an encoder, and then restore the original image through a decoder to achieve the goal. However, its encoder retains all information as much as possible during the compression process, which makes the feature extraction and reconstruction results susceptible to interference from irrelevant factors. Therefore, the method of this embodiment draws on the idea of ​​CMC to construct an encoder to extract features, so that the extracted features are closer to the essence.

[0073] When designing the decoder part of this embodiment, its network structure needs to be adapted to the encoder, and the flipped form of the AlexNet variant network is selected, which is symmetrical with the encoder structure. Specifically, all the convolution and downsampling operations in the encoder are replaced with transposed convolution and upsampling operations. Each layer performs upsampling, halving the number of channels, and nonlinear activation processing in turn, and finally outputs a reconstructed image with the same size as the original image block. This design makes the reconstruction of the image simple and efficient. In the previous pre-training stage, the encoder has been able to extract the essential texture representation from the Lab image block, and the decoder only needs to focus on gradually restoring the original image through transposed convolution and upsampling operations. The network structure of the decoder is as follows Figure 8 As shown, it includes a fully connected layer (FC), a reconstructed feature layer, and 5 transposed convolutional layers (DeConv).

[0074] The transposed convolution is a parameterized upsampling process learned through back-propagation. Its essence is to use a learnable convolution kernel to slide on the low-resolution feature map and fill the gaps, thereby gradually restoring the high-resolution features. Upsampling uses bilinear interpolation to determine the new pixel value by calculating the weighted average of adjacent pixels. The weight is determined by the relative distance between the target pixel and the surrounding four original pixels. Bilinear interpolation can strike a balance between smoothness and computational efficiency, effectively suppress checkerboard artifacts, and ensure the continuity of wood texture and the natural transition of defect boundaries.

[0075] Reconstruction-based anomaly detection methods usually require a loss function. This loss function is not only used to guide the network to learn the ability to reconstruct the original image during the training phase, but also used to calculate the residual map between the reconstructed image and the original image during the test phase, thereby marking the abnormal position. Commonly used reconstruction loss indicators include mean square error MSE, cross entropy loss, and structural similarity index SSIM. Among them, cross entropy loss is mainly used to process binary or probability distribution data to measure the difference between two probability distributions, so it is not suitable for use in the method of this embodiment. The mean square error MSE is achieved by calculating the error sample by sample and then taking the average. Specifically in image processing, it compares the errors of two images pixel by pixel, which is suitable for application in the method of this embodiment. The calculation formula of the mean square error MSE is: (5); In formula (5): Represents the original image, i.e., the target image; represents the reconstructed image; Indicates the height and width of the image size; Indicates the original image At pixel position The intensity value at Represents the reconstructed image At this pixel position , calculate the difference in pixel intensity values ​​at each corresponding position of the two images, take the square of the difference and sum it up, and finally find the average.

[0076] The pixel-by-pixel calculation makes the mean square error (MSE) very sensitive to pixel value differences, and it is easy to identify the location of anomalies. However, due to its oversensitivity and the neglect of the structural information of the image, when there is a slight offset between the reconstructed image and the original image, these errors caused by alignment problems will also be treated as anomalies by the mean square error (MSE). In addition, this method is difficult to detect defective areas that have changed visually but whose pixel values ​​remain roughly the same. Bergmann et al. pointed out that methods such as the mean square error (MSE) assume that adjacent pixels are independent, but this assumption does not fully conform to the structural characteristics of actual images, resulting in their poor performance in detecting structural differences. The structural similarity index (SSIM) can improve this problem well.

[0077] The structural similarity index SSIM was proposed by Wang et al. in 2004. It aims to better evaluate the visual quality of an image by integrating these factors. Compared with traditional pixel-level error metrics, the structural similarity index SSIM pays more attention to the structural characteristics of the image.

[0078] The structural similarity index SSIM defines a distance metric to measure the similarity between two K×K image patches p and q, which takes into account the brightness similarity. (p,q), contrast similarity (p,q) and structural similarity (p, q), which is applied to the method of this embodiment to calculate , and , the formula is as follows: (6); In formula (6): , , ∈R is a user-defined weight factor, which is used to control the influence of brightness, contrast and structure, and is usually set to 1.

[0079] Among them, brightness similarity for: (7); In formula (7): , Respectively represent the original image and reconstructed image The mean of is a small constant, usually set to 0.01, used to prevent the denominator from being zero.

[0080] Contrast Similarity for: (8); In formula (8): , Respectively represent the original image and reconstructed image The standard deviation of is a small constant, usually set to 0.03.

[0081] Structural similarity for: (9); In formula (9): Indicates the original image and reconstructed image The covariance between .

[0082] Substituting the above three calculation formulas into the calculation formula of the structural similarity index SSIM, we can get: (10).

[0083] When the original image and reconstructed image When two images are exactly the same, =1; on the contrary, when the original image and reconstructed image When two images are completely different in brightness, contrast, and structure, =-1; therefore, the value range of SSIM is [−1,1].

[0084] However, SSIM is based on local window calculation and is less sensitive to defects in global distribution. In addition, the overall brightness distribution of the reconstructed image is offset from that of the original image. SSIM may not be able to fully reflect such problems.

[0085] Therefore, the method of this embodiment uses the weighted sum of the mean square error MSE and the structural similarity index SSIM as the final loss function to train the decoder, that is, (11); In formula (11): represents the final loss function of the decoder; is the weight coefficient, which is used to control the loss weight of the mean square error MSE and the structural similarity index SSIM, which is set here as .

[0086] In this way, a balance can be achieved between pixel-level accuracy and structural consistency, thereby improving the overall performance of the decoder.

[0087] Finally, a defect marking module of error fusion and adaptive threshold segmentation is constructed to calculate the residual map and locate defects.

[0088] After the decoder is trained, the network can realize the complete process from input image to reconstructed original image. After the decoder outputs the reconstructed image, it is necessary to calculate the residual image between the reconstructed image and the input target image to highlight the defect area, so as to facilitate the marking of the defect location.

[0089] This embodiment method designs a defect marking module with error fusion and adaptive threshold segmentation. Specifically, the mean square error MSE and the structural similarity index SSIM are used to calculate the residual images of the reconstructed image and the target image respectively, and then the two calculated residual images are fused into one residual image. Finally, with the help of adaptive threshold segmentation technology, defect positioning is achieved more accurately.

[0090] That is, first calculate the MSE loss and SSIM loss between the reconstructed image and the target image, calculate the residual map, and then perform normalization and weighted summation to obtain the fused residual map : (12); In formula (12): is the fusion residual map; is the weight coefficient, which is used to control the weights of the two residual images in the fused residual image; Represents the residual graph obtained by calculating the MSE loss; Represents the residual graph obtained by calculating the SSIM loss; Represents the maximum value of the residual graph obtained by calculating the MSE loss, Indicates the maximum value of the residual map obtained by calculating the SSIM loss, which is used to normalize to the range [0,1].

[0091] Finally, in order to convert the pixel residual value into a binary defect mask, this embodiment adopts an adaptive threshold algorithm based on image statistics. Perform Gaussian smoothing filtering to eliminate isolated noise points; then calculate the dynamic threshold ,in for The mean of the pixel residual values ​​in , is the standard deviation, is the threshold coefficient, which is used to dynamically adjust the detection sensitivity. The larger the value, the higher the calculated threshold The larger it is, the stricter the detection is. Finally, the pixel area where the pixel residual value exceeds the threshold is marked as a defect.

[0092] Based on the above, the construction and training process of the wood surface defect detection model is completed, and then it is deployed in the actual wood surface defect detection application.

[0093] S3, input the preprocessed wood image to be detected into the trained wood surface defect detection model, and after the encoder and decoder perform encoding and decoding operations in sequence, obtain a wood reconstructed image, and calculate the residual between the preprocessed wood image to be detected and the wood reconstructed image through the defect marking module to realize wood defect detection and pixel-level positioning.

[0094] This step is to deploy the trained wood surface defect detection model to the actual application of wood surface defect detection. Fig. 9 , which shows a complete wood surface defect detection process.

[0095] That is, firstly, the wood image to be detected is preprocessed based on the preprocessing operation in step S1, converted into a Lab image, and then input into the trained wood surface defect detection model. After encoding by its encoder, the corresponding feature vector is output to the decoder, and the decoder performs the decoding operation and outputs the wood reconstructed image. Finally, the residual map between the preprocessed wood image to be detected and the wood reconstructed image is calculated by the defect marking module, and it is determined whether there is a pixel residual value exceeding the dynamic threshold. If so, the pixel area where the pixel residual value exceeds the threshold is marked as a defect, thereby realizing wood defect detection and pixel-level positioning.

[0096] In summary, the wood surface defect detection method based on multi-view anomaly detection of the present invention constructs an encoder module based on the idea of ​​CMC to extract image features; at the same time, a decoder module is built according to the principle of the decoder to reconstruct the original image. By comparing the reconstructed image with the original image, the defect position is marked, thereby efficiently completing the wood defect detection task.

[0097] The innovative design of the present invention organically combines the excellent feature extraction capability of the CMC method with the pixel-level anomaly detection capability of the reconstruction anomaly detection method. By decoupling the essential features of wood texture from interference information, and combining pixel-level residual analysis, the present invention achieves accurate detection and positioning of defects without manual labeling, fundamentally avoiding the reliance of traditional methods on massive labeled data. In addition, the encoder built based on the CMC concept focuses on extracting the essential features of normal textures, which enables the model to have good generalization capabilities even when facing unseen wood species, thereby providing a low-cost and highly generalized defect detection solution for the wood processing industry.

[0098] Embodiment 2, This embodiment provides a device for implementing a wood surface defect detection method based on multi-view anomaly detection, the device comprising: A sample acquisition module is used to acquire an original wood image data set and preprocess the original wood image therein to obtain a target image, wherein the original wood image is an RGB image and the target image is a Lab image; Model building modules for wood surface defect detection models, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether a pixel residual value in the fused residual map exceeds a dynamic threshold, so as to mark a pixel area exceeding the dynamic threshold as a defect; The defect detection module is used to detect the preprocessed wood image to be detected using the trained wood surface defect detection model. After the encoder and decoder perform encoding and decoding operations in sequence, a wood reconstructed image is obtained. The residual between the preprocessed wood image to be detected and the wood reconstructed image is calculated through the defect marking module to realize wood defect detection and pixel-level positioning.

[0099] Embodiment 3, This embodiment also provides an electronic device, including: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enable the at least one processor to execute the wood surface defect detection method based on multi-view anomaly detection as described above.

[0100] In this embodiment, the electronic device may include, but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and the like.

[0101] Embodiment 4, This embodiment also provides a machine-readable storage medium storing executable instructions, which, when executed, enable the machine to execute the wood surface defect detection method based on multi-view anomaly detection as described above.

[0102] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes that implement the functions of any of the above-mentioned embodiments are stored, and a computer or processor of the system or device can read and execute instructions stored in the readable storage medium.

[0103] In this case, the program code itself read from the machine-readable medium can realize the function of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0104] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code may be downloaded from a server computer or a cloud via a communication network.

[0105] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0107] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0109] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A wood surface defect detection method based on multi-view anomaly detection, characterized in that: The method comprises: S1, obtaining an original wood image dataset and preprocessing the original wood images therein to obtain a target image, wherein the original wood image is an RGB image, and the target image is a Lab image; S2. Construct a wood surface defect detection model, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether a pixel residual value in the fused residual map exceeds a dynamic threshold, so as to mark a pixel area exceeding the dynamic threshold as a defect; S3, input the preprocessed wood image to be detected into the trained wood surface defect detection model, and after the encoder and decoder perform encoding-decoding operations in sequence, obtain a wood reconstructed image, and calculate the residual between the wood reconstructed image and the preprocessed wood image to be detected through the defect marking module to achieve wood defect detection and pixel-level positioning.

2. The wood surface defect detection method based on multi-view anomaly detection according to claim 1 is characterized in that: In S1, the original wood images in the original wood image dataset are preprocessed, specifically including: The original wood image is cropped into a number of image blocks of 512×512 pixels, and an overlapping strategy with a step size of 256 pixels is adopted to ensure texture continuity and completely cover the defect area; Perform data augmentation operations on the cropped image blocks, including horizontal flipping, vertical flipping, and rotation; Use the color module in Python's skimage library to convert the augmented image block into a Lab image block to obtain the target image.

3. The wood surface defect detection method based on multi-view anomaly detection according to claim 1 is characterized in that: In S2, positive and negative samples are constructed based on the target image, specifically: the L view and ab view of the same Lab image block in the target image are used as positive sample pairs, and the L view and ab view between different Lab image blocks are used as negative sample pairs.

4. The wood surface defect detection method based on multi-view anomaly detection according to claim 3 is characterized in that: In S2, the encoder is trained by contrastive learning using the positive and negative samples, specifically including: The encoder extracts the representation vectors of the L views in the positive and negative samples respectively and the representation vector of ab view , and calculate the characterization vector and the characterization vector The similarity between: (1); In formula (1): Indicates L view and ab view The similarity between corresponding representation vectors; and Represents L view and ab view Through two encoders and The extracted feature representation; To control the hyperparameters of the similarity scale, by adjusting The value of controls the sensitivity of similarity. The smaller the value, the higher the sensitivity. is an exponential function; Construct a contrast loss function and maximize the positive sample pair ( , ) similarity, minimize the negative sample pair ( , )and( , ), where i≠j, optimize the contrast loss function to optimize the training encoder; Among them, anchoring the L view, comparing the contrast loss of the ab view for: (2); In formula (2): represents the L view of the first image; represents the ab view of the first image; is the similarity between a pair of positive samples; To include anchor points With positive samples The similarity of Negative samples The sum of similarities; It means taking the average of all possible sample pairs, including positive and negative samples; Anchoring the ab view, comparing the contrast loss of the L view for: (3); Then the contrast loss function of the encoder is: (4)。 5. The wood surface defect detection method based on multi-view anomaly detection according to claim 1 is characterized in that: In S2, a loss function is constructed by using mean square error MSE and structural similarity index SSIM to train the decoder, specifically: (11); In formula (11): Represents the input target image; represents the reconstructed image output by the decoder; represents the final loss function of the decoder; is the weight coefficient, which is used to control the loss weight of the mean square error MSE and the structural similarity index SSIM, and is set to .

6. The wood surface defect detection method based on multi-view anomaly detection according to claim 1 is characterized in that: In S2, a residual image between the reconstructed image and the target image is obtained to obtain a fused residual image, specifically: The MSE loss and SSIM loss between the reconstructed image and the target image are calculated respectively, and the corresponding residual graphs are obtained respectively, and then they are normalized and weighted summed to obtain the fused residual graph. : (12); In formula (12): is the fusion residual map; is the weight coefficient, which is used to control the weight of the residual graph corresponding to the MSE loss and SSIM loss in the fused residual graph; Represents the residual graph obtained by calculating the MSE loss; Represents the residual graph obtained by calculating the SSIM loss; Represents the maximum value of the residual graph obtained by calculating the MSE loss, Indicates the maximum value of the residual map obtained by calculating the SSIM loss, which is used to normalize to the range [0,1].

7. The wood surface defect detection method based on multi-view anomaly detection according to claim 6 is characterized in that: The step S2 also includes calculating a dynamic threshold, namely: For the fusion residual map Perform Gaussian smoothing filtering to eliminate isolated noise points; Calculating dynamic thresholds ,in, is the fusion residual graph The mean of the residual values ​​of the pixels in ; is the standard deviation of the pixel residual value; is the threshold coefficient, which is used to dynamically adjust the detection sensitivity.

8. A device for implementing a wood surface defect detection method based on multi-view anomaly detection, characterized in that: The device comprises: A sample acquisition module is used to acquire an original wood image data set and preprocess the original wood image therein to obtain a target image, wherein the original wood image is an RGB image and the target image is a Lab image; Model building modules for wood surface defect detection models, including: Constructing an encoder of the wood surface defect detection model based on an AlexNet variant network, constructing positive and negative samples based on the target image, and using the positive and negative samples to train the encoder through contrastive learning; A decoder of the wood surface defect detection model is constructed based on a flipped AlexNet variant network, parameters of the trained encoder are frozen, and the output of the encoder is used as the input of the decoder. A loss function is constructed through mean square error MSE and structural similarity index SSIM to train the decoder, and the decoder outputs a reconstructed image; Constructing a defect marking module for error fusion and adaptive threshold segmentation, the defect marking module is used to obtain a residual map between the reconstructed image and the target image to obtain a fused residual map, and to determine whether a pixel residual value in the fused residual map exceeds a dynamic threshold, so as to mark a pixel area exceeding the dynamic threshold as a defect; The defect detection module is used to detect the preprocessed wood image to be detected by using the trained wood surface defect detection model. After the encoder and the decoder perform encoding-decoding operations in sequence, a wood reconstructed image is obtained. The residual between the wood reconstructed image and the preprocessed wood image to be detected is calculated by the defect marking module to realize wood defect detection and pixel-level positioning.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enable the at least one processor to execute the wood surface defect detection method based on multi-view anomaly detection as described in any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores executable instructions, which, when executed, enable the machine to perform the wood surface defect detection method based on multi-view anomaly detection as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Texture surface defect detection method and system

    CN110969606A

  • Color texture fabric defect detection method based on generative adversarial network

    CN113989224A

  • Multi-view generation method based on comparative learning

    CN112598775A

  • Display panel Mura defect evaluation method and system and readable storage medium

    CN112954304A

  • Super-pixel-based flexible IC substrate color change defect detection method and device

    CN112991302A

Cited By

  • Real-time monitoring method for engine production line assembly

    CN120599548A

  • Underwater building defect detection method based on adaptive deep learning model training

    CN120823491A

  • Wood surface defect detection method based on multi-view coding and feature memory bank

    CN121708008A

  • A Wood Surface Defect Detection Method Based on Multi-View Encoding and Feature Memory

    CN121708008B

  • Artificial board defect online detection and classification system based on machine vision

    CN121788514A