Wooden door and window inner and outer frame size detection method based on semantic segmentation network

By using the W-HRNet model based on semantic segmentation network, combined with parallel multi-resolution architecture and wavelet transform, the time-consuming, labor-intensive, and error-prone problems of detecting the inner and outer frame dimensions of wooden doors and windows are solved, achieving efficient and accurate automated detection.

CN120876357APending Publication Date: 2025-10-31NORTHEAST FORESTRY UNIV

Patent Information

Application Number
CN202510763483.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the current technology, the measurement of the inner and outer frame dimensions of wooden doors and windows mainly relies on manual measurement, which is time-consuming, labor-intensive, and prone to human error, making it difficult to adapt to diverse size specifications.

Method used

A semantic segmentation network-based approach is adopted, which utilizes the W-HRNet network model combined with a parallel multi-resolution architecture and wavelet transform to segment the inner and outer frame contours of wooden doors and windows. Automated detection is achieved through image preprocessing, sample annotation, model training and optimization.

Benefits of technology

It significantly improves the accuracy and robustness of the inspection of the inner and outer frame dimensions of wooden doors and windows, reduces labor costs, and realizes automated inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876357A_ABST
    Figure CN120876357A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic segmentation network-based wooden door and window inner and outer frame size detection method, and belongs to the technical field of wooden door and window inner and outer frame size detection. The method comprises the following steps: obtaining an original image of a wood door and window in a to-be-detected area and carrying out image preprocessing; constructing an image-label pair data set; importing the training set and the verification set into a W-HRNet network model for training and optimization; a final result is obtained for precision evaluation; obtaining a segmentation result of wood door and window inner and outer frame contours; and calculating geometrical characteristic parameters of inner and outer frames of the wooden door and window. According to the method, the spatial detail expression capability is enhanced, automatic detection of the sizes of the inner and outer frames of the wooden door and window is realized, the labor cost is remarkably reduced, and the accuracy and robustness of size detection of the inner and outer frames of the wooden door and window are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks, belonging to the technical field of inner and outer frame dimension detection for wooden doors and windows. Background Technology

[0002] Inspecting the dimensions of the inner and outer frames of wooden doors and windows is a crucial step in industrial production, significantly impacting product quality control and production efficiency. Accurate dimensional inspection not only ensures products meet specifications and standards but also enhances production line automation and reduces production costs.

[0003] Currently, in the wooden door and window manufacturing industry, the measurement of inner and outer frame dimensions mainly relies on manual measurement. Operators need to use measuring tapes or other measuring tools to measure the dimensions of the inner, middle, and outer frames of wooden doors and windows one by one. This method is not only time-consuming and labor-intensive, but also prone to human error. In actual production, wooden doors and windows come in a wide variety of sizes, ranging from hundreds to thousands of millimeters, from small to large, which poses a significant challenge to traditional manual measurement methods. Summary of the Invention

[0004] To address the problems existing in the background technology, the present invention provides a method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks, the method comprising the following steps: S1: Acquire the original image of the wooden doors and windows in the area to be tested and perform image preprocessing; S2: Based on the structural features of the wooden door and window frames in the area to be tested, determine the range and number of sample annotations, and select the original sample images from the preprocessed wooden door and window images in S1 for human-computer interactive annotation to construct an image-label pair dataset; S3: Perform multi-class labeling on the frame structure of wooden doors and windows, randomly divide the labeled sample dataset into training set, validation set and test set according to a certain proportion, and import the training set and validation set into the improved W-HRNet network model for training and optimization. S4: Perform contour segmentation of the inner and outer frames of wooden doors and windows on the W-HRNet network model trained and optimized in S3, extract the model results and perform calculations to obtain the final result, and then evaluate the accuracy of the final result through the test set. If the accuracy requirement is met, proceed to S5; otherwise, return to S2 to check the sample drawing and adjust and optimize each network model until the requirement is met. S5: Apply the W-HRNet network model that passed the accuracy evaluation in S4 to the detection of the inner and outer frame contours of wooden doors and windows in actual production. Input the image of the wooden door and window to be tested into the model for inference, obtain the segmentation results of the inner and outer frame contours of the wooden door and window, and output the segmentation results in the form of contour coordinate point sequence. S6: Based on the contour coordinate point sequence output by S5, calculate the geometric feature parameters of the inner and outer frames of wooden doors and windows, and store the calculation results in the database to realize the automated detection of the contours of the inner and outer frames of wooden doors and windows.

[0006] Furthermore, step S1 includes the following steps: S101: Use an industrial camera to acquire images of wooden doors and windows in the area to be tested; S102: Perform internal and external parameter calibration on the industrial camera, establish the camera imaging geometric model, and eliminate radial and tangential distortion of the lens; S103: Perform preprocessing operations on the calibrated image; Furthermore, the preprocessing operation described in S103 includes image grayscale conversion using the RGB three-channel weighted average method, noise reduction processing using Gaussian filtering, and contrast enhancement processing.

[0007] Furthermore, the W-HRNet network model described in S3 includes The basic network HRNet consists of two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3. After each convolutional layer, BN and ReLU operations are performed to reduce the resolution of the input image to 1 / 4 of the original image. Parallel multi-resolution architecture: It contains four branches, each of which uses multiple 3×3 convolutions. After convolution, BN and ReLU operations are performed. Each branch corresponds to a different resolution, and the number of feature map channels in each branch increases layer by layer. Feature exchange and fusion are achieved between branches through upsampling and downsampling. Furthermore, the decoder of the underlying network HRNet employs a progressive feature fusion method, including the following steps: S301: Upsample the 32x downsampled features and use transposed convolution to increase their resolution to 16x; concatenate the upsampled 32x features with the original 16x features to form preliminary fused features; S302: Upsample the 16x downsampled features and use transposed convolution to increase their resolution to 8x; then concatenate the upsampled 16x features with the original 8x features to form further fused features; S303: Retain the features of three independent paths: the path with 32x downsampling features upsampled to 8x scale, the path with 16x downsampling features upsampled to 8x scale, and the original 8x downsampling feature path. The features of the three independent paths are then concatenated to form a multi-scale fused feature. S304: In the process of multi-scale fusion feature fusion, WTConv2d is used in the last_layer to enhance the spatial detail expression capability, and CBAM is combined in the first branch x[0] and the second branch x[1] to improve the semantic information extraction capability. S305: Use the optimized model as the initial model and iteratively train it on a new sample set until the model performance reaches the expected effect.

[0008] Furthermore, the resolutions corresponding to the four branches of the parallel multi-resolution architecture are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original wood image, respectively. Under the hrnetv2_w18 configuration, the number of feature map channels of each branch are 18, 36, 72 and 144, respectively.

[0009] Furthermore, the feature exchange and fusion includes the following steps: S3-1: Use two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3, followed by BN and ReLU operations.

[0010] S3-2: After passing through three bottleneck modules, the output of S3-1 is divided into 2 / 3 / 4 parallel branches in stages. Each stage is downsampled and the number of channels is increased through the transition layer. S3-3: In each stage, different branches exchange and fuse features through upsampling and downsampling.

[0011] Furthermore, the model training described in S3 includes the following steps: S3A: Parameter initialization settings: S3B: Model Training: Input the wooden window dataset into the W-HRNet network model and choose whether to freeze the training. S3C: Parameter Update: Employs the cross-entropy loss function, selects whether to use Focal Loss to address class imbalance, and supports setting different weight coefficients for different classes; calculates the loss value between the prediction result and the label, and updates the network parameters through backpropagation; The formula for calculating the cross-entropy loss function is as follows:

[0012] in: Indicates the number of categories; This represents the weight coefficient of the i-th category; Indicates the true label; This indicates the probability that the model predicts the pixel belongs to the i-th class; When using Focal Loss, the cross-entropy loss function becomes:

[0013] in: It is a modulation factor used to reduce the weight of easily classified samples and increase the weight of difficult-to-classify samples; S3D: Output Model: The weights of the W-HRNet network model are saved every five epochs and evaluated on the training set; training stops when the number of training epochs reaches a set value, and the W-HRNet network model weight file is saved in the logs folder.

[0014] Furthermore, the accuracy evaluation described in S4 is calculated using precision, exact rate, recall, score, and mean intersection-over-union ratio, as shown in the following formulas:

[0015]

[0016]

[0017]

[0018]

[0019] in: ACC stands for accuracy. TP represents the number of pixels correctly classified as the target category; TN represents the number of pixels correctly classified as non-target categories; FP represents the number of pixels that were misclassified as the target category; FN represents the number of pixels misclassified as non-target categories; Precision, or P, represents accuracy. Recall, or R, represents the recall rate. F1 represents a fraction; mIoU represents the average crossover-union ratio; k represents the category.

[0020] Compared with the prior art, the beneficial effects of the present invention are: This invention applies an improved W-HRNet network model to the segmentation of the inner and outer frame contours of wooden doors and windows. By combining a parallel multi-resolution architecture with wavelet transform, and simultaneously incorporating an attention mechanism into the deep learning model for wooden door and window frame contour segmentation, it considers both the W-HRNet network model's ability to capture overall multi-scale information and the wavelet transform, effectively increasing the receptive field of convolution and providing crucial information for contour recognition and size detection of wooden doors and windows. This enhances the ability to represent spatial details, enables automated detection of the inner and outer frame dimensions of wooden doors and windows, significantly reduces labor costs, and improves the accuracy and robustness of the size detection. Attached Figure Description

[0021] Figure 1 This is a flowchart of the present invention; Figure 2 This is a schematic diagram of the structure of the W-HRNet network model of the present invention; Figure 3 This is a schematic diagram of the multi-scale fusion module in the W-HRNet network model of this invention; Figure 4 This is a schematic diagram of the semantic segmentation results for wooden doors and windows according to the present invention; Figure 5 This is a schematic diagram of the detection results of the inner and outer frame contours of wooden doors and windows according to the present invention. Detailed Implementation

[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0023] A method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks, the method comprising the following steps: S1: Acquire the original image of the wooden doors and windows in the area to be tested and perform image preprocessing; S2: Based on the structural features of the wooden door and window frames in the area to be tested, determine the range and number of sample annotations. Then, select 1210 original sample images from the preprocessed wooden door and window images in S1 and use Labelme software to perform human-computer interactive annotation on the original sample images to construct the image-label pair dataset required for deep learning. S3: In Labelme software, the frame structure of wooden doors and windows is labeled in multiple categories. The frame is divided into three categories: outer frame, middle frame, and inner frame. The labeled sample dataset is randomly divided into training set, validation set, and test set in a certain ratio (8:1:1). The training set and validation set are then imported into the improved W-HRNet network model for training and optimization. S4: Perform contour segmentation of the inner and outer frames of wooden doors and windows on the W-HRNet network model trained and optimized in S3, extract the model results and perform calculations to obtain the final result, and then evaluate the accuracy of the final result through the test set. If the accuracy requirement is met, proceed to S5; otherwise, return to S2 to check the sample drawing and adjust and optimize each network model until the requirement is met. S5: Apply the W-HRNet network model that passed the accuracy evaluation in S4 to the detection of the inner and outer frame contours of wooden doors and windows in actual production. Input the image of the wooden door and window to be tested into the model for inference, obtain the segmentation results of the inner and outer frame contours of the wooden door and window, and output the segmentation results in the form of contour coordinate point sequence. S6: Based on the contour coordinate point sequence output by S5, calculate the geometric feature parameters of the inner and outer frames of wooden doors and windows, including key data such as the positional relationship and size of the inner and outer frames, and store the calculation results in the database to realize the automated detection of the contours of the inner and outer frames of wooden doors and windows.

[0024] Furthermore, step S1 includes the following steps: S101: Use an industrial camera to acquire images of wooden doors and windows in the area to be tested; the acquisition parameters of the industrial camera are set as follows: image resolution of 5472×3648 pixels, frame rate of 20fps, pixel size of 24μm and equipped with LED supplementary light source to ensure image acquisition quality.

[0025] S102: Perform internal and external parameter calibration on the industrial camera, establish the camera imaging geometric model, and eliminate radial and tangential distortion of the lens; S103: Perform preprocessing operations on the calibrated image to improve image quality; Furthermore, the preprocessing operation described in S103 includes image grayscale conversion using the RGB three-channel weighted average method, noise reduction processing using Gaussian filtering, and contrast enhancement processing.

[0026] Furthermore, the W-HRNet network model described in S3 includes The basic network HRNet consists of two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3. Each convolutional layer is followed by BN (batch normalization) and ReLU (corrected linear unit activation function) operations to reduce the resolution of the input image to 1 / 4 of the original image, thereby reducing the computational cost of the network. Parallel multi-resolution architecture: It consists of four branches, each of which uses multiple 3×3 convolutions. After convolution, Batch Normalization (BN) and ReLU operations are performed. Each branch corresponds to a different resolution, and the number of feature map channels in each branch increases progressively. Feature exchange and fusion are achieved between branches through upsampling and downsampling. Specifically, low-resolution feature maps are upsampled using bilinear interpolation, while high-resolution feature maps are downsampled using convolutions with a stride of 2 and a kernel size of 3×3. Features of different resolutions are added pixel by pixel after fusion and then activated by ReLU.

[0027] Furthermore, the decoder of the underlying network HRNet employs a progressive feature fusion method, including the following steps: S301: Upsample the 32x downsampled features (i.e., the lowest resolution features) and use transposed convolution (ConvTranspose2d) to upsample them to a 16x scale; concatenate the upsampled 32x features with the original 16x features (torch.cat) to form preliminary fused features; S302: Upsample the 16x downsampled features and use transposed convolution to increase their resolution to 8x; concatenate the upsampled 16x features with the original 8x features (torch.cat) to form further fused features; S303: Retain the features of three independent paths: the path with 32x downsampling features upsampled to 8x scale, the path with 16x downsampling features upsampled to 8x scale, and the original 8x downsampling feature path. Concatenate the features of the three independent paths (torch.cat) to form a multi-scale fused feature. S304: In the process of multi-scale fusion feature fusion, WTConv2d is used in the last_layer to enhance the spatial detail expression capability, while CBAM (attention mechanism) is combined in the first branch x[0] and the second branch x[1] to improve the semantic information extraction capability; S305: Use the optimized model as the initial model and iteratively train it on a new sample set until the model performance reaches the expected effect.

[0028] Furthermore, the resolutions corresponding to the four branches of the parallel multi-resolution architecture are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original wood image, respectively. Under the hrnetv2_w18 configuration, the number of feature map channels of each branch are 18, 36, 72 and 144, respectively.

[0029] Furthermore, the feature exchange and fusion includes the following steps: S3-1: Two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3 are used. The first convolutional layer converts the input three-channel RGB to 64 channels, while the second convolutional layer keeps the input 64 channels unchanged. Batch normalization (BN) and ReLU operations are performed after the convolution. After the first convolutional layer, the input image is added to the convolutional output to form the first residual connection. After the second convolutional layer, no additional residual connections are added to simplify the network structure and maintain the stability of feature propagation.

[0030] S3-2: After passing the output of S3-1 through three bottleneck modules (1×1-3×3-1×1 convolution, number of channels 64→256), it forms 2 / 3 / 4 parallel branches in stages. Each stage is downsampled and the number of channels is increased through transition layers. Specifically: Entering Stage 2, two parallel branches are formed: the first branch maintains 1 / 4 resolution and 18 channels, while the second branch downsamples to 1 / 8 resolution and 36 channels through a transition layer. Each branch uses the BasicBlock module for feature extraction. Entering Stage 3, three parallel branches are formed: a third branch is added based on Stage 2, downsampling to 1 / 16 resolution and 72 channels through a transition layer. The resolution of the first two branches remains unchanged, with 18 and 36 channels respectively. Entering Stage 4, four parallel branches are formed: a fourth branch is added based on Stage 3, downsampling to 1 / 32 resolution and 144 channels through a transition layer. The resolution of the first three branches remains unchanged, with 18, 36, and 72 channels respectively. S3-3: In each stage, different branches exchange and fuse features through upsampling and downsampling.

[0031] For high-resolution to low-resolution feature fusion, a 3×3 convolution with a stride of 2 is used for downsampling; for low-resolution to high-resolution feature fusion, a transposed convolution is used for upsampling to high resolution. Features of different resolutions are fused by concatenation (torch.cat), and the fused features are then activated using ReLU.

[0032] Furthermore, the model training described in S3 includes the following steps: S3A: Parameter initialization settings: During training, the image size was adjusted to 360×640. The training epochs were set to 100, the initial learning rate to 4e-3, and the batch size to 4. The learning rate was adjusted using either the step or cosine algorithm, with a minimum learning rate of 0.01 times the initial learning rate. During training, either the Adam optimizer (initial learning rate 5e-4) or the SGD optimizer (initial learning rate 4e-3) could be used for parameter optimization, with a momentum coefficient of 0.9. The network structure employed a single residual connection, transposed convolutional upsampling, and feature concatenation strategy, with the last convolutional layer using wavelet transform (WTConv2d).

[0033] S3B: Model Training: Input the wooden window dataset into the W-HRNet network model, and choose whether to perform frozen training. During frozen training, freeze the backbone network parameters and only fine-tune other parts. S3C: Parameter Update: Employs the cross-entropy loss function, selects whether to use Focal Loss to address class imbalance, and supports setting different weight coefficients for different classes ([0.1931, 0.8503, 5.9156, 0.6768]); calculates the loss value between the prediction result and the label, and updates the network parameters through backpropagation; The formula for calculating the cross-entropy loss function is as follows:

[0034] in: Indicates the number of categories; This represents the weight coefficient of the i-th category; Indicates the true label; This indicates the probability that the model predicts the pixel belongs to the i-th class; When using Focal Loss, the cross-entropy loss function becomes:

[0035] in: It is a modulation factor used to reduce the weight of easily classified samples and increase the weight of difficult-to-classify samples; S3D: Output Model: The weights of the W-HRNet network model are saved every five epochs and evaluated on the training set; training stops when the number of training epochs reaches a set value, and the W-HRNet network model weight file is saved in the logs folder.

[0036] Furthermore, the accuracy evaluation described in S4 is calculated using precision, exact rate, recall, score, and mean intersection-over-union ratio, as shown in the following formulas:

[0037]

[0038]

[0039]

[0040]

[0041] in: ACC stands for accuracy. TP represents the number of pixels correctly classified as the target category; TN represents the number of pixels correctly classified as non-target categories; FP represents the number of pixels that were misclassified as the target category; FN represents the number of pixels misclassified as non-target categories; Precision, or P, represents accuracy. Recall, or R, represents the recall rate. F1 represents a fraction; mIoU represents the average crossover-union ratio; k represents the category.

[0042] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0043] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks, characterized in that: The method includes the following steps: S1: Acquire the original image of the wooden doors and windows in the area to be tested and perform image preprocessing; S2: Based on the structural features of the wooden door and window frames in the area to be tested, determine the range and number of sample annotations, and select the original sample images from the preprocessed wooden door and window images in S1 for human-computer interactive annotation to construct an image-label pair dataset; S3: Perform multi-class labeling on the frame structure of wooden doors and windows, randomly divide the labeled sample dataset into training set, validation set and test set according to a certain proportion, and import the training set and validation set into the improved W-HRNet network model for training and optimization. S4: Perform contour segmentation of the inner and outer frames of wooden doors and windows on the W-HRNet network model trained and optimized in S3, extract the model results and perform calculations to obtain the final result, and then evaluate the accuracy of the final result through the test set. If the accuracy requirement is met, proceed to S5; otherwise, return to S2 to check the sample drawing and adjust and optimize each network model until the requirement is met. S5: Apply the W-HRNet network model that passed the accuracy evaluation in S4 to the detection of the inner and outer frame contours of wooden doors and windows in actual production. Input the image of the wooden door and window to be tested into the model for inference, obtain the segmentation results of the inner and outer frame contours of the wooden door and window, and output the segmentation results in the form of contour coordinate point sequence. S6: Based on the contour coordinate point sequence output by S5, calculate the geometric feature parameters of the inner and outer frames of wooden doors and windows, and store the calculation results in the database to realize the automated detection of the contours of the inner and outer frames of wooden doors and windows.

2. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 1, characterized in that: S1 includes the following steps: S101: Use an industrial camera to acquire images of wooden doors and windows in the area to be tested; S102: Perform internal and external parameter calibration on the industrial camera, establish the camera imaging geometric model, and eliminate radial and tangential distortion of the lens; S103: Perform preprocessing operations on the calibrated image.

3. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 2, characterized in that: The preprocessing operations described in S103 include image grayscale conversion using the RGB three-channel weighted average method, noise reduction using Gaussian filtering, and contrast enhancement.

4. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 1, characterized in that: The W-HRNet network model described in S3 includes The basic network HRNet consists of two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3. After each convolutional layer, BN and ReLU operations are performed to reduce the resolution of the input image to 1 / 4 of the original image. Parallel multi-resolution architecture: It contains four branches, each of which uses multiple 3×3 convolutions. After convolution, BN and ReLU operations are performed. Each branch corresponds to a different resolution, and the number of feature map channels in each branch increases layer by layer. Feature exchange and fusion are achieved between branches through upsampling and downsampling.

5. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 4, characterized in that: The decoder of the underlying network HRNet employs a progressive feature fusion method, including the following steps: S301: Upsample the 32x downsampled features and use transposed convolution to increase their resolution to 16x; concatenate the upsampled 32x features with the original 16x features to form preliminary fused features; S302: Upsample the 16x downsampled features and use transposed convolution to increase their resolution to 8x; then concatenate the upsampled 16x features with the original 8x features to form further fused features; S303: Retain the features of three independent paths: the path with 32x downsampling features upsampled to 8x scale, the path with 16x downsampling features upsampled to 8x scale, and the original 8x downsampling feature path. The features of the three independent paths are then concatenated to form a multi-scale fused feature. S304: In the process of multi-scale fusion feature fusion, WTConv2d is used in the last_layer to enhance the spatial detail expression capability, and CBAM is combined in the first branch x[0] and the second branch x[1] to improve the semantic information extraction capability. S305: Use the optimized model as the initial model and iteratively train it on a new sample set until the model performance reaches the expected effect.

6. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 5, characterized in that: The four branches of the parallel multi-resolution architecture correspond to resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original timber image, respectively. Under the hrnetv2_w18 configuration, the number of feature map channels for each branch are 18, 36, 72, and 144, respectively.

7. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 6, characterized in that: The feature exchange and fusion includes the following steps: S3-1: Use two consecutive convolutional layers with a stride of 2 and a kernel size of 3×3, and then perform BN and ReLU operations after convolution; S3-2: After passing through three bottleneck modules, the output of S3-1 is divided into 2 / 3 / 4 parallel branches in stages. Each stage is downsampled and the number of channels is increased through the transition layer. S3-3: In each stage, different branches exchange and fuse features through upsampling and downsampling.

8. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 1, characterized in that: The model training described in S3 includes the following steps: S3A: Parameter initialization settings: S3B: Model Training: Input the wooden window dataset into the W-HRNet network model and choose whether to freeze the training. S3C: Parameter Update: Employs the cross-entropy loss function, selects whether to use Focal Loss to address class imbalance, and supports setting different weight coefficients for different classes; calculates the loss value between the prediction result and the label, and updates the network parameters through backpropagation; The formula for calculating the cross-entropy loss function is as follows: in: Indicates the number of categories; This represents the weight coefficient of the i-th category; Indicates the true label; This indicates the probability that the model predicts the pixel belongs to the i-th class; When using Focal Loss, the cross-entropy loss function becomes: in: It is a modulation factor used to reduce the weight of easily classified samples and increase the weight of difficult-to-classify samples; S3D: Output Model: The weights of the W-HRNet network model are saved every five epochs and evaluated on the training set; training stops when the number of training epochs reaches a set value, and the W-HRNet network model weight file is saved in the logs folder.

9. The method for detecting the inner and outer frame dimensions of wooden doors and windows based on semantic segmentation networks according to claim 1, characterized in that: The accuracy evaluation described in S4 is calculated using precision, exact rate, recall, score, and mean crossover ratio, as shown in the following formulas: in: ACC stands for accuracy. TP represents the number of pixels correctly classified as the target category; TN represents the number of pixels correctly classified as non-target categories; FP represents the number of pixels that were misclassified as the target category; FN represents the number of pixels misclassified as non-target categories; Precision, or P, represents accuracy. Recall, or R, represents the recall rate. F1 represents a fraction; mIoU represents the average crossover-union ratio; k represents the category.

Citation Information

Patent Citations

  • Deep hole surface defect detection and CV size calculation method

    CN118037664A

Cited By

  • Semantic segmentation-based ore drawing machine stockpiling detection method and system

    CN121475049A