X-ray single projection imaging wood type intelligent identification method and system and medium
By combining X-ray phase-contrast single-projection imaging with the improved EfficientNetV2 deep learning model, the problems of low efficiency and inconvenient equipment in wood species identification are solved, achieving non-destructive, efficient, and accurate wood species identification.
Patent Information
- Application Number
- CN202511218597.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies for wood species identification suffer from low efficiency, high destructiveness, and susceptibility to surface treatments. In particular, traditional X-ray CT scanning is time-consuming and the equipment is not portable, making it difficult to achieve non-destructive and efficient identification.
By employing X-ray phase-contrast single-projection imaging combined with an improved EfficientNetV2 deep learning model, single-projection images of different positions on the wood are acquired through multiple translations. These images are then preprocessed and used to train the deep learning model. The attention mechanism is optimized to improve feature extraction capabilities, thereby achieving non-destructive and efficient identification.
It achieves efficient and non-destructive identification of wood species with high accuracy, adapts to different thicknesses and resolutions, and is suitable for wood samples of common furniture thicknesses, significantly improving identification efficiency.
Smart Images

Figure CN121095653A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the fields of X-ray imaging, agricultural technology, cultural heritage protection, image processing and artificial intelligence, in particular to an X-ray single projection imaging wood species intelligent identification method, system and medium, which is based on phase-contrast X-ray single projection imaging and is used for efficient and non-destructive wood species intelligent identification. BACKGROUND
[0002] Accurate identification of wood species is of great importance in various academic and commercial fields, including cultural heritage protection, import and export inspection and quarantine, and paleontological research. In paleontological research, identification of wood or charcoal samples helps to gain a deeper understanding of the historical evolution of forests. Similarly, in archaeology and art history, wood material identification enhances the understanding of the history, use and optimal preservation methods of cultural relics. In addition, with the growing demand for high-value wood in furniture, flooring and handicrafts, illegal trade practices in the market have increased, i.e. using inferior materials to replace profits, which highlights the need for efficient and reliable wood identification technology.
[0003] Currently, the identification of wood species mainly depends on macroscopic identification on site and microscopic identification by professionals in the laboratory, for example, see [S. Carlquist, Comparative wood anatomy: systematic, ecological, and evolutionary aspects of dicotyledon wood. Springer Science & Business Media, 2013]. This process requires a lot of professional knowledge, a long identification period, and destructive testing of the original material. See [X. Pan, Z. Yu, and Z. Yang, "A Multi-Scale Convolutional Neural Network Combined with a Portable Near-Infrared Spectrometer for the Rapid, Non-Destructive Identification of Wood Species," Forests, vol. 15, no. 3, p. 556, 2024.], near-infrared spectral imaging can accurately identify wood species by reflecting the molecular group or elemental information inside the wood. However, when analyzing processed wood products, it becomes challenging to distinguish the signal from the wood itself from the signal from the processing-related substances. Chemical additives used in the wood processing process interfere with the spectral signal, leading to potential inaccurate identification. With the advancement of deep learning algorithms and computer hardware, the feature extraction capability of deep learning-based intelligent recognition of sequence and image signals has been developed. This meets the requirements of efficient and real-time wood identification. Intelligent recognition systems based on computer vision have made significant progress in general image recognition. Combining artificial intelligence algorithms for non-destructive testing of wood has become a trend.
[0004] In recent years, many algorithms have been developed for identifying wood. Artificial intelligence has been combined to identify wood from visible light images [D. Herrera-Poyatos et al., “Deep Learning methodology for the identification of wood species using high-resolution macroscopic images,” in 2024 International Joint Conference on Neural Networks (IJCNN), 2024: IEEE, pp. 1-8.]. This visible light-based identification method, although convenient in some cases, is highly susceptible to surface treatment due to its high dependence on surface information. Therefore, they are easily misled by treatment methods such as staining, veneering, and painting, making accurate identification of counterfeit wood products complex.
[0005] To overcome these challenges, non-destructive methods that capture internal structural information of wood are crucial. X-rays, due to their high penetration ability, can non-destructively detect the internal microstructure of wood. However, traditional absorption contrast imaging has difficulty in forming clear images of wood texture structure, as wood is mainly composed of low-Z elements. In recent years, X-ray phase contrast imaging methods have been developed to achieve high-contrast imaging of the internal microstructure of weakly absorbing samples.
[0006] Although some studies have explored wood identification using X-ray computed tomography (CT) [Z. Zhang et al., “A-RepVGG: Research on Classification Algorithms based on Deep Learning and Wood CT Images,” 2024.], there are some practical limitations to this technology. Compared with X-ray single projection imaging, CT is very time-consuming, requiring multi-angle data acquisition to reconstruct three-dimensional images. This process involves a large amount of data processing, resulting in prolonged scanning time. In contrast, single projection imaging can be completed in just a few tens to a few hundred milliseconds, significantly improving identification efficiency, especially when dealing with large volume samples. In terms of device portability, CT systems are usually bulky and complex, requiring installation in professional laboratories or inspection sites. This setup poses challenges for on-site inspections or inspections of large wood products. In contrast, single projection X-ray imaging systems are relatively compact and flexible, facilitating easier deployment in different environments. SUMMARY
[0007] The present application aims to provide an X-ray single projection imaging wood species intelligent identification method, system and medium, so as to realize non-destructive and efficient wood species identification.
[0008] In order to achieve the above-mentioned purpose, the present application provides a wood species intelligent identification method based on phase contrast X-ray single projection imaging, comprising:
[0009] S1: different wood is used as a sample, and an X-ray phase contrast imaging system is used to obtain single projection images of different positions of the sample by multiple translations, as an image data set;
[0010] S2: the single projection images are preprocessed to obtain a preprocessed image data set;
[0011] S3: the preprocessed image data set is used to train a deep learning model; the deep learning model is an improved EfficientNetV2 model, and the improved EfficientNetV2 model is based on the traditional EfficientNetV2 model with FusedMBConv blocks and MBConv blocks, and a global attention module is connected in series after the convolution operation of each FusedMBConv block, which is used to perform channel and spatial attention weighting on the feature map, and the parameters and compression ratio are optimized, thereby reducing the original calculation amount; in each MBConv block, a high-efficiency channel attention mechanism module is connected in series while the depth separable convolution is retained, which is used to enhance the inter-channel attention;
[0012] S4: a sample to be tested is obtained, and a single projection image of one position is taken, which is preprocessed, and the preprocessed image is input into the trained deep learning model for sample classification.
[0013] The X-ray phase contrast imaging system comprises a micro-focus X-ray source, a sample translation stage and a plane detector arranged in sequence, the sample translation stage is used to install the sample, and the sample is wood.
[0014] The X-ray source, the sample translation stage and the plane detector are coaxially arranged, and the sample translation stage and the plane detector can move along the optical axis; the shape of the sample is processed into a wood board with different thicknesses, the direction of the cut surface is according to the conventional use direction of the wood, and the surface of the sample is fixed on the sample translation stage of the X-ray phase contrast imaging system in parallel to the detector panel.
[0015] In the step S1, the obtained projection images are taken as an image data set, wherein two samples of each kind of wood are assigned as a test data set, and the rest are assigned as a training data set and a validation data set; in the step S3, the preprocessed training data set and the validation data set are used to train and validate the deep learning model, and the trained and validated deep learning model is used for sample classification.
[0016] The preprocessing includes flat field and dark field correction, edge trimming, and random cropping.
[0017] The loss function of the deep learning model is set as a cross-entropy loss; during model training, an Adam optimizer with a dynamic learning rate is adopted, and the initial learning rate is 0.0001, and the learning rate is reduced from 0.0001 to 0.00001.
[0018] In the step S4, the classification result obtained by classifying the sample includes wood species and unknown species; the wood species intelligent identification method further includes a step S5 of generating a wood species identification based on the classification result and warning for unknown species.
[0019] In another aspect, the present application provides a wood species intelligent identification system based on phase contrast X-ray single projection imaging, comprising: an X-ray phase contrast imaging system configured to take wood as a sample and obtain single projection images of the sample; an image preprocessing module for preprocessing the single projection images; a deep learning classification module loaded with a trained deep learning model, wherein the deep learning model adopts an improved EfficientNetV2 model and is used for sample classification, the improved EfficientNetV2 model is based on the traditional EfficientNetV2 model with FusedMBConv blocks and MBConv blocks, and after the convolution operation of each FusedMBConv block, a global attention module is connected in series for channel and spatial attention weighting of the feature map, and the parameters and compression ratio thereof are optimized to reduce the original calculation amount; in each MBConv block, a high-efficiency channel attention mechanism module is connected in series while the depth separable convolution is retained, for enhancing the inter-channel attention; and a result output module generates a wood species identification according to the classification result obtained by classifying the sample and warns for unknown species.
[0020] In another aspect, the present application provides a computer readable storage medium storing a computer program, which realizes the wood species intelligent identification method based on phase contrast X-ray single projection imaging as described above when the computer program is executed by a processor.
[0021] The wood species intelligent recognition method based on phase contrast X-ray single projection imaging of the application improves the recognition accuracy and recognition efficiency by combining X-ray single projection imaging with an improved deep learning network model. By using an optimized attention mechanism to enhance the EfficientnetV2 deep learning model, complex texture features can be effectively extracted from a single X-ray projection image, thereby improving recognition efficiency, recognition accuracy and computational efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a flowchart of the wood species intelligent recognition method based on phase contrast X-ray single projection imaging of the application.
[0023] Figure 2 is a measurement principle diagram of the wood species intelligent recognition method based on phase contrast X-ray single projection imaging of the application.
[0024] Figure 3 is a network architecture diagram of the improved EfficientNetV2 model.
[0025] Figure 3 In the figure, Conv: standard convolution; Fused-MBConv: a kind of fused and optimized MBConv (Mobile Inverted Bottleneck Convolution) module; ECA-MBConv: a kind of MBConv (Mobile Inverted Bottleneck Convolution) module combined with ECA (Efficient Channel Attention) attention mechanism; Depthwise Conv: depthwise convolution; ECA: Efficient Channel Attention attention mechanism module; Residual: residual layer; GAP: global average pooling (Global Average Pooling); FC: fully connected layer (Fully Connected Layer); GAM: global attention mechanism module (Global Attention Module); Input Feature: input feature map; Output Feature: output feature map; Pooling: pooling layer; ReLU: rectified linear unit (Rectified Linear Unit) activation function; Sigmoid: a kind of activation function. DETAILED DESCRIPTION
[0026] The application will be further described in connection with specific embodiments. It should be understood that the following embodiments are only used to illustrate but not to limit the scope of the application.
[0027] The wood species intelligent identification method based on phase contrast X-ray single projection imaging of the application combines X-ray single projection imaging with an improved deep learning model (EfficientNetV2 network architecture), realizes efficient and non-destructive identification by capturing the internal microstructure texture features of wood, replaces the traditional method of relying on surface features or time-consuming CT scanning, and can realize accurate classification of wood species only with a single X-ray projection image, and can realize intelligent identification non-destructively and efficiently. The attention mechanism of the EfficientnetV2 deep learning model is improved, which significantly enhances the ability of the model to identify complex wood textures, thereby realizing higher identification accuracy. The method is suitable for wood samples of common furniture thickness, and the change in thickness does not affect the identification accuracy.
[0028] Because wood is mainly composed of low-Z materials, the phase contrast imaging of X-rays is beneficial to enhance the contrast of the internal texture of wood. When X-rays pass through an object, the phase term is usually 2 to 3 orders of magnitude larger than the absorption term, which indicates that the sensitivity of phase contrast X-ray imaging is much higher than that of traditional absorption imaging. The use of X-ray phase change to enhance the contrast of weakly absorbing materials (such as internal structures of wood) is superior to traditional absorption imaging.
[0029] As shown in Figure 1 The wood species intelligent identification method based on phase contrast X-ray single projection imaging of the application includes the following steps:
[0030] Step S1: Different woods are used as samples respectively, and an X-ray phase contrast imaging system is used to obtain single projection images of different positions of the samples by multiple translations, as an image data set;
[0031] After obtaining the single projection images of the samples at different positions, the projection images have internal texture features of the wood. Therefore, by combining with deep learning, accurate classification of wood can be realized only with a single X-ray projection image, and the method can be implemented non-destructively and efficiently. That is, after the database (i.e., the image data set) is built, a single projection image of a sample to be measured at any position can be obtained to identify what kind of wood it is.
[0032] As shown in Figure 2 The wood species intelligent identification method based on phase contrast X-ray single projection imaging of the application uses an X-ray phase contrast imaging system to obtain single projection images. It supports single projection imaging, and the data acquisition time is ≤100 ms.
[0033] The X-ray phase contrast imaging system comprises a microfocus X-ray source 10, a sample translation stage 20 for mounting a sample 21, which is wood, and a flat detector 30 arranged in sequence. The X-ray source 10, the sample translation stage 20 and the flat detector 30 are coaxially arranged, and the sample translation stage 20 and the flat detector 30 are movable along the optical axis to adjust the magnification of the imaging system.
[0034] The microfocus X-ray source 10 is preferably a microfocus transmission tungsten target X-ray tube (X-ray WorX GmbH, XWT-225-THE Plus) with a focal spot of 2 pm to ensure the spatial coherence of the phase contrast imaging.
[0035] The flat detector 30 is an X-ray flat panel detector comprising a CsI scintillator (DALSA 6K) and a complementary metal-oxide-semiconductor (CMOS) detector (scintillator in front of the CMOS and fiber faceplate) to achieve large field of view and high resolution imaging. The flat detector 30 has an expandable field of view of up to 114 mm x 146 mm, i.e. a pixel array of 2304 pixel x 2940 pixel, with a single pixel size of 49.5 pm.
[0036] The microfocus X-ray source 10 is fixed, and the sample translation stage 20 is used for four-dimensional motion of X, Y, Z (three-dimensional translation) and W (one-dimensional rotation).
[0037] To meet the training and testing requirements of deep learning neural networks, several experimental woods are selected, such as Dalbergia barbata, Dalbergia cochinchinensis, Dalbergia grandiflora, Dalbergia hivaoana, Dalbergia nigra, Dalbergia odorifera, Dalbergia pinnata, Dalbergia sissoo and Pinus sp. These species are widely used in furniture manufacturing, construction and cultural heritage. Among them, Dalbergia grandiflora and Dalbergia odorifera belong to the genus Dalbergia. They are two pairs of species that are very similar to each other but have significant price differences on the market. The shape of the sample is machined into several wood boards with thicknesses ranging from 5 to 15 mm, and the direction of the sample section is according to the direction commonly used for wood, i.e. cut in the longitudinal direction. The obtained single projection images are used as an image dataset, in which two samples of each wood are selected and assigned as a test dataset, and the rest are assigned as a training dataset and a validation dataset. In this way, it is ensured that these test images come from separate wood samples, which are completely different from the wood samples corresponding to the images in the training and validation datasets.
[0038] In order to make the probability distribution of the collected image dataset as reasonable as possible, the surface of the sample is fixed parallel to the detector panel on the sample translation stage 20 of the X-ray phase contrast imaging system to collect X-ray phase contrast images with appropriate spatial resolution magnified by the imaging system geometry.
[0039] The image obtained by X-ray projection is a complex texture structure superimposed by multi-layer information of the board in thickness. There is no obvious microstructure difference between different samples, which makes it extremely difficult to distinguish directly by manual visual inspection. In this embodiment, the size of the sample is much larger than the field of view of the area detector 30, so every time a projection image is taken, the sample translation stage 20 moves the sample perpendicular to the X-ray direction by a certain distance, and then takes the next projection image. In this way, step-by-step scanning is performed to generate thousands of projection images for each sample to ensure that as much area position as possible is collected.
[0040] The step S1 can also include data augmentation of the image dataset by using at least one of various processing methods such as horizontal / vertical flipping, random rotation, brightness and contrast adjustment, flipping, etc., to achieve cross-modal data adaptation.
[0041] In this embodiment, when acquiring the single projection image of the sample, the imaging conditions are as follows:
[0042] The voltage of the microfocus X-ray source 10 is 20-225kV, and the current is 20-1000μA;
[0043] The distance between the sample and the microfocus X-ray source 10 is 110mm, and the distance between the area detector 30 and the sample is 300mm; the data acquisition time is ≤100ms.
[0044] In this embodiment, the geometric magnification is adjusted by the distance from the sample to the light source and the distance from the detector to the light source point, and the geometric magnification is currently maintained at about 3. The imaging distance is relatively fixed, so that a fixed geometric magnification is maintained to meet the conditions when the image dataset is taken. Different sizes of samples are collected by translation to collect single projection images at different positions of the sample.
[0045] Step S2: pre-processing the single projection image to obtain a pre-processed image dataset;
[0046] Wherein, in order to better focus on the characteristic texture information of the sample projection image and eliminate other interference factors, image preprocessing is necessary. The preprocessing includes flat field and dark field correction, edge trimming, random cropping.
[0047] The step S1 also includes collecting a dark field image without a sample and without X-rays and a flat field image without a sample and with X-rays for flat field and dark field correction of the projection image.
[0048] Flat field and dark field correction: aims to ensure that the generated image is best prepared for subsequent analysis. Flat field and dark field correction is applied to minimize the impact of background lighting and detector characteristic changes on the projection image. The formula for flat field correction is:
[0049]
[0050] where I c (x,y) denotes the corrected projection image, I(x,y) denotes the original projection image, I f (x,y) is a flat-field image, I d (x,y) is a dark-field image.
[0051] Edge cropping: refers to edge cropping of the wood sample, cutting off the part of the texture adjacent to the edge of the wood block, only retaining the internal texture, in order to eliminate the influence of the edge of the wood block due to cutting, damage or edge thickness variation and other factors. These irrelevant details will interfere with the deep learning model, resulting in inaccurate wood species identification. By focusing the attention of the deep learning model on the internal texture of the wood, edge cropping can improve the recognition accuracy. In the present invention, edge cropping uses a conventional threshold method rather than a model, the reason being that the accuracy of the cropping is not particularly high.
[0052] Random cropping: Direct projection images are usually large, resulting in information redundancy, which affects the efficiency of deep learning and recognition. Therefore, image random cropping refers to randomly cropping a specific size image (512x512) from different positions within a large size image to generate a pre-processed image dataset. Thus, when detecting, the image of the sample to be detected is input, the image preprocessing includes automatically performing image random cropping, the size of the image is still 512x512, and the cropped image is then recognized to give a final recognition result. Random cropping is used to ensure that all processed images have consistent sizes, which is crucial for effectively training neural networks, as it requires input data to have uniform shape and size.
[0053] Step S3: training the deep learning model using the pre-processed image dataset;
[0054] where the pre-processed training dataset and validation dataset are used to train and validate the deep learning model, and the trained and validated deep learning model is used for sample classification. During model training, the data of two samples are shielded and not used for training, but are used as a test dataset.
[0055] As Figure 3As shown, the deep learning model is an improved EfficientNetV2 model (namely I-EfficientNetV2), which is constructed by the following way: on the basis of the traditional EfficientNetV2 model with FusedMBConv block and MBConv block, a global attention module (GAM) is connected in series after the convolution operation of each FusedMBConv block, which is used for channel and spatial attention weighting of the feature map, and the parameters and compression ratio thereof are optimized, thereby reducing the original calculation amount; while the depth separable convolution is retained in each MBConv block, an efficient channel attention mechanism module (ECA) is connected in series, which is used for enhancing the inter-channel attention.
[0056] Therefore, the improved EfficientNetV2 model comprises a convolution layer, three optimized MBConv modules (Fused-MBConv) combined with a global attention module (GAM), three MBConv modules combined with an efficient channel attention mechanism module (ECA) and an AFDS layer (namely an average pooling layer, a fully connected layer, a Dropout layer and a Softmax layer) connected in sequence. MBConv refers to Mobile Inverted Bottleneck Convolution.
[0057] It should be noted that in the present application, the main purpose of replacing SE with ECA is to reduce model complexity, reduce parameter amount and floating point operation amount, and at the same time enhance inter-channel attention and improve model performance. However, in the first three MBConv modules, SE is replaced with a global attention module (GAM) instead of an ECA module, which not only retains the original channel information expression ability of the SE module in the shallow layer, avoids performance degradation due to premature simplification of attention mechanism, but also increases the global spatial information expression ability. And by optimizing the GAM parameters and compression ratio, the original calculation amount is reduced and the calculation efficiency is improved. The ECA module realizes local cross-channel interaction through one-dimensional convolution, which may not be able to fully capture the global channel relationship when the number of channels is small, thereby affecting the performance.
[0058] By replacing the SE (Squeeze-and-Excitation) channel attention module of the EfficientNetV2 model with an ECA module, the present application solves the feature confusion problem caused by texture superposition, significantly improves the feature extraction ability for complex superimposed textures, and at the same time reduces the calculation complexity.
[0059] The ECA module realizes feature weight distribution through a local cross-channel interaction strategy, adopts one-dimensional convolution to dynamically capture cross-channel interaction, and avoids dimension reduction.
[0060] In order to meet the intelligent identification requirements of the complex texture structure in the wood projection image, the neural network used in the application is based on the EfficientnetV2 [M.X. Tan and Q.V. Le, "EfficientNetV2: Smaller Models and Faster Training," in International Conference on Machine Learning (ICML), Electr Network, Jul 18-24 2021, vol. 139, SAN DIEGO: Jmlr-Journal Machine Learning Research, in Proceedings of Machine Learning Research, 2021, pp. 7102-7110.] network, and by appropriately improving the attention mechanism thereof, an efficient channel attention mechanism called ECA [Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, "ECA-Net: Efficient channel attention for deep convolutional neural networks," in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2020, pp. 11534-11542.] is introduced into the EfficientnetV2 network, and the identification efficiency and accuracy thereof can be further improved. The ECA is a local cross-channel interaction strategy without dimension reduction. The improved network structure is as shown in Figure 1
[0061] The AFDS layer refers to an average pooling layer, a fully connecting layer, a dropout layer, and a softmax layer. The average pooling layer compresses the extracted feature map information into a one-dimensional vector. The dropout layer is used to prevent overfitting by randomly deactivating a portion of neurons (i.e., setting their weights and biases to zero) during the training phase, thereby enhancing the robustness of the network. The softmax layer maps the one-dimensional output vector to a value between 0 and 1, producing a result similar to a probability distribution. The output value S i According to formula (2), the mapping is as follows:
[0062]
[0063] where e i is the value of the i-th neuron before mapping, and the maximum value of i is the number of categories of the network.
[0064] Since the network is designed to solve the classification problem of multiple types of wood samples, the loss function of the deep learning model is set as the cross entropy loss. The loss function is shown in formula (3):
[0065]
[0066] where Y is the true distribution of the data set, y is the predicted distribution, j is the index of the wood sample type, and M is the number of wood sample types in the library. When the loss function reaches the minimum value, it indicates that the y distribution is close to the Y distribution, and the network training is complete.
[0067] During model training, the Adam optimizer with dynamic learning rate is used, with an initial learning rate of 0.0001 and a learning rate decreasing from 0.0001 to 0.00001.
[0068] Subsequently, the model prediction ability is verified by selecting wood samples of different species from the pre-processed validation data set to test the performance of the network model. For example, the model verification result is as follows: 100% accuracy is achieved for 7 out of 9 types of wood. Only slight confusion occurs between two very similar species, Dalbergia cochinchinensis and Dalbergia hivaoana, but the overall identification accuracy is still over 98.5%. The test results show that the proposed method can achieve high recognition accuracy when applied to new wood samples.
[0069] Accuracy = True Positive / (True Positive + False Positive); Recall = True Positive / (True Positive + False Negative); True Positive: the image input into the prediction network is certain wood, and the model also predicts it as certain wood. False Positive: the image input into the prediction network is not certain wood, but the model incorrectly classifies it as certain wood. False Negative (FN): the image input into the prediction network is certain wood, but the model incorrectly classifies it as other wood.
[0070] It can be concluded that the network model performs well in terms of prediction accuracy when identifying different species of wood samples, and can be applied in practice. In addition, the present application is tested at different resolutions, verifying that the method of the present application has a certain range of tolerance for low-resolution images, i.e. single projection data collected at low magnification (corresponding to lower imaging spatial resolution, with a reduction in resolution of 16 times) can still be effectively classified using the trained deep learning model described above. At the same time, the deep learning model can maintain high recognition accuracy on wood of different thicknesses, which has considerable advantages in practical applications; in the wood processing industry, due to inconsistent cutting or different processing requirements, the thickness of wood often changes, and despite these differences, the method of the present application consistently achieves accurate intelligent identification of wood species; therefore, in non-destructive testing scenarios, such as inspecting wooden artifacts in cultural heritage protection - where the thickness of the wood may be unknown or highly variable - the method of the present application is still reliable, and this result highlights the strong adaptability of the model to changes in wood thickness.
[0071] Step S4: obtaining a sample to be tested and taking a single projection image at one position, pre-processing the image, and inputting the pre-processed image into the trained deep learning model for classification of the sample.
[0072] The classification result obtained by classifying the sample includes the wood species and the unknown species.
[0073] The multi-position projection probability fusion algorithm is used to quantitatively obtain the fusion confidence of multiple projection positions, and to improve the classification accuracy of the wood species.
[0074] The multi-position projection probability fusion algorithm is to collect different projection images of the same wood sample at different positions, respectively predict the classification results (i.e. obtain different classification results and corresponding probabilities), and statistically obtain the fusion confidence Psim according to the classification results and corresponding probabilities predicted by all projection images.
[0075] The formula of the fusion confidence Psim is:
[0076]
[0077] wherein j is the index of the sample library class, k is any test sample image, N is the total number of test sample images, and y[j][k] is the prediction distribution from the deep learning network.
[0078] The identification of the unknown class in step S4 includes calculating the fusion confidence Psim of the unknown class with each wood species in the image dataset. If none of the fusion confidences Psim meets the threshold condition, it is determined as an unknown class and the database expansion process is triggered.
[0079] Step S5: generating wood species identification based on the classification result and warning for unknown classes.
[0080] In another aspect, the present application provides a wood species intelligent identification system based on phase contrast X-ray single projection imaging, comprising: an X-ray phase contrast imaging system configured to take wood as a sample and obtain a single projection image of the sample; an image preprocessing module for preprocessing the single projection image; a deep learning classification module loaded with a trained deep learning model, wherein the deep learning model uses an improved EfficientNetV2 model and is used for classifying the sample, and the improved EfficientNetV2 model is based on the traditional EfficientNetV2 model with FusedMBConv blocks and MBConv blocks, and after the convolution operation of each FusedMBConv block, a global attention module is connected in series for channel and spatial attention weighting of the feature map, and the parameters and compression ratio thereof are optimized to reduce the original calculation amount; in each MBConv block, a high-efficiency channel attention mechanism module is connected in series while the depth separable convolution is retained, for enhancing the inter-channel attention; and a result output module generates wood species identification according to the classification result obtained by classifying the sample and warns for unknown classes.
[0081] In another aspect, the present application further provides a computer readable storage medium storing a computer program, which, when executed by a processor, implements the wood species intelligent identification method based on phase contrast X-ray single projection imaging described above.
[0082] Embodiment one:
[0083] The selected experimental samples were collected from 40 different species of wood, including: Dalbergia bariensis, Dalbergia cochinchinensis, Dalbergia latifolia, Dalbergia melanoxylon, Pterocarpus macrocarpus, Pterocarpus erinaceus, Pterocarpus erinaceus, Pterocarpus datura, Pterocarpus angolanus, Pterocarpus santalinus, Pterocarpus indicus, Dalbergia retusa, Cassia fistula, Cassia fasciata, Cassia fasciata 'Blanc' ... These species are widely used in furniture making, construction, and cultural heritage, but some woods look similar and are easily confused, leading to counterfeiting and the sale of inferior goods. For example, Dalbergia bariensis and Dalbergia cochinchinensis both belong to the rosewood family, with similar density and appearance, making them easy to confuse; Pterocarpus macrocarpus and Pterocarpus erinaceus both belong to the Pterocarpus genus and are similar in appearance; various types of Guibourtia genus wood have similar appearances and can only be identified through laboratory testing; many of these are species that are very similar to each other but have significant price differences in the market. The samples were machine-processed into several squares with thicknesses ranging from 5-15 mm to meet the training and testing needs of deep learning neural networks. The experimental parameters were set as follows: tube voltage 50-100kV, current 100μA, exposure time 100ms, distance from X-ray source to sample (SDD) 110mm, and distance from sample to detector 300mm. This setting achieved sufficient phase contrast enhancement, resulting in an effective pixel size of 13.3μm after geometric magnification of the image. The resulting projected image visualized the distribution of sieve tubes, vessels, and fiber structures within the wood. The wood sample was fixed on the sample stage parallel to the detector panel. The image obtained by X-ray projection is a complex texture structure composed of multiple layers of information superimposed on the thickness of the wood. There were no obvious microstructural differences between different samples, making direct visual identification extremely difficult. Because the sample size was much larger than the detector's field of view, the sample stage was moved a certain distance after each projected image was captured before capturing the next image. This gradual scanning generated thousands of projected images for each sample, ensuring that as many areas as possible were collected. In addition, dark-field images without samples and without X-rays, as well as flat-field images without samples but with X-rays, were acquired for projected image correction. Image preprocessing is necessary to better focus the features of the sample projection image and eliminate other interfering factors. This includes flat-field and dark-field correction, edge trimming, and random cropping (the cropped image is a specific size). This preprocessing sequence aims to ensure that the generated image is optimally prepared for subsequent analysis. Flat-field correction is applied to minimize the impact of artifacts introduced by changes in background illumination and detector characteristics on the projection image. The flat-field correction algorithm given in Equation (1) is used to remove the background, where I... c(x,y) denotes the corrected projection image, I(x,y) denotes the original projection image, I f (x,y) is a flat-field image, I d (x,y) is a dark-field image.
[0084]
[0085] In addition, the wood samples are edge-trimmed, which cuts off the part of the texture adjacent to the edge of the wood block and only keeps the internal texture. This is to eliminate the influence of the wood block edge due to cutting, damage or edge thickness variation, etc. These irrelevant details will interfere with the deep learning model and lead to inaccurate wood species identification. By focusing the model's attention on the internal texture of the wood, edge trimming can improve the accuracy of identification. In addition, the direct projection image is usually large, resulting in information redundancy, which affects the efficiency of deep learning and identification. Therefore, images of a specific size (512x512) are randomly cropped from different positions within the large size image to generate the training set. This preprocessing ensures that all processed images have consistent sizes, which is crucial for effectively training neural networks, as it requires input data to have uniform shape and size.
[0086] The collected projection images are divided into: training data set, validation data set and test data set. The projection images of two samples in each wood species (320 groups of image modules are collected for each sample) are used as the test data set, and the remaining samples (320 groups of image modules are collected for each sample) are used as the training and validation data set, which ensures that these test images come from separate wood samples, which are completely different from the images used in the training and validation data set. Then the wood sample projection images in the training data set and the validation data set are used to train the neural network, and the wood sample projection images in the test data set are used to test the generalization performance of the trained model. In order to meet the intelligent recognition requirements of the complex texture structure in the wood projection image, the neural network used in the present application is based on the EfficientnetV2 [M. X. Tan and Q. V. Le, "EfficientNetV2: Smaller Models and Faster Training," in International Conference on Machine Learning (ICML), Electr Network, Jul 18-24 2021, vol. 139, SAN DIEGO: Jmlr-Journal Machine Learning Research, in Proceedings of Machine Learning Research, 2021, pp. 7102-7110.] network, and by appropriately improving its attention mechanism, an efficient channel attention mechanism called ECA [Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, "ECA-Net: Efficient channel attention for deep convolutional neural networks," in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2020, pp. 11534-11542.] is introduced into the EfficientnetV2 network, and the recognition efficiency and accuracy can be further improved. The ECA module is a local cross-channel interaction strategy without dimension reduction. The improved network structure is as shown in Figure 1 .
[0087] The AFDS layers of the head are AveragePooling, FullyConnecting, Dropout, and Softmax, where AveragePooling compresses the extracted feature map information into a one-dimensional vector. Dropout is a technique used in training phase of the network to prevent overfitting by randomly deactivating a portion of neurons (i.e., setting their weights and biases to zero), thereby enhancing the robustness of the network. Softmax maps the one-dimensional vector of the final output to values between 0 and 1, producing a result similar to a probability distribution, where the output value S i According to formula (2), the mapping is:
[0088]
[0089] where e i is the value of the i-th neuron before mapping, and the maximum value of i is the number of classes of the network.
[0090] Generally, a trained model achieves a high classification accuracy on the training set. However, due to overfitting, it often performs poorly on the test set. To avoid overfitting and improve the generalization ability of the machine learning model, image enhancement techniques are used in this embodiment. During the wood sample image enhancement process, different types of projection images are artificially created through various processing methods or their combinations such as brightness enhancement, contrast enhancement, flipping, etc. Random brightness enhancement and contrast adjustment roughly simulate the projection images obtained under different exposure times and energies; the random brightness step involves randomly increasing or decreasing pixel values by 100-500. The random contrast adjustment step redistributes the gray scale histogram of the image, and the redistributed minimum and maximum gray scale values are randomly set to 90%-110% of their original values. These enhancement techniques are randomly applied to the images in the training dataset, effectively increasing the diversity of the training images, ensuring that even images that are initially very similar become more diverse after enhancement. The projection images in the test dataset remain unchanged to accurately evaluate the generalization ability of the model. The validation dataset data is used to evaluate the convergence efficiency of the model during training.
[0091] The wood types involved in the training dataset and the validation dataset are a total of 40 classes. Since the network is designed to solve the classification problem of multi-type wood samples, the loss function is set to CrossEntropy Loss, as shown in formula (3).
[0092]
[0093] Wherein, Y is the real distribution of the data set, y is the predicted distribution, j is the index of the wood sample type, and M is the number of wood sample types in the library. The loss function reaches the minimum value, indicating that the y distribution is close to the Y distribution, indicating that the network training is completed.
[0094] The prediction ability of the model was then tested by selecting samples of different species from the pre-processed test data set to test the performance of the network model. The test data set consists of new wood samples that have never been used in the network training process, and is used to confirm the universality of the network algorithm in wood species identification after training. In this embodiment, the model verification result is: 100% accuracy is achieved for 12 of the 40 types of wood. Only slight confusion occurs between two very similar species, Dalbergia cochinchinensis and Dalbergia hivaoana, but the overall identification accuracy is still more than 98.5%. That is, when the input image pixel size of the improved EfficientNetV2 model (i.e. I-EfficientNetV2 model) is 53.2 μm, the recognition accuracy is ≥98%. The test results show that the method can achieve high recognition accuracy when applied to new wood samples.
[0095] In order to verify the effectiveness of the model in actual application scenarios, in addition to testing different types of wood board samples, the study also tested actual furniture samples - high-value wood bed cabinet samples. This test collected a bed cabinet made of Dalbergia hivaoana, and the test used the bed cabinet drawer panel and the upper top panel of the bed cabinet for testing. By collecting different projection images at different positions, the projection collection method is the same as the collection method of the wood board standard sample in the paper. The collected X-ray projection data is directly used as the test data set, and the previously trained network model is used to identify the wood species. The accuracy of the identification result is 98.3%.
[0096] In summary, the wood species intelligent identification method based on phase contrast X-ray single projection imaging of the present application improves the recognition accuracy and recognition efficiency by combining X-ray single projection imaging with an improved deep learning network model. By using an optimized attention mechanism to enhance the feature extraction ability of the network model, the recognition accuracy of the network model is improved, and the recognition efficiency is also improved. The EfficientNetV2 deep learning model effectively extracts complex texture features from a single X-ray projection image, thereby improving recognition accuracy and computational efficiency. The validation results show that the model achieves an overall recognition accuracy of over 98% in the classification of 40 wood species, with some wood species achieving 100% accuracy. The proposed method can avoid some limitations of existing wood identification methods. Traditional methods based on macroscopic and microscopic visible light features and near-infrared imaging largely rely on surface characteristics. Therefore, when the surface of the wood changes due to factors such as staining, veneering, or wear and tear, their accuracy will decrease. In contrast, the present method uses the penetrating power of X-rays to extract internal microscopic structure features, providing a more comprehensive and reliable basis for species identification. In addition, compared to CT-based techniques that require multi-angle scanning and costly 3D reconstruction, this method only needs to collect X-ray single projection images, thereby significantly improving efficiency. This simplification reduces complexity and resource consumption while improving adaptability to practical applications. Even with reduced image resolution, the model maintains an identification accuracy of over 98%. This finding has important practical significance as it indicates that high-resolution imaging is not necessary for high recognition accuracy. In cases where high-resolution imaging is limited by device constraints or considerations of image acquisition efficiency, the proposed method remains effective due to its optimized algorithm framework.
[0097] The above is only a preferred embodiment of the present application, not to limit the scope of the present application, the above embodiment of the present application can be made various changes. Any simple, equivalent changes and modifications made in accordance with the content of the claims and the description of the present application, all fall within the scope of the present application. The present application is not described in detail, all are conventional technical content.
Claims
1. A method for intelligent identification of wood species based on phase-contrast X-ray single-projection imaging, characterized in that, include: Step S1: Using different types of wood as samples, the X-ray phase contrast imaging system was used to repeatedly translate and acquire single-projection images of different positions of the samples, which were then used as an image dataset. Step S2: Preprocess the single projection image to obtain a preprocessed image dataset; Step S3: Train the deep learning model using the preprocessed image dataset; the deep learning model is an improved EfficientNetV2 model. Based on the traditional EfficientNetV2 model with FusedMBConv and MBConv blocks, the improved EfficientNetV2 model adds a global attention module after the convolution operation of each FusedMBConv block. This module performs channel and spatial attention weighting on the feature map and optimizes its parameters and compression ratio, reducing the original computational cost. While retaining depthwise separable convolutions in each MBConv block, an efficient channel attention mechanism module is added to enhance inter-channel attention. Step S4: Acquire the sample to be tested and take a single-projection image of one position, preprocess it, and input the preprocessed image into the trained deep learning model for sample classification.
2. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 1, characterized in that, The X-ray phase contrast imaging system includes a microfocus X-ray source, a sample translation stage, and a surface detector arranged in sequence. The sample translation stage is used to mount the sample, which is wood.
3. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 2, characterized in that, The X-ray source, sample translation stage, and surface detector are arranged coaxially, and the sample translation stage and surface detector can move along the optical axis; The sample was shaped into a wooden board of varying thickness, with the cut surfaces oriented in the direction conventionally used for wood. The sample surface was fixed to the sample translation stage of the X-ray phase contrast imaging system with the detector panel parallel to the surface.
4. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 1, characterized in that, In step S1, the acquired projected images are used as an image dataset, in which two samples of each type of wood are selected and assigned as a test dataset, and the rest are assigned as a training dataset and a validation dataset. In step S3, the preprocessed training dataset and validation dataset are used to train and validate the deep learning model, which is then used to classify samples.
5. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 1, characterized in that, The preprocessing includes flat and dark field correction, edge trimming, and random cropping.
6. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 1, characterized in that, The loss function of the deep learning model is set to cross-entropy loss; during model training, the Adam optimizer with dynamic learning rate is used, with an initial learning rate of 0.0001 and the learning rate decreasing from 0.0001 to 0.00001.
7. The intelligent wood species identification method based on phase-contrast X-ray single-projection imaging according to claim 1, characterized in that, In step S4, the classification results obtained from the sample classification include wood species and unknown species; The intelligent identification method for wood species further includes step S5: generating wood species identifiers based on the classification results and issuing warnings for unknown species.
8. A smart wood species identification system based on phase-contrast X-ray single-projection imaging, characterized in that, include: The X-ray phase-contrast imaging system is configured to use wood as a sample to acquire a single-projection image of the sample. The image preprocessing module is used to preprocess single-projection images; The deep learning classification module is equipped with a trained deep learning model. This model uses an improved EfficientNetV2 model for sample classification. The improved EfficientNetV2 model, based on the traditional EfficientNetV2 model with FusedMBConv and MBConv blocks, adds a global attention module after the convolution operation of each FusedMBConv block. This module performs channel and spatial attention weighting on the feature map and optimizes its parameters and compression ratio, reducing the original computational cost. Furthermore, while retaining depthwise separable convolutions in each MBConv block, an efficient channel attention mechanism module is added to enhance inter-channel attention. The results output module generates wood species identifiers based on the classification results obtained from the sample classification and provides early warnings for unknown species.
9. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the intelligent wood species identification method based on phase-contrast X-ray single-projection imaging as described in any one of claims 1-7.