A method for measuring the crease degree of primary cured tobacco based on unsupervised depth estimation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2026-08-11
AI Technical Summary
前一种方法由于需要测量烟叶初烤前后的面积,而烟叶初烤过程是在不同地区进行的,导致测量过程复杂、效率低、成本高;后一种方法受到自身携带的叶脉、叶片纹理影响,测量结果不准确
[0032] This invention provides a method for measuring the wrinkle degree of newly cured tobacco based on unsupervised depth estimation. It determines the degree of wrinkle by recognizing the unevenness of the tobacco leaf surface wrinkles as different depth information at each pixel on the depth map. Compared to manual measurement, this eliminates sensory differences and, compared to ordinary image texture acquisition methods, depth estimation provides a three-dimensional representation of the tobacco leaf's wrinkle degree. Furthermore, unsupervised depth estimation only requires training with the geometric constraints of image sequences or stereo images as supervision, eliminating the need for expensive labeled data. This method is low-cost, widely applicable, and easy to operate. This invention establishes a describable method for measuring the wrinkle degree of newly cured tobacco and utilizes the wrinkle degree as a factor to differentiate tobacco leaf parts.
Smart Images

Figure CN116433598B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco leaf physical index measurement technology, specifically to a method for measuring the wrinkleness of freshly cured tobacco based on unsupervised depth estimation. Background Technology
[0002] The wrinkling of newly cured tobacco leaves is a change in appearance caused by the contraction of cell tissue due to changes in temperature and moisture. Since the physicochemical properties of cell tissues differ in different parts of the tobacco leaf during growth, different parts will exhibit varying degrees of wrinkling after curing. The national standard for flue-cured tobacco, GB 2635-1992, states that the grade of tobacco leaves is mainly determined by three factors: part of the leaf, grade, and color. Among these, differences in appearance between parts of the tobacco leaf can be identified through features such as vein pattern, leaf shape, degree of wrinkling, and color. In the field of machine vision, features such as texture, edge, wrinkling depth, and RGB values can be extracted accordingly. Therefore, wrinkling is one of the important factors in determining the part of the tobacco leaf and plays a crucial role in tobacco leaf grading.
[0003] Currently, the measurement of wrinkles in newly cured tobacco mainly relies on visual judgment by tobacco grading experts. However, differences in sensory perception among experts lead to variations in the grading process. Existing intelligent measurement methods primarily measure the area of the tobacco leaf before and after wrinkling, using the rate of change between the two for quantitative characterization, or indirectly represent the wrinkles by studying related characteristics and elements. The former method requires measuring the area of the tobacco leaf before and after initial curing, which takes place in different regions, resulting in a complex, inefficient, and costly process. The latter method is affected by the leaf veins and texture, leading to inaccurate results.
[0004] In summary, while the existing methods for measuring the wrinkles of first-cured tobacco can achieve the measurement objective, their measurement results are not accurate enough, and the methods need further improvement. Summary of the Invention
[0005] This invention provides a method for measuring the wrinkle degree of newly cured tobacco based on unsupervised depth estimation. First, the leaf surface wrinkling morphology that occurs after curing is defined as the "wrinkle degree of newly cured tobacco," and the evaluation parameter for the wrinkle degree is set as the surface pixel arithmetic mean deviation R. a Maximum difference R between surface pixels z Then, by improving the unsupervised depth estimation network model, the wrinkle depth map of the prepared tobacco leaf samples is extracted. The Roberts operator is then used to extract wrinkle texture information from the depth map, and the surface pixel arithmetic mean deviation R is obtained. a Maximum difference R between surface pixels z And using the surface pixel arithmetic mean deviation Ra Maximum difference R between surface pixels z Clustering is performed to determine the location of tobacco leaves.
[0006] To achieve the above-mentioned technical objectives, the present invention is implemented through the following technical solution:
[0007] A method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation includes the following steps:
[0008] S1: The leaf wrinkling morphology that occurs during the initial curing of tobacco leaves is called "initial curing tobacco wrinkling degree," and the evaluation parameter for the initial curing tobacco wrinkling degree is set as the surface pixel arithmetic mean deviation R. a Maximum difference R between surface pixels z ;
[0009] S2: Prepare tobacco leaf samples by differentiating parts based on the degree of wrinkling of the first-cured tobacco;
[0010] S3: Obtain a frontal local image of a tobacco leaf sample and construct a training dataset and a test dataset for unsupervised depth estimation;
[0011] S4: Construct an improved unsupervised depth estimation network model, and use this model to process local images of tobacco leaves to obtain the depth map of wrinkles in the first-cured tobacco.
[0012] S5: Use the Roberts operator to extract the texture of the initial flue-cured tobacco wrinkles in the depth map, and obtain the surface pixel arithmetic mean deviation R of the initial flue-cured tobacco wrinkle texture map. a Maximum difference R between surface pixels z ;
[0013] S6: Use the Birch algorithm to calculate the arithmetic mean deviation R of surface pixels. a Maximum difference R between surface pixels z Clustering was performed to classify the tobacco leaf parts corresponding to the degree of wrinkling in the initial curing into upper, middle, and lower categories (B, C, X).
[0014] Preferably, the "wrinkle degree of initially cured tobacco" refers to the geometric shape characteristics of the peaks and valleys with spacing on the surface of the initially cured tobacco leaf; the surface pixel arithmetic mean deviation R a The surface pixel arithmetic mean deviation R represents the arithmetic mean of the absolute values of the differences between the pixel values at each point in the depth map and the measurement reference point. It reflects the overall situation of the changes in the depth of wrinkles in the depth map. a A larger value indicates a greater degree of peak-valley fluctuation in wrinkle formation, and vice versa; the maximum surface pixel value R... zThe depth map represents the highest or lowest peak and valley pixel values on the folds, reflecting the largest folds in a tobacco leaf image. The tobacco leaf samples are composed of samples from various regions of Yunnan Province, categorized by tobacco leaf grading experts based on the fold degree of the newly cured tobacco, including three parts: upper, middle, and lower (B, C, X). The partial images of the front of the tobacco leaf are obtained by an industrial camera. These partial images are only the result of a single shot; in practical applications, multiple partial images of the tobacco leaf can be obtained to represent the entire tobacco leaf. The improved unsupervised depth estimation network model introduces a convolutional modulation module and a normalization fusion module, and a skip connection is designed between the decoder and encoder. The fold depth map changes from bright yellow to dark blue based on the principle that the distance between the object and the camera in the depth map changes from near to far.
[0015] Preferably, the tobacco leaf samples are all determined by experts from the Yunnan Provincial Tobacco Quality Supervision and Inspection Station based on the factor of the wrinkle degree of the initial flue-cured tobacco. The samples are then captured under natural light conditions using an industrial camera of model AO-HD206 on a self-made tobacco leaf image acquisition platform. A total of 8,598 partial images of the front tobacco leaves at three different locations (B, C, and X) are collected, including 8,250 training images and 348 test images.
[0016] Preferably, the improved unsupervised depth estimation network model uses a U-shaped network encoding and decoding segmentation network as a reference. The network inputs the video stream from the actual scene into the convolutional layer to generate high-level encoded information, and then uses multi-scale fusion upsampling decoding to restore the video stream image resolution. The reprojection matrix is calculated using the video stream image, resampling is performed, and the error at a higher input resolution is calculated. The decoder is based on the Swing Transformer backbone network and integrates a convolutional modulation module and a multi-layer perceptron (MLP). In the decoding stage, Normalization-based Attention Module (NAM) is used to achieve adaptive selection of feature channels and supplement contextual information between the decoder and encoder through skip connections. In the upsampling process, NAM uses matching feature patches to pass the features lost during downsampling through skip connections, making up for the loss of detailed information in the decoded features, thereby improving decoding efficiency and effect, and making the reconstructed high-resolution image as accurate as possible.
[0017] Preferably, the encoder is constructed based on a SwingTransformer convolutional neural network. In each stage, a convolutional modulation layer is used instead of the basic attention encoding layer of the SwingTransformer. The essence of the convolutional modulation layer remains a self-attention layer, resulting in higher memory efficiency and better suitability for deep feature extraction in weakly textured scenes. Simultaneously, a multi-layer perceptron (MLP) based on grouped convolutional cross-location fusion is designed. The input features are viewed as composed of multiple patches of the same size. By regrouping and transforming, cross-location information exchange is achieved within each patch, thereby balancing the loss of feature edge information.
[0018] The encoder, drawing inspiration from Conv2Former, uses dynamic weights A, composed of feature query values Q and feature key values K, as input. It then extracts weights A directly from the input layer x using a convolutional layer, transforming the vector multiplication of A and V into a Hadamard matrix multiplication. For example... Figure 1 As shown in the convolutional modulation module, the attention weights A are extracted and activated through depthwise separable convolution and the GELU function. Similar to the feature weights in the convolutional layer, the weights are updated after each iteration. Since static weights do not undergo large-scale dynamic changes with the input values, this is more conducive to the convergence of the neural network. Therefore, the principle of the convolutional modulation layer is as follows:
[0019] A = DConv k×k (W1X),V=Conv 1×1 (W2X)
[0020] X′=A⊙V
[0021] In the formula, ⊙ represents the Hadamard product, W1 and W2 are the weight matrices of A and V, respectively, and Conv 1×1 DConv represents a 1×1 convolution. k×k This represents a depthwise separable convolution with kernel k, where X′∈R C×H×W The convolutional modulation module is the output feature. In the convolutional modulation operation, DConv... k×k Associating the feature space location (H,W) with all pixels within a k×k squared region, the computational complexity of the dynamic weight A is changed from exponential to linear superposition through convolutional modulation. When a larger convolutional kernel k is used, the area of focus of the static attention weight is slightly weaker than that of the dynamic weight, thus enabling the network to extract features better in weak texture backgrounds.
[0022] Because the shrinking of the attention region of the convolutional modulation module leads to the loss of information in the feature edge regions, the encoder also designed a multi-layer perceptron (MLP) based on grouped convolutional cross-location fusion. Borrowing some ideas from the MLP-Mixer, it treats the input features as composed of multiple patches of the same size, and achieves cross-location information exchange between each patch through regrouping and transformation, thereby balancing the loss of feature edge information. For example... Figure 1 As shown in the MLP module, for the MLP input feature x∈R C×H×W Grouping by p×p patches yields S features x∈R. C×S Then, dimensionality reduction is performed through grouped convolutions to reduce the computational complexity of the linear layers. At this point, the absolute position of each patch remains unchanged, but the linear layers will perform calculations vertically according to the patch. Finally, the features processed by the linear layers are restored and skip connections are made with the original features. The specific process is as follows:
[0023]
[0024] In the formula, S represents the number of feature groups, and p is the pixel size of the patch; Liner and GConv k×k These are linear layers and grouped convolutions, respectively.
[0025] Preferably, the decoder embeds a normalized fusion layer, which performs an efficient normalized attention (NAM) fusion between the previous layer features and the shallow features. By allocating attention weights, the shallow features are superimposed on the decoded features, thereby achieving the decoding supplement of detailed information.
[0026] Among them, the decoder is as follows Figure 2 As shown, x1 represents the deep feature input through the upsampling layer, x2 represents the shallow feature input through skip connections, and y represents the output feature of normalized fusion. Normalized fusion includes two branches: x1 features are compressed via a convolutional modulation module for subsequent fusion operations; x2 uses normalized attention to input shallow features according to channel weight ratios, thus enhancing important information in the shallow features. The implementation of normalized attention is as follows:
[0027]
[0028] Among them B in ∈R N×B×1×H×W Let N be single-channel feature matrices of batch size B, where x2∈R B×C×H×W The transformation yields μ B and σ B B inThe mean and standard deviation of a single mini-batch, ε is the minimum value to prevent collapse, γ i These are the eigenvalues of a single channel; the balanced eigenma matrix B can be obtained from the above formula. out and channel weight coefficient w i Unlike the original NAM, since the state inside the shallow features is no longer considered, the affine transformation parameters are removed here. Therefore, the normalization fusion process is as follows:
[0029] y = DConv 1×1 [CM(x1),Sigmiod(W γ BN(x2))]
[0030] In the formula, CM and Sigmoid are the convolutional modulation module and activation function, respectively; W γ It is all the weights of x2 w i The vector formed.
[0031] The beneficial effects of this invention are:
[0032] This invention provides a method for measuring the wrinkle degree of newly cured tobacco based on unsupervised depth estimation. It determines the degree of wrinkle by recognizing the unevenness of the tobacco leaf surface wrinkles as different depth information at each pixel on the depth map. Compared to manual measurement, this eliminates sensory differences and, compared to ordinary image texture acquisition methods, depth estimation provides a three-dimensional representation of the tobacco leaf's wrinkle degree. Furthermore, unsupervised depth estimation only requires training with the geometric constraints of image sequences or stereo images as supervision, eliminating the need for expensive labeled data. This method is low-cost, widely applicable, and easy to operate. This invention establishes a describable method for measuring the wrinkle degree of newly cured tobacco and utilizes the wrinkle degree as a factor to differentiate tobacco leaf parts.
[0033] The method for measuring the wrinkle degree of newly flue-cured tobacco provided in this invention was used to test and classify a dataset of 348 tobacco leaf samples, including 118 upper tobacco (B), 116 middle tobacco (C), and 114 lower tobacco (X). This dataset was determined by tobacco grading experts in Yunnan Province based on the wrinkle degree factor. The classification of tobacco leaf parts using the wrinkle degree of newly flue-cured tobacco is highly effective, with a classification accuracy of 86.44% for upper tobacco (B), 93.1% for middle tobacco (C), and 89.94% for lower tobacco (X), resulting in an overall accuracy of 89.83%. These experimental results demonstrate that the unsupervised depth estimation method of this invention for obtaining the wrinkle degree of newly flue-cured tobacco is highly effective in determining the tobacco leaf part based on its wrinkle degree, providing a more objective method for this assessment. Attached Figure Description
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a diagram of the overall network framework;
[0036] Figure 2 It is a normalized fusion structure diagram;
[0037] Figure 3 This is an example image of a dataset of samples of freshly roasted tobacco with wrinkles;
[0038] Figure 4 This is a schematic diagram of the cross-section of the pleats of the first-cured tobacco.
[0039] Figure 5 This is a diagram showing the estimated depth of wrinkles in the first-cured tobacco.
[0040] Figure 6 This is an image showing the extracted texture of the folds in freshly roasted tobacco.
[0041] Figure 7 This is a graph showing the results of cluster analysis;
[0042] Figure 8 These are the experimental results of the method of this invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example 1
[0045] A method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation includes the following steps:
[0046] S1: The leaf wrinkling morphology that occurs during the initial curing of tobacco leaves is called "initial curing tobacco wrinkling degree," and the evaluation parameter for the initial curing tobacco wrinkling degree is set as the surface pixel arithmetic mean deviation R. a Maximum difference R between surface pixels z ;
[0047] S2: Prepare tobacco leaf samples by differentiating parts based on the degree of wrinkling of the first-cured tobacco;
[0048] S3: Obtain a frontal local image of a tobacco leaf sample and construct a training dataset and a test dataset for unsupervised depth estimation;
[0049] S4: Construct an improved unsupervised depth estimation network model, and use this model to process local images of tobacco leaves to obtain the depth map of wrinkles in the first-cured tobacco.
[0050] S5: Use the Roberts operator to extract the texture of the initial flue-cured tobacco wrinkles in the depth map, and obtain the surface pixel arithmetic mean deviation R of the initial flue-cured tobacco wrinkle texture map. a Maximum difference R between surface pixels z ;
[0051] S6: Use the Birch algorithm to calculate the arithmetic mean deviation R of surface pixels. a Maximum difference R between surface pixels z Clustering was performed to classify the tobacco leaf parts corresponding to the degree of wrinkling in the initial curing into upper, middle, and lower categories (B, C, X).
[0052] Example 2
[0053] Method for measuring the wrinkleness of freshly cured tobacco based on Example 1:
[0054] There is currently no unified definition for the wrinkle degree of early-cured tobacco in the tobacco industry. It is generally considered that the different wrinkling conditions that occur in different parts of the tobacco leaf after curing due to the physicochemical properties of cell tissues in different parts of the leaf during growth are considered as wrinkles in early-cured tobacco. The wrinkles in early-cured tobacco are mainly characterized by leaf folds and undulations. From a three-dimensional perspective, a coordinate system is established on the early-cured tobacco, with the horizontal axis as the x-axis, the vertical axis as the z-axis, and the undulation direction of the wrinkles as the y-axis. The wrinkle degree is then truncated along a direction parallel to the x-axis, and the wrinkle degree section is displayed in both the x and y directions. Figure 4 As shown, for ease of calculation, this paper refers to the geometric shape features of the peaks and valleys with spacing on the surface of the first-cured tobacco leaf as the first-cured tobacco wrinkle degree. Each point in the wrinkle degree depth map corresponds one-to-one with a pixel, and different pixel values are presented according to the distance of the target object from the camera. The pixel values of the wrinkle degree texture map generated from the depth map also generate different pixel values according to the first-cured tobacco wrinkle degree, thus reflecting the changes in the leaf surface. Therefore, the pixel values on the xz plane are obtained for calculating the first-cured tobacco wrinkle degree evaluation parameters. The position of the pixel value is represented by both the x-axis and z-axis, and the magnitude of the pixel value is represented by the y-axis. Since the shape characteristics of the first-cured tobacco wrinkles are similar to the microscopic shape of the surface roughness of industrial products, the method for determining the first-cured tobacco wrinkle degree refers to the parameter selection for surface roughness in the evaluation surface structure parameters and their numerical series of the Product Geometry Specification (GPS) GB / T1031-2009. This paper sets the evaluation parameter for the first-cured tobacco wrinkle degree as the surface pixel arithmetic mean deviation R. a Maximum difference R between surface pixelsz Surface pixel arithmetic mean deviation R a The surface pixel arithmetic mean deviation R represents the arithmetic mean of the absolute values of the differences between the pixel values at each point in the depth map and the measurement reference point. It reflects the overall situation of the changes in the depth of wrinkles in the depth map. a A larger value indicates a greater degree of peak and valley fluctuation in wrinkle formation, and vice versa; the maximum value R of surface pixels z The highest or lowest peak-to-valley pixel value in the depth map represents the maximum wrinkle value in an image of a tobacco leaf. Combining two initial flue-cured tobacco wrinkle assessment parameters reflects the overall wrinkle variation in initial flue-cured tobacco, which is the initial flue-cured tobacco wrinkle degree. Since the measurement base point for surface roughness of industrial products is zero, while the measurement base point for initial flue-cured tobacco wrinkle degree is the tobacco leaf plane, the position of each tobacco leaf's surface changes in the three-dimensional coordinate system. Therefore, the base point for each tobacco leaf is different and cannot be uniformly defined. Since the initial flue-cured tobacco wrinkle degree measurement assesses the degree of wrinkle variation, the base point for initial flue-cured tobacco wrinkle degree is set as the average surface pixel value. The fluctuation of the average pixel value indicates the degree of wrinkle.
[0055] The arithmetic mean deviation of surface pixels R a With the maximum value R of surface pixels z The calculation formula is shown in the formula below:
[0056]
[0057]
[0058]
[0059] In the formula, X represents the average pixel value; n represents the number of pixel values; X i |X| represents the pixel value at point i. max This represents the maximum pixel value in the image based on the average surface pixel count. Represents the pixel value at point i and the average pixel value on the surface. The gap between them.
[0060] Example 3
[0061] Method for measuring the wrinkleness of freshly cured tobacco based on Example 1:
[0062] Using a U-shaped network encoding and decoding segmentation network as a reference, the network inputs the video stream from the real-world scene into the convolutional layer to generate high-level encoded information, and then uses multi-scale fusion upsampling decoding to restore the video stream image resolution. The reprojection matrix is calculated using the video stream image, resampling is performed, and the error at a higher input resolution is calculated. The decoder is based on the SwingTransformer backbone network and integrates a convolutional modulation module and a multi-layer perceptron (MLP). During the decoding stage, Normalization-based Attention Module (NAM) is used to adaptively select feature channels and supplement contextual information between the decoder and encoder through skip connections. In upsampling, NAM uses matching feature patches to pass features lost during downsampling through skip connections, compensating for the loss of detailed information in the decoded features, thus improving decoding efficiency and quality, and making the reconstructed high-resolution image as accurate as possible.
[0063] The encoding layer employs a Swing Transformer-based convolutional neural network. In this invention, the encoding layer uses a convolutional modulation layer in each stage instead of the basic Swing Transformer attention encoding layer, but the essence of the convolutional modulation layer remains a self-attention layer. In conventional self-attention layers such as ViT and Swing Transformer, the self-attention mechanism involves the feature query value Q, the feature key value K, and the queried feature value V, with attention weights A = Softmax(QK). T When the input feature x∈R C×H×W When the calculated attention weight is A∈R C×C While A brings broad spatial information attention as a feature, its computational complexity also increases significantly with the increase of C. Borrowing ideas from Conv2Former, this invention uses the dynamic weights A, composed of Q and K, as input weights, and extracts weights A directly from the input layer x using a convolutional layer. Simultaneously, the vector multiplication of A and V is transformed into a Hadamard matrix multiplication. For example... Figure 1 As shown in the convolutional modulation module, the attention weights A are extracted and activated through depthwise separable convolution and the GELU function. Similar to the feature weights in the convolutional layer, the weights are updated after each iteration. Since static weights do not undergo large-scale dynamic changes with the input values, this is more conducive to the convergence of the neural network. Therefore, the principle of the convolutional modulation layer is as follows:
[0064] A = DConv k×k (W1X),V=Conv 1×1 (W2X)
[0065] X′=A⊙V
[0066] In the formula, ⊙ represents the Hadamard product, W1 and W2 are the weight matrices of A and V, respectively, and Conv 1×1 DConv represents a 1×1 convolution. k×k This represents a depthwise separable convolution with kernel k, where X′∈R C×H×W The convolutional modulation module is the output feature. In the convolutional modulation operation, DConv... k×k This approach associates the feature space location (H, W) with all pixels within a k×k squared region. Through convolutional modulation, the computational complexity of the dynamic weights A is transformed from an exponential increase to a linear summation. When using a larger convolutional kernel k, the area of focus for the static attention weights is slightly weaker than that of the dynamic weights. Therefore, this convolutional modulation module can establish relationships through convolution, exhibiting higher memory efficiency compared to self-attention, which is beneficial for deep feature extraction in weakly textured scenes.
[0067] Because the shrinking of the attention region of the convolutional modulation module leads to the loss of information in feature edge regions, this invention designs a multi-layer perceptron (MLP) based on grouped convolutional cross-location fusion. Borrowing some ideas from the MLP-Mixer, the input features are considered as composed of multiple patches of the same size. By regrouping and transforming these patches, cross-location information exchange is achieved within each patch, thereby balancing the loss of feature edge information. Figure 1 As shown in the MLP module, for the MLP input feature x∈R C×H×W Grouping by p×p patches yields S features x∈R. C×S Then, dimensionality reduction is performed through grouped convolutions to reduce the computational complexity of the linear layers. At this point, the absolute position of each patch remains unchanged, but the linear layers will perform calculations vertically according to the patch. Finally, the features processed by the linear layers are restored and skip connections are made with the original features. The specific process is as follows:
[0068]
[0069] In the formula, S represents the number of feature groups, and p is the pixel size of the patch; Liner and GConv k×k These are linear layers and grouped convolutions, respectively.
[0070] During network decoding, excessively deep networks often lead to the loss of shallow line feature information. Although high-level semantic information is concentrated in deep features, for some deep learning tasks that require shallow information, the positional contour information contained in the shallow layers is more important. For depth estimation, the relative positions between pixels are crucial, as this leads to further enrichment of the details of disparity edges. Therefore, this invention embeds a normalized fusion layer to decode features. The normalized fusion layer performs an efficient normalized attention (NAM) fusion between the features from the previous layer and the shallow features, superimposing the shallow features into the decoded features by assigning attention weights, thereby supplementing the decoding of detailed information. Figure 2 As shown, x1 represents the deep feature input through the upsampling layer, x2 represents the shallow feature input through skip connections, and y represents the output feature of normalized fusion. Normalized fusion includes two branches: x1 features are compressed via a convolutional modulation module for subsequent fusion operations; x2 uses normalized attention to input shallow features according to channel weight ratios, thus enhancing important information in the shallow features. The implementation of normalized attention is as follows:
[0071]
[0072] Among them B in ∈R N×B×1×H×W Let N be single-channel feature matrices of batch size B, where x2∈R B×C×H×W The transformation yields μ B and σ B B in The mean and standard deviation of a single mini-batch, ε is the minimum value to prevent collapse, γ i These are the eigenvalues of a single channel; the balanced eigenma matrix B can be obtained from the above formula. out and channel weight coefficient w i Unlike the original NAM, since the state inside the shallow features is no longer considered, the affine transformation parameters are removed here. Therefore, the normalization fusion process is as follows:
[0073] y = DConv 1×1 [CM(x1),Sigmiod(W γ BN(x2))]
[0074] Where CM and Sigmoid are the convolutional modulation module and activation function, respectively; W γ It is all the weights of x2 w i The vector formed.
[0075] Example 4
[0076] This invention utilizes flue-cured tobacco samples from various regions of Yunnan Province provided by the Yunnan Provincial Tobacco Quality Supervision and Inspection Station to construct an image dataset. The flue-cured tobacco samples used were all determined by experts from the Yunnan Provincial Tobacco Quality Supervision and Inspection Station based on the factor of wrinkle degree. The samples were acquired using an industrial camera (model AO-HD206) on a self-made tobacco leaf image acquisition platform under natural light conditions. The dataset includes 8598 partial images of the frontal tobacco leaves at three different locations within the same grade: B2F, C2F, and X2F. This includes 8250 images in the training set and 348 images in the test set. The test set contains 118 images of upper tobacco (B), 116 images of middle tobacco (C), and 114 images of lower tobacco (B). Examples of tobacco leaf samples are shown below. Figure 3 As shown in the figure. The method of this invention is used to process a local image of tobacco leaf folds, and the resulting image shows the estimated depth of folds in the initially cured tobacco. Figure 5 As shown; based on the color variation of the depth map of leaf wrinkles, which ranges from dark blue to bright yellow, the effect of extracting the wrinkle texture of the first-cured tobacco using the Roberts operator is shown in the image below. Figure 6 As shown; finally, the Birch method was used to cluster the wrinkle degree α of the wrinkle texture map, and the clustering results are as follows. Figure 7 As shown; the accuracy of classification based on wrinkle degree after obtaining the wrinkle texture of tobacco leaves using the method of this invention is as follows: Figure 8 As shown, the classification accuracy of the upper tobacco (B) reaches 86.44%, the classification accuracy of the middle tobacco (C) reaches 93.1%, and the classification accuracy of the lower tobacco (X) reaches 89.94%, with an overall accuracy of 89.83%, which can meet the grading requirements in practical applications.
[0077] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0078] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation, characterized in that, Includes the following steps: S1: The leaf wrinkling morphology that occurs during the initial curing of tobacco leaves is called "initial curing tobacco wrinkling degree," and the evaluation parameter for the initial curing tobacco wrinkling degree is set as the arithmetic mean deviation of surface pixels. Maximum difference between surface pixels ; S2: Prepare tobacco leaf samples by differentiating parts based on the degree of wrinkling of the first-cured tobacco; S3: Obtain a frontal local image of a tobacco leaf sample and construct a training dataset and a test dataset for unsupervised depth estimation; S4: Construct an improved unsupervised depth estimation network model and use this model to process local images of tobacco leaves to obtain the depth map of wrinkles in the first-cured tobacco. S5: Use the Roberts operator to extract the texture of the initial flue-cured tobacco wrinkles in the depth map and obtain the surface pixel arithmetic mean deviation of the initial flue-cured tobacco wrinkle texture map. Maximum difference between surface pixels ; S6: Use the Birch algorithm to calculate the arithmetic mean deviation of surface pixels. Maximum difference between surface pixels Clustering was performed to classify the flue-cured tobacco leaves into upper, middle, and lower sections based on the degree of wrinkling. The surface pixel arithmetic mean deviation The arithmetic mean of the absolute values of the differences between the pixel values at each point in the depth map and the measurement reference point; the maximum difference between surface pixels. This represents the highest or lowest peak or valley pixel value on the fold in the depth map; The improved unsupervised depth estimation network model takes the video stream from the real scene as input into the convolutional layer to generate high-level coding information, and then uses multi-scale fusion upsampling decoding to restore the video stream image resolution. The reprojection matrix is calculated using video stream images, and the error at higher input resolution is calculated after resampling. The encoder is based on the Swing Transformer backbone network and integrates the convolutional modulation module and the multilayer perceptron (MLP). During the decoding stage, the decoder uses Normalized Attention (NAM) to achieve adaptive selection of feature channels and supplements contextual information between the decoder and encoder through skip connections. In the upsampling stage, NAM passes the features lost during the downsampling process through matching feature patches in the form of skip connections to make up for the loss of detailed information of the decoded features. The encoder is built on the Swing Transformer convolutional neural network. In each stage, a convolutional modulation module is used to replace the Swing Transformer basic attention encoding layer. The essence of the convolutional modulation module is still a self-attention layer. At the same time, a multilayer perceptron based on grouped convolution and cross-position fusion is designed. The input features are viewed as consisting of multiple equally sized patches. By regrouping and transforming these patches, cross-positional information exchange is achieved, thereby balancing the loss of edge information. This applies to the MLP input features. according to Grouping patches of different sizes yields S features. Then, dimensionality reduction is performed through grouped convolutions to reduce the computational complexity of the linear layers. At this point, the absolute position of each patch remains unchanged, but the linear layers will perform calculations vertically according to the patch. Finally, the features processed by the linear layers are restored and skip connections are made with the original features. The specific process is as follows: In the formula, S represents the number of feature groups, and p is the pixel size of the patch; and These are linear layers and grouped convolutions, respectively.
2. The method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation according to claim 1, characterized in that, The encoder uses dynamic weights A, composed of feature query value Q and feature key value K, as input, and employs a convolutional layer to directly input the weights from the input layer. Weight A is extracted from V, and the vector multiplication of A and V is transformed into a Hadamard matrix product. The attention weight A is extracted and activated through depthwise separable convolution and the GELU function. Similar to the feature weights in the convolutional layer, the weights are updated after each iteration. The principle of the convolutional modulation module is as follows: In the formula, For Hadamard product, and These are two weight matrices, A and V. express convolution, This represents a depthwise separable convolution with kernel k. This represents the output characteristics of the convolution modulation module.
3. The method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation according to claim 1, characterized in that, The decoder embeds a normalized fusion layer, which performs normalized attention fusion on the features of the previous layer and the shallow features. By allocating attention weights, the shallow features are superimposed on the decoded features, thereby achieving the decoding supplement of detailed information.
4. The method for measuring the wrinkleness of newly cured tobacco based on unsupervised depth estimation according to claim 1, characterized in that, The normalization fusion of the decoder includes , Two branches and the fusion result y, This indicates that the deep features are input through the upsampling layer. y represents the shallow feature input through skip connections, and y represents the output feature of normalized fusion. in The features are compressed by the convolutional modulation module to facilitate subsequent fusion operations. Shallow features are input according to channel weight ratios using normalized attention; the implementation of normalized attention is as follows: in This represents N single-channel feature matrices with a batch size of B, where... Obtained by transformation; and express The mean and standard deviation of a single mini-batch. To prevent the minimum value from crashing, These are the eigenvalues of a single channel; the balanced eigenma matrix can be obtained from the above formula. and channel weight coefficient Unlike the original NAM, since the state inside the shallow features is no longer considered, the affine transformation parameters are removed here; therefore, the normalization fusion process is as follows: In the formula and These are the convolutional modulation module and the activation function, respectively. yes All weights The vector formed.