Method for obtaining region and confirming area of overlapping cut tobacco mask image
Through the improved Mask-RCNN network detection model, combined with DenseNet121 and U-FPN network, the problem of detection and segmentation of overlapping tobacco images is solved, and accurate calculation of tobacco component determination and tobacco quality evaluation are achieved.
Patent Information
- Application Number
- CN202211493859.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-11-25
AI Technical Summary
The prior art is difficult to accurately identify and segment overlapping tobacco wire images, which affects the calculation accuracy of the tobacco wire component doping ratio, especially the phenomenon of overlapping and stacking tobacco wires on the tobacco quality inspection line.
The improved Mask-RCNN network detection model is adopted, combined with DenseNet121 and U-FPN networks, and the anchor optimization parameters in the RPN proposed network are set, and the image feature extraction and target candidate box generation are achieved to accurately detect and segment the overlapping tobacco area.
The detection accuracy of overlapping tobacco images is improved, and the area proportion, length and width parameters of different tobacco wires can be accurately obtained, and the problems of overlapping tobacco wires in the determination of tobacco components and tobacco quality evaluation are solved.
Smart Images

Figure CN115861213B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tobacco image processing, and in particular to a method for acquiring and confirming the area of an overlapping tobacco mask image. Background Art
[0002] The guidelines for implementation of Articles 9 and 10 of the World Health Organization's Framework Convention on Tobacco Control (FCTC) require tobacco product manufacturers and importers to disclose the ingredients of tobacco products to government authorities, including the type of tobacco and the blending ratio of each type of tobacco. Tobacco manufacturers are also required to have equipment and methods for testing and measuring tobacco ingredients. The ratio of tobacco blends (leaf cuts, stem cuts, expanded leaf cuts, and reconstituted tobacco) in cigarettes has a significant impact on the smoke characteristics, physical indicators, and sensory quality of cigarettes. Therefore, high-precision and efficient tobacco type identification and component determination are of great significance for ensuring the quality of tobacco blending processes, homogeneous production, examining formula design, and verifying the authenticity of tobacco products.
[0003] Currently, tobacco component detection has been extensively studied, primarily using manual and instrumental methods. Manual sorting involves visually selecting the cut tobacco, stem cuts, expanded cuts, and reconstituted cuts from the blended tobacco, weighing and calculating the proportions of each type. This method is inefficient and subject to significant subjective influence. Instrumental methods include RGB analysis, hyperspectral imaging, near-infrared spectroscopy, thermal analysis, cigarette smoke analysis, anhydrous acetone analysis, and machine vision.
[0004] Among them, Kou Xiaoteng et al. measured the RGB mean values of tobacco powder made from different ratios of cut tobacco and stems, established a polynomial regression model between the stem-cut blending ratio and the RGB mean, and proposed a method for predicting the proportion of stem-cut components in cut tobacco based on RGB image processing. Mei Jifan et al. used the spectral data of all pixels in a sample's hyperspectral image to perform pixel-by-pixel component discrimination, and then used the average spectral data of all pixels to perform sample-by-sample component discrimination. They proposed a method for determining cut tobacco components based on hyperspectral imaging technology. Li Ruili et al. collected near-infrared spectral data from cut tobacco samples with different component ratios, established an infrared spectral model using the PLS method, and proposed a method for predicting the uniformity of cut tobacco blending based on infrared spectroscopy. Zhang Yaping et al. used thermogravimetric analysis to measure the similarity of the thermogravimetric reaction conversion rate curves of formulated cut tobacco samples taken at different time points. Based on the coefficient of variation between the similarities, they developed a method for calculating the uniformity of cut tobacco blending. Ye Hongyin et al. analyzed and compared mainstream smoke indicators and conventional chemical composition of cut tobacco under different blending methods, and concluded that the composition of cut tobacco blending varies. Lin Hui et al., using the characteristic of expanded tobacco (which has a much higher floating rate in anhydrous acetone than other types of tobacco), developed a method for determining the component ratios of expanded tobacco leaves. Dong Hao et al., using machine vision to acquire tobacco images, established a feature database and correlation equations based on pixel variance, contrast, entropy, angular second moment, and four texture eigenvalues in RGB and HSV color spaces for different tobacco types, allowing them to determine tobacco type. However, these detection methods suffer from various issues, including destructive testing, long testing cycles, demanding testing conditions, and incomplete detection of tobacco types.
[0005] In recent years, deep learning methods based on machine vision have provided advanced and efficient image processing solutions for image classification, object detection, and image segmentation. Currently, deep learning methods based on machine vision have been used for tobacco image classification. Gao Zhenyu et al. proposed a tobacco recognition method based on convolutional neural networks, targeting the structural differences between various tobacco types. However, the accuracy of the test set differed significantly from that of the training set, and the model exhibited a degree of overfitting, resulting in low generalization ability. Zhong Yu et al. established a recognition model based on a residual neural network and optimized the model's pre-trained weights, optimization algorithm, and learning rate. The results showed that the trained model achieved accuracy and recall rates exceeding 96%. Niu Qunfeng et al. optimized the ResNet50 network by adding a multi-scale structure, modifying the number of block stackings in the stage layer, and adjusting the focal loss function. Experimental results demonstrated a classification accuracy of 96.56%.
[0006] In the aforementioned studies of tobacco image classification methods, the image features of individual tobacco strands within a blended tobacco were analyzed to identify the tobacco type. However, in actual quality inspection lines, blended tobacco inevitably contains overlapping and stacked tobacco strands. Object detection and segmentation methods for overlapping and stacked tobacco strands are rarely studied. However, the identification and component determination of overlapping and stacked tobacco strands directly impacts the accuracy of calculating the blending ratio of tobacco components, making research of this nature crucial.
[0007] Currently, machine vision-based object detection and segmentation methods for overlapping images have been extensively studied in several fields. Dandan Wang et al. achieved accurate segmentation for apples in orchards, taking into account the overlapping and occlusion phenomena, the varying background colors of apples, and the similarity between unripe apples and background leaves. Using Mask-RCNN as the main network architecture, they added an attention mechanism to the backbone network to enhance feature extraction, achieving a mean average segmentation accuracy (MAP) of 91.7%. Yang Yu et al. achieved accurate segmentation and picking point location for strawberries with overlapping and hidden objects, as well as under varying lighting conditions. Using Mask RCNN as the overall algorithm architecture, combined with a visual localization algorithm for strawberry picking points, they achieved precise strawberry picking, achieving an average detection accuracy of 95.78%. Jiahao Qi et al. used SSD as the network architecture for cabin auxiliary equipment with overlapping and dense occlusion. They used a focal loss function to balance positive and negative samples in the dataset to improve the model's detection of densely occluded and overlapping cabin equipment, achieving a mean average recognition accuracy (MAP) of 78.95%. Daizhou Wen et al. used a CNN as the primary architecture to correlate the relationships between bubbles detected in two frames for precise segmentation, achieving 85% accuracy. Hao Wu et al. used a residual U-Net network to identify overlapping immunohistochemistry (IHC)-positive cells, detecting 86.04% of the overlapping cells.
[0008] The detection and segmentation of overlapping tobacco strands presents significant challenges due to the small size, diverse shapes, and complex physical and morphological characteristics of individual tobacco strands. In particular, the macroscopic differences between expanded and leaf tobacco strands are minimal, making machine vision-based single-strand identification and classification challenging. Furthermore, the four different types of tobacco strands can overlap in 24 different ways, and self-entanglement within the tobacco strands can easily become indistinguishable from the overlapping types. This poses significant challenges to the detection and segmentation of overlapping tobacco strands, as well as the subsequent calculation of their component areas. Summary of the Invention
[0009] In view of the above problems, the present invention is proposed to provide a method for acquiring regions and confirming areas of overlapping tobacco mask images that overcomes the above problems or at least partially solves the above problems.
[0010] To achieve the above object, the technical solution adopted by the present invention is:
[0011] In a first aspect, an embodiment of the present invention provides a method for obtaining and confirming the area of an overlapping tobacco mask image, comprising the following steps:
[0012] Acquire an image of tobacco to be detected;
[0013] The tobacco image is identified and segmented using a pre-trained improved Mask-RCNN network detection model; the improved Mask-RCNN network detection model uses a DenseNet121 and U-FPN network, and sets anchor optimization parameters in an RPN proposal network;
[0014] According to the detection results of the improved Mask-RCNN network detection model, the corresponding area ratio, length and width parameters of different tobacco leaves in the tobacco leaf image are determined.
[0015] Furthermore, the corresponding area ratio, length and width parameters of different tobacco shreds in the tobacco shred image are determined, including: obtaining a mask image of overlapping tobacco shreds based on the detection results, and drawing a fitted overlapping area; and calculating the corresponding area ratio, length and width parameters of different tobacco shreds respectively.
[0016] Furthermore, the calculation of the corresponding area proportions of different tobacco shreds includes:
[0017] Grayscale processing and binarization are performed on the mask image, and the number of tobacco outlines is calculated and counted to determine the area of the unobstructed tobacco object Area1 and the obstructed tobacco object Area2; the mask image is an overlapping image of the two tobacco shreds;
[0018] Performing a cyclic judgment on multiple contours of the obstructed tobacco, finding two contours based on their areas, and fitting the overlapping area of the tobacco based on the minimum rectangle inside the contours, the inscribed circle of the rectangle, and the circumtangent of the inscribed circle of the rectangle;
[0019] Generate a mask image for the overlapping area of tobacco obtained by fitting, perform mask operation on the unobstructed image contour, find the overlapping area, and determine the actual overlapping area of tobacco
[0020] Combined with the standard block for calculating pixel area, the actual area of the unobstructed tobacco object and the actual area of the obstructed tobacco object are calculated respectively.
[0021] Furthermore, the lengths of different tobacco pieces are calculated, including:
[0022] Performing calculus calculations on the unobstructed tobacco object and the obstructed tobacco object, performing segmented fitting center processing, and measuring along the tobacco centerline from the beginning to the end;
[0023] According to the overlapping area determined when calculating the corresponding area proportions of different tobacco shreds, the obscured tobacco object is divided into three segments for processing, and the pixel length of the obscured tobacco object is obtained;
[0024] For the unobstructed tobacco object, the center line is fitted to obtain the pixel length of the unobstructed tobacco object;
[0025] According to the respective pixel lengths and in combination with the pixel length calculation standard block, the actual length of the unobstructed tobacco object and the actual length of the obstructed tobacco object are calculated respectively.
[0026] Furthermore, the width of different tobacco shreds is calculated, including:
[0027] Performing calculus calculations on the unobstructed tobacco object and the obstructed tobacco object, respectively, performing segmented fitting of the maximum inscribed circle, measuring along the inscribed circle of the tobacco from the beginning to the end, and calculating the average value of the diameter;
[0028] According to the overlapping area determined when calculating the corresponding area proportions of different tobacco shreds, the obscured tobacco object is divided into three segments, and the width means of the three segments are averaged to obtain the pixel width of the obscured tobacco object;
[0029] For the unobstructed tobacco objects, the largest inscribed circle is fitted, and the mean of all inscribed circle diameters is used as the pixel width of the unobstructed tobacco objects;
[0030] According to the respective pixel widths, combined with the pixel width calculation standard block, the actual width of the unobstructed tobacco object and the actual width of the obstructed tobacco object are calculated respectively.
[0031] Furthermore, before the improved Mask-RCNN network detection model obtained through pre-training detects the tobacco image, the method further includes:
[0032] Acquire multiple types of overlapping cut tobacco images through a cut tobacco vibration experiment; the overlapping cut tobacco images are composed of at least any two of cut tobacco leaves, cut stems, expanded cut tobacco leaves, and reconstituted cut tobacco overlapped in different orders;
[0033] Processing the multiple types of overlapping tobacco images with an OpenCV algorithm to obtain cropped images;
[0034] The experimental image annotation tool sets labels for the cropped overlapping images and generates corresponding mask images as training datasets;
[0035] The improved Mask-RCNN network is iteratively trained using the training data set to obtain an improved Mask-RCNN network detection model.
[0036] Furthermore, the improved Mask-RCNN network detection model obtained through pre-training recognizes and segments the tobacco image, including:
[0037] Extracting image features corresponding to the tobacco image using the DenseNet121 and U-FPN networks;
[0038] The image features are used as input to the proposal network RPN to generate a series of target candidate boxes;
[0039] The series of target candidate frames are passed through FC and fully convolutional network FCN to complete the recognition and segmentation of overlapping tobacco images and output the segmented image.
[0040] Furthermore, the DenseNet121 and U-FPN networks are used to extract image features corresponding to the tobacco image, including:
[0041] Extracting feature information of the tobacco image using one convolutional layer and four feature extraction layers in the DenseNet121 network;
[0042] And perform pooling operation on the feature information extracted by each feature extraction layer to obtain four feature return values;
[0043] The four feature return values are respectively downsampled and horizontally connected to generate four first feature maps of different scales;
[0044] The four first feature maps of different scales are respectively input into the four feature layers of the U-FPN network for fusion to obtain image features corresponding to the tobacco image.
[0045] Furthermore, the four first feature maps of different scales are respectively input into the four feature layers of the U-FPN network for fusion to obtain image features corresponding to the tobacco image; including:
[0046] The first feature maps of the four different scales are respectively subjected to 1×1 convolution, fused with the up-sampled features, and then subjected to 3×3 convolution, and are input into the corresponding feature layer of the U-FPN network to obtain P2, P3, P4 and P5 features;
[0047] A bottom-up and rating feature reuse structure is added, and shallow layer information is sequentially transferred to the upper feature layer for fusion to obtain image features corresponding to the tobacco image.
[0048] Furthermore, a bottom-up and rating feature reuse structure is added to sequentially pass shallow layer information to the upper feature layer for fusion, thereby obtaining image features corresponding to the tobacco image, including:
[0049] The P2 feature is added to the P3 feature through two 3×3 convolutions and the first feature map C3 of the second scale is added through 3×3 convolution to obtain the fused enhanced P3 feature;
[0050] The enhanced P3 feature is added to the P4 feature through two 3×3 convolutions and the first feature map C4 of the third scale through 3×3 convolution to obtain the fused enhanced P4 feature;
[0051] The enhanced P4 feature is added to the P5 feature through two 3×3 convolutions and the first feature map C5 of the fourth scale through 3×3 convolution to obtain the fused enhanced P5 feature;
[0052] The enhanced P5 feature is added to the feature map C6 obtained by fusing the first feature maps of four different scales through two 3×3 convolutions to obtain the fused P6 feature.
[0053] Furthermore, the image features are used as input to the proposal network RPN to generate a series of target candidate boxes, including:
[0054] Set multiple reference anchor frames at each position of the fused enhanced P3, P4, P5 features and P6 features;
[0055] The proposal network RPN adjusts the anchor box according to the feature selection of each stage and outputs the image of the region of interest.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention provides a method for acquiring and confirming the area of overlapping tobacco mask images, comprising the following steps: acquiring a tobacco image to be inspected; identifying and segmenting the tobacco image using a pre-trained improved Mask-RCNN network detection model; employing a DenseNet121 and U-FPN network in the improved Mask-RCNN network detection model, and setting anchor optimization parameters in the RPN proposal network; and determining the corresponding area percentages, length, and width parameters of different tobacco shreds in the tobacco image based on the detection results of the improved Mask-RCNN network detection model. The present invention uses Mask-RCNN as the main network framework, employs a DenseNet121 and U-FPN network, and sets anchor optimization parameters in the RPN proposal network to accurately detect the tobacco image to be inspected, thereby determining the corresponding area percentages, length, and width parameters of different tobacco shreds in the tobacco image. By acquiring and calculating overlapping areas, and calculating tobacco length and width parameters, the method can assess tobacco quality, whole-shred rate, and broken-shred rate, addressing issues affecting tobacco component determination and tobacco quality. It helps to solve the problems of type recognition and component calculation of overlapping tobacco in blended tobacco and overlapping image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A flowchart of a method for obtaining and confirming the area of an overlapping tobacco mask image provided by an embodiment of the present invention;
[0059] Figure 2 Provides a training process diagram of an improved Mask-RCNN network detection model for an embodiment of the present invention;
[0060] Figure 3 Schematic diagram of the structure of four types of tobacco;
[0061] Figure 4 Schematic diagram of an image acquisition system for tobacco vibration experiments provided by an embodiment of the present invention;
[0062] Figure 5 A flowchart of overlapping tobacco image preprocessing provided by an embodiment of the present invention;
[0063] Figure 6 A schematic diagram of overlapping tobacco leaves segmented by an example according to an embodiment of the present invention;
[0064] Figure 7 The overall framework diagram of the improved Mask-RCNN network provided by the embodiment of the present invention;
[0065] Figure 8 Backbone network framework diagram provided by the embodiment of the present invention;
[0066] Figure 9A U-FPN network framework diagram provided by an embodiment of the present invention;
[0067] Figure 10 A schematic diagram illustrating the principle of calculating the pixel area of two overlapping tobacco shreds according to an embodiment of the present invention;
[0068] Figure 11 A schematic diagram illustrating the length calculation principle of two overlapping tobacco shreds provided in an embodiment of the present invention;
[0069] Figure 12 This is a schematic diagram of the calculation principle of the width of two overlapping tobacco shreds provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.
[0071] In the description of the present invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front end," "rear end," "both ends," "one end," "the other end," and the like, indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limiting the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0072] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "provided with," "connected," etc., should be understood in a broad sense. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection or an electrical connection; it may refer to a direct connection or an indirect connection through an intermediate medium; it may refer to internal communication between two components. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0073] Reference Figure 1 As shown, the present invention provides a method for obtaining and confirming the area of an overlapping tobacco mask image, comprising the following steps:
[0074] S1. Obtaining an image of tobacco to be detected;
[0075] S2. Recognizing and segmenting the tobacco image using a pre-trained improved Mask-RCNN network detection model; the improved Mask-RCNN network detection model uses a DenseNet121 and U-FPN network, and sets anchor point optimization parameters in the RPN proposal network;
[0076] S3. Determine the corresponding area ratio, length, and width parameters of different tobacco shreds in the tobacco shred image based on the detection results of the improved Mask-RCNN network detection model.
[0077] In step S2, Mask-RCNN is used as the main network framework. By changing the convolutional network and FPN in the backbone network to Densenet121 and U-FPN respectively, and optimizing the size and aspect ratio of the anchor boxes in RPN, an improved Mask-RCNN network detection model is proposed.
[0078] The Mask-RCNN network was selected as the main architecture for the instance segmentation network. The ResNet50 in the Mask-RCNN backbone was replaced with a DenseNet121. This effectively improved the ability to extract subtle features from shallow layers of overlapping tobacco, even in situations where overlapping tobacco has numerous overlapping types, features vary slightly, and features in overlapping areas are easily confused.
[0079] The proposed U-FPN structure adds upsampling and C2 and C3 horizontal connections to the original FPN. This enhances the utilization of shallow-layer information and extracted small tobacco features when small overlapping tobacco objects have rich shallow-layer features but deep-layer features contain less information about small objects.
[0080] Optimize the Anchors parameter in the RPN, designing the size and aspect_ratios to be suitable for overlapping small tobacco objects. This improves the extraction and candidate box capabilities of overlapping small tobacco objects, reducing missed detections and redundant computations.
[0081] The improved Mask-RCNN network (Densenet121, U-FPN, Anchors parameters) constructed achieved an accuracy of 90.2% in target detection and 89.1% in instance segmentation on the overlapping tobacco dataset, which is better than other similar segmentation networks.
[0082] In one embodiment, before detecting the tobacco image using the pre-trained improved Mask-RCNN network detection model, referring to Figure 2 As shown, the method further includes:
[0083] S21. Acquire multiple types of overlapping cut tobacco images through a cut tobacco vibration experiment; the overlapping cut tobacco images are composed of at least any two types of cut tobacco, cut stems, expanded cut tobacco, and reconstituted cut tobacco overlapping in different orders;
[0084] S22, processing the multiple types of overlapping tobacco images with an OpenCV algorithm to obtain a cropped image;
[0085] S23, the experimental image annotation tool sets labels for the cropped overlapping images and generates corresponding mask images as training data sets;
[0086] S24. Iteratively train the improved Mask-RCNN network using the training data set to obtain an improved Mask-RCNN network detection model.
[0087] like Figure 3 As shown, a cigarette contains four types of tobacco: leaf cuts (a), stem cuts (b), expanded leaf cuts (c) and reconstituted tobacco (d). In the tobacco vibration experiment, in order to ensure shooting accuracy and stability, an industrial camera and a manual focus lens can be used. The computer and the industrial camera are connected by a network cable to ensure the transmission speed and stability of the tobacco pictures. A ring light source is selected to ensure uniform brightness within the shooting field of view and eliminate the influence of tobacco shadows. Among them, industrial camera: Hikvision industrial camera MV-CE100-30GM / GC 10 million color camera can be selected; matching 12mm lens MVL-HF1224M-10MP; light source: ring LED light source with high brightness ring R120-80-25 angle; industrial camera bracket: industrial camera bracket CCD experimental bracket fine adjustment; universal light source lighting stand 600mm threaded rod fine adjustment knob; darkroom: three-dimensional lighting room made of four photographic reflectors. The overall image acquisition system is as follows Figure 4 shown.
[0088] For example, a total of 920 overlapping tobacco images were captured using the image acquisition system, with a single image size of 3840×2748. The overlapping tobacco images came from four types of tobacco (G for stems, P for expanded leaves, Y for leaves, and Z for reconstituted tobacco), totaling 24 overlapping types. For different types of tobacco overlap, the ratio of the number of captured tobacco images is different. For the overlap of the same type of tobacco (for example, GG for stems-stems), the ratio of self-winding, adhesion, and stack is 1:1:2. For the overlap of different types of tobacco (for example, GP for stems-expanded leaves), the ratio of adhesion and stack is 1:1. Considering that the captured data set is difficult to obtain and there are many types of overlapping types, the samples are relatively small, and the ratio of the training set to the test set is 8:2. Multiple types of overlapping tobacco images can be obtained through step S21.
[0089] In step S22, if the overlapping tobacco images obtained by the image acquisition system are small and contain a lot of background information, the training and prediction time of the segmentation model will be increased. The image preprocessing algorithm process is to process the overlapping tobacco images with the opencv algorithm, find the minimum circumscribed circle of the object, and perform contour cutting. The algorithm process is shown in Figure 5 As shown in the figure, the image preprocessing algorithm can reduce the invalid background information in the image while retaining the foreground information, thereby significantly reducing the image size.
[0090] In step S23, for example, the image standard tool Labelme is used to set labels on the pre-processed overlapping tobacco image to generate the corresponding mask image. Then, the code is made using the official coco dataset to create an overlapping tobacco dataset of coco data type. The four types of tobacco areas in the image are marked respectively, and the rest of the areas are defaulted to background. Taking the overlapping type in GP as an example, the marked image Figure 6 As shown in Figure 2, part a represents the original image, part b represents the mask image for instance segmentation, and part c represents the visualization of the mask image.
[0091] Finally, after the processing of the above steps S21-23, a training data set is formed. Step S24 uses the training data set to iteratively train the improved Mask-RCNN network to obtain an improved Mask-RCNN network detection model.
[0092] In one embodiment, detecting the tobacco image using a pre-trained improved Mask-RCNN network detection model includes:
[0093] S201, using the DenseNet121 and U-FPN networks to extract image features corresponding to the tobacco image;
[0094] S202: Using the image features as input to a proposal network (RPN) to generate a series of target candidate boxes.
[0095] S203 : The series of target candidate frames are subjected to FC and a fully convolutional network (FCN) to complete the recognition and segmentation of the overlapping tobacco images, and output a segmented image.
[0096] In this embodiment, the improved Mask-RCNN network structure consists of Backbone (CNN and FPN), RPN, RoiAlign, FCN and FC layer modules. The improved Mask-RCNN overall framework is as follows: Figure 7As shown in the figure. The model input is an overlapping tobacco image of size 500×500 pixels. The backbone network uses the DenseNet121+U-FPN combination to extract features and obtain a feature map. The feature map output from the backbone is then fed into the proposal network (RPN) to generate a region of interest. Subsequently, the ROI output from the RPN is mapped to extract the corresponding overlapping tobacco features in the shared feature map. Finally, the overlapping tobacco image is recognized and classified through FC and a fully convolutional network (FCN). The output of the model is the type of overlapping tobacco.
[0097] Step S201 specifically includes:
[0098] S2011, extracting feature information of the tobacco image using one convolutional layer and four feature extraction layers in the DenseNet121 network;
[0099] S2012, performing a pooling operation on the feature information extracted by each feature extraction layer to obtain four feature return values;
[0100] S2013, performing different downsampling and horizontal concatenation on the four feature return values to generate four first feature maps of different scales;
[0101] S2014. The four first feature maps of different scales, i.e., the first feature map of the first scale, the first feature map of the second scale, the first feature map of the third scale, and the first feature map of the fourth scale, are respectively input into the four feature layers of the U-FPN network, and image features corresponding to the tobacco image are obtained.
[0102] Due to the different morphological features of overlapping tobacco shreds and the small feature differences between different types of tobacco shreds, it is difficult to extract features from different types of tobacco shreds, and even more difficult to extract features from overlapping tobacco shreds and their overlapping areas. In Mask-RCNN, the use of Resnet50 to extract features at different levels of the input overlapping image has poor results. Increasing the number of DenseNet121 network layers can enhance the ability to extract small target detail information from different types of tobacco shreds. The use of dense connections between layers and multiple rounds of shallow information reuse can effectively extract subtle difference features and overlapping area features from shallow information in different types of overlapping tobacco shreds. Therefore, the use of DenseNet121 can effectively extract tiny features from small-sized overlapping tobacco shreds or large-sized overlapping tobacco shreds, enhancing the feature extraction capability of the overall overlapping tobacco shreds and, to a certain extent, solving the problem of shallow feature loss.
[0103] DenseNet121 is set as four feature extraction layers, namely Dense Block 1, Dense Block 2, Dense Block 3 and Dense Block 4. The four feature return values undergo different downsampling times (2, 3, 4, 5) and horizontal connections, and finally form a new type of Backbone, such as Figure 8 shown.
[0104] In step S2014, the first feature maps of four different scales are fused corresponding to the four feature layers of the input U-FPN network to obtain image features corresponding to the tobacco image. Specifically, the fusion process includes:
[0105] 1) The first feature maps of the four different scales are respectively subjected to 1×1 convolution and fused with the upsampled features, and then subjected to 3×3 convolution, and are input into the corresponding feature layer of the U-FPN network to obtain P2, P3, P4 and P5 features;
[0106] 2) Add a bottom-up and rating feature reuse structure, and sequentially transmit shallow layer information to the upper feature layer for fusion to obtain the image features corresponding to the tobacco image.
[0107] Due to the low resolution of small overlapping tobacco objects, after extraction by the CNN network, the shallow feature maps have a smaller receptive field and contain more small object details. The deep feature maps have a larger receptive field but contain less small object information. Although the top-down and same-layer connected structure of the FPN can merge deep and shallow features to meet the subsequent requirements of overlapping tobacco classification and detection, it still cannot make up for the problem of fully utilizing the subtle features in the shallow feature map.
[0108] The shallow feature information of small target detection is rich and important. Therefore, based on the top-down structure, a bottom-up and horizontal feature reuse structure is added to pass the shallow information to each feature layer (P3, P4, P5, P6), thereby enhancing the effective use of shallow feature information. The modified FPN network structure U-FPN, such as Figure 9 shown.
[0109] The P2 feature is added to the P3 feature through two 3×3 convolutions and the first feature map C3 of the second scale is added through 3×3 convolution to obtain the fused enhanced P3 feature;
[0110] The enhanced P3 feature is added to the P4 feature through two 3×3 convolutions and the first feature map C4 of the third scale through 3×3 convolution to obtain the fused enhanced P4 feature;
[0111] The enhanced P4 feature is added to the P5 feature through two 3×3 convolutions and the first feature map C5 of the fourth scale through 3×3 convolution to obtain the fused enhanced P5 feature;
[0112] The enhanced P5 feature is added to the feature map C6 obtained by fusing the first feature maps of four different scales through two 3×3 convolutions to obtain the fused P6 feature.
[0113] by Figure 9 Taking P3 in the example, under the premise of ensuring the number of feature layers in each layer, P2 is added to P3 through 3×3 / 2Conv and C3 is added to P3 through 3×3Conv. This can realize the reuse of the shallow information of the P2 layer and the shallow information of the C3 layer into P3, so as to enhance the fusion of the overall shallow feature information in P3.
[0114] In step S202, the image features are used as input to the proposal network (RPN) to generate a region of interest image. This includes setting multiple reference anchor frames at each position of the fused enhanced P3, P4, P5, and P6 features. The proposal network (RPN) adjusts the anchor frames based on the feature selection at each stage and outputs the region of interest image.
[0115] In Mask-RCNN, the scale and ratio of anchors are [128, 256, 512] and [1:1, 1:2, 2:1], respectively. Nine reference anchors are set for each position on the feature map. RPN selects and adjusts the anchor output ROIs based on the features of each stage. Taking P2 as an example, the feature map size of the P2 layer is 256×256, with a stride of 4. In this way, each pixel on P2 will generate a 4×4 anchor anchor box with an area of 16 based on the current coordinates. Based on the scale and ratio of the anchor, candidate boxes of three sizes and three shapes are generated for each pixel. After two layers of convolution, foreground and background classification is performed, and the offset with the candidate box is regressed. If the overlap with the target is greater than or equal to 0.7, it is foreground, and if the overlap with the target is less than or equal to 0.3, it is background. The remaining boxes are removed.
[0116] In Mask-RCNN, the smallest anchor scale is 128×128, but overlapping tobacco leaves contain many small objects, some of which are much smaller than this scale, making it impossible to detect these objects. Ideally, the smaller the object, the more numerous and denser the anchors should be to cover all candidate regions. The larger the object, the fewer and sparser the anchors should be, otherwise high overlap will cause redundant computation. However, the anchor parameters set in Mask-RCNN, when targeting small overlapping tobacco leaves, result in fewer and sparse anchors for small objects, while detecting larger objects with more and denser anchors. In this case, by appropriately adjusting the scale and ratio of the anchors, the detection performance of small objects can be significantly improved without significantly increasing the computational load.
[0117] In one embodiment, in the above step S3, the corresponding area ratio, length and width parameters of different tobacco shreds in the tobacco shred image are determined, including: obtaining a mask image of overlapping tobacco shreds based on the detection results, and drawing a fitted overlapping area; and calculating the corresponding area ratio, length and width parameters of different tobacco shreds respectively.
[0118] In the above embodiment, the detection model based on the improved Mask-RCNN network effectively achieves target detection and instance segmentation for overlapping tobacco shreds of various shapes and overlapping forms, and obtains the outlines of each overlapping tobacco target. Based on this, the pixel area and respective area proportion of the corresponding tobacco shreds can be calculated and obtained from the mask image using the OpenCV algorithm. However, the overlapping area of the obscured tobacco shreds cannot be obtained. The lack of the area of the overlapping area of the obscured tobacco shreds will directly lead to errors in the calculation of the individual areas of different types of tobacco shreds and the total area statistics of the tobacco shreds during the subsequent determination of tobacco components.
[0119] The algorithm considers the overlapping tobacco area by using the improved Mask-RCNN network to generate a mask image of overlapping tobacco, determine the obscured tobacco, draw the fitted overlapping area based on the distribution of the obscured overlapping tobacco, and use the fitted overlapping area and the unobstructed tobacco to determine the actual overlapping area.
[0120] Among them, the morphological detection of overlapping tobacco shreds (two tobacco shreds overlapping each other) mainly includes three aspects: the pixel area of each of the two tobacco shreds, the actual length of each of the two tobacco shreds, and the width of each of the two tobacco shreds.
[0121] 1. Calculate the pixel area of each of the two tobacco leaves. The specific steps are as follows:
[0122] (1) Determine the obscured tobacco object.
[0123] First, the mask image is grayscale processed, binarized, and the number of tobacco outlines is calculated and counted. The single outline is the unobstructed tobacco (see Figure 10The unobstructed tobacco in the figure), the multi-contour is the obstructed tobacco (see Figure 10 The pixel area of the unobstructed tobacco is Area1, and the total area of the obstructed tobacco (due to obstruction, the obstructed tobacco object is divided into two) is Area2. Calculate the pixel area of the standard block (the actual area is Area_standard. The specific pixel length is
[0124] (2) Fitting the overlapping area of tobacco.
[0125] Perform cyclic judgment on multiple contours of the obscured tobacco, and find two of them according to the size of the contour area. For the two contours, construct the minimum rectangle inside the contour respectively, and then use the center of the rectangle as the circle point and the side length as the diameter to draw the minimum inscribed circle round_1, whose center is (x1, y1) and diameter is d1. Similarly, draw the minimum inscribed circle round_2 of the second contour, whose center is (x2, y2) and diameter is d2. Connect the centers of the two inscribed circles (x1, y1) and (x2, y2) and draw a straight line L0. By extending the straight line L0, draw the tangent lines L1 and L2 of the same diameter as the circles for round_1 and round_2. Draw the fitting overlapping area trapezoid_1 according to L1 and L2 (see Figure 10 ).
[0126] (3) Determine the actual tobacco overlap area.
[0127] Generate a mask image for the fitting area, perform mask operation on the unobstructed image contour, and find the overlapping area (see Figure 10 Finally, the pixel count of the overlapping area is calculated as
[0128] (4) Calculation of respective areas and conversion to actual areas.
[0129] The actual area of unobstructed tobacco is:
[0130]
[0131] The actual area of the blocked tobacco is:
[0132]
[0133] 2. Calculate the length of each of the two tobacco shreds. The specific steps are as follows:
[0134] (1) Determine the obscured tobacco object.
[0135] First, the mask image is grayscale processed, binarized, and the number of tobacco outlines is calculated and counted. The single outline is the unobstructed tobacco (see Figure 11 The unobstructed tobacco in the figure), the multi-contour is the obstructed tobacco (see Figure 11 Calculate the standard block (the actual length of the pixel area is Length_standard. The specific pixel length is
[0136] (2) Length calculation.
[0137] By introducing the idea of calculus, we can perform segmented fitting of the center of the curved object. Starting from the center line of the tobacco, we measure along the center line from the beginning to the end. The length value obtained in this way is more accurate than other methods.
[0138] The specific steps are:
[0139] a. Search for the outline of the tobacco, capture its largest form, and rotate it so that the rectangle circumscribing the largest outline is horizontal (the rectangle's length is x1 and its width is y1). Create an X / Y coordinate system based on the lower right corner of the rotated rectangle, with the longest distance in the coordinate system being x1 and the highest distance being y1.
[0140] b. The tobacco image in the coordinate system is segmented multiple times.
[0141] c. The tobacco image is divided into m parts Each of the m images contains a partial contour. Find the center point of the contour. Loop through the m images and find all the center points.
[0142] d. Perform curve fitting on all the center points, which is the fitting center curve of the curved tobacco.
[0143] (3) Calculation of pixel length and actual length of the obscured tobacco.
[0144] According to the calculation method of the pixel area of each of the two tobacco leaves, the overlapping area can be obtained. The specific process is:
[0145] a. Using the length calculation method, fit the center curve of the two obscured tobacco outlines and obtain the pixel lengths. The corresponding pixel lengths are Length_1_1 and Length_1_2 respectively.
[0146] b. Using the length calculation method, fit the center curve of the two obscured tobacco contours and obtain the pixel lengths. The corresponding pixel lengths are Length_1_3.
[0147] c. Then the total pixel length of the obscured tobacco is:
[0148] Length_1=Length_1_1+Length_1_2+Length_1_3.
[0149] d. Then the actual pixel length of the obscured tobacco is:
[0150]
[0151] (4) Calculation of the pixel length and actual length of the unobstructed tobacco.
[0152] The specific process is:
[0153] Since there's no occlusion and only one object, the total pixel length of the unobstructed tobacco is calculated. Using the length calculation method, a central curve is fitted to the outline of the unobstructed tobacco, and the pixel length is obtained. The corresponding total pixel length is Length_2.
[0154] b. Then the actual pixel length of the obscured tobacco is:
[0155]
[0156] 3. Calculate the width of each of the two tobacco shreds. The specific steps are as follows:
[0157] (1) Determine the obscured tobacco object.
[0158] First, the mask image is grayscale processed, binarized, and the number of tobacco outlines is calculated and counted. The single outline is the unobstructed tobacco (see Figure 12 The unobstructed tobacco in the figure), the multi-contour is the obstructed tobacco (see Figure 12 Calculate the pixel area of the standard block (actual width is Width_standard. The specific pixel width is
[0159] (2) Width calculation.
[0160] By incorporating the concept of calculus, we perform segmented fitting of the maximum inscribed circle of the curved object. Starting from multiple maximum inscribed circles of the tobacco outline, we measure along these inscribed circles from the beginning to the end. This method yields a more accurate width value than other methods.
[0161] The specific steps are:
[0162] a. Search for the outline of the tobacco, capture its largest form, and rotate it so that the rectangle circumscribing the largest outline is horizontal (the rectangle's length is x1 and its width is y1). Create an X / Y coordinate system based on the lower right corner of the rotated rectangle, with the longest distance in the coordinate system being x1 and the highest distance being y1.
[0163] b. The tobacco image in the coordinate system is segmented multiple times.
[0164] c. The tobacco image is divided into n parts Each of the n images contains a partial contour. Find the largest inscribed circle in the contour. Loop through the n images and find all the largest inscribed circles.
[0165] d. Count the diameters of all the largest inscribed circles and calculate the average, which is the average width of the bent tobacco.
[0166] (3) Calculation of pixel width and actual width of the obscured tobacco.
[0167] According to the calculation method of the pixel areas of the two tobacco leaves, the overlapping area can be obtained.
[0168] The specific process is:
[0169] a. Using the width calculation method, perform multi-end fitting on the two outlines of the obscured tobacco to obtain the maximum inscribed circle and pixel diameter. The corresponding average pixels are Width_1_1 and Width_1_2, respectively.
[0170] b. Using the width calculation method, perform multi-end fitting on the two outlines of the obscured tobacco to obtain the maximum inscribed circle and pixel width. The corresponding pixel widths are Width_1 and Width_3 respectively.
[0171] c. Then the average total pixel width of the obscured tobacco is:
[0172]
[0173] d. The actual pixel width of the obscured tobacco is:
[0174]
[0175] (4) Calculation of the pixel width and actual width of the unobstructed tobacco.
[0176] The specific process is:
[0177] Since there's no occlusion and only one object, the total pixel width of the unobstructed tobacco is calculated. Using the width calculation method, we determine the maximum inscribed circle of the unobstructed tobacco outline at multiple ends and then average the pixel width. The corresponding total pixel average is Width_2.
[0178] b. Then the actual pixel width of the obscured tobacco is:
[0179]
[0180] In this embodiment, the identification and calculation of overlapping areas, along with tobacco length and width parameters, can be used to assess tobacco quality, whole tobacco percentage, and broken tobacco percentage, addressing issues affecting tobacco component determination and quality. The average detection rate for overlapping tobacco area increased from 81.2% to 90%, effectively avoiding negative optimization in the identification and calculation of overlapping areas. This allows for accurate tobacco classification and component area calculation.
[0181] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for obtaining and confirming the area of an overlapping tobacco mask image, characterized in that: The following steps are involved: Acquire an image of tobacco to be detected; The tobacco image is identified and segmented using a pre-trained improved Mask-RCNN network detection model; the improved Mask-RCNN network detection model uses a DenseNet121 and U-FPN network, and sets anchor optimization parameters in an RPN proposal network; According to the detection results of the improved Mask-RCNN network detection model, a mask image of the overlapping tobacco shreds is obtained, and a fitting overlapping area is drawn; Calculate the corresponding area ratio, length and width parameters of different tobacco shreds respectively; The calculation of the corresponding area ratios of different tobacco shreds includes: Grayscale processing and binarization are performed on the mask image, and the number of tobacco outlines is calculated and counted to determine the area of the unobstructed tobacco object Area1 and the obstructed tobacco object Area2; the mask image is an overlapping image of the two tobacco shreds; Performing a cyclic judgment on multiple contours of the obstructed tobacco, finding two contours based on their areas, and fitting the overlapping area of the tobacco based on the minimum rectangle inside the contours, the inscribed circle of the rectangle, and the circumtangent of the inscribed circle of the rectangle; Generate a mask image for the overlapping area of tobacco obtained by fitting, perform mask operation on the unobstructed image contour, find the overlapping area, and determine the actual tobacco overlapping area Area OverlappedArea ; Combined with the standard block for calculating pixel area, the actual area of the unobstructed tobacco object and the actual area of the obstructed tobacco object are calculated respectively.
2. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 1, characterized in that: Calculate the length of different tobacco pieces, including: Performing calculus calculations on the unobstructed tobacco object and the obstructed tobacco object, performing segmented fitting center processing, and measuring along the tobacco centerline from the beginning to the end; According to the overlapping area determined when calculating the corresponding area proportions of different tobacco shreds, the obscured tobacco object is divided into three segments for processing, and the pixel length of the obscured tobacco object is obtained; For the unobstructed tobacco object, the center line is fitted to obtain the pixel length of the unobstructed tobacco object; According to the respective pixel lengths and in combination with the pixel length calculation standard block, the actual length of the unobstructed tobacco object and the actual length of the obstructed tobacco object are calculated respectively.
3. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 2, characterized in that: Calculate the width of different tobacco pieces, including: Performing calculus calculations on the unobstructed tobacco object and the obstructed tobacco object, respectively, performing segmented fitting of the maximum inscribed circle, measuring along the inscribed circle of the tobacco from the beginning to the end, and calculating the average value of the diameter; According to the overlapping area determined when calculating the corresponding area proportions of different tobacco shreds, the obscured tobacco object is divided into three segments, and the width means of the three segments are averaged to obtain the pixel width of the obscured tobacco object; For the unobstructed tobacco objects, the largest inscribed circle is fitted, and the mean of all inscribed circle diameters is used as the pixel width of the unobstructed tobacco objects; According to the respective pixel widths, combined with the pixel width calculation standard block, the actual width of the unobstructed tobacco object and the actual width of the obstructed tobacco object are calculated respectively.
4. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 1, characterized in that: Before the improved Mask-RCNN network detection model obtained through pre-training detects the tobacco image, the method further includes: Acquire multiple types of overlapping cut tobacco images through a cut tobacco vibration experiment; the overlapping cut tobacco images are composed of at least any two of cut tobacco leaves, cut stems, expanded cut tobacco leaves, and reconstituted cut tobacco overlapped in different orders; Processing the multiple types of overlapping tobacco images with an OpenCV algorithm to obtain cropped images; The experimental image annotation tool sets labels for the cropped overlapping images and generates corresponding mask images as training datasets; The improved Mask-RCNN network is iteratively trained using the training data set to obtain an improved Mask-RCNN network detection model.
5. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 4, characterized in that: The improved Mask-RCNN network detection model obtained through pre-training recognizes and segments the tobacco image, including: Extracting image features corresponding to the tobacco image using the DenseNet121 and U-FPN networks; The image features are used as input to the proposal network RPN to generate a series of target candidate boxes; The series of target candidate frames are passed through FC and fully convolutional network FCN to complete the recognition and segmentation of overlapping tobacco images and output the segmented image.
6. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 5, characterized in that: Extracting image features corresponding to the tobacco image using the DenseNet121 and U-FPN networks includes: Extracting feature information of the tobacco image using one convolutional layer and four feature extraction layers in the DenseNet121 network; And perform pooling operation on the feature information extracted by each feature extraction layer to obtain four feature return values; The four feature return values are respectively downsampled and horizontally connected to generate four first feature maps of different scales; The four first feature maps of different scales are respectively input into the four feature layers of the U-FPN network for fusion to obtain image features corresponding to the tobacco image.
7. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 6, characterized in that: The four first feature maps of different scales are respectively input into the four feature layers of the U-FPN network for fusion to obtain image features corresponding to the tobacco image; including: The first feature maps of the four different scales are respectively subjected to 1×1 convolution, fused with the up-sampled features, and then subjected to 3×3 convolution, and are input into the corresponding feature layer of the U-FPN network to obtain P2, P3, P4 and P5 features; A bottom-up and rating feature reuse structure is added, and shallow layer information is sequentially transferred to the upper feature layer for fusion to obtain image features corresponding to the tobacco image.
8. The method for obtaining and confirming the area of an overlapping tobacco mask image according to claim 7, characterized in that: Add a bottom-up and rating feature reuse structure, and sequentially pass shallow layer information to the upper feature layer for fusion to obtain the image features corresponding to the tobacco image, including: The P2 feature is added to the P3 feature through two 3×3 convolutions and the first feature map C3 of the second scale is added through 3×3 convolution to obtain the fused enhanced P3 feature; The enhanced P3 feature is added to the P4 feature through two 3×3 convolutions and the first feature map C4 of the third scale through 3×3 convolution to obtain the fused enhanced P4 feature; The enhanced P4 feature is added to the P5 feature through two 3×3 convolutions and the first feature map C5 of the fourth scale through 3×3 convolution to obtain the fused enhanced P5 feature; The enhanced P5 feature is added to the feature map C6 obtained by fusing the first feature maps of four different scales through two 3×3 convolutions to obtain the fused P6 feature.
Citation Information
Patent Citations
Mask R-CNN-based synchronous identification method for multiple tobacco leaf parts
CN114972154A