AI-based Key Index Prediction Method for Silk Reeling Process and Cigarette Packing Quality Assurance Method
Through AI-based methods, combined with deep learning networks and custom loss functions, the joint prediction and detection of multiple key quality indicators in the tobacco silk making and rolling process is achieved, solving the problem of high-precision automated quality control in the existing technology, and improving the prediction accuracy and image segmentation accuracy.
Patent Information
- Application Number
- CN202510743973.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The prior art is difficult to achieve high-precision and automated quality control in the process of tobacco silk making and rolling, especially in image segmentation, feature extraction and model training, and traditional methods are difficult to deal with the association relationship of multiple key quality indicators at the same time.
Using an AI-based method, by collecting silk making and wrapping process data, preprocessing and feature extraction, combining deep learning networks and custom loss functions, multi-task branch learning and extrusion enhancement axial attention mechanism is introduced to realize the joint prediction and detection of multiple key quality indicators.
It significantly improves the prediction accuracy and quality assurance capabilities during the silk making and rolling process, enhances the attention mechanism and generalization capabilities of the model, and improves the accuracy of image segmentation and the effectiveness of feature extraction.
Smart Images

Figure CN120318219B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an AI-based method for predicting key indicators of a silk-making process and ensuring the quality of winding and packaging. Background Art
[0002] In modern tobacco shred and packaging production, the prediction and quality assurance of key indicators are crucial factors influencing the quality of the final product. Due to the complex texture, color, and morphological characteristics of tobacco leaves and cut tobacco, traditional detection and prediction methods rely on manual experience or simple statistical models, making it difficult to achieve high-precision, automated quality control. With the advancement of artificial intelligence (AI) and deep learning technologies, intelligent analysis methods based on computer vision have gradually become a key technical means to improve the stability of tobacco shred processes and packaging quality. However, despite the great potential of these technologies, they still face numerous challenges in practical application. For example, in image segmentation, traditional methods struggle to accurately segment fine structures such as tobacco particles and stems, hindering subsequent quality assessment. In feature extraction, traditional methods often fail to fully capture the macrotextural characteristics of tobacco leaves and cut tobacco, resulting in insufficient discriminatory power of prediction models for key indicators. Furthermore, in model training, traditional loss functions and attention mechanisms struggle to effectively optimize prediction model performance, limiting their practical application.
[0003] Furthermore, tobacco shredding and packaging involve multiple key quality indicators, often with complex interdependencies. Traditional single-task learning methods struggle to handle multiple tasks simultaneously, limiting improvements in overall quality assurance capabilities.
[0004] Therefore, there is an urgent need for an AI prediction and quality assurance method that can handle multiple tasks simultaneously, has high-precision prediction capabilities and efficient computing performance. Summary of the Invention
[0005] In view of the fact that traditional detection and prediction methods rely on manual experience or simple statistical models and are difficult to achieve high-precision and automated quality control, and the shortcomings of traditional computer vision-based intelligent analysis methods in image segmentation, feature extraction, loss function design and attention mechanism, this invention proposes an AI-based method for predicting key indicators of the silk-making process and ensuring the quality of winding and packaging. The method aims to improve prediction accuracy, optimize feature extraction, and enhance the model's attention mechanism by combining artificial intelligence and deep learning technology to address the problems of predicting key indicators of the silk-making process and ensuring the quality of winding and packaging.
[0006] To achieve the above objectives, the following technical solutions are adopted:
[0007] An AI-based method for predicting key indicators of the silk-making process and ensuring the quality of the winding package, including:
[0008] Collect data from the tobacco making and wrapping processes, including physical parameters, equipment status data, and tobacco leaf / cut tobacco images from the tobacco making process, as well as cigarette parameters, equipment data, and packaging images from the wrapping process;
[0009] Preprocessing and feature extraction are performed on the silk-making process data and the winding process data respectively to obtain a silk-making process feature vector and a winding process feature vector;
[0010] The silk-making process feature vector and the wrapping process feature vector are fed into a deep learning network and trained using a custom loss function and a squeeze-enhanced axial attention mechanism; wherein the deep learning network comprises an encoder-decoder main network architecture, wherein the encoder comprises four stages, each stage having two 3×3 depthwise separable convolutional layers, followed by a 2×2 maximum pooling layer and a ReLU activation function, with the number of filters doubling at each stage for extracting visual features; the decoder comprises multiple separable convolutional layers and upsampling layers, upsampling the encoder output through a transposed convolution operation, and combining features from the encoder to achieve image reconstruction and segmentation;
[0011] Multi-task branch learning is introduced on the basis of the main network architecture of the deep learning network, and it is expanded into a multi-task joint learning framework that combines shared feature extraction with multi-task branches. The shared backbone network extracts common features, and multi-task branches are set up to respectively process classification tasks, regression tasks, target detection tasks and image segmentation tasks, so as to realize the joint prediction and detection of multiple key quality indicators in the silk making and winding processes.
[0012] Furthermore, the pre-processing of the silk making process data and the wrapping process data respectively includes:
[0013] The time series data of the silk-making process data and the winding process data are cleaned and normalized, and the image data are subjected to color enhancement, contrast enhancement and image segmentation.
[0014] Furthermore, the feature extraction of the pre-processed silk-making process data includes:
[0015] Extracting features related to key indicators from the physical parameters of the silk-making process to obtain the physical parameter characteristics of silk-making, including: extracting the heating rate and humidity fluctuation range by analyzing the temperature and humidity change curves;
[0016] Extract features related to key indicators from the equipment operation data of the silk-making process to obtain equipment operation characteristics, including speed and fault status;
[0017] Extract image features from tobacco leaf and tobacco cut images to obtain tobacco cut image features, including color features, macro texture features, and shape features;
[0018] The silk-making physical parameter characteristics, the equipment operation characteristics and the silk-making image characteristics are integrated and dimension reduction is performed to obtain the silk-making process feature vector.
[0019] Furthermore, the feature extraction of the pre-processed packaging process data includes:
[0020] Extract features related to packaging quality, including cigarette weight stability;
[0021] Extracting the sealing quality characteristics of packaging from the operating parameters of packaging equipment;
[0022] Extracting image features from the package image to obtain package image features, including cigarette arrangement regularity features and packaging pattern integrity features;
[0023] The cigarette weight stability feature, the sealing quality feature and the wrapping image feature are integrated and dimensionality reduction is performed to obtain the wrapping process feature vector.
[0024] Furthermore, the encoder structure includes:
[0025] The 3×3 depthwise separable convolutional layer at each stage is composed of a BiConvLSTM module and a SwinTransformer module in series;
[0026] The bidirectional convolutional long short-term memory network module is used to divide the input feature map into blocks of size m×n and perform conditional position encoding on each block. Convolution operations are integrated into the input-to-state and state-to-state transformations, and the hidden states of the forward and backward paths are fused through channel splicing to enhance the network's ability to learn complex patterns and relationships.
[0027] The Swain transformer block is used to divide the input feature map into non-overlapping s×s patches as "tokens", which are projected to the specified dimension through a linear embedding layer; window-based multi-head self-attention (W-MSA) is calculated between patches, with a window size of w×w, satisfying the relationship: ;Introduce a cross-window connection mechanism to achieve information interaction between adjacent regions through shifting window division;
[0028] The encoder uses densely connected convolution to achieve cross-layer feature reuse by concatenating the output of the previous convolutional layer with the output of all previous layers along the channel dimension and then inputting them into the next layer;
[0029] The bidirectional convolutional long short-term memory network module and the Swain transformer module are integrated into the skip connection and connected to the corresponding stage of the decoder, so as to realize the refined extraction and effective fusion of features and improve the model's ability to model multi-scale and long-distance dependencies.
[0030] Furthermore, the decoder structure includes:
[0031] Each upsampling layer is followed by two groups of separable convolution units connected in series. Each group of separable convolution units contains: a 3×3 depth-wise separable convolution layer, in which the number of filters is halved at each stage; a 1×1 point-by-point convolution layer for channel dimension feature reorganization; a dynamic upsampling kernel generation module, which generates transposed convolution kernel weights based on the feature map of the corresponding stage of the encoder. The dynamic upsampling kernel generates a transposed convolution kernel based on the encoder features, enabling the upsampling process to adaptively adjust the reconstruction strategy for different tobacco morphologies (such as bent stems and fragments), thereby improving the accuracy of segmentation edges.
[0032] Among them, the decoder output layer introduces a feature calibration mechanism to weight the reconstructed features through spatial attention weights.
[0033] Furthermore, the custom loss function includes a weighted combination of Dice loss, Jaccard loss, boundary loss and MSE:
[0034]
[0035] in, They are the weighted coefficients of Dice loss, Jaccard loss, boundary loss and MSE, which are used to adjust the influence weight of each loss item;
[0036] The squeeze-enhanced axial attention mechanism compresses features in the row / column direction and fuses position embedding with local convolution enhancement to enhance the feature learning ability of the neural network in different axes.
[0037] Furthermore, the specific implementation of the squeeze-enhanced axial attention mechanism includes the following steps:
[0038] (1) Bidirectional axial compression:
[0039] Input feature map Squeeze operations are performed along the row and column directions, mapping to a single feature marker for each row, where is the image height, is the image width, is the number of channels, and the features after row-wise compression are obtained and column-wise compressed features ;
[0040] Calculate the bidirectional attention weight and get the row-wise attention weight and column-wise attention weights ;in, , ;
[0041] Generate bidirectional attention features, that is, row-wise attention features and column-wise attention features ;
[0042] (2) Position embedding compensation:
[0043] The positional encoding PE is extracted from the original input feature X through 1×1 convolution:
[0044]
[0045] Fusion of positional encoding with bidirectional attention features:
[0046]
[0047] in, The enhanced axial attention feature It is the axial attention feature; Represents the dimension expansion operation, which is used to convert the row-wise attention features and column-wise attention features The dimensions are uniformly adjusted to the same spatial size H×W×C as the original input feature map;
[0048] (3) Spatial detail enhancement:
[0049] Perform 3×3 convolution on the original input feature X to extract local detail features ;
[0050] Fusion of global attention and local features through detail weight gating:
[0051]
[0052] in, is the Sigmoid function, which is used to normalize the weights so that between; Represents element-by-element multiplication, and the final output feature The complexity is .
[0053] Furthermore, wherein the pair of input feature maps The extrusion operations are performed along the row and column directions respectively, including:
[0054] Row-wise compression: sum the features of each row to get the row compression features , and through the learnable matrix Generate a query ,key ,value , the specific formula is:
[0055]
[0056] in, Represents the result of summing over the row dimension, i.e., the row compression feature; are the projection matrices of query, key, and value in the row direction respectively;
[0057] Column-wise compression: Sum the features of each column to get the column-wise compression features , and through the learnable matrix Generate a query ,key ,value , the specific formula is:
[0058]
[0059] in, Represents the result of summing over the column dimension, i.e., column compression feature; The projection matrices for query, key, and value in the column direction respectively.
[0060] Furthermore, the multi-task branches include: a silk making process task branch, which is used to output the silk length distribution uniformity score, tobacco width and thickness classification, stem content percentage, and blade status classification; and a wrapping process task branch, which is used to output the sealing integrity classification, filter offset value, trademark clarity score and foreign matter detection box.
[0061] Compared with the prior art, the present invention achieves the following beneficial effects:
[0062] 1. This invention combines artificial intelligence and deep learning technologies, employing a bidirectional convolutional long short-term memory (BiConvLSTM) module and a Swin Transformer module, combined with densely connected convolutions, to effectively extract and fuse multi-scale and long-range dependent features. This feature extraction approach not only improves the model's ability to process complex images and time series data, but also enhances the model's focus on key areas, improving overall prediction performance.
[0063] 2. This invention introduces multi-task branching learning, expanding the deep learning network into a multi-task joint learning framework that combines shared feature extraction with multi-task branching. This framework enables the joint prediction and detection of multiple key quality indicators in the silk making and winding processes, significantly improving overall quality assurance and detection efficiency. Furthermore, the multi-task branching design enables the model to be specifically optimized for different tasks, improving its generalization and practicality.
[0064] 3. The present invention introduces a custom loss function and a squeeze-enhanced axial attention mechanism. This loss function can perform refined modeling for key indicators in the silk-making process. It not only takes into account the traditional error calculation method, but also combines industry characteristics to enhance the sensitivity to key variables, thereby improving the accuracy and robustness of the model for quality assessment. The squeeze-enhanced axial attention mechanism strengthens the feature learning ability of the neural network in different axes through squeezing and expansion operations, so that the model can more accurately capture the information of key parts when processing the winding quality analysis task, thereby improving the overall prediction performance and the interpretability of the model. The experimental results show the superiority of the present invention in image segmentation and feature prediction, which significantly improves the prediction accuracy of key indicators of the silk-making process and the winding process.
[0065] 4. This invention employs an enhanced level set algorithm for image segmentation, overcoming the shortcomings of traditional level set models in boundary localization and segmentation accuracy. By replacing the traditional Gaussian kernel with a bilateral kernel, robustness to image noise is enhanced, improving the accuracy and reliability of image segmentation, and enabling more precise segmentation of fine structures such as tobacco particles and stems.
[0066] 5. This invention proposes a method for extracting macrotexture features: Texture features are key to quality analysis of cut tobacco and tobacco stems. To this end, this invention proposes a texture feature extraction method that combines the characteristics of Schmid filters and Gabor filters. This method can more effectively capture macrotexture information on the tobacco surface, thereby improving feature differentiation and facilitating subsequent quality assessment and optimization of prediction models.
[0067] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0069] Figure 1 This is a schematic diagram of the specific steps of an AI-based method for predicting key indicators of a silk-making process and ensuring the quality of a package, according to an embodiment of the present invention;
[0070] Figure 2 This is a flow chart of an AI-based method for predicting key indicators of a silk-making process and ensuring the quality of a package, according to an embodiment of the present invention. DETAILED DESCRIPTION
[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0072] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0073] Figure 1 The figure shows a schematic diagram of the specific steps of an AI-based method for predicting key indicators of the silk-making process and ensuring the quality of the winding package. Figure 2 The figure shows a flow chart of a method for predicting key indicators of silk making process and ensuring the quality of winding and packaging based on AI. Figure 1 and Figure 2 As shown, an AI-based method 100 for predicting key indicators of a silk-making process and ensuring the quality of a package, comprising:
[0074] S110: Collecting tobacco making process data and tobacco wrapping process data, including physical parameters, equipment status data, and tobacco leaf / cut tobacco images of the tobacco making process, and cigarette parameters, equipment data, and packaging images of the tobacco wrapping process;
[0075] The step S110 specifically includes:
[0076] S111: Collecting silk making process data
[0077] Various sensors are used to collect physical parameters during the tobacco-making process, such as temperature, humidity, pressure, and flow rate. The equipment control system also captures equipment operating status data, such as speed, operating time, and fault alarms. Industrial cameras are installed at various stages of the tobacco-making process, such as leaf selection, shredding, and drying, to capture images of tobacco leaves and shredded tobacco, detecting impurities, leaf shape, tobacco color, and moisture content.
[0078] S112: Collecting data on the packaging process
[0079] Collects data on the operating parameters of the wrapping equipment, quality data on the packaging materials, and physical indicators of the cigarettes, such as weight, circumference, and length. During the wrapping process, machine vision is used to collect image data on the cigarette appearance, packaging integrity, and label printing quality.
[0080] S120: Preprocessing and feature extraction are performed on the silk-making process data and the winding process data respectively to obtain a silk-making process feature vector and a winding process feature vector;
[0081] The step S120 specifically includes:
[0082] S121: Data Preprocessing
[0083] S1211: Time Series Data Preprocessing
[0084] The time series data collected by sensors during the silk-making and wrapping processes is preprocessed in the same manner. Data cleaning removes noise, outliers, and duplicates. Normalization maps data of varying types and ranges to specific intervals, facilitating subsequent data analysis and model training.
[0085] S1212: Image data preprocessing
[0086] The image data collected by machine vision in the silk-making process and the winding process are pre-processed, and the pre-processing methods of the silk-making process image data and the winding process image data are consistent.
[0087] S12121: Color Enhancement
[0088] Color enhancement is achieved by mixing the original image with its grayscale version in a certain ratio, as shown in the following formula:
[0089]
[0090] in, is the color-enhanced image; is the original color image; is the grayscale version of the original image; is the mixing ratio parameter, which controls the degree of color enhancement. The larger the value, the closer the image is to a grayscale image. For example, to balance the effect of color fidelity and contrast enhancement, you can choose .
[0091] S12122: Contrast Enhancement
[0092] First, the image is segmented into non-overlapping blocks. The pixel intensity histogram and cumulative distribution function of each block are calculated. The cliplimit is set to control the degree of contrast enhancement. The pixel intensities are transformed, and finally, the enhanced blocks are reassembled to obtain the output image. By enhancing local image contrast, reducing noise, and highlighting image details, a better image foundation is provided for subsequent analysis.
[0093] S12123: Image Segmentation Technology
[0094] The embodiment of the present invention uses an enhanced level set algorithm for image segmentation, which improves the traditional level set model and more accurately segments the image area. The enhanced level set evolution equation is:
[0095]
[0096] in, is the level set function The rate of change over time t characterizes the evolution speed of image segmentation boundaries (such as tobacco / stem edges); is the level set function, representing the implicit surface of the boundary; is the smoothing parameter, which is used to control the smoothness of the surface; It is the Laplace operator of the level set function, which is used to smooth the surface; is the level set function The gradient operator is used to calculate the rate of change of the boundary normal direction; is the energy term coefficient, which controls the boundary advancement speed; is an external force field that depends on the image gradient; is the penalty term parameter, which is used to adjust the impact of the regional term; is the edge detection function, generally taken , used to maintain boundaries, is the image gradient, which is calculated as: , used to quantify the intensity of local grayscale changes; among them, , is the gradient component calculated by the Sobel operator.
[0097] This technology provides key feature information by adaptively adjusting parameters and replacing the Gaussian kernel with a bilateral kernel, thereby improving the accuracy and reliability of image segmentation. The bilateral kernel calculation formula is:
[0098]
[0099] in, is the center pixel coordinate of the current bilateral response calculation; is the radius of the kernel window, and the window size is It's a pixel Grayscale or color value; It is the weighted local mean within the window, usually normalized and weighted by the same bilateral weights; is the bilateral energy term of the current pixel, which is used to reflect the local change intensity of the area; is the bilateral weight, a weight coefficient combining spatial distance and pixel similarity, and is often defined as:
[0100]
[0101] in, is the control space weight, is to control the pixel similarity weight, for example, , .
[0102] S122: Feature Extraction
[0103] S1221: Silk Making Process Feature Extraction
[0104] Extract key features from the physical parameters of the tobacco-making process and equipment operating data. For example, analyze temperature and humidity curves to extract features such as heating rate and humidity fluctuation range. Extract color, texture, and shape features from tobacco leaf and cut tobacco images. For example, extract macro-texture features using the frequency and spatial characteristics of wavelet transforms and extract leaf shape features using edge detection algorithms.
[0105] Among them, the macro texture feature method is as follows:
[0106] The image is decomposed using a two-dimensional discrete wavelet transform and a Discrete Meyer (dmey) filter to decompose the image into approximate (LL) and detail (LH, HL, HH) subbands to capture information at different frequencies. The gradient map of the image is calculated based on a Schmid Gabor-like filter to extract macro texture features. The formula is:
[0107]
[0108] in, is the filter function; are the pixel coordinates in the image; Used to control the scale of the filter; Used to control the direction of the filter; is the center frequency of the filter.
[0109] Macro feature vector calculation uses 13 pairs of filters of different scales to generate multiple edge maps , and calculate its sorting histogram. For each edge graph , quantify its local texture pattern by computing the sorted histogram:
[0110]
[0111] in, Indicates the The histogram of the edge map The count value of each bin; are the width and height of the image, and the summation is performed over the entire image; It is After the filter is applied, the pixel edge strength of the location; is the Kronecker delta function, which is true when the expression in the brackets is true (i.e., the value is ) is 1, otherwise it is 0; : Histogram Number, usually Take 13; : filter index.
[0112] Connect the histograms of each edge graph and finally get:
[0113]
[0114] Get a dimension of The macro texture feature vector is used to comprehensively characterize the texture structure of tobacco leaves or tobacco.
[0115] S1222: Fusion of silk-making process features
[0116] The features extracted from the physical parameters and equipment operation data of the tobacco production process are combined with the features extracted from the tobacco leaf and cut tobacco images. Since the physical parameters, equipment operation data, and image feature dimensions are different, the various features are standardized. The standardized physical parameter features, equipment operation features, and image features are spliced together:
[0117]
[0118] in, Physical parameter characteristics of silk making (such as the rate of change of temperature and humidity); Equipment operating characteristics (such as speed, fault status); It is the silk making image features (such as color features, macro texture features (color histogram), shape features).
[0119] Since the dimension after feature fusion is high, principal component analysis (PCA) is used to reduce the dimension, and the result is is the final feature vector after dimensionality reduction, which is used for subsequent model training.
[0120] S1223: Feature extraction of the packaging process
[0121] Extract features related to packaging quality, such as calculating weight standard deviation and coefficient of variation from cigarette weight data, extracting features related to packaging speed stability and sealing quality from packaging equipment operating parameters, and extracting features such as cigarette arrangement neatness and packaging pattern integrity from packaging images.
[0122] The coefficient of variation of weight is To characterize the relative dispersion of cigarette weight:
[0123]
[0124] in, For the The weight of a cigarette; is the weight standard deviation; is the average weight of all cigarettes; is the sample size.
[0125] Calculate the sealing quality features and use image gradient changes to detect the integrity of the sealing edge:
[0126]
[0127] in, It is a characteristic indicator of sealing quality, indicating the total gradient intensity of the sealing area. A larger value indicates a more pronounced edge and a clearer and more complete seal. A lower value indicates possible blurring or damage. Is the image at position The gradient value at , which indicates the degree of grayscale change, is calculated using the Sobel operator (with Directions as an example):
[0128]
[0129] in, Represents the pixel intensity value of the image at the specified coordinate location.
[0130] Hough transform is used to detect cigarette edge lines and calculate the arrangement deviation. The smaller the value, the more orderly the cigarettes are arranged:
[0131]
[0132] in, is the number of detected cigarette edge points; is the coordinate of the edge point of the cigarette; and are the parameters of the fitted straight line; is the arrangement deviation.
[0133] The structural similarity index (SSIM) of the packaging pattern region is calculated to measure the completeness. The closer the structural similarity index is to 1, the more complete the packaging pattern is:
[0134]
[0135] in, is the structural similarity index; For standard packaging images, To detect packaging images; is the image mean; is the variance; is the covariance; is a constant (used to avoid zero denominator).
[0136] S1224: Fusion of wrapping process features
[0137] Step S1224 is used to merge and combine the features extracted in step S1223. In order to make the features of different dimensions have the same weight, Min-Max normalization is performed. The weight features, equipment operation features and image features are merged and combined:
[0138]
[0139] Represents the weight stability characteristics of cigarettes; Represents the sealing quality characteristics; It represents the neatness of cigarette arrangement; SSIM represents the integrity of packaging pattern.
[0140] Since the dimension after feature fusion is high, principal component analysis (PCA) is used to reduce the dimension, and the result is is the final feature vector after dimensionality reduction, which is used for subsequent model training.
[0141] S130: The silk-making process feature vector and the wrapping process feature vector are input into a deep learning network, and trained using a custom loss function and a squeeze-enhanced axial attention mechanism; wherein the deep learning network is composed of an encoder-decoder as a main network architecture, the encoder having four stages, each stage having two 3×3 depthwise separable convolutional layers, followed by a 2×2 maximum pooling layer and a ReLU activation function, and the number of filters is doubled at each stage for extracting visual features; the decoder is composed of multiple separable convolutional layers and upsampling layers, and the encoder output is upsampled by a transposed convolution operation, and the features from the encoder are combined to achieve image reconstruction and segmentation;
[0142] S131: Building a Deep Learning Network Architecture
[0143] The result obtained in step S120 A deep learning network is trained. The deep learning network consists of an encoder-decoder architecture. The encoder consists of four stages, each of which has two 3×3 depthwise separable convolutional layers, followed by a 2×2 max pooling layer and a ReLU activation function. The number of filters doubles at each stage to extract visual features. In addition, the encoder uses densely connected convolutions and cross-layer feature reuse (the concept of "collective knowledge") to improve the model's ability to express complex features, while alleviating the problems of vanishing and exploding gradients and making training more stable. The decoder consists of multiple separable convolutional layers and upsampling layers. The encoder output is upsampled through a transposed convolution operation and combined with the features from the encoder to achieve image reconstruction and segmentation.
[0144] Specifically, the encoder architecture includes the following: The 3×3 depthwise separable convolutional layer at each stage is composed of a bidirectional convolutional long short-term memory (BiConvLSTM) module and a Swin Transformer module in series. The BiConvLSTM incorporates convolution operations in both input-to-state and state-to-state transformations, uses conditional positional encoding to divide the input into multiple blocks, and dynamically generates positional encoding information by reshaping and convolving each block. This enhances the model's understanding of spatial location and addresses the spatial correlation limitations of traditional LSTMs. It captures temporal dependencies and processes data through both forward and backward paths, effectively integrating information and improving the network's ability to learn complex patterns and relationships. The Swin Transformer module divides the input image into non-overlapping patches known as "tokens," which are projected to a specified dimension via a linear embedding layer. The bidirectional convolutional long short-term memory network module and the Swing transformer module are integrated into the skip connection to achieve refined feature extraction and effective fusion, improving the model's ability to model multi-scale and long-distance dependencies.
[0145] Each encoder stage executes in sequence: a bidirectional convolutional long short-term memory network module processes spatiotemporal features; a Swain transformer module extracts global attention features; and a densely connected convolution fuses cross-layer features.
[0146] Furthermore, the bidirectional convolutional long short-term memory network module implements spatiotemporal feature fusion through the following operations:
[0147] (1) Divide the input feature map into blocks of size m×n and perform conditional position encoding on each block:
[0148]
[0149] in, For the The input features of the block, Code dynamically generated positions;
[0150] (2) Incorporating 3×3 depthwise convolution operations into the input-to-state and state-to-state transformations, the hidden states of the forward and backward paths are fused through channel splicing to enhance the network’s ability to learn complex patterns and relationships. The process of fusion of the hidden states of the forward and backward paths through channel splicing is:
[0151]
[0152] Furthermore, the Swain transformer block is used to divide the input feature map into non-overlapping patches as "tokens", which are projected to the specified dimension through a linear embedding layer. Specifically, the following processing is performed:
[0153] (1) Divide the feature map into non-overlapping s×s patches and project them to dimension d via linear embedding;
[0154] (2) Calculate window-based multi-head self-attention (W-MSA) between patches, with a window size of w×w, satisfying the relationship:
[0155]
[0156] (3) A cross-window connection mechanism is introduced to achieve information interaction between adjacent regions through shifting window division.
[0157] By setting the window size to , ensuring that adjacent windows within an s×s patch overlap by at least one pixel, forcing cross-region information exchange and addressing the problem of texture correlation across tobacco regions in tobacco images. Transposed convolution kernels are generated based on encoder features, enabling the upsampling process to adaptively adjust the reconstruction strategy for different tobacco morphologies (such as bent stems and fragments), improving segmentation edge accuracy.
[0158] The encoder uses densely connected convolution to achieve cross-layer feature reuse by concatenating the output of the previous convolutional layer with the output of all previous layers along the channel dimension and then inputting them into the next layer. The concatenation formula is:
[0159]
[0160] in, Represents the input features of the kth layer.
[0161] Decoder structure: Each upsampling layer is followed by two sets of separable convolution units connected in series. Each set of separable convolution units contains: (1) a 3×3 depthwise separable convolution layer, where the number of filters is halved at each stage; (2) a 1×1 pointwise convolution layer for channel dimension feature reorganization; and (3) a dynamic upsampling kernel generation module that generates the transposed convolution kernel weights based on the feature maps of the corresponding encoder stage:
[0162]
[0163] in, Features passed to the encoder through skip connections;
[0164] The decoder output layer introduces a feature calibration mechanism to weight the reconstructed features through spatial attention weights:
[0165]
[0166] in, It is the reconstructed feature map generated by the decoder network. Its size is the same as the input image (e.g., H×W×C) and contains spatial details (e.g., tobacco stem edges, cigarette arrangement).
[0167] Module Interaction: The bidirectional convolutional long short-term memory network module and the Swain transformer module are integrated into the skip connection and connected to the corresponding stage of the decoder to achieve refined feature extraction and effective fusion, and improve the model's ability to model multi-scale and long-distance dependencies. The fusion method is:
[0168] a. Perform channel compression on the output features of the bidirectional convolutional long short-term memory network module, retaining the first 50% of the channels;
[0169] b. Spatially interpolate the output features of the Swain transformer module to the target resolution;
[0170] c. Add the two to the current features of the decoder element by element, and control the information flow through the gate weight:
[0171]
[0172] Among them, α and β are learnable parameters, with initial values set to 0.3 and automatically optimized through back propagation; The output features of the BiConvLSTM module contain spatiotemporal dependency information (such as the temporal changes in tobacco texture). The output features of the SwinTransformer module contain global attention information (such as long-range correlation of tobacco stems); is the intermediate feature map of the current stage of the decoder, containing upsampled spatial details (such as segmentation edges).
[0173] S132: Constructing a Deep Learning Network Loss Function
[0174] In order to optimize the learning effect of the model and improve the accuracy, the loss function is composed of Dice loss, Jaccard intersection-over-union loss, boundary loss, and mean square error (MSE) loss weighted by the following formula:
[0175]
[0176] in, They are the weighted coefficients of Dice loss, Jaccard loss, boundary loss and MSE, which are used to adjust the influence weight of each loss item; optional, , .
[0177] Dice loss measures the degree of overlap between the model prediction results and the true labels:
[0178]
[0179] in, Respectively represent The true label and predicted value of each pixel; Indicates the total number of pixels; is a small value to avoid the denominator being zero.
[0180] Jaccard intersection-over-union loss is used to measure the intersection-over-union ratio of the predicted results and the true labels:
[0181]
[0182] Among them, the molecule is the intersection of the prediction and the label; the denominator Jaccard loss is more sensitive to small targets than Dice loss and can enhance the model's attention to details.
[0183] It is the boundary loss, which aims to optimize the accuracy of the target edge and is defined as follows:
[0184]
[0185] in, Respectively represent The gradient of the predicted segmentation result of each pixel and the true segmentation label is used to capture the boundary information.
[0186] is the mean squared error loss, which is used to calculate the pixel difference between the model output and the real image:
[0187]
[0188] The mean squared error loss function makes the model output closer to the real image in terms of overall pixel distribution.
[0189] S133: Introducing squeeze-enhanced axial attention mechanism
[0190] The squeeze-enhanced axial attention mechanism uses an adaptive over-detail weighted gating method to efficiently compute and aggregate global information, while enhancing local spatial details and reducing time complexity. This mechanism compresses features in the row / column direction and fuses position embedding with local convolutional enhancement to enhance the neural network's ability to learn features along different axes. Its core components include a squeeze operation, axial attention calculation, position embedding compensation, and spatial detail enhancement.
[0191] The specific implementation of the squeeze-enhanced axial attention mechanism includes the following steps:
[0192] (1) Bidirectional axial compression:
[0193] Input feature map The squeezing operation is performed along the row direction and the column direction respectively, where is the image height, is the image width, is the number of channels, including:
[0194] Row-wise compression: Squeeze the input features in the row direction and map them to a single feature marker for each row. Specifically, sum the input feature map X along the row direction (height dimension), compress all row features of each column into one value, and obtain the row compression feature. , and through the learnable matrix Generate a query ,key ,value , the specific formula is:
[0195]
[0196] in, Represents the result of summing over the row dimension, i.e., the row compression feature; are the projection matrices of query, key, and value in the row direction respectively;
[0197] Column compression: Squeeze the input features in the column direction and map them to a single feature marker in each column. Specifically, sum the input feature map X along the column direction (width dimension) and compress all column features of each row into one value to obtain the column compression feature.
[0198] , and through the learnable matrix Generate a query ,key ,value , the specific formula is:
[0199]
[0200] in, Represents the result of summing over the column dimension, i.e., row compression feature; The projection matrices for query, key, and value in the column direction respectively.
[0201] Through the above input feature map Squeeze operations along the row and column directions respectively, map to a single feature marker in each row, and obtain the compressed features in the row direction and column-wise compressed features .
[0202] Calculate the bidirectional attention weight and get the row-wise attention weight and column-wise attention weights ;in, ;
[0203] Generate bidirectional attention features, that is, row-wise attention features and column-wise attention features ;
[0204] The row-wise attention features are obtained by compressing the input feature map along the row (height dimension) , the width (W) and channel (C) information is retained, but the height dimension is compressed to 1.
[0205] The column-wise attention features are obtained by compressing the input feature map along the column (width dimension) , the height (H) and channel (C) information is retained, but the width dimension is compressed to 1.
[0206] because The spatial dimensions of the two are inconsistent , direct addition is mathematically infeasible, so it is necessary to adjust the dimension through Reshape to achieve fusion.
[0207] (2) Position embedding compensation:
[0208] The positional encoding PE is extracted from the original input feature X through 1×1 convolution:
[0209]
[0210] Represents the original input features, and retains local detail information after fusion; represent Convolution, used to map position information to the embedding space;
[0211] Fusion of positional encoding with bidirectional attention features:
[0212]
[0213] in, is the enhanced axial attention feature, It is the axial attention feature; Represents the dimension expansion operation, which is used to convert the row-wise attention features and column-wise attention features The dimensions are uniformly adjusted to the same spatial size H×W×C as the original input feature map for subsequent fusion and position encoding compensation; Because the position information is lost during the extrusion process, a learnable position embedding is introduced to compensate for it. Adding to the axial attention feature , and obtain the enhanced axial attention feature , which can restore the spatial position information lost due to the compression operation and ensure that the model can distinguish tobacco details in different areas (such as tobacco stem position and cutting edge).
[0214] Assume that the input tobacco image size is 256×256×64 (H×W×C), and row compression is: ,go through After the operation, it is expanded to 256×256×64. Column compression: ,go through After this operation, the image is expanded to 256×256×64. Fusion: After adding the two, the features at each spatial location contain global statistical information for the entire image row and column. Position encoding is then added to further enhance local position perception and improve the accuracy of locating broken tobacco or stem-containing areas.
[0215] This method uses dimensionality expansion to restore row and column-wise global attention features to their original spatial dimensions, enabling cross-directional information fusion and providing a foundation for positional encoding compensation. This dimensionality expansion operation reduces computational complexity while ensuring the model effectively captures global horizontal and vertical dependencies in tobacco images, thereby improving the accuracy of quality predictions (such as stem content and tobacco length uniformity).
[0216] (3) Spatial detail enhancement:
[0217] For the original input features Perform 3×3 convolution to extract local detail features ;
[0218] Fusion of global attention and local features through detail weight gating:
[0219]
[0220] in, Indicates the importance weight of local details that should be preserved at each spatial location; Is the Sigmoid function, used to normalize the weights so that between; Represents element-by-element multiplication (Hadamard product), and the final output feature The complexity is O(HWC).
[0221] The standard global attention computation complexity is , and the squeeze-enhanced axial attention mechanism reduces to , significantly reducing the computational effort on high-resolution images while maintaining the ability to aggregate global information.
[0222] Through step S130, a main encoder-decoder network architecture is constructed. The encoder-decoder serves as the core structure of the backbone network, assuming the core functions of general feature extraction and representation learning. The encoder extracts general features (such as a global representation of tobacco texture and the temporal dependencies of device parameters) through multi-stage convolutional layers (such as BiConvLSTM modules and SwinTransformer modules) and attention mechanisms, which are then shared with multiple task branches. The decoder, as one of the task branches, restores the high-dimensional features extracted by the encoder to higher-spatial-resolution outputs (such as segmentation masks and reconstructed images) through upsampling and separable convolution operations. This serves specific tasks (such as tobacco segmentation and seal integrity inspection). The decoder focuses on tasks requiring high-resolution output (such as segmentation), while other task branches directly utilize the encoder features for lightweight computation (such as classification or regression).
[0223] S140: Multi-task branch learning is introduced on the basis of the main network architecture of the deep learning network, and it is expanded into a multi-task joint learning framework that combines shared feature extraction with multi-task branches. The shared backbone network extracts common features, and multi-task branches are set up to process classification tasks, regression tasks, target detection tasks and image segmentation tasks respectively, so as to realize the joint prediction and detection of multiple key quality indicators in the silk making and winding process.
[0224] This step S140 introduces multi-task branch learning. By employing a deep learning architecture of "shared feature extraction + multi-task branching," the backbone network extracts universal feature information with good generalization capabilities. On this basis, multiple task branch heads are configured to handle tasks such as classification, regression, object detection, and image segmentation. Based on shared features, each branch is optimized for a specific task, enabling the joint prediction and detection of multiple key quality indicators in the silk-making and winding processes, significantly improving overall quality assurance and detection efficiency.
[0225] Furthermore, shared feature extraction is performed: the feature output of the encoder (such as the intermediate layer features of the encoder in step S130) is directly reused as the common input of multiple task branches to avoid repeated calculations.
[0226] Multi-task branches: Based on shared features, independent output heads are designed for different tasks (such as classification, regression, and detection). For example, the decoder itself can serve as one of the task branches (such as image segmentation); other branches (such as stem content prediction and blade state classification) are directly connected to the encoder's intermediate feature layer and processed through fully connected layers or lightweight networks. The multi-task branches are connected as follows: the silk-making task branch (such as stem content prediction) is connected to the encoder's third-stage output features; the wrapping task branch (such as seal detection) is connected to the encoder's fourth-stage output features.
[0227] The multi-task branch extends the functionality of the architecture. By sharing encoder features, it enables multi-objective joint optimization at a low computational cost. For example, tobacco length uniformity scoring (a regression task) and stem content prediction (a classification task) in the tobacco-making process can share encoder-extracted tobacco texture features; while seal integrity determination (a segmentation task) and filter position deviation detection (a regression task) in the wrapping process can share encoder-extracted packaging edge features.
[0228] Finally, the network model predicts and detects the key indicators of tobacco production, including the following prediction results: (1) tobacco length distribution uniformity index. The network automatically analyzes the tobacco image and predicts the uniformity score of the cut length distribution, which is highly correlated with the actual measured tobacco length consistency rate. (2) tobacco width and thickness distribution. By segmenting the tobacco area in the image, the variation range of its cross-sectional width and thickness is counted; the output is a continuous value indicator or classification label (such as qualified / thin / thick). (3) stem content prediction. Based on texture features and color distribution, the number of visible stems in the image is identified and estimated; the output is a percentage of the stem content prediction value. (4) tobacco body forming rate. The network output prediction value represents the probability that the tobacco maintains structural integrity during the forming process. (5) cutting blade condition index. The blade condition is indirectly predicted, with the edge clarity and cut shape in the image as a reference; the output is a blade condition classification label (normal / dulled / severely dulled).
[0229] The prediction and detection of key indicators of the package include: (1) Package seal integrity determination, outputting binary or multi-classification results, indicating whether there are problems such as missing seals or warping in the seal area. (2) Filter position deviation detection, the model automatically measures the deviation of the filter position from the standard position and outputs the offset value or offset level. (3) Trademark printing clarity scoring, through image texture and edge information analysis, outputs the print pattern clarity index. (4) Foreign matter / stain detection, using the target detection module to output the location information and type classification results of suspected foreign matter.
[0230] In summary, the present invention combines artificial intelligence and deep learning technologies to propose a series of innovative solutions for the key indicator prediction and quality assurance issues of the silk-making and packaging processes, significantly improving the prediction accuracy, optimizing the feature extraction capability, enhancing the model attention mechanism, and realizing multi-task joint learning and real-time efficient detection, providing strong technical support for quality control in the tobacco industry.
[0231] Test and comparative experiments: Tested on a cigarette factory dataset:
[0232] Dataset: 100,000 tobacco images (including manually annotated tobacco stem areas and the true stem content rate).
[0233] Comparison models: Baseline 1: traditional level set algorithm (Gaussian kernel); Baseline 2: U-Net + ResNet50; Baseline 3: the proposed method.
[0234] Evaluation indicators: Dice coefficient (segmentation accuracy); prediction error of stem rate (absolute percentage error, APE); single image inference time (ms).
[0235] Experimental results: see Table 1.
[0236] Table 1 Experimental results
[0237]
[0238] Test experiments show that the Dice coefficient of the proposed method is improved by 24% (0.75→0.93), and the prediction error of the stem rate is reduced by 71% (8.7%→2.5%). The inference speed is 2 times faster than that of U-Net, meeting the real-time detection requirements of the production line (≤30ms).
[0239] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.
[0240] It should also be noted that, in the embodiments of the present application, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements.
[0241] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the embodiments of the present application may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown in the embodiments of the present application, but rather will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.
Claims
1. An AI-based method for predicting key indicators of silk-making process and ensuring the quality of winding and packaging, characterized in that: include: Collect data from the tobacco making and wrapping processes, including physical parameters, equipment status data, and tobacco leaf / cut tobacco images from the tobacco making process, as well as cigarette parameters, equipment data, and packaging images from the wrapping process; Preprocessing and feature extraction are performed on the silk-making process data and the winding process data respectively to obtain a silk-making process feature vector and a winding process feature vector; The silk-making process feature vector and the wrapping process feature vector are fed into a deep learning network and trained using a custom loss function and a squeeze-enhanced axial attention mechanism; wherein the deep learning network comprises an encoder-decoder main network architecture, wherein the encoder comprises four stages, each stage having two 3×3 depthwise separable convolutional layers, followed by a 2×2 maximum pooling layer and a ReLU activation function, with the number of filters doubling at each stage for extracting visual features; the decoder comprises multiple separable convolutional layers and upsampling layers, upsampling the encoder output through a transposed convolution operation, and combining features from the encoder to achieve image reconstruction and segmentation; The custom loss function includes a weighted combination of Dice loss, Jaccard loss, boundary loss and MSE: in, They are the weighted coefficients of Dice loss, Jaccard loss, boundary loss and MSE, which are used to adjust the influence weight of each loss item; The squeeze-enhanced axial attention mechanism compresses features in the row / column direction and fuses position embedding with local convolution enhancement to enhance the feature learning ability of the neural network in different axes. Multi-task branch learning is introduced on the basis of the main network architecture of the deep learning network, and it is expanded into a multi-task joint learning framework that combines shared feature extraction with multi-task branches. The shared backbone network extracts common features, and multi-task branches are set up to respectively process classification tasks, regression tasks, target detection tasks and image segmentation tasks, so as to realize the joint prediction and detection of multiple key quality indicators in the silk making and winding processes.
2. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 1 is characterized in that: in, Preprocessing the silk making process data and the winding process data respectively includes: The time series data of the silk-making process data and the winding process data are cleaned and normalized, and the image data are subjected to color enhancement, contrast enhancement and image segmentation.
3. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 2, characterized in that: in, Feature extraction of pre-processed silk-making process data includes: Extracting features related to key indicators from the physical parameters of the silk-making process to obtain the physical parameter characteristics of silk-making, including: extracting the heating rate and humidity fluctuation range by analyzing the temperature and humidity change curves; Extract features related to key indicators from the equipment operation data of the silk-making process to obtain equipment operation characteristics, including speed and fault status; Extract image features from tobacco leaf and tobacco cut images to obtain tobacco cut image features, including color features, macro texture features, and shape features; The silk-making physical parameter characteristics, the equipment operation characteristics and the silk-making image characteristics are integrated and dimension reduction is performed to obtain the silk-making process feature vector.
4. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 3 is characterized in that: in, Feature extraction of pre-processed packaging process data includes: Extract features related to packaging quality, including cigarette weight stability; Extracting the sealing quality characteristics of packaging from the operating parameters of packaging equipment; Extracting image features from the package image to obtain package image features, including cigarette arrangement regularity features and packaging pattern integrity features; The cigarette weight stability feature, the sealing quality feature and the wrapping image feature are integrated and dimensionality reduction is performed to obtain the wrapping process feature vector.
5. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 1 is characterized in that: in, The encoder structure includes: The 3×3 depth-wise separable convolutional layer at each stage is composed of a bidirectional convolutional long short-term memory network BiConvLSTM module and a Swin Transformer module in series; The bidirectional convolutional long short-term memory network module is used to divide the input feature map into blocks of size m×n and perform conditional position encoding on each block. Convolution operations are integrated into the input-to-state and state-to-state transformations, and the hidden states of the forward and backward paths are fused through channel splicing to enhance the network's ability to learn complex patterns and relationships. The Swain transformer module is used to divide the input feature map into non-overlapping s×s patches as "tokens", which are projected to the specified dimension through a linear embedding layer; the window-based multi-head self-attention W-MSA is calculated between the patches, with a window size of w×w, satisfying the relationship: ;Introduce a cross-window connection mechanism to achieve information interaction between adjacent regions through shifting window division; The encoder uses densely connected convolution to achieve cross-layer feature reuse by concatenating the output of the previous convolutional layer with the output of all previous layers along the channel dimension and then inputting them into the next layer; The bidirectional convolutional long short-term memory network module and the Swain transformer module are integrated into the skip connection and connected to the corresponding stage of the decoder, so as to realize the refined extraction and effective fusion of features and improve the model's ability to model multi-scale and long-distance dependencies.
6. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 5, characterized in that: in, The decoder structure includes: Each upsampling layer is followed by two sets of separable convolutional units connected in series. Each set of separable convolutional units consists of: a 3×3 depthwise separable convolutional layer with the number of filters halved at each stage; a 1×1 pointwise convolutional layer for channel-dimensional feature reorganization; and a dynamic upsampling kernel generation module that generates transposed convolution kernel weights based on the feature maps of the corresponding encoder stage. Among them, the decoder output layer introduces a feature calibration mechanism to weight the reconstructed features through spatial attention weights.
7. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 1 is characterized in that: in, The specific implementation of the squeeze-enhanced axial attention mechanism includes the following steps: (1) Bidirectional axial compression: Input feature map Squeeze operations are performed along the row and column directions, mapping to a single feature marker for each row, where is the image height, is the image width, is the number of channels, and the features after row-wise compression are obtained and column-wise compressed features Calculate the bidirectional attention weight and get the row-wise attention weight and column-wise attention weights ;in, ; Generate bidirectional attention features, that is, row-wise attention features and column-wise attention features ; (2) Position embedding compensation: The positional encoding PE is extracted from the original input feature X through 1×1 convolution: Fusion of positional encoding with bidirectional attention features: in, is the enhanced axial attention feature, It is the axial attention feature; Represents the dimension expansion operation, which is used to convert the row-wise attention features and column-wise attention features The dimensions are uniformly adjusted to the same spatial size H×W×C as the original input feature map; (3) Spatial detail enhancement: Perform 3×3 convolution on the original input feature X to extract local detail features ; Fusion of global attention and local features through detail weight gating: in, is the Sigmoid function, which is used to normalize the weights so that between; Represents element-by-element multiplication, and the final output feature The complexity is O(HWC).
8. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 7, characterized in that: in, The input feature map The extrusion operations are performed along the row and column directions respectively, including: Row-wise compression: sum the features of each row to get the row compression features , and through the learnable matrix Generate a query , the specific formula is: in, Represents the result of summing over the row dimension, i.e., the row compression feature; are the projection matrices of query, key, and value in the row direction respectively; Column-wise compression: Sum the features of each column to get the column-wise compression features , and through the learnable matrix Generate a query , the specific formula is: in, Represents the result of summing over the column dimension, i.e., row compression feature; The projection matrices for query, key, and value in the column direction respectively.
9. The AI-based method for predicting key indicators of silk-making process and ensuring quality of winding and packaging according to claim 1, characterized in that: in, The multi-task branches include: a silk making process task branch, which is used to output the silk length distribution uniformity score, tobacco width and thickness classification, stem content percentage, and blade status classification; and a wrapping process task branch, which is used to output the sealing integrity classification, filter offset value, trademark clarity score and foreign object detection box.
Citation Information
Patent Citations
Quasi-circular object recognition counting detection algorithm based on machine vision and deep learning
CN111523535A
Cigarette brand recommendation method based on knowledge graph in new retail mode
CN115618108A