Classification method and device, classification model training method and device and electronic equipment

By obtaining and processing the chemical composition and category characteristic data of primary tobacco leaves and inputting them into the classification model, the problems of low efficiency and high subjectivity of traditional manual classification are solved, and efficient and accurate automatic classification of tobacco leaves grades are achieved.

CN120145152APending Publication Date: 2025-06-13CHINA TOBACCO FUJIAN IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510259754.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The classification of traditional primary tobacco leaves ratings relies on manual observation, is inefficient and subjective, making it difficult to ensure the accuracy and consistency of classification.

Method used

By obtaining the chemical composition characteristic data and category characteristic data of the primary tobacco leaves, performing embedding representation processing and adding classification marks, inputting them into the classification model for automatic hierarchical classification.

Benefits of technology

The efficiency and accuracy of the classification of the grade of primary tobacco leaves is improved, the risk of manual misjudgment is reduced, and the automatic classification of the grade of tobacco leaves is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145152A_ABST
    Figure CN120145152A_ABST
Patent Text Reader

Abstract

The invention relates to a classification method and device, a classification model training method and device and electronic equipment, and relates to the technical field of tobacco leaf processing. The classification method comprises the steps that first feature data of flue-cured tobacco leaves are acquired, and the first feature data comprise chemical component feature data and category feature data of the flue-cured tobacco leaves; performing embedding representation processing on the first feature data to obtain second feature data; adding a classification mark in the second feature data to obtain third feature data; and inputting the third feature data into the classification model to obtain the grade of the flue-cured tobacco leaves.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of tobacco processing, and particularly relates to a classification method, a classification device, a training method of a classification model, a training device of a classification model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The primary baking of tobacco leaves is a process of converting freshly harvested tobacco leaves in the field into raw tobacco with certain quality, style, and grade standards. This link can effectively remove excess moisture and help the tobacco leaves retain their flavor and aroma, laying a good foundation for subsequent threshing and redrying and fermentation processes. Summary of the Invention

[0003] The inventors found that: for the tobacco leaves after primary baking, industrial grading is required to make the chemical components of the tobacco leaves of the same grade after grading more consistent, reduce the nicotine coefficient of variation of the tobacco leaves before processing, enhance the nicotine stability of the feeding tobacco leaves, and ensure the stability and consistency of the quality of the processed cut tobacco. In the traditional tobacco leaf processing link, the grade classification of the tobacco leaves after primary baking mainly relies on manual observation of the appearance of the tobacco leaves after primary baking. The grading workers will roughly classify the tobacco leaves after primary baking into different grade categories according to seven appearance grade factors such as maturity, leaf structure, identity, oil content, chroma, length, and damage.

[0004] Therefore, how to ensure the efficiency and reliability of the process of classifying the primary baked tobacco leaves (tobacco leaves after primary baking) is a problem to be solved.

[0005] In view of this, the present disclosure provides a classification method. According to some embodiments of the first aspect of the present disclosure, a classification method is provided, including: obtaining first feature data of the primary baked tobacco leaves, where the first feature data includes chemical component feature data and category feature data of the primary baked tobacco leaves; performing an embedding representation process on the first feature data to obtain second feature data; adding a classification mark to the second feature data to obtain third feature data; and inputting the third feature data into a classification model to obtain the grade of the primary baked tobacco leaves.

[0006] In some embodiments, performing an embedding representation process on the first feature data to obtain second feature data includes: performing a first embedding representation process on the chemical component feature data to obtain first sub-feature data; performing a second embedding representation process on the category feature data to obtain second sub-feature data; and obtaining the second feature data according to the first sub-feature data and the second sub-feature data.

[0007] In some embodiments, performing a first embedding representation process on chemical composition feature data to obtain first sub-feature data includes: calculating the product of the chemical composition feature data and a first vector to obtain a first product; calculating the sum of the first product and first feature bias data to obtain the first sub-feature data, where the first vector and the first feature bias data are determined during the training process of the embedding representation model.

[0008] In some embodiments, performing a second embedding representation process on category feature data to obtain second sub-feature data includes: determining one-hot representation data corresponding to the category feature data according to the category feature data; calculating the product of the one-hot representation data and a second vector to obtain a second product; calculating the sum of the second product and second feature bias data to obtain the second sub-feature data, where the second vector and the second feature bias data are determined during the training process of the embedding representation model.

[0009] In some embodiments, obtaining first feature data of the primary-cured tobacco leaves includes: performing normalization processing on the original feature data of the primary-cured tobacco leaves to obtain the first feature data.

[0010] In some embodiments, performing normalization processing on the original feature data of the primary-cured tobacco leaves to obtain the first feature data includes: calculating the median and interquartile range of the original feature data of the primary-cured tobacco leaves; performing normalization processing on the original feature data of the primary-cured tobacco leaves according to the original feature data, median, and interquartile range of the primary-cured tobacco leaves to obtain the first feature data.

[0011] In some embodiments, the category feature data includes at least one of the production area and tobacco leaf variety of the primary-cured tobacco leaves.

[0012] In some embodiments, the classification model includes one or more layers of Transformer encoders and a linear neural network classifier, where, in the case where the classification model includes multiple layers of Transformer encoders, the multiple layers of Transformer encoders are connected in series.

[0013] According to some embodiments of the second aspect of the present disclosure, there is provided a method for training a classification model, including: obtaining first feature training data of primary-cured tobacco leaves, where the first feature training data includes chemical composition feature data and category feature data of the primary-cured tobacco leaves; determining grade annotation information of the first feature training data according to the first feature training data; performing an embedding representation process on the first feature training data to obtain second feature training data; adding a classification marker to the second feature training data to obtain third feature training data; training the classification model according to the third feature training data and the grade annotation information.

[0014] According to some embodiments of the third aspect of the present disclosure, a classification device is provided, including: a first acquisition unit configured to acquire first feature data of initially baked tobacco leaves, where the first feature data includes chemical composition feature data and category feature data of the initially baked tobacco leaves; a first embedding representation unit configured to perform an embedding representation process on the first feature data to obtain second feature data; a first addition unit configured to add a classification label to the second feature data to obtain third feature data; and an input unit configured to input the third feature data into a classification model to obtain the grade of the initially baked tobacco leaves.

[0015] According to some embodiments of the fourth aspect of the present disclosure, a training device for a classification model is provided, including: a second acquisition unit configured to acquire first feature training data of initially baked tobacco leaves, where the first feature training data includes chemical composition feature data and category feature data of the initially baked tobacco leaves; a determination unit configured to determine grade annotation information of the first feature training data according to the first feature training data; a second embedding representation unit configured to perform an embedding representation process on the first feature training data to obtain second feature training data; a second addition unit configured to add a classification label to the second feature training data to obtain third feature training data; and a training unit configured to train the classification model according to the third feature training data and the grade annotation information.

[0016] According to some embodiments of the fifth aspect of the present disclosure, an electronic device is provided, including: a memory and a processor coupled to the memory, where the processor is configured to execute the classification method in any of the above embodiments or the training method of the classification model in any of the above embodiments based on instructions stored in the memory.

[0017] According to some embodiments of the sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the classification method in any of the above embodiments or the training method of the classification model in any of the above embodiments is implemented.

[0018] According to some embodiments of the seventh aspect of the present disclosure, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the classification method in any of the above embodiments or the training method of the classification model in any of the above embodiments is implemented.

[0019] In the above embodiments, the first feature data including the chemical composition characteristic data and the category characteristic data of the primary roasted tobacco leaves is obtained, the first feature data is subjected to an embedding representation process and a classification label is added to obtain the third feature data, and finally the third feature data is input into the classification model to obtain the grade of the primary roasted tobacco leaves, ensuring the efficiency and reliability of the process of classifying the primary roasted tobacco leaves. When obtaining the feature data of the primary roasted tobacco leaves, the chemical composition characteristic data and the category characteristic data of the primary roasted tobacco leaves are considered, and relatively complete feature data is obtained, improving the accuracy and reliability of classifying the grade of the primary roasted tobacco leaves through the classification model. In addition, by classifying the grade of the primary roasted tobacco leaves through the classification model, useful feature information can be extracted from a large amount of third feature data to achieve automatic classification of the tobacco leaf grade, which is faster than manual classification, improving the efficiency of the process of classifying the primary roasted tobacco leaves, and also reducing the risk of errors caused by manual classification of the grade of the primary roasted tobacco leaves. Furthermore, through the embedding representation process of the first feature data, a data format beneficial to model processing can be obtained, improving the efficiency of classifying the grade of the primary roasted tobacco leaves. By adding a classification label to the second feature data, the global information in the second feature data can be integrated, improving the accuracy of classifying the grade of the primary roasted tobacco leaves. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings forming a part of the specification illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0021] With reference to the drawings, the present disclosure can be more clearly understood from the following detailed description.

[0022] Figure 1 Schematic diagrams showing some embodiments of the classification method of the present disclosure.

[0023] Figure 2 Schematic diagrams showing some embodiments of the embedding representation process of the present disclosure.

[0024] Figure 3 Schematic diagrams showing some other embodiments of the classification method of the present disclosure.

[0025] Figure 4 Schematic diagrams showing some embodiments of the training method of the classification model of the present disclosure.

[0026] Figure 5 Schematic diagrams showing some further embodiments of the classification method of the present disclosure.

[0027] Figure 6 Schematic diagrams showing some embodiments of the classification device of the present disclosure.

[0028] Figure 7Schematic diagrams showing some embodiments of a training apparatus for a classification model of the present disclosure.

[0029] Figure 8 Schematic diagrams showing some embodiments of an electronic device of the present disclosure. Detailed implementation manners

[0030] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0031] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.

[0032] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present disclosure and its application or use.

[0033] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the description.

[0034] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0035] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0036] Currently, the method for classifying the grades of initially cured tobacco leaves relying on manual observation has the following drawbacks: on the one hand, the efficiency of manual grade classification is relatively low. Especially in the case of large-scale tobacco leaf production, it requires a large amount of human and time costs. With the rapid development of the cigarette industry, the traditional grade classification method relying on manual observation is increasingly difficult to meet the fast and efficient production rhythm of the modern tobacco processing industry. On the other hand, due to the inevitable subjectivity of the traditional grade classification method relying on manual observation, different grade classification workers may have differences in the grade judgment of the same initially cured tobacco leaf, resulting in difficulties in ensuring the accuracy and consistency of the grade classification of initially cured tobacco leaves, and further affecting the subsequent tobacco processing formula and the quality stability of cigarette products.

[0037] Artificial intelligence technology is a current research hotspot. If the grading of flue-cured tobacco leaves can be assisted or even replaced by workers through deep learning algorithms, the efficiency and reliability of the grading of flue-cured tobacco leaves can be effectively guaranteed, which has great and urgent practical significance for reducing manual misjudgment and ensuring the consistency of cigarette quality.

[0038] Regarding how to ensure the efficiency and reliability of the process of classifying flue-cured tobacco leaves (tobacco leaves after primary baking), it is as follows.

[0039] Figure 1 Schematic diagrams showing some embodiments of the classification method of the present disclosure.

[0040] As Figure 1 shown, the classification method includes step 110 to step 140, and this classification method is executed by a classification device.

[0041] In step 110, first feature data of the flue-cured tobacco leaves is obtained, where the first feature data includes chemical composition feature data and category feature data of the flue-cured tobacco leaves.

[0042] For example, the chemical composition feature data of the flue-cured tobacco leaves can be chemical components such as nicotine content and total sugar. The chemical composition feature data is a numerical value, and the category feature data is not a numerical value but a category option.

[0043] In some embodiments, the chemical composition characteristic data of the initially baked tobacco leaves may be, for example, total alkaloids, reducing sugars, total sugars, total nitrogen, potassium, chlorine, pH value, starch, dichloromethane extract, solanesol, sulfate radical, phosphate radical, magnesium, calcium, neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, rutin, oxalic acid, malonic acid, succinic acid, malic acid, citric acid, vanillic acid, myristic acid, palmitic acid, linoleic acid, oleic acid + linolenic acid, stearic acid, arachidic acid, aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, proline, glucosamine (Glu-An), fructosylaminobutyric acid (Fru-Amb), fructosylhistidine (Fru-His), fructosylproline (Fru-Pro), fructosylvalline (Fru-Val), fructosylthreonine (Fru-Thr), fructosylglycine (Fru-Gly), fructosylalanine (Fru-Ala), fructosylasparagine (Fru-Asn), fructosylaspartic acid (Fru-Asp), fructosylglutamine (Fru-Gln), fructosylglutamic acid (Fru-Glu), fructosylisoleucine (Fru-Ile), fructosylleucine (Fru-Leu), fructosyltyrosine (Fru-Tyr), fructosylphenylalanine (Fru-Phe), fructosyltryptophan (Fru-Trp), neophytadiene, etc.

[0044] For example, the chemical composition characteristic data of the initially baked tobacco leaves can be determined by means of near-infrared scanning of the initially baked tobacco leaves.

[0045] For example, the category characteristic data of the initially baked tobacco leaves may be the production area and tobacco leaf variety of the initially baked tobacco leaves, etc.

[0046] By taking into account the chemical composition characteristic data and category characteristic data of the initially baked tobacco leaves in the process of obtaining the first characteristic data, the scientificity and accuracy of the grade classification of the initially baked tobacco leaves are effectively improved, promoting the further development of the tobacco leaf processing industry in terms of quality control and resource utilization. It provides more objective data support for grading.

[0047] Figure 2 A schematic diagram showing other embodiments of the classification method of the present disclosure.

[0048] As Figure 2 shown, the first characteristic data 21 of the initially baked tobacco leaves is composed of the chemical composition characteristic data x num and the category characteristic data x cat of the initially baked tobacco leaves.

[0049] In step 120, embedding representation processing is performed on the first characteristic data to obtain second characteristic data.

[0050] By performing embedding representation processing on the first feature data, a data format suitable for model processing can be obtained, improving the efficiency and accuracy of the classification of the classification model.

[0051] As Figure 2 shown, after performing embedding representation processing on the first feature data, the second feature data 22 is obtained.

[0052] In step 130, a classification label is added to the second feature data to obtain the third feature data.

[0053] For example, the classification label can be a CLS label.

[0054] As Figure 2 shown, after adding the CLS label to the second feature data, the third feature data 23 is obtained, where the CLS label is added to the front end of the second feature data 22.

[0055] In step 140, the third feature data is input into the classification model to obtain the grade of the initial-cured tobacco leaves.

[0056] For example, the classification model includes one or more layers of Transformer encoders and a linear neural network classifier, where, in the case where the classification model includes multiple layers of Transformer encoders, the multiple layers of Transformer encoders are connected in series.

[0057] As Figure 2 shown, when the third feature data 23 is input into the classification model, first, the third feature data will pass through one or more layers of Transformer encoders to obtain the feature data 24 processed by the Transformer encoder, and finally the feature data 24 processed by the Transformer encoder is sent to the classifier to obtain the classification result (i.e., the grade) corresponding to the first feature data. .

[0058] In the above embodiments, the first feature data including the chemical composition feature data and the category feature data of the primary-cured tobacco leaves is obtained, the first feature data is subjected to an embedding representation process and a classification label is added to obtain the third feature data, and finally the third feature data is input into the classification model to obtain the grade of the primary-cured tobacco leaves, ensuring the efficiency and reliability of the process of classifying the primary-cured tobacco leaves. When obtaining the feature data of the primary-cured tobacco leaves, the chemical composition feature data and the category feature data of the primary-cured tobacco leaves are considered, and relatively complete feature data is obtained, improving the accuracy and reliability of classifying the grade of the primary-cured tobacco leaves through the classification model. In addition, by classifying the grade of the primary-cured tobacco leaves through the classification model, useful feature information can be extracted from a large amount of third feature data to realize the automatic classification of the tobacco leaf grade, which is faster than manual classification, improving the efficiency of the process of classifying the primary-cured tobacco leaves, and also reducing the risk of errors caused by manual classification of the grade of the primary-cured tobacco leaves. In addition, through the embedding representation process of the first feature data, a data format beneficial to model processing can be obtained, improving the efficiency of classifying the grade of the primary-cured tobacco leaves. By adding a classification label to the second feature data, the global information in the second feature data can be integrated, improving the accuracy of classifying the grade of the primary-cured tobacco leaves.

[0059] For step 110, the first feature data of the primary-cured tobacco leaves is obtained. For example, the first feature data can be obtained by normalizing the original feature data of the primary-cured tobacco leaves.

[0060] In some embodiments, the median and interquartile range of the original feature data of the primary-cured tobacco leaves are calculated; according to the original feature data, median and interquartile range of the primary-cured tobacco leaves, the original feature data of the primary-cured tobacco leaves is normalized to obtain the first feature data, as follows.

[0061] For example, the original feature data set X = {X 1 , X 2 ,..., X n}, where X i = {x i-num , x i-cat} represents that the original feature data includes the chemical composition feature data x i-num and the category feature data x i-cat of the primary-cured tobacco leaves, and n is the number of specific original feature data in the original feature data set. Calculate the median Median(X) and interquartile range IQR(X) of the original feature data set. The normalized original feature data is X i-scaled = (X i - Median(X)) / (IQR(X)), i = 1, 2, 3... n.

[0062] By normalizing the original characteristic data of the initially baked tobacco leaves, the characteristics of different magnitudes are unified into a comparable range, reducing the risk that the classification model overemphasizes the characteristics with large magnitudes and ignores the characteristics with small magnitudes due to the characteristic data of different dimensions and magnitudes.

[0063] For step 120, perform an embedding representation process on the first characteristic data. For example, the chemical composition characteristic data can be subjected to a first embedding representation process to obtain first sub-characteristic data; the category characteristic data can be subjected to a second embedding representation process to obtain second sub-characteristic data; and the second characteristic data is obtained based on the first sub-characteristic data and the second sub-characteristic data.

[0064] Figure 3 A schematic diagram showing some embodiments of the embedding representation process of the present disclosure.

[0065] As Figure 3 shown, perform an embedding representation process on the first characteristic data to obtain an embedding representation T of d dimensions i (i.e., the second characteristic data), and the embedding representation process is as follows.

[0066] (1)

[0067] (2)

[0068] Among them, b i represents the bias of the i-th first characteristic data, and f i represents a function for converting the first characteristic data into an embedding representation, which is mainly used to convert the dimension of the first characteristic data into d dimensions. x i is a set of numerical values, including one or more chemical composition characteristic data and one or more category characteristic data. The first characteristic data includes two types of characteristic data, one is chemical composition characteristic data, and the other is category characteristic data. R d is used to indicate a d-dimensional vector, that is, f i is mainly used to convert x i into a d-dimensional vector. Specifically, f i generates the embedding representation of the first characteristic data by calculating the product of the relevant data of the first characteristic data (for the chemical composition characteristic data, the relevant data is the chemical composition characteristic data itself; for the category characteristic data, the relevant data is the one-hot representation corresponding to the category characteristic data) and the vector.

[0069] In some embodiments, for the first embedding representation processing of chemical composition feature data, it is as follows: calculate the product of the chemical composition feature data and the first vector to obtain the first product; calculate the sum of the first product and the first feature bias data to obtain the first sub-feature data, where the first vector and the first feature bias data are determined during the training process of the embedding representation model.

[0070] For example, the product of the chemical composition feature data and the first vector can be the product obtained by element-wise multiplication of the chemical composition feature data and the first vector.

[0071] (3)

[0072] Among them, represents the embedding representation corresponding to the chemical composition feature data, is the first vector, which is a d-dimensional data, represents the chemical composition feature data, indicates that the embedding representation corresponding to the chemical composition feature data is a d-dimensional vector.

[0073] In some embodiments, for the second embedding representation processing of category feature data, it is as follows: according to the category feature data, determine the one-hot representation data corresponding to the category feature data; calculate the product of the one-hot representation data and the second vector to obtain the second product; calculate the sum of the second product and the second feature bias data to obtain the second sub-feature data, where the second vector and the second feature bias data are determined during the training process of the embedding representation model.

[0074] For example, the product of the category feature data and the second vector can be the product obtained by element-wise multiplication of the category feature data and the second vector.

[0075] (4)

[0076] Among them, represents the embedding representation corresponding to the category feature data, is the second vector, which is a d-dimensional data, is the one-hot representation corresponding to the category feature data. For example, if there are three category feature data for the first-cured tobacco leaves, then the first category feature data can be [1, 0, 0], the second category feature data can be [0, 1, 0], and the third category feature data can be [0, 0, 1].

[0077] (5)

[0078] where n (num) represents the number of chemical composition feature data, n (cat)It represents the quantity of categorical feature data. After performing embedding representation processing on the first feature data, second feature data of n×d dimensions is obtained. Additionally, the stack function is mainly used to stack the embedding representations corresponding to multiple first feature data to obtain the second feature data.

[0079] As Figure 3 shown, x (num) is chemical composition feature data (for example, 0.4, 0.2, 1.8, etc.), W (num) is the first vector, and b (num) is the first feature bias of the chemical composition feature data. By calculating the product of x (num) and W (num) and then adding b (num) , the embedding representation of the chemical composition feature data can be obtained. Similarly, x 1 (cat) and x 2 (cat) are categorical feature data, W 1 (cat) and W 2 (cat) are the second vectors, and b 1 (cat) and b 2 (cat) are the second feature biases. First, by determining the one-hot representations corresponding to x 1 (cat) and x 2 (cat) , and then by calculating the product of the one-hot representations and W 1 (cat) or W 2 (cat) and then adding b 1 (cat) or b 2 (cat) , the embedding representation of the categorical feature data can be obtained. Finally, by stacking the embedding representations of one or more chemical composition feature data and the embedding representations of one or more categorical feature data, second feature data of n×d dimensions can be obtained.

[0080] By adopting different embedding representation processing methods for chemical composition feature data and categorical feature data, it is possible to better extract the features of different feature data, which helps to improve the efficiency and accuracy of the classification model for hierarchical classification. By introducing feature biases during the process of embedding the first feature data, it can help the embedding representation model find the optimal activation threshold and improve the overall performance of the embedding representation model.

[0081] In some embodiments, the first vector, the second vector, the first feature bias, and the second feature bias are parameters of the embedding representation model, and the first vector, the second vector, the first feature bias, and the second feature bias are all determined during the training process of the embedding representation model. Specifically, the input of the embedding representation model is the first feature training data, and the output is the embedding representation corresponding to the first feature training data. According to the embedding representation corresponding to the first feature training data and the target embedding representation of the first feature training data, the embedding representation model is trained to make the embedding representation corresponding to the first feature training data approximate the target embedding representation of the first feature training data, so as to optimize the first vector, the second vector, the first feature bias, and the second feature bias.

[0082] For step 130, classification tokens are added to the second feature data. For example, the stack function can be used. Among them, the classification token (CLS token) can effectively integrate the global information of the second feature data and serve as the core input of the subsequent classification model, as follows.

[0083] (6)

[0084] For step 140, the third feature data is input into the classification model. The third feature data will go through multiple layers of Transformer encoders. Among them, the input of the latter layer of Transformer encoder is the output of the previous layer of Transformer encoder, and the input of the first layer of Transformer encoder is T 0 . The relationship between the output of the previous layer of Transformer encoder and the input of the latter layer of Transformer encoder is shown in formula (7), where O t is the processing function of the Transformer encoder, C t is the output of the latter layer, and C t-1 is the input of the previous layer.

[0085] (7)

[0086] (8)

[0087] Among them, A is the attention weight matrix of the Transformer encoder, representing the correlation between the embedding representations of each feature data including the classification token and the embedding representations of other feature data. is the query matrix, is the key matrix, and d k is the dimension of the key vector, and n is the number of feature data input to the Transformer encoder.

[0088] The classification model is made interpretable through the attention weight matrix, that is, it can clearly know which indicators play a key role in classification. By taking the square root of the key vector, the data is balanced, which helps the convergence of the classification model.

[0089] After being processed by multiple layers of Transformer encoders, the third feature data containing the [CLS] token is obtained , where L represents the number of layers of the Transformer encoder. It is input into a classifier composed of a fully connected neural network (Linear) using the Softmax activation function to obtain the final classification result , as shown in formula (9).

[0090] (9)

[0091] Among them, LayerNorm is the layer normalization operation.

[0092] Figure 4 A schematic diagram showing some embodiments of the training method of the classification model of the present disclosure.

[0093] As Figure 4 shown, the training method of the classification model includes steps 410 to 450, and the training method of the classification model is executed by the training device of the classification model.

[0094] In step 410, the first feature training data of the primary roasted tobacco leaves is obtained, where the first feature training data includes the chemical composition feature data and the category feature data of the primary roasted tobacco leaves.

[0095] For example, the category feature data of the primary roasted tobacco leaves can be the production area and tobacco variety of the primary roasted tobacco leaves, etc.

[0096] In step 420, according to the first feature training data, the grade annotation information of the first feature training data is determined.

[0097] For example, the training set formed by the first feature training data of the primary roasted tobacco leaves and the corresponding grade annotation information can be , where y i is the grade annotation information corresponding to the first feature training data x i .

[0098] For example, the grade annotation information can be carried out by professional sorting workers. In order to reduce the error of manual annotation, the grade annotation information can be cross-validated by two or more independent sorting workers, and some first feature training data with low annotation consistency can be eliminated.

[0099] In step 430, the first feature training data is subjected to an embedding representation process to obtain the second feature training data.

[0100] In step 440, classification labels are added to the second feature training data to obtain the third feature training data.

[0101] In step 450, the classification model is trained according to the third feature training data and the grade annotation information.

[0102] In some embodiments, the training method of the classification model includes parameter setting, number of iteration rounds, and accuracy analysis.

[0103] For example, the data set is randomly divided and divided into a training set, a validation set, and a test set in a ratio of 8:1:1. Appropriate training parameters are selected to train the classification model with the training set. The dimension of the second feature data is 32; the number of Transformer encoder layers is set to 6; the number of Transformer encoder attention heads is set to 6, and the random inactivation rate of the self-attention layer is set to 0.1; the number of neurons in the linear neural network classifier is 64; the model learning rate is 1e-5; the number of rounds of model iteration is 30; the size of each batch of training set data is 64. In the training process, the input third feature data T is passed through the layers of the classification model using the forward propagation algorithm, and the output result of the classification model is calculated and obtained. , and the error of the network is calculated using backpropagation and the gradient, which are propagated backward from the output layer to the input layer, and the weights and biases of the model are updated using the stochastic gradient descent optimization algorithm. After the iteration is completed, the trained classification model is obtained.

[0104] After the model training is completed, the test set is used to evaluate the classification performance, and the evaluation criterion is the weighted F1 value (weight-F1), which represents the accuracy of the classification model. In this embodiment, the weighted F1 value of the grade classification of the flue-cured tobacco leaves after initial curing reaches 0.9308.

[0105] In the above embodiment, the chemical composition feature data and the category feature data of the flue-cured tobacco leaves are considered in the feature training data for training the classification model, that is, the feature data in two aspects are considered, which improves the accuracy of the classification model for the grade classification of the flue-cured tobacco leaves. By performing embedding representation processing on the first feature training data, a data format that is beneficial for the classification model to process can be obtained, which improves the efficiency of the grade classification of the flue-cured tobacco leaves. By adding classification labels to the second feature training data, the global information in the second feature training data can be integrated, which improves the accuracy of the classification model for the grade classification of the flue-cured tobacco leaves.

[0106] Regarding step 410, the first feature training data of the flue-cured tobacco leaves is obtained. For example, the first feature training data can be obtained by normalizing the original feature training data of the flue-cured tobacco leaves.

[0107] In some embodiments, calculate the median and interquartile range of the original feature training data of the initially baked tobacco leaves; according to the original feature training data, median and interquartile range of the initially baked tobacco leaves, perform normalization processing on the original feature training data of the initially baked tobacco leaves to obtain the first feature training data.

[0108] For step 430, perform embedding representation processing on the first feature training data. For example, the first sub-feature training data can be obtained by performing first embedding representation processing on the chemical composition feature data; the second sub-feature training data can be obtained by performing second embedding representation processing on the category feature data; according to the first sub-feature training data and the second sub-feature training data, obtain the second feature training data.

[0109] In some embodiments, for performing first embedding representation processing on the chemical composition feature data, the specific steps are as follows: calculate the product of the chemical composition feature data and the first vector to obtain the first product; calculate the sum of the first product and the first feature bias data to obtain the first sub-feature training data, where the first vector and the first feature bias data are determined during the training process of the embedding representation model.

[0110] In some embodiments, for performing second embedding representation processing on the category feature data, the specific steps are as follows: according to the category feature data, determine the one-hot representation data corresponding to the category feature data; calculate the product of the one-hot representation data and the second vector to obtain the second product; calculate the sum of the second product and the second feature bias data to obtain the second sub-feature training data, where the second vector and the second feature bias data are determined during the training process of the embedding representation model.

[0111] Figure 5 A schematic diagram showing still some other embodiments of the classification method of the present disclosure.

[0112] As Figure 5 shown, the classification method includes steps 510 to 550.

[0113] In step 510, obtain the original feature data and original feature training data of the initially baked tobacco leaves, and determine the grade annotation information of the original feature training data.

[0114] In step 520, perform normalization processing on the original feature data and the original feature training data to obtain the first feature data and the first feature training data.

[0115] In step 530, perform embedding representation processing on the first feature data and the first feature training data to obtain the second feature data and the second feature training data.

[0116] In step 540, train the classification model according to the second feature training data to obtain the trained classification model.

[0117] For example, classification labels are added to the second feature training data to obtain third feature training data, and the classification model is trained using the third feature training data to obtain the trained classification model.

[0118] In step 550, the grade of the flue-cured tobacco leaves is determined using the second feature data and the trained classification model.

[0119] For example, classification labels are added to the second feature data to obtain third feature data, and then the third feature data is input into the trained classification model to determine the grade of the flue-cured tobacco leaves.

[0120] In the above embodiments, first feature data including chemical composition feature data and category feature data of flue-cured tobacco leaves is obtained, embedding representation processing and adding classification labels are performed on the first feature data to obtain third feature data, and finally the third feature data is input into the classification model to obtain the grade of the flue-cured tobacco leaves, ensuring the efficiency and reliability of the process of classifying flue-cured tobacco leaves. When obtaining the feature data of flue-cured tobacco leaves, the chemical composition feature data and category feature data of flue-cured tobacco leaves are considered, and relatively complete feature data is obtained, improving the accuracy and reliability of classifying the grade of flue-cured tobacco leaves through the classification model. In addition, by classifying the grade of flue-cured tobacco leaves through the classification model, useful feature information can be extracted from a large amount of third feature data to realize automatic classification of tobacco leaf grades, which is faster than manual classification, improving the efficiency of the process of classifying flue-cured tobacco leaves, and also reducing the risk of errors caused by manual classification of the grade of flue-cured tobacco leaves. Furthermore, through the embedding representation processing of the first feature data, a data format beneficial to model processing can be obtained, improving the efficiency of classifying the grade of flue-cured tobacco leaves, and by adding classification labels to the second feature data, the global information in the second feature data can be integrated, improving the accuracy of classifying the grade of flue-cured tobacco leaves.

[0121] Figure 6 Schematic diagrams showing some embodiments of the classification device of the present disclosure.

[0122] As Figure 6 shown, the classification device 60 includes a first acquisition unit 61, a first embedding representation unit 62, a first addition unit 63, and an input unit 64.

[0123] The first acquisition unit 61 is configured to acquire first feature data of flue-cured tobacco leaves, where the first feature data includes chemical composition feature data and category feature data of flue-cured tobacco leaves.

[0124] For example, the category feature data of flue-cured tobacco leaves may be the production area and tobacco leaf variety of flue-cured tobacco leaves, etc.

[0125] In some embodiments, the first acquisition unit 61 is further configured to normalize the original feature data of the initially baked tobacco leaves to obtain first feature data.

[0126] In some embodiments, the first acquisition unit 61 is further configured to calculate the median and interquartile range of the original feature data of the initially baked tobacco leaves; and normalize the original feature data of the initially baked tobacco leaves according to the original feature data, median and interquartile range of the initially baked tobacco leaves to obtain first feature data.

[0127] The first embedding representation unit 62 is configured to perform an embedding representation process on the first feature data to obtain second feature data.

[0128] In some embodiments, the first embedding representation unit 62 is further configured to perform a first embedding representation process on the chemical composition feature data to obtain first sub-feature data; perform a second embedding representation process on the category feature data to obtain second sub-feature data; and obtain second feature data according to the first sub-feature data and the second sub-feature data.

[0129] In some embodiments, the first embedding representation unit 62 is further configured to calculate the product of the chemical composition feature data and the first vector to obtain a first product; and calculate the sum of the first product and the first feature bias data to obtain first sub-feature data, where the first vector and the first feature bias data are determined during the training process of the embedding representation model.

[0130] In some embodiments, the first embedding representation unit 62 is further configured to determine one-hot representation data corresponding to the category feature data according to the category feature data; calculate the product of the one-hot representation data and the second vector to obtain a second product; and calculate the sum of the second product and the second feature bias data to obtain second sub-feature data, where the second vector and the second feature bias data are determined during the training process of the embedding representation model.

[0131] The first addition unit 63 is configured to add a classification marker to the second feature data to obtain third feature data.

[0132] The input unit 64 is configured to input the third feature data into a classification model to obtain the grade of the initially baked tobacco leaves.

[0133] For example, the classification model includes one or more layers of Transformer encoders and a linear neural network classifier, where, in the case where the classification model includes multiple layers of Transformer encoders, the multiple layers of Transformer encoders are connected in series.

[0134] In the above embodiments, the first feature data including the chemical composition feature data and the category feature data of the primary-cured tobacco leaves is obtained, the first feature data is subjected to an embedding representation process and a classification label is added to obtain the third feature data, and finally the third feature data is input into the classification model to obtain the grade of the primary-cured tobacco leaves, ensuring the efficiency and reliability of the process of classifying the primary-cured tobacco leaves. When obtaining the feature data of the primary-cured tobacco leaves, the chemical composition feature data and the category feature data of the primary-cured tobacco leaves are considered, and relatively complete feature data is obtained, improving the accuracy and reliability of classifying the grade of the primary-cured tobacco leaves through the classification model. In addition, by classifying the grade of the primary-cured tobacco leaves through the classification model, useful feature information can be extracted from a large amount of third feature data to realize the automatic classification of the tobacco leaf grade, which is faster than manual classification, improving the efficiency of the process of classifying the primary-cured tobacco leaves, and also reducing the risk of errors caused by manual classification of the grade of the primary-cured tobacco leaves. Furthermore, through the embedding representation process of the first feature data, a data format beneficial to model processing can be obtained, improving the efficiency of classifying the grade of the primary-cured tobacco leaves. By adding a classification label to the second feature data, the global information in the second feature data can be integrated, improving the accuracy of classifying the grade of the primary-cured tobacco leaves.

[0135] Figure 7 Schematic diagrams showing some embodiments of the training device of the classification model of the present disclosure.

[0136] As Figure 7 shown, the training device 70 of the classification model includes a second acquisition unit 71, a determination unit 72, a second embedding representation unit 73, a second addition unit 74, and a training unit 75.

[0137] The second acquisition unit 71 is configured to acquire the first feature training data of the primary-cured tobacco leaves, where the first feature training data includes the chemical composition feature data and the category feature data of the primary-cured tobacco leaves.

[0138] For example, the category feature data of the primary-cured tobacco leaves may be the production area and tobacco leaf variety of the primary-cured tobacco leaves, etc.

[0139] In some embodiments, the second acquisition unit 71 is further configured to perform normalization processing on the original feature training data of the primary-cured tobacco leaves to obtain the first feature training data.

[0140] In some embodiments, the second acquisition unit 71 is further configured to calculate the median and interquartile range of the original feature training data of the primary-cured tobacco leaves; according to the original feature training data, median and interquartile range of the primary-cured tobacco leaves, perform normalization processing on the original feature training data of the primary-cured tobacco leaves to obtain the first feature training data.

[0141] A determination unit 72, configured to determine the grade annotation information of the first feature training data according to the first feature training data.

[0142] A second embedding representation unit 73, configured to perform an embedding representation process on the first feature training data to obtain second feature training data.

[0143] In some embodiments, the second embedding representation unit 73 is further configured to perform a first embedding representation process on the chemical composition feature data to obtain first sub-feature training data; perform a second embedding representation process on the category feature data to obtain second sub-feature training data; and obtain the second feature training data according to the first sub-feature training data and the second sub-feature training data.

[0144] In some embodiments, the second embedding representation unit 73 is further configured to calculate the product of the chemical composition feature data and the first vector to obtain a first product; calculate the sum of the first product and the first feature bias data to obtain the first sub-feature training data, where the first vector and the first feature bias data are determined during the training process of the embedding representation model.

[0145] In some embodiments, the second embedding representation unit 73 is further configured to determine the one-hot representation data corresponding to the category feature data according to the category feature data; calculate the product of the one-hot representation data and the second vector to obtain a second product; calculate the sum of the second product and the second feature bias data to obtain the second sub-feature training data, where the second vector and the second feature bias data are determined during the training process of the embedding representation model.

[0146] A second addition unit 74, configured to add a classification marker to the second feature training data to obtain third feature training data.

[0147] A training unit 75, configured to train the classification model according to the third feature training data and the grade annotation information.

[0148] In the above embodiments, the chemical composition feature data and the category feature data of the primary-cured tobacco leaves are considered in the feature training data for training the classification model, that is, the feature data in two aspects are considered, which improves the accuracy of the classification model for classifying the grades of primary-cured tobacco leaves. By performing an embedding representation process on the first feature training data, a data format beneficial to the processing of the classification model can be obtained, which improves the efficiency of classifying the grades of primary-cured tobacco leaves. By adding a classification marker to the second feature training data, the global information in the second feature training data can be integrated, which improves the accuracy of the classification model for classifying the grades of primary-cured tobacco leaves.

[0149] Figure 8 A schematic diagram showing some embodiments of the electronic device of the present disclosure.

[0150] As Figure 8 shown, the electronic device 80 of this embodiment includes: a memory 81 and a processor 82 coupled to the memory 81. The processor 82 is configured to execute the classification method in any of the foregoing embodiments or the training method of the classification model in any of the foregoing embodiments based on the instructions stored in the memory 81.

[0151] The memory 81 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.

[0152] The electronic device 80 may further include an input / output interface 83, a network interface 84, a storage interface 85, etc. These interfaces 83, 84, 85 and the memory 81 and the processor 82 may be connected through a bus 86, for example. Among them, the input / output interface 83 provides a connection interface for input / output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, and a speaker. The network interface 84 provides a connection interface for various networking devices. The storage interface 85 provides a connection interface for external storage devices such as an SD card and a USB flash drive.

[0153] In the above embodiment, the first feature data including the chemical composition feature data and the category feature data of the primary-cured tobacco leaves is obtained, the first feature data is subjected to an embedding representation process and a classification label is added to obtain the third feature data, and finally the third feature data is input into the classification model to obtain the grade of the primary-cured tobacco leaves, which ensures the efficiency and reliability of the process of classifying the primary-cured tobacco leaves. When obtaining the feature data of the primary-cured tobacco leaves, the chemical composition feature data and the category feature data of the primary-cured tobacco leaves are considered, and relatively complete feature data is obtained, improving the accuracy and reliability of classifying the grade of the primary-cured tobacco leaves through the classification model. In addition, by classifying the grade of the primary-cured tobacco leaves through the classification model, useful feature information can be extracted from a large amount of third feature data to realize the automatic classification of the tobacco leaf grade, which is faster than manual classification, improving the efficiency of the process of classifying the primary-cured tobacco leaves, and also reducing the risk of errors caused by manual classification of the grade of the primary-cured tobacco leaves. In addition, through the embedding representation process of the first feature data, a data format beneficial to model processing can be obtained, improving the efficiency of classifying the grade of the primary-cured tobacco leaves. By adding a classification label to the second feature data, the global information in the second feature data can be integrated, improving the accuracy of classifying the grade of the primary-cured tobacco leaves.

[0154] In some embodiments, a computer program product is protected, including a computer program or instructions, which when executed by a processor implement the above classification method or the training method of the classification model. The computer program product includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a classification device or a training device of the classification method, or installed from a storage device, or installed from a ROM. When the computer program is executed by a CPU, the above functions defined in the method of the embodiments of the present disclosure are executed.

[0155] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0156] So far, the classification method, device, training method of the classification model, device, and electronic device of the present disclosure have been described in detail. In order to avoid obscuring the concept of the present disclosure, some details well known in the art are not described. Those skilled in the art can fully understand how to implement the technical solutions disclosed here based on the above description.

[0157] The methods and systems of the present disclosure can be implemented in many ways. For example, the methods and systems of the present disclosure can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps of the method is only for illustration, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.

[0158] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A classification method comprising: Acquiring first characteristic data of freshly-cured tobacco leaves, wherein the first characteristic data includes chemical component characteristic data and category characteristic data of the freshly-cured tobacco leaves; Performing embedding representation processing on the first feature data to obtain second feature data; adding a classification mark to the second characteristic data to obtain third characteristic data; The third characteristic data is input into a classification model to obtain the grade of the first-cured tobacco leaves.

2. The classification method according to claim 1, wherein: The embedding representation processing is performed on the first feature data to obtain the second feature data, comprising: Performing a first embedding representation process on the chemical composition characteristic data to obtain first sub-characteristic data; Performing a second embedding representation process on the category feature data to obtain second sub-feature data; The second feature data is obtained according to the first sub-feature data and the second sub-feature data.

3. The classification method according to claim 2, wherein: The performing a first embedding representation process on the chemical composition characteristic data to obtain the first sub-characteristic data comprises: Calculating the product of the chemical composition characteristic data and the first vector to obtain a first product; The first product and the first feature bias data are added to obtain the first sub-feature data, wherein the first vector and the first feature bias data are determined by the embedding representation model during the training process.

4. The classification method according to claim 2, wherein: The performing a second embedding representation process on the category feature data to obtain second sub-feature data comprises: Determine, according to the category feature data, one-hot representation data corresponding to the category feature data; Calculate the product of the one-hot representation data and the second vector to obtain a second product; The second product and the second feature bias data are added to obtain the second sub-feature data, wherein the second vector and the second feature bias data are determined by the embedding representation model during the training process.

5. The classification method according to any one of claims 1 to 4, wherein: The first characteristic data of the first-cured tobacco leaves are obtained as follows: The original characteristic data of the first-cured tobacco leaves are normalized to obtain the first characteristic data.

6. The classification method according to claim 5, wherein: The normalization process is performed on the original characteristic data of the first-cured tobacco leaves to obtain the first characteristic data, which includes: Calculating the median and interquartile range of the original characteristic data of the freshly cured tobacco leaves; The original characteristic data of the freshly-cured tobacco leaves are normalized according to the original characteristic data of the freshly-cured tobacco leaves, the median and the interquartile range to obtain the first characteristic data.

7. The classification method according to any one of claims 1 to 4, wherein: The category characteristic data includes at least one of the production area and tobacco variety of the first-cured tobacco leaves.

8. The classification method according to any one of claims 1 to 4, wherein: The classification model includes one or more layers of Transformer encoders and a linear neural network classifier, wherein, when the classification model includes the multi-layer Transformer encoder, the multi-layer Transformer encoder is connected in series.

9. A method for training a classification model, comprising: Acquire first feature training data of freshly-cured tobacco leaves, wherein the first feature training data includes chemical component feature data and category feature data of the freshly-cured tobacco leaves; Determining level labeling information of the first feature training data according to the first feature training data; Performing embedding representation processing on the first feature training data to obtain second feature training data; Adding a classification mark to the second feature training data to obtain third feature training data; The classification model is trained according to the third feature training data and the grade labeling information.

10. A classification device comprising: A first acquisition unit is configured to acquire first characteristic data of the freshly cured tobacco leaves, wherein the first characteristic data includes chemical component characteristic data and category characteristic data of the freshly cured tobacco leaves; A first embedding representation unit is configured to perform embedding representation processing on the first feature data to obtain second feature data; A first adding unit is configured to add a classification mark to the second feature data to obtain third feature data; The input unit is configured to input the third feature data into the classification model to obtain the grade of the first-cured tobacco leaves.

11. A training device for a classification model, comprising: A second acquisition unit is configured to acquire first feature training data of freshly-cured tobacco leaves, wherein the first feature training data includes chemical component feature data and category feature data of the freshly-cured tobacco leaves; a determining unit, configured to determine level labeling information of the first feature training data according to the first feature training data; A second embedding representation unit is configured to perform embedding representation processing on the first feature training data to obtain second feature training data; A second adding unit is configured to add a classification tag to the second feature training data to obtain third feature training data; The training unit is configured to train the classification model according to the third feature training data and the grade labeling information.

12. An electronic device, comprising: Memory; and A processor coupled to the memory, the processor being configured to execute the classification method according to any one of claims 1 to 8 or the classification model training method according to claim 9 based on instructions stored in the memory.

13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the classification method described in any one of claims 1 to 8 or the classification model training method described in claim 9.

14. A computer program product, comprising computer instructions, which, when executed by a processor, implement the classification method described in any one of claims 1 to 8 or the classification model training method described in claim 9.