Tobacco leaf quality digital evaluation method and system based on multi-source data fusion

CN120688905APending Publication Date: 2025-09-23CHINA TOBACCO GUANGXI IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510502882.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing tobacco leaf quality assessment methods rely on single-dimensional data analysis and lack the interactive integration of multi-source data. As a result, the evaluation indicators are unable to cover the multi-dimensional linkage between the plant's physiological state, chemical composition and final sensory evaluation. Data traceability is difficult and the scoring labels are incomplete, which affects the stability of model training and the accuracy of output results. It is difficult to meet the needs of large-scale planting management for insights into regional quality distribution.

Method used

By collecting and binding the agronomic indicators of tobacco plants with the reflectance values ​​of the hyperspectral instrument, combined with the near-infrared spectral data of the flue-cured tobacco leaves, agronomic morphological spectral structure set and near-infrared chemical fusion feature group are generated. The intersection is identified and the null values ​​are eliminated. The label binding relationship is established, and a label alignment data set is generated. Finally, the field-level quality grade distribution analysis is carried out.

Benefits of technology

It realizes the multi-dimensional data integration of tobacco leaf quality assessment, improves the traceability and integrity of data, improves the accuracy and consistency of assessment results, and supports accurate judgment of inter-field management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688905A_ABST
    Figure CN120688905A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of quality evaluation, in particular to a tobacco leaf quality digital evaluation method and system based on multi-source data fusion, and the method comprises the following steps: collecting tobacco plant agronomic indexes and spectral data, numbering and binding, obtaining a baked sample and detecting chemical components, splicing agronomic and chemical parameters to establish a fusion vector, and evaluating the quality of tobacco leaves. And performing label verification and correction in combination with sensory scores, generating sample grades in a classified manner, and summarizing and outputting grade distribution information according to field parcels. According to the method, agricultural indexes such as the leaf area, the SPAD value, the leaf size and the plant height are bound with hyperspectral band reflectivity data, score missing and abnormal items are eliminated, sensory labels have consistency and representativeness, and the sensory labels can be obtained based on parameter input of a specific band reflection value, a carbon-nitrogen ratio and a full-spectrum intensity interval in combination with aroma quality and coordination scoring. Layered label classification is carried out and spatial distribution is summarized, so that the availability and parameter consistency of samples are improved, and a quality evaluation result has a spatial visual attribute.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of quality assessment, and in particular to a digital tobacco leaf quality assessment method and system based on multi-source data fusion. Background Art

[0002] The field of quality assessment technology primarily focuses on the quantitative or qualitative evaluation of the physical properties, chemical composition, and functional performance of products, raw materials, or process outputs. The goal is to quickly and objectively determine the quality without destroying the sample structure. This field widely utilizes mathematical modeling, pattern recognition, statistical analysis, image processing, and machine learning techniques, combined with spectral analysis, imaging, and sensing technologies to obtain multi-dimensional sample information and establish a quantifiable evaluation index system. By constructing standardized quality evaluation models, the evaluation process is automated, standardized, and digitized, improving the accuracy and consistency of evaluations. This field is widely used in a variety of industries, including agriculture, food, medicine, and industrial manufacturing.

[0003] The digital tobacco leaf quality assessment method aims to digitally and quantitatively evaluate tobacco leaf appearance, internal composition, and overall grade using image processing, spectral recognition, and multivariate statistical modeling. This method, through standardized data collection and modeling, implements tobacco leaf quality scoring. It can be widely applied in tobacco procurement, grading, processing, and quality control processes, achieving efficient, objective, and consistent tobacco leaf quality assessment and improving overall quality management across the tobacco industry chain.

[0004] Traditional evaluation methods rely on the analysis of a single data dimension, often using image processing or single-line processing of spectral data. These methods lack mechanisms for interactive fusion of multi-source data, resulting in evaluation metrics that fail to capture the multi-dimensional linkages between plant physiological state, chemical composition, and final sensory evaluation. In terms of sample management, a lack of clear plant numbering links makes data traceability difficult and prone to sample mismatches and label bias. In the construction of scoring labels, there is a lack of systematic handling of missing and outlier sensory scores. Label data is incomplete and unbalanced, impacting model training stability and output accuracy. The ability to analyze macro-quality distribution at the field level is also relatively weak. Existing models often score at the individual sample level, failing to effectively map to the field spatial scale and failing to meet the needs for regional quality distribution insights required for large-scale crop management. For example, during the tobacco procurement phase, the lack of accurate field-level quality information can easily lead to irrational or misjudgment of grading, resulting in errors in procurement prices and reduced adaptability to subsequent processing. These issues have limited the refinement and standardization of existing evaluation technologies in practical applications. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a digital tobacco quality evaluation method and system based on multi-source data fusion.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a digital tobacco leaf quality assessment method based on multi-source data fusion, comprising the following steps:

[0007] S1: Based on the numbered sample plants in the tobacco field, collect the plant's agronomic indicators and use a hyperspectral instrument to obtain the corresponding leaf band reflectance values. Based on the sample number and the regional division number, the agronomic indicators and the spectral band reflectance are combined to generate an agronomic morphological spectral structure set.

[0008] S2: Based on the agronomic morphology spectral structure set, obtain flue-cured tobacco leaves with consistent tobacco plant numbers, detect key chemical component indicators, map and combine near-infrared spectral curves with chemical component indicators, and mark the source time and field information to obtain a near-infrared chemical fusion feature group;

[0009] S3: Based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, identifying tobacco plant numbers with intersections, performing null value filling and inconsistent item removal on multiple vector fields, and performing sample availability screening on the merged feature vector set to obtain a fusion indicator vector set;

[0010] S4: Based on the fusion indicator vector set, a label binding relationship is established between samples and sensory scores under the same flue-cured tobacco batch, samples with missing scores are eliminated, and outliers in the score range are corrected to generate a usable data structure and obtain a label alignment dataset.

[0011] As a further solution of the present invention, the agronomic morphology spectrum structure set includes leaf growth trend identification, reflection band energy combination parameters, plant posture normalization variables, SPAD value spectral response group and morphological feature spectrum mapping index, the near-infrared chemical fusion feature group includes chemical component spectrum correspondence coefficient, near-infrared band principal component factor, sample identification structure under regional number, content intensity curve structure and chemical indicator classification label, the fusion indicator vector set includes the primary key index table after feature splicing, field standardization result matrix, null value completion record set, feature field consistency label set and sample screening number list, the label alignment data set includes aroma dimension quality mapping table, coordination dimension scoring structure, sensory label and sample primary key mapping comparison table, abnormal label screening mark set and scoring structure validity confirmation record.

[0012] As a further solution of the present invention, the steps for obtaining the agronomic morphology spectral structure set are specifically as follows:

[0013] S101: Based on the numbered sample plants in the tobacco field, the leaf area, SPAD value, leaf length, leaf width, and plant height of the plants are collected, and the field index number of each indicator is bound to the tobacco plant number to construct an agronomic parameter information set for each tobacco plant under the number, establish a mapping value between the corresponding number and the agronomic variable, and generate agronomic variable mapping information;

[0014] S102: Based on the agronomic variable mapping information, the band reflectance data recorded by the hyperspectral data acquisition instrument corresponding to the tobacco plant number is called, and based on the docking position of the band response value between the time axis and the number index, the spectral band data structure corresponding to each tobacco plant number is constructed, and the band position sequence is unified to generate a reflectance band comparison set;

[0015] S103: Call the reflectance band control set and agronomic variable mapping information, perform position matching in the variable dimension based on the agronomic parameters with the same number and the corresponding band reflectance values, combine the agronomic indicators and the connected multi-segment spectral values ​​into a set of vector structures, and establish a combined parameter set within the partition with the tobacco plant number as the index unit to generate an agronomic morphological spectral structure set.

[0016] As a further solution of the present invention, the steps for obtaining the near-infrared chemical fusion feature group are specifically as follows:

[0017] S201: Based on the agronomic morphology spectral structure set, flue-cured tobacco leaves with consistent tobacco plant numbers are obtained, near-infrared reflectance spectrum data corresponding to each sample is collected, near-infrared reflectance of each tobacco leaf sample is divided into bands according to the number sequence, and main reflectance band values ​​within the continuous interval are extracted to generate a near-infrared main band value group;

[0018] S202: Based on the near-infrared main band value group, the moisture content, total sugar, reducing sugar, total nitrogen, nicotine, potassium, and protein mass percentage of the corresponding tobacco leaf sample are detected, the binding relationship between the post-curing sample number and the tobacco plant number at the time of collection is called, an index corresponding to the sample numbers is established, and the chemical composition data is added to the original band value record structure as a variable field to generate a component reflection correlation information array;

[0019] S203: Call the component reflection association information, mark the acquisition time and field number information of each data record, aggregate the field division number and sample reflection data structure, and place the sample records in the same area into independent data segments. Combine the reflection spectrum segments and the aggregated data structure of the component indicators to establish a near-infrared chemical fusion feature group.

[0020] As a further solution of the present invention, the steps of obtaining the fusion index vector set are specifically as follows:

[0021] S301: Based on the agronomic morphological spectral structure set and the near-infrared chemical fusion feature group, identifying record entries with the same tobacco plant number, performing field splicing on agronomic parameters and chemical indicators according to the number consistency, constructing a parameter set containing multiple indicator dimensions, and generating a sample parameter splicing value;

[0022] S302: calling the sample parameter splicing value, unifying the unit system of agronomic variables and chemical indicators based on the primary key number of each record, performing interval normalization processing on the reflectance, sugar concentration, and element content vector fields, calculating the sample availability ratio between all fields, and determining whether there are vacancies and scale inconsistencies between fields, calculating the sample field consistency measure, eliminating the field set whose consistency measure is lower than the availability threshold, and generating field consistency screening information;

[0023] S303: Filter information based on the field consistency, extract the remaining valid field groups, construct a primary key vector structure of the retained fields under each sample, encapsulate the vectors according to the tobacco plant numbers and store them as partitioned data units, and establish a fusion indicator vector set.

[0024] As a further solution of the present invention, the formula for calculating the sample field consistency metric is:

[0025]

[0026] Among them, U i represents the field consistency measure of the i-th sample, A ij represents the normalized agronomic parameter effectiveness index value of the i-th sample on the j-th field, B ij Indicates the normalized chemical composition of the i-th sample in the j-th field, C ij Indicates the unit standard error value of the i-th sample on the j-th field, D ij It represents the measurement interval error value of the i-th sample on the j-th field, i represents the index number of the sample, j represents the index number of the field, and m represents the total number of fields in the sample parameter set.

[0027] As a further solution of the present invention, the steps of obtaining the label alignment dataset are specifically as follows:

[0028] S401: Based on the fusion index vector set, sensory score data of the flue-cured tobacco leaf sample matching each sample number is obtained, and the aroma quality, aroma quantity, pungency, aftertaste, and coordination score values ​​are extracted. The score data structure is field-bound with the original sample number, and a mapping structure between the sample primary key and multiple scores is constructed to generate sensory score binding information;

[0029] S402: Call the sensory score binding information, filter the numbered records with missing score data, index and remove the missing samples, and perform a difference ratio determination on the maximum and minimum values ​​of the retained sample score records, correct the score records that exceed the upper limit of the score setting range and are lower than the lower limit, and generate a valid sample score interval value;

[0030] S403: According to the valid sample score interval value, the score field is filled into the corresponding sample primary key field in the fusion indicator vector set, the field structure is adjusted in position and the naming is unified, a label structured data form is constructed, and a label alignment data set is established.

[0031] As a further embodiment of the present invention, the method further comprises:

[0032] S5: calling the label alignment dataset, selecting the main reflectance band value, carbon-nitrogen related chemical ratio value and near-infrared full spectrum intensity interval value reflecting the spectral structure characteristics of the plant as key input items, combining the coordination and aroma quality scores in the sensory label, constructing a scoring grade identifier by grade interval, performing a corresponding classification operation on each sample feature and the scoring grade, identifying the label grade range to which the sample belongs, and summarizing the spatial distribution with samples from the same field to generate field-level tobacco leaf quality grade distribution information;

[0033] The field-level tobacco leaf quality grade distribution information specifically includes a grade spatial distribution layer, a sensory label density block map, a set of grade segment boundaries between fields, a fusion index and label grade comparison table, and a hot zone archive number index.

[0034] As a further solution of the present invention, the steps for obtaining the field-level tobacco leaf quality grade distribution information are specifically as follows:

[0035] S501: Based on the label alignment dataset, extract the main reflection band value, carbon-nitrogen ratio value, and near-infrared full spectrum intensity interval value from the feature vector of each sample, combine the harmony and aroma quality score items in the sensory score, normalize the input data by field, construct an input sample vector group, and generate a structure score input value;

[0036] S502: Calling the structure score input value, performing corresponding level matching on the input sample vectors according to the score interval set by the coordination and fragrance quality score, calculating the numerical deviation of the attribution level of each vector and comparing it with the label interval spacing, calculating the level fit of each sample in the label space, matching and classifying according to the level fit and the preset interval, and establishing the sample score level value;

[0037] The formula for calculating the rank fit of each sample in the label space is:

[0038]

[0039] Among them, R k Indicates the fitting degree of the score of the kth sample, E kl represents the normalized value of the lth input field of the kth sample, F kl represents the central score value of the lth item in the target label level interval of the kth sample, G k represents the standard deviation of the kth sample input field, H k Indicates the standard deviation of the score label interval, M k It represents the ratio of the kth sample label score to the corresponding input field, where k is the sample index, l is the field dimension index, and p is the total number of input fields;

[0040] S503: Based on the sample scoring grade value, the spatial position of the sample grade label is paired according to the field number to which each sample belongs, and the number of multiple grade samples in the same area is counted and accumulated, and a spatial distribution block structure is generated according to the grade group to establish the field-level tobacco leaf quality grade distribution information.

[0041] A digital tobacco leaf quality assessment system based on multi-source data fusion, the digital tobacco leaf quality assessment system based on multi-source data fusion is used to implement the digital tobacco leaf quality assessment method based on multi-source data fusion, the system comprising:

[0042] The agronomic morphology extraction module collects agronomic indicators of the numbered sample plants in the tobacco field and uses a hyperspectral instrument to obtain the corresponding leaf band reflectance values. Based on the sample number and the regional division number, the agronomic indicators and the spectral band reflectance are combined to generate an agronomic morphology spectral structure set.

[0043] The chemical feature fusion module obtains flue-cured tobacco leaves with consistent tobacco plant numbers based on the agronomic morphology spectral structure set, detects key chemical component indicators, maps and combines near-infrared spectral curves with chemical component indicators, and marks the source time and field information to obtain a near-infrared chemical fusion feature group;

[0044] The indicator vector screening module identifies tobacco plant numbers with intersections based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, performs null value filling and inconsistent item removal on multiple vector fields, and performs sample availability screening on the merged feature vector set to obtain a fusion indicator vector set;

[0045] The label alignment processing module establishes a label binding relationship between samples and sensory scores under the same flue-cured tobacco batch based on the fusion index vector set, removes samples with missing scores, corrects outliers in the score range, generates a usable data structure, and obtains a label alignment dataset;

[0046] The quality grade distribution module calls the label alignment dataset, selects the main reflection band value, carbon-nitrogen related chemical ratio value and near-infrared full spectrum intensity interval value reflecting the spectral structure characteristics of the plant as key input items, combines the coordination and aroma quality scores in the sensory label, constructs the score grade identification by grade interval, identifies the label grade range to which the sample belongs, and generates field-level tobacco leaf quality grade distribution information.

[0047] Compared with the prior art, the advantages and positive effects of the present invention are:

[0048] In the present invention, by accurately sampling the plants with numbers in the field stage and binding agronomic indicators such as leaf area, SPAD value, leaf size and plant height with hyperspectral band reflectance data, the physiological state and spectral characteristics of each plant are clarified, and an agronomic morphological spectral structure set with traceability is established, which enhances the verifiability of the data source. By performing near-infrared spectroscopy detection on tobacco leaves after baking from plants with the same number and synchronously collecting chemical composition data, the physical and chemical properties are organically integrated to generate a clearly labeled fusion feature parameter set. In the process of establishing sample labels, by eliminating missing scores and abnormal items, The sensory labels are made consistent and representative, and a label-aligned data set is formed. Based on the parameter input of specific band reflectance values, carbon-nitrogen ratio values ​​and full-spectrum intensity ranges, combined with the aroma quality and coordination scores, hierarchical label classification is performed and the spatial distribution is summarized, which realizes the quality grade coverage analysis at the field level and improves the data integrity and comparability. Through the cross-fusion and unified standardization of data from multiple sources, the availability and parameter consistency of samples are improved. The correspondence between label binding and spatial distribution is constructed in the classification output, so that the quality assessment results have spatial visual attributes, which is conducive to the accurate judgment of management differences between fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0050] Figure 2 This is a detailed flow chart of S1 of the present invention;

[0051] Figure 3 This is a detailed flow chart of S2 of the present invention;

[0052] Figure 4 This is a detailed flow chart of S3 of the present invention;

[0053] Figure 5 This is a detailed flow chart of S4 of the present invention;

[0054] Figure 6 This is a detailed flow chart of S5 of the present invention;

[0055] Figure 7 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0057] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0058] See also Figure 1 The present invention provides a technical solution: a digital evaluation method for tobacco leaf quality based on multi-source data fusion, comprising the following steps:

[0059] S1: Based on the numbered sample plants in the tobacco field, agronomic indicators of the plants are collected, including leaf area, SPAD value, leaf size and plant height. The corresponding leaf band reflectance values ​​are obtained using a hyperspectral instrument. Based on the sample number and the regional division number, the plant number partition is called, and the agronomic indicators and spectral band reflectance are parameterized to generate an agronomic morphological spectral structure set.

[0060] S2: Based on the agronomic morphology spectral structure set, obtain flue-cured tobacco leaves with consistent tobacco plant numbers, collect near-infrared reflectance data, and detect key chemical composition indicators, including moisture content, total sugar, reducing sugar, total nitrogen, nicotine, potassium, and protein. Based on the binding relationship between the flue-cured number and the field collection number, map and combine the near-infrared spectral curves with the chemical composition indicators, and mark the source time and field information to obtain the near-infrared chemical fusion feature group;

[0061] S3: Based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, identify the tobacco plant numbers with intersections, horizontally splice the parameter sets, establish a parameter vector set with the sample number as the primary key, unify the units and scales, implement null value filling and inconsistent item removal for multiple vector fields, and screen the sample availability of the merged feature vector set to obtain the fusion indicator vector set;

[0062] S4: Based on the fusion indicator vector set, the sensory scores of the corresponding flue-cured tobacco leaves are obtained. According to the aroma quality, aroma volume, pungency, aftertaste and harmony scores, a label binding relationship is established between the samples and the sensory scores of the same flue-cured tobacco batch. The label scores are verified to ensure that they cover the entire sample set. By removing samples with missing scores and correcting outliers in the score range, a usable data structure is generated to obtain a label-aligned dataset.

[0063] S5: Call the label alignment dataset, and based on each fused feature vector, select the main reflectance band value, carbon-nitrogen related chemical ratio value and near-infrared full spectrum intensity interval value that reflect the spectral structure characteristics of the plant as key input items, combine the coordination and aroma quality scores in the sensory label, and construct the score grade identification by level interval. Perform corresponding classification operations on each sample feature and the score grade, identify the label grade range to which the sample belongs, and make a spatial distribution summary with the samples from the same field to generate field-level tobacco leaf quality grade distribution information.

[0064] The agronomic morphological spectral structure set includes leaf growth trend identification, reflectance band energy combination parameters, plant posture normalization variables, SPAD value spectral response group and morphological feature spectral mapping index. The near-infrared chemical fusion feature group includes the correspondence coefficient between chemical component spectra, near-infrared band principal component factors, sample identification structure under regional number, content intensity curve structure and chemical indicator classification label. The fusion indicator vector set includes the primary key index table after feature splicing, field standardization result matrix, null value completion record set, feature field consistency label set and sample screening number list. The label alignment dataset includes the aroma dimension quality mapping table, the coordination dimension scoring structure, the sensory label and sample primary key mapping comparison table, the abnormal label screening mark set and the scoring structure validity confirmation record. The field-level tobacco leaf quality grade distribution information specifically includes the grade spatial distribution layer, the sensory label density block map, the grade segment boundary set between fields, the fusion indicator and label grade comparison table and the hot zone archive number index.

[0065] See also Figure 2 ,The specific steps for obtaining the agronomic morphology spectrum structure set are:

[0066] S101: Based on the numbered sample plants in the tobacco field, the leaf area, SPAD value, leaf length, leaf width, and plant height of the plants are collected, and the field index number of each indicator is bound to the tobacco plant number to construct an agronomic parameter information set for each tobacco plant under the number, establish a mapping value between the corresponding number and the agronomic variable, and generate agronomic variable mapping information;

[0067] According to the numbered sample plants in the tobacco field, each plant number is read in turn, and the individual tobacco plants are indexed by the plant number. The agronomic parameter data fields corresponding to the plant number are read one by one. The leaf area is obtained by extracting the leaf edge contour after image segmentation and converting it based on the pixel ratio. The SPAD value is measured at the center of the middle leaf expansion surface by a handheld SPAD meter and averaged. The leaf length and leaf width are respectively obtained by image processing to identify the main vein length and the maximum vertical width of the leaf. The plant height is measured and recorded with a ruler from the ground starting point to the end of the longest leaf at the top. After the collection is completed, the above leaf area, SPAD value, leaf length, leaf width and plant height are mapped in sequence into the parameter field table with the number as the primary key, and the field insertion operation is performed according to the number. The plant agronomic information records with the number as the primary index are established one by one. By constructing a vector structure of a five-dimensional agronomic variable set, a one-way mapping relationship from the number to the parameter group is formed, and then the complete number-agronomic variable index mapping information is generated. For example, a tobacco plant numbered T012 has a leaf area of ​​188.6cm after image segmentation. 2 The SPAD value measured by the SPAD instrument is 38.4. The leaf length and leaf width measured by image recognition are 25.1 cm and 13.7 cm respectively. The plant height measured by the manual ruler is 92.5 cm. The parameter group is combined into (188.6, 38.4, 25.1, 13.7, 92.5) and assigned to the T012 entry to finally obtain the agronomic variable mapping information.

[0068] S102: Based on the agronomic variable mapping information, the band reflectance data recorded by the hyperspectral data acquisition instrument corresponding to the tobacco plant number is called. Based on the docking position of the band response value between the time axis and the number index, the spectral band data structure corresponding to each tobacco plant number is constructed, and the band position sequence is unified to generate a reflectance band comparison set.

[0069] According to the mapping information of agronomic variables, the number set of each tobacco plant in the agronomic variable is first extracted, the number list is traversed one by one, and the timestamp entry associated with the same number is found in the hyperspectral original record data set. The band reflectance data of the corresponding time node is read, and the target data group is obtained by calling the two-dimensional cross-index position corresponding to the timestamp and number in the reflectance record matrix. The band range is from 400nm to 1000nm, and the reflectance value is recorded at an interval of 5nm, which constitutes 121 band dimensions. For example, the number T012 corresponds to the time point 20240712T1032. The reflectance vector obtained by searching is [r1, r2, ..., r 121 ], and number each band uniformly from B1 to B 121, and perform serial number correction operations on their position sorting, establish a number index mapping from T012 to the band vector, and complete the structured construction from the number to the band reflectance value by traversing all sample numbers. Finally, with the number as the primary key and the band sequence number as the dimension field, a reflectance band comparison set corresponding to the individual tobacco plant is constructed.

[0070] S103: Calling the light reflectance band comparison set and agronomic variable mapping information, performing position matching in the variable dimension based on the agronomic parameters with the same number and the corresponding band reflectance values, combining the agronomic indicators and the connected multiple spectral values ​​into a set of vector structures, and establishing a set of combined parameters within the partition using the tobacco plant number as the index unit to generate an agronomic morphological spectral structure set;

[0071] Call the reflectance band control set and agronomic variable mapping information, and read the agronomic parameter vector (leaf area, SPAD value, leaf length, leaf width, plant height) and reflectance vector (B1 to B 121 ), and then the two vectors are connected in parallel according to the number to generate a combined vector structure. The length of each vector group is 126 dimensions, of which the first 5 dimensions are agronomic parameters and the last 121 dimensions are band reflectance values. The combined vector is inserted into the mapping table with the number as the key. During the data structure construction process, it is necessary to determine whether the SPAD value of the agronomic parameter is in the normal range. The reference range of the SPAD value is set to 35.0-45.0. When the SPAD value is greater than 45.0, it is marked as a high chlorophyll sample, and less than 35.0 is marked as a low chlorophyll sample. For example, the SPAD value of sample T012 is 38.4, which is within the standard range and marked as a normal category. The combined vector is represented as (188.6, 38.4, 25.1, 13.7, 92.5, r1, r2, ..., r 121 ), and finally add the combined vectors of all samples into the sample vector table with the number as the key to generate the agronomic morphological spectral structure set of all tobacco plants in the partition.

[0072] See also Figure 3 The specific steps for obtaining the near-infrared chemical fusion feature group are as follows:

[0073] S201: Based on the agronomic morphology spectral structure set, obtain flue-cured tobacco leaves with consistent tobacco plant numbers, collect near-infrared reflectance spectrum data corresponding to each sample, divide the near-infrared reflectance of each tobacco leaf sample into bands according to the number sequence, and extract the main reflectance band value in the continuous interval to generate a near-infrared main band value group;

[0074] Based on the agronomic morphology spectral structure set, the coding field of each numbered tobacco plant is first extracted, and the tobacco leaf sample with the same code is matched in the post-baking sample number set through the field index. The entity of the tobacco leaf sample after baking is obtained, and the full-band scan is performed on it using a near-infrared reflectance spectrometer. The original reflectance spectral vector obtained by sampling at 1nm intervals in the range of 400nm to 2500nm is collected. After acquisition, the spectral curve of each tobacco leaf sample is segmented according to a fixed interval. The specific division method is to divide the band into five continuous segments, namely segment one is 400-799nm, segment two is 800-1199nm, segment three is 1200-1599nm, segment four is 1600-1999nm, and segment five is 2000-2500nm. In each segment, each band is called The corresponding reflectance value is calculated by averaging the reflectance values ​​of all bands within the segment to obtain the mean value, and the difference between the reflectance of each band and the mean value is compared. The band with the smallest difference is selected as the main reflectance band of the segment by setting the judgment condition. The specific judgment formula is ΔRi = |Ri–μ|, where Ri is the reflectance value of the i-th band, μ is the mean reflectance of all bands in the segment, and the band with the smallest difference ΔRi is the main band. For example, the tobacco leaf numbered T056 has a band reflectance mean of 0.438 in the band 1200–1599nm of segment three. When the reflectance at band 1326nm is 0.441 and ΔR is 0.003, which is smaller than the ΔR of other bands, then the band is set as the main band of segment three. In this way, one main band is selected from each of the five segments, and finally a near-infrared main band value group containing five elements is formed.

[0075] S202: Based on the near-infrared main band value group, the moisture content, total sugar, reducing sugar, total nitrogen, nicotine, potassium, and protein mass percentage of the corresponding tobacco leaf sample are detected. The binding relationship between the post-curing sample number and the tobacco plant number at the time of collection is called to establish a corresponding index between the sample numbers. The chemical composition data is added to the original band value record structure as a variable field to generate a component reflection correlation information array.

[0076] Based on the near-infrared main band value group, the five main band reflectance values ​​corresponding to each sample were called and input into the analysis module of the physical and chemical detection device respectively. By referring to the set standard detection process, the sample mass was weighed after constant temperature drying and the moisture content was obtained by thermogravimetry. The mass percentage was expressed as the ratio of the mass difference before and after water evaporation to the initial mass. The total sugar and reducing sugar contents were detected by chemical titration, the total nitrogen mass percentage was detected by Kjeldahl method, the nicotine content was determined by gas chromatography, atomic absorption spectrometry was used to determine the potassium mass percentage, and the protein content was measured by Bradford protein quantification method. The above data were processed into numerical values ​​expressed in percentages, and reasonable interval reference values ​​for the seven chemical index values ​​were set, such as moisture content of 10%-16%, total sugar content of 18%-28%, reducing sugar of 14%-22%, and total nitrogen of 1. .4%–2.4%, nicotine 1.8%–3.0%, potassium 1.6%–2.8%, and protein 5.0%–9.0%. The obtained data is bound to the aforementioned tobacco plant number field. By comparing the sample numbers in the agronomic structure set, a one-to-one match is performed with the flue-cured tobacco leaf sample numbers. After the corresponding fields are merged, the seven chemical component values ​​are added as new variable fields to the original main band value group record structure. For example, the main band group corresponding to number T056 is (0.362, 0.448, 0.441, 0.425, 0.395). The detected moisture content is 13.2%, total sugar 25.6%, reducing sugar 19.3%, total nitrogen 2.1%, nicotine 2.4%, potassium 2.3%, and protein 7.1%. These seven values ​​are appended to the band value group record, and the component reflection association information matrix of this data is finally completed.

[0077] S203: Recalling component reflectance association information, marking the acquisition time and field number information of each data record, aggregating the field division number and sample reflectance data structure, placing sample records in the same area into independent data segments, combining the reflection spectrum segments and the aggregated data structure of the component indicators, and establishing a near-infrared chemical fusion feature group;

[0078] Call the component reflection association information, first read the collection time field and field number field in each record, confirm that the time information format is YYYYMMDD and record the daily sample collection batch number, extract the field number and establish a field area index table, classify the records according to the field number field, and group all sample records under the same number through the field number grouping operation to form an independent data segment. Each data segment contains several record data, and each record data consists of five near-infrared main band values ​​and seven component mass percentages. After combination, a 12-dimensional vector structure is generated, and all the data segments are sorted by unified numbering. All sample data are aggregated within each field to ensure that the records are arranged in chronological order according to the collection time. For example, the field numbered FD105 has 10 records, which are numbered FD105_01 to FD105_10 after being sorted according to the collection time field 20240710 to 20240715. The corresponding sample data structures are (0.362, 0.448, 0.441, 0.425, 0.395, 13.2, 25.6, 19.3, 2.1, 2.4, 2.3, 7.1), respectively. These data are merged to form the near-infrared chemical fusion feature group of the field.

[0079] See also Figure 4 , the specific steps for obtaining the fusion indicator vector set are:

[0080] S301: Based on the agronomic morphological spectral structure set and the near-infrared chemical fusion feature group, record entries with the same tobacco plant number are identified, and the agronomic parameters and chemical indicators are field-joined according to the number consistency to construct a parameter set containing multiple indicator dimensions and generate sample parameter splicing values;

[0081] Based on the agronomic morphology spectrum structure set and the near-infrared chemical fusion feature group, the primary key number of each record in the agronomic morphology spectrum structure set was first extracted, and the sample data entry with the same number was found in the near-infrared chemical fusion feature group. The matching operation of the number field was performed to ensure that the primary key numbers in the two types of data structures were consistent. Then, the agronomic parameter vector (leaf area, SPAD value, leaf length, leaf width, plant height, 121-dimensional visible-near-infrared spectral reflectance value) and the chemical indicator vector (five-segment main band reflectance, water content, total sugar, reducing sugar, total nitrogen, nicotine, potassium, protein) were field-spliced ​​using the sample number as the index unit. The splicing order was to splice the agronomic parameters first, then Concatenate the chemical indicator fields to form a complete multi-category indicator sample vector. After concatenation, record the sample number and store the combined vector. For example, for sample number T023, the agronomic parameters are (201.2, 40.3, 27.5, 14.2, 95.1) and the 121-dimensional reflectance value. The chemical main band is (0.362, 0.448, 0.441, 0.425, 0.395), and the component values ​​are (13.2, 25.6, 19.3, 2.1, 2.4, 2.3, 7.1). The combined result is a vector structure with a total length of 134 dimensions. Repeat the above concatenation process for each primary key number to form a unified format for the sample parameter concatenation value.

[0082] S302: Call the sample parameter splicing value, unify the unit system of agronomic variables and chemical indicators based on the primary key number of each record, and perform interval normalization on the reflectance, sugar concentration, and element content vector fields. Calculate the sample availability ratio between all fields, determine whether there are vacancies and scale inconsistencies between fields, calculate the sample field consistency measure, eliminate the field set whose consistency measure is lower than the availability threshold, and generate field consistency screening information;

[0083] The formula for calculating the sample field consistency measure is:

[0084]

[0085] Among them, U i represents the field consistency measure of the i-th sample, A ij represents the normalized agronomic parameter effectiveness index value of the i-th sample on the j-th field, B ij Indicates the normalized chemical composition of the i-th sample in the j-th field, C ij Indicates the unit standard error value of the i-th sample on the j-th field, D ij It represents the measurement interval error value of the i-th sample on the j-th field, i represents the index number of the sample, j represents the index number of the field, and m represents the total number of fields in the sample parameter set;

[0086] This indicator is used to evaluate the availability and structural consistency between agronomic data and chemical indicators at the field level during the sample fusion process, that is, whether the multi-source heterogeneous data has formed a high-quality unified vector that can participate in modeling after fusion. ij ·B ij Indicates that agronomic parameters and chemical indicators are valid for the current field; Represents the comprehensive fluctuation of measurement error; the higher the validity and the smaller the error, the larger the value of this item, indicating that the field has a high degree of structural consistency and reliability in the sample, so U i The larger the value is, the better the structural consistency of all fields of the sample is, and the more suitable it is for subsequent analysis and modeling.

[0087] To call the sample parameter splicing value, first extract the primary key number in each record, and then call the agronomic parameter field and chemical indicator field in turn. Standardize the field content, unify the leaf area to cm2, the leaf length, leaf width and plant height to cm, maintain the original unit of the SPAD value, unify the sugar and other chemical components to the mass percentage unit, and maintain the original value range of the main band reflectance and 121-dimensional reflectance. After completing the unit standardization, perform the minimum-maximum normalization operation on all numerical fields according to the normalization formula:

[0088]

[0089] Each field is normalized, and the value range is as follows: Taking plant height as an example, the maximum value is 122.5 cm and the minimum value is 68.3 cm. If the plant height of sample T023 is 95.1 cm, then its normalized value is (95.1-68.3) / (122.5-68.3)=26.8 / 54.2≈0.494. Other fields are processed similarly. After normalization, the availability ratio of each sample field is calculated one by one. For example, if there are 134 fields in the sample and there are 132 valid values, the availability ratio is 132 / 134≈0.985. If it is less than the set threshold (for example, 0.90), it is marked as unavailable. Then, the consistency between fields is measured, and the following formula is called:

[0090]

[0091] Assume that m = 3 fields are selected from sample T023 for consistency measurement (the fields are plant height, total sugar, and protein), and the example calculation is performed with field index j = 1, 2, 3:

[0092] Field 1 (plant height):

[0093] A i1 =0.494 (normalized agronomic value);

[0094] B i1=1.000 (the component value after normalization is set to 1.000);

[0095] C i1 =0.005 (standard error, after unit conversion);

[0096] D i1 =0.010 (measurement interval error);

[0097] The calculation item is

[0098] Field 2 (Total Sugars):

[0099] A i2 =0.820, B i2 =0.784, C i2 =0.015, D i2 =0.012.

[0100] The calculation item is

[0101] Field 3 (Protein):

[0102] A i3 =0.650, B i3 =0.697, C i3 =0.008, D i3 =0.006.

[0103] The calculation item is

[0104] Add the consistency contributions of the three fields and take the absolute value:

[0105] U T023 =|44.17+33.47+45.30|=122.94;

[0106] If the consistency threshold is set to 100.0, then the sample field consistency metric of T023 is greater than the threshold, and the sample is judged to be a set of well-consistent fields and can participate in the subsequent data structure packaging process. Conversely, if the calculation result is less than 100.0, the relevant fields will be recorded in the elimination field set. Ultimately, a field consistency screening information table is formed to record whether each field passes the measurement requirements. For example, if fields 1, 2, and 3 of T023 all pass, but field 8 fails, field 8 is eliminated, and fields 1, 2, and 3 are retained. They will be extracted and encapsulated in the field vector set structure later.

[0107] S303: Filter information based on field consistency, extract the remaining valid field groups, construct a primary key vector structure for the retained fields under each sample, encapsulate the vectors by tobacco plant number and store them as partitioned data units, and establish a fusion indicator vector set;

[0108] Based on the field consistency screening information, the retained field index in the screening results is read, and the field values ​​marked as valid under each sample are extracted in sequence. The field primary key vector structure of the sample is constructed. The retention order is agronomic fields first and component fields last. The field index is stored in the form of a structure array. For example, T023 only retains the leaf area, plant height, main bands three and four, nicotine, and protein fields. The vector structure is (201.2, 95.1, 0.441, 0.425, 2.4, 7.1). A 6-dimensional vector result bound to the T023 number is established. Then, based on the sample's primary key number, all valid field vector structures are partitioned and grouped according to the area to which the number belongs. For example, if the field number of T023 is FD105, its structure data will be assigned to the FD105 partition. Repeat this step to perform encapsulation and allocation for all retained field samples. Finally, the valid field set of samples under each field is stored as a fusion indicator vector set.

[0109] See also Figure 5 ,The steps for obtaining the label alignment dataset are as follows:

[0110] S401: Based on the fusion indicator vector set, obtain the sensory score data of the flue-cured tobacco leaf sample matching each sample number, extract the aroma quality, aroma quantity, pungency, aftertaste, and harmony score values, and bind the score data structure to the original sample number by field, construct a mapping structure between the sample primary key and multiple scores, and generate sensory score binding information;

[0111] Based on the fusion indicator vector set, the primary key number field in each vector structure is first called to match the registered sensory score samples in the sensory evaluation record dataset one by one. The corresponding score information is extracted through the number field matching operation, and the five score values ​​of aroma quality, aroma quantity, pungency, aftertaste, and harmony are extracted from the score field. All scores are standardized expert scores with a score range of 0 to 10. Each record contains the corresponding five field values. For example, the sensory score record of tobacco leaf numbered T045 is 7 points for aroma quality, 8 points for aroma quantity, 5 points for pungency, 6 points for aftertaste, and 7 points for harmony. The five-dimensional score vector structure (7, 8, 5, 6, 7) is bound to the fusion indicator vector number T045 through field mapping operation, forming a direct mapping relationship between primary key T045 and sensory score. After all primary key numbers are matched, a complete score information table is formed, which is then integrated into the score binding dataset with the sample number as the index field to generate sensory score binding information.

[0112] S402: Call the sensory score binding information, filter the numbered records with missing score data, index and remove the missing samples, and determine the difference ratio between the maximum and minimum values ​​of the retained sample score records. Correct the score records that exceed the upper limit of the score setting range and are lower than the lower limit, and generate a valid sample score interval value.

[0113] Call the sensory score binding information, first read whether there is an empty score field in each record, traverse the five score fields of aroma quality, aroma quantity, pungency, aftertaste, and coordination, perform primary key number indexing on sample records with empty values ​​and remove the record, and no longer participate in subsequent score analysis operations. For example, if the aftertaste score of number T051 is missing, the entire sample record T051 is marked as unavailable. After removal, perform maximum and minimum value statistical operations on each score field in the retained record. For example, the maximum value of the aroma quality score is 9 points and the minimum value is 4 points. The calculated difference ratio is (9-4) / 10=0.5. The five score fields are then All difference ratio calculations are performed, and it is determined whether there are outlier fields. The upper limit of the score setting range is defined as 10 points and the lower limit is 0 points. If the score of a record exceeds the range, such as the value of the fragrance volume field is 11 points, it will be corrected to 10 points; if the irritation score value of a record is -1, it will be corrected to 0 points; the specific correction process is: if the score value is greater than 10, it is assigned to 10, if it is less than 0, it is assigned to 0, and after correction, it is reinserted into the score field, keeping the difference ratio in the field within the range of [0,1], that is, the maximum difference value does not exceed the width of the score interval, and finally a set of valid sample score interval values ​​with all score values ​​complete and legal is obtained.

[0114] S403: Based on the valid sample score interval value, the score field is filled into the corresponding sample primary key field in the fusion indicator vector set, the field structure is adjusted and the naming is unified, a label structured data format is constructed, and a label alignment dataset is established;

[0115] According to the valid sample score interval value, the score field is written into the corresponding position of the fusion indicator vector set according to the sample primary key number field by number index. Based on the structured vector array, the five score fields of aroma quality, aroma quantity, pungency, aftertaste, and coordination are appended to the end of each vector. The unified field naming method is "score_aroma_quality", "score_aroma_quantity",

[0116] "score_stimulation", "score_aftertaste", and "score_balance" are sorted in this order to the last field of the fusion indicator vector structure to form a standardized label post-field structure. For example, the original fusion indicator vector numbered T045 is (201.2, 0.441, 2.4, 7.1, ...), and the adjusted structure is (201.2, 0.441, 2.4, 7.1, ..., 7, 8, 5, 6, 7), where the score field is located in the last five digits. After completing the score field insertion, field position adjustment, and naming unification operations for all sample structures, the fusion indicator and score label are combined into a unified data structure and stored as a sample label structured data table to finally construct the label alignment dataset.

[0117] See also Figure 6 The specific steps for obtaining the field-level tobacco leaf quality grade distribution information are as follows:

[0118] S501: Based on the label alignment dataset, extract the main reflection band value, carbon-nitrogen ratio value, and near-infrared full spectrum intensity interval value from the feature vector of each sample. Combined with the harmony and aroma quality score items in the sensory score, the input data is normalized by field, and an input sample vector group is constructed to generate the structure score input value.

[0119] Based on the label alignment dataset, the primary key number field of each sample is first read, and the main reflection band value in the fusion index vector is extracted from it. Five band values ​​are selected as the input feature field, and then the carbon-nitrogen ratio is read from the chemical indicator field. This value is calculated by dividing the total carbon mass percentage by the total nitrogen mass percentage. For example, the carbon content of sample T072 is 38.0% and the total nitrogen is 1.9%, then the carbon-nitrogen ratio is 38.0 / 1.9≈20.00. Then the near-infrared full spectrum reflectance vector (121 dimensions) is extracted, and its intensity interval value is calculated. The operation method is to extract the maximum, minimum and mean values ​​of the reflectance and combine them into a three-dimensional interval feature field. For example, the maximum reflectance is 0.521, the minimum is 0.305, and the mean is 0.412, then the near-infrared intensity interval feature vector (0.521, 0.305, 0.412) is generated. ), then read the sensory score data in the label field, extract the two scores of harmony and aroma quality, for example, the harmony score is 8 points, and the aroma quality score is 7 points. The above five main band values, one carbon-nitrogen ratio, three near-infrared intensities, and two sensory scores, a total of 11 fields, are merged into the original input vector structure by column. The minimum-maximum normalization operation is performed on all numerical fields, and the normalization interval is set to [0,1]. All numerical values ​​are normalized based on the maximum and minimum values ​​of the sample field. For example, if the maximum value of a field is 30.0, the minimum is 10.0, and the current value is 20.0, then the normalization result is (20.0-10.0) / (30.0-10.0)=0.5. The normalized values ​​of all fields are calculated in this way, and finally a standardized input sample vector group with a length of 11 dimensions for each sample is generated to form the structural score input value.

[0120] S502: Calling the input value of the structure score, performing corresponding grade matching on the input sample vectors according to the score interval set by the harmony and fragrance quality score, calculating the numerical deviation of the attribute grade of each vector and comparing it with the label interval spacing, calculating the grade fit of each sample in the label space, matching and classifying according to the grade fit and the preset interval, and establishing the sample score grade value;

[0121] The formula for calculating the rank fit of each sample in the label space is:

[0122]

[0123] Among them, R k Indicates the fitting degree of the score of the kth sample, E kl represents the normalized value of the lth input field of the kth sample, F kl represents the central score value of the lth item in the target label level interval of the kth sample, G k represents the standard deviation of the kth sample input field, H k Indicates the standard deviation of the score label interval, M kIt represents the ratio of the kth sample label score to the corresponding input field, where k is the sample index, l is the field dimension index, and p is the total number of input fields;

[0124] This indicator is used to measure whether the feature vector of each tobacco leaf sample is highly consistent with its sensory rating range, that is, whether it has clear label representativeness. It is the basis for building a grading prediction model and spatial distribution deduction. The numerator is the deviation sum of the feature and the grade center value, which represents the "distance" from the label. The denominator is the joint uncertainty of the input and the label interval, and finally multiplied by M k Strengthen the actual fitting ability of input fields and labels. Therefore, R k The smaller the value, the more the sample's characteristic performance fits the rating level, and the more suitable it is for building feature-based label prediction and spatial distribution models.

[0125] Call the structure score input value, read the harmony and aroma quality scores in each record in sequence according to the sample number index, set the sensory score classification standard to three categories: Grade A is 9-10 points, Grade B is 7-8 points, and Grade C is 5-6 points. Match each sample to the grade according to its score against the above score interval. For example, a sample with a harmony of 8 and an aroma quality of 7 is classified as Grade B. Then calculate the difference between the normalized value of the sample input field and the center value of the grade target label interval to calculate the grade fit of each sample in the label space. Call the formula:

[0126]

[0127] Assume p = 11 fields, and assume that the normalized value of sample T072 at a certain field l = 3 is 0.51, and the corresponding level B has a central score of 0.55 in this field. Then the deviation of this field is |0.51-0.55| = 0.04, and the total deviation is 0.42 when accumulated item by item. The standard deviation of the input field of sample T072 is G k =0.09, the target grade label’s score interval span standard deviation is H k =0.07, then the denominator is Correspondence ratio M k It is defined as the mean score divided by the field mean. For example, if the mean score is 7.5 and the field mean is 6.2, then M k =7.5 / 6.2≈1.210, substituting all the values ​​into the equation:

[0128]

[0129] If the fitting threshold is set to 5.0, and the grade fitting degree R_k of sample T072 is judged to be lower than the threshold, the sample will be classified into the target grade B. Otherwise, it will be considered to be classified into grade A or grade C. The fitting degree is compared with the center fitting distance of each grade, and the smallest one is selected to be classified into that grade. Finally, a clear sample score grade value is generated for each sample record.

[0130] S503: Based on the sample grade values, the spatial position of the sample grade labels is paired according to the field number to which each sample belongs, and the number of samples of multiple grades in the same area is counted and accumulated. A spatial distribution block structure is generated by grade group to establish field-level tobacco leaf quality grade distribution information;

[0131] According to the sample score grade value, the field number field of each sample is read according to the primary key number, the sample grade value is mapped to the field area to which it belongs, and a mapping index structure between the field number and the sample grade is established. For all sample records under the same field number, the quantity distribution of each grade value is counted. For example, field FD102 contains a total of 25 samples, of which 6 are grade A, 13 are grade B, and 6 are grade C. The counting structure is {A:6, B:13, C:6}. According to the above statistical results, a spatial grade distribution unit is generated, with the field number as the primary index and the grade statistics as the field item. A spatial distribution block data structure is constructed, and the field structure data of each field is combined into an overall current spatial grade mapping table, ultimately forming field-level tobacco leaf quality grade distribution information.

[0132] See also Figure 7 A digital tobacco leaf quality evaluation system based on multi-source data fusion is provided. The digital tobacco leaf quality evaluation system based on multi-source data fusion is used to execute the digital tobacco leaf quality evaluation method based on multi-source data fusion. The system includes:

[0133] The agronomic morphology extraction module collects agronomic indicators of the numbered sample plants in the tobacco field and uses a hyperspectral instrument to obtain the corresponding leaf band reflectance values. Based on the sample number and the regional division number, the agronomic indicators and the spectral band reflectance are combined to generate an agronomic morphology spectral structure set.

[0134] The chemical feature fusion module obtains flue-cured tobacco leaves with consistent plant numbers based on the agronomic morphology spectral structure set, detects key chemical component indicators, maps and combines near-infrared spectral curves with chemical component indicators, and marks the source time and field information to obtain a near-infrared chemical fusion feature group;

[0135] The indicator vector screening module identifies intersecting tobacco plant numbers based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, fills in null values ​​and removes inconsistent items in multiple vector fields, and screens the sample availability of the merged feature vector set to obtain a fused indicator vector set.

[0136] The label alignment processing module establishes a label binding relationship between samples and sensory scores in the same flue-cured tobacco batch based on the fusion indicator vector set, removes samples with missing scores, corrects outliers in the score range, generates a usable data structure, and obtains a label alignment dataset.

[0137] The quality grade distribution module calls the label alignment dataset, selects the main reflection band value reflecting the spectral structure characteristics of the plant, the carbon-nitrogen related chemical ratio value and the near-infrared full spectrum intensity interval value as key input items, combines the coordination and aroma quality scores in the sensory label, and constructs the score grade identification layered by grade interval. It performs corresponding classification operations on each sample feature and the score grade, identifies the label grade range to which the sample belongs, and summarizes the spatial distribution with the samples from the same field to generate field-level tobacco leaf quality grade distribution information.

[0138] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A digital evaluation method for tobacco leaf quality based on multi-source data fusion, characterized in that: The following steps are involved: S1: Based on the numbered sample plants in the tobacco field, collect the plant's agronomic indicators and use a hyperspectral instrument to obtain the corresponding leaf band reflectance values. Based on the sample number and the regional division number, the agronomic indicators and the spectral band reflectance are combined to generate an agronomic morphological spectral structure set. S2: Based on the agronomic morphology spectral structure set, obtain flue-cured tobacco leaves with consistent tobacco plant numbers, detect key chemical component indicators, map and combine near-infrared spectral curves with chemical component indicators, and mark the source time and field information to obtain a near-infrared chemical fusion feature group; S3: Based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, identifying tobacco plant numbers with intersections, performing null value filling and inconsistent item removal on multiple vector fields, and performing sample availability screening on the merged feature vector set to obtain a fusion indicator vector set; S4: Based on the fusion indicator vector set, a label binding relationship is established between samples and sensory scores under the same flue-cured tobacco batch, samples with missing scores are eliminated, and outliers in the score range are corrected to generate a usable data structure and obtain a label alignment dataset.

2. The digital evaluation method for tobacco leaf quality based on multi-source data fusion according to claim 1 is characterized in that: The agronomic morphology spectrum structure set includes leaf growth trend identification, reflection band energy combination parameters, plant posture normalization variables, SPAD value spectrum response group and morphological feature spectrum mapping index; the near-infrared chemical fusion feature group includes chemical component spectrum correspondence coefficient, near-infrared band principal component factor, sample identification structure under regional number, content intensity curve structure and chemical indicator classification label; the fusion indicator vector set includes the primary key index table after feature splicing, field standardization result matrix, null value completion record set, feature field consistency label set and sample screening number list; the label alignment data set includes aroma dimension quality mapping table, coordination dimension scoring structure, sensory label and sample primary key mapping comparison table, abnormal label screening mark set and scoring structure validity confirmation record.

3. The digital evaluation method for tobacco leaf quality based on multi-source data fusion according to claim 2 is characterized in that: The steps for obtaining the agronomic morphology spectrum structure set are specifically as follows: S101: Based on the numbered sample plants in the tobacco field, the leaf area, SPAD value, leaf length, leaf width, and plant height of the plants are collected, and the field index number of each indicator is bound to the tobacco plant number to construct an agronomic parameter information set for each tobacco plant under the number, establish a mapping value between the corresponding number and the agronomic variable, and generate agronomic variable mapping information; S102: Based on the agronomic variable mapping information, the band reflectance data recorded by the hyperspectral data acquisition instrument corresponding to the tobacco plant number is called, and based on the docking position of the band response value between the time axis and the number index, the spectral band data structure corresponding to each tobacco plant number is constructed, and the band position sequence is unified to generate a reflectance band comparison set; S103: Call the reflectance band control set and agronomic variable mapping information, perform position matching in the variable dimension based on the agronomic parameters with the same number and the corresponding band reflectance values, combine the agronomic indicators and the connected multi-segment spectral values ​​into a set of vector structures, and establish a combined parameter set within the partition with the tobacco plant number as the index unit to generate an agronomic morphological spectral structure set.

4. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 3, characterized in that: The steps for obtaining the near-infrared chemical fusion feature group are specifically as follows: S201: Based on the agronomic morphology spectral structure set, flue-cured tobacco leaves with consistent tobacco plant numbers are obtained, near-infrared reflectance spectrum data corresponding to each sample is collected, near-infrared reflectance of each tobacco leaf sample is divided into bands according to the number sequence, and main reflectance band values ​​within the continuous interval are extracted to generate a near-infrared main band value group; S202: Based on the near-infrared main band value group, the moisture content, total sugar, reducing sugar, total nitrogen, nicotine, potassium, and protein mass percentage of the corresponding tobacco leaf sample are detected, the binding relationship between the post-curing sample number and the tobacco plant number at the time of collection is called, an index corresponding to the sample numbers is established, and the chemical composition data is added to the original band value record structure as a variable field to generate a component reflection correlation information array; S203: Call the component reflection association information, mark the acquisition time and field number information of each data record, aggregate the field division number and sample reflection data structure, and place the sample records in the same area into independent data segments. Combine the reflection spectrum segments and the aggregated data structure of the component indicators to establish a near-infrared chemical fusion feature group.

5. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 4, characterized in that: The steps for obtaining the fusion index vector set are specifically as follows: S301: Based on the agronomic morphological spectral structure set and the near-infrared chemical fusion feature group, identifying record entries with the same tobacco plant number, performing field splicing on agronomic parameters and chemical indicators according to the number consistency, constructing a parameter set containing multiple indicator dimensions, and generating a sample parameter splicing value; S302: calling the sample parameter splicing value, unifying the unit system of agronomic variables and chemical indicators based on the primary key number of each record, performing interval normalization processing on the reflectance, sugar concentration, and element content vector fields, calculating the sample availability ratio between all fields, and determining whether there are vacancies and scale inconsistencies between fields, calculating the sample field consistency measure, eliminating the field set whose consistency measure is lower than the availability threshold, and generating field consistency screening information; S303: Filter information based on the field consistency, extract the remaining valid field groups, construct a primary key vector structure of the retained fields under each sample, encapsulate the vectors according to the tobacco plant numbers and store them as partitioned data units, and establish a fusion indicator vector set.

6. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 5, characterized in that: The formula for calculating the sample field consistency metric is: Among them, U i represents the field consistency measure of the i-th sample, A ij represents the normalized agronomic parameter effectiveness index value of the i-th sample in the j-th field, B ij Indicates the normalized chemical composition of the i-th sample in the j-th field, C ij Indicates the unit standard error value of the i-th sample on the j-th field, D ij It represents the measurement interval error value of the i-th sample on the j-th field, i represents the index number of the sample, j represents the index number of the field, and m represents the total number of fields in the sample parameter set.

7. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 6, characterized in that: The steps for obtaining the label alignment dataset are specifically as follows: S401: Based on the fusion index vector set, sensory score data of the flue-cured tobacco leaf sample matching each sample number is obtained, and the aroma quality, aroma quantity, pungency, aftertaste, and coordination score values ​​are extracted. The score data structure is field-bound with the original sample number, and a mapping structure between the sample primary key and multiple scores is constructed to generate sensory score binding information; S402: Call the sensory score binding information, filter the numbered records with missing score data, index and remove the missing samples, and perform a difference ratio determination on the maximum and minimum values ​​of the retained sample score records, correct the score records that exceed the upper limit of the score setting range and are lower than the lower limit, and generate a valid sample score interval value; S403: According to the valid sample score interval value, the score field is filled into the corresponding sample primary key field in the fusion indicator vector set, the field structure is adjusted in position and the naming is unified, a label structured data form is constructed, and a label alignment data set is established.

8. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 7, characterized in that: The method further comprises: S5: calling the label alignment dataset, selecting the main reflectance band value, carbon-nitrogen related chemical ratio value and near-infrared full spectrum intensity interval value reflecting the spectral structure characteristics of the plant as key input items, combining the coordination and aroma quality scores in the sensory label, constructing a scoring grade identifier by grade interval, performing a corresponding classification operation on each sample feature and the scoring grade, identifying the label grade range to which the sample belongs, and summarizing the spatial distribution with samples from the same field to generate field-level tobacco leaf quality grade distribution information; The field-level tobacco leaf quality grade distribution information specifically includes a grade spatial distribution layer, a sensory label density block map, a set of grade segment boundaries between fields, a fusion index and label grade comparison table, and a hot zone archive number index.

9. The method for digital evaluation of tobacco leaf quality based on multi-source data fusion according to claim 8, characterized in that: The specific steps for obtaining the field-level tobacco leaf quality grade distribution information are as follows: S501: Based on the label alignment dataset, extract the main reflection band value, carbon-nitrogen ratio value, and near-infrared full spectrum intensity interval value from the feature vector of each sample, combine the harmony and aroma quality score items in the sensory score, normalize the input data by field, construct an input sample vector group, and generate a structure score input value; S502: Calling the structure score input value, performing corresponding level matching on the input sample vectors according to the score interval set by the coordination and fragrance quality score, calculating the numerical deviation of the attribution level of each vector and comparing it with the label interval spacing, calculating the level fit of each sample in the label space, matching and classifying according to the level fit and the preset interval, and establishing the sample score level value; The formula for calculating the rank fit of each sample in the label space is: Among them, R k Indicates the fitting degree of the score of the kth sample, E kl represents the normalized value of the lth input field of the kth sample, F kl represents the central score value of the lth item in the target label level interval of the kth sample, G k represents the standard deviation of the kth sample input field, H k Indicates the standard deviation of the score label interval, M k It represents the ratio of the kth sample label score to the corresponding input field, where k is the sample index, l is the field dimension index, and p is the total number of input fields; S503: Based on the sample scoring grade value, the spatial position of the sample grade label is paired according to the field number to which each sample belongs, and the number of multiple grade samples in the same area is counted and accumulated, and a spatial distribution block structure is generated according to the grade group to establish the field-level tobacco leaf quality grade distribution information.

10. A digital tobacco leaf quality assessment system based on multi-source data fusion, characterized in that: The system is used to implement the digital evaluation method for tobacco leaf quality based on multi-source data fusion according to any one of claims 1 to 9, and the system comprises: The agronomic morphology extraction module collects agronomic indicators of the numbered sample plants in the tobacco field and uses a hyperspectral instrument to obtain the corresponding leaf band reflectance values. Based on the sample number and the regional division number, the agronomic indicators and the spectral band reflectance are combined to generate an agronomic morphology spectral structure set. The chemical feature fusion module obtains flue-cured tobacco leaves with consistent tobacco plant numbers based on the agronomic morphology spectral structure set, detects key chemical component indicators, maps and combines near-infrared spectral curves with chemical component indicators, and marks the source time and field information to obtain a near-infrared chemical fusion feature group; The indicator vector screening module identifies tobacco plant numbers with intersections based on the agronomic morphological spectral structure set and the infrared chemical fusion feature group, performs null value filling and inconsistent item removal on multiple vector fields, and performs sample availability screening on the merged feature vector set to obtain a fusion indicator vector set; The label alignment processing module establishes a label binding relationship between samples and sensory scores under the same flue-cured tobacco batch based on the fusion index vector set, removes samples with missing scores, corrects outliers in the score range, generates a usable data structure, and obtains a label alignment dataset; The quality grade distribution module calls the label alignment dataset, selects the main reflection band value, carbon-nitrogen related chemical ratio value and near-infrared full spectrum intensity interval value reflecting the spectral structure characteristics of the plant as key input items, combines the coordination and aroma quality scores in the sensory label, constructs the score grade identification by grade interval, identifies the label grade range to which the sample belongs, and generates field-level tobacco leaf quality grade distribution information.