Tobacco spectral data processing method and device, electronic equipment, medium and product
By acquiring near-infrared spectral data of tobacco leaves and determining two-dimensional similarity using directional and shape similarity components, the problem of manual judgment in traditional tobacco leaf substitution methods is solved, enabling rapid and accurate screening of tobacco leaf substitutes and improving the efficiency and scientific nature of tobacco raw material matching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TOBACCO ZHEJIANG IND CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional methods of tobacco leaf substitution rely on manual judgment, which leads to high workload, strong subjectivity, low efficiency, and difficulty in finding the optimal substitute tobacco leaf.
By acquiring near-infrared spectral data of the tobacco leaves to be evaluated and the target tobacco leaves, two-dimensional similarity is determined using directional similarity components and shape similarity components, enabling rapid and accurate screening of alternative tobacco leaves.
It has achieved automation, standardization, and precision in tobacco leaf similarity determination and substitution screening, improving the efficiency and scientific nature of tobacco leaf raw material matching.
Smart Images

Figure CN122065047A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of tobacco raw material quality evaluation technology, and in particular to a method, apparatus, electronic device, medium and product for processing tobacco spectral data. Background Technology
[0002] Tobacco leaf substitution is the most important part of cigarette formulation. When a specific type of tobacco leaf runs out, it is necessary to find a substitute tobacco leaf and ensure that the overall quality of the cigarettes remains stable.
[0003] In related technologies, traditional tobacco substitution still relies primarily on manual judgment, based on tobacco chemical indicators combined with blending experience and sensory evaluation. However, this method requires repeated manual evaluation, comparison, and adjustment, which leads to problems such as high workload, strong subjectivity, low efficiency, and difficulty in finding the optimal substitute tobacco. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, medium, and product for processing tobacco leaf spectral data, so as to determine the two-dimensional similarity of tobacco leaves based on the directional similarity component and shape similarity component of the near-infrared spectral data of tobacco leaves, so as to achieve the effect of rapid and accurate screening of alternative tobacco leaves.
[0005] According to one aspect of the present invention, a method for processing tobacco leaf spectral data is provided, the method comprising:
[0006] Acquire first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of the target tobacco leaf;
[0007] For at least one of the tobacco leaves to be evaluated, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated; wherein the two-dimensional similarity is determined by the directional similarity component and the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0008] Based on at least one of the two-dimensional similarities, similar tobacco leaves for replacing the target tobacco leaf are determined from at least one of the tobacco leaves to be evaluated.
[0009] According to another aspect of the present invention, a tobacco leaf spectral data processing apparatus is provided, the apparatus comprising:
[0010] The spectral data acquisition module is used to acquire first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of the target tobacco leaf.
[0011] A two-dimensional similarity determination module is used to determine the two-dimensional similarity between at least one of the tobacco leaves to be evaluated and the target tobacco leaf based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated; wherein the two-dimensional similarity is determined by the directional similarity component and the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data;
[0012] A tobacco leaf screening module is used to determine, based on at least one of the two-dimensional similarities, similar tobacco leaves to replace the target tobacco leaf from at least one of the tobacco leaves to be evaluated.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When one or more programs are executed by one or more processors, the one or more processors implement a tobacco leaf spectral data processing method as described in any of the embodiments of this disclosure.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute any of the tobacco leaf spectral data processing methods of the present invention.
[0018] According to another aspect of the present disclosure, a computer program product is provided, which, when executed by a processor, implements a tobacco leaf spectral data processing method as described in any of the embodiments of the present disclosure.
[0019] The technical solution of this disclosure provides a common data foundation for subsequent spectral feature extraction and shape similarity component calculation based on piecewise derivative smoothing by acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf. This ensures the consistency and comparability of the spectral comparison analysis between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, for at least one tobacco leaf to be evaluated, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the first and second near-infrared spectral data. This two-dimensional similarity is determined based on the directional and shape similarity components between the first and second near-infrared spectral data, achieving a multi-dimensional and accurate characterization of the spectral features of the tobacco leaves and significantly improving the accuracy and reliability of the matching determination between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, by determining similar tobacco leaves from at least one tobacco leaf to be evaluated to replace the target tobacco leaf based on at least one two-dimensional similarity, the precision and efficiency of tobacco leaf quality matching are achieved, providing a scientific and quantifiable basis for the selection and substitution of tobacco raw materials. The technical solution of this disclosure solves the problems of high workload, strong subjectivity, low efficiency and difficulty in finding the best alternative tobacco leaves caused by the manual judgment method in related technologies. It realizes the determination of the two-dimensional similarity of tobacco leaves based on the directional similarity component and shape similarity component of the near-infrared spectral data of tobacco leaves, so as to quickly and accurately screen alternative tobacco leaves. It realizes the automation, standardization and precision of tobacco leaf similarity judgment and alternative screening, and greatly improves the efficiency and scientificity of tobacco raw material matching.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a method for processing tobacco leaf spectral data provided in this embodiment of the present disclosure;
[0023] Figure 2 A flowchart illustrating a method for processing tobacco leaf spectral data provided in this embodiment of the present disclosure;
[0024] Figure 3A flowchart illustrating an optional embodiment of a tobacco leaf spectral data processing method provided in this disclosure;
[0025] Figure 4 This is a schematic diagram of the structure of a tobacco leaf spectral data processing device provided in an embodiment of the present disclosure;
[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0030] It should be noted that in the field of tobacco raw material quality evaluation technology, near-infrared spectroscopy analysis technology, with its advantages of speed, non-destructive nature, and low cost, is widely used in quantitative analysis of chemical composition and sensory indicators, identification of origin and part of tobacco plant, analysis of flavor characteristics, and digital formulation substitution. As a type of fingerprint spectrum, near-infrared spectroscopy can comprehensively reflect the differences in the types and contents of chemical components between two types of tobacco leaves at the molecular level.
[0031] Currently, spectral similarity assessment methods mainly employ cosine similarity, correlation coefficient, and spectral distance. Cosine similarity treats the spectrum as a high-dimensional vector, assessing the angle between two high-dimensional vectors to determine their directional consistency in high-dimensional space. Correlation coefficient, in its algorithm, is a centered cosine similarity, representing the angular relationship between two vectors deviating from their central values. Theoretically, it treats two fingerprint spectra as multiple two-dimensional discrete points and examines the linear relationship between these data points. Some studies have introduced the concept of spectral distance into near-infrared spectral similarity comparison, essentially calculating the area enclosed by two spectral curves. However, these methods have limited applicability and must be combined with spectral preprocessing. Furthermore, near-infrared spectral fingerprints are ordered curves with temporal or spatial scales, exhibiting local curve shape characteristics resulting from sequential structures. Methods such as cosine similarity, correlation coefficient, and spectral distance can only characterize the numerical relationship between corresponding points, losing information about the arrangement order of spectral points, thus failing to comprehensively reflect the similarity between two near-infrared spectral curves.
[0032] Based on this, this embodiment introduces a two-dimensional similarity assessment method based on near-infrared spectral directional similarity and shape similarity, building upon traditional near-infrared spectral directional similarity. This method, established by mapping near-infrared spectral direction and shape similarity, can comprehensively characterize the consistency of two near-infrared spectra and achieve the transformation from spectral similarity to tobacco leaf similarity. This method possesses rigorous mathematical guidance, clear mathematical meaning, and strong interpretability, enabling it to scientifically and reliably characterize the similarity of different tobacco leaves, thus providing technical support for the digital substitution of tobacco leaf formulations.
[0033] Figure 1 This is a flowchart illustrating a method for processing tobacco leaf spectral data according to an embodiment of this disclosure. This embodiment is applicable to screening tobacco leaves to identify similar tobacco leaves used to replace target tobacco leaves. This method can be executed by a tobacco leaf spectral data processing device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers or servers. Figure 1 As shown, the method in this embodiment includes:
[0034] S110, acquire at least one first near-infrared spectral data of the tobacco leaf to be evaluated and a second near-infrared spectral data of the target tobacco leaf.
[0035] The tobacco leaves to be evaluated can refer to the tobacco raw materials that are to be assessed for similarity to determine whether they can replace the target tobacco leaves in the cigarette leaf blend formulation. Optionally, the tobacco leaves to be evaluated can be at least one of the following: tobacco leaves in the cigarette company's inventory, newly purchased tobacco leaves, tobacco leaves from different products, and tobacco leaves of different grades. The tobacco leaves to be evaluated can be processed according to preset sample preparation standards to ensure the stability and comparability of their near-infrared spectral data and avoid interference with the similarity assessment results due to differences in sample condition. The number of tobacco leaves to be evaluated can be one or more. Regardless of whether the number of tobacco leaves to be evaluated is one or more, the near-infrared spectral data obtained after acquiring near-infrared spectra of the tobacco leaves to be evaluated is the first near-infrared spectral data. The first near-infrared spectral data can refer to the set of spectral response values acquired using a near-infrared spectrometer on the processed tobacco leaves to be evaluated under preset conditions. The first near-infrared spectral data can contain multiple data points, each data point corresponding to the spectral absorption response value of the corresponding tobacco leaf to be evaluated at a specific wavenumber. The first near-infrared spectral data can refer to the absorption spectrum of the tobacco leaves to be evaluated acquired in the near-infrared band. First near-infrared spectral data reflects the types, content, and molecular structure of chemical components in the tobacco leaf being evaluated at the molecular level, serving as the core data carrier characterizing the quality characteristics of the tobacco leaf. The target tobacco leaf can refer to the benchmark tobacco leaf currently used in the cigarette leaf blend where alternative raw materials need to be found; it is the reference object for similarity assessment. Generally, the quality characteristics of the target tobacco leaf (such as chemical composition, sensory characteristics, and blend compatibility) are already clear. Its near-infrared spectral data serves as the benchmark for similarity calculation, used for comparative analysis with the first near-infrared spectral data of the tobacco leaf being evaluated, thereby screening out performance-matching alternative tobacco leaves. Second near-infrared spectral data can refer to the set of spectral response values collected using the same model of near-infrared spectrometer as the one used to collect the first near-infrared spectral data, under completely identical acquisition conditions for the target tobacco leaf processed according to the same sample preparation standards. Second near-infrared spectral data can also refer to the absorption spectrum of the target tobacco leaf collected in the near-infrared band. The data structure of the second near-infrared spectral data (including the number of data points, wavenumber correspondence, and response value dimension) is completely matched with that of the first near-infrared spectral data, ensuring that the two have a basis for direct comparison and avoiding systematic errors introduced due to differences in acquisition conditions.
[0036] In this embodiment, after determining at least one tobacco leaf to be evaluated, in order to acquire near-infrared spectra of the tobacco leaf to be evaluated, the tobacco leaf to be evaluated can be prepared according to a preset sample preparation standard, and the prepared tobacco leaf sample can be acquired by near-infrared spectra to obtain the first near-infrared spectral data of the tobacco leaf to be evaluated.
[0037] Optionally, acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated includes: preparing a sample of the tobacco leaf to be evaluated for at least one tobacco leaf to be evaluated, drying the tobacco leaf sample to be evaluated at a preset drying temperature and a preset drying time, and grinding the dried tobacco leaf sample to be evaluated into powder to obtain a tobacco leaf powder sample; using a Fourier transform near-infrared spectrometer to perform infrared scanning on the tobacco leaf powder sample, and acquiring spectral data in a preset spectral band to obtain first near-infrared spectral data corresponding to the tobacco leaf to be evaluated.
[0038] Sample preparation refers to a series of standardized processing operations performed on the tobacco leaves to be evaluated in order to obtain stable and comparable near-infrared spectral data. This can be used to eliminate the interference of differences in the physical state of the tobacco leaves (such as moisture, particle size, and morphology) on spectral detection. Optionally, sample preparation includes at least one step such as removing tobacco stems, cutting and crushing, and uniform mixing. Sample preparation can ensure the consistency of the processing effects in subsequent drying and grinding steps, laying the foundation for the reliability of spectral data. The tobacco leaf sample to be evaluated can refer to the intermediate product of tobacco leaves with a uniform morphology and no obvious impurities (such as tobacco stems and debris) after the sample preparation step. The morphology of the tobacco leaf sample to be evaluated is usually broken tobacco leaf fragments (rather than powder), which is a transitional form connecting the tobacco leaves to be evaluated with the subsequent drying and grinding steps, ensuring uniform moisture loss during the drying process and obtaining a powder sample with a uniform particle size after grinding. The preset drying temperature can refer to a predetermined constant temperature used to control the moisture content of the tobacco leaf sample to be evaluated. Optionally, the preset drying temperature includes 30℃, 35℃, 40℃, and 45℃. Preferably, the preset drying temperature can be 40℃. It should be noted that higher drying temperatures lead to the loss of low-boiling-point volatile aroma compounds in the tobacco samples being evaluated; conversely, lower drying temperatures may fail to reduce the moisture content of the samples. Therefore, to avoid excessively high temperatures causing the loss of low-boiling-point volatile aroma compounds while efficiently reducing the moisture content to a stable range, a preset drying temperature of 40°C can be set. The preset drying time refers to a predetermined, fixed duration for drying the tobacco samples at the preset drying temperature. Optionally, the preset drying time can include 7 hours, 8 hours, 9 hours, and 10 hours, etc. Preferably, the preset drying time can be 8 hours. It should be noted that longer drying times at the preset drying temperature result in the loss of low-boiling-point volatile aroma compounds, increasing unnecessary costs and potentially affecting the repeatability of spectral data; shorter drying times, due to insufficient and uneven moisture removal, may obscure the spectral signals of the target chemical components, reducing the comparability and reliability of the spectral data. Both excessively long and short drying times can interfere with subsequent similarity assessment results. Therefore, in order to ensure that the moisture content of tobacco leaf samples from different parts and production areas is stable within the range of 5%-6%, to minimize the interference of moisture on the spectrum, to achieve homogenization of the moisture content of different types of tobacco leaf samples to be evaluated, and to ensure the repeatability and comparability of subsequent spectral acquisition, the preset drying time can be set to 8 hours.
[0039] The dried tobacco leaf sample to be evaluated refers to a sample that has been dried at a preset temperature and time, resulting in a stable low moisture content (e.g., 5%-6%) without moisture fluctuations. The physical state of the dried sample remains that of dry tobacco leaf fragments, and its moisture content meets the requirements for near-infrared spectral acquisition. It can be directly used in subsequent grinding steps, avoiding spectral data distortion due to residual moisture. The tobacco powder sample refers to a uniform powder obtained by grinding the dried tobacco leaf sample and then sieving it through a 60-mesh sieve (250μm aperture). The core characteristics of the tobacco powder sample are uniform particle size (ensuring consistent diffuse reflectance during spectral acquisition), absence of large particles, and the ability for uniform near-infrared light penetration and full utilization of the tobacco's chemical components, ensuring the accuracy and representativeness of the spectral response values. A Fourier transform near-infrared spectrometer is the core instrument for acquiring near-infrared spectral data from tobacco powder samples. The working principle of a Fourier transform near-infrared spectrometer is understandable. It generates interference light through an interferometer, illuminates the sample, and receives the reflected / transmitted light signals. The Fourier transform converts the time-domain signal into a frequency-domain (wavenumber) signal, outputting spectral response values at different wavenumbers. Fourier transform near-infrared spectrometers possess stable wavenumber accuracy and response sensitivity, ensuring that the acquired spectral data accurately reflects the chemical composition characteristics of tobacco leaves. Infrared scanning refers to placing a tobacco powder sample in a sample cup, allowing it to settle naturally, and then placing the sample cup in the detection optical path of the Fourier transform near-infrared spectrometer. The spectrometer then irradiates the sample with near-infrared light and acquires signals according to preset parameters (such as the number of scans and resolution). The core of infrared scanning is to cause near-infrared light to interact with the chemical components (such as sugars, nicotine, and total nitrogen) in the tobacco powder sample, resulting in molecular vibrational absorption and generating spectral signals related to the type and content of the chemical components. The preset spectral band refers to the pre-defined wavenumber range used to acquire near-infrared spectral data. Optionally, the preset spectral band can be 3800-10000 per centimeter (3800-10000cm). -1 The value represents the number of vibrations of light waves per centimeter. The preset spectral bands cover the molecular vibrational (combination and overtone) absorption peaks of core chemical components in tobacco leaves, such as sugars, nicotine, and total nitrogen, which can comprehensively reflect the chemical composition characteristics of tobacco leaves and provide effective data support for subsequent similarity assessment.
[0040] In one embodiment, after identifying at least one type of tobacco leaf to be evaluated, the tobacco leaf can be processed according to a preset sample preparation procedure (removing stems, cutting and crushing, and uniformly mixing) to obtain a sample of tobacco leaf to be evaluated that is free of obvious impurities and has a uniform morphology. Further, the sample of tobacco leaf to be evaluated is placed in a preset drying temperature environment and dried for a preset drying time to ensure that the moisture content of the sample remains stable at a low level, resulting in a dried sample of tobacco leaf to be evaluated. Further, the dried sample of tobacco leaf to be evaluated is ground using a grinding device and passed through a 60-mesh sieve to obtain a tobacco leaf powder sample with uniform particle size. Further, the tobacco leaf powder sample is placed in a sample cup and naturally compacted. The sample cup is then placed in the detection optical path of a Fourier transform near-infrared spectrometer. The Fourier transform near-infrared spectrometer is used to irradiate the tobacco leaf powder sample with near-infrared light and acquire signals according to preset parameters (such as the number of scans and resolution). Spectral response data under a preset spectral band are acquired, ultimately obtaining first near-infrared spectral data that reflects the chemical composition characteristics of the tobacco leaf to be evaluated.
[0041] It should be noted that the target tobacco leaf can be processed using the first near-infrared spectral data acquisition method described above to obtain the second near-infrared spectral data of the target tobacco leaf. This embodiment will not repeat the description here.
[0042] S120. For at least one tobacco leaf to be evaluated, determine the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated.
[0043] Two-dimensional similarity refers to a comprehensive evaluation index constructed by fusing directional similarity components and shape similarity components, used to comprehensively characterize the degree of similarity between the tobacco leaf to be evaluated and the target tobacco leaf. In this embodiment, two-dimensional similarity can be determined by the directional similarity components and shape similarity components between the first near-infrared spectral data and the second near-infrared spectral data. The directional similarity component can be used to indicate the overall trend consistency of the first and second near-infrared spectral data. The directional similarity component can refer to a quantitative index of the degree of consistency of the overall numerical trend of the two high-dimensional vectors in high-dimensional space when the first and second near-infrared spectral data are regarded as high-dimensional vectors respectively. The directional similarity component can be used to describe the macroscopic trend similarity of the two spectral curves, including the overall upward and downward trend of the spectrum, the relative distribution position of peaks and valleys, and the consistency of the overall fluctuation amplitude of the values. Generally, the higher the directional similarity component, the more consistent the overall numerical trend of the two spectral curves. The shape similarity component can be used to indicate the local contour consistency of the first and second near-infrared spectral data. The shape similarity component can refer to a quantitative index of the degree of consistency of the spectral curves corresponding to the first and second near-infrared spectral data in terms of local contour features. Shape similarity components can be used to describe the microscopic similarity between two spectral curves, including changes in the slope of peaks and valleys, the location of inflection points, the steepness of local fluctuations, and the degree of matching between peak and valley shapes, directly reflecting the local morphological features of the spectral curves. Generally, a higher shape similarity component indicates that the local contour features of the two spectral curves are more consistent.
[0044] In this embodiment, after acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf, for the at least one tobacco leaf to be evaluated, the directional similarity component between the first and second near-infrared spectral data can be determined based on the first and second near-infrared spectral data of the tobacco leaf to be evaluated and the target tobacco leaf; and the shape similarity component between the first and second near-infrared spectral data of the tobacco leaf to be evaluated and the target tobacco leaf can be determined based on the directional and shape similarity components between the first and second near-infrared spectral data. Furthermore, the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf can be determined based on the directional and shape similarity components between the first and second near-infrared spectral data.
[0045] S130. Based on at least one two-dimensional similarity, identify similar tobacco leaves from at least one tobacco leaf to be evaluated for use as a substitute for the target tobacco leaf.
[0046] Among them, similar tobacco leaves refer to tobacco leaves selected from at least one type of tobacco leaves to be evaluated based on the quantitative results of two-dimensional similarity, and whose two-dimensional similarity meets the preset standards. Generally, similar tobacco leaves have a high degree of consistency with the target tobacco leaves in terms of chemical composition and spectral characteristics, and can replace the target tobacco leaves in the cigarette leaf blend, while ensuring the stability of the cigarette product style and the uniformity of quality.
[0047] In this embodiment, after obtaining the two-dimensional similarity corresponding to at least one tobacco leaf to be evaluated, the at least one two-dimensional similarity can be compared with a preset standard. Further, tobacco leaves whose corresponding two-dimensional similarity meets the preset standard can be screened out, and these screened tobacco leaves can be used as similar tobacco leaves to replace the target tobacco leaves. The preset standard can be flexibly set according to the formulation requirements and quality control thresholds of the cigarette product. Optionally, the preset standard includes at least one of the following: a two-dimensional similarity greater than or equal to a preset similarity threshold; a preset number of two-dimensional similarities ranked first after sorting by two-dimensional similarity from largest to smallest. The preset number can include 1, 3, or 5, etc.
[0048] In one implementation, after obtaining the two-dimensional similarity corresponding to at least one tobacco leaf to be evaluated, the at least one two-dimensional similarity can be sorted in descending order to obtain sorted two-dimensional similarity. Further, a predetermined number of two-dimensional similarities can be selected from the sorted two-dimensional similarities, and the tobacco leaf to be evaluated corresponding to the selected two-dimensional similarities can be determined as a similar tobacco leaf to replace the target tobacco leaf.
[0049] The technical solution of this disclosure provides a common data foundation for subsequent spectral feature extraction and shape similarity component calculation based on piecewise derivative smoothing by acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf. This ensures the consistency and comparability of the spectral comparison analysis between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, for at least one tobacco leaf to be evaluated, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the first and second near-infrared spectral data. This two-dimensional similarity is determined based on the directional and shape similarity components between the first and second near-infrared spectral data, achieving a multi-dimensional and accurate characterization of the spectral features of the tobacco leaves and significantly improving the accuracy and reliability of the matching determination between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, by determining similar tobacco leaves from at least one tobacco leaf to be evaluated to replace the target tobacco leaf based on at least one two-dimensional similarity, the precision and efficiency of tobacco leaf quality matching are achieved, providing a scientific and quantifiable basis for the selection and substitution of tobacco raw materials. The technical solution of this disclosure solves the problems of high workload, strong subjectivity, low efficiency and difficulty in finding the best alternative tobacco leaves caused by the manual judgment method in related technologies. It realizes the determination of the two-dimensional similarity of tobacco leaves based on the directional similarity component and shape similarity component of the near-infrared spectral data of tobacco leaves, so as to quickly and accurately screen alternative tobacco leaves. It realizes the automation, standardization and precision of tobacco leaf similarity judgment and alternative screening, and greatly improves the efficiency and scientificity of tobacco raw material matching.
[0050] Figure 2 This is a schematic flowchart illustrating another method for processing tobacco leaf spectral data provided in this embodiment. The technical solution of this embodiment can be combined with other embodiments; for the same or related parts, descriptions of other embodiments can be used, and will not be repeated here. Figure 2 As shown, the method in this embodiment may specifically include:
[0051] S210, acquire at least one first near-infrared spectral data of the tobacco leaf to be evaluated and a second near-infrared spectral data of the target tobacco leaf.
[0052] S220. For at least one type of tobacco leaf to be evaluated, perform spectral segmentation on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated, respectively, to obtain multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data.
[0053] Spectral segmentation refers to the process of dividing complete first and second near-infrared spectral data into multiple continuous sub-intervals based on the molecular vibrational absorption characteristics (differences in combination and overtone absorption) and interference levels (different degrees of scattering, noise, and baseline drift) of different bands in the near-infrared spectrum. The first spectral band refers to the multiple continuous sub-intervals of spectral data obtained after spectral segmentation of the first near-infrared spectral data. Each first spectral band corresponds to a specific molecular vibrational absorption type (combination, first overtone, second overtone), and water peaks and unstable regions have been removed. The second spectral band refers to the multiple continuous sub-intervals of spectral data obtained after segmenting the second near-infrared spectral data according to the exact same segmentation rules as the first near-infrared spectral data (same cutoff point, same band range). The number of bands and the corresponding wavenumber range of the second spectral band correspond one-to-one with the first spectral band.
[0054] In this embodiment, after obtaining the first near-infrared spectral data of the tobacco leaf to be evaluated and the second near-infrared spectral data of the target tobacco leaf, in order to improve the accuracy of the subsequent similarity determination results, spectral regions can be removed from the first and second near-infrared spectral data respectively, eliminating strong absorption water peaks and unstable spectral regions from the near-infrared spectral data. Furthermore, based on the spectral region removal results, multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data can be determined respectively. The spectral segmentation process will be specifically explained below using the first near-infrared spectral data as an example.
[0055] Optionally, the method for determining multiple first spectral bands corresponding to the first near-infrared spectral data includes: determining the first peak start point and the first peak end point corresponding to the first wavenumber based on a peak lookup function, a first wavenumber, and the first near-infrared spectral data; averaging the first peak start points corresponding to at least one tobacco leaf to be evaluated to obtain the first peak cutoff point corresponding to the first wavenumber; and averaging the first peak end points corresponding to at least one tobacco leaf to be evaluated to obtain the second peak cutoff point corresponding to the first wavenumber; and averaging the first peak cutoff points corresponding to at least one tobacco leaf to be evaluated; and determining the first peak cutoff point based on the peak lookup function, the second wavenumber, and the first near-infrared spectral data. The process involves determining the starting point and ending point of the second peak corresponding to the second wavenumber; averaging the starting points of the second peak corresponding to at least one tobacco leaf to be evaluated to obtain the cutoff point of the third peak corresponding to the second wavenumber; and averaging the ending points of the second peak corresponding to at least one tobacco leaf to be evaluated to obtain the cutoff point of the fourth peak corresponding to the second wavenumber; and segmenting the first near-infrared spectral data according to the preset bands to be removed, the first peak cutoff point, the second peak cutoff point, the third peak cutoff point, and the fourth peak cutoff point to obtain multiple first spectral bands corresponding to the first near-infrared spectral data.
[0056] The peak lookup function refers to an algorithmic function used to locate absorption peaks near a specific wavenumber in near-infrared spectral data. The peak lookup function can identify the peak point, the starting points on both sides of the peak, and the ending points in the spectral curve, providing precise location information for subsequently determining the peak intercept. The first wavenumber can refer to a pre-defined, strong absorption wavenumber that needs to be prioritized for removal from the near-infrared spectral data. The first wavenumber can be the core search target of the peak lookup function, used to locate the boundary range of this strong absorption peak. Optionally, the first wavenumber can be 5000 per centimeter (5000cm). -1 The wavenumber corresponds to the combination frequency absorption peak of the OH bonds in water molecules. The first peak start point can be identified in the first near-infrared spectral data using a peak lookup function, at the valley point to the left of the first wavenumber absorption water peak (i.e., the starting boundary point where the absorption peak begins to rise), representing the starting point of the left range of this strong absorption water peak. The first peak end point can be identified in the first near-infrared spectral data using a peak lookup function, at the valley point to the right of the first wavenumber absorption water peak (i.e., the termination boundary point where the absorption peak begins to fall and becomes stable), representing the ending point of the right range of this strong absorption water peak.
[0057] The first peak intercept point can refer to a unified location point obtained by arithmetically averaging the first peak starting points corresponding to at least one tobacco leaf being evaluated. Determining the first peak intercept point can be used to eliminate peak positioning errors in individual tobacco leaf samples and to determine the standardized removal boundary to the left of the first wavenumber absorption water peak. The second peak intercept point can refer to a unified location point obtained by arithmetically averaging the first peak ending points corresponding to at least one tobacco leaf being evaluated. The second peak intercept point can serve as the standardized removal boundary to the right of the first wavenumber absorption water peak, jointly defining the removal range of the strong absorption water peak with the first peak intercept point. The second wavenumber can refer to a pre-defined second strong absorption wavenumber that needs to be prioritized for removal in the near-infrared spectral data. The second wavenumber can be the core retrieval target of the peak lookup function, used to locate the boundary range of the strong absorption water peak. Optionally, the second wavenumber can be 7000 per centimeter (7000cm). -1The wavenumber corresponds to the overtone absorption peak of the OH bond in water molecules. The second peak start point can be identified in the first near-infrared spectral data using a peak lookup function, at the valley point to the left of the second wavenumber absorption peak (i.e., the starting boundary point where the absorption peak begins to rise), representing the starting point of the left range of this strong absorption peak. The second peak end point can be identified in the first near-infrared spectral data using a peak lookup function, at the valley point to the right of the second wavenumber absorption peak (i.e., the ending boundary point where the absorption peak begins to decline and stabilizes), representing the ending point of the right range of this strong absorption peak. The third peak cutoff point can be the unified position point obtained by arithmetically averaging the second peak start points corresponding to at least one tobacco leaf being evaluated. The third peak cutoff point can serve as the standardized elimination boundary to the right of the second wavenumber absorption peak. The fourth peak cutoff point can be the unified position point obtained by arithmetically averaging the second peak end points corresponding to at least one tobacco leaf being evaluated. The fourth peak intercept can serve as a standardized rejection boundary to the right of the second wavenumber absorption water peak, jointly defining the rejection range of this strong absorption water peak with the third peak intercept. The preset rejection bands can refer to pre-defined invalid bands that need to be removed from the first near-infrared spectral data, excluding the strong absorption peak regions corresponding to the first and second wavenumbers. Optionally, the preset rejection bands can be the wavenumber region corresponding to the first 50 data points of the first near-infrared spectral data, used to eliminate instability interference from the initial instrument acquisition. That is, the preset rejection bands can be 1-50 data points.
[0058] It should be noted that the strong absorption peak regions of water in the first and second near-infrared spectral data are removed because: on the one hand, the strong absorption peaks of water molecules in the near-infrared band can easily mask the characteristic signals of target components such as proteins and sugars, affecting the accuracy of subsequent smoothing and similarity calculations; on the other hand, the moisture content fluctuates greatly due to environmental influences, and removing the water peaks can eliminate this irrelevant variable, allowing the spectral data to focus on the intrinsic quality characteristics of tobacco leaves, thus improving the accuracy and reliability of similar tobacco leaf identification.
[0059] In this embodiment, the first peak start point and the first peak end point can be determined directly based on the output of the peak lookup function; alternatively, based on the peak lookup function, the first wavenumber, and the first near-infrared spectral data, the first peak value of the first wavenumber absorption peak is determined, and the position point to the left of the first peak value that is at a preset wavenumber range (the wavenumber of this position point is less than the wavenumber of the first peak value) is determined as the first peak start point, and the position point to the right of the first peak value that is at a preset wavenumber range (the wavenumber of this position point is greater than the wavenumber of the first peak value) is determined as the first peak end point. The preset wavenumber range can include 400 per centimeter, 500 per centimeter, or 600 per centimeter, etc. Similarly, the determination method for the second peak start point and the second peak end point is the same as that for the first peak start point and the first peak end point; at least one of the above determination methods can be used to determine the second peak start point and the second peak end point.
[0060] In one embodiment, the first wavenumber and first near-infrared spectral data can be processed according to a peak lookup function to obtain the first peak point of the first wavenumber absorption peak. Further, a point to the left of the first peak point, at a predetermined wavenumber range, can be determined as the first peak start point, and a point to the right of the first peak point, at a predetermined wavenumber range, can be determined as the first peak end point. Further, the average of the first peak start points corresponding to at least one tobacco leaf to be evaluated can be calculated, and the resulting average value can be determined as the first peak cutoff point corresponding to the first wavenumber. And, the average of the first peak end points corresponding to at least one tobacco leaf to be evaluated can be calculated, and the resulting average value can be determined as the second peak cutoff point corresponding to the first wavenumber. Further, the second wavenumber and first near-infrared spectral data can be processed according to a peak lookup function to obtain the second peak point of the second wavenumber absorption peak. Further, a point to the left of the second peak point, at a predetermined wavenumber range, can be determined as the second peak start point, and a point to the right of the second peak point, at a predetermined wavenumber range, can be determined as the second peak end point. Furthermore, the starting points of the second peak corresponding to at least one tobacco leaf to be evaluated can be averaged, and the resulting average value can be determined as the third peak intercept point corresponding to the second wavenumber. Similarly, the ending points of the second peak corresponding to at least one tobacco leaf to be evaluated can be averaged, and the resulting average value can be determined as the fourth peak intercept point corresponding to the second wavenumber. Further, the first near-infrared spectral data can be spectrally segmented according to preset bands to be removed, the first peak intercept point, the second peak intercept point, the third peak intercept point, and the fourth peak intercept point. Preset bands to be removed are then removed from the first near-infrared spectral data; the spectral region composed of the first and second peak intercept points is removed from the first near-infrared spectral data; and the spectral region composed of the third and fourth peak intercept points is removed from the first near-infrared spectral data. This results in multiple first spectral bands corresponding to the first near-infrared spectral data.
[0061] For example, assuming the first near-infrared spectral data has 1609 data points, the preset band to be removed can be 1-50 data points, and the first peak cutoff point is... The second peak intercept point is The third peak intercept point is The fourth peak intercept point is Furthermore, after segmenting the first near-infrared spectral data according to the preset bands to be removed, the first peak cutoff point, the second peak cutoff point, the third peak cutoff point, and the fourth peak cutoff point, the data point positions of the multiple first spectral bands are as follows: 51- , - , -1609. Furthermore, these three first spectral bands correspond to: 3800-5500 per centimeter, the combination absorption band of the molecule; 5500-7500 per centimeter, the first harmonic absorption band of the molecule; and 7500-10000 per centimeter, the second harmonic absorption band of the molecule.
[0062] S230. Based on multiple first spectral bands and multiple second spectral bands, determine the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0063] In this embodiment, after obtaining multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data, the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data can be determined based on the multiple first spectral bands and the multiple second spectral bands.
[0064] Optionally, based on multiple first spectral bands and multiple second spectral bands, the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data is determined, including: for the multiple first spectral bands, using a least squares fitting algorithm to smooth the first spectral bands and the corresponding second spectral bands respectively, to obtain smoothed first spectral bands and smoothed second spectral bands; using a correlation coefficient algorithm to calculate the correlation coefficient between the smoothed first spectral bands and the smoothed second spectral bands, as the band directional similarity component corresponding to the first spectral bands; and determining the band directional similarity component corresponding to the multiple first spectral bands as the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0065] The least squares fitting algorithm can refer to a smoothing algorithm used to reduce spectral noise and preserve effective signals. Optionally, the least squares fitting algorithm can be the Savitzky-Golay fitting algorithm (SG fitting algorithm for short). The core logic of the least squares fitting algorithm is to fit the data points within a local moving window of the spectral data using a quadratic polynomial, minimizing the sum of squared errors between the fitted curve and the original spectrum, thereby achieving denoising and smoothing and avoiding noise interference with the accuracy of similarity calculations. The second spectral band corresponding to the first spectral band can refer to the sub-interval spectral data among multiple second spectral bands that completely overlaps with the wavenumber range of a single first spectral band and corresponds to the molecular vibrational absorption characteristics (e.g., if the first spectral band is "51 to the first peak intercept point -1", then the corresponding second spectral band is a sub-interval within the same wavenumber range in the second near-infrared spectrum). The smoothed first spectral band can refer to the spectral data obtained after denoising a single first spectral band using the least squares fitting algorithm. The smoothed first spectral band has eliminated high-frequency noise interference while preserving the overall numerical trend characteristics of the spectrum. The smoothed second spectral band can refer to the spectral data obtained after denoising the second spectral band corresponding to the first spectral band using the least squares fitting algorithm, and its data structure is completely matched with the smoothed first spectral band.
[0066] The correlation coefficient algorithm can refer to an algorithm used to quantify the consistency of the overall numerical trend between two spectral bands. Optionally, the correlation coefficient algorithm can be the Pearson correlation coefficient algorithm. The band direction similarity component can refer to the quantized value calculated by the correlation coefficient algorithm for a single first spectral band and its corresponding second spectral band. The band direction similarity component can be used to characterize the consistency of the local overall numerical trend between the two spectra within the band (such as the rising and falling trends of the spectrum within the band, and the similarity of the relative distribution of peaks and valleys).
[0067] In this embodiment, the method of smoothing the first spectral band using the least squares fitting algorithm is the same as the method of smoothing the second spectral band using the least squares fitting algorithm. The process of smoothing the first spectral band will be described in detail below.
[0068] Optionally, the first spectral band is smoothed using a least squares fitting algorithm to obtain a smoothed first spectral band. This includes: traversing the first spectral band using a sliding window of a second preset size; for each sliding window traversed, inputting the spectral response values corresponding to multiple data points within the sliding window into a pre-constructed quadratic polynomial, and outputting the spectral response value corresponding to the sliding window; and determining the smoothed first spectral band based on the spectral response values corresponding to multiple sliding windows.
[0069] The second preset size sliding window can refer to a pre-defined fixed-length data window used for local data fitting of the first spectral band. Optionally, the second preset size can be 4, 5, or 6 data points in length, etc. Sliding traversal refers to the process of moving the second preset size sliding window sequentially from the starting data point of the first spectral band according to wavenumber (moving 1 data point at a time) until all data points of the band are covered. The core purpose is to perform segment-by-segment local fitting of the spectral data to achieve smooth noise reduction across the entire band. The spectral response value can refer to the quantized value of absorption intensity (such as absorbance value) at a specific wavenumber in the near-infrared spectral data.
[0070] In one implementation, for multiple first spectral bands, a sliding window of a second preset size can be used to traverse the first spectral bands. Further, for each traversed sliding window, the spectral response values corresponding to multiple data points within the sliding window can be input into a pre-constructed quadratic polynomial, and the output value is used as the second smoothed spectral response value corresponding to the sliding window. Further, after obtaining the second smoothed spectral response values corresponding to multiple sliding windows, the multiple second smoothed spectral response values can be arranged according to the traversal order of the sliding windows. To ensure that the number of spectral points before and after smoothing is the same, a preset number of spectral points can be interpolated at the beginning and end of the arranged second smoothed spectral response values using a preset interpolation algorithm. Further, the smoothed first spectral band can be determined based on the multiple interpolated second smoothed spectral response values.
[0071] For example, the spectral response value corresponding to the sliding window can be determined using the following formula:
[0072]
[0073] in, Indicates the first The spectral response values corresponding to each sliding window; Indicates the first The spectral response value corresponding to the first data point within a sliding window; Indicates the first The spectral response value corresponding to the second data point within the sliding window; Indicates the first The spectral response value corresponding to the third data point within the sliding window; Indicates the first The spectral response value corresponding to the fourth data point within the sliding window; Indicates the first The spectral response value corresponding to the fifth data point within the sliding window.
[0074] In this embodiment, after obtaining the smoothed first spectral band and the smoothed second spectral band, the correlation coefficient algorithm can be used to calculate the correlation coefficient between the smoothed first spectral band and the smoothed second spectral band to obtain the band direction similarity component corresponding to the first spectral band.
[0075] In one embodiment, a first average response value corresponding to the smoothed first spectral band can be determined based on the smoothed first spectral band, and a second average response value corresponding to the smoothed second spectral band can be determined based on the smoothed second spectral band. Further, a correlation coefficient between the smoothed first spectral band and the smoothed second spectral band can be determined based on the spectral response value corresponding to each data point in the smoothed first spectral band, the first average response value, the spectral response value corresponding to each data point in the smoothed second spectral band, and the second average response value, and the obtained correlation coefficient can be used as the band direction similarity component corresponding to the first spectral band. Further, the band direction similarity components corresponding to multiple first spectral bands can be determined as the direction similarity components between the first near-infrared spectral data and the second near-infrared spectral data.
[0076] For example, the band direction similarity component corresponding to the first spectral band can be determined using the following formula:
[0077]
[0078] in, This indicates the band direction similarity component corresponding to the first spectral band; This represents the total number of data points in the smoothed first spectral band (or the smoothed second spectral band); Indicates the first spectral band after smoothing. The spectral response values corresponding to each data point; This represents the first average response value corresponding to the first spectral band; Indicates the second spectral band after smoothing. The spectral response values corresponding to each data point; This represents the second average response value corresponding to the second spectral band.
[0079] S240. Based on multiple first spectral bands and multiple second spectral bands, determine the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0080] In this embodiment, after obtaining multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data, the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data can be determined based on the multiple first spectral bands and the multiple second spectral bands.
[0081] Optionally, based on multiple first spectral bands and multiple second spectral bands, the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data is determined, including: for the multiple first spectral bands, a least squares fitting algorithm and a cubic polynomial are used to perform derivative smoothing on the first spectral bands to obtain processed first spectral bands; and a least squares fitting algorithm and a cubic polynomial are used to perform derivative smoothing on the second spectral bands corresponding to the first spectral bands to obtain processed second spectral bands; a correlation coefficient algorithm is used to calculate the correlation coefficient between the processed first spectral bands and the processed second spectral bands, as the band shape similarity component corresponding to the first spectral bands; and the band shape similarity component corresponding to the multiple first spectral bands is determined as the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0082] Here, a cubic polynomial can refer to a pre-constructed third-order polynomial model, whose formula is as follows: ,in, This is the spectral response value; This parameter represents the relative position of spectral data points within the sliding window and can be used to establish positional relationships between data points within the sliding window. Indicates the parameters to be fitted; This represents the fitting error. For example, in the case where the sliding window length is 15 data points, .
[0083] In this embodiment, the method of using the least squares fitting algorithm and cubic polynomial to perform derivative smoothing on the first spectral band is the same as the method of using the least squares fitting algorithm and cubic polynomial to perform derivative smoothing on the second spectral band. The process of performing derivative smoothing on the first spectral band will be described in detail below.
[0084] Optionally, the first spectral band is smoothed by taking the derivative using a least squares fitting algorithm and a cubic polynomial to obtain the processed first spectral band. This includes: sliding the first spectral band through a sliding window of a first preset size; for each sliding window, determining the first smoothed spectral response value corresponding to the sliding window based on the spectral response values corresponding to multiple data points in the sliding window and a pre-determined sequence of weighting coefficients corresponding to the multiple data points; wherein the weighting coefficients are determined according to the least squares fitting algorithm and the cubic polynomial; and determining the processed first spectral band based on the first smoothed spectral response values corresponding to the multiple sliding windows.
[0085] The first preset size sliding window can refer to a pre-defined fixed-length data window used for local data fitting of the first spectral band. Optionally, the first preset size can be 14, 15, or 16 data points, etc. Sliding traversal refers to the process of moving the sliding window of the first preset size sequentially from the starting data point of the first spectral band according to wavenumber (moving one data point at a time) until all data points of the band are covered. The core purpose is to perform segment-by-segment local fitting of the spectral data to achieve smooth denoising across the entire band. The weighting coefficient sequence can refer to a fixed coefficient sequence derived based on the least squares fitting algorithm and cubic polynomial derivation. The number of elements in this sequence is consistent with the number of data points within the sliding window. The weighting coefficient sequence is only related to the window size and relative position parameters of the sliding window, and is independent of the first spectral band; it can be predetermined and reused.
[0086] Optionally, the steps for determining the weight coefficient sequence corresponding to multiple data points include: constructing multiple cubic polynomial equations based on the multiple data points within the sliding window; constructing a design matrix corresponding to the sliding window based on the multiple polynomial equations; processing the design matrix using a least squares fitting algorithm to obtain a cubic polynomial model corresponding to the sliding window and a weight coefficient matrix corresponding to the sliding window; calculating the first derivative of the cubic polynomial model; and determining the weight coefficient sequence corresponding to the multiple data points within the sliding window based on the derivative result and the weight coefficient matrix.
[0087] Here, a cubic polynomial equation can refer to an equation of the form [equation missing] constructed based on the relative index of data points within a sliding window and the spectral response value. The equation, where, Indicates the first One data point; Indicates the relationship with the first The relative position parameters corresponding to each data point; Indicates the parameters to be fitted; Indicates the fitting error; Indicates the relationship with the first The spectral response values corresponding to each data point. The design matrix can be defined by the relative positions of each data point within the sliding window. A matrix consisting of powers 0 to 3 is of the form: The number of rows equals the number of data points within the sliding window, and the number of columns equals the number of parameters to be fitted in the cubic polynomial equation. The design matrix can be composed of the coefficients of multiple cubic polynomial equations. For example, assuming there are 15 data points within the sliding window, the design matrix can be a 15-row, 4-column matrix, with each row in the form of... The cubic polynomial corresponding to the sliding window can refer to the cubic polynomial with specific parameter values obtained after solving it using the least squares fitting algorithm, which can accurately fit the local features of the spectral data within the sliding window. The least squares fitting algorithm can refer to a data algorithm that solves the cubic polynomial equation corresponding to the design matrix by minimizing the sum of squared fitting errors to find the optimal parameters. The weight coefficient matrix can refer to the coefficient matrix obtained during the least squares fitting process. The number of columns in this coefficient matrix equals the number of data points within the sliding window, and the number of rows equals the number of parameters to be fitted in the cubic polynomial equation. Each row of elements corresponds to the calculated weight of a parameter to be fitted in the cubic polynomial equation. The first derivative calculation can refer to the operation of directly calculating the first derivative of the cubic polynomial corresponding to the sliding window, which can be used to extract information about the slope change of the spectral curve and characterize the local shape features of the spectrum. The weight coefficients can refer to the elements extracted from the weight coefficient matrix that correspond to the target fitting parameters in the cubic polynomial equation. The number of these elements can be equal to the number of data points within the sliding window. The target fitting parameters used to extract the weight coefficients can be determined based on the derivative results of the cubic polynomial.
[0088] In one implementation, a sliding window of a first preset size is used to traverse the first spectral band. For any sliding window reached, the middle data point (midpoint) within the sliding window can be set as the relative position origin, and its corresponding relative position parameter value is a first value. Further, based on the relative position origin, negative integers (e.g., from -1 to -7) are sequentially assigned to all data points to the left of the sliding window, and positive integers (e.g., from 1 to 7) are sequentially assigned to all data points to the right of the sliding window. Thus, the relative position parameters of multiple data points within the sliding window can be obtained. Further, the relative position parameters of the multiple data points within the sliding window, along with the spectral response values extracted from the first spectral band corresponding to the multiple data points, can be substituted into a cubic polynomial to construct multiple cubic polynomial equations corresponding to the sliding window. Then, based on the coefficients of the multiple cubic polynomial equations, a design matrix corresponding to the sliding window is constructed. Further, a least-squares fitting algorithm is used to process the design matrix and the multiple cubic polynomial equations to solve for the optimal fitting parameters of the cubic polynomial, thereby determining the cubic polynomial corresponding to the sliding window and obtaining the weight coefficient matrix. Furthermore, the cubic polynomial undergoes first-order derivative processing to obtain the derivative formula corresponding to the cubic polynomial. Substituting the relative position parameter values corresponding to the points in the sliding window into the derivative formula yields the derivative result, which includes the slope of the points in the sliding window being equal to the target fitting parameter in the cubic polynomial. Further, based on the derivative result, a row of elements corresponding to the target fitting parameter can be extracted from the weight coefficient matrix. This extracted row of elements serves as the weight coefficient sequence corresponding to multiple data points within the sliding window. This weight coefficient sequence can include the weight coefficient corresponding to each data point.
[0089] Furthermore, a sliding window of a first preset size is used to traverse the first spectral band. For each sliding window traversed, the spectral response value corresponding to each data point within the sliding window can be obtained from the first spectral band. Further, for multiple data points, the spectral response value corresponding to each data point can be multiplied by a pre-determined weighting coefficient to obtain the value to be superimposed corresponding to the data point. Further, the values to be superimposed corresponding to multiple data points are added together, and the resulting value is then compared with the first smoothed spectral response value corresponding to the sliding window. Further, after the sliding window has completed a full traversal of the first spectral band, the first smoothed spectral response values corresponding to all sliding windows can be arranged in wavenumber order of the first spectral band to obtain the arranged first smoothed spectral response values. Further, to ensure that the number of spectral points before and after smoothing is the same, a preset number of spectral points can be interpolated at the beginning and end of the arranged first smoothed spectral response values using a preset interpolation algorithm. Further, the processed first spectral band can be determined based on the multiple interpolated first smoothed spectral response values.
[0090] For example, taking a first preset size of 15 data points as an example, the relative position parameters and spectral response values corresponding to the 15 data points within the current sliding window are obtained. The relative position parameters corresponding to the 15 data points can be: Substituting the relative position parameters and spectral response values corresponding to the 15 data points into the general form of the cubic polynomial: This yields 15 cubic polynomial equations. For example, When, its corresponding cubic polynomial equation is: Furthermore, based on the 15 cubic polynomial equations, a 15×4 design matrix is constructed. Each line is in the form of Furthermore, minimize the sum of squared fitting errors of the 15 cubic polynomial equations. Design matrix Substituting the spectral response value of each data point within the sliding window into the parameter calculation formula of the least squares fitting algorithm: ,in, This represents a matrix consisting of the spectral response values corresponding to 15 data points (i.e., 15...). Furthermore, solving the above parameter-finding formula yields a set of definite parameters. This yields the cubic polynomial corresponding to the current sliding window (i.e., the fitted cubic polynomial). Furthermore, solving the above parameter calculation formula yields the weight coefficient matrix. Furthermore, for the well-fitted cubic polynomial... Find the first derivative to obtain the derivative formula. The relative position parameter of the point in the current sliding window. Substituting into the derivative formula, at this time... , The derivative formula simplifies to: Furthermore, the weighting coefficient matrix It is a 4×15 coefficient matrix, which is related to the window size and relative position parameters of the sliding window. The weight coefficient matrix has 4 rows, corresponding to the... The corresponding weight coefficient sequence, and The corresponding weight coefficient sequence, and The corresponding weight coefficient sequence and Corresponding weight coefficient sequence Furthermore, it can be based on Extract the elements of the second row from the weight coefficient matrix to form the weight coefficient sequence corresponding to multiple data points. Further, the following formula can be obtained: .
[0091] In this embodiment, after obtaining the processed first spectral band and the processed second spectral band, the correlation coefficient algorithm can be used to calculate the correlation coefficient between the processed first spectral band and the processed second spectral band to obtain the band shape similarity component corresponding to the first spectral band.
[0092] In one embodiment, a third average response value corresponding to the first spectral band can be determined based on the processed first spectral band, and a fourth average response value corresponding to the second spectral band can be determined based on the processed second spectral band. Further, a correlation coefficient between the processed first spectral band and the processed second spectral band can be determined based on the spectral response value, the third average response value, the spectral response value, and the fourth average response value corresponding to each data point in the processed first spectral band, and the obtained correlation coefficient can be used as the band shape similarity component corresponding to the first spectral band. Further, the band shape similarity components corresponding to multiple first spectral bands can be determined as the shape similarity components between the first near-infrared spectral data and the second near-infrared spectral data.
[0093] For example, the band shape similarity component corresponding to the first spectral band can be determined using the following formula:
[0094]
[0095] in, This represents the component whose band shape is similar to that of the first spectral band. This indicates the total number of data points in the first spectral band (or the second spectral band) after processing; Indicates the first spectral band after processing. The spectral response values corresponding to each data point; This represents the third average response value corresponding to the first spectral band; Indicates the second spectral band after processing. The spectral response values corresponding to each data point; This represents the fourth average response value corresponding to the second spectral band.
[0096] S250. Determine the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the directional similarity component and the shape similarity component.
[0097] In this embodiment, after obtaining the directional similarity component and the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data, the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf can be determined based on the directional similarity component and the shape similarity component.
[0098] Optionally, the directional similarity component includes band directional similarity components corresponding to multiple first spectral bands; the shape similarity component includes band shape similarity components corresponding to multiple first spectral bands; determining the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the directional similarity component and the shape similarity component includes: for multiple first spectral bands, determining the band two-dimensional similarity corresponding to the first spectral bands based on the band directional similarity component and the band shape similarity component; determining the information entropy weight corresponding to the first spectral bands based on the first information entropy corresponding to the multiple first spectral bands and the second information entropy corresponding to the multiple second spectral bands; and weighted summing the information entropy weights corresponding to the multiple first spectral bands and the band two-dimensional similarity to obtain the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf.
[0099] Among them, the two-dimensional band similarity refers to the comprehensive quantitative index obtained by fusing the band direction similarity component and the band shape similarity component corresponding to a single first spectral band. The two-dimensional band similarity can simultaneously characterize the overall trend and local shape similarity of two spectra within the first spectral band. The first information entropy refers to the quantitative value calculated for the smoothed first spectral band based on Shannon's information entropy theory. The first information entropy characterizes the richness of information in the spectral data within the first spectral band. The larger the first information entropy, the more abundant the information on the chemical components of tobacco leaves contained in the first spectral band. The second information entropy refers to the quantitative value calculated for the smoothed second spectral band based on Shannon's information entropy theory. The information entropy weight refers to the weight coefficient of the first spectral band determined based on the proportion of the information entropy of a single first spectral band to the total information entropy of all first spectral bands.
[0100] In one implementation, for multiple first spectral bands, the product between the band direction similarity component and the band shape similarity component corresponding to the first spectral band can be determined. The square root of this product is then taken, and the resulting value is determined as the two-dimensional band similarity corresponding to the first spectral band. Further, smoothed first spectral bands and smoothed second spectral bands can be obtained. Based on the spectral response value corresponding to each data point within the smoothed first spectral band, a first information entropy corresponding to the first spectral band is determined. Similarly, based on the spectral response value corresponding to each data point within the smoothed second spectral band, a second information entropy corresponding to the second spectral band is determined. Further, multiple first information entropies and multiple second information entropies can be added together to obtain the total information entropy. For multiple first spectral bands, a second spectral band corresponding to the first spectral band can be determined. The first information entropy corresponding to the first spectral band and the second information entropy corresponding to the second spectral band are added together to obtain a first segmented information entropy. The ratio between the first segmented information entropy and the total information entropy is determined, and this ratio is used as the information entropy weight corresponding to the first spectral band. Furthermore, after obtaining the information entropy weights corresponding to multiple first spectral bands, the product between the information entropy weight corresponding to each first spectral band and the corresponding band two-dimensional similarity can be determined to obtain the superimposed similarity corresponding to multiple first spectral bands. The multiple superimposed similarities are added together, and the similarity obtained after addition is used as the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf.
[0101] For example, taking a scenario where both the first and second spectral bands have three elements, the two-dimensional similarity of the band corresponding to the first spectral band can be determined using the following formula: ;in, Indicates the relationship with the first Two-dimensional similarity of the bands corresponding to the first spectral band. ; Indicates the relationship with the first The band direction similarity components corresponding to the first spectral band; Indicates the relationship with the first The band shape similarity components corresponding to the first spectral band.
[0102] Furthermore, the first information entropy can be determined using the following formula: ; ;in, Indicates the relationship with the first The first information entropy corresponding to the first spectral band; This represents the total number of data points in the first spectral band after smoothing. Indicates the first spectral band after smoothing. The spectral response values corresponding to each data point; Indicates the first The normalized spectral response values of each data point.
[0103] Furthermore, the information entropy weights can be determined using the following formula: ;in, Indicates the relationship with the first The information entropy weights corresponding to the first spectral band; Indicates the relationship with the first The first spectral band corresponding to the first The second information entropy weight of the second spectral band; The first information entropy corresponding to the first first spectral band, the first information entropy corresponding to the second first spectral band, and the first information entropy corresponding to the third first spectral band are represented sequentially. The second information entropy corresponding to the first second spectral band, the second information entropy corresponding to the second second spectral band, and the second information entropy corresponding to the third second spectral band are represented sequentially.
[0104] Furthermore, the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf can be determined using the following formula:
[0105]
[0106] in, This indicates the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf; This represents the two-dimensional similarity of the band to the first spectral band. This represents the information entropy weight corresponding to the first first spectral band; This represents the two-dimensional similarity of the band to the second first spectral band. This represents the information entropy weight corresponding to the second first spectral band; This represents the two-dimensional similarity of the band to the third first spectral band. This represents the information entropy weight corresponding to the third first spectral band.
[0107] S260. Based on at least one two-dimensional similarity, identify similar tobacco leaves from at least one tobacco leaf to be evaluated for use as a substitute for the target tobacco leaf.
[0108] The technical solution of this disclosure, by performing spectral segmentation on the first and second near-infrared spectral data of at least one tobacco leaf to be evaluated, obtains multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data; further, based on the multiple first and second spectral bands, a directional similarity component between the first and second near-infrared spectral data is determined; and based on the multiple first and second spectral bands, a shape similarity component between the first and second near-infrared spectral data is determined; further, based on the directional and shape similarity components, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined, thereby achieving refined and multi-dimensional quantitative analysis of the spectral characteristics of tobacco leaves, significantly improving the accuracy and reliability of tobacco leaf similarity determination.
[0109] Figure 3 This is a flowchart illustrating an optional embodiment of a tobacco leaf spectral data processing method provided in this disclosure. Specific implementation details of this method can be found in the following embodiments. Technical features that are the same as or similar to those in the above embodiments will not be repeated here.
[0110] See Figure 3 The method in this embodiment specifically includes the following steps:
[0111] First, first near-infrared spectral data of at least one stocked tobacco leaf and second near-infrared spectral data of the target tobacco leaf are collected. Further, spectral region removal is performed on both the first and second near-infrared spectral data to remove water peak regions and unstable spectral regions. Further, after spectral region removal, the first near-infrared spectral data can be divided into multiple first spectral bands; and after spectral region removal, the second near-infrared spectral data can be divided into multiple second spectral bands. Further, for at least one first spectral band, the SG smoothing algorithm is used to smooth both the first spectral band and the corresponding second spectral band to obtain the directional similarity component corresponding to the first spectral band. And, for at least one first spectral band, the derivative SG smoothing algorithm is used to perform derivative smoothing on both the first spectral band and the corresponding second spectral band to obtain the shape similarity component corresponding to the first spectral band. Furthermore, based on the smoothed first spectral band and the smoothed second spectral band, the information entropy weight corresponding to the first spectral band is determined. Also, based on the directional similarity component and shape similarity component corresponding to the first spectral band, the band two-dimensional similarity is determined. Further, the information entropy weights and band two-dimensional similarities corresponding to multiple first spectral bands are weighted and summed to obtain the two-dimensional similarity between the stockpiled tobacco leaves and the target tobacco leaves.
[0112] Figure 4 This is a schematic diagram of a tobacco leaf spectral data processing device provided in an embodiment of this disclosure. Figure 4 As shown, the tobacco leaf spectral data processing device includes: a spectral data acquisition module 410, a two-dimensional similarity determination module 420, and a tobacco leaf screening module 430. The spectral data acquisition module 410 is used to acquire first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf. The two-dimensional similarity determination module 420 is used to determine the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf, based on the first and second near-infrared spectral data of the tobacco leaf to be evaluated; wherein the two-dimensional similarity is determined by the directional similarity component and the shape similarity component between the first and second near-infrared spectral data. The tobacco leaf screening module 430 is used to determine similar tobacco leaves from the at least one tobacco leaf to be evaluated for use as a substitute for the target tobacco leaf, based on at least one two-dimensional similarity.
[0113] The technical solution of this disclosure provides a common data foundation for subsequent spectral feature extraction and shape similarity component calculation based on piecewise derivative smoothing by acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf. This ensures the consistency and comparability of the spectral comparison analysis between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, for at least one tobacco leaf to be evaluated, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the first and second near-infrared spectral data. This two-dimensional similarity is determined based on the directional and shape similarity components between the first and second near-infrared spectral data, achieving a multi-dimensional and accurate characterization of the spectral features of the tobacco leaves and significantly improving the accuracy and reliability of the matching determination between the tobacco leaf to be evaluated and the target tobacco leaf. Furthermore, by determining similar tobacco leaves from at least one tobacco leaf to be evaluated to replace the target tobacco leaf based on at least one two-dimensional similarity, the precision and efficiency of tobacco leaf quality matching are achieved, providing a scientific and quantifiable basis for the selection and substitution of tobacco raw materials. The technical solution of this disclosure solves the problems of high workload, strong subjectivity, low efficiency and difficulty in finding the best alternative tobacco leaves caused by the manual judgment method in related technologies. It realizes the determination of the two-dimensional similarity of tobacco leaves based on the directional similarity component and shape similarity component of the near-infrared spectral data of tobacco leaves, so as to quickly and accurately screen alternative tobacco leaves. It realizes the automation, standardization and precision of tobacco leaf similarity judgment and alternative screening, and greatly improves the efficiency and scientificity of tobacco raw material matching.
[0114] In some embodiments of this disclosure, optionally, the two-dimensional similarity determination module 420 includes: a spectral segmentation submodule, a directional similarity component determination submodule, a shape similarity component determination submodule, and a two-dimensional similarity determination submodule. The spectral segmentation submodule is used to perform spectral segmentation on the first near-infrared spectral data and the second near-infrared spectral data respectively, obtaining multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data; the directional similarity component determination submodule is used to determine the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data based on the multiple first spectral bands and the multiple second spectral bands; the shape similarity component determination submodule is used to determine the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data based on the multiple first spectral bands and the multiple second spectral bands; and the two-dimensional similarity determination submodule is used to determine the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the directional similarity component and the shape similarity component.
[0115] In some embodiments of this disclosure, optionally, the directional similarity component determination submodule includes: a spectral band smoothing unit, a correlation coefficient calculation unit, and a directional similarity component determination unit. Specifically, the spectral band smoothing unit is used to smooth multiple first spectral bands using a least-squares fitting algorithm, respectively, the first spectral bands and the corresponding second spectral bands, to obtain smoothed first spectral bands and smoothed second spectral bands; the correlation coefficient calculation unit is used to calculate the correlation coefficient between the smoothed first spectral band and the smoothed second spectral band using a correlation coefficient algorithm, as the directional similarity component of the band corresponding to the first spectral band; and the directional similarity component determination unit is used to determine the directional similarity component of the band corresponding to the multiple first spectral bands as the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0116] In some embodiments of this disclosure, optionally, the shape similarity component determination submodule includes: a spectral band differentiation and smoothing unit, a correlation coefficient calculation unit, and a shape similarity component determination unit. Specifically, the spectral band differentiation and smoothing unit is used to perform differentiation and smoothing processing on multiple first spectral bands using a least squares fitting algorithm and a cubic polynomial to obtain processed first spectral bands, and to perform differentiation and smoothing processing on second spectral bands corresponding to the first spectral bands using a least squares fitting algorithm and a cubic polynomial to obtain processed second spectral bands; the correlation coefficient calculation unit is used to calculate the correlation coefficient between the processed first spectral band and the processed second spectral band using a correlation coefficient algorithm, as the shape similarity component of the band corresponding to the first spectral band; and the shape similarity component determination unit is used to determine the shape similarity component of the band corresponding to the multiple first spectral bands as the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data.
[0117] In some embodiments of this disclosure, optionally, the spectral band derivative smoothing unit is specifically used to perform sliding traversal of the first spectral band using a sliding window of a first preset size; for each sliding window traversed, a first smoothed spectral response value corresponding to the sliding window is determined based on the spectral response values corresponding to multiple data points in the sliding window and a pre-determined weight coefficient sequence corresponding to the multiple data points; wherein, the weight coefficient sequence is determined according to the least squares fitting algorithm and a cubic polynomial; the weight coefficient sequence includes weight coefficients corresponding to each data point; and the processed first spectral band is determined based on the spectral response values corresponding to the multiple sliding windows.
[0118] In some embodiments of this disclosure, optionally, the step of determining the weight coefficient sequence corresponding to multiple data points includes: constructing multiple cubic polynomial equations based on multiple data points within a sliding window; constructing a design matrix corresponding to the sliding window based on the multiple polynomial equations; processing the design matrix using a least squares fitting algorithm to obtain a cubic polynomial model corresponding to the sliding window and a weight coefficient matrix corresponding to the sliding window; performing first-order derivative processing on the cubic polynomial model; and determining the weight coefficient sequence corresponding to the multiple data points within the sliding window based on the derivative result and the weight coefficient matrix.
[0119] In some embodiments of this disclosure, optionally, the directional similarity component includes band directional similarity components corresponding to multiple first spectral bands; the shape similarity component includes band shape similarity components corresponding to multiple first spectral bands; the two-dimensional similarity determination submodule includes: a band two-dimensional similarity determination unit, an information entropy weight determination unit, and a two-dimensional similarity determination unit. Specifically, the band two-dimensional similarity determination unit is used to determine the band two-dimensional similarity corresponding to multiple first spectral bands based on the band directional similarity components and band shape similarity components corresponding to the first spectral bands; the information entropy weight determination unit is used to determine the information entropy weight corresponding to the first spectral bands based on the information entropy corresponding to the multiple first spectral bands and the information entropy corresponding to the multiple second spectral bands; and the two-dimensional similarity determination unit is used to perform a weighted summation of the information entropy weights corresponding to the multiple first spectral bands and the band two-dimensional similarity to obtain the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf.
[0120] In some embodiments of this disclosure, optionally, the spectral segmentation submodule includes: a first peak start point determination unit, a first peak cutoff point determination unit, a second peak start point determination unit, a third peak cutoff point determination unit, and a first spectral band determination unit. The first peak start point determination unit is used to determine the first peak start point and the first peak end point corresponding to the first wavenumber based on a peak query function, a first wavenumber, and first near-infrared spectral data. The first peak cutoff point determination unit is used to average the first peak start points corresponding to at least one tobacco leaf to be evaluated to obtain the first peak cutoff point corresponding to the first wavenumber; and to average the first peak end points corresponding to at least one tobacco leaf to be evaluated to obtain the second peak cutoff point corresponding to the first wavenumber. The second peak start point determination unit is used to determine the second peak start point corresponding to the second wavenumber based on a peak query function, a second wavenumber, and first near-infrared spectral data. The system includes a second peak start point and a second peak end point corresponding to the second wave number; a third peak cutoff point determination unit, used to average the second peak start points corresponding to at least one tobacco leaf to be evaluated to obtain a third peak cutoff point corresponding to the second wave number; and to average the second peak end points corresponding to at least one tobacco leaf to be evaluated to obtain a fourth peak cutoff point corresponding to the second wave number; and a first spectral band determination unit, used to perform spectral segmentation on the first near-infrared spectral data according to the preset bands to be eliminated, the first peak cutoff point, the second peak cutoff point, the third peak cutoff point, and the fourth peak cutoff point to obtain multiple first spectral bands corresponding to the first near-infrared spectral data.
[0121] Optionally, in some embodiments of this disclosure, the spectral data acquisition module 410 includes a powder sample preparation unit and a spectral data determination unit. The powder sample preparation unit is used to prepare a sample of at least one type of tobacco leaf to be evaluated, dry the sample at a preset drying temperature and time, and grind the dried sample into powder to obtain a tobacco powder sample. The spectral data determination unit is used to perform infrared scanning on the tobacco powder sample using a Fourier transform near-infrared spectrometer, acquiring spectral data in a preset spectral band to obtain first near-infrared spectral data corresponding to the tobacco leaf to be evaluated.
[0122] The tobacco leaf spectral data processing device provided in this embodiment can execute the tobacco leaf spectral data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0123] It is worth noting that the various units and modules included in the above-mentioned tobacco leaf spectral data processing device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0124] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0125] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0126] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0127] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as tobacco leaf spectral data processing methods.
[0128] In some embodiments, the tobacco leaf spectral data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the tobacco leaf spectral data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the tobacco leaf spectral data processing method by any other suitable means (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] Computer programs used to implement the tobacco leaf spectral data processing method of this disclosure can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] This disclosure provides a computer-readable storage medium storing computer instructions for causing a processor to execute a tobacco leaf spectral data processing method, comprising: acquiring first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of a target tobacco leaf; determining a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the first near-infrared spectral data and the second near-infrared spectral data, for at least one of the tobacco leaves to be evaluated; wherein the two-dimensional similarity is determined by a directional similarity component and a shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data; and determining a similar tobacco leaf from the at least one tobacco leaf to be evaluated for replacing the target tobacco leaf based on the at least one two-dimensional similarity.
[0132] In the context of this disclosure, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0135] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0136] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of embodiments of this disclosure.
[0137] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the tobacco leaf spectral data processing method of any embodiment of this disclosure.
[0138] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0139] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for processing spectral data of tobacco leaves, characterized in that, include: Acquire first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of the target tobacco leaf; For at least one of the tobacco leaves to be evaluated, a two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated; wherein the two-dimensional similarity is determined based on the directional similarity component and the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data. Based on at least one of the two-dimensional similarities, similar tobacco leaves for replacing the target tobacco leaf are determined from at least one of the tobacco leaves to be evaluated.
2. The method for processing tobacco leaf spectral data according to claim 1, characterized in that, The step of determining the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated includes: The first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated are respectively segmented into multiple first spectral bands corresponding to the first near-infrared spectral data and multiple second spectral bands corresponding to the second near-infrared spectral data; Based on multiple first spectral bands and multiple second spectral bands, determine the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data; and Based on multiple first spectral bands and multiple second spectral bands, determine the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data; The two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf is determined based on the directional similarity component and the shape similarity component.
3. The method for processing tobacco leaf spectral data according to claim 2, characterized in that, The step of determining the directional similarity component between the first near-infrared spectral data and the second near-infrared spectral data based on multiple first spectral bands and multiple second spectral bands includes: For multiple first spectral bands, the least squares fitting algorithm is used to smooth the first spectral band and the second spectral band corresponding to the first spectral band respectively, to obtain smoothed first spectral bands and smoothed second spectral bands; The correlation coefficient algorithm is used to calculate the correlation coefficient between the smoothed first spectral band and the smoothed second spectral band, which is used as the band direction similarity component corresponding to the first spectral band. The directional similarity components corresponding to the multiple first spectral bands are determined as the directional similarity components between the first near-infrared spectral data and the second near-infrared spectral data.
4. The method for processing tobacco leaf spectral data according to claim 2, characterized in that, The step of determining the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data based on multiple first spectral bands and multiple second spectral bands includes: For multiple first spectral bands, the least squares fitting algorithm and cubic polynomial are used to perform derivative smoothing on the first spectral bands to obtain the processed first spectral bands. Also, the least squares fitting algorithm and cubic polynomial are used to perform derivative smoothing on the second spectral bands corresponding to the first spectral bands to obtain the processed second spectral bands. The correlation coefficient algorithm is used to calculate the correlation coefficient between the processed first spectral band and the processed second spectral band, which is used as the band shape similarity component corresponding to the first spectral band. The band shape similarity components corresponding to multiple first spectral bands are determined as the shape similarity components between the first near-infrared spectral data and the second near-infrared spectral data.
5. The method for processing tobacco leaf spectral data according to claim 4, characterized in that, The process of smoothing the first spectral band by employing a least-squares fitting algorithm and a cubic polynomial to obtain the processed first spectral band includes: The first spectral band is traversed by a sliding window of a first preset size; For each of the traversed sliding windows, a first smoothed spectral response value corresponding to the sliding window is determined based on the spectral response values corresponding to multiple data points in the sliding window and a pre-determined weight coefficient sequence corresponding to the multiple data points; wherein, the weight coefficient sequence is determined according to the least squares fitting algorithm and a cubic polynomial; the weight coefficient sequence includes weight coefficients corresponding to each data point; The first spectral band after processing is determined based on the spectral response values corresponding to the multiple sliding windows.
6. The method for processing tobacco leaf spectral data according to claim 5, characterized in that, The steps for determining the weight coefficient sequence corresponding to the multiple data points include: Based on multiple data points within the sliding window, construct multiple cubic polynomial equations, and construct a design matrix corresponding to the sliding window based on the multiple polynomial equations. The design matrix is processed using a least squares fitting algorithm to obtain a cubic polynomial model corresponding to the sliding window and a weight coefficient matrix corresponding to the sliding window. The first derivative of the cubic polynomial model is calculated, and the weight coefficient sequence corresponding to the multiple data points within the sliding window is determined based on the derivative result and the weight coefficient matrix.
7. The method for processing tobacco leaf spectral data according to claim 2, characterized in that, The directional similarity component includes band directional similarity components corresponding to multiple first spectral bands; the shape similarity component includes band shape similarity components corresponding to multiple first spectral bands. Determining the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf based on the directional similarity component and the shape similarity component includes: For multiple first spectral bands, a two-dimensional similarity of the bands corresponding to the first spectral bands is determined based on the band direction similarity component and the band shape similarity component. The information entropy weight corresponding to the first spectral band is determined based on the first information entropy corresponding to the plurality of first spectral bands and the second information entropy corresponding to the plurality of second spectral bands; The information entropy weights corresponding to multiple first spectral bands and the two-dimensional similarity of the bands are weighted and summed to obtain the two-dimensional similarity between the tobacco leaf to be evaluated and the target tobacco leaf.
8. The method for processing tobacco leaf spectral data according to claim 2, characterized in that, The method for determining the multiple first spectral bands corresponding to the first near-infrared spectral data includes: Based on the peak lookup function, the first wave number, and the first near-infrared spectral data, determine the first peak start point and the first peak end point corresponding to the first wave number; The average of the starting points of the first peak corresponding to at least one of the tobacco leaves to be evaluated is used to obtain the first peak intercept point corresponding to the first wave number; and the average of the ending points of the first peak corresponding to at least one of the tobacco leaves to be evaluated is used to obtain the second peak intercept point corresponding to the first wave number. Based on the peak lookup function, the second wavenumber, and the first near-infrared spectral data, determine the start point and end point of the second peak corresponding to the second wavenumber; The average of the starting points of the second peak corresponding to at least one of the tobacco leaves to be evaluated is used to obtain the cutoff point of the third peak corresponding to the second wave number; and the average of the ending points of the second peak corresponding to at least one of the tobacco leaves to be evaluated is used to obtain the cutoff point of the fourth peak corresponding to the second wave number. The first near-infrared spectral data is segmented according to the preset bands to be removed, the first peak intercept point, the second peak intercept point, the third peak intercept point, and the fourth peak intercept point to obtain multiple first spectral bands corresponding to the first near-infrared spectral data.
9. The method for processing tobacco leaf spectral data according to claim 1, characterized in that, The acquisition of first near-infrared spectral data of at least one tobacco leaf to be evaluated includes: For at least one type of tobacco leaf to be evaluated, a sample preparation is performed on the tobacco leaf to be evaluated to obtain a sample of tobacco leaf to be evaluated. The sample of tobacco leaf to be evaluated is dried at a preset drying temperature and a preset drying time, and the dried sample of tobacco leaf to be evaluated is ground into powder to obtain a tobacco leaf powder sample. The tobacco powder sample was scanned by a Fourier transform near-infrared spectrometer to collect spectral data in a preset spectral band, so as to obtain the first near-infrared spectral data corresponding to the tobacco leaf to be evaluated.
10. A tobacco leaf spectral data processing device, characterized in that, include: The spectral data acquisition module is used to acquire first near-infrared spectral data of at least one tobacco leaf to be evaluated and second near-infrared spectral data of the target tobacco leaf. A two-dimensional similarity determination module is used to determine the two-dimensional similarity between at least one of the tobacco leaves to be evaluated and the target tobacco leaf based on the first near-infrared spectral data and the second near-infrared spectral data of the tobacco leaf to be evaluated; wherein the two-dimensional similarity is determined by the directional similarity component and the shape similarity component between the first near-infrared spectral data and the second near-infrared spectral data; A tobacco leaf screening module is used to determine, based on at least one of the two-dimensional similarities, similar tobacco leaves to replace the target tobacco leaf from at least one of the tobacco leaves to be evaluated.