Method and system for identifying purity of honey based on spectral data

By dividing spectral data into independent chemical semantic sub-windows for principal component analysis and Mahalanobis distance fusion, the problem of high false negative rate of low concentration samples in honey adulteration detection is solved, and honey purity identification with high sensitivity and high accuracy is achieved.

CN122259507APending Publication Date: 2026-06-23SHAANXI HONEYCOMB ECOLOGICAL AGRICULTURAL TECHNOLOGY DEVELOPMENT CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI HONEYCOMB ECOLOGICAL AGRICULTURAL TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2026-05-28
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing technologies for detecting adulterated honey suffer from problems such as high false negative rates for low-concentration adulterated samples and poor resistance to interference, especially in complex testing conditions where it is difficult to achieve high sensitivity and high accuracy in identification.

Method used

By dividing spectral data into independent chemical semantic sub-windows for principal component analysis, and combining Mahalanobis distance and adaptive weight fusion, weak adulterant signals are eliminated from being diluted by full-spectrum background noise, thus improving anti-interference capability.

Benefits of technology

It improves the sensitivity and accuracy of identifying low-concentration adulterated honey, enhances the robustness of the system under complex detection conditions, and ensures the objectivity and reliability of the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122259507A_ABST
    Figure CN122259507A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of spectral analysis, and relates to a honey purity identification method and system based on spectral data. The method comprises the following steps: collecting an original spectrum of a to-be-tested honey sample and performing serial pretreatment to obtain a pretreated spectrum matrix; dividing the pretreated spectrum matrix into two independent chemical semantic sub-windows, independently performing principal component analysis to extract principal component scores; calculating Mahalanobis distances of the to-be-tested honey sample relative to a pure honey reference distribution in each principal component space; performing cross-window statistical standardization on the Mahalanobis distances of each window, and fusing based on a pre-determined adaptive weight to obtain a comprehensive spectral deviation value; comparing the comprehensive spectral deviation value with a discrimination threshold value, and outputting a honey purity identification result. The application decouples sugar component information and moisture information, avoids dilution of low-concentration adulteration signals in full-spectrum background noise, and improves the recognition accuracy of low-concentration adulterated honey.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of spectral analysis technology, specifically relating to a method and system for identifying the purity of honey based on spectral data. Background Technology

[0002] Honey purity identification is a crucial step in ensuring food safety and regulating market order, directly impacting consumers' health and the healthy development of the honey industry. In recent years, C-4 plant syrup adulteration has become the most common method of adulteration in the honey market due to its low cost and high similarity in appearance to pure honey. Low-concentration adulteration, due to its weak characteristic signals, has always been a technical challenge in industrial quality inspection and market supervision.

[0003] Traditional methods for detecting adulteration in honey include isotope ratio mass spectrometry (IRMS) and high-performance liquid chromatography (HPLC). While these methods offer high accuracy, they suffer from drawbacks such as long detection cycles, complex operations, expensive equipment, and the need for specialized personnel, failing to meet the demands of rapid quality control in industrial production lines and on-site market sampling. Near-infrared spectroscopy, with its advantages of being non-destructive, rapid, and requiring no sample pretreatment, is gradually becoming the mainstream technology in honey adulteration detection. Existing technologies are mostly based on a unified modeling approach across the entire spectrum, directly inputting full-spectrum data into a discriminant model for feature extraction and classification. However, this method forcibly mixes information from sugar-sensitive and moisture-sensitive bands, causing weak signals from low-concentration adulteration to be severely diluted by the full-spectrum background noise, ultimately resulting in a persistently high false negative rate for low-concentration adulterated samples.

[0004] To address this issue, Chinese patent application CN121347441A discloses a method for detecting adulterated honey. This method employs near-infrared spectroscopy, using the ratio of absorbance at 1450 nm (water molecule absorption peak) to 1490 nm (sugar molecule absorption peak) as the criterion. A portable detection device is designed to enable rapid on-site detection of adulterated honey. However, the technical solution disclosed in this patent application relies solely on the absorbance information from two isolated wavelengths for ratio calculation. It fails to fully explore and utilize the rich and complementary chemical component information contained in other bands of the spectrum. In complex industrial testing conditions, such as unavoidable instrument baseline drift, environmental and sample temperature fluctuations, and light scattering noise, relying solely on isolated band signals is highly susceptible to interference from occasional noise, making it difficult to ensure highly sensitive and stable identification of low-concentration adulterated samples with weak signal characteristics. Summary of the Invention

[0005] The purpose of this invention is to propose a method and system for identifying the purity of honey based on spectral data, in order to solve the technical problems of existing technologies being easily diluted by full-spectrum background noise when extracting weak adulterant signals, or having poor resistance to physical interference due to relying on only a very few isolated wavelengths, resulting in insufficient sensitivity and accuracy in identifying low-concentration adulterated honey under actual complex working conditions.

[0006] To address the above problems, the technical solution proposed in this invention for a honey purity identification method based on spectral data is as follows: Methods for identifying honey purity based on spectral data include: S1. Collect the original spectrum of the honey sample to be tested and preprocess it to obtain the preprocessed spectral matrix; S2. Divide the preprocessed spectral matrix into a first chemical semantic sub-window corresponding to the first preset band and a second chemical semantic sub-window corresponding to the second preset band. Perform principal component analysis independently on the first chemical semantic sub-window and the second chemical semantic sub-window to obtain the first principal component score vector and the second principal component score vector. S3. Based on the pre-established pure honey reference distribution, calculate the first Mahalanobis distance of the first principal component score vector corresponding to the first chemical semantic sub-window and the second Mahalanobis distance of the second principal component score vector corresponding to the second chemical semantic sub-window. S4. Perform cross-window statistical standardization on the first and second Mahalanobis distances respectively, and then perform weighted fusion with pre-determined adaptive weights to obtain the comprehensive spectral deviation value of the honey sample to be tested. S5. Compare the comprehensive spectral deviation value with the preset discrimination threshold, and output the purity identification result of the honey sample to be tested.

[0007] This invention decouples sugar and moisture information, originally mixed in the full-spectrum covariance matrix, into independent low-dimensional principal component spaces by dividing the preprocessed spectral matrix into two independent chemical semantic sub-windows and performing principal component analysis independently. This structural design allows weak local band variations caused by low-concentration adulteration to be concentratedly expressed in specific sub-spaces, thus avoiding the problem of weak adulteration signals being diluted by full-spectrum background noise. Furthermore, since the principal components extracted are within a sub-window with a certain bandwidth, the risk of single-point wavelengths being easily affected by local physical interference and directly causing judgment failure is avoided, improving the model's sensitivity, anti-interference ability, and accuracy in identifying low-concentration adulterated honey.

[0008] Furthermore, the process of acquiring the original spectrum of the honey sample to be tested and preprocessing it to obtain a preprocessed spectral matrix includes: performing standard normal variable transformation on the original spectrum to eliminate scattering baseline drift caused by physical factors; and performing convolution smoothing on the spectrum after standard normal variable transformation to obtain a preprocessed spectral matrix.

[0009] This invention employs a preprocessing sequence of standard normal variable transformation and convolutional smoothing, which eliminates scattering baseline drift caused by physical factors while preserving the detailed features of spectral peaks.

[0010] Further, the extraction process of the first principal component score vector and the second principal component score vector includes: calculating the cumulative variance contribution rate of each principal component in the first chemical semantic sub-window, selecting the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold, and extracting the first principal component score vector; calculating the cumulative variance contribution rate of each principal component in the second chemical semantic sub-window, selecting the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold, and extracting the second principal component score vector.

[0011] This invention selects the minimum number of principal components required for the cumulative variance contribution rate to first exceed a preset cumulative variance contribution rate threshold. This allows the retention of most chemically relevant variation information while discarding noisy principal components with extremely low contribution rates. This selection criterion avoids interference from noise dimensions on subsequent Mahalanobis distance calculations.

[0012] Furthermore, the calculation methods for the first Mahalanobis distance and the second Mahalanobis distance include: Obtain the first mean vector and first covariance matrix of the pure honey training sample group in the first chemical semantic sub-window, and calculate the first Mahalanobis distance by combining the first principal component score vector; obtain the second mean vector and second covariance matrix of the pure honey training sample group in the second chemical semantic sub-window, and calculate the second Mahalanobis distance by combining the second principal component score vector.

[0013] Furthermore, the cross-window statistical standardization processing of the first Mahalanobis distance and the second Mahalanobis distance includes: obtaining the mean and standard deviation of the first Mahalanobis distance of the pure honey training sample group in the first chemical semantic sub-window, and performing standardization processing on the first Mahalanobis distance using the mean and standard deviation of the first Mahalanobis distance; obtaining the mean and standard deviation of the second Mahalanobis distance of the pure honey training sample group in the second chemical semantic sub-window, and performing standardization processing on the second Mahalanobis distance using the mean and standard deviation of the second Mahalanobis distance.

[0014] Since the principal component dimensions of the two chemical semantic sub-windows differ, direct addition would lead to scale inconsistencies. This invention utilizes the Mahalanobis distance mean and Mahalanobis distance standard deviation of the pure honey training sample group to perform standardization processing, thereby giving the distance metric a unified basis.

[0015] Furthermore, the method for obtaining the comprehensive spectral deviation value of the honey sample to be tested includes: The mean of the standardized Mahalanobis distance of the known adulterated training sample group in the first chemical semantic sub-window is used as the first discriminant separation index, and the mean of the standardized Mahalanobis distance in the second chemical semantic sub-window is used as the second discriminant separation index. Based on the first and second discriminant separation indices, the first adaptive weight of the first chemical semantic sub-window and the second adaptive weight of the second chemical semantic sub-window are calculated respectively. The first and second Mahalanobis distances after standardization are weighted and summed using the first and second adaptive weights to obtain the comprehensive spectral deviation value.

[0016] This invention uses the standardized Mahalanobis distance mean of each chemical semantic sub-window to the known adulterated training sample group as the discrimination separation index to allocate adaptive weights, abandoning the practice of relying on manual experience to set parameters. This method uses the measured statistical separation of each window to the adulterated samples as the fusion contribution ratio, so that the chemical semantic space with richer discrimination information dominates in the fusion result, further amplifying the effective adulteration features.

[0017] Furthermore, the preset discrimination threshold is determined as follows: Calculate the mean and standard deviation of the comprehensive spectral deviation of the pure honey training sample group. Add the mean of the comprehensive spectral deviation to the standard deviation of the comprehensive spectral deviation by a preset multiple and use the result as the preset discrimination threshold.

[0018] Furthermore, the step of comparing the comprehensive spectral deviation value with a preset discrimination threshold and outputting the purity identification result of the honey sample to be tested includes: When the overall spectral deviation value is greater than the preset discrimination threshold, the honey sample to be tested is determined to be adulterated honey; when the overall spectral deviation value is less than or equal to the preset discrimination threshold, the honey sample to be tested is determined to be pure honey; the ratio of the overall spectral deviation value to the preset discrimination threshold is calculated, and the ratio is output as a relative deviation reference value.

[0019] Furthermore, the identification method also includes: monitoring the sample temperature of the honey sample in real time when collecting the original spectrum of the honey sample to be tested; when the sample temperature deviates from the preset calibration temperature range, triggering a temperature warning and pausing the output of purity identification results.

[0020] The technical solution of the honey purity identification system based on spectral data proposed in this invention is as follows: A honey purity identification system based on spectral data includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, the honey purity identification method based on spectral data described in any of the above technical solutions is implemented.

[0021] Compared with the prior art, the present invention has at least the following beneficial effects: On the one hand, this invention effectively avoids the blind dilution of weak, low-concentration adulterant signals by full-spectrum high-frequency background noise through a decoupling mechanism of physical semantic windowing and independent dimensionality reduction mapping. Simultaneously, compared to traditional methods that rely solely on a few isolated wavelength points, this invention utilizes rich and multidimensional chemically complementary information within specific local bands, enhancing the system's robustness against instrument baseline drift, environmental temperature fluctuations, and complex impurity interference, thereby improving the sensitivity and accuracy of identifying low-concentration adulterated honey under complex detection conditions. On the other hand, this invention introduces a cross-window statistical standardization and adaptive dynamic weight fusion mechanism. The discrimination process is based on objective statistical laws, endowing the model with the ability to adaptively resist occasional noise interference in local frequency bands, further ensuring the objectivity and reliability of the identification results. Attached Figure Description

[0022] Figure 1 This is a flowchart of the honey purity identification method based on spectral data of the present invention; Figure 2 This is a schematic diagram of the hardware structure of the honey purity identification system based on spectral data according to the present invention. Detailed Implementation

[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0024] Specific embodiments of the honey purity identification method based on spectral data proposed in this invention: like Figure 1 As shown, the method for identifying honey purity based on spectral data includes the following steps: S1. Spectral acquisition and preprocessing.

[0025] For the honey samples to be tested, the raw spectra were first acquired using a Fourier transform near-infrared spectrometer. To ensure optical path stability and repeatability of test results, a liquid fiber optic transmissive / reflective probe was used for contact measurement during the spectral acquisition stage. Specific instrument hardware scanning parameters were set as follows: scanning range of 12,000 to 4,000 wavenumbers, optical path length of 2 mm, 32 superimposed scans per test, and spectral resolution of 8 wavenumbers. Before performing batch acquisition, a temperature baseline correction operation was performed according to the instrument's factory-preset calibration coefficients to eliminate systematic baseline shifts in the spectrum caused by ambient temperature fluctuations. This combination of instrument hardware parameters and calibration procedures ensures a reasonable scan time while maintaining a high signal-to-noise ratio, thus laying a solid hardware data foundation for rapid testing on the production line.

[0026] After the raw spectra are acquired, a two-step cascaded preprocessing operation is performed in a specific order. The first step is to perform a standard normal variable transformation on the raw spectra. The purpose of this step is to eliminate scattering baseline drift caused by physical factors such as differences in the size of microscopic particles within the honey sample and inconsistent light transmission paths. The standard normal variable transformation is performed independently for each raw spectrum. The logic of this transformation is to eliminate additive baseline shift and multiplicative optical path interference by centering and scaling the statistical characteristics of the entire spectrum. Specifically, the absorbance value corresponding to each wavenumber point on the raw spectrum is subtracted from the arithmetic mean of the absorbance of that spectrum across the entire wavelength range, and then divided by the standard deviation of the absorbance of that spectrum across the entire wavelength range. After the above mathematical transformation, the overall absorbance mean of each spectrum is converted to 0, and the corresponding standard deviation is uniformly normalized to 1, thereby establishing an absorbance benchmark that is at an absolutely uniform scale and statistically comparable among different batches of honey samples.

[0027] After completing the standard normal variable transformation, the second step is to perform convolutional smoothing on the transformed spectrum. The Savitzky-Gorye polynomial convolutional smoothing algorithm is used here. To balance noise reduction and signal fidelity, the specific smoothing parameters are set to a sliding window of 7 points and a fitting polynomial order of 3. The underlying logic of convolutional smoothing is to utilize the mathematical properties of local polynomial least squares fitting to perform a weighted average of adjacent data points within the sliding window. This processing method can effectively suppress high-frequency random instrument noise in the spectral signal while preserving the width and height of the original honey spectral peaks to the greatest extent possible. This specific parameter configuration avoids over-smoothing caused by an excessively large window or an excessively low order, preventing physical shifts in the peak positions or geometric distortions in the peak shapes of key absorption peaks.

[0028] The two preprocessing steps described above must be executed in a predetermined order and cannot be reversed. If convolutional smoothing is performed first in the actual algorithm code, followed by standard normal transformation, the small changes in the local band mean introduced during the smoothing operation will directly interfere with the calculation benchmarks of the mean and standard deviation of the full-spectrum statistics during the standard normal transformation stage. This will lead to unpredictable systematic numerical biases in the standardized results of the entire spectrum. The correct execution logic must be: after inputting the original spectrum, perform standard normal transformation first, followed by convolutional smoothing, to finally obtain the preprocessed spectral matrix.

[0029] In the data structure of the preprocessed spectral matrix, the row dimension corresponds to each independent honey sample to be tested, and the column dimension corresponds to each wavenumber point collected. The preprocessed spectral matrix will be used as the input source for the subsequent chemical semantic windowing processing.

[0030] S2, Chemical Semantic Windowing and Independent Principal Component Mapping.

[0031] Existing near-infrared spectroscopy identification methods typically treat the entire spectrum, containing thousands of wavenumber points, as an indivisible high-dimensional feature vector, directly inputting it into the discrimination model. The drawback of this unified full-spectrum modeling approach is that it forcibly mixes and compresses spectral information reflecting different chemical components into a single covariance matrix. Specifically, the sugar component-sensitive bands are treated indiscriminately with the water and hydroxyl-sensitive bands in the full-spectrum data. When a honey sample is adulterated with extremely low concentrations of plant syrup, this trace adulteration only causes a weak absorbance change in a very localized spectral band. If this weak local change signal is directly superimposed on the full-spectrum background noise composed of hundreds or even thousands of irrelevant wavenumber points, the weak change signal will be completely diluted during the eigenvalue decomposition of the full-spectrum covariance matrix. This causes the model to fail to effectively capture these key adulteration features in the leading principal components, ultimately resulting in a precipitous drop in the model's accuracy in identifying low-concentration adulterated honey.

[0032] To address the issue of weak signals being diluted by full-spectrum noise, this invention, based on the absorption mechanism of the main chemical components of honey in the near-infrared band, divides the preprocessed spectral matrix output from the preceding steps into two chemically independent sub-windows.

[0033] First, the preprocessed spectral matrix is ​​divided into first chemical semantic sub-windows corresponding to the first preset wavelength band. The range of the first preset wavelength band is set to 5800 to 6200 wavenumbers. This spectral range is the first overtone absorption region of carbon-hydrogen bonds in sugar molecules such as fructose and glucose. This region has extremely high physical sensitivity to differences in the types and relative proportions of sugar components. The relative proportions of the overtone absorption peak positions and peak intensities of fructose and glucose in natural pure honey show significant and clearly distinguishable differences from the corresponding characteristics of common plant syrups. Therefore, the first preset wavelength band is the core region for identifying plant syrup incorporation. Typically, the center of the first overtone absorption of carbon-hydrogen bonds is located around 5900 wavenumbers. This invention uses this center as a reference and reserves a buffer bandwidth of approximately 200 wavenumbers on both sides. This setting not only fully covers the characteristic peak position differences of different sugar molecules in this overtone region, but also avoids interference from the carbon-hydrogen bond combination frequency region below 5600 wavenumbers, where the signal-to-noise ratio is low.

[0034] Secondly, the preprocessed spectral matrix is ​​divided into second chemical semantic sub-windows corresponding to the second preset wavelength band. The range of the second preset wavelength band is set to 4800 to 5200 wavenumbers. This spectral range is the combination absorption region of oxygen-hydrogen bonds, which can comprehensively reflect the absolute water content and complex molecular environment information of hydroxyl compounds in the honey system. When plant syrup is added to honey, it will disrupt and reshape the hydrogen bond network structure of hydroxyl groups in the original honey system, thereby causing changes in absorption characteristics that can be stably detected by the instrument in this specific wavelength band. The second preset wavelength band and the first preset wavelength band complement each other in terms of chemical information dimension. The selection of the second preset wavelength band is also based on the objectively known absorption center position, that is, the oxygen-hydrogen bond combination absorption center is usually located around 5000 wavenumbers. A buffer bandwidth of about 200 wavenumbers is reserved on both sides of this center, which aims to cover the shift in the combination peak position caused by subtle changes in the hydrogen bond network structure, and to avoid including the carbon-hydrogen-oxygen combination frequency region below 4600 wavenumbers, which is strongly affected by water absorption, in the calculation range. In actual industrial applications, operators can adaptively fine-tune the sub-window boundaries within a range of ±100 wavenumbers from the absorption center, based on the measured difference spectrum of pure honey and adulterated honey from the spectrometer used, to match the spectral resolution and signal-to-noise ratio hardware characteristics of different instruments.

[0035] After segmenting the physical bands, principal component analysis was performed independently on the first and second chemical semantic sub-windows. Within each sub-window, covariance matrix calculation and eigenvalue decomposition were performed on the preprocessed spectral data to obtain a series of principal components arranged in descending order of eigenvalues. To determine the number of principal components to retain, the cumulative variance contribution rate of each principal component in each chemical semantic sub-window was calculated. The cumulative variance contribution rate quantifies the proportion of the original spectral data variation explained by the top few principal components relative to the total variation. The specific calculation formula is as follows:

[0036] In this formula, Indicates the preceding The cumulative variance contribution rate of each principal component This represents the eigenvalue obtained by performing eigenvalue decomposition on the covariance matrix within a specific chemical semantic sub-window. The eigenvalues ​​are arranged in descending order, and the magnitude of these eigenvalues ​​directly represents the eigenvalues. The amount of effective spectral variance information contained in each principal component; This indicates the total number of wavenumber points contained in the chemical semantic sub-window, i.e., the total input dimension of the local analysis space; This indicates the number of principal components that are currently being extracted when calculating the cumulative contribution rate.

[0037] After calculating the cumulative variance contribution rate of each principal component, the system will The values ​​are compared one by one with a preset cumulative variance contribution rate threshold. In this standard procedure, the preset cumulative variance contribution rate threshold is set to 95%. For the first chemical semantic sub-window, the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold are selected, and the first principal component score vector is extracted based on these retained principal components. Similarly, for the second chemical semantic sub-window, the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold are selected, and the second principal component score vector is extracted accordingly.

[0038] The physical meaning of this selection criterion is to ensure that the minimum number of principal components selected can independently explain more than 95% of the spectral variation information of the corresponding sub-window. This not only fully preserves most of the effective variation information closely related to the chemical composition of the sample, but also discards the noisy principal components corresponding to the smallest eigenvalues, avoiding irreversible mathematical interference from the instrument noise dimension on the subsequent calculation of Mahalanobis distance.

[0039] Through the aforementioned chemical semantic windowing strategy and independent principal component mapping process, this invention physically decouples the sugar component information and moisture information mixed in the full-spectrum covariance matrix. This operation reduces the dimensionality of high-dimensional, strongly coupled complex spectral data into two physically meaningful and completely independent low-dimensional principal component spaces. This allows the subtle local changes caused by extremely low concentrations of plant syrup adulteration in specific spectral absorption regions to be highly concentrated numerically expressed in the low-dimensional principal component space corresponding to the first chemical semantic sub-window, thereby blocking the path of key subtle adulteration signals being diluted by full-spectrum high-frequency background noise.

[0040] S3, Calculation of Mahalanobis distance in dual-window mode.

[0041] The Mahalanobis distance calculation is based on a pre-constructed reference distribution of pure honey. The system first acquires a training sample group of pure honey. In the principal component space corresponding to the first chemical semantic sub-window, the first mean vector and the first covariance matrix of this training sample group are calculated. Using the same calculation logic, the second mean vector and the second covariance matrix of the training sample group are calculated in the principal component space corresponding to the second chemical semantic sub-window. Within these two independent spaces, the first and second mean vectors describe the central location of the pure honey sample group at a multi-dimensional geometric level. The first and second covariance matrices reflect the statistical dispersion of normal pure honey samples along various principal component directions and the potential correlation structure between different principal component features.

[0042] For any input honey sample to be tested, the system extracts its preprocessed spectral data within the first and second chemical semantic sub-windows, respectively. These preprocessed data are then projected onto their corresponding principal component loading matrices for linear combination operations, thereby obtaining the first principal component score vector in the first chemical semantic sub-window and the second principal component score vector in the second chemical semantic sub-window for the honey sample to be tested.

[0043] Subsequently, the system calculates the first Mahalanobis distance by combining the first principal component score vector, the first mean vector, and the first covariance matrix, and calculates the second Mahalanobis distance by combining the second principal component score vector, the second mean vector, and the second covariance matrix. The general mathematical formula for calculating these two Mahalanobis distances is as follows:

[0044] In this formula, This indicates that the honey sample to be tested was in the [number]th [year]. The Mahalanobis distance corresponding to each chemical semantic sub-window, where the parameters are... The value can be 1 or 2, corresponding to the calculation scenarios of the first and second Mahalanobis distances, respectively. (Subscript) The distance is calculated using Mahalanobis distance. This indicates that the honey sample to be tested was in the [number]th [year]. The principal component score vectors of each chemical semantic sub-window represent the specific coordinates of the sample to be tested in space. This indicates that the training sample group of pure honey is in the th... The mean vector of each chemical semantic sub-window. The corresponding covariance matrix is... Its inverse matrix, superscript This represents the transpose of the matrix. The formula mathematically normalizes the physical scale along each principal component direction by introducing the inverse of the covariance matrix into the quadratic form structure. This metric effectively utilizes the internal variance characteristics of pure honey samples to offset the imbalance in discrimination weights caused by differences in the absolute magnitude of variance between different principal components. It also eliminates collinearity interference between features of different principal components, making the final distance metric more objective and reasonable in a multidimensional statistical sense.

[0045] A mathematical prerequisite for performing the Mahalanobis distance calculation is that the covariance matrix must be full-rank and completely invertible. To satisfy this necessary condition, the absolute number of pure honey training samples used to calculate the mean vector and covariance matrix must be greater than the principal component dimension of the corresponding chemical semantic sub-window. In actual near-infrared spectroscopy measurement environments, the objective existence of instrument background electrical noise and physical scattering of particles within the honey sample means that the pure honey training samples inevitably have non-zero statistical dispersion in each principal component direction. As long as the number of training samples is sufficient, the covariance matrix will naturally exhibit a full-rank state, eliminating the need to forcibly introduce artificial mathematical compensation constants at the algorithm's underlying level. To ensure the stability and reliability of the identification model, the system will enforce an automatic verification procedure before solidifying the underlying model parameters, comparing the number of pure honey training samples with the principal component dimension. If it is found that the number of samples fails to meet the basic condition of being greater than the principal component dimension, the system will proactively refuse to complete the model solidification process at the physical code level and issue a clear warning to the on-site operators, forcibly requiring the addition of a sufficient number of pure honey training samples with diverse coverage before restarting model training.

[0046] By independently deriving and calculating Mahalanobis distance in their respective chemical semantic subspaces, the system can ensure that the local data deviation caused by the incorporation of extremely low concentrations of plant syrup in a specific spectral space is completely independent and the numerical value is greatly amplified.

[0047] S4, Cross-window statistical standardization and adaptive weight fusion.

[0048] Because the principal component space dimensions corresponding to the first and second chemical semantic sub-windows are usually not equal, the directly calculated first and second Mahalanobis distances exhibit objective systematic differences in their statistical distribution scales. Directly adding these two distance values, which are at different statistical levels, will inevitably lead to the high-dimensional data irrationally dominating the final judgment result. To eliminate this statistical incomparability across spatial dimensions, cross-window statistical standardization is performed on the first and second Mahalanobis distances. Specifically, the mean and standard deviation of the first Mahalanobis distance for the pure honey training sample group in the first chemical semantic sub-window are obtained, as are the mean and standard deviation of the second Mahalanobis distance for the pure honey training sample group in the second chemical semantic sub-window. Then, the Mahalanobis distances for each chemical semantic sub-window are standardized using the standard score transformation formula, as follows:

[0049] In this formula, This represents the standardized Mahalanobis distance, when the parameter... When the value is 1 or 2, it corresponds to the first Mahalanobis distance and the second Mahalanobis distance after standardization, respectively. This represents the original distance actually calculated for the honey sample to be tested within the corresponding chemical semantic sub-window, i.e., the first Mahalanobis distance or the second Mahalanobis distance. This represents the mean Mahalanobis distance of the pure honey training sample group within the corresponding chemical semantic sub-window. This represents the standard deviation of the Mahalanobis distance of the pure honey training sample group within the corresponding chemical semantic sub-window.

[0050] After mathematical transformation, the standardized first and second Mahalanobis distances are uniformly translated into the relative deviation of the tested honey sample from the center of the pure honey population. For pure honey samples, the expected value of this standardized value will approach 0 infinitely. However, for abnormal samples adulterated with plant syrup, the expected value of this standardized value will be significantly greater than 0, and the magnitude of the deviation is strictly positively correlated with the actual amount of adulteration. To ensure the smooth execution of the algorithm, the system automatically verifies whether the standard deviation of the Mahalanobis distance in each window is greater than 0 before solidifying the underlying model parameters. Only when the pure honey sample group used for training covers a sufficiently rich variety of honey sources, geographical distribution, and collection batches, and the differences in natural chemical composition between samples are objectively reflected, will this standard deviation be greater than 0. If an extreme anomaly of a standard deviation equal to 0 occurs during the calculation, the system will directly determine that the diversity of the current training set samples is insufficient, refuse to complete the subsequent model solidification process, and force the operator to supplement pure honey training samples covering a wider range of variations before restarting the training.

[0051] After eliminating data scale differences between different windows, the system performs a weighted fusion of the processed first Mahalanobis distance and the processed second Mahalanobis distance using pre-determined adaptive weights. Specifically, the process is as follows: First, the system obtains the mean standardized Mahalanobis distance calculated in the first chemical semantic sub-window for the known adulterated training sample group, and directly defines it as the first discriminant separation index. Similarly, the system obtains the mean standardized Mahalanobis distance calculated in the second chemical semantic sub-window for the same known adulterated training sample group, and defines it as the second discriminant separation index. The formula for calculating the discriminant separation index is as follows:

[0052] In the above formula, This represents the discrimination and resolution index corresponding to the chemical semantic sub-window. This represents the overall arithmetic mean of the standardized Mahalanobis distances of the known adulterated training sample group within this specific window space. A larger value for this discrimination and separation index indicates that the spectral features extracted from this specific spectral band window can further separate the adulterated sample group from the normal pure honey sample group; in other words, the window's contribution to the measured discrimination of the physical behavior of plant syrup adulteration is stronger. Based on the numerical comparison of the first and second discrimination and separation indices, the system calculates the first adaptive weight of the first chemical semantic sub-window and the second adaptive weight of the second chemical semantic sub-window, respectively. The formula for allocating the adaptive weights is as follows:

[0053] In this formula, This represents the adaptive weights of the corresponding chemical semantic sub-windows, calculated from the data. This represents the first discriminant separation index. This represents the second discriminant separation index. The formula calculates the relative percentage of each window's discriminant separation index to the total separation, dynamically and significantly tilting the final discriminant weights towards the chemical semantic sub-window with the stronger measured physical separation capability. As another core engineering safeguard mechanism, before formally calculating the adaptive weight values, the system first verifies whether the sum of the first and second discriminant separation indices is strictly greater than 0. If the sum equals 0, it means that the currently collected training sample set has lost its spatial separation capability for any adulterated samples. In this case, the system determines that the training set completely fails to meet the basic data specifications and outputs a severe warning signal regarding training set quality. This mechanism provides a fallback mechanism at the algorithm's physical level, eliminating the need to manually introduce any mathematical compensation constants lacking physical basis.

[0054] After obtaining the adaptive weight allocation scheme, the system uses the first adaptive weight to perform a multiplicative weighted sum on the standardized first Mahalanobis distance, and simultaneously uses the second adaptive weight to perform a multiplicative weighted sum on the standardized second Mahalanobis distance. These two weighted values ​​are then summed to obtain a comprehensive spectral deviation value that fully reflects the true state of the honey sample being tested. The formula for this fusion calculation is as follows:

[0055] In the formula This represents the final calculated composite spectral deviation value. Indicates the first adaptive weight. This represents the second adaptive weight. This represents the first Mahalanobis distance after standardization. This represents the second Mahalanobis distance after standardization. This cross-window collaborative dynamic weighting fusion algorithm design enables the weak, low-concentration plant syrup adulteration signal captured in a specific chemical semantic subspace to occupy an absolute dominant position in the final fusion discriminant value due to its extremely high dynamic adaptive weight, thus blocking the path of weak feature signals being blindly diluted by mean in the fusion of full-spectrum background data.

[0056] S5. Output of purity determination results.

[0057] The final step involves comparing the calculated comprehensive spectral deviation value with the system's preset discrimination threshold to output the final purity identification result. Regarding the determination of the preset discrimination threshold, during the initial offline model training phase, the system comprehensively calculates the mean and standard deviation of the comprehensive spectral deviation values ​​for the pure honey training sample group. Subsequently, the system sets the mean comprehensive spectral deviation value plus a preset multiple of the standard deviation as the preset discrimination threshold for that batch of models.

[0058] The selection of the preset multiplier is strongly constrained by both statistical distribution laws and on-site quality inspection logic. Typically, the multiplier is set to three standard deviations. This value, under the physical assumption of a normal distribution, covers the vast majority of normal, pure honey samples, thus achieving a balance between reducing the false positive rate of normal batches of pure honey and ensuring the detection rate of low-concentration, extremely difficult-to-adulterate samples. During the identification and testing process on the actual production line, when the system calculates through forward inference that the comprehensive spectral deviation value of the honey sample to be tested is greater than the preset discrimination threshold, the honey sample is immediately determined to be adulterated honey. When the comprehensive spectral deviation value is less than or equal to the preset discrimination threshold, the honey sample is determined to be pure honey.

[0059] After calculating the comprehensive spectral deviation value of the honey sample to be tested, the system not only logically compares this value with a preset discrimination threshold to output a purity identification result—that is, directly determining whether the honey sample is pure honey or adulterated honey—for rapid sorting of qualified and unqualified honey at the enterprise's front-end quality inspection line, but also simultaneously calculates and outputs a second level of information with more quantitative reference value. Specifically, the system calculates the ratio of the comprehensive spectral deviation value to the preset discrimination threshold and outputs this calculated ratio as a relative deviation reference value. The formula for calculating the relative deviation reference value is as follows:

[0060] In this formula, This represents the relative deviation reference value that the system ultimately outputs to the enterprise terminal. This represents the comprehensive spectral deviation value of the honey sample to be tested, calculated in the preceding computational steps. This represents the preset discrimination threshold determined by the system during the offline model training phase. This formula establishes a one-dimensional scale with the preset discrimination threshold as the absolute reference. When the honey sample being tested is normal pure honey, its comprehensive spectral deviation value will inevitably be within the reasonable fluctuation range of natural variation. Therefore, the calculated relative deviation reference value is usually less than or equal to 1. Conversely, when the honey sample being tested is adulterated with abnormal impurities such as plant syrup, this relative deviation reference value will inevitably be greater than 1. More importantly, the magnitude of this ratio quantifies the severity of the honey sample's deviation from the geometric center of normal pure honey; a larger ratio value means a relatively higher concentration of adulterants within the sample.

[0061] This invention effectively avoids the dilution of weak, low-concentration adulterant signals by mean-averaging in full-spectrum background noise by decoupling spectral features into independent chemical semantic sub-windows for dimensionality reduction analysis. This improves the model's sensitivity and accuracy in identifying low-concentration adulterated honey. Furthermore, this invention combines cross-window statistical standardization with an adaptive dynamic weight fusion mechanism, making the discrimination process more consistent with objective physical and statistical laws, and enhancing the system's resistance to instrument drift and impurity interference. In addition, the accompanying real-time sample temperature monitoring and anomaly interception mechanism further ensures the rigor of the testing conditions, and the entire testing process requires no complex sample pretreatment or additional expensive hardware costs.

[0062] An embodiment of the honey purity identification system based on spectral data provided by this invention: At the hardware physical level, such as Figure 2 As shown, the honey purity identification system based on spectral data includes: a sample container for loading the honey to be tested, a temperature sensor for real-time monitoring of the sample temperature, an optical fiber probe for non-destructive acquisition of contact signals, an FT-NIR spectrometer, and a processing unit.

[0063] The FT-NIR spectrometer is a Fourier transform near-infrared spectrometer; the processing unit is an embedded data processing unit containing a large-capacity memory and a high-performance processor. The memory internally stores and stores all core underlying mathematical model parameters generated offline from a large standard training set. This pre-stored parameter list comprehensively covers the standard normal variable transformation statistics used to eliminate the original spectral baseline drift, the convolution kernel parameters used for high-frequency spectral noise reduction, the principal component loading matrices corresponding to the first and second chemical semantic sub-windows, the first and second mean vectors of the pure honey training sample group in their respective principal component spaces, and the corresponding inverse matrices of the first and second covariance matrices, the mean and standard deviation standardization parameters for the first and second Mahalanobis distances used for cross-dimensional unified scale, the first and second adaptive weights determined based on measured discriminant separation, and the preset discrimination threshold used as the final redline discrimination benchmark. When the entire testing equipment is deployed to the actual testing site, the processor only needs to call these fixed physical parameters in memory and perform a one-way forward inference calculation task based on the original spectrum of the honey sample to be tested that is currently collected in real time.

[0064] Furthermore, because near-infrared spectral signals are inherently extremely sensitive to the molecular thermal motion and physical state within the test system, even minute temperature fluctuations can cause shifts in the absorption peak positions. To prevent such environmental interference, the system must continuously monitor the current sample temperature in real time using an integrated temperature sensor when acquiring the raw spectrum of the honey sample. The system processor has a built-in set of pre-emptive environmental parameter hard verification logic in its underlying code. When the real-time monitored sample temperature deviates from the system's preset calibration temperature range (typically 20±2℃), the processor will immediately trigger an audible and visual temperature warning. Within the same millisecond-level timeframe of triggering the warning, the processor will block all subsequent downstream spectral data flow at the software level and suspend the output of any purity identification results. At this time, the device's interactive control screen will clearly prompt the inspector to place the abnormally warm sample back into a constant-temperature environment for thermal equilibrium treatment until the sensor confirms that the sample's internal temperature has completely returned to within the preset calibration temperature range before the inspector is allowed to restart the complete spectral acquisition and detection process.

[0065] On-site quality control operators do not need complex academic backgrounds in modern analytical chemistry or advanced data statistics. They only need to follow the standard to perform extremely simple sample physical loading and one-click start instructions to obtain accurate honey purity identification results in a very short response time.

[0066] To ensure the robust and reliable discrimination performance of the aforementioned algorithm model in actual industrial quality inspection scenarios, the training set used to offline construct the underlying benchmark parameters is built according to specific physical and statistical specifications. Specifically, the physical collection of the pure honey training sample group must cover representative different plant nectar sources, such as common varieties like acacia honey, jujube honey, and locust honey, while also spanning different geographical origins and seasonal collection batches in both spatial and temporal dimensions. To meet the high-dimensional spatial computation requirements of the underlying statistics, the absolute number of pure honey samples used for model calibration must be no less than 50. The core purpose of setting this minimum number and sample diversity requirement is to ensure that the first and second covariance matrices extracted in the preceding operations, i.e., in the first... Each chemical semantic sub-window is uniformly represented as The covariance matrix parameters can accurately describe the multidimensional statistical dispersion structure of normal pure honey populations in their respective independent chemical semantic spaces, avoiding algorithm failures such as singular states of the covariance matrix or excessively large matrix condition numbers caused by overly singular features of training samples.

[0067] After establishing the pure honey benchmark, the construction of the known adulteration training sample group also needs to follow the concentration gradient design logic. The adulterants used in the adulteration samples must comprehensively cover the C-4 plant syrups most commonly used in industrial adulteration, such as high-fructose corn syrup and rice syrup. In terms of concentration, the mass fraction of artificially added plant syrup must completely cover the extremely broad range of 5% to 30%. Within this overall range, there is a mandatory structural constraint: the number of samples in the low adulteration ratio range (5% to 10%) must account for more than 40% of the total number of all known adulteration training samples in that batch. This non-uniform concentration distribution design directly serves the accurate calculation of the discrimination resolution index, ensuring that the final calculated... Numerical values ​​accurately reflect the model's actual physical separation and purification capabilities when faced with the most subtle adulteration ranges. This dataset structure design avoids the risk of statistical misjudgment; that is, if the training set is filled with highly concentrated adulterated samples with extremely obvious characteristics, the large deviations generated by these high-concentration samples will severely dominate and inflate the statistical accuracy. The overall statistical mean is used to cause the system to seriously overestimate the actual discrimination efficiency of a certain sub-window for the extremely low concentration range when allocating weights.

[0068] To ensure the physical accuracy of the aforementioned fine concentration gradient, the laboratory preparation process of the adulterated training samples strictly adhered to standard gravimetric procedures. Personnel physically mixed a known absolute mass of C-4 plant syrup, precisely weighed using a balanced scale, with a similarly precisely weighed known absolute mass of standard pure honey, according to a pre-set target mass fraction adulteration ratio. After the high-viscosity fluids were thoroughly and uniformly mixed using mechanical equipment, the sample container containing the mixture was left to stand under controlled conditions for at least two hours before being transferred to a spectrometer for near-infrared spectral acquisition. This standing time threshold was set to allow sufficient time for the syrup and honey mixture, which involves complex intermolecular forces, to reach thermodynamic equilibrium, avoiding repeatability errors in spectral data due to inhomogeneous mixing.

[0069] Furthermore, the memory stores computer program instructions, which, when executed by the processor, enable the honey purity identification method based on spectral data as described in the above embodiments. The honey purity identification system based on spectral data also includes other components well-known to those skilled in the art, such as a communication bus and communication interface; their configuration and functions are known in the art and will not be described further here.

[0070] While various embodiments of the invention have been shown and described in this specification, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention.

Claims

1. A method for identifying the purity of honey based on spectral data, characterized in that, include: S1. Collect the original spectrum of the honey sample to be tested and preprocess it to obtain the preprocessed spectral matrix; S2. Divide the preprocessed spectral matrix into a first chemical semantic sub-window corresponding to the first preset band and a second chemical semantic sub-window corresponding to the second preset band. Perform principal component analysis independently on the first chemical semantic sub-window and the second chemical semantic sub-window to obtain the first principal component score vector and the second principal component score vector. S3. Based on the pre-established pure honey reference distribution, calculate the first Mahalanobis distance of the first principal component score vector corresponding to the first chemical semantic sub-window and the second Mahalanobis distance of the second principal component score vector corresponding to the second chemical semantic sub-window. S4. Perform cross-window statistical standardization on the first and second Mahalanobis distances respectively, and then perform weighted fusion with pre-determined adaptive weights to obtain the comprehensive spectral deviation value of the honey sample to be tested. S5. Compare the comprehensive spectral deviation value with the preset discrimination threshold, and output the purity identification result of the honey sample to be tested.

2. The method for identifying honey purity based on spectral data according to claim 1, characterized in that, The process of collecting the original spectrum of the honey sample to be tested and preprocessing it to obtain a preprocessed spectral matrix includes: performing standard normal variable transformation on the original spectrum to eliminate scattering baseline drift caused by physical factors; and performing convolution smoothing on the spectrum after standard normal variable transformation to obtain the preprocessed spectral matrix.

3. The method for identifying honey purity based on spectral data according to claim 1, characterized in that, The extraction process of the first principal component score vector and the second principal component score vector includes: Calculate the cumulative variance contribution rate of each principal component in the first chemical semantic sub-window, select the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold, and extract the first principal component score vector; calculate the cumulative variance contribution rate of each principal component in the second chemical semantic sub-window, select the minimum number of principal components required for the cumulative variance contribution rate to first exceed the preset cumulative variance contribution rate threshold, and extract the second principal component score vector.

4. The method for identifying honey purity based on spectral data according to claim 1, characterized in that, The methods for calculating the first Mahalanobis distance and the second Mahalanobis distance include: Obtain the first mean vector and first covariance matrix of the pure honey training sample group in the first chemical semantic sub-window, and calculate the first Mahalanobis distance by combining the first principal component score vector; obtain the second mean vector and second covariance matrix of the pure honey training sample group in the second chemical semantic sub-window, and calculate the second Mahalanobis distance by combining the second principal component score vector.

5. The method for identifying honey purity based on spectral data according to claim 4, characterized in that, The cross-window statistical standardization processing of the first Mahalanobis distance and the second Mahalanobis distance includes: obtaining the mean and standard deviation of the first Mahalanobis distance of the pure honey training sample group in the first chemical semantic sub-window, and performing standardization processing on the first Mahalanobis distance using the mean and standard deviation of the first Mahalanobis distance; obtaining the mean and standard deviation of the second Mahalanobis distance of the pure honey training sample group in the second chemical semantic sub-window, and performing standardization processing on the second Mahalanobis distance using the mean and standard deviation of the second Mahalanobis distance.

6. The method for identifying honey purity based on spectral data according to claim 5, characterized in that, The method for obtaining the comprehensive spectral deviation value of the honey sample to be tested includes: The mean of the standardized Mahalanobis distance of the known adulterated training sample group in the first chemical semantic sub-window is used as the first discriminant separation index, and the mean of the standardized Mahalanobis distance in the second chemical semantic sub-window is used as the second discriminant separation index. Based on the first and second discriminant separation indices, the first adaptive weight of the first chemical semantic sub-window and the second adaptive weight of the second chemical semantic sub-window are calculated respectively. The first and second Mahalanobis distances after standardization are weighted and summed using the first and second adaptive weights to obtain the comprehensive spectral deviation value.

7. The method for identifying honey purity based on spectral data according to claim 6, characterized in that, The preset discrimination threshold is determined as follows: Calculate the mean and standard deviation of the comprehensive spectral deviation of the pure honey training sample group. Add the mean of the comprehensive spectral deviation to the standard deviation of the comprehensive spectral deviation by a preset multiple and use the result as the preset discrimination threshold.

8. The method for identifying honey purity based on spectral data according to claim 1, characterized in that, The step of comparing the comprehensive spectral deviation value with a preset discrimination threshold and outputting the purity identification result of the honey sample to be tested includes: When the overall spectral deviation value is greater than the preset discrimination threshold, the honey sample to be tested is determined to be adulterated honey; when the overall spectral deviation value is less than or equal to the preset discrimination threshold, the honey sample to be tested is determined to be pure honey; the ratio of the overall spectral deviation value to the preset discrimination threshold is calculated, and the ratio is output as a relative deviation reference value.

9. The method for identifying honey purity based on spectral data according to claim 1, characterized in that, Also includes: The sample temperature of the honey sample is monitored in real time while collecting the original spectrum of the honey sample to be tested. When the sample temperature deviates from the preset calibration temperature range, a temperature warning is triggered and the output of purity identification results is paused.

10. A honey purity identification system based on spectral data, characterized in that, It includes a memory and a processor. The memory stores computer program instructions, which, when executed by the processor, implement the honey purity identification method based on spectral data as described in any one of claims 1-9.