Tea quality detection method based on combination of bimodal spectrum and AG-FDN
By combining dual-modal spectroscopy with the AG-FDN model, the polarization reflectance spectrum of tea leaves and the fluorescence transmission spectrum of tea liquor are integrated, which solves the limitation of single-modal testing in existing tea quality detection, and realizes three-dimensional characterization of tea quality and rapid and accurate quality grading.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG ECONOMIC & TRADE POLYTECHNIC
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-24
AI Technical Summary
Existing tea quality testing technologies rely on single-modal spectroscopy, which makes it difficult to fully reflect the complex quality characteristics of tea, resulting in limited detection accuracy and generalization ability. Furthermore, traditional methods suffer from problems such as large subjective errors, strong destructiveness, and long processing times.
By employing dual-modal spectroscopy combined with the AG-FDN model, the polarization reflectance spectrum of tea leaves and the fluorescence transmission spectrum of tea infusion are integrated. Through attention-guided feature decoupling and fusion, the physical characteristics and biochemical components of tea leaves are decoupled and fused.
It improves the accuracy and robustness of tea quality testing, reduces the risk of misjudgment due to physical damage or biochemical variations, automates and accelerates quality grading, and provides interpretable decision-making basis.
Smart Images

Figure CN121917469A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tea quality testing technology, specifically a tea quality testing method combining dual-modal spectroscopy and AG-FDN. Background Technology
[0002] As a globally important economic crop and beverage ingredient, tea's quality directly impacts its market value and consumer experience. Traditional tea quality evaluation primarily relies on sensory assessment and chemical testing. However, sensory assessment is easily influenced by the subjective experience of the assessors, resulting in large fluctuations in results and low efficiency. While chemical testing can quantitatively analyze components, it requires sample destruction, is time-consuming, and costly. With the development of spectroscopic technology and machine learning, non-destructive testing based on spectral images has gradually become a research hotspot. Multispectral imaging technology can capture reflectance or transmission spectra at different wavelengths to extract the physical and biochemical characteristics of tea, and combine this with deep learning models to achieve automated grading. However, existing spectroscopic detection methods mostly focus on single modes, making it difficult to comprehensively reflect the complex quality characteristics of tea. Furthermore, insufficient research on the mechanisms of feature decoupling and fusion limits detection accuracy and generalization ability.
[0003] Traditional tea quality testing techniques have several limitations: First, sensory evaluation relies on human experience, and different tea tasters may have different grading opinions for the same tea, especially in subtle quality distinctions, where subjective errors are significant. Second, chemical testing requires solvent extraction and chromatographic separation to analyze components such as tea polyphenols and amino acids. While this provides quantitative data, the destructive nature of the samples makes them unusable, and the long testing cycle makes it difficult to meet the real-time testing needs of large-scale production lines. Furthermore, existing spectroscopic detection techniques mostly use single-modal data, ignoring the complementarity between tea leaf appearance and tea liquor characteristics. For example, polarization reflectance spectroscopy can reflect the integrity of leaf cell structure but cannot directly reveal the content of water-soluble components; while fluorescence spectroscopy of tea liquor can characterize soluble substances, it is difficult to assess physical damage to leaves. This single-modal analysis results in insufficient model recognition of complex quality problems, especially when detecting mixed defects, leading to a high misjudgment rate. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a tea quality detection method combining dual-modal spectroscopy and AG-FDN. This method integrates external polarization reflectance spectroscopy and tea infusion fluorescence transmission spectroscopy, combined with attention-guided AG-FDN model, to achieve the decoupling and fusion of physical characteristics and biochemical components of tea.
[0005] To solve the above-mentioned technical problems, this invention provides the following technical solution: a method for detecting tea quality using dual-modal spectroscopy combined with AG-FDN, the specific steps of which are as follows:
[0006] S100, Dual-modal sample preparation: Tea samples of the same variety but different origins and grades were selected from the main tea-producing areas and prepared by combing and laying them flat in white ceramic petri dishes; tea soup samples were prepared by treating the samples in a constant temperature water bath and filtering them for a constant time according to the standard tea-to-water ratio, water temperature, and brewing time.
[0007] S200, Dual-modal spectral image acquisition: A dual-modal system including a multispectral imager, a three-dimensional stage, and a stable light source is constructed. The multispectral imager acquires polarization reflectance spectra of tea leaves and fluorescence emission spectra under specific excitation light. For tea infusion, polarization transmission spectra are acquired with the same parameters, and fluorescence transmission spectra are acquired under the same excitation light.
[0008] S300, dual-modal spectral preprocessing: For polarization images, denoising is performed using a 3×3 Gaussian filter kernel with a standard deviation of 1.5, followed by dark current and white balance correction, and segmentation using the Otsu algorithm and morphological opening operation; For fluorescence images, denoising is performed using a db4 wavelet basis, 3-level decomposition, and soft thresholding, followed by dark current and fluorescence spectral correction, and segmentation using the Otsu algorithm and morphological closing operation.
[0009] S400, Dual-modal Feature Processing and Model Training: An attention-guided AG-FDN model is built, preprocessed dual-modal data is input, the features of the micro-physical properties of the appearance, the physical properties of the tea soup, the biochemical components of the appearance, and the dissolved components of the tea soup are decoupled, fused through an attention mechanism, combined with grade labels, and the AG-FDN model is trained.
[0010] S500, Tea Quality Testing and Result Output: The tea to be tested is processed according to the steps of dual-modal sample preparation, dual-modal spectral image acquisition and dual-modal spectral preprocessing, and input into the AG-FDN model to output the quality grade. At the same time, the abnormal regions of physical structure / biochemical components are identified through the feature anomaly quantification algorithm.
[0011] Furthermore, in S100, during the preparation of the dual-modal sample, when preparing the tea infusion sample, the standard brewing operation is performed according to the following classification:
[0012] Green tea: Weigh 3.0g or 5.0g of representative tea sample, corresponding to a 150mL or 250mL tea tasting cup, with a tea-to-water mass-to-volume ratio of 1:50. Pour in boiling water at 100℃, cover and time for 4 minutes. After brewing, filter the tea using a stainless steel filter within 3 seconds and transfer the tea soup into a 10mm optical path transparent quartz cuvette.
[0013] Black tea, white tea, yellow tea, and strip / curled oolong tea: Weigh 3.0g or 5.0g of tea sample, tea-to-water ratio 1:50, pour in 100℃ boiling water, cover and time for 5 minutes, the rest of the operation is the same as green tea;
[0014] For round / curled / granular oolong tea: Weigh 3.0g or 5.0g of tea sample, with a tea-to-water ratio of 1:50, pour in boiling water at 100℃, cover and time for 6 minutes, and follow the same procedure as for green tea.
[0015] Black tea: Brew twice. For the first brew, weigh out 3.0g or 5.0g of tea sample, pour in boiling water at 100℃, cover and steep for 2 minutes, then filter. For the second brew, cover and time for 5 minutes. Combine the two brews and follow the same procedure as for green tea.
[0016] Compressed tea: Adjust the first brewing time according to the degree of compression, 2-5 minutes; for the second brewing, time it for 5-8 minutes; the rest of the operation is the same as for black tea.
[0017] Scented tea: After removing petals and impurities, weigh 3.0g of tea sample. For the first brew, cover and steep for 3 minutes. For the second brew, cover and steep for 5 minutes. The rest of the process is the same as for green tea.
[0018] Tea bag: Place one tea bag sample in a 150mL tea tasting cup, pour in 100℃ boiling water, cover and steep for 3 minutes. Then, lift the tea bag once at the 3rd minute and once at the 4th minute, with a 1-minute interval. After 5 minutes, drain the tea and assess the integrity of the tea bag leaves.
[0019] Powdered tea: Weigh 0.6g of tea sample and place it in a 240mL tea tasting bowl. Pour 150mL of boiling water into the bowl and steep for 3 minutes, stirring with a tea whisk throughout the process. After steeping, take the tea liquor directly for testing.
[0020] Throughout all operations, the water temperature is controlled by a constant temperature water bath with an accuracy of ±0.1℃; the filter screen is made of stainless steel with uniform pore size; the quartz cuvettes must be pre-washed with boiling water and dried to ensure that no residual impurities affect the spectral transmittance.
[0021] Furthermore, in S200, the polarization angle adjustment accuracy of the multispectral imager in dual-modal spectral image acquisition is ≤ ±0.5°, the fluorescence excitation light wavelength adjustment accuracy is ≤ ±2nm; the ambient light intensity in the dark room is ≤3lx, and the background reflectivity is ≤1%.
[0022] Furthermore, in S200, the multispectral imager in the dual-modal spectral image acquisition switches bands sequentially in a wavelength range of 400-1000nm with a step size of 5nm. Under each band, images are acquired using polarization angles of 0°, 45°, 90°, and 135°, and each polarization angle is repeatedly acquired 3 times. The specific excitation light is a light source with a wavelength of 450nm-470nm, and the intensity of the excitation light can be adjusted within the range of 80mW-120mW according to the chlorophyll content in the tea sample, with an adjustment accuracy of ±5mW.
[0023] Furthermore, in S300, the specific steps of segmentation using the Otsu algorithm and morphological closing operation in the dual-modal spectral preprocessing are as follows: Select the fluorescence transmission image at a wavelength of 600nm from the fluorescence image, and let the image grayscale value set be... ,in It represents the individual gray values in the gray value set, where L is the total number of gray levels. The Otsu algorithm is used to calculate the segmentation threshold T, and the calculation formula is: ,in It is the threshold variable used to calculate the segmentation threshold. These represent the pixel percentages of the tea infusion region and the cuvette region, respectively, segmented by threshold t. The grayscale variances of the foreground and background are represented; the image is segmented into a binary image based on a threshold T. , Let be the image coordinates, where Represents the pixel area of the tea soup. Representing the background region pixels; performing morphological closing operation on the binary image B using a 3×3 square structuring element S, first performing a dilation operation, the formula is: , The binary image after dilation operation in coordinates The pixel value at that location is dilated once. These are the coordinates within the structuring element; the erosion operation is performed using the following formula: , The binary image after the closing operation is in coordinates The pixel value at that location was eroded once to obtain a complete effective detection area for the tea soup.
[0024] Furthermore, in S400, the attention-guided AG-FDN model in dual-modal feature decoupling and model training includes an input layer, a feature decoupling layer, a feature fusion layer, and a classification output layer connected in sequence.
[0025] Input layer: Receives preprocessed polarization image features and fluorescence image features; the dimensions of the polarization image features are [batch_size, height, width, polar_channel], where batch_size is the number of samples in each training iteration, set to 32; height and width are the image dimensions, uniformly adjusted to 224×224; polar_channel is the number of feature channels in the polarization image, determined based on the imaging wavelength and the number of polarization angles; the dimensions of the fluorescence image features are [batch_size, height, width, number of fluorescence channels], the number of fluorescence channels being determined based on the excitation wavelength and the number of acquisition wavelengths.
[0026] Feature decoupling layer: The polarization spectral features are decoupled into the microscopic physical features of the shape and the physical features of the tea soup through the convolutional layer; the fluorescence spectral features are decoupled into the biochemical component features of the shape and the dissolved component features of the tea soup. The convolutional kernel size is 3×3, the number of convolutional kernels is 64, the stride is 1, and the padding is 1.
[0027] Feature fusion layer: Employs a cross-modal attention fusion mechanism to calculate the correlation weights between polarization and fluorescence features, fuses the decoupled multimodal features, and generates a fused feature vector;
[0028] Classification output layer: The fused feature vector passes through two fully connected layers. The first fully connected layer has 512 neurons and uses ReLU as the activation function. The second fully connected layer has the number of neurons corresponding to the number of tea quality grades and uses softmax as the activation function. The fused features are then transformed into specific quality grades and output through a quality grade mapping algorithm.
[0029] Furthermore, in S400, the specific steps for implementing dual-modal feature fusion weight allocation in dual-modal feature decoupling and model training are as follows: First, for each feature channel obtained after decoupling, extract the signal distribution information corresponding to the polarization feature and fluorescence feature in that channel; then compare the degree of difference in signal distribution between the polarization feature and fluorescence feature within the same feature channel, and according to the rule that the smaller the difference, the higher the reference value of that channel for tea quality detection, initially assign basic weights to each feature channel; then, calculate the basic weights of all feature channels and perform normalization processing to obtain the final fusion weights of each feature channel; finally, according to the final fusion weights, weightedly integrate the polarization feature and fluorescence feature of the corresponding feature channel to obtain the fused feature vector.
[0030] Furthermore, in S400, during dual-modal feature decoupling and model training, the calculation formula for the quality level mapping algorithm is as follows: ,in It refers to the output quality level. It is the Sigmoid function, and the formula is: , These are the input values for the Sigmoid function. For temperature coefficient, It is the decoupled first A fusion feature, It is a deviation correction term. It is the number of fused features. It is the first The fusion weights of each feature channel.
[0031] Furthermore, in S400, the specific steps for training the AG-FDN model in the bimodal feature decoupling and model training are as follows: The preprocessed bimodal image data is divided into a training set, a validation set, and a test set in a 7:2:1 ratio; data augmentation processing is performed on the training set data, including random rotation, random flipping, and random brightness adjustment; the Adam optimizer parameters are set with a learning rate of 0.001, which decays to 0.9 every 10 training epochs, and a weight decay coefficient of 0.0001; the cross-entropy loss function is used to calculate the loss between the predicted value and the true level label, and the loss function formula is: ,in, The total number of tea quality grades, For real labels, The model predicts that the sample belongs to the first... The probability of a class, where the sample belongs to class r. ,otherwise Using a batch size of 32 bits, the training set data was input into the AG-FDN model for training, which lasted for 100 rounds. After each round of training, the model performance was evaluated using a validation set. During training, if the performance on the validation set did not improve for 5 consecutive rounds, the training was stopped early; otherwise, the model parameters were saved after 100 rounds of training.
[0032] Furthermore, in the S500 process of rapid detection and output of tea quality, the calculation formula for the feature anomaly quantification algorithm is as follows: ,in Image coordinates Anomaly at the location, G represents the mean and standard deviation of the fusion features of normal samples. It is a Gaussian space weight. It is to avoid the minimum value where the denominator is 0.
[0033] Compared with existing technologies, this dual-modal spectroscopy combined with AG-FDN method for tea quality detection has the following advantages:
[0034] I. This invention integrates the multi-dimensional features of tea leaf shape polarization reflectance spectrum and tea infusion fluorescence transmission spectrum by constructing a dual-modal sample preparation and spectral image acquisition system. Traditional detection methods often rely on single-modal data, which are easily affected by environmental interference or feature limitations. However, this method achieves a three-dimensional characterization of tea quality by decoupling the microscopic physical features of the shape, the physical features of the tea infusion, biochemical components, and dissolved components. Combined with a cross-modal attention fusion mechanism, the model can dynamically capture the correlation weight between polarization and fluorescence features, improving the robustness of feature expression. The shape polarization reflectance spectrum can reflect the integrity of leaf cell structure, while the tea infusion fluorescence spectrum can reveal the content and distribution of water-soluble components. The complementarity of the two effectively reduces the risk of misjudgment due to physical damage or biochemical variation, and significantly improves the accuracy of quality grading.
[0035] Second, this invention introduces KL divergence and joint entropy to calculate the dual-modal difference attention weights, enabling the model to accurately locate microscopic defect areas such as external cracks and abnormal turbidity of tea soup, providing interpretable decision-making basis for quality control. At the same time, dynamic learning rate decay and data augmentation techniques are used during training to enhance the model's adaptability to different lighting conditions and sample diversity. Compared with traditional methods that require multiple independent detections, this solution achieves full automation of the preparation-collection-analysis-output process, and avoids overfitting by stopping training early, ensuring stability and economy in industrial scenarios.
[0036] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0038] Figure 1 Flowchart for dual-modal spectral AG-FDN tea quality testing;
[0039] Figure 2 Flowchart for dual-modal data processing and model application;
[0040] Figure 3 This is a schematic diagram of the AG-FDN model structure and feature processing. Detailed Implementation
[0041] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structure, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0042] Example 1:
[0043] Longjing tea quality testing
[0044] Dual-modal sample preparation: Longjing tea samples of the same variety but from different origins and grades were selected from the main production area of Longjing to ensure that the samples cover the common quality differences in this production area, providing a comprehensive reference benchmark for subsequent testing; during the preparation of morphological samples, the tea leaves were laid flat in white ceramic petri dishes to ensure that the tea leaves do not overlap and are evenly distributed, such as... Figure 2 As shown, this reduces interference from tea leaves obscuring each other on subsequent spectral image acquisition, resulting in more realistic and accurate acquisition of morphological features. Tea samples are prepared by weighing tea leaves and purified water at a tea-to-water ratio of 1:45, with the water temperature controlled at 95℃. A constant temperature water bath is used to maintain the water temperature accuracy within ±0.1℃. Precise temperature control avoids the impact of temperature fluctuations on the dissolution of substances within the tea leaves, ensuring the stability of the tea liquor composition. The brewing time is strictly controlled at 4 minutes to allow the soluble components in the tea leaves to dissolve fully and stably. After brewing, the tea is filtered through a stainless steel filter within 3 seconds. Rapid filtration reduces additional contact between the tea liquor and tea leaves, preventing further changes in composition. The tea liquor is then transferred to a transparent quartz cuvette with a 10mm optical path. A uniform optical path specification provides consistent optical conditions for subsequent spectral detection, ensuring the comparability of the detection data. Figure 1 As shown.
[0045] Dual-modal spectral image acquisition: A dual-modal system including a multispectral imager, a 3D stage, and a stable light source was constructed. Acquisition was conducted in a darkroom environment. The extremely low ambient light and background reflectivity minimized interference from external light on the spectral signal, ensuring that the acquired spectral information came solely from the tea sample itself. For the tea's appearance, the multispectral imager switched bands in a 400-1000nm wavelength range with 5nm steps. This wavelength range covers the characteristic spectra of the tea's appearance and internal components. The 5nm step size ensured both the richness of spectral information and prevented excessive data processing burden due to too many bands. Each band was divided into 0°, Polarization reflectance spectra were acquired at polarization angles of 45°, 90°, and 135°. Acquisition at different polarization angles can capture the differences in the microstructure of the tea surface. Each polarization angle was repeatedly acquired three times, and the average of multiple acquisitions was taken to reduce random errors. An excitation wavelength of 460nm was selected. The targeted excitation wavelength can effectively excite the fluorescence of biochemical components in tea that are related to quality. The intensity can be adjusted to adapt to samples with different chlorophyll contents to ensure that the fluorescence signal is moderate. Fluorescence emission spectra were acquired. High-precision wavelength adjustment ensured the accuracy of fluorescence detection. For tea infusion, the same parameters as those used for external acquisition were used to acquire polarization transmission spectra and fluorescence transmission spectra.
[0046] Dual-modal spectral preprocessing: For polarized images, a 3×3 Gaussian filter kernel is used for noise reduction. This filtering method can smooth image noise while preserving image edge information well, avoiding the loss of useful features. Dark current and white balance corrections are applied to eliminate the influence of imaging device noise and uneven illumination on the image, making the image grayscale values more realistically reflect the sample features. Segmentation is performed using the Otsu algorithm and morphological opening operations. The Otsu algorithm can automatically determine the optimal segmentation threshold to effectively separate the tea leaf area from the background, while the morphological opening operation can remove small noise and small protrusions in the segmented image, making the tea leaf area more distinct. The outline is clearer. For the fluorescence image, a 3-level decomposition is performed using the db4 wavelet basis, and soft thresholding is used for denoising. Wavelet transform denoising can remove noise while preserving the detailed features of the fluorescence signal, which is suitable for processing high-frequency noise in fluorescence images. Dark current and fluorescence spectrum correction are used to eliminate the influence of equipment noise and spectral drift on the fluorescence signal. The Otsu algorithm and morphological closing operation are used for segmentation. The Otsu algorithm separates the tea soup area from the cuvette background, and the morphological closing operation can fill small holes in the tea soup area, making the effective detection area of tea soup more complete and providing an accurate area range for subsequent feature extraction.
[0047] The specific steps for segmentation using the Otsu algorithm and morphological closing operation are as follows: Select a fluorescence transmission image with a wavelength of 600nm from the fluorescence image, and let the image grayscale value set be... ,in It represents the individual gray values in the gray value set, where L is the total number of gray levels. The Otsu algorithm is used to calculate the segmentation threshold T, and the calculation formula is: ,in It is the threshold variable used to calculate the segmentation threshold. These represent the pixel percentages of the tea infusion region and the cuvette region, respectively, segmented by threshold t. The grayscale variances of the foreground and background are represented; the image is segmented into a binary image based on a threshold T. , Let be the image coordinates, where Represents the pixel area of the tea soup. Representing the background region pixels; performing morphological closing operation on the binary image B using a 3×3 square structuring element S, first performing a dilation operation, the formula is: , The binary image after dilation operation in coordinates The pixel value at that location is dilated once. These are the coordinates within the structuring element; the erosion operation is performed using the following formula: , The binary image after the closing operation is in coordinates The pixel value at that location was eroded once to obtain a complete effective detection area for the tea soup.
[0048] Bimodal feature processing and model training: An attention-guided AG-FDN model was built, and preprocessed bimodal data was input. This model is specifically designed for bimodal data and can effectively mine quality-related features in the appearance and tea soup.
[0049] The model's input layer receives polarization and fluorescence image features. The uniform image size and defined number of channels ensure a standardized data format, facilitating model processing.
[0050] The feature decoupling layer uses a 3×3 convolutional kernel to decouple polarization spectral features into shape microscopic physical features and tea infusion physical features, and fluorescence spectral features into shape biochemical component features and tea infusion dissolved component features. The decoupling process can decompose complex spectral features into targeted single-type features, which is convenient for subsequent separate analysis.
[0051] The feature fusion layer employs a cross-modal attention fusion mechanism to calculate the correlation weights between polarization and fluorescence features. The specific steps are as follows: First, for each feature channel obtained after decoupling, the signal distribution information corresponding to the polarization and fluorescence features in that channel is extracted. Then, the difference in signal distribution between the polarization and fluorescence features within the same feature channel is compared. Based on the rule that "the smaller the difference, the higher the reference value of that channel for tea quality detection," a basic weight is initially assigned to each feature channel. Subsequently, the basic weights of all feature channels are statistically analyzed and normalized to obtain the final fusion weight for each feature channel. Finally, according to this final fusion weight, the polarization and fluorescence features of the corresponding feature channels are weighted and integrated to obtain the fused feature vector. This integrates multi-dimensional information from both the appearance and the tea liquor, improving the model's ability to discriminate quality.
[0052] The classification output layer processes the fused feature vector through two fully connected layers. These fully connected layers further extract higher-order features through non-linear transformations. The softmax activation function converts the output into probabilities for each level, facilitating the determination of the final level. A quality level mapping algorithm then transforms the output into specific quality levels. This algorithm establishes a precise correspondence between the fused features and the actual quality levels. The calculation formula for the quality level mapping algorithm is as follows: ,in It refers to the output quality level. It is the Sigmoid function, and the formula is: , These are the input values for the Sigmoid function. For temperature coefficient, It is the decoupled first A fusion feature, It is a deviation correction term. It is the number of fused features. It is the first The fusion weights of each feature channel.
[0053] The specific steps for model training are as follows: The preprocessed bimodal image data is divided into training, validation, and test sets in a 7:2:1 ratio. This reasonable ratio ensures the model has sufficient training samples while allowing for evaluation of generalization ability through the validation and test sets. Data augmentation is performed on the training set using random rotation, flipping, and brightness adjustments to increase sample diversity and prevent overfitting. The Adam optimizer is set up; it efficiently updates model parameters, and learning rate decay and weight decay prevent overfitting during training. The cross-entropy loss function is used to calculate the loss, with the formula: ,in, The total number of tea quality grades, For real labels, The model predicts that the sample belongs to the first... The probability of a class, where the sample belongs to class r. ,otherwise This function can effectively measure the difference between the predicted value and the true label, and guide the optimization of model parameters. It is trained for 100 rounds with a batch size of 32. The performance is evaluated with a validation set in each round. If the performance does not improve for 5 consecutive rounds, it is stopped early to avoid invalid training. Finally, the model parameters are saved to obtain a detection model with stable performance.
[0054] After processing the Longjing tea to be tested according to the above-mentioned dual-modal sample preparation, spectral image acquisition and preprocessing steps, it was input into the trained AG-FDN model. The model achieved an overall accuracy of 98.5% in the tea quality detection task. Specifically, the accuracy rate for identifying Grade 1 (best) Longjing tea was 99.2%, Grade 2 was 98.8%, Grade 3 was 97.6%, Grade 4 was 96.9%, and Grade 5 (worst) was 97.3%. It can quickly use the learned feature patterns to judge the quality of tea and output the quality grade, realizing rapid detection of tea quality. At the same time, the feature anomaly quantification algorithm identifies abnormal areas of the physical structure or biochemical components of tea. The algorithm has an accuracy rate of 96.8% in identifying damaged areas of leaves and 95.7% in identifying areas with abnormal chlorophyll content. By calculating the degree of deviation between the features and normal samples, abnormal areas can be accurately located, providing a reference for further analysis and improvement of tea quality.
[0055] In summary, this embodiment focuses on the quality testing of Longjing tea, strictly adhering to the dual-modal testing process of the invention patent. From the sample preparation stage, it ensures the standardization of the appearance and tea liquor samples, laying a reliable foundation for subsequent testing. Spectral acquisition utilizes multi-parameter control and precise equipment setup to obtain spectral information reflecting tea quality. The preprocessing step leverages various algorithms to optimize image quality and improve feature extraction accuracy. The AG-FDN model, through feature decoupling, fusion, and training, achieves accurate discrimination of Longjing tea quality grades. Finally, the output results are presented and abnormal regions are identified. The entire process fully utilizes the dual-modal technology and algorithms in the patent, enabling rapid, objective, and comprehensive testing of Longjing tea quality, providing an effective means for Longjing quality control.
[0056] Example 2:
[0057] Example of Longjing Tea Quality Testing
[0058] From the main production area of Longjing tea, 50 samples of each of the same variety "Longjing No. 43" and six grades (from premium to fifth grade) were selected to ensure that the samples covered 21 different production areas and common quality defects, providing a comprehensive quality difference benchmark for model training. For the preparation of the morphological samples, 5g of tea leaves were placed in a 10cm diameter white ceramic petri dish and gently combed with a soft brush to ensure that the tea leaves were laid out in a single layer without overlapping, avoiding mutual occlusion of leaves and affecting the accurate capture of morphological texture and color in subsequent spectral images. For the preparation of the tea infusion samples, a tea-to-water ratio of 1:42 was used, weighing 3g of tea leaves and 126mL of pure... Pour purified water into a 500mL beaker and place the beaker in a constant temperature water bath. Set the water temperature to 92℃ and maintain it with a temperature feedback system of ±0.1℃ to ensure stable dissolution of soluble components from the tea leaves. Strictly control the brewing time to 4 minutes. Immediately after the timer expires, filter the tea using a 0.2mm pore size stainless steel filter within 3 seconds to prevent continuous contact between the tea and leaves from altering the composition. Then, slowly transfer the filtered tea into a 10mm path length transparent quartz cuvette to avoid air bubbles affecting subsequent spectral transmission detection, thus completing the preparation of the dual-modal sample. Figure 3 As shown.
[0059] A dual-modal detection system consisting of a multispectral imager, a three-dimensional motorized stage, and a stable halogen tungsten light source was constructed. The entire acquisition process was conducted in a darkroom environment. A darkroom shielding layer and an ambient light monitor ensured that the ambient light intensity was ≤3 lx. A low-reflectivity matte black pad was laid on the stage surface to minimize interference from background light on the spectral signal. For tea leaf shape samples, a ceramic petri dish containing the sample was placed in the center of the three-dimensional stage. The stage height was adjusted to maintain a 30 cm distance between the imager lens and the sample. The multispectral imager switched wavelengths sequentially within the 400-1000 nm wavelength range in 5 nm increments. Polarization reflectance spectra were acquired at polarization angles of 0°, 45°, 90°, and 135° for each wavelength band. The polarization angle was adjusted using the imager's built-in polarizer wheel, with an adjustment accuracy of ≤±0.5°. To reduce the impact of random noise, the data was repeatedly sampled three times for each polarization angle. Then, an LED excitation light with a wavelength of 465 nm was selected. The intensity of the excitation light was adjusted based on the previously measured chlorophyll content of the samples, with an adjustment accuracy of ±5 mW. Fluorescence emission spectra of the tea leaves were then collected, with the excitation light wavelength adjustment accuracy ≤ ±2 nm. For the tea infusion samples, a quartz cuvette containing the tea infusion was fixed on a special fixture on the stage, ensuring the center of the cuvette was aligned with the optical axis of the imager lens. Using the same wavelength range, step size, and polarization angle parameters as the tea leaf samples, polarization transmission spectra of the tea infusion were collected. Simultaneously, the 465 nm excitation light used for the tea leaf samples was used, with the excitation light intensity finely adjusted according to the color depth of the tea infusion, to collect the fluorescence transmission spectrum of the tea infusion. This completed the dual-modal spectral image acquisition. All acquired data were automatically stored in ENVI format for subsequent preprocessing and analysis.
[0060] The acquired polarization reflectance and transmission spectral images were first preprocessed to optimize image quality. For the polarization image, a 3×3 Gaussian filter kernel was used for noise reduction. A Gaussian function was used to apply a weighted average to the area surrounding each pixel, smoothing random noise while preserving key features such as the tea leaf edges and the tea liquor interface. Dark current correction was then performed by subtracting the dark current image acquired by the imager under complete darkness from the original polarization image to eliminate the influence of thermal noise from the sensor. Finally, white balance correction was applied, using the standard white area of a white ceramic culture dish as a reference to adjust the gain of the RGB channels of the image. The color reference of polarization images collected from different batches is consistent. Finally, the Otsu algorithm is used to automatically calculate the segmentation threshold, separating the tea leaf shape area from the ceramic petri dish background and the tea soup area from the quartz cuvette background. Morphological opening operations are then performed on the segmented binary images to remove small noise points, resulting in clear effective polarization image regions. For fluorescence images, due to the weak fluorescence signal and susceptibility to high-frequency noise interference, a 3-level wavelet decomposition using the db4 wavelet basis is employed to decompose the image into low-frequency approximation coefficients and high-frequency detail coefficients. A soft thresholding method is used to process the high-frequency detail coefficients, setting the threshold T based on the image noise level. The calculation formula is as follows: After removing the high-frequency components corresponding to the noise, the image is reconstructed using inverse wavelet transform; subsequently, dark current correction and fluorescence spectral correction are performed; finally, the Otsu algorithm is used to calculate the segmentation threshold, separating the effective region and background in the fluorescence image, and then morphological closing operation is performed on the binary image, with the formula as follows: , The binary image after dilation operation in coordinates The pixel value at that location is dilated once. These are the coordinates within the structuring element; the erosion operation is performed using the following formula: , The binary image after the closing operation is in coordinates The pixel value at the location was eroded once to obtain a complete effective detection area of the tea soup, thus completing the dual-modal spectral preprocessing.
[0061] An attention-guided AG-FDN model was constructed. The overall architecture of the model includes an input layer, a feature decoupling layer, a feature fusion layer, and a classification output layer. The input layer receives preprocessed polarization image features and fluorescence image features. The polarization image feature dimension is [32, 224, 224, polar_channel], the batch_size is set to 32, and the image size is uniformly adjusted to 224×224. The polar_channel is determined to be 484 based on the number of imaging wavelengths (121 bands in 400-1000nm) and the number of polarization angles (4). The image feature dimensions are [32, 224, 224, number of fluorescence channels]. The number of fluorescence channels is determined to be 121 based on the excitation wavelength (1) and the number of acquisition wavelengths (121). The feature decoupling layer uses a 3×3 convolutional kernel to construct a convolutional layer. Through convolution operations, the polarization spectral features are decoupled into the microscopic physical features of the shape and the physical features of the tea infusion. At the same time, the fluorescence spectral features are decoupled into the biochemical component features of the shape and the dissolved component features of the tea infusion, realizing the separation and extraction of different types of features. The feature fusion layer adopts a cross-modal attention fusion mechanism to calculate the correlation weight between polarization features and fluorescence features. Specific steps are as follows. The process is as follows: First, for each feature channel obtained after decoupling, the signal distribution information corresponding to the polarization feature and fluorescence feature under that channel is extracted. Then, the difference in signal distribution between the polarization feature and fluorescence feature within the same feature channel is compared. According to the rule that "the smaller the difference, the higher the reference value of that channel for tea quality detection", a basic weight is initially assigned to each feature channel. Subsequently, the basic weights of all feature channels are calculated and normalized to obtain the final fusion weight of each feature channel. Finally, according to the final fusion weight, the polarization feature and fluorescence feature of the corresponding feature channel are weighted and integrated to obtain the fused feature vector. Then, the four types of features after decoupling—exterior microphysical, tea soup physical, exterior biochemical components, and tea soup dissolved components—are fused into a unified fused feature vector. The classification output layer contains two fully connected layers. The first fully connected layer has 512 neurons and uses the ReLU activation function to extract high-order features from the fused feature vector. The number of neurons in the second fully connected layer is consistent with the number of quality grades of Longjing tea (6 grades). The softmax activation function is used to convert the output into the probability of the sample belonging to each grade. Finally, the quality grade mapping algorithm is used. The calculation formula of the quality grade mapping algorithm is: By combining the Sigmoid function with the deviation correction term, the probability value is transformed into a specific quality grade label.
[0062] During the model training phase, the preprocessed bimodal image data was randomly divided into training, validation, and test sets in a 7:2:1 ratio. Data augmentation was performed on the training set data, including random rotation, horizontal random flipping (probability 0.5), and random brightness adjustment, to increase sample diversity and prevent model overfitting. The Adam optimizer parameters were set with an initial learning rate of 0.001. Every 10 training epochs, the learning rate was adjusted to 0.9 using a learning rate decay strategy, while the weight decay coefficient was set to 0.0001 to prevent model parameter degradation. Excessive size can lead to overfitting. The cross-entropy loss function is used to calculate the loss between the model's predicted value and the true level label of the sample. The training set data is input into the AG-FDN model with a batch size of 32 for training. A total of 100 training cycles are set. After each training cycle, the accuracy and F1 score of the model are evaluated using the validation set data. If the accuracy of the validation set does not improve for 5 consecutive cycles, an early stop training strategy is triggered to avoid ineffective training. If the performance of the validation set is still improving after 100 training cycles, the model parameters of the final cycle are saved, and the AG-FDN model training is completed.
[0063] For the Longjing tea to be tested, the above-mentioned dual-modal sample preparation, dual-modal spectral image acquisition, and dual-modal spectral preprocessing steps were strictly followed: First, the tea leaves were laid flat to prepare shape samples; then, tea liquor samples were prepared by brewing and filtering under standard conditions. Next, polarization and fluorescence spectral images of the shape and tea liquor were acquired. After preprocessing such as Gaussian filtering, wavelet denoising, correction, and segmentation, feature data meeting the model input requirements were obtained. The preprocessed feature data was input into the trained AG-FDN model. The model, through feature decoupling, cross-modal attention fusion, and classification output, outputs the quality grade of the tea within 0.5 seconds. Validated on the test set, the model achieved an overall accuracy of 98.3% in grading the quality of Longjing tea, with an accuracy of 99.1% for premium Longjing and 98.7% for first-grade Longjing, meeting the requirements for industrial-grade detection accuracy. Simultaneously, the model invoked a feature anomaly quantification algorithm. The calculation formula for the feature anomaly quantification algorithm is as follows: Using the mean and standard deviation of the fusion features of normal samples as a benchmark, and combined with Gaussian spatial weights, the anomaly degree (A(x,y)) of each pixel of the test sample is calculated. When the anomaly degree is >2.5, the region is determined to be an abnormal region. The abnormal location is marked with a red box in the output result through the image annotation function, such as "damaged leaf edge area" and "local turbidity area of tea soup". This provides an intuitive decision basis for quality control in the tea production process and completes the entire rapid detection process of Longjing tea quality.
[0064] In summary, this embodiment focuses on the quality testing of Longjing tea. Multiple samples from the main production area were selected and standardized to prepare dual-modal samples of appearance and tea liquor. A system was built in a darkroom to accurately acquire dual-modal spectral images. Data was preprocessed and optimized using Gaussian filtering and wavelet denoising. Then, features were decoupled and fused using the AG-FDN model, and the model was trained using cross-entropy loss function, ultimately achieving rapid output of the quality grade of the Longjing tea to be tested and identification of abnormal areas. The overall accuracy is high, providing an efficient and objective testing solution for the quality control of Longjing tea.
[0065] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for tea quality detection using dual-modal spectroscopy combined with AG-FDN, characterized in that, The specific steps of this method are as follows: S100, Dual-modal sample preparation: Tea samples of the same variety but different origins and grades were selected from the main tea-producing areas and prepared into shape samples by combing and laying them flat in white ceramic petri dishes; Tea soup samples were prepared by temperature control and timed filtration according to the standard tea-to-water ratio, water temperature, and brewing time. S200, Dual-modal spectral image acquisition: A dual-modal system including a multispectral imager, a stage, and a stable light source is constructed. The multispectral imager acquires polarization reflectance spectra of the tea leaves and fluorescence emission spectra under a specific excitation light. For the tea infusion, polarization transmission spectra are acquired with the same parameters, and fluorescence transmission spectra are acquired under the same excitation light. S300, dual-modal spectral preprocessing: denoising and correcting polarization images before segmentation; Fluorescence images are segmented after wavelet denoising and correction to optimize image quality; S400, Dual-modal Feature Processing and Model Training: An attention-guided AG-FDN model is built, preprocessed dual-modal data is input, the features of the micro-physical properties of the appearance, the physical properties of the tea soup, the biochemical components of the appearance, and the dissolved components of the tea soup are decoupled, fused through an attention mechanism, combined with grade labels, and the AG-FDN model is trained. S500, Tea Quality Testing and Result Output: The tea to be tested is processed according to the steps of dual-modal sample preparation, dual-modal spectral image acquisition and dual-modal spectral preprocessing, and input into the AG-FDN model to output the quality grade. At the same time, the abnormal regions of physical structure / biochemical components are identified through the feature anomaly quantification algorithm.
2. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 1, characterized in that, In the S100 dual-modal sample preparation, when preparing the tea soup sample, the tea-to-water ratio is 1:40-1:50, the water temperature is 80℃-100℃, the brewing time is 4-5 minutes, the water temperature is controlled by a temperature control device, the filtration is completed shortly after the brewing time ends, and the tea soup is transferred into a transparent cuvette.
3. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 1, characterized in that, In the S200 dual-modal spectral image acquisition, the polarization angle adjustment accuracy and fluorescence excitation wavelength adjustment accuracy of the multispectral imager meet the detection requirements; the acquisition environment is a dark room environment with low ambient light and low background reflectivity.
4. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 1, characterized in that, In the S200 dual-modal spectral image acquisition, the multispectral imager switches bands sequentially according to a wavelength range of 400-1000nm and a preset step size. Image acquisition is performed using multiple polarization angles in each band, and each polarization angle is repeatedly acquired multiple times. The specific excitation light is a light source with a wavelength of 450nm-470nm, and the intensity of the excitation light can be adjusted according to the chlorophyll content in the tea sample within the range of 80mW-120mW.
5. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 4, characterized in that, In the S300 dual-modal spectral preprocessing, the segmentation operation for the fluorescence image is as follows: First, the fluorescence transmission image corresponding to the 600nm wavelength in the fluorescence image is selected, and the segmentation threshold is determined by the Otsu algorithm. This threshold is used to distinguish the pixel ratio of the tea soup area and the cuvette area. Based on this threshold, the image is divided into a binary image, where the part with a pixel value of 1 in the binary image corresponds to the tea soup area, and the part with a pixel value of 0 corresponds to the background area. Then, morphological closing operation is performed on the binary image. A 3×3 square structuring element is used. First, a dilation operation is performed, and then an erosion operation is performed to finally obtain the complete effective detection area of the tea soup.
6. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 4, characterized in that, In the S400 dual-modal feature decoupling and model training, the attention-guided AG-FDN model includes an input layer, a feature decoupling layer, a feature fusion layer, and a classification output layer connected in sequence. Input layer: Receives preprocessed polarization image features and fluorescence image features; Feature decoupling layer: The polarization spectral features are decoupled into the microscopic physical features of the shape and the physical features of the tea soup through the convolutional layer; the fluorescence spectral features are decoupled into the biochemical component features of the shape and the dissolved component features of the tea soup. Feature fusion layer: Employs a cross-modal attention fusion mechanism to calculate the correlation weights between polarization and fluorescence features, fuses the decoupled multimodal features, and generates a fused feature vector; Classification output layer: The fused feature vector passes through two fully connected layers, and the fused features are transformed into specific quality levels and output through a quality level mapping algorithm.
7. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 6, characterized in that, In the S400, the specific steps for implementing the weight allocation of dual-modal feature fusion in the decoupling and model training are as follows: First, for each feature channel obtained after decoupling, extract the signal distribution information corresponding to the polarization feature and fluorescence feature in that channel; then compare the degree of difference between the signal distribution of polarization feature and fluorescence feature in the same feature channel, and according to the rule that the smaller the difference, the higher the reference value of that channel for tea quality detection, initially allocate basic weights to each feature channel. Then, the basic weights of all feature channels are calculated and normalized to obtain the final fusion weights of each feature channel. Finally, according to the final fusion weight, the polarization features and fluorescence features of the corresponding feature channels are weighted and integrated to obtain the fused feature vector.
8. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 6, characterized in that, In the S400, during dual-modal feature decoupling and model training, the calculation formula for the quality level mapping algorithm is as follows: ,in It refers to the output quality level. It is the Sigmoid function, and the formula is: , These are the input values for the Sigmoid function. For temperature coefficient, It is the decoupled first A fusion feature, It is a deviation correction term. It is the number of fused features. It is the first The fusion weights of each feature channel.
9. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 1, characterized in that, In S400, the specific steps for training the AG-FDN model in the dual-modal feature decoupling and model training are as follows: the preprocessed dual-modal image data is divided into a training set, a validation set, and a test set according to a preset ratio; and the training set data is subjected to data augmentation processing. Set the optimizer parameters and use the cross-entropy loss function to calculate the loss between the predicted value and the true level label. The loss function formula is as follows: ,in, The total number of tea quality grades, For real labels, The model predicts that the sample belongs to the first... The probability of a class, where the sample belongs to class r. ,otherwise The training set data is input into the AG-FDN model for training. After each round of training, the model performance is evaluated using the validation set. During training, if the preset stopping condition is met, the training is stopped early; otherwise, after completing the preset number of training rounds, the model parameters are saved.
10. The tea quality detection method based on dual-modal spectroscopy combined with AG-FDN according to claim 1, characterized in that, In the S500 process for rapid detection and output of tea quality, the calculation formula for the feature anomaly quantification algorithm is as follows: ,in Image coordinates Anomaly at the location, G represents the mean and standard deviation of the fusion features of normal samples. It is a Gaussian space weight. It is to avoid the minimum value where the denominator is 0.