Grading method and system for quality detection of radix paeoniae rubra decoction pieces, terminal and medium

By introducing an attention mechanism and a hybrid denoising model, combined with a Transformer-based quality prediction model for Paeonia lactiflora, the problems of spectral data redundancy and noise interference in the quality detection of Paeonia lactiflora slices are solved. This enables non-destructive, rapid, and accurate quality grading of Paeonia lactiflora slices, promoting the standardization and modernization of the traditional Chinese medicine industry.

CN120913705APending Publication Date: 2025-11-07TIANJIN UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510919886.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing near-infrared hyperspectral detection methods for the quality testing of Paeonia lactiflora slices suffer from redundant spectral data and noise interference, resulting in poor grading effects and accuracy.

Method used

A hybrid denoising model combining attention mechanisms, generative adversarial networks, and convolutional autoencoders from deep learning, along with the Transformer architecture, is used to construct a quality prediction model for Paeonia lactiflora. Near-infrared hyperspectral imaging technology is then used to accurately predict the content of components and origin of Paeonia lactiflora slices and to classify their quality.

Benefits of technology

It achieves non-destructive, rapid, and accurate quality grading of Paeonia lactiflora slices, solves the problems of noise interference and information redundancy in the spectral analysis of complex Chinese medicine systems, is suitable for large-scale promotion, protects the integrity of Chinese medicinal materials, and improves the standardization and modernization of the Chinese medicine industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913705A_ABST
    Figure CN120913705A_ABST
Patent Text Reader

Abstract

The invention provides a grading method and system for quality detection of radix paeoniae rubra decoction pieces, a terminal and a medium. The method comprises the following steps: carrying out near-infrared hyperspectral data acquisition on a plurality of radix paeoniae rubra decoction piece samples from different producing areas; establishing a radix paeoniae rubra quality database based on the near-infrared hyperspectral data, and associating the data in the radix paeoniae rubra quality database with corresponding radix paeoniae rubra component content labels, producing area labels and manually marked quality grading labels; inputting near-infrared hyperspectral data of the radix paeoniae rubra decoction pieces to be detected into the trained mixed denoising model; and inputting the preprocessed near-infrared hyperspectral data into the trained quality prediction model for grading prediction. According to the grading method and system for quality detection of the radix paeoniae rubra decoction pieces, the terminal and the medium, the mixed denoising model based on the generative adversarial network and the convolutional auto-encoder is constructed, and accurate prediction and quality grading of the component content and the production place of the radix paeoniae rubra decoction pieces are realized in combination with a Transform architecture.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of traditional Chinese medicine quality control, and particularly relates to a grading method and system for quality detection of red peony root slices, a terminal and a medium. BACKGROUND

[0002] Red peony root is a commonly used traditional Chinese medicinal material, has the effects of clearing heat and cooling blood, removing blood stasis and relieving pain, and is widely used in the field of traditional Chinese medicine clinical treatment and traditional Chinese medicine pharmacy. As a commonly used traditional Chinese medicinal material in the field of traditional Chinese medicine clinical treatment and traditional Chinese medicine pharmacy, the quality of red peony root is affected by many factors such as production place, growth environment and harvesting and processing, and the difference in the content of effective components directly relates to the efficacy and safety. Although traditional detection methods such as high performance liquid chromatography (HPLC) and gas chromatography (GC) are accurate, they are complex to operate, time-consuming and costly, and need destructive pretreatment, and thus cannot meet the demand of large-scale rapid detection. Near-infrared hyperspectral technology has gradually been applied in the quality detection of agricultural products, food and medicines due to its advantages of rapidness, non-destructiveness and no need for pretreatment. However, the spectral data of traditional Chinese medicine complex system have problems such as information redundancy, noise interference (such as baseline drift and stray light) and insufficient model generalization ability, which affect the grading effect and grading accuracy of red peony root slices, and thus efficient processing and feature extraction technology is needed to solve the above problems. SUMMARY

[0003] Therefore, the application aims to provide a grading method, system, terminal and medium for quality detection of red peony root slices to solve the problem that the existing near-infrared hyperspectral detection applied to the quality detection of red peony root slices has poor grading effect and grading accuracy due to the existence of redundant information in the spectral data of red peony root slices.

[0004] To achieve the above object, the technical scheme of the application is as follows.

[0005] In a first aspect, the application provides a grading method for quality detection of red peony root slices, comprising the following steps.

[0006] Near-infrared hyperspectral data of red peony root slice samples from multiple different production places are collected, and an attention mechanism in deep learning is introduced in the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral features of the region containing more active ingredients;

[0007] A red peony root quality database is established based on the near-infrared hyperspectral data, and the data in the red peony root quality database are associated with corresponding red peony root component content labels, production place labels and artificial marking quality grading labels;

[0008] A hybrid denoising model is obtained by fusing a generative adversarial network and a convolutional autoencoder, and the hybrid denoising model is trained using the data in the red peony root quality database to obtain the trained hybrid denoising model;

[0009] construct a quality prediction model based on Transformers, and train the quality prediction model using data in the Radix Paeonia quality database to obtain a trained quality prediction model;

[0010] Obtain near-infrared hyperspectral data of the to-be-tested Radix Paeonia decoction pieces, and input the near-infrared hyperspectral data of the to-be-tested Radix Paeonia decoction pieces into the trained hybrid denoising model to obtain preprocessed near-infrared hyperspectral data.

[0011] Input the preprocessed near-infrared hyperspectral data into the trained quality prediction model for grading prediction.

[0012] Further, the near-infrared hyperspectral data of the Radix Paeonia decoction pieces samples from multiple different origins are collected, and the attention mechanism in deep learning is introduced in the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral characteristics of the region containing more active ingredients, including:

[0013] Obtain Radix Paeonia decoction pieces samples from multiple different origins, record batches and origins, and perform near-infrared hyperspectral data collection after performing split processing, to obtain near-infrared hyperspectral images;

[0014] The near-infrared hyperspectral images are processed using the YOLOv8 target detection algorithm to automatically identify the outline of the Radix Paeonia decoction pieces and delineate the region of interest, to obtain a feature map;

[0015] The CNN is used to perform multi-scale feature extraction on the feature map and attention weight calculation to obtain the attention weight of each feature channel; each feature channel corresponds to a Radix Paeonia component.

[0016] According to the attention weight, the collection resource is adjusted, the image sampling frequency and spectral signal intensity of the key region are dynamically improved based on the attention weight, and the key region data and non-key region data are weighted and fused based on the attention weight, to obtain near-infrared hyperspectral data.

[0017] Further, the CNN is used to perform multi-scale feature extraction on the feature map and attention weight calculation to obtain the attention weight of each feature channel; each feature channel corresponds to a Radix Paeonia component, including:

[0018] The CNN is used to perform multi-scale feature extraction on the feature map; wherein, the convolution layer: 3x3 convolution kernel and 5x5 convolution kernel are used to extract features in parallel, 3x3 kernel convolution is used to capture microscopic texture, and 5x5 convolution kernel is used to capture macroscopic spectral trend; secondly, the pooling layer: maximum pooling and average pooling are used alternately;

[0019] After obtaining the global feature vector through global average pooling, the global feature vector is input into two fully connected layers to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a Chai Pae component.

[0020] Further, the adjusting the collection resource according to the attention weight includes: based on the attention weight, dynamically improving the image sampling frequency and the spectral signal intensity of the key region, and based on the attention weight, weighting and fusing the key region data and the non-key region data to obtain near-infrared hyperspectral data.

[0021] According to the attention weight, the sampling frequency of the key region is dynamically improved, so that the sampling frequency of the active ingredient enrichment region is improved by at least 60%, and the spectral signal intensity is improved by at least 45%.

[0022] The key region data and the non-key region data are fused according to the attention weight to improve the spectral signal-to-noise ratio of the effective component to more than 25dB.

[0023] Further, the fusion generative adversarial network and the convolutional autoencoder obtain a hybrid denoising model, and the hybrid denoising model is trained by using data in the Chai Pae quality database to obtain a trained hybrid denoising model.

[0024] The generator is constructed based on a full convolutional neural network structure, the input is a random noise vector, and the output is spectral data of the same dimension as the near-infrared hyperspectral data; wherein the generator includes multiple transpose convolutional layers, which gradually expand the feature map size through multiple transpose convolutional layer operations to generate signals similar to the original spectrum; secondly, the generator also includes a hidden layer ReLU activation to enhance the nonlinear mapping to capture the nonlinear change of the characteristic peak of the Chai Pae component; thirdly, the generator also includes an output layer Tanh activation to map the spectral value to [-1, 1] to adapt to the normalized real spectral range;

[0025] Spectral prior constraints are embedded between each transpose convolutional layer in the generator to force the peak shape of the generated spectral data in the Chai Pae component feature area to be consistent with the real data, and through a gradient penalty term to ensure the physical interpretability of the generated spectral data, suppress meaningless noise generation, and realize accurate recovery of the spectral dimension and strengthening of the component features;

[0026] The discriminator is constructed based on a full convolutional neural network structure, the input is the spectral data generated by the generator or the real clean spectral data, and the output is a scalar to represent the probability that the input data is real data; wherein the generator includes multiple convolutional layers, the activation function uses LeakyReLU in the hidden layer, and Sigmoid is used in the output layer, and multiple convolutional layers extract features step by step; secondly, the loss function optimization adopts a binary cross-entropy with weight, and the discrimination error of the Chai Pae key component band is given a weight of 2 times;

[0027] construct a generative adversarial network based on the generator and the discriminator, and train the generative adversarial network using data in the database of red peony root quality;

[0028] construct an encoder, wherein the encoder includes multiple convolutional layers to perform convolutional operations on the input noisy spectral data, gradually reducing the feature map size and extracting data features;

[0029] construct a decoder, wherein the decoder is strictly symmetrical to the encoder to restore the spectral dimension layer by layer and ensure that the receptive field of each convolutional layer matches the encoder;

[0030] construct a convolutional autoencoder based on the encoder and the decoder, and train the convolutional autoencoder using data in the database of red peony root quality; wherein the convolutional autoencoder further includes an output layer Tanh activation to map the reconstructed spectrum to [-1, 1] and ensure consistency with the output range of the convolutional autoencoder;

[0031] fuse the trained generative adversarial network and the convolutional autoencoder to obtain a trained hybrid denoising model; wherein the hybrid denoising model first performs preliminary denoising on the near-infrared hyperspectral data through the generative adversarial network, and then performs secondary denoising through the convolutional autoencoder.

[0032] Further, the quality prediction model based on Transformers is constructed, and the quality prediction model is trained using data in the database of red peony root quality to obtain a trained quality prediction model, comprising:

[0033] construct a quality prediction model based on Transformers, which includes an input layer, a Transformer module, and a multi-task output layer; wherein the Transformer module is composed of multiple Transformer blocks, and each Transformer block includes a multi-head attention layer, a feedforward neural network layer, and a layer normalization layer;

[0034] divide the data in the database of red peony root quality into a training set, a validation set, and a test set according to a predetermined ratio, determine the corresponding red peony root component content label, origin label, and artificial marked quality grading label of the data, normalize the red peony root component content label, use one-hot encoding for the origin label to convert each origin into a unique vector representation, and encode the artificial marked quality grading label to obtain a grading ordinal vector;

[0035] define a loss function, use a mean square error loss function for component content prediction, use a cross-entropy loss function for origin prediction, and use a cross-entropy loss function for grading prediction;

[0036] training the quality prediction model using an optimizer to obtain a trained quality prediction model.

[0037] Further, the quality prediction model based on Transformers is constructed, and the quality prediction model comprises an input layer, a Transformer module, and a multi-task output layer, wherein the Transformer module is composed of a plurality of Transformer blocks stacked together, and each Transformer block comprises a multi-head attention layer, a feedforward neural network layer, and a layer normalization layer, and the input layer directly takes the preprocessed near-infrared hyperspectral data as input to retain complete spectral information.

[0038] The input layer directly takes the preprocessed near-infrared hyperspectral data as input to retain complete spectral information.

[0039] The multi-head attention layer decomposes the input feature dimension into a plurality of subspaces, and uses different sub-controls to focus, gather and cover different red peony root component peak regions, so that the quality prediction model can capture long-distance spectral band correlation in parallel to capture the collaborative features of the component peaks; wherein the multi-head attention mechanism realizes the cross-modal association between the spectral features and the quality labels through conditional input: in generating the query vector, in addition to the spectral data, the origin one-hot encoding and the hierarchical ordinal vector are synchronously input, the modal fusion is realized through a weight matrix, and then the attention score is calculated; secondly, the output of the multi-head attention layer is obtained by concatenating the multi-subspace attention results and then performing linear transformation;

[0040] The feedforward neural network layer processes the output of the multi-head attention layer, and the feedforward neural network comprises two fully connected layers, and a ReLU activation function is used in the middle.

[0041] The layer normalization layer is used before and after the multi-head attention layer and the feedforward neural network layer.

[0042] The multi-task output layer converts the output of the Transformer module into a fixed-length vector through global average pooling to retain the global statistical features of the spectral data, and then performs red peony root component content prediction, origin prediction and hierarchical prediction through a fully connected layer; wherein a linear regression output is used for component content prediction, a softmax function is used for origin prediction to output the prediction probability of each origin, and a softmax function is used for hierarchical prediction to output the prediction probability of each hierarchical level.

[0043] In a second aspect, the embodiment of the present application further provides a grading system for quality detection of red peony root decoction pieces, which comprises a grading device, a near-infrared hyperspectral camera, a light source, a sample disc and a conveyor belt; the grading device comprises:

[0044] The collection module is configured to collect near-infrared hyperspectral data of red peony root samples from different origins, and introduce an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage, so that the trained model can identify spectral characteristics of regions containing more active ingredients.

[0045] The association module is configured to establish a red peony root quality database based on the near-infrared hyperspectral data, and associate data in the red peony root quality database with corresponding red peony root component content labels, origin labels, and artificial marking quality grading labels.

[0046] The fusion module is configured to fuse a generative adversarial network and a convolutional autoencoder to obtain a hybrid denoising model, and train the hybrid denoising model using data in the red peony root quality database to obtain a trained hybrid denoising model.

[0047] The construction module is configured to construct a quality prediction model based on Transformers, and train the quality prediction model using data in the red peony root quality database to obtain a trained quality prediction model.

[0048] The acquisition module is configured to acquire near-infrared hyperspectral data of a red peony root sample to be tested, and input the near-infrared hyperspectral data of the red peony root sample to be tested into the trained hybrid denoising model to obtain preprocessed near-infrared hyperspectral data.

[0049] The grading module is configured to input the preprocessed near-infrared hyperspectral data into the trained quality prediction model for grading prediction.

[0050] In a third aspect, an embodiment of the present application also provides a terminal, comprising:

[0051] one or more processors;

[0052] a storage device configured to store one or more programs;

[0053] a display configured to display results;

[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement the grading method for red peony root sample quality detection as described above.

[0055] In a fourth aspect, the present application also provides a storage medium containing computer executable instructions for executing the grading method for red peony root sample quality detection as described above when executed by a computer processor.

[0056] Compared with the prior art, the grading method for red peony root sample quality detection, system, terminal, and medium of the present application have the following advantages:

[0057] (1) The grading method, system, terminal and medium for detecting the quality of red peony root decoction pieces, specifically a red peony root decoction piece grading method combining near-infrared hyperspectral imaging technology and deep learning algorithm, which realizes accurate prediction of the ingredient content and origin of red peony root decoction pieces and quality grading by combining a Transformer architecture through constructing a generative adversarial network (GAN) and a convolutional autoencoder (CAE) hybrid denoising model based on a full convolutional neural network (FCN). The application of this research and development technology can effectively promote the standardization and modernization process of the traditional Chinese medicine industry and provide protection for improving the market competitiveness of traditional Chinese medicinal materials.

[0058] (2) The grading method, system, terminal and medium for detecting the quality of red peony root decoction pieces also provide an efficient grading method combining an attention mechanism and a hybrid deep learning architecture, which realizes integrated nondestructive testing of origin traceability, quantitative determination of ingredient content and quality grading by combining near-infrared hyperspectral imaging with a Transformer network, solves the noise interference and information redundancy problems of spectral analysis of complex traditional Chinese medicine systems, and establishes a standardized detection process. Therefore, the present application solves the problem of damage to samples by traditional detection methods, further protects the integrity of traditional Chinese medicinal materials, and is suitable for large-scale promotion and use due to its nondestructive characteristics. BRIEF DESCRIPTION OF DRAWINGS

[0059] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings in which:

[0060] Figure 1 A flowchart of a grading method for detecting the quality of red peony root decoction pieces according to embodiment one of the present application;

[0061] Figure 2 A structure diagram of a grading system for detecting the quality of red peony root decoction pieces according to embodiment two of the present application;

[0062] Figure 3 A structure diagram of a grading device in a grading system for detecting the quality of red peony root decoction pieces according to embodiment two of the present application;

[0063] Figure 4 A structure diagram of a face fake image detection terminal according to embodiment three of the present application. DETAILED DESCRIPTION

[0064] The application will be described in further detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are intended for explanation only and are not limiting of the application. In addition, it should be noted that only the parts related to the application are shown in the drawings for ease of description.

[0065] Embodiment one

[0066] Figure 1 A flowchart of a grading method for detecting the quality of red peony root slices according to embodiment one of the application. The method described in this embodiment can be used for detecting and grading red peony root slices. The application proposes a complete scheme based on attention mechanism data collection, FCN-GAN and CAE hybrid denoising, and Transformer multi-task prediction for the near-infrared spectral data characteristics of red peony root slices, filling the gap in the detection of multiple quality indicators of traditional Chinese medicine slices in the prior art. See Figure 1 The grading method specifically includes the following steps:

[0067] Step 101: Collecting near-infrared hyperspectral data of red peony root slice samples from multiple different origins, and introducing an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage to enable the trained model to identify spectral characteristics of regions containing more active ingredients.

[0068] Specifically, red peony root slices from Sichuan, Inner Mongolia, Shaanxi, Shanxi, Gansu and other origins are collected, batch numbers and origins are recorded and then packaged, and then hyperspectral data is collected. Subsequently, an attention mechanism in deep learning is introduced during the near-infrared hyperspectral data collection stage. When collecting near-infrared hyperspectral images of red peony root slices, the model can automatically focus on the key parts of the red peony root slices according to the pre-trained feature patterns. By analyzing the distribution characteristics of active ingredients of red peony root slices from different origins, the trained model can identify spectral characteristics of regions containing more active ingredients. During the collection process, more collection resources are allocated to these key regions.

[0069] Step 102: Establishing a red peony root quality database based on the near-infrared hyperspectral data, and associating corresponding red peony root ingredient content labels, origin labels and artificial labeled quality grading labels with the data in the red peony root quality database.

[0070] Specifically, the red peony root quality database is established, including the appearance characteristics, physical property data, and multiple ingredient content data of red peony root slices from different origins, for training a quality prediction model.

[0071] Optionally, based on the quality data, the quality grading standard is specified as follows: first grade: the production area is high-quality, the content of paeoniflorin is ≥5%, the content of paeoniflorin lactone is ≥0.5%, and the content of other harmful ingredients is lower than the specified threshold; second grade: the production area is medium, the content of paeoniflorin is 3%-5%, the content of paeoniflorin lactone is 0.3%-0.5%, and each index meets the basic pharmaceutical requirements; third grade: the production area is general, the content of paeoniflorin is <3%, the content of paeoniflorin lactone is <0.3%, or there is a certain impurity exceeding the standard. For example, when manually marking the quality grading label, the following table can be referred to:

[0072]

[0073]

[0074] Step 103, fuse the generated adversarial network and the convolutional autoencoder to obtain a hybrid denoising model, and train the hybrid denoising model using data in the red peony root quality database to obtain a trained hybrid denoising model.

[0075] Specifically, based on the deep learning hybrid denoising preprocessing, the generative adversarial network (GAN) and the convolutional autoencoder (CAE) are fused to utilize the advantages of both to denoise the near-infrared hyperspectral data. GAN generates noise-free data through generator and discriminator adversarial learning, and CAE uses convolutional layers to automatically extract data features to achieve self-encoding denoising. Combining the two can more effectively remove noise, retain key spectral information, and improve data quality.

[0076] In actual application process, a large amount of near-infrared hyperspectral data collected by attention mechanism can be collected, including spectral information of red peony root slices of different origins and batches. The data is labeled to mark noise level, sample origin, batch, and other key information. The data is divided into a training set (70%), a validation set (15%), and a test set (15%). The training set is used for model training, the validation set monitors the model training process to prevent overfitting, and the test set evaluates the model performance.

[0077] Step 104, construct a quality prediction model based on Transformers, and train the quality prediction model using data in the red peony root quality database to obtain a trained quality prediction model.

[0078] Specifically, the pre-processed data needs to be further arranged. It is divided into a training set, a validation set, and a test set, and the division ratio can be set to 70%, 15%, and 15%. The corresponding red peony root component content labels (such as paeoniflorin and paeoniflorin lactone content), origin labels (such as Sichuan and Inner Mongolia origin labels), and manually marked quality grading labels (which can be developed according to component content, origin, and traditional quality standards, such as first grade, second grade, etc.) are determined.

[0079] Step 105, acquire the near-infrared hyperspectral data of the to-be-tested red root of radix paeoniae and input the near-infrared hyperspectral data of the to-be-tested red root of radix paeoniae into the mixed denoising model after training to obtain the preprocessed near-infrared hyperspectral data.

[0080] In actual application, the mixed denoising model includes GAN rough denoising and CAE fine repair, and the data after GAN preliminary denoising is input into the trained CAE for secondary denoising. GAN fits the real spectral distribution through adversarial learning, solving the global mapping of "noise-noise".

[0081] CAE focuses on the local spectral band of active ingredients (such as the range of 2000 nm ± 50 nm) by using the 32-dimensional bottleneck features of the encoder, and further suppresses the scattering noise through pixel-level reconstruction of transposed convolution, so that the signal-to-noise ratio of the key area is above 35 dB. After CAE processing, further optimized denoising spectral data is obtained. This mixed denoising process fully utilizes the advantages of GAN and CAE, GAN can preliminarily remove noise and generate approximate real spectral data, CAE can fine denoising and reconstruction based on data features, and the combination of the two can more effectively remove noise. When training jointly, the CAE loss is back-propagated to the GAN, forming a "generation-reconstruction-feedback" closed loop, and the key spectral information related to the quality of red root of radix paeoniae is retained.

[0082] Step 106, input the preprocessed near-infrared hyperspectral data into the trained quality prediction model for classification prediction.

[0083] Preferably, in the above step 101, near-infrared hyperspectral data of red root of radix paeoniae samples from multiple different producing areas is collected, and an attention mechanism in deep learning is introduced in the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral features of the region containing more active ingredients, which specifically includes the following steps:

[0084] Step 1011, acquire red root of radix paeoniae samples from multiple different producing areas, record batches and producing areas and perform near-infrared hyperspectral data collection after performing split processing, to obtain near-infrared hyperspectral images.

[0085] In actual application, red root of radix paeoniae from Sichuan, Inner Mongolia, Shaanxi, Shanxi and Gansu can be collected, batched, and spectral images are collected using a near-infrared hyperspectral camera (such as InnoSpectra VNIR-1000, wavelength range 900-1700 nm) with a resolution of 64x64 pixels / pill.

[0086] Step 1012, process the near-infrared hyperspectral image using the YOLOv8 target detection algorithm to automatically identify the outline of the red root of radix paeoniae and delineate the region of interest to obtain a feature map.

[0087] Specifically, the YOLOv8 target detection algorithm can be used to automatically identify the outline of the red peony root decoction piece, to delimit a 64x64 pixel region of interest (ROI), to exclude background interference, and to position the accuracy rate ≥99%.

[0088] In actual application, first, 5000 images of red peony root decoction pieces can be collected, covering five major producing areas such as Sichuan and Inner Mongolia, including different forms such as slices (thickness 1-3 mm) and segments (length 1-5 cm), with a resolution of 1280x720 pixels, fixed light source (6500K white light) and focal length (30 cm) during collection to reduce illumination and viewing angle deviation. The open source tool LabelImg is used to manually label the outline of the decoction piece, and a Pascal VOC format data set is generated, with the label category being "red peony root decoction piece" and the annotation box adhering to the edge of the decoction piece (error ≤2 pixels).

[0089] Then, the red peony root decoction piece images are randomly rotated (±15°), scaled (0.8-1.2 times), and horizontally flipped to enhance the model's robustness to different angles of placement. Gaussian noise (σ=0.05) is added, and brightness / contrast (±20%) is adjusted to simulate complex lighting environments. The training set has 4000 images, the verification set has 500 images, and the test set has 500 images, with a class balance rate of 1:1.

[0090] After that, the YOLOv8n lightweight model (with about 3.5M parameters) is used, which balances accuracy and speed, and the industrial-grade GPU (RTX4090) inference speed ≥100FPS, meeting the real-time detection requirements. For small decoction piece targets (average pixel area ratio 15%-25%), the neck feature fusion path is adjusted, the shallow feature weight is enhanced, and the small target outline recognition ability is improved. Anchor box parameter retraining: based on the decoction piece annotation data, the anchor box prior is calculated, the initial anchor box size is optimized (average width-to-height ratio 1.2:1, suitable for elliptical / irregular shapes of decoction pieces), and the positioning regression error is reduced. Loss function adjustment: strengthen the weight ratio of classification loss (CE Loss) and positioning loss (CIoU Loss) (3:2), focus on optimizing the boundary box adhesion, and the sample proportion of CIoU ≥0.95 is increased to 98%. AdamW (weight decay 0.0005), learning rate cosine annealing (initial 1e-3, decay to 1e-5), train for 300 epochs, and the mAP@0.5 of the verification set is stable at 99.2%.

[0091] Finally, ROI positioning is performed. The specific implementation steps of ROI positioning are as follows: input image is resized to 640*640 pixels, the aspect ratio is kept, black pixels (RGB=0,0,0) are filled at the edges to avoid distortion, and pixel values are normalized to [0,1]. The backbone network (CSPDarknet) outputs 3 layers of feature maps (80*80, 40*40, 20*20), which correspond to detection of targets of different scales, respectively. The neck (PAFPN) fuses multi-scale features, and the detection head outputs the coordinates (x,y,w,h) of the boundary box of the medicinal piece. The predicted box with a confidence of greater than or equal to 0.95 is subjected to post-processing. The threshold settings are as follows: the IoU threshold is 0.5, the confidence threshold is 0.9, overlapping boxes are filtered, and only one optimal box is kept for a single medicinal piece. The 640*640 predicted box is mapped back to the original image coordinates, the center coordinates (xc,yc) are calculated, and a 64*64 pixel ROI is drawn with the center coordinates as the center, so that the coverage rate of the medicinal piece body is greater than or equal to 95%.

[0092] In step 1013, multi-scale feature extraction is performed on the feature map by using a CNN, and attention weight calculation is performed to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a red peony root component.

[0093] Specifically, first, multi-scale feature extraction is performed on the feature map by using a CNN; wherein the convolutional layer: 3*3 and 5*5 convolutional kernels are used in parallel to extract features, the 3*3 kernel convolution is used to capture microscopic textures, and the 5*5 convolutional kernel is used to capture macroscopic spectral trends; second, the pooling layer: maximum pooling and average pooling are used alternately; then the global feature vector is obtained by global average pooling, and the global feature vector is input into two fully connected layers to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a red peony root component.

[0094] In actual application, CNN feature extraction: (1) Convolutional layer: 3*3 and 5*5 convolutional kernels are used in parallel to extract features, 3*3 kernels capture microscopic textures (such as cell structure), and 5*5 kernels capture macroscopic spectral trends (such as component distribution gradient). (2) Pooling layer: maximum pooling (window 3*3, retaining peak features) and average pooling (window 3*3, smoothing baseline drift) are used alternately, and a 128-dimensional feature map is output.

[0095] The operation formula of the pooling layer is as follows: the pooling layer is used to reduce the dimension of data and reduce the amount of calculation. Taking maximum pooling as an example, assuming that the input feature map is X and the output feature map is Y, and the pooling window size is s*s, the calculation formula of maximum pooling is as follows:

[0096]

[0097] where (i, j) is the coordinate of the element in the output feature map Y, and the formula means that the maximum value of the elements in the window is taken as the value of the output feature map Y at each position (i, j) on the input feature map X with a pooling window of size s x s sliding at each position. The average pooling is to calculate the average value of the elements in the window as the output, and the formula is:

[0098]

[0099] The pooling operation effectively reduces the data dimension while retaining important feature information, enabling the model to process data faster and reducing the risk of overfitting.

[0100] Attention weight calculation: (1) Global average pooling: calculate the channel-level feature vector for the ROI feature map (H x W x C, C = 288 bands) Get z e R 288 (2) Fully connected layer weighting: first layer (ReLU activation): x1 = ReLU (W1z + b1); second layer (Sigmoid activation): a c = Sigmoid (W2x1 + b2), generating attention weight a c e [0, 1], the weight of the key component band (such as 2000-2300 nm paeoniflorin peak) is greater than or equal to 0.8, and the weight of the non-characteristic band is less than or equal to 0.3.

[0101] In attention calculation, first, the global feature vector is obtained by global average pooling. The input feature map is F, with a size of H x W x C (H is the height, W is the width, and C is the number of channels), and the global average pooling calculation is:

[0102]

[0103] where z c is the value of the cth channel in the global feature vector. Through global average pooling, the feature map of each channel is compressed into a value, obtaining a global feature vector z with a length of C. Then, the global feature vector z is input into two fully connected layers. The weight of the first fully connected layer is W1, the bias is b1, and the output is x1; the weight of the second fully connected layer is W2, the bias is b2, and the final output is x2. Then:

[0104] x1 = σ1 (W1z + b1)

[0105] x2 = σ2 (W2x1 + b2)

[0106] where σ1 and σ2 are activation functions, the activation function of the first fully connected layer can be selected as ReLU function, and the activation function of the second fully connected layer uses Sigmoid function. That is:

[0107] σ1(x) = max(0,x)

[0108]

[0109] Finally, the attention weight 'a' for each feature channel is obtained. c ,Right now:

[0110] a c =x2(c)

[0111] These weights represent the degree of attention the model pays to different feature channels and are used for subsequent data collection.

[0112] During data acquisition, the acquisition resources are adjusted according to attention weights. Originally, the sampling frequency for the entire image was f0, but for an attention weight of a... c For the region, the adjusted sampling frequency is f, which can be adjusted using the following formula:

[0113] f = f0 + α × a c

[0114] Where α is a scaling factor used to control the adjustment range of the sampling frequency. This formula represents the adjustment based on the attention weight α. c The size of the sample is increased to improve the sampling frequency of key parts, so that more data can be collected from key parts, thereby obtaining richer spectral information.

[0115] Step 1014: Adjust the acquired resources according to the attention weight, dynamically increase the image sampling frequency and spectral signal intensity of key areas based on the attention weight, and perform weighted fusion of key area data and non-key area data based on the attention weight to obtain near-infrared hyperspectral data.

[0116] Specifically, firstly, the sampling frequency of key regions is dynamically increased according to the attention weight, so that the sampling frequency of active ingredient-rich regions is increased by at least 60% and the spectral signal intensity is increased by at least 45%; then, the data of key regions and non-key regions are fused according to the attention weight to improve the spectral signal-to-noise ratio of effective ingredients to above 25dB.

[0117] In practical applications, according to weight a c Dynamically increasing the sampling frequency in key areas increases the sampling frequency in areas rich in active ingredients (such as the phloem) by 60% and the spectral signal intensity by 45%, as shown by the formula: f = f0 + α × a c The key area data D1 and non-key area data D2 are merged according to their weights, using the formula: D = a c ×D1+(1-a c )×D2, improves the spectral signal-to-noise ratio of active ingredients to over 25dB.

[0118] Preferably, the above step 103, the fusion generates a hybrid denoising model of the generative adversarial network and the convolutional autoencoder, and the hybrid denoising model is trained by using data in the red peony root quality database to obtain a trained hybrid denoising model, and specifically includes the following steps:

[0119] Step 1031, constructing a generator based on a full convolutional neural network structure, inputting a random noise vector and outputting spectral data of the same dimension as the near-infrared hyperspectral data; wherein the generator includes multiple transpose convolutional layers, gradually expands the feature map size through multiple transpose convolutional layer operations, and generates a signal similar to the original spectrum; secondly, the generator also includes a hidden layer ReLU activation to enhance the nonlinear mapping to capture the nonlinear changes of the characteristic peaks of the red peony root components; thirdly, the generator also includes an output layer Tanh activation to map the spectral values to [-1, 1] to adapt to the normalized real spectral range.

[0120] In actual application, a full convolutional neural network (FCN) structure is adopted, the input is a random noise vector, and the output is spectral data of the same dimension as the near-infrared hyperspectral data. The generator is composed of multiple transpose convolutional layers, which gradually expand the feature map size through transpose convolutional operations to generate a signal similar to the original spectrum. For example, 4 layers of transpose convolution (kernel 4x4, step 2) are adopted to recover the spectral dimension layer by layer (100-dimensional noise→256→128→64→288 bands), and the embodiment outputs a 288-dimensional spectral vector, which directly matches the dimension of the near-infrared spectral data. The hidden layer ReLU activation enhances the nonlinear mapping to capture the nonlinear changes of the characteristic peaks of the paeoniflorin, lactone glycosides and other components (such as the second derivative characteristics at 2000 nm). The output layer Tanh activation maps the spectral values to [-1, 1] to adapt to the normalized real spectral range (the traditional Chinese medicine spectrum is distributed in [-0.8, 0.8] after pretreatment, with a mean value ±1σ).

[0121] Step 1032, embedding spectral prior constraints between each transpose convolutional layer in the generator to force the generated spectral data to have peak shapes consistent with the real data in the characteristic regions of the red peony root components, and ensuring the physical interpretability of the generated spectral data through a gradient penalty term to suppress the generation of meaningless noise and achieve accurate recovery of the spectral dimension and strengthening of the component characteristics.

[0122] For example, the spectral prior constraints are embedded between the transpose convolutional layers: the generated spectrum is forced to have peak shapes consistent with the real data in the 2000-2300 nm (paeoniflorin characteristic region) and 1650-1750 nm (lactone glycoside characteristic region), the physical interpretability of the generated spectrum is ensured through a gradient penalty term (Gradient Penalty) to suppress the generation of meaningless noise, thereby achieving accurate recovery of the spectral dimension and strengthening of the component characteristics.

[0123] Step 1033, constructing a discriminator based on a full convolutional neural network structure, the input being spectral data generated by the generator or real clean spectral data, and the output being a scalar representing the probability that the input data is real data; wherein the generator includes multiple convolutional layers, the activation function using LeakyReLU in the hidden layer and Sigmoid in the output layer, and the features being extracted step by step by the multiple layers of convolution; secondly, the loss function optimization adopts a binary cross-entropy with weights, and the discrimination error of the key ingredient waveband of red-leaf radix paeoniae is given a weight of 2 times.

[0124] For example, the discriminator is constructed based on a full convolutional neural network, the input being spectral data generated by the generator or real clean spectral data, and the output being a scalar representing the probability that the input data is real data. The discriminator includes multiple convolutional layers, the activation function using LeakyReLU in the hidden layer and Sigmoid in the output layer, and the features being extracted step by step by the three layers of convolution (kernel 3x3, LeakyReLU activation).

[0125] The first layer (64 cores): capturing the starch-based spectral features in the 1000-1500 nm range, and distinguishing between the decoction pieces and the counterfeits;

[0126] The second layer (128 cores): focusing on the active ingredient waveband in the 1650-2300 nm range, and identifying the peak shift caused by the difference in the place of production (for example, the peak intensity at 2050 nm of red-leaf radix paeoniae produced in Sichuan is 15% higher than that produced in Inner Mongolia);

[0127] The third layer (256 cores): global feature fusion, outputting a discrimination probability of 0-1, and the identification accuracy of the baseline drift noise reaching 92%.

[0128] Loss function optimization: adopting a binary cross-entropy with weights: the discrimination error of the key ingredient waveband (such as 2000 nm) is given a weight of 2 times, and the formula is as follows:

[0129]

[0130] where w c is the channel attention weight, which strengthens the discrimination ability of the active ingredient region, improves the discrimination accuracy compared with the traditional equal-weight loss, and realizes the specific discrimination of traditional Chinese medicine noise.

[0131] Step 1034, constructing a generative adversarial network based on the generator and the discriminator, and training the generative adversarial network by using the data in the red-leaf radix paeoniae quality database.

[0132] In practical application, the training process is as follows: the real data contains original spectra with baseline drift and scattering noise (simulating the actual acquisition noise), the generated data is the noise-free spectrum initial value generated by the noise vector, and the discriminator is forced to learn the mapping relationship of the "noise-noise-free" spectrum pair, which is different from the limitation of the prior art which only uses clean data for training.

[0133] The generator and the discriminator are trained at the same time. The generator tries to generate spectrum data that makes the discriminator misjudge as real data, and the discriminator tries to distinguish real data from generated data. A batch of noisy spectrum data and real clean spectrum data are randomly extracted in the training set, and the loss of the generator and the discriminator is calculated. The generator loss adopts a binary cross-entropy loss function, and the formula is:

[0134]

[0135] Wherein, N is the batch size, z i is a random noise vector, G(z i ) is the spectrum data generated by the generator, and D(G(z i )) is the judgment probability of the discriminator for the generated data.

[0136] Step 1035, constructing an encoder, wherein the encoder includes a plurality of convolutional layers to perform convolution operation on the input noisy spectrum data, gradually reducing the feature map size and extracting data features.

[0137] Specifically, the encoder (Encoder) is constructed as follows: composed of a plurality of convolutional layers, performing convolution operation on the input noisy spectrum data, gradually reducing the feature map size and extracting data features. The convolution kernel size, step and padding of the convolutional layer are adjusted according to the data characteristics. A 3-layer 5x5 convolution kernel (step 2, padding 2) is adopted, which is different from the 3x3 kernel of the traditional CAE, and is suitable for the wide range of characteristics of traditional Chinese medicine spectrum (such as the wide peak of paeoniflorin at 2000nm): the first layer (128 kernels): extracting the basic characteristics of the full wave band of 1000-2500nm, capturing the baseline trend; the second layer (64 kernels): focusing on the active ingredient zone of 1600-2200nm, strengthening the feature peak separation of lactone glycosides and paeoniflorin; the third layer (32 kernels): compressed to 32-dimensional bottleneck features, retaining the position (such as peak position) and intensity (such as peak height) information of key spectral bands, and removing the low-frequency interference of irrelevant components such as starch. And through the ReLU activation advantage to avoid the gradient saturation of Sigmoid, ensure the details of the weak absorption peak such as 1650nm are retained, and the peak shape recognition is improved compared with the Tanh activation of the prior art.

[0138] Step 1036, constructing a decoder, wherein the decoder is strictly symmetrical with the encoder to realize the restoration of the spectrum dimension layer by layer and ensure that the receptive field of each convolutional layer matches the encoder.

[0139] Specifically, the decoder is constructed: strictly symmetrical with the encoder, 3-layer transposed convolution (kernel 5x5, stride 2, padding 2) reduces the spectral dimension layer by layer (32→64→128→288 bands), ensures that the receptive field of each convolution layer matches the encoder, and avoids feature misplacement (such as the 2000nm peak position error≤1nm after decoding). The output layer Tanh activation: maps the reconstructed spectrum to [-1, 1], consistent with the output range of the GAN, facilitating cascading processing, and more suitable for the normalization characteristics of traditional Chinese medicine spectra than the linear output of the prior art.

[0140] Step 1037, based on the encoder and decoder, a convolutional autoencoder is constructed, and the convolutional autoencoder is trained using data in the red peony root quality database; wherein the convolutional autoencoder further comprises an output layer Tanh activation to map the reconstructed spectrum to [-1, 1] and ensure consistency with the output range of the convolutional autoencoder.

[0141] In actual application, the training process is: taking noisy spectral data as input and original clean spectral data as target output, using mean square error (MSE) loss function to train CAE, the formula is:

[0142]

[0143] wherein, x i is the reconstructed spectral data output by CAE. The parameters of the encoder and the decoder are adjusted by the back propagation algorithm to minimize the error between the reconstructed data and the original clean data. In the training process, the model performance is evaluated using the validation set, and the hyperparameters are adjusted according to the validation set loss to prevent overfitting. An additional component peak constraint term is added:

[0144] Step 1038, fuse the trained generative adversarial network and convolutional autoencoder to obtain a trained hybrid denoising model; wherein the hybrid denoising model first denoises the near-infrared hyperspectral data through the generative adversarial network, and then denoises the data through the convolutional autoencoder.

[0145] In actual application, the trained hybrid denoising model can also be evaluated, and the specific evaluation process is: using a test set to evaluate the performance of the hybrid denoising model. Peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are used as evaluation indicators. The PSNR calculation formula is:

[0146]

[0147] wherein, is the maximum pixel value of the spectral data, MSE is the mean square error of the data after denoising and the original clean data. The higher the PSNR value, the closer the data after denoising is to the original clean data. The SSIM calculation formula is:

[0148]

[0149] wherein μ x , μ y are the mean values of x and y respectively, are the standard deviations of x and y respectively, σ xy is the covariance of x and y, and C1 and C2 are constants to avoid zero denominator. The closer the SSIM value is to 1, the higher the structural similarity between the data after denoising and the original clean data. By comparing the PSNR and SSIM values of the data before and after denoising, the denoising effect of the mixed denoising model is evaluated to determine whether the model meets the expected performance requirements.

[0150] Preferably, the step 104 of constructing a quality prediction model based on Transformers, and training the quality prediction model using the data in the red-processed radix paeoniae quality database to obtain a trained quality prediction model, specifically includes the following steps:

[0151] Step 1041, constructing a quality prediction model based on Transformers, the quality prediction model comprising an input layer, a Transformer module, and a multi-task output layer; wherein the Transformer module is composed of multiple Transformer blocks, and each Transformer block contains a multi-head attention layer, a feedforward neural network layer, and a layer normalization layer.

[0152] The above input layer: directly takes the pre-processed near-infrared hyperspectral data as input to retain complete spectral information. Specifically, the processed near-infrared hyperspectral data (288-dimensional spectral vector) can be directly taken as input to retain complete spectral information and avoid information loss caused by manual feature selection in traditional methods.

[0153] The above multi-head attention layer: decomposes the input feature dimension into multiple subspaces, and uses different sub-controls to focus, gather and cover different red-processed radix paeoniae component peak regions, so that the quality prediction model can capture long-distance spectral band correlation in parallel to capture the collaborative features of component peaks; wherein the multi-head attention mechanism realizes the cross-modal association of spectral features and quality labels through conditional input: when generating the query vector, in addition to the spectral data, the origin one-hot encoding and the hierarchical ordinal vector are also input synchronously, the modal fusion is realized through the weight matrix, and then the attention score is calculated; secondly, the output of the multi-head attention layer is obtained by concatenating the multiple subspace attention results and then performing linear transformation.

[0154] Because in traditional hyperspectral data processing, models such as CNN are limited by local convolution kernels and can only capture local features in adjacent wavelength regions (such as spectral band changes within a 10 nm range), while the effective components of Chishao decoction pieces (such as paeoniflorin and lactone glycosides) have characteristic peaks distributed in non-continuous wavelength regions (such as 350 nm apart between 2000 nm and 1650 nm), and long-distance dependent modeling is needed to reveal their collaborative variation rules.

[0155] Therefore, the present embodiment breaks through the input feature dimension d model = 512 is divided into 8 subspaces (h = 8, each head dimension d k = d v = 64). The first and second heads focus on the 2000 ± 50 nm paeoniflorin main peak region, capturing the variation characteristics of its intensity and width; the third and fourth heads focus on the 1650 ± 30 nm lactone glycoside shoulder peak region, identifying subtle shifts related to the origin (such as the peak of Sichuan-produced decoction pieces being 5 nm to the left); the fifth to eighth heads cover the 1000-1500 nm starch-based spectrum, suppressing irrelevant noise interference with active ingredients. This division allows the model to capture long-distance spectral band correlations of more than 300 nm in parallel, improving the ability to capture component peak collaborative features compared to single attention head global average modeling. Among them, the number of heads is h, the input feature dimension is d model , and the dimension of each head is d

[0156] In actual application, unlike traditional Transformers that only process single modal data, the multi-head attention mechanism in the present embodiment achieves cross-modal association of spectral features and quality labels through conditional input: when generating the query vector Q, in addition to spectral data, the origin one-hot encoding and hierarchical ordinal vector are also input simultaneously, and modal fusion is achieved through the weight matrix W q . For example, when processing samples from Sichuan production areas, the attention weight of the first head will automatically strengthen the association between the 2000 nm peak and the "high-quality production area" label, forming a "production area-component" exclusive mapping path. Through end-to-end training, the weight matrix W k , W v adaptively adjusts the contribution of different wavelengths to the quality index. Experiments show that after training, the attention weight of the 2000 nm paeoniflorin peak and the origin label is higher than that of random initialization, directly associating high content characteristics of high-quality production areas. For the input feature vector x, the query vector Q, the key vector K, and the value vector V are obtained through linear transformation, i.e. Q = W q x, K = W k x, V = W v x, where W q , W k , and W v are learnable weight matrices.

[0157] Then the attention score is calculated: The output of multi-head attention is the concatenation of the h attention results followed by a linear transformation: MultiHead(Q, K, V) = W o ([head1;... ; head h ]), where head i is the attention result of the i-th head, W o is the weight matrix of the linear transformation.

[0158] The above feed-forward neural network layer: processes the output of the multi-head attention layer, and the feed-forward neural network includes two fully connected layers with ReLU activation function in between.

[0159] In actual application, the feed-forward neural network layer: processes the output of the multi-head attention, and the feed-forward neural network includes two fully connected layers with ReLU activation function in between. The formula is

[0160] FFN(z) = W2(ReLU(W1z + b1)) + b2

[0161] where W1, W2 are weight matrices, and b1, b2 are biases. This design enhances the hierarchical expression of spectral features, especially for the detailed features of weak absorption peaks (such as the shoulder peak of lactone glycoside at 1650 nm).

[0162] The above layer normalization layer: uses layer normalization before and after the multi-head attention layer and the feed-forward neural network layer.

[0163] Specifically, the layer normalization layer: uses layer normalization before and after the multi-head attention layer and the feed-forward neural network layer, and the formula is:

[0164]

[0165] where μ and σ 2 are the mean and variance of the sample feature dimension, γ and β are learnable parameters, and ∈ is a small constant to prevent the denominator from being zero.

[0166] The above multi-task output layer: converts the output of the Transformer module into a fixed-length vector through global average pooling, preserving the global statistical features of the spectral data, and then performs red sage root component content prediction, origin prediction and grading prediction through fully connected layers; among them, linear regression is used for component content prediction to output prediction values, softmax function is used for origin prediction to output prediction probabilities of each origin, and softmax function is used for grading prediction to output prediction probabilities of each grading.

[0167] In practical applications, the multi-task output layer transforms the Transformer module output into a fixed-length vector using global average pooling, preserving the global statistical characteristics of the spectral data (such as baseline trends and characteristic peak intensity distributions). Then, fully connected layers are used to perform component content prediction, origin prediction, and classification prediction, respectively. For component content prediction, linear regression is used to output the predicted values. For origin prediction, the softmax function is used to output the predicted probability of each origin, i.e. For graded prediction, the softmax function is also used to output the predicted probability for each grade, i.e. Among them W content W origin W grade and b content b origin b grade These are all learnable weights and biases.

[0168] Step 1042: Divide the data in the Paeonia lactiflora quality database into training set, validation set and test set according to a preset ratio, and determine the corresponding Paeonia lactiflora component content label, place of origin label and artificially marked quality grading label. Then, normalize the Paeonia lactiflora component content label, use one-hot encoding for the place of origin label to convert each place of origin into a unique vector representation, and encode the artificially marked quality grading label to obtain the grading ordinal vector.

[0169] In practical applications, normalization is used for ingredient content labels. A commonly used normalization formula is:

[0170]

[0171] Where x is the original component content value, x min and x max These are the minimum and maximum values ​​of the component's content in the training set, respectively. Origin labels use one-hot encoding, transforming each origin into a unique vector representation. Manually labeled quality grading tags are also encoded for model recognition.

[0172] Step 1043: Define the loss function. For component content prediction, use the mean squared error loss function; for origin prediction, use the cross-entropy loss function; and for grading prediction, use the cross-entropy loss function.

[0173] In practical applications, for component content prediction, the mean squared error (MSE) loss function is used, as shown in the following formula:

[0174]

[0175] Where N is the number of samples. is a predicted ingredient content value, is a true ingredient content value.

[0176] For origin prediction, the cross-entropy loss function is used, and the formula is as follows:

[0177]

[0178] where C ′ is the number of origin categories, is the true origin label, is the predicted origin probability.

[0179] For grading prediction, the cross-entropy loss function is also used, and the formula is as follows:

[0180]

[0181] where G is the number of grading categories, is the true grading label, is the predicted grading probability. The total loss function is L = L content + L origin + L grade .

[0182] Step 1044, training the quality prediction model using an optimizer to obtain a trained quality prediction model.

[0183] In actual application, the model is trained using an optimizer (such as the Adam optimizer). The steps for the Adam optimizer to update the parameters θ are as follows: first, calculate the gradient Then calculate the first-order moment estimate and the second-order moment estimate, m t = β1m t-1 + (1-β1)g t , Then correct the bias of the first-order moment estimate and the second-order moment estimate Finally, update the parameters where α ′ is the learning rate.

[0184] The grading method for quality detection of red peony root decoction pieces described in this embodiment, by collecting near-infrared hyperspectral data on red peony root decoction piece samples from multiple different origins, and introducing an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral characteristics of the region containing more active ingredients; based on the near-infrared hyperspectral data, a red peony root quality database is established, and the data in the red peony root quality database are associated with corresponding red peony root ingredient content labels, origin labels and artificial marked quality grading labels; a hybrid denoising model is obtained by fusing a generative adversarial network and a convolutional autoencoder, and the data in the red peony root quality database are used to train the hybrid denoising model to obtain the trained hybrid denoising model; a quality prediction model based on Transformers is constructed, and the data in the red peony root quality database are used to train the quality prediction model to obtain the trained quality prediction model; the near-infrared hyperspectral data of the red peony root decoction pieces to be tested are obtained, and the near-infrared hyperspectral data of the red peony root decoction pieces to be tested are input into the trained hybrid denoising model to obtain preprocessed near-infrared hyperspectral data; the preprocessed near-infrared hyperspectral data are input into the trained quality prediction model for grading prediction.

[0185] Compared with the prior art, the method described in this embodiment constructs a generative adversarial network (GAN) based on a full convolutional neural network (FCN) and a convolutional autoencoder (CAE) hybrid denoising model, and realizes accurate prediction of red peony root decoction piece ingredient content and origin and quality grading by combining a Transformer architecture.

[0186] The method described in this embodiment provides an efficient grading method that fuses an attention mechanism and a hybrid deep learning architecture to solve the problems of traditional red peony root decoction piece quality detection methods, such as complicated operation, strong destructiveness, and insufficient model generalization ability, etc., realizes integrated non-destructive detection of origin tracing, ingredient content quantitative determination and quality grading by combining near-infrared hyperspectral imaging with a Transformer network, and solves the noise interference and information redundancy problems of spectral analysis of complex Chinese medicine systems, and establishes a standardized detection process.

[0187] Embodiment Two

[0188] Figure 2 The structure diagram of the grading system for quality detection of red peony root decoction pieces described in Embodiment Two of the present application. Figure 2 A block diagram of an exemplary system suitable for implementing an embodiment of the present application is shown. Figure 2 The system shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present application, such as Figure 2As shown, the grading system for detecting the quality of red peony root slices includes a grading device 201, a near-infrared hyperspectral camera 202, a light source 203, a sample disc 204, and a conveyor belt 205.

[0189] In actual application, the red peony root slices to be detected can be placed in the sample disc 204 and moved to the lower side of the near-infrared hyperspectral camera 202 through the conveyor belt 205, and the light source 203 is used for irradiation in cooperation with the near-infrared hyperspectral camera 202 to collect near-infrared hyperspectral images, and finally transmitted to the grading device 201 for data processing and grading prediction.

[0190] Figure 3 FIG. 2 is a structural schematic diagram of the grading device in the grading system for detecting the quality of red peony root slices according to Embodiment Two of the present application, Figure 3 The device shown is only an example and should not limit the functions and use range of the embodiments of the present application, such as Figure 3 As shown, the grading device for detecting the quality of red peony root slices includes:

[0191] The acquisition module 2011 is configured to collect near-infrared hyperspectral data of red peony root slice samples from multiple different origins, and introduce an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral characteristics of the region containing more active ingredients.

[0192] The association module 2012 is configured to establish a red peony quality database based on the near-infrared hyperspectral data, and associate the data in the red peony quality database with corresponding red peony component content labels, origin labels, and artificial marking quality grading labels.

[0193] The fusion module 2013 is configured to fuse a generative adversarial network and a convolutional autoencoder to obtain a hybrid denoising model, and train the hybrid denoising model using data in the red peony quality database to obtain a trained hybrid denoising model.

[0194] The construction module 2014 is configured to construct a quality prediction model based on Transformers, and train the quality prediction model using data in the red peony quality database to obtain a trained quality prediction model.

[0195] The acquisition module 2015 is configured to acquire near-infrared hyperspectral data of the red peony root slices to be detected, and input the near-infrared hyperspectral data of the red peony root slices to be detected into the trained hybrid denoising model to obtain preprocessed near-infrared hyperspectral data.

[0196] The grading module 2016 is configured to input the preprocessed near-infrared hyperspectral data into the trained quality prediction model for grading prediction.

[0197] The grading system for quality detection of red peony root decoction pieces provided by the embodiment of the present application can execute the grading method for quality detection of red peony root decoction pieces provided by any embodiment of the present application, and has the function modules and beneficial effects corresponding to the execution method.

[0198] Embodiment three

[0199] Figure 4 A structural schematic diagram of a face fake image detection terminal provided for the third embodiment of the present application is shown in the figure; Figure 4 A block diagram of an exemplary terminal system suitable for implementing embodiments of the present application is shown. Figure 4 The terminal system shown is merely an example, and should not bring any limitation to the function and use range of the embodiments of the present application.

[0200] As shown in the figure, Figure 4 The terminal 12 is shown in the form of a general purpose computing device. The components of the terminal 12 can include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processing unit 16.

[0201] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus architectures including an industry standard architecture (ISA), micro-channel architecture (MAC), enhanced ISA (EISA), Video Electronics Standards Association (VESA) local bus, and a peripheral component interconnect (PCI) bus.

[0202] The terminal 12 typically includes a variety of computer system readable media. Such media can be any available media that is accessible by the terminal 12 and includes both volatile and non-volatile media, removable and non-removable media.

[0203] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The terminal 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be provided for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Figure 4 Not shown is typically called a "hard disk drive" (HDD). Although Figure 4A disk drive, a floppy disk drive, a CD-ROM drive, a DVD-ROM drive, or other removable media drive, a flash memory card drive, a multimedia arcade game drive, and / or a hard disk drive can be provided for reading from and writing to a removable, nonvolatile magnetic media (e.g., a "floppy disk"), and to a removable, nonvolatile optical media (e.g., an optical CD-ROM, DVD-ROM, or other optical media). In these instances, each drive can be connected to the system bus 18 by one or more data media interfaces. The memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.

[0204] The program / utility 40 having a set (at least one) of program modules 42 can be stored in memory 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, may

[0205] The terminal 12 can also be communicatively coupled to one or more external devices 14, such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with the terminal 12; and / or one or more devices that enable the terminal 12 to communicate with one or more other computing devices. Such communication can be via input / output (I / O) interfaces 22. Furthermore, the terminal 12 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the public network, such as the Internet, via network adapter 20. As depicted, network adapter 20 communicates with the other components of terminal 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with terminal 12. Such as, but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0206] The processing unit 16 performs various function applications and data processing by running programs stored in the system memory 28, such as implementing the grading method for detecting the quality of red peony root slices provided by the embodiments of the application.

[0207] Embodiment Four

[0208] The embodiment four of the application also provides a storage medium containing computer executable instructions, which when executed by a computer processor, are used to perform any of the grading methods for detecting the quality of red peony root slices provided by the above embodiments.

[0209] The computer storage medium of the embodiments of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.

[0210] The computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave, in which computer readable program code is embodied. Such propagated data signals can take a wide variety of forms, including but not limited to electro-magnetic signals, optical signals, or any suitable combination thereof. Computer readable signal medium can also be any computer readable medium that is not a storage medium, that is capable of storing the program for use by or in connection with the instruction execution system, apparatus, or device.

[0211] The program code embodied on the computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the above.

[0212] The computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0213] Note that the above merely describes preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.

Claims

1. A grading method for detecting the quality of red peony root decoction pieces, characterized in that, The application relates to a method for predicting the quality of Radix Paeoniae Rubra based on near-infrared hyperspectral data. The method comprises the following steps: near-infrared hyperspectral data of Radix Paeoniae Rubra samples from different origins are collected, and an attention mechanism in deep learning is introduced in the near-infrared hyperspectral data collection stage, so that the trained model can identify the spectral characteristics of the region containing more active ingredients; a Radix Paeoniae Rubra quality database is established based on the near-infrared hyperspectral data, and the data in the Radix Paeoniae Rubra quality database are associated with corresponding Radix Paeoniae Rubra ingredient content labels, origin labels and artificial marking quality grading labels; a hybrid denoising model is obtained by fusing a generative adversarial network and a convolutional autoencoder, and the hybrid denoising model is trained using the data in the Radix Paeoniae Rubra quality database to obtain a trained hybrid denoising model; a quality prediction model based on Transformers is constructed, and the quality prediction model is trained using the data in the Radix Paeoniae Rubra quality database to obtain a trained quality prediction model; near-infrared hyperspectral data of a to-be-tested Radix Paeoniae Rubra sample are obtained, and the near-infrared hyperspectral data of the to-be-tested Radix Paeoniae Rubra sample are input into the trained hybrid denoising model to obtain preprocessed near-infrared hyperspectral data; 2. The method of claim 1, wherein, the preprocessed near-infrared hyperspectral data are input into the trained quality prediction model for grading prediction. The method for collecting near-infrared hyperspectral data of Radix Paeoniae Rubra samples from different origins and introducing an attention mechanism in deep learning in the near-infrared hyperspectral data collection stage so that the trained model can identify the spectral characteristics of the region containing more active ingredients comprises the following steps: Radix Paeoniae Rubra samples from different origins are obtained, and the batches and origins are recorded and subjected to a split packaging process, and then near-infrared hyperspectral data are collected to obtain near-infrared hyperspectral images; a YOLOv8 target detection algorithm is used to process the near-infrared hyperspectral images, automatically identify the outline of the Radix Paeoniae Rubra sample, and demarcate a region of interest to obtain a feature map; CNN is used to perform multi-scale feature extraction on the feature map and perform attention weight calculation to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a Radix Paeoniae Rubra ingredient; 3. The method of claim 2, wherein, the collected resources are adjusted according to the attention weight, the image sampling frequency and the spectral signal intensity of the key region are dynamically improved based on the attention weight, and the key region data and the non-key region data are weighted and fused based on the attention weight to obtain near-infrared hyperspectral data. The method for using CNN to perform multi-scale feature extraction on the feature map and performing attention weight calculation to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a Radix Paeoniae Rubra ingredient, comprises the following steps: CNN is used to perform multi-scale feature extraction on the feature map; wherein a convolution layer: 3*3 convolution kernels and 5*5 convolution kernels are used to extract features in parallel, the 3*3 kernel convolution is used to capture microscopic textures, and the 5*5 convolution kernel is used to capture macroscopic spectral trends; secondly, a pooling layer: maximum pooling and average pooling are used alternately; after a global feature vector is obtained through global average pooling, the global feature vector is input into two fully connected layers to obtain the attention weight of each feature channel; wherein each feature channel corresponds to a Radix Paeoniae Rubra ingredient.

4. The method of claim 3, wherein, The adjusting the collection resource according to the attention weight, dynamically improving the image sampling frequency and spectral signal intensity of the key region based on the attention weight, and weighting and fusing the key region data and the non-key region data based on the attention weight to obtain near-infrared hyperspectral data, comprising: Dynamically improving the key region sampling frequency based on the attention weight, so that the active ingredient enrichment region sampling frequency is improved by at least 60%, and the spectral signal intensity is improved by at least 45%; Fusing the key region data and the non-key region data according to the attention weight to improve the effective component spectral signal-to-noise ratio to more than 25dB.

5. The method of claim 1, wherein, The fusion generative adversarial network and the convolutional autoencoder obtain a hybrid denoising model, and the data in the red-processed radix paeoniae quality database is used to train the hybrid denoising model to obtain a trained hybrid denoising model, comprising: A generator is constructed based on a full convolutional neural network structure, the input is a random noise vector, and the output is spectral data with the same dimension as the near-infrared hyperspectral data; wherein the generator includes multiple transpose convolutional layers, which gradually expand the feature map size through multiple transpose convolutional layer operations to generate signals similar to the original spectrum; secondly, the generator also includes a hidden layer ReLU activation to enhance the non-linear mapping to capture the non-linear changes of the characteristic peaks of the red-processed radix paeoniae components; thirdly, the generator also includes an output layer Tanh activation to map the spectral values to [-1, 1] to adapt to the normalized real spectral range; Spectral prior constraints are embedded between each transpose convolutional layer in the generator to force the generated spectral data to have the same peak shape as the real data in the characteristic region of the red-processed radix paeoniae components, and gradient penalty terms are used to ensure the physical interpretability of the generated spectral data, suppress meaningless noise generation, and achieve accurate recovery of the spectral dimension and strengthening of the component characteristics; A discriminator is constructed based on a full convolutional neural network structure, the input is the spectral data generated by the generator or the real clean spectral data, and the output is a scalar representing the probability that the input data is real data; wherein the generator includes multiple convolutional layers, the activation function uses LeakyReLU in the hidden layer and Sigmoid in the output layer, and multiple convolutional layers are used to extract features step by step; secondly, the loss function optimization adopts a binary cross-entropy with weight, and the discrimination error of the red-processed radix paeoniae key component band is given a weight of 2 times; A generative adversarial network is constructed based on the generator and the discriminator, and the data in the red-processed radix paeoniae quality database is used to train the generative adversarial network; An encoder is constructed, wherein the encoder includes multiple convolutional layers to perform convolution operations on the input noisy spectral data, gradually reduce the feature map size, and extract data features; A decoder is constructed, wherein the decoder is strictly symmetrical to the encoder to realize layer-by-layer restoration of the spectral dimension and ensure that the receptive field of each convolutional layer matches the encoder; A convolutional autoencoder is constructed based on the encoder and the decoder, and the data in the red-processed radix paeoniae quality database is used to train the convolutional autoencoder; wherein the convolutional autoencoder also includes an output layer Tanh activation to map the reconstructed spectrum to [-1, 1] to ensure consistency with the output range of the convolutional autoencoder; The generated adversarial network and the convolutional autoencoder after the fusion training obtain a trained hybrid denoising model; wherein the hybrid denoising model is used to preliminarily denoise the near-infrared hyperspectral data through the generated adversarial network, and then is used to secondarily denoise through the convolutional autoencoder.

6. The method of claim 1, wherein, The quality prediction model based on Transformers is constructed, and data in the Radix Paeonia quality database is used to train the quality prediction model to obtain a trained quality prediction model, including: The quality prediction model based on Transformers is constructed, and the quality prediction model includes an input layer, a Transformer module, and a multi-task output layer; wherein the Transformer module is composed of multiple Transformer blocks, and each Transformer block includes a multi-head attention layer, a feedforward neural network layer, and a layer normalization layer; Data in the Radix Paeonia quality database is divided into a training set, a validation set, and a test set according to a preset ratio, and corresponding Radix Paeonia component content labels, origin labels, and artificial marked quality grading labels are determined, and then the Radix Paeonia component content labels are normalized, the origin labels are one-hot encoded, each origin is converted into a unique vector representation, and the artificial marked quality grading labels are encoded to obtain a grading ordinal vector; A loss function is defined, and for component content prediction, a mean square error loss function is used, for origin prediction, a cross-entropy loss function is used, and for grading prediction, a cross-entropy loss function is used. An optimizer is used to train the quality prediction model to obtain a trained quality prediction model.

7. The method of claim 6, wherein, The quality prediction model based on Transformers is constructed, and the quality prediction model includes an input layer, a Transformer module, and a multi-task output layer; wherein the Transformer module is composed of multiple Transformer blocks, and each Transformer block includes a multi-head attention layer, a feedforward neural network layer, and a layer normalization layer, including: The input layer directly uses preprocessed near-infrared hyperspectral data as input to retain complete spectral information; The multi-head attention layer decomposes the input feature dimension into multiple subspaces, and uses different sub-controls to focus, gather, and cover different Radix Paeonia component peak regions, so that the quality prediction model can capture long-distance spectral band correlations in parallel to capture component peak collaborative features; wherein the multi-head attention mechanism realizes cross-modal association between spectral features and quality labels through conditional input: when generating a query vector, in addition to spectral data, origin one-hot encoding and grading ordinal vectors are simultaneously input, modal fusion is realized through a weight matrix, and then attention scores are calculated; secondly, the output of the multi-head attention layer is obtained by concatenating multiple subspace attention results and then performing linear transformation; The feedforward neural network layer processes the output of the multi-head attention layer, and the feedforward neural network includes two fully connected layers with a ReLU activation function in between; Layer normalization: layer normalization is used before and after the multi-head attention layer and the feed-forward neural network layer; Multi-task output layer: the output of the Transformer module is converted into a fixed-length vector through global average pooling, retaining the global statistical features of the spectral data, and then through a fully connected layer, the content of the red peony root ingredient prediction, origin prediction and grading prediction are carried out respectively; among them, for the content of the ingredient prediction, the linear regression output is used to predict the value, for the origin prediction, the softmax function is used to output the prediction probability of each origin, and for the grading prediction, the softmax function is used to output the prediction probability of each grading.

8. A grading system for detecting the quality of red peony root decoction pieces, characterized in that, The grading device comprises: The acquisition module is configured to collect near-infrared hyperspectral data of red peony root samples from different origins, and introduce an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage, so that the trained model can identify spectral features of regions containing more active ingredients; The association module is configured to establish a red peony root quality database based on the near-infrared hyperspectral data, and associate the data in the red peony root quality database with corresponding red peony root ingredient content labels, origin labels and artificial marking quality grading labels; The fusion module is configured to fuse a generative adversarial network and a convolutional autoencoder to obtain a hybrid denoising model, and train the hybrid denoising model using data in the red peony root quality database to obtain a trained hybrid denoising model; The construction module is configured to construct a quality prediction model based on Transformers, and train the quality prediction model using data in the red peony root quality database to obtain a trained quality prediction model; The acquisition module is configured to collect near-infrared hyperspectral data of red peony root samples from different origins, and introduce an attention mechanism in deep learning during the near-infrared hyperspectral data collection stage, so that the trained model can identify spectral features of regions containing more active ingredients; The grading module is configured to input the preprocessed near-infrared hyperspectral data into the trained quality prediction model for grading prediction.

9. A terminal, characterized by comprising: Comprise: One or more processors; Storage device for storing one or more programs; Display for displaying results; When the one or more programs are executed by the one or more processors, the one or more processors implement the grading method for red peony root quality detection according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to perform the grading method for red peony root quality detection according to any one of claims 1-7.

Citation Information

Cited By

  • Cattail pollen content prediction method and system based on hyperspectral double-flow multi-scale CNN (Convolutional Neural Network)

    CN121527631A

  • Pulmonary vein isolation using a novel balloon-based ablation system

    CN121527631B