A schisandra chinensis quality nondestructive detection method and system based on bimodal fusion
By employing a dual-modal fusion method for Schisandra chinensis quality testing, which combines hyperspectral imaging and image recognition technologies, the problem of incomplete information in Schisandra chinensis quality testing has been solved, enabling rapid and accurate non-destructive testing and improving testing efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST FORESTRY UNIV
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for detecting the quality of Schisandra chinensis suffer from incomplete information based on a single modality. Traditional detection methods are inefficient and destructive, making it difficult to achieve rapid, accurate, and non-destructive testing.
A detection method based on dual-modal fusion is adopted, which combines hyperspectral imaging and image recognition technology. The feature sensitivity is enhanced by a convolutional block attention module, and a bidirectional cross-attention mechanism is used to promote the deep information flow and complementarity between image and spectral features. Finally, a highly discriminative joint representation is generated through an adaptive fusion module.
This method enables rapid and non-destructive testing of Schisandra chinensis quality, improves the accuracy and stability of testing, overcomes the shortcomings of incomplete information from a single modality, and enhances the quality of feature characterization and discrimination ability.
Smart Images

Figure CN122487256A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of Schisandra chinensis quality testing, and in particular relates to a non-destructive testing method and system for Schisandra chinensis quality based on dual-modal fusion. Background Technology
[0002] Schisandra chinensis, a traditional Chinese medicine and health food ingredient, has garnered significant attention due to its unique pharmacological effects and wide range of applications. Its lignans, organic acids, and other active substances endow it with multiple benefits, including antioxidant, liver-protective, blood sugar-regulating, and anti-fatigue properties. In recent years, with the increasing demand for natural products, the cultivation and trade of Schisandra chinensis have expanded continuously, and its quality differences have significantly impacted its efficacy and market value.
[0003] However, traditional quality testing methods rely heavily on manual experience or chemical analysis, which suffers from limitations such as low efficiency, high cost, and difficulty in achieving rapid and non-destructive identification. Therefore, it is necessary to introduce modern information technology and emerging intelligent algorithms to establish efficient, accurate, and interpretable quality testing methods for Schisandra chinensis, providing technical support for the standardization and industrialization of medicinal materials.
[0004] Various detection methods have been explored for the quality evaluation of Schisandra chinensis. Early methods relied primarily on manual identification and experience, but this approach is subjective, makes repeatability difficult, and is only suitable for rapid detection of small samples. With the development of analytical chemistry, high-performance liquid chromatography (HPLC) and gas chromatography (GC) have been widely used for the qualitative and quantitative analysis of lignans, polysaccharides, and organic acids in Schisandra chinensis. However, these methods have complex procedures, demanding sample pretreatment, and are often destructive, making them unsuitable for large-scale, rapid online detection.
[0005] In recent years, optical detection techniques such as near-infrared spectroscopy have also been applied to the quality analysis of Schisandra chinensis. For example, by combining near-infrared spectroscopy with chemometric methods, rapid and non-destructive determination of the polysaccharide content, a key active ingredient in Schisandra chinensis, has been successfully achieved. These methods offer advantages such as being non-destructive, rapid, and information-rich, but they still have limitations in distinguishing complex and similar samples and revealing deeper characteristics. Given these limitations, developing a novel quality detection method that combines rapidity, accuracy, and non-destructiveness has become a crucial problem urgently needing to be solved in this field.
[0006] Hyperspectral Imaging (HSI) combines the advantages of traditional imaging and spectral analysis, enabling the simultaneous acquisition of continuous spectral information and spatial distribution characteristics of target samples, thus achieving multi-dimensional information characterization of samples. Compared with traditional chemical detection methods, HSI is not only faster and more non-destructive, but also obtains rich physical and chemical property information without complex sample pretreatment. Therefore, it is widely used in agricultural product quality control, food testing, and quality evaluation of traditional Chinese medicine. Typical studies have shown that HSI exhibits high accuracy and feasibility in tasks such as medicinal material identification, origin traceability, authenticity verification, and quantitative analysis of active ingredients. Meanwhile, image recognition technology, especially deep learning-based image recognition methods, can automatically extract spatial features such as texture, morphology, and color of samples, demonstrating excellent performance in large-scale classification and pattern recognition tasks. With the continuous development of algorithms such as convolutional neural networks and attention mechanisms, image recognition has become an important technological support for promoting intelligent detection of food and traditional Chinese medicine.
[0007] 2.4 Shortcomings of existing technologies While hyperspectral imaging (HSI) and image recognition technologies have shown great potential in the field of quality inspection of traditional Chinese medicine (TCM) materials, both have limitations when applied independently. On the one hand, although HSI provides rich spectral information, its inherent high dimensionality and information redundancy can easily lead to problems such as model overfitting and low computational efficiency. On the other hand, while traditional image recognition methods can intuitively capture macroscopic features such as the morphology and texture of samples, they are limited in their ability to reveal differences in internal chemical composition and deep quality attributes. Therefore, no single technology can fully characterize the multidimensional quality features of Schisandra chinensis. Given the natural complementarity of the two technologies in information representation, achieving effective fusion of spectral and image features is considered a key approach to overcome the bottleneck of single technologies and improve detection accuracy and reliability. However, early multimodal fusion strategies, such as simple feature concatenation or element-wise summation, often ignore the complex intrinsic relationships between heterogeneous data, limiting the upper limit of model performance. In recent years, research trends have shifted towards designing more refined interaction mechanisms to achieve deep fusion. For example, when Liu et al. were conducting non-destructive testing of trace mycotoxins in red ginseng, they combined hyperspectral imaging technology with a deep learning model containing an attention module to achieve pixel-level accurate evaluation of the intrinsic quality of Chinese medicinal materials. Ma et al. proposed a spectral-image feature fusion convolutional neural network (S-IFCNN), which achieved efficient traceability of the geographical origin of wolfberry by fusing spectral and image data from hyperspectral images, with an accuracy rate of 91.99%.
[0008] In summary, a non-destructive testing method and system for Schisandra chinensis quality based on dual-modal fusion is proposed. Summary of the Invention
[0009] In view of this, the present invention aims to propose a non-destructive testing method and system for Schisandra chinensis quality based on dual-modal fusion, so as to solve the problems of incomplete single-modal information, difficulty in effectively fusing multi-modal features, low efficiency and destructiveness of traditional Schisandra chinensis quality testing methods.
[0010] To achieve the above objectives, as one aspect of the present invention, the following technical solution is adopted to provide a non-destructive testing method for Schisandra chinensis quality based on dual-modal fusion, comprising the following steps: S1. Sample preparation: Select Schisandra chinensis samples from various producing areas, remove impurities and inferior fruits, mix them, and classify them into three grades according to national standards. S2. Data Acquisition: Each Schisandra chinensis fruit is subjected to hyperspectral scanning. After data processing, hyperspectral image data is obtained, and the corresponding spectral curve data is extracted. S3. After extracting features from the processed image data and curve data respectively, the feature sensitivity of each branch is enhanced by the convolutional block attention module, so that the network focuses on the most discriminative information. S4. The modal interaction module promotes the flow and complementarity of information between enhanced images and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. S5. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power. S6. Integrate S3 to S5 into a bimodal cross-attention network model; S7. Quality Inspection: Collect the hyperspectral curve data and image data of the Schisandra chinensis sample to be tested, and input them into the model in S6 to obtain the quality type.
[0011] Furthermore, in S2, the data processing procedure includes: S21. Hyperspectral Image Correction: Acquire blackboard reference images and whiteboard reference images, and correct the original hyperspectral image data and curve data by using the original spectral reflectance, black mark reflectance and white mark reflectance; S22. Curve data preprocessing: The Savitzky-Golay smoothing algorithm is used to denoise and smooth the spectral curve of the corrected reflectance spectrum to obtain the preprocessed hyperspectral curve data. S23. Image data preprocessing: After adjusting the images to a uniform pixel size, data augmentation strategies are used to increase sample diversity.
[0012] Furthermore, in S3, the preprocessed hyperspectral curve data is used to extract local spectral characteristics using multiple one-dimensional convolutional layers, residual connections are used to alleviate the gradient vanishing problem, and the feature map is compressed into a fixed-dimensional feature vector; the preprocessed image data is used to capture the color, texture, and morphological features of the Schisandra chinensis sample using a lightweight convolutional neural network.
[0013] Furthermore, in S3, the convolutional block attention module includes a channel attention structure and a spatial attention structure. The feature vectors and color, texture and morphological features are first modeled by the channel attention module to highlight the key feature channels, and then the spatial attention module focuses on the key areas of the feature map, completing the dual weighting of channel and spatial dimensions, and finally outputting enhanced and focused image data and curve data.
[0014] Furthermore, in S4, the modal interaction module acquires the enhanced image data and curve data as Query and Key / Value respectively, calculates the relevance weights through dot product, and generates a weighted representation; then, the roles are exchanged to complete the reverse interaction.
[0015] Furthermore, in S5, the adaptive fusion module performs weighted fusion of features at the global level, the local level, and the interaction level, respectively.
[0016] Furthermore, the global level assigns weights to the two modalities using the softmax function; the local level employs a dimension-wise sigmoid gating mechanism to assign dynamic weights to each feature dimension, achieving fine-grained modal balance.
[0017] Furthermore, the interaction layer balances the global and local fusion results through global interaction weights, adds cross-modal interaction features, and finally generates a joint feature representation for classification.
[0018] Furthermore, the cross-modal interaction feature is a feature vector obtained by concatenating image features and spectral features after S4 interaction.
[0019] As another aspect of the present invention, a non-destructive testing system for Schisandra chinensis quality based on dual-modal fusion is characterized by comprising: The dual-branch feature extraction module extracts features from the processed image data and curve data respectively; Convolutional block attention module enhances the feature sensitivity of each branch; The modal interaction module promotes information flow and complementarity between enhanced image and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power.
[0020] Beneficial effects: 1. This invention uses a dual-modal fusion strategy to simultaneously acquire hyperspectral image data and spectral curve data of Schisandra chinensis, characterizing sample quality from two dimensions: spatial morphology and spectral characteristics. This overcomes the shortcomings of incomplete information from a single modality and enables rapid and non-destructive detection of Schisandra chinensis quality.
[0021] 2. In the feature extraction stage, this invention introduces a convolutional block attention module. Through dual weighting of channel attention and spatial attention, it adaptively enhances key features in the image branch and spectral branch, enabling the network to focus on the most discriminative information and improving the quality of feature representation and discrimination ability.
[0022] 3. This invention constructs a modal interaction module through a bidirectional Cross-Attention mechanism, realizing deep information flow and complementarity between image features and spectral features, ensuring that multimodal features are fully integrated at a deep level, and effectively mining the intrinsic correlation between heterogeneous data.
[0023] 4. This invention employs a multi-level dynamic weight allocation strategy to construct an adaptive fusion module, which performs weighted fusion of features at the global, local, and interactive levels. Combined with residual connections and normalization operations, it generates a joint representation with high discriminative power, thereby improving the accuracy of quality classification and the stability of the model. Attached Figure Description
[0024] Figure 1 This is a flowchart of the present invention; Figure 2 The average spectral curves of Schisandra chinensis of three different qualities are shown in the figure. Figure 3 This is a diagram of the overall network structure described in this invention; Figure 4 This is a structural diagram of the spectral feature extraction module described in this invention; Figure 5 This is a structural diagram of the image feature extraction module of the present invention; Figure 6 This is a structural diagram of the convolutional block attention module of the present invention; Figure 7 This is a structural diagram of the modal interaction module of the present invention; Figure 8 This is a structural diagram of the adaptive fusion module of the present invention; Figure 9 This is an example of the debugging and running results of the model of the present invention; Figure 10 This is a loss curve diagram of the present invention; Figure 11 This is an accuracy curve of the present invention; Figure 12 This is the confusion matrix diagram of the present invention; Figure 13 The hyperspectral imaging system built for this invention; Figure 14 This is one of the flowcharts illustrating the method of human-computer interaction between the interactive control module and the user in the embodiments of this application; Figure 15 This is one of the application scenario diagrams of the human-computer interaction interface in the embodiments of this application; Figure 16 This is the second flowchart illustrating the method of human-computer interaction between the interactive control module and the user in the embodiments of this application; Figure 17 This is the second schematic diagram of the application scenario of the human-computer interaction interface in the embodiments of this application; Figure 18 This is the third schematic diagram of the application scenario of the human-computer interaction interface in the embodiments of this application; Figure 19 This is the third flowchart illustrating the method of human-computer interaction between the interactive control module and the user in the embodiments of this application; Figure 20 This is the fourth flowchart illustrating the method of human-computer interaction between the interactive control module and the user in the embodiments of this application; Figure 21 This is the fourth schematic diagram of the application scenario of the human-computer interaction interface in the embodiments of this application; Figure 22 This is the fifth illustration of an application scenario of the human-computer interaction interface in the embodiments of this application; Figure 23 This is the sixth illustration of the application scenario of the human-computer interaction interface in the embodiments of this application; Figure 24 This is the seventh schematic diagram of the application scenario of the human-computer interaction interface in the embodiments of this application; Figure 25 This is the initial user interface for the software system. Figure 26 Analysis results for first-grade Schisandra chinensis; Figure 27 The analysis results are for second-grade Schisandra chinensis. Figure 28 The analysis results are for third-grade Schisandra chinensis. Detailed Implementation Specific implementation method one: Referring to the accompanying drawings, this embodiment provides a non-destructive testing method for Schisandra chinensis quality based on dual-modal fusion, comprising the following steps: S1. Sample preparation: Select Schisandra chinensis samples from various producing areas, remove impurities and inferior fruits, mix them, and classify them into three grades according to national standards. Mature fruits of Schisandra chinensis were selected as the experimental subjects. To ensure the breadth and representativeness of the sample sources, samples were collected from six major producing areas in northern China, including the Greater Khingan Mountains and Yichun in Heilongjiang Province, Changbai Mountain and Tonghua in Jilin Province, Oroqen in Inner Mongolia Autonomous Region, and Fushun in Liaoning Province.
[0026] Before processing, all samples underwent manual screening to remove impurities and inferior fruits. To construct a comprehensive quality assessment model applicable to different production areas, samples from various regions were uniformly mixed. The grading process strictly followed the national standard "Commodity Specifications and Grades of Chinese Medicinal Materials—Schisandra" (T / CACM 1021.42-2018). Grading was primarily based on the size, color, and plumpness of the fruits, ultimately classifying all samples into three grades: first-class, second-class, and third-class. In subsequent data analysis, these three grades were coded as labels "0" (first-class), "1" (second-class), and "2" (third-class), respectively.
[0027] S2. Data Acquisition: Each Schisandra chinensis fruit is subjected to hyperspectral scanning. After data processing, hyperspectral image data is obtained, and the corresponding spectral curve data is extracted. This invention uses a laboratory-built hyperspectral imaging system to acquire hyperspectral images of Schisandra chinensis samples, such as... Figure 13 As shown. This system can capture images of 224 consecutive wavelengths in the visible and near-infrared short-wave regions (397.66–1003.81 nm). It consists of an FX10 hyperspectral camera (Spectral Imaging Ltd., Oulu, Finland), a light source, a moving platform, a computer, and related control software. The spectral resolution is 5.5 nm. Before acquiring hyperspectral images of the Schisandra chinensis sample, the halogen lamp is turned on, the camera and the electrically controlled moving platform are connected, and the system is allowed to warm up for 10 minutes while the relevant acquisition parameters are set. The Schisandra chinensis sample is placed in a designated area on the moving platform, and the movement is controlled by the software, allowing the camera, perpendicular to the direction of movement, to scan the sample line by line, ultimately obtaining the hyperspectral image data of the Schisandra chinensis sample.
[0028] To eliminate dark current noise from the equipment and the effects of uneven illumination and sensor response differences, the original hyperspectral image needs to be black-and-white corrected to convert it from the original radiance values to relative reflectance, which more accurately reflects the physicochemical properties of the sample. The correction process involves acquiring blackboard and whiteboard reference images and standardizing the original hyperspectral image accordingly. First, under the same environmental conditions as the sample acquisition, a whiteboard reference image is acquired using a standard whiteboard, whose reflectance is close to 100%. Then, the light source is turned off, and the lens is completely covered with an opaque black lens cap to acquire a blackboard reference image. Finally, the acquired data is corrected using the original whiteboard and blackboard images. The correction formula is as follows:
[0029] In the formula: Represents the relative reflectance of the spectrum. Indicates the original spectral reflectance. Indicates the reflectivity of the black label. This indicates the white label reflectance.
[0030] To effectively filter out random noise and baseline drift in the original spectral curve data while retaining the main spectral features, this invention employs the Savitzky-Golay (SG) smoothing algorithm to preprocess the corrected reflectance spectrum. SG filtering uses a fixed odd-numbered sliding window, fitting a polynomial of a specified order within the window using the least squares method, and replacing the original value with the fitted value at the center of the window, thus achieving noise reduction and smoothing of the spectral curve. The average spectral curve of the preprocessed Schisandra chinensis sample in the wavelength range of 397.66–1003.81 nm is shown below. Figure 2 As shown.
[0031] The spectral characteristics in the visible light region (400–780 nm) are mainly related to the pigment components in Schisandra chinensis fruit. In the 400–550 nm range, reflectance is generally at an extremely low level, attributed to the strong absorption of anthocyanins and other red pigments, which is also the direct cause of the fruit's deep red color. The spectral reflectance in the red edge region (680–750 nm) increases sharply. The position, slope, and amplitude of the red edge are extremely sensitive to the fruit's internal biochemical parameters (such as pigment content and cell structure). The reflectance in the near-infrared region (780–1000 nm) is mainly affected by scattering from the fruit's internal cell structure and the absorption of chemical bonds (such as OH, CH) vibrations in organic matter such as water and sugars. The curve rises and flattens after entering this region, forming a high-reflectance plateau. The level of reflectance is related to the cell density and water-holding capacity of the pulp. Near 970 nm, the spectral curve shows a slight downward slope, corresponding to the absorption peak of the stretching vibration of OH bonds in water. The spectral curves of Schisandra chinensis samples exhibit clearly distinguishable class characteristics in the pigment absorption band and red edge position in the visible light region, as well as in the reflectance platform related to moisture and internal structure in the near-infrared region.
[0032] To effectively extract discriminative features from images and improve the model's generalization ability, this invention implements a series of standardized preprocessing procedures for all input RGB images. First, all image sizes are uniformly adjusted to 224×224 pixels to conform to the input specifications of the pre-trained backbone network. Simultaneously, data augmentation strategies are employed to expand sample diversity and reduce the risk of overfitting. Specific operations include random horizontal and vertical flips with a probability of 0.5, random rotations within a ±15 degree range, and slight color jitter covering brightness, contrast, saturation, and hue, aiming to enhance the model's robustness to different viewing angles, object orientations, and lighting variations.
[0033] The overall architecture of the Dual-Modal Cross-Attention Network (DMCA-Net) proposed in this invention is as follows: Figure 3As shown, the network mainly consists of the following modules: a two-branch feature extraction module, a convolutional block attention module (CBAM), a modal interaction module, and an adaptive fusion module. In the feature extraction stage, to enhance the feature sensitivity of each branch, we integrated a convolutional block attention module (CBAM) at the end of both feature extractors, adaptively adjusting the feature response weights for channel and spatial dimensions, allowing the network to focus on the most discriminative information. The modal interaction module promotes information flow and complementarity between image and spectral features through a bidirectional Cross-Attention mechanism, ensuring full fusion of multimodal features at a deep level. Finally, an adaptive fusion module introduces a multi-level dynamic weight allocation strategy, combined with residual connections and normalization operations to ensure training stability, ultimately generating a highly discriminative joint representation for the classification and prediction of Schisandra chinensis quality.
[0034] S3. After extracting features from the processed image data and curve data respectively, the feature sensitivity of each branch is enhanced by the convolutional block attention module, so that the network focuses on the most discriminative information. like Figure 4 As shown, the spectral feature extraction module adopts a one-dimensional convolutional neural network (1D-CNN) structure, specifically including: 1. Input layer: receiving preprocessed spectral curve data (224 dimensions); 2. Convolutional layer: multiple one-dimensional convolutional layers for extracting local spectral features; 3. Residual connection: introducing residual connections to alleviate the gradient vanishing problem; 4. Global average pooling: compressing the feature map into a fixed-dimensional feature vector. After the spectral curve data is input, it first enters the Initial Block, where linear transformation, batch normalization, and ReLU activation are performed to complete the initial feature mapping. Then, it sequentially passes through three Residual Blocks, where fully connected layers, batch normalization, ReLU activation, and Dropout regularization are used to deeply extract and fuse features. Finally, the CBAM attention module enhances key features, and the processed result is output.
[0035] After residual stacking, the network outputs a 512-dimensional feature vector, which is considered a highly condensed representation of the original spectral curve data. This vector will then be fed into the attention and modal interaction modules for further processing.
[0036] like Figure 5 As shown, the image feature extraction module uses a lightweight convolutional neural network (such as ShuffleNetV2 or MobileNet) as the backbone network to effectively capture the color, texture, and morphological features of Schisandra chinensis samples while maintaining high efficiency, providing high-quality image representations for subsequent cross-modal interaction and adaptive fusion. After image data is input, it enters the ShuffleNetV2 x0.5 backbone network to complete feature extraction, and then the CBAM module enhances key features, finally outputting the processing results.
[0037] like Figure 6 As shown, to enhance the discriminative ability during feature extraction, this invention introduces a Convolutional Block Attention Module (CBAM) in both the image and spectral branches. This module consists of two substructures: Channel Attention and Spatial Attention, which can adaptively allocate weights across different dimensions of the feature map.
[0038] The CBAM module's workflow is as follows: For the input feature map, the channel attention module first models the weights of each channel, highlighting key feature channels. Then, the spatial attention module focuses on the key regions of the feature map, completing a dual weighting of channel and spatial dimensions. Finally, it outputs an enhanced and focused optimized feature map. This design enables the network to adaptively focus on the most discriminative spectral bands and image regions, providing more discriminative feature representations for subsequent modal interactions and fusion.
[0039] S4. The modal interaction module promotes the flow and complementarity of information between enhanced images and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. like Figure 7 As shown, to fully leverage the complementarity between images and hyperspectral features, this invention designs a modal interaction module, the core of which is a bidirectional Cross-Attention mechanism. Unlike simple stitching or unidirectional attention, this module can establish bidirectional information flow between images and spectral features, enabling image features to actively focus on key bands in the spectrum, while spectral features can also perceive discriminative regions in the image, thereby achieving deep modal collaboration.
[0040] Specifically, let the image features be... Spectral characteristics are During the interaction, one modality feature is first used as the Query and the other as the Key / Value. Relevance weights are calculated using a dot product to generate a weighted representation. Then, the roles are switched to complete the reverse interaction. Here, Q represents the query vector, K represents the key vector, and V represents the value vector.
[0041] For example, when using an image as the query and a spectrum as the key-value pair:
[0042] The image features are: Spectral characteristics are , This represents the interactive features of an image modality. The attention calculation formula is:
[0043] The length of the eigenvector, It is the transpose of the Key matrix, and Attention(Q,K,V) represents the attention mechanism function. Conversely, the same process applies when using the spectrum as the Query.
[0044] Finally, the enhanced modal features are represented as follows:
[0045]
[0046] in Indicates the enhanced image features, This represents the enhanced spectral features. Simultaneously, a cross-modal interaction feature term is defined. :
[0047] This indicates a splicing operation. This interactive feature is added as additional information in the subsequent adaptive fusion module to further enhance the discriminative power of modality integration.
[0048] This design not only ensures efficient interaction between modalities but also intuitively reflects the correspondence between images and spectral features through the interaction weight matrix, demonstrating good interpretability. Experimental results show that this interaction mechanism significantly improves the discriminative power of multimodal feature fusion and the generalization ability of the model in the Schisandra chinensis quality classification task.
[0049] S5. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power. like Figure 8 As shown, the adaptive fusion module employs a global-local two-level fusion strategy: weighted fusion of features at the global, local, and interaction levels. First, at the global level, the module assigns weights to the two modalities using the softmax function. Let the image features be... Spectral characteristics are Global fusion is represented as:
[0050] in, Indicates global fusion features, , The weights represent those learned by the network, reflecting the relative contributions of different modalities at the global semantic level.
[0051] At the local level, the module employs a dimension-wise sigmoid gating mechanism to assign dynamic weights to each feature dimension, achieving fine-grained modal balance:
[0052] Where represents the weight of each feature dimension. This indicates element-wise multiplication. This represents the globally fused features. This design ensures that the model can adaptively select the most discriminative modal features in a specific dimension.
[0053] At the interaction level, the module introduces a global interaction weight. This is used to balance the global and local fusion results, while also adding cross-modal interaction feature terms. :
[0054] in Indicates the final feature, =0.1 is an empirically set constant used to suppress excessive amplification of interaction terms.
[0055] This module enables the model to fully leverage the complementary advantages of spectral and image processing, while maintaining robustness and interpretability during multimodal information integration. Experimental results demonstrate that the adaptive fusion module plays a crucial role in improving classification accuracy and reducing performance fluctuations, providing a more robust feature representation for Schisandra chinensis quality detection.
[0056] S6. Integrate S3 to S5 into a bimodal cross-attention network model; steps S3 to S5 are the process of constructing the bimodal cross-attention network model.
[0057] S7. Quality Inspection: Hyperspectral curve data and image data of the Schisandra chinensis sample to be tested are collected and input into the model in S6 to determine the quality type. Specifically: First, prepare a batch of Schisandra chinensis samples to be tested. After removing impurities, collect their hyperspectral curve data and image data separately, and match them one-to-one with the samples. Then, simultaneously input both types of data into the network input, and the network determines the quality type of the sample. Specific Implementation Method Two: A non-destructive testing system for Schisandra chinensis quality based on dual-modal fusion, comprising: The dual-branch feature extraction module extracts features from the processed image data and curve data respectively; Convolutional block attention module enhances the feature sensitivity of each branch; The modal interaction module promotes information flow and complementarity between enhanced image and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power. Specific implementation method three: To comprehensively evaluate the performance of the proposed model in the Schisandra chinensis quality detection task, this invention employs multi-dimensional evaluation metrics, including: classification accuracy, precision, recall, and F1-score, with the following formulas:
[0059]
[0060]
[0061]
[0062] Where TP, FP, and FN represent the first, second, and third... The true positives, false positives, and false negatives of each class, where K is the total number of classes and N is the total number of samples.
[0063] The hardware configuration used in this invention is an AMD Ryzen 7 8745H with Radeon 780M Graphics paired with an NVIDIA GeForce RTX 4060 Laptop GPU. During model training, we employed optimized hyperparameter configurations: using the AdamW optimizer (learning rate 1e-3, weight decay 1e-4) combined with a cosine annealing learning rate scheduling strategy, selecting labeled smooth cross-entropy loss for regression training, setting the batch size to 32, training epochs to 100, and introducing an early stopping mechanism (patience value of 20 epochs). Simultaneously, data standardization and 10-fold cross-validation were used to ensure the stability and generalization ability of the model training. The model execution process is as follows: Figure 9 As shown.
[0064] This invention evaluates the performance of a bimodal cross-attention network model, and the results are shown in Table 1 and... Figures 10 to 12 .
[0065] Table 1 Model Evaluation Indicators
[0066] Training process analysis: As can be seen from the training and validation loss curves, the model converges rapidly within the first 10 epochs, with the training loss decreasing from the initial 1.2 to approximately 0.5, and then stabilizing around 20 epochs. Simultaneously, the accuracy curve rises rapidly and plateaus after 30 epochs, with both the training and validation accuracy remaining around 98%, showing a highly consistent trend without significant bifurcation. This indicates that the model converges well within the data scale of this invention, without serious overfitting or underfitting issues.
[0067] Classification performance analysis: The confusion matrix further reveals the model's discrimination performance across different categories. Overall, all three categories of samples were identified with high accuracy, with the first category showing the most stable prediction, exhibiting only one misclassification. A small amount of confusion existed between categories one and two, indicating a high degree of similarity in appearance or spectral features between these two categories, increasing the difficulty for the model to distinguish them. Nevertheless, the overall misclassification rate was less than 3%, demonstrating that the model still possesses strong discriminative ability in differentiating highly similar categories.
[0068] Conclusion: The training / validation curves and confusion matrix results fully validate the effectiveness and reliability of the proposed model. The model can achieve stable convergence in a relatively small number of iterations, and the accuracy of the training set and the validation set are highly consistent, indicating that the optimization strategy adopted can effectively avoid overfitting and underfitting. The final classification accuracy is close to 98%, indicating that the model has a strong discriminative ability in extracting and fusing image and spectral features; Although there was some confusion between the second-grade and third-grade samples, the overall misclassification rate was less than 3%, showing that the model still maintained a strong ability to distinguish between highly similar categories. The fact that almost all first-class samples were correctly identified further demonstrates that the model can achieve stable and reliable classification performance in categories with high feature discrimination.
[0069] These results demonstrate that the proposed model not only exhibits excellent convergence efficiency and classification accuracy, but also demonstrates good robustness and generalization ability under complex sample distributions, providing strong support for its application in practical Schisandra chinensis quality detection.
[0070] To systematically evaluate the contribution of each proposed module to the overall performance, ablation experiments were conducted on a dual-branch network, and the results are shown in Table 2. Overall, the model performance continuously improves with the gradual introduction of modules, fully demonstrating the important role of each component in feature learning and multimodal fusion.
[0071] Table 2 Ablation Experiment Results of the Optimized Model
[0072] analyze: (1) Single module contribution Baseline model: 0.97 M parameters, accuracy 95.30%. Introducing CBAM alone: Accuracy improved to 95.93%, indicating that embedding channels and spatial attention effectively improved the quality of single-modal representation. Introducing modal interaction separately: the accuracy also reached 95.93%, confirming the necessity of deep modal interaction before fusion, which can effectively uncover deep complementary information between modalities. Introducing adaptive fusion separately: accuracy improved to 95.62%, indicating that the dynamic weighting strategy dependent on sample input is more effective than simple feature concatenation, and can achieve more refined information integration. (2) Modular combination effect CBAM + Modal Interaction: Accuracy jumped to 97.34%, with 1.21 M parameters, indicating that the high-quality single-modal features refined by CBAM significantly improved the efficiency and depth of subsequent cross-modal interactions. CBAM + Adaptive Fusion: Accuracy increased to 97.97%, likely due to the enhanced features of CBAM enabling the adaptive fusion module to more accurately assess the importance of each modality, thereby allocating optimal fusion weights. Modal interaction + adaptive fusion: 97.18% accuracy, demonstrating the high compatibility between features derived from deep interaction and dynamic fusion strategies. (3) Complete model The complete model integrating all three modules achieved a maximum accuracy of 98.59%, while the number of parameters was only 1.26 M, representing a significant improvement of 3.29 percentage points compared to the baseline model, thus demonstrating the overall effectiveness of the framework of this invention.
[0073] Conclusion: The ablation experiment results verified the independent and synergistic effects of each module. Among them, the modal interaction module and the adaptive fusion module made the most significant contributions to the performance improvement and were the core elements that enabled the model to outperform the baseline structure.
[0074] Model Deployment: To further verify the actual performance of the proposed improved algorithm, this invention integrates the model with an embedded development board to build a real-time detection visualization interface for Schisandra chinensis samples, and systematically evaluates the detection effect based on existing datasets.
[0075] Hardware: Considering the model characteristics and actual development costs of this invention, the Jetson Orin Nano CLB development kit provides an ideal hardware platform for the edge deployment of the Schisandra chinensis quality detection model. The device is equipped with an NVIDIA Ampere architecture GPU (1024 cores, 32 Tensor Cores) and a 6-core Arm Cortex-A78AE CPU, boasting up to 40 TOPS of INT8 inference performance, enabling efficient operation of the lightweight CNN-Transformer model constructed in this invention. It supports mainstream deep learning frameworks such as PyTorch, and with configurable power consumption of 10-25W, it enables real-time spectral analysis and qualitative analysis of quality levels at the edge, meeting the real-time requirements of production line quality monitoring and providing reliable computing power support for the practical application of complex deep neural networks in the field of spectral analysis.
[0076] Software component: The software system is developed based on Python and integrates a model inference engine, graphical interface and data processing module, and supports real-time logging of detailed running information.
[0077] Figures 14 to 24 As shown Please see Figure 14 In some implementations, the interactive control module is further configured with a human-computer interaction interface, and the method for the interactive control module to perform human-computer interaction with the user includes: Step 0001: In response to the user's first operation in the first area of the human-computer interaction interface, import the hyperspectral curve data and image data of the Schisandra chinensis sample into the Schisandra chinensis quality non-destructive testing system; The hyperspectral curve data and image data are pre-acquired based on a preset spectral acquisition device; Step 0002: In response to the user's second operation in the first area, display the hyperspectral curve data, image data, and qualitative analysis results corresponding to the selected Schisandra chinensis sample in the second area of the human-computer interaction interface; Step 0003: In response to the user's third operation in the first area, display the hyperspectral curve data and image data corresponding to the selected Schisandra chinensis sample, as well as the quantitative analysis results, in the second area.
[0078] It should be noted that, under normal circumstances, step 0001 should be executed before steps 0002 and 0003. Steps 0002 and 0003 can be executed according to the user's actual needs through corresponding operations. The execution order and whether they are executed depend on the user's needs. Figure 14 The process shown is for illustrative purposes only and should not be construed as limiting.
[0079] Specifically, based on the above implementation methods, a human-computer interaction interface (HCI) is provided. This interface is generally displayed to the user through a computer device's monitor. The user performs operations on the HCI using a mouse, keyboard, or voice control, thereby achieving human-computer interaction. For details on the specific implementation of HCI, please refer to [link to relevant documentation]. Figure 15 , Figure 15 The overall structure of the human-computer interaction interface is illustrated exemplarily. The first area is the functional area where the user interacts with the front-end module through operations. This area can include interactive components such as buttons, text boxes, and drop-down menus, and is typically set to gray as the background color. The second area is the results display area, used to present qualitative or quantitative analysis results to the user through charts or other methods. Specifically, it can be represented as a display window set within the first area. In this interface, the user can interact with the front-end module by performing operations on the interactive components included in the first interface, such as buttons, text boxes, and drop-down menus, thereby obtaining interactive feedback or qualitative / quantitative analysis results of the Schisandra chinensis sample from the human-computer interaction interface.
[0080] Furthermore, in some implementations, please refer to... Figure 16 Step 0001 specifically includes: Step 00011: In response to the user's click operation on the first button in the first area, import hyperspectral curve data and image data into the Schisandra chinensis quality non-destructive testing system; Step 00012: Display the first text message in the first area to indicate to the user that the hyperspectral curve data and image data have been imported.
[0081] Specifically, please refer to Figure 17 For applying pre-acquired hyperspectral curve data and image data to the Schisandra chinensis quality non-destructive testing method in the above embodiments, for example, the user can import the pre-acquired hyperspectral curve data and image data into the backend module of the Schisandra chinensis quality non-destructive testing system by operating the first button on the first interface. Figure 17 As shown, the first interface has a first button. After the user clicks the first button, the human-computer interaction interface will display a pop-up checkbox, allowing the user to select the file paths corresponding to the hyperspectral curve data and image data. After the user selects the paths, the files corresponding to the hyperspectral curve data and image data are imported into the backend module. Simultaneously, the frontend module displays text information below the first button to indicate to the user that the spectrum import is complete, for example... Figure 18 The situation is shown.
[0082] Furthermore, in some implementations, please refer to Figure 19 Step 0002 specifically includes: Step 00021: In response to the user's selection operation of the drop-down menu in the first area, select the Schisandra chinensis sample; Step 00022: In response to the user's click on the second button in the first area, display the hyperspectral curve data, image data, and qualitative analysis results corresponding to the selected Schisandra chinensis sample in the second area.
[0083] Similarly, in some implementations, please refer to Figure 20 Step 0003 specifically includes: Step 00031: In response to the user's selection operation of the drop-down menu in the first area, select the Schisandra chinensis sample; Step 00032: In response to the user's click on the third button in the first area, display the hyperspectral curve data, image data, and quantitative analysis results corresponding to the selected Schisandra chinensis sample in the second area.
[0084] Specifically, based on the above implementation method, after successfully importing hyperspectral curve data and image data, the user can select the Schisandra chinensis sample to be analyzed using the drop-down menu in the first area, and then control the computer device to execute the non-destructive testing method for Schisandra chinensis quality in the above implementation method by clicking the corresponding model button (corresponding to the second button or the third button). Finally, the front-end module displays the qualitative / quantitative analysis results corresponding to the model selected by the user in the second area in the form of multiple charts.
[0085] For example, please refer to Figure 21 Users can extract data from various sample fields in the imported file by clicking the expand button in the drop-down menu. The data will then be displayed in the drop-down menu. Users can select the sample to be analyzed by clicking the drop-down menu. The interface after selection will then appear as shown. Figure 22 As shown.
[0086] Then, users can perform qualitative or quantitative analysis by clicking the buttons corresponding to different models, for example, in Figure 22 Based on the scenario described, when the user clicks the second button, the front-end module controls the back-end module to execute the non-destructive testing method for Schisandra chinensis quality described in the above implementation. The qualitative analysis model is used to determine the qualitative analysis results corresponding to the sample. Then, the hyperspectral curve data, image data charts, and qualitative analysis charts corresponding to the sample are displayed in the second area, such as... Figure 23As shown. The hyperspectral curve data and image data charts are displayed in the upper left of the second area. The qualitative analysis charts generally include two parts: one is the qualitative analysis result chart, which is used to characterize the qualitative analysis of the Schisandra chinensis sample quality grade, and is displayed in the upper right of the second area; the other is the prediction confidence chart, which characterizes the prediction confidence of the qualitative analysis and is displayed in the lower left of the second area.
[0087] For example, in Figure 24 As shown, when the user selects the corresponding sample and clicks the third button, the quantitative analysis model determines the quantitative analysis results for the sample. Then, the second area displays the hyperspectral curve data, image data charts, and quantitative analysis charts corresponding to the sample, such as... Figure 24 As shown. The hyperspectral curve data and image data charts are displayed in the upper left of the second area. The quantitative analysis charts generally include two parts: one is the quantitative analysis result chart, which is used to characterize the qualitative analysis results of the Schisandra chinensis sample quality grade and is displayed in the upper right of the second area; the other is the prediction confidence chart, which characterizes the prediction confidence of the quantitative analysis and is displayed in the lower left of the second area.
[0088] It should be noted that the display positions of hyperspectral curve data and image data charts and qualitative / quantitative analysis charts in the second area can be adjusted according to user habits. The above implementation is only an illustrative example and should not be construed as limiting.
[0089] Furthermore, for the qualitative / quantitative analysis models corresponding to the second and third buttons in the above embodiments, in addition to the preferred analysis model of the Schisandra chinensis quality non-destructive testing method in the above embodiments, other qualitative / quantitative analysis models in the current related technologies can also be selected. For details, please refer to Tables 1 and 2 in the above embodiments for other qualitative / quantitative analysis models. If multiple qualitative / quantitative analysis models are deployed simultaneously on the computer device, a second (third) button can be set for each analysis model for the user to select.
[0090] Furthermore, in some implementations, the method for the interactive control module to perform human-computer interaction with the user also includes: In response to the second operation, the hyperspectral curve data and image data, the parameter information of the qualitative analysis model, and the qualitative analysis results are displayed in text form in the text box in the first area; or in response to the third operation, the hyperspectral curve data and image data, the parameter information of the quantitative analysis model, and the quantitative analysis results are displayed in text form in the text box.
[0091] For details, please continue reading Figure 23 as well as Figure 24For example, when the user clicks the second or third button, the human-computer interface, in addition to displaying the qualitative / quantitative analysis results in a chart format in the second area, will simultaneously display the selected sample, its spectral characteristics, the qualitative / quantitative analysis results, and the prediction confidence level in a text box located at the bottom of the first area. However, unlike the chart in the second area, the text-based qualitative / quantitative analysis results generally only include the prediction result with the highest probability. Thus, the user can directly copy the text displayed in the text box to save it as the analysis result of the current qualitative / quantitative analysis.
[0092] The electronic device in this application includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the non-destructive testing method for Schisandra chinensis quality described in the above embodiment is implemented.
[0093] The computer-readable storage medium in the embodiments of this application stores a computer program, which, when executed by one or more processors, implements the non-destructive testing method for Schisandra chinensis quality described in the above embodiments.
[0094] like Figures 25 to 28 As shown, by inputting bimodal data of Schisandra chinensis samples, the system can quickly distinguish the quality of various samples. To more directly understand the model's detection performance, the detection results of the model on the Schisandra chinensis dataset are presented in the figure. These metrics include not only the final predicted category and its confidence distribution, but also the total number of model parameters and single inference latency. The total number of parameters reflects the model's complexity, while the inference latency measures its computational efficiency on the target hardware. These performance metrics, providing immediate feedback, along with the clear classification results, together provide direct and valuable empirical data for evaluating the model's deployment performance.
[0095] This invention successfully constructed and validated a bimodal deep learning framework called DMCA-Net for rapid and non-destructive quality detection of Schisandra chinensis. This framework, through an innovative dual-branch structure and the organic integration of CBAM attention, modal interaction, and adaptive fusion modules, achieves deep collaborative modeling of sample images and hyperspectral information. Experimental results show that this method achieves a classification accuracy of 98.59%, significantly outperforming all single-modal and traditional baseline methods, fully validating its effectiveness and robustness. Ablation experiments further reveal that each module contributes to performance, especially the adaptive fusion mechanism, which is key to achieving efficient feature integration. In summary, this invention provides an efficient and feasible technical solution for intelligent quality evaluation of Schisandra chinensis and offers valuable methodological reference for intelligent detection tasks of related agricultural products, demonstrating broad application prospects.
[0096] The embodiments of the present invention disclosed above are merely illustrative of the invention. These embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A non-destructive testing method for Schisandra chinensis quality based on dual-modal fusion, characterized in that, Includes the following steps: S1. Sample preparation: Select Schisandra chinensis samples from various producing areas, remove impurities and inferior fruits, mix them, and classify them into three grades according to national standards. S2. Data Acquisition: Each Schisandra chinensis fruit is subjected to hyperspectral scanning. After data processing, hyperspectral image data is obtained, and the corresponding spectral curve data is extracted. S3. After extracting features from the processed image data and curve data respectively, the feature sensitivity of each branch is enhanced by the convolutional block attention module, so that the network focuses on the most discriminative information. S4. The modal interaction module promotes the flow and complementarity of information between enhanced images and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. S5. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power. S6. Integrate S3 to S5 into a bimodal cross-attention network model; S7. Quality Inspection: Collect the hyperspectral curve data and image data of the Schisandra chinensis sample to be tested, and input them into the model in S6 to obtain the quality type.
2. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 1, characterized in that, In step S2, the data processing procedure includes: S21. Hyperspectral Image Correction: Acquire blackboard reference images and whiteboard reference images, and correct the original hyperspectral image data and curve data by using the original spectral reflectance, black mark reflectance and white mark reflectance; S22. Curve data preprocessing: The Savitzky-Golay smoothing algorithm is used to denoise and smooth the spectral curve of the corrected reflectance spectrum to obtain the preprocessed hyperspectral curve data. S23. Image data preprocessing: After adjusting the images to a uniform pixel size, data augmentation strategies are used to increase sample diversity.
3. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 2, characterized in that, In step S3, the preprocessed hyperspectral curve data is used to extract local spectral characteristics using multiple one-dimensional convolutional layers, and residual connections are used to alleviate the gradient vanishing problem. The feature map is compressed into a fixed-dimensional feature vector. The preprocessed image data is then used with a lightweight convolutional neural network to capture the color, texture, and morphological features of the Schisandra chinensis sample.
4. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 3, characterized in that, In S3, the convolutional block attention module includes a channel attention structure and a spatial attention structure. The feature vectors and color, texture and morphological features are first modeled by the channel attention module to highlight the key feature channels. Then, the spatial attention module focuses on the key areas of the feature map, completing the dual weighting of the channel and spatial dimensions, and finally outputting enhanced and focused image data and curve data.
5. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 4, characterized in that, In step S4, the modal interaction module obtains the enhanced image data and curve data as Query and Key / Value respectively, calculates the relevance weights through dot product, and generates a weighted representation. Then they switched roles and completed the reverse interaction.
6. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 5, characterized in that, In S5, the adaptive fusion module performs weighted fusion of features at the global level, the local level, and the interaction level, respectively.
7. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 6, characterized in that: The global level assigns weights to the two modalities using the softmax function; the local level employs a dimensional sigmoid gating mechanism to assign dynamic weights to each feature dimension, achieving fine-grained modal balance.
8. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 7, characterized in that: The interaction layer balances the global and local fusion results through global interaction weights, adds cross-modal interaction features, and finally generates a joint feature representation for classification.
9. The method for non-destructive testing of Schisandra chinensis quality based on dual-modal fusion according to claim 8, characterized in that: The cross-modal interaction feature is a feature vector obtained by concatenating image features and spectral features after S4 interaction.
10. A non-destructive testing system for Schisandra chinensis quality based on dual-modal fusion, characterized in that, include: The dual-branch feature extraction module extracts features from the processed image data and curve data respectively; Convolutional block attention module enhances the feature sensitivity of each branch; The modal interaction module promotes information flow and complementarity between enhanced image and spectral features through a bidirectional Cross-Attention mechanism, ensuring that multimodal features are fully integrated at a deep level. The adaptive fusion module uses a multi-level dynamic weight allocation strategy to process the interactive features, and combines residual connections and normalization operations to generate a joint representation with high discriminative power.