Intelligent identification method and system for traditional Chinese medicinal materials based on cross-modal comparative learning and adaptive attention fusion
By integrating cross-modal contrastive learning with adaptive attention, the visual, odor, and textual information of Chinese medicinal materials are deeply integrated, solving the problems of strong subjectivity in identification results, difficulty in identifying easily confused varieties, and insufficient robustness due to modal deficiencies in existing technologies. This achieves intelligent identification of Chinese medicinal materials with high accuracy and strong adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 李佳育
- Filing Date
- 2025-11-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to effectively combine visual, odor, and textual information of Chinese medicinal materials for intelligent identification, resulting in highly subjective identification results, difficulties in standardization, challenges in identifying easily confused varieties, insufficient utilization of odor information, and a lack of modality loss robustness.
We employ a cross-modal contrastive learning and adaptive attention fusion method, which maps three modalities of data—image, odor, and text—to a unified feature space through image, odor, and text encoders. We then perform adaptive weighted fusion through contrastive learning and attention mechanisms, combined with multi-task learning and modality missing robustness training, to achieve deep integration and identification of multimodal information.
It achieves high accuracy and strong generalization ability in the identification of Chinese medicinal materials, and can maintain high identification accuracy even in the absence of modalities. It is adaptable to different application scenarios and conforms to the "observation, auscultation, inquiry and palpation" concept of traditional Chinese medicine.
Smart Images

Figure CN121834645A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and traditional Chinese medicine identification, and particularly relates to a traditional Chinese medicinal material intelligent identification method and system based on cross-modal contrast learning and adaptive attention fusion. BACKGROUND
[0002] Traditional Chinese medicinal material identification is a core link of the traditional Chinese medicine industry chain, and is directly related to the quality safety and clinical efficacy of traditional Chinese medicine. Traditional identification relies on sensory experience such as "looking at color, smelling smell, tasting taste, and touching texture", and has the following technical bottlenecks:
[0003] First, strong subjectivity and difficulty in standardization. Different identification experts have different judgments on the same medicinal material, making it difficult to form unified and quantitative quality standards, which is not conducive to large-scale production and quality control.
[0004] Second, long personnel training cycle and scarce expert resources. It takes more than ten years of practical accumulation to train a qualified traditional Chinese medicinal material identifier, and the existing number of experts cannot meet the needs of industrial development.
[0005] Third, it is difficult to identify confusing varieties. Some traditional Chinese medicinal materials have similar appearances but different efficacies, and it is difficult to distinguish them by vision alone, such as honeysuckle and golden honeysuckle, Sichuan fritillary and flat fritillary.
[0006] Fourth, the use of odor information is not sufficient. Volatile components of traditional Chinese medicinal materials are important identification basis, but traditional methods are difficult to objectively quantify and digitally store odor characteristics.
[0007] In recent years, artificial intelligence technology has made some progress in the field of traditional Chinese medicinal material recognition. Existing technical solutions mainly include:
[0008] Image recognition method based on computer vision: convolutional neural network is used to extract visual features from medicinal material images for classification. However, single image modality is difficult to distinguish similar-looking confusing varieties, and is sensitive to light and shooting angle.
[0009] Odor recognition method based on electronic nose: gas sensor array is used to collect the electrical signal response of medicinal material volatile components, and machine learning algorithm is used for classification. However, electronic nose signal is easily affected by environmental temperature and humidity, and it is difficult to reflect the morphological characteristics of medicinal materials.
[0010] Auxiliary identification method based on text knowledge: rules system is constructed by using pharmacopoeia description and expert knowledge. However, the rules are rigid and difficult to handle fuzzy boundary conditions.
[0011] The above single-modal methods each have limitations and are difficult to comprehensively use multiple sensory information to make judgments like experienced identification experts. At present, there is no intelligent identification technology scheme for traditional Chinese medicinal materials that deeply fuses image, smell, and text modal information and has modal missing robustness. SUMMARY
[0012] In view of the deficiencies of the prior art, the present application provides an intelligent identification method and system for traditional Chinese medicinal materials based on cross-modal contrast learning and adaptive attention fusion. The present application innovatively deeply aligns and fuses three heterogeneous modalities of visual images, smell sensing, and text knowledge, and introduces a modal dropout training strategy to enhance the robustness of the system to modal missing, thereby realizing intelligent identification of traditional Chinese medicinal materials with high accuracy and strong generalization ability.
[0013] To achieve the above object, the present application adopts the following technical solutions:
[0014] An intelligent identification method for traditional Chinese medicinal materials based on cross-modal contrast learning and adaptive attention fusion, comprising the following steps:
[0015] S1. Multi-modal data acquisition step: image data, smell time series data, and text description data of traditional Chinese medicinal material samples are respectively acquired to construct a three-modal data set;
[0016] S2. Single-modal feature encoding step: image encoder, smell encoder, and text encoder are respectively used to map each modality data to a unified dimensional feature space to obtain image feature vectors, smell feature vectors, and text feature vectors;
[0017] S3. Cross-modal contrast alignment step: a contrast learning loss function is used to pull the three-modal features of the same medicinal material sample closer to each other in the semantic space and push the features of different medicinal material samples farther away from each other, thereby realizing cross-modal semantic alignment;
[0018] S4. Adaptive attention fusion step: a cross-modal attention mechanism is used to calculate the correlation weight between modalities, and adaptive weighted fusion is performed in combination with modality confidence estimation to generate multi-modal fusion features;
[0019] S5. Multi-task joint learning step: based on the fusion features, multi-task prediction such as medicinal material category classification, origin discrimination, quality grade evaluation, etc. is simultaneously performed, and the feature representation capability is enhanced through multi-task learning;
[0020] S6. Identification result output step: the multi-task prediction results are comprehensively combined to output the category, origin, quality grade, and confidence of the traditional Chinese medicinal material.
[0021] Further, in the S1 step, the image data acquisition includes: using an image acquisition device to obtain a multi-view color image of the traditional Chinese medicinal material sample, with an image resolution of no less than 224x224 pixels and a color space in RGB format.
[0022] Further, in the S1 step, the odor time-series data acquisition includes: placing the quantitative traditional Chinese medicinal material sample in a sampling cavity of an odor acquisition device, and after the odor volatilization is balanced, using an odor detection device to continuously acquire odor response signals at a fixed sampling frequency to form a multi-channel time-series data matrix.
[0023] Further, the odor detection device includes at least one of the following types: (1) Professional electronic nose device: including a PEN3 type portable electronic nose of AIRSENSE company in Germany, a FOX series electronic nose of AlphaMOS company in France, a Heracles II type or a Heracles NEO type ultra-fast gas-phase electronic nose; (2) Gas chromatography-mass spectrometry (GC-MS): using one-dimensional gas chromatography separation combined with mass spectrometry detection for accurate analysis of volatile component composition, and outputting one-dimensional chromatogram data; (3) Comprehensive two-dimensional gas chromatography-mass spectrometry (GCxGC-MS) or comprehensive two-dimensional gas chromatography-time-of-flight mass spectrometry (GCxGC-TOF-MS): using two chromatographic columns with different polarities in series, realizing two-dimensional separation through a modulator, and outputting a two-dimensional chromatogram image (in the form of a thermal map), with a peak capacity of 10 to 50 times that of ordinary GC-MS, and capable of detecting hundreds to thousands of volatile compounds; (4) Gas chromatography-ion mobility spectrometry (GC-IMS): used for rapid odor fingerprint analysis, with a sensitivity of ppb level; (5) Multi-channel gas sensor array: including at least three of metal oxide semiconductor sensors, electrochemical sensors, photoionization sensors, and environmental sensors, with a detection channel number of no less than 8.
[0024] Further, in the S1 step, the text description data includes at least two of the following text information of the traditional Chinese medicinal material: origin, nature and taste, meridian, efficacy, indication, character description, microscopic characteristics, and physicochemical identification.
[0025] Further, in the S2 step, the image encoder uses a pre-trained Vision Transformer (ViT) model, and the processing procedure includes: (a) Image patching (Patch Embedding): dividing the input image into image blocks of a fixed size (such as 16x16 or 14x14 pixels), and mapping each image block into an embedding vector through linear projection; (b) Positional Encoding: Adding learnable positional embeddings to preserve spatial position information; (c) Class Label: Adding learnable [CLS] token at the beginning of the sequence; (d) Transformer Encoding: Processing through multi-layer Multi-Head Self-Attention and Feed-Forward Network (FFN); (e) Feature Extraction: Projecting the output of [CLS] position to a d-dimensional image feature vector.
[0026] Further, in the S2 step, the odor encoder adopts a time-series Transformer structure, and the processing procedure includes: (a) Time-series difference preprocessing: Calculating the first-order difference Δx_t = x_t - x_{t-1} of adjacent time steps to enhance the dynamic change information of the odor; (b) Sequence segmentation: Dividing long time-series data into sub-sequences using sliding window or non-overlapping block method; (c) Positional Encoding: Adding sinusoidal positional encoding or learnable positional embeddings for each time step; (d) Multi-head self-attention encoding: Extracting time-series features through L-layer Transformer encoder, each layer containing Multi-Head Self-Attention and FFN; (e) Sequence pooling: Using global average pooling or [CLS] Token pooling to obtain a d-dimensional odor feature vector.
[0027] Further, in the S2 step, the text encoder adopts a pre-trained language model such as BERT (Bidirectional Encoder Representations from Transformers) or its Chinese variant, and the processing procedure includes: (a) Tokenization: Converting text description into WordPiece or BPE token sequence; (b) Special token: Adding [CLS] token at the beginning of the sequence and [SEP] token at the end of the sequence; (c) Embedding layer: Mapping token IDs to word embeddings, and adding position embeddings and paragraph embeddings; (d) Transformer encoding: Processing through multi-layer bidirectional self-attention mechanism; (e) Feature extraction: The hidden state at the [CLS] position is mapped to a d-dimensional text feature vector through a projection layer.
[0028] Further, in the S3 step, the cross-modal contrast alignment adopts the InfoNCE contrast loss function (Noise Contrastive Estimation). For training samples with a batch size of N, the loss function is defined as follows: Image-odor contrast loss: L_io = -(1 / N)∑_{i=1}^{N}log[exp(sim(z_i^img,z_i^odor) / τ) / Σ_{j=1}^{N}exp(sim(z_i^img,z_j^odor) / τ)] Image-text contrast loss: L_it = -(1 / N)∑_{i=1}^{N}log[exp(sim(z_i^img,z_i^txt) / τ) / Σ_{j=1}^{N}exp(sim(z_i^img,z_j^txt) / τ)] Odor-text contrast loss: L_ot = -(1 / N)∑_{i=1}^{N}log[exp(sim(z_i^odor,z_i^txt) / τ) / Σ_{j=1}^{N}exp(sim(z_i^odor,z_j^txt) / τ)] Where z^img, z^odor, z^txt are the L2 normalized image, odor, and text feature vectors, respectively; sim(u,v) = u·v / (||u||·||v||) is the cosine similarity function; τ is the temperature hyperparameter, which controls the smoothness of the distribution, and the typical value is 0.07. Total contrast loss: L_contrast = (L_io + L_it + L_ot) / 3
[0029] Further, in the S4 step, the cross-modal attention mechanism includes: (a) Calculate the attention of the image to the odor: Take the image feature as the query Q, and the odor feature as the key K and the value V. Calculate the image-odor interaction feature through the scaled dot-product attention; (b) Calculate the attention of the image to the text: Take the image feature as the query Q, and the text feature as the key K and the value V. Calculate the image-text interaction feature; (c) Calculate the attention of the odor to the text: Take the odor feature as the query Q, and the text feature as the key K and the value V. Calculate the odor-text interaction feature.
[0030] Further, the calculation formula of the scaled dot-product attention (Scaled Dot-Product Attention) is as follows: Attention(Q, K, V) = softmax(QK^T / √d_k)·V where Q ∈ R^{n×d_k} is the query matrix, K ∈ R^{m×d_k} is the key matrix, V ∈ R^{m×d_v} is the value matrix, d_k is the dimension of the key vector, and dividing by √d_k is used to prevent the softmax gradient from disappearing due to the large result of the dot product.
[0031] Further, in the S4 step, the modal confidence estimation includes adding a confidence prediction head for each modal branch to output the reliability scores w_img, w_odor, and w_txt of the modal features, with a value range of 0 to 1; the adaptive fusion weight is normalized by softmax: α = softmax([w_img, w_odor, w_txt]).
[0032] Further, the modal confidence estimation can also use an uncertainty perception method: output the mean μ_m and variance σ_m of the feature distribution for each modal branch 2 use an uncertainty-based weighting formula for fusion: α_m = (1 / σ_m 2 ) / Σ_k(1 / σ_k 2 ), so that the modal with smaller variance (high certainty) obtains a larger weight.
[0033] Further, in the S4 step, the adaptive weighted fusion includes: fusion feature F_fusion = α_img×f_img + α_odor×f_odor + α_txt×f_txt+ CrossAttn(f_img, f_odor, f_txt) where α_img, α_odor, and α_txt are the normalized modal weights, and CrossAttn is the cross-modal attention interaction feature.
[0034] Further, in the S5 step, the multi-task joint learning includes: main task: medicinal material category classification, loss function L_cls = CrossEntropy(pred_cls, label_cls) auxiliary task 1: origin discrimination, loss function L_origin = CrossEntropy(pred_origin, label_origin) auxiliary task 2: quality grade evaluation, loss function L_grade = CrossEntropy(pred_grade, label_grade) Total training loss: L_total = L_cls + λ_1 * L_origin + λ_2 * L_grade + λ_3 * L_contrast Wherein, λ_1, λ_2, λ_3 are weight coefficients of each task.
[0035] Further, the medicinal material category classification can adopt a prototype learning method, comprising: (a) Prototype initialization: maintain a learnable prototype vector P_c for each medicinal material category, c = 1, 2,..., C, wherein C is the total number of categories; (b) Similarity calculation: calculate the similarity s_c between the fusion feature F_fusion and each category prototype: sim(F_fusion, P_c); (c) Classification prediction: based on the similarity, calculate the category probability p_c through softmax: p_c = exp(s_c / τ) / Σ_kexp(s_k / τ); (d) Prototype update: update the prototype through momentum during training: P_c ← m·P_c + (1-m)·f_c, wherein f_c is the average feature of the category sample, and m is a momentum coefficient. The prototype learning method has better interpretability and small sample learning ability.
[0036] Further, the method further comprises a modal missing robustness training step: in the training stage, randomly mask one or two modal inputs with a preset probability, forcing the model to learn the ability to make accurate predictions under incomplete input conditions.
[0037] Further, the modal missing robustness training comprises: (a) Single modal dropout: randomly set the feature vector of a certain modal to zero or replace it with a learnable missing label vector with a probability p_1; (b) Double modal dropout: randomly set the feature vectors of two modal to zero or replace them with a probability p_2; (c) Dynamic completion of missing modal: using the features of the available modal, the feature representation of the missing modal is predicted through a cross-modal generation network (such as an MLP or a lightweight Transformer decoder) to realize dynamic completion.
[0038] The application also provides a traditional Chinese medicinal material intelligent identification system based on cross-modal contrast learning and adaptive attention fusion, comprising:
[0039] A multi-modal data acquisition module comprising an image acquisition unit, an odor acquisition unit and a text storage unit, for acquiring three-modal data of traditional Chinese medicinal material samples;
[0040] Image encoding module: using a pre-trained visual Transformer model to encode image data into a feature vector;
[0041] Odor encoding module: according to the odor data type, using a time series Transformer model, a one-dimensional convolutional neural network or a two-dimensional convolutional neural network to encode odor data into a feature vector;
[0042] Text encoding module: using a pre-trained language model to encode text description into a feature vector;
[0043] Cross-modal alignment module: using a contrastive learning method to realize semantic space alignment of three modal features;
[0044] Adaptive fusion module: including a cross-modal attention sub-module and a confidence estimation sub-module, realizing adaptive weighted multi-modal feature fusion;
[0045] Multi-task prediction module: including a category classification head, an origin discrimination head and a quality evaluation head, for multi-task joint prediction;
[0046] Result output module: integrating various prediction results to output an identification report.
[0047] Further, the odor collection unit includes one of the following two implementation modes or a combination thereof: Mode one, professional laboratory equipment scheme: (a) Electronic nose based on metal oxide sensor array: such as PEN3 type of AIRSENSE Analytics company in Germany, equipped with 10 MOS sensors (W1C, W5S, W3C, W6S, W5C, W1S, W1W, W2S, W2W, W3S), respectively sensitive to aromatic, nitrogen oxide, ammonia, hydrogen, alkyl aromatic, methane, sulfide, ethanol, etc.; (b) Electronic nose based on fast gas chromatography: such as Heracles II type or Heracles NEO type of Alpha MOS company in France, using double chromatographic column (MXT-5 non-polar column and MXT-1701 medium polar column) and double FID detector, analysis time about 2 minutes; (c) Gas chromatography-mass spectrometry (GC-MS): used for precise qualitative and quantitative analysis of volatile components, outputting one-dimensional chromatographic data; (d) Comprehensive two-dimensional gas chromatography-mass spectrometry (GCxGC-MS) or comprehensive two-dimensional gas chromatography-time-of-flight mass spectrometry (GCxGC-TOF-MS): two chromatographic columns with different polarities (such as non-polar DB-5ms and medium-polar DB-17ms) are connected in series, and comprehensive two-dimensional separation is achieved through a thermal modulator or a flow modulator, outputting a two-dimensional chromatogram image (a heat map with the first-dimensional retention time as the horizontal axis, the second-dimensional retention time as the vertical axis, and signal intensity as the color), and the peak capacity can reach 5000 to 10000, which can simultaneously detect hundreds to thousands of volatile compounds, and is suitable for comprehensive characterization of volatile oil of complex traditional Chinese medicinal materials; (e) Gas chromatography-ion mobility spectrometry (GC-IMS): such as FlavourSpec from G.A.S, Germany, which has both chromatographic separation capability and ppb-level detection sensitivity of ion mobility spectrometry; Mode two, low-cost portable solution: (a) Constant-temperature sealed sampling chamber: volume of 50 to 500 milliliters, equipped with temperature control device, temperature control accuracy ±1℃; (b) Multi-channel gas sensor array: including at least 8 detection channels, covering different gas types; (c) Analog-to-digital conversion and data acquisition circuit: sampling accuracy of 12 bits or more, sampling frequency of 1 to 10 Hz; (d) Microcontroller and communication interface.
[0048] Further, the multi-channel gas sensor array in the low-cost portable solution includes at least three of the following sensor types: (a) Environmental sensor: such as Bosch BME680, which can simultaneously detect temperature, humidity, air pressure and VOC total amount; (b) Metal oxide semiconductor (MOS) gas sensor: such as Hanwei MQ series (MQ-3 detects ethanol, MQ-5 detects flammable gas, and MQ-9 detects carbon monoxide) or Japan Figaro TGS series (TGS2600 detects air quality, TGS2602 detects VOC and ammonia, and TGS2620 detects organic solvent vapor); (c) Volatile organic compound (VOC) specific sensor: such as Weisheng WSP2110, MS1100, or Sensirion SGP30 / SGP40; (d) Multi-channel integrated gas sensor: such as Seeed Grove multi-channel gas sensor V2. Advantages
[0049] Compared with the prior art, the present application has the following advantages:
[0050] First, three-modal deep fusion. The application initiatively fuses three heterogeneous modal information of image, smell and text, fully simulates the comprehensive identification process of human experts "looking, smelling and knowing", and significantly improves the recognition ability of easy mixed varieties.
[0051] Second, cross-modal contrast learning alignment. The contrast learning method is used to map heterogeneous modalities to a unified semantic space, so that the multi-modal representations of the same medicinal material are close to each other, and effective cross-modal information integration is realized.
[0052] Third, adaptive attention fusion and uncertainty perception. The cross-modal attention mechanism is used to dynamically calculate the correlation between modalities, and the adaptive weighted fusion is realized by combining the modal confidence and uncertainty estimation. When the quality of certain modal data is poor, the weight of the data is automatically reduced, thereby improving the robustness and accuracy of the system.
[0053] Fourth, multi-task joint learning and prototype learning. Multiple related tasks such as medicinal material category, origin and quality are simultaneously learned, the generalization ability of feature representation is enhanced through knowledge transfer between tasks, and the prototype vector is maintained for each medicinal material category by combining the prototype learning method, thereby improving the explainability and small sample learning ability of classification.
[0054] Fifth, modal missing robustness and dynamic completion. Through the modal dropout training strategy and missing modal dynamic completion mechanism, the system can maintain high identification accuracy when part of the modal data is missing, thereby meeting the scene requirements of equipment failure or incomplete data in actual application.
[0055] Sixth, wide device adaptability. The smell collection module supports professional laboratory equipment such as PEN3, Heracles II, GC-MS, GCxGC-MS and GC-IMS to obtain high-precision analysis results, and also supports low-cost portable sensor array solutions to meet the needs of on-site rapid detection, thereby adapting to different application scenarios. In particular, the two-dimensional gas chromatography-mass spectrometry (GCxGC-MS) can output two-dimensional chromatographic images, the information amount is greatly improved, and hundreds to thousands of volatile compounds can be detected.
[0056] Seventh, in line with traditional Chinese medicine cognition. The three-modal fusion architecture of the application is in line with the "looking, smelling, asking and cutting" concept of traditional medicinal material identification, and is more easily accepted and used by practitioners in the traditional Chinese medicine industry. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 is the overall flowchart of the method of the application;
[0058] Figure 2 is the overall architecture diagram of the system of the application;
[0059] Figure 3Schematic diagram of cross-modal contrastive learning and attention fusion mechanism;
[0060] Figure 4 Hardware structure diagram of odor collection unit;
[0061] Figure 5 Schematic diagram of modal missing robustness training strategy. DETAILED DESCRIPTION
[0062] The application will be described in further detail below with reference to the drawings and specific embodiments.
[0063] Embodiment 1: Three-modal intelligent identification method for traditional Chinese medicinal materials
[0064] As shown in the figure, the intelligent identification method for traditional Chinese medicinal materials provided in the embodiment includes the following steps: Figure 1
[0065] Step one, multi-modal data collection
[0066] Image collection: Use a camera with more than 12 million pixels or a smartphone to take pictures of the traditional Chinese medicinal material samples in a standardized lighting environment. Collect 3 to 5 images of each sample from different angles, and adjust the resolution to 224x224 pixels. Use the RGB color space.
[0067] Odor collection: Choose one of the following schemes according to actual conditions: Scheme A (electronic nose equipment): Take 1 to 5 grams of traditional Chinese medicinal material samples and place them in a headspace sampling bottle. Use a PEN3 electronic nose or Heracles II ultra-fast gas-phase electronic nose for detection to obtain 10 to hundreds of channel odor response time series data. Scheme B (GC-MS equipment): Take an appropriate amount of traditional Chinese medicinal material samples, extract volatile components using headspace solid-phase microextraction (HS-SPME) or steam distillation, and analyze using a gas chromatograph-mass spectrometer to obtain one-dimensional chromatogram data. Scheme C (GCxGC-MS equipment): Take an appropriate amount of traditional Chinese medicinal material samples, extract volatile components using headspace solid-phase microextraction, and analyze using a two-dimensional gas chromatograph-mass spectrometer to obtain a two-dimensional chromatogram image (heat map). The image has the first dimension retention time as the horizontal axis, the second dimension retention time as the vertical axis, and the signal intensity represented by color depth. Scheme D (portable equipment): Place 5 to 10 grams of traditional Chinese medicinal material samples in a 200-milliliter constant-temperature sealed glass container, set the temperature to 25±1 degrees Celsius, and let the odor fully volatilize and balance for 3 to 5 minutes. Start the multi-channel gas sensor array to continuously collect data for 10 minutes at a sampling frequency of 1 hertz to obtain a multi-channel time series data matrix.
[0068] Text collection: Extract the standardized description text of each medicinal material from authoritative literature such as the Chinese Pharmacopoeia, Chinese Herbal Medicine Standards, and Chinese Herbal Medicine Dictionary, including information such as origin, nature and flavor, meridian tropism, efficacy and indication, and character description.
[0069] Step two, data preprocessing
[0070] Image preprocessing: Perform image normalization (scale pixel values to the 0-1 interval), random cropping, random horizontal flipping, color jitter, and other data augmentation operations.
[0071] Odor preprocessing: Different preprocessing methods are used according to the type of odor data: (a) For electronic nose time series data: Perform baseline correction (subtract air background signal), first-order time series difference, sliding window segmentation (window size 100, step size 50), and zero-mean unit variance standardization; (b) For one-dimensional GC-MS chromatographic data: Perform baseline correction, peak detection and alignment, and normalization processing; (c) For two-dimensional GCxGC-MS chromatographic images: Convert two-dimensional chromatographic data to heat map image format, perform image normalization, and resize (e.g., resize to 224x224 pixels).
[0072] Text preprocessing: Use Chinese word segmentation tools for word segmentation, add special marker symbols, and convert to the input format required by the language model.
[0073] Step three, feature encoding
[0074] Image feature encoding: Use the EVA-02-Large pre-trained model as the image encoder. Input a 224x224x3 image, pass it through image patching (patch size=14), linear embedding, adding a learnable class marker, and a 12-layer Transformer encoder, take the output corresponding to the class marker, and map it to a 768-dimensional image feature vector f_img through a projection layer.
[0075] Odor feature encoding: Select the corresponding encoder according to the type of odor data: (a) For electronic nose time series data: Use a time series Transformer encoder, input the odor time series matrix, add position encoding after first-order difference preprocessing, pass it through a multi-layer self-attention Transformer encoder, and use average pooling to obtain a 768-dimensional odor feature vector f_odor; (b) For one-dimensional GC-MS chromatographic data: Use a one-dimensional convolutional neural network (1D-CNN) for feature extraction; (c) For two-dimensional GCxGC-MS chromatogram images: the same or similar two-dimensional convolutional neural network or visual Transformer as the image encoder is used for feature extraction.
[0076] Text feature encoding: Chinese-BERT-wwm pre-training model is used. After processing the segmented text sequence through the 12-layer Transformer encoder, the hidden state at the CLS position is taken, and the projection layer is mapped to a 768-dimensional text feature vector f_txt.
[0077] Step four, cross-modal contrast alignment
[0078] As shown on the left, three sets of InfoNCE contrast loss are used to realize cross-modal alignment: Figure 3 For N medicinal material samples in the same batch, the image-odor feature pair, image-text feature pair, and odor-text feature pair of the same sample are taken as positive sample pairs, and the pairing between different samples is taken as negative sample pairs. Through contrast learning, the three-modal features of the same medicinal material are clustered in a 768-dimensional semantic space, and the features of different medicinal materials are far away from each other. The temperature parameter τ is set to 0.07.
[0079] Step five, adaptive attention fusion
[0080] As shown on the right, the cross-modal attention fusion process is as follows: Figure 3 (a) Image-odor interaction: taking f_img as Query, f_odor as Key and Value, the attention output f_img_odor of image to odor is calculated; (b) Image-text interaction: taking f_img as Query, f_txt as Key and Value, the attention output f_img_txt is calculated; (c) Odor-text interaction: taking f_odor as Query, f_txt as Key and Value, the attention output f_odor_txt is calculated; (d) Confidence estimation: each modality is predicted through an independent two-layer MLP to obtain confidence scores w_img, w_odor, and w_txt; (e) Adaptive fusion: F_fusion = softmax([w_img, w_odor, w_txt])·[f_img, f_odor, f_txt] + concat(f_img_odor, f_img_txt, f_odor_txt) The final fusion feature dimension is 768x4 = 3072.
[0081] Step six, multi-task joint learning
[0082] Based on the fusion feature F_fusion, three task heads are connected respectively: Class classification head: 3072→1024→N_cls (number of medicinal material categories), using cross-entropy loss; Origin discrimination head: 3072→512→N_origin (number of origins), using cross-entropy loss; Quality evaluation head: 3072→512→N_grade (number of quality grades), using cross-entropy loss. Total loss function: L_total = L_cls + 0.3×L_origin + 0.2×L_grade + 0.5×L_contrast
[0083] Step seven, modal missing robustness training
[0084] As shown in Figure 5 , a modal dropout strategy is introduced during training: Randomly set the image modal feature to zero with a probability of 20%; Randomly set the smell modal feature to zero with a probability of 20%; Randomly set the text modal feature to zero with a probability of 10%; Shield two modalities at the same time with a probability of 10%. The shielded modalities are replaced by learnable missing marker vectors, so that the model learns to make predictions under incomplete input.
[0085] Step eight, identification result output
[0086] In the inference stage, the available modal data of the medicinal material to be identified is input into the system, and the following outputs are output: Medicinal material category: the category with the highest probability and its confidence; Origin judgment: the origin with the highest probability and its confidence; Quality grade: predicted quality grade and its confidence; Comprehensive identification report: contains all the above information and contribution analysis of each modality.
[0087] Example 2: Three-modal intelligent identification system for traditional Chinese medicinal materials
[0088] As shown in Figure 2 , the intelligent identification system for traditional Chinese medicinal materials provided in this embodiment includes the following components:
[0089] Image acquisition unit: including industrial camera or smartphone camera, ring light, sample tray, image acquisition software.
[0090] Odor collection unit: the following options can be selected according to the application scenario: Option A (laboratory high-precision option): (a) PEN3 portable electronic nose of AIRSENSE Analytics, Germany: equipped with an array of 10 metal oxide semiconductor (MOS) sensors, sensor numbers and detection characteristics are as follows: W1C is sensitive to aromatic compounds, W5S is sensitive to nitrogen oxides and ozone, W3C is sensitive to ammonia and aromatic amines, W6S is sensitive to hydrogen, W5C is sensitive to alkane aromatic compounds, W1S is sensitive to short-chain alkanes and methane, W1W is sensitive to sulfides and terpenes, W2S is sensitive to ethanol and some aromatic compounds, W2W is sensitive to organic sulfides and aromatic compounds, W3S is sensitive to long-chain alkanes; (b) Heracles II or Heracles NEO ultra-fast gas chromatography electronic nose of Alpha MOS, France: using double chromatographic column technology (non-polar MXT-5 column and medium-polar MXT-1701 column), equipped with a double hydrogen flame ionization detector (FID), analysis time is about 120 seconds; (c) Comprehensive two-dimensional gas chromatography-mass spectrometry (GCxGC-MS) or comprehensive two-dimensional gas chromatography-time-of-flight mass spectrometry (GCxGC-TOF-MS): using two different polarity chromatographic columns in series and a thermal modulator to achieve comprehensive two-dimensional separation, outputting a two-dimensional chromatogram image (thermal map), with a peak capacity of up to 5000 to 10000, and capable of simultaneously detecting hundreds to thousands of volatile compounds; (d) Gas chromatography-ion mobility spectrometry (GC-IMS): such as FlavourSpec of G.A.S, Germany, combining the separation capability of gas chromatography and the ppb-level detection sensitivity of ion mobility spectrometry; (e) Gas chromatography-mass spectrometry (GC-MS): used for precise qualitative and quantitative analysis of volatile components, outputting one-dimensional chromatographic data; Option B (portable field option): (a) Constant temperature sealed sampling chamber: a borosilicate glass or stainless steel container with a volume of 100 to 300 milliliters, equipped with a semiconductor refrigeration or heating temperature control module, with a temperature control accuracy of ±0.5°C; (b) Multi-channel gas sensor array: using the following sensor combination, with 8 to 16 detection channels: - Bosch BME680 environmental sensor: integrated temperature, humidity, air pressure, VOC detection; - MQ series metal oxide semiconductor sensors: MQ-3 (ethanol), MQ-5 (flammable gas / LPG), MQ-9 (carbon monoxide); - Japanese Figaro TGS series sensors: TGS2600 (air quality), TGS2602 (VOC / ammonia), TGS2620 (organic solvents); - Figaro WSP2110 or MS1100 type VOC sensor; - Seeed Grove multi-channel gas sensor V2 (based on GM-102B / 302B / 502B / 702B); (c) Signal conditioning circuit: includes instrumentation amplifier, active low-pass filter (cutoff frequency 10 Hz), 12-bit or higher precision ADC; (d) Microcontroller: Arduino Mega 2560, ESP32 or STM32 series; (e) Communication interface: USB, WiFi or Bluetooth. The total hardware cost of Scheme B is about 300 to 800 RMB.
[0091] Text storage unit: stores a structured Chinese medicinal material description text database, covering standard descriptions of more than 500 commonly used Chinese medicinal materials.
[0092] Computing processing unit: runs deep learning inference program, which can use a workstation equipped with GPU or edge computing device. Model inference time is less than 100 milliseconds per sample.
[0093] Human-computer interaction interface: provides sample information input, photograph triggering, data visualization, identification report generation and other functions.
[0094] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A multimodal intelligent identification method for traditional Chinese medicinal materials, characterized in that, Includes the following steps: S1. Multimodal data acquisition steps: Collect image data, odor data, and text description data of Chinese medicinal material samples respectively; S2, Single-modal feature encoding steps: Extract features from each modal data using an image encoder, an odor encoder, and a text encoder respectively to obtain image feature vectors, odor feature vectors, and text feature vectors; S3. Cross-modal semantic alignment step: The cross-modal alignment method is used to map the features of different modalities to a unified semantic space, so that the multimodal features of the same medicinal material sample are close to each other in the semantic space; S4. Multimodal feature fusion step: The aligned multimodal features are fused to generate multimodal fused features; S5. Classification and identification steps: Classify and identify the medicinal materials based on the fused features, and output the identification results of the medicinal materials.
2. The method according to claim 1, characterized in that, In step S3, cross-modal semantic alignment adopts a contrastive learning method, which uses a contrastive learning loss function to bring the three-modal features of the same medicinal material sample closer together and push the features of different medicinal material samples further apart.
3. The method according to claim 1, characterized in that, In step S4, the multimodal feature fusion adopts an adaptive weighted fusion method, which includes: calculating the correlation information between modalities through a cross-modal interaction mechanism; and performing adaptive weighting in combination with modal confidence estimation to generate multimodal fused features.
4. The method according to claim 3, characterized in that, The cross-modal interaction mechanism employs an attention mechanism, including: calculating the attention interaction between image features and odor features; calculating the attention interaction between image features and text features; and calculating the attention interaction between odor features and text features.
5. The method according to claim 4, characterized in that, The attention interaction adopts the scaled dot product attention mechanism, and the calculation formula is: Attention(Q,K,V)=softmax(QK transpose / √d_k)×V, where Q is the query matrix, K is the key matrix, V is the value matrix, and d_k is the dimension of the key vector.
6. The method according to claim 2, characterized in that, The contrastive learning method employs the InfoNCE contrastive loss function, which includes image-odor contrastive loss, image-text contrastive loss, and odor-text contrastive loss.
7. The method according to claim 1, characterized in that, In step S1, odor data acquisition uses odor detection equipment to detect the volatile components of Chinese medicinal material samples, forming multi-channel time-series data, one-dimensional chromatographic data, or two-dimensional chromatographic image data.
8. The method according to claim 7, characterized in that, The odor detection device includes at least one of the following types: (1) Electronic nose devices: including portable electronic noses based on metal oxide sensor arrays or ultra-fast gas phase electronic noses based on gas chromatography principles; (2) Gas chromatography-mass spectrometry (GC-MS); (3) Two-dimensional gas chromatography-mass spectrometry (GC×GC-MS) or two-dimensional gas chromatography-time-of-flight mass spectrometry (GC×GC-TOF-MS); (4) Gas chromatography-ion mobility spectrometry (GC-IMS); (5) Multi-channel gas sensor array: including at least two types of metal oxide semiconductor sensors, electrochemical sensors, photoionization sensors, and environmental sensors, with no less than 8 detection channels.
9. The method according to claim 8, characterized in that, When the odor detection device is a two-dimensional gas chromatography-mass spectrometry instrument, the odor data is a two-dimensional chromatographic image, and the odor encoder uses an image feature extraction network to extract features from the two-dimensional chromatographic image.
10. The method according to claim 1, characterized in that, In step S1, the text description data includes at least two types of textual information from the following categories: source, properties, meridian tropism, efficacy, indications, morphological description, microscopic features, and physicochemical identification features of the Chinese medicinal materials.
11. The method according to claim 1, characterized in that, In step S2, the image encoder employs a deep neural network model, including convolutional neural networks, visual Transformers, or other image feature extraction networks.
12. The method according to claim 1, characterized in that, In step S2, the odor encoder selects the appropriate neural network structure based on the odor data type: (a) When the odor data is multi-channel time-series data, a time-series feature extraction network is used; (b) When the odor data is one-dimensional chromatographic data, a one-dimensional feature extraction network is used; (c) When the odor data is a two-dimensional chromatographic image, an image feature extraction network is used.
13. The method according to claim 1, characterized in that, In step S2, the text encoder uses a natural language processing model to extract features from the text description.
14. The method according to claim 3, characterized in that, The modality confidence estimation includes: setting up a confidence prediction network for each modality branch and outputting a reliability score for the modality feature; the adaptive weighted fusion is performed by normalizing and weighting based on the confidence scores of each modality.
15. The method according to claim 14, characterized in that, The modal confidence estimation also includes uncertainty-aware estimation: the mean and variance of the output feature distribution for each modal branch are fused using an uncertainty-based weighting formula, so that the modality with smaller variance receives greater weight.
16. The method according to claim 1, characterized in that, The method also includes a multi-task joint learning step: based on the fusion features, predictions are made for at least two tasks in medicinal material classification, origin identification, and quality grade assessment.
17. The method according to claim 1, characterized in that, The classification and recognition in step S5 adopts a prototype learning method, which includes: maintaining a learnable prototype vector for each medicinal material category; calculating the similarity between the fused features and the prototypes of each category; and performing classification prediction based on the similarity.
18. The method according to claim 1, characterized in that, The method also includes a modality missing robustness training step: during the training phase, one or more modal inputs are randomly masked with a preset probability, enabling the model to learn the ability to make predictions under incomplete input conditions.
19. The method according to claim 18, characterized in that, The modality missing robustness training includes: single-modal dropout, which randomly sets the feature vector of a certain modality to zero or replaces it with a learnable missing label vector with a first probability; dual-modal dropout, which randomly sets the feature vectors of two modalities to zero or replaces them with a second probability; and missing modality dynamic completion, which uses the features of available modalities to predict the feature representation of the missing modality through a cross-modal generation network.
20. A multimodal intelligent identification system for traditional Chinese medicinal materials, characterized in that, include: Multimodal data acquisition module: used to acquire image data, odor data, and text description data of Chinese medicinal material samples; Feature encoding module: used to extract features from each modality of data to obtain the feature vector of each modality; Cross-modal alignment module: used to map features from different modalities to a unified semantic space; Multimodal fusion module: Used to fuse multimodal features and generate fused features; Classification and prediction module: used to output the identification results of Chinese medicinal materials based on fusion features.
21. The system according to claim 20, characterized in that, The cross-modal alignment module uses a contrastive learning method to achieve semantic space alignment of trimodal features; the multimodal fusion module includes a cross-modal attention submodule and a confidence estimation submodule, which are used for adaptive weighted multimodal feature fusion.
22. The system according to claim 20, characterized in that, The multimodal data acquisition module includes an image acquisition unit, an odor acquisition unit, and a text storage unit; the odor acquisition unit includes one or a combination of the following: Option 1, Professional Equipment Solution: Use at least one of the following: electronic nose device, gas chromatography-mass spectrometry, full two-dimensional gas chromatography-mass spectrometry, or gas chromatography-ion mobility spectrometry. Option 2, portable device solution: includes a constant temperature sealed sampling chamber, a multi-channel gas sensor array, signal conditioning and analog-to-digital conversion circuit, microcontroller and communication interface.
23. The system according to claim 22, characterized in that, The electronic nose device includes a portable electronic nose based on a metal oxide sensor array or an ultra-fast gas chromatographic electronic nose based on the principle of gas chromatography; the full two-dimensional gas chromatograph-mass spectrometer outputs two-dimensional chromatographic image data; the multi-channel gas sensor array includes at least three types of environmental sensors, alcohol sensors, combustible gas sensors, carbon monoxide sensors, volatile organic compound sensors, and complex gas sensors.
24. The system according to claim 20, characterized in that, The system also includes a multi-task prediction module, which is used to simultaneously predict at least two of the following tasks: medicinal material classification, origin identification, and quality grade assessment.
25. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 19.
26. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1 to 19.