A Traditional Chinese Medicine Identification and Analysis System Based on Cluster Analysis

Through the collaborative design of multimodal hypergraph spatiotemporal encoder and drug-related causal intervention mechanism, the problem of insufficient adaptability of Chinese medicine identification technology in multimodal data fusion and dynamic scenarios is solved, and the multi-dimensional improvement and dynamic adaptability of Chinese medicine identification technology are achieved.

CN120067776BActive Publication Date: 2025-07-22CHANGCHUN UNIV OF CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510554266.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing traditional Chinese medicine identification technology has shortcomings in multimodal data fusion and dynamic scenario adaptability, and it is difficult to achieve the needs of multi-dimensional feature integration, drug properties knowledge guidance and dynamic continuous learning.

Method used

A Chinese medicine identification and analysis system based on cluster analysis is adopted, and a multimodal hypergraph spatiotemporal encoder and a drug-related causal intervention mechanism is combined with traditional Chinese medicine knowledge graph and causal reasoning to build a cross-modal feature space, and a dynamic incremental processing module is used for data update and cluster structure evolution.

Benefits of technology

It has achieved multi-dimensional improvement in traditional Chinese medicine identification technology, optimized cross-modal fusion accuracy, improved consistency between clustering results and pharmacopoeia standards, enhanced dynamic adaptability, and reduced model update cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067776B_ABST
    Figure CN120067776B_ABST
Patent Text Reader

Abstract

The present invention discloses a traditional Chinese medicine identification and analysis system based on clustering analysis. The present invention relates to the technical field of drug analysis and identification, and includes the following modules: a data acquisition module, which is used to obtain multi-modal data of traditional Chinese medicine and process asynchronous stream data. The multi-modal data includes microscopic images, spectral data, metabolomics data, and electronic nose / tongue sensing data. The processing of asynchronous stream data includes timestamp marking, multi-device data calibration, and feature caching. This traditional Chinese medicine identification and analysis system based on clustering analysis realizes a leapfrog improvement of traditional Chinese medicine identification technology in multiple dimensions and levels through the collaborative design of a multi-modal hypergraph spatio-temporal encoder and a medicinal property causal intervention mechanism, solves the core contradiction between traditional clustering algorithms and domain knowledge, and provides a general framework for cross-modal dynamic learning in the field of medical informatics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pharmaceutical analysis and identification, and specifically to a traditional Chinese medicine identification and analysis system based on cluster analysis. Background Art

[0002] With the application of instrumental analysis technology in the detection of traditional Chinese medicine, although chromatography, spectroscopy, etc. can provide some objective data, they are limited to single or a few indicators and cannot reflect the holistic characteristics of the "synergistic effect of multiple components" of traditional Chinese medicine. In this context, the traditional Chinese medicine identification technology based on cluster analysis has emerged, attempting to explore the classification rules of medicinal materials through unsupervised learning. However, there are many deep-seated technical defects in the existing systems. First, the foundation of cross-modal data fusion is weak. The modern detection of traditional Chinese medicine forms a situation where multi-source heterogeneous data coexist, such as microscopic images, near-infrared spectra, metabolomics data, and electronic nose / tongue data. These data belong to different feature spaces, with significant differences in dimension, distribution characteristics, and semantic levels. Traditional clustering algorithms rely on single-modal or simple fusion strategies, ignoring cross-modal non-linear associations, resulting in clustering results deviating from the true attributes of medicinal materials. Second, traditional Chinese medicine theory is separated from data-driven clustering. The classification of traditional Chinese medicine needs to consider both data characteristics and medicinal property knowledge. Traditional clustering algorithms only rely on data similarity metrics and lack the embedding of prior knowledge of traditional Chinese medicine. For example, "cold-natured" and "hot-natured" medicinal materials may be mis-clustered due to similar chemical compositions, while medicinal materials with synergistic effects but different appearances are wrongly separated. Existing semi-supervised clustering also cannot systematically integrate the classification rules in ancient books and is difficult to achieve semantic alignment between clustering results and abstract medicinal property attributes. Third, the adaptability of static clustering models to dynamic real-world scenarios is poor. Traditional Chinese medicine identification faces dynamic challenges such as the discovery of new varieties, environmental changes, and the evolution of processing techniques. Traditional clustering systems use fixed parameters. When new data exceeds the range, it is necessary to manually reset the parameters and retrain the entire model, with high maintenance costs and unable to reflect the evolution of the attributes of medicinal materials in real time. Incremental clustering algorithms are prone to causing feature space distortion and cluster structure collapse when dealing with multi-modal asynchronous updates. Although existing patented technologies have improvements, they have not broken through the bottleneck of multi-modal fusion and solved the problem of model adaptability in dynamic scenarios, and are difficult to meet the complex requirements of traditional Chinese medicine identification for multi-dimensional feature integration, medicinal property knowledge guidance, and dynamic continuous learning. Summary of the Invention

[0003] (I) Technical Problems to be Solved

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a traditional Chinese medicine identification and analysis system based on cluster analysis, which solves the problems that although existing patented technologies have improvements, they have not broken through the bottleneck of multi-modal fusion and solved the problem of model adaptability in dynamic scenarios, and are difficult to meet the complex requirements of traditional Chinese medicine identification for multi-dimensional feature integration, medicinal property knowledge guidance, and dynamic continuous learning.

[0005] (II) Technical Solutions

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A traditional Chinese medicine identification and analysis system based on clustering analysis, comprising:

[0007] A data acquisition module, configured to obtain multi-modal data of traditional Chinese medicine, including microscopic images, spectral data, metabolomics data, and electronic nose / tongue sensing data, and perform timestamp marking and caching on asynchronous stream data;

[0008] A multi-modal hypergraph spatio-temporal encoder, which constructs a cross-modal feature space based on a dynamic hyper-edge generation algorithm and realizes non-linear fusion of multi-source heterogeneous data through hypergraph spatio-temporal convolution;

[0009] A property causal intervention clustering module, which combines the traditional Chinese medicine knowledge graph with causal reasoning to generate a clustering objective function constrained by medicine properties and performs two-stage intervention on cluster drift;

[0010] A dynamic incremental processing module, which adopts an elastic computing architecture to realize incremental feature update and cluster structure evolution of multi-modal stream data and supports edge-cloud collaborative computing.

[0011] Preferably, the data acquisition module includes:

[0012] An asynchronous stream data interface, configured to receive modal data with different sampling frequencies and align the lagging modal features through a timestamp compensation algorithm;

[0013] A multi-device calibration unit, which uses a differentiable distribution alignment algorithm to perform dynamic domain adaptation on heterogeneous sensor data and eliminate feature scale offsets caused by device accuracy differences;

[0014] A feature preprocessing unit, which performs local texture enhancement on microscopic images, baseline correction and noise filtering on spectral data, and standard normalization processing on metabolomics data.

[0015] Preferably, the construction method of the multi-modal hypergraph spatio-temporal encoder includes:

[0016] a: Semantic hyper-edge generation step: Based on the medicinal material efficacy classification rules recorded in the Chinese Pharmacopoeia, multi-modal feature nodes belonging to the same efficacy category are forcibly connected to form a hyper-edge structure guided by prior knowledge;

[0017] b: Data hyper-edge generation step: The auto-encoder model is used to learn the non-linear association of cross-modal feature distributions, and cross-modal data-driven hyper-edges are automatically generated according to feature similarity;

[0018] c: Time hyper-edge generation step: For asynchronously arriving stream data, a time series prediction network is introduced to compensate and predict the features of delayed modalities, and time dimension compensation hyper-edges are generated;

[0019] d: Hypergraph spatio-temporal convolution step: Design a dual-channel convolution kernel. Among them, the first channel captures the high-order spatial correlations of cross-modal features, and the second channel dynamically adjusts the weight ratio of historical features through a temporal decay mechanism for joint spatio-temporal dimensional modeling.

[0020] Preferably, the implementation of the drug property causal intervention clustering module includes:

[0021] e: Drug property knowledge causal graph construction step: Transform the hierarchical classification rules of "nature, flavor, and meridian tropism" and "synergistic efficacy" in traditional Chinese medicine classics into causal graph nodes, and quantify the causal influence intensity between nodes through an attention mechanism; that is, transform the medicinal material classification rules in "Compendium of Materia Medica" into graph nodes, and quantify the causal intensity of "nature, flavor → meridian tropism → efficacy" through the graph attention network GAT.

[0022] f: Causal constraint clustering optimization step: Introduce the semantic constraint term of the drug property causal graph into the deep clustering objective function to force the clustering result to be consistent with the causal logic of the medicinal material efficacy classification; that is, jointly optimize the deep clustering loss, the KL divergence constraint of drug property knowledge, and the dynamic time warping DTW drift penalty term.

[0023] g: Two-stage drift intervention step: Detect the cluster semantic offset in the rough screening stage, and trigger the GAN to reconstruct the abnormal cluster features and compare them with the pharmacopoeia data in the fine screening stage to distinguish real drift from noise; that is, detect the cluster semantic offset based on the drug property causal graph in the first stage, and distinguish real drug property drift from data noise through feature reconstruction and comparison with the pharmacopoeia data in the second stage, and correct the abnormal clusters.

[0024] Preferably, the dynamic incremental processing module includes:

[0025] h: Drift detection and response unit: Calculate the trajectory difference between the newly added data and the historical cluster centroid based on the dynamic time warping DTW algorithm. If it exceeds the threshold, trigger cluster splitting; that is, analyze the trajectory difference between the newly added data and the historical cluster centroid through the dynamic time warping algorithm, and trigger cluster splitting or merging operations when the difference exceeds the preset threshold.

[0026] i: Elastic feature update unit: Dynamically allocate modal computing resources through differentiable neural architecture search DNAS, and give priority to processing high information entropy modalities; that is, dynamically allocate computing resources and give priority to processing key modal data according to the real-time calculated modal information entropy and drug property knowledge requirements.

[0027] j: Edge-cloud collaboration mechanism: Deploy binary MobileNetV3 at the edge to extract local features, and run the full-modal hypergraph encoder in the cloud and regularly send down the distilled model; that is, deploy a lightweight feature extraction model at the edge, reduce the dimension and compress the high-dimensional data and upload it to the cloud; after the cloud completes hypergraph encoding and clustering calculations, regularly send down the simplified model optimized by knowledge distillation to the edge.

[0028] Preferably, the dual-stage drift intervention step g includes:

[0029] g1: Coarse-grained semantic drift detection: Calculate the semantic attribute offset of the current cluster center in the medicinal property causal graph. If the weight changes of attributes such as "taste and nature", "meridian tropism", "cold nature", "hot nature", etc. exceed ±15%, it is marked as a high-risk cluster.

[0030] g2: Fine-grained feature reconstruction verification: Perform GAN reconstruction on the features of the high-risk cluster, and calculate the Wasserstein distance between the reconstructed data and the same type of medicinal materials recorded in the pharmacopoeia. If the distance value is greater than the preset threshold and confirmed by expert blind evaluation, it is determined as a real medicinal property drift and a new cluster is created, that is: perform adversarial generation network reconstruction on the original features of the high-risk abnormal cluster, compare the distribution similarity between the reconstructed features and the historical data of the same type of medicinal materials recorded in the pharmacopoeia. If the similarity is lower than the preset threshold and confirmed by expert evaluation, it is determined as a real medicinal property drift and a new cluster is created.

[0031] Preferably, the calculation method of the time decay factor of the hypergraph spatio-temporal convolution is as follows:

[0032] ;

[0033] where α: a learnable parameter that controls the steepness of the decay curve. The initial value is 0.1 and is optimized by backpropagation. β: a learnable parameter that controls the time sensitivity of the decay rate. The initial value is 0.01 and is dynamically adjusted during training, t current is the current system timestamp, unit: millisecond, t feature is the feature generation timestamp, representing the time point when data collection or calculation is completed. The weight of historical features decays exponentially with the delay time; : the time decay factor, which is used to dynamically adjust the weight of historical features in the hypergraph spatio-temporal convolution.

[0034] That is, the time decay mechanism in step d of the construction method of the multi-modal hypergraph spatio-temporal encoder is realized in the following way: According to the delay duration between the feature generation timestamp and the current system time, dynamically adjust the weight ratio of historical features in the convolution calculation. The longer the delay time, the greater the weight decay amplitude; adopt a learnable non-linear parameterized form to express the decay function, which can adapt to the timeliness differences of different modal data.

[0035] Preferably, it further includes an internal consistency verification unit, which is configured as:

[0036] k: For the data of the same batch, generate features by using hypergraph encoding and traditional tandem encoding respectively, calculate the Wasserstein distribution distance. If the distance reduction of hypergraph encoding is ≥ 40%, it is determined that the fusion is effective, that is: for the multi-modal data of the same medicinal material batch, generate feature vectors by using hypergraph encoding and traditional tandem encoding respectively, and verify the cross-modal fusion effect of hypergraph encoding through the distribution distance metric algorithm;

[0037] l: Project the clustering cluster center into the sub-space of the medicinal property knowledge graph. If the cosine similarity with the corresponding medicinal property node < 0.85, trigger the causal intervention module; that is: project the clustering cluster center into the embedding space of the medicinal property knowledge graph, calculate the semantic similarity between the cluster center and the corresponding medicinal property node. If it is lower than the preset threshold, trigger the causal intervention module to re-optimize the clustering result.

[0038] Preferably, it further includes an external validity verification unit, which is configured as:

[0039] o: Bidirectionally match the clustering result with the Chinese Pharmacopoeia, and retrieve the implicit associated descriptions in the classics for the unmatched clusters, that is: bidirectionally match the system clustering result with the medicinal material classification in the Chinese Pharmacopoeia, and automatically retrieve the implicit associated descriptions in the classics for the unmatched clustering clusters to supplement the explanation;

[0040] p: Calculate the Kappa consistency coefficient through expert blind evaluation. If ≥ 0.75 and the misjudged cases are verified as real drifts through GAN reconstruction, output the final identification result, that is: organize multiple rounds of expert blind evaluation verification, conduct consistency analysis on the system output result and the independent classification result of senior pharmacists. If it meets the preset consistency standard and the misjudged cases are verified as real drifts through feature reconstruction, confirm the effectiveness of the system.

[0041] Preferably, the system is extended and applied to the scenarios of origin traceability of Chinese medicinal materials, quality monitoring of processing technology and analysis of incompatibility of traditional Chinese medicine, including:

[0042] q: Origin traceability of Chinese medicinal materials, construct an origin feature fingerprint library by clustering and analyzing the multi-modal feature differences of medicinal materials from different origins;

[0043] r: Quality monitoring of processing technology, analyze the dynamic changes of chemical components and physical properties during the processing process, and identify process deviations;

[0044] s: Analysis of incompatibility of traditional Chinese medicine, mine the potential conflict patterns of medicinal material combinations based on the medicinal property causal graph, and provide compatibility safety warnings.

[0045] (III) Beneficial effects

[0046] The present invention provides a traditional Chinese medicine identification and analysis system based on clustering analysis. It has the following beneficial effects:

[0047] The traditional Chinese medicine identification and analysis system based on clustering analysis realizes a leapfrog improvement of traditional Chinese medicine identification technology in multiple dimensions and at multiple levels through the collaborative design of a multimodal hypergraph spatio-temporal encoder and a medicinal property causal intervention mechanism. At the technical performance level, the cross-modal fusion accuracy is significantly optimized. The hypergraph encoder breaks through the limitations of traditional feature concatenation or kernel fusion through the dynamic generation mechanism of semantic hyperedges, data hyperedges, and time hyperedges, reduces the feature alignment error between microscopic images and spectral data, and improves the recognition accuracy of key components in metabolomics data. The clustering interpretability is enhanced. The two-stage intervention mechanism driven by the medicinal property causal graph improves the Kappa consistency coefficient between the clustering results and the classification standards of the Chinese Pharmacopoeia, and successfully identifies various local medicinal materials not recorded in the pharmacopoeia. After pharmacological experiments verify their efficacy, it promotes the dynamic expansion of the traditional Chinese medicine knowledge system. The dynamic adaptability is comprehensively improved. The elastic incremental architecture combined with edge-cloud collaborative computing reduces the model update time and the inference energy consumption at the edge, solves the core contradiction of the disconnection between traditional clustering algorithms and domain knowledge, and provides a general framework for cross-modal dynamic learning in the field of medical informatics. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the overall framework of the present invention;

[0049] Figure 2 It is a schematic flowchart of the construction method of the multimodal hypergraph spatio-temporal encoder of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Please refer to Figure 1 and Figure 2 The present invention provides a technical solution: a traditional Chinese medicine identification and analysis system based on clustering analysis, including the following modules:

[0052] A data acquisition module is used to obtain multi-modal data of traditional Chinese medicine and process asynchronous stream data. The multi-modal data includes microscopic images, spectral data, metabolomics data, and electronic nose / tongue sensing data. Processing asynchronous stream data includes timestamp marking, multi-device data calibration, and feature caching. Among them, in the process of obtaining multi-modal data in the data acquisition module, for the acquisition of microscopic images, a high-resolution microscope Olympus BX53 is used to collect images of medicinal material sections, 200 images are collected in each batch, with a resolution of 512×512, and they are stored in PNG format; for the acquisition of spectral data, a near-infrared spectrometer OceanInsight HR4000 is used to scan medicinal material powders, with a wavelength range of 200 - 1100 nm and a sampling interval of 1 nm, generating spectral curve data; for the acquisition of metabolomics data, a liquid chromatography-mass spectrometry (LC-MS) instrument is used to detect medicinal material extracts, recording the mass-to-charge ratio (m / z) and peak area, generating a data matrix in CSV format; for the acquisition of electronic nose / tongue data, a sensor array is used to collect volatile components and taste electrochemical signals of medicinal materials, 10 seconds of data is collected for each sample, and the sampling frequency is 100 Hz.

[0053] The processing of asynchronous stream data includes:

[0054] Timestamp marking: Add a timestamp accurate to milliseconds to each batch of data and store it in the metadata field;

[0055] Caching mechanism: Use a Redis database to cache the modal data that arrives late. For example, if the metabolomics data is delayed by 1 hour, set the TTL survival time to 24 hours.

[0056] A multi-modal hypergraph spatio-temporal encoder constructs a cross-modal unified feature space based on a dynamically generated hyperedge structure and performs non-linear fusion of multi-source heterogeneous data through spatio-temporal coupled convolution operations. Among them, the dynamic hyperedge generation includes semantic hyperedges, data hyperedges, and time hyperedges. Semantic hyperedges, based on the list of "heat-clearing" medicinal materials in the Chinese Pharmacopoeia, such as Coptis chinensis and Lonicera japonica, forcefully connect their microscopic image, spectral, and metabolomics data nodes into hyperedges; data hyperedges, train an improved Wasserstein autoencoder with a hidden layer dimension of 256 and a LeakyReLU activation function, input microscopic images (LBP texture features) and spectral data (first derivative features), output a cross-modal correlation weight matrix, and generate data-driven hyperedges; time hyperedges, for metabolomics data with a delay of more than 1 hour, predict missing features through an LSTM-Transformer hybrid network (2 LSTM layers, 4 Transformer heads) and generate compensatory hyperedges.

[0057] The hypergraph spatio-temporal convolution process includes a spatial convolution kernel and a temporal decay factor; the spatial convolution kernel is a 3×3 convolution kernel with a stride of 1 and a padding method of "same" to extract cross-modal spatial correlation features; for the temporal decay factor, the initial weight decay rate β = 0.01, and based on the time difference Δt between the timestamp generated according to the feature and the current time, the historical feature weight is dynamically adjusted to 1 / (1 + β·Δt).

[0058] The drug property causal intervention clustering module combines the traditional Chinese medicine knowledge graph with dynamic causal reasoning to generate a clustering objective function constrained by drug properties and performs semantic-level intervention on the cluster drift during the clustering process;

[0059] Among them, the drug property causal intervention clustering module:

[0060] Construction of the drug property knowledge causal graph:

[0061] Take "nature and flavor" (cold, hot, warm, cool), "meridian tropism" (lung meridian, liver meridian, etc.), and "efficacy" (clearing heat, tonifying qi, etc.) as graph nodes;

[0062] Calculate the influence intensity between nodes through a graph attention network (GAT, number of heads 4), such as the weight of "cold nature → heat-clearing efficacy" is 0.92.

[0063] Causal constraint clustering optimization:

[0064] Setting of the objective function weights: α = 0.6 (clustering loss), β = 0.3 (causal constraint), γ = 0.1 (drift penalty);

[0065] Select the Adam optimizer, learning rate 0.001, and the number of training epochs is 100.

[0066] The dynamic incremental processing module uses an elastic computing architecture for incremental feature update and cluster structure evolution of multi-modal flow data, supporting adaptive resource allocation between the edge and the cloud; among them, the dynamic incremental processing module includes drift detection and edge-cloud collaboration;

[0067] Drift detection: Set the length of the dynamic time warping (DTW) window to 10 and calculate the trajectory difference between the newly added data and the historical cluster centroid; among them, the difference threshold is set to 0.3, and if the threshold is exceeded, cluster splitting is triggered.

[0068] Edge-cloud collaboration: Deploy binary MobileNetV3 at the edge, with a model size of 8MB, to extract local texture features of microscopic images; the cloud issues a distilled model every 6 hours, with the TinyBERT architecture and the number of parameters reduced to 1 / 10.

[0069] The data acquisition module includes:

[0070] An asynchronous streaming data interface, configured to receive modal data with different sampling frequencies and align the lagging modal features through a timestamp compensation algorithm. Among them, the timestamp compensation algorithm performs cubic spline interpolation on the metabolomics data that arrives late, based on the data at adjacent timestamps, such as spectral data at times t-1 and t+1, to fill in the missing features. The interpolated data is smoothed in time series through an LSTM unit, where the hidden layer dimension of the LSTM unit is 128;

[0071] A multi-device calibration unit that uses a differentiable distribution alignment algorithm to perform dynamic domain adaptation on heterogeneous sensor data and eliminate the feature scale shift caused by device accuracy differences; including: differentiable Wasserstein distance: calculating the distribution difference of data collected by different electronic nose devices, and the formula is:

[0072] ;

[0073] Optimize sensor parameters (such as gain coefficients) through gradient descent to minimize the distribution difference between devices, with 100 iterations and a learning rate of 0.01;

[0074] Among them, P: the data distribution of the reconstructed features, such as the GAN-generated data of the high-risk cluster; Q: the historical data distribution of the same type of medicinal materials in the pharmacopoeia database; Γ(P,Q): the set of all possible joint distributions of P and Q; γ: an instance in the joint distribution, (x,y): a data pair sampled from γ, x comes from P, and y comes from Q, : the Euclidean distance between samples x and y, inf: the infimum, indicating finding the smallest possible expected distance;

[0075] The feature preprocessing unit performs local texture enhancement on microscopic images, baseline correction and noise filtering on spectral data, and standard normalization processing on metabolomics data; among them, the microscopic image processing in the feature preprocessing unit includes local texture enhancement and background removal; local texture enhancement uses the CLAHE algorithm (contrast-limited adaptive histogram equalization), with a block size of 8×8 and a contrast limit of 2.0; background removal extracts the medicinal material area through Otsu threshold segmentation, fills the holes, and saves it as a mask image. The spectral data processing in the feature preprocessing unit includes baseline correction and noise filtering; baseline correction uses the AsymmetricLeastSquares algorithm, with a smoothing parameter of 1e5 and an asymmetry factor of 0.05; noise filtering uses the Savitzky-Golay filter, with a window length of 11 and a polynomial order of 3. The standardization of metabolomics data in the feature preprocessing unit is performed through Z-score normalization, calculating the mean μ and standard deviation σ for each feature dimension to generate a standardization matrix.

[0076] The construction method of the multi-modal hypergraph spatio-temporal encoder includes:

[0077] a: Semantic hyperedge generation step: Based on the medicinal material efficacy classification rules recorded in the Chinese Pharmacopoeia, multi-modal feature nodes belonging to the same efficacy category are forcibly connected to form a hyperedge structure guided by prior knowledge. The process includes: parsing the electronic version of the Chinese Pharmacopoeia (XML format), extracting medicinal material efficacy classification labels, such as "diaphoretic class" and "qi-tonifying class"; for each efficacy category, forcibly connecting its corresponding multi-modal feature nodes into hyperedges. The multi-modal feature nodes include microscopic images, spectra, and metabolomic data; the hyperedge weight is initialized to 1.0 and dynamically adjusted through backpropagation during training.

[0078] b: Data hyperedge generation step: Through an autoencoder model, learn the non-linear association of cross-modal feature distributions, and automatically generate cross-modal data-driven hyperedges according to feature similarity. The process includes: through an improved Wasserstein autoencoder, the encoder structure is a 4-layer fully connected layer, with input dimensions 1024→512→256→128, the activation function is LeakyReLU, and the negative slope is 0.2; the decoder structure is a symmetric 4-layer fully connected layer, and the output dimension is the same as the input; the reconstruction loss is MSE + Wasserstein distance loss, with a weight of 0.5. Input microscopic images and spectral data, and output a cross-modal association matrix with dimensions 256×300. If the elements in the matrix are greater than the threshold of 0.7, data hyperedges are generated. Among them, for microscopic images: LBP features, with a dimension of 256; for spectral data: first derivative features, with a dimension of 300.

[0079] c: Temporal hyperedge generation step: For streaming data that arrives asynchronously, introduce a temporal prediction network to compensate and predict the features of the delayed modality, and generate temporal dimension compensation hyperedges. The process includes:

[0080] LSTM-Transformer hybrid network:

[0081] LSTM part: 2-layer LSTM, with a hidden layer dimension of 256, processing the time series of delayed metabolomic data;

[0082] Transformer part: 4-head self-attention mechanism, with a positional encoding dimension of 256;

[0083] Output layer: fully connected layer (dimension 256→metabolomic feature dimension 128), predicting missing features.

[0084] For data with a delay exceeding 1 hour, generate temporal compensation hyperedges, with the weight initialized to 0.8 and dynamically adjusted according to the prediction error.

[0085] d: Steps of hypergraph spatio-temporal convolution: Design a two-channel convolution kernel. Among them, the first channel captures the high-order spatial correlations of cross-modal features, and the second channel dynamically adjusts the weight ratio of historical features through a time decay mechanism for joint spatio-temporal dimensional modeling;

[0086] The steps of hypergraph spatio-temporal convolution include:

[0087] Spatial convolution kernel: A 3×3 convolution kernel with a stride of 1 and 64 output channels to extract cross-modal spatial features;

[0088] Time decay factor: Calculate the weight decay coefficient as 1 / (1 + 0.01·Δt) according to the time difference Δt between the feature generation timestamp and the current time;

[0089] After multiplying the convolution result by the decay coefficient, the final feature is output through the ReLU activation function.

[0090] Through the semantic, data, and time hyper-edge dynamic generation mechanism, multi-modal-spatio-temporal coupling is realized, breaking through the limitations of traditional multi-modal fusion that only relies on feature concatenation or kernel methods.

[0091] The implementation of the drug property causal intervention clustering module includes:

[0092] e: Steps of constructing the drug property knowledge causal graph: Convert the hierarchical classification rules of "taste, property, and meridian tropism" and "synergistic efficacy" in traditional Chinese medicine classics into causal graph nodes, and quantify the causal influence intensity between nodes through an attention mechanism;

[0093] The process of constructing the drug property knowledge causal graph includes:

[0094] Extract the medicinal material attributes (taste, property, meridian tropism, efficacy) from the electronic version of Compendium of Materia Medica, and construct triples, such as "Coptis chinensis → cold nature → clearing heat";

[0095] Use the Graph Attention Network (GAT, number of heads 4) to calculate the influence weights between nodes. For example, the weight of "cold nature → clearing heat" is 0.92, and the weight of "bitter taste → clearing heat" is 0.85;

[0096] Store it in the format of an adjacency matrix (number of nodes N×N), and the matrix elements are weight values.

[0097] f: Steps of causal constraint clustering optimization: Introduce the semantic constraint term of the drug property causal graph into the deep clustering objective function to force the clustering result to be consistent with the causal logic of the medicinal material efficacy classification; The process is as follows:

[0098] Causal constraint clustering objective function:

[0099] Deep clustering loss: Adopt the deep embedding clustering DEC loss, including KL divergence constraint;

[0100] KL divergence constraint of drug property knowledge: Calculate the KL divergence between the cluster center and the causal graph nodes. The formula is:

[0101] ;

[0102] where L causal : The causal constraint term of drug property in the clustering objective function, which is used to force the clustering result to align with the traditional Chinese medicine theory knowledge. C: The total number of clusters, such as the number of clusters preset or generated dynamically. P c : The actual distribution of the drug properties of the c-th cluster, which is obtained by statistically analyzing the drug property attributes of the samples within the cluster. For example, the proportion of "cold nature" is 0.7. Q c : The theoretical distribution of the c-th cluster in the drug property causal graph, which is predefined according to the Chinese Pharmacopoeia or traditional Chinese medicine classics. For example, the theoretical proportion of "cold nature" is 0.85. D KL : Kullback-Leibler divergence, which is used to measure the difference between the actual distribution and the theoretical distribution;

[0103] Drift penalty term: Calculate the trajectory difference between the new data and the historical clusters through dynamic time warping (DTW). When the difference value exceeds the threshold of 0.3, the penalty coefficient γ = 0.1.

[0104] g: Two-stage drift intervention step: Detect the semantic drift of the cluster based on the drug property causal graph in the first stage, and distinguish the real drug property drift from the data noise by comparing the feature reconstruction with the pharmacopoeia data in the second stage, and correct the abnormal clusters.

[0105] The two-stage drift intervention process is as follows:

[0106] Coarse screening stage: Calculate the attribute offset of the cluster center in the causal graph. For example, the weight of "cold nature" drops by 15%; if the offset exceeds ±15%, it is marked as a high-risk cluster.

[0107] Fine screening stage: For the features of the high-risk clusters, use a GAN to generate reconstructed data. Generator: 4-layer fully connected. Discriminator: 3-layer CNN; Calculate the Wasserstein distance between the reconstructed data and the same type of medicinal materials in the pharmacopoeia, and set the threshold to 0.25; if the distance > 0.25 and confirmed by experts, it is determined as a real drift and a new cluster is created.

[0108] Drug property causal intervention transforms the traditional Chinese medicine theory into a causally constrained term that can be optimized mathematically, solving the problem of the disconnection between traditional clustering algorithms and domain knowledge.

[0109] The dynamic incremental processing module includes:

[0110] h: Drift detection and response unit: Analyze the trajectory difference between the new data and the centroid of the historical clusters through the dynamic time warping algorithm, and trigger the cluster splitting or merging operation when the difference exceeds the preset threshold;

[0111] The drift detection and response process is as follows:

[0112] Dynamic Time Warping (DTW): Calculate the minimum path cost between the new data sequence X new and the historical cluster centroid sequence X hist :

[0113] ;

[0114] X new : The feature sequence of the new data, with dimensions n×d, where n is the number of time steps and d is the feature dimension. X hist : The feature sequence of the historical cluster centroid, with dimensions consistent with X new . W: The set of alignment paths, representing the time correspondence of feature points between X new and X hist . x i : The feature vector at the i-th time step in X new . yj: The feature vector at the j-th time step in X hist . ‖·‖: Euclidean distance, used to calculate the difference between feature vectors;

[0115] Set the window constraint to 10, and trigger cluster splitting when the path cost exceeds 0.3.

[0116] Cluster splitting operation: Divide the new data into sub-clusters, recalculate the centroids, and verify the semantic rationality of the new clusters through the drug property causal graph.

[0117] i: Elastic feature update unit: Dynamically allocate computing resources and preferentially process key modal data according to the real-time calculated modal information entropy and drug property knowledge requirements;

[0118] The elastic feature update process includes:

[0119] Modal information entropy calculation: Calculate the Shannon entropy for each modal data:

[0120] ; H(X): Information entropy, n: The total number of possible value categories of modal data X, i: Index variable, x i : The i-th value category of modal data X, p(x i ): The probability that modal data X takes the value x i ;

[0121] Modal data with information entropy higher than the threshold, such as H>2.5, is preferentially processed.

[0122] Differentiable Neural Architecture Search (DNAS): The search space includes convolutional kernel sizes (3×3, 5×5), number of channels (64, 128), activation functions (ReLU, LeakyReLU); Search for the optimal subgraph every 24 hours, and the search takes about 30 minutes.

[0123] j: Edge-cloud collaboration mechanism: Deploy a lightweight feature extraction model at the edge, reduce the dimension of high-dimensional data and compress it before uploading to the cloud; after the cloud completes hypergraph encoding and clustering calculations, regularly send the simplified model optimized by knowledge distillation to the edge.

[0124] The edge-cloud collaboration mechanism includes:

[0125] Edge deployment: Binary MobileNetV3: Convert floating-point weights to 8-bit integers, and compress the model size to 8MB; process microscopic images in real time at 15 frames per second and upload feature vectors with a dimension of 256.

[0126] Cloud model distribution: Generate a simplified model, TinyBERT, with 4 layers and 4 heads every 6 hours through knowledge distillation; the size of the distributed model is 20MB, and the loading time at the edge is <1 second.

[0127] The elastic incremental architecture realizes dynamic allocation of modal computing resources through DNAS, which is different from the fixed modal weights of traditional federated learning solutions.

[0128] The two-stage drift intervention step g includes:

[0129] g1: Coarse-grained semantic drift detection: Calculate the semantic attribute offset of the current clustering cluster center in the medicinal property causal graph. If the weight change of core attributes such as "taste" and "meridian tropism" exceeds the preset threshold, it is marked as a high-risk abnormal cluster;

[0130] The process of coarse-grained semantic drift detection includes:

[0131] Calculation of semantic attribute offset: Extract the medicinal property attribute vector of the current clustering cluster center, such as the weight of "cold nature" is 0.7 and the weight of "meridian tropism - lung meridian" is 0.8; obtain the corresponding theoretical attribute vector of this cluster from the medicinal property causal graph, such as the weight of "cold nature" is 0.85 and the weight of "meridian tropism - lung meridian" is 0.9;

[0132] Calculate the offset percentage:

[0133] , where V current : The actual medicinal property attribute vector of the current clustering cluster center, V theory : The theoretical medicinal property attribute vector of the corresponding medicinal material category in the medicinal property causal graph;

[0134] If the offset rate of core attributes such as "cold nature" and "hot nature" exceeds ±15%, it is marked as a high-risk cluster.

[0135] Marking and caching of high-risk clusters: Store the feature vectors and metadata of high-risk clusters: timestamp, source device, in the temporary database MySQL, and the retention period is 48 hours.

[0136] g2: Fine-grained feature reconstruction verification: Reconstruct the original features of high-risk abnormal clusters using a generative adversarial network, and compare the distribution similarity between the reconstructed features and the historical data of the same type of medicinal materials recorded in the pharmacopoeia. If the similarity is lower than the preset threshold and is confirmed by expert evaluation, it is determined as a real medicinal property drift and a new cluster is created.

[0137] The fine-grained feature reconstruction verification process includes:

[0138] GAN reconstruction network training:

[0139] Generator: A 4-layer fully connected network with input dimensions 256→512→256→128, and the activation function is LeakyReLU;

[0140] Discriminator: A 3-layer convolutional network with the number of channels 64→128→256, and the output is the probability of the authenticity of the reconstructed data;

[0141] Loss function: WassersteinGAN loss, and the gradient penalty coefficient λ = 10.

[0142] Reconstructed feature comparison: Input the original features of the high-risk cluster to generate reconstructed features; Extract the historical data of the same type of medicinal materials from the pharmacopoeia database, such as 100 sets of metabolome data of "Coptis chinensis"; Calculate the Wasserstein distance between the reconstructed features and the historical data: ; If the distance value > 0.25 (preset threshold), it is determined as a real drift.

[0143] Expert verification process: Anonymize the reconstructed data and the original data and submit them to 3 senior traditional Chinese medicine pharmacists; The experts independently evaluate the medicinal property attributes, such as the intensity of "cold nature". If at least 2 people confirm the attribute change, it is finally determined as a real drift.

[0144] New cluster creation and knowledge update: Create an independent new cluster, associate a new label, such as "Angelica sinensis from Gansu"; Update the medicinal property causal graph, add new nodes or adjust the edge weights, such as the weight of "Angelica sinensis from Gansu → meridian tropism - liver meridian" is 0.75.

[0145] The time decay mechanism in step d of the construction method of the multi-modal hypergraph spatio-temporal encoder is implemented as follows: According to the delay duration between the feature generation timestamp and the current system time, dynamically adjust the weight ratio of historical features in the convolution calculation. The longer the delay time, the greater the weight decay amplitude; Use a learnable non-linear parameterized form to express the decay function, which can adapt to the timeliness differences of different modal data.

[0146] The implementation process includes time difference calculation, non-linear decay function design, weight dynamic adjustment, and training and optimization;

[0147] Among them, the time difference calculation includes: Obtain the feature generation timestamp tfeature such as 20231010120000000 and the current system time t current ; calculate the latency time (in milliseconds), Δt: latency time, t current : current system time, t feature : feature generation timestamp.

[0148] The initial attenuation function in the design of the non-linear attenuation function is set to:

[0149]

[0150] where is a learnable parameter, with an initial value ; optimize through backpropagation , and the loss function is the feature reconstruction error MSE.

[0151] The weight dynamic adjustment includes: for the historical feature vector F hist , calculate the attenuated weight:

[0152] ;

[0153] If △t = 3600000ms (1 hour), , then λ = 0.72, and the historical feature weight decays by 28%.

[0154] Training and optimization include: during the training of the hypergraph encoder, jointly optimize the convolutional kernel parameters and the attenuation function parameters; after each round of training, adjust the learning rate according to the validation set loss, with an initial value of 0.001 and an attenuation factor of 0.1 / 10 rounds.

[0155] It also includes an internal consistency check unit, which is configured as:

[0156] k: For the multi-modal data of the same medicinal material batch, generate feature vectors using hypergraph encoding and traditional concatenated encoding respectively, and verify the cross-modal fusion effect of hypergraph encoding through the distribution distance metric algorithm;

[0157] l: Project the cluster center to the embedding space of the medicinal property knowledge graph, calculate the semantic similarity between the cluster center and the corresponding medicinal property node, and if it is lower than the preset threshold, trigger the causal intervention module to re-optimize the clustering result.

[0158] The internal consistency check unit includes hypergraph fusion effect verification and medicinal property semantic alignment verification, and the implementation process includes cross-modal fusion effect verification, medicinal property semantic alignment verification, and verification result feedback;

[0159] The cross-modal fusion effect verification includes:

[0160] Feature Generation Comparison: For the medicinal material data of the same batch, such as 50 groups of "ginseng" samples, the hypergraph encoder is used to generate the feature vector F hyper , and the traditional tandem encoder generates F concat ; The tandem coding method is to directly splice the microscopic image (256 dimensions), spectrum (300 dimensions), and metabolome (128 dimensions) data into a 684-dimensional vector.

[0161] Distribution Distance Calculation: The Wasserstein distance is used to measure the distribution difference between the two feature vectors; if the distribution distance of the hypergraph encoding drops by ≥40% compared to the traditional method, it is determined that the fusion is effective. For example, if it drops from 1.2 to 0.7, the fusion is determined to be effective.

[0162] The inspection of the alignment of medicinal property semantics includes:

[0163] Knowledge Graph Projection: The cluster center vector C c is input into the graph embedding model of the medicinal property causal graph, such as Node2Vec, with a dimension of 128; the corresponding medicinal property node embedding vector E node is obtained, such as the embedding of the "heat-clearing" node.

[0164] Similarity Calculation: Calculate the cosine similarity:

[0165] ; C c : The cluster center vector, E node : The medicinal property node embedding vector, Sim: The cosine similarity; if the similarity < 0.85 (the preset threshold), the causal intervention module is triggered to re-optimize the clustering objective function.

[0166] The feedback of the verification result includes: If the verification of three consecutive batches of data fails, the system automatically triggers the retraining of the full-scale model; when retraining, the weight β of the medicinal property causal constraint is increased to 0.5 to strengthen the semantic alignment.

[0167] It also includes an external validity verification unit, which is configured as:

[0168] o: Bidirectionally match the system clustering results with the medicinal material classification in the Chinese Pharmacopoeia, and automatically retrieve the implicit associated descriptions in the classics to supplement the explanations for the unmatched clusters;

[0169] p: Organize multiple rounds of expert blind reviews for verification, conduct a consistency analysis of the system output results and the independent classification results of senior pharmacists. If the preset consistency standard is reached and the misjudged cases are verified as real drifts through feature reconstruction, the system validity is confirmed.

[0170] The external validity verification unit includes pharmacopoeia matching and expert blind review. The implementation process includes two-way pharmacopoeia matching and expert blind review verification. Among them, the two-way pharmacopoeia matching includes forward screening and reverse screening. Forward screening is to compare the clustering results, such as: cluster label "heat-clearing category - Coptis chinensis", with the medicinal material classification entries recorded in the Chinese Pharmacopoeia one by one. If the match is successful, it is recorded as "verified". If the match fails, such as the new cluster "unknown category - A", it enters the reverse screening process.

[0171] Reverse screening is to retrieve paragraphs containing similar descriptions, such as "bitter and cold in nature, returning to the heart meridian", from the full-text pharmacopoeia database, such as TCM-ID, for the unmatched clusters. The TF-IDF algorithm is used to calculate the text similarity. If the highest similarity > 0.6, it is associated as "implicit match".

[0172] Expert blind review verification includes data anonymization processing, consistency analysis, and misjudgment case handling. Data anonymization processing includes: randomly shuffling the order of the clustering results, hiding the algorithm information, and generating an evaluation report in PDF format, including medicinal material images, spectral curves, and clustering labels, such as "Cluster-12".

[0173] Consistency analysis includes: 3 experts independently classify and record their judgments on the medicinal properties, such as "cold in nature, returning to the lung meridian"; calculate the Fleiss’ Kappa coefficient:

[0174] ;

[0175] Among them, P o is the actual agreement rate, representing the consistency ratio of the actual classification results among experts. The calculation formula:

[0176] ; Among them, P o : actual agreement rate, N is the number of samples, k is the number of classification categories, nij is the number of experts who classify the i-th sample into the j-th category, and M is the total number of experts;

[0177] P e is the expected random agreement rate, representing the expected value of consistency when experts classify randomly. The calculation formula:

[0178] ; Among them, P e : expected random agreement rate, p j is the proportion of the j-th category in all classifications, k: consistency coefficient, with a value range of [-1, 1]. The larger the value, the higher the consistency among experts;

[0179] If k ≥ 0.75, that is, strong consistency, it is determined that the system is effective.

[0180] The handling of misjudgment cases includes: for cases where there is a discrepancy between the expert and the system, such as Cluster-12 being judged as "hot" by the expert, triggering GAN reconstruction verification; if the distance between the reconstructed features and the pharmacopoeia data is > 0.25, it is submitted to the Pharmacopoeia Committee for re-review.

[0181] Set internal double verification and external double screening to ensure the credibility of the technical solution and avoid the defect of relying solely on a single verification method.

[0182] The system is extended and applied to the following scenarios:

[0183] q: Traceability of the origin of Chinese medicinal materials. By clustering and analyzing the multi-modal feature differences of medicinal materials from different origins, a fingerprint library of origin features is constructed;

[0184] r: Quality monitoring of the processing technology. Analyze the dynamic changes of chemical components and physical properties during the processing, and identify process deviations;

[0185] s: Analysis of incompatibility of traditional Chinese medicine. Based on the causal graph of medicinal properties, mine the potential conflict patterns of medicinal combinations and provide warnings for compatibility safety.

[0186] It should be further noted that in the specific implementation process, the system is extended and applied to the traceability of the origin of medicinal materials, the monitoring of the processing technology, and the analysis of incompatibility. The implementation process includes the traceability of the origin of medicinal materials, the quality monitoring of the processing technology, and the analysis of incompatibility of traditional Chinese medicine; among them, the traceability of the origin of medicinal materials includes the construction of a fingerprint library of features, accuracy testing, and a traceability query interface; the construction of the fingerprint library of features includes: collecting multi-modal data of medicinal materials from different origins, such as Angelica sinensis from Yunnan vs. Gansu; a hypergraph encoder generates origin feature vectors, and a discriminant model for origin is trained through an SVM classifier; the accuracy testing includes: for 200 batches of data, the discriminant accuracy of the origin reaches 96.3%. The traceability query interface includes: the user uploads an image of the medicinal material, and the system returns the probability distribution of the origin, such as "Yunnan: 82%, Gansu: 18%".

[0187] The quality monitoring of the processing technology includes dynamic process analysis: real-time collection of temperature, humidity sensor data and microscopic images during the processing; generating hypergraph features every 5 minutes and monitoring the trajectory of the cluster centroid; if the centroid deviation exceeds the threshold (DTW > 0.3), trigger a process anomaly alarm. Case: During the processing of batch D of "Rheum palmatum", it is detected that the weight of the "purgative effect" decreases by 20%, and the system prompts "over-fried".

[0188] The analysis of incompatibility of traditional Chinese medicine includes the mining of medicinal property conflicts and a visualization interface; the mining of medicinal property conflicts includes: inputting a compatibility plan, such as "Aconite + Pinellia ternata", and extracting the nodes of the causal graph of the medicinal properties of each medicinal material; calculating the conflict score between nodes through a graph neural network (GNN):

[0189] ;

[0190] where w is the medicinal property weight, A is the conflict relationship matrix, such as the "Eighteen Incompatibilities" rule, : conflict score; u, v: respectively represent two nodes in the medicinal property causal diagram; w u , w v : the weight values of nodes u and v, reflecting the importance or intensity of the node in the current compatibility scheme; if the conflict score > 0.5, a red warning is generated. The visualization interface includes: displaying the medicinal material compatibility network diagram, highlighting the conflict edges, such as the "Fuzi - Banxia" edge is marked in red.

[0191] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0192] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A traditional Chinese medicine identification and analysis system based on cluster analysis, characterized in that It includes the following modules: A data acquisition module, which is used to obtain multi-modal data of traditional Chinese medicine and process asynchronous stream data. The multi-modal data includes microscopic images, spectral data, metabolomics data, and electronic nose / tongue sensing data. The processing of asynchronous stream data includes timestamp marking, multi-device data calibration, and feature caching; A multi-modal hypergraph spatio-temporal encoder, which constructs a cross-modal unified feature space based on a dynamically generated hyperedge structure and performs non-linear fusion of multi-source heterogeneous data through spatio-temporal coupled convolution operations; A medicinal property causal intervention clustering module, which combines the traditional Chinese medicine knowledge graph with dynamic causal reasoning to generate a clustering objective function constrained by medicinal properties and performs semantic-level intervention on cluster drift during the clustering process; A dynamic incremental processing module, which uses an elastic computing architecture for incremental feature update and cluster structure evolution of multi-modal stream data, and supports adaptive resource allocation between the edge and the cloud.

2. The Chinese medicine identification and analysis system based on cluster analysis according to claim 1, characterized in that: The data acquisition module includes: An asynchronous stream data interface, configured to receive modal data with different sampling frequencies and align the lagging modal features through a timestamp compensation algorithm; A multi-device calibration unit, which uses a differentiable distribution alignment algorithm to perform dynamic domain adaptation on heterogeneous sensor data and eliminate feature scale offsets caused by device precision differences; A feature preprocessing unit, which performs local texture enhancement on microscopic images, baseline correction and noise filtering on spectral data, and standard normalization processing on metabolomics data.

3. The Chinese medicine identification and analysis system based on cluster analysis according to claim 2, characterized in that: The construction method of the multi-modal hypergraph spatio-temporal encoder includes: a: Based on the medicinal efficacy classification rules recorded in the Chinese Pharmacopoeia, forcibly connect multi-modal feature nodes belonging to the same efficacy category to form a hyperedge structure guided by prior knowledge; b: Learn the non-linear association of cross-modal feature distributions through an autoencoder model, and automatically generate cross-modal data-driven hyperedges according to feature similarity; c: For asynchronously arriving stream data, introduce a time series prediction network to compensate and predict the features of delayed modalities, and generate time dimension compensation hyperedges; d: Design a dual-channel convolutional kernel. Among them, the first channel captures the high-order spatial association of cross-modal features, and the second channel dynamically adjusts the weight ratio of historical features through a time decay mechanism for joint spatio-temporal dimension modeling.

4. The Chinese medicine identification and analysis system based on clustering analysis according to claim 3, characterized in that: The implementation of the medicinal property causal intervention clustering module includes: e: Convert the hierarchical classification rules of "nature, flavor, and meridian tropism" and "efficacy synergy" in traditional Chinese medicine classics into causal graph nodes, and quantify the causal influence intensity between nodes through an attention mechanism; f: Introduce a semantic constraint term of the medicinal property causal graph into the deep clustering objective function to force the clustering result to be consistent with the causal logic of medicinal material efficacy classification; g: Detect cluster semantic drift based on the medicinal property causal graph in the first stage, and distinguish real medicinal property drift from data noise by comparing feature reconstruction with pharmacopoeia data in the second stage, and correct abnormal clusters.

5. The Chinese medicine identification and analysis system based on clustering analysis according to claim 4, wherein: The dynamic incremental processing module includes: h: A drift detection and response unit: Analyze the trajectory differences between newly added data and historical cluster centroids through a dynamic time warping algorithm, and trigger cluster splitting or merging operations when the differences exceed a preset threshold; The drift detection and response process is as follows: Dynamic Time Warping (DTW): Calculate the new data sequence X new and the historical cluster centroid sequence X hist of the minimum path cost: ; X new : The feature sequence of the newly added data, with a dimension of n×d, where n is the number of time steps and d is the feature dimension, X hist : The feature sequence of the historical cluster centroids, with the same dimension as X new being consistent, W: The set of alignment paths, representing the temporal correspondence of the feature points between X new and X hist , x i : X new The feature vector at the i-th time step in, y j : X hist The feature vector at the j-th time step in, ‖·‖: Euclidean distance, used to calculate the difference between feature vectors; Set the window constraint to 10, and trigger cluster splitting when the path cost exceeds 0.3; Cluster splitting operation: divide the newly added data into sub-clusters, recalculate the centroids, and verify the semantic rationality of the new clusters through the medicine property causal graph; i: Elastic feature update unit: dynamically allocate computing resources and preferentially process key modal data according to the real-time calculated modal information entropy and medicine property knowledge requirements; j: Edge-cloud collaboration mechanism: deploy a lightweight feature extraction model at the edge, reduce the dimension and compress the high-dimensional data and then upload it to the cloud; after the cloud completes hypergraph encoding and clustering calculation, regularly send the simplified model optimized by knowledge distillation to the edge.

6. The Chinese medicine identification and analysis system based on clustering analysis according to claim 5, characterized in that: In the medicine property causal intervention clustering module, g includes: g1: Calculate the semantic attribute offset of the current clustering cluster center in the medicine property causal graph. If the weight change of the core attribute exceeds the preset threshold, mark it as a high-risk abnormal cluster; g2: Reconstruct the original features of the high-risk abnormal cluster through an adversarial generation network, compare the distribution similarity between the reconstructed features and the historical data of the same type of medicinal materials recorded in the pharmacopoeia. If the similarity is lower than the preset threshold and is confirmed by expert evaluation, it is determined as a real medicine property drift and a new cluster is created.

7. The Chinese medicine identification and analysis system based on cluster analysis according to claim 6, characterized in that: The time decay mechanism in the hypergraph spatio-temporal encoder step is implemented as follows: Dynamically adjust the weight ratio of historical features in the convolution calculation according to the delay duration between the feature generation timestamp and the current system time. The longer the delay time, the greater the weight decay amplitude; The decay function adopts a learnable non-linear parameterized form, which can adapt to the timeliness differences of different modal data.

8. The Chinese medicine identification and analysis system based on clustering analysis according to claim 7, characterized in that: It also includes an internal consistency verification unit, which is configured as: k: Generate feature vectors for the multi-modal data of the same medicinal material batch by using hypergraph encoding and traditional serial encoding respectively, and verify the cross-modal fusion effect of hypergraph encoding through the distribution distance metric algorithm; l: Project the clustering cluster center into the embedding space of the medicine property knowledge graph, calculate the semantic similarity between the cluster center and the corresponding medicine property node. If it is lower than the preset threshold, trigger the causal intervention module to re-optimize the clustering result.

9. The Chinese medicine identification and analysis system based on clustering analysis according to claim 8, characterized in that: It also includes an external validity verification unit, which is configured as: o: Bidirectionally match the system clustering result with the medicinal material classification in the Chinese Pharmacopoeia, and automatically retrieve the implicit association descriptions in the classics to supplement the explanation for the unmatched clustering clusters; p: Organize multiple rounds of expert blind evaluation verification, conduct a consistency analysis between the system output result and the independent classification result of senior pharmacists. If the preset consistency standard is reached and the misjudged cases are verified as real drifts through feature reconstruction, confirm the system validity.

10. A traditional Chinese medicine identification and analysis system based on cluster analysis according to claim 9, characterized in that: The system is extended and applied to the following scenarios: q: Traceability of the origin of Chinese medicinal materials, construct an origin feature fingerprint library by clustering and analyzing the multi-modal feature differences of medicinal materials from different origins; r: Quality monitoring of the processing technology, analyze the dynamic changes of chemical components and physical properties during the processing process, and identify process deviations; s: Analysis of incompatibility of traditional Chinese medicine combinations, mine potential conflict patterns of medicinal material combinations based on the medicine property causal graph, and provide compatibility safety warnings.

Citation Information

Patent Citations

  • Interstitial lung disease knowledge base construction method, device and system

    CN118428460A

  • Traditional Chinese medicine material authenticity identification platform

    CN119470828A