Traditional Chinese medicine analysis and identification method and system based on clustering analysis

Through dynamic weight allocation and high-order nonlinear collaborative characterization technology, the problem of multimodal data fusion failure in traditional Chinese medicine identification is solved, and the high accuracy and robustness of traditional Chinese medicine identification is achieved, and the quality standardization and intelligent analysis of traditional Chinese medicine is supported.

CN120387104AActive Publication Date: 2025-07-29CHANGCHUN UNIV OF CHINESE MEDICINE

Patent Information

Application Number
CN202510873723.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The prior art cannot effectively handle the dynamic correlation and coordinated characterization of multimodal data in Chinese medicine identification, which leads to the difficulty of adaptively adjusting the weight ratio of different modes when the model faces uneven sample distribution, noise interference or migration of origin, resulting in blurred cluster boundaries and an increase in the rate of misjudgment.

Method used

Through dynamic weight allocation mechanism and higher-order nonlinear collaborative representation technology, a cross-modal attention matrix is generated by a dual-flow deep convolution network and a bidirectional long and short-term memory network, combined with high-order tensor fusion space and adversarial generation training, dynamically adjust the modal weights, and an adaptive nuclear spectral clustering algorithm generates Chinese medicine category labels, filters out abnormal samples and optimizes the model through dual-channel verification.

Benefits of technology

It significantly improves the identification accuracy in complex scenarios such as differences in related species and processed products, reduces the misjudgment rate, ensures the robustness of the model under noise interference and data distribution offset, and provides technical support for standardizing the quality of traditional Chinese medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387104A_ABST
    Figure CN120387104A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine analysis and identification method and system based on clustering analysis, and relates to the technical field of traditional Chinese medicine analysis, and the method comprises the following steps: S1, collecting a chemical fingerprint spectrum, a microscopic image, metabonomics data and geographical indication information of a traditional Chinese medicine sample; according to the traditional Chinese medicine analysis and identification method and system based on clustering analysis, through a dynamic weight distribution mechanism and a high-order nonlinear collaborative characterization technology, the core problem of multi-modal data fusion failure in traditional Chinese medicine identification is effectively solved; the dynamic weight distribution network adaptively adjusts the contribution degree of each modal based on real-time entropy fluctuation and sample density, so that the identification precision under complicated scenes such as related species and processed product difference is remarkably improved; the high-order tensor fusion space is combined with adversarial generation training, the characterization limitation of a linear kernel function on component-form-effect nonlinear correlation is broken through, and the clustering boundary better fits the overall characteristics of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traditional Chinese medicine drug analysis, and specifically provides a method and system for analyzing and identifying traditional Chinese medicine based on cluster analysis. Background Art

[0002] Traditional methods for analyzing and identifying traditional Chinese medicine have long faced the core challenge of the failure of multi-source heterogeneous data fusion. As a complex material system, the identification of traditional Chinese medicine requires comprehensive multi-dimensional information such as chemical components, microscopic structure, metabolic characteristics, and production area environment. However, the existing technical framework based on cluster analysis cannot effectively handle the dynamic association and collaborative representation of such cross-modal data. Traditional methods usually integrate data features from different sources in a preset weight or simple superposition manner, ignoring the dynamic change law of the contribution degree between data modalities in the scenario of traditional Chinese medicine identification. For example, when the identification object is closely related species, the spatial topological features of microscopic images may be more discriminative than chemical fingerprints; while in the analysis of the differences in processing techniques, the dynamic changes of metabolomics become the key discriminant basis. The existing static fusion mechanism makes it difficult for the model to adaptively adjust the weight ratio of different modalities in the face of uneven sample distribution, processing noise interference, or production area migration, resulting in blurred clustering boundaries and increased misjudgment rates. More seriously, the non-linear coupling relationship between chemical components and morphological features cannot be accurately modeled by linear kernel functions or fixed similarity metrics. For example, the concentration gradient change of the effective components of medicinal materials may show a high-order non-linear association with the topological evolution of its microscopic structure, while the existing clustering algorithms can only capture shallow linear features, leading to the deviation of the identification results from the overall characteristics of traditional Chinese medicine "composition-morphology-efficacy". This rigid problem of the data fusion mechanism has become the core bottleneck restricting the in-depth application of cluster analysis technology in the field of traditional Chinese medicine identification. Summary of the Invention

[0003] (I) Technical Problems to be Solved

[0004] Aiming at the deficiencies of the prior art, the present invention provides a method and system for analyzing and identifying traditional Chinese medicine based on cluster analysis, which solves the core technical problem of the failure of dynamic weight allocation and non-linear collaborative representation of multi-modal data.

[0005] (II) Technical Solutions

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for analyzing and identifying traditional Chinese medicine based on cluster analysis, comprising the following steps:

[0007] S1: Collect the chemical fingerprint spectra, microscopic images, metabolomics data, and geographical indication information of traditional Chinese medicine samples. Perform wavelet packet transform on the chemical fingerprint spectra to extract frequency domain features and eliminate baseline drift. Perform superpixel segmentation on the microscopic images to construct a cell wall topological structure diagram. Use the sliding window method to extract time series dynamic features from the metabolomics data. Perform spatial interpolation on the geographical indication information to generate an environmental factor distribution matrix;

[0008] S2: Respectively extract the frequency domain features of chemical fingerprints and the topological features of microscopic images through a two-stream deep convolutional network. Use a bidirectional long short-term memory network to analyze the implicit correlation between metabolomics data and geographical indication information, and generate a cross-modal attention matrix;

[0009] S3: Based on the entropy value fluctuations and sample distribution densities of each modal feature vector calculated in real time, generate modal weight coefficients through a dynamic gating mechanism, and dynamically adjust the fusion ratio of chemical fingerprint features, microscopic image features, and metabolomics features;

[0010] S4: Map the weighted multi-modal features to a preset high-order tensor fusion space, construct pseudo-samples through adversarial generative training, and learn the non-linear mapping function between chemical component concentrations and microscopic structure topologies;

[0011] S5: Based on the local density distribution of samples in the high-order tensor space, use an adaptive kernel spectral clustering algorithm to generate traditional Chinese medicine category labels, and reverse decode the fusion features through a variational autoencoder to screen out abnormal samples whose reconstruction errors exceed the preset threshold to trigger weight reallocation. It should be further noted that the traditional Chinese medicine drug analysis and identification method based on clustering analysis includes dual-channel verification and model iteration: apply adversarial perturbations to the input data to generate controllable noise, monitor the singular value changes of the Jacobian matrix of the clustering results to adjust the robustness threshold; backpropagate the verification results to the dynamic weight allocation network and high-order representation engine through a two-way feedback optimization loop to realize the iterative update of model parameters.

[0012] Preferably, the implementation of the dynamic gating mechanism in step S3 includes:

[0013] Real-time monitor the information entropy change rate of the frequency domain features of chemical fingerprints to generate a first weight correction factor;

[0014] Calculate the sample distribution density through a kernel density estimator to generate a second weight correction factor;

[0015] Input the first weight correction factor and the second weight correction factor into a multi-layer perceptron to output modal weight coefficients.

[0016] Preferably, the learning of the non-linear mapping function in step S4 includes:

[0017] Construct a deep metric learning network to generate pseudo-samples with changing chemical component concentration gradients through a generative adversarial network;

[0018] In the high-order tensor fusion space, enforce the physical relevance between the clustering boundary and the topological evolution path of the microstructure.

[0019] Preferably, the implementation of the adaptive kernel spectral clustering algorithm in step S5 includes:

[0020] Dynamically adjust the bandwidth parameter of the Gaussian kernel function according to the density gradient of the samples in the high-order tensor space;

[0021] Determine the optimal number of clusters by calculating the eigenvalue decay rate of the Laplacian matrix.

[0022] Preferably, step S5 further includes:

[0023] Apply adversarial perturbations to the input data to generate controllable noise;

[0024] Monitor the change in the singular value of the Jacobian matrix of the clustering result and dynamically adjust the robustness threshold of the model to noise.

[0025] A traditional Chinese medicine drug analysis and identification system based on clustering analysis, comprising:

[0026] A multi-modal data acquisition module integrating a spectrometer, a microscopic imaging device and an environmental sensor, used to obtain chemical fingerprint spectra, microscopic images and geographical indication information;

[0027] A dynamic weight allocation module with a dual-stream deep convolutional network and a bidirectional long short-term memory network built in, used to generate a cross-modal attention matrix and dynamic weight coefficients;

[0028] A high-order collaborative representation module containing a tensor fusion engine and a generative adversarial network, used to construct a non-linear mapping function;

[0029] A clustering verification module deploying an adaptive kernel spectral clustering algorithm and a variational autoencoder, used to generate class labels and screen abnormal samples.

[0030] Preferably, the hardware cooperation relationship in the system includes: an entropy value fluctuation monitor in the dynamic weight allocation module is connected to the output end of the dual-stream deep convolutional network to calculate the change rate of the information entropy of the chemical fingerprint frequency domain features, and a kernel density estimator is connected to the output end of the bidirectional long short-term memory network to generate a sample distribution density correction factor; the topological constraint unit of the high-order collaborative representation module is electrically connected to the tensor fusion engine to enforce the consistency of the clustering boundary.

[0031] Preferably, the verification mechanism is implemented as follows: The input end of the adversarial perturbation generator of the clustering verification module is connected to the multi-modal data acquisition module to apply controllable noise, and the input end of the singular value monitoring unit is connected to the output end of the adaptive kernel spectral clustering algorithm to adjust the robustness threshold, ensuring the stability of the model under noise interference.

[0032] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps of the method for analyzing and identifying traditional Chinese medicine based on cluster analysis.

[0033] (III) Beneficial effects

[0034] The present invention provides a method and system for analyzing and identifying traditional Chinese medicine based on cluster analysis, having the following beneficial effects:

[0035] (I). The method and system for analyzing and identifying traditional Chinese medicine based on cluster analysis effectively solve the core problem of multi-modal data fusion failure in traditional Chinese medicine identification through the dynamic weight allocation mechanism and the high-order non-linear collaborative representation technology. The dynamic weight allocation network adaptively adjusts the contribution degree of each modality based on real-time entropy value fluctuations and sample density, significantly improving the identification accuracy in complex scenarios such as differences between closely related species and processed products; the high-order tensor fusion space combines adversarial generation training to break through the representation limitation of the linear kernel function for the non-linear correlation of "composition - morphology - efficacy", making the clustering boundary more conform to the overall characteristics of traditional Chinese medicine. The dual-channel verification mechanism ensures the robustness of the model under noise interference and data distribution deviation through double screening of feature reconstruction and adversarial stability, reducing the misjudgment rate compared with traditional methods, and providing reliable technical support for the standardization of traditional Chinese medicine quality.

[0036] (II). The method and system for analyzing and identifying traditional Chinese medicine based on cluster analysis achieve the leap from experience dependence to intelligent analysis in traditional Chinese medicine identification through hardware modular design and in-depth algorithm collaboration. The built-in dynamic gating unit and topological constraint module in the system transform traditional identification experience into quantifiable physical rules, retaining the overall view of traditional Chinese medicine while endowing the model with interpretability; the adaptive kernel parameter optimization and two-way feedback mechanism significantly reduce the manual parameter adjustment cost, making the method highly applicable in scenarios such as grass-roots drug quality inspection and rapid screening in cross-border trade. In addition, the multi-modal data fusion framework provides an extensible technical platform for the traceability of traditional Chinese medicine resources and the research on the material basis of drug efficacy, promoting the deep integration of digitalization and intelligentization in the traditional Chinese medicine industry. Description of the drawings

[0037] Figure 1 It is a schematic framework diagram of the whole of the present invention;

[0038] Figure 2 It is a control logic timing diagram of the present invention. Detailed implementation manners

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] See also Figure 1 and Figure 2 The present invention provides a technical solution: a method for analyzing and identifying traditional Chinese medicine based on cluster analysis, comprising the following steps:

[0041] S1: Collect chemical fingerprints, microscopic images, metabolomics data, and geographical indication information of traditional Chinese medicine samples. Perform wavelet packet transform on the chemical fingerprints to extract frequency domain features and eliminate baseline drift. Perform superpixel segmentation on the microscopic images to construct cell wall topology maps. Use the sliding window method to extract time series dynamic features from the metabolomics data. Perform spatial interpolation on the geographical indication information to generate an environmental factor distribution matrix.

[0042] S2: A two-stream deep convolutional network is used to extract the frequency domain features of chemical fingerprints and the topological features of microscopic images. A bidirectional long short-term memory network is used to analyze the implicit correlation between metabolomics data and geographical indication information, generating a cross-modal attention matrix.

[0043] S3: Based on the real-time calculated entropy fluctuations of each modality’s feature vectors and the sample distribution density, a dynamic gating mechanism is used to generate modality weight coefficients and dynamically adjust the fusion ratio of chemical fingerprint features, microscopic image features, and metabolomics features.

[0044] S4: Map the weighted multimodal features to a preset high-order tensor fusion space, construct pseudo samples through adversarial generative training, and learn the nonlinear mapping function between chemical component concentration and microstructure topology; it should be further explained that the construction of the high-order tensor fusion space includes: mapping chemical fingerprints, microscopic images and metabolomics features into third-order tensors respectively, and generating a joint representation matrix through tensor contraction operations, in which the chemical fingerprint tensor is expanded along the frequency dimension, the microtopology tensor is expanded along the spatial dimension, and the metabolomics tensor is expanded along the time dimension.

[0045] The tensor dimension definition and fusion operation rules in the high-order tensor fusion space include third-order tensor construction and tensor fusion operations, where:

[0046] The third-order tensor construction includes:

[0047] Chemical fingerprint tensor , expanded along the frequency dimension F, the number of channels C is the number of frequency bands after wavelet packet transform;

[0048] Microscopic image tensor , unfolded along the spatial dimension S, with the number of superpixels being P;

[0049] Metabolomics tensor , unfolded along the time dimension T, with the number of metabolite species being M.

[0050] Tensor fusion operation:

[0051] Through tensor contraction (TensorContraction), T c , T m , T b Perform element-wise multiplication along the shared dimension (such as the channel dimension) to generate a joint representation matrix , and the formula is:

[0052] .

[0053] It should be further noted that for the implementation of adversarial generative training: the generator of the generative adversarial network receives a random noise vector and the chemical component concentration gradient of the real sample, and outputs a pseudo-microscopic structure topology map; the discriminator simultaneously receives the fusion features of the real sample and the generated sample, and forces the generator to learn the non-linear mapping relationship between chemistry and morphology through the adversarial loss function.

[0054] S5: Based on the local density distribution of the samples in the high-order tensor space, an adaptive kernel spectral clustering algorithm is used to generate traditional Chinese medicine category labels, and the fusion features are reversely decoded through a variational autoencoder, and abnormal samples with a reconstruction error exceeding a preset threshold are screened to trigger weight reallocation. It should be further noted that the traditional Chinese medicine drug analysis and identification method based on clustering analysis includes dual-channel verification and model iteration: applying adversarial perturbations to the input data to generate controllable noise, monitoring the change in the singular value of the Jacobian matrix of the clustering results to adjust the robustness threshold; through a two-way feedback optimization loop, the verification results are backpropagated to the dynamic weight allocation network and the high-order representation engine to realize iterative update of the model parameters.

[0055] It should be further noted that in the specific implementation process, the parameter optimization of the adaptive kernel spectral clustering includes: according to the local density distribution of the samples in the high-order tensor space, calculating the density gradient direction within the neighborhood of each sample, and dynamically adjusting the Gaussian kernel bandwidth along the gradient direction; the eigenvalue decay rate of the Laplacian matrix is determined by calculating the cumulative energy ratio of the first k eigenvalues, and when the energy ratio exceeds the set threshold, the number of clusters is stopped from increasing.

[0056] In dynamically adjusting the Gaussian kernel bandwidth according to the density gradient, the calculation method of the density gradient is as follows:

[0057] Local density calculation:

[0058] For sample xi , whose local density ρ i is defined as:

[0059] ;

[0060] where σ is the current Gaussian kernel bandwidth.

[0061] Density gradient direction:

[0062] Calculate the density gradient ▽ρ of sample x i : i :

[0063] ;

[0064] Bandwidth parameter update rule:

[0065] Adjust the bandwidth σ according to the gradient direction:

[0066] ;

[0067] where η is the learning rate and ||▽ρ|| is the normalized value of the gradient norm.

[0068] It should be further noted that in the specific implementation process, the specific implementation process of the dual-channel verification mechanism includes the following steps:

[0069] S21: Implementation of the feature reconstruction verification channel: The encoder of the variational autoencoder receives the fused features and generates latent vectors, and the decoder reconstructs the latent vectors into the original multi-modal data; calculate the difference in mutual information entropy between the reconstructed data and the original data. If the difference exceeds the preset threshold, it is determined as an abnormal sample and triggers the retraining of the dynamic weight allocation module.

[0070] S22: Implementation of the adversarial stability verification channel: The adversarial perturbation generator adds a mixed perturbation of Gaussian noise and salt-and-pepper noise to the input data to generate adversarial samples; input the adversarial samples into the clustering model and calculate the maximum singular value change rate of the Jacobian matrix of the clustering results. If the change rate exceeds the dynamic adjustment threshold, the noise sensitivity parameter of the model is reduced.

[0071] S23: Iterative logic of the two-way feedback optimization loop: The feature reconstruction verification result updates the decoder weights of the variational autoencoder through backpropagation, and at the same time freezes the convolutional network parameters of the dynamic weight allocation module; the adversarial stability verification result selectively updates the discriminator parameters of the generative adversarial network through the gradient masking strategy to avoid overfitting of the model to noisy data.

[0072] The implementation of the dynamic gating mechanism in step S3 includes:

[0073] Real-time monitor the change rate of the information entropy of the chemical fingerprint frequency domain features to generate the first weight correction factor;

[0074] Calculate the sample distribution density through a kernel density estimator to generate a second weight correction factor;

[0075] Input the first weight correction factor and the second weight correction factor into a multi-layer perceptron to output the modal weight coefficient.

[0076] It should be further noted that in the specific implementation process, the learning of the non-linear mapping function in step S4 includes:

[0077] Construct a deep metric learning network and generate pseudo-samples with changing chemical component concentration gradients through a generative adversarial network;

[0078] In the high-order tensor fusion space, enforce the physical relevance between the clustering boundary and the topological evolution path of the microstructure.

[0079] It should be further noted that in the specific implementation process, the implementation of the adaptive kernel spectral clustering algorithm in step S5 includes:

[0080] Dynamically adjust the bandwidth parameter of the Gaussian kernel function according to the density gradient of the samples in the high-order tensor space;

[0081] Determine the optimal number of clusters by calculating the eigenvalue decay rate of the Laplacian matrix.

[0082] Step S5 also includes:

[0083] Apply adversarial perturbations to the input data to generate controllable noise;

[0084] Monitor the change in the singular value of the Jacobian matrix of the clustering result and dynamically adjust the robustness threshold of the model to noise.

[0085] A traditional Chinese medicine drug analysis and identification system based on clustering analysis, comprising:

[0086] A multi-modal data acquisition module, integrating a spectrometer, a microscopic imaging device and an environmental sensor, for obtaining chemical fingerprint spectra, microscopic images and geographical indication information;

[0087] A dynamic weight allocation module, built-in with a two-stream deep convolutional network and a bidirectional long short-term memory network, for generating a cross-modal attention matrix and dynamic weight coefficients;

[0088] A high-order collaborative representation module, including a tensor fusion engine and a generative adversarial network, for constructing a non-linear mapping function;

[0089] A clustering verification module, deploying an adaptive kernel spectral clustering algorithm and a variational autoencoder, for generating class labels and screening out abnormal samples.

[0090] It should be further noted that in the specific implementation process, the loss function of the generative adversarial network GAN is designed as follows:

[0091] Generator loss function:

[0092] ;

[0093] where c is the concentration gradient of chemical components, Contour is the regularization term of the clustering silhouette coefficient, and λ is the balance factor.

[0094] Discriminator loss function:

[0095] .

[0096] It should be further noted that in the specific implementation process, the hardware cooperation relationship in the system includes: the entropy fluctuation monitor in the dynamic weight allocation module is connected to the output end of the dual-stream deep convolutional network to calculate the information entropy change rate of the chemical fingerprint frequency domain features, and the kernel density estimator is connected to the output end of the bidirectional long short-term memory network to generate the sample distribution density correction factor; the topological constraint unit of the high-order collaborative representation module is electrically connected to the tensor fusion engine to enforce the consistency of the clustering boundary.

[0097] Implementation of the verification mechanism: The input end of the adversarial perturbation generator of the clustering verification module is connected to the multi-modal data acquisition module to apply controllable noise, and the input end of the singular value monitoring unit is connected to the output end of the adaptive kernel spectral clustering algorithm to adjust the robustness threshold to ensure the stability of the model under noise interference.

[0098] It should be further noted that in the specific implementation process, for the reconstruction error exceeding the threshold and dynamically adjusting the robustness threshold, the threshold setting standard and feedback mechanism are clarified as follows:

[0099] Threshold setting of the feature reconstruction verification channel:

[0100] Calculation formula for the mutual information entropy difference threshold e:

[0101] ;

[0102] where is the mean value of the mutual information entropy of the normal samples in the training set, is the standard deviation.

[0103] Dynamic adjustment rule of the adversarial stability verification channel:

[0104] Let the maximum singular value change rate of the Jacobian matrix be △λ, and the update formula for the robustness threshold θ is:

[0105] ;

[0106] Among them, γ is the adjustment step size. When △λ > θ old the tolerance of the model to noise is increased.

[0107] Weight update strategy of the feedback optimization loop:

[0108] When the feature reconstruction verification is triggered, only the parameters of the multi-layer perceptron in the dynamic weight allocation module are updated, and the weights of the convolutional network are frozen;

[0109] When the adversarial stability verification is triggered, the gradient masking strategy is adopted, and only the parameters of the last fully connected layer of the discriminator of the generative adversarial network are updated.

[0110] It should be further noted that in the specific implementation process, the detailed implementation of dynamic weight allocation and cross-modal association modeling includes:

[0111] S11: Feature extraction of the two-stream deep convolutional network: In the frequency-domain feature extraction of chemical fingerprints, the frequency-domain signal after wavelet packet transform is input into the first convolutional network, and the concentration gradient changes of components in different frequency bands are captured layer by layer through multi-layer convolutional kernels; in the topological feature extraction of microscopic images, graph convolution operations are performed on the cell wall structure diagram after superpixel segmentation to aggregate the texture information of adjacent superpixels to construct a global topological graph.

[0112] S112: Association modeling of the bidirectional long short-term memory network: The time-series dynamic features of metabolomics data and the spatial interpolation matrix of geographical indication information are spliced into a spatio-temporal joint vector and input into the bidirectional long short-term memory network. The cross-period association patterns between metabolite concentrations and origin environmental factors (such as temperature, humidity) are captured through forward and backward time gating units.

[0113] S13: Implementation of the dynamic gating mechanism: The sliding window statistics of the information entropy change rate of the chemical fingerprint frequency-domain features are calculated, and its standard deviation is used as the first weight correction factor; the kernel density estimation is used for the sample distribution density to generate the second weight correction factor reflecting data sparsity; the two factors are input into the multi-layer perceptron, and the activation function uses the Sigmoid function to output the modal weight coefficient ranging from 0 to 1.

[0114] Among them, in the dynamic gating mechanism, the calculation formulas of the first weight correction factor W1 and the second weight correction factor W2 are as follows: Among them, the formula of the first weight correction factor:

[0115] ;

[0116] Among them, E t is the information entropy of the chemical fingerprint frequency-domain features within the current time window, σ is the standard deviation function, and the denominator is the maximum standard deviation of the historical entropy values, which is used for normalization.

[0117] Formula of the second weight correction factor:

[0118] ;

[0119] Among them, K is the Gaussian kernel function, h is the bandwidth parameter, and N is the number of samples, which is used to quantify the sparsity of the sample distribution.

[0120] The output of the multi-layer perceptron includes: After splicing W1 and W2, they are input into a three-layer fully connected network. The activation function is Sigmoid, and the output modal weight coefficients α, β, and γ are output, corresponding to the chemical, microscopic, and metabolic modalities, satisfying α + β + γ = 1.

[0121] It should be further noted that in the specific implementation process, the hardware cooperation and data flow control of the system module include the following content:

[0122] Sa: Hardware linkage of the multi-modal data acquisition module: The spectrometer and the microscopic imaging device achieve timing alignment of data acquisition through synchronous trigger signals. The spatial interpolation data of the environmental sensor is matched with the metabolomics data through timestamps to ensure the spatio-temporal consistency of the multi-modal data.

[0123] Sb: Real-time calculation of the dynamic weight allocation module: The output end of the two-stream deep convolutional network is connected to the entropy fluctuation monitor to calculate the change rate of the information entropy of the chemical fingerprint frequency domain features in real time; the hidden state output end of the bidirectional long short-term memory network is connected to the kernel density estimator to generate a sample distribution density correction factor.

[0124] Sc: Implementation of topological constraints in the high-order collaborative representation module: The output end of the tensor fusion engine is connected to the topological constraint unit, which enforces the clustering boundary through preset microscopic structure evolution rules (such as the negative correlation between cell wall thickness and component concentration). If the clustering result conflicts with the rules, it triggers the reallocation of the feature fusion ratio.

[0125] It should be further noted that in the specific implementation process, the key control logic of the traditional Chinese medicine drug analysis and identification system based on clustering analysis includes:

[0126] Priority control of dynamic weight allocation: When the sample distribution density is lower than the preset critical value, the weight coefficient of the microscopic image features is automatically increased to give priority to ensuring the morphological discrimination ability of easily confused medicinal materials; when it is detected that there are violent fluctuations in the metabolomics data, it triggers the correlation strengthening calculation of geographical indication information to suppress noise interference.

[0127] Collaborative mechanism of adversarial training and clustering optimization: The silhouette coefficient of the clustering result is introduced as a regularization term in the training loss function of the generative adversarial network, forcing the pseudo-samples generated by the generator to conform to the distribution law of the current clustering structure; at the same time, the clustering module dynamically expands the clustering boundary according to the density distribution of the generated samples to avoid the model falling into local optima.

[0128] Abnormal sample screening and model self-repair: After the feature reconstruction verification channel detects an abnormal sample, the local model retraining process is automatically started: Only normal samples are used to update the perceptron parameters of the dynamic weight allocation module, and the adversarial network parameters of the high-order representation module are frozen to ensure the stability of the core representation ability of the model during the iteration process.

[0129] Through the dynamic weight allocation mechanism and the high-order nonlinear collaborative representation technology, the core problem of multi-modal data fusion failure in traditional Chinese medicine identification is effectively solved. The dynamic weight allocation network adaptively adjusts the contribution degree of each modality based on the real-time entropy value fluctuation and sample density, significantly improving the identification accuracy in complex scenarios such as the differences between closely related species and processed products; The high-order tensor fusion space combines adversarial generation training, breaking through the representation limitation of the linear kernel function on the non-linear correlation of "composition-morphology-efficacy", making the clustering boundary more in line with the overall characteristics of traditional Chinese medicine. The dual-channel verification mechanism ensures the robustness of the model under noise interference and data distribution deviation through double screening of feature reconstruction and adversarial stability, reducing the misjudgment rate compared with traditional methods, and providing reliable technical support for the standardization of traditional Chinese medicine quality.

[0130] Through the deep coordination of hardware modular design and algorithms, the leap from experience-dependent to intelligent analysis in traditional Chinese medicine identification is realized. The built-in dynamic gating unit and topological constraint module in the system transform traditional identification experience into quantifiable physical rules, not only retaining the holistic view of traditional Chinese medicine but also endowing the model with interpretability; The adaptive kernel parameter optimization and two-way feedback mechanism greatly reduce the cost of manual parameter adjustment, making the method highly universal in scenarios such as grass-roots drug quality inspection and rapid screening in cross-border trade. In addition, the multi-modal data fusion framework provides an extensible technical platform for the traceability of traditional Chinese medicine resources and the study of the material basis of pharmacodynamic effects, promoting the deep integration of digitalization and intelligentization in the traditional Chinese medicine industry.

[0131] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0132] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for analyzing and identifying traditional Chinese medicine based on cluster analysis, characterized in that, The following steps are involved: S1: Collect chemical fingerprints, microscopic images, metabolomics data, and geographical indication information of traditional Chinese medicine samples, perform wavelet packet transform on the chemical fingerprints to extract frequency domain features and eliminate baseline drift, perform superpixel segmentation on the microscopic images to construct cell wall topology maps, use a sliding window method to extract time series dynamic features from the metabolomics data, and perform spatial interpolation on the geographical indication information to generate an environmental factor distribution matrix; S2: A two-stream deep convolutional network is used to extract the frequency domain features of chemical fingerprints and the topological features of microscopic images. A bidirectional long short-term memory network is used to analyze the implicit correlation between metabolomics data and geographical indication information, generating a cross-modal attention matrix. S3: Based on the real-time calculated entropy fluctuations of each modality’s feature vectors and the sample distribution density, a dynamic gating mechanism is used to generate modality weight coefficients and dynamically adjust the fusion ratio of chemical fingerprint features, microscopic image features, and metabolomics features. S4: Map the weighted multimodal features to a preset high-order tensor fusion space, construct pseudo samples through adversarial generative training, and learn the nonlinear mapping function between chemical component concentration and microstructural topology; S5: Based on the local density distribution of samples in the high-order tensor space, an adaptive kernel spectral clustering algorithm is used to generate TCM category labels. The fusion features are then inversely decoded through a variational autoencoder to screen out abnormal samples whose reconstruction error exceeds a preset threshold to trigger weight redistribution.

2. The method for analyzing and identifying traditional Chinese medicine based on cluster analysis according to claim 1, wherein: The implementation of the dynamic gating mechanism in step S3 includes: Real-time monitoring of the information entropy change rate of the chemical fingerprint frequency domain characteristics to generate a first weight correction factor; Calculate the sample distribution density through the kernel density estimator to generate the second weight correction factor; The first weight correction factor and the second weight correction factor are input into a multilayer perceptron, and a modal weight coefficient is output.

3. The method for analyzing and identifying traditional Chinese medicine based on cluster analysis according to claim 2, wherein: The learning of the nonlinear mapping function in step S4 includes: Build a deep metric learning network and generate pseudo samples with chemical component concentration gradient changes through generative adversarial networks; In the high-order tensor fusion space, the physical correlation between cluster boundaries and microstructure topological evolution paths is enforced.

4. The method for analyzing and identifying traditional Chinese medicine based on cluster analysis according to claim 3, wherein: The implementation of the adaptive kernel spectral clustering algorithm in step S5 includes: Dynamically adjust the bandwidth parameter of the Gaussian kernel function according to the density gradient of the sample in the high-order tensor space; The optimal number of clusters is determined by calculating the eigenvalue decay rate of the Laplace matrix.

5. A method for analyzing and identifying traditional Chinese medicine drugs based on cluster analysis according to claim 4, characterized in that: The step S5 further includes: Apply adversarial perturbations to input data to generate controllable noise; Monitor the changes in the singular values of the Jacobian matrix of the clustering results and dynamically adjust the robustness threshold of the model to noise.

6. A traditional Chinese medicine drug analysis and identification system based on cluster analysis, characterized in that include: Multimodal data acquisition module, integrating spectrometer, microscopic imaging device and environmental sensor, used to obtain chemical fingerprint, microscopic image and geographical indication information; Dynamic weight allocation module, with built-in two-stream deep convolutional network and bidirectional long short-term memory network, is used to generate cross-modal attention matrix and dynamic weight coefficients; A high-level collaborative representation module, including a tensor fusion engine and a generative adversarial network, for constructing nonlinear mapping functions; The clustering verification module deploys an adaptive kernel spectral clustering algorithm and a variational autoencoder to generate category labels and screen abnormal samples.

7. The Chinese medicine drug analysis and identification system based on cluster analysis according to claim 6, characterized in that: The dynamic weight allocation module further includes: An entropy value fluctuation monitor, connected to the output end of the two-stream deep convolutional network, for calculating the information entropy change rate of the chemical fingerprint frequency domain features in real time; A kernel density estimator, connected to the output end of the bidirectional long short-term memory network, for generating a sample distribution density correction factor.

8. The Chinese medicine drug analysis and identification system based on clustering analysis according to claim 7, characterized in that: The high-order collaborative representation module further includes: A topological constraint unit, electrically connected to the tensor fusion engine, for forcibly constraining the consistency between the clustering boundary and the microscopic structure topological evolution path.

9. The Chinese medicine drug analysis and identification system based on cluster analysis according to claim 8, wherein: The clustering verification module further includes: An adversarial perturbation generator, with its input end connected to the multi-modal data acquisition module and its output end connected to the adaptive kernel spectral clustering algorithm, for applying controllable noise; A singular value monitoring unit, with its input end connected to the output end of the adaptive kernel spectral clustering algorithm, for adjusting the robustness threshold.

10. A computer-readable storage medium storing a computer program, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method for identifying traditional Chinese medicinal materials by combining infra-red spectra with cluster analysis

    CN101532954A

  • Stem cell differentiation degree evaluation method and system based on image analysis

    CN118279912A

  • Health care data processing method and system based on intelligent perception

    CN120032846A

  • Traditional Chinese medicine identification and analysis system based on clustering analysis

    CN120067776A

  • Prediction system, method and apparatus for abnormal brain connectivity, and readable storage medium

    WO2023077603A1

Cited By

  • AI digital fingerprint construction method and system for traditional Chinese medicinal material synthetic biological characteristics

    CN122048393A

  • Food foreign matter detection method and system based on information entropy dynamic routing and graph convolution

    CN122176696A