Chinese herbal medicine identification method and system based on multi-modal large model, device, medium

By combining image and text data preprocessing with a multimodal large model approach, and integrating them into consistent multimodal data for the identification of Chinese herbal medicines, the problem of insufficient identification in small sample cases of traditional methods is solved, and higher accuracy and efficiency in counterfeiting detection are achieved.

CN120071070BActive Publication Date: 2026-04-07BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for identifying counterfeit Chinese herbal medicines rely on manual testing and traditional analytical techniques, which are particularly inadequate in small sample situations, affecting the accuracy of identification.

Method used

A counterfeit detection method based on a multimodal large model is adopted. Image data is processed by spatial matching and pixel normalization, and text data is processed by standardization. Multi-class consistent image and text data are integrated into consistent multimodal data, which is then input into the multimodal large model for quality identification.

Benefits of technology

It improves the ability to identify counterfeits in small sample situations, enhances the accuracy and comprehensiveness of counterfeit detection, reduces labor costs and operational difficulty, and has higher counterfeit detection precision and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071070B_ABST
    Figure CN120071070B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, system, device, and medium for identifying counterfeit Chinese herbal medicines based on a multimodal large model, belonging to the field of data processing technology. The method includes: responding to the target multimodal data being image data, performing registration processing on the image data based on spatial matching to obtain multi-class spatially registered image data; and performing pixel normalization processing on the multi-class spatially registered image data to obtain multi-class consistent image data; responding to the target multimodal data being text data, performing standardization processing on the text data to obtain multi-class consistent text data; using the multi-class consistent image data and multi-class consistent text data as target consistent multimodal data; and inputting the target consistent multimodal data into a target multimodal large model to obtain the quality identification result of the target Chinese herbal medicine. This disclosure can improve the identification capability under small sample conditions, thereby improving the accuracy of counterfeit identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of data processing technology, and more specifically, relates to a method, system, device, and medium for identifying counterfeit Chinese herbal medicines based on a multimodal large model. Background Technology

[0002] Currently, the methods for identifying counterfeit Chinese herbal medicines mainly rely on manual testing and traditional analytical techniques. Existing methods primarily include: empirical identification techniques, which rely on the appearance (shape, color, odor, taste) of the medicinal materials, using sensory methods such as sight, touch, smell, and taste (with water and fire tests added when necessary) to achieve identification; microscopic identification, which uses instruments such as stereomicroscopes to observe the subtle characteristics of the surface or cross-section of the medicinal materials to identify their quality; and molecular identification techniques, which analyze the DNA or RNA information of the medicinal materials, using molecular biology methods to identify the type, origin, and quality of the materials. However, traditional methods for identifying counterfeit Chinese herbal medicines rely on large amounts of labeled data and have insufficient identification capability in small sample situations, affecting the accuracy of identification. Summary of the Invention

[0003] The purpose of this disclosure is to provide a method, system, device, and medium for identifying counterfeit Chinese herbal medicines based on a multimodal large model, so as to improve the identification capability under small sample conditions and thus improve the accuracy of identification.

[0004] A first aspect of this disclosure provides a method for identifying counterfeit traditional Chinese medicine based on a multimodal large model, comprising:

[0005] In response to the fact that the target multimodal data is image data, the image data is registered based on spatial matching to obtain multi-class spatially registered image data; and pixel normalization is performed on the multi-class spatially registered image data to obtain multi-class consistent image data.

[0006] Since the target multimodal data is text data, the text data is standardized to obtain multi-class consistent text data;

[0007] The multi-class consistent image data and the multi-class consistent text data are used as target consistent multimodal data; the target multimodal data is obtained by collecting target Chinese herbal medicines in different ways;

[0008] The target consistency multimodal data is input into the target multimodal large model to obtain the quality identification results of the target Chinese herbal medicine.

[0009] A second aspect of this disclosure provides a system for identifying counterfeit traditional Chinese medicine based on a multimodal large model, comprising:

[0010] The image preprocessing module is used to perform registration processing on the image data based on spatial matching in response to the target multimodal data being image data, to obtain multi-class spatially registered image data; and to perform pixel normalization processing on the multi-class spatially registered image data to obtain multi-class consistent image data.

[0011] The text preprocessing module is used to standardize the text data in response to the target multimodal data being text data, so as to obtain consistent text data of multiple types;

[0012] A consistency multimodal data determination module is used to use the multi-class consistent image data and the multi-class consistent text data as target consistency multimodal data; the target multimodal data is obtained by collecting target Chinese herbal medicines according to different methods;

[0013] The authenticity identification module is used to input the target consistency multimodal data into the target multimodal large model to obtain the quality identification result of the target Chinese herbal medicine.

[0014] A third aspect of this disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for identifying counterfeit Chinese herbal medicines based on a multimodal large model.

[0015] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for identifying counterfeit Chinese herbal medicines based on a multimodal large model.

[0016] The beneficial effects of the multimodal large model-based method, system, device, and medium for identifying counterfeit Chinese herbal medicines provided in this disclosure are as follows:

[0017] This disclosure ensures the consistency and comparability of target multimodal data by processing image data through spatial matching and pixel normalization, and by standardizing text data, providing a reliable foundation for subsequent quality identification. Simultaneously, this disclosure can flexibly handle different forms of target multimodal data, effectively processing and analyzing both images and text, improving the accuracy and comprehensiveness of counterfeit detection. Subsequently, using a large-scale multimodal model for quality identification can fully utilize the multidimensional characteristics of traditional Chinese medicine (TCM) to achieve a more comprehensive and in-depth counterfeit detection analysis. Compared to traditional single-modal counterfeit detection methods, this method has higher counterfeit detection accuracy and stronger generalization ability. Furthermore, this disclosure is efficient and convenient; through automated data preprocessing and model identification, it can significantly improve the efficiency of TCM counterfeit detection and reduce labor costs and operational difficulty. Therefore, this disclosure can improve identification capabilities in small sample situations, thereby improving the accuracy of counterfeit detection. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a method for identifying counterfeit Chinese herbal medicines based on a multimodal large model, provided in an embodiment of this disclosure.

[0020] Figure 2 A structural block diagram of a Chinese herbal medicine authentication system based on a multimodal large model provided in an embodiment of this disclosure;

[0021] Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of the present disclosure. Detailed Implementation

[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, so as to provide a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will understand that this disclosure may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this disclosure with unnecessary detail.

[0023] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0024] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for identifying counterfeit Chinese herbal medicines based on a multimodal large model, according to an embodiment of this disclosure. The method includes:

[0025] S101: In response to the target multimodal data being image data, the image data is registered based on spatial matching to obtain multi-class spatially registered image data; and pixel normalization is performed based on the multi-class spatially registered image data to obtain multi-class consistent image data.

[0026] In this embodiment, the image data may include RGB images, hyperspectral images, and microscopic images. The image data can be acquired using a high-resolution camera, a hyperspectral imager, or a microscope, and can reflect the overall, cross-sectional, and detailed texture features of the target Chinese herbal medicine from multiple dimensions.

[0027] Specifically, image data is registered based on spatial matching to obtain multiple types of spatially registered image data, including:

[0028] Select multiple pairs of corresponding feature points from RGB and hyperspectral images;

[0029] The homography affine transformation matrix for each pair of feature points is calculated using the least squares method.

[0030] Using RGB images as a reference, hyperspectral images are transformed based on homography affine transformation matrix to complete registration processing and obtain multi-class spatial registration image data.

[0031] Image registration involves finding identical or similar feature points in different images, establishing correspondences based on these points, and determining spatial transformation parameters between images. Based on the spatial matching results, geometric transformations are performed on different images to align identical objects or regions in their spatial positions. Multi-class spatial registration image data is a collection of images from different perspectives and of different types (e.g., global and local, macroscopic and microscopic) that have been registered and are spatially aligned. The homography affine transformation matrix describes the spatial transformation relationship from RGB images to hyperspectral images.

[0032] Pixel normalization is the process of transforming the pixel values ​​of images in spatially registered image data of different types to a unified range (such as [0,1] or [-1,1]), eliminating the problem of different pixel value ranges caused by differences in shooting equipment and environment. Multi-type consistent image data refers to various types of image data that have achieved consistency in pixel value range, data format, etc. after pixel normalization.

[0033] S102: In response to the target multimodal data being text data, the text data is standardized to obtain multi-class consistent text data.

[0034] In this embodiment, the text data may include three-dimensional morphological data and pharmacopoeia textual description data. The text data can be obtained by using a structured light scanner to acquire the three-dimensional morphological data of the target Chinese herbal medicine, and the surface unevenness features can be analyzed to effectively identify counterfeiting methods such as mold pressing.

[0035] Standardizing text data involves cleaning, formatting, and vocabulary standardization to ensure consistency across different sources.

[0036] Multi-class consistent text data refers to various types of text data that have been standardized and are unified in terms of format, semantic expression, and other aspects.

[0037] S103: Use multi-class consistent image data and multi-class consistent text data as target consistent multimodal data; target multimodal data is obtained by collecting target Chinese herbal medicines in different ways.

[0038] In this embodiment, the target multimodal data refers to a dataset acquired through various acquisition methods for the target Chinese herbal medicine, which may include information in different dimensions. For example, appearance images acquired using a high-resolution camera can visually represent the shape and color; hyperspectral images acquired using a hyperspectral imaging unit can also capture the surface reflectance spectrum of the medicinal material; microscopic images acquired using a microscopic imaging unit can present microscopic texture details; and descriptive text of medicinal properties extracted from pharmacopoeias and other sources can provide textual information such as chemical composition and efficacy.

[0039] The target consistency multimodal data is a multimodal dataset that integrates multi-class consistent image data and multi-class consistent text data, making them consistent in terms of format, scale, and semantics, and can be used collaboratively for subsequent analysis.

[0040] This embodiment first preprocesses the image data and text data separately, and then integrates the two types of data. The principle is that spatial registration of image data aligns different images in spatial location, facilitating comprehensive analysis; pixel normalization unifies the range of pixel values, eliminating numerical differences between images and benefiting subsequent model processing. Standardization of text data ensures standardized expression, improving text quality and consistency. Finally, integrating the two types of consistent data provides a coordinated and unified multimodal input for a large multimodal model, leveraging the complementary advantages of multimodal data to improve the accuracy of quality identification of traditional Chinese medicine.

[0041] Target-consistent multimodal data refers to preprocessed multimodal data from different sources and types that are unified in terms of format, scale, space, and location, and can be used collaboratively for subsequent processing.

[0042] S104: Input the target consistency multimodal data into the target multimodal large model to obtain the quality identification results of the target Chinese herbal medicine.

[0043] In this embodiment, the target multimodal large model is a trained multimodal large model capable of fusing data information from multiple different modalities. By learning features and patterns in the data, it possesses the ability to analyze and judge the target Chinese herbal medicine, used to identify the quality of the target Chinese herbal medicine, i.e., whether the target Chinese herbal medicine is genuine or counterfeit, thus achieving the authentication of Chinese herbal medicine. The quality of the target Chinese herbal medicine is its authenticity, i.e., the assessment of whether it is genuine. This embodiment can construct the target multimodal large model based on the Transformer architecture.

[0044] Specifically, the steps in this embodiment can be as follows:

[0045] First, the preprocessed target-consistent multimodal data is input into the target multimodal large model. This step enables the target multimodal large model to acquire multimodal information for analysis, preparing for subsequent feature extraction and quality assessment.

[0046] Then, after receiving the input data, the target multimodal large model utilizes its internal multi-layer neural network structure and various algorithms to extract features from different modalities. In this embodiment, the features from different modalities are then fused through a fusion mechanism (such as weighted fusion) to learn the intrinsic connections and complementary information between the multimodal data, thereby forming a comprehensive feature representation of the target Chinese herbal medicine.

[0047] Finally, the target multimodal big model compares the learned feature representations of the target Chinese herbal medicines with the knowledge and patterns about the quality of Chinese herbal medicines learned in the training data. By calculating similarity, probability and other methods, it judges the quality of the target Chinese herbal medicines and outputs the identification results, such as whether it is a genuine product or a counterfeit product.

[0048] As can be seen from the above, this disclosure ensures the consistency and comparability of target multimodal data by processing image data through spatial matching and pixel normalization, and by standardizing text data, providing a reliable foundation for subsequent quality identification. Simultaneously, this disclosure can flexibly handle different forms of target multimodal data, effectively processing and analyzing both images and text, improving the accuracy and comprehensiveness of counterfeit detection. Subsequently, using a large-scale target multimodal model for quality identification can fully utilize the multidimensional characteristics of traditional Chinese medicine, achieving a more comprehensive and in-depth counterfeit detection analysis. Compared to traditional single-modal counterfeit detection methods, this method has higher counterfeit detection accuracy and stronger generalization ability. Furthermore, this disclosure is efficient and convenient; through automated data preprocessing and model identification, it can significantly improve the efficiency of traditional Chinese medicine counterfeit detection, reducing labor costs and operational difficulty. Therefore, this disclosure can improve identification capabilities in small sample situations, thereby improving the accuracy of counterfeit detection.

[0049] In one embodiment of this disclosure, before performing pixel normalization processing based on multi-class spatially registered image data, the method further includes:

[0050] Perform Fourier transform on each type of spatially registered image data to obtain a spectrogram;

[0051] Calculate the total energy in the high-frequency region of the spectrum to obtain the total energy value; the high-frequency region is the region with a frequency greater than or equal to the frequency threshold.

[0052] If the total energy value is greater than or equal to the first energy threshold, then noise reduction processing is performed on the multi-class spatially registered image data.

[0053] In this embodiment, the Fourier transform is a mathematical transformation method that converts a time-domain or spatial-domain signal into a frequency-domain signal. In image processing, a two-dimensional Fourier transform can convert an image from the spatial domain to the frequency domain, obtaining a spectrum, thereby allowing analysis of the image's frequency components.

[0054] The high-frequency region refers to the area in the spectrum where the frequency is greater than or equal to a set frequency threshold, and it is related to the noise information in the image. The total energy value is a numerical value obtained by calculating and summing the energy values ​​corresponding to each frequency component in the high-frequency region of the spectrum, reflecting the energy distribution in the high-frequency region. The first energy threshold is a pre-set reference value used to determine whether the energy in the high-frequency region of the spectrum is too high. When the total energy value is greater than or equal to this threshold, the high-frequency region is considered to have abnormal energy and contains a lot of noise.

[0055] Noise reduction is the process of manipulating an image to reduce or remove noise and improve image quality.

[0056] Specifically, noise reduction processing is performed on spatially registered images of various types, including:

[0057] Since each type of spatial registration image is an RGB spatial registration image, noise reduction processing is performed on the RGB spatial registration image based on the bilateral filtering algorithm;

[0058] Since each type of spatial registration image is a hyperspectral spatial registration image, noise reduction processing is performed on the hyperspectral spatial registration image based on the spectral smoothing filtering algorithm;

[0059] Since each type of spatial registration image is a microscope spatial registration image, noise reduction processing is performed on the microscope spatial registration images based on the Wiener filtering algorithm.

[0060] In this embodiment, noise reduction processing of the RGB spatially registered image is performed based on a bilateral filtering algorithm, including:

[0061] Determine the spatial domain standard deviation of the bilateral filtering algorithm;

[0062] The neighborhood range value of each pixel in the RGB spatially registered image is determined based on the spatial domain standard deviation.

[0063] The weighted average value of the pixels within the neighborhood range of each pixel is calculated to obtain the filtered value of that pixel.

[0064] The noise reduction process is completed by mapping the filtered values ​​of multiple pixels.

[0065] The spatial domain standard deviation of the bilateral filtering algorithm can be set empirically.

[0066] Denoising processing of hyperspectral spatially registered images based on spectral smoothing filtering algorithms includes:

[0067] Determine the size of the filtering window;

[0068] For each pixel in the hyperspectral spatially registered image, the window is slid across the spectral curve according to the size of the filtering window, and the average pixel value within the filtering window is calculated to obtain the smoothed spectral curve.

[0069] The smoothed spectral curves of each pixel are combined and reconstructed to obtain a denoised hyperspectral spatially registered image.

[0070] In this embodiment, determining the filter window size includes:

[0071] Determine the initial filter window size;

[0072] In response to the total energy value being greater than or equal to the second energy threshold, the initial filter window size will be increased by the first step size.

[0073] In response to a total energy value that is less than the second energy threshold but greater than or equal to the first energy threshold, the initial filter window size will be reduced by a second step size.

[0074] The second energy threshold is greater than the first energy threshold.

[0075] The second energy threshold, the first step length, and the second step length are all set based on experience.

[0076] In this embodiment, the bilateral filtering algorithm is a non-linear filtering method that combines the spatial proximity and pixel value similarity of the image. While smoothing the image, it can better preserve the edge information of the image, avoiding the edge blurring problem caused by traditional filtering methods during the noise reduction process.

[0077] Spectral smoothing filtering algorithm is a filtering method that processes the spectral curve of each pixel in a hyperspectral image. It calculates the values ​​of adjacent bands on the spectral curve to smooth the spectral curve, reduce noise interference, and highlight spectral features.

[0078] The Wiener filtering algorithm is a linear filtering method based on statistical characteristics. It uses the minimum mean square error as the criterion and filters images contaminated by noise according to the noise statistics and autocorrelation function of the image. While suppressing noise, it preserves as much detail information as possible in the image.

[0079] As can be seen from the above, this embodiment obtains the spectrum through Fourier transform, which can intuitively analyze the frequency characteristics of image data; secondly, it calculates the total energy of high-frequency regions and sets a threshold for judgment, effectively identifying and filtering out image data with large noise interference; finally, it performs preprocessing and noise reduction on images that meet the conditions, which helps to improve the accuracy and stability of subsequent pixel normalization processing, reduce the negative impact of noise on image analysis results, and thus enhance the overall quality of image data and the reliability of subsequent applications.

[0080] In one embodiment of this disclosure, the training process of the target multimodal large model includes:

[0081] Semantic matching was performed on historically consistent multimodal data and the pharmacopoeia knowledge base to obtain multiple data pairs;

[0082] Feature extraction is performed on multiple data pairs to obtain multi-modal features;

[0083] A multimodal pre-trained model is obtained by pre-training features of multiple modalities;

[0084] The parameters of the multimodal pre-trained model are adjusted based on an incremental parameter update strategy to obtain the target multimodal large model.

[0085] In this embodiment, a pharmacopoeia knowledge base is introduced, which allows for the training and optimization of the model using the rich identification knowledge contained in the pharmacopoeia. Through this method, the model can make judgments based on authoritative standards during the identification process, improving the accuracy and reliability of counterfeit detection.

[0086] The semantic information contained in historical multimodal data is compared with the drug knowledge in the pharmacopoeia knowledge base to find semantically related data and combine them into data pairs. For example, image data describing the appearance of a drug is paired with text data describing the appearance of the drug in the pharmacopoeia.

[0087] For each data pair containing different modalities, appropriate feature extraction algorithms are used. For example, for text data, word embedding and other methods can be used to extract features; for image data, convolutional neural networks can be used to extract features, thereby obtaining feature representations for multiple modalities.

[0088] The extracted multimodal features are input into the Transformer architecture and pre-trained using large-scale multimodal data. By optimizing the model's parameters, the model can learn the correlations and feature representations between different modalities, thus acquiring preliminary multimodal data processing capabilities.

[0089] Based on the pre-trained model, the model is fine-tuned using new multimodal data. Following an incremental parameter update strategy, the parameter update amount is calculated based on the new data, and the model parameters are gradually updated to enable the model to better adapt to specific tasks and data, ultimately obtaining the target multimodal large model.

[0090] As can be seen from the above, this embodiment achieves deep semantic matching of information by integrating historically consistent multimodal data and a pharmacopoeia knowledge base, effectively improving the accuracy and reliability of the model. The feature extraction step ensures the comprehensive capture of features from multiple modalities, providing a rich information foundation for the model. Pre-training further enhances the model's generalization ability, enabling it to cope with complex and ever-changing task scenarios. The incremental parameter update strategy ensures that the model continuously optimizes itself during continuous learning.

[0091] In one embodiment of this disclosure, a multimodal pre-trained model is obtained by pre-training multimodal features, including:

[0092] Multi-modal features are divided into positive and negative samples;

[0093] Global optimization is performed on positive and negative samples based on the global contrastive loss function.

[0094] Local optimization is performed on positive and negative samples based on the local contrastive loss function;

[0095] The initial parameters of the multimodal large model are determined based on global optimization and local optimization.

[0096] The initial parameters of the multimodal large model are used to determine the multimodal pre-trained model.

[0097] In this embodiment, positive samples are feature combinations that have similar semantics or represent the same type of things in multimodal features, that is, these samples are similar in semantics or category.

[0098] Negative samples, in contrast to positive samples, are combinations of multimodal features with different semantics or representing different types of things, used to enable models to learn to distinguish between different things or semantics.

[0099] The global contrastive loss function is used to measure the differences between multimodal features in the overall space. By maximizing the similarity between positive samples and minimizing the similarity between negative samples, the model learns the global distribution pattern of features.

[0100] The formula for calculating the global contrastive loss function is:

[0101]

[0102]

[0103]

[0104] in, The value of the global contrastive loss function. For the sample The feature vector of , whose positive sample feature vector is . The negative sample feature vector is , , This is a function for calculating the cosine similarity between two feature vectors. For negative sample weight coefficients, ; and Both are used to iterate through the indices of negative samples. The number of negative samples, through analysis of... Iterate through the samples to find the maximum similarity between the negative and positive samples. ,right Traversal is used to calculate the summation term in the denominator. In order to determine negative samples Weighting coefficients ; To dynamically adjust parameters, they can be adjusted according to the training rounds. Adjustment, These are the initial adjustment parameters, set based on experience. To control the rate of decrease of the adjustment parameter, .

[0105] The local contrast loss function focuses on measuring the differences of multimodal features in local regions, helping the model capture local details of features and enabling the model to better distinguish features that are similar locally but different overall.

[0106] The formula for calculating the local contrastive loss function is:

[0107]

[0108] Among them, the feature vector is divided into Each of the local regions is represented as follows: , , , , To adjust the parameters, they can be set based on experience to adjust the model's sensitivity to local similarity differences.

[0109] Positive and negative samples are input into the model, and the difference between the model's output feature representation and the expected representation is calculated using a global contrastive loss function. Through backpropagation, the model's parameters are adjusted so that the features of positive samples are closer together in the feature space, while the features of negative samples are further apart, thus allowing the model to learn the global distribution pattern of multimodal features.

[0110] Building upon global optimization, positive and negative samples are input into the model again, and the loss value is calculated using a local contrastive loss function. This function focuses on the local parts of the features, and the model parameters are further adjusted through backpropagation, enabling the model to capture the local similarities and differences of multimodal features and improve the model's ability to distinguish feature details.

[0111] In this embodiment, through multiple iterative optimizations of the global and local contrastive loss functions, the model parameters gradually converge to a better state. These optimized parameters are then determined as the initial parameters of the multimodal large model. The obtained initial parameters are loaded into a pre-designed multimodal model architecture, thereby constructing a multimodal pre-trained model.

[0112] As can be seen from the above, this embodiment explicitly divides multi-modal features into positive and negative samples, which helps the model to more accurately understand the correlations and differences between data. The use of a global contrastive loss function enhances the model's ability to grasp the overall data distribution, achieving global optimization. Simultaneously, the introduction of a local contrastive loss function further improves the model's sensitivity to detailed features, achieving local optimization. This embodiment ensures the accuracy and robustness of the initial parameters of the multimodal large model.

[0113] In one embodiment of this disclosure, before adjusting the initial parameters of the multimodal pre-trained model based on an incremental parameter update strategy, the method further includes:

[0114] Calculate the distance between each modal feature and each genuine prototype cluster;

[0115] The probability of authenticity for each modality feature is determined based on distance.

[0116] If the probability of a genuine product is less than a first probability threshold, this type of modality feature is identified as a counterfeit product.

[0117] If the probability of a product being genuine is greater than or equal to the first probability threshold, the modality feature is determined to be genuine.

[0118] In this embodiment, the incremental parameter update strategy does not update all parameters at once during model training. Instead, it updates the model parameters gradually and in small amounts based on new data, enabling the model to continuously adapt to new data and optimize performance. Therefore, few-shot learning is introduced to classify samples, that is, to determine the authenticity of each modality feature based on few-shot learning.

[0119] Few-shot learning is a learning method that enables a model to learn effective features and patterns when the number of samples is small, thereby enabling it to accurately classify or predict new samples.

[0120] Modal features are information extracted from different modal data (such as text, images, audio, etc.) that can represent the essential characteristics of that modal data.

[0121] The genuine product prototype cluster is a set of cluster centers representing typical features of genuine products, obtained through cluster analysis of known modal features. The genuine product probability represents the likelihood that a certain modal feature belongs to a genuine product. The first probability threshold is a pre-defined probability critical value used to determine whether a modal feature is genuine or counterfeit.

[0122] Specifically, firstly, using a few-shot learning method, the similarity between modal features and genuine prototype clusters is measured by calculating the Euclidean distance. Then, the distance is converted into a genuine probability using a Gaussian function, allowing this similarity to be intuitively represented in probabilistic form. Finally, based on a pre-set first probability threshold, the modal feature is determined to be either genuine or counterfeit.

[0123] As can be seen from the above, this embodiment improves the accuracy and robustness of model training by determining the authenticity of each type of modal feature. Furthermore, this embodiment leverages the advantages of few-shot learning, enabling efficient modal feature authenticity determination with limited data resources, thus providing a foundation for subsequent incremental parameter updates of the model.

[0124] In one embodiment of this disclosure, the initial parameters of a multimodal pre-trained model are adjusted based on an incremental parameter update strategy to obtain a target multimodal large model, including:

[0125] Based on gradient magnitude, target parameters in the initial parameters whose correlation with counterfeit / genuine products is greater than the first correlation threshold are identified.

[0126] The target parameters are divided into global parameters and local parameters;

[0127] Fix the global parameters and fine-tune the local parameters until the error value of the target contrast loss function is less than the first error threshold, and obtain the target local parameters;

[0128] The target multimodal large model is determined based on global parameters and target local parameters.

[0129] In this embodiment, the gradient magnitude is a numerical value that measures the size of the gradient, reflecting the potential magnitude of parameter updates. The counterfeit / genuine product correlation is the degree of association between the model parameters and counterfeit or genuine product features in the data; the higher the correlation, the greater the influence of the parameters on the counterfeit or genuine product features. The counterfeit / genuine product correlation can be calculated based on cosine similarity.

[0130] The first correlation threshold is a pre-set critical value used to determine whether the correlation between a parameter and the counterfeit / genuine product is high enough. The target parameter is selected from the initial parameters that have a correlation with the counterfeit / genuine product greater than the first correlation threshold.

[0131] Global parameters are those that play a crucial role in the overall performance of the model and the macroscopic feature representation of multimodal data. They are kept constant during fine-tuning to avoid disrupting the general features already learned by the model. Local parameters are those related to the model's learning and representation of local detailed features. They can be adjusted during fine-tuning to optimize the model's ability to process specific local features.

[0132] The formula for calculating the target-contrast loss function is:

[0133]

[0134] in, The values ​​of the comparison loss function are compared to the target. These are the weight coefficients of the global contrastive loss function. These are the weighting coefficients of the local contrastive loss function. When the focus is on global optimization, When focusing on local optima, .

[0135] The first error threshold is a pre-set error critical value. When the error value of the target contrast loss function is less than this threshold, the model training is considered to have reached the accuracy requirement.

[0136] The target multimodal large model is a trained model obtained after adjusting and optimizing the parameters of a multimodal pre-trained model. Specifically, the fixed global parameters and the target local parameters obtained through fine-tuning are integrated to update the parameters of the multimodal pre-trained model, thereby obtaining the target multimodal large model, which can better perform multimodal data-related tasks.

[0137] As can be seen from the above, this embodiment employs an incremental parameter update strategy to adjust the initial parameters of the multimodal pre-trained model, which can accurately locate target parameters highly correlated with the authenticity of the data, thereby effectively improving model performance. This embodiment subdivides the target parameters into global and local categories and processes them separately, preserving the overall characteristics of the model while achieving fine-tuning of specific information. By fixing the global parameters and fine-tuning the local parameters, the model ensures that it maintains its original advantages while making adaptive adjustments to counterfeit / genuine features. This embodiment improves the model's recognition accuracy.

[0138] Corresponding to the multimodal large model-based method for identifying counterfeit Chinese herbal medicines in the above embodiments, Figure 2 This is a structural block diagram of a multimodal large model-based system for authenticating traditional Chinese medicine, provided as an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiment of this disclosure are shown. References Figure 2 The Chinese herbal medicine authentication system 20 based on a multimodal large model includes: an image preprocessing module 21, a text preprocessing module 22, a consistency multimodal data determination module 23, and a authenticity identification module 24.

[0139] Among them, the image preprocessing module 21 is used to perform registration processing on the image data based on spatial matching in response to the target multimodal data being image data, to obtain multi-class spatially registered image data; and to perform pixel normalization processing on the multi-class spatially registered image data to obtain multi-class consistent image data;

[0140] Text preprocessing module 22 is used to standardize the text data in response to the target multimodal data being text data, so as to obtain multi-class consistent text data;

[0141] The consistency multimodal data determination module 23 is used to use multi-class consistent image data and multi-class consistent text data as target consistent multimodal data; the target multimodal data is obtained by collecting target Chinese herbal medicines in different ways;

[0142] The authenticity identification module 24 is used to input the target consistency multimodal data into the target multimodal large model to obtain the quality identification results of the target Chinese herbal medicine.

[0143] In one embodiment of this disclosure, the Chinese herbal medicine authentication system 20 based on a multimodal large model further includes: a noise reduction processing module, used to perform Fourier transform on each type of spatially registered image data to obtain a spectrum diagram;

[0144] Calculate the total energy in the high-frequency region of the spectrum to obtain the total energy value; the high-frequency region is the region with a frequency greater than or equal to the frequency threshold.

[0145] If the total energy value is greater than or equal to the first energy threshold, then noise reduction processing is performed on the multi-class spatially registered image data.

[0146] In one embodiment of this disclosure, the noise reduction processing module is specifically used to perform noise reduction processing on the RGB spatial registration image based on a bilateral filtering algorithm in response to each type of spatial registration image being an RGB spatial registration image;

[0147] Since each type of spatial registration image is a hyperspectral spatial registration image, noise reduction processing is performed on the hyperspectral spatial registration image based on the spectral smoothing filtering algorithm;

[0148] Since each type of spatial registration image is a microscope spatial registration image, noise reduction processing is performed on the microscope spatial registration images based on the Wiener filtering algorithm.

[0149] In one embodiment of this disclosure, the Chinese herbal medicine authentication system 20 based on a multimodal large model further includes: a model training module, used to perform semantic matching on historically consistent multimodal data and a pharmacopoeia knowledge base to obtain multiple data pairs;

[0150] Feature extraction is performed on multiple data pairs to obtain multi-modal features;

[0151] A multimodal pre-trained model is obtained by pre-training features of multiple modalities;

[0152] The parameters of the multimodal pre-trained model are adjusted based on an incremental parameter update strategy to obtain the target multimodal large model.

[0153] In one embodiment of this disclosure, the model training module is specifically used to divide multi-modal features into positive samples and negative samples;

[0154] Global optimization is performed on positive and negative samples based on the global contrastive loss function.

[0155] Local optimization is performed on positive and negative samples based on the local contrastive loss function;

[0156] The initial parameters of the multimodal large model are determined based on global optimization and local optimization.

[0157] The initial parameters of the multimodal large model are used to determine the multimodal pre-trained model.

[0158] In one embodiment of this disclosure, the Chinese herbal medicine authentication system 20 based on a multimodal large model further includes: a small sample learning module, used to calculate the distance between each type of modal feature and each type of genuine product prototype cluster;

[0159] The probability of authenticity for each modality feature is determined based on distance.

[0160] If the probability of a genuine product is less than a first probability threshold, this type of modality feature is identified as a counterfeit product.

[0161] If the probability of a product being genuine is greater than or equal to the first probability threshold, the modality feature is determined to be genuine.

[0162] In one embodiment of this disclosure, the model training module is further configured to identify target parameters in the initial parameters whose correlation with counterfeit / genuine products is greater than a first correlation threshold based on the gradient magnitude.

[0163] The target parameters are divided into global parameters and local parameters;

[0164] Fix the global parameters and fine-tune the local parameters until the error value of the target contrast loss function is less than the first error threshold, and obtain the target local parameters;

[0165] The target multimodal large model is determined based on global parameters and target local parameters.

[0166] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of the present disclosure. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above system embodiments, for example... Figure 2 The functions of the image preprocessing module 21, text preprocessing module 22, consistency multimodal data determination module 23, and authenticity identification module 24 are shown.

[0167] It should be understood that, in the embodiments of this disclosure, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0168] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0169] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0170] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this disclosure can execute the implementation methods described in the first and second embodiments of the method for identifying counterfeit Chinese herbal medicines based on a multimodal large model provided in the embodiments of this disclosure, or they can execute the implementation methods of the electronic devices described in the embodiments of this disclosure, which will not be repeated here.

[0171] In another embodiment of this disclosure, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to implement these processes. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0172] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0173] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0174] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0175] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0176] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this disclosure, depending on actual needs.

[0177] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0178] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this disclosure, and these modifications or substitutions should all be covered within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for identifying counterfeit Chinese herbal medicines based on a multimodal large model, characterized in that, include: Since the target multimodal data is image data, the image data is registered based on spatial matching to obtain multi-class spatially registered image data; Pixel normalization is then performed on the multi-class spatially registered image data to obtain multi-class consistent image data; Since the target multimodal data is text data, the text data is standardized to obtain multi-class consistent text data; The multi-class consistent image data and the multi-class consistent text data are used as target consistent multimodal data; The target multimodal data is obtained by collecting data from the target Chinese herbal medicines in different ways; The target consistency multimodal data is input into the target multimodal large model to obtain the quality identification results of the target Chinese herbal medicine; The training process of the target multimodal large model includes: Semantic matching was performed on historically consistent multimodal data and the pharmacopoeia knowledge base to obtain multiple data pairs; Feature extraction is performed on the multiple data pairs to obtain multi-modal features; The multi-modal features are divided into positive samples and negative samples; Global optimization is performed on the positive and negative samples based on the global contrastive loss function. The formula for calculating the global contrastive loss function is as follows: in, The value of the global contrastive loss function. For the sample The feature vector of , whose positive sample feature vector is . The negative sample feature vector is , , This is a function for calculating the cosine similarity between two feature vectors. For negative sample weight coefficients; and Both are used to iterate through the indices of negative samples. The number of negative samples; To dynamically adjust parameters, For training rounds, These are the initial adjustment parameters. The hyperparameter used to control the rate of decrease of the adjustment parameter; Local optimization is performed on the positive and negative samples based on the local contrastive loss function. The formula for calculating the local contrast loss function is as follows: Among them, the feature vector is divided into Each of the local regions is represented as follows: , , , , To adjust the parameters; The initial parameters of the multimodal large model are determined based on the global optimization and the local optimization. The multimodal pre-trained model is determined based on the initial parameters of the aforementioned multimodal large model; Calculate the distance between each modal feature and each genuine prototype cluster; The probability of authenticity for each modality feature is determined based on the distance. In response to the probability of authenticity being less than a first probability threshold, this type of modal feature is determined to be counterfeit. In response to the probability of authenticity being greater than or equal to a first probability threshold, this type of modal feature is determined to be genuine. Based on the gradient magnitude, target parameters in the initial parameters whose correlation with counterfeit / genuine products is greater than a first correlation threshold are identified. The target parameters are divided into global parameters and local parameters. The global parameters are those that play a key role in the overall performance of the model and the macroscopic feature representation of multimodal data, while the local parameters are those that are related to the model's learning and representation of local detailed features. The global parameters are fixed, and the local parameters are fine-tuned until the error value of the target contrast loss function is less than the first error threshold, thus obtaining the target local parameters; The target multimodal large model is determined based on the global parameters and the target local parameters.

2. The method for authenticating traditional Chinese medicine based on a multimodal large model as described in claim 1, characterized in that, Before performing pixel normalization processing based on the multi-class spatially registered image data, the method further includes: Perform Fourier transform on each type of spatially registered image data to obtain a spectrogram; The total energy value is obtained by calculating the sum of the energy in the high-frequency region of the spectrum; the high-frequency region is the region with a frequency greater than or equal to a frequency threshold. If the total energy value is greater than or equal to the first energy threshold, then the multi-class spatial registration image data is subjected to noise reduction processing.

3. The method for identifying counterfeit Chinese herbal medicines based on a multimodal large model as described in claim 2, characterized in that, The noise reduction processing of the multi-type spatially registered image data includes: Since each type of spatial registration image is an RGB spatial registration image, noise reduction processing is performed on the RGB spatial registration image based on the bilateral filtering algorithm; Since each type of spatial registration image is a hyperspectral spatial registration image, noise reduction processing is performed on the hyperspectral spatial registration image based on the spectral smoothing filtering algorithm; Since each type of spatial registration image is a microscope spatial registration image, noise reduction processing is performed on the microscope spatial registration images based on the Wiener filtering algorithm.

4. A system for authenticating traditional Chinese medicine based on a multimodal large model, characterized in that, include: The image preprocessing module is used to perform registration processing on the image data based on spatial matching in response to the target multimodal data being image data, to obtain multi-class spatially registered image data; Pixel normalization is then performed on the multi-class spatially registered image data to obtain multi-class consistent image data; The text preprocessing module is used to standardize the text data in response to the target multimodal data being text data, so as to obtain consistent text data of multiple types; A consistency multimodal data determination module is used to take the multi-class consistent image data and the multi-class consistent text data as target consistent multimodal data; The target multimodal data is obtained by collecting data from the target Chinese herbal medicines in different ways; The authenticity identification module is used to input the target consistency multimodal data into the target multimodal large model to obtain the quality identification result of the target Chinese herbal medicine; The model training module is used to perform semantic matching between historically consistent multimodal data and the pharmacopoeia knowledge base to obtain multiple data pairs; Feature extraction is performed on the multiple data pairs to obtain multi-modal features; The multi-modal features are divided into positive samples and negative samples; Global optimization is performed on the positive and negative samples based on the global contrastive loss function. The formula for calculating the global contrastive loss function is as follows: in, The value of the global contrastive loss function. For the sample The feature vector of , whose positive sample feature vector is . The negative sample feature vector is , , This is a function for calculating the cosine similarity between two feature vectors. For negative sample weight coefficients; and Both are used to iterate through the indices of negative samples. The number of negative samples; To dynamically adjust parameters, For training rounds, These are the initial adjustment parameters. The hyperparameter used to control the rate of decrease of the adjustment parameter; Local optimization is performed on the positive and negative samples based on the local contrastive loss function. The formula for calculating the local contrast loss function is as follows: Among them, the feature vector is divided into Each of the local regions is represented as follows: , , , , To adjust the parameters; The initial parameters of the multimodal large model are determined based on the global optimization and the local optimization. The multimodal pre-trained model is determined based on the initial parameters of the aforementioned multimodal large model; Calculate the distance between each modal feature and each genuine prototype cluster; The probability of authenticity for each modality feature is determined based on the distance. In response to the probability of authenticity being less than a first probability threshold, this type of modal feature is determined to be counterfeit. In response to the probability of authenticity being greater than or equal to a first probability threshold, this type of modal feature is determined to be genuine. Based on the gradient magnitude, target parameters in the initial parameters whose correlation with counterfeit / genuine products is greater than a first correlation threshold are identified. The target parameters are divided into global parameters and local parameters. The global parameters are those that play a key role in the overall performance of the model and the macroscopic feature representation of multimodal data, while the local parameters are those that are related to the model's learning and representation of local detailed features. The global parameters are fixed, and the local parameters are fine-tuned until the error value of the target contrast loss function is less than the first error threshold, thus obtaining the target local parameters; The target multimodal large model is determined based on the global parameters and the target local parameters.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Soft attention pedestrian re-identification model, method, device and system based on reinforcement learning

    CN116597502A

  • Drug image recognition management method based on multi-modal learning

    CN119150078A

  • Long-tail target detection and fine-grained large model pre-training method, storage medium and electronic equipment

    CN119649167A