CLIP-guided multi-modal fusion microscopic identification method
Through the CLIP-guided multimodal fusion microscopic identification method, the self-cycle prompt target detection network and multimodal feature fusion network are used to accurately locate and classify microscopic images under the condition of no labeled data, solving the problems of low microscopic image identification efficiency and high cost of labeled data acquisition in the existing technology, and achieving efficient authenticity identification and quality detection of medicinal plants.
Patent Information
- Application Number
- CN202510069792.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing microscopic image identification methods have problems such as insufficient use of professional knowledge, differences between the domains of standard microscopic feature images and scanned images, and high cost of obtaining large amounts of labeled data, which limits the performance of deep learning models.
A CLIP-guided multimodal fusion microscopic identification method is proposed, including a self-cycle prompt object detection network and a multimodal feature fusion network, generating an initial pseudomask through the SAM model, combining the multi-level similarity feature contour deduplication method to locate key areas in the microscopic image under the condition of labeling data, and fusing it with the CLIP model with the pharmacopoeia text features to extract fine-grained local features and position-related features to achieve accurate classification of microscopic images.
Under the condition of no labeled data, the positioning ability of the key areas of the microscopic image is effectively improved, the precise classification of microscopic images is realized, and a highly effective technical solution for authenticity identification and quality detection of medicinal plants is provided.
Smart Images

Figure CN120014638A_ABST
Abstract
Description
Technical Field
[0001] The invention is designed for microscopic identification of medicinal plants Background Art
[0002] Microscopic identification of medicinal plants is a key method for authenticity identification and quality assessment of traditional Chinese medicines. It can provide a scientific basis for medicinal material identification through the analysis of microscopic features in microscopic images. However, traditional microscopic identification methods rely on manual operation and expert experience, and have the disadvantages of low efficiency and strong subjectivity, which makes it difficult to meet the actual needs of large-scale medicinal material testing. Therefore, the study of microscopic feature identification methods based on automated technology has become a research focus in recent years.
[0003] Existing microscopic image identification methods mainly rely on deep learning models to achieve automated processing through feature extraction and classification. Although this has improved the analysis efficiency to a certain extent, it still faces the following difficulties: first, professional knowledge is not fully utilized; second, there are domain differences between standard microscopic feature images and scanned images; third, labeled data is scarce and costly to obtain, which further limits the performance of deep learning models.
[0004] In view of the above problems, this paper proposes a CLIP-guided multimodal fusion microscopic identification method, which includes a microscopic feature localization module and a fine-grained classification module. In the microscopic feature localization module, a self-circular prompting target detection network is proposed, and the SAM model is used to generate the initial pseudo mask of the microscopic image. Combined with the multi-level similarity feature contour deduplication method, the key areas in the microscopic image are accurately located under the condition of unlabeled data; in the fine-grained classification module, a multimodal feature fusion network is constructed, and the fine-grained local features of the microscopic features are extracted through a multi-layer convolution module. The semantic features of the pharmacopoeia text are modeled in combination with the CLIP model to extract the location-related features of the identification points, and the dynamic feature fusion module is used to deeply fuse the features, thereby achieving accurate classification of microscopic images.
[0005] By combining self-supervised learning with multimodal comparative learning, the present invention breaks through the technical bottleneck of traditional microscopic image analysis methods in positioning and classification, effectively improves the positioning capability of key areas of microscopic images under the condition of unlabeled data, and provides an efficient technical solution for authenticity identification and quality inspection of medicinal plants. Summary of the invention
[0006] The purpose of the present invention is to solve the problem of misjudgment caused by unclear feature distinction in the identification process of Chinese patent medicine ingredients, and to propose a CLIP-guided multimodal fusion microscopic identification method.
[0007] The above invention objectives are mainly achieved through the following technical solutions:
[0008] S1. Scan the Chinese patent medicine sample with a slice scanner to obtain a panoramic image, then convert the panoramic image into a common format image, and then use a fixed-size window to cut the image into small images;
[0009] S2. Use the self-loop prompt target detection network to obtain the location information of the microscopic features of Chinese herbal medicine: First, use the SAM model to generate the initial pseudo-mask of the microscopic features, then use the contour deduplication module of the multi-level similarity features of the self-loop prompt target detection network to fine-screen the mask, and use the finely screened mask as the prompt information to generate the mask of the next cycle, and realize the iterative optimization of the pseudo-mask, so as to accurately locate the microscopic features of Chinese herbal medicine without labeled data. The steps are as follows:
[0010] (1) Generate an initial pseudo mask of the microscopic image using the SAM model to provide preliminary location information for the target area in the microscopic image;
[0011] (2) Extract contours from the binary mask in the microscopic image and select the contour with the largest area as the main contour;
[0012] (3) Based on the Fourier descriptor, the global shape similarity between the main contour and the stored contour is calculated. If the Euclidean distance is less than the preset threshold ∈ D , then the global shapes are judged to be similar, and the calculation formula is as follows:
[0013]
[0014] d D =||D new -D stored ||<∈ D (2)
[0015] In the formula is the Fourier transform coefficient, D k is the normalized descriptor.
[0016] (4) Based on the Hu moment feature, calculate the similarity of the local details of the main contour and the stored contour. If the Euclidean distance is less than the preset threshold ∈ H , then it is judged that the local details are similar, and the calculation formula is as follows:
[0017] H i =∑ p+q=i μ pq (i=1,…,7) (3)
[0018] d H =||H new -H stored ||<∈ H (4)
[0019] Where μpq represents the central moment of the image, H i represents the rotation invariant of the Hu moment.
[0020] (5) Based on the contour area ratio, the scale consistency between the main contour and the stored contour is judged. If the area ratio is greater than the preset threshold ∈ A , then the scale is considered consistent, and the calculation formula is as follows:
[0021]
[0022] If all three conditions above are met, the current main contour is marked as repeated and filtered; if any condition is not met, the current main contour is added to the valid mask set.
[0023] (6) Using the position information after fine screening in steps (3)-(5) as the input prompt for the next round, and gradually optimizing the feature area of the pseudo mask;
[0024] (7) When the Dice loss in the loop iteration does not decrease significantly for five consecutive rounds, the final optimized feature region position is output;
[0025] S3. Extract fine-grained local features and position-related features of microscopic features of Chinese herbal medicines: Taking the feature area of the microscopic image as input, feature extraction is performed through multi-layer convolution, and combined with the texture description of the Chinese characters in the pharmacopoeia, the microscopic image and the pharmacopoeia text information are fused through the CLIP model, so that the image features can be rich in the association information between the identification points:
[0026] (1) processing the feature area of the microscopic image by a feature extraction module, wherein the feature extraction module includes an MBConv multi-layer convolution module and a CLIP module;
[0027] (2) The multi-layer convolution module extracts fine-grained local features of feature regions through deep convolution;
[0028] (3) In the CLIP module, the feature vector of the text description in the pharmacopoeia is extracted and mapped to the same semantic space with the extracted image features. By sharing the semantic space, the image features are given semantic information about the positional relationship of the identification points;
[0029] (4) The fine-grained local features generated by the multi-layer convolution module and the CLIP module and the location-related features of the discriminant points are input into the dynamic feature fusion module for feature fusion;
[0030] S4. Build a dynamic feature fusion module: group the features of the multi-layer convolution module and the CLIP module through a multi-head attention mechanism, and use an adaptive gating mechanism to weight and optimize the importance of features in different channels;
[0031] (1) In the multi-head channel attention mechanism, the input fusion features are divided into multiple subspaces, and the position correlation features of each channel are extracted through the global context information. The calculation formula of the position correlation features is as follows:
[0032]
[0033] In the formula represents the global average pooling operation, is the input channel feature, c i Represents the extracted position association information, which is used to describe the position relationship between the identification points in the channel.
[0034] (2) Based on the position association information, the attention weight of the channel is generated through multi-layer linear transformation and activation function to adjust the importance of the channel features. The calculation formula of the attention weight is as follows:
[0035]
[0036] Where σ is the Sigmoid activation function, δ is the ReLU activation function, and W 1,i ,W 2,i , is the linear transformation matrix, b 1,i and b 2,i is the bias term.
[0037] (3) The input features are weighted by channel attention weights to strengthen key features and suppress redundant information. The calculation formula is as follows:
[0038]
[0039] Where ⊙ represents the channel-by-channel weighted operation, ensuring that the optimized features contain more task-related information.
[0040] (4) In the adaptive gating mechanism, the gating factor is generated by channel normalization operation, and the calculation formula of the gating factor is as follows:
[0041]
[0042] Where IN(·) represents the channel normalization function, σ is the Sigmoid activation function, and β is the gating factor, which is used to dynamically adjust the proportion of local detail features and position-related features.
[0043] (5) Dynamically weight the optimized features through the gating factor, and the calculation formula is as follows:
[0044]
[0045] Where ⊙ represents the channel-by-channel weighted operation.
[0046] S5. Calculate the loss function: CLIP loss calculates the similarity between image features and text features, and classification loss calculates the difference between the model output and the true label. The final loss is composed of CLIP loss and classification loss to optimize the model parameters.
[0047] (1) CLIP loss calculates the similarity between image features and text features to enhance the interaction between semantic information and image information. The formula is as follows:
[0048]
[0049] In the formula represents the cosine similarity calculation, and They are image features and text features respectively.
[0050] (2) Classification loss calculates the difference between the model output and the true label. The formula is as follows:
[0051]
[0052] In the formula represents the fusion feature of the i-th sample, is the cross entropy loss function.
[0053] (3) Combined with classification loss function Similarity loss function to CLIP The classification performance of the optimized model is calculated as follows:
[0054]
[0055] Where λ MB and λ clip is the gating score corresponding to the fine-grained local features and position-related features.
[0056] Effects of the Invention
[0057] The present invention provides a CLIP-guided multimodal fusion microscopic identification method. The method first generates an initial pseudo mask of a microscopic image using a SAM model, and iteratively optimizes the pseudo mask through a self-loop prompt target detection network to accurately locate the key feature area in the microscopic image; then, the visual features of the feature area are extracted through a multi-layer convolution module, and the associated features containing the position relationship of the identification points are extracted through the CLIP model combined with the text identification standards in the pharmacopoeia; finally, a dynamic feature fusion module is constructed, and the dynamic feature fusion module groups and processes the visual features and associated features through a multi-head channel attention mechanism, and dynamically adjusts the importance of the feature channel using an adaptive gating mechanism to enhance the ability to capture fine-grained features; finally, the final classification label is obtained through a classification network to complete the identification task of the microscopic components of Chinese patent medicines. Experiments show that the present invention has the following advantages: (1) the key feature area is accurately located through an adaptive cyclic prompt mechanism, improving the positioning ability under unlabeled data conditions; (2) the visual features of the feature area are aligned with the text features using the CLIP model, so that the image features contain the position relationship of the identification points; (3) the dynamic feature fusion module performs collaborative fusion of features, so that the final features include both fine-grained local information and associated information with the position relationship of the identification points. The present invention is suitable for the authenticity identification and quality assessment of medicinal plants, and provides an efficient microscopic image analysis method. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Flow chart of CLIP-guided multimodal fusion microscopic identification method;
[0059] Figure 2 Illustration of the self-loop prompt target detection network structure;
[0060] Figure 3 Illustration of the multimodal fusion classification network structure;
[0061] Figure 4 Illustration of the structure of the dynamic feature fusion module;
[0062] Figure 5 Illustration of the effect of microscopic characteristics identification of Chinese patent medicine; Specific implementation methods Specific implementation method one:
[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] like Figure 1 As shown in the figure, the CLIP-guided multimodal fusion microscopic identification method includes three modules: data processing, microscopic feature localization, and fine-grained classification. The data processing module mainly divides the scanned image into small images, the microscopic feature localization module mainly locates the characteristic area of the Chinese patent medicine through the self-loop prompt target detection network, and the fine-grained classification module classifies the characteristic area through the fine-grained classification network guided by CLIP. The specific steps are as follows:
[0066] S1. Scan the Chinese patent medicine sample with a slice scanner to obtain a panoramic image, then convert the panoramic image into a common format image, and then use a fixed-size window to cut the image into small images;
[0067] S2. Use the self-loop prompt target detection network to obtain the location information of the microscopic features of Chinese herbal medicine: First, use the SAM model to generate the initial pseudo-mask of the microscopic features, then use the contour deduplication module of the multi-level similarity features of the self-loop prompt target detection network to fine-screen the mask, and use the finely screened mask as the prompt information to generate the mask of the next cycle, so as to achieve iterative optimization of the pseudo-mask, so as to accurately locate the microscopic features of Chinese herbal medicine without labeled data;
[0068] S3. Extract fine-grained local features and position-related features of microscopic features of Chinese herbal medicines: Taking the feature area of the microscopic image as input, feature extraction is performed through multi-layer convolution, and combined with the Chinese text texture description of the pharmacopoeia, the microscopic image and pharmacopoeia text information are fused through the CLIP model, so that the image features can be rich in the association information between the identification points;
[0069] S4. Construct a dynamic feature fusion module: The dynamic feature fusion module groups the features from the feature extraction module and the CLIP model through a multi-head attention mechanism, and uses an adaptive gating mechanism to weight and optimize the importance of features of different channels;
[0070] S5. Model training and optimization: CLIP loss calculates the similarity between image features and text features, and classification loss calculates the difference between model output and true label. The sum of CLIP loss and classification loss constitutes the final loss to optimize model parameters.
[0071] S6. Result prediction: The microscopic features of the Chinese herbal medicine to be classified are input into the classification network. After feedforward propagation, the classification network calculates the probability of each category and finally outputs the category label corresponding to the image;
[0072] The embodiments of the present invention are described in detail below:
[0073] The embodiment of the present invention uses scanned images of 10 medicinal plant slices and applies the algorithm of the present invention to implement identification and analysis, which is specifically implemented as follows.
[0074] S1. Scan the Chinese patent medicine sample with a slice scanner to obtain a panoramic image, then convert the panoramic image into a common format image, and then use a fixed-size window to cut the image into small images;
[0075] like Figure 2 As shown, the self-loop prompt target detection network training includes the following steps:
[0076] S2. Use the self-loop prompt target detection network to obtain the location information of the microscopic features of Chinese herbal medicine: First, use the SAM model to generate the initial pseudo-mask of the microscopic features, then use the contour deduplication module of the multi-level similarity features of the self-loop prompt target detection network to fine-screen the mask, and use the finely screened mask as the prompt information to generate the mask of the next cycle, and realize the iterative optimization of the pseudo-mask, so as to accurately locate the microscopic features of Chinese herbal medicine without labeled data. The steps are as follows:
[0077] (1) Input the image and extract the feature map through the image encoder of the SAM model;
[0078] (2) Initial pseudo-mask generation: The mask decoder decodes the features to generate an initial pseudo-mask. The initial pseudo-mask generation formula is as follows:
[0079] M0=σ(Decoder(F prompt )),M0∈[0,1] H×W (14)
[0080] Where σ is the Sigmoid activation function, and the output M0 is the initial pseudo mask.
[0081] (3) The initial pseudo-mask is binarized to divide the pixel values into target area or background area. Then, the target area contour is extracted for subsequent screening and optimization;
[0082] (5) The contours of the target area are screened and deduplicated using Fourier descriptors, Hu moments, and area ratios, retaining only the most relevant areas;
[0083] The low-frequency component of the Fourier descriptor is used to describe the overall shape characteristics of the contour, which can effectively distinguish targets with similar shapes but different sizes or positions. The calculation formula is:
[0084]
[0085] Where c n is the complex representation of contour points, and N is the number of contour points.
[0086] In order to capture the local detail features of the contour, the Hu moment based on the image center moment is introduced. The Hu moment can effectively represent the local geometric features of the contour, and its calculation formula is:
[0087]
[0088] Where μ pq Represents the central moment of the image.
[0089] Finally, the area ratio is used as the judgment basis. By calculating the ratio of the target area contour area to the background area, it is further judged whether the target area meets the scale consistency requirements of the expected area.
[0090] Combining the above multi-level features, the contour features are compared and screened, and only the target area that meets all the above conditions is retained, thereby obtaining a more accurate microscopic image target area.
[0091] (6) The mask position information after the current round of fine screening is input into the SAM model as prompt information to generate the next round of masks, and the mask quality is evaluated by Dice loss to optimize the positioning of the target area;
[0092] First, the binary mask generated from the current round Extract the boundary prompt box B t .
[0093] Next, change prompt box B t Passed to the SAM model together with the input image I to generate a new round of mask M t+1 , the process is as follows:
[0094] M t+1 =SAM(I,Bt) (17)
[0095] In order to evaluate the effect of each round of optimization, the Dice loss function is introduced to measure the generated mask M t+1 With label M true The matching degree is:
[0096]
[0097] The numerator is the intersection of the generated mask and the true label, and the denominator is the sum of the two. The smaller the Dice loss, the closer the generated mask is to the target area. The entire optimization is performed in a loop until the Dice loss does not change significantly in 5 consecutive iterations. At this time, the model is considered to have converged and the final optimized mask M is output. * =M T .
[0098] like Figure 3 The training of the fine-grained classification network shown consists of the following steps:
[0099] S3. Extract fine-grained local features and position-related features of microscopic features of Chinese herbal medicines: Take the feature area of the microscopic image as input, extract features through multi-layer convolution, and combine it with the texture description of Chinese characters in the pharmacopoeia. Use the CLIP model to fuse the microscopic image and pharmacopoeia text information, so that the image features can be rich in the association information between the identification points. The steps are as follows:
[0100] (1) The fine-grained local features of the feature area are extracted through a multi-layer convolution module. The formula is as follows:
[0101] Y = X + ReLU6 (Conv 1x1 (DepthwiseConv(Conv 1x1 (X)))) (19)
[0102] Where X is the input feature map, Conv 1x1 is a 1x1 convolution, DepthwiseConv is a depthwise convolution, and ReLU6 is an activation function. ReLU6 limits the output range to [0,6] to avoid over-activation.
[0103] (2) Extract the visual features of the microscopic image through the image encoder of the CLIP model. Let I be the input microscopic image, then the extracted visual features F img It can be expressed as:
[0104] F img =ImageEncoder(I) (20)
[0105] (3) Input the text description of microscopic identification of medicinal plants in the pharmacopoeia into the text encoder of the CLIP model to extract the semantic features of the microscopic image. Let T be the text description of the microscopic structure in the pharmacopoeia. The extracted text embedding can be expressed as:
[0106] F text =TextEncoder(T) (21)
[0107] (4) The visual feature F img and semantic features F text Embedded into the same multimodal space, multimodal feature alignment is performed through contrastive learning mechanism. The formula is as follows:
[0108]
[0109] Where cos() is the cosine similarity, which is used to measure the similarity between image and text embedding. N is the number of samples in the training batch.
[0110] like Figure 4 The dynamic feature fusion module shown includes the following steps:
[0111] S4. Construct a dynamic feature fusion module: The dynamic feature fusion module groups the features from the feature extraction module and the CLIP model through a multi-head attention mechanism, and uses an adaptive gating mechanism to weight and optimize the importance of features of different channels. The steps are as follows:
[0112] (1) The position association information of each feature channel is extracted through the global average pooling operation. The calculation formula is as follows:
[0113]
[0114] In the formula is the input channel feature, H and W are the height and width of the feature map respectively, c i Indicates the location association information of the channel.
[0115] (2) Generate channel weight α through the attention module i , used to adjust the importance of channel features, the specific formula is:
[0116] α i =σ(W2·δ(W1·c i +b1)+b2) (24)
[0117] Where W1 and W2 are linear transformation matrices, b1 and b2 are bias terms, δ is the ReLU activation function, and σ is the Sigmoid activation function.
[0118] (3) Through the channel-by-channel attention weight α i Perform weighted operations on input features to generate features The calculation formula is as follows:
[0119]
[0120] Where ⊙ represents the element-by-element multiplication operation.
[0121] (4) The weighted feature channels are input into the adaptive gating mechanism, which dynamically adjusts the importance of the feature channels based on the local context information. It generates the gating factor β, which is calculated as follows:
[0122] β=σ(W gate ·δ(IN(x f ))) (26)
[0123] Where, IN() is the normalization operation, W gate is the weight matrix of the gating mechanism, δ is the ReLU activation function, and σ is the Sigmoid activation function.
[0124] (5) Use the gating factor to dynamically weight the optimized features to further balance the importance of detail features and position-related features. The calculation formula is as follows:
[0125]
[0126] Where ⊙ represents the channel-by-channel weighted operation.
[0127] S5. Model training and optimization: CLIP loss calculates the similarity between image features and text features, and classification loss calculates the difference between model output and true label. The sum of CLIP loss and classification loss constitutes the final loss to optimize model parameters. The steps are as follows:
[0128] (1) CLIP loss calculates the similarity between image features and text features. The formula is as follows:
[0129]
[0130] In the formula represents the cosine similarity calculation, and They are image features and text features respectively.
[0131] (2) Classification loss calculates the difference between the model output and the true label. The formula is as follows:
[0132]
[0133] In the formula represents the fusion feature of the i-th sample, is the cross entropy loss function.
[0134] (3) Combined with classification loss function Similarity loss function to CLIP The classification performance of the optimized model is calculated as follows:
[0135]
[0136] In the formula, λ MB and λ clip is the gating score corresponding to the fine-grained local features and position-related features.
[0137] S6. Result prediction: The microscopic features of the Chinese herbal medicine to be classified are input into the classification network. After feedforward propagation, the classification network calculates the probability of each category and finally outputs the category label corresponding to the image;
[0138] The final effect is as follows Figure 5 As shown in the figure, it can be seen that the feature area can be accurately located and the type of feature can be accurately identified.
[0139] The present invention may also have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of the present invention.
Claims
1. A CLIP-guided multimodal fusion microscopic identification method characterized by: The steps include: S1. Scan the Chinese patent medicine sample with a slice scanner to obtain a panoramic image, then convert the panoramic image into a common format image, and then use a fixed-size window to cut the image into small images; S2. Use the self-loop prompt target detection network to obtain the location information of the microscopic features of Chinese herbal medicine: First, use the SAM model to generate the initial pseudo-mask of the microscopic features, then use the contour deduplication module of the multi-level similarity features of the self-loop prompt target detection network to fine-screen the mask, and use the finely screened mask as the prompt information to generate the mask of the next cycle, so as to achieve iterative optimization of the pseudo-mask, so as to accurately locate the microscopic features of Chinese herbal medicine without labeled data; S3. Extract fine-grained local features and position-related features of microscopic features of Chinese herbal medicines: Taking the feature area of the microscopic image as input, feature extraction is performed through multi-layer convolution, and combined with the texture description of the Chinese characters in the pharmacopoeia, the microscopic image and the pharmacopoeia text information are fused through the CLIP model, so that the image features contain the association information between the identification points; S4. Build a dynamic feature fusion module: group the features from the feature extraction module and the CLIP model through a multi-head attention mechanism, and use an adaptive gating mechanism to weight and optimize the features of different channels; S5. Model training and optimization: CLIP loss calculates the similarity between image features and text features, and classification loss calculates the difference between model output and true label. The sum of CLIP loss and classification loss constitutes the final loss to optimize model parameters. S6. Result prediction: The microscopic features of the Chinese herbal medicine to be classified are input into the classification network. After feedforward propagation, the classification network calculates the probability of each category and finally outputs the category label corresponding to the feature.
2. A CLIP-guided multimodal fusion microscopic identification method as claimed in claim 1, characterized in that: The method of obtaining the position information of the microscopic features of Chinese herbal medicine by using the self-loop hint target detection network described in step S2 is as follows: first, the initial pseudo mask of the microscopic features is generated using the SAM model, and then the pseudo mask is iteratively optimized by the self-loop hint target detection network, so as to accurately locate the microscopic features of Chinese herbal medicine without labeled data. The method comprises the following steps: S21, using the SAM model to generate an initial pseudo mask of the microscopic image, providing preliminary position information for the target area in the microscopic image; S22, introduce a self-loop prompt target detection network, use the position information generated in the current round as the input prompt for the next round, and gradually optimize the boundaries and feature areas of the pseudo mask; S23. During the loop prompting process, the contour deduplication method of multi-level similarity features is used to screen and optimize the repeated masks to ensure that each target area is marked only once; S24. When the Dice loss has not been significantly improved for five consecutive rounds in the loop iteration, the final optimized Chinese herbal medicine microscopic feature region mask is output.
3. A CLIP-guided multimodal fusion microscopic identification method as claimed in claim 1, characterized in that: The method for deduplicating contours of multi-level similarity features described in step S2 comprises the following steps: S25, extracting contours from the binary mask of the characteristic region of the Chinese herbal medicine, and selecting the contour with the largest area as the main contour; S26, extracting multi-level similarity features of the main contour, including Fourier descriptor, Hu moment feature and contour area ratio; S27, based on the Fourier descriptor, calculate the global shape similarity between the main contour and the stored contour, if the Euclidean distance is less than the preset threshold ∈ D , then the global shapes are judged to be similar, and the calculation formula is as follows: d D =||D new -D stored ||<∈ D (2) In the formula is the Fourier transform coefficient, D k is the normalized descriptor; S28, based on the Hu moment feature, calculate the similarity of the local details of the main contour and the stored contour, if the Euclidean distance is less than the preset threshold ∈ H , then it is judged that the local details are similar, and the calculation formula is as follows: H i =∑ p+q=i m pq (i=1,…,7) (3) d H =||H new -H stored ||<∈ H (4) Where μ pq represents the central moment of the image, H = (H1, H2, ..., H7), H i represents the rotation invariant of Hu moment; S29, based on the contour area ratio, determine the scale consistency between the new contour and the stored contour, if the area ratio is greater than the preset threshold ∈ A , then the scale is considered consistent, and the calculation formula is as follows: If all three conditions above are met, the current contour is marked as duplicate and filtered; if any condition is not met, the current main contour is added to the valid mask set.
4. A CLIP-guided multimodal fusion microscopic identification method as claimed in claim 1, characterized in that: In step S3, the fine-grained local features and position-related features of the microscopic features of Chinese herbal medicine are extracted: the feature region of the microscopic image is used as input, feature extraction is performed through the feature extraction module, and the microscopic image and the pharmacopoeia text information are fused through the CLIP model in combination with the texture description of the pharmacopoeia, so that the image features can be rich in the association information between the identification points. The steps include: S31, processing the microscopic features of Chinese herbal medicine through a feature extraction module, wherein the feature extraction module includes a multi-layer convolution module and a CLIP module; S32, multi-layer convolution module extracts fine-grained local features of the image through deep convolution; S33. In the CLIP module, extract the feature vector of the text description in the pharmacopoeia, map it and the extracted image features to the same semantic space, and make the image features have global correlation information including the position relationship of the identification points by sharing the semantic space; S34. Input the features generated by the multi-layer convolution module and the CLIP module into the dynamic feature fusion module for collaborative fusion.
5. The dynamic fusion classification method for microscopic identification of medicinal plants according to claim 1, characterized in that: The construction of the dynamic feature fusion module described in step S4: the features from the feature extraction module and the CLIP model are grouped and processed through the multi-head attention mechanism, and the importance of features of different channels is weighted and optimized using the adaptive gating mechanism. The steps are as follows: S41. In the multi-head channel attention mechanism, the input fusion features are divided into multiple subspaces, and the contextual features of each channel are extracted by identifying the point position association information. The calculation formula of the contextual features is as follows: In the formula represents the global average pooling operation, is the input channel feature, c i Represents the extracted location association information; S42. Based on the extracted position association information, the attention weight of the channel is generated through multi-layer linear transformation and activation function to adjust the importance of the channel feature. The calculation formula is as follows: Where σ is the Sigmoid activation function, δ is the ReLU activation function, and W 1,i ,W 2,i , is the linear transformation matrix, b 1,i and b 2,i is the bias term; S43. Use channel attention weights to perform weighted operations on input features to strengthen key features and suppress redundant information. The calculation formula is as follows: Where ⊙ represents the channel-by-channel weighted operation; S44. In the adaptive gating mechanism, a channel normalization operation is used to generate a gating factor, which is calculated as follows: Where IN(·) represents the channel normalization function, σ is the Sigmoid activation function, and β is the gating factor; S45. Use the gating factor to dynamically weight the optimized features to further balance the importance of detail features and position-related features. The calculation formula is as follows: Where ⊙ represents the channel-by-channel weighted operation.
6. The dynamic fusion classification method for microscopic identification of medicinal plants according to claim 1, characterized in that: Step S5 calculates the loss function: CLIP loss calculates the similarity between image features and text features, and classification loss calculates the difference between model output and true label. The final loss is composed of CLIP loss and classification loss to optimize model parameters. S51, CLIP loss calculates the similarity between image features and text features to strengthen the interaction between semantic information and image information. The formula is as follows: In the formula represents the cosine similarity calculation, and They are the feature representations of images and texts respectively; S52, classification loss calculates the difference between the model output and the true label. The formula is as follows: In the formula represents the fusion feature of the i-th sample, is the cross entropy loss function; S53. Combined classification loss function Similarity loss function to CLIP The classification performance of the optimized model is calculated as follows: In the formula, λ MB and λ clip is the gating score corresponding to the fine-grained local features and position-related features.
Citation Information
Patent Citations
Method and device for detecting medicinal material components of Chinese patent medicine based on YOLOX model
CN114299492A
Pollen image classification method and system based on regional information fusion
CN115953611A
Traditional Chinese medicine microscopic recognition method and system based on multidimensional channel attention mechanism
CN116524495A
Diatom microscopic image automatic classification method based on comparative learning and deep learning
CN116681940A
Pollen image classification method based on convolutional neural network and multi-scale cavity attention fusion
CN117496260A
Cited By
Method and system for rapidly identifying pathogenic bacteria
CN120689871A
Traditional Chinese medicine identification method and system based on large model, electronic equipment and storage medium
CN121306603A