Fine-grained recognition method and device using multi-modal morphological perception, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610999547.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本申请旨在解决现有基于深度学习的中草药识别技术中主要存在的以下缺陷:样本数据获取困难与长尾分布问题以及细微形态特征难以有效捕捉
[0015]This application provides a fine-grained identification method, apparatus, electronic device, and storage medium utilizing multimodal morphology perception. The aim of this application is to accurately capture subtle morphological differences in medicinal materials using a fine-grained feature extraction module based on multimodal morphology perception, and to significantly improve the accuracy and interpretability of identifying similar morphological herbs with subtle features, achieving an accuracy improvement of up to 15.68%.
Smart Images

Figure CN122780929A_ABST
Abstract
Description
Technical Field
[0001] The disclosure relates to the field of data processing technology, and more specifically to a fine-grained identification method and apparatus, electronic device and storage medium utilizing multimodal morphological perception. Background Technology
[0002] Due to limitations imposed by species' geographical distribution, seasonal growth cycles, and the specialized nature of collection, high-quality labeled samples of Chinese herbal medicines are scarce, with counterfeit samples being particularly rare. This results in a typical long-tail distribution of data, and insufficient sample size in a few categories affects the model's generalization ability. Traditional data augmentation methods (including geometric transformations, pixel perturbations, occlusion erasure, and sample blending) can only make minor changes within the neighborhood of the original data distribution and cannot generate samples with new semantic content. Generative augmentation methods based on generative adversarial networks or diffusion models generally face problems such as unstable training, insufficient detail fidelity, and lack of domain knowledge guidance. They struggle to achieve controllable semantic innovation while ensuring morphological accuracy, potentially introducing morphological artifacts that do not match the real medicinal materials, thus impairing the model's recognition performance.
[0003] Furthermore, the key distinguishing features of traditional Chinese medicinal herbs (such as leaf margin serration density and angle, vein branching pattern and vein sequence, and trichome distribution and density) exhibit subtle differences, ranging from millimeter-level to pixel-level fine-grained variations. This places extremely high demands on feature extraction capabilities. Existing deep neural networks, influenced by "visual inertia," tend to prioritize easily distinguishable features such as leaf color, overall outline, and light distribution, while neglecting subtle morphological features crucial for classification decisions. Traditional local attention mechanisms, while capable of locating local image regions, lack domain knowledge guidance and cannot transform diagnostic features defined in authoritative literature such as the Pharmacopoeia of the People's Republic of China, such as vein density, edge serrations, and trichome distribution, into structured network guidance signals. Existing methods often employ single-scale feature extraction approaches, making it difficult to simultaneously capture macroscopic morphology and microscopic texture, resulting in insufficient recognition accuracy when distinguishing morphologically similar traditional Chinese medicinal herb species. Summary of the Invention
[0004] This application aims to address the following main shortcomings in existing deep learning-based Chinese herbal medicine identification technologies: difficulty in obtaining sample data, long-tail distribution problems, and difficulty in effectively capturing subtle morphological features.
[0005] According to exemplary embodiments of the present disclosure, a fine-grained recognition method utilizing multimodal morphology perception is provided. The method may include the following steps: a multi-scale-multi-frequency-domain parallel morphological feature extractor can extract multi-scale features of the input sample image using a multi-scale feature extraction branch and extract multi-frequency-domain features of the input sample image using a multi-frequency-domain feature extraction branch, and then generate a global morphological feature map based on the multi-scale features and multi-frequency-domain features; and a triple attention-guided fusion unit can inject the extracted multi-scale features and multi-frequency-domain features into the backbone network, and can apply channel filtering attention, spatial localization attention, and frequency-domain calibration attention to generate an output feature map.
[0006] According to embodiments of this disclosure, the multi-scale feature extraction branch can employ three lightweight dilated convolutional branches with different dilation rates to process the input sample image in parallel to obtain multi-scale features. The three different dilation rates can be set to d=1, d=2, and d=3, respectively, and the multi-scale features can include macroscopic shape features, medium-scale texture features, and microscopic detail features. The multi-frequency domain feature extraction branch can introduce a learnable dual-tree complex wavelet transform high-frequency branch to extract multi-frequency domain information. After concatenating the frequency domain responses in various directions, feature enhancement is performed through a learnable downsampling layer and a small convolutional neural network to obtain the multi-frequency domain features.
[0007] According to embodiments of this disclosure, the step of generating a global morphological feature map may include: concatenating multi-scale features and multi-frequency domain features along the channel dimension, encoding them into a global description vector through fusion convolution and global average pooling, and then generating a global morphological feature map through projection using a two-layer perceptron.
[0008] According to embodiments of this disclosure, channel-selective attention can refer to performing global average pooling on the global morphological feature map extracted by a multi-scale-multi-frequency-domain parallel morphological feature extractor, transforming the features through a fully connected layer, and generating channel attention weights by normalization using a sigmoid function. Spatial localization attention can refer to performing average pooling, max pooling, and edge pooling operations on the input global morphological feature map to obtain corresponding features, concatenating the corresponding features along the channel dimension, and then generating spatial attention weights by normalization using a 3×3 convolution and a sigmoid function. Frequency-domain calibration attention refers to calculating cross-frequency-domain attention weights based on the global morphological feature map using a lightweight scale attention network.
[0009] According to embodiments of this disclosure, the step of generating an output feature map may include multiplying the channel attention weights, spatial attention weights, and cross-frequency domain attention weights with the global morphological feature map channel by channel to obtain three attention features, and then concatenating the three attention features by channel dimension and performing 1×1 convolution to fuse and reduce the dimensionality to generate the output feature map.
[0010] According to embodiments of this disclosure, the fine-grained recognition method may further include: a dynamic hard sample comparison optimization unit can construct a dynamic boundary cosine distance loss function based on the output feature map and the global morphological feature map, calculate the hard positive sample loss and the hard negative sample loss, and dynamically update the boundary threshold with each training round.
[0011] According to embodiments of this disclosure, the dynamic hard sample contrast optimization unit can concatenate the output feature map and the global morphological feature map along the feature dimension, project them onto a unified 128-dimensional contrast space through a small multilayer perceptron, and normalize them to distribute them on a unit hypersphere. It can calculate the cosine distance matrix of all sample pairs within a batch, select the k pairs with the largest distance from all similar sample pairs, and calculate the sum of squared distances as the hard positive sample loss, where the value of k is set to 10% of the number of similar samples in the current batch. It can also select the k pairs with the smallest distance from all dissimilar sample pairs and calculate the sum of squared distances between them and the dynamic boundary as the hard negative sample loss.
[0012] According to embodiments of this disclosure, a fine-grained recognition device utilizing multimodal morphology perception is provided. The device may include: a multi-scale-multi-frequency-domain parallel morphological feature extractor, which extracts multi-scale features of the input sample image using a multi-scale feature extraction branch and extracts multi-frequency-domain features of the input sample image using a multi-frequency-domain feature extraction branch, and then generates a global morphological feature map based on the multi-scale features and multi-frequency-domain features; and a triple attention-guided fusion unit, which injects the extracted multi-scale features and multi-frequency-domain features into a backbone network, and applies channel filtering attention, spatial localization attention, and frequency domain calibration attention to generate an output feature map.
[0013] According to embodiments of this disclosure, a computer-readable storage medium storing a computer program is provided, which, when executed by a processor, implements the fine-grained identification method described above.
[0014] According to an embodiment of this disclosure, an electronic device is provided, the electronic device comprising: a processor; and a memory storing a computer program, wherein when the computer program is executed by the processor, the fine-grained recognition method described above is implemented.
[0015] This application provides a fine-grained identification method, apparatus, electronic device, and storage medium utilizing multimodal morphology perception. The aim of this application is to accurately capture subtle morphological differences in medicinal materials using a fine-grained feature extraction module based on multimodal morphology perception, and to significantly improve the accuracy and interpretability of identifying similar morphological herbs with subtle features, achieving an accuracy improvement of up to 15.68%. Attached Figure Description
[0016] These and / or other embodiments will become apparent and more readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 This is a flowchart of a semantically driven and morphology-aware data augmentation method according to exemplary embodiments of the present disclosure; Figure 2 This is a flowchart of a fine-grained recognition method utilizing multimodal morphology perception according to exemplary embodiments of the present disclosure; Figure 3 This is a schematic diagram of the structure of a fine-grained recognition module for multimodal morphology perception according to exemplary embodiments of the present disclosure; and Figure 4 This is a schematic diagram of the structure of an electronic device according to one or more exemplary embodiments. Detailed Implementation
[0017] The invention will now be described more fully below with reference to the accompanying drawings, various embodiments of which are illustrated. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. The same reference numerals throughout denote the same elements.
[0018] It will be understood that while terms such as "first," "second," "third," etc., may be used herein to distinguish similar objects, they are not necessarily used to describe a specific order or sequence. Furthermore, in describing exemplary embodiments, methods and / or processes may have been presented herein as a specific sequence of steps. However, a method or process should not be limited to the specific order of steps described herein, to the extent that the method or process does not depend on the specific order of steps described herein. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims concerning the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of this application.
[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. "At least one of A and B" means "A and / or B". As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will also be understood that when the terms "comprising" and / or variations thereof or "including" and / or variations thereof are used in this specification, it indicates the presence of the stated features, areas, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, areas, integrals, steps, operations, elements, components, and / or groups thereof.
[0020] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will also be understood that terms (such as those defined in a general dictionary) shall be interpreted as having the same meaning as they have in the relevant field and in the context of this disclosure, and shall not be interpreted in an idealized or overly formalized sense unless so clearly defined herein.
[0021] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0022] Figure 1 This is a flowchart of a semantically driven and morphologically aware data augmentation method according to an exemplary embodiment of the present disclosure.
[0023] like Figure 1 As shown, the above-mentioned data augmentation method based on semantic-driven and morphological awareness can generally include the following operational steps.
[0024] In operation step S100, a multi-source semantic prompting construction unit is constructed. A multimodal large language model is used to analyze the visual features of existing Chinese medicinal herb images, extracting macroscopic descriptions. Through semantic structuring processing, the extracted information is transformed into structured prompt word templates, forming a structural semantic guidance vector.
[0025] In step S200, a morphologically faithful diffusion generation unit is constructed. The basic diffusion model is then subjected to domain-adaptive fine-tuning on a Chinese herbal medicine image dataset (e.g., the Chinese Herbal Medicine Database (TCMP-300)) using the DreamBooth method, enabling the diffusion model to fully learn the visual distribution characteristics of Chinese herbal medicines. The structured semantic guidance vector constructed in step S100 is injected into the latent space of the basic diffusion model as a guiding condition, driving the model to generate Chinese herbal medicine sample images with high morphological fidelity.
[0026] The operation steps S100 and S200 will be explained in detail below with reference to specific examples and embodiments.
[0027] In an exemplary embodiment according to this application, in operation step S100, a large language model such as DPT-4 or LLaMa can be used to perform visual feature analysis on existing Chinese medicinal material images (e.g., Chinese herbal medicine data TCMP-300), extracting macroscopic descriptions including the main morphology, leaf margin features, vein patterns, color distribution, and background context of the Chinese medicinal materials to generate basic image description text. Furthermore, the generated basic image description text can be deeply integrated with standardized morphological descriptions (including leaf shape, leaf margin serrations, leaf vein type and density, leaf texture, and trichomes) contained in the *Pharmacopoeia of the People's Republic of China*, as well as the experience and knowledge of Chinese medicine experts. Through semantic structuring processing, a structured prompt word template is constructed to generate a structured semantic guidance vector for subsequent diffusion models. This structured prompt word template can be professionally reviewed and refined by a team of Chinese medicinal material experts to ensure the accuracy and professionalism of the morphological description.
[0028] Structured prompt word templates can employ a dual-channel structure of positive and negative prompt words. Positive prompt words can include basic quality requirements (e.g., image sharpness, compositional rationality), core morphological features (e.g., accurate descriptions of key identification features such as leaf shape, leaf margin, leaf veins, and trichomes), and pharmacopoeia-standardized descriptions. Negative prompt words can be used to explicitly exclude common distorted imaging artifacts, erroneous morphological features, and background interference. For example, negative prompt words can include characteristics of counterfeit products, imaging artifacts, characteristics of adulterants, and erroneous morphologies.
[0029] In an exemplary embodiment according to this application, in operation step S200, based on the Stable Diffusion (or simply SD) model as the basic diffusion model, DreamBooth is used to perform domain-adaptive fine-tuning on a dataset of Chinese herbal medicine images (e.g., TCMP-300), enabling the diffusion model to fully learn the visual distribution features of Chinese herbal medicines, thereby generating new images (i.e., new samples) containing specific objects. The visual distribution features of Chinese herbal medicines may include key visual features such as vein information and leaf edge information. Among them, vein information may include details such as the branching pattern, vein sequence type, and vein density of leaf veins, and leaf edge information may include features such as the density, angle, shape, and edge smoothness of leaf edge serrations. DreamBooth can fine-tune the entire SD model on images of specific objects (e.g., original slice images of medicinal materials) so that it learns to generate new images containing that specific object. The objective function of the fine-tuning is optimized based on the denoising diffusion probability model L (Equation 1).
[0030] Formula 1: In Equation 1 above, c is the corresponding text description of the original medicinal material slice image. It is a text feature vector encoded by the CLIP (Constrastive Language-Image Pre-training) text encoder. This is noise predicted by U-Net. To reduce computational cost, VAE is used to compress the image into a low-dimensional latent space (z0=E(x)) using the structured semantic guidance vector constructed in operation step S100 as guidance conditions, and then noise is gradually added to the latent space to obtain... Where x refers to the original image of the medicinal material slice. This is achieved by minimizing the difference between the actual noise and the predicted noise. With the 2-norm, the SD model learns to recover noise from noisy latent representations guided by text, thereby generating new samples.
[0031] Furthermore, the generation of new samples employs a multimodal generation model, allowing for the determination of the optimal parameter combination through extensive experimentation. The quality of the generated new samples can be ensured through experimental parameter optimization. Parameters may include: setting the redraw intensity to the range of 0.5 to 0.6 to balance preserving the original morphology with introducing reasonable variations, thereby avoiding image distortion; setting the sampling step count to 25 to 35 steps to ensure clear detail quality in the generated image (e.g., texture, cross-sectional structure, vein morphology, leaf edge features); and setting the cue word relevance coefficient to 10 to 15 to ensure that the generated image effectively follows semantic guidance and visual features learned from public datasets, while avoiding image distortion.
[0032] After image generation (i.e., new samples) is completed, a dual quality filtering mechanism is used for selection: First, in the automatic selection stage, the structural similarity index and the cosine similarity of the deep feature space are calculated, with selection thresholds set to 0.75 and 0.80, respectively. Then, a manual review stage is conducted, where a team of traditional Chinese medicine experts scores the morphological accuracy, texture realism, and completeness of professional features, retaining generated samples with a total score greater than or equal to 25 / 30. The selected generated samples (i.e., valid generated images) can be used for training subsequent discrimination models, effectively solving the problem of small sample scarcity. This effectively expands the original dataset from approximately 1930 images to 21694 images, increasing the sample size per class by about 11 times, with more than 2300 images for each species. This effectively alleviates the long-tail problem and significantly improves the model's generalization ability.
[0033] Figure 2 This is a flowchart of a fine-grained recognition method utilizing multimodal morphology perception according to exemplary embodiments of the present disclosure.
[0034] S300 of the fine-grained recognition method using multimodal morphology perception according to an exemplary embodiment of the present disclosure may include: a multi-scale-multi-frequency-domain parallel morphological feature extractor extracts multi-scale features of the input sample image using a multi-scale feature extraction branch and extracts multi-frequency-domain features of the input sample image using a multi-frequency-domain feature extraction branch, and then generates a global morphological feature map based on the multi-scale features and multi-frequency-domain features (S310); a triple attention-guided fusion unit injects the extracted multi-scale features and multi-frequency-domain features as prior knowledge into the backbone network, and applies channel filtering attention, spatial localization attention, and frequency domain calibration attention to generate an output feature map, thereby realizing feature modulation of multi-level multi-granularity morphological perception (S320); and a dynamic hard sample comparison optimization unit constructs a dynamic boundary cosine distance loss function based on the output feature map and the global morphological feature map, calculates the hard positive sample loss and the hard negative sample loss, and dynamically updates the boundary threshold with each training round (S330).
[0035] Although the exemplary embodiments described above describe a fine-grained recognition method S300 that may include steps S310 (operating with a multi-scale-multi-frequency domain parallel morphological feature extractor), S320 (operating with a triple attention-guided fusion unit), and S330 (operating with a dynamic hard sample contrast optimization unit), this disclosure is not limited thereto. For example, the fine-grained recognition method S300 according to the exemplary embodiments may include both step S310 (operating with a multi-scale-multi-frequency domain parallel morphological feature extractor) and step S320 (operating with a triple attention-guided fusion unit). Alternatively, the fine-grained recognition method S300 according to the exemplary embodiments may include only step S310 (operating with a multi-scale-multi-frequency domain parallel morphological feature extractor).
[0036] In the fine-grained recognition method S300, the fine-grained feature extraction module of multimodal morphology perception can be used to accurately capture the subtle morphological differences of medicinal materials. Its goal is to capture millimeter-level morphological differences (leaf margin serrations, leaf vein angles) as well as multidimensional features such as color, texture, structure, and edges, and to construct a highly discriminative feature space.
[0037] A multi-module morphology-aware fine-grained feature extraction module may include a multi-scale-multi-frequency domain parallel morphological feature extractor. In an exemplary embodiment, the multi-module morphology-aware fine-grained feature extraction module may include a multi-scale-multi-frequency domain parallel morphological feature extractor and a triple attention-guided fusion unit; however, this disclosure is not limited thereto. In an exemplary embodiment, the multi-module morphology-aware fine-grained feature extraction module may include a multi-scale-multi-frequency domain parallel morphological feature extractor, a triple attention-guided fusion unit, and a dynamic hard sample contrast optimization unit.
[0038] In an exemplary embodiment, the multi-scale-multi-frequency-domain parallel morphological feature extractor can utilize two parallel branches, including a multi-scale feature extraction branch and a multi-frequency-domain feature extraction branch, to achieve multi-dimensional extraction of morphological features.
[0039] In step S310, the multi-scale feature extraction branch can employ three lightweight dilated convolution branches with different dilation rates to process the input sample image in parallel to obtain multi-scale features. Specifically, the three lightweight dilated convolution branches with different dilation rates are used to capture the macroscopic shape features, medium-scale texture features, and microscopic detail features of the medicinal material image, respectively. The three different dilation rates can be set to d=1, d=2, and d=3, respectively, with a kernel size of 3×3 and 16 output channels. In other words, the multi-scale features can include macroscopic shape features, medium-scale texture features, and microscopic detail features.
[0040] The multi-frequency domain feature extraction branch can introduce a learnable dual-tree complex wavelet transform (DT-CWT) high-frequency branch to extract multi-directional (θ=1,…,6) multi-frequency domain information, thereby capturing high-frequency detail information that is difficult to perceive by traditional convolution. High-frequency detail information can include the sharpness of edge serrations, the fine structure of leaf vein branch points, and the distribution density of trichomes. After concatenating the frequency domain responses from various directions, feature enhancement is performed through a learnable downsampling layer and a small convolutional neural network to obtain the multi-frequency domain feature F_freq. Finally, the multi-scale feature F_m and the multi-frequency domain feature F_freq are concatenated along the channel dimension, and then encoded into a global descriptive vector through fusion convolution and global average pooling. This vector is then projected through a two-layer perceptron to generate a 128-dimensional global morphological feature vector (global morphological feature map) F_global.
[0041] In step S320, the triple attention-guided fusion unit can inject the multi-scale-multi-frequency domain global morphological feature vector F_global generated by the multi-scale-multi-frequency domain parallel morphological feature extractor as prior knowledge into the multi-level layers of a backbone network such as ResNet or MobileNet. Specifically, channel filtering attention, spatial localization attention, and frequency domain calibration attention are applied in each of the 2nd, 3rd, and 4th layers of the deep backbone network. Exemplary embodiments are not limited to this; for example, channel filtering attention, spatial localization attention, and frequency domain calibration attention can be applied to multiple other layers of the backbone network.
[0042] Channel filtering attention refers to the process of performing global average pooling on the global morphological feature map F_global extracted by the multi-scale-multi-frequency domain parallel morphological feature extractor, transforming the features through two fully connected layers (with a reduction ratio r set to 16), and generating channel attention weights by normalization using the Sigmoid function, thereby filtering out semantic channels that are sensitive to key identification features such as leaf edge serrations and leaf vein morphology.
[0043] Spatial localization attention refers to applying average pooling, max pooling, and edge pooling operations to the input morphological feature map F_global. Edge pooling uses the Sobel operator to extract the edges of the input morphological feature map F_global to enhance the accuracy of spatial weights. The features obtained from these three pooling operations are concatenated along the channel dimension, then normalized using a 3×3 convolution and the Sigmoid function to generate spatial attention weights. This allows for precise localization of dense morphological feature regions on the leaf, such as the serrated edges of the leaf margin, vein intersections, and trichome distribution areas—key physical regions.
[0044] Frequency domain calibrated attention refers to the process of using a lightweight scale attention network to calculate cross-frequency domain attention weights based on the global morphological feature map F_global extracted by a multi-scale-multi-frequency domain parallel morphological feature extractor, thereby dynamically enhancing or suppressing texture information in specific frequency bands.
[0045] The weights of the three attention methods are multiplied channel-by-channel by the input global morphological feature map F_global to obtain three attention features. These three attention features are then concatenated along the channel dimension and fused and dimensionality-reduced using a 1×1 convolution to generate the final output feature map (or output feature map) F_b. As a result, multi-level, multi-granularity morphological perception feature modulation is achieved.
[0046] In step S330, the dynamic hard sample comparison optimization unit constructs a dynamic boundary cosine distance loss function based on the output feature map F_b obtained in step S320 and the global morphological feature map F_global obtained in step S310, calculates the hard positive sample loss and hard negative sample loss, and dynamically updates the boundary threshold with each training round.
[0047] The dynamic hard sample contrast optimization unit can concatenate the final output feature map F_b of the backbone network of the triple attention-guided fusion unit with the global morphological feature map F_global output by the multi-scale-multi-frequency domain parallel morphological feature extractor along the feature dimension, project it onto a unified 128-dimensional contrast space through a small multilayer perceptron, and then perform... 2. Normalization makes it distributed on a unit hypersphere.
[0048] Then, the dynamic hard sample contrast optimization unit can calculate the cosine distance matrix D of all sample pairs within a batch. The distance calculation formula is D_ij = 1 - f_i·f_j, where i represents sample i, j represents sample j, f_i represents the feature of sample i, f_j represents the feature of sample j, and D_ij represents the distance between sample i and sample j. The sample loss is divided into two parts: one is the hard positive sample loss, and the other is the hard negative sample loss.
[0049] Based on the cosine distance matrix D, select the top k pairs with the largest distance from all similar sample pairs (k is set to 10% of the number of similar samples in the current batch, with a minimum of 1), calculate the sum of squared distances between them as the hard positive sample loss L_intra, and force the model to reduce the distance between the most difficult-to-distinguish similar samples.
[0050] Furthermore, based on the cosine distance matrix D, the top k pairs with the smallest distances are selected from all outlier pairs, and the sum of the squares of their distances to the dynamic boundary m(t) is calculated as the hard negative sample loss L_inter. The dynamic boundary m(t) is adjusted using a cosine scheduling strategy, increasing with each training round, with an initial value of 0.3 and a final value of 0.7, penalizing outlier pairs whose distances are below the boundary.
[0051] The overall objective function L_total is a linear combination of the standard cross-entropy classification loss and the contrastive loss. Where α is the contrastive loss weight coefficient, which can be 0.2, and L_ce refers to the standard cross-entropy classification loss function. This dynamic hard sample contrastive optimization unit optimizes the feature space geometry, effectively mitigating classification confusion caused by intra-class differences and inter-class similarities.
[0052] In addition, exemplary embodiments of this disclosure also provide a fine-grained identification device that can be used for traditional Chinese medicine. Figure 3 A schematic diagram of the structure of a fine-grained identification device according to an exemplary embodiment of the present disclosure is shown.
[0053] According to an exemplary embodiment of the present disclosure, a multimodal morphology-aware fine-grained feature extraction module 300 can be used to capture morphological differences of Chinese herbal medicines. The multimodal morphology-aware fine-grained feature extraction module 300 may include at least one of a multi-scale-multi-frequency domain parallel morphological feature extractor 310, a triple attention-guided fusion unit 320, and a dynamic hard sample comparison optimization unit 330.
[0054] In an exemplary embodiment, the multi-scale-multi-frequency-domain parallel morphological feature extractor 310 extracts multi-scale features using a multi-scale feature extraction branch and extracts multi-frequency-domain features using a multi-frequency-domain feature extraction branch, and then generates a global morphological feature vector F_global based on the multi-scale features and the multi-frequency-domain features.
[0055] In an exemplary embodiment, the triple attention-guided fusion unit 320 uses the extracted global morphological feature vector F_global as prior knowledge and injects it into multiple layers of the backbone network. Furthermore, channel filtering attention, spatial localization attention, and frequency domain calibration attention are applied in each of the 2nd, 3rd, and 4th layers of the backbone network to generate the final output feature map F_b, thereby achieving multi-level, multi-granularity morphological perception feature modulation. The exemplary embodiment is not limited to this; channel filtering attention, spatial localization attention, and frequency domain calibration attention can be applied in each of the other multiple layers of the backbone network.
[0056] In an exemplary embodiment, the dynamic hard sample comparison optimization unit 330 constructs a dynamic boundary cosine distance loss function based on the global morphological feature vector F_global generated by the multi-scale-multi-frequency domain parallel morphological feature extractor 310 and the final output feature map F_b output by the triple attention-guided fusion unit 320, calculates the hard positive sample loss and the hard negative sample loss, and dynamically updates the boundary threshold with each training round, thereby constructing a feature distribution that is compact within classes and separated between classes.
[0057] The multi-scale-multi-frequency domain parallel morphological feature extractor 310, the triple attention-guided fusion unit 320, and the dynamic hard sample comparison optimization unit 330 are basically the same as the multi-scale-multi-frequency domain parallel morphological feature extractor, the triple attention-guided fusion unit, and the dynamic hard sample comparison optimization unit described in the above fine-grained recognition method using multimodal morphological perception, so they will not be described in detail here.
[0058] Figure 4 A schematic diagram of the structure of an electronic device according to one or more exemplary embodiments is shown. Figure 3 As shown, the electronic device 400 may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions or code in the memory 430 to execute a fine-grained identification method for traditional Chinese medicine. The processor 410 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 410 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0059] In an exemplary embodiment, a computer-readable storage medium may also be provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the fine-grained identification method as described in the exemplary embodiment above. The computer-readable storage medium may be, for example, a memory including instructions. Optionally, the computer-readable storage medium may be: a read-only memory (ROM), a random access memory (RAM), a random access programmable read-only memory (PROM), an electrically erasable programmable read-only memory (EEPROM), a dynamic random access memory (DRAM), a static random access memory (SRAM), flash memory, non-volatile memory, a CD-ROM, a CD-R, a CD+R, a CD-RW, a CD+RW, a DVD-ROM, a DVD-R, a DVD+R, a DVD-RW, a DVD+RW, a DVD-RAM, a BD-ROM, a BD-R, or a BD-R... LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0060] According to exemplary embodiments of the present disclosure, a computer program product may also be provided, the computer program product including computer-executable instructions, which, when executed by at least one processor, implement the fine-grained identification method of traditional Chinese medicine according to exemplary embodiments of the present disclosure.
[0061] This application provides a fine-grained recognition method, apparatus, storage medium, and computer program product utilizing multimodal morphology perception. The aim of this application is to expand the traditional Chinese medicine dataset by constructing a semantically driven morphology-fidelity data enhancement module and a morphology-fidelity diffusion generation unit, and to accurately capture subtle morphological differences in medicinal materials using a multimodal morphology-fidelity-aware fine-grained feature extraction module, thereby constructing a highly discriminative feature space.
[0062] Although this application has been specifically shown and described with reference to some embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope defined by the claims.
Claims
1. A fine-grained recognition method utilizing multimodal morphology perception, characterized in that, The method includes the following steps: A multi-scale-multi-frequency domain parallel morphological feature extractor utilizes a multi-scale feature extraction branch to extract multi-scale features from the input sample image and a multi-frequency domain feature extraction branch to extract multi-frequency domain features from the input sample image. Based on the multi-scale and multi-frequency domain features, a global morphological feature map is generated. The triple attention-guided fusion unit injects the extracted multi-scale and multi-frequency domain features into the backbone network, and applies channel-selective attention, spatial localization attention, and frequency domain calibration attention to generate the output feature map.
2. The fine-grained identification method according to claim 1, characterized in that: The multi-scale feature extraction branch uses three lightweight dilated convolution branches with different dilation rates to process the input sample image in parallel to obtain multi-scale features. The three different dilation rates are set to d=1, d=2, and d=3, respectively, and the multi-scale features include macroscopic shape features, medium-scale texture features, and microscopic detail features. The multi-frequency domain feature extraction branch introduces a learnable dual-tree complex wavelet transform high-frequency branch to extract multi-frequency domain information. After concatenating the frequency domain responses in each direction, feature enhancement is performed through a learnable downsampling layer and a small convolutional neural network to obtain multi-frequency domain features.
3. The fine-grained identification method according to claim 2, characterized in that: The steps for generating a global morphological feature map include: concatenating multi-scale features and multi-frequency domain features along the channel dimension, encoding them into a global description vector through fusion convolution and global average pooling, and then generating a global morphological feature map through projection using a two-layer perceptron.
4. The fine-grained identification method according to claim 1, characterized in that: Channel filtering attention refers to the process of performing global average pooling on the global morphological feature map extracted by the multi-scale-multi-frequency domain parallel morphological feature extractor, transforming the features through a fully connected layer, and generating channel attention weights by normalization using the Sigmoid function. Spatial localization attention refers to performing average pooling, max pooling, and edge pooling operations on the input global morphological feature map to obtain corresponding features, concatenating the corresponding features along the channel dimension, and then normalizing them through 3×3 convolution and the Sigmoid function to generate spatial attention weights. and Frequency domain calibrated attention refers to calculating cross-frequency domain attention weights using a lightweight scale attention network based on the global morphological feature map.
5. The fine-grained identification method according to claim 1, characterized in that, The steps to generate the output feature map include multiplying the channel attention weights, spatial attention weights, and cross-frequency domain attention weights with the global morphological feature map channel by channel to obtain three attention features. Then, the three attention features are concatenated along the channel dimension and fused and reduced in dimensionality through a 1×1 convolution to generate the output feature map.
6. The fine-grained identification method according to claim 1, characterized in that, The method further includes: a dynamic hard sample comparison optimization unit constructs a dynamic boundary cosine distance loss function based on the output feature map and the global morphological feature map, calculates the hard positive sample loss and the hard negative sample loss, and dynamically updates the boundary threshold with each training round.
7. The fine-grained identification method according to claim 6, characterized in that, The dynamic hard sample contrast optimization unit concatenates the output feature map and the global morphological feature map along the feature dimension, projects them onto a unified 128-dimensional contrast space through a small multilayer perceptron, and normalizes them to distribute them on a unit hypersphere. Calculate the cosine distance matrix for all sample pairs within a batch. Select the top k pairs with the largest distance from all similar sample pairs, and calculate the sum of squared distances between them as the hard positive sample loss. Here, the value of k is set to 10% of the number of similar samples in the current batch. Select the top k pairs with the smallest distance from all outlier pairs, and calculate the sum of squared distances between them and the dynamic boundary as the hard negative sample loss.
8. A fine-grained recognition device utilizing multimodal morphology perception, characterized in that, The device includes: A multi-scale-multi-frequency domain parallel morphological feature extractor extracts multi-scale features from the input sample image using a multi-scale feature extraction branch and multi-frequency domain features from the input sample image using a multi-frequency domain feature extraction branch. Based on the multi-scale and multi-frequency domain features, a global morphological feature map is generated. The triple attention-guided fusion unit injects the extracted multi-scale and multi-frequency domain features into the backbone network, and applies channel-selective attention, spatial localization attention, and frequency domain calibration attention to generate the output feature map.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the fine-grained recognition method as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: processor; Memory, which stores computer programs When the computer program is executed by the processor, it implements the fine-grained recognition method as described in any one of claims 1 to 7.