Traditional Chinese medicine decoction piece quality detection and quantitative measurement method and system based on multi-modal large model and multi-task learning

By employing a multimodal large model and multi-task learning approach, the problems of intelligent and standardized quality inspection of Chinese herbal medicine pieces were solved, enabling adaptive quality inspection and precise quantitative measurement of Chinese herbal medicine pieces, and outputting interpretable quality inspection reports.

CN121901984APending Publication Date: 2026-04-21ZHEJIANG CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG CANCER HOSPITAL
Filing Date
2026-01-14
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current methods for quality testing of Chinese herbal medicine pieces suffer from low levels of intelligence and lack of standardization. Traditional methods rely on human experience, while machine vision technology has limitations, making it difficult to convert text from Chinese herbal medicine books into executable computer vision instructions, and quantitative measurement is also challenging.

Method used

We employ a method based on multimodal large model and multitask learning. We extract standard rules from books related to traditional Chinese medicine decoction pieces through multimodal large model, acquire visual features using a multitask learning backbone network with hard parameter sharing mechanism, and combine it with SAM 3 model for pixel-level semantic segmentation and quantitative measurement to generate quality inspection reports.

Benefits of technology

It achieves intelligent and standardized quality inspection of Chinese herbal medicine pieces, can adaptively update identification standards without retraining the model, possesses human expert-level cognitive ability, accurately identifies the characteristics and quantitatively measures Chinese herbal medicine pieces, and outputs interpretable quality inspection reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901984A_ABST
    Figure CN121901984A_ABST
Patent Text Reader

Abstract

The invention provides a traditional Chinese medicine decoction piece quality detection and quantitative measurement method and system based on a multi-modal large model and multi-task learning, and relates to the technical field of traditional Chinese medicine decoction piece detection. The method comprises the following steps: firstly, extracting semantic quality standards and feature descriptions from related traditional Chinese medicine identification books, standards and maps by utilizing a knowledge retrieval enhancement generation technology; secondly, a multi-task learning backbone network based on a hard parameter sharing mechanism is constructed, and efficient extraction and cross-task collaboration of image features are achieved; thirdly, dynamic prompts are generated according to the extracted semantic features in combination with a low-rank adaptive fine-tuning SAM 3 model, pixel-level segmentation and quantitative measurement are conducted on the fine biological features of the traditional Chinese medicine decoction pieces, and finally logical reasoning is conducted through visual question and answer agency of a multi-mode large model in combination with quantitative data and a standard threshold value to judge the grade of the traditional Chinese medicine decoction pieces. According to the method, the crossing from qualitative experience identification to quantitative intelligent analysis is realized, and the accuracy and interpretability of industrial quality inspection of the traditional Chinese medicine decoction pieces are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, computer vision, and industrial automation testing technology, specifically to a method and system for quality testing and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning. Background Technology

[0002] As the material basis for TCM clinical prescriptions, the quality of prepared Chinese medicinal herbs directly affects medical safety and efficacy. Traditionally, the quality evaluation of prepared Chinese medicinal herbs relies primarily on the "morphological identification" of experienced pharmacists. This involves comprehensively judging the authenticity and quality of medicinal materials through visual inspection (shape, color, cross-sectional characteristics), tactile examination (texture, weight), olfactory examination (odor), and gustatory examination (taste). For example, high-quality Astragalus membranaceus should have the characteristics of "chrysanthemum heart" (radiating texture on the cross-section) and "golden well and jade railing" (yellowish-white wood and pale yellow bark); Salvia miltiorrhiza is considered best when it is "purplish-red in color and has cinnabar dots on the cross-section."

[0003] The current quality control of Chinese herbal medicine decoction pieces faces three major challenges: First, there is a break in the inheritance of identification experience, with a shortage of experts and strong subjectivity in manual testing, making it difficult to unify industrial production standards; second, traditional machine vision technology has limitations, with related models relying on labeled data and suffering from defects such as lack of semantic understanding, difficulty in quantitative measurement, and long-tailed data distribution, which cannot match the abstract descriptions and precise testing requirements of pharmacopoeias and other related Chinese herbal medicine books; third, although there are emerging technologies such as multimodal large models, SAM technology, and multi-task learning, an intelligent quality inspection technology that can convert text in pharmacopoeias and other related Chinese herbal medicine books into visual instructions and combine high-precision segmentation to achieve quantitative testing has not yet been formed.

[0004] In summary, there is currently a lack of intelligent quality inspection technology for Chinese herbal medicine slices that can convert relevant Chinese medicine book text standards into executable computer vision instructions and combine them with high-precision segmentation models for quantitative measurement. Summary of the Invention

[0005] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method and system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, which solves the problems of low intelligence and low standardization in existing quality inspection of traditional Chinese medicine decoction pieces.

[0006] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: Firstly, this application proposes a method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, the method comprising: Based on a multimodal large model, standard rules for quality testing and quantitative measurement of Chinese herbal medicines are extracted from relevant books on Chinese herbal medicines. These standard rules include key identification points of Chinese herbal medicines. The pre-acquired images of Chinese herbal medicine slices are input into a multi-task learning backbone network based on a hard parameter sharing mechanism to obtain the visual features of the Chinese herbal medicine slices. Based on the key identification points, a large visual model is used to perform pixel-level semantic segmentation of Chinese herbal medicine pieces to obtain quantitative measurement results of Chinese herbal medicine pieces. The visual features and quantitative measurement results are input into a pre-built reasoning agent, and combined with the standard rules, the reasoning results of traditional Chinese medicine decoction pieces are generated.

[0007] In one embodiment, the multi-task learning backbone network based on the hard parameter sharing mechanism includes: a shared feature encoder and multiple specific task heads; the multiple specific task heads include at least task heads for performing the herbal medicine category classification task, defect region localization task, and global semantic feature mapping task individually or in cooperation.

[0008] In a preferred embodiment, the shared encoder includes a Swing Transformer.

[0009] In a preferred embodiment, the loss weights of tasks corresponding to multiple specific task heads are dynamically adjusted by introducing the GradNorm algorithm. in, Total loss function; 、 、 These represent the weights for the classification task, the contrastive learning task, and the segmentation task, respectively. 、 , These represent the cross-entropy loss function, the supervised contrastive loss function, and the Dice loss function, respectively.

[0010] In one embodiment, supervised contrastive learning is performed in the shared feature space of the multi-task learning backbone network of the hard parameter sharing mechanism. By maximizing the similarity between the features of high-quality medicinal slices of the same type and the standard semantic embedding, and minimizing the similarity between them and the features of counterfeit or inferior products, a discriminative visual feature vector is generated.

[0011] Preferably, supervised contrastive learning employs a hierarchical loss function: in, For anchor point sample features, For the positive sample set, To compare the sample sets, Temperature coefficient; Anchor point sample index; Anchor point sample set; Positive sample index; Represents the feature vector of a positive sample; This represents the feature vector of the comparison sample.

[0012] In one embodiment, the large visual model includes SAM 3.

[0013] In a preferred embodiment, a low-rank adaptation strategy is used to fine-tune the SAM 3 model: the parameters of the SAM 3 image encoder are kept frozen, and a low-rank adaptation matrix is ​​injected into the Transformer layer of the mask decoder.

[0014] In one embodiment, the quantitative measurement includes doping rate calculation, characteristic density calculation, and size measurement.

[0015] Secondly, this application also proposes a system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, the system comprising: The standard rule extraction module is configured to extract standard rules for quality testing and quantitative measurement of Chinese herbal medicines from books related to Chinese herbal medicines based on a multimodal large model. The standard rules include key identification points of Chinese herbal medicines. The visual feature acquisition module is configured to input pre-acquired images of Chinese herbal medicine slices into a multi-task learning backbone network based on a hard parameter sharing mechanism to acquire the visual features of the Chinese herbal medicine slices. The quantitative measurement result acquisition module is configured to guide the visual large model to perform pixel-level semantic segmentation of Chinese herbal medicine slices based on the key identification points, and acquire the quantitative measurement results of Chinese herbal medicine slices. The reasoning result generation module is configured to input the visual features and quantitative measurement results into a pre-built visual question-answering reasoning agent, and combine them with the standard rules to generate reasoning results for traditional Chinese medicine decoction pieces.

[0016] (III) Beneficial Effects This invention provides a method and system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning. Compared with the prior art, it has the following advantages: 1. This application constructs a quality inspection system for traditional Chinese medicine (TCM) decoction pieces with human expert-level cognitive capabilities by deeply integrating multi-task learning, large-scale segmentation models, and multimodal knowledge reasoning technologies. It utilizes MLLM to understand complex TCM decoction piece identification text standards, SAM3 to achieve precise quantification of irregular biological characteristics, and MTL to ensure efficient and robust feature extraction. This application effectively solves the problems of difficulty in quantifying the "doping rate" and algorithmically describing the "properties" of TCM decoction pieces, providing strong technical support for the intelligent and standardized quality inspection of TCM decoction pieces.

[0017] 2. This application is based on RAG's standard digitization and feature deconstruction, and does not directly rely on manually hard-coded rules. Instead, it uses a multimodal large model to extract content from books related to traditional Chinese medicine decoction pieces, and uses OCR technology to transform unstructured text into structured data, extracting key identification points (such as "radial texture" and "concentric rings") and quantitative indicators (such as "diameter > 0.5cm" and "impurities < 3%)" and other standard rules for subsequent reasoning. This approach enables adaptive testing of the quality inspection and quantitative measurement standards for traditional Chinese medicine decoction pieces. When the identification standards for traditional Chinese medicine decoction pieces are updated, only the knowledge base document needs to be updated, without retraining the visual model, which has extremely high flexibility.

[0018] 3. This application utilizes the powerful pixel-level segmentation capabilities of SAM3, combined with low-rank adaptation (LoRA) technology for fine-tuning with a small number of samples, to adapt it to the unique biological textures of traditional Chinese medicine (TCM) decoction pieces (such as distinguishing between xylem and phloem). Simultaneously, by converting textual features extracted from relevant TCM books (such as "cinnabar dots") into SAM3 text prompts, automatic segmentation of specific micro-morphologies is achieved. Based on the segmentation mask, quantitative data such as impurity area (for doping rate) and cross-sectional structure ratio are accurately calculated, rather than relying solely on classification probabilities. This also solves the challenges of "quantitative measurement" and "irregular feature recognition" of TCM decoction pieces.

[0019] 4. This application introduces a hard parameter-shared multi-task learning (MTL) backbone network, which includes a shared feature extractor and serves three downstream tasks: (a) a global classification task (determining the type of medicinal material); (b) a contrastive learning task (based on supervised contrastive loss, bridging the gap between homogeneous medicinal materials and widening the gap between heterogeneous medicinal materials); and (c) a visual cue generation task (generating an embedding to guide SAM3 segmentation). An adaptive loss weight balancing mechanism is introduced to resolve gradient conflict issues in multi-task training.

[0020] 5. In this application, tasks such as defect region localization and global semantic feature mapping are not completed independently by a single task head, but rely on cross-task collaboration of multiple task heads—the contrastive learning head ensures the discriminative power of global features, the segmentation cue head connects to the downstream segmentation task, and finally achieves a closed loop of "feature extraction - defect localization - quantitative measurement". Multiple task heads cooperate with each other to complete the task. In addition, the collaboration of the three task heads is achieved through GradNorm dynamic loss weight balancing (total loss = classification loss + contrastive loss + segmentation loss), ensuring that "defect region localization", "global semantic mapping" and "category classification" are optimized simultaneously, avoiding single task dominating training.

[0021] 6. This application can output quality inspection reports containing reasoning logic (e.g., "determined to be unqualified because the detected rhizome residue rate reached 5%, exceeding the standard limit of 2%), which is highly interpretable and meets the requirements of pharmaceutical industry GMP for data integrity. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of the method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the parsing of the description of traditional Chinese medicine decoction pieces (Astragalus membranaceus) into a JSON object in an embodiment of the present invention; Figure 3 This is a schematic diagram of visual question answering (VQA) interaction in an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Definitions: Computable-identifiable atoms are a concept that combines computational theory and formal verification. At its core, it refers to a class of "atomic-level" objects whose properties or identities can be definitively determined through algorithms (computable functions). In this application, "atom" does not refer to the atom in chemistry, but rather to an indivisible basic unit.

[0026] Example 1: Firstly, this invention proposes a method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, see [link to relevant documentation]. Figure 1 The method includes: S1. Based on a multimodal large model, standard rules for quality testing and quantitative measurement of Chinese herbal medicines are extracted from books related to Chinese herbal medicines. The standard rules include key identification points of Chinese herbal medicines. S2. Input the pre-acquired images of Chinese herbal medicine slices into a multi-task learning backbone network based on a hard parameter sharing mechanism to obtain the visual features of the Chinese herbal medicine slices. S3. Based on the key identification points, guide the visual large model to perform pixel-level semantic segmentation of Chinese herbal medicine pieces and obtain quantitative measurement results of Chinese herbal medicine pieces. S4. Input the visual features and quantitative measurement results into the pre-constructed visual question-answering reasoning agent, and combine them with the standard rules to generate reasoning results for traditional Chinese medicine decoction pieces.

[0027] The following is in conjunction with the appendix Figure 1-3 The following details the implementation process of an embodiment and its preferred embodiment of the present invention, along with explanations of the specific steps S1-S4.

[0028] S1. Based on a multimodal large model, standard rules for quality testing and quantitative measurement of Chinese herbal medicine slices are extracted from relevant books on Chinese herbal medicine slices. The standard rules include key identification points of Chinese herbal medicine slices.

[0029] This step is used to construct a multimodal standard knowledge base for the quality testing and quantitative measurement of Chinese herbal medicine pieces, and to extract standard rules and characteristics related to the identification, testing, and measurement of Chinese herbal medicine pieces.

[0030] Unstructured text and image data were collected from relevant standards and authoritative books on Chinese herbal medicine identification. Optical character recognition (OCR) and natural language processing (NLP) technologies were used to extract the morphological descriptions (such as color, texture, and cross-sectional features), quantitative thresholds (such as diameter, impurity limits, and ash content requirements) and identification terms of various Chinese herbal medicine slices, and a domain knowledge graph based on a vector database was constructed.

[0031] In a preferred embodiment, this step does not merely involve text-level storage of the identification standards for traditional Chinese medicine decoction pieces, but rather uses a multimodal large model to parse the morphological descriptions, identification terms, and quantitative thresholds into computable discriminative atoms. These computable discriminative atoms include at least: (1) Structural atoms: Geometric topological features used to describe the cross-sectional texture morphology (such as radial cracks and concentric rings); (2) Proportional atoms: used to describe the area or voxel proportion of different tissue regions in the whole image of medicinal slices; (3) Threshold type atom: used to describe the executable judgment rules of quantitative constraints (such as "the proportion of impurity region ≤ 3%)".

[0032] Each identification atom is further bound to a corresponding computational constraint function and decision logic, enabling the text standard to directly drive the subsequent visual segmentation and quantitative measurement process, thereby achieving a closed-loop mapping of 'text standard → visual computation → quantitative decision'.

[0033] Specifically, the above steps include: 1) Obtain the raw data.

[0034] Unstructured text and image data were collected from relevant standards for the identification of Chinese medicinal herbs and authoritative books on the identification of Chinese medicinal herbs. The data obtained included, but was not limited to, PDF scans, XML documents, and high-resolution microscopic images.

[0035] 2) Utilize multimodal large models for content extraction.

[0036] Different large models are used to process different raw data, including: Use an OCR engine (such as PaddleOCR) to recognize text content; Utilize large language models (such as GPT-4o or the open-source Qwen-VL) to perform semantic chunking on the text. For example, parse the description of "Astragalus membranaceus" into a JSON object, specifically as follows: Figure 2 As shown.

[0037] For the images in the atlas, MLLM is used to generate detailed visual captions, such as: "The image shows the typical features of a cross-section of Astragalus membranaceus, with obvious radial fissures (chrysanthemum heart) and dark brown cambium rings."

[0038] 3) Vectorized storage of extracted content.

[0039] The parsed text blocks and image descriptions are transformed into high-dimensional vectors through an embedding model (such as BGE-M3) and stored in a vector database (such as Milvus) so that relevant standards can be quickly retrieved later through RAG (Retrieval-Augmented Generation).

[0040] It should be noted that the standard digitization and feature deconstruction based on RAG can achieve adaptive standards for the quality detection and quantitative measurement of Chinese herbal medicine pieces. When the relevant identification standards for Chinese herbal medicine pieces are updated, only the knowledge base document needs to be updated, without retraining the visual model, which has extremely high flexibility. In addition, in some embodiments, a few-shot learning interface can be set up, allowing users to upload a small number (1-5 images) of novel counterfeits or rare diseases and their descriptions. The detection logic can be updated in real time using the context learning capability of the multimodal large model without retraining the backbone network.

[0041] S2. Input the pre-acquired images of Chinese herbal medicine slices into a multi-task learning backbone network based on a hard parameter sharing mechanism to obtain the visual features of the Chinese herbal medicine slices.

[0042] When acquiring images of Chinese herbal medicine slices, an industrial-grade multi-angle imaging system is used to obtain high-resolution images of the Chinese herbal medicine slices under inspection in the visible light (RGB) and near-infrared (NiR) bands, covering dorsal, ventral, and cross-sectional views.

[0043] After acquiring high-dimensional multispectral images, the acquired images are input into a multi-task learning backbone network based on a hard parameter sharing mechanism to obtain the visual features of the Chinese herbal medicine slices to be inspected. However, to achieve real-time detection on an industrial production line, it is essential to avoid deploying a separate model for each detection task. This embodiment designs a hard parameter sharing multi-task learning (MTL) network architecture, which effectively solves this problem.

[0044] The multi-task learning backbone network based on the hard parameter sharing mechanism includes a shared feature encoder and multiple task-specific heads. The shared encoder simultaneously supports three downstream tasks, with each task head undertaking a specific task objective. The multiple task-specific heads are used to execute the backbone network's target tasks: medicinal herb classification, defect region localization, and global semantic feature mapping.

[0045] In one embodiment, the multi-task learning backbone network uses a Swing Transformer or ResNet as a shared encoder and introduces an adaptive gradient weighting mechanism. This mechanism dynamically adjusts the weights of each task in backpropagation by monitoring the gradient change rates of classification loss, regression loss, and contrastive loss, in order to suppress gradient conflicts and prevent a single task from dominating the training process.

[0046] It should be noted that the Swin Transformer-Large (a large-parameter version of the Swin Transformer series) is used as the encoder with shared hard parameters. This structure excels at capturing visual features at different scales, recognizing both overall contours and preserving fine textures. The input image is resized to 1024 x 1024 resolution.

[0047] In one embodiment, the multiple specific task heads specifically include: a classification head, a contrast learning head, and a segmentation prompt head. These multiple specific task heads need to cover the classification of medicinal slices, defect region localization, and global semantic feature mapping. The above three task heads fully cover this requirement through a combination of "direct implementation + indirect support." Specifically, the classification head directly implements "medicinal slice classification"; the segmentation prompt head directly implements "global semantic feature mapping"; and the contrast learning head indirectly implements the core support task of "defect region localization" through feature discrimination optimization. The specific correspondence between the task heads and the target task is shown in Table 1 below. Table 1. Correspondence between target tasks and task headers 1) Task header corresponding to the task of classifying medicinal slices: Classification Head The explicit function of the classification head is to output the probability of the medicinal material type (P_cls), which directly matches the "medicinal herb type classification" task. The medicinal herb type classification head consists of a global average pooling (GAP) layer and a fully connected layer, outputting the probability of the medicinal material type (P_cls). This enables the identification of different types of Chinese medicinal materials.

[0048] 2) Task header for defect region localization: Segmentation prompt head + contrast learning head The core objective of "defect region localization" is to identify defective areas such as impurities, mold, insect infestation, and hollowness in traditional Chinese medicine decoction pieces. This objective needs to be achieved through pixel-level segmentation by the subsequent SAM 3 model. In this embodiment, the core function of the segmentation prompt head is to generate dense image embeddings, which are directly used as image encoding inputs for the subsequent SAM 3 model, or as visual prompts for SAM 3 (such as "segment all foreign objects, rhizomes, and non-medicinal stems"). Based on this visual prompt, the SAM 3 model combines defect descriptions in standard rules (such as "residual rhizomes," "moldy spots," "impurities," "insect-infested areas," and "hollow areas") to achieve accurate localization and segmentation of defective areas (e.g., "segment all foreign objects, rhizomes, and non-medicinal stems"), and outputs a mask for the defective areas. The quantitative measurement module calculates parameters such as the area and proportion of the defective areas to complete "defect region localization + quantification."

[0049] "Dense feature map embedding" is essentially a structured mapping of global semantic information of medicinal material images (including global semantics such as texture, structure, and local features), which directly corresponds to the "global semantic feature mapping task". The core purpose of this feature embedding is "to serve as the image encoding input or visual cue for the subsequent SAM 3 model". The task of SAM 3 is pixel-level semantic segmentation (such as segmenting medicinal parts / defect areas), which requires the structured information provided by global semantic feature mapping. Therefore, the core task of the segmentation cue head is to complete the extraction and mapping of global semantic features to support downstream segmentation tasks.

[0050] It should be noted that defect region localization is not achieved independently and directly by a single task head, but rather indirectly through the collaboration of a "contrast learning head + segmentation prompt head + SAM 3 model": the contrast learning head optimizes feature discrimination, making the features of defective regions (such as counterfeit products or impurities) significantly different from those of genuine products; the segmentation prompt head provides visual guidance to SAM 3, offering precise visual cues; and based on these features and cues, SAM 3 ultimately completes pixel-level localization (segmentation) of the defect region. In other words, the segmentation prompt head is the core pre-support for defect region localization, while the contrast learning head is the "feature optimization core" for defect region localization, corresponding to the core supporting role of this task.

[0051] 3) The corresponding task head for the global semantic feature mapping task: contrastive learning head + segmentation prompt head (dual task head collaboration).

[0052] The core of "global semantic feature mapping" is to transform the visual information of Chinese herbal medicine slice images into feature embeddings with global semantic associations, providing structured feature support for subsequent segmentation and reasoning, and supporting "feature aggregation of similar high-quality slices and feature separation of heterogeneous / inferior slices", corresponding to the synergistic effect of the two task heads: The Contrastive Head's core function is to map features extracted by the shared encoder onto a 128-dimensional hypersphere using Supervised Contrastive Loss. This achieves the goal of "bringing closer the feature distance between homogeneous / same-grade medicinal slices and widening the feature distance between heterogeneous / counterfeit products," thus enhancing the discriminative power of the features. This head directly performs "differential mapping of global semantic features," for example, distinguishing between "wild astragalus" and "cultivated astragalus," and "genuine salvia miltiorrhiza" and "counterfeit red astragalus," generating feature vectors with global semantic discriminative power. This design is used to distinguish between "counterfeit products with extremely similar appearances (such as wild ginseng and cultivated ginseng)" or "medicinal materials of different grades / origins." Unlike "defect region localization," the core of this design is identifying non-genuine features in medicinal materials (such as defects in counterfeit or inferior products). The Contrastive Head optimizes the feature space distribution, providing highly discriminative feature support for the subsequent localization of defect regions (i.e., indirectly assisting in the localization of defect regions through feature differences). This also demonstrates a strong correlation between the Contrastive Head and "defect region localization" (identifying defects in counterfeit / inferior products).

[0053] It's important to note that the Contrastive Head is a non-linear projection head that maps features onto a 128-dimensional unit hypersphere. Furthermore, supervised contrastive loss is applied during contrastive learning. The output of the Contrastive Head is not directly the similarity value between the detected herbal medicine slice and the most similar sample in the database. Instead, it is a 128-dimensional unit hypersphere feature vector (also called an "embedding vector") after non-linear projection. This vector serves as a standardized identification feature carrier for distinguishing between "high-quality products of the same type" and "counterfeit products of different types." The similarity score is a derived result calculated based on this feature vector, not the direct output of the Contrastive Head.

[0054] The segmentation prompt head, which assists in mapping, not only serves SAM 3 segmentation but also contains semantic information about the global structure of the medicinal slices (such as the ratio of bark to wood and the distribution of cross-sectional texture). It complements the feature vectors generated by the contrastive learning head, together forming a complete "global semantic feature mapping" result, which supports the comprehensive judgment of the subsequent VQA inference agent.

[0055] It should be noted that the core of global semantic feature mapping is "generating structured feature embeddings" (serving subsequent segmentation), which is directly output by the segmentation prompt head; while the core of contrastive learning head is "optimizing feature space distribution" (enhancing discriminative power). Its output is a feature vector projected onto a 128-dimensional hypersphere. The purpose is to help distinguish between genuine / counterfeit products and different grades, and indirectly provide a feature basis for the location of defective areas, rather than directly completing "global semantic mapping".

[0056] For images within a batch, if they belong to the same origin or grade (positive sample pair), their feature distance is reduced; if they belong to different origins or are counterfeit (negative sample pair), their distance is increased. This is crucial for distinguishing counterfeit products that are "extremely similar in appearance" (such as wild ginseng and cultivated ginseng).

[0057] Segmentation Prompt Head: Performs the defect region localization task, assisting the subsequent large visual model (such as the SAM 3 model) in locating defect regions and providing visual guidance. The segmentation prompt head generates dense image embeddings, which can be directly used as image encoding input for the subsequent SAM 3 model, or as visual prompts for SAM 3.

[0058] In one embodiment, to prevent the classification task (which is relatively simple) from converging too quickly and inhibiting the learning of fine-grained features, an adaptive gradient weighting mechanism is introduced to achieve adaptive loss balance. Specifically, the GradNorm algorithm is used to dynamically adjust the loss weights of each task. .

[0059] in: The core optimization objective is represented by the total loss function for multi-task learning, which is also the sum of the weighted losses for classification, contrastive learning, and segmentation tasks. By integrating the losses from multiple tasks, the shared feature encoder (such as the Swing Transformer) is guided to learn common visual features that are suitable for all downstream tasks. This ensures that the model is optimized simultaneously in the three core tasks of "medicinal slice classification, fine-grained identification, and pixel-level segmentation," ensuring collaborative learning among tasks and ultimately improving the accuracy of quality inspection and quantitative measurement.

[0060] 、 、 All three represent task weight coefficients, which are loss weight coefficients dynamically adjusted based on the GradNorm algorithm, updated in real time with the L2 norm of the gradients for each task during training. The core purpose of these three task weight coefficients is to address the "gradient conflict" problem in multi-task training (e.g., simple classification tasks may converge quickly, inhibiting feature learning in fine-grained contrastive learning or segmentation tasks), ensuring that all tasks iterate and optimize at a similar rate, and avoiding a single task dominating the training process. (Classification Weight) represents the loss weight for the classification task, used to balance the cross-entropy loss. The contribution percentage of the total loss to the "Classification of Herbal Pieces" task, which is used to adjust the contribution percentage of the total loss to the "Classification of Herbal Pieces" task. (Contrastive Weight) represents the loss weight for the contrastive learning task, used to balance the supervised contrastive loss. The contribution percentage of ) (Segmentation Weight) represents the loss weight for the segmentation task, used to balance the Dice loss. The contribution percentage of the "pixel-level semantic segmentation" task is used to adjust the loss of the "pixel-level semantic segmentation" task.

[0061] , , Representing the loss function for each task: (Cross-Entropy Loss): Corresponding to the task of classifying medicinal herbs (executed by the "classification head"); it measures the difference between the probability distribution of the medicinal herb types predicted by the model and the true labels (such as "Astragalus", "Salvia miltiorrhiza", "counterfeit Astragalus membranaceus"), optimizing the model's classification accuracy for major categories of medicinal herbs. For example, it is used to train the model to distinguish different types of Chinese medicinal herbs (such as Astragalus membranaceus and Salvia miltiorrhiza), and to identify genuine / counterfeit, high-quality / low-quality medicinal herbs, which is the core loss for achieving "qualitative identification".

[0062] (Supervised Contrastive Loss): Corresponding to fine-grained contrastive learning tasks (executed by the "contrastive learning head" of a multi-task network), it enhances the discriminative power of feature vectors by "bringing the feature distance of similar high-quality medicinal slices closer together and widening the feature distance of heterogeneous / inferior medicinal slices." It solves the fine-grained differentiation problem of "same name, different origin" (same variety, different place of origin) and "similar morphological counterfeits" (such as wild ginseng and cultivated ginseng) in traditional Chinese medicine medicinal slices, making the features learned by the model more targeted.

[0063] (Dice Loss): For pixel-level semantic segmentation tasks (executed by the "segmentation hint head" to provide feature support for the subsequent SAM 3 model), it is used to measure the degree of overlap between the segmentation mask predicted by the model (such as "xylem", "phloem", "impurity region", or "chrysanthemum heart" of Astragalus membranaceus, "cinnabar dot" of Salvia miltiorrhiza, and impurity region) and the real labeled mask, and optimize the accuracy of segmentation boundaries. This method optimizes the segmentation of irregular biological features of traditional Chinese medicine decoction pieces (such as the "chrysanthemum heart" of Astragalus membranaceus and the "cinnabar dots" of Salvia miltiorrhiza). Compared with traditional cross-entropy loss, it is more suitable for handling segmentation tasks with imbalanced samples (such as a small proportion of impurity regions), and provides an accurate mask basis for quantitative indicators such as "doping rate calculation and feature density measurement".

[0064] It should be noted that the Adaptive Gradient Weighting mechanism dynamically adjusts the weights of each task in backpropagation by monitoring the gradient change rates of classification loss, regression loss, and contrastive loss, thereby achieving adaptive loss balance, suppressing gradient conflicts, and preventing a single task from dominating the training process.

[0065] In a preferred embodiment, supervised contrastive learning is performed within a shared feature space. This generates discriminative visual feature vectors by maximizing the similarity between features of high-quality herbal slices of the same type and standard semantic embeddings, while minimizing their similarity to features of counterfeit or inferior products. The supervised contrastive learning employs a hierarchical loss function. in: For anchor point sample features, anchor point samples i The high-dimensional feature vector (128-dimensional in the document) is extracted and projected by the multi-task learning backbone network (Swin Transformer).

[0066] For all samples with anchor points i A set of samples of the same type and grade constitutes the entire set of positive samples. In the detection of traditional Chinese medicine decoction pieces, this specifically refers to the anchor sample. i A collection of samples of the same type and of the same quality level (such as anchor points) i If it is classified as "second-grade Danshen", then... (For all second-grade Danshen samples), used to enhance "feature aggregation of similar high-quality samples".

[0067] For all samples with anchor points iThe set of samples that do not meet the "same species, same grade" requirement encompasses all types of negative samples. In the detection of Chinese herbal medicine slices, there are three types of negative samples: 1. Samples from different origins (same name, different production areas, such as different production areas of the same species of Astragalus); 2. Counterfeit products with similar characteristics (such as the counterfeit "Gansu Danshen" of Salvia miltiorrhiza); 3. Samples of different grades (such as first-grade vs. third-grade), used to achieve "heterogeneous sample feature separation".

[0068] Temperature coefficient is a hyperparameter controlling the gradient smoothness of the contrastive loss function, used to adjust the discriminative sensitivity of feature similarity, thereby enhancing the model's fine-grained ability to distinguish between "same name, different location" or "similar phenotypic counterfeits". This parameter is typically used in the detection of traditional Chinese medicine decoction pieces. ∈[0.01, 0.1]: The smaller the value, the more stringent the distinction between similar features (suitable for distinguishing "extremely similar counterfeit products", such as wild ginseng and cultivated ginseng). The larger the value, the smoother the feature similarity calculation (suitable for handling "slight trait differences", such as Astragalus membranaceus from different origins of the same grade).

[0069] The index identifier representing the anchor sample is a unique identifier for a single sample in the training batch. In this embodiment, it refers to one sample of Chinese herbal medicine to be detected (such as one slice of Astragalus membranaceus or one piece of Salvia miltiorrhiza), which serves as the "benchmark sample" for comparative learning.

[0070] This refers to the set of all anchor point samples in the training batch, that is, the total number of samples participating in the comparison calculation in the current training iteration. For example, in this embodiment, it is the set of Chinese herbal medicine samples to be trained in a certain batch (such as a batch containing 50 Astragalus membranaceus samples, 30 Salvia miltiorrhiza samples, and 20 counterfeit samples).

[0071] Represents the set of positive samples The single sample index in is related to the anchor sample. i In this embodiment, the reference to samples that meet the "same type and same level" condition is related to the anchor sample. i Samples of prepared Chinese medicinal herbs belonging to the same type and quality grade (e.g., when the anchor point is "first-class Astragalus membranaceus"), (Referring to other first-grade Astragalus samples).

[0072] Indicates positive samples p The high-dimensional feature vector extracted and projected by the backbone network has dimensions and Consistent (128-dimensional), with anchor sample iSample feature vectors with the same core characteristics (such as the feature vector of first-grade Astragalus membranaceus with "distinct chrysanthemum heart and sufficient powder"); Represents the set of negative samples The high-dimensional feature vector obtained by extracting and projecting a single sample from the backbone network, with dimensions equal to... Consistent with anchor point samples. i The sample feature vectors that differ include: samples from different origins (e.g., Astragalus membranaceus from Gansu and Inner Mongolia); counterfeit samples (e.g., counterfeit Astragalus membranaceus "Golden Chicken"); samples of different grades (e.g., samples with the anchor point being first grade, It is a second-grade / general grade Astragalus membranaceus.

[0073] S3. Based on the key identification points, guide the visual large model to perform pixel-level semantic segmentation of Chinese herbal medicine pieces and obtain quantitative measurement results of Chinese herbal medicine pieces.

[0074] This step generates natural language prompts (Text Prompts) based on the key identification points extracted in step S1, namely the feature descriptions of Chinese herbal medicine slices (such as "yellowish-white bark", "pale yellow xylem", "with radial fissures"). These prompts are then input into a finely tuned Segment Anything Model 3 (SAM 3) model, guiding the model to perform pixel-level semantic segmentation of specific anatomical structures (such as the cambium ring, xylem, and phloem) and foreign objects (such as non-medicinal parts and mold spots). Quantitative measurements of the Chinese herbal medicine slices are then performed based on masks of the segmented contours.

[0075] In this embodiment, pixel-level semantic segmentation is not completed in one go, but is generated iteratively through a standard constraint segmentation mechanism.

[0076] Specifically, the computable distinguishable atoms corresponding to the key distinguishing points are transformed into segmentation constraints and dynamically injected during the multi-round inference process of the SAM3 model: (1) In the initial segmentation stage, the visual big model generates candidate segmentation results based on global semantic features; (2) During the constraint verification stage, the candidate segmentation results are matched with the threshold constraints of the identification atoms to identify regions that do not meet the standard conditions; (3) In the correction and guidance stage, regions that do not meet the standard conditions are embedded as new visual cues to guide the segmentation model to perform local refinement.

[0077] By using the above methods, the standard rules can be used to constrain the segmentation process in real time, so that the final segmentation results meet the quantitative requirements of the quality standards for Chinese herbal medicine slices at the pixel level.

[0078] The specific steps include: 1) Text-driven segmentation: Using the feature words extracted in step S1, dynamic prompts are constructed. For example, for a slice of Salvia miltiorrhiza, the system generates the prompt: "Segment the red cinnabarpoints in the cross section". Based on the prompt, SAM3 accurately outlines the contour of each "cinnabarpoint" in the image and generates a binary mask.

[0079] 2) Quantitative Measurement: Based on the actual needs of quality testing and quantitative measurement of different Chinese herbal medicine slices, one or more quantitative measurements can be performed, including but not limited to doping rate calculation, characteristic density calculation, and size measurement. Specifically, the calculation process is as follows: Doping rate calculation: The standard definitions of "medicinal parts" and "non-medicinal parts" (such as "removed stem" and "removed pit") corresponding to the batch to be inspected are identified and defined by a multimodal large model. SAM 3 was used to segment all discrete objects in the image and classify them according to texture features into "genuine tissue", "residual stem base (root tip)", "subsidiary impurities (such as mud and sand)" and "foreign matter"; the prompt was "Segment all foreign matter, reed, and non-medicinal stems". The SAM 3 model outputs an impurity mask.

[0080] Calculate the sum of pixel areas A_{impurity} of non-medicinal parts and impurities and the total pixel area A_{total} of the sample. The doping rate is calculated as follows: or, Doping rate = Both of the above calculation formulas can be used to calculate the doping rate because the pixel area and the mask are equivalent concepts in this scenario (the core function of the mask is to define the region, and then the pixel area is obtained by counting the number of pixels covered by the mask).

[0081] Finally, the doping rate R value was compared with the limit threshold specified in relevant books to obtain the quantitative measurement results of Chinese herbal medicine slices.

[0082] Characteristic density calculation: For Astragalus membranaceus, calculate the proportion of the "chrysanthemum heart" fissure area. If the proportion is too large (e.g., >20%) or too small (e.g., <1%), it may indicate that the medicinal material is decayed or has not been grown for long enough.

[0083] Size measurement: The diameter and thickness are calculated based on the minimum bounding rectangle of the mask, the pixel units are converted to physical units (mm), and compared with the standards in relevant books.

[0084] It should be noted that although SAM 3 has a strong zero-sample capability, the biological structure of Chinese herbal medicine pieces (such as phloem, xylem, and oil cells) differs greatly from that of general objects. Fine-tuning is necessary to solve the problems of "quantitative measurement" and "irregular feature recognition" of Chinese herbal medicine pieces.

[0085] In one embodiment, the fine-tuning of the SAM 3 model employs a low-rank adaptation (LoRA) strategy. This strategy keeps the massive image encoder parameters of SAM 3 frozen to preserve its generalization ability, injecting a low-rank adaptation (LoRA) matrix into the Transformer layer of the mask decoder. Training only these few LoRA parameters (approximately 1% of the total parameters) allows the model to adapt to both microscopic and macroscopic textures of traditional Chinese medicine (TCM), and it is trained using a dedicated dataset containing annotations of micro-features of TCM (such as the "golden well jade railing" feature of Astragalus membranaceus and the "purple oil" feature of Salvia miltiorrhiza) to adapt to the segmentation of irregular biological edges.

[0086] S4. Input the visual features and quantitative measurement results into the pre-constructed visual question-answering reasoning agent, and combine them with the standard rules to generate reasoning results for traditional Chinese medicine decoction pieces.

[0087] This step inputs the visual features and quantitative measurement results into a pre-built visual question-answering reasoning agent, and combines them with the standard rules. The reasoning result is not a single pass / fail judgment, but a quality reasoning chain generated by the visual question-answering reasoning agent.

[0088] The inference chain includes at least: (1) Triggered identification atom identifier; (2) The corresponding quantitative measurement results; (3) Comparison with standard threshold; (4) Final quality assessment conclusion.

[0089] The output of this reasoning chain enables the test results to be traceable and interpretable, meeting the pharmaceutical industry's requirements for the auditability of the quality control process.

[0090] Specifically: A multimodal large model (MLLM) visual question answering (VQA) reasoning agent is constructed to achieve comprehensive rating of Chinese herbal medicine slices. The visual question answering (VQA) reasoning agent performs reasoning through the "thought chain" (CoT) model: "Detection of feature A -> SAM measures diameter as B -> Standard rule requires first-class product diameter > C -> Conclusion: Meets first-class product standard".

[0091] The geometric parameters (area, perimeter, percentage), doping rate (area of ​​non-medicinal parts / total area), texture density index, etc. of the segmented region obtained in step S3 are used as inputs along with the visual feature vector generated in step S2; the agent performs logical comparison according to the standard rules retrieved in step S1, and generates a quality inspection report that meets the specifications.

[0092] In one embodiment, the multimodal large model performs a "Chain-of-Thought" (CoT) reasoning process including: Step 1: Input a high-resolution image of the current medicinal slices and identify the prominent features visible in the image (such as "white cross-section" and "sufficient powderiness"). Step 2: Call the quantitative measurement results of SAM 3 (such as "diameter 1.2cm", "peel ratio 35%", impurity rate = 0.5%, insect-infested area = 0%); call the visual feature vector extracted by the MTL network; Step 3: Search the knowledge base for the grading standards of this variety (e.g., "First-class products require a diameter > 1.0cm and sufficient powder content"). Step 4: MLLM (such as LLaVA-NeXT or a finely tuned Qwen-VL-Chat) combines textual standard rules with visual features, quantitative measurement results and other evidence to make inferences and logically compare the measured values ​​with the standard values; Step 5: Determine the grade comprehensively and generate a natural language explanation (e.g., "Classified as first-class product because the diameter meets the standard and there are no insects").

[0093] Example of the interaction process: Figure 3 As shown.

[0094] Furthermore, MLLM can also generate descriptions of images for manual verification. For example: "The bark of this sample is yellowish-white, the wood is pale yellow, the cambium rings are clear and dark brown, and obvious radial textures are visible, which matches the characteristics of 'golden well and jade railing'." This descriptive capability solves the black-box problem of traditional CV models that can only output "category iD", enhancing the trustworthiness.

[0095] The following describes the implementation process of Example 1 and its preferred embodiment using the detection logic of typical medicinal materials: Case 1: Quality Inspection and Quantitative Measurement of Astragalus membranaceus Feature extraction: The MTL network distinguishes between wild and cultivated varieties through contrastive learning (wild varieties have denser bark).

[0096] SAM 3 task: Segment the xylem, phloem, and cambium ring.

[0097] Quantitative indicator: Calculate the bark-to-wood area ratio (phloem area / xylem area). High-quality Astragalus membranaceus has a thicker bark, and the ratio should be within a specific range (e.g., 1:2 to 1:3).

[0098] VQA determination: Combining bark-to-wood ratio and diameter data, determine whether it is "core-removed" or "sulfur-fumigated" Astragalus membranaceus (sulfur fumigation will cause abnormal white color, which can be identified by MLLM through color histogram analysis).

[0099] Case Study 2: Quality Inspection and Quantitative Measurement of Salvia miltiorrhiza Feature extraction: Identify surface color (purple-red to dark red).

[0100] SAM 3 task: Segment the scattered "cinnabar dots" (dot-like vascular bundles) on the cross-section.

[0101] Quantitative indicator: Cinnabar spot density (numbers / cm²). Studies have shown that cinnabar spot density is positively correlated with tanshinone content.

[0102] VQA assessment: If the cinnabar spots are sparse or black in color, it indicates that the content of active ingredients is low or that the product is moldy. It is recommended to downgrade or reject the product.

[0103] Example 2: This embodiment, based on Embodiment 1 and its preferred embodiments described above, further includes: Based on the reasoning results, control commands are generated to drive automated sorting equipment to perform physical separation of Chinese herbal medicine pieces.

[0104] One purpose of quality testing and quantitative measurement of Chinese herbal medicine (TCM) decoction pieces is to classify them into grades (e.g., superior, general, unqualified) based on the results of quality testing and quantitative measurement, thereby guiding the sorting of TCM decoction pieces. To intelligently adapt to existing automated sorting equipment, this embodiment outputs the grade of the decoction pieces (e.g., superior, general, unqualified) based on the reasoning results, and generates control commands to drive the automated sorting equipment to perform physical separation.

[0105] Example 3: Secondly, the present invention also provides a system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, the system comprising: The standard rule extraction module is configured to extract standard rules for quality testing and quantitative measurement of Chinese herbal medicines from books related to Chinese herbal medicines based on a multimodal large model. The standard rules include key identification points of Chinese herbal medicines. The visual feature acquisition module is configured to input pre-acquired images of Chinese herbal medicine slices into a multi-task learning backbone network based on a hard parameter sharing mechanism to acquire the visual features of the Chinese herbal medicine slices. The quantitative measurement result acquisition module is configured to guide the visual large model to perform pixel-level semantic segmentation of Chinese herbal medicine slices based on the key identification points, and acquire the quantitative measurement results of Chinese herbal medicine slices. The reasoning result generation module is configured to input the visual features and quantitative measurement results into a pre-built visual question-answering reasoning agent, and combine them with the standard rules to generate reasoning results for traditional Chinese medicine decoction pieces.

[0106] In practical implementation, the standard rule extraction module is configured to parse standard documents in PDF and Word formats, obtain standard rules for quality testing and quantitative measurement of Chinese herbal medicine pieces, and convert text and illustrations into vector indexes. In practical implementation, the visual feature acquisition module includes an optical imaging unit: containing a multispectral camera array and a coaxial light source, used to eliminate surface reflections and capture the internal texture of the medicinal slices, thereby acquiring images of the medicinal slices.

[0107] In practical implementation, the quantitative measurement result acquisition module includes a prompt generation engine: it is used to dynamically generate segmentation prompts for specific Chinese herbal medicine pieces (such as Astragalus membranaceus, Salvia miltiorrhiza, and Paeonia lactiflora) based on the preliminary classification results.

[0108] In addition, in one embodiment, in order to generate control commands based on the reasoning results to drive the automated sorting equipment to perform the physical separation of Chinese herbal medicine pieces, the above system is also equipped with a pneumatic sorting actuator: receiving grading commands and controlling high-frequency air valves to reject unqualified products.

[0109] It is understood that the system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning provided in this embodiment of the invention corresponds to the aforementioned method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning. Explanations, examples, and beneficial effects related to this system can be found in the corresponding content of the method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, and will not be repeated here. It should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, characterized in that, The method includes: Based on a multimodal large model, standard rules for quality testing and quantitative measurement of Chinese herbal medicines are extracted from relevant books on Chinese herbal medicines. These standard rules include key identification points of Chinese herbal medicines. The pre-acquired images of Chinese herbal medicine slices are input into a multi-task learning backbone network based on a hard parameter sharing mechanism to obtain the visual features of the Chinese herbal medicine slices. Based on the key identification points, a large visual model is used to perform pixel-level semantic segmentation of Chinese herbal medicine pieces to obtain quantitative measurement results of Chinese herbal medicine pieces. The visual features and quantitative measurement results are input into a pre-built reasoning agent, and combined with the standard rules, the reasoning results of traditional Chinese medicine decoction pieces are generated.

2. The method as described in claim 1, characterized in that, The multi-task learning backbone network based on the hard parameter sharing mechanism includes: a shared feature encoder and multiple specific task heads; the multiple specific task heads include at least task heads for performing the herbal medicine category classification task, defect region localization task, and global semantic feature mapping task individually or in cooperation.

3. The method as described in claim 2, characterized in that, The shared encoder includes the Swing Transformer.

4. The method as described in claim 2, characterized in that, By introducing the GradNorm algorithm, the loss weights of tasks corresponding to multiple specific task heads are dynamically adjusted: in, Represents the total loss function; 、 、 These represent the weights for the classification task, the contrastive learning task, and the segmentation task, respectively. 、 , These represent the cross-entropy loss function, the supervised contrastive loss function, and the Dice loss function, respectively.

5. The method as described in claim 1, characterized in that, Supervised contrastive learning is performed in the shared feature space of the multi-task learning backbone network of the hard parameter sharing mechanism. By maximizing the similarity between the features of high-quality medicinal slices of the same type and the standard semantic embedding, and minimizing the similarity between them and the features of counterfeit or inferior products, a visual feature vector with discriminative power is generated.

6. The method as described in claim 5, characterized in that, Supervised contrastive learning employs a hierarchical loss function: in, For anchor point sample features, For the positive sample set, To compare the sample sets, Temperature coefficient; Indicates the anchor sample index; Represents the set of anchor point samples; Indicates the positive sample index; Represents the feature vector of a positive sample; This represents the feature vector of the comparison sample.

7. The method as described in claim 1, characterized in that, The large visual model includes SAM 3.

8. The method as described in claim 7, characterized in that, A low-rank adaptation strategy is used to fine-tune the SAM 3 model: keep the parameters of the SAM 3 image encoder frozen, and inject a low-rank adaptation matrix into the Transformer layer of the mask decoder.

9. The method as described in claim 1, characterized in that, The quantitative measurements include doping rate calculation, characteristic density calculation, and size measurement.

10. A system for quality detection and quantitative measurement of traditional Chinese medicine decoction pieces based on multimodal large model and multi-task learning, characterized in that, The system includes: The standard rule extraction module is configured to extract standard rules for quality testing and quantitative measurement of Chinese herbal medicines from books related to Chinese herbal medicines based on a multimodal large model. The standard rules include key identification points of Chinese herbal medicines. The visual feature acquisition module is configured to input pre-acquired images of Chinese herbal medicine slices into a multi-task learning backbone network based on a hard parameter sharing mechanism to acquire the visual features of the Chinese herbal medicine slices. The quantitative measurement result acquisition module is configured to guide the visual large model to perform pixel-level semantic segmentation of Chinese herbal medicine slices based on the key identification points, and acquire the quantitative measurement results of Chinese herbal medicine slices. The reasoning result generation module is configured to input the visual features and quantitative measurement results into a pre-built visual question-answering reasoning agent, and combine them with the standard rules to generate reasoning results for traditional Chinese medicine decoction pieces.