Chinese herbal medicine image recognition method, device and equipment based on multi-modal model

Through the image recognition method based on multimodal model, the problem of slow training process of Chinese herbal image recognition algorithms is solved, and fast annotation, online training and model deployment are achieved, efficiency and accuracy are improved, and cost is reduced.

CN120125987APending Publication Date: 2025-06-10YUNNAN BAIYAO GRP MEDICINE E-COMMERCE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510038056.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The training process of Chinese herbal image recognition algorithm is long and slow, resulting in a long model iteration optimization cycle, difficult deployment, difficult to quickly apply and implement, and high time and cost.

Method used

The image recognition method based on multimodal model is adopted, including data acquisition module, labeling data module, model online training module and model packaging and deployment module to realize rapid labeling, online training and model deployment.

Benefits of technology

It greatly improves the efficiency and recognition accuracy of traditional training models, simplifies the deployment process of Chinese herbal medicine image recognition and multimodal model, and reduces time and capital costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125987A_ABST
    Figure CN120125987A_ABST
Patent Text Reader

Abstract

The invention relates to a Chinese herbal medicine image recognition method, device and equipment based on a multi-modal model, and the method comprises the steps: carrying out the pre-marking through a marking data module according to image data obtained by a data collection module, and carrying out the online training of an image class task model and a multi-modal task model through a model online training module according to the marked image data, amp is trained through the model; the evaluation module tests a model training result and outputs a model file, and the model file is converted into a deployable model structure and a parameter file through the model packaging and deploying module according to the output model file. According to the method, the problem of image recognition in a traditional Chinese medicine specific scene and the tedious Chinese herbal medicine introduction text and Chinese herbal medicine image modal alignment process are simplified, the multi-modal model is trained online, the effect is checked, the Chinese herbal medicine image data annotation is constructed, the traditional model training efficiency and the recognition accuracy are greatly improved, and the traditional Chinese medicine recognition efficiency is improved. Meanwhile, amp is recognized for other Chinese herbal medicine images; and a new platform solution is provided for similar work such as multi-modal model deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a method, device, and equipment for Chinese herbal medicine image recognition based on a multi-modal model. Background Art

[0002] There are problems such as a long training process, slow time, a long model iteration and optimization cycle, and difficult deployment in traditional deep learning models for the recognition of Chinese herbal medicines, which are numerous in variety. For example, for the image recognition algorithm of Panax notoginseng Chinese herbal medicine, the entire algorithm, from data collection, data annotation, model training, model deployment to model online, requires 3 - 6 months. Among them, the time period for data collection and annotation is 2 - 3 months. It is difficult for the algorithm to be quickly applied and implemented, greatly consuming time and capital costs. Summary of the Invention

[0003] This application provides a method for Chinese herbal medicine image recognition based on a multi-modal model, which is characterized by including: According to the image data obtained by the data collection module, use the annotation data module for pre-annotation; According to the annotated image data, use the model online training module to online train an image class task model and a multi-modal task model, and test the model training results through the model training & evaluation module and output a model file; According to the output model file, convert it into a deployable model structure and parameter file through the model encapsulation and deployment module.

[0004] Optionally, the method for Chinese herbal medicine image recognition based on a multi-modal model is characterized in that: The image class task model is divided into a detection model and a segmentation model according to the business and application scenarios and generates a model file; The multi-modal task model uses a neural network to classify and recognize image data, and then inputs it into LLMs and outputs a model file; The model training & evaluation module tests the basetestcase of the image class model and the multi-modal model according to the code and generates a model file.

[0005] Optionally, the multi-modal task model uses a neural network to classify and recognize an image, and then inputs it into LLMs and outputs a model file, including: The multi-modal task model obtains the feature vector of the image data through the Image Encoder, and converts the image data into a displayable pixel format through the Image Dncoder; The feature vector and the labeled text obtain a vector that shows the internal connection and key features between the display image and the text through Q-Former, and then capture the relationships between elements through Transformer to generate an output sequence. The output sequence is formed into an output result using Fully Connected; According to the generated displayable pixel format and the output result, Flatten is used to convert the multi-dimensional input data into a one-dimensional vector, and the vector is input into LLMs to generate a multi-modal model.

[0006] Optionally, the method for identifying Chinese herbal medicine images based on a multi-modal model is characterized in that: According to different application scenarios, the PC side uses an API encapsulation port, and the mobile side uses a "mobile model" deployment interface; The Chinese herbal medicine image recognition platform provides tools for converting the model into a mobile device and deployment code basecase.

[0007] Optionally, the image task model is divided into a detection model and a segmentation model according to business and application scenarios and generates model files, including: The detection model is used to determine the image category and the position of the object in the image; The segmentation model is used to detect the objects in the image and generate a pixel-level mask for each detected object to distinguish the boundaries of different objects.

[0008] Optionally, for the image data obtained by the data acquisition module, pre-annotation is performed using the annotation data module, including: The annotation of the image data and the model is to pre-annotate the unannotated data through existing image models, text models, and multi-modal models, and provide a function for manually modifying the annotation.

[0009] Optionally, the model training & evaluation module generates model files by online testing the base testcases of the image model and the multi-modal model according to the code, including: The online code testing is to perform self-testing using the platform code or upload the code by oneself for testing.

[0010] This application also provides a device for identifying Chinese herbal medicine images based on a multi-modal model, characterized in that the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the method for identifying Chinese herbal medicine images based on a multi-modal model according to any one of claims 1 to 7.

[0011] Optionally, the device for Chinese herbal medicine image recognition based on a multimodal model is characterized in that the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the method for Chinese herbal medicine image recognition based on a multimodal model according to any one of claims 1 to 7 is implemented.

[0012] This application also provides an electronic device, which is characterized by including: a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method for Chinese herbal medicine image recognition based on a multimodal model according to any one of claims 1 to 7 is implemented.

[0013] The beneficial effects of this application are as follows: The problem of image recognition in specific scenarios of traditional Chinese medicine, as well as the cumbersome process of aligning the text of Chinese herbal medicine introductions and the modalities of Chinese herbal medicine images, are simplified. Fast annotation is designed, a multimodal model is trained online, the model is quickly deployed at the application end, and a system for Chinese herbal medicine image data annotation, online model training & viewing of training effects, to model storage and quick deployment to the application end is constructed, greatly improving the efficiency of traditional training models and the recognition accuracy. At the same time, it provides a new platform solution for similar work such as other Chinese herbal medicine image recognition & multimodal model deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following briefly introduces the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0015] Figure 1 Shows a flowchart of a method for Chinese herbal medicine image recognition based on a multimodal model disclosed in this application; Figure 2 Shows a module diagram of a Chinese herbal medicine image recognition platform disclosed in this application; Figure 3 Shows a structure diagram of a multimodal basemodel disclosed in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following will detail various exemplary embodiments, features, and aspects of this application with reference to the drawings. The same reference numerals in the drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, unless otherwise specified, the drawings do not have to be drawn to scale.

[0017] Among them, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.

[0018] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior or better than other embodiments.

[0019] In addition, for a better illustration of this application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that this application can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of this application.

[0020] This application is a method for image recognition. In this method, the user can upload an image of Chinese herbal medicine into the system through a software interface. The system will analyze and process the uploaded image, extract the information of the Chinese herbal medicine in the image, and label the Chinese herbal medicine in the image, so as to realize the recognition and analysis of the Chinese herbal medicine image.

[0021] As Figure 1 shown, a method for Chinese herbal medicine image recognition based on a multi-modal model according to an embodiment of this application specifically includes the following steps: S100, according to the image data obtained by the data acquisition module, use the annotation data module for pre-annotation.

[0022] In this step, the data acquisition module is used to receive the uploaded image data and store these image data to realize the centralized storage and management of captured, crawled, purchased, and free data, and then transmit the stored image data into the annotation data module. In the annotation data module, the unannotated data is pre-annotated through the existing image data, text data, and multi-modal model, and a function for manually modifying annotations is provided to achieve one-to-one correspondence between the data and the annotations.

[0023] S200, according to the annotated image data, use the model online training module to online train the image class task model and the multi-modal task model, and test the model training result through the model training & evaluation module and output the model file.

[0024] In this step, the model online training module is used to receive the pre-annotated image data for online training of the model. According to different business and model application scenarios, an image task model and a multi-modal task model are trained respectively. The image task model generates an image detection model and an image segmentation model for image data detection, classification, and image segmentation. The multi-modal task model processes the image data and the pre-annotated text data to generate a multi-modal model.

[0025] And the model training & evaluation module is used to receive the generated model and automatically test the model training effect online. This module provides base testcases for image classification tasks, detection tasks, segmentation tasks, and multi-modal tasks, and can perform online patrol of test results. It also provides the function of custom uploading code to test the model results.

[0026] S300. According to the output model file, it is converted into a deployable model structure and parameter file through the model encapsulation and deployment module.

[0027] In this step, the model encapsulation and deployment module receives the trained model, quickly converts it into a deployable model structure and parameter file, and according to different application scenarios, the PC side uses the API encapsulation interface, and the mobile side uses the "mobile model" deployment interface. The system provides tools for converting the model to the mobile side and the basecase of deployment code. After the model structure and parameter file are generated, they can be transmitted to the application side for model online launch.

[0028] The above content can use the existing image data, text data, and multi-modal model through the annotation data module to achieve rapid annotation, and use the model online training module and the model training & evaluation module to achieve online training of the multi-modal model and online viewing of the training effect, and use the model encapsulation and deployment module to achieve rapid model deployment on the application side. At the same time, it greatly improves the efficiency of traditional training models and the recognition accuracy, providing a new platform solution for similar work such as other Chinese herbal medicine image recognition & multi-modal model deployment.

[0029] Such as Figure 2 shown, the Chinese herbal medicine image recognition platform includes the following content: First, the purpose of the data acquisition module is to collect and store data, centralizing and managing the data collected through shooting, web crawling, procurement, and free datasets. Secondly, the purpose of the data annotation module is to pre-annotate the unannotated data using existing image models, text models, and multi-modal models for the data that needs to be annotated, and provide a function for manual annotation modification, with the data and annotations corresponding one by one. Then, the purpose of the model online training module is to train the model online, combining different business and model application scenarios to train image-based task models and multi-modal task models online. Next, the purpose of the model testing & evaluation module is to automatically test the model training effect online. This module provides base test cases for image classification tasks, detection, segmentation tasks, and multi-modal tasks, and can conduct online patrol tests on the results, and also provides the function of custom uploading code to test the model results. Finally, the purpose of the model encapsulation and deployment module is to quickly convert the trained model into a deployable model structure and parameter file, and according to different application scenarios, use the API encapsulation interface for the PC side and the "mobile model" deployment interface for the mobile side. The platform provides tools for converting the model into a mobile version and the base case of the deployment code.

[0030] Among them, the image-based task models are divided into detection models and segmentation models according to business and application scenarios and generate model files. Among them, the multi-modal task model uses neural networks to classify and identify image data, then inputs it into LLMs and outputs model files; the model training & evaluation module conducts online tests on the base test cases of image-based models and multi-modal models according to the code and generates model files.

[0031] Specifically, the purpose of the image-based task model is to train the model online, and different functional models can be generated according to different business and model application scenarios, such as the Panax notoginseng inspection model, the Panax notoginseng classification model, and the Panax notoginseng image segmentation model. Among them, the functions of the Panax notoginseng detection model include detecting the appearance and morphology of Panax notoginseng, and the system detects the color, shape, size, and surface features of Panax notoginseng to evaluate whether it meets the standards; the functions of the Panax notoginseng classification model include constructing a Panax notoginseng feature dictionary, obtaining the contour vector map of Panax notoginseng, and calculating the main data of Panax notoginseng in each dimension based on the contour vector map, and then comparing these data with the indicators in the Panax notoginseng feature dictionary to achieve the classification of Panax notoginseng; the functions of the Panax notoginseng image segmentation model are reflected in the disease identification and quality control of Panax notoginseng. Through image segmentation technology, the disease spots are separated from the complex background, improving the accuracy and effect of disease spot segmentation, and can more accurately identify the disease spots on the Panax notoginseng leaves.

[0032] As Figure 3 shown, the multi-modal basemodel structure diagram includes the following content: The multi-modal task model obtains the feature vector of the image data through the Image Encoder, and converts the image data into a displayable pixel format through the Image Dncoder. The feature vector and the annotated text obtain the vector of the internal connection and key features between the display image and the text through the Q-Former, and then capture the relationship between elements through the Transformer to generate an output sequence. The output sequence is formed into an output result using the Fully Connected. According to the generated displayable pixel format and the output result, the multi-dimensional input data is converted into a one-dimensional vector through Flatten and input into the LLMs to generate the multi-modal model.

[0033] Among them, the Image input is the input of the image data, and the image data is sent to the Image Encoder and the Image Dncoder.

[0034] Among them, the Image Encoder (image encoder) is mainly responsible for converting the image data into a representation that is easier for the model to process, that is, converting it into a high-dimensional feature vector, which can be used for subsequent image classification, segmentation and other tasks, and sent to the Q-Former for processing. The Image Dncoder (image decoder) is an interface or tool for converting the encoded image data into a pixel format that can be displayed and processed in the application. Its main function is to decode the image data into the RGBA pixel format for processing.

[0035] Among them, the Text Input is the text input, and the pre-annotation information of the image in the annotation data module is sent to the Q-Former for processing.

[0036] Among them, the Q-Former is a neural network structure used in the multi-modal large model, and its function is to integrate the obtained image vector and the annotated text for cross-modal semantic understanding, and send the integration result to the Transformer.

[0037] Among them, the queries of the Transformer interact with each other through the self-attention layer under the action of the Q-Former, and interact with the frozen image features through the cross-attention layer. At the same time, these queries can also interact with the text through the same self-attention layer. Therefore, it can better process and understand the information from different modalities, generate an output sequence and send it to the Fully Connected.

[0038] Among them, Fully Connected is the fully connected layer. When the output sequence of the Transformer is input into the fully connected layer, each vector will perform a dot product operation with the weight matrix of the fully connected layer respectively, and add a bias term. Then, the operation result is non-linearly transformed through an activation function (such as ReLU, sigmoid, softmax, etc.) to obtain the output of the fully connected layer.

[0039] Among them, the main function of Flatten is to flatten the input multi-dimensional data, that is, the displayable pixel format converted by the Image Dncoder and the output result generated by the Fully Connected, into one-dimensional data. This process does not change the total number of elements in the data, only changes the data shape.

[0040] Finally, the generated one-dimensional data is input into the LLMs to form and output a multi-modal base model.

[0041] Specifically, according to different application scenarios, the PC side uses the API encapsulation port, and the mobile side uses the "mobile model" deployment interface. The Chinese herbal medicine image recognition platform provides tools for converting the model into a mobile side and the deployment code basecase.

[0042] Among them, the deployment of the interface is executed by the model encapsulation and deployment module. The PC side needs to select the corresponding technology stack and programming language, write and encapsulate the API service. The mobile side needs to select the corresponding mobile technology (such as React Native, Flutter, native Android / iOS development, etc.), design the interface according to the requirements of the mobile application, and implement the interface logic to interact with the backend service.

[0043] Specifically, the detection model is to determine the image category and the position of the object in the image; the segmentation model is to detect the objects in the image and generate a pixel-level mask for each detected object to distinguish the boundaries of different objects.

[0044] Among them, the detection model takes an image as input, analyzes and processes the image through a neural network. The detection model can identify the Chinese herbal medicines in the image, assign a category label to each Chinese herbal medicine, and also determine the specific positions of these objects in the image, marked by bounding boxes.

[0045] In addition, the segmentation model not only performs pixel-level classification, but also assigns each pixel to a specific semantic category and a specific object instance. It can distinguish different instances of Chinese herbal medicines so that they are assigned different labels at the pixel level.

[0046] Specifically, the annotation of the image data and the model is to pre-annotate the unannotated data through existing image models, text models, and multi-modal models, and provide a function for manual modification of annotations.

[0047] Among them, model pre-annotation is to annotate the image data and the model according to the existing model, using a pre-trained large model or an automated annotation tool to reduce the workload of manual annotation and improve the annotation efficiency. At the same time, a function for manual modification of annotations is provided to correct the incorrectly annotated Chinese herbal medicine data.

[0048] Specifically, the online code testing is to use the platform code for self-testing or upload the code by oneself for testing.

[0049] Among them, the platform code is the test code pre-installed in the model testing & evaluation module, and the platform code can be directly used for automatic testing. At the same time, the model testing & evaluation module also supports users to upload other test codes to conduct other tests on the generated model.

[0050] The above steps use a device for Chinese herbal medicine image recognition based on a multi-modal model, which is characterized in that the device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the method for Chinese herbal medicine image recognition based on a multi-modal model described in any one of the above.

[0051] The above steps use an electronic device, which is characterized by including a memory and a processor, and a computer program is stored in the memory, and the processor implements the method for Chinese herbal medicine image recognition based on a multi-modal model described in any one of the above when executing the computer program.

[0052] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary technical personnel in the technical field to understand the embodiments disclosed herein.

Claims

1. A method for Chinese herbal medicine image recognition based on a multimodal model, characterized in that: include: According to the image data acquired by the data acquisition module, the data annotation module is used to pre-annotate; Based on the labeled image data, use the model online training module to train the image task model and multimodal task model online, and use the model training & evaluation module to test the model training results and output the model file; According to the output model file, it is converted into a deployable model structure and parameter file through the model packaging and deployment module.

2. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 1, characterized in that: The image task model is divided into a detection model and a segmentation model according to the business and application scenarios, and a model file is generated; The multimodal task model uses a neural network to classify and recognize image data, then inputs LLMs and outputs a model file; The model training & evaluation module tests the basetestcase of the image model and the multimodal model online according to the code and generates the model file.

3. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 1, characterized in that: The multimodal task model uses a neural network to classify and recognize images, then inputs LLMs and outputs a model file, including: The multimodal task model obtains a feature vector of the image data through an Image Encoder, and converts the image data into a displayable pixel format through an Image Dncoder; The feature vector and the annotated text are used to obtain the vector of the intrinsic connection and key features between the display image and the text through Q-Former, and the relationship between the elements is captured through Transformer to generate an output sequence, and the output sequence is formed into an output result using FullyConnected; According to the generated displayable pixel format and output results, the multi-dimensional input data is converted into a one-dimensional vector through Flatten and input into LLMs to generate a multimodal model.

4. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 1, characterized in that: Depending on the application scenario, the PC side uses the API encapsulation port, and the mobile side uses the "mobile side model" deployment interface; The Chinese herbal medicine image recognition platform provides a basecase for model conversion into mobile tools and deployment code.

5. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 2, characterized in that: The image task model is divided into a detection model and a segmentation model according to the business and application scenarios and generates a model file, including: The detection model is used to determine the image category and the location of the object in the image; The segmentation model detects objects in an image and generates a pixel-level mask for each detected object to distinguish the boundaries of different objects.

6. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 1, characterized in that: The image data acquired by the data acquisition module is pre-annotated using the annotation data module, including: The image data and model are labeled by pre-labeling the unlabeled data through the existing image model, text model, and multimodal model, and providing a manual modification labeling function.

7. The method for Chinese herbal medicine image recognition based on a multimodal model as claimed in claim 2, characterized in that: The model training & evaluation module tests the base testcase of the image model and the multimodal model online according to the code and generates the model file, including: The code online test is to use the platform code for self-testing or upload the code by yourself for testing.

8. A device for Chinese herbal medicine image recognition based on a multimodal model, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the method for Chinese herbal medicine image recognition based on a multimodal model as described in any one of claims 1 to 7.

9. The device for Chinese herbal medicine image recognition based on multimodal model according to claim 8, characterized in that The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for Chinese herbal medicine image recognition based on a multimodal model as described in any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A memory and a processor, wherein a computer program is stored in the memory, wherein the processor implements the method for Chinese herbal medicine image recognition based on a multimodal model as described in any one of claims 1 to 7 when executing the computer program.