Spore-pollen image recognition method and system based on multi-modal dialogue large model

By constructing a pollen image recognition system based on a multimodal dialogue model, the problems of traditional pollen identification relying on expert experience and the lack of interactive capabilities in single-modal models are solved. This system achieves efficient and accurate automatic pollen identification and analysis, and possesses professional interactive capabilities.

CN122415453APending Publication Date: 2026-07-17HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HENAN UNIVERSITY
Filing Date
2026-03-27
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional pollen identification methods rely heavily on expert experience, resulting in low efficiency. Furthermore, existing single-modal deep learning models struggle to integrate multi-source information and lack interactive capabilities, leading to insufficient accuracy and efficiency in pollen identification.

Method used

A pollen image recognition method based on a multimodal dialogue model is adopted. Through a multi-stage image enhancement process, a hierarchical label smoothing strategy, a pollen recognition rule base, and domain-specific prompt word engineering, combined with efficient parameter fine-tuning technology, a GLM4v multimodal dialogue model is constructed to achieve automatic recognition and analysis of pollen images.

Benefits of technology

It significantly improves the accuracy and efficiency of pollen identification, especially the fine-grained classification accuracy of families, genera, and species, and has natural language interaction capabilities, providing professional morphological identification basis and reasoning process, thus enhancing the interpretability and practicality of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415453A_ABST
    Figure CN122415453A_ABST
Patent Text Reader

Abstract

本发明属于图像识别与人工智能技术领域,公开一种基于多模态对话大模型的孢粉图像识别方法与系统,该方法包括:获取孢粉图像并进行数据增强处理;对数据增强处理后的孢粉图像进行标注,标注内容包括软标签;构建孢粉识别规则库,用于存储和管理孢粉形态学知识、分类规则、特征匹配模板及历史识别案例;利用标注后的孢粉多模态数据和构建的孢粉识别规则库,基于多模态对话大模型构建孢粉图像识别模型;通过构建的孢粉图像识别模型对孢粉图像进行识别。本发明通过将通用多模态大模型与孢粉学专业领域知识深度融合,建立了层次标签平滑策略和领域特化的提示词工程框架,结合参数高效微调技术,实现了基于多模态特征深度融合的孢粉智能识别。
Need to check novelty before this filing date? Find Prior Art