A multi-modal large model-based semantic enhancement insect fine-grained identification method

CN122531068BActive Publication Date: 2026-09-18SICHUAN AGRI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611016020.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-18
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

该类方法能够利用文本语义,但通常没有针对昆虫分类层级、易混类别和差异属性建立稳定约束,容易受到类别名称、提示词和背景噪声影响

Benefits of technology

[0056] 1. By using category semantic knowledge units, insect classification levels, morphological attributes, and easily confused category difference attributes are uniformly organized, enabling multimodal large models to utilize taxonomic prior knowledge in insect identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531068B_ABST
    Figure CN122531068B_ABST
Patent Text Reader

Abstract

The application relates to a multi-modal large model-based semantic enhancement insect fine-grained identification method, and belongs to the image recognition field.The method comprises the following steps: constructing an insect image sample set and a category semantic knowledge unit, and generating multi-text type category semantic enhancement descriptions; constructing a category multi-modal prototype library, and generating a candidate category set; reading a difference attribute set, and extracting local visual evidence; organizing multi-modal large model identification input, and calculating initial identification scores, classification level consistency scores, difference enhancement scores and evidence support degree scores; performing score effectiveness judgment, normalization fusion and long-tail reweighting on each score to obtain final category scores; calculating an identification reliability score according to the candidate category with the highest final category score to obtain a final result including multi-class data. The application enhances the explainability, traceability and reviewability of the identification result by outputting key morphological attributes, difference attributes, local visual evidence, reference evidence and unverified attributes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition, and more particularly to a semantically enhanced fine-grained insect recognition method based on a multimodal large model. Background Technology

[0002] Insect identification is a fundamental aspect of agricultural, forestry, and ecological surveys. With the development of image acquisition equipment, mobile terminals, field monitoring equipment, and publicly available insect datasets, image-based automatic insect identification technology has been extensively studied. Existing methods typically employ convolutional neural networks, visual Transformers, visual state-space models, or visual fundamental models to extract insect image features before classifying them.

[0003] In coarse-grained recognition scenarios, the aforementioned methods can achieve certain results. However, in fine-grained insect recognition, different insect species may only differ in local areas such as wing veins, markings, antennae, abdominal color, and leg structure. Even within the same category, significant variations can occur due to differences in posture, life cycle, lighting, background, shooting angle, and image quality. Therefore, classification methods relying solely on overall image features are prone to confusion between similar categories. In recent years, multimodal large-scale models have developed the ability to jointly understand images and text, simultaneously processing insect images, category names, Latin names, classification levels, and morphological descriptions. This provides a foundation for introducing taxonomic and morphological knowledge into insect recognition. However, when general-purpose multimodal large-scale models are directly applied to fine-grained insect recognition, they still tend to remain at the level of overall image-text matching or overall image description, lacking local visual evidence to verify the key differences in easily confused categories.

[0004] Currently, existing technologies mainly fall into three categories: The first category comprises insect recognition methods based on pure visual deep learning models. These methods use convolutional neural networks, visual Transformers, or basic visual models to extract image features and output insect categories through classifiers. These methods rely heavily on sample size and image feature quality, and are sensitive to long-tailed categories, local fine-grained differences, and cross-domain image variations. The second category consists of open category recognition methods based on image-text models or multimodal models. These methods use category names or descriptions as textual prompts and perform image-text similarity matching with insect images. While these methods can utilize textual semantics, they typically lack stable constraints for insect classification levels, easily confused categories, and differential attributes, making them susceptible to the influence of category names, prompt words, and background noise. The third category is recognition methods assisted by knowledge bases or knowledge graphs. These methods can introduce category attributes or higher-level classification relationships, but most methods merely treat knowledge as global text or post-processing constraints, lacking a process to transform differential attributes into local visual evidence verification that is locatable, answerable, and scoreable.

[0005] Therefore, existing technologies have at least the following technical problems: 1. Confusion easily occurs between similar insect categories. Different insect categories may be highly similar in overall appearance, with key differences often concentrated in local areas such as forewings, hindwings, antennae, abdomen, legs, wing veins, or markings. Existing purely visual classification methods may not be able to reliably focus on these fine-grained areas. 2. Insufficient utilization of hierarchical knowledge in insect taxonomy. Insect categories naturally have hierarchical relationships such as order, family, genus, and species, but existing identification methods usually only output the final category label, without fully utilizing the upper-level classification path to constrain candidate categories, easily leading to cross-level misjudgments. 3. Lack of evidence loop for easily confused category difference attributes. Even when attribute text is introduced, existing methods mostly perform overall text matching, failing to transform the difference attributes into attribute-level discrimination problems, and failing to establish matching relationships with corresponding local visual parts. 4. Long-tailed categories and target domain offset affect recognition stability. Publicly available insect datasets often have an imbalance in the number of category samples, and field monitoring images may differ from training images in background, lighting, scale, and sharpness, causing the model to favor categories with more samples or similar imaging conditions. 5. The interpretability and verifiability of the identification results are insufficient. Existing systems usually only provide the category name and confidence level, lacking key discriminant attributes, local evidence, unverified attributes, and low-confidence verification markers, which makes it difficult to meet the requirements for traceable results in agricultural, forestry, and ecological surveys. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a semantically enhanced fine-grained insect recognition method based on a multimodal large model, thus solving the deficiencies of the prior art.

[0007] The objective of this invention is achieved through the following technical solution: a semantically enhanced fine-grained insect recognition method based on a multimodal large model, the method comprising:

[0008] S1. Construct an insect image sample set and category semantic knowledge units, and generate multi-text type category semantic enhanced descriptions;

[0009] S2. Construct a multi-modal prototype library of categories and generate a set of candidate categories based on the images of insects to be identified;

[0010] S3. For the candidate categories and their easily confused categories in the candidate category set, read the set of differential attributes, and extract local visual evidence from the insect image to be identified based on the observable visual parts corresponding to each differential attribute.

[0011] S4. Organize a multimodal large model to identify inputs and calculate initial identification scores, then calculate classification hierarchy consistency scores, difference enhancement scores, and evidence support scores;

[0012] S5. The classification hierarchy consistency score, difference enhancement score, and evidence support score are evaluated for score validity, normalized and fused, and long-tail reweighted to obtain the final category score.

[0013] S6. Calculate the recognition reliability score based on the candidate category with the highest final category score to obtain the final result including data from multiple categories.

[0014] The construction of the insect image sample set and category semantic knowledge unit includes:

[0015] Obtain insect image samples and their category labels from publicly available insect datasets, images collected by monitoring equipment, or images annotated by experts. Perform invalid image removal, label unification, duplicate sample removal, and low-quality sample labeling on the samples. Then divide the samples into training set, validation set, and test set according to a preset ratio.

[0016] For each insect category, a category semantic knowledge unit is constructed. The category semantic knowledge unit includes at least the category label, category name, Latin scientific name, classification hierarchy information, morphological attribute description, easily confused category set, differential attribute set, observable visual parts, attribute source confidence, and attribute matching confidence threshold.

[0017] The generated multi-text type category semantic enhancement description includes:

[0018] The category semantic knowledge units are converted into text representations used by the multimodal large model. Five types of text are generated for each category, including classification hierarchical path text, scientific name text, Chinese name text, morphological attribute text, and easily confused category difference attribute text.

[0019] The five types of text are encoded separately and then weighted and fused to obtain category semantic vectors. When a category lacks a difference attribute field, the difference attribute text is set to empty text or back text, and the weight corresponding to the difference attribute text is set to 0. Then the remaining weights are normalized again.

[0020] The constructed category of multimodal prototype library includes:

[0021] A category multimodal prototype library is constructed using category visual prototypes, category semantic prototypes, category difference attribute prototypes, and category hierarchical path prototypes.

[0022] For category visual prototypes, visual features of each sample image under the same category are extracted using a visual encoding function. Then, the visual features are clustered, and representative samples whose distance from the cluster centroid is within a set range are selected. The category visual prototype is obtained by averaging the features of the representative samples.

[0023] The category semantic prototype is composed of the obtained category semantic vector;

[0024] The prototype of the category difference attribute is obtained by text encoding the difference attribute.

[0025] The category-level path prototype is obtained from the category-level path text encoding or one-hot encoding vector.

[0026] The step of generating a candidate category set based on the image of the insect to be identified includes:

[0027] After receiving the image of the insect to be identified, the system generates the visual features, image semantic vector, and upper-level classification prediction path of the image of the insect to be identified. Then, it generates the visual prototype recall category set, the semantic prototype recall category set, the classification level path recall category set, and the easily confused category expansion set, respectively. The system then takes the union of these sets to obtain the candidate category set.

[0028] The upper-level classification prediction path is obtained by a multimodal large model or an independent auxiliary classifier based on the image of the insect to be identified. It includes at least one of the three levels: order, family, and genus. For any predicted level among order, family, and genus, if the level contains multiple possible results, the highest-scoring level is retained. Each result forms a corresponding upper-level classification candidate result. Then, the seed-level categories that match the above upper-level classification candidate results form a classification-level path recall category set.

[0029] If the union of the visual prototype recall category set, the semantic prototype recall category set, the classification hierarchy path recall category set, and the easily confused category extension set is empty, then the entire category set will be used as the candidate category set. If the number of candidate categories exceeds the preset upper limit, then the top K candidate categories will be retained after being sorted by at least two of the following: visual similarity, semantic similarity, and hierarchy path consistency.

[0030] S3 specifically includes:

[0031] Read candidate categories set of difference attributes and determine each difference attribute The contrast categories, directions of difference, and observable visual parts ;

[0032] Images of insects to be identified Coarse localization of insect targets is performed to obtain the insect target region. If the insect target region does not exist, or the confidence of the insect target region is lower than the preset target threshold, the entire image of the insect to be identified is taken as the target region, and a label indicating insufficient image quality is output.

[0033] Locate the observable visual parts corresponding to the differences in attributes within the insect target area;

[0034] Based on observable visual parts The localization result is used to crop a local region of the image, and this local region is recorded as local visual evidence. If the same difference attribute corresponds to multiple observable visual parts, then multiple local region images are cropped separately, and the local region image with the highest location confidence or the highest matching score is retained.

[0035] For each difference attribute Corresponding local visual evidence Visibility is assessed if the local area simultaneously meets the following conditions: the location confidence level is not lower than a preset location threshold, the area of ​​the local area is not less than a preset proportion of the insect target area, the effective discriminable area proportion of the local area is not lower than a preset visibility proportion threshold, and the clarity index of the local area is not lower than a preset clarity threshold; then the visibility is marked. Set as If any of the above conditions are not met, then visibility will be marked. Set as And the difference attributes corresponding to this local visual evidence. The property is marked as unusable. The effective discriminable area ratio is the ratio of the area of ​​the local region that is not occluded and can present texture, edge or color features to the total area of ​​the local region. The sharpness index is represented by the Laplacian variance of the local region. When the Laplacian variance is lower than the preset sharpness threshold, the local region is judged to be blurry.

[0036] The organization's multimodal large model identification input and calculation of the initial identification score include:

[0037] The images of the insects to be identified, local visual evidence, candidate category sets, category semantic enhancement descriptions, classification hierarchy information, and attribute-level discrimination problems are organized into a multimodal large model recognition input;

[0038] The initial recognition score is obtained through either a comparative image-text matching method or a generative image-text matching method.

[0039] The calculation process for the classification hierarchy consistency score includes:

[0040] The consistency between the candidate category and the upper-level classification hierarchy prediction path is calculated to obtain the classification hierarchy consistency score. ,in, Indicates candidate category The hierarchical consistency score, , , Representing candidate categories Matching results with the upper-level predicted paths at the order, family, and genus levels. , , They are respectively , , The weights;

[0041] The calculation process for the difference enhancement score includes:

[0042] For candidate categories Read its set of differences and for each difference attribute Local visual evidence Match the attribute descriptions to obtain a matching score. Attribute source confidence Attribute matching confidence and attribute visibility marker The difference enhancement score was obtained. ,when If candidate category c does not have a valid difference attribute, the difference enhancement score will be increased. Items marked as invalid will not have their difference enhancement enabled during score fusion, and the fusion weights for other valid items will be renormalized. Indicates the stabilizing factor;

[0043] The calculation process for the evidence support score includes:

[0044] Based on a multimodal prototype library, the consistency between the image to be identified and the candidate category prototypes is measured to obtain the evidence support score. ,in, Indicates candidate category Evidence support score Represents the visual features of the image to be identified. This represents the image semantic vector corresponding to the image to be identified. This represents the difference attribute evidence vector corresponding to the image to be identified. This represents the prediction vector from the higher-level classification hierarchy of the image to be identified. , , , Representing categories The category visual prototype, category semantic prototype, category difference attribute prototype, and category hierarchy path prototype. , , , They are respectively , , and The weight, This represents the normalized similarity function;

[0045] If a prototype is missing, the weight corresponding to that prototype is reset to 0, and the weights of the remaining prototypes are renormalized. If all prototypes are missing, then... Marked as invalid score item.

[0046] S5 includes:

[0047] For any scoring item In the candidate category set Internal calculation of maximum value and minimum value ,like The normalization result is expressed as ,like This indicates that the score item does not have a distinguishing effect within the current candidate set, and is therefore marked as an invalid score item. (For score items that enhance difference...) If no valid visible difference attribute exists for any of the candidate categories, the score item will also be marked as invalid.

[0048] Let the set of valid scores be denoted as ,when When not empty, the basic fusion score is represented as , Indicates the scoring item The fusion weights are determined such that when a score item is invalid, its weight is not included in the fusion, and the weights of other valid score items are renormalized according to their original proportions. When empty, use the initial recognition score within the candidate category set. As a base fusion score, if the initial recognition score also lacks discriminatory power, then... Set to uniform distribution and trigger low-confidence check flags;

[0049] The basic fusion score is reweighted based on the frequency of samples by category to obtain the reweighting coefficient. The final category score is then obtained as follows: ,when When the value is 0, the final category score of each candidate category in the candidate category set is set to 0. Alternatively, it can revert to the normalized result of the initial recognition score, where, Indicate category The corresponding long-tail weighted coefficients, Indicate category The basic fusion score.

[0050] S6 includes:

[0051] Based on the candidate category with the highest final category score The reliability score for identification is expressed as follows: ,in, This represents the score interval between the highest-scoring candidate category and the second-highest-scoring candidate category. express The effective visibility ratio of the corresponding difference attribute and These represent the normalized evidence support score and the hierarchical consistency score, respectively. , , , They are respectively , , and The reliability score weighting;

[0052] when Less than the preset review threshold When low confidence is reached, output a low confidence check flag and list the reasons for the low confidence.

[0053] The candidate category with the highest final category score As the final predicted category, the corresponding classification hierarchy path is output, based on the recognition reliability score. Evidence support score , effective visibility ratio of difference attributes The number of unverified attributes and the reasons for low confidence are used to generate an evidence sufficiency status. The evidence sufficiency status includes at least two types: sufficient evidence and insufficient evidence. When the identification reliability score is not lower than the preset review threshold, and the key difference attributes are visible and the evidence support meets the preset requirements, the evidence sufficiency status is sufficient evidence. When there are key difference attributes that are not visible, insufficient evidence support, classification hierarchical path conflicts, or low confidence review markers, the evidence sufficiency status is insufficient evidence.

[0054] The final output includes the final predicted category, classification hierarchy path, final category score, recognition reliability score, evidence sufficiency status, low confidence verification marker, key morphological attributes, easily confused category difference attributes, attribute-level discrimination problem, local visual evidence, reference evidence, and unverified attributes.

[0055] The present invention has the following advantages:

[0056] 1. By using category semantic knowledge units, insect classification levels, morphological attributes, and easily confused category difference attributes are uniformly organized, enabling multimodal large models to utilize taxonomic prior knowledge in insect identification.

[0057] 2. Enhance the semantic description through multiple text type categories, and reduce the semantic bias caused by a single category name or a single descriptive text by utilizing classification path, Latin scientific name, Chinese name, morphological attribute and difference attribute respectively.

[0058] 3. By generating a candidate category set through visual prototype recall, semantic prototype recall, classification hierarchical path recall, and expansion and merging of easily confused categories, the risk of the correct category being excluded in the early candidate screening is reduced.

[0059] 4. By converting easily confused category differences into attribute-level discrimination problems and extracting corresponding local visual evidence, the recognition process can focus on fine-grained discrimination areas such as forewings, antennae, abdomen, and wing veins, thereby improving the ability to distinguish between similar categories.

[0060] 5. By using classification hierarchy consistency scores, cross-level misjudgments that conflict with the upper-level prediction paths are reduced.

[0061] 6. By constraining the visibility of differential attributes, the confidence of attribute sources, and the confidence of attribute matching, we can avoid attributes that are not visible, have insufficient source reliability, or have low model judgment confidence from having an excessive impact on the final recognition results.

[0062] 7. By judging the validity of scores, using min-max normalization, and redistributing weights, inconsistencies in the dimensions of different scoring items and problems with the denominator are avoided. The issue of invalid evidence being forcibly incorporated into the fusion process.

[0063] 8. By using a long-tail reweighting mechanism, the classification bias caused by a large number of samples in the head category is mitigated, and the chance of the tail category being correctly identified is increased.

[0064] 9. Through the low-confidence review mechanism, the system can output a review mark when the scores are close, key parts are not visible, evidence is insufficient, or there is a hierarchical conflict, reducing the risk of incorrect identification results being used directly.

[0065] 10. Enhance the interpretability, traceability, and verifiability of identification results by outputting key morphological attributes, difference attributes, local visual evidence, reference evidence, and unverified attributes. Attached Figure Description

[0066] Figure 1 This is a schematic diagram of the process of the present invention.

[0067] Figure 2 A schematic diagram illustrating the construction of category semantic knowledge units and category multimodal prototype libraries.

[0068] Figure 3 This diagram illustrates candidate recall, differential attribute matching, score fusion, and result output. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application provided below with reference to the accompanying drawings is not intended to limit the scope of protection of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. The present invention will be further described below with reference to the accompanying drawings.

[0070] This invention relates to a semantically enhanced fine-grained insect recognition method based on a multimodal large model. Instead of directly classifying insect images as a whole using the multimodal large model, this method first constructs category semantic knowledge units containing classification hierarchy, morphological attributes, and easily confused category difference attributes. Then, it generates multi-text type category semantically enhanced descriptions and a category multimodal prototype library. During recognition, a candidate category set is formed through visual prototype recall, semantic prototype recall, classification hierarchy path recall, and easily confused category expansion. Subsequently, the easily confused category difference attributes are converted into attribute-level discrimination problems, and corresponding local visual evidence is extracted. Finally, the initial recognition score, hierarchy consistency score, difference enhancement score, and evidence support score are evaluated for validity, unified normalization, weighted fusion, long-tail reweighting, and low-confidence verification. This advances the recognition process from overall image-text matching to a fine-grained recognition process with classification constraints, local evidence, and verification mechanisms.

[0071] like Figure 1 As shown, it specifically includes the following:

[0072] Step 1, as follows Figure 2 As shown, an insect image sample set and category semantic knowledge units are constructed;

[0073] The system acquires insect image samples and their category labels from publicly available insect datasets, images collected by monitoring equipment, or images annotated by experts. It then removes invalid images, unifies labels, removes duplicate samples, and marks low-quality samples. Finally, it divides the samples into training, validation, and test sets according to a preset ratio. Low-quality samples refer to images that, although containing the target insect and with generally usable category labels, suffer from issues such as blurriness, overexposure, underexposure, low resolution, small target size, severe occlusion, invisible key morphological parts, strong background interference, or insufficient labeling credibility. These issues make it difficult to stably observe key morphological attributes such as color, texture, wing shape, antennae morphology, pattern distribution, and wing vein characteristics.

[0074] Subsequently, a category semantic knowledge unit is constructed for each insect category. The category semantic knowledge unit can be stored in a table, key-value pair, structured text, or JSON structure, and should include at least the category label, category name, Latin scientific name, taxonomic hierarchy information, morphological attribute description, set of easily confused categories, set of differential attributes, observable visual parts, attribute source confidence, and attribute matching confidence threshold.

[0075] The classification hierarchy information includes at least one of the following: order, family, genus, and species; the morphological attribute description includes at least two of the following: color, texture, wing shape, antennae morphology, body length, pattern distribution, wing vein characteristics, leg morphology, abdominal morphology, and life cycle stage; the differential attribute description is organized in easily confused category pairs, and each differential attribute includes at least four of the following: differential attribute name, candidate category, comparison category, direction of difference, observable visual part, observability level, attribute-level discrimination question template, positive evidence description, negative evidence description, attribute source confidence, and attribute matching confidence threshold.

[0076] When some fields are missing in the public dataset, they can be supplemented by domain expert annotation, taxonomic literature extraction, or public entomological literature retrieval; fields that still cannot be obtained are marked with unknown placeholders and a fallback strategy is adopted in subsequent template generation or score fusion.

[0077] Step 2: Generate semantically enhanced descriptions for multiple text type categories;

[0078] This step converts the category semantic knowledge unit into a text representation that can be used by the multimodal large model. Five types of text are generated for each category, including: (1) Classification hierarchical path text: describing the order, family, genus and species path of the category; (2) Scientific name text: describing the Latin scientific name of the category; (3) Chinese name text: describing the Chinese name or common name of the category; (4) Morphological attribute text: describing the observable morphological features of the category such as color, texture, wing shape, antennae, body length, spots, and wings; (5) Difference attribute text of easily confused categories: describing the key differences, observable parts and directions of differences between the category and easily confused categories.

[0079] The above multiple text types are encoded separately and then weighted and fused to obtain the category semantic vector. ,in, Indicate category Category semantic vector, , , , , Representing categories The classification hierarchy path text, scientific name text, Chinese name text, morphological attribute text, and difference attribute text. This represents a text encoding function. , , , , They represent , , , , The fusion weights are all not less than 0 and satisfy the following conditions: .

[0080] When a category lacks a difference attribute field, Set to empty text or backtext, and Set the weights to 0, and then normalize the remaining weights again.

[0081] Step 3: Build a category-based multimodal prototype library;

[0082] This step provides a traceable reference for candidate category recall, evidence support assessment, and result interpretation. The category multimodal prototype library includes at least a category visual prototype, a category semantic prototype, a category difference attribute prototype, and a category hierarchical path prototype.

[0083] For category visual prototypes, the system first uses the visual encoding function. Visual features are extracted from sample images within the same category. These visual features are then clustered, and representative samples closest to the cluster centroids are selected. Finally, the average of the representative sample features is used to obtain the category visual prototype. ,in, Indicate category Category visual prototype, Indicate category The representative sample set, This indicates the number of samples. This represents a sample image.

[0084] If category No image samples are available, or Then Marked as missing, this item is not enabled in scoring items involving visual prototypes.

[0085] Category semantic prototype The category semantic vector obtained from step 2 Composition, namely:

[0086]

[0087] Category Difference Attribute Prototype Obtained by text encoding of the difference attribute, i.e., category Difference attribute text Input text encoding function ,get:

[0088]

[0089] Category Hierarchy Path Prototype By category Classification hierarchy path Specifically, the system can assemble hierarchical path text by arranging the names of order, family, genus, etc., in a fixed order, and input it into the text encoding function. This yields the hierarchical path encoding:

[0090]

[0091] Alternatively, the system can pre-build a fixed dictionary for all possible nodes at each level and categorize them. Set to the position of the node belonging to the corresponding level. The remaining positions are set to Then, these are concatenated according to the hierarchical order of order, family, and genus to obtain the hierarchical path one-hot vector. Here, one-hot refers to a "one-hot encoding" method, meaning that in a fixed-length vector, only values ​​taken from the specified values ​​are used. The position indicates the hierarchy node to which this category belongs; other positions take values... When information at a certain level is missing, the corresponding vector at that level remains intact. Alternatively, a category hierarchy path prototype can be built using only known levels.

[0092] For category-based multimodal prototype libraries, reference image numbers, attribute entry sources, expert annotation records, taxonomic literature sources, DNA barcodes, or molecular identification records can also be saved. These reference evidences are not required to be fully entered during the reasoning stage, but are instead used as optional reference items for category prototype or evidence support calculation.

[0093] Step 4, as follows Figure 3 As shown, a set of candidate categories is generated based on the image of the insect to be identified;

[0094] This step corresponds to the reasoning stage, where the system receives the image of the insect to be identified. Then, the visual features, semantic vectors, and upper-level classification path prediction of the insect image to be identified are generated first. Then, the visual prototype recall category set, semantic prototype recall category set, classification path recall category set, and easily confused category extension set are generated respectively. The union of the above sets is used to obtain the candidate category set.

[0095] The upper-level classification prediction path is obtained by a multimodal large model or an independent auxiliary classifier based on the image of the insect to be identified, and includes at least one of the three levels: order, family, and genus. This prediction result does not include the final species-level category to be determined, to avoid cyclically correcting the final score using the final category itself. For any predicted level among order, family, and genus, if that level contains multiple possible results, the highest-scoring result for that level is retained. Each result forms a corresponding upper-level classification candidate result; then, the species-level categories that match the above upper-level classification candidate results form a classification-level path recall category set.

[0096] The candidate category set is represented as ,in, This represents the set of categories obtained from visual prototype recall, i.e., the set of categories obtained based on the similarity between the visual features of the image to be identified and the visual prototypes of the categories. One category; This represents the set of categories obtained from semantic prototype recall, i.e., the top categories obtained based on the similarity between the semantic description of the image to be identified or the semantic vector of the image and the semantic prototype of the category. One category; This represents the set of categories retrieved by the classification hierarchy path, that is, the set of categories that match the predicted path of the upper classification hierarchy. This represents an extended set of easily confused categories.

[0097] The category set obtained from visual prototype recall is defined as follows: The set of categories obtained by semantic prototype recall is defined as follows: The extended set of easily confused categories is defined as follows: ,in, Indicates the image to be recognized Visual features; This represents the image semantic vector corresponding to the image to be identified; Indicate category The corresponding set of easily confused categories; This represents the normalized similarity function.

[0098] like , , and If the union of all categories is empty, then the set of all categories will be empty. This serves as a set of candidate categories. If the number of candidate categories exceeds a preset limit, the categories are sorted based on at least two of the following criteria: visual similarity, semantic similarity, and hierarchical path consistency, and the top categories are retained. There are 10 candidate categories.

[0099] Step 5: Generate attribute-level discrimination questions and extract local visual evidence;

[0100] This step involves reading the set of differential attributes for candidate categories and their easily confused categories within the candidate category set. And based on each difference attribute Corresponding observable visual parts Extracting local visual evidence from images of insects to be identified Local visual evidence refers to a local region image obtained by locating or cropping from an image of an insect to be identified; the vector encoded from this local region image is called the local evidence feature.

[0101] Furthermore, local visual evidence is generated through the following sub-steps:

[0102] Step 501: Read candidate classes set of difference attributes and determine each difference attribute The contrast categories, directions of difference, and observable visual parts .

[0103] Step 502: Image of the insect to be identified Coarse localization of the insect target is performed to obtain the insect target region. Coarse target localization can be achieved using target detection models, saliency map localization, instance segmentation models, or multimodal large model region localization. If the insect target region does not exist, or the confidence level of the insect target region is lower than a preset target threshold, the entire image of the insect to be identified is taken as the target region, and a label indicating insufficient image quality is output.

[0104] Step 503: Locate the observable visual parts corresponding to the differential attributes within the insect target area. For the forewings, hindwings, antennae, abdomen, legs, wing veins, and markings, a part detection model, semantic segmentation model, keypoint localization model, attention heatmap, or a preset insect part template can be used for localization. When using a preset insect part template, the system first establishes a normalized coordinate system based on the major axis, minor axis, and head-tail direction of the insect target area, and then crops the corresponding areas according to the relative positions of different parts in the normalized coordinate system.

[0105] Step 504: Based on observable visual areas The localization result is used to crop a local region of the image, and this local region is recorded as local visual evidence. If the same difference attribute corresponds to multiple observable visual parts, then multiple local region images are cropped separately, and the local region image with the highest location confidence or the highest matching score is retained. Alternatively, a weighted average of the matching scores of multiple local region images can be taken.

[0106] Step 505, for each difference attribute Corresponding local visual evidence Visibility assessment is performed. If the local area simultaneously meets the following conditions: the location confidence level is not lower than a preset location threshold, the area of ​​the local area is not less than a preset proportion of the insect target area, the effective discriminable area proportion of the local area is not lower than a preset visibility proportion threshold, and the clarity index of the local area is not lower than a preset clarity threshold, then the visibility is marked. Set as If any of the above conditions are not met, then visibility will be marked. Set as And the difference attributes corresponding to this local visual evidence. This is marked as an unavailable attribute.

[0107] The effective discriminable area ratio is the ratio of the area of ​​a local region that is not occluded and can present texture, edge or color features to the total area of ​​that local region; the sharpness index is represented by the Laplacian variance of that local region, and when the Laplacian variance is lower than the preset sharpness threshold, the local region is judged to be blurry.

[0108] Meanwhile, the system converts the difference attributes into attribute-level discrimination problems, such as: (1) Does the middle of the forewing in the image have a clear and continuous kidney-shaped spot? (2) Does the middle of the forewing in the image have a clear and continuous kidney-shaped spot? (3) Is the color of the base of the abdomen in the image significantly darker than the rest of the abdomen?

[0109] When the corresponding visual part is not visible, the difference attribute is marked as an unavailable attribute, and its influence is reduced or canceled during score fusion.

[0110] Step 6: Organize the multimodal large model recognition input and calculate the initial recognition score;

[0111] This step organizes the insect image to be identified, local visual evidence, candidate category set, category semantic enhancement description, classification hierarchy information, and attribute-level discrimination problem into a multimodal large model recognition input.

[0112] The visual input includes images of the insects to be identified. And several pieces of local visual evidence .

[0113] The text input includes semantically enhanced descriptions of candidate categories, classification paths, morphological attributes, differential attributes, attribute-level discrimination questions, and output format constraints.

[0114] The recognition prompt template can be described as follows: The insect in the image belongs to one of the following candidate categories. Please first determine whether the key morphological attributes of each candidate category are visible, then answer the attribute-level discrimination questions one by one, and output the category matching score, difference attribute matching score, attribute visibility, attribute matching confidence and discrimination criteria for each candidate category.

[0115] Initial recognition score It can be obtained through comparative image-text matching or generative image-text matching.

[0116] Furthermore, the comparative image-text matching method is as follows:

[0117] ,

[0118]

[0119] in, Indicates candidate category Unnormalized score for contrastive image-text matching; Represents the set of candidate categories Any one-time traversal category The unnormalized score of the comparative image-text matching is used for normalized denominator calculation. Indicate category semantic enhancement description, This represents the cross-modal matching temperature coefficient, and .

[0120] Generative image-text matching method is as follows:

[0121] ,

[0122]

[0123] in, Indicates candidate category Unnormalized score for generative image-text matching; Represents the set of candidate categories Any one-time traversal category The unnormalized score of generative image-text matching is used for normalized denominator calculation; Indicate category Name or category number in A token at each position, where a token represents the smallest text unit used by the multimodal large model when processing text, and can be a character, word, subword or symbol, now translated as word element; Indicate category The name or category number located at the first The token sequence preceding each position Indicate category Name or category number in Tokens for each location; The length of the token indicating the category name or category number; This represents the conditional probability output of a large multimodal model.

[0124] It should be noted that this invention does not limit the specific multimodal large model structure; any multimodal large model that supports image and text input, image and text matching, region question answering, or candidate text scoring can be used as an implementation method.

[0125] Step 7: Calculate the classification hierarchy consistency score;

[0126] This step is used to reduce cross-level misclassification. The system uses the upper-level classification prediction path obtained in step 4 to calculate the consistency between the candidate category and the upper-level classification prediction path.

[0127] Classification hierarchy consistency score is represented as ,in, Indicates candidate category The hierarchical consistency score, , , Representing candidate categories The matching result with the upper-level predicted path at the three levels of order, family, and genus is set to 1 when there is a match and 0 when there is no match. , , They are respectively , , The weights are all not less than 0 and satisfy the following conditions: .

[0128] The closer a classification level is to the species level, the higher its weight. Therefore, it is set... For example, it is advisable , , If a prediction for a certain level is missing, the corresponding weight will be reset to 0, and the weights of the remaining levels will be normalized again.

[0129] Step 8: Calculate the difference enhancement score;

[0130] This step is used to establish a matching relationship between easily confused category difference attributes and local visual evidence. For candidate categories The system reads its set of differential attributes. and for each difference attribute Local visual evidence The system matches the attribute descriptions to obtain the matching score, attribute source confidence, attribute matching confidence, and attribute visibility flag.

[0131] The difference enhancement score is represented as ,in, Indicating local visual evidence and differential attributes The normalized matching score between them, with a value range of 100. ; Indicates the difference attribute The confidence level of the attribute source, with a value range of . ; Indicates the difference attribute The attribute matching confidence score, with a value range of 100%. ; Indicates the difference attribute The visibility marker, taken when visible. It cannot be taken when it is not visible. , It is a stability factor greater than 0, used to avoid the denominator being 0.

[0132] when or candidate categories When no valid difference attribute exists, increase the difference enhancement score. Mark as invalid score items, disable difference enhancement during score fusion, and renormalize the fusion weights of other valid score items.

[0133] Attribute-level discrimination results may include: candidate category, comparison category, difference attribute name, observable visual part, attribute-level discrimination question, answer result, matching score, visibility marker, and attribute matching confidence.

[0134] Step 9: Calculate the evidence support score;

[0135] This step, based on a category-based multimodal prototype library, measures the consistency between the image to be identified and the candidate category prototypes. The evidence support score is represented as... ,in, Indicates candidate category Evidence support score; Represents the visual features of the image to be identified; This represents the image semantic vector corresponding to the image to be identified; This represents the differential attribute evidence vector corresponding to the image to be identified; This represents the prediction vector of the upper-level classification hierarchy of the image to be identified; , , , Representing categories Category visual prototype, category semantic prototype, category difference attribute prototype, and category hierarchy path prototype; , , , They are respectively , , and The weights are all not less than 0 and satisfy the following conditions: .

[0136] If a certain prototype is missing, for example, if there are not enough image samples for the tail category, it will cause... If a prototype is missing, its corresponding weight is reset to 0, and the weights of the remaining available prototypes are renormalized. If all reference prototypes are missing, then... Marked as invalid score item.

[0137] Step 10: Perform score validity assessment, normalization fusion, and long-tail reweighting;

[0138] This step addresses issues such as inconsistent score dimensions, invalid score items, denominators of 0, and long-tail category bias. The system does not directly add the scores. , , and Instead, it first performs validity checks and normalization within the candidate category set for each score item.

[0139] For any scoring item In the candidate category set Internal calculation of maximum value and minimum value ,like The normalization result is expressed as ,like This indicates that the score item does not have a distinguishing effect within the current candidate set, and is therefore marked as an invalid score item. (Regarding difference enhancement scores...) If no valid visible difference attribute exists for any of the candidate categories, the score item will also be marked as invalid.

[0140] Let the set of valid scores be denoted as ,when When not empty, the basic fusion score is represented as ,in, Indicates the scoring item The fusion weights satisfy .

[0141] When a scoring item is invalid, its weight is not included in the merging process, and the weights of the other valid scoring items are renormalized according to their original proportions. When empty, use the initial recognition score within the candidate category set. As the base fusion score; if the initial recognition score also lacks discriminatory power, then... Set to uniform distribution and trigger low-confidence verification flags.

[0142] To mitigate long-tail class bias, the system reweights the base fusion score based on the frequency of class samples. The reweighting coefficient is expressed as follows: ,in, This represents a truncation function, used to restrict a value to [...]. Within the range, This represents the average number of samples per class in the training set, and satisfies... , This represents the total number of samples in the training set; This represents the number of categories in the entire category set; Indicate category The number of samples in the training set or category sample library; A smoothing factor greater than 0; It is a long-tailed regulating factor, and ; and Let represent the lower and upper bounds of the reweighting coefficients, respectively, and satisfy . .

[0143] The final category score is represented as When the denominator is 0, the final category score of each candidate category in the candidate category set is set to 0. Alternatively, it can revert to the normalized result of the initial recognition score.

[0144] in, Indicates candidate category The final category score; Indicates the candidate category to be calculated; This represents the set of candidate categories obtained from the preceding steps; Represents the set of candidate categories Any candidate category used for summation and normalization; Indicate category The corresponding long-tail weighting coefficients; Indicate category The corresponding long-tail weighting coefficients; Indicate category Basic fusion score; Indicate category Basic fusion score; This represents the summation of all candidate categories within the candidate category set; the denominator represents the sum of the basic fusion scores of all candidate categories after long-tail reweighting, which is used to normalize the final category scores.

[0145] Step 11: Calculate the reliability of identification and output the low centroid verification mark;

[0146] This step is used to avoid forcibly outputting high-confidence results when evidence is insufficient. The system selects the candidate category with the highest final category score. Calculate the reliability score for identification .

[0147] Furthermore, the reliability score is represented as ,in, This represents the score interval between the highest-scoring candidate category and the second-highest-scoring candidate category. If there is only one candidate category, then... ; express The effective visibility ratio of the corresponding difference attribute; and These represent the normalized evidence support score and the hierarchical consistency score, respectively. , , , They are respectively , , and The reliability score weights are non-negative and normalized summed among the valid items. .

[0148] The effective visibility ratio is expressed as: .

[0149] when Less than the preset review threshold When the system outputs a low-confidence review flag, it lists the reasons for the low confidence. Reasons for low confidence include, but are not limited to: Top 1 and Top 2 scores are close, key difference attributes are not visible, classification hierarchy path conflicts, insufficient evidence support, low image quality, or lack of sufficient reference evidence in the category library.

[0150] The system will select the candidate category with the highest final category score. The system uses the final predicted category as the classification level and outputs the corresponding hierarchical path. It also considers the recognition reliability score. Evidence support score , effective visibility ratio of difference attributes The number of unverified attributes and the reasons for low confidence are used to generate an evidence sufficiency status. Evidence sufficiency status includes at least two types: sufficient evidence and insufficient evidence. When the identification reliability score is not lower than a preset review threshold, and key difference attributes are visible and the evidence support meets preset requirements, the evidence sufficiency status is sufficient. When key difference attributes are not visible, evidence support is insufficient, there are classification hierarchy path conflicts, or low confidence review markers, the evidence sufficiency status is insufficient.

[0151] The final output includes the final predicted category, classification hierarchy path, final category score, recognition reliability score, evidence sufficiency status, low confidence verification marker, key morphological attributes, easily confused category difference attributes, attribute-level discrimination problem, local visual evidence, reference evidence, and unverified attributes.

[0152] Step 12: Target domain calibration and closed-loop update;

[0153] This step is used to adapt to domain differences between laboratory specimen images, publicly available dataset images, field monitoring images, and images taken with a mobile phone. The system collects a small number of insect image samples from the target domain in the deployment scenario. After expert verification or high-confidence recognition results, these samples are used to update at least one of the following: candidate recall threshold, attribute matching confidence threshold, fusion weight, verification threshold, and category multimodal prototype library.

[0154] The target domain calibration parameter set is denoted as The target domain calibration sample set is denoted as The parameters can be selected through the following objectives:

[0155] ,

[0156] in, Represents the set of calibration parameters in all possible target domains. In the middle, find one to make The parameter combination that yields the maximum value. It can be accuracy, macro-average F1, candidate recall, low-confidence false positive rate, or a weighted combination thereof. This objective function is only used to select the threshold and fusion parameters and does not require retraining of the multimodal large model ontology.

[0157] When computational resources are limited in the deployment scenario, a multimodal large model can be used as the teacher model to generate soft labels or attribute evidence labels for a lightweight visual model or a lightweight attribute matching model. Then, an edge recognition model is obtained through knowledge distillation. This distillation method is an optional implementation and does not affect the basic process of recognition using a multimodal large model in this invention.

[0158] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and improvements, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A semantically enhanced fine-grained insect recognition method based on a multimodal large model, characterized in that: The method includes: S1. Construct an insect image sample set and category semantic knowledge units, and generate multi-text type category semantic enhanced descriptions; S2. Construct a multi-modal prototype library of categories and generate a set of candidate categories based on the images of insects to be identified; S3. For the candidate categories and their easily confused categories in the candidate category set, read the set of differential attributes, and extract local visual evidence from the insect image to be identified based on the observable visual parts corresponding to each differential attribute. S4. Organize a multimodal large model to identify inputs and calculate initial identification scores, then calculate classification hierarchy consistency scores, difference enhancement scores, and evidence support scores; S5. The classification hierarchy consistency score, difference enhancement score, and evidence support score are evaluated for score validity, normalized and fused, and long-tail reweighted to obtain the final category score. S6. Calculate the recognition reliability score based on the candidate category with the highest final category score to obtain the final result including data from multiple categories; The constructed category of multimodal prototype library includes: A category multimodal prototype library is constructed using category visual prototypes, category semantic prototypes, category difference attribute prototypes, and category hierarchical path prototypes. For category visual prototypes, visual features of each sample image under the same category are extracted using a visual encoding function. Then, the visual features are clustered, and representative samples whose distance from the cluster centroid is within a set range are selected. The category visual prototype is obtained by averaging the features of the representative samples. The category semantic prototype is composed of the obtained category semantic vector; The prototype of the category difference attribute is obtained by text encoding the difference attribute. The category-level path prototype is obtained from the category-level path text encoding or one-hot encoding vector. The step of generating a candidate category set based on the image of the insect to be identified includes: After receiving the image of the insect to be identified, the system generates the visual features, image semantic vector, and upper-level classification prediction path of the image of the insect to be identified. Then, it generates the visual prototype recall category set, the semantic prototype recall category set, the classification level path recall category set, and the easily confused category expansion set, respectively. The system then takes the union of these sets to obtain the candidate category set. The upper-level classification prediction path is obtained by a multimodal large model or an independent auxiliary classifier based on the image of the insect to be identified. It includes at least one of the three levels: order, family, and genus. For any predicted level among order, family, and genus, if the level contains multiple possible results, the highest-scoring level is retained. Each result forms a corresponding upper-level classification candidate result. Then, the seed-level categories that match the above upper-level classification candidate results form a classification-level path recall category set. If the union of the visual prototype recall category set, the semantic prototype recall category set, the classification hierarchy path recall category set, and the easily confused category extension set is empty, then all category sets will be used as candidate category sets. If the number of candidate categories exceeds the preset upper limit, then the top K candidate categories will be retained after being sorted by at least two of visual similarity, semantic similarity, and hierarchy path consistency. S3 specifically includes: Read candidate categories set of difference attributes and determine each difference attribute The contrast categories, directions of difference, and observable visual parts ; Images of insects to be identified Coarse localization of insect targets is performed to obtain the insect target region. If the insect target region does not exist, or the confidence of the insect target region is lower than the preset target threshold, the entire image of the insect to be identified is taken as the target region, and a label indicating insufficient image quality is output. Locate the observable visual parts corresponding to the differences in attributes within the insect target area; Based on observable visual parts The localization result is used to crop a local region of the image, and this local region is recorded as local visual evidence. If the same difference attribute corresponds to multiple observable visual parts, then multiple local region images are cropped separately, and the local region image with the highest location confidence or the highest matching score is retained. For each difference attribute Corresponding local visual evidence Visibility is assessed if the local area simultaneously meets the following conditions: the location confidence level is not lower than a preset location threshold, the area of ​​the local area is not less than a preset proportion of the insect target area, the effective discriminable area proportion of the local area is not lower than a preset visibility proportion threshold, and the clarity index of the local area is not lower than a preset clarity threshold; then the visibility is marked. Set as If any of the above conditions are not met, then visibility will be marked. Set as And the difference attributes corresponding to this local visual evidence. Recorded as unusable attribute; where, the effective discriminable area ratio is the ratio of the area of ​​the local region that is not occluded and can present texture, edge or color features to the total area of ​​the local region. The sharpness index is represented by the Laplacian variance of the local region. When the Laplacian variance is lower than the preset sharpness threshold, the local region is judged to be blurry. The organization's multimodal large model identification input and calculation of the initial identification score include: The images of the insects to be identified, local visual evidence, candidate category sets, category semantic enhancement descriptions, classification hierarchy information, and attribute-level discrimination problems are organized into a multimodal large model recognition input; The initial recognition score is obtained through either a comparative image-text matching method or a generative image-text matching method. The calculation process for the classification hierarchy consistency score includes: The consistency between the candidate category and the upper-level classification hierarchy prediction path is calculated to obtain the classification hierarchy consistency score. ,in, Indicates candidate category The hierarchical consistency score, , , Representing candidate categories Matching results with the upper-level predicted paths at the order, family, and genus levels. , , They are respectively , , The weights; The calculation process for the difference enhancement score includes: For candidate categories Read its set of differences and for each difference attribute Local visual evidence Match the attribute descriptions to obtain a matching score. Attribute source confidence Attribute matching confidence and attribute visibility marker The difference enhancement score was obtained. ,when If candidate category c does not have a valid difference attribute, the difference enhancement score will be increased. Items marked as invalid will not have their difference enhancement enabled during score fusion, and the fusion weights for other valid items will be renormalized. Indicates the stabilizing factor; The calculation process for the evidence support score includes: Based on a multimodal prototype library, the consistency between the image to be identified and the candidate category prototypes is measured to obtain the evidence support score. ,in, Indicates candidate category Evidence support score Represents the visual features of the image to be identified. This represents the image semantic vector corresponding to the image to be identified. This represents the difference attribute evidence vector corresponding to the image to be identified. This represents the prediction vector from the higher-level classification hierarchy of the image to be identified. , , , Representing categories The category visual prototype, category semantic prototype, category difference attribute prototype, and category hierarchy path prototype. , , , They are respectively , , and The weight, This represents the normalized similarity function; If a prototype is missing, the weight corresponding to that prototype is reset to 0, and the weights of the remaining prototypes are renormalized. If all prototypes are missing, then... Marked as invalid score item.

2. The semantically enhanced fine-grained insect recognition method based on a multimodal large model according to claim 1, characterized in that: The construction of the insect image sample set and category semantic knowledge unit includes: Obtain insect image samples and their category labels from publicly available insect datasets, images collected by monitoring equipment, or images annotated by experts. Perform invalid image removal, label unification, duplicate sample removal, and low-quality sample labeling on the samples. Then divide the samples into training set, validation set, and test set according to a preset ratio. For each insect category, a category semantic knowledge unit is constructed. The category semantic knowledge unit includes at least the category label, category name, Latin scientific name, classification hierarchy information, morphological attribute description, easily confused category set, differential attribute set, observable visual parts, attribute source confidence, and attribute matching confidence threshold.

3. The semantically enhanced fine-grained insect recognition method based on a multimodal large model according to claim 1, characterized in that: The generated multi-text type category semantic enhancement description includes: The category semantic knowledge units are converted into text representations used by the multimodal large model. Five types of text are generated for each category, including classification hierarchical path text, scientific name text, Chinese name text, morphological attribute text, and easily confused category difference attribute text. The five types of text are encoded separately and then weighted and fused to obtain category semantic vectors. When a category lacks a difference attribute field, the difference attribute text is set to empty text or back text, and the weight corresponding to the difference attribute text is set to 0. Then the remaining weights are normalized again.

4. The semantically enhanced fine-grained insect recognition method based on a multimodal large model according to claim 1, characterized in that: S5 includes: For any scoring item In the candidate category set Internal calculation of maximum value and minimum value ,like The normalization result is expressed as ,like This indicates that the score item does not have a distinguishing effect within the current candidate set, and is therefore marked as an invalid score item. (For score items that enhance difference...) If no valid visible difference attribute exists for any of the candidate categories, the score item will also be marked as invalid. Let the set of valid scores be denoted as ,when When not empty, the basic fusion score is represented as , Indicates the scoring item The fusion weights are determined such that when a score item is invalid, its weight is not included in the fusion, and the weights of other valid score items are renormalized according to their original proportions. When empty, use the initial recognition score within the candidate category set. As a base fusion score, if the initial recognition score also lacks discriminatory power, then... Set to uniform distribution and trigger low-confidence check flags; The basic fusion score is reweighted based on the frequency of samples by category to obtain the reweighting coefficient. The final category score is then obtained as follows: ,when When the value is 0, the final category score of each candidate category in the candidate category set is set to 0. Alternatively, it can revert to the normalized result of the initial recognition score, where, Indicate category The corresponding long-tail weighted coefficients, Indicate category The basic fusion score.

5. The semantically enhanced fine-grained insect recognition method based on a multimodal large model according to claim 4, characterized in that: S6 includes: Based on the candidate category with the highest final category score The reliability score for identification is expressed as follows: ,in, This represents the score interval between the highest-scoring candidate category and the second-highest-scoring candidate category. express The effective visibility ratio of the corresponding difference attribute and These represent the normalized evidence support score and the hierarchical consistency score, respectively. , , , They are respectively , , and The reliability score weighting; when Less than the preset review threshold When low confidence is reached, output a low confidence check flag and list the reasons for the low confidence. The candidate category with the highest final category score As the final predicted category, the corresponding classification hierarchy path is output, based on the recognition reliability score. Evidence support score , effective visibility ratio of difference attributes The number of unverified attributes and the reasons for low confidence are used to generate an evidence sufficiency status. The evidence sufficiency status includes at least two types: sufficient evidence and insufficient evidence. When the identification reliability score is not lower than the preset review threshold, and the key difference attributes are visible and the evidence support meets the preset requirements, the evidence sufficiency status is sufficient evidence. When there are key difference attributes that are not visible, insufficient evidence support, classification hierarchical path conflicts, or low confidence review markers, the evidence sufficiency status is insufficient evidence. The final output includes the final predicted category, classification hierarchy path, final category score, recognition reliability score, evidence sufficiency status, low confidence verification marker, key morphological attributes, easily confused category difference attributes, attribute-level discrimination problem, local visual evidence, reference evidence, and unverified attributes.

Citation Information

Patent Citations

  • Hierarchical image classification method and device based on semantic knowledge guidance

    CN117975151A

  • Image recognition method and device, equipment, storage medium and program product

    CN118658035A

  • RAG-based voucher classification method, medium and equipment

    CN121166928A