A zero-shot object counting method, system, device and storage medium

CN121505360BActive Publication Date: 2026-06-23CHONGQING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2025-12-16
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing computer vision models neglect the potential inherent semantic relationships between feature channels in zero-shot target counting tasks, resulting in limited model representation capabilities and interpretability. Traditional methods cannot effectively utilize the semantic correlation between channels.

Method used

The implicit attribute channel grouping mechanism (IACG) is adopted to dynamically allocate feature channels to semantic groups through prototype vectors and soft allocation strategies. Combined with cross-group attention mechanism (CGAM) and intra-group consistency loss (IGCL), an efficient information exchange channel is established to enhance the collaborative interaction between channels.

Benefits of technology

It improves the counting accuracy and adaptability of zero-sample target counting, demonstrating excellent adaptability across diverse visual tasks and data distributions, and enhancing the quality and interpretability of feature representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505360B_ABST
    Figure CN121505360B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision, and relates to a zero-shot object counting method, system, device and storage medium; wherein the zero-shot object counting method comprises: obtaining a text prompt word of a to-be-identified object and a to-be-identified image; text feature extraction is performed on the text prompt word to obtain a text feature; visual feature extraction is performed on the to-be-identified image to obtain a visual feature; the feature channels of the visual feature are dynamically allocated to a plurality of semantic groups through an implicit attribute channel grouping mechanism to obtain grouped features; a cross-group attention mechanism is used to process the grouped features to obtain enhanced grouped features; the text feature and the enhanced grouped features are fused to obtain fused features, and an object density prediction map is generated according to the fused features; the number of objects is calculated according to the object density prediction map to obtain an object counting result. The present application can improve the object counting precision.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Text-guided zero sample target counting method, program product and electronic equipment

    CN120147298A