Text-guided zero sample target counting method, program product and electronic equipment
By using a multi-level polar dual attention module in the target count, establishing a bidirectional correlation and processing positive and negative polarity information, the one-way limitation in the information transmission process is solved and the accuracy of density estimation is improved.
Patent Information
- Application Number
- CN202510321233.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art has unidirectional limitations in the information transmission process in the target count, resulting in insufficient expression ability of complex relationships between modals, which in turn affects the accuracy of density estimation.
A multi-level polar dual attention module is adopted to establish a fine-grained bidirectional relationship through visual polarity cross attention processing and text polarity cross attention processing, ensuring the mutual transmission of information from text to vision and from vision to text, and separating the positive and negative polarity parts of the query vector and key vector, and calculating the similarity between homopolarity and heteropolarity.
It enhances the complementarity and consistency of features, solves the one-way limitation of the traditional cross-attention mechanism, improves the expression ability of cross-modal relationships, and improves the accuracy of the density estimation graph of the target.
Smart Images

Figure CN120147298A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a text-guided zero-shot object counting method, program product, and electronic device. Background Art
[0002] Object counting, as an important task in the field of computer vision, has broad application prospects in fields such as crowd, agriculture, transportation, and healthcare. The zero-shot object counting algorithm (also called class-agnostic object counting method) aims to construct a paradigm-free object counter and can specify the objects to be counted. It assumes that the model has not seen samples of the target class during the training phase, but by leveraging the prior knowledge and representation ability of the model, it can infer and count objects of new classes. The zero-shot object counting algorithm usually utilizes the semantic and feature representations learned by the model during the training phase, combined with the description information or attribute information of the target class, to estimate the count of objects of the new class.
[0003] In the related art, Jiang et al. proposed a counting method CLIP-Count that utilizes a pre-trained vision-language model. This method can simultaneously understand images and related text descriptions through the pre-trained vision-language model. Additionally, CLIP-Count introduced an image patch-text contrastive loss to guide the model to learn informative image patch-level visual representations, enabling the model to better understand different objects in the image and generate more accurate density predictions, achieving end-to-end density map prediction in a zero-shot manner based on multi-modal (text and vision).
[0004] In multi-modal learning, efficiently fusing information from different modalities (such as vision and text) is crucial for improving model performance. The traditional cross-attention mechanism realizes the interaction between modalities by calculating the attention weights between a query of one modality and the key and value of another modality. However, these methods are often limited in capturing complex, non-linear relationships between modalities, especially in dealing with negative polarity partial information and modality specificity, resulting in inaccurate object density estimation. Summary of the Invention
[0005] This application aims to at least solve the technical problems existing in the prior art and provides a text-guided zero-shot object counting method, program product, and electronic device.
[0006] In a first aspect, the present application provides a text-guided zero-shot object counting method, including: obtaining a query image and an object description text, inputting the query image and the object description text into a trained object counting model, and the object counting model outputs a density estimation map of the object; the object counting model includes: a visual feature extraction module for obtaining visual features of the query image; a text feature extraction module for obtaining text features of the object description text; a multi-level polar dual attention module including a plurality of cascaded feature fusion modules at different levels; wherein, each feature fusion module at least one level performs visual polar cross-attention processing on the visual-to-text fusion features or visual features output by the feature fusion module at the previous level to obtain the visual-to-text fusion features at this level, and performs text polar cross-attention processing on the text-to-visual fusion features or text features output by the feature fusion module at the previous level to obtain the text-to-visual fusion features at this level; a decoder that decodes the visual-to-text fusion features output by the feature fusion module at least one level to obtain a density estimation map of the object.
[0007] In a second aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method provided in the first aspect of the present application.
[0008] In a third aspect, the present application provides an electronic device, the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a text-guided zero-shot object counting method provided in the first aspect of the present application.
[0009] The beneficial technical effects of the present invention: The visual feature extraction module is used to obtain visual features of the query image, the text feature extraction module is used to obtain text features of the object description text, and the multi-level polar dual attention module uses visual polar cross-attention processing and text polar cross-attention processing to process visual features and text features respectively, which can establish a fine-grained bidirectional association between text features and visual features, ensuring that information is not only transmitted from text to vision, but also fed back from vision to text, enhancing the complementarity and consistency of features, and solving the one-way limitation in the information transmission process of the traditional cross-attention mechanism. At the same time, in visual polar cross-attention processing and text polar cross-attention processing, the positive and negative polarity partial information of the query vector and the key vector are separated and the similarity of the same polarity and the similarity of the different polarity are calculated, and the negative polarity partial information is retained, improving the expression ability of the cross-modal relationship of the visual-to-text fusion features, thereby improving the accuracy of the density estimation map of the object obtained by the decoder. Description of the Drawings
[0010] Figure 1 It is a schematic structural diagram of the target counting model in a preferred embodiment of the present invention;
[0011] Figure 2 It is a schematic structural diagram of the target counting model in another preferred embodiment of the present invention;
[0012] Figure 3 It is a schematic diagram of the connection relationship and internal structure of the visual feature extraction module and the decoder in another preferred embodiment of the present invention;
[0013] Figure 4 It is a schematic process diagram of visual polarity cross-attention processing and text polarity cross-attention processing in another preferred embodiment of the present invention;
[0014] Figure 5 It is a schematic diagram of fine-tuning the text encoder in another preferred embodiment of the present invention;
[0015] Figure 6 It is a schematic diagram of the contrast learning alignment embedding hidden space in another preferred embodiment of the present invention;
[0016] Figure 7 It is a graph of the qualitative comparison results between the target counting model of the present invention and the existing CLIP-Count model in another preferred embodiment of the present invention;
[0017] Figure 8 It is a schematic structural diagram of an electronic device in another preferred embodiment of the present invention. Specific Embodiments
[0018] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.
[0019] The execution entities of a text-guided zero-shot object counting method provided by the present invention include, but are not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, a text-guided zero-shot object counting method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0020] The present invention provides a text-guided zero-shot object counting method. In a preferred embodiment, the method includes:
[0021] Step S1, obtain a query image and a target description text. The target description text is a text including a target category. The target category can be a living object or an inanimate object, such as cells, people, animals, cups, fruits, grapes, tablets, etc. Exemplarily, Figure 8 as shown, when it is necessary to count the number of cups in the query image, the target description text can be "cups".
[0022] Step S2, input the query image and the target description text into a trained object counting model.
[0023] Step S3, the object counting model outputs a density estimation map of the target, that is, a density estimation heat map. The pixel value of each pixel point in the density estimation heat map represents whether there is a target at that pixel point. Exemplarily, the geometric center points of each target in the density estimation heat map present high-brightness points.
[0024] Please refer to Appendix Figure 1 and 2 , the object counting model includes:
[0025] A visual feature extraction module, which is used to obtain the visual features of the query image.
[0026] Exemplarily, the visual feature extraction module includes a visual encoder, and the visual encoder can adopt the visual encoder of CLIP. To obtain more efficient visual features and train any image set without relevant metadata, the visual encoder preferably adopts the DINOv2 visual encoder. Specifically, the image patch-level feature map output by the DINOv2 visual encoder is used as the visual feature.
[0027] A text feature extraction module for obtaining the text features of the target description text.
[0028] Exemplarily, the text feature extraction module can be a BERT encoder or a text encoder of CLIP. Preferably, the text feature extraction module adopts the text encoder of CLIP, which connects images and natural language through large-scale pre-training on image-text pair data and is more easily aligned with visual features.
[0029] A multi-level polarity dual attention module, including a plurality of cascaded feature fusion modules at different levels; wherein, each feature fusion module at least at one level performs visual polarity cross-attention processing on the visual-to-text fusion features or visual features output by the feature fusion module at the previous level to obtain the visual-to-text fusion features at this level, and performs text polarity cross-attention processing on the text-to-visual fusion features or text features output by the feature fusion module at the previous level to obtain the text-to-visual fusion features at this level.
[0030] In this embodiment, each of at least one feature fusion module has a text input end, a text output end, a visual input end, and a visual output end. Input the text features to the text input end of the feature fusion module at the first level, input the visual features to the visual input end of the feature fusion module at the first level, the text output end of the feature fusion module at the first level outputs the text-to-visual fusion features to the text input end of the feature fusion module at the second level, the visual output end of the feature fusion module at the first level outputs the visual-to-text fusion features to the visual input end of the feature fusion module at the second level, and the feature fusion modules at multiple levels are connected in turn by analogy. For adjacent feature fusion modules, the text output end of the feature fusion module at the previous level is correspondingly connected to the text input end of the feature fusion module at the next level, and the visual output end of the feature fusion module at the previous level is directly or indirectly correspondingly connected to the visual input end of the feature fusion module at the next level.
[0031] A decoder that decodes the visual-to-text fusion features output by at least one level of feature fusion module to obtain the target density estimation map.
[0032] Exemplarily, the decoder is an existing decoder based on a CNN network. It can decode the visual-to-text fusion features output by the feature fusion module at the last level to obtain the target density estimation map. It can also decode the comprehensive fusion features obtained by fusing the visual-to-text fusion features output by the feature fusion modules at the last level and one or more other (non-last) levels to obtain the target density estimation map. Please see the appendix Figure 3The right - hand Decoder structure can also connect the visual - to - text fusion features output by the feature fusion modules of the last layer and more than one (non - last) other layer to each node of the decoder network in a residual form respectively, so as to obtain the target density estimation map.
[0033] In this embodiment, preferably, the feature fusion module of each layer in at least one layer includes:
[0034] A visual - to - text polarity cross - attention unit that performs visual - polarity cross - attention processing on the visual features or the visual - to - text fusion features output by the feature fusion module of the previous layer to obtain the visual - to - text attention features of this layer. For the feature fusion module of the first layer, its visual - to - text polarity cross - attention unit performs visual - polarity cross - attention processing on the visual features. For the feature fusion module of non - first layers, its visual - to - text polarity cross - attention unit performs visual - polarity cross - attention processing on the visual - to - text fusion features output by the feature fusion module of the previous layer.
[0035] A text - to - visual polarity cross - attention unit that performs text - polarity cross - attention processing on the text features or the text - to - visual fusion features output by the feature fusion module of the previous layer to obtain the text - to - visual attention features of this layer. For the feature fusion module of the first layer, its text - to - visual polarity cross - attention unit performs text - polarity cross - attention processing on the text features. For the feature fusion module of non - first layers, its text - to - visual polarity cross - attention unit performs text - polarity cross - attention processing on the text - to - visual fusion features output by the feature fusion module of the previous layer.
[0036] A first fusion unit that performs a first fusion process on the visual - to - text attention features and the text - to - visual attention features to obtain visual - to - text fusion features.
[0037] A second fusion unit that performs a second fusion process on the visual - to - text attention features and the text - to - visual attention features to obtain text - to - visual fusion features.
[0038] In this embodiment, please refer to Figure 3 , the visual - to - text polarity cross - attention unit and the text - to - visual polarity cross - attention unit are together denoted as PDCA (PolaDualCrossAttention). The first fusion unit and the second fusion unit are feed - forward neural networks FNN with two or more layers. Specifically, the first fusion unit or the second fusion unit can be two or more cascaded fully - connected layers.
[0039] In this embodiment, please refer to Figure 4 , the process of the visual - to - text polarity cross - attention unit performing visual - polarity cross - attention processing includes:
[0040] Step A1: Use the visual feature or the visual-to-text fusion feature output by the upper-level feature fusion module as the visual query vector vis_q, and use the text feature or the text-to-visual fusion feature output by the upper-level feature fusion module as the visual key vector vis_k and the visual value vector vis_v.
[0041] Step A2: Adjust the amplitudes of the visual query vector vis_q and the visual value vector vis_v with the first feature scaling factor scale1. Preferably, the first feature scaling factor scale1 is a learning parameter, learned during the training of the target counting model to dynamically adjust the feature amplitudes. Preferably, the visual query vector vis_q and the visual value vector vis_v are first reshaped into multiple heads and then the amplitude adjustment is performed.
[0042]
[0043] Step A3: Perform positive and negative polarity decomposition and enhancement processing on the amplitude-adjusted visual query vector vis_q' to obtain the first query positive polarity enhanced feature Q1 + and the first query negative polarity enhanced feature Q1 - ; perform positive and negative polarity decomposition and enhancement processing on the visual key vector vis_k to obtain the first key positive polarity enhanced feature K1 + and the first key negative polarity enhanced feature K1 - .
[0044] In this embodiment, the positive and negative polarity decomposition of the vector means taking the vector itself as the positive polarity part, and defining the result after bitwise inversion of the vector (such as multiplying each element by -1) as the negative polarity part of the vector.
[0045] Traditional cross-attention often ignores the negative value information in the query-key pair due to non-negative feature mapping (such as the ReLU activation function), which limits the model's ability to express complex relationships between modalities. To alleviate this problem, this application decomposes the query and the key into positive and negative polarity parts, and performs enhancement processing on the positive and negative polarity parts respectively.
[0046] Further preferably, a learnable power function is introduced for enhancement processing, and the specific calculation formula is as follows:
[0047] Q1 + = ReLU(vis_q') 1+α1·sigmoid ( p1
[0048] Q1 - = ReLU(-vis_q') 1+α1·sigmoid ( p1
[0049] K1+ = ReLU(vis_k) 1+α1·sigmoid ( p1
[0050] K1 - = ReLU(-vis_k) 1+α1·sigmoid ( p1 )
[0051] Wherein, -vis_q' is obtained by bitwise inversion of vis_q', and -vis_k is obtained by bitwise inversion of vis_k. Exemplarily, bitwise inversion means that if a number is 1, then its bitwise inversion is -1. p1 represents the first power exponent parameter, and α1 is the first weight parameter used to control the feature scaling intensity. Both p1 and α1 are learnable parameters and are learned during the training of the target counting model. ReLU(·) represents the ReLU activation function.
[0052] Step A4, calculate the first positive similarity S1 through the first query positive polarity enhanced feature Q1 + and the first key positive polarity enhanced feature K1 + ; calculate the first negative similarity S1 through the first query negative polarity enhanced feature Q1 pos and the first key negative polarity enhanced feature K1 - and the first key negative polarity enhanced feature K1 - Calculate the first negative similarity S1 neg .
[0053] In this embodiment, the first positive similarity S1 pos or the first negative similarity S1 neg can be calculated by Euclidean distance or cosine similarity. Preferably, the first positive similarity S1 pos or the first negative similarity S1 neg is calculated using the following formula:
[0054] S1 pos = (Q1 + ·K1 +
[0055] S1 neg = (Q1 - ·K1 - )
[0056] Wherein, "·" represents the dot product operator.
[0057] Step A5, perform activation function processing on the difference between the first positive similarity S1 pos and the first negative similarity S1 neg to obtain the first attention weight attn1. The activation function is preferably softmax(), specifically:
[0058] attn1 = softmax(S1 pos - S1 neg )
[0059] Step A6. Obtain the visual-to-text attention feature vis at this level based on the first attention weight attn1 and the visually adjusted value vector vis_v'. context Specifically:
[0060] vis context = Dropout(attn1)·vis_v′.
[0061] Among them, Dropout(·) represents the dropout operation, which is used to prevent overfitting during training.
[0062] In this embodiment, preferably, see Figure 4 , the process of the text-to-visual polarity cross-attention unit performing text polarity cross-attention processing includes:
[0063] Step B1. Use the text feature or the text-to-visual fusion feature output by the feature fusion module of the previous level as the text query vector txt_q, and use the visual feature or the visual-to-text fusion feature output by the feature fusion module of the previous level as the text key vector txt_k and the text value vector txt_v.
[0064] Step B2. Adjust the amplitudes of the text query vector txt_q and the text value vector txt_v with the second feature scaling factor scale2. Preferably, the second feature scaling factor scale2 is a learning parameter, which is learned during the training of the target counting model to dynamically adjust the feature amplitude. Preferably, the text query vector txt_q and the text value vector txt_v are first reshaped in multiple heads and then the amplitude is adjusted.
[0065]
[0066] Step B3. Perform positive and negative polarity decomposition and enhancement processing on the amplitude-adjusted text query vector txt_q' to obtain the second query positive polarity enhanced feature Q2 + and the second query negative polarity enhanced feature Q2 - ; perform positive and negative polarity decomposition and enhancement processing on the text key vector txt_k to obtain the second key positive polarity enhanced feature K2 + and the second key negative polarity enhanced feature K2 - .
[0067] Traditional cross-attention often ignores negative value information in query-key pairs due to non-negative feature mapping (such as ReLU), which limits the model's ability to express complex relationships between modalities. To alleviate this problem, this application decomposes the query and key into positive and negative polarity parts and performs enhancement processing on the positive and negative polarity parts respectively.
[0068] Further preferably, a learnable power function is introduced for enhancement processing, and the specific calculation formula is as follows:
[0069] Q2 + = ReLU(txt_q') 1+α2·sigmoid(p2)
[0070] Q2 - = ReLU(-txt_q') 1+α2·sigmoid(p2)
[0071] K2 + = ReLU(txt_k) 1+α2·sigmoid(p2)
[0072] K2 - = ReLU(-txt_k) 1+α2·sigmoid(p2)
[0073] Among them, -txt_q' is obtained by bitwise inversion of txt_q', and -txt_k is obtained by bitwise inversion of txt_k. p2 represents the second power exponent parameter, and α2 is the second weight parameter used to control the feature scaling intensity. Both p2 and α2 are learnable parameters and are learned during the training of the target counting model.
[0074] Step B4, calculate the second positive similarity S2 through the second query positive polarity enhanced feature Q2 + and the second key positive polarity enhanced feature K2 + ; calculate the second negative similarity S2 through the second query negative polarity enhanced feature Q2 pos ; and the second key negative polarity enhanced feature K2 - and preferably: - calculate the second negative similarity S2 neg :
[0075] S2 pos = (Q2 + · K2 + )
[0076] S2 neg = (Q2 - · K2 - )
[0077] Step B5, for the second positive similarity S2 pos and the second negative similarity S2 negThe difference is processed by an activation function to obtain the second attention weight attn2. The activation function is preferably softmax(), specifically:
[0078] attn2 = softmax(S2 pos - S2 neg ).
[0079] Step B6: Based on the second attention weight attn2 and the amplitude-adjusted text value vector txt_v', obtain the text-to-visual attention feature txt at this level context :
[0080] txt context = Dropout(attn2)·txt_v'.
[0081] In this embodiment, through the visual polarity cross-attention processing in steps A1 - A6 and the text polarity cross-attention processing in steps B1 - B6, the polarity dual cross-attention mechanism is realized. The PDCA unit composed of the visual-to-text polarity cross-attention unit and the text-to-visual polarity cross-attention unit can establish a fine-grained bidirectional association between text and visual features, significantly enhancing the fusion and understanding ability of multi-modal information. This dual information flow allows text and visual features to interact in two directions, ensuring that information is not only transmitted from text to vision but also fed back from vision to text, enhancing the complementarity and consistency of features and solving the one-way limitation in the information transmission process of the traditional cross-attention mechanism. At the same time, by separating the positive and negative components of the query and key and calculating the same-polarity and different-polarity similarities, negative value information is retained, improving the expression ability of cross-modal relationships. The use of power functions and the setting of learnable p1, p2, α1, and α2 implement a dynamic scaling mechanism to reduce the entropy of the attention distribution and improve the discrimination of strong and weak signals, which helps to improve the accuracy of target counting.
[0082] In a preferred embodiment, please refer to the attached Figure 3 , the feature fusion module in each level of at least one level further includes a visual multi-head self-attention unit ( Figure 3 (the MHSA on the right in (a)) and a text multi-head self-attention unit ( Figure 3 (the MHSA on the left in (a))).
[0083] The visual multi-head self-attention unit performs multi-head self-attention processing on the visual features or the visual-to-text fusion features output by the feature fusion module of the previous layer to obtain visual self-attention features. Specifically, for the feature fusion module of the first layer, its visual multi-head self-attention unit performs multi-head self-attention processing on the visual features. For the feature fusion module that is not the first layer, its visual multi-head self-attention unit performs multi-head self-attention processing on the visual-to-text fusion features output by the feature fusion module of the previous layer.
[0084] At this time, the visual-to-text polarity cross-attention unit performs visual polarity cross-attention processing on the visual self-attention features to obtain the visual-to-text attention features of this layer.
[0085] The text multi-head self-attention unit performs multi-head self-attention processing on the text features or the text-to-visual fusion features output by the feature fusion module of the previous layer to obtain text self-attention features. Specifically, for the feature fusion module of the first layer, its text multi-head self-attention unit performs multi-head self-attention processing on the text features. For the feature fusion module that is not the first layer, its text multi-head self-attention unit performs multi-head self-attention processing on the text-to-visual fusion features output by the feature fusion module of the previous layer.
[0086] At this time, the text-to-visual polarity cross-attention unit performs text polarity cross-attention processing on the text self-attention features to obtain the text-to-visual attention features of this layer.
[0087] In this embodiment, the visual multi-head self-attention unit and the text multi-head self-attention unit are used to reduce the overfitting risk, improve the expression ability of features, and accelerate the calculation efficiency.
[0088] In this embodiment, the process of the visual-to-text polarity cross-attention unit performing visual polarity cross-attention processing on the visual self-attention features is similar to steps A1 - A6 of the above preferred embodiment, only replacing step A1 with the following steps, and steps A2 - A6 remain unchanged:
[0089] The new step A1 uses the visual self-attention features as the visual query vector vis_q, and the text self-attention features as the visual key vector vis_k and the visual value vector vis_v.
[0090] In this embodiment, the process of the text-to-visual polarity cross-attention unit performing text polarity cross-attention processing on the text self-attention features is similar to steps B1 - B6 of the above preferred embodiment, only replacing step B1 with the following steps, and steps B2 - B6 remain unchanged:
[0091] A new step B1, using the text self-attention feature as the text query vector txt_q, and using the visual self-attention feature as the text key vector txt_k and the text value vector txt_v.
[0092] In a preferred embodiment, please refer to the attached Figure 3 , there is also a convolutional upsampling module (Conv&2xup) connected between the feature fusion modules of two adjacent levels. The convolutional upsampling module performs convolution and upsampling on the visual-to-text fusion feature output by the feature fusion module of the previous level in two adjacent levels, and uses the processing result as the input of the feature fusion module of the next level in two adjacent levels.
[0093] In this embodiment, the convolutional upsampling module includes a convolutional unit and an upsampling unit connected in sequence. The convolutional unit has a skip connection and is used to further process the visual-to-text fusion feature output by the feature fusion module of the previous level and capture local spatial information. The upsampling unit is preferably a 2× bilinear interpolation layer (i.e., 2-fold upsampling) and is used to double the resolution of the feature map output by the convolutional unit. The motivation for such a design is to capture the relationship between text and image in a finer-grained manner, enhancing the complementarity and consistency of features.
[0094] In this embodiment, any convolutional upsampling module is used to perform convolution and upsampling on the visual-to-text fusion feature output by the feature fusion module of its previous level, and input the processing result into the visual multi-head self-attention unit or the visual-to-text polarity cross-attention unit of the feature fusion module located in its next level.
[0095] In a preferred embodiment, please refer to the attached Figure 2 , the visual feature extraction module includes: a visual encoder that processes the query image to obtain a visual embedding; a visual adapter that transforms the visual embedding to a specified dimension to obtain a visual feature.
[0096] And / or, the text feature extraction module includes: a text encoder that processes the target description text to obtain a text embedding; a text adapter that transforms the text embedding to a specified dimension to obtain a text feature.
[0097] In this embodiment, since the visual encoder (such as DINOv2) and the text encoder (such as CLIP) are pre-trained using different training methods on datasets with different distributions, this results in the image patch embeddings extracted by the visual encoder and the text embeddings extracted by the text encoder possibly being in different feature latent spaces. Therefore, visual adapters and text adapters are set to transform the visual embedding and the text embedding to the same specified dimension respectively, facilitating subsequent operations such as comparison and fusion.
[0098] In this embodiment, preferably, both the visual adapter and the text adapter adopt a MultiLayer Perceptron (MLP). Exemplarily, the dimension of the visual embedding input to the visual adapter is 768, the dimension of the visual features output by the visual adapter is 512, the dimension of the text embedding input to the text adapter is 512, and the dimension of the text features output by the text adapter is 512.
[0099] In this embodiment, the visual adapter Av and the text adapter At are expressed as:
[0100]
[0101] where ε I represents the visual embedding, ε t represents the text embedding, d i represents the visual embedding dimension, d t represents the text embedding dimension, d e represents the set dimension.
[0102] In a preferred embodiment, please refer to Appendix Figure 3 (b). The decoder includes a plurality of cascaded decoding convolutional upsampling modules and an output convolutional layer, and also includes more than one addition unit. More than one addition unit are respectively connected in series between more than one pair of adjacent decoding convolutional upsampling modules, and each addition unit corresponds to a feature fusion module at one level; specifically, the first addition unit corresponds to the feature fusion module at the second level, the second addition unit corresponds to the feature fusion module at the third level, and so on.
[0103] The feature fusion module at the first level outputs visual-to-text fusion features to the input end of the first decoding convolutional upsampling module;
[0104] The addition unit is used to add the output features of the previous decoding convolutional upsampling module of the addition unit and the visual-to-text fusion features output by the corresponding-level feature fusion module, and input the addition result into the next decoding convolutional upsampling module of the addition unit.
[0105] In this embodiment, the decoding convolutional upsampling module includes a decoding convolution and a decoding upsampling layer connected in sequence. The decoding upsampling layer is preferably a 2× interpolation layer. The convolution kernel of the decoding convolution is preferably 3×3, the convolution kernel of the output convolutional layer is preferably 1×1, and the output convolutional layer usually has a Sigmoid activation function. Each time passing through a decoding convolutional upsampling module will double the spatial resolution and halve the dimension.
[0106] Exemplarily, taking Figure 3Taking (b) as an example, the multi-level polar dual attention module includes feature fusion modules at two levels, namely the first level and the second level. A convolutional upsampling module is arranged between the feature fusion modules at the first level and the second level. The decoder includes 4 cascaded decoding convolutional upsampling modules and 1 output convolutional layer. The decoder is provided with 1 addition unit, which is located between the first and the second decoding convolutional upsampling modules and corresponds to the feature fusion module at the second level.
[0107] In the above example, in order to decode feature maps at two different scales simultaneously, the decoder first uses the first decoding convolutional upsampling module to perform convolution and 2× interpolation (the first decoding convolutional upsampling module of the decoder) on the coarse-grained visual-to-text fusion feature M output by the feature fusion module at the first level, so that its dimension is the same as that of the fine-grained visual-to-text fusion feature M output by the feature fusion module at the second level. c Then, before the second decoding convolutional upsampling module, it is fused with M through summation to obtain M f : f f‘ :
[0108] M f‘ = Lerp(σ(Conv 3×3 (M c ))) + σ(Conv 3×3 (M f ))
[0109] where Lerp(.) represents bilinear interpolation and σ(.) is the activation function. The input M f‘ is sent to the second decoding convolutional upsampling module.
[0110] In this embodiment, the decoder can decode visual-to-text fusion features at different scales, which can improve the accuracy of the density estimation map of the target.
[0111] In a preferred embodiment, the training process of the target counting model includes:
[0112] The first-stage training is to fine-tune the visual encoder, visual adapter, text encoder, and text adapter in a self-supervised learning manner. During the training, the visual features and text features are aligned in the high-dimensional latent space through the contrastive learning loss (InfoNCEloss).
[0113] In this embodiment, the visual encoder uses DINOv2, and the image patch feature map output by DINOv2 is selected as the visual embedding. The text encoder uses the text encoder of CLIP.
[0114] Construct a training set Among them, the i-th sample includes a training image a class description text and a binary mapping map In this mapping map, at the center of each object described by t i in X i is set to 1, and other positions are 0. i is the sample index. The sum of Y i can be calculated to count the number of objects described by t i in X i as shown in the following formula, where p and q are the indices of pixel points in the mapping map:
[0115] Count(X i ,t i ) = ∑ p,q (Y i ) p,q
[0116] Similarly, a validation set and a test set are constructed. The training set does not overlap with the validation set and the test set:
[0117] In this embodiment, the steps for fine-tuning the text encoder of CLIP are as follows: as Figure 5 shown, first freeze the network parameters of the existing text encoder of CLIP, add multiple learnable prompt vectors to the text embedding, each prompt vector having the same dimension as the text embedding ε t . Use the text dataset to train the text encoder of CLIP. During training, the prompt vectors will participate in the interaction process with all other vectors in the TransformerBlock, and guide the downstream task by learning a new implicit representation. Finally, extract features from the last end vector (EndOfTextToken, EOT_TOKEN is the maximum value in each sequence) to obtain the final text embedding, completing the fine-tuning of the text encoder of CLIP.
[0118] In this embodiment, the steps for fine-tuning the visual encoder of DINOv2 are as follows: also add multiple prompt vectors to the visual embedding, and use the image set to train the visual encoder of DINOv2. During training, the prompt vectors will participate in the interaction process with all other vectors.
[0119] In this embodiment, the first stage of training process is that the visual encoder, visual adapter, text encoder and text adapter form a backbone network, and the backbone network is iteratively trained using the training set until the fine-tuning training stop condition is reached, and the image block-text contrast loss is calculated in each training, and the weight of the backbone network is fine-tuned by the image block-text contrast loss. The fine-tuning training stop condition is not limited to the number of training times reaching the preset maximum number of fine-tuning times or the image block-text contrast loss converges. The trained backbone network is tested on the validation set. Validated on the test set According to the mask value corresponding to the image block in the mapping image, define and The image patch embedding set for positive and negative samples, image patch-text contrast loss The calculation formula is as follows:
[0120]
[0121] Among them, w i represents the weight of the i-th positive sample, s(.,.) represents the cosine similarity, and τ=0.07 represents the temperature coefficient. Contrastive loss function By pulling the text embedding closer t Visual embedding with positive samples The distance between them pushes the text embedding ε t Visual Embedding with Negative Samples The distance between them aligns the image block embedding space with the text embedding space, such as Figure 6 shown.
[0122] In the second phase of training, after fine-tuning the visual feature extraction module and the text feature extraction module, the network parameters of the two modules are frozen, and the learnable parameters in the multi-level polarity dual attention module and the network parameters of the decoder are updated during training. The second phase of training uses supervised learning and uses the training set to iteratively train the target counting model. i As the true density map of the sample, in each training, for the i-th sample, the density estimation map of the target output by the decoder is used With Map Compare pixel by pixel and find the mean square error To continue adjusting the parameters of the decoder and the multi-level polar dual attention module. i By summing pixel by pixel, we can get the estimated number of items of interest.
[0123] The performance of the target counting model provided in this application is experimentally verified as follows:
[0124] Experiments were conducted on the FSC-147, CARPK, and ShanghaiTech datasets. The mean absolute error (MAE) and root mean square error (RMSE) were used as evaluation metrics. The experimental comparison models were several state-of-the-art few-shot, reference-free, and zero-shot class-agnostic object counting methods, including FamNet, CFOCNet, CounTR, LOCA, RepRPN-C, RCC, Xuetal., and CLIP-Count. The experimental results show that the object counting model (Ours) provided by this application has relatively impressive performance and generality.
[0125] 1. Quantitative analysis and comparison on FSC-147
[0126] The object counting model (Ours) of this application was compared with state-of-the-art class-agnostic object counting methods, and the quantitative results were summarized in Table 1. Overall, the object counting model (Ours) of this application performed better than state-of-the-art zero-shot class-agnostic counting methods on FSC-147 and was competitive compared with reference-free methods. It is worth noting that compared with reference-free methods, zero-shot object counting aims to solve more challenging scenarios because the model needs to understand the correspondence between text and visual features.
[0127] Table 1 Quantitative analysis results on FSC-147
[0128]
[0129] 2. Quantitative analysis on CARPK
[0130] The object counting model (Ours) of this application was tested on the CARPK dataset to evaluate its generalization ability across datasets. Specifically, the model was trained on FSC-147 and directly evaluated on the test set of CARPK without any fine-tuning. For fair comparison, Jiang et al. trained RCC on FSC-147 using the same visual backbone (Vit-B) instead of their proposed FSC-133. As shown in Table 2, the object counting model (Ours) of this application showed a certain performance improvement compared with the representative reference-free counting method RCC and the zero-shot counting method CLIP-Count.
[0131] Table 2 Cross-dataset evaluation on CARPK
[0132]
[0133] Indicates retraining with the same backbone network
[0134] 3. Quantitative Analysis on ShanghaiTech
[0135] An existing CLIP-based crowd counting method called CrowdCLIP evaluated the performance of the model across datasets by training a specific-class object model on a part of the ShanghaiTech dataset and testing on another part. Similar experiments were conducted to directly evaluate the performance of the target counting model (Ours) of this application and RCC, CLIP-Count in the crowd counting task (i.e., without using any part of the ShanghaiTech dataset for fine-tuning). To ensure fair comparison, the backbone network of RCC was also modified to be consistent with the target counting model (Ours) of this application. The experimental results are shown in Table 3.
[0136] Table 3 Cross-dataset Evaluation on ShanghaiTech
[0137]
[0138]
[0139] Indicates retraining with the same backbone network
[0140] 4. Qualitative Analysis Comparison
[0141] On the FSC-147 dataset, the target counting model (Ours) of this application was compared with the current state-of-the-art zero-shot counting method CLIP-Count, and it was found that their performances were similar in sparse scenarios, and the target counting model (Ours) of this application had certain advantages in dense scenarios, such as Figure 7 . Specifically, the counting boundaries obtained by the target counting model (Ours) of this application are clearer, and it can locate the high-density regions of the target center with high fidelity. In contrast, the density prediction in CLIP-Count may show a non-concentrated pattern.
[0142] The present invention also discloses a computer program product, including a computer program, which when executed by a processor implements the steps of the above-mentioned text-guided zero-shot target counting method provided by the present invention. The computer program product should be understood as a software product that mainly implements its solution through a computer program, such as a program product integrated in the cloud or a software library.
[0143] The present invention also discloses an electronic device. In one embodiment, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein,
[0144] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to execute a text-guided zero-shot object counting method provided by the present invention.
[0145] As Figure 8 shown, it is a schematic structural diagram of an electronic device for a text-guided zero-shot object counting method provided by an embodiment of the present invention. The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a text-guided zero-shot object counting method program.
[0146] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing a text-guided zero-shot object counting method, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device and process data.
[0147] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of a text-guided zero-shot object counting method program, etc., but also to temporarily store data that has been output or will be output.
[0148] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to implement the connection communication between the memory 11 and at least one processor 10, etc.
[0149] The communication interface 13 is used for the communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.
[0150] Figure 8 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 8 the shown structure does not constitute a limitation on the electronic device, and it can include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0151] For example, although not shown, the electronic device can also include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source can also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device can also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0152] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0153] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0154] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A text-guided zero-shot target counting method, characterized in that: include: Obtain a query image and a target description text, input the query image and the target description text into a trained target counting model, and the target counting model outputs a density estimation map of the target; The target counting model includes: A visual feature extraction module, used to obtain visual features of the query image; A text feature extraction module is used to obtain text features of the target description text; A multi-level polarity dual attention module, comprising a plurality of cascaded levels of feature fusion modules; wherein the feature fusion module of each level in at least one level performs visual polarity cross attention processing on the visual-to-text fusion features or visual features output by the feature fusion module of the previous level to obtain the visual-to-text fusion features of the current level, and performs text polarity cross attention processing on the text-to-visual fusion features or text features output by the feature fusion module of the previous level to obtain the text-to-visual fusion features of the current level; The decoder decodes the visual-to-text fusion features output by the feature fusion module of at least one level to obtain a density estimation map of the target.
2. A text-guided zero-sample target counting method as claimed in claim 1, characterized in that: The feature fusion module of each level in at least one level includes: The visual-to-text polarity cross attention unit performs visual polarity cross attention processing on the visual features or the visual-to-text fusion features output by the feature fusion module of the previous level to obtain the visual-to-text attention features of this level; The text-to-visual polarity cross attention unit performs text-to-visual fusion features output by the feature fusion module of the previous level with text-polarity cross attention to obtain the text-to-visual attention features of this level. A first fusion unit performs a first fusion process on the visual-to-text attention feature and the text-to-visual attention feature to obtain a visual-to-text fusion feature; The second fusion unit performs a second fusion process on the visual-to-text attention feature and the text-to-visual attention feature to obtain a text-to-visual fusion feature.
3. A text-guided zero-sample target counting method as claimed in claim 2, characterized in that: The feature fusion module of each level in at least one level further includes a visual multi-head self-attention unit and a textual multi-head self-attention unit; The visual multi-head self-attention unit performs multi-head self-attention processing on the visual features or the visual-to-text fusion features output by the feature fusion module of the previous level to obtain the visual self-attention features; The visual-to-text polarity cross-attention unit performs visual polarity cross-attention processing on the visual self-attention features to obtain the visual-to-text attention features of this level; The text multi-head self-attention unit performs multi-head self-attention processing on the text features or the text-to-visual fusion features output by the feature fusion module of the previous level to obtain the text self-attention features; The text-to-visual polarity cross-attention unit performs text-polarity cross-attention processing on the text self-attention features to obtain the text-to-visual attention features of this level.
4. A text-guided zero-sample target counting method as claimed in claim 2 or 3, characterized in that: The process of the visual-to-text polarity cross-attention unit performing visual polarity cross-attention processing includes: Using visual features or visual self-attention features or visual-to-text fusion features output by a feature fusion module of a previous layer as a visual query vector, and using text features or text self-attention features or text-to-visual fusion features output by a feature fusion module of a previous layer as a visual key vector and a visual value vector; Adjusting the magnitude of the visual query vector and the visual value vector by a first feature scaling factor; Performing positive and negative polarity decomposition and enhancement processing on the amplitude-adjusted visual query vector to obtain a first query positive polarity enhancement feature and a first query negative polarity enhancement feature; performing positive and negative polarity decomposition and enhancement processing on the visual key vector to obtain a first key positive polarity enhancement feature and a first key negative polarity enhancement feature; Calculating a first positive similarity by using the first query positive polarity enhancement feature and the first key positive polarity enhancement feature; calculating a first negative similarity by using the first query negative polarity enhancement feature and the first key negative polarity enhancement feature; Performing activation function processing on the difference between the first positive similarity and the first negative similarity to obtain a first attention weight; The visual-to-text attention features of this level are obtained based on the first attention weight and the magnitude-adjusted visual value vector.
5. A text-guided zero-sample target counting method as claimed in claim 2 or 3, characterized in that: The process of the text-to-visual polarity cross-attention unit performing text polarity cross-attention processing includes: Using text features or text self-attention features or text-to-visual fusion features output by a feature fusion module of a previous layer as a text query vector, and using visual features or visual self-attention features or visual-to-text fusion features output by a feature fusion module of a previous layer as a text key vector and a text value vector; Adjusting the magnitude of the text query vector and the text value vector by a second feature scaling factor; Performing positive and negative polarity decomposition and enhancement processing on the amplitude-adjusted text query vector to obtain a second query positive polarity enhancement feature and a second query negative polarity enhancement feature; performing positive and negative polarity decomposition and enhancement processing on the text key vector to obtain a second key positive polarity enhancement feature and a second key negative polarity enhancement feature; Calculating a second positive similarity by using the second query positive polarity enhancement feature and the second key positive polarity enhancement feature; calculating a second negative similarity by using the second query negative polarity enhancement feature and the second key negative polarity enhancement feature; Performing activation function processing on the difference between the second positive similarity and the second negative similarity to obtain a second attention weight; The text-to-visual attention features of this level are obtained based on the second attention weight and the amplitude-adjusted text value vector.
6. A text-guided zero-sample target counting method as claimed in claim 1, 2 or 3, characterized in that: A convolution upsampling module is also connected between the feature fusion modules of two adjacent levels. The convolution upsampling module performs convolution and upsampling processing on the visual-to-text fusion features output by the feature fusion module of the previous level in the two adjacent levels, and uses the processing results as the input of the feature fusion module of the next level in the two adjacent levels.
7. A text-guided zero-sample target counting method as claimed in claim 1, 2 or 3, characterized in that: The visual feature extraction module comprises: Visual encoder, which processes the query image to obtain visual embedding; Visual adapter, which converts visual embedding to specified dimensions to obtain visual features; And / or, the text feature extraction module includes: Text encoder, which processes the target description text to obtain text embedding; Text adapter, converts text embedding to the specified dimension to obtain text features.
8. A text-guided zero-sample target counting method as claimed in claim 1, 2 or 3, characterized in that: The decoder includes a plurality of cascaded decoding convolution upsampling modules and an output convolution layer, and also includes one or more adding units, wherein the one or more adding units are respectively connected in series between one or more adjacent pairs of decoding convolution upsampling modules, and each adding unit corresponds to a feature fusion module of a level; The first-level feature fusion module outputs the visual-to-text fusion feature to the input of the first decoding convolution upsampling module; The adding unit is used to add the output features of the previous decoding convolution upsampling module of the adding unit and the visual to text fusion features output by the feature fusion module of the corresponding level, and input the addition result into the next decoding convolution upsampling module of the adding unit.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of any one of the methods of claims 1 to 8 are implemented.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, A memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor so that the at least one processor can execute a text-guided zero-sample target counting method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Video scene text detection method and system based on deep learning, medium and equipment
CN109919025A
Universal category object counting method, device and equipment and storage medium
CN118397301A
Reference image segmentation method and system based on multi-level feature fusion
CN119049058A
Channel Fusion for Vision-Language Representation Learning
US20240119713A1
Knowledge fusion multi-modal interaction method and apparatus based on improved alignment method
WO2025025290A1
Cited By
Zero sample target counting method, system and device and storage medium
CN121505360A
A zero-shot object counting method, system, device and storage medium
CN121505360B