Object counting method and device for enhancing text guidance by utilizing frequency characteristics
By introducing an adaptive frequency selection module into the text-guided object counting method to process the frequency domain features, combined with the spatial domain features, the problem that existing methods ignore in frequency domain analysis is solved, and the performance and accuracy of object counting are improved.
Patent Information
- Application Number
- CN202510295981.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing text-guided object counting method mainly runs in the spatial domain, ignoring the effect of frequency domain analysis on object counting, resulting in insufficient performance when capturing insignificant periodic patterns or subtle changes in the spatial domain.
By introducing K cascading adaptive frequency selection modules into the decoder, the frequency domain amplitude spectrum and phase spectrum of the visual feature are filtered using the adaptive frequency selector, the relevant frequency components are dynamically emphasized, and they are converted back to the spatial domain, and decoded in combination with the spatial domain features.
Effectively combining spatial domain and frequency domain features, improving the performance of object counting, achieving more robust and detailed frequency domain feature representation, and enhancing the accuracy of object positioning and counting.
Smart Images

Figure CN120219899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer image technology, and in particular, to a method and device for enhancing text-guided object counting by using frequency features. Background Art
[0002] Text-guided object counting aims to count specific objects in an image according to a text description and output a density estimation map, where a point is marked at the center of each object in the density estimation map. In existing text-guided object counting schemes, a visual encoder is used to extract visual features of the image, a text encoder is used to obtain text embeddings of the text description, the visual features and the text embeddings are fused to obtain fused features, and the fused features are input into a decoder to obtain the density estimation map. Existing text-guided object counting schemes mainly operate in the spatial domain and often ignore the role of frequency domain analysis in object counting. In fact, frequency domain features can provide complementary information and can capture periodic patterns or subtle changes in data that are not obvious in the spatial domain. If the decoder can combine these frequency domain features for decoding processing, it will help improve object counting performance. Summary of the Invention
[0003] This application aims to at least solve the technical problems existing in the prior art and provides a method and device for enhancing text-guided object counting by using frequency features.
[0004] In a first aspect, this application provides a method for enhancing text-guided object counting by using frequency features, including: inputting a query image into a visual encoder to obtain visual features; inputting an object description text into a text encoder to obtain text embeddings; using a fusion network to fuse the visual features and the text features to obtain fused features; inputting the fused features into a decoder to obtain a density estimation map of the object, where the decoder includes K cascaded adaptive frequency selection modules and 1 output layer, both K and k are positive integers, k ∈ [1, K], and the k-th adaptive frequency selection module includes a first convolutional unit, an adaptive frequency selector, and a first upsampling unit connected in sequence; wherein, the adaptive frequency selector of the k-th adaptive frequency selection module filters the output features of the first convolutional unit in the amplitude spectrum and the phase spectrum in the frequency domain respectively, and converts the filtering result to the spatial domain.
[0005] The above technical solution: The visual features of the query image and the text embedding of the object description text are fused to obtain a fused feature, which contains rich information in both the text and image modalities. The decoder decodes based on the fused feature. The decoder includes K cascaded adaptive frequency selection modules and 1 output layer. The adaptive frequency selector filters the output features of the first convolutional unit in the amplitude spectrum and phase spectrum in the frequency domain respectively, considering the different roles of the amplitude spectrum and phase spectrum in the object counting task, to achieve a more robust and detailed frequency domain feature representation, dynamically emphasizes relevant frequency components during the decoding process, effectively combines the spatial domain and frequency domain features, and achieves accurate object localization and accurate counting.
[0006] In a preferred embodiment of the present application, the adaptive frequency selector of the k-th adaptive frequency selection module includes: a transformation unit that transforms the output features of the first convolutional unit in the k-th adaptive frequency selection module into the frequency domain to obtain a first frequency domain representation; a decomposition unit that decomposes the first frequency domain representation into an amplitude spectrum and a phase spectrum; a weighting unit that respectively performs weighting processing on the amplitude spectrum and the phase spectrum to obtain a weighted amplitude spectrum and a weighted phase spectrum; a synthesis unit that combines the weighted amplitude spectrum and the weighted phase spectrum to obtain a second frequency domain representation; an inverse transformation module that performs spatial domain conversion and residual processing on the second frequency domain representation to obtain the input features of the first upsampling unit.
[0007] The above technical solution: The first frequency domain representation is decomposed into an amplitude spectrum and a phase spectrum, and the amplitude spectrum and the phase spectrum are respectively filtered with different filtering parameters to adaptively emphasize or suppress the characteristic frequency components. The extracted frequency domain detail features (the second frequency domain representation) are residually connected with the spatial domain features (the input features of the first upsampling unit) through the inverse transformation module, effectively combining the spatial domain features and the frequency domain features.
[0008] In a preferred embodiment of the present application, the weighting unit filters the amplitude spectrum with an amplitude weight and an amplitude bias to obtain a weighted amplitude spectrum, and filters the phase spectrum with a phase weight and a phase bias to obtain a weighted phase spectrum.
[0009] The above technical solution: Linear transformation is performed on the amplitude spectrum and the phase spectrum to facilitate flexibly capturing the additive relationship between the amplitude spectrum and the phase spectrum. The amplitude weight, amplitude bias, phase weight, and phase bias can be learned during model training and can adaptively emphasize or suppress specific frequency components.
[0010] In a preferred embodiment of the present application, the inverse transformation module includes: an inverse transformation unit that transforms the second frequency domain representation to the spatial domain to obtain a first spatial domain representation; an activation function unit that performs an activation function process on the first spatial domain representation to obtain a second spatial domain representation; and a residual connection unit that performs a residual connection between the second spatial domain representation and the output features of the first convolution unit of the k-th adaptive frequency selection module to obtain the input features of the first upsampling unit of the k-th adaptive frequency selection module.
[0011] The above technical solution: introduces non-linear modeling through the activation function, and uses the residual connection unit to perform a residual connection between the second spatial domain representation and the output features of the first convolution unit of the k-th adaptive frequency selection module, ensuring that the model can utilize the adjusted spatial features (the second spatial domain representation) and the original spatial features (the output features of the first convolution unit), and can retain spatial information and promote the gradient flow during training.
[0012] In a preferred embodiment of the present application, the visual features include local patch embeddings and global CLS tokens; the fusion network includes a first three-stream attention fusion module that performs cross-attention processing on the local patch embeddings, global CLS tokens, and text embeddings to obtain a first combined feature, and the fusion feature includes the first combined feature.
[0013] The above technical solution: The global CLS token contains the global context information of the query image. However, existing object counting methods usually ignore the role of the global CLS token, resulting in poor fusion of the global and local features of the query image, affecting the accuracy of object counting. In the present application, the local patch embeddings and the global CLS tokens are used together as visual features to be fused with the text embeddings, which can capture rich and complementary information across multiple modalities. This fusion enhances the model's ability to accurately locate and count the specified objects.
[0014] In a preferred embodiment of the present application, the first three-stream attention fusion module includes: a first cross-attention unit that performs cross-attention operations with the local patch embeddings as the query, the global CLS tokens as the key and value to obtain a first cross-attention feature; a second cross-attention unit that performs cross-attention operations with the local patch embeddings as the query, the text embeddings as the key and value to obtain a second cross-attention feature; and a first weighted fusion unit that performs weighted fusion on the local patch embeddings, the first cross-attention feature, and the second cross-attention feature to obtain a first combined feature.
[0015] The above technical solution: Two cross-attention units, namely the first cross-attention unit and the second cross-attention unit, are used to fuse the local patch embeddings with the global CLS token and the text embedding respectively. The first weighted fusion unit is used to balance the contributions of the outputs of the two cross-attention units, so that the first combined feature effectively combines spatial and text information, enhancing multimodal information fusion. This fusion helps to enhance the model's ability to accurately locate and count specified objects.
[0016] In a preferred embodiment of the present application, the fusion network further includes: an upsampling processing module that performs upsampling processing on the first combined feature output by the first three-stream attention fusion module to obtain an upsampled combined feature; a second three-stream attention fusion module that performs cross-attention processing on the upsampled combined feature, the global CLS token, and the text embedding to obtain a second combined feature. The fusion feature includes the first combined feature and the second combined feature, and the first combined feature serves as the input feature of the first adaptive frequency selection module. The decoder further includes an addition unit located after the k-th adaptive frequency selection module. The addition unit adds the output feature of the k-th adaptive frequency selection module and the second combined feature to obtain an added feature, and uses the added feature as the input feature of the (k + 1)-th adaptive frequency selection module or the input feature of the output layer.
[0017] The above technical solution: The first combined feature is upsampled by the upsampling processing module, and the upsampled combined feature, the global CLS token, and the text embedding are further fused by the second three-stream attention fusion module to obtain information with a different level from the first combined feature, which can enrich the feature representation. By adding the output feature of the k-th adaptive frequency selection module and the second combined feature through the addition unit, the features in the decoding process can be enriched to improve the model's ability to accurately locate and count specified objects.
[0018] In a preferred embodiment of the present application, the second three-stream attention fusion module includes: a third cross-attention unit that performs cross-attention operation with the upsampled combined feature as the query, the global CLS token as the key and value to obtain a third cross-attention feature; a fourth cross-attention unit that performs cross-attention operation with the upsampled combined feature as the query, the text embedding as the key and value to obtain a fourth cross-attention feature; a second weighted fusion unit that performs weighted fusion on the upsampled combined feature, the third cross-attention feature, and the fourth cross-attention feature to obtain a second combined feature.
[0019] The above technical solution: Two cross-attention units, namely the third cross-attention unit and the fourth cross-attention unit, are used to fuse the upsampled combined features with the global CLS token and the text embedding respectively. The second weighted fusion unit is used to balance the contributions of the outputs of the two cross-attention units, enabling the second combined feature to effectively combine spatial and text information, enhancing multimodal information fusion, and this fusion helps to enhance the model's ability to accurately locate and count specified objects.
[0020] In a preferred embodiment of the present application, the upsampling processing module includes a cascaded second convolutional unit and a second upsampling unit.
[0021] In a second aspect, the present application provides an apparatus for enhancing text-guided object counting using frequency features, which is used to implement a method for enhancing text-guided object counting using frequency features in the first aspect. The apparatus includes: a visual feature acquisition module that inputs a query image into a visual encoder to obtain visual features; a text embedding acquisition module that inputs an object description text into a text encoder to obtain text embeddings; a fused feature acquisition module that uses a fusion network to fuse visual features and text features to obtain fused features; a decoding module that inputs the fused features into a decoder to obtain a density estimation map of the object. The decoder includes K cascaded adaptive frequency selection modules and 1 output layer, where K and k are both positive integers, k ∈ [1, K]. The k-th adaptive frequency selection module includes a first convolutional unit, an adaptive frequency selector, and a first upsampling unit connected in sequence. Among them, the adaptive frequency selector of the k-th adaptive frequency selection module performs filtering processing on the amplitude spectrum and phase spectrum of the output features of the first convolutional unit in the frequency domain respectively, and converts the filtering result to the spatial domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic flowchart of an object counting method in a preferred embodiment of the present invention;
[0023] Figure 2 is a model framework diagram relied on by an object counting method in a preferred embodiment of the present invention;
[0024] Figure 3 is a schematic network structure diagram of a first three-stream attention fusion module in a preferred embodiment of the present invention;
[0025] Figure 4 is a visualization result of the object counting method provided by the present invention on the FSC-147 and CARPK datasets. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0027] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0028] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "mounted", "connected", and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0029] The execution subject of a method for enhancing text-guided object counting using frequency features provided by the present invention includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, a method for enhancing text-guided object counting using frequency features provided by the present invention can be executed by software or hardware installed on a terminal device or a server device. The software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0030] The present invention discloses a method for enhancing text-guided object counting using frequency features. In a preferred embodiment, please refer to Figure 1 , including:
[0031] Step S1, input a query image into a visual encoder to obtain visual features.
[0032] In this embodiment, the query image is an image for which object counting is required, and the objects can be living targets and inanimate targets. Exemplarily, the query image can be Figure 2 the image of flying birds in Figure 4 , as well as the images of moving vehicles, egg images, grape images, etc. in Figure 2 . Referring to , when the query image Image is input into the visual encoder p , a visual embedding can be obtained, and the visual embedding is two-dimensionally reshaped to obtain local embedding patches ε p . Exemplarily, the visual features include local embedding patches ε
[0033] Step S2: Input the object description text into the text encoder to obtain a text embedding.
[0034] In this embodiment, the object description text is a text that describes the objects to be counted and may include the object names. Exemplarily, Figure 2 the object description text in Figure 4 is "A photo of{birds}", and the object description texts in Figure 2 are "eggs", "apples", "cars", etc. The text encoder can be an existing BERT encoder or the text encoder of CLIP (Contrastive Language-Image Pretraining). Exemplarily, after the object description text is input into the text encoder t , a text embedding ε
[0035] Step S3: Use a fusion network to fuse the visual features and the text features to obtain a fused feature.
[0036] Exemplarily, the fusion network is not limited to fusing the visual features and the text features by direct splicing, element-level fusion, cross-attention, etc.
[0037] Step S4: Input the fused feature into the decoder to obtain the density estimation map Y pred of the object. Please refer to Figure 2 , and the decoder( Figure 2The Decode in [ ] includes K cascaded adaptive frequency selection modules and 1 output layer, where both K and k are positive integers, k ∈ [1, K]. The k-th adaptive frequency selection module (AFS Block) includes a first convolutional unit, an adaptive frequency selector (AFS), and a first upsampling unit connected in sequence. Among them, the adaptive frequency selector of the k-th adaptive frequency selection module filters the output features of the first convolutional unit in the amplitude spectrum and phase spectrum in the frequency domain respectively, and converts the filtering result to the spatial domain. The first upsampling unit is preferably but not limited to 2-fold upsampling. The features are continuously refined through K cascaded adaptive frequency selection modules to improve the accuracy of the density estimation map.
[0038] In this embodiment, the value range of K can be [1, 10], preferably 4. The structures of each adaptive frequency selection module (AFS Block) are the same. Specifically, the adaptive frequency selector converts the output features of the first convolutional unit to the frequency domain and decomposes them into an amplitude spectrum and a phase spectrum. The filtering process is to emphasize and / or suppress the amplitude and phase corresponding to specific frequencies in the amplitude spectrum and phase spectrum, enabling the model to focus on the features related to the counting task. The filtering parameters in the filtering process can be obtained by learning during model training.
[0039] In a preferred embodiment, please refer to Figure 2 , the adaptive frequency selector AFS of the k-th adaptive frequency selection module includes:
[0040] A transformation unit that transforms the output features of the first convolutional unit in the k-th adaptive frequency selection module (denoted as being spatial domain features) to the frequency domain to obtain a first frequency domain representation That is:
[0041]
[0042] Among them, represents converting the spatial domain features to the frequency domain, not limited to the discrete Fourier transform method.
[0043] A decomposition unit that decomposes the first frequency domain representation into an amplitude spectrum and a phase spectrum The specific decomposition process is:
[0044]
[0045] Among them, represents the real part of the first frequency domain representation , represents the imaginary part of the first frequency domain representation . The amplitude spectrum Capture the intensity and phase spectrum of each frequency component Encode the phase information of each frequency component.
[0046] Weighting unit, respectively for the amplitude spectrum and the phase spectrum Perform weighting processing to obtain a weighted amplitude spectrum and a weighted phase spectrum Specifically, use the learnable amplitude weight W mag to perform weighting processing on the amplitude spectrum Use the learnable phase weight W phase to perform weighting processing on the phase spectrum The amplitude weight W mag and the phase weight W phase Can be learned during model training to achieve adaptive emphasis or suppression of specific frequency components.
[0047] Synthesis unit, combine the weighted amplitude spectrum and the weighted phase spectrum to obtain the second frequency domain representation The specific combination method is as follows:
[0048]
[0049] Inverse transformation module, perform spatial domain conversion and residual processing on the second frequency domain representation to obtain the input features of the first upsampling unit.
[0050] In this embodiment, preferably, please refer to Figure 2 , the weighting unit uses the amplitude weight W mag and the amplitude bias b mag to perform filtering processing on the amplitude spectrum to obtain a weighted amplitude spectrum Use the phase weight W phase and the phase bias b phase to perform filtering processing on the phase spectrum to obtain a weighted phase spectrum The specific filtering processing process is expressed as:
[0051]
[0052] In this embodiment, use different weighting weights and biases to perform weighting processing on the amplitude spectrum and the phase spectrum respectively. The amplitude weight W mag , the amplitude bias b mag , the phase weight W phase and the phase bias b phaseBoth can be learned during model training, adaptively selecting frequency components related to the counting task in the amplitude spectrum and phase spectrum respectively, highlighting the different roles of the amplitude spectrum and phase spectrum in the counting task, and achieving a more robust and detailed feature representation.
[0053] In this embodiment, the Adaptive Frequency Selector (AFS) enhances feature representation by dynamically adjusting frequency components. In computer vision, high frequencies capture fine details, while low frequencies provide broader context information. Spatial domain methods often struggle to fully utilize these patterns. The Adaptive Frequency Selector (AFS) addresses this issue by adaptively emphasizing or suppressing specific frequencies using learnable parameters, enabling the model to focus on task-related features.
[0054] In this embodiment, preferably, refer to Figure 2 , the inverse transform module includes:
[0055] An inverse transform unit that transforms the second frequency domain representation to the spatial domain to obtain the first spatial domain representation which is expressed as:
[0056]
[0057] where, represents transforming the frequency domain representation to the spatial domain, preferably but not limited to using the inverse discrete Fourier transform method.
[0058] An activation function unit that performs activation function processing on the first spatial domain representation to obtain the second spatial domain representation σ(·) is the activation function, preferably but not limited to the GLUE function.
[0059] A residual connection unit that performs a residual connection on the second spatial domain representation and the output feature of the first convolutional unit of the k-th adaptive frequency selection module to obtain the input feature of the first upsampling unit of the k-th adaptive frequency selection module
[0060]
[0061] In a preferred embodiment, for the problem that existing object counting methods usually ignore the role of the global CLS token, resulting in poor fusion of the global and local features of the query image and affecting the object counting accuracy, the visual features are set to include the local patch embedding ε p and the global CLS token ε cls . The global CLS token ε clsThe CLS (Classification) global embedding for the visual encoder is a special token. Its core function is to aggregate the global information of the query image. Through the self-attention mechanism of the visual encoder, the CLS embedding gradually fuses the information of all local image patches, and finally forms a vector that can represent the overall content of the image.
[0062] In this embodiment, please refer to Figure 2 and Figure 3 , the fusion network includes a first three-stream attention fusion module. The first three-stream attention fusion module performs cross-attention processing on the local patch embedding, the global CLS token, and the text embedding to obtain a first combined feature. The fusion feature includes the first combined feature. At this time, the first combined feature is used as the input feature of the first adaptive frequency selection module (AFS Block). The specific structure of the fusion network (Feature Interaction) at this time only includes the solid line part in Figure 2 .
[0063] In this embodiment, please refer to Figure 3 , the first three-stream attention fusion module includes:
[0064] The first cross-attention unit ( Figure 3 The Cross-Attention on the left in p ), using the local patch embedding ε cls as the query Q, the global CLS token ε
[0065] as the key K and value V to perform cross-attention operation to obtain the first cross-attention feature ( Figure 3 The Cross-Attention on the right in p ), using the local patch embedding ε t as the query Q, the text embedding ε
[0066] The first weighted fusion unit, which performs weighted fusion on the local patch embedding ε p , the first cross-attention feature and the second cross-attention feature to obtain the first combined feature ε mix :
[0067]
[0068] Among them, α represents the first fusion weight, which is learned during model training to balance contributions and can take values in the range of [0, 1].
[0069] In a preferred embodiment, to enrich the features in the decoding process and improve the model's ability to accurately locate and count specified objects, refer to the attached Figure 2 dotted part of the Feature Interaction in the figure. The Feature Interaction also includes:
[0070] An upsampling processing module that performs upsampling on the first combined feature output by the first three-stream attention fusion module to obtain an upsampled combined feature ε mix-up .
[0071] A second three-stream attention fusion module that performs cross-attention processing on the upsampled combined feature ε mix-up , the global CLS marker ε cls and the text embedding ε t to obtain a second combined feature ε'. mix At this time, the fused features include the first combined feature ε mix and the second combined feature ε'. mix The first combined feature ε mix serves as the input feature of the first adaptive frequency selection module.
[0072] Please refer to Figure 2 in the figure. The decoder also includes an addition unit located after the k-th adaptive frequency selection module. The addition unit adds the output feature of the k-th adaptive frequency selection module and the second combined feature ε' mix to obtain an added feature, and uses the added feature as the input feature of the (k + 1)-th adaptive frequency selection module (k + 1 ≤ K) or the input feature of the output layer. Preferably, the addition unit is located between the first adaptive frequency selection module and the second adaptive frequency selection module, which facilitates the added second combined feature to continuously refine features through subsequent multi-level adaptive frequency selection modules to improve the accuracy of the density estimation map. When the addition unit is located after the last adaptive frequency selection module, the added feature is used as the input feature of the output layer.
[0073] In this embodiment, the Feature Interaction includes a solid part and a dotted part. Specifically, please refer to Figure 2 in the figure. The upsampling processing module includes a cascaded second convolutional unit and a second upsampling unit. The input end of the second convolutional unit is connected to the output end of the first three-stream attention fusion module, and the output end of the second convolutional unit is connected to the input end of the second upsampling unit. The second upsampling unit is preferably but not limited to 2-fold upsampling.
[0074] In this embodiment, preferably, the second three-stream attention fusion module adopts a network structure similar to that of the first three-stream attention fusion module, specifically including:
[0075] The third cross-attention unit Using the upsampled combined feature ε mix-up As the query Q, and using the global CLS token ε cls As the key K and value V to perform cross-attention operation to obtain the third cross-attention feature
[0076] The fourth cross-attention unit Using the upsampled combined feature ε mix-up As the query Q, and using the text embedding ε t As the key K and value V to perform cross-attention operation to obtain the fourth cross-attention feature
[0077] The second weighted fusion unit, which performs weighted fusion on the upsampled combined feature ε mix-up , the third cross-attention feature and the fourth cross-attention feature to obtain the second combined feature ε' mix :
[0078]
[0079] where β represents the second fusion weight, which is learned during model training and is used to balance the contributions, and the value range can be [0,1].
[0080] The present invention also discloses an apparatus for enhancing text-guided object counting using frequency features, which is used to implement the above-mentioned method for enhancing text-guided object counting using frequency features. In a preferred embodiment, the apparatus includes:
[0081] A visual feature acquisition module, which inputs a query image into a visual encoder to obtain visual features;
[0082] A text embedding acquisition module, which inputs an object description text into a text encoder to obtain text embeddings;
[0083] A fusion feature acquisition module, which uses a fusion network to fuse visual features and text features to obtain fusion features;
[0084] The decoding module inputs the fused features into a decoder to obtain the density estimation map of the object. The decoder includes K cascaded adaptive frequency selection modules and 1 output layer. Both K and k are positive integers, and k ∈ [1, K]. The k-th adaptive frequency selection module includes a first convolutional unit, an adaptive frequency selector, and a first upsampling unit connected in sequence. Among them, the adaptive frequency selector of the k-th adaptive frequency selection module filters the output features of the first convolutional unit in the amplitude spectrum and phase spectrum in the frequency domain respectively, and converts the filtering result to the spatial domain.
[0085] In this embodiment, the visual feature acquisition module, the text embedding acquisition module, the fused feature acquisition module, and the decoding module correspond to step S1, step S2, step S3, and step S4 in the above method for enhancing text-guided object counting using frequency features, respectively, and will not be elaborated here.
[0086] The following introduces the training and verification processes of the model based on the method for enhancing text-guided object counting using frequency features provided by the present invention.
[0087] 1. Construct the architecture of the model (denoted as the FSCNet model) for implementing the method for enhancing text-guided object counting using frequency features:
[0088] Please refer to Figure 2 , the model includes the INOv2 ViTB / 14 pre-trained visual encoder and the Googlebert-base-uncased pre-trained text encoder in the encoding part. Freeze the network parameters of the pre-trained visual encoder and the pre-trained text encoder. The fusion network includes a first three-stream attention fusion module, an upsampling processing module, and a second three-stream attention fusion module, and construct a decoder network (Decode) as shown in Figure 2 .
[0089] 2. Training parameter settings
[0090] To balance the importance of each stream in the first three-stream attention fusion module and the second three-stream attention fusion module, initialize the learnable parameters, the first fusion weight α and the second fusion weight β, to 0.5.
[0091] Dataset Preparation. To evaluate the counting ability of the model, the FSC-147 dataset was used, which is the first large-scale dataset designed for category-agnostic counting and contains 6,135 diverse images from 147 categories. To evaluate the generalization ability of the model, the CARPK dataset was used, which contains 1,448 bird's-eye view images of parking lots with a total of 89,777 vehicles. The images in the dataset were resized to 384×384 and data augmentation was applied according to the CLIP-Count protocol for consistent comparison. A point annotation map was pre-annotated for each image in the dataset, where each object center was marked with a point, and an object description text corresponding to each image was generated.
[0092] Evaluation Metric Determination: The mean absolute error (MAE) and root mean square error (RMSE) were used to evaluate the model performance. A loss function was constructed through the weighted sum of the mean absolute error (MAE) and root mean square error (RMSE).
[0093] Training Settings: The model was trained using the AdamW optimizer with a learning rate configured to 1×10 -4 and a weight decay of 1×10 -2 . A StepLR scheduler was used to reduce the learning rate by a decay factor of 0.33 every 100 epochs. The training was conducted on an NVIDIA RTX4090 GPU for 200 epochs with a batch size of 32.
[0094] 3. Model Training
[0095] The constructed FSCNet model was trained using the dataset, the point annotation map and the object description text of each image in the dataset, and the training parameters set above. The loss function was calculated during the training process, and the model parameters other than the visual encoder and text encoder in the FSCNet model, such as the amplitude weight, amplitude bias, phase weight, phase bias, the first fusion weight α and the second fusion weight β, etc., were optimized through gradient descent according to the loss function. The training was stopped when the training stop condition was reached, and the trained FSCNet model was obtained. The training stop condition is preferably but not limited to the number of training times reaching the preset maximum number of training times, or the value of the loss function converging.
[0096] 4. Comparative Verification of the Object Counting Ability of the Trained FSCNet Model
[0097] Table 1 below shows the quantitative comparison results of the zero-shot counting performance of the FSCNet model obtained by training the present invention and the current relatively advanced object counting methods on the FSC-147 dataset. FSCNet is superior to previous methods in various evaluation metrics, especially in terms of validation and test RMSE. Utilizing the text-guided method and a powerful model architecture, FSCNet has achieved significant improvements in the zero-shot setting. Compared with the previously best-performing method VA-Count, the average MAE has increased by 19.75% and the average RMSE has increased by 39.86%.
[0098] It is worth noting that the performance of FSCNet is close to that of the few-shot method LOCA, which is generally considered the upper limit of text-guided methods. This result highlights the effectiveness of FSCNet in handling various counting tasks without the need for additional visual samples, thus narrowing the gap between zero-shot and few-shot methods.
[0099] Table 1 Quantitative comparison results of zero-shot counting performance on the FSC-147 dataset
[0100]
[0101] To evaluate the generalization ability of FSCNet on new datasets, experiments were conducted on the CARPK dataset, and the quantitative results are shown in Table 2. For a fair comparison, like previous methods, FSCNet was trained on the FSC-147 dataset and directly evaluated on the CARPK dataset without any fine-tuning. From the experimental results, the MAE of FSCNet is 10.08 and the RMSE is 12.20. FSCNet has significantly superior performance compared to existing object counting methods, demonstrating its strong generalization ability on different datasets.
[0102] Table 2 Cross-dataset evaluation of the CARPK dataset
[0103]
[0104]
[0105] By overlaying the predicted density estimation map on the input query image, the qualitative results on the FSC-147 dataset were visualized, as Figure 4 (a) shows. These visualizations highlight the ability of the FSCNet model to accurately capture the spatial distribution of objects in different scenarios. In addition, Figure 4(b) shows the qualitative results of the FSCNet model on the CARPK dataset, demonstrating the effectiveness of the object counting method provided by this application in object localization and counting for various categories, shapes, sizes, and densities. The density estimation map proves the robustness of the object counting method provided by this application, effectively handling challenging scenarios such as overlapping objects and variations in object scale and appearance. These results highlight the generality of the object counting method provided by this application across different datasets and object distributions.
[0106] It can be seen that the frequency-selective counting network (FSCNet) introduced by the object counting method and device provided by this application solves the key limitations of existing methods by leveraging spatial and frequency domain features. By integrating the three-stream attention fusion module (TSAFM) and the adaptive frequency selector (AFS), FSCNet effectively combines global, local, and textual features while dynamically emphasizing the frequency components relevant to the counting task. Experimental results on the FSC-147 and CARPK benchmarks show that FSCNet not only achieves state-of-the-art accuracy but also improves robustness and adaptability in different scenarios. These advancements highlight the potential of frequency domain analysis and multimodal fusion in improving counting accuracy, establishing FSCNet as a new benchmark for future text-guided object counting research.
[0107] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation", "one preferred implementation", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0108] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A method for object counting using frequency features to enhance text guidance, characterized in that: include: Input the query image to the visual encoder to obtain visual features; Input object description text to text encoder to get text embedding; The fusion network is used to fuse visual features and text features to obtain fusion features; The fused features are input into a decoder to obtain a density estimation map of the object. The decoder includes K cascaded adaptive frequency selection modules and 1 output layer, K and k are both positive integers, k∈[1,K], and the kth adaptive frequency selection module includes a first convolution unit, an adaptive frequency selector, and a first upsampling unit connected in sequence; wherein the adaptive frequency selector of the kth adaptive frequency selection module filters the amplitude spectrum and the phase spectrum of the output features of the first convolution unit in the frequency domain, respectively, and converts the filtering results to the spatial domain.
2. The object counting method using frequency feature enhanced text guidance as claimed in claim 1, characterized in that: The adaptive frequency selector of the kth adaptive frequency selection module includes: A transformation unit, which transforms the output feature of the first convolution unit in the kth adaptive frequency selection module into the frequency domain to obtain a first frequency domain representation; a decomposition unit, decomposing the first frequency domain representation into an amplitude spectrum and a phase spectrum; A weighting unit performs weighted processing on the amplitude spectrum and the phase spectrum to obtain a weighted amplitude spectrum and a weighted phase spectrum; A synthesis unit, combining the weighted amplitude spectrum and the weighted phase spectrum to obtain a second frequency domain representation; The inverse transform module performs spatial domain transformation and residual processing on the second frequency domain representation to obtain the input features of the first upsampling unit.
3. The object counting method using frequency feature enhanced text guidance as claimed in claim 2, characterized in that: The weighting unit performs filtering processing on the amplitude spectrum using the amplitude weight and the amplitude offset to obtain a weighted amplitude spectrum, and performs filtering processing on the phase spectrum using the phase weight and the phase offset to obtain a weighted phase spectrum.
4. The object counting method using frequency feature enhanced text guidance as claimed in claim 2, characterized in that: The inverse transformation module comprises: an inverse transform unit, converting the second frequency domain representation into the spatial domain to obtain the first spatial domain representation; An activation function unit, performing activation function processing on the first spatial domain representation to obtain a second spatial domain representation; The residual connection unit performs a residual connection on the second spatial domain representation and the output features of the first convolution unit of the kth adaptive frequency selection module to obtain the input features of the first upsampling unit of the kth adaptive frequency selection module.
5. The object counting method using frequency feature enhanced text guidance according to any one of claims 1 to 4, characterized in that: The visual features include local patch embedding and global CLS labeling; The fusion network includes a first three-stream attention fusion module, which performs cross-attention processing on local patch embedding, global CLS tagging and text embedding to obtain a first combined feature, and the fused feature includes the first combined feature.
6. The object counting method using frequency feature enhanced text guidance as claimed in claim 5, characterized in that: The first three-stream attention fusion module includes: The first cross-attention unit uses the local patch embedding as the query and the global CLS tag as the key and value to perform a cross-attention operation to obtain the first cross-attention feature; The second cross attention unit uses the local patch embedding as the query and the text embedding as the key and value to perform a cross attention operation to obtain the second cross attention feature; The first weighted fusion unit performs weighted fusion on the local patch embedding, the first cross-attention feature and the second cross-attention feature to obtain a first combined feature.
7. The object counting method using frequency feature enhanced text guidance as claimed in claim 5, characterized in that: The fusion network also includes: An upsampling processing module performs upsampling processing on the first combined feature output by the first three-stream attention fusion module to obtain an upsampled combined feature; A second three-stream attention fusion module performs cross-attention processing on the upsampled combined feature, the global CLS tag and the text embedding to obtain a second combined feature, wherein the fused feature includes the first combined feature and the second combined feature, and the first combined feature is used as an input feature of the first adaptive frequency selection module; The decoder also includes an adding unit located after the kth adaptive frequency selection module, which adds the output feature of the kth adaptive frequency selection module and the second combined feature to obtain an added feature, and uses the added feature as an input feature of the k+1th adaptive frequency selection module or an input feature of the output layer.
8. The object counting method using frequency feature enhanced text guidance as claimed in claim 5, characterized in that: The second and third stream attention fusion modules include: The third cross attention unit uses the above sampled combined feature as the query and the global CLS tag as the key and value to perform a cross attention operation to obtain the third cross attention feature; A fourth cross attention unit, performing a cross attention operation with the above sampled combined feature as a query and the text embedding as a key and a value to obtain a fourth cross attention feature; The second weighted fusion unit performs weighted fusion on the upsampled combined feature, the third cross-attention feature and the fourth cross-attention feature to obtain the second combined feature.
9. The object counting method using frequency feature enhanced text guidance as claimed in claim 7 or 8, characterized in that: The up-sampling processing module includes a cascaded second convolution unit and a second up-sampling unit.
10. An object counting device guided by frequency feature enhanced text, used to implement the object counting method guided by frequency feature enhanced text according to any one of claims 1 to 9, characterized in that: include: The visual feature acquisition module inputs the query image into the visual encoder to obtain visual features; The text embedding acquisition module inputs the object description text to the text encoder to obtain the text embedding; The fusion feature acquisition module uses the fusion network to fuse visual features and text features to obtain fusion features; A decoding module inputs the fused features into a decoder to obtain a density estimation map of the object. The decoder includes K cascaded adaptive frequency selection modules and an output layer, where K and k are both positive integers, k∈[1,K]. The kth adaptive frequency selection module includes a first convolution unit, an adaptive frequency selector, and a first upsampling unit connected in sequence; wherein the adaptive frequency selector of the kth adaptive frequency selection module performs filtering processing on the amplitude spectrum and phase spectrum of the output features of the first convolution unit in the frequency domain, respectively, and converts the filtering processing results into the spatial domain.
Citation Information
Patent Citations
High-resolution dense target counting method based on convolutional neural network
CN113239904A
Vision-based cable structure health monitoring method and system
CN113421224A
Speaker counting method and device based on deep learning, equipment and storage medium
CN113903328A
Crowd counting system and method based on cross-modal feature alignment fusion
CN117315428A
Space domain-frequency domain combined enhanced steel member defect size intelligent measuring system
CN118761959A
Cited By
Extreme illumination-oriented visible light and infrared image fusion method, system and device based on text-guided spatial frequency domain interaction and medium
CN121214128A