Target counting method and device based on multi-modal large model
Through the target counting method based on multimodal large model, the number prompt words are used to generate fusion features, which solves the problem of the dependence of the existing technology on sample samples and performance degradation in the new environment, and achieves higher adaptability and efficiency.
Patent Information
- Application Number
- CN202510184100.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-01
AI Technical Summary
The existing target counting technology is highly dependent on sample samples and has significantly reduced performance in new environments, making it difficult to meet the generality and flexibility requirements of highly automated scenarios.
Using the target counting method based on the multimodal large model, a target counting model including prompt word generation, feature extraction, feature fusion and feature decoding modules is constructed, and a number of prompt words are used to generate fusion features to predict the predicted value of the target number.
It improves the adaptability and performance of the model to target counting tasks, reduces the cost of manual intervention, enhances the practicality of the model in multiple scenarios, breaks the dependence on target samples, and significantly improves the deployment efficiency of the general counting algorithm.
Smart Images

Figure CN120234556A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly relates to a target counting method and device based on a multimodal large model. Background Art
[0002] As one of the important research directions in computer vision, target counting has been widely applied in fields such as industrial production, smart agriculture, retail, and traffic monitoring. For example, in a production line, an object counting algorithm can monitor the production quantity and speed in real time; in the breeding industry, it is used to monitor the number of animals in real time; in crowded public places, it is used to count the number of people to ensure safety. However, most existing target counting technologies are specific to certain categories. Usually, users need to provide sample examples of the target category to clarify the counting object, and then count the number of instances semantically similar to the example in the picture as the prediction result. This example-dependent approach is not practically feasible in scenarios with a high degree of automation, especially in dynamic and diverse application environments. In addition, existing technologies have insufficient robustness in cross-domain scenarios, and their performance drops significantly when migrated to new scenarios, usually requiring retraining the model, which further limits their practicality.
[0003] In summary, the main defects of the current technology include a high dependence on example samples and a significant drop in performance in new environments, which make it difficult for existing methods to meet the requirements of generality and flexibility in highly automated scenarios. Summary of the Invention
[0004] This application provides a target counting method based on a multimodal large model to solve the problems of high dependence on example samples and significant drop in performance in new environments in existing target counting technologies.
[0005] Correspondingly, this application also provides a target counting device based on a multimodal large model, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.
[0006] To solve the above technical problems, this application discloses a target counting method based on a multimodal large model, and the method includes:
[0007] Construct a target counting model; the target counting model includes a prompt word generation module, a feature extraction module, a feature fusion module, and a feature decoding module;
[0008] In the prompt word generation module, generate a number prompt word according to the category text and the number of targets corresponding to the query image;
[0009] Use the feature extraction module to extract the text feature of the number prompt word and the image feature of the query image;
[0010] The text features and image features are fused using a feature fusion module to generate fused features;
[0011] The fused features are predicted using a feature decoding module to obtain a predicted density map, and a predicted value of the target quantity is obtained based on the predicted density map.
[0012] This application also discloses an object counting device based on a multimodal large model, and the device includes:
[0013] A prompt word generation module, configured to generate a quantity prompt word according to the category text corresponding to the query image and the number of targets;
[0014] A feature extraction module, configured to extract the text features of the quantity prompt word and the image features of the query image;
[0015] A feature fusion module, configured to fuse the text features and the image features to generate fused features;
[0016] A feature decoding module, configured to predict the fused features to obtain a predicted density map, and obtain a predicted value of the target quantity based on the predicted density map.
[0017] This application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, one or more methods described in this application are implemented.
[0018] This application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods described in this application are implemented.
[0019] In this application, the language modality is introduced. The category of the target to be counted is indicated by the quantity prompt word. The text features of the quantity prompt word and the image features of the query image are extracted using the feature extraction module, effectively improving the model's ability to extract quantity-related features, and fundamentally enhancing the model's adaptability and performance for the target counting task. Then, the text features and the image features are fused to generate fused features, a predicted density map is generated according to the fused features, and a predicted value of the target quantity is obtained based on the predicted density map. It can efficiently predict the quantity and distribution of the target, thereby reducing the cost of manual intervention and improving the practicality of the model in multiple scenarios. This application breaks the dependence of traditional counting techniques on target samples, significantly improves the deployment efficiency of general counting algorithms in automated systems, and makes target counting more convenient. Moreover, due to the diversity of the expression of the quantity prompt word, this method has stronger generalization ability and can adapt to the target counting requirements of any category, providing a solid technical support for complex and diverse target counting tasks.
[0020] Additional aspects and advantages of the present application will be given in the following description section, which will become apparent from the following description, or can be learned through the practice of the present application. Description of the Drawings
[0021] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0022] Figure 1 is a flowchart of the object counting method based on a multimodal large model provided for an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of generating number prompting words provided for an embodiment of the present application;
[0024] Figure 3 is an overall framework diagram of the object counting method based on a multimodal large model provided for an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of the T2C adapter provided for an embodiment of the present application;
[0026] Figure 5 is a schematic structural diagram of the object counting device based on a multimodal large model provided for an embodiment of the present application;
[0027] Figure 6 is a schematic structural diagram of the electronic device provided for an embodiment of the present application. Detailed Embodiments
[0028] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present application and should not be construed as limiting the present application.
[0029] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0030] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0031] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this. For the technical problems existing in the prior art, the object counting method and device based on a multimodal large model provided by the present application are intended to solve at least one of the technical problems of the prior art.
[0032] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0033] The embodiments of the present application provide a possible implementation manner, such as Figure 1 As shown, a flowchart of an object counting method based on a multimodal large model is provided. This solution can be executed by any electronic device. Optionally, it can be executed on the server side or the terminal device.
[0034] As Figure 1 shown, the method may include the following steps:
[0035] Step 101, construct an object counting model; the object counting model includes a prompt word generation module, a feature extraction module, a feature fusion module, and a feature decoding module;
[0036] Step 102, in the prompt word generation module, generate a number prompt word according to the category text corresponding to the query image and the number of objects.
[0037] In the embodiments of the present application, the number prompt word is a text description method that combines the object category and the number of objects, which can clarify the object category to be counted and provide number-related information.
[0038] Step 103: Use the feature extraction module to extract the text features of the number cue word and the image features of the query image.
[0039] Step 104: Use the feature fusion module to fuse the text features and the image features to generate fused features.
[0040] In the embodiment of the present application, the vision-language model extracts features, using the number cue word and the query image as inputs. The number cue word injects counting awareness into the encoder of the vision-language model by providing number-related information, which can effectively improve the ability to encode and extract visual features related to the number of objects, significantly enhance the accuracy of the counting task, and make it more suitable for general category object counting tasks. In the feature fusion module, the category-related information in the number cue word is fused with the image features to obtain fused features.
[0041] Step 105: Use the feature decoding module to predict a predicted density map from the fused features, and obtain a predicted value of the target quantity based on the predicted density map.
[0042] In the embodiment of the present application, given a query image I ∈ R H×W×3 , a category text C, a target number a p and a density map D, the aim is to predict a density map D * with the smallest error from D. First, according to the category text C and the target number a p , a corresponding number cue word P Q is generated. After generating the cue word, the vision-language model is used to extract the text feature F Q and the image feature F I respectively, and then the text feature and the image feature are fused to obtain the fused feature F V . Then a preset decoder is used to predict the density map D * ∈ R H×W×1 of the target category. Summing the pixel values of the density map can obtain the predicted value of the target quantity.
[0043] In the embodiments of the present application, a language modality is introduced. The number of target categories to be counted is indicated by number prompts. The text features of the number prompts and the image features of the query image are extracted using a feature extraction module, effectively improving the model's ability to extract number-related features and fundamentally enhancing the model's adaptability and performance for the target counting task. Subsequently, the text features and image features are fused to generate fused features, a prediction density map is generated based on the fused features, and a predicted value of the target quantity is obtained based on the prediction density map. It can efficiently predict the quantity and distribution of targets, thereby reducing the cost of manual intervention and improving the practicality of the model in multiple scenarios. The embodiments of the present application break the dependence of traditional counting techniques on target samples, significantly improving the deployment efficiency of general counting algorithms in automated systems and making target counting more convenient. Moreover, due to the diversity of the expressions of number prompts, this method has stronger generalization ability and can adapt to the target counting requirements of any category, providing solid technical support for complex and diverse target counting tasks.
[0044] In an alternative embodiment, the number prompts include factual text prompts constructed using the correct number of targets and counterfactual text prompts constructed using incorrect numbers of targets.
[0045] In the embodiments of the present application, according to the category text C and the number of targets a p , the corresponding number prompt P is generated Q . In the training phase, the embodiments of the present application use a specific template "{a photo of [number] [class]}" to generate number prompts, where [number] is replaced by the number of targets and [class] is replaced by the category text. The number prompts are divided into two categories:
[0046] · Factual text prompts: Number prompts that are consistent with the actual number of targets in the query image.
[0047] · Counterfactual text prompts: Number prompts that are inconsistent with the actual number of targets in the query image.
[0048] For example Figure 2 as shown, the factual text prompt generated is "A photo of 15 balloons", and the counterfactual text prompts generated are "A photo of 13 balloons", "A photo of 17 balloons".
[0049] An image-text pair composed of a query image and a factual text prompt is called a positive sample pair (positive sample), while an image-text pair composed of a counterfactual text prompt is called a negative sample pair (negative sample). These sample pairs are used to optimize the model performance during training, and number awareness is injected into the model by distinguishing positive and negative sample pairs.
[0050] In an alternative embodiment, the feature extraction module includes a text encoder and an image encoder; the text features of the number prompt and the image features of the query image are extracted using the feature extraction module, including:
[0051] Use the text encoder to extract features from the number prompt to obtain text features;
[0052] Use the image encoder to extract features from the query image to obtain image features;
[0053] Among them, the text encoder uses the BERT model, and the image encoder uses the DINOv2 model.
[0054] As Figure 3 shown, in the embodiment of the present application, factual text prompts and counterfactual text prompts are first generated, where the factual text prompt is such as "A photo of 14 kiwis", and the counterfactual text prompts are such as "A photo of 10 kiwis", "A photo of 12 kiwis", "A photo of 16 kiwis", "A photo of 18 kiwis". After generating the above number prompts, the text features F T are extracted through the text encoder Ψ of the vision-language model Q ∈R 1×d . At the same time, the image features I are extracted from the query image I using the image encoder Ψ where p represents the image patch size and d represents the dimension of the embedding feature.
[0055] In an alternative embodiment, the text features and the image features are fused using the feature fusion module to generate fused features, including:
[0056] Extract category-related information from the text features;
[0057] Use the cross-attention mechanism to inject the category-related information into the image features to obtain fused features.
[0058] As Figure 3 shown, in order to accurately predict the target quantity of a specified category, the embodiment of the present application needs to inject category information into the image feature F IIn the training process, to reduce the computational complexity, category-related information is extracted from the numerical prompts because the [class] tag is used when constructing the numerical prompts; in the inference process, the template {aphoto of [class]} is used to construct category prompts and extract category-related information F C Inject the category-related information F C into the image feature F I to obtain the fused feature In the embodiment of the present application, the feature fusion of two modalities is realized through the cross-attention module. In the cross-attention module, the cross-attention mechanism (Cross-Attention) is used to inject the category-related information F C into the image feature F I to obtain the fused feature F V , and its formula is expressed as:
[0059] F V = F I + MLP&LN(CA(F I , F C ))
[0060] Among them, CA(x, y) represents the cross-attention calculation with x as the query, y as the key and value. MLP&LN(x) represents the non-linear mapping operation through the multi-layer perceptron (MLP) after layer normalization (LayerNormalization) of x. Through the above mechanism, the category information is effectively incorporated into the image feature, and the generated fused feature has stronger category perception ability.
[0061] In an optional embodiment, the feature decoding module includes a two-stream adaptive counting decoder; the two-stream adaptive counting decoder includes a CNN stream, a Transformer stream, a T2C adapter and a gating network;
[0062] Using the feature decoding module to predict the fused feature to obtain the predicted density map, including:
[0063] Use the CNN stream to capture the local detail information in the fused feature to obtain the first intermediate feature;
[0064] Use the Transformer stream to extract the global information in the fused feature to obtain the second intermediate feature;
[0065] Use the T2C adapter to transfer the second intermediate feature to the CNN stream and fuse it with the first intermediate feature to obtain the updated first intermediate feature;
[0066] Generate a first density map and a second density map based on the updated first intermediate feature and second intermediate feature respectively;
[0067] Fuse the first density map and the second density map through a gating network to obtain the predicted density map of the target class.
[0068] As Figure 3 shown, the embodiment of the present application designs a Dual-stream Adaptive Counting decoder (DAC-decoder) to predict the density map D * ∈R H×W×1 of the target class. The DAC-decoder consists of two branches (CNN stream and Transformer stream, and the Transformer stream is hereinafter simply referred to as the Trans stream), a T2C (Transformer-to-CNN) adapter, and a gating network. The CNN stream is good at capturing local detail information, while the Trans stream focuses on extracting global information. By combining the advantages of the two streams, the DAC-decoder can significantly improve the accuracy and robustness of density map prediction. The fused feature F V is input into the two decoding branches simultaneously, and the first intermediate feature and the second intermediate feature are obtained respectively. Subsequently, the T2C adapter transfers the global information of the Trans stream to the CNN stream to enhance its ability to understand global counting information. The outputs of the two streams generate the predicted density maps and through a 1×1 convolutional layer. Finally, the two density maps are fused through the gating network to obtain the final prediction result D * .
[0069] In an optional embodiment, as Figure 4 shown, the T2C adapter includes a Channel-Excitation block (CE) and a Cross-Attention block (CA);
[0070] Using the T2C adapter to transfer the second intermediate feature into the CNN stream and fuse it with the first intermediate feature to obtain the updated first intermediate feature, including:
[0071] In the channel weight calibration module, use a multi-layer perceptron to project the second intermediate feature onto the channel dimension of the first intermediate feature in the CNN stream to generate channel recalibration weights;
[0072] In the cross-attention module, adjust the first intermediate feature and the second intermediate feature through the cross-attention mechanism to obtain cross-attention features;
[0073] Generate adaptive features based on channel recalibration weights and cross-attention features;
[0074] Perform an element-wise addition operation on the adaptive features and the first intermediate features to generate updated first intermediate features.
[0075] In the embodiment of the present application, the CE module projects the second intermediate features of the Trans stream to the channel dimension of the first intermediate features of the CNN stream to generate channel recalibration weights. Specifically, it is expressed as:
[0076]
[0077] Among them, represents an element-wise multiplication operation at the channel level, and MLP·ReLU·MLP(·) represents a mapping projection operation to align the channel recalibration weights with the dimension of the first intermediate features of the CNN stream, and ReLU is an activation function.
[0078] The CA module further adjusts the features through a cross-attention mechanism, using as the query (Query), as the key (Key) and value (Value), and the formula is:
[0079]
[0080] Among them, LN is a layer normalization function, and CA represents cross-attention calculation.
[0081] Through the channel weight calibration module and the cross-attention module, adaptive features F T2C = F CA + F CE can be generated. Finally, an element-wise addition operation is performed on the adaptive features and the original first intermediate features of the CNN stream to generate updated first intermediate features This design ensures that even if the Trans stream does not provide effective adaptation information, the CNN stream can still maintain its original features.
[0082] In the embodiment of the present application, the two-stream adaptive counting decoder integrates the global information capture ability of the Transformer stream and the local detail extraction advantage of the CNN stream, thereby improving the robustness and applicability of the model in cross-domain applications, making the entire method not only able to accurately predict the number of targets, but also have strong generalization ability and flexibility.
[0083] In an alternative embodiment, the first density map and the second density map are fused through a gating network to obtain a predicted density map of the target class, including:
[0084] Dynamically allocate weights for the first density map and the second density map through a gating network, and fuse the first density map and the second density map based on the weights to obtain a final predicted density map: where w is the weight predicted by the gating network based on the original fused feature F V The gating network takes the fused feature as input and realizes the complementary advantages of the two decoding streams by learning reasonable weight allocation to generate the final predicted density map.
[0085] In an optional embodiment, after using a feature decoding module to predict a predicted density map from the fused feature and obtaining a predicted value of the target quantity based on the predicted density map, the method further includes:
[0086] Optimizing the target counting model through a loss function;
[0087] where the loss function includes a counting loss function, a two-stream cross-number ranking loss function, and an image-text number alignment loss function;
[0088] The counting loss function is used to constrain the difference between the predicted density map and the true density map;
[0089] The two-stream cross-number ranking loss function is used to rank the sum of pixel values in the sub-regions of the first density map and the second density map to establish a relative relationship constraint of local density values;
[0090] The image-text number alignment loss function is used to establish a similarity comparison relationship constraint between negative sample pairs and positive sample pairs; the positive sample pair is a sample pair generated according to the factual text prompt and the query image; the negative sample pair is a sample pair generated according to the counterfactual text prompt and the query image.
[0091] In the embodiments of the present application, the counting loss is used as the main optimization target, the two-stream cross-number ranking loss focuses on optimizing the performance of the two-stream adaptive counting decoder, and the image-text number alignment loss further improves the encoder's perception ability of number information. By designing the synergistic effect of multiple loss functions, the prediction performance and robustness of the model are improved.
[0092] In terms of the counting loss function, the mean square error (MSE) of the pixels of the true density map D and the predicted density map D * , and is used as the counting loss for optimization, and the formula is:
[0093]
[0094] Among them, λ1 and λ2 respectively represent the weights of the losses between the prediction results of the Transformer stream and the CNN stream (i.e., the corresponding second density map and the first density map) and the ground truth density map, which are used to balance the influence among the three. This loss function ensures the accuracy of the model in the density map estimation task by directly constraining the difference between the predicted density map and the ground truth density map.
[0095] In terms of the two-stream cross-number ranking loss function, in order to further utilize the complementarity of the two-stream decoder, the embodiment of the present application proposes a two-stream cross-number ranking loss L rank , as Figure 3 shown, this loss establishes a relative relationship constraint of local density values by sorting the sum of pixel values in the sub-regions of the predicted density maps of the Transformer stream and the CNN stream.
[0096] Specifically, first, the true local density value ranking needs to be obtained. The ground truth density map D is evenly divided into n non-overlapping sub-regions, and the sum of pixel values in each sub-region is calculated and sorted in descending order to obtain the sorting index. Using this index and the density maps predicted by the two streams (divided in the same way as the ground truth density map), two sequences and
[0097] For any two adjacent elements v i and v i+1 from the same sequence, the v i ranked in front should theoretically be not less than v i+1 , that is, v i ≥v i+1 . This ranking constraint exists within the CNN stream and the Transformer stream; introducing this ranking constraint between the two streams can further strengthen the cross-stream counting consistency. The embodiment of the present application calls the loss function derived from this ranking constraint the two-stream cross-number ranking loss, which is defined as:
[0098]
[0099] where l is the sequence interval at which the loss is generated. In the embodiment of the present application, l = 5 is taken. If l is too small, a small error in the prediction may flip the order and lead to an inconsistency between the predicted order and the true order, making it too sensitive; when l is too large, this ranking constraint is easily satisfied, and the loss function is difficult to play a role.
[0100] Through the constraint of the L rank loss, the two-stream decoder can better capture the local distribution characteristics of the target and the global counting relationship, thereby further improving the accuracy and stability of the model prediction results.
[0101] In the embodiments of the present application, the dual-stream adaptive counting decoder fuses the global features of the Transformer stream and the local features of the CNN stream, and uses the dual-stream cross-number sorting loss function to ensure the collaborative optimization of the two streams, thereby obtaining more accurate and robust counting results.
[0102] In terms of the image-text quantity alignment loss function, in the embodiments of the present application, the prompt word constructed with the correct target quantity is called the factual text prompt word, which forms a positive sample pair with the query image; while the prompt words constructed with other incorrect quantities are counterfactual text prompt words, which form a negative sample pair with the query image. In fact, the negative sample pair is not a completely negative sample pair, and there is a positive category description in the counterfactual description. Therefore, this kind of prompt is considered to be a sufficiently difficult negative sample, and the learning of difficult negative samples is crucial for improving the model performance. Therefore, the quantity prompt used in the embodiments of the present application can improve the extraction of quantity-related features by the image encoder through the image-text quantity alignment loss function, thereby further improving the counting performance of the model. Specifically, the similarity s of the positive sample pair (image and factual prompt) P should be higher than the similarity of the negative sample pair (image and counterfactual prompt) At the same time, the similarity of the negative sample pair should not be completely zero, but is sorted according to the difference between the counterfactual quantity prompt word and the correct quantity: the larger the difference, the lower the similarity of the negative sample pair.
[0103] Assume that the correct target quantity is a P , and a set of counterfactual quantities is constructed through the value interval Δ Δ is a suitable interval determined by a P . According to the difference from a P , the negative sample pair similarity sequence is sorted in ascending order Then the alignment loss is:
[0104]
[0105] This alignment mechanism effectively transmits the quantity information in the text prompt word to the image representation of the model by carefully depicting the positive and negative sample relationships.
[0106] In the embodiments of the present application, the overall loss is designed as the weighted sum of the three losses mentioned above:
[0107] L = L count + μ1L rank + μ2L align
[0108] Among them, μ1 and μ2 are the weights of the two-stream cross-number ranking loss and the image-text number alignment loss, which are used to adjust the impact of each part on the overall optimization. Through multiple constraints on counting accuracy, sorting relationship and number alignment, the model can complete the target counting task more accurately and have stronger generalization performance.
[0109] like Figure 2 and Figure 3 As shown, the embodiment of the present application generates positive and negative image-text pairs through number prompt words and query images, and through comparative learning of the image-text quantity alignment loss function in the feature space of the visual language model encoder, effectively improves the ability of the image encoder to extract number-related features, improves the accuracy of the decoder's predicted density map (DensityMap), and thus improves the counting performance.
[0110] The method in the embodiment of the present application shows excellent performance in the field of general target counting, which not only broadens the application boundary of target counting technology, but also provides new development directions for fields such as warehouse automation counting, especially in industrial production, smart logistics, modern agriculture and other fields. By reducing example dependence, optimizing decoder design and innovating number prompt word mechanism, the embodiment of the present application lays a solid foundation for target counting technology in more scenarios in the future, and at the same time promotes the further integration and development of computer vision and deep learning technology in automation systems.
[0111] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application also provides a target counting device based on a multimodal large model, such as Figure 5 As shown, the device comprises:
[0112] A prompt word generation module 501 is used to generate a number prompt word according to the category text and the number of targets corresponding to the query image;
[0113] A feature extraction module 502, for extracting text features of the number prompt words and image features of the query image;
[0114] A feature fusion module 503 is used to fuse text features and image features to generate fused features;
[0115] The feature decoding module 504 is used to predict the fused features to obtain a predicted density map, and obtain a predicted value of the target quantity based on the predicted density map.
[0116] In the embodiments of the present application, a language modality is introduced. The target category to be counted is indicated by a number prompt word. The text features of the number prompt word and the image features of the query image are extracted using a feature extraction module, effectively improving the model's ability to extract number-related features, fundamentally enhancing the model's adaptability and performance for the target counting task. Subsequently, the text features and image features are fused to generate fused features, a predicted density map is generated based on the fused features, and a predicted value of the target quantity is obtained based on the predicted density map. It can efficiently predict the quantity and distribution of the target, thereby reducing the cost of manual intervention and improving the practicality of the model in multiple scenarios. The embodiments of the present application break the dependence of traditional counting techniques on target samples, significantly improving the deployment efficiency of general counting algorithms in automated systems and making target counting more convenient. Moreover, due to the diversity of the expressions of the number prompt words, this method has stronger generalization ability and can adapt to the target counting requirements of any category, providing a solid technical support for complex and diverse target counting tasks.
[0117] The target counting device based on the multi-modal large model provided by the embodiments of the present application can implement Figures 1 to 4 each process implemented in the method embodiments. To avoid repetition, it will not be elaborated here.
[0118] The target counting device based on the multi-modal large model in the embodiments of the present application can execute the target counting method based on the multi-modal large model provided by the embodiments of the present application. Their implementation principles are similar. The actions performed by each module and unit in the target counting device based on the multi-modal large model in the embodiments of the present application correspond to the steps in the target counting method based on the multi-modal large model in the embodiments of the present application. For the detailed functional descriptions of each module of the target counting device based on the multi-modal large model, reference can specifically be made to the descriptions in the corresponding target counting method based on the multi-modal large model shown in the foregoing text, which will not be elaborated here.
[0119] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the target counting method based on the multi-modal large model shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the target counting method based on the multi-modal large model provided by the present application breaks the dependence of traditional counting techniques on target samples, significantly improving the deployment efficiency of general counting algorithms in automated systems and making target counting more convenient. Moreover, due to the diversity of the expressions of the number prompt words, this method has stronger generalization ability and can adapt to the target counting requirements of any category, providing a solid technical support for complex and diverse target counting tasks.
[0120] In an optional embodiment, an electronic device is also provided, such asFigure 6 As shown Figure 6 As shown, the electronic device 600 may be a server, including: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.
[0121] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0122] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0123] The memory 603 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0124] The memory 603 is used to store the application program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.
[0125] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0126] The server provided by this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0127] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0128] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0129] It should be noted that the above-mentioned computer-readable storage medium in the present application may also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0130] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately without being assembled into the electronic device.
[0131] The above-mentioned computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.
[0132] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the target counting method and apparatus based on a multimodal large model provided in the above various alternative implementation manners.
[0133] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, executed as a stand-alone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0135] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the prompt word generation module can also be described as "a prompt word generation module for generating number prompt words according to the category text corresponding to the query image and the target number".
[0136] The above description is only the preferred embodiment of the present application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A target counting method based on a multimodal large model, characterized in that: The method comprises: Constructing a target counting model; the target counting model includes a prompt word generation module, a feature extraction module, a feature fusion module and a feature decoding module; In the prompt word generation module, a number prompt word is generated according to the category text and the number of targets corresponding to the query image; Using the feature extraction module to extract text features of the number prompt word and image features of the query image; Using the feature fusion module to fuse the text feature and the image feature to generate a fusion feature; The feature decoding module is used to predict the fused features to obtain a predicted density map, and a predicted value of the target quantity is obtained based on the predicted density map.
2. The target counting method based on multimodal large model according to claim 1 is characterized in that: The feature extraction module includes a text encoder and an image encoder; the step of extracting the text features of the number prompt word and the image features of the query image by using the feature extraction module includes: Using the text encoder to extract features from the number prompt words to obtain the text features; Using the image encoder to extract features from the query image to obtain the image features; Among them, the text encoder adopts the BERT model, and the image encoder adopts the DINOv2 model.
3. According to the target counting method based on the multimodal large model of claim 1, the step of fusing the text feature and the image feature using the feature fusion module to generate a fusion feature comprises: Extract category-related information from text features; The cross-attention mechanism is used to inject category-related information into image features to obtain fused features.
4. The target counting method based on multimodal large model according to claim 1 is characterized in that: The feature decoding module includes a dual-stream adaptive counting decoder; the dual-stream adaptive counting decoder includes a CNN stream, a Transformer stream, a T2C adapter and a gating network; The step of using the feature decoding module to predict the fusion feature to obtain a prediction density map includes: Use the CNN stream to capture local detail information in the fused feature to obtain a first intermediate feature; Using the Transformer stream to extract global information from the fused features, to obtain the second intermediate features; Using the T2C adapter, the second intermediate feature is transferred to the CNN stream, and is fused with the first intermediate feature to obtain an updated first intermediate feature; Generate a first density map and a second density map based on the updated first intermediate feature and the second intermediate feature respectively; The first density map and the second density map are fused through the gating network to obtain a predicted density map of the target category.
5. The target counting method based on multimodal large model according to claim 4 is characterized in that: The T2C adapter includes a channel weight calibration module and a cross attention module; The method of using the T2C adapter to transfer the second intermediate feature to the CNN stream and fusing the second intermediate feature with the first intermediate feature to obtain an updated first intermediate feature includes: In the channel weight calibration module, a multi-layer perceptron is used to project the second intermediate feature onto the channel dimension of the first intermediate feature in the CNN stream to generate a channel recalibration weight; In the cross-attention module, the first intermediate feature and the second intermediate feature are adjusted by a cross-attention mechanism to obtain a cross-attention feature; Generate an adaptation feature according to the channel recalibration weight and the cross attention feature; The adapted feature is added to the first intermediate feature block by block to generate an updated first intermediate feature.
6. The target counting method based on multimodal large model according to claim 4 is characterized in that: The number prompt words include factual text prompt words constructed using the correct target number and counterfactual text prompt words constructed using the incorrect target number.
7. The target counting method based on multimodal large model according to claim 6 is characterized in that: After the feature decoding module is used to predict the fusion feature to obtain a predicted density map, and a predicted value of the target quantity is obtained based on the predicted density map, the method further includes: Optimizing the target counting model through a loss function; The loss function includes a counting loss function, a dual-stream cross-number ranking loss function, and an image-text quantity alignment loss function; The counting loss function is used to constrain the difference between the predicted density map and the true density map; The dual-stream crossover number sorting loss function is used to sort the sum of pixel values in the sub-regions of the first density map and the second density map, and establish a relative relationship constraint of local density values; The image-text quantity alignment loss function is used to establish a similarity comparison relationship constraint between a negative sample pair and a positive sample pair; the positive sample pair is a sample pair generated based on the factual text prompt word and the query image; the negative sample pair is a sample pair generated based on the counterfactual text prompt word and the query image.
8. A target counting device based on a multimodal large model, characterized in that: The device comprises: A prompt word generation module, used to generate a number prompt word according to the category text and the number of targets corresponding to the query image; A feature extraction module, used for extracting text features of the number prompt word and image features of the query image; A feature fusion module, used for fusing the text feature and the image feature to generate a fusion feature; The feature decoding module is used to predict the fused features to obtain a predicted density map, and obtain a predicted value of the target quantity based on the predicted density map.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.